Online analysis instrument nonlinear error correction method based on manifold learning
By calculating the distortion rate of the manifold tangent space and the distortion weighted distance, the problem of nonlinear error correction of online analytical instruments in multi-physics coupling environment is solved, realizing high-precision measurement correction and accurate reconstruction of data manifold, thus meeting the real-time and reliability requirements of industrial process control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QINGDAO SANHUATAI ENG TECH CO LTD
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-12
AI Technical Summary
Existing online analytical instruments are ineffective in nonlinear error correction under multi-physics coupled environments. Traditional manifold learning algorithms cannot accurately perceive the physical stress and historical memory effect of drastic environmental changes on sensors, leading to misjudgment of neighboring points and affecting the accuracy and reliability of measurement values.
By calculating the distortion rate of the manifold tangent space and the distortion-weighted distance, an error correction method based on manifold learning is constructed. The original physical quantity readings and auxiliary environmental parameters are integrated to construct the distortion-weighted distance, which replaces the Euclidean distance. The local linear embedding algorithm is used to map to a low-dimensional space, and a prediction model is established by combining the least squares support vector machine to obtain the corrected measurement values.
It improves the generalization ability and measurement accuracy of the nonlinear error correction model, ensures that the local linear embedding algorithm is correctly deployed in the high distortion region, and realizes accurate reconstruction of the data manifold and real-time correction of the measured values.
Smart Images

Figure CN122020125A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology. More specifically, this invention relates to a method for correcting nonlinear errors in online analytical instruments based on manifold learning. Background Technology
[0002] Online analytical instruments, as key equipment for industrial process control, are widely used in petrochemical, steel smelting and environmental monitoring fields. The real-time performance and accuracy of their measurement data directly determine the level of optimized control and safe operation of production processes.
[0003] In real-world industrial applications, analytical instruments are constantly exposed to complex environments involving the coupling of multiple physical fields. During operation, sensors not only respond to changes in the concentration of the analyte but are also inevitably subjected to the combined impact of external disturbances such as sudden changes in chamber temperature, fluctuations in sample gas pressure, unstable carrier gas flow, and changes in ambient humidity. These environmental parameters often exhibit dynamic and nonlinear characteristics, rather than being constant values under ideal conditions.
[0004] However, existing error correction techniques still have limitations in handling such problems. Especially when environmental parameters fluctuate drastically in industrial settings, sensors often enter highly nonlinear response regions, causing severe bending or folding of the data manifold describing the system state. In this specific scenario, the Euclidean distance metric used in traditional manifold learning algorithms cannot perceive the physical stress and historical memory effects caused by drastic environmental changes on the sensor. This can easily lead the algorithm to erroneously traverse manifold bends when searching for reference neighborhoods for correction, misclassifying sample points with similar spatial coordinates but different actual physical states as similar neighbors. This neighborhood selection error caused by metric failure prevents the error prediction model from accurately reconstructing the current nonlinear drift characteristics, ultimately leading to the failure of instrument measurement correction and severely impacting the reliability of monitoring data. Summary of the Invention
[0005] To address the aforementioned technical problem of poor nonlinear error correction performance in online analytical instruments, this invention provides a nonlinear error correction method for online analytical instruments based on manifold learning, comprising: Obtain a high-dimensional feature vector set containing original physical quantity readings and auxiliary environmental parameters; calculate the manifold tangent space distortion rate of each dimension of the high-dimensional feature vector; the manifold tangent space distortion rate is positively correlated with the environmental coupling stress index; the environmental coupling stress index is positively correlated with the product of the signal transient response entropy, environmental sensitivity coefficient, and time change rate of the auxiliary environmental parameters of the corresponding dimension data; the signal transient response entropy includes the Shannon entropy of the probability distribution of all values of the corresponding dimension data within a set window, and is positively correlated with the ratio of the variance to the mean of all values within the window; calculate the distortion rate of each high-dimensional feature vector. The comprehensive manifold tangent space distortion rate is the arithmetic mean of the manifold tangent space distortion rates of all dimensions of data under the same high-dimensional feature vector. A distortion-weighted distance is constructed based on the distance between the high-dimensional feature vector set and the product of the corresponding comprehensive manifold tangent space distortion rates. A local linear embedding algorithm is used, replacing the Euclidean distance with the distortion-weighted distance, to map the high-dimensional feature vector set to a low-dimensional space, obtaining a low-dimensional feature vector sequence. A prediction model is trained based on the low-dimensional feature vector sequence. Real-time raw physical quantity readings are obtained, and corrected measurement values are acquired based on the real-time raw physical quantity readings and the prediction model.
[0006] This invention calculates the comprehensive manifold tangent space distortion rate, reflecting the curvature of the data structure, by integrating original physical quantity readings with auxiliary environmental parameters, and constructs a distortion-weighted distance to replace the Euclidean distance. This allows the algorithm to sense the physical stress in the data space when searching for nearest neighbors, actively avoiding pseudo-nearest neighbors in high-distortion regions. This ensures that the locally linear embedding algorithm can correctly unfold high-dimensional data manifolds affected by environmental coupling, improving the generalization ability and measurement accuracy of the nonlinear error correction model under multi-physics coupling environments.
[0007] Preferably, the transient response entropy of the signal satisfies the expression: ; In the formula, The transient response entropy of the signal at the i-th sampling point in the c-th dimension; Indicates the length of the sliding time window; Indicates the sliding time window of the i-th sampling point. The probability distribution density of the c-th dimension data at each sampling point; This represents the variance of the c-th dimension data within the sliding time window of the i-th sampling point; This represents the mean of the data in the c-th dimension within the sliding time window of the i-th sampling point; Represents the natural logarithm function; This represents a minute value.
[0008] This invention combines the disorder and fluctuation of signal values, enabling real-time assessment of the complexity and uncertainty of the internal state in the early stages when environmental parameters have not yet caused significant reading drift but the sensor has already experienced response oscillations or instability. This provides an accurate basis for subsequent assessment of the degree of manifold distortion.
[0009] Preferably, the environmental coupling stress index satisfies the expression: ; In the formula, The environmental coupling stress index represents the c-th dimension of the data at the i-th sampling point; The transient response entropy of the signal at the i-th sampling point in the c-th dimension; Indicates the environmental sensitivity coefficient; This represents the time rate of change vector of the normalized auxiliary environmental parameter sequence components at the i-th sampling point; The sign indicating the magnitude of a vector; This represents the normalization function.
[0010] This invention not only considers the severity of environmental changes, but also combines the environmental sensitivity coefficient and the sensor's own response state, thereby quantifying the magnitude of the nonlinear impact or stress on the sensor's measurement performance caused by external environmental changes at a specific moment, and accurately identifying the key regions in the data manifold that are most severely affected by environmental interference and most prone to nonlinear abrupt changes.
[0011] Preferably, the manifold tangent space distortion rate satisfies the expression: ; In the formula, This represents the manifold tangent space distortion rate of the c-th dimension of the data at the i-th sampling point; The environmental coupling stress index represents the c-th dimension of the data at the i-th sampling point; Indicates instantaneous distortion weight; Indicates the cumulative distortion weight; Indicates the number of sampling points in a memory cycle; The environmental coupling stress index represents the c-th dimension data of the im-th sampling point in the memory period of the i-th sampling point.
[0012] This invention employs a manifold tangent space distortion rate calculation model that includes instantaneous distortion weights and cumulative distortion weights. By introducing an integral term within the memory period, it is possible to identify data samples in the recovery or lag phases, preventing manifold structure evaluation biases caused by neglecting historical heat accumulation or component residual effects, and ensuring accurate capture of dynamic process errors.
[0013] Preferably, the distortion-weighted distance satisfies the expression: ; In the formula, This represents the distortion-weighted distance between the i-th sampling point and the j-th sampling point; Indicates the first High-dimensional feature vector of each sampling point With the High-dimensional feature vector of each sampling point The square of the Euclidean distance between them; and Let represent the combined manifold tangent space distortion rates of the i-th sampling point and the j-th sampling point, respectively; The symbol representing the magnitude of a vector.
[0014] This invention constructs a distortion-weighted distance formula and dynamically corrects the traditional Euclidean distance using the distortion rate of the manifold tangent space. When the instrument is in a highly distorted state with strong nonlinearity and significant environmental interference, this invention amplifies the computational distance between sample points, forcing the neighborhood search algorithm to be more cautious in these dangerous areas. It tends to select samples with truly similar physical states as neighbors, rather than samples that are only geometrically close but have drastically different physical meanings, thus fundamentally solving the short-circuit problem in manifold learning.
[0015] Preferably, the prediction model trained based on low-dimensional feature vector sequences includes: The prediction model is a least squares support vector machine with a radial basis function kernel. The prediction model takes the low-dimensional feature vector sequence as input and the difference between the original physical quantity readings and the standard true values at the corresponding sampling time as output.
[0016] Preferably, the step of obtaining real-time raw physical quantity readings and, based on the real-time raw physical quantity readings and the prediction model, obtaining corrected measurement values includes: During the real-time operation of the instrument, the real-time raw physical quantity reading sequence and the real-time auxiliary environmental parameter sequence are acquired to construct a real-time high-dimensional feature vector; based on the real-time high-dimensional feature vector, a real-time low-dimensional feature vector is constructed; and the prediction error is obtained. The corrected measurement value is obtained by subtracting the real-time raw physical quantity reading sequence of the instrument from the prediction error.
[0017] Preferably, the step of constructing a real-time low-dimensional feature vector based on a real-time high-dimensional feature vector includes: Calculate the distortion-weighted distance between the real-time high-dimensional feature vector and each vector in the high-dimensional feature vector set, and select... The nearest neighbor, This is the default value; Calculate the reconstructed weight vector of the real-time high-dimensional feature vector as a linear representation of its nearest neighbor; use the reconstructed weight vector to perform a linear weighted combination of the low-dimensional feature vector sequence corresponding to the nearest neighbor to obtain the real-time low-dimensional feature vector.
[0018] Preferably, obtaining the prediction error includes: Input the real-time low-dimensional feature vector into the error prediction model and output the prediction error.
[0019] Preferably, obtaining the memory cycle includes: Place the online analytical instrument in a constant temperature environment and wait for the reading to stabilize; Adjust the set temperature to raise the ambient temperature. Record the change in temperature from the start of the temperature change until the output signal amplitude reaches its final steady state. The duration of time elapsed is used as the memory cycle.
[0020] The beneficial effects of this invention are as follows: (1) This invention proposes the distortion rate of the manifold tangent space and constructs the distortion weighted distance accordingly. This distance metric can sense the curvature of the data space and penalize the path in the high distortion region when searching for the neighborhood, thereby eliminating pseudo-nearest neighbors that are geometrically close but physically different, ensuring that the local linear embedding algorithm can unfold the manifold along the real physical evolution path and improving the accuracy of nonlinear feature extraction. (2) This invention uses an improved manifold learning algorithm to map high-dimensional features to a low-dimensional space, thereby eliminating data redundancy and noise, and retaining the intrinsic features that can essentially reflect the error change law. Combined with the least squares support vector machine to establish a prediction model, it overcomes the dependence of traditional neural networks on a large number of training samples and can obtain extremely high generalization ability under small sample conditions. (3) Through the online reconstruction mechanism based on distortion weighted distance, the system can quickly calculate the low-dimensional projection of real-time data and output the correction value, which meets the requirements of industrial process control for the real-time performance and high reliability of analytical instrument data. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating a nonlinear error correction method for online analytical instruments based on manifold learning according to the present invention; Figure 2 It is a box plot illustrating the distribution of normalized data across various dimensions; Figure 3 This is a schematic diagram illustrating the comparison between the actual value of the error residual and the model prediction. Detailed Implementation
[0022] This invention discloses a method for correcting nonlinear errors in online analytical instruments based on manifold learning, referring to... Figure 1 This includes steps S1-S4: S1: Obtain the original physical quantity reading sequence and the auxiliary environmental parameter sequence through the data acquisition module; perform timestamp alignment and normalization on the original physical quantity reading sequence and the auxiliary environmental parameter sequence to construct a high-dimensional feature vector set.
[0023] It should be noted that during the operation of online analytical instruments in industrial fields, their output signals depend not only on the concentration of the analyte but also on the multiple coupling effects of external physical fields such as the temperature of the measurement chamber, the pressure of the sample gas, the flow rate of the carrier gas, and the ambient humidity. Traditional error correction methods often assume that environmental factors act independently or perform only simple linear compensation, ignoring the complex nonlinear drift of sensor response characteristics under the combined effects of multiple physical fields. In order to decouple the true instrument state from the high-dimensional data space, this invention first constructs a multi-source heterogeneous dataset containing internal response information and external environmental driving information, integrating discrete single physical quantities into a high-dimensional feature vector describing the overall state of the system, providing a holographic data foundation for subsequent mining of nonlinear manifold structures.
[0024] Specifically, the data acquisition module obtains the original physical quantity reading sequence and the auxiliary environmental parameter sequence; the original physical quantity reading sequence and the auxiliary environmental parameter sequence are then time-stamp aligned and normalized to construct a high-dimensional feature vector set, including: The system acquires in real time the raw physical quantity reading sequences output by the internal detection unit of the online analytical instrument. These raw physical quantity reading sequences include, but are not limited to, the spectral intensity sequences of the spectrometer, the detector voltage response sequences of the gas chromatograph, and the current response sequences of the electrochemical analyzer.
[0025] The auxiliary environmental parameter sequence of the instrument at the time of measurement is collected synchronously. The auxiliary environmental parameter sequence includes the measurement chamber temperature sequence, sample gas pressure sequence, carrier gas flow rate sequence, and ambient humidity sequence.
[0026] The original physical quantity reading sequence and the auxiliary environmental parameter sequence contain multi-dimensional data. The multi-dimensional data is time-aligned to ensure that each set of data corresponds to the same physical moment. The max-min normalization method is used to map each dimension of data to the 0-1 interval, obtaining normalized data for each dimension and eliminating the influence of different physical dimensions. The normalized data for each dimension is then used to construct a high-dimensional feature vector set, denoted as... ,in Let i be the high-dimensional feature vector of the i-th sampling point. This represents the total number of sampling points.
[0027] It should be noted that, as Figure 2The box plot shows the distribution range, median, and dispersion of the features in each dimension after normalization, demonstrating the standardization effect of the features after data preprocessing.
[0028] Thus, a high-dimensional feature vector set has been obtained.
[0029] S2: Calculate the signal transient response entropy of the original physical quantity reading sequence in each high-dimensional eigenvector, and derive the environmental coupling stress index by combining the time change rate of the auxiliary environmental parameter sequence; perform an integral transformation on the environmental coupling stress index to determine the manifold tangent space distortion rate.
[0030] It should be noted that the core assumption of manifold learning algorithms is that high-dimensional data is distributed on a low-dimensional manifold structure, and that this manifold is linear in its local neighborhood, i.e., a tangent space approximation. However, in real-world applications, when environmental parameters fluctuate drastically, the input-output characteristics of sensors enter a strongly nonlinear region, causing the data manifold to bend or fold dramatically in local areas. In such cases, traditional Euclidean distance cannot accurately measure the similarity between samples. Therefore, this invention introduces the concepts of stress and strain from physics. By calculating the entropy within the signal and the driving force of the external environment, it constructs a manifold tangent space distortion rate that can dynamically characterize the degree of manifold bending, thereby indicating which regions require special processing by the algorithm.
[0031] Specifically, the transient response entropy of the original physical quantity reading sequences in each high-dimensional eigenvector is calculated, and the environmental coupling stress exponent is derived by combining the time change rate of the auxiliary environmental parameter sequence, including: It should be noted that when a sensor is in the region dominated by nonlinear error, its microscopic response is often accompanied by oscillations. To assess this instability, this invention introduces the concept of information entropy. The higher the entropy value, the greater the degree of disorder within the signal, meaning the greater the likelihood that the signal is in an unstable state.
[0032] Select length as A sliding time window, for example, The preferred number of sampling points is 50.
[0033] The probability distribution density of arbitrary sampling points in arbitrary dimension data within a sliding time window is obtained using the histogram statistical method: the normalized data value range of the dimension data, i.e., the interval from 0 to 1, is divided into K equally spaced statistical intervals. The number of sampling points falling into each statistical interval within the sliding time window is counted, the frequency of the statistical interval to which each sampling point belongs is calculated, and the frequency is used as the probability distribution density of the sampling point in the dimension data within the sliding time window.
[0034] The transient response entropy of data from any sampling point in any dimension satisfies the expression: ; In the formula, The transient response entropy of the signal at the i-th sampling point in the c-th dimension; Indicates the length of the sliding time window; Indicates the sliding time window of the i-th sampling point. The probability distribution density of the c-th dimension data at each sampling point; This represents the variance of the c-th dimension data within the sliding time window of the i-th sampling point; This represents the mean of the data in the c-th dimension within the sliding time window of the i-th sampling point; Represents the natural logarithm function; This represents a tiny value, used to avoid a denominator of 0. For example, .
[0035] In the formula, Shannon entropy represents the c-th dimension of the data at the i-th sampling point, and represents the degree of dispersion of the signal amplitude distribution; It represents the relative fluctuation amplitude of the signal relative to its average level; when the sensor is in the linear operating region, the signal is stable and the transient response entropy is low; when it enters the nonlinear saturation region or the interference region, the signal amplitude distribution is discrete and the fluctuation is aggravated, resulting in an increase in the transient response entropy of the signal.
[0036] It should be noted that the transient response entropy of a signal only reflects the internal state, while nonlinear drift is usually driven by changes in the external environment. Considering that drastic changes in environmental parameters can exacerbate the nonlinear response of the sensor, this invention constructs an environmental coupling stress index to describe the linear superposition effect of the rate of environmental change on signal instability.
[0037] The environmental sensitivity coefficient is set by fitting historical calibration data; for example, the environmental sensitivity coefficient is 0.8.
[0038] The environmental coupling stress index of arbitrary dimension data at any sampling point satisfies the expression: ; In the formula, The environmental coupling stress index represents the c-th dimension of the data at the i-th sampling point; The transient response entropy of the signal at the i-th sampling point in the c-th dimension; Indicates the environmental sensitivity coefficient; This represents the time rate of change vector of the normalized auxiliary environmental parameter sequence components at the i-th sampling point; The sign indicating the magnitude of a vector; This represents the normalization function.
[0039] In the formula, Represents the internal state. Representing external driving force, when the external environment changes drastically and the internal signal entropy value is high at the same time, the environmental coupling stress exponent increases, indicating that the data point is in a strong nonlinear coupling region, that is, the manifold structure is under strong compression or stretching.
[0040] It should be noted that the degree of manifold curvature is not only related to the current instantaneous stress, but also has a historical memory effect. For example, the effect of sustained high temperatures on sensors is cumulative. Therefore, this invention introduces an integral term to evaluate this cumulative distortion.
[0041] Preferably, the integral transformation of the environmental coupling stress index is performed to determine the manifold tangential space distortion rate, including: Set the instantaneous distortion weight and the cumulative distortion weight. For example, the instantaneous distortion weight is 0.4 and the cumulative distortion weight is 0.6. It should be noted that the sum of the instantaneous distortion weight and the cumulative distortion weight is preferably 1, and the specific values of the instantaneous distortion weight and the cumulative distortion weight can be adjusted according to the response time constant of the instrument. If the instrument has a fast response speed, increase the value of the instantaneous distortion weight; if the instrument has a large thermal inertia, increase the value of the cumulative distortion weight.
[0042] During the instrument calibration phase, the instrument is placed in a constant temperature environment. After the reading stabilizes, the ambient temperature setpoint is quickly adjusted to raise the ambient temperature by 10°C within 10 seconds. The time elapsed from the moment the temperature begins to change until the amplitude change of the output signal reaches 95% of the final steady-state change is recorded. This time is used as the memory period. For example, the memory period is 3000 sampling points.
[0043] The distortion rate of the manifold tangent space of data at any sampling point and in any dimension satisfies the expression: ; In the formula, This represents the manifold tangent space distortion rate of the c-th dimension of the data at the i-th sampling point; The environmental coupling stress index represents the c-th dimension of the data at the i-th sampling point; Indicates instantaneous distortion weight; Indicates the cumulative distortion weight; Indicates the number of sampling points in a memory cycle; Indicates the first The environmental coupling stress index of the c-th dimension data of the im-th sampling point in the memory period of each sampling point.
[0044] In the formula, The average environmental coupling stress index represents the data of the c-th dimension within the memory period of the i-th sampling point, reflecting the cumulative effect of the environmental influence on the sensor over a period of time. Indicates the first The weighted sum of the environmental coupling stress index at each sampling point and the historical average environmental coupling stress index within the memory period reflects the overall curvature of the manifold structure. A larger value indicates a more pronounced curvature. The higher the degree of manifold curvature near a sampling point, the greater the error that will occur when using Euclidean distance to find neighboring points.
[0045] Thus, the distortion rate of the manifold tangent space was obtained.
[0046] S3: Construct a distortion-weighted distance metric using the distortion rate of the manifold tangent space of each high-dimensional feature vector, replacing the Euclidean distance search for the nearest neighbor of each high-dimensional feature vector; calculate the local reconstruction weight matrix based on the nearest neighbor, and map the high-dimensional feature vector set to the low-dimensional embedding space.
[0047] It should be noted that the classic Locally Linear Embedding (LLE) algorithm relies on Euclidean distance to determine nearest neighbors, which is effective when the manifold is flat. However, in regions with high distortion, the manifold undergoes severe curvature, and points with close Euclidean distances may have large geodesic distances on the manifold—a problem known as the short-circuiting problem. Therefore, this invention utilizes the distortion rate of the manifold tangent space to construct a physically aware weighted distance metric. When the distortion rate of a sample point is high, its calculated distance to other points is artificially amplified, forcing the algorithm to be more cautious when searching its neighborhood, tending to select points in directions where the manifold structure changes more gently, or shrinking the search radius when the curvature is large, thus ensuring that the mathematical assumption of local linearity still holds at the physical level.
[0048] Specifically, a distortion-weighted distance metric is constructed using the distortion rate of the manifold tangent space of each high-dimensional feature vector, replacing the Euclidean distance search for the nearest neighbor of each high-dimensional feature vector, including: Calculate the arithmetic mean of the manifold tangent space distortion rates for all dimensions of data at the i-th sampling point, denoted as . , which is the overall manifold tangent space distortion rate of the i-th sampling point.
[0049] The distortion-weighted distance between any two sampling points satisfies the expression: ; In the formula, This represents the distortion-weighted distance between the i-th sampling point and the j-th sampling point; Indicates the first High-dimensional feature vector of each sampling point With the High-dimensional feature vector of each sampling point The square of the Euclidean distance between them; and Let represent the combined manifold tangent space distortion rates of the i-th sampling point and the j-th sampling point, respectively; The symbol representing the magnitude of a vector.
[0050] In the formula, when the combined manifold tangent space distortion rate of the two sampling points is large, the product term will significantly amplify the metric distance between them, indicating that in regions with strong nonlinearity, the distance between points is greater than the geometric distance, thereby avoiding the incorrect selection of neighbors across the curved parts of the manifold.
[0051] Select each high-dimensional feature vector based on the distortion weighted distance. The nearest neighbors form the nearest neighborhood of the corresponding high-dimensional feature vector. For example, .
[0052] Preferably, the local reconstruction weight matrix is calculated based on the nearest neighbor, and the high-dimensional feature vector set is mapped to the low-dimensional embedding space, including: Calculate the local reconstruction weight matrix so that each high-dimensional feature vector can be linearly represented by high-dimensional feature vectors in its neighborhood set, and minimize the reconstruction error.
[0053] It should be noted that the calculation of local reconstruction weights aims to capture the local geometric structure of the data manifold. The specific calculation process is as follows: for the i-th high-dimensional feature vector... Construct the local covariance matrix within its nearest neighborhood, and obtain the weight vector by solving a system of constrained linear equations. The weight vectors must satisfy the condition that the sum of all weight coefficients in the weight vector is 1. The weight vectors of all sample points are combined to form a global reconstructed weight matrix.
[0054] Mapping the high-dimensional feature vector set to a low-dimensional embedding space while preserving the local reconstruction weight matrix, yields a low-dimensional feature vector sequence, denoted as . ,in Let i be the low-dimensional feature vectors corresponding to the i sampling points. This indicates the number of sampling points.
[0055] It should be noted that the mapping process aims to preserve the local topological structure of the data, meaning that the linear reconstruction relationship between a sample point and its neighbors in the high-dimensional space must still hold in the low-dimensional space. Specifically, this is achieved by finding a set of low-dimensional embedding coordinates such that each sample point can still be linearly represented by its nearest neighbors with the same local reconstruction weights in the low-dimensional space, while minimizing the global reconstruction error. Computationally, this is done by constructing a global alignment matrix based on the local reconstruction weight matrix, performing eigenvalue decomposition on this matrix, and selecting the eigenvectors corresponding to the smallest number of non-zero eigenvalues as the final low-dimensional eigenvectors.
[0056] Thus, the low-dimensional embedding space has been obtained.
[0057] S4: Establish a least squares support vector machine error prediction model using a low-dimensional feature vector sequence as input; collect real-time data and calculate the distortion rate of the comprehensive manifold tangent space, then input the distortion-weighted distance projection into the error prediction model to obtain the prediction error and complete the output correction.
[0058] It should be noted that, after manifold dimensionality reduction, the high-dimensional original measurement data and environmental data are compressed into low-dimensional feature vectors. These low-dimensional feature vectors essentially remove noise and redundancy, preserving the topological structure of the instrument's operating state. At this point, by using a least-squares support vector machine with strong few-sample learning capabilities to establish a mapping model from low-dimensional state features to the actual measurement error, high-precision nonlinear regression can be achieved. During the online operation phase, real-time data also undergoes physical feature extraction and manifold projection, enabling the calibration process to adaptively follow changes in the environment and instrument state.
[0059] Specifically, a least squares support vector machine error prediction model is established using a low-dimensional feature vector sequence as input, including: The low-dimensional feature vector sequence is used as input, and the residual between the instrument measurement value and the standard true value at the corresponding sampling time is used as output. The least squares support vector machine is trained to obtain the error prediction model.
[0060] Preferably, the least squares support vector machine uses a radial basis kernel function, and its kernel parameters and regularization parameters are determined through cross-validation.
[0061] It's important to note that after model training, real-time collected data cannot be directly input into the model. This is because real-time data is high-dimensional and also affected by manifold distortion. Feature extraction and projection must be performed using the same physical logic as during the training phase to ensure the accuracy of the prediction results.
[0062] Preferably, real-time data is collected and the comprehensive manifold tangent space distortion rate is calculated. After distortion-weighted distance projection, the data is input into the error prediction model to obtain the prediction error and complete the output correction, including: During the real-time operation of the instrument, the real-time raw physical quantity reading sequence and the real-time auxiliary environment parameter sequence are acquired to construct a real-time high-dimensional feature vector.
[0063] Calculate the distortion-weighted distance between the real-time high-dimensional feature vector and each vector in the high-dimensional feature vector set, and select... The nearest neighbor, for example, ; Calculate the reconstructed weight vector of the real-time high-dimensional feature vector as a linear representation of its nearest neighbor; use the reconstructed weight vector to perform a linear weighted combination of the low-dimensional feature vector sequence corresponding to the nearest neighbor to obtain the real-time low-dimensional feature vector.
[0064] Input the real-time low-dimensional feature vector into the error prediction model and output the prediction error; subtract the current original measurement reading of the instrument from the prediction error to obtain the corrected measurement value.
[0065] It should be noted that, as Figure 3 This diagram illustrates the comparison between the actual value of the error residual and the model's predicted value. The horizontal axis represents the sampling point number, and the vertical axis represents the prediction error. It shows the fitting accuracy of the error prediction model for nonlinear errors.
[0066] This completes the nonlinear error correction of online analytical instruments based on manifold learning.
[0067] While various embodiments of the invention have been shown and described in this specification, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention.
Claims
1. A method for correcting nonlinear errors in online analytical instruments based on manifold learning, characterized in that, include: Obtain a high-dimensional feature vector set containing raw physical quantity readings and auxiliary environmental parameters; Calculate the manifold tangent space distortion rate of each dimension of the data for each high-dimensional feature vector; the manifold tangent space distortion rate is positively correlated with the environmental coupling stress index; the environmental coupling stress index is positively correlated with the product of the signal transient response entropy, environmental sensitivity coefficient, and time change rate of auxiliary environmental parameters of the corresponding dimension data; the signal transient response entropy includes the Shannon entropy of the probability distribution of all values of the corresponding dimension data within a set window, and is positively correlated with the ratio of the variance to the mean of all values within the window; Calculate the comprehensive manifold tangent space distortion rate of each high-dimensional feature vector, where the comprehensive manifold tangent space distortion rate is the arithmetic mean of the manifold tangent space distortion rates of all dimensions of data under the same high-dimensional feature vector; construct a distortion-weighted distance based on the distance between the high-dimensional feature vector set and the product of the corresponding comprehensive manifold tangent space distortion rates; use the local linear embedding algorithm and replace the Euclidean distance with the distortion-weighted distance to map the high-dimensional feature vector set to a low-dimensional space to obtain a low-dimensional feature vector sequence; A prediction model is trained based on a low-dimensional feature vector sequence; real-time raw physical quantity readings are obtained; and corrected measurement values are obtained based on the real-time raw physical quantity readings and the prediction model.
2. The method for correcting nonlinear errors in online analytical instruments based on manifold learning according to claim 1, characterized in that, The transient response entropy of the signal satisfies the expression: ; In the formula, The transient response entropy of the signal at the i-th sampling point in the c-th dimension; Indicates the length of the sliding time window; Indicates the sliding time window of the i-th sampling point. The probability distribution density of the c-th dimension data at each sampling point; This represents the variance of the c-th dimension data within the sliding time window of the i-th sampling point; This represents the mean of the data in the c-th dimension within the sliding time window of the i-th sampling point; Represents the natural logarithm function; This represents a minute value.
3. The method for correcting nonlinear errors in online analytical instruments based on manifold learning according to claim 1, characterized in that, The environmental coupling stress index satisfies the expression: ; In the formula, The environmental coupling stress index represents the c-th dimension of the data at the i-th sampling point; The transient response entropy of the signal at the i-th sampling point in the c-th dimension; Indicates the environmental sensitivity coefficient; This represents the time rate of change vector of the normalized auxiliary environmental parameter sequence components at the i-th sampling point; The sign indicating the magnitude of a vector; This represents the normalization function.
4. The method for correcting nonlinear errors in online analytical instruments based on manifold learning according to claim 1, characterized in that, The tangent space distortion rate of the manifold satisfies the following expression: ; In the formula, This represents the manifold tangent space distortion rate of the c-th dimension of the data at the i-th sampling point; The environmental coupling stress index represents the c-th dimension of the data at the i-th sampling point; Indicates instantaneous distortion weight; Indicates the cumulative distortion weight; Indicates the number of sampling points in a memory cycle; The environmental coupling stress index represents the c-th dimension data of the im-th sampling point in the memory period of the i-th sampling point.
5. The method for correcting nonlinear errors in online analytical instruments based on manifold learning according to claim 1, characterized in that, The distortion-weighted distance satisfies the expression: ; In the formula, This represents the distortion-weighted distance between the i-th sampling point and the j-th sampling point; Indicates the first High-dimensional feature vector of each sampling point With the High-dimensional feature vector of each sampling point The square of the Euclidean distance between them; and Let represent the combined manifold tangent space distortion rates of the i-th sampling point and the j-th sampling point, respectively; The symbol representing the magnitude of a vector.
6. The method for correcting nonlinear errors in online analytical instruments based on manifold learning according to claim 1, characterized in that, The prediction model trained based on low-dimensional feature vector sequences includes: The prediction model is a least squares support vector machine with a radial basis function kernel. The prediction model takes the low-dimensional feature vector sequence as input and the difference between the original physical quantity readings and the standard true values at the corresponding sampling time as output.
7. The method for correcting nonlinear errors in online analytical instruments based on manifold learning according to claim 1, characterized in that, The process of obtaining real-time raw physical quantity readings and, based on the real-time raw physical quantity readings and the prediction model, obtaining corrected measurement values includes: During the real-time operation of the instrument, the real-time raw physical quantity reading sequence and the real-time auxiliary environment parameter sequence are acquired to construct a real-time high-dimensional feature vector; based on the real-time high-dimensional feature vector, a real-time low-dimensional feature vector is constructed; and the prediction error is obtained. The corrected measurement value is obtained by subtracting the real-time raw physical quantity reading sequence of the instrument from the prediction error.
8. The method for correcting nonlinear errors in online analytical instruments based on manifold learning according to claim 7, characterized in that, The construction of real-time low-dimensional feature vectors based on real-time high-dimensional feature vectors includes: Calculate the distortion-weighted distance between the real-time high-dimensional feature vector and each vector in the high-dimensional feature vector set, and select... The nearest neighbor, This is the default value; Calculate the reconstructed weight vector of the real-time high-dimensional feature vector as a linear representation of its nearest neighbor; use the reconstructed weight vector to perform a linear weighted combination of the low-dimensional feature vector sequence corresponding to the nearest neighbor to obtain the real-time low-dimensional feature vector.
9. The method for correcting nonlinear errors in online analytical instruments based on manifold learning according to claim 7, characterized in that, The acquisition of prediction error includes: Input the real-time low-dimensional feature vector into the error prediction model and output the prediction error.
10. The method for correcting nonlinear errors in online analytical instruments based on manifold learning according to claim 4, characterized in that, The acquisition of the memory cycle includes: Place the online analytical instrument in a constant temperature environment and wait for the reading to stabilize; Adjust the set temperature to raise the ambient temperature. Record the change in temperature from the start of the temperature change until the output signal amplitude reaches its final steady state. The duration of time elapsed is used as the memory cycle.