Method and system for rapid evaluation of bulk ctd sensor conductivity drift performance

CN122506468BActive Publication Date: 2026-09-11QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610991663.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-11
Estimated Expiration
2046-07-06

AI Technical Summary

Technical Problem

(1)设备成本高昂

Benefits of technology

(1)不依赖高精度外部基准,显著降低评估成本

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122506468B_ABST
    Figure CN122506468B_ABST
Patent Text Reader

Abstract

The application discloses a kind of batch CTD sensor conductivity drift performance fast evaluation method and system, it is related to marine observation instrument test and calibration technical field.The method comprises: the conductivity time series data of multiple temperature salinity depth sensors is synchronously collected;Statistical weight of each sensor is estimated based on the historical error variance in sliding window, and improved robust fuzzy C mean fusion algorithm is used to generate virtual true value, and virtual true value sequence is biased and corrected based on group median;The deviation sequence of each sensor relative to corrected virtual true value is calculated, and long-term drift rate and initial bias are estimated by linear regression estimation;Quantitative score is carried out from four dimensions, and comprehensive score is calculated by weighting, and drift grade is judged.The application does not need to rely on high-precision external reference, supports batch parallel evaluation, can quickly output objective quantitative score and diagnostic conclusion, and is suitable for CTD sensor batch production detection and on-site rapid screening scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of marine observation instrument testing and calibration technology, and in particular to a rapid evaluation method and system for the conductivity drift performance of batch CTD sensors. Background Technology

[0002] Conductivity-Temperature-Depth (CTD) sensors are core instruments for ocean observation, widely used in marine environmental monitoring, marine resource exploration, and climate research. Among these, the accuracy of conductivity measurement directly affects the accuracy of seawater salinity calculation, making it one of the most critical performance indicators of CTD sensors. However, in practical use, the measurement accuracy of conductivity sensors is highly susceptible to drift caused by factors such as electrode contamination, biofouling, environmental temperature changes, and component aging, leading to a decline in the quality of observational data. Unlike temperature and depth parameters, conductivity measurement lacks a long-term stable physical benchmark and is significantly affected by factors such as electrode contamination and aging. It is the parameter with the fastest drift and the most difficult to calibrate among CTD sensor parameters, and its calibration problem is more complex and challenging than that of temperature and depth. Therefore, accurately evaluating the conductivity drift performance of CTD sensors is of great significance for ensuring the reliability of ocean observation data and supporting marine scientific research and engineering applications.

[0003] Currently, CTD sensor conductivity calibration and drift assessment mainly rely on the following two types of technical solutions: Option 1: Comparative calibration method based on a high-precision constant-temperature water bath. This method involves simultaneously placing the CTD sensor under test and a high-precision reference CTD sensor in a high-precision constant-temperature water bath. Temperature and salinity control maintains a stable environment within the bath. Conductivity measurements of the tested sensor and the reference sensor are read and compared one by one, and the deviation is calculated to determine the degree of drift. This method offers high accuracy, is traceable to national standards, and is currently the mainstream calibration method.

[0004] Option 2: Periodic calibration based on standard solutions. This method uses standard seawater or a standard solution with known conductivity. The CTD sensor is placed in the standard solution to measure its conductivity. The difference between the measured value and the standard value is compared to determine the sensor's measurement accuracy and drift status. This method is commonly used for periodic calibration in laboratories.

[0005] However, the above-mentioned existing technical solutions still have the following shortcomings: (1) High equipment cost. The purchase cost of high-precision constant temperature water bath and high-precision reference CTD sensor equipment is expensive, and they require regular metrological verification and maintenance, resulting in high operating costs. It is difficult to popularize them as a high-frequency daily screening method in ordinary production processes; (2) The evaluation efficiency is low. The constant temperature control, salinity balance and sequential comparison of each sensor are time-consuming. The evaluation cycle of a single sensor is usually tens of minutes to several hours, which is difficult to meet the high-throughput evaluation needs of dozens to hundreds of sensors in the scenarios of mass production, factory testing and rapid on-site screening. (3) Lack of an effective initial screening mechanism. Before entering the high-precision calibration process, there is a lack of low-cost, objective, and quantifiable rapid initial screening methods, which can easily lead to sensors with obvious performance abnormalities occupying expensive high-precision calibration resources and causing resource waste; (4) Insufficient interpretability of evaluation results. Traditional methods usually only output a single deviation value or error range, which is difficult to intuitively reflect the multi-dimensional characteristics of sensor drift. On-site engineers find it difficult to quickly determine whether the sensor should be released directly, continue to be observed, calibrated and used, or be reworked or discarded based on a single indicator;

[0006] To address the shortcomings of the existing technologies, it is necessary to propose a rapid evaluation method that can support parallel processing of multiple sensors, does not rely on high-precision external hardware benchmarks, can generate reference benchmarks based on the sensor group's own data, and output objective quantitative scores, drift levels, and diagnostic conclusions. Summary of the Invention

[0007] To address the aforementioned technical problems, this invention provides a rapid evaluation method and system for the conductivity drift performance of batch CTD sensors. This system enables parallel processing of multiple sensors and rapid output of objective quantitative scores and drift levels without relying on high-precision external benchmarks and a strictly constant temperature environment. As a result, it effectively reduces evaluation costs, improves evaluation efficiency, and provides an intuitive basis for engineering decisions.

[0008] To achieve the above objectives, the technical solution of the present invention is as follows: A rapid evaluation method for conductivity drift performance of batch CTD sensors includes the following steps: Step S1: Synchronously collect the conductivity time series data of multiple temperature, salinity and depth sensors deployed in the test environment, and perform text cleaning, outlier filtering and quality screening on the raw conductivity data to obtain the preprocessed effective conductivity data. Step S2: Based on the historical error variance within the sliding window, estimate the statistical weights of each sensor online, and use an improved robust fuzzy C-means fusion algorithm to fuse the conductivity data of all sensors at the same time to generate virtual true values. Then, perform bias correction on the virtual true value sequence based on the population median. Step S3: Calculate the deviation sequence of each sensor relative to the calibrated virtual true value, perform univariate linear regression on the deviation sequence of each sensor, and estimate its long-term drift rate and initial bias. Step S4: Quantify the scores from four dimensions: bias, drift rate, stability, and range, calculate the weighted comprehensive score, and automatically determine and output the drift level.

[0009] In the above scheme, step S1 specifically includes: S1.1: Simultaneously collect conductivity time-series data from no less than 20 temperature, salinity, and depth sensors at a sampling frequency of no less than 1Hz, with a continuous collection time of no less than 10 minutes; S1.2: Identify and extract the pure conductivity data column from the raw data column of each sensor, and take the arithmetic mean of multiple conductivity measurement columns at the same time as the final conductivity measurement value of the sensor at time t, denoted as . ; S1.3: Outlier filtering is performed based on the 3σ criterion, and the global mean of conductivity values ​​of all sensors at all times is calculated. and global standard deviation , will satisfy or The data points were identified as outliers and removed. S1.4: Perform raw sensor data quality screening and calculate the global median of conductivity values ​​for all sensors. and the median of each sensor ,when The sensor was determined to have severe initial attenuation, and its initial weight was reduced in subsequent fusion processes.

[0010] In the above scheme, the virtual truth value generation in step S2 specifically includes the following sub-steps: S2.1: Online estimation of historical error variance: Set the length of the sliding window The sensor is updated online using the following formula. At any moment Historical error variance : ; Among them, error , The virtual truth value of the previous moment. The value range is 20-100; S2.2: Calculation of statistical weights: Based on the historical error variance, the sensor value is calculated using the following formula. At any moment Fusion weights , ; in, This represents the total number of sensors; S2.3: Improved Robust Fusion of Fusion in Fusion of ... Set the number of clusters Design the coupling coefficient between the membership index and the sensor weights. : ; in, This is the adjustment coefficient; Initialize normal class centers using the weighted average and the upper quartile, respectively. and noise center ; Then, the membership matrix is ​​updated alternately. and cluster center Minimize the objective function: ; in, For the first Data points Belongs to the The membership degree of each cluster, For the first Data points To the cluster center Euclidean distance, For the first Membership index coupling coefficient of each data point For the first The centers of each cluster are identified; after iterative convergence, the cluster center with the largest sum of membership degrees is taken as the virtual truth value at the current time step. .

[0011] In the above scheme, the membership update formula in step S2.3 is: ; in, For the first The data point belongs to the th data point The membership degree of each cluster, and All values ​​represent the Euclidean distance from the data point to the corresponding cluster center. For the first Membership index coupling coefficient of each data point; The formula for updating cluster centers is: ; in, For the first Cluster centers, For the first The measured values ​​of each sensor; When the maximum change in the membership matrix between two consecutive iterations is less than a preset threshold The iteration terminates at this point: ; in, This represents the membership value updated in the current iteration. This is the membership value from the previous iteration. At the same time, the maximum number of iterations is set to 100 to prevent the algorithm from failing to converge.

[0012] In the above scheme, step S2, which involves bias correction of the virtual truth sequence based on the population median, specifically includes the following sub-steps: S2.4.1: Calculate the global median of conductivity values ​​for all sensors at all times: ; in, The total number of sensors, This represents the total number of sampling times. S2.4.2: Calculate the median of the dummy truth sequence: ; The correction amount is calculated according to the following formula. : ; in, The median is the corrected intensity coefficient; S2.4.3: Apply global bias correction to the virtual truth sequence: , ; in, For a moment The corrected virtual truth value.

[0013] In the above scheme, step S3 specifically includes: S3.1: Calculate each sensor At each sampling time Deviation: ; The deviations at all times are arranged in chronological order to form a deviation sequence. ; S3.2: Establish a univariate linear regression model for the deviation sequence of each sensor: ; in, For long-term drift rate, For the initial bias, The random error term follows a normal distribution with a mean of 0. Long-term drift rate estimated using the least squares method and initial bias : ; ; in, The mean value at each sampling time. This is the mean of the biased sequence.

[0014] In the above scheme, step S4, which involves quantitative scoring from four dimensions, specifically includes: Bias score is used to assess the initial systematic bias of the sensor; ; Drift score is used to assess the long-term drift trend of a sensor; ; Stability score is used to assess the dispersion of sensor measurements; ; The numerator is the standard deviation of the biased sequence. ; Range scoring is used to assess whether there are localized, drastic jumps; ; in, That is, the range of the biased sequence. ; Overall Score: .

[0015] In the above scheme, the drift level determination criterion in step S4 is: Overall Score The time frame is determined to be "stable," which in engineering terms means that the sensor is performing well and can be used directly. The timeframe indicates "slight drift," which in engineering terms means the performance is acceptable and it is recommended to monitor it during use. The timeframe indicates "moderate drift," which in engineering terms means a significant performance degradation and suggests calibration before use. The time was determined to be "severe drift", which means that the performance is unqualified and it is recommended to rework or discard it.

[0016] In the above scheme, in step S2, for sensors marked as having severe initial attenuation in step S1.4, their fusion weights are multiplied by an attenuation factor. , The range of values ​​is .

[0017] A rapid evaluation system for the conductivity drift performance of batch CTD sensors implementing the method described above, comprising: Multi-channel synchronous data acquisition module: used to synchronously acquire conductivity data from multiple temperature, salinity and depth sensors at a sampling rate of not less than 1Hz, supports multiple sensor interface protocols and automatically identifies connections; Data preprocessing module: used to realize automatic identification of conductivity columns, text cleaning, outlier filtering based on the 3σ criterion, and quality screening of raw sensor data; Historical error variance and weight calculation module: used to perform historical error calculation, sliding window variance estimation, statistical weight calculation, low-confidence sensor weight attenuation and weight normalization processing; Improved robust fuzzy C-means fusion module: used to realize online weight update based on historical error variance, robust fuzzy C-means fusion with membership degree and weight coupling, and generation of virtual truth value; Virtual Truth Median Correction Module: Used to calculate the population median and the virtual truth median, as well as to correct the overall bias of the virtual truth sequence; The deviation accumulation and regression analysis module is used to perform deviation sequence calculation, univariate linear regression analysis, and long-term drift rate estimation. Multi-dimensional quantitative scoring and report generation module: used to realize four-dimensional drift scoring, comprehensive score calculation, drift level determination, and export of visualized results including deviation curves and evaluation reports.

[0018] Through the above technical solution, the present invention provides a rapid evaluation method and system for the conductivity drift performance of batch CTD sensors, which has the following beneficial effects: (1) It does not rely on high-precision external benchmarks, which significantly reduces the evaluation cost. This invention generates virtual true values ​​by fusing group data from multiple CTD sensors under test, eliminating the need to rely on high-precision reference CTD sensors and high-precision constant temperature water baths in the initial screening stage. This avoids the costs of purchasing, maintaining, and periodically verifying high-precision reference equipment, and significantly reduces the hardware and operation and maintenance costs of batch sensor drift performance evaluation. (2) Supports batch parallel evaluation, which greatly improves evaluation efficiency. This invention can simultaneously acquire and process conductivity time-series data from no fewer than 20 CTD sensors. Through a multi-channel synchronous acquisition and parallel computing architecture, it overcomes the efficiency bottleneck of traditional methods that rely on serial comparison of each sensor. It is suitable for mass production, factory testing, and rapid on-site screening scenarios, reducing the evaluation time for a single batch from tens of minutes to several hours using traditional methods to minutes, significantly improving the initial screening efficiency of batch sensors. (3) Construct an objective and quantitative multi-dimensional drift assessment system This invention quantifies and scores sensor performance across four dimensions: bias, drift rate, stability, and range. Each dimension's score range is [0,1], and a unified evaluation is achieved through weighted composite scores. Compared to traditional methods that only output a single deviation value or rely on human experience, this invention improves the objectivity, comparability, and repeatability of the evaluation results. (4) The evaluation results are intuitive and interpretable, which facilitates engineering decision-making. This invention outputs a four-dimensional score, a comprehensive score, and a clear drift level, classifying sensors into four levels: "stable," "slight drift," "moderate drift," and "severe drift," each with a clear engineering meaning. Field engineers can quickly determine whether a sensor should be released directly, monitored during use, used after calibration, or returned for repair, providing an intuitive and clear basis for engineering decisions. (5) The algorithm is robust and suitable for complex data environments. This invention combines outlier filtering based on the 3σ criterion, low-confidence sensor weight attenuation, online estimation of historical error variance using a sliding window, and an improved robust fuzzy C-means fusion algorithm incorporating noise, forming a multi-layered data quality control system. It effectively reduces the impact of abnormal sensors, transient noise, and local jumps on the virtual truth generation results, maintaining the stability of the evaluation results even when a small number of severely degraded sensors exist in a multi-sensor group, demonstrating broad engineering applicability. (6) Provides an effective pre-screening method for high-precision calibration This invention can serve as a rapid preliminary screening step before entering the traditional high-precision calibration process, identifying and diverting sensors with significantly abnormal performance in advance, thus preventing them from occupying expensive high-precision calibration resources. Through a two-stage evaluation architecture of "coarse screening - fine calibration," optimal allocation of calibration resources is achieved, improving the economy and efficiency of the overall calibration process. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0020] Figure 1 This is a schematic flowchart of a rapid evaluation method for the conductivity drift performance of batch CTD sensors disclosed in an embodiment of the present invention. Figure 2 This is a schematic diagram of a rapid evaluation system for the conductivity drift performance of batch CTD sensors disclosed in an embodiment of the present invention. Detailed Implementation

[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0022] This invention provides a rapid method for evaluating the conductivity drift performance of batch CTD sensors, such as... Figure 1 As shown, the process includes the following steps: synchronous acquisition and data preprocessing (S1), virtual truth generation and median correction (S2), deviation sequence calculation and global regression analysis (S3), and multi-dimensional quantitative scoring and comprehensive judgment (S4).

[0023] Step S1: Synchronous Acquisition and Data Preprocessing Step S1 is used to acquire the raw conductivity time series data of multiple CTD sensors under test, and to clean, filter and screen the data to provide a reliable data foundation for subsequent fusion processing.

[0024] S1.1: Sensor Deployment and Data Acquisition Installed in a constant temperature water bath or a regular water bath One temperature, salinity, and depth sensor to be measured. To ensure the validity of population statistics, the number of sensors... No fewer than 20. The conductivity-sensitive ends of each sensor are placed at similar heights and in similar flow field positions to reduce measurement differences introduced by uneven temperature and salinity distribution within the water tank.

[0025] Through a multi-channel synchronous data acquisition module, at a uniform sampling frequency Conductivity data from each sensor are collected synchronously. Sampling frequency. The frequency should be no less than 1Hz to ensure the capture of the time-varying characteristics of the conductivity signal. Continuous acquisition duration. At least 10 minutes are required to obtain a sufficient number of data samples for statistical analysis and drift assessment.

[0026] After collection, obtain OK, The original data matrix of the column, where This represents the total number of sampling points. This represents the total number of sensors.

[0027] S1.2: Conductivity Data Cleaning and Extraction Identify and extract the pure conductivity data column from the raw data columns of each sensor. Since the raw data file may contain multiple parameter columns such as timestamp, temperature, depth, and salinity, it is necessary to automatically identify the conductivity data column based on the field name or column identifier.

[0028] Each sensor One or more pure numerical conductivity measurement columns can be obtained (e.g., the same sensor may output multiple conductivity measurement channels). The arithmetic mean of these columns at the same time is taken as the value of the sensor at time [time value missing]. The final conductivity measurement value is denoted as .

[0029] S1.3: Based on 3 Outlier filtering criteria Calculate the global mean of the effective conductivity measurements of all sensors at all time points. and global standard deviation : ; ; in, The total number of sensors, This represents the total number of sampling times. For the first Each sensor at time The measured conductivity value.

[0030] For any data point If any of the following conditions are met, it is considered an outlier and will be removed: ; The above 3 The criterion is based on the assumption of normal distribution. Under normal distribution, the probability of data points exceeding the mean ± 3 standard deviations is about 0.27%, which can be regarded as a low probability event and is therefore judged as outliers.

[0031] S1.4: Quality screening of raw sensor data Calculate the global median of conductivity values ​​for all sensors. : ; For sensors Calculate the median : ; The sensor is considered to have severe initial attenuation if the following conditions are met: ; The basis for the above judgment criteria is that when the median conductivity measurement of a certain sensor is lower than 70% of the global median of all sensors, it indicates that the measured value of the sensor has deviated significantly from the population center level, and there may be problems such as severe electrode contamination, sensor damage, or serious inaccuracy in factory calibration.

[0032] Sensors marked as having severe initial attenuation will be assigned lower initial weights in subsequent fusion processes, with their fusion weights multiplied by the attenuation factor. , The range of values ​​is However, its original measurement remains unchanged and still participates in subsequent data processing and evaluation processes.

[0033] II. Step S2: Virtual Truth Generation and Median Correction Step S2 is the core step of this invention. This step utilizes data from the entire sensor group to generate a high-precision virtual truth value online, independent of any external hardware benchmark. And bias correction is performed on it based on the population median.

[0034] S2.1: Online estimation of historical error variance Define sensor At any moment Measurement error The difference between its measured value and the fusion result from the previous time step: ; in, This is the virtual truth value from the previous time step. For the initial time step... Since there is no The initial variance of all sensors can be set to a very small positive number (e.g., This ensures the algorithm starts up correctly.

[0035] The error variance of each sensor is updated online using a sliding window mechanism. Sliding window length The value range is from 20 to 100, and in this embodiment, it is taken as... Sliding window length The physical meaning is the number of historical data points used to estimate the error variance; Taking values ​​that are too small will cause the variance estimate to be overly sensitive to instantaneous noise. Taking values ​​that are too large will cause the variance estimate to lag in response to changes in sensor weights. The updated formula is as follows: ; in, For sensors At any moment Measurement error, and They are the current time and The recursive formula is based on the squared error value from a previous time step. The physical meaning of this formula is: the current error variance estimate equals the estimate from the previous time step, plus the squared error contribution from newly entered windows, minus the squared error contribution from exited windows. This achieves online recursive updating of variance, avoiding redundant calculations of all historical data and improving computational efficiency.

[0036] S2.2: Calculation of Statistical Weights Based on the historical error variance, the sensor value is calculated using the following formula. At any moment Fusion weights : ; in, The total number of sensors, For sensors At any moment The historical error variance is used. The physical meaning of this weighting formula is: the smaller the error variance of a sensor, the more reliable its measurement value, and the greater its weight in the fusion; conversely, the larger the error variance of a sensor, the lower the reliability of its measurement value, and the smaller its weight. The denominator is the sum of the reciprocals of the weights of all sensors, used to normalize the weights to ensure that the sum of the weights of all sensors is 1.

[0037] For sensors marked as having severe initial attenuation in step S1.4, their fusion weights are multiplied by the attenuation factor. ( ( ), in order to reduce its impact on virtual truth generation.

[0038] S2.3: Improved Robust Fusion of Fusion ... Set the number of clusters These correspond to the "normal class" and the "noise class," respectively. The coupling coefficient between the membership index and the sensor weights is designed. : ; in, For sensors The fusion weight, The adjustment coefficient is taken as [value] in this embodiment. The coupling coefficient The physical meaning is: when the sensor weight When the value is large (i.e., the reliability is high), When the membership index is close to 1, it resembles hard clustering, with clear classification boundaries; when the sensor weights When the value is small (i.e., the reliability is low), near A smaller positive number (a membership index) allows the data point to have a more ambiguous category affiliation, reducing its impact on the fusion result.

[0039] Normal class center The initial estimate is a weighted average of all measurements: ; in, For the first The current measurement value of each sensor. Set its fusion weight.

[0040] Noise Class Center The initial estimate is taken from the upper quartile (i.e., the 75th percentile) of all measurements to ensure that the initial noise class center is located in the high-end region of the data distribution, thus distinguishing it from the normal class center.

[0041] The initial range of the noise class is adaptively determined based on the distribution of normal sensors: ; in, This is the radius adjustment coefficient, in this embodiment... The value is 2. The value range is usually from 1.5 to 3.0. The number of clusters, The normal class center is defined by this formula, which means that the initial radius variance of the noise class is proportional to the average deviation of the normal sensor around the normal class center. Used to adjust the initial size of the noise class, so that the noise class can adaptively cover data points that deviate from the normal range.

[0042] Define the objective function Measure the sum of weighted distances between each data point and its cluster center: ; in, The number of clusters, The total number of sensors, For the first The data point (i.e. the th data point) The measurement value of the first sensor belongs to the first sensor. The membership degree of each cluster, For the first The membership index coupling coefficient of each data point (its meaning is as described above) same), For the first Data points To the Cluster centers Euclidean distance, For the first Cluster centers.

[0043] By alternating updates to the membership matrix and cluster center Minimize objective function .

[0044] The membership update formula is: ; in, For the first Data points To the cluster center Euclidean distance, For the first Data points To the cluster center Euclidean distance, For the first The membership index coupling coefficient for each data point is defined as described above. same.

[0045] The formula for updating cluster centers is: ; in, For the first Cluster centers, For the first The data point belongs to the th data point The membership degree of each cluster, For the first Membership index coupling coefficient of each data point For the first The formula represents the measurements from each sensor. The physical meaning of this formula is that the cluster center is the weighted average of all data points according to their membership degree; data points with higher membership degrees have a greater influence on the location of the cluster center.

[0046] When the maximum change in the membership matrix between two consecutive iterations is less than a preset threshold The iteration terminates at this point: ; in, This represents the membership value updated in the current iteration. This is the membership value from the previous iteration. This is the convergence threshold. The maximum number of iterations is also set to 100 to prevent the algorithm from failing to converge.

[0047] After iterative convergence, compare the sum of the membership degrees of the two clusters. and The class with the largest sum of membership degrees is marked as the "normal class", and the cluster center of this class is taken as the virtual truth value at the current moment. : .

[0048] S2.4: Median bias correction of virtual truth values Before calibration, calculate the global median of conductivity values ​​for all sensors at all times. : ; in, The total number of sensors, This represents the total number of sampling times.

[0049] Then calculate the virtual truth sequence. the median of : ; Calculate the correction amount : ; in, This is the median correction intensity coefficient. The physical meaning of is the intensity factor that controls the degree of correction: The larger the value, the stronger the correction, and the greater the degree to which the virtual true value sequence as a whole approaches the population median; Pick The range can strike a balance between effectively eliminating systematic offsets and preserving reasonable differences introduced by the fusion algorithm, avoiding overcorrection.

[0050] Finally, a global bias correction is applied to the virtual truth sequence: ; in, For a moment The corrected virtual ground truth, in physical terms, refers to the virtual ground truth after median bias correction. This ensures that the overall level of the virtual ground truth sequence aligns with the central trend of the sensor population, reducing the systematic bias in the fusion algorithm caused by a small number of biased sensors or initial weight values. This will serve as a reference benchmark for subsequent deviation calculations.

[0051] III. Step S3: Calculation of Deviation Sequence and Global Regression Analysis S3.1: Calculation of Deviation Sequence For each sensor Calculate its value at each sampling time. Measured values Compared with the corrected virtual true value The difference is used to obtain the deviation: ; The deviations at all moments, arranged in chronological order, constitute the sensor. Deviation sequence: ; deviation sequence The physical meaning is: sensor At any moment The degree of systematic deviation relative to the group reference benchmark; a positive value indicates that the sensor measurement is higher than the virtual true value, and a negative value indicates that it is lower than the virtual true value.

[0052] S3.2: Global Regression Analysis A univariate linear regression analysis was performed on the deviation sequence of each sensor to establish a regression model: ; in, Sampling time ( ), Long-term drift rate represents the rate of change of sensor deviation over time, with units of conductivity value / sampling point. Its physical meaning is the amount of drift of the sensor per unit time. The initial system bias represents the systematic deviation of the sensor from the virtual true value at the initial moment, expressed in units of conductivity. Its physical meaning is the inherent initial measurement error of the sensor. The random error term follows a normal distribution with a mean of 0, representing random fluctuations in the deviation sequence that cannot be explained by a linear trend.

[0053] Long-term drift rate estimated using the least squares method and initial bias .

[0054] Drift rate The estimation formula is: ; in, The mean value across all sampling times. For sensors The mean of the deviation sequence. The numerator represents the covariance between time and deviation, and the denominator represents the variance between time and deviation. The ratio of the two is the regression slope. .

[0055] Initial bias The estimation formula is: ; The physical meaning of this formula is: the regression line passes through the point... The intercept is derived from the slope and the mean point. That is, the initial time ( The deviation estimate of ).

[0056] IV. Step S4: Multi-dimensional quantitative scoring and comprehensive judgment Step S4 quantifies the drift performance of each sensor from four dimensions: bias, drift rate, stability, and range, and determines the drift level by weighted comprehensive score.

[0057] S4.1: Global Four-Dimensional Drift Score Each sensor is quantitatively scored based on the following four dimensions, with the score range for each dimension being [value range missing]. A higher score indicates better performance in that dimension.

[0058] (1) Bias score

[0059] Bias score is used to assess the initial systematic bias of the sensor, based on the initial systematic bias. calculate: ; in, For sensors The initial bias is expressed in terms of conductivity (S / m). The threshold of 0.001 is the reference allowable deviation in conductivity (S / m). The physical meaning of this formula is as follows: when the absolute value of the initial bias does not exceed 0.001, the bias score is 1 (full score); when the absolute value of the initial bias reaches or exceeds 0.001, the bias score drops to 0 (lowest score), and intermediate values ​​are deducted linearly.

[0060] (2) Drift score

[0061] Drift score is used to assess the long-term drift trend of the sensor, based on the regression slope. calculate: ; in, For sensors The long-term drift rate is expressed as conductivity value per sampling point (S / m / point). A threshold of 0.0001 is the reference limit for the drift rate. The physical meaning of this formula is: when the absolute value of the drift rate does not exceed 0.0001, the drift score is 1 (full marks); when the absolute value of the drift rate reaches or exceeds 0.0001, the drift score drops to 0 (lowest score), with intermediate values ​​reduced linearly.

[0062] (3) Stability score

[0063] Stability scores are used to assess the dispersion of sensor measurements, based on the standard deviation of the deviation sequence. calculate: ; in, For sensors The standard deviation of a biased series is calculated using the following formula: ; The physical meaning of is the degree of dispersion of the deviation sequence around its mean, reflecting the random fluctuation range of the sensor measurement values. The threshold of 0.0001 is the reference limit for the standard deviation. The physical meaning of this formula is: when the standard deviation does not exceed 0.0001, the stability score is 1 (full score); when the standard deviation reaches or exceeds 0.0001, the stability score drops to 0 (minimum score), and intermediate values ​​are deducted linearly.

[0064] (4) Range scoring

[0065] Range scoring is used to assess the range of a biased sequence, reflecting the presence of local drastic jumps, based on the range of the biased sequence. calculate: ; in, For sensors The range of the deviation sequence, i.e., the difference between the maximum and minimum deviations, is expressed in conductivity values ​​(S / m). Physically, it represents the range of the deviation sequence, reflecting the overall fluctuation amplitude of the sensor's deviation throughout the measurement period and is sensitive to drastic local jumps. A threshold of 0.003 is the reference limit for the range. The physical meaning of this formula is: when the deviation range does not exceed 0.003, the range score is 1 (full marks); when the deviation range reaches or exceeds 0.003, the range score drops to 0 (lowest score), with intermediate values ​​reduced linearly.

[0066] Overall score: ; The weights 0.3, 0.4, 0.2, and 0.1 represent the weighting coefficients for the four dimensions: bias, drift rate, stability, and range, respectively, with a sum of 1. The weighting coefficients are set as follows: drift rate is the most direct indicator of long-term sensor performance degradation, therefore it is assigned the highest weight of 0.4; bias reflects the initial accuracy of the sensor and is a primary concern in engineering applications, thus it is assigned a weight of 0.3; stability reflects the random noise level of the sensor, thus it is assigned a weight of 0.2; and range reflects extreme fluctuations and is assigned a weight of 0.1 as an auxiliary indicator. Overall Score The range of values ​​is A higher score indicates better sensor drift performance.

[0067] S4.2: Drift Level Determination The drift level is determined based on the final overall score, and the determination criteria are shown in Table 1: Table 1 Drift Level Determination Criteria

[0068] S4.3: Output Results Output a four-dimensional score (bias score) for each sensor. Drift rating Stability score Scope of scoring ), Overall score The drift level is visualized and exported as a deviation graph and evaluation report.

[0069] V. System Implementation Examples This invention also provides a rapid evaluation system for the conductivity drift performance of batch CTD sensors implementing the above method, such as... Figure 2 As shown, it includes the following modules: 1. Multi-channel synchronous data acquisition module This module is used to simultaneously acquire conductivity data from no fewer than 20 temperature, salinity, and depth sensors at a sampling rate of no less than 1 Hz. It supports multiple sensor interface protocols (such as RS-232, RS-485, USB, Ethernet, etc.), and can automatically identify and connect multiple sensors to achieve synchronous triggering of data acquisition and real-time transmission.

[0070] 2. Data Preprocessing Module The functions configured to perform step S1 include automatic identification and extraction of conductivity columns, text cleaning, and 3D-based... The module performs outlier filtering and quality screening of raw sensor data. It can automatically identify the conductivity field in the raw data file, extract pure numerical data, remove outliers exceeding limits, and mark sensors with severe initial attenuation.

[0071] 3. Historical Error Variance and Weight Calculation Module The functions used to perform steps S2.1 and S2.2 include historical error calculation, sliding window variance estimation, statistical weight calculation, low-confidence sensor weight attenuation, and weight normalization. This module receives the effective conductivity data output by the data preprocessing module in real time and updates the error variance and fusion weights of each sensor online.

[0072] 4. Improved Robust Fusion Module for Fusion with Fuzzy C-Means The module is configured to perform steps S2.1 to S2.3, including online weight updates based on historical error variance, robust fuzzy C-means fusion coupling membership and weights, and generation of virtual ground truth. This module sets the number of clusters. The objective function is optimized through alternating iterations, and the virtual truth value at the current time step is output. .

[0073] 5. Virtual True Value Median Correction Module The module is configured to perform the function of step S2.4, including calculating the population median and the median of the virtual truth value, and correcting the overall bias of the virtual truth value sequence. This module outputs the corrected virtual truth value. .

[0074] 6. Accumulated Deviation and Regression Analysis Module The module is configured to perform step S3, which includes deviation sequence calculation, univariate linear regression analysis, and long-term drift rate estimation. This module performs regression analysis independently for each sensor and outputs the long-term drift rate. and initial bias .

[0075] 7. Multi-dimensional quantitative scoring and report generation module The function configured to perform step S4 includes four-dimensional drift scoring, comprehensive score calculation, drift level determination, and export of visualized results including deviation curves and evaluation reports. This module can generate a four-dimensional score radar chart for each sensor, a deviation-over-time curve, and a detailed report containing all evaluation indicators, facilitating viewing and analysis by engineers.

[0076] VI. Application Examples The method and system of the present invention will be described below through a specific application example.

[0077] In the factory testing of a certain batch of CTD sensors, there were a total of 30 sensors under test (i.e. (), installed in a standard constant temperature water bath. Through a multi-channel synchronous data acquisition module, with... The sampling frequency is continuously collected for 600 seconds (i.e. Second, This yields a raw conductivity data matrix with 30 rows and 600 columns.

[0078] After data cleaning, outlier filtering, and quality screening in steps S1.2 to S1.4, all data from the 30 sensors were valid, and no sensor was marked as having severe initial attenuation.

[0079] In step S2, the length of the sliding window is set. adjustment coefficient Median correction intensity coefficient Fusion generates a virtual truth sequence. And corrected to obtain .

[0080] In step S3, the bias sequence for each sensor is calculated and a univariate linear regression analysis is performed. Taking a certain sensor as an example, its initial bias... Long-term drift rate Standard deviation of the bias sequence Deviation range .

[0081] In step S4, the scores for each dimension of the sensor are calculated: Bias score: ; Drift rating: ; Stability rating: ; Range rating: ; Overall Score: ; The sensor's overall score ,belong The range is classified as "moderate drift," which in engineering terms means a significant performance degradation. Calibration is recommended before use.

[0082] The evaluation of all 30 sensors was completed within minutes of data acquisition (mainly due to 600 seconds of data acquisition time plus approximately 5 seconds of algorithm computation time), which significantly improved evaluation efficiency compared to the traditional method of comparing each sensor one by one (approximately 20 minutes per sensor, totaling approximately 600 minutes for 30 sensors).

[0083] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A rapid evaluation method for the conductivity drift performance of batch CTD sensors, characterized in that, Includes the following steps: Step S1: Synchronously collect the conductivity time series data of multiple temperature, salinity and depth sensors deployed in the test environment, and perform text cleaning, outlier filtering and quality screening on the raw conductivity data to obtain the preprocessed effective conductivity data. Step S2: Based on the historical error variance within the sliding window, estimate the statistical weights of each sensor online, and use an improved robust fuzzy C-means fusion algorithm to fuse the conductivity data of all sensors at the same time to generate virtual true values. Then, perform bias correction on the virtual true value sequence based on the population median. Step S3: Calculate the deviation sequence of each sensor relative to the calibrated virtual true value, perform univariate linear regression on the deviation sequence of each sensor, and estimate its long-term drift rate and initial bias. Step S4: Quantify the scores from four dimensions: bias, drift rate, stability, and range, calculate the weighted comprehensive score, and automatically determine and output the drift level; Step S1 specifically includes: S1.1: Simultaneously collect conductivity time-series data from no less than 20 temperature, salinity, and depth sensors at a sampling frequency of no less than 1Hz, with a continuous collection time of no less than 10 minutes; S1.2: Identify and extract the pure conductivity data column from the raw data column of each sensor, and take the arithmetic mean of multiple conductivity measurement columns at the same time as the final conductivity measurement value of the sensor at time t, denoted as . ; S1.3: Outlier filtering is performed based on the 3σ criterion, and the global mean of conductivity values ​​of all sensors at all times is calculated. and global standard deviation , will satisfy or The data points were identified as outliers and removed. S1.4: Perform raw sensor data quality screening and calculate the global median of conductivity values ​​for all sensors. and the median of each sensor ,when The sensor was determined to have severe initial attenuation, and its initial weight was reduced in subsequent fusion processes. The virtual truth value generation in step S2 specifically includes the following sub-steps: S2.1: Online estimation of historical error variance: Set the length of the sliding window The sensor is updated online using the following formula. At any moment Historical error variance : ; Among them, error , The virtual truth value of the previous moment. The value range is 20-100; S2.2: Calculation of statistical weights: Based on the historical error variance, the sensor value is calculated using the following formula. At any moment Fusion weights , ; in, This represents the total number of sensors; S2.3: Improved Robust Fusion of Fusion in Fusion of ... Set the number of clusters Design the coupling coefficient between the membership index and the sensor weights. : ; in, This is the adjustment coefficient; Initialize normal class centers using the weighted average and the upper quartile, respectively. and noise center ; Then, the membership matrix is ​​updated alternately. and cluster center Minimize the objective function: ; in, For the first Data points Belongs to the The membership degree of each cluster, For the first Data points To the cluster center Euclidean distance, For the first Membership index coupling coefficient of each data point For the first The centers of each cluster are identified; after iterative convergence, the cluster center with the largest sum of membership degrees is taken as the virtual truth value at the current time step. .

2. The method according to claim 1, characterized in that, In step S2.3, the membership update formula is: ; in, For the first The data point belongs to the th data point The membership degree of each cluster, and All values ​​represent the Euclidean distance from the data point to the corresponding cluster center. For the first Membership index coupling coefficient of each data point; The formula for updating cluster centers is: ; in, For the first Cluster centers, For the first Measurement values ​​from each sensor; When the maximum change in the membership matrix between two consecutive iterations is less than a preset threshold The iteration terminates at this point: ; in, This represents the membership value updated in the current iteration. This is the membership value from the previous iteration. At the same time, the maximum number of iterations is set to 100 to prevent the algorithm from failing to converge.

3. The method according to claim 1, characterized in that, In step S2, the bias correction of the virtual truth sequence based on the population median specifically includes the following sub-steps: S2.4.1: Calculate the global median of conductivity values ​​for all sensors at all times: ; in, The total number of sensors, This represents the total number of sampling times. S2.4.2: Calculate the median of the dummy truth sequence: ; The correction amount is calculated according to the following formula. : ; in, The median is the corrected intensity coefficient; S2.4.3: Apply global bias correction to the virtual truth sequence: , ; in, For a moment The corrected virtual truth value.

4. The method according to claim 1, characterized in that, Step S3 specifically includes: S3.1: Calculate each sensor At each sampling time Deviation: ; The deviations at all times are arranged in chronological order to form a deviation sequence. ; S3.2: Establish a univariate linear regression model for the deviation sequence of each sensor: ; in, For long-term drift rate, For the initial bias, The random error term follows a normal distribution with a mean of 0. Long-term drift rate estimated using the least squares method and initial bias : ; ; in, The mean value at each sampling time. This is the mean of the biased sequence.

5. The method according to claim 1, characterized in that, Step S4 involves quantitative scoring across four dimensions, specifically including: Bias score is used to assess the initial systematic bias of the sensor; ; Drift score is used to assess the long-term drift trend of a sensor; ; Stability score is used to assess the dispersion of sensor measurements; ; The numerator is the standard deviation of the biased sequence. ; Range scoring is used to assess whether there are localized, drastic jumps; ; in, That is, the range of the biased sequence. ; Overall Score: 。 6. The method according to claim 1, characterized in that, The drift level determination criteria in step S4 are as follows: Overall Score The time frame is determined to be "stable," which in engineering terms means that the sensor is performing well and can be used directly. The timeframe indicates "slight drift," which in engineering terms means the performance is acceptable and it is recommended to monitor it during use. The timeframe indicates "moderate drift," which in engineering terms means a significant performance degradation and suggests calibration before use. The time was determined to be "severe drift", which means that the performance is unqualified and it is recommended to rework or discard it.

7. The method according to claim 1, characterized in that, In step S2, for sensors marked as having severe initial attenuation in step S1.4, their fusion weights are multiplied by an attenuation factor. , The range of values ​​is .

8. A rapid evaluation system for the conductivity drift performance of batch CTD sensors implementing the method of any one of claims 1 to 6, characterized in that, include: Multi-channel synchronous data acquisition module: used to synchronously acquire conductivity data from multiple temperature, salinity and depth sensors at a sampling rate of not less than 1Hz, supports multiple sensor interface protocols and automatically identifies connections; Data preprocessing module: used to realize automatic identification of conductivity columns, text cleaning, outlier filtering based on the 3σ criterion, and quality screening of raw sensor data; Historical error variance and weight calculation module: used to perform historical error calculation, sliding window variance estimation, statistical weight calculation, low-confidence sensor weight attenuation and weight normalization processing; Improved robust fuzzy C-means fusion module: used to realize online weight update based on historical error variance, robust fuzzy C-means fusion with membership degree and weight coupling, and generation of virtual truth value; Virtual Truth Median Correction Module: Used to calculate the population median and the virtual truth median, as well as to correct the overall bias of the virtual truth sequence; The deviation accumulation and regression analysis module is used to perform deviation sequence calculation, univariate linear regression analysis, and long-term drift rate estimation. Multi-dimensional quantitative scoring and report generation module: used to realize four-dimensional drift scoring, comprehensive score calculation, drift level determination, and export of visualized results including deviation curves and evaluation reports.

Citation Information

Patent Citations

  • Data correction method of ocean temperature-salinity-depth sensor

    CN118211103A

  • Argo conductivity data correction method based on thermal-halocline internal wave influence

    CN119510904A