A Method for Identifying Temperature-Salinity-Depth Data Affected by Turbulence in a Portable Underwater Glider
By using XGBoost algorithm model and feature matrix construction technology, the problem of difficult identification of temperature and salt deep data under the influence of turbulence by portable underwater gliders is solved, and the accurate identification of data affected by turbulence and the authenticity of data is achieved.
Patent Information
- Application Number
- CN202510058953.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-01-15
AI Technical Summary
When a portable underwater glider operates in the ocean, the temperature and salt depth data have errors or singular values due to turbulence, and it is difficult for the prior art to accurately identify data affected by turbulence.
Using the XGBoost algorithm model, a characteristic matrix is constructed by calculating the strain rate tensor, vortex tensor, pulsating strain rate tensor and pulsating vortex tensor of data points outside the threshold range of motion speed, and a model is trained to identify data affected by turbulence.
Accurate identification of data affected by turbulence is achieved, error judgments are avoided during the temperature-salt jump layer, and the authenticity of temperature-salt deep data is improved.
Smart Images

Figure CN119474888B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital signal processing, and particularly to a method for identifying temperature, salinity, and depth data affected by turbulence in a portable underwater glider. Background Art
[0002] Due to the development of marine technology, portable underwater gliders are more widely used, mainly for measuring ocean temperature, salinity, and depth data. However, when an underwater glider operates in the ocean, it will pass through the thermocline and be affected by turbulence during movement, which easily causes errors or singular values in the obtained temperature, salinity, and depth data.
[0003] Currently, for the data identification of portable underwater gliders, the existence of spikes or singular values is mainly determined by conventional statistical methods, and the temperature, salinity, and depth data are processed by median filtering and smoothing filtering. However, the temperature, salinity, and depth data processed by the above methods can only identify outliers with large fluctuations in the thermocline and cannot identify data affected by turbulence. And if the above traditional methods are used to identify data points affected by turbulence, not only will the correct data when passing through the thermocline be misidentified, but also the problem of distortion of temperature, salinity, and depth data due to inaccurate parameter settings will occur. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a method for identifying temperature, salinity, and depth data affected by turbulence in a portable underwater glider, so as to accurately identify data affected by turbulence.
[0005] To achieve the above object, the technical solution of the present invention is as follows:
[0006] A method for identifying temperature, salinity, and depth data affected by turbulence in a portable underwater glider includes the following steps:
[0007] Step 1, obtaining the temperature, salinity, and depth data collected by the portable underwater glider and performing preprocessing;
[0008] Step 2, calculating the movement speed of the portable underwater glider underwater according to the obtained depth data and determining the speed threshold;
[0009] Step 3, for data points outside the speed threshold range, calculating the strain rate tensor, vorticity tensor, pulsating strain rate tensor, and pulsating vorticity tensor, and jointly constructing a feature matrix with the temperature, salinity, and depth data;
[0010] Step 4, model training: inputting the constructed feature matrix into the XGBoost algorithm model, the model outputs the identified data affected by turbulence, using the real temperature, salinity, and depth data affected by turbulence to calculate the loss, and adjusting the model parameters through an optimization algorithm to obtain a trained model;
[0011] Step 5, evaluate the evaluation performance of the model at different depths through recall rate, precision rate, and F1 score metrics;
[0012] Step 6, input the temperature, salinity, and depth data collected by the portable underwater glider to be recognized into the qualified model, and the model outputs the data affected by turbulence.
[0013] In the above solution, the salinity data is the collected conductivity data, and the depth data is the collected pressure data.
[0014] In the above solution, in step 1, the preprocessing method is as follows: fill in the missing data by linear interpolation, divide the overall temperature, conductivity, and pressure data into up and down segments, respectively obtain the two segments of data for diving and surfacing, eliminate obvious abnormal data points, eliminate the conductivity data exceeding 0 - 6 S / m, eliminate the pressure data less than or equal to 0 dbar, and eliminate the temperature data outside -2.5 - 40 °C.
[0015] In the above solution, in step 2, the calculation formula for the movement speed of the portable underwater glider underwater is as follows:
[0016] ;
[0017] Wherein, represents the movement speed of the portable underwater glider, respectively represent the depth data at the current moment and the depth data at the previous moment, represents the sampling frequency.
[0018] In the above solution, the method for determining the speed threshold range is as follows: according to the obtained movement speed , calculate the average value u and variance α of the speed, and the speed threshold range S is: u - α ≤ S ≤ u + α. The data points outside the speed threshold range refer to the data points less than u - α and greater than u + α.
[0019] In the above solution, the calculation formula for the strain rate tensor is as follows:
[0020] ;
[0021] The calculation formula for the vorticity tensor is as follows:
[0022] ;
[0023] The calculation formula for the pulsating strain rate tensor is as follows:
[0024] ;
[0025] The calculation formula for the pulsating vorticity tensor is as follows:
[0026] ;
[0027] wherein, represent two different directions on the space coordinate system; are respectively the velocity components in the direction of; are respectively the space components in the direction of; represents the rate of change of the velocity component ; represents the rate of change of the velocity component ; represents the rate of change of the space component ; represents the rate of change of the space component ; and are respectively the pulsating parts of the velocity component , that is, the velocity pulsation components; represents the rate of change of the velocity pulsation component , represents the rate of change of the velocity pulsation component .
[0028] In the above solution, each time the model identifies the input data, at each node of each decision tree, XGBoost will evaluate all available features and select a feature for classification to achieve the minimum value of the objective function; starting from an empty decision tree, all data are at the root node. For the current node, XGBoost will traverse all features and possible split points and calculate the split gain Gain of each split point:
[0029] ;
[0030] wherein, are respectively the sum of the first-order derivatives of the left and right child nodes after splitting, are respectively the sum of the second-order derivatives of the left and right child nodes after splitting; is the regularization parameter used to control the complexity of the model; is the regularization parameter used to control the number of leaf nodes;
[0031] If the split gain Gain is positive, it can be split. If it is negative or zero, no splitting will be performed.
[0032] In a further technical solution, the objective function is as follows:
[0033] ;
[0034] where n represents the total number of samples, and i represents the i-th sample, is the recognition result of the (t - 1)-th decision tree for turbulence, represents the eigenvalue of the i-th sample, is the prediction function of the t-th decision tree; represents the error function, is the actually observed turbulence state of the i-th sample, is the regularization function, and the formula is as follows:
[0035] ;
[0036] where T represents the number of decision trees, j represents the j-th decision tree, represents the node value of the j-th decision tree;
[0037] By splitting and partitioning the nodes of the decision tree, the minimum node value can be obtained, so that the objective function achieves the optimal solution.
[0038] Through the above technical solution, a method for identifying turbulence-affected temperature, salinity, and depth data of a portable underwater glider provided by the present invention has the following beneficial effects:
[0039] 1. The present invention uses the XGBoost algorithm model to identify turbulence data. This algorithm will reconstruct the decision tree for the data points that do not perfectly identify turbulence. The features selected during the splitting of each decision tree are different, which is to better fit the residuals. Each decision tree will present the final recognition result. However, due to the existence of the learning rate, even if a certain decision tree is not accurately recognized, it will not dominate the entire result. The final result is jointly constructed by the recognition of all decision trees. Therefore, this method can effectively and accurately identify the data points affected by turbulence;
[0040] 2. Compared with the conventional data statistical method, due to the setting of the speed threshold, this method can only process the data points outside the speed threshold range, and can avoid misjudgment when passing through the thermocline;
[0041] 3. The present invention selects 7 feature quantities to construct the feature matrix, which can provide a more comprehensive description of the flow characteristics, help the model identify turbulence from multiple angles, and avoid the problem of inaccurate turbulence identification caused by selecting a single feature quantity. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.
[0043] Figure 1 is a schematic flow chart of a method for identifying turbulence-affected temperature, salinity, and depth data of a portable underwater glider disclosed in an embodiment of the present invention;
[0044] Figure 2 A flowchart for a decision tree;
[0045] Figure 3 The temperature data obtained from the portable underwater glider at 0-100 meters is obtained by the XGBoost algorithm model, which is affected by turbulence.
[0046] Figure 4 The conductivity data obtained by the portable underwater glider from 0-100 meters are used to obtain the data points affected by turbulence through the XGBoost algorithm model. DETAILED DESCRIPTION
[0047] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0048] The present invention provides a method for identifying temperature, salinity and depth data of turbulence effects on a portable underwater glider. Figure 1 As shown, the following steps are included:
[0049] Step 1, obtaining temperature, salinity and depth data collected by a portable underwater glider and preprocessing them;
[0050] The salinity data collected by the portable underwater glider is conductivity data, and the depth data is pressure data.
[0051] The preprocessing method is as follows: fill in the missing data through linear interpolation, divide the overall temperature, conductivity and pressure data into up and down segments, obtain the diving and floating data respectively, and eliminate obvious abnormal data points: eliminate conductivity data exceeding 0-6S / m, eliminate pressure data less than or equal to 0dbar, and eliminate temperature data outside of -2.5-40℃, to ensure that the data for analysis have no obvious abnormal values.
[0052] Step 2, calculating the underwater movement speed of the portable underwater glider according to the acquired depth data, and determining a speed threshold;
[0053] The calculation formula for the underwater speed of a portable underwater glider is as follows:
[0054] ;
[0055] in, represents the moving speed of the portable underwater glider, Respectively represent the depth data of the current moment and the depth data of the previous moment, Indicates the sampling frequency.
[0056] Since the speed of the portable underwater glider does not change significantly without the influence of turbulence, the obtained movement speed , calculate the average value u and variance α of the speed, and the speed threshold range S is: u-α≤S≤u+α. Data points outside the speed threshold range refer to data points that are less than u-α and greater than u+α.
[0057] The speed threshold setting is not fixed. The threshold can be changed by judging the performance of the algorithm through the recall rate, precision rate and F1 score. By introducing the speed threshold setting, it can be ensured that the mutation point when passing through the thermohaline layer will not affect the judgment of whether it is turbulence or not, which can effectively ensure the authenticity of the data.
[0058] Step 3: For data points outside the velocity threshold range, the strain rate tensor, vortex tensor, pulsating strain rate tensor and pulsating vortex tensor are calculated, and a feature matrix is constructed together with the temperature, salinity and depth data.
[0059] The selection of characteristic quantities in the XGBoot algorithm uses a combination of original data and flow invariants. The original data are temperature, pressure, and conductivity. The invariants are strain rate tensor and vortex tensor, pulsating strain rate tensor, and pulsating vortex tensor.
[0060] The strain rate tensor contains the tensile and shear deformation information of the fluid. The calculation formula of the strain rate tensor is as follows:
[0061] ;
[0062] The vortex tensor describes the rotational motion of the fluid element. The calculation formula of the vortex tensor is as follows:
[0063] ;
[0064] The pulsating strain rate tensor is used to describe the local deformation rate in turbulence. The calculation formula of the pulsating strain rate tensor is as follows:
[0065] ;
[0066] The pulsating vortex tensor is used to describe the local rotation rate in turbulence. The calculation formula of the pulsating vortex tensor is as follows:
[0067] ;
[0068] in, Represents two different directions in the spatial coordinate system; They are The velocity component in the direction; They are The spatial component of direction; Represents the velocity component Rate of change; Represents the velocity component Rate of change; Represents the spatial component Rate of change; Represents the spatial component Rate of change; And Are respectively the pulsating parts of the velocity components That is, the velocity pulsation components; Represents the velocity pulsation component Rate of change, Represents the velocity pulsation component Rate of change.
[0069] Step 4, Model training: Input the constructed feature matrix into the XGBoost algorithm model. The model outputs the data identified as being affected by turbulence. Use the real data of temperature, salinity, and depth affected by turbulence to calculate the loss, and adjust the model parameters through an optimization algorithm to obtain a trained model.
[0070] Each time the model identifies the input data, at each node of each decision tree, XGBoost evaluates all available features and selects a feature for classification to achieve the minimum value of the objective function; starting from an empty decision tree, all data are at the root node. For the current node, XGBoost traverses all features and possible split points, and calculates the split gain Gain for each split point:
[0071] ;
[0072] Among them, Are respectively the sums of the first-order derivatives of the left and right child nodes after splitting, Are respectively the sums of the second-order derivatives of the left and right child nodes after splitting; Is a regularization parameter used to control the complexity of the model; Is a regularization parameter used to control the number of leaf nodes;
[0073] The split gain is an index to measure the splitting effect, which measures the reduction in loss before and after splitting of a certain node. It is calculated based on the reduction in the objective function. If the split gain Gain is positive, it can be split; if it is negative or zero, no splitting is performed.
[0074] The objective function is as follows:
[0075] ;
[0076] Among them, n represents the total number of samples, i represents the i-th sample, is the recognition result of the (t-1)-th decision tree for turbulence, that is, the sum of the prediction results of the previous decision trees; represents the eigenvalue of the i-th sample, is the prediction function of the t-th decision tree; represents the error function, is the actually observed turbulence state of the i-th sample, is the regularization function, and the formula is as follows:
[0077] ;
[0078] where T represents the number of decision trees, j represents the j-th decision tree, represents the node value of the j-th decision tree;
[0079] The regularization function can limit the maximum depth of the decision tree to ensure the complexity of the algorithm, ensure that the result can be calculated with the minimum complexity, and prevent overfitting. After simplification, the objective function can be regarded as a quadratic function. By splitting and partitioning the nodes of the decision tree, the minimum node value can be obtained, so that the objective function can achieve the optimal solution.
[0080] The final output of each decision tree in this algorithm is to output a probability value between 0 and 1 for each data point according to the two states of turbulence and non-turbulence, indicating the possibility that the point is affected by turbulence. This probability value is obtained by accumulating the prediction results of all decision trees. Finally, in order to reasonably record the data points affected by turbulence, the data points affected by turbulence are finally determined by setting the threshold. Finally, according to the setting, label 1 means affected by turbulence, and label 0 means not affected by turbulence. It is stored in a table containing temperature, salinity, and depth data.
[0081] As Figure 2 shown, for example, at the root node, the strain rate tensor can provide the maximum splitting gain, and this feature is used for node splitting, splitting into child node 1 and child node 2. The data points in child node 1 are turbulent, and the data in child node 2 are non-turbulent. Whether it is the data points in node 1 or node 2, they will be split according to the feature that can provide the maximum splitting gain. According to Figure 2 shown, in the splitting process, non-turbulent points may contain turbulence, and turbulent points may contain non-turbulence. This is why different features among the 7 features (such as feature 2 or feature 3) are used for continuous splitting and refinement until the data points end or reach the maximum depth of the decision tree (the maximum depth and complexity of the decision tree are controlled by the regularization function). Each decision tree is constructed for the data points that the previous decision tree fails to perfectly identify, that is, fitting the residuals, continuously optimizing the previous decision tree, and achieving the best recognition.
[0082] Step 5, evaluate the performance of the model at different depths through the metrics of recall, precision, and F1 score.
[0083] Since the temperature-salinity-depth data has depth dependence, the performance of the model at different depth layers needs to be considered when evaluating the model. For example, the recall, precision, and F1 score of the model at different depth layers can be analyzed to see if the model can accurately capture the turbulent effects at each depth level. When evaluating, the data needs to be divided into different depth intervals, and the recall, precision, and F1 score within each interval are calculated separately. Then, while ensuring the overall evaluation is qualified, ensure that the evaluation of each depth layer is also qualified.
[0084] ;
[0085] R is the recall rate, TP is the number of data points correctly predicted as turbulent by the model, and FN is the number of turbulent data points mispredicted as non-turbulent by the model.
[0086] ;
[0087] P is the precision rate, and FP is the number of non-turbulent data points mispredicted as turbulent by the model.
[0088] ;
[0089] The F1 score is the harmonic mean of the precision and recall rates, which comprehensively considers the balance between precision and recall.
[0090] By comprehensively evaluating the recall, precision, and F1 scores obtained from the temperature-salinity-depth data at different depth levels and the overall recall, precision, and F1 scores, it is possible to effectively avoid misjudgments of turbulent data points caused by different temperature-salinity-depth characteristics at different depths.
[0091] Step 6, input the temperature, salinity, and depth data collected by the portable underwater glider to be identified into the qualified model, and the model outputs the data affected by turbulence. As Figure 3 and Figure 4 shown, the data points affected by turbulence obtained from the temperature and conductivity data of the portable underwater glider at 0 - 100 meters by the XGBoost algorithm model; where the circles represent the data points affected by turbulence. The data affected by turbulence is processed using a third-order low-pass filter to reduce the influence of turbulence on the temperature-salinity-depth data.
[0092] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for identifying temperature, salinity and depth data of turbulence effects on a portable underwater glider, characterized in that: The steps include: Step 1, obtaining temperature, salinity and depth data collected by a portable underwater glider and preprocessing them; Step 2, calculating the underwater movement speed of the portable underwater glider according to the acquired depth data, and determining a speed threshold range; Step 3, for data points outside the velocity threshold range, the strain rate tensor, vortex tensor, pulsating strain rate tensor and pulsating vortex tensor are calculated, and a feature matrix is constructed together with the temperature, salinity and depth data; Step 4, model training: input the constructed feature matrix into the XGBoost algorithm model, the model outputs the identified data affected by turbulence, uses the real data of temperature, salinity, and depth affected by turbulence to calculate the loss, and adjusts the model parameters through the optimization algorithm to obtain a trained model; Step 5: Evaluate the recognition performance of the model at different depths through recall, precision, and F1 score indicators; Step 6, the temperature, salinity and depth data collected by the portable underwater glider to be identified are input into the qualified model, and the model outputs the data affected by turbulence.
2. The method for identifying temperature, salinity and depth data of turbulence effects of a portable underwater glider according to claim 1, characterized in that: The salinity data is the collected conductivity data, and the depth data is the collected pressure data.
3. The method for identifying turbulence-affected temperature, salinity and depth data of a portable underwater glider according to claim 2, characterized in that: In step 1, the preprocessing method is as follows: fill in the missing data by linear interpolation, divide the overall temperature, conductivity and pressure data into up and down segments, obtain the diving and floating data respectively, eliminate obvious abnormal data points, eliminate the conductivity data exceeding 0 to 6 S / m, eliminate the pressure data less than or equal to 0 dbar, and eliminate the temperature data outside the temperature range of -2.5 to 40 °C.
4. The method for identifying turbulence-affected temperature, salinity and depth data of a portable underwater glider according to claim 2, characterized in that: In step 2, the calculation formula for the underwater movement speed of the portable underwater glider is as follows: ; in, represents the moving speed of the portable underwater glider, Respectively represent the depth data of the current moment and the depth data of the previous moment, Indicates the sampling frequency.
5. The method for identifying temperature, salinity and depth data of turbulence effects of a portable underwater glider according to claim 4, characterized in that: The method for determining the speed threshold range is as follows: Based on the obtained motion speed , calculate the average value u and variance α of the speed, the speed threshold range S is: u-α≤S≤u+α, and the data points outside the speed threshold range are data points less than u-α and greater than u+α.
6. The method for identifying temperature, salinity and depth data of turbulence effects of a portable underwater glider according to claim 1, characterized in that: The strain rate tensor is calculated as follows: ; The calculation formula of the vortex tensor is as follows: ; The calculation formula of the pulsating strain rate tensor is as follows: ; The calculation formula of the pulsation vortex tensor is as follows: ; in, Represents two different directions in the spatial coordinate system; They are The velocity component in the direction; They are The spatial component of direction; Represents the velocity component The rate of change of Represents the velocity component The rate of change of Represents spatial component The rate of change of Represents spatial component The rate of change of and The velocity components are The pulsating part, namely the velocity pulsating component; Indicates the velocity pulsation component The rate of change, Indicates the velocity pulsation component The rate of change.
7. The method for identifying temperature, salinity and depth data of turbulence effects of a portable underwater glider according to claim 1, characterized in that: Each time the model recognizes the input data, at each node of each decision tree, XGBoost will evaluate all available features and select a feature for classification to achieve the minimum value of the objective function; starting from an empty decision tree, all data are at the root node. For the current node, XGBoost will traverse all features and possible split points and calculate the split gain Gain of each split point: ; in, are the sum of the first-order derivatives of the left and right child nodes after splitting, are the sum of the second-order derivatives of the left and right child nodes after splitting; is a regularization parameter used to control the complexity of the model; is a regularization parameter used to control the number of leaf nodes; If the split gain Gain is positive, splitting will occur; if it is negative or zero, no splitting will occur.
8. The method for identifying temperature, salinity and depth data of turbulence effects of a portable underwater glider according to claim 7, characterized in that: The objective function is as follows: ; Where n represents the total number of samples, i represents the i-th sample, is the recognition result of turbulence by the t-1th decision tree; represents the eigenvalue of the i-th sample, is the prediction function of the t-th decision tree; represents the error function, is the actual observed turbulence state of the i-th sample, is the regularization function, and the formula is as follows: ; Where T represents the number of decision trees, j represents the jth decision tree, Represents the node value of the jth decision tree; By splitting and dividing the nodes of the decision tree, the minimum node value can be obtained, so that the objective function can achieve the optimal solution.
Citation Information
Patent Citations
Marine abnormal mesoscale vortex identification method and device and storage medium
CN115310051A
Mesoscale vortex observation method based on AUG reinforcement learning
CN119066982A