A method for constructing a marine dissolved oxygen concentration reconstruction model based on Argo temperature-salinity profiles

The marine dissolved oxygen concentration reconstruction model based on Argo temperature-salinity profile solves the problem of sparse and uneven marine dissolved oxygen observation data, and realizes efficient and accurate dissolved oxygen concentration prediction and low-cost sampling method, which is suitable for large-scale off-site sampling.

CN116864026BActive Publication Date: 2025-12-12AEROSPACE INFORMATION RES INST CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310062810.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-18
Publication Date
2025-12-12
Estimated Expiration
2043-01-18

AI Technical Summary

Technical Problem

Existing technologies have resulted in sparse and unevenly distributed ocean dissolved oxygen concentration observation data, leading to insufficient understanding of the global ocean dissolved oxygen distribution and its influence by physical processes. Furthermore, traditional shipborne observations are costly and highly susceptible to extreme weather conditions.

Method used

The marine dissolved oxygen concentration reconstruction model based on the Argo temperature-salinity profile is constructed by screening dissolved oxygen data of the same month in previous years from Argo, performing interpolation calculations and regression prediction model construction, using temperature-salinity data to correct dissolved oxygen concentration, constructing a spatiotemporal classification and zoning of marine dissolved oxygen, and improving data density and prediction accuracy.

Benefits of technology

It improves the problem of data sparsity in the Argo profile database at different latitude and longitude coordinates, increases the accuracy of dissolved oxygen concentration prediction, reduces workload and cost, improves sampling efficiency and flexibility, and adapts to large-scale off-site sampling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116864026B_ABST
    Figure CN116864026B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of marine information, and more particularly to a method for constructing a marine dissolved oxygen concentration reconstruction model based on an Argo temperature-salinity profile, comprising the following steps: obtaining monthly scale dissolved oxygen data at an arbitrary target depth; constructing a marine dissolved oxygen spatio-temporal classification and partition at different target depths; constructing a regression prediction model for the corresponding spatio-temporal classification and partition; completing the reconstruction of monthly scale temperature-salinity data at the corresponding spatio-temporal classification and partition and target depth, and constructing a marine dissolved oxygen concentration reconstruction model. The present application overcomes the defect of the discrete and sparse data in the Argo system at each latitude and longitude coordinate on the target depth plane through spatial interpolation; a regression prediction model is constructed using the coupling relationship between temperature and salinity and dissolved oxygen concentration at the same position, a marine dissolved oxygen concentration reconstruction model is constructed based on the reconstruction of dissolved oxygen concentration using temperature and salinity at the corresponding position, and the influence of data sparseness on interpolation accuracy is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of ocean information, and particularly relates to a method for constructing a marine dissolved oxygen concentration reconstruction model based on an Argo temperature-salinity profile. BACKGROUND

[0002] Marine dissolved oxygen (DOXY) is oxygen dissolved in water, which provides the necessary biochemical environment for marine life and is an important material for marine life activities. The dissolved oxygen concentration in seawater is not only a main index for measuring seawater quality, evaluating marine ecological environment, and an important basis for marine scientific experiments and resource exploration, but also a necessary parameter for understanding marine biogeochemical processes, global climate change and marine carbon cycle.

[0003] The existing global dissolved oxygen concentration spatial grid data is mainly based on ship measurement, anchor buoy, and underwater intelligent detection equipment, and the data has poor continuous updating capability. Among them, ship measurement is the most important way, such as water sample collection by CTD water sampler, and then chemical analysis titration method is used to determine the continuous (discrete) water column sample. The ship measurement method in the prior art has the defects of too low sampling rate and insufficient spatio-temporal resolution, and the time cost and economic cost of shipborne observation are very high. The ship measurement method is limited by extreme weather, and in the conditions of severe sea conditions, polar sea areas and the like, the observation data is very scarce. At present, the observation data of marine dissolved oxygen is still relatively sparse and unevenly distributed in time and space, which greatly limits people's understanding and understanding of the spatio-temporal distribution of global marine dissolved oxygen and its influence on physical processes. Therefore, there is an urgent need in the market for dissolved oxygen observation and analysis data processing methods covering more time and space ranges. SUMMARY

[0004] In view of the above analysis, the present application aims to provide a method for constructing a marine dissolved oxygen concentration reconstruction model based on an Argo temperature-salinity profile, to solve at least one of the existing technical problems.

[0005] The main purpose of the present application is achieved by the following technical solutions:

[0006] The present application provides a method for constructing a marine dissolved oxygen concentration reconstruction model based on an Argo temperature-salinity profile, comprising:

[0007] Based on the Argo dissolved oxygen profile data of the same month in previous years, climate state monthly dissolved oxygen data is obtained by screening; based on the climate state monthly dissolved oxygen data, monthly dissolved oxygen data of any target depth is obtained; the climate state monthly dissolved oxygen data contains dissolved oxygen data of different months and different depths in previous years;

[0008] The plane corresponding to each target depth is divided into a plurality of spatial units by using latitude and longitude coordinates. Based on the monthly scale dissolved oxygen data of the target depth, the dissolved oxygen concentrations of different spatial units at the same target depth are calculated by first interpolation to obtain the dissolved oxygen concentration of the center point of the spatial unit, which is taken as the dissolved oxygen concentration of the spatial unit. Based on the dissolved oxygen concentrations of each spatial unit, the spatial units with similar dissolved oxygen concentrations at the target depth are classified into the same region to complete the spatio-temporal classification and partition of the marine dissolved oxygen at the target depth. The different target depths are partitioned to construct the spatio-temporal classification and partition of the marine dissolved oxygen at different target depths.

[0009] The monthly scale temperature and salinity data at the same target depth are obtained by screening the Argo temperature and salinity profile data of the same month in previous years. The monthly scale temperature and salinity data and the monthly scale dissolved oxygen concentration data at the same spatio-temporal classification partition, the same target depth and the same latitude and longitude coordinates are selected to construct a regression prediction model of the spatio-temporal classification partition with temperature and salinity as variables and dissolved oxygen concentration as dependent variable.

[0010] Based on the monthly scale temperature and salinity data of the sampling points, the correction value of the dissolved oxygen concentration at the target depth of each sampling point in the corresponding spatio-temporal classification partition at different latitudes and longitudes is obtained from the regression prediction model. The sampling points are the data sampling points of the Argo system at different depths and latitude and longitude coordinates. Based on the correction value of the dissolved oxygen concentration at the target depth, the second interpolation calculation is performed to obtain the corrected dissolved oxygen concentration of the center point of the spatial unit, complete the reconstruction of the monthly scale temperature and salinity data in the corresponding spatio-temporal classification partition and target depth, and construct the marine dissolved oxygen concentration reconstruction model containing different target depths and different spatio-temporal partitions.

[0011] Preferably, the climate state monthly scale dissolved oxygen data is obtained, comprising:

[0012] Obtaining Argo annual dissolved oxygen profile data;

[0013] Obtaining the annual climate state monthly scale dissolved oxygen data Doxy of each month from the Argo annual dissolved oxygen profile data by taking month as the screening parameter. M (h); wherein M is 1-12, and h is the depth.

[0014] Preferably, the first interpolation calculation of the dissolved oxygen concentrations of different spatial units at the same target depth comprises: taking the center point of the spatial unit as the target point, taking the data points near the center point of the spatial unit as the adjacent sampling points, and giving different weights to the adjacent sampling points and the center point of the spatial unit based on the distance between the adjacent sampling points and the center point of the spatial unit to obtain the dissolved oxygen concentration of the center point of the spatial unit.

[0015] Preferably, the dissolved oxygen concentration of the target point satisfies the following formula:

[0016]

[0017] wherein i is the number of the adjacent sampling point, n is the number of the adjacent sampling points designated to perform interpolation, di is the distance from the i-th adjacent sampling point to the target point, Doxyi represents the dissolved oxygen concentration of the adjacent sampling point i designated to perform interpolation. i i wherein di is the distance from the i-th adjacent sampling point to the target point, Doxyi represents the dissolved oxygen concentration of the adjacent sampling point i designated to perform interpolation.

[0018] Preferably, the regression prediction model taking the temperature and salinity as the independent variables and the dissolved oxygen concentration as the dependent variable comprises:

[0019] Based on the monthly scale dissolved oxygen data of the same spatio-temporal classification partition and target depth and the monthly scale temperature and salinity data of the same target depth and the same month, a test set and a training set containing the independent variables of temperature, salinity, longitude, latitude, year and the dependent variable of dissolved oxygen concentration are constructed, wherein the data units of the test set and the training set contain one-to-one corresponding dependent variables and independent variables.

[0020] The training set is divided into k training subsets by random sampling, the training subsets are cut, the minimum mean square value of the training subset data after cutting is selected as the optimal cutting point, and a decision tree model of k decision sub-trees with different splitting structures is constructed; the mean value of the output values of each data point in each decision sub-tree at the optimal cutting point of each decision tree is obtained to obtain the prediction value of each decision tree; and the mean value of the prediction values of the k decision trees is taken as the prediction output by the regression prediction model.

[0021] The test set is substituted into the random forest regression model to obtain the prediction value of the dissolved oxygen concentration, and the model is evaluated in terms of accuracy based on the prediction value of the dissolved oxygen concentration and the true value of the dissolved oxygen concentration in the test set; if the accuracy meets the preset requirement, the training of the marine dissolved oxygen concentration reconstruction model of the partition is completed, and if the requirement is not met, the maximum depth of the decision tree and the number of decision trees are adjusted until the accuracy of the model meets the preset requirement.

[0022] Preferably, the construction of the decision tree model of k decision sub-trees with different splitting structures comprises:

[0023] The optimal cutting point corresponding to the minimum mean square value of the two parts after cutting of the training subset is obtained, wherein the mean square value RMSE satisfies the formula:

[0024]

[0025] R1 and R2 satisfy:

[0026] R1(j, s) = {x | x (j) ≤ s}, R2(j, s) = {x | x (j) > s}; ​

[0027] wherein, x is an input feature value, y is an output value; c1 is an average value of the left sub-region output; c2 is an average value of the right sub-region output; j is an optimal split variable, s is an optimal split point, x (j) represents a feature value of x=j; R1 is a left sub-region after splitting, R2 is a right sub-region after splitting; y i represents an output value of the i-th data point; min represents taking a minimum value of the min function adjacent to the right region;

[0028] Based on the average value of the output value of each data point in each decision sub-tree in each decision tree, the prediction value of each decision tree is obtained, and the decision tree prediction value c k satisfies:

[0029]

[0030] wherein, y i represents an output value of the i-th data point, N k represents the number of data points in the k-th training subset, c k represents the average value of the output value of the decision tree corresponding to the k-th training subset.

[0031] The average value of the k decision tree prediction values is taken as the prediction value output by the regression prediction model, and the average value c - satisfies:

[0032]

[0033] wherein, l≤k, represents in the l-th training subset, c l represents the average value of the output value of the decision tree corresponding to the l-th training subset.

[0034] Preferably, the optimal split point corresponding to the minimum value of the mean square value of the two parts after splitting of the training subset is obtained, comprising:

[0035] Any variable in the training subset is selected for multiple splitting, and the optimal split point s1 of the variable is obtained; the root mean square error RMSE1 corresponding to the optimal split point satisfies:

[0036]

[0037] wherein, c1 is an average value of the left sub-region output; c2 is an average value of the right sub-region output; R1 is a left sub-region after splitting, R2 is a right sub-region after splitting; y i represents an output value of the i-th data point; min represents taking a minimum value of the min function adjacent to the right region;

[0038] According to the same method, the rest of the variables in the training subset are sequentially cut, and the mean square deviation is calculated, the optimal cut point of each variable is obtained, the mean square deviations corresponding to the optimal cut points of each variable are compared, and the optimal cut point corresponding to the minimum mean square deviation is taken as the optimal cut point s of the training subset, and the corresponding variable is taken as the optimal cut variable j.

[0039] Preferably, the adjusting the maximum depth of the decision tree and the number of decision trees comprises a cross-validation method, and the cross-validation method comprises:

[0040] The training set is further divided into a training part, a validation part and a test part; the training part is used for model training, the validation part is used for adjusting parameters, and the test part is used for measuring the performance of the model; the training set is further divided into k training subsets, and the training subsets are evenly divided from the middle position;

[0041] Each training subset is validated once, and the remaining k-1 training subset data are used as the training part to obtain k models;

[0042] The average of the root mean square errors of the final validation parts of the k models is taken as the performance index corresponding to the k models, and the performance index corresponding to the decision tree parameters is recorded;

[0043] The k models are traversed, and the above steps are repeated, and the decision tree parameters corresponding to the optimal performance index are taken as the optimal parameters.

[0044] Preferably, the performance index corresponding to the model is a root mean square error; and the decision tree parameters comprise a maximum depth of the decision tree and a number of decision trees.

[0045] Preferably, the random sampling method comprises Bootstrap sampling.

[0046] Compared with the prior art, the present application can at least realize one of the following beneficial effects:

[0047] (1) The present application uses Argo temperature and salinity profile data to screen monthly scale temperature and salinity data at the same target depth, and uses the monthly scale temperature and salinity data to reconstruct the dissolved oxygen concentration of the spatial unit in the same partition, which improves the problem of large deviation between the interpolation calculation and the actual value of the dissolved oxygen concentration of the spatial unit caused by the local over-sparse data of the Argo profile database on different latitude and longitude coordinates, and improves the prediction accuracy of the dissolved oxygen concentration.

[0048] (2) The present application divides the plane corresponding to each target depth into different spatial units by using latitude and longitude coordinates, and obtains the dissolved oxygen concentration of different spatial units at the target depth by interpolating the dissolved oxygen concentration of the target depth with the dissolved oxygen concentration of different spatial units at the same target depth, and then obtains the dissolved oxygen concentration of different spatial units at different target depths, greatly making up for the lack of data sparsity of the existing Argo profile database at different latitude and longitude coordinates.

[0049] (3) The present application partitions the data points based on the dissolved oxygen concentration obtained by the first interpolation calculation of the dissolved oxygen concentration of different spatial units at the same target depth, and constructs a temperature-salinity-dissolved oxygen prediction model in the region with close dissolved oxygen concentration; on the basis of introducing the interpolation calculation correction of temperature-salinity-dissolved oxygen, the number of prediction models to be constructed is reduced, the workload is reduced, and the efficiency is improved.

[0050] (4) The present application uses Bootstrap random sampling which can repeat sampling, can obtain a data set with better uniformity, so that the data units with different variables in the training set and the test set are more uniformly distributed, avoiding the adverse effects of similar variable data sets on the prediction accuracy of decision trees; at the same time, since Bootstrap can statistically analyze sample variance, the smaller the variance, the better the uniformity of sampling, and the result of random sampling is easy to quantitatively characterize.

[0051] (5) The present application introduces a spatial weight related to the distance, taking into account the distance factor of the spatial unit, which is helpful to reduce the error of data correlation evaluation and improve the accuracy of interpolation calculation.

[0052] (6) The present application uses an Argo system as a representative of a self-powered active buoy sampling system, which has lower use cost than the traditional ship sampling method, and the sampling method is more flexible, which can adapt to the task scene of large-scale sampling at the same time in different places; compared with the traditional passive buoy sampling, the active buoy can complete larger area sampling through active movement, has higher sampling efficiency and resistance to ocean current system, and can complete efficient and accurate sampling in the specified area.

[0053] In the present application, the above-mentioned technical solutions can be combined with each other to realize more preferred combination schemes. Other features and advantages of the present application will be described in the subsequent specification, and some advantages will become apparent from the specification, or will be understood by implementing the present application. The purpose and other advantages of the present application can be achieved and obtained by the contents specifically pointed out in the specification, examples and drawings. BRIEF DESCRIPTION OF DRAWINGS

[0054] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification, illustrate embodiments of the application and are not intended to limit the scope of the application, and the same reference numerals designate the same elements throughout the accompanying drawings.

[0055] Figure 1 A flow chart of a method for reconstructing a marine dissolved oxygen concentration model based on an Argo temperature and salinity profile according to an embodiment of the present application;

[0056] Figure 2 A distribution map of Argo marine dissolved oxygen concentration at 1000 dbar in March 2015 according to an embodiment of the present application;

[0057] Figure 3 A distribution map of Argo marine temperature and salinity data at 1000 dbar in March 2015 according to an embodiment of the present application;

[0058] Figure 4 An interpolation result of marine dissolved oxygen climate data in March according to an embodiment of the present application;

[0059] Figure 5 A zoned map of marine dissolved oxygen concentration in March according to an embodiment of the present application;

[0060] Figure 6 A distribution map of marine dissolved oxygen concentration at 1000 dbar in March 2005 according to an embodiment of the present application;

[0061] Figure 7 A statistical chart of relative errors of a test set of each zoned model according to an embodiment of the present application;

[0062] Figure 8 A comparison chart of true values and reconstructed values of a test set of each zoned model according to an embodiment of the present application;

[0063] Figure 9 A statistical chart of relative errors of reconstructed values and measured values of dissolved oxygen according to an embodiment of the present application. DETAILED DESCRIPTION

[0064] The preferred embodiments of the present application will be described in detail below with reference to the drawings, which form a part of this specification, and together with the embodiments of the present application illustrate the principles of the present application, but are not intended to limit the scope of the present application.

[0065] In order to better illustrate the technical solutions of the present application, the following technical terms are explained:

[0066] Absolute error

[0067] Relative error

[0068] Relative error

[0069] Relative error is the percentage of absolute error in the true value, that is, relative error = | measured value - true value | / true value.

[0070] Mean square error

[0071] Mean square error is the expected value reflecting the difference between the estimated value and the true value, which is often used to evaluate the degree of change of data and the accuracy of predicted data.

[0072] Root mean square error

[0073] Root of mean square error.

[0074] Mean absolute error

[0075] Mean absolute error is the average of the absolute values of the deviations of all individual observations from the arithmetic mean; mean absolute error can avoid the problem of error offset, so it can accurately reflect the size of the actual prediction error.

[0076] Coefficient of determination

[0077] Coefficient of determination is the explained variation / total variation; the higher the coefficient of determination, the higher the degree of explanation, and the better the effect of the regression model.

[0078] Minimum percentage error

[0079] Minimum percentage error = | measured value - true value | / true value * 100%.

[0080] Maximum depth of decision tree

[0081] The maximum depth of the decision tree refers to the number of splits obtained by splitting the original data set to obtain a decision tree that meets the accuracy requirements.

[0082] Number of decision trees

[0083] The number of decision trees refers to the number of minimum level decision trees that meet the accuracy requirements after splitting in the decision tree.

[0084] dbar

[0085] Dbar is a unit of ocean depth, 1 dbar is equal to 1 m.

[0086] Based on the technical problems existing in the prior art, the present application proposes a marine dissolved oxygen concentration reconstruction model based on global marine biogeochemical buoy profile data.

[0087] Argo (Array for real-time geostrophic oceanography) is the only real-time observation system for global stereoscopic observation of the upper ocean in the field of biogeochemistry, which provides 240,000 dissolved oxygen profile data in the global range, and provides an important data basis for understanding and analyzing the current characteristics and trends of global ocean dissolved oxygen.

[0088] Argo adopts a powered active float sampling system, and operates in the mode of "submersion-preset depth layer drift-submersion-ascending measurement-surface positioning and data transmission". Argo floats can only measure dissolved oxygen at different depths at a latitude and longitude location during the ascending stage in one operation cycle, but due to the design of the Argo sampling scheme, it is difficult for the Argo system to avoid the problem of sparse data at different latitude and longitude coordinates.

[0089] Meanwhile, since the main source of ocean dissolved oxygen is photosynthesis of atmospheric and phytoplankton, temperature and salinity are the main factors affecting the concentration of ocean dissolved oxygen. Generally, the higher the temperature and salinity, the lower the dissolved oxygen concentration. There is a certain spatiotemporal coupling relationship between the temperature, salinity and dissolved oxygen concentration at different depths of the water area. Argo floats also observe the temperature and salinity profile while observing dissolved oxygen. At the same period, the amount of temperature and salinity data in the Argo database reaches 2.2 million, which is much more than the amount of dissolved oxygen data (22 million). In order to further improve the accuracy of the evaluation of the dissolved oxygen concentration of the adjacent area, the temperature and salinity data at different depths of the water area are used as an important basis for the evaluation of the dissolved oxygen concentration of the adjacent area.

[0090] An ocean dissolved oxygen interpolation method based on profile spatial characteristics is established, and a standard observation layer depth global climatological monthly ocean dissolved oxygen spatial grid product is developed, which provides a data basis for carrying out global ocean ecological health assessment and sustainable development management.

[0091] Based on the problems found by the inventors in the research process, the present application discloses a method for constructing an ocean dissolved oxygen concentration reconstruction model based on Argo temperature and salinity profiles, which comprises the following steps:

[0092] Step 1: Based on the Argo dissolved oxygen profile data of the same month in previous years, the climatological monthly dissolved oxygen data is obtained by screening; based on the climatological monthly dissolved oxygen data, the monthly dissolved oxygen data of any target depth is obtained; the climatological monthly dissolved oxygen data comprises dissolved oxygen data of different months and different depths in previous years;

[0093] Specifically, the method for obtaining the climatological monthly dissolved oxygen data is as follows: accessing the Argo dissolved oxygen profile data of previous years, taking the month as a screening parameter, and screening the climatological monthly dissolved oxygen data of the Mth month from the Argo dissolved oxygen profile data of previous years to obtain the climatological monthly dissolved oxygen data DoxyM (h), wherein M takes 1-12, and h is the depth. The Argo dissolved oxygen profile data provides dissolved oxygen data Doxy(Y, M, h) at different depths from January 1, 2010 to December 31, 2021; the monthly dissolved oxygen data over the years is obtained by screening with the month as the classification parameter: Doxy1(h), …, Doxy M (h); wherein Y is the year.

[0094] Specifically, the climatological monthly dissolved oxygen data contains monthly dissolved oxygen data at multiple depths, which is determined by the Argo annual dissolved oxygen profile data collection method.

[0095] Specifically, the acquisition of monthly dissolved oxygen data at any target depth includes: selecting the monthly dissolved oxygen data closest to the target depth from the climatological monthly dissolved oxygen data as the monthly dissolved oxygen data at the target depth.

[0096] For example, the present application selects multiple target depths as shown in Table 1:

[0097] Table 1 Target depth standard depth for 2000 meters above sea level ocean element observation

[0098] Standard layer Depth (m) Standard layer Depth (m) Standard layer Depth (m) 1 10 10 200 19 1000 2 20 11 250 20 1100 3 30 12 300 21 1200 4 40 13 400 22 1300 5 50 14 500 23 1400 6 75 15 600 24 1500 7 100 16 700 25 1750 8 125 17 800 26 2000 9 150 18 900

[0099] For example, the Argo buoy system collects data in the following way, which consists of five stages to form a cycle:

[0100] ① Sinking stage;

[0101] ② Stagnant layer drifting stage;

[0102] ③ Sinking stage again;

[0103] ④ Rising measurement stage;

[0104] ⑤ Drifting, while communicating with satellite stage.

[0105] For example, the Argo dissolved oxygen profile data obtained from the Argo buoy measurement selects the data point closest to the depth with the same latitude and longitude coordinates as the target depth in Table 1, as shown in Table 2, and the depth of the above-mentioned data point is:

[0106] Table 2 Data point depth closest to the target depth for 2000 meters above sea level ocean element observation

[0107] Standard layer Depth (m) Standard layer Depth (m) Standard layer Depth (m) 1 13 10 193 19 999 2 22 11 245 20 1080 3 32 12 308 21 1190 4 39 13 386 22 1320 5 52 14 489 23 1430 6 78 15 568 24 1520 7 97 16 634 25 1779 8 120 17 789 26 2005 9 149 18 987

[0108] Step 2: Divide the plane corresponding to each target depth into a plurality of spatial units using latitude and longitude coordinates, perform the first interpolation calculation based on the monthly scale dissolved oxygen data of the target depth, obtain the dissolved oxygen concentration of the center point of the spatial unit, and take it as the dissolved oxygen concentration of the spatial unit; based on the dissolved oxygen concentration of each spatial unit, the spatial units with close dissolved oxygen concentration at the target depth are divided into the same type of region, and the spatio-temporal classification zoning of marine dissolved oxygen at the target depth is completed; different target depth zoning processing is performed to construct the spatio-temporal classification zoning of marine dissolved oxygen at different target depths.

[0109] Specifically, the dissolved oxygen concentration at the same target depth under a plurality of latitude and longitude coordinates is divided into spatial units according to latitude and longitude coordinates, and the dissolved oxygen concentration of the center point of the spatial unit is calculated by interpolation based on the dissolved oxygen concentration of the data points at the same target depth near the center point of the spatial unit in the monthly scale dissolved oxygen data and the distance between the nearby data points and the center of the spatial unit.

[0110] Specifically, the dissolved oxygen concentration isosurface with fixed interval is set, the isosurface is divided into several level ranges with different dissolved oxygen concentrations according to fixed interval, the spatial units are divided into different levels based on the dissolved oxygen concentration of the spatial units, and the spatial units with the same level are divided into the same type of region.

[0111] When implemented, the isosurface is divided into levels with different dissolved oxygen concentrations according to fixed interval: 0~r, r~2r, …, (v-1)r~vr; wherein, r is the interval of the dissolved oxygen concentration isosurface, and v is any natural number greater than or equal to 1.

[0112] It should be noted that the regions with close dissolved oxygen concentration have similar environments, and the temperature and salinity have close effects on dissolved oxygen, so the same classification region is uniformly studied, and the same prediction model is constructed.

[0113] It should be noted that the spatial units with close dissolved oxygen concentration refer to the spatial units with the same isosurface level and adjacent latitude and longitude coordinates, that is, the latitude and longitude of the spatial units with close dissolved oxygen concentration are adjacent, and the dissolved oxygen concentration of the spatial units meets: as an example, the dissolved oxygen concentration of the spatial units with close dissolved oxygen concentration is ∈[(v-1)r, vr].

[0114] Step 3: Based on the Argo temperature and salinity profile data of the same month in previous years, monthly scale temperature and salinity data of the same target depth are obtained by screening; monthly scale temperature and salinity data and monthly scale dissolved oxygen concentration data at the same spatio-temporal classification zoning, the same target depth and the same latitude and longitude coordinates are selected, and a regression prediction model of the spatio-temporal classification zoning is constructed with temperature and salinity as variables and dissolved oxygen concentration as dependent variable.

[0115] It should be noted that the above space-time classification partition refers to classifying regions with adjacent latitude and longitude coordinates and close dissolved oxygen concentration, and dividing into space-time classification same class regions.

[0116] It should be noted that the regression prediction model is affected by the space-time classification partition, the target depth and the month, and is different with the change of the space-time classification partition, the target depth and the month.

[0117] Specifically, the monthly scale dissolved oxygen data is obtained by the same method, and the monthly scale temperature and salinity data is obtained from the Argo temperature and salinity profile data.

[0118] It should be noted that in order to construct the prediction model, the above monthly scale temperature and salinity data and monthly scale dissolved oxygen concentration data are selected as temperature and salinity concentration data and dissolved oxygen concentration data obtained at the same time at the same position of the same sampling point by the Argo buoy system; for constructing the regression prediction model, the data points used in steps 1-3 are derived from Argo sampling points with both temperature and salinity concentration data and dissolved oxygen concentration data.

[0119] At the same time, since the dissolved oxygen concentration measurement condition is more demanding, the number of effective dissolved oxygen concentration data in the Argo database is still much smaller than that of temperature and salinity concentration data, and the present application is based on the coupling relationship between the dissolved oxygen concentration of the same sampling data point and the temperature and salinity data of the same sampling point, and based on the existing rich temperature and salinity data, the missing dissolved oxygen concentration data of adjacent regions is reconstructed to obtain the corrected value of the dissolved oxygen concentration data of the sampling point of the adjacent region.

[0120] Specifically, based on the two kinds of data of the sampling points with dissolved oxygen concentration and temperature and salinity data in a certain space-time classification partition, a regression prediction model is constructed with temperature and salinity data as variables and dissolved oxygen concentration as dependent variable, and the regression prediction model corresponds to the space-time classification partition grade one by one; the temperature and salinity data of the sampling points with only temperature and salinity data in the space-time classification partition are substituted into the regression prediction model corresponding to the space-time classification partition to obtain the corrected value of the dissolved oxygen concentration data of the sampling points with only temperature and salinity data.

[0121] Specifically, the random forest regression method is used to construct the regression prediction model.

[0122] In implementation, data points of the same space-time classification partition, the same target depth and the same month are selected to construct the prediction model, with dissolved oxygen concentration as dependent variable, temperature and salinity in the monthly scale temperature and salinity data of the same target depth and the same month as variables to construct the regression prediction model.

[0123] Step 4: Based on the monthly temperature-salinity data sampling points, the correction value of the dissolved oxygen concentration of each sampling point at the target depth in the corresponding spatio-temporal classification partition is obtained from the regression prediction model; the sampling points are data sampling points of the Argo system at different depths and latitude-longitude coordinates; based on the correction value of the dissolved oxygen concentration at the target depth, a second interpolation calculation is performed to obtain the corrected dissolved oxygen concentration of the center point of the spatial unit, complete the reconstruction of the monthly temperature-salinity data of the corresponding spatio-temporal classification partition and target depth, and construct the marine dissolved oxygen concentration reconstruction model containing different target depths and different spatio-temporal partitions.

[0124] Specifically, the temperature, salinity, latitude-longitude coordinates, and year of the monthly temperature-salinity data sampling points are substituted into the regression prediction model corresponding to the spatio-temporal classification partition and target depth to obtain the correction value of the dissolved oxygen concentration of the sampling points.

[0125] Specifically, based on the correction value of the dissolved oxygen concentration of the data points at the same target depth near the center point of the spatial unit in the monthly dissolved oxygen data, and the correction value of the dissolved oxygen concentration of the center point of the spatial unit by interpolating the distance between the nearby data points and the center of the spatial unit.

[0126] When implemented, the correction value of the dissolved oxygen concentration of the sampling points is the same as the first interpolation calculation of the dissolved oxygen concentration of the center point of the spatial unit in step 2.

[0127] Specifically, the second interpolation calculation in step 4 to obtain the corrected dissolved oxygen concentration of the center point of the spatial unit includes: taking the center point of the spatial unit as the target point, and taking the data points near the center point of the spatial unit in the monthly dissolved oxygen data at the same depth as the adjacent sampling points; based on the different weights given to the distance between the adjacent sampling points and the center point of the spatial unit, the mean value of the correction values of the dissolved oxygen concentrations of the adjacent sampling points at the same target depth is obtained to obtain the corrected dissolved oxygen concentration of the center point of the spatial unit.

[0128] Specifically, it includes: taking the center point of the spatial unit as the target point, and taking the data points near the center point of the spatial unit in the monthly dissolved oxygen data at the same depth as the adjacent sampling points, based on the different weights given to the distance between the adjacent sampling points and the center point of the spatial unit, to obtain the dissolved oxygen concentration of the center point of the spatial unit

[0129] It should be noted that the correction value of the dissolved oxygen concentration obtained from the regression prediction model is based on the correction of the temperature-salinity data related to the dissolved oxygen, which greatly improves the prediction accuracy compared to the first dissolved oxygen interpolation calculation result.

[0130] Specifically, based on the regression prediction model, the marine dissolved oxygen concentration at different depths and different spatio-temporal partitions is obtained, and the marine dissolved oxygen concentration reconstruction model containing different target depths and different spatio-temporal partitions is constructed.

[0131] Compared with the prior art, the present application divides the plane corresponding to each target depth into different spatial units by using latitude and longitude coordinates, performs first interpolation calculation on the dissolved oxygen concentrations of different spatial units at the same target depth by using the dissolved oxygen concentration of the target depth, and obtains the dissolved oxygen concentrations of different spatial units at different target depths, thereby greatly making up for the lack of sparseness of data at different latitude and longitude coordinates in the existing Argo profile database.

[0132] Compared with the prior art, the present application obtains monthly scale temperature and salinity data at the same target depth by using Argo temperature and salinity profile data, and reconstructs the dissolved oxygen concentrations of spatial units in the same partition by using the monthly scale temperature and salinity data, thereby improving the problem of large deviation between the interpolation calculation and the actual value of the dissolved oxygen concentration of the spatial unit caused by the local excessive sparseness of data at different latitude and longitude coordinates in the Argo profile database, and improving the prediction accuracy of the dissolved oxygen concentration.

[0133] Compared with the prior art, the present application partitions data points based on the dissolved oxygen concentrations obtained by performing first interpolation calculation on the dissolved oxygen concentrations of different spatial units at the same target depth, and constructs a prediction model of temperature and salinity on dissolved oxygen in regions with close dissolved oxygen concentrations; on the basis of introducing interpolation calculation correction of temperature and salinity on dissolved oxygen, the number of prediction models is limited, the workload is reduced, and the efficiency is improved.

[0134] Compared with the prior art, the present application adopts an active float sampling system represented by the Argo system, which has lower use cost than the traditional ship sampling method, and the sampling method is more flexible and can adapt to the task scenario of large-scale sampling at different places at the same time; compared with the traditional passive float sampling, the active float can complete sampling in a larger area through active movement, has higher sampling efficiency and resistance to ocean current system, and can complete efficient and accurate sampling in a given area.

[0135] Specifically, the first interpolation calculation on the dissolved oxygen concentrations of different spatial units at the same target depth in step 2 comprises: taking the spatial unit center point as a target point, taking the data points near the spatial unit center point of the same depth monthly scale dissolved oxygen data as adjacent sampling points, giving different weights based on the distances between the adjacent sampling points and the spatial unit center point, and obtaining the dissolved oxygen concentration of the spatial unit center point.

[0136] Specifically, the dissolved oxygen concentration of the target point satisfies the following formula:

[0137]

[0138] Wherein, i is the number of adjacent sampling points, n is the number of adjacent sampling points specified to perform interpolation, d iDoxy is the distance from the i-th nearest sampling point to the target point. i This represents the dissolved oxygen concentration at the neighboring sampling point i where interpolation was performed.

[0139] Specifically, methods for obtaining dissolved oxygen concentration at nearby sampling points include:

[0140] Using the center of the k-th spatial unit as the center, number the data points in the monthly dissolved oxygen data at the same target depth from near to far as C1, C2, ..., C... x x represents the data point number, and at least three sampling points with the smallest x value are selected as neighboring sampling points.

[0141] Compared with existing technologies, this invention introduces distance-related spatial weights and takes into account the distance factor of spatial units, which helps to reduce errors in data correlation assessment and improve the accuracy of interpolation calculation.

[0142] For example, such as Figure 4 As shown, the interpolation calculation results of dissolved oxygen concentration in different spatial units at the same target depth show that: after interpolation calculation, the originally scattered and sparse dissolved oxygen data points become uniformly distributed and continuous data points; the color depth in the figure more intuitively represents the level of dissolved oxygen concentration.

[0143] Specifically, step 2, which involves constructing a spatiotemporal classification and zoning of ocean dissolved oxygen at the target depth, includes: dividing isosurfaces into different levels of dissolved oxygen concentration at fixed intervals: 0~r, r~2r, …, (v-1)r~vr, and dividing them into different regions according to the different levels of dissolved oxygen concentration, where r is the isosurface interval of dissolved oxygen concentration and v is any natural number greater than or equal to 1.

[0144] For example, such as Figure 5 As shown, the isosurface spacing is set to 40 μmol / kg, and the isosurfaces are divided into different levels of dissolved oxygen concentration: 0–40 μmol / kg, 40–80 μmol / kg, …, 280–320 μmol / kg; further, based on the different levels of dissolved oxygen concentration, different regions are divided and marked with different colors and assigned numbers; the color intensity in the figure can intuitively indicate that the dissolved oxygen concentration belongs to the same category of region.

[0145] Specifically, regarding the random forest regression method described in step 3: The random forest regression method is suitable for multi-variable regression analysis. It uses a regression tree analysis method to randomly split the training dataset into multiple data subsets. The middle subset is split, and the minimum sum of the mean squared errors of the two split parts is used as the optimal split point. A decision tree model with decision subtrees of different split structures is constructed. The constructed decision tree model is then validated using test set data to confirm the accuracy of the decision tree model's data prediction.

[0146] If the accuracy meets the preset requirement, output the decision tree model as the final model;

[0147] If the accuracy does not meet the preset requirement, continue to split the two parts of the data subset after splitting to form decision sub-trees, find the optimal splitting point of the decision sub-trees, and construct a new decision tree model; use the test set data to verify the constructed decision tree model, confirm the prediction accuracy of the decision tree model for the data, and continue until the accuracy meets the preset requirement or the maximum depth of the decision tree and the number of decision trees in the forest reach the preset value, and output the decision tree model as the final model.

[0148] The regression prediction model constructed in step 3 with temperature and salinity as variables and dissolved oxygen concentration as dependent variable comprises:

[0149] S301: Based on the monthly scale dissolved oxygen data of the same spatio-temporal classification partition and target depth, and the monthly scale temperature and salinity data of the same target depth and the same month, a test set and a training set containing variables temperature, salinity, longitude, latitude, year and dependent variable dissolved oxygen concentration are constructed; wherein the data units of the test set and the training set contain one-to-one corresponding dependent variables and variables;

[0150] Specifically, since the Argo buoy can simultaneously obtain measured temperature and salinity data and dissolved oxygen concentration at the same spatio-temporal position, the monthly scale dissolved oxygen data and the monthly scale temperature and salinity data have corresponding data points with the same spatio-temporal classification partition, the same target depth, the same month, and the same longitude and latitude. The monthly scale dissolved oxygen data and the corresponding dissolved oxygen concentration data in the monthly scale temperature and salinity data are one-to-one corresponding, and are split into multiple data units containing one-to-one corresponding variables and dissolved oxygen concentration.

[0151] S302: The training set is divided into k training subsets by random sampling, and the training subsets are cut. The mean square value of the two parts of the training subset data after cutting is selected as the optimal cutting point, and a decision tree model with k decision sub-trees with different splitting structures is constructed. The mean value of the output values of each data point in each decision sub-tree at the optimal cutting point of each decision tree is obtained to obtain the prediction value of each decision tree. The mean value of the prediction values of the k decision trees is taken as the prediction output of the regression prediction model.

[0152] S303: The test set is substituted into the random forest regression model to obtain the prediction value of the dissolved oxygen concentration. Based on the prediction value of the dissolved oxygen concentration and the true value of the dissolved oxygen concentration in the test set, the accuracy of the model is evaluated. If the accuracy meets the preset requirement, the training of the partitioned marine dissolved oxygen concentration reconstruction model is completed. If the requirement is not met, the maximum depth of the decision tree and the number of decision trees are adjusted until the model accuracy meets the preset requirement.

[0153] It should be noted that the trained decision tree model has two characteristic parameters of maximum depth of decision tree and number of decision trees in addition to variable type, and the above two parameters and variables are used as indexes of training quality of decision tree model, which jointly determine the prediction accuracy of decision tree model.

[0154] It should be noted that the training set training process is that the variables and dependent variables (output values) of the training data are known, and the regression prediction model identifies the labeled training data during training and adjusts the characteristic parameters in the model to adjust the prediction accuracy of the regression prediction model.

[0155] S301 constructs the test set and the training set containing the variables of temperature, salinity, longitude, latitude, year and the dependent variable of dissolved oxygen concentration, including:

[0156] S3011: obtaining data points with the same longitude and latitude from the monthly scale dissolved oxygen data, monthly scale temperature and salinity data in the same spatio-temporal classification partition, the same target depth and the same month; the data points generate information containing the variables of temperature, salinity, longitude, latitude, year and dissolved oxygen concentration;

[0157] S3012: generating a first data unit from any one of the variables of temperature, salinity, longitude, latitude and year and the corresponding dissolved oxygen concentration from the data points, and sequentially generating a second data unit, …, a fifth data unit from the remaining variables and the corresponding dissolved oxygen concentration, and the dependent variable and the variable in each data unit correspond one by one.

[0158] S302 includes Bootstrap sampling, specifically, containing the steps of:

[0159] (1) using repeated sampling technique to extract a certain number of repeated sampling from the original sample;

[0160] (2) calculating the statistical quantity T to be estimated according to the extracted sample;

[0161] (3) repeating the above N times to obtain N statistical quantities T;

[0162] (4) calculating the sample variance of the above N statistical quantities T to estimate the variance of the statistical quantity T.

[0163] Compared with the prior art, the Bootstrap random sampling with repeated sampling can obtain a data set with better uniformity, so that the data units with different variables in the training set and the test set are more uniformly distributed, avoiding the adverse effect of similar variable data set on the prediction accuracy of decision tree; at the same time, since Bootstrap can statistically analyze the sample variance, the smaller the variance, the better the uniformity of sampling, and the result of random sampling is easy to quantitatively characterize.

[0164] The decision tree model of S302 constructing k decision sub-trees with different split structures comprises:

[0165] S3021, obtaining an optimal split point corresponding to a minimum value of a mean square value of two parts after splitting the training subset, wherein the mean square value RMSE satisfies the formula:

[0166]

[0167] R1 and R2 satisfy:

[0168] R1(j, s) = {x | x (j) ≤ s}, R2(j, s) = {x | x (j) > s};

[0169] Wherein, x is an input feature value, y is an output value; c1 is the average value of the output of the left sub-region; c2 is the average value of the output of the right sub-region; j is the optimal split variable, s is the optimal split point, x(j) represents the feature value x=j; R1 is the left sub-region after splitting, R2 is the right sub-region after splitting; y i represents the output value of the i th data point; min represents the minimum value of the min function adjacent to the right region;

[0170] S3022, based on the average of the output values of each data point in each decision sub-tree in each decision tree, obtaining a prediction value of each decision tree, the prediction value of the decision tree c k satisfies:

[0171]

[0172] Wherein, y i represents the output value of the i th data point, N k represents the number of data points in the k th training subset, c k represents the average of the output values of the decision trees corresponding to the k th training subset.

[0173] S3023, taking the average of the prediction values of the k decision trees as the prediction value output by the regression prediction model, the average of the prediction values of the k decision trees c - satisfies:

[0174]

[0175] Wherein, l ≤ k, represents the l th training subset, c l represents the average of the output values of the decision trees corresponding to the l th training subset.

[0176] Specifically, S3021 obtains the optimal split point corresponding to the minimum value of the mean square value of the two parts after splitting the training subset, comprising the following steps:

[0177] (1) Select any variable in the training subset for multiple splitting, and obtain the optimal split point s1 of the variable; the mean square error RMSE1 corresponding to the optimal split point satisfies:

[0178]

[0179] Wherein, c1 is the average value of the left sub-region output; c2 is the average value of the right sub-region output; R1 is the left sub-region after splitting, R2 is the right sub-region after splitting; y i represents the output value of the i-th data point; min represents the minimum value of the min function adjacent to the right region;

[0180] (2) The remaining variables in the training subset are sequentially split according to the same method, and the mean square error is calculated, the optimal split point of each variable is obtained, and the mean square error corresponding to the optimal split point of each variable is compared, the optimal split point corresponding to the minimum mean square error is taken as the optimal split point s of the training subset, and the corresponding variable is taken as the optimal split variable j.

[0181] The precision evaluation in S303 includes using one or more of the minimum percentage error, mean square error or decision coefficient to evaluate the precision of the regression prediction model.

[0182] Preferably, the precision evaluation selects the minimum percentage error.

[0183] Specifically, the minimum percentage error is less than 10%; the maximum depth of the decision tree is 300-500.

[0184] The adjustment of the maximum depth of the decision tree and the number of decision trees in S303 includes using the cross-validation method.

[0185] Specifically, the steps of the cross-validation method include:

[0186] S3031: Continue to divide the training set into training, validation and test parts; wherein the training part is used for model training, the validation part is used for adjusting parameters, and the test part is used to measure the performance of the model; the training set is further divided into k training subsets, and the training subsets are evenly divided at the middle position;

[0187] S3032: Each training subset is validated once, and the remaining k-1 training subset data is used as the training part to obtain k models;

[0188] S3033: The average of the root mean square errors of the final validation part of the k models is taken as the performance index corresponding to the k models, and the performance index corresponding to the decision tree parameters is recorded;

[0189] S3034: Traverse the k models, repeat the above steps, and select the optimal parameters corresponding to the optimal performance index as the optimal parameters.

[0190] Preferably, the performance index corresponding to the model is root mean square error; and the decision tree parameters include maximum depth of the decision tree and the number of decision trees.

[0191] As an example, the application discloses a method for evaluating the precision of a random forest regression model of a certain partition by using a test set, and the evaluation indexes include one or more of minimum percentage error, mean square error, root mean square error, mean absolute error and determination coefficient; specifically including:

[0192] (1) The data set is divided into a training set and a test set at a ratio of 4:1, and the input features include temperature, salinity, longitude, latitude and year, and the dependent variable is dissolved oxygen concentration;

[0193] (2) The Bootstrap sampling method is used to randomly divide the data set into k training sets, each of which contains data units with one-to-one corresponding variables and dependent variable dissolved oxygen concentration; one decision tree is trained for each training set, and the value of k is the number of trees in the random forest regression parameter; in the training process, the left and right training subsets are obtained by splitting each training set; for data units with the same variable, the mean square deviation of the left and right training subsets of the variable is calculated, and the splitting point corresponding to the minimum mean square deviation is obtained as the optimal splitting point of the variable; the optimal splitting points of all variables are obtained in sequence, the mean square deviations corresponding to all variable splitting points are compared, and the variable with the minimum mean square deviation is selected as the optimal splitting variable of the training set; one decision tree is constructed for each training set to obtain k decision trees;

[0194] (3) The mean value of the dissolved oxygen concentrations of the data units in each training set is taken as the output value of the decision tree corresponding to the training set; and the mean value of the output values of the k decision trees is taken as the predicted value output by the random forest regression prediction model;

[0195] (4) The precision of the model is evaluated by using one or more of the minimum percentage error, mean square error, root mean square error, mean absolute error and determination coefficient through the test set data; the random forest regression prediction model is input with the temperature, salinity, longitude, latitude and year in the test set data to obtain a predicted value, and the predicted value and the measured value in the test set are subjected to precision analysis; if the minimum percentage error is less than or equal to 10%, the training of the reconstruction model of the marine dissolved oxygen concentration of the partition is completed, and if the minimum percentage error is greater than 10%, the cross-validation method is used to automatically adjust the maximum depth of the decision tree and the number of decision trees until the maximum depth of the decision tree reaches 500; and the training of the random forest regression prediction model is completed.

[0196] As an example, as Figure 7As shown in the results of the accuracy evaluation of the random forest regression prediction model for each partition, the predicted values of the random forest regression prediction model for each partition are uniformly distributed around the measured values, showing an approximately normal distribution, and the model has good prediction accuracy. Figure 8 As shown in the results of the accuracy evaluation of the random forest regression prediction model for each partition, the predicted values of the random forest regression prediction model for each partition are uniformly distributed around the measured values, showing an approximately normal distribution, and the model has good prediction accuracy.

[0197] Specifically, the second interpolation calculation in step 4 is performed to obtain the corrected dissolved oxygen concentration of the spatial unit center point, wherein the corrected dissolved oxygen concentration Doxy' of the spatial unit center point satisfies:

[0198]

[0199] where i is the number of adjacent sampling points, n is the number of adjacent sampling points for which interpolation is performed, di is the distance from the i-th adjacent sampling point to the target point, Doxy' represents the corrected dissolved oxygen concentration of the adjacent sampling point i for which interpolation is performed. i i

[0200] Further, in order to evaluate the accuracy of the regression prediction model for each partition, the present application also discloses a precision correction method for reconstructing dissolved oxygen concentration data based on temperature-salinity profiles, comprising:

[0201] (1) Reconstructing the dissolved oxygen data of previous years using the Argo temperature-salinity profile data of previous years, and the reconstructed data set is identified as: O TS ; obtaining the Argo dissolved oxygen measured data of previous years, identified as: O Argo ; mapping O _Argo and O TS data to a 1°x1° grid.

[0202] (2) Comparing the two grid values corresponding to the spatial position based on the absolute error and the maximum percentage error.

[0203] Specifically, the historical data selects the Argo temperature-salinity profile data from 2010 to 2021 and the dissolved oxygen data from 2010 to 2021.

[0204] As shown in the results of the accuracy evaluation of the random forest regression prediction model for each partition, the predicted values of the random forest regression prediction model for each partition are uniformly distributed around the measured values, showing an approximately normal distribution, and the model has good prediction accuracy. Figure 9 As shown in the results of the accuracy evaluation of the random forest regression prediction model for each partition, the predicted values of the random forest regression prediction model for each partition are uniformly distributed around the measured values, showing an approximately normal distribution, and the model has good prediction accuracy.​​

[0205] The above description is only the preferred embodiment of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application.

Claims

1. A method for constructing a marine dissolved oxygen concentration reconstruction model based on Argo temperature-salinity profiles, characterized in that, The application relates to a method for reconstructing ocean dissolved oxygen concentration, and belongs to the technical field of ocean data reconstruction. The method comprises the following steps: Based on the Argo dissolved oxygen profile data of the same month in different years, climate state monthly scale dissolved oxygen data is obtained through screening; Based on the climate state monthly scale dissolved oxygen data, monthly scale dissolved oxygen data of any target depth is obtained; the climate state monthly scale dissolved oxygen data contains dissolved oxygen data of different months and different depths in different years; Each target depth is divided into multiple spatial units by using the latitude and longitude coordinates, the first interpolation calculation is performed on the dissolved oxygen concentrations of different spatial units under the same target depth based on the monthly scale dissolved oxygen data of the target depth, the dissolved oxygen concentration of the center point of the spatial unit is obtained, and the dissolved oxygen concentration of the spatial unit is taken as the dissolved oxygen concentration of the spatial unit; based on the dissolved oxygen concentrations of the spatial units, the spatial units with close dissolved oxygen concentrations under the target depth are divided into the same type of regions, and the spatial and temporal classification and partition of the ocean dissolved oxygen under the target depth are completed; different target depths are partitioned, and the spatial and temporal classification and partition of the ocean dissolved oxygen under different target depths are constructed; Based on the Argo temperature and salinity profile data of the same month in different years, monthly scale temperature and salinity data of the same target depth is obtained through screening; Monthly scale temperature and salinity data and monthly scale dissolved oxygen concentration data of the same spatial and temporal classification and partition, the same target depth and the same latitude and longitude coordinates are selected, a regression prediction model of the spatial and temporal classification and partition is constructed, the temperature and salinity are taken as variables, and the dissolved oxygen concentration is taken as a dependent variable; Based on the monthly scale temperature and salinity data of the monthly scale temperature and salinity data sampling points, the correction value of the dissolved oxygen concentration of each sampling point at the target depth in the corresponding spatial and temporal classification and partition is obtained from the regression prediction model; 2. The construction method of claim 1, wherein, The sampling points are data sampling points obtained by the Argo system at different depths and latitude and longitude coordinates; based on the correction value of the dissolved oxygen concentration at the target depth, the second interpolation calculation is performed to obtain the corrected dissolved oxygen concentration of the center point of the spatial unit, the reconstruction of the monthly scale temperature and salinity data of the corresponding spatial and temporal classification and partition and the target depth is completed, and the ocean dissolved oxygen concentration reconstruction model containing different target depths and different spatial and temporal partitions is constructed. The method for reconstructing ocean dissolved oxygen concentration comprises the following steps: The monthly climatological monthly scale dissolved oxygen data Doxy is obtained by screening the annual dissolved oxygen profile data of Argo by month as a screening parameter M (h); wherein M is 1-12, and h is depth.

3. The construction method of claim 1, wherein, Obtaining Argo dissolved oxygen profile data in different years; 4. The construction method according to claim 3, characterized in that, The first interpolation calculation on the dissolved oxygen concentrations of different spatial units under the same target depth comprises the following steps: taking the center point of the spatial unit as a target point, taking the data points near the center point of the spatial unit as adjacent sampling points, and giving different weights based on the distances between the adjacent sampling points and the center point of the spatial unit to obtain the dissolved oxygen concentration of the center point of the spatial unit. ; where i is the index of the neighboring sampling point, n is the number of neighboring sampling points designated to perform interpolation, d i is the distance from the i-th neighboring sampling point to the target point, Doxy i represents the dissolved oxygen concentration of the neighboring sampling point i designated to perform interpolation.

5. The Argo-based reconstruction model construction method of the ocean dissolved oxygen concentration according to claim 1, characterized in that, The dissolved oxygen concentration of the target point satisfies the following formula: The regression prediction model of the temperature and salinity as variables and the dissolved oxygen concentration as a dependent variable comprises the following steps: Based on the monthly scale dissolved oxygen data of the same spatial and temporal classification and partition and target depth and the monthly scale temperature and salinity data of the same target depth and the same month, a test set and a training set containing variables temperature, salinity, longitude, latitude, year and dependent variable dissolved oxygen concentration are constructed; the data units of the test set and the training set contain one-to-one dependent variables and variables. The training set is divided into k training subsets by random sampling, the training subsets are cut, the minimum mean square value of the data of the two parts of the training subsets after cutting is selected as the optimal cutting point, and a decision tree model of k decision sub-trees with different split structures is constructed; the average of the output values of each data point in each decision sub-tree at the optimal cutting point of each decision tree is obtained to obtain the prediction value of each decision tree; and the average of the prediction values of the k decision trees is taken as the prediction output by the regression prediction model; The test set is substituted into the random forest regression model to obtain the prediction value of the dissolved oxygen concentration, and the model is evaluated in terms of accuracy based on the prediction value of the dissolved oxygen concentration and the true value of the dissolved oxygen concentration in the test set; if the accuracy meets the preset requirement, the training of the partitioned marine dissolved oxygen concentration reconstruction model is completed, and if the requirement is not met, the maximum depth of the decision tree and the number of decision trees are adjusted until the accuracy of the model meets the preset requirement.

6. The construction method of claim 5, wherein, The decision tree model of k decision sub-trees with different split structures is constructed, including: The optimal cutting point corresponding to the minimum mean square value of the two parts after cutting of the training subset is obtained, wherein the mean square value RMSE satisfies the formula: ; R1 and R2 satisfy: ; wherein x is an input feature value, y is an output value; c1 is an average value of the left sub-region output; c2 is an average value of the right sub-region output; j is an optimal split variable, s is an optimal split point, x (j) represents a feature value of x = j; R1 is a left sub-region after splitting, R2 is a right sub-region after splitting; y i represents an output value of the i th data point; min represents taking a minimum value of the min function adjacent to the right region; Based on the mean value of the output values of the data points in each decision subtree in each decision tree, obtain the prediction value of each decision tree, the decision tree prediction value c k satisfies: ; where y i represents the output value of the i-th data point, N k represents the number of data points in the k-th training subset, c k represents the mean value of the output values of the decision tree corresponding to the k-th training subset; The mean of the k decision tree prediction values is taken as the prediction value output by the regression prediction model, and the mean c of the k decision tree prediction values satisfies: - satisfies: ; wherein, l≤k, denotes the lth training subset, c l denotes the mean value of the output value of the decision tree corresponding to the lth training subset.

7. The construction method of claim 6, wherein, The optimal cutting point corresponding to the minimum mean square value of the two parts after cutting of the training subset is obtained, including: Any variable in the training subset is selected for multiple cutting, and the optimal cutting point s1 of the variable is obtained; the mean square error RMSE1 corresponding to the optimal cutting point satisfies: ; wherein, c1 is the average value of the left sub-region output; c2 is the average value of the right sub-region output; R1 is the left sub-region after cutting, R2 is the right sub-region after cutting; y i represents the output value of the i-th data point; min represents taking the minimum value of the adjacent right region of the min function; The remaining variables in the training subset are cut in the same way, and the mean square error is calculated to obtain the optimal cutting point of each variable. The optimal cutting point corresponding to the minimum mean square error is taken as the optimal cutting point s of the training subset, and the corresponding variable is taken as the optimal cutting variable j.

8. The construction method of claim 5, wherein, The adjustment of the maximum depth of the decision tree and the number of decision trees includes a cross-validation method, and the cross-validation method includes: The training set is further divided into a training part, a validation part and a test part; the training part is used for model training, the validation part is used for adjusting parameters, and the test part is used for measuring the performance of the model; the training set is further divided into k training subsets, and the training subsets are evenly divided from the middle position; Each training subset is validated once, and the remaining k-1 training subset data is used as the training part to obtain k models; The average of the root mean square errors of the final validation parts of the k models is taken as the performance index corresponding to the k models, and the decision tree parameters corresponding to the performance index are recorded; The k models are traversed, and the above steps are repeated to select the optimal decision tree parameters corresponding to the optimal performance index as the optimal parameters.

9. The construction method according to claim 8, wherein, The performance index corresponding to the model is the root mean square error; and the decision tree parameters include the maximum depth of the decision tree and the number of decision trees.

10. The construction method of claim 5, wherein, The random sampling method includes Bootstrap sampling.

Citation Information

Patent Citations

  • Underwater three-dimensional temperature-salinity parallel forecasting method based on LightGBM model

    CN113946978A

  • Atmospheric carbon dioxide column concentration high coverage reconstruction method

    CN114974453A