GNSS Coordinate Time Series Prediction Method Based on Genetic Algorithm Optimized SVM

By using a genetic algorithm optimized SVM model in GNSS coordinate timing prediction, combined with geophysical effect data, the problems of noise impact, missing value processing and calculation complexity in the prior art are solved, and high-precision GNSS coordinate timing prediction is achieved.

CN119939258BActive Publication Date: 2025-06-24SHANDONG JIANZHU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510422190.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-24
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The prior art has noise influence in GNSS coordinate timing prediction, difficulty in processing missing values, complex calculations of traditional SVM algorithms and noise-sensitive, and insufficient utilization of geophysical effect data and GNSS historical data, resulting in low prediction accuracy.

Method used

The support vector machine (SVM) model based on genetic algorithm optimization is used to predict the timing of GNSS coordinates with geophysical effect data. Search the best parameter combination of SVM models through genetic algorithms, build GA-SVM models, and fuse multi-source data for training to improve prediction accuracy.

Benefits of technology

It significantly improves the accuracy of GNSS coordinate timing analysis and prediction, reduces the workload of manual parameter adjustment, enhances the generalization ability and training effect of the model, and is suitable for applications such as identification of precursor signals of crustal motion, geological disaster warning, etc.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939258B_ABST
    Figure CN119939258B_ABST
Patent Text Reader

Abstract

The present invention discloses a GNSS coordinate time series prediction method based on genetic algorithm optimized SVM, belonging to the technical field of geodetic deformation monitoring, and comprising the following steps: S1. Collect the geophysical effect data and GNSS time series of multiple GNSS stations in a certain area, and process them to obtain the complete time series of each GNSS station; S2. Construct a GA-SVM model: use the genetic algorithm to optimize the support vector machine model to obtain the GA-SVM model that meets each GNSS station; S3. Conduct model evaluation and prediction. The present invention uses four characteristic data such as geophysical effects to predict GNSS, and uses GA for optimization, which can perform multi-objective optimization, solve the parameter optimization problem in high-dimensional data, and improve the performance of the model on high-dimensional data; at the same time, it provides sufficient training and verification data for the model, improves the training effect and prediction accuracy of the model, and enables the model to have higher accuracy and generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of geodetic deformation monitoring, and more particularly to a GNSS coordinate time series prediction method based on genetic algorithm optimized SVM. Background Art

[0002] In the "Next Generation GNSS" program of the European Space Agency, the goal of GNSS coordinate time series prediction with enhanced physical models is to reduce the 7-day coordinate prediction error from centimeter level to millimeter level. Conducting research on high-precision GNSS coordinate time series prediction methods has important and far-reaching significance for aspects such as identification of crustal movement precursor signals, geological disaster warning, and maintenance of the Earth reference frame. Geophysical effects (such as temperature, tidal and non-tidal effects, atmospheric and hydrological loads, etc.) can cause millimeter- to centimeter-level periodic or long-term displacements on the Earth's surface. Relevant research on predicting and analyzing GNSS coordinate time series by combining GNSS deformation data with geophysical effect data is still lacking.

[0003] In such prediction methods, the GNSS coordinate time series has noise, which will have various effects on time series analysis, modeling, and prediction. And there are a certain number of missing values in the coordinate time series of each GNSS station, which results in poor prediction effects when using traditional algorithms; the traditional support vector machine algorithm has a complex calculation process, low calculation efficiency for large-scale data, and high sensitivity to noise, and it is difficult to optimize parameters. The parameters of the traditional support vector machine model rely on manual determination, and the parameter selection error is relatively large.

[0004] Moreover, traditional predictions ignore the constraints of geophysical effect data and GNSS historical data, and do not fully integrate multi-source data with GNSS deformation data, which seriously hinders the rapid development of GNSS real-time high-precision deformation monitoring technology and the popularization and application of its deformation data.

[0005] Based on this, a GNSS coordinate time series prediction method based on genetic algorithm optimized SVM is proposed. Summary of the Invention

[0006] The purpose of the present invention is to provide a GNSS coordinate time series prediction method based on genetic algorithm optimized SVM to solve the problems in the background art.

[0007] To achieve the above purpose, the present invention provides a GNSS coordinate time series prediction method based on genetic algorithm optimized SVM, including the following steps:

[0008] S1. Data Processing: Collect the geophysical effect data and GNSS time series of multiple GNSS stations in a certain area to obtain the original time series of each GNSS station. Then, eliminate gross errors and interpolate and complete the original time series to obtain the complete time series of each GNSS station, and divide it into a training set and a test set.

[0009] S2. Construct a GA-SVM Model: Use the training set data as input, optimize the support vector machine model using the genetic algorithm, select the optimal parameter combination and import it into the support vector machine model for training to obtain a GA-SVM model that meets each GNSS station.

[0010] S3. Model Evaluation and Prediction: Predict the GA-SVM model using the training set and the test set respectively. Through multiple cross-validations, obtain the final GA-SVM model for GNSS coordinate time series prediction.

[0011] Preferably, in S1, the acquisition method of the original time series is as follows: Obtain the daily vertical crustal deformation time series and geophysical effect data of multiple GNSS stations respectively, process the geophysical effect data to obtain the daily geophysical effect data at the location of the GNSS station, and use the daily vertical crustal deformation time series and the daily geophysical effect data as input features, and the vertical crustal deformation GNSS as a single output feature to obtain the original time series.

[0012] Preferably, the geophysical effect data includes meteorological data, hydrological data, and sea level change data of each station.

[0013] Preferably, the specific steps of S1 are as follows:

[0014] S11. Fill in missing values and standardize the original time series data of each GNSS station: First, use the random forest regression model to process the original time series. The random forest model is:

[0015] ;

[0016] where, is the predicted value, is the predicted value of the th decision tree, D is the number of decision trees. For the missing values in the original time series, use other columns as features to train the random forest model, predict the missing values and fill them in;

[0017] After filling in the missing values, standardize the data to convert it into a form with a mean of 0 and a standard value of 1, expressed as:

[0018] ;

[0019] Among them, is the time series after filling in the missing values, is the mean value of the data, is the standard deviation, is a small constant;

[0020] S12. After the time series after processing is subjected to gross error rejection, a time series without gross errors and noise is obtained. The specific steps are as follows:

[0021] (1) Use the singular spectrum analysis method to decompose the standardized time series to obtain the reconstructed time series, and then use the Butterworth low-pass filter to process the reconstructed time series to extract the trend component of the time series;

[0022] Among them, the filtering formula is:

[0023] ;

[0024] Among them, , are the filter coefficients, is the input signal, is the trend component, is the discrete time sequence number, M is the maximum delay order of Z is the maximum delay order;

[0025] (2) Use the interquartile range method to start gross error detection. First, calculate the interquartile range, which is expressed as: ;

[0026] Among them, is the upper quartile, is the lower quartile;

[0027] Then determine the upper and lower bounds of the data. The data nodes that exceed the upper and lower bounds are considered to be gross errors;

[0028] Among them, the lower bound is:

[0029] ;

[0030] The upper bound is:

[0031] ;

[0032] (3) After completing the gross error detection, perform seasonal component calculation and noise calculation. Among them, the seasonal analysis calculation formula is expressed as:

[0033] ;

[0034] The noise calculation formula is expressed as:

[0035] ;

[0036] Among them, is the original time series after standardization, is the reconstructed time series;

[0037] S13. Use the Kriging interpolation method and the Kalman filter - random forest interpolation model to interpolate and supplement the data of each site obtained in S12 respectively.

[0038] Preferably, in step (1) of S12, the specific steps of using the singular spectrum analysis method to decompose the time series and obtain the reconstructed time series are as follows:

[0039] 1) Trajectory matrix construction: Convert the time series X processed in S12 into a trajectory matrix H, expressed as:

[0040] ;

[0041] Among them, is the window length, is the number of columns of the trajectory matrix;

[0042] 2) Singular value decomposition: Perform singular value decomposition on the trajectory matrix H to obtain:

[0043] ;

[0044] Among them, is the left singular vector matrix, is the singular value matrix, is the right singular value vector matrix, represents the transpose of V;

[0045] 3) Component selection and reconstruction: Select the main components according to the variance contribution rate of the singular values. The calculation formula of the variance contribution rate is expressed as:

[0046] ;

[0047] Among them, is the th singular value, is the total number of singular values, represents the j th singular value;

[0048] After selecting the components, use the first components selected , , to reconstruct the time series, expressed as:

[0049] .

[0050] Preferably, the specific steps of S2 are as follows:

[0051] S21. First, perform genetic encoding, convert the parameter penalty factor and the radial basis function parameters into genetic operation objects, use binary encoding, set the encoding range as [a, b], and the encoding precision as , and the encoding length as , satisfying the formula:

[0052] ;

[0053] Then initialize the population, set the population size W , the maximum number of iterations T , and globally search in the 2D space;

[0054] S22. Perform adaptive adjustment of the crossover probability, mutation probability, the number of elite parent populations, and the number of offspring populations;

[0055] S23. Make a process decision: retain elite parent populations, among which populations flow to the updated population, and populations perform the next operation;

[0056] If the termination principle is satisfied, jump out of the GA algorithm and perform SVM model training;

[0057] If the termination principle is not satisfied, continue with the genetic operations;

[0058] S24. Perform selection, mutation, and crossover operations in sequence to obtain offspring populations;

[0059] S25. Perform loop iteration: obtain offspring populations and return to the updated population, converge with populations and return to step S22 to perform loop operations until the required parameter combination is selected.

[0060] Preferably, the specific steps of S22 are as follows:

[0061] 1) Calculate the individual fitness using the fitness function, and the fitness function is expressed as:

[0062] ;

[0063] Among them, is the fitness function, is the number of samples included in the training set, is the predicted value, is the actual value;

[0064] 2) Adaptively adjust the crossover probability, mutation probability, the number of elite parent populations, and the number of offspring populations;

[0065] Based on the adjustment of individual fitness, determine the maximum fitness and minimum fitness of individuals in the population, then the crossover probability and mutation probability are expressed as:

[0066] ;

[0067] ;

[0068] where, is the crossover probability, is the mutation probability, is the maximum fitness of individuals in the population, is the average fitness of the population; , , , are respectively preset constants, is the individual fitness;

[0069] The number of elite parent populations and the number of offspring populations are determined according to the population size and algorithm settings. Among them, represents the number of elite parent populations, which is a certain proportion of the population size, The number of offspring populations is set according to the expected population update size and adjusted based on empirical formulas and experimental effects.

[0070] Preferably, the specific steps of S24 are:

[0071] 1) Selection operation: For populations, calculate the probability of being selected according to the fitness ratio, expressed as:

[0072] ;

[0073] where, represents the number of individuals, is the probability of an individual being selected;

[0074] 2) Mutation operation: Add a random number within a certain range to the original gene value to obtain the mutated gene value, expressed as:

[0075] ;

[0076] where, is the mutated gene value, is the original gene value, is a random number;

[0077] 3) Crossover operation. Randomly select two populations from populations, select the crossover point for crossover to obtain a new population, denoted as:

[0078] ;

[0079] ;

[0080] Among them, the two populations are A and B, , ; is the crossover point, , represent the populations after crossover.

[0081] Preferably, in S3, the root mean square error, mean absolute error, and goodness of fit are used to perform cross-validation on the test results;

[0082] Among them, the root mean square error is expressed as:

[0083] ;

[0084] The mean absolute error is expressed as:

[0085] ;

[0086] The goodness of fit is expressed as:

[0087] ;

[0088] Among them, represents the average value of the actual predicted values.

[0089] Therefore, a GNSS coordinate time series prediction method based on genetic algorithm optimized SVM of the present invention has the following beneficial effects:

[0090] (1) The present invention uses the genetic algorithm (GA) to search for the support vector machine (SVM) model parameters, finds the best parameter combination for constructing the SVM model, constructs the GA-SVM optimal prediction model, and fuses four types of feature data such as geophysical effects to train the GA-SVM model, greatly improving the accuracy of GNSS coordinate time series analysis and prediction, and can be applied to the identification of crustal movement precursor signals, geological disaster early warning, climate change, etc.

[0091] (2) In the present invention, GA is used to optimize SVM. By simulating natural selection and genetic mechanisms, it can search for the optimal solution globally, avoiding SVM from falling into local optima; it can reduce manual intervention. GA can automatically optimize the parameters of SVM (such as kernel function, regularization parameter, etc.), reducing the workload of manual parameter tuning.

[0092] (3) The GA-SVM model uses four types of characteristic data such as geophysical effects to predict GNSS. The amount of data is huge. Using GA for optimization can perform multi-objective optimization, solve the parameter optimization problem in high-dimensional data, and improve the performance of the model on high-dimensional data. At the same time, it provides sufficient training and validation data for the model, which can improve the training effect and prediction accuracy of the model, making the model have higher precision and generalization ability.

[0093] The technical solution of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings

[0094] Figure 1 It is the flowchart of the embodiment of the present invention;

[0095] Figure 2 It is the data construction process diagram in step S1 of the embodiment of the present invention;

[0096] Figure 3 It is the gross error image of the detection of four random GNSS stations in the embodiment of the present invention. Among them, (a) is LHAI, (b) is PCJM, (c) is XIAN, and (d) is YANT;

[0097] Figure 4 It is the interpolation result diagram of four stations in the embodiment of the present invention. Among them, (a) is LHAI, (b) is PCJM, (c) is XIAN, and (d) is YANT;

[0098] Figure 5 It is the construction flowchart of the genetic algorithm optimized support vector machine regression model in the embodiment of the present invention;

[0099] Figure 6 It is the model structure diagram of the support vector machine in the embodiment of the present invention;

[0100] Figure 7 It is the comparison diagram of the prediction results of the CORS station geodetic height time series training set based on the GA-SVM algorithm in the embodiment of the present invention. Among them, (a) is DONT, (b) is LHAI, (c) is PCJM, (d) is QNYN, (e) is SHAT, (f) is XIAG, (g) is YANT, and (h) is ZJYH;

[0101] Figure 8This is a comparison graph of the predicted results of the CORS station geodetic height time series test set based on the GA-SVM algorithm in the embodiment of the present invention. Among them, (a) is DONT, (b) is XIAG, (c) is PCJM, (d) is QNYN, (e) is SHAT, (f) is LHAI, (g) is YANT, and (h) is ZJYH;

[0102] Figure 9 This is a comparison graph of the predicted values and true values of the GA-SVM training set in the embodiment of the present invention. Among them, (a) is DONT, (b) is LHAI, (c) is PCJM, (d) is QNYN, (e) is SHAT, (f) is XIAG, (g) is YANT, and (h) is ZJYH;

[0103] Figure 10 This is a comparison graph of the predicted values and true values of the GA-SVM test set in the embodiment of the present invention. Among them, (a) is DONT, (b) is LHAI, (c) is PCJM, (d) is QNYN, (e) is SHAT, (f) is XIAG, (g) is YANT, and (h) is ZJYH;

[0104] Figure 11 This is a predicted error curve graph of the test set in the embodiment of the present invention. Among them, (a) is DONT, (b) is LHAI, (c) is PCJM, (d) is QNYN, (e) is SHAT, (f) is XIAG, (g) is YANT, and (h) is ZJYH. Detailed implementation manners

[0105] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0106] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.

[0107] Embodiment

[0108] In this embodiment, the above prediction method is used to perform prediction analysis on 38 GNSS stations in the Zhejiang region, as Figure 1 shown. The specific steps are as follows:

[0109] S1. Data processing:

[0110] 1) Acquisition of original data: The Zhejiang Institute of Surveying and Mapping provided the daily vertical crustal deformation time series (GNSS time series) of 38 GNSS stations in the Zhejiang region from January 1, 2015 to December 31, 2017. The data of 8 stations were selected as the prediction verification data for this invention.

[0111] Acquisition of geophysical effect data: 659 daily meteorological stations in China were obtained from the China Meteorological website, and the meteorological stations in the study area were screened out. MATLAB was used to read and process the temperature and air pressure daily meteorological data of the meteorological stations from January 1, 2015 to December 31, 2017; NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) provided the daily raster image data of terrestrial water storage and groundwater storage from January 1, 2015 to December 31, 2017 (hydrological data), where the hydrological time series is the sum of the two time series of terrestrial water storage and groundwater storage; the grid data of sea level difference time values from January 1, 2015 to December 31, 2017 (sea level change data) was obtained from the Aviso official website, but there were missing data in some time periods.

[0112] Processing of geophysical effect data: For the temperature and air pressure data, inverse distance interpolation was performed on the daily meteorological data using Arcmap, and the data of the meteorological stations were interpolated to the GNSS stations by the interpolation method; for the hydrological data, Arcmap was used to extract the raster image from the surface element to the GNSS stations; for the sea level change data, Matlab was used to extract the time value raster image data, and then Arcmap was used to extract the raster image data from the surface element to the GNSS stations. Since there are missing values in this type of data, least squares method was used for difference filling to obtain the complete sea level change time value data, and finally the daily meteorological data was obtained by taking the average value.

[0113] Through the above operations, the daily geophysical effect data of temperature, air pressure, hydrology, and sea level change at the location of the GNSS stations can be obtained.

[0114] Data construction: Let time t represent a certain day. In this embodiment, the four geophysical effect data of temperature (t), air pressure (t), hydrology (t), and sea level change (t) and the GNSS time series are used as the five input features for studying the genetic algorithm optimized support vector machine regression prediction model for GNSS coordinate time series analysis and prediction, and the vertical crustal deformation GNSS (t) is used as the single output feature of the model, obtaining the model import data as shown in Figure 2 And 80% of each station is used as the training set data, and 20% is used as the test set data.

[0115] Process the above data as follows:

[0116] S11. Import the data models of each site, and perform missing value filling and standardization on the original time series data of each GNSS site: First, use the random forest regression model to process the original time series. The random forest model is:

[0117] ;

[0118] where, is the predicted value, is the predicted value of the th decision tree, and D is the number of decision trees;

[0119] For the missing values in the original time series, use other columns as features, train the random forest model, predict and fill the missing values;

[0120] After completing the filling of missing values, standardize the data so that it is converted into a form with a mean of 0 and a standard value of 1, expressed as:

[0121] ;

[0122] where, is the mean of the data, is the standard deviation, is a small constant, which is 10 in this embodiment (-10) ; used to avoid division by zero error.

[0123] S12. Eliminate gross errors from the processed time series to obtain a time series without gross errors and noise. The specific steps are as follows:

[0124] (1) Use the singular spectrum analysis method to decompose the time series to obtain the reconstructed time series. The steps are as follows:

[0125] 1) Trajectory matrix construction: Convert the time series X processed in S12 into a trajectory matrix H, expressed as:

[0126] ;

[0127] where, is the window length, is the number of columns of the trajectory matrix;

[0128] 2) Singular value decomposition: Perform singular value decomposition on the trajectory matrix H to obtain:

[0129] ;

[0130] where, is the left singular vector matrix, is the singular value matrix (diagonal matrix), is the right singular value vector matrix;

[0131] 3) Component selection and reconstruction: Select the main components according to the variance contribution rate of the singular values. The calculation formula of the variance contribution rate is expressed as:

[0132] ;

[0133] where, is the th singular value, is the total number of singular values;

[0134] After selecting the components, use the first components , , to reconstruct the time series, expressed as:

[0135] .

[0136] Then use the Butterworth low-pass filter to process the reconstructed time series and extract the trend component of the time series;

[0137] where the filtering formula is:

[0138] ;

[0139] where, , are the filter coefficients, is the input signal, is the trend component, is the discrete time sequence number, M is the maximum delay order, Z is the maximum delay order;

[0140] (2) Use the interquartile range method to start gross error detection. First, calculate the interquartile range, expressed as: ;

[0141] where, is the upper quartile, which is the 75th percentile in this embodiment, is the lower quartile, which is the 25th percentile in this embodiment;

[0142] Then determine the upper and lower bounds of the data. The data nodes outside the upper and lower bounds are identified as gross errors;

[0143] where the lower bound is:

[0144] ;

[0145] The upper limit is:

[0146] ;

[0147] As Figure 3 shown, the detected gross error images obtained by randomly selecting four stations LHAI, PCJM, XIAG, and YANT.

[0148] (3) After completing the gross error detection, seasonal component calculation and noise calculation are performed. Among them, the seasonal analysis calculation formula is expressed as:

[0149] ;

[0150] The noise calculation formula is expressed as:

[0151] ;

[0152] Among them, is the reconstructed time series; through the above steps, the decomposition, gross error detection and optimization of the time series are realized, and the time series without gross errors and noise is obtained.

[0153] S13. The Kriging interpolation method and the Kalman filter - random forest interpolation model are used to interpolate and supplement the data of each station obtained in S12 respectively, and the results are shown in Table 1.

[0154] Table 1 Number of GNSS to be interpolated

[0155] ;

[0156] In this embodiment, the Kriging interpolation method (Kriging) and the Kalman filter - random forest interpolation model (KFI - RF) are used for interpolation respectively. The data of each station obtained in S12 are imported into the two interpolation models respectively. Table 2 below records the interpolation accuracies of the four stations LHAI, PCJM, XIAG, and YANT in the two models.

[0157] Table 2 Comparison of interpolation accuracies of KFI - RF and Kriging for the selected stations

[0158] ;

[0159] It can be seen from Table 2 that the root mean square error (RMSE) of using the Kriging interpolation method is on average reduced by 38.725% compared with the Kalman filter - random forest interpolation method, and the mean absolute error (MAE) is on average reduced by 48.4875%. Therefore, the results obtained by interpolating with the Kriging interpolation model are selected as the experimental data for subsequent prediction. As Figure 4It is the result map interpolated using the Kriging interpolation model by randomly selecting four stations, namely LHAI, PCJM, XIAG, and YANT.

[0160] S2. Construct the GA-SVM model: Use the training set data as input, optimize the support vector machine model using the genetic algorithm, select the optimal parameter combination and import it into the support vector machine model for training the support vector machine model to obtain the GA-SVM model that meets each GNSS station. Each station corresponds to a GA-SVM model. As Figure 5 shown in the flowchart for constructing the GA-SVM model.

[0161] First of all, the support vector machine (SVM) is a supervised learning algorithm commonly used in classification and regression analysis. The basic idea is to find an optimal hyperplane in the feature space to separate samples of different classes. SVM attempts to find a hyperplane such that sample points of different classes are as far away from the hyperplane as possible, while sample points of the same class are as close to it as possible. Such a hyperplane is called the maximum margin hyperplane and has good generalization ability. When the dataset in the existing space is linearly inseparable, the SVM algorithm can map the data to a high-dimensional space, transform the linearly inseparable dataset into a linearly separable dataset, and then create a hyperplane in this space such that the distance from the points closest to the hyperplane in the two groups of data to the hyperplane is the largest. The model structure of the support vector machine is as Figure 6 shown.

[0162] Extensive tests on the SVM model have shown that when constructing the SVM model, the accurate setting of its parameter penalty factor c and radial basis function parameter g has a great impact on the prediction accuracy and calculation speed of the SVM model. For the reasonable values of these two parameters c and g, finding their optimal combination can shorten the model calculation time and greatly improve the prediction accuracy. The traditional SVM algorithm has a high computational complexity, is sensitive to parameter selection, requires a large amount of manual search and trial-and-error calculations for model hyperparameters, and has low computational efficiency and is prone to model overfitting, resulting in the need to optimize the model accuracy. Therefore, using the GA optimization algorithm can effectively solve the problem of parameter selection optimization, can increase the global search ability, improve the generalization ability of the model, and thus can quickly obtain the global optimal parameters (c, g).

[0163] The specific steps for GA to optimize SVM are as follows:

[0164] S21. First, perform genetic coding, convert the problem parameters into genetic operation objects, use binary coding, set the coding range as [a, b], the coding precision as , and the coding length as , satisfying the formula:

[0165] ;

[0166] Then initialize the population and set the population size W , the maximum number of iterations T , and globally search in a 2D space;

[0167] S22. Perform adaptive adjustment of the crossover probability, mutation probability, the number of elite parent populations, and the number of offspring populations, as follows:

[0168] 1) Calculate the individual fitness. Usually, the fitness function can be the performance metric of the SVM model on the validation set. Here, the root mean square error is selected. First, use the fitness function to calculate the individual fitness, and the fitness function is expressed as:

[0169] ;

[0170] Among them, is the fitness function, is the number of samples contained in the training set, is the predicted value, is the real value;

[0171] 2) Adaptively adjust the crossover probability, mutation probability, the number of elite parent populations, and the number of offspring populations;

[0172] Based on the adjustment of individual fitness, determine the maximum fitness and minimum fitness of individuals in the population. Then, the crossover probability and mutation probability are expressed as:

[0173] ;

[0174] ;

[0175] Among them, is the crossover probability, is the mutation probability, is the maximum fitness of individuals in the population, is the average fitness of the population; , , , are respectively preset constants, is the individual fitness;

[0176] The number of elite parent populations and the number of offspring populations are determined according to the population size and algorithm settings. Among them, represents the number of elite parent populations, which is a certain proportion of the population size, the number of offspring populations, which is set according to the expected population update size and adjusted based on empirical formulas and experimental results.

[0177] S23. Make process decisions: Retain elite parent populations, among which populations flow to the updated population, populations proceed to the next step;

[0178] If the termination principle is satisfied, exit the GA algorithm and perform SVM model training;

[0179] If the termination principle is not satisfied, continue with the genetic operations;

[0180] S24. Perform selection, mutation, and crossover operations in sequence to obtain offspring populations;

[0181] 1) Selection operation: For populations, calculate the probability of being selected according to the fitness proportion, expressed as:

[0182] ;

[0183] Among them, represents the number of individuals, is the probability of an individual being selected;

[0184] 2) Mutation operation: Add a random number within a certain range to the original gene value to obtain the mutated gene value, expressed as:

[0185] ;

[0186] Among them, is the mutated gene value, is the original gene value, is the random number

[0187] 3) Crossover operation: Randomly select two populations from populations, select the crossover point for crossover to obtain a new population, expressed as:

[0188] ;

[0189] ;

[0190] Among them, the two populations are A and B, , ; is the crossover point, , represent the populations after crossover.

[0191] S25. Perform loop iteration: Obtain offspring populations and return them to the updated population, and The individuals in a population converge and return to step S22 for loop operations until parameter combinations that meet the requirements are screened out, and then SVM model training is performed to obtain the GA-SVM model that meets each site.

[0192] S3. Model evaluation and prediction: The GA-SVM model is used to make predictions on the training set and the test set respectively, and the model prediction results are exported. Through the accurate calculation of the root mean square error, mean absolute error, and goodness of fit,

[0193] Among them, the root mean square error is expressed as:

[0194] ;

[0195] The mean absolute error is expressed as:

[0196] ;

[0197] The goodness of fit is expressed as:

[0198] ;

[0199] Among them, represents the average of the actual predicted values.

[0200] After the calculation is completed, the comparison of the training set prediction results and the test set prediction results is plotted. As shown in Figure 7 and Figure 8 , the red line and the blue line represent the results of the training set and the test set respectively. It can be intuitively seen from the image that the GA-SVM model has a good fitting effect and generalization ability, and it can also test whether the model is overfitted. Obviously, the model for verifying the data is good.

[0201] As shown in Figure 9 and Figure 10 , the predicted values vs. the true values of the training set and the test set are compared. This type of image is simple and powerful, which can better analyze the error distribution, verify the generalization ability, and can assist in model tuning and result interpretation, and can help better understand and improve the model.

[0202] As shown in Figure 11 , it is the test set prediction error graph of the GA-SVM model, which shows the difference between the true value and the predicted value. The prediction errors of the selected sites are all within 7mm.

[0203] Table 3 RMSE prediction accuracy table

[0204] ;

[0205] Table 4 R 2 prediction accuracy table

[0206] ;

[0207] Table 5 MAE Prediction Accuracy Table

[0208] ;

[0209] Predictions were made using the traditional support vector machine algorithm (SVM), the particle swarm optimization support vector machine algorithm (PSO - SVM), and the genetic algorithm optimized random forest algorithm (GA - RF) respectively, and compared with the prediction method of the present invention. The prediction accuracy tables are shown in Table 3, Table 4, and Table 5. The RMSE in the training set of the GA - SVM model is within 1.5 mm, and the MAE is within 1.1 mm, and R 2 can reach 0.99; the RMSE in the test set is within 1.6 mm, and the MAE is within 1.3 mm, and R 2 can reach 0.95, and the lowest also reaches 0.85.

[0210] It can be seen from the prediction accuracy that the R of the traditional SVM algorithm and the constructed PSO - SVM and GA - RF models 2 is unstable, and there are overfitting and underfitting phenomena, which reflects that the prediction accuracy of the GA - SVM model on the test set is much higher than that of the original SVM model, the particle swarm optimization support vector machine algorithm (PSO - SVM), and the genetic algorithm optimized random forest algorithm (GA - RF). The improvement effect of the GA - SVM model on the test set is the most significant, and its goodness of fit R 2 has even achieved a leap - like growth. It can be seen that the GA - SVM model has stronger generalization ability and can make more accurate predictions.

[0211] Therefore, the GNSS coordinate time - series prediction method based on genetic algorithm optimized SVM of the present invention uses four characteristic data such as geophysical effects to predict GNSS, and uses GA for optimization, which can perform multi - objective optimization, solve the parameter optimization problem in high - dimensional data, and improve the performance of the model on high - dimensional data; at the same time, it provides sufficient training and verification data for the model, improves the training effect and prediction accuracy of the model, and makes the model have higher accuracy and generalization ability.

[0212] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements do not make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A GNSS coordinate time series prediction method based on genetic algorithm optimized SVM, characterized in that: The following steps are involved: S1. Data processing: Collect geophysical effect data and GNSS time series from multiple GNSS stations in a certain area to obtain the original time series of each GNSS station. Then remove the gross errors and interpolate the original time series to obtain the complete time series of each GNSS station, and divide them into training set and test set. The specific steps are as follows: S11. Fill missing values ​​and standardize the original time series data of each GNSS station: First, use the random forest regression model to process the original time series. The random forest model is: ; in, is the predicted value, It is The predicted value of a decision tree, D is the number of decision trees. For missing values ​​in the original time series, other columns are used as features to train a random forest model, predict missing values ​​and fill them in; After completing the missing value filling, the data is standardized to convert it into a form with a mean of 0 and a standard value of 1, expressed as: ; in, To fill in the missing values ​​of the time series, is the mean of the data, is the standard deviation, is a small constant; S12, removing gross errors from the processed time series to obtain a time series without gross errors and noise. The specific steps are as follows: (1) Use the singular spectrum analysis method to decompose the standardized time series to obtain the reconstructed time series. The specific steps are as follows: 1) Trajectory matrix construction: The time series processed by S12 X Converted to trajectory matrix H, expressed as: ; in, is the window length, is the number of columns of the trajectory matrix; 2) Singular value decomposition: Perform singular value decomposition on the trajectory matrix H and obtain: ; in, is the left singular vector matrix, is the singular value matrix, is the right singular value vector matrix, represents the transpose of V; 3) Component selection and reconstruction: The main components are selected according to the variance contribution rate of the singular value. The calculation formula of the variance contribution rate is expressed as: ; in, It is singular values, is the total number of singular values, Indicates singular values; After selecting the ingredients, use the selected Ingredients , , To reconstruct the time series, it is expressed as: ; Then, the reconstructed time series is processed using a Butterworth low-pass filter to extract the trend component of the time series; The filtering formula is: ; in, , are the filter coefficients, is the input signal, is the trend component, is a discrete time sequence number, M yes The maximum delay order of Z yes Maximum delay order; (2) Use the interquartile range method to start gross error detection. First, calculate the interquartile range, which is expressed as: ; in, is the upper quartile, is the lower quartile; Then determine the upper and lower limits of the data, and the data nodes beyond the upper and lower limits are considered to be gross errors; The lower limit is: ; The upper limit is: ; (3) After the gross error detection is completed, the seasonal component calculation and noise calculation are performed. The seasonal analysis calculation formula is expressed as: ; The noise calculation formula is expressed as: ; in, is the original time series after standardization, is the reconstructed time series; S13, using Kriging interpolation method and Kalman filter-random forest interpolation model to interpolate and supplement the data of each station obtained in S12; S2. Constructing a GA-SVM model: Taking the training set data as input, optimizing the support vector machine model using a genetic algorithm, selecting the optimal parameter combination and importing it into the support vector machine model, training the support vector machine model, and obtaining a GA-SVM model that meets the needs of each GNSS site; S3. Model evaluation and prediction: The GA-SVM model is predicted using the training set and the test set respectively. After multiple cross-validations, the final GA-SVM model is obtained to perform GNSS coordinate time series prediction.

2. The GNSS coordinate time series prediction method based on genetic algorithm optimized SVM according to claim 1, characterized in that: In S1, the original time series is obtained by respectively obtaining the daily time series of crustal vertical deformation and the geophysical effect data of multiple GNSS stations, processing the geophysical effect data to obtain the daily geophysical effect data of the location of the GNSS station, using the daily time series of crustal vertical deformation and the daily geophysical effect data as input features, and using the crustal vertical deformation GNSS as a single output feature to obtain the original time series.

3. The GNSS coordinate time series prediction method based on genetic algorithm optimized SVM according to claim 2 is characterized by: The geophysical effect data include meteorological data, hydrological data, and sea level change data of each site.

4. The GNSS coordinate time series prediction method based on genetic algorithm optimized SVM according to claim 1, characterized in that: The specific steps of S2 are: S21, first perform genetic coding, convert the parameter penalty factor and radial basis function parameters into genetic operation objects, use binary coding, set the coding range to [a, b], and the coding accuracy to , the encoding length is , satisfying the formula: ; Then initialize the population and set the population size W , maximum number of iterations T , global search in 2D space; S22, adaptively adjust the crossover probability, mutation probability, elite parent population number, and offspring population number; S23, make process decision: retain elite parent populations, among which The population flows to the renewal population, The next step is to proceed to the next population; If the termination principle is met, the GA algorithm is exited and the SVM model training is performed; If the termination principle is not met, the genetic operation continues; S24, perform selection, mutation, and crossover operations in sequence, and obtain offspring population; S25, perform loop iteration: get The offspring population returns to the updated population, and The populations are merged and the process returns to step S22 to repeat the loop operation until a parameter combination that meets the requirements is screened out.

5. The GNSS coordinate time series prediction method based on genetic algorithm optimized SVM according to claim 4 is characterized in that: The specific steps of S22 are: 1) Use the fitness function to calculate the individual fitness. The fitness function is expressed as: ; in, is the fitness function, is the number of samples included in the training set, is the predicted value, is a real value; 2) Adaptively adjust the crossover probability, mutation probability, and the number of elite parent populations and the number of offspring populations; Based on the adjustment of individual fitness, the maximum and minimum fitness of individuals in the population are determined, and the crossover probability and mutation probability are expressed as: ; ; in, is the crossover probability, is the mutation probability, is the maximum fitness of individuals in the population, is the average fitness of the population; , , , are preset constants, is the individual fitness; The number of elite parent populations and the number of offspring populations are determined according to the population size and algorithm settings.

6. The GNSS coordinate time series prediction method based on genetic algorithm optimized SVM according to claim 5, characterized in that: The specific steps of S24 are: 1) Select operation: The probability of being selected is calculated according to the fitness ratio of a population, expressed as: ; in, represents the number of individuals, is the probability of an individual being selected; 2) Mutation operation: The original gene value plus the random number is used to obtain the mutated gene value, which is expressed as: ; in, is the gene value after mutation, is the original gene value, is a random number; 3) Crossover operation, in A population randomly selects two populations, selects the intersection point for crossover to obtain a new population, which is expressed as: ; ; Among them, there are two populations, A and B. , ; is the intersection point, , Represents the population after crossover.

7. The GNSS coordinate time series prediction method based on genetic algorithm optimized SVM according to claim 1, characterized in that: In S3, the test results are cross-validated using root mean square error, mean absolute error and goodness of fit; The root mean square error is expressed as: ; The mean absolute error is expressed as: ; The goodness of fit is expressed as: ; in, Represents the average of the actual predicted values.

Citation Information

Patent Citations

  • UWB positioning method based on particle swarm optimization support vector machine and autonomous integrity monitoring

    CN115469267A

  • Continuous casting slab surface longitudinal crack process optimization method based on support vector machine and genetic algorithm

    CN118296476A