GNSS coordinate time sequence prediction method for optimizing SVM based on genetic algorithm
By using a support vector machine model optimized by genetic algorithm in GNSS coordinate timing prediction, the geophysical effect data is fused, and the problems of noise impact, missing values and calculation complexity in the prior art are solved, and high-precision GNSS coordinate timing prediction is achieved.
Patent Information
- Application Number
- CN202510422190.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The prior art has noise influence, missing value problems, high computational complexity, noise sensitivity and difficult parameter tuning in GNSS coordinate timing prediction, and has not fully integrated geophysical effect data and GNSS historical data.
Using a support vector machine (SVM) model based on genetic algorithm optimization, the GNSS coordinate timing is predicted through data processing and model construction.
The accuracy of GNSS coordinate timing analysis and prediction is improved, the workload of manual parameter adjustment is reduced, and the model's performance and generalization ability on high-dimensional data is improved.
Smart Images

Figure CN119939258A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of geodetic deformation monitoring, and in particular to a GNSS coordinate time series prediction method based on genetic algorithm optimized SVM. Background Art
[0002] In the European Space Agency's "Next Generation GNSS" program, the goal of physical model-enhanced GNSS coordinate time series prediction is to reduce the 7-day coordinate prediction error from centimeters to millimeters. Research on high-precision prediction methods for GNSS coordinate time series has important and far-reaching significance for the identification of precursor signals of crustal movement, geological disaster warning, and maintenance of the earth's reference frame. Geophysical effects (such as temperature, tidal and non-tidal effects, atmospheric and hydrological loads, etc.) can cause periodic or long-term displacements of the surface at the millimeter to centimeter level. There is still a lack of research on the prediction and analysis of GNSS coordinate time series by combining GNSS deformation data with geophysical effect data.
[0003] In this type of prediction method, there is noise in the GNSS coordinate time series, which will have many impacts on time series analysis, modeling and prediction, and there is a certain amount of missing in the coordinate time series of each GNSS station, which leads to poor prediction results when using traditional algorithms. The traditional support vector machine algorithm is complex to calculate, has low calculation efficiency for large-scale data, is highly sensitive to noise, and is difficult to tune parameters. The parameters of the traditional support vector machine model are manually determined, and the parameter selection error is large.
[0004] Moreover, traditional predictions ignore the constraints of geophysical effect data and GNSS historical data, and fail to fully integrate multi-source data with GNSS deformation data, which seriously hinders the rapid development of GNSS real-time high-precision deformation monitoring technology and the promotion and application of its deformation data.
[0005] Based on this, a GNSS coordinate time series prediction method based on genetic algorithm optimized SVM is proposed. Summary of the invention
[0006] The purpose of the present invention is to provide a GNSS coordinate time series prediction method based on genetic algorithm optimization SVM to solve the problems in the background technology.
[0007] To achieve the above object, the present invention provides a GNSS coordinate time series prediction method based on genetic algorithm optimized SVM, comprising the following steps: S1. Data processing: Collect geophysical effect data and GNSS time series of multiple GNSS stations in a certain area to obtain the original time series of each GNSS station, then remove gross errors and perform interpolation to complete the original time series to obtain the complete time series of each GNSS station, and divide them into training set and test set; S2. Constructing a GA-SVM model: Taking the training set data as input, optimizing the support vector machine model using a genetic algorithm, selecting the optimal parameter combination and importing it into the support vector machine model, training the support vector machine model, and obtaining a GA-SVM model that meets the needs of each GNSS site; S3. Model evaluation and prediction: The GA-SVM model is predicted using the training set and the test set respectively. After multiple cross-validations, the final GA-SVM model is obtained to perform GNSS coordinate time series prediction.
[0008] Preferably, in S1, the original time series is obtained by respectively obtaining the daily time series of crustal vertical deformation and geophysical effect data of multiple GNSS stations, processing the geophysical effect data to obtain the daily geophysical effect data of the location of the GNSS station, using the daily time series of crustal vertical deformation and the daily geophysical effect data as input features, and using the crustal vertical deformation GNSS as a single output feature to obtain the original time series.
[0009] Preferably, the geophysical effect data include meteorological data, hydrological data, and sea level change data of each site.
[0010] Preferably, the specific steps of S1 are: S11. Fill missing values and standardize the original time series data of each GNSS station: First, use the random forest regression model to process the original time series. The random forest model is: ; in, is the predicted value, It is The predicted value of a decision tree, D is the number of decision trees. For missing values in the original time series, other columns are used as features to train a random forest model, predict missing values and fill them in; After completing the missing value filling, the data is standardized to convert it into a form with a mean of 0 and a standard value of 1, expressed as: ; in, To fill in the missing values of the time series, is the mean of the data, is the standard deviation, is a small constant; S12, removing gross errors from the processed time series to obtain a time series without gross errors and noise. The specific steps are as follows: (1) Use the singular spectrum analysis method to decompose the standardized time series to obtain the reconstructed time series, and then use the Butterworth low-pass filter to process the reconstructed time series to extract the trend component of the time series; The filtering formula is: ; in, , are the filter coefficients, is the input signal, is the trend component, is a discrete time sequence number, M yes The maximum delay order of Z yes Maximum delay order; (2) Use the interquartile range method to start gross error detection. First, calculate the interquartile range, which is expressed as: ; in, is the upper quartile, is the lower quartile; Then determine the upper and lower limits of the data, and the data nodes beyond the upper and lower limits are considered to be gross errors; The lower limit is: ; The upper limit is: ; (3) After the gross error detection is completed, the seasonal component calculation and noise calculation are performed. The seasonal analysis calculation formula is expressed as: ; The noise calculation formula is expressed as: ; in, is the original time series after standardization, is the reconstructed time series; S13. Use Kriging interpolation method and Kalman filter-random forest interpolation model to interpolate and supplement the data of each station obtained in S12.
[0011] Preferably, in step (1) of S12, the specific steps of using the singular spectrum analysis method to decompose the time series to obtain the reconstructed time series are: 1) Trajectory matrix construction: The time series X processed by S12 is converted into the trajectory matrix H, which is expressed as: ; in, is the window length, is the number of columns of the trajectory matrix; 2) Singular value decomposition: Perform singular value decomposition on the trajectory matrix H and obtain: ; in, is the left singular vector matrix, is the singular value matrix, is the right singular value vector matrix, represents the transpose of V; 3) Component selection and reconstruction: The main components are selected according to the variance contribution rate of the singular value. The calculation formula of the variance contribution rate is expressed as: ; in, It is singular values, is the total number of singular values, Indicates j singular values; After selecting the ingredients, use the selected Ingredients , , To reconstruct the time series, it is expressed as: .
[0012] Preferably, the specific steps of S2 are: S21, first perform genetic coding, convert the parameter penalty factor and radial basis function parameters into genetic operation objects, use binary coding, set the coding range to [a, b], and the coding accuracy to , the encoding length is , satisfying the formula: ; Then initialize the population and set the population size W , maximum number of iterations T , global search in 2D space; S22, adaptively adjust the crossover probability, mutation probability, elite parent population number, and offspring population number; S23, make process decision: retain elite parent populations, among which The population flows to the renewal population, The next step is to proceed to the next population; If the termination principle is met, the GA algorithm is exited and the SVM model training is performed; If the termination principle is not met, the genetic operation continues; S24, perform selection, mutation, and crossover operations in sequence, and obtain offspring population; S25, perform loop iteration: get The offspring population returns to the updated population, and The populations are merged and the process returns to step S22 to repeat the loop operation until a parameter combination that meets the requirements is screened out.
[0013] Preferably, the specific steps of S22 are: 1) Use the fitness function to calculate the individual fitness. The fitness function is expressed as: ; in, is the fitness function, is the number of samples included in the training set, is the predicted value, is a real value; 2) Adaptively adjust the crossover probability, mutation probability, and the number of elite parent populations and the number of offspring populations; Based on the adjustment of individual fitness, the maximum and minimum fitness of individuals in the population are determined, and the crossover probability and mutation probability are expressed as: ; ; in, is the crossover probability, is the mutation probability, is the maximum fitness of individuals in the population, is the average fitness of the population; , , , are preset constants, is the individual fitness; The number of elite parent populations and the number of offspring populations are determined according to the population size and algorithm settings, where: represents the number of elite parent populations, which is a certain proportion of the population size. The number of offspring populations is set according to the expected population update scale and adjusted based on empirical formulas and experimental results.
[0014] Preferably, the specific steps of S24 are: 1) Select operation: The probability of being selected is calculated according to the fitness ratio of a population, expressed as: ; in, represents the number of individuals, is the probability of an individual being selected; 2) Mutation operation: The original gene value is added with a random number within a certain range to obtain the mutated gene value, expressed as: ; in, is the gene value after mutation, is the original gene value, is a random number; 3) Crossover operation, in A population randomly selects two populations, selects the intersection point for crossover to get a new population, which is expressed as: ; ; There are two populations, A and B. , ; is the intersection point, , Represents the population after crossover.
[0015] Preferably, in S3, the test results are cross-validated using root mean square error, mean absolute error and goodness of fit; The root mean square error is expressed as: ; The mean absolute error is expressed as: ; The goodness of fit is expressed as: ; in, Represents the average of the actual predicted values.
[0016] Therefore, the present invention provides a GNSS coordinate time series prediction method based on genetic algorithm optimization SVM, which has the following beneficial effects: (1) The present invention uses a genetic algorithm (GA) to search for support vector machine (SVM) model parameters, finds the best parameter combination for constructing the SVM model, constructs a GA-SVM optimal prediction model, and integrates four kinds of feature data such as geophysical effects to train the GA-SVM model, which greatly improves the accuracy of GNSS coordinate time series analysis and prediction, and can be applied to crust movement precursor signal identification, geological disaster warning, climate change, etc.
[0017] (2) In the present invention, GA is used to optimize SVM. By simulating natural selection and genetic mechanisms, it can search for the optimal solution in a global range and avoid SVM falling into local optimality. It can reduce human intervention. GA can automatically optimize the parameters of SVM (such as kernel function, regularization parameter, etc.), reducing the workload of manual parameter adjustment.
[0018] (3) The GA-SVM model uses four feature data such as geophysical effects to predict GNSS. The amount of data is huge. GA can be used for optimization to perform multi-objective optimization, solve the parameter optimization problem in high-dimensional data, and improve the performance of the model on high-dimensional data. At the same time, sufficient training and verification data is provided for the model, which can improve the training effect and prediction accuracy of the model, making the model more accurate and generalizable.
[0019] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flow chart of an embodiment of the present invention; Figure 2 This is a diagram of the data construction process in step S1 of an embodiment of the present invention; Figure 3 These are gross error images detected by four random GNSS stations according to an embodiment of the present invention, where (a) is LHAI, (b) is PCJM, (c) is XIAN, and (d) is YANT; Figure 4 The interpolation result diagrams of four stations in the embodiment of the present invention, wherein (a) is LHAI, (b) is PCJM, (c) is XIAN, and (d) is YANT; Figure 5 A flowchart of constructing a genetic algorithm optimized support vector machine regression model according to an embodiment of the present invention; Figure 6 A model structure diagram of a support vector machine according to an embodiment of the present invention; Figure 7 This is a comparison chart of the prediction results of the CORS station geodetic time series training set based on the GA-SVM algorithm according to an embodiment of the present invention, where (a) is DONT, (b) is LHAI, (c) is PCJM, (d) is QNYN, (e) is SHAT, (f) is XIAG, (g) is YANT, and (h) is ZJYH; Figure 8This is a comparison chart of the prediction results of the CORS station geodetic time series test set based on the GA-SVM algorithm according to an embodiment of the present invention, where (a) is DONT, (b) is XIAG, (c) is PCJM, (d) is QNYN, (e) is SHAT, (f) is LHAI, (g) is YANT, and (h) is ZJYH; Fig. 9 1 is a comparison diagram of the predicted value and the true value based on the GA-SVM training set according to an embodiment of the present invention, wherein (a) is DONT, (b) is LHAI, (c) is PCJM, (d) is QNYN, (e) is SHAT, (f) is XIAG, (g) is YANT, and (h) is ZJYH; Fig.10 1 is a comparison chart of the predicted value and the true value based on the GA-SVM test set according to an embodiment of the present invention, wherein (a) is DONT, (b) is LHAI, (c) is PCJM, (d) is QNYN, (e) is SHAT, (f) is XIAG, (g) is YANT, and (h) is ZJYH; Fig.11 1 is a test set prediction error curve diagram of an embodiment of the present invention, where (a) is DONT, (b) is LHAI, (c) is PCJM, (d) is QNYN, (e) is SHAT, (f) is XIAG, (g) is YANT, and (h) is ZJYH. DETAILED DESCRIPTION
[0021] The technical solution of the present invention is further described below through the accompanying drawings and embodiments.
[0022] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.
[0023] Example This embodiment uses the above prediction method to perform prediction analysis on 38 GNSS stations in Zhejiang Province. Figure 1 As shown, the specific steps are as follows: S1. Data processing: 1) Original data acquisition: Zhejiang Provincial Institute of Surveying and Mapping provided the daily time series (GNSS time series) of crustal vertical deformation from January 1, 2015 to December 31, 2017 at 38 GNSS stations in Zhejiang Province, from which data from 8 stations were selected as the prediction verification data for this invention.
[0024] Acquisition of geophysical effect data: 659 daily meteorological stations in my country were obtained from the China Meteorological Website and the meteorological stations in the study area were screened out. MATLAB was used to read and process the daily meteorological data of temperature and air pressure from January 1, 2015 to December 31, 2017. NASA Goddard Earth Sciences Data and Information Servives Center (GES DISC) provided daily raster image data (hydrological data) of terrestrial water reserves and groundwater reserves from January 1, 2015 to December 31, 2017, where the hydrological time series is the sum of the terrestrial water reserves and groundwater reserves. The sea level difference time grid data (sea level change data) from January 1, 2015 to December 31, 2017 was obtained from the Aviso official website, but the data was missing in some time periods.
[0025] Geophysical effect data processing: For temperature and pressure data, Arcmap is used to perform inverse distance difference on the daily meteorological data, and the data of the meteorological station is interpolated to the GNSS station by interpolation; for hydrological data, Arcmap is used to extract raster images from surface elements to GNSS stations; for sea level change data, Matlab is used to extract time-value raster image data, and then Arcmap is used to extract raster image data from surface elements to GNSS stations. Since this type of data has missing values, the least squares method is used to perform difference supplement to obtain complete sea level change time-value data, and finally the daily meteorological data is obtained by averaging.
[0026] Through the above operations, we can obtain the daily data of four geophysical effects: temperature, air pressure, hydrology, and sea level change at the location of the GNSS station.
[0027] Data construction: Time t represents a certain day. This embodiment uses four geophysical effect data, temperature (t), air pressure (t), hydrology (t), and sea level change (t), and GNSS time series as five input features of the genetic algorithm optimized support vector machine regression prediction model for GNSS coordinate time series analysis and prediction, and uses crustal vertical deformation GNSS (t) as the single output feature of the model. Figure 2 The model imports data and uses the first 80% of each site as the training set data and the last 20% as the test set data.
[0028] The above data are processed as follows: S11. Import the data model of each station, fill missing values and standardize the original time series data of each GNSS station: First, use the random forest regression model to process the original time series. The random forest model is: ; in, is the predicted value, It is The predicted value of a decision tree, D is the number of decision trees; For missing values in the original time series, use other columns as features, train a random forest model, predict missing values and fill them; After completing the missing value filling, the data is standardized to convert it into a form with a mean of 0 and a standard value of 1, expressed as: ; in, is the mean of the data, is the standard deviation, is a small constant, which is 10 in this embodiment. (-10) ; Used to avoid division by zero errors.
[0029] S12, removing gross errors from the processed time series to obtain a time series without gross errors and noise. The specific steps are as follows: (1) Use the singular spectrum analysis method to decompose the time series and obtain the reconstructed time series. The steps are as follows: 1) Trajectory matrix construction: The time series X processed by S12 is converted into the trajectory matrix H, which is expressed as: ; in, is the window length, is the number of columns of the trajectory matrix; 2) Singular value decomposition: Perform singular value decomposition on the trajectory matrix H and obtain: ; in, is the left singular vector matrix, is the singular value matrix (diagonal matrix), is the right singular value vector matrix; 3) Component selection and reconstruction: The main components are selected according to the variance contribution rate of the singular value. The calculation formula of the variance contribution rate is expressed as: ; in, It is singular values, is the total number of singular values; After selecting the ingredients, use the selected Ingredients , , To reconstruct the time series, it is expressed as: .
[0030] Then, the reconstructed time series is processed using a Butterworth low-pass filter to extract the trend component of the time series; The filtering formula is: ; in, , are the filter coefficients, is the input signal, is the trend component, is a discrete time sequence number, M is The maximum delay order of Z yes Maximum delay order; (2) Use the interquartile range method to start gross error detection. First, calculate the interquartile range, which is expressed as: ; in, is the upper quartile, in this embodiment the 75th percentile, is the lower quartile, which is the 25th percentile in this embodiment; Then determine the upper and lower limits of the data, and the data nodes beyond the upper and lower limits are considered to be gross errors; The lower limit is: ; The upper limit is: ; like Figure 3 As shown in Figure 1, the gross error images detected were obtained by randomly selecting four stations: LHAI, PCJM, XIAG, and YANT.
[0031] (3) After the gross error detection is completed, the seasonal component calculation and noise calculation are performed. The seasonal analysis calculation formula is expressed as: ; The noise calculation formula is expressed as: ; in, is the reconstructed time series; through the above steps, the decomposition, gross error detection and optimization of the time series are realized, and a time series without gross errors and noise is obtained.
[0032] S13. Kriging interpolation method and Kalman filter-random forest interpolation model were used to interpolate and supplement the data of each station obtained in S12. The results are shown in Table 1.
[0033] Table 1 Number of GNSSs that need to be interpolated ;
[0034] In this embodiment, Kriging interpolation method (Kriging) and Kalman filter-random forest interpolation model (KFI-RF) are used for interpolation respectively. The data of each station obtained by S12 are imported into the two interpolation models respectively. Table 2 below records the interpolation accuracy of LHAI, PCJM, XIAG, and YANT of the four stations in the two models. Table 2 Comparison of KFI-RF and Kriging interpolation accuracy at selected sites ;
[0035] It can be seen from Table 2 that the root mean square error (RMSE) of the Kriging interpolation method is reduced by an average of 38.725% and the mean absolute error (MAE) is reduced by an average of 48.4875% compared with the Kalman filter-random forest interpolation method. Therefore, the results obtained by the Kriging interpolation model are selected as the experimental data for subsequent predictions. Figure 4 This is the interpolation result map of four randomly selected stations LHAI, PCJM, XIAG, and YANT using the Kriging interpolation model.
[0036] S2. Construct GA-SVM model: Take the training set data as input, use genetic algorithm to optimize the support vector machine model, select the optimal parameter combination and import it into the support vector machine model, train the support vector machine model, and obtain a GA-SVM model that meets each GNSS station. Each station corresponds to a GA-SVM model, such as Figure 5 Shown is the construction flow chart of the GA-SVM model.
[0037] First of all, support vector machine (SVM) is a supervised learning algorithm, which is often used for classification and regression analysis. The basic idea is to find an optimal hyperplane in the feature space to separate samples of different categories. SVM tries to find a hyperplane so that sample points of different categories are as far away from the hyperplane as possible, while making sample points of the same category as close to it as possible. Such a hyperplane is called a maximum margin hyperplane and has good generalization ability. When the data set in the existing space is linearly inseparable, the SVM algorithm can map the data to a high-dimensional space, convert the linearly inseparable data set into a linearly separable data set, and then create a hyperplane in this space so that the distance between the points closest to the hyperplane of the two sets of data is the largest. The model structure of the support vector machine is as follows: Figure 6 shown.
[0038] After a large number of tests on the SVM model, it is shown that when constructing the SVM model, the accurate setting of the parameter penalty factor c and the radial basis function parameter g has a great impact on the prediction accuracy and calculation speed of the SVM model. For the reasonable values of the two parameters c and g, finding their optimal combination can shorten the model calculation time and greatly improve the prediction accuracy. The traditional SVM algorithm has high computational complexity and is sensitive to parameter selection. It requires a lot of manual search and trial and error calculations for model hyperparameters. In addition, the computational efficiency is low and the model is prone to overfitting, resulting in the need to optimize the model accuracy. Therefore, the use of the GA optimization algorithm can effectively solve the problem of parameter selection optimization, increase the global search capability, and improve the generalization ability of the model, so that the global optimal parameters (c, g) can be quickly obtained.
[0039] The specific steps of GA optimization SVM are as follows: S21, first perform genetic coding, convert the problem parameters into genetic operation objects, use binary coding, set the coding range to [a, b], and the coding accuracy to , the encoding length is , satisfying the formula: ; Then initialize the population and set the population size W , maximum number of iterations T , global search in 2D space; S22, adaptively adjust the crossover probability, mutation probability, elite parent population number, and offspring population number, as follows: 1) Calculate individual fitness. Usually, the fitness function can be the performance index of the SVM model on the validation set. Here, the root mean square error is selected. First, use the fitness function to calculate the individual fitness. The fitness function is expressed as: ; in, is the fitness function, is the number of samples included in the training set, is the predicted value, is a real value; 2) Adaptively adjust the crossover probability, mutation probability, and the number of elite parent populations and the number of offspring populations; Based on the adjustment of individual fitness, the maximum and minimum fitness of individuals in the population are determined, and the crossover probability and mutation probability are expressed as: ; ; in, is the crossover probability, is the mutation probability, is the maximum fitness of individuals in the population, is the average fitness of the population; , , , are preset constants, is the individual fitness; The number of elite parent populations and the number of offspring populations are determined according to the population size and algorithm settings, where: represents the number of elite parent populations, which is a certain proportion of the population size. The number of offspring populations is set according to the expected population update scale and adjusted based on empirical formulas and experimental results.
[0040] S23, make process decision: retain elite parent populations, among which The population flows to the renewal population, The next step is to proceed to the next population; If the termination principle is met, the GA algorithm is exited and the SVM model training is performed; If the termination principle is not met, the genetic operation continues; S24, perform selection, mutation, and crossover operations in sequence, and obtain offspring population; 1) Select operation: The probability of being selected is calculated according to the fitness ratio of a population, expressed as: ; in, represents the number of individuals, is the probability of an individual being selected; 2) Mutation operation: The original gene value is added with a random number within a certain range to obtain the mutated gene value, expressed as: ; in, is the gene value after mutation, is the original gene value, is a random number 3) Crossover operation, in A population randomly selects two populations, selects the intersection point for crossover to get a new population, which is expressed as: ; ; There are two populations, A and B. , ; is the intersection point, , Represents the population after crossover.
[0041] S25, perform loop iteration: get The offspring population returns to the updated population, and The populations are merged and the process returns to step S22, and a loop operation is performed until a parameter combination that meets the requirements is screened out, and SVM model training is performed to obtain a GA-SVM model that meets the requirements of each site.
[0042] S3. Model evaluation and prediction: The GA-SVM model is predicted using the training set and the test set respectively, and the model prediction results are exported. The root mean square error, mean absolute error and goodness of fit are accurately calculated. The root mean square error is expressed as: ; The mean absolute error is expressed as: ; The goodness of fit is expressed as: ; in, Represents the average of the actual predicted values.
[0043] After the calculation is completed, the prediction results of the training set and the test set are compared, as shown in the figure below. Figure 7 , Figure 8 As shown in the figure, the red line and the blue line represent the results of the training set and the test set respectively. From the image, we can intuitively see that the GA-SVM model has a good fitting effect and generalization ability, and we can test whether the model is overfitting. Obviously, the model of the verification data is good.
[0044] like Fig. 9 , Fig.10 As shown in the figure, the predicted values of the training set vs. the true values and the predicted values of the test set vs. the true values are compared. This type of graph is simple and powerful, and it can better analyze the error distribution, verify the generalization ability, and assist in model tuning and result interpretation, which can help to better understand and improve the model.
[0045] like Fig.11 The following is a plot of the prediction error of the test set of the GA-SVM model, showing the difference between the true value and the predicted value. The prediction errors of the selected stations are all within 7 mm.
[0046] Table 3 RMSE prediction accuracy table ;
[0047] Table 4 R 2 Measuring accuracy table ;
[0048] Table 5 MAE prediction accuracy table ;
[0049] The traditional support vector machine algorithm (SVM), particle swarm optimized support vector machine algorithm (PSO-SVM) and genetic algorithm optimized random forest algorithm (GA-RF) were used for prediction and compared with the prediction method of the present invention. The prediction accuracy is shown in Tables 3, 4 and 5. In the training set of the GA-SVM model, the RMSE is within 1.5 mm, the MAE is within 1.1 mm, and the R 2 can reach 0.99; in the test set, RMSE is within 1.6mm, MAE is within 1.3mm, R 2 It can reach 0.95, and the lowest is 0.85.
[0050] It can be seen from the prediction accuracy that the R of the traditional SVM algorithm and the constructed PSO-SVM and GA-RF models are 2 Unstable, with overfitting and underfitting, which shows that the prediction accuracy of the GA-SVM model on the test set is much higher than that of the original SVM model, the particle swarm optimized support vector machine algorithm (PSO-SVM), and the genetic algorithm optimized random forest algorithm (GA-RF). The GA-SVM model has the most significant improvement effect on the test set, and its goodness of fit R 2 It has even achieved a leapfrog growth. It can be seen that the GA-SVM model has stronger generalization ability and can make more accurate predictions.
[0051] Therefore, the present invention provides a GNSS coordinate time series prediction method based on genetic algorithm optimized SVM, which uses four characteristic data such as geophysical effects to predict GNSS and uses GA for optimization. It can perform multi-objective optimization, solve the parameter optimization problem in high-dimensional data, and improve the performance of the model on high-dimensional data. At the same time, it provides sufficient training and verification data for the model, improves the training effect and prediction accuracy of the model, and makes the model have higher accuracy and generalization ability.
[0052] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A GNSS coordinate time series prediction method based on genetic algorithm optimized SVM, characterized in that: The following steps are involved: S1. Data processing: Collect geophysical effect data and GNSS time series of multiple GNSS stations in a certain area to obtain the original time series of each GNSS station, then remove gross errors and perform interpolation to complete the original time series to obtain the complete time series of each GNSS station, and divide them into training set and test set; S2. Constructing a GA-SVM model: Taking the training set data as input, optimizing the support vector machine model using a genetic algorithm, selecting the optimal parameter combination and importing it into the support vector machine model, training the support vector machine model, and obtaining a GA-SVM model that meets the needs of each GNSS site; S3. Model evaluation and prediction: The GA-SVM model is predicted using the training set and the test set respectively. After multiple cross-validations, the final GA-SVM model is obtained to perform GNSS coordinate time series prediction.
2. The GNSS coordinate time series prediction method based on genetic algorithm optimized SVM according to claim 1, characterized in that: In S1, the original time series is obtained by respectively obtaining the daily time series of crustal vertical deformation and the geophysical effect data of multiple GNSS stations, processing the geophysical effect data to obtain the daily geophysical effect data of the location of the GNSS station, using the daily time series of crustal vertical deformation and the daily geophysical effect data as input features, and using the crustal vertical deformation GNSS as a single output feature to obtain the original time series.
3. The GNSS coordinate time series prediction method based on genetic algorithm optimized SVM according to claim 2 is characterized by: The geophysical effect data include meteorological data, hydrological data, and sea level change data of each site.
4. The GNSS coordinate time series prediction method based on genetic algorithm optimized SVM according to claim 1, characterized in that: The specific steps of S1 are: S11. Fill missing values and standardize the original time series data of each GNSS station: First, use the random forest regression model to process the original time series. The random forest model is: ; in, is the predicted value, It is The predicted value of a decision tree, D is the number of decision trees. For missing values in the original time series, other columns are used as features to train a random forest model, predict missing values and fill them in; After completing the missing value filling, the data is standardized to convert it into a form with a mean of 0 and a standard value of 1, expressed as: ; in, To fill in the missing values of the time series, is the mean of the data, is the standard deviation, is a small constant; S12, removing gross errors from the processed time series to obtain a time series without gross errors and noise. The specific steps are as follows: (1) Use the singular spectrum analysis method to decompose the standardized time series to obtain the reconstructed time series, and then use the Butterworth low-pass filter to process the reconstructed time series to extract the trend component of the time series; The filtering formula is: ; in, , are the filter coefficients, is the input signal, is the trend component, is a discrete time sequence number, M yes The maximum delay order of Z yes Maximum delay order; (2) Use the interquartile range method to start gross error detection. First, calculate the interquartile range, which is expressed as: ; in, is the upper quartile, is the lower quartile; Then determine the upper and lower limits of the data, and the data nodes beyond the upper and lower limits are considered to be gross errors; The lower limit is: ; The upper limit is: ; (3) After the gross error detection is completed, the seasonal component calculation and noise calculation are performed. The seasonal analysis calculation formula is expressed as: ; The noise calculation formula is expressed as: ; in, is the original time series after standardization, is the reconstructed time series; S13. Use Kriging interpolation method and Kalman filter-random forest interpolation model to interpolate and supplement the data of each station obtained in S12.
5. The GNSS coordinate time series prediction method based on genetic algorithm optimized SVM according to claim 4 is characterized in that: In step (1) of S12, the specific steps of using the singular spectrum analysis method to decompose the time series to obtain the reconstructed time series are: 1) Trajectory matrix construction: The time series processed by S12 X Converted to trajectory matrix H, expressed as: ; in, is the window length, is the number of columns of the trajectory matrix; 2) Singular value decomposition: Perform singular value decomposition on the trajectory matrix H and obtain: ; in, is the left singular vector matrix, is the singular value matrix, is the right singular value vector matrix, represents the transpose of V; 3) Component selection and reconstruction: The main components are selected according to the variance contribution rate of the singular value. The calculation formula of the variance contribution rate is expressed as: ; in, It is singular values, is the total number of singular values, Indicates singular values; After selecting the ingredients, use the selected Ingredients , , To reconstruct the time series, it is expressed as: 。 6. The GNSS coordinate time series prediction method based on genetic algorithm optimized SVM according to claim 1, characterized in that: The specific steps of S2 are: S21, first perform genetic coding, convert the parameter penalty factor and radial basis function parameters into genetic operation objects, use binary coding, set the coding range to [a, b], and the coding accuracy to , the encoding length is , satisfying the formula: ; Then initialize the population and set the population size W , maximum number of iterations T , global search in 2D space; S22, adaptively adjust the crossover probability, mutation probability, elite parent population number, and offspring population number; S23, make process decision: retain elite parent populations, among which The population flows to the renewal population, The next step is to proceed to the next population; If the termination principle is met, the GA algorithm is exited and the SVM model training is performed; If the termination principle is not met, the genetic operation continues; S24, perform selection, mutation, and crossover operations in sequence, and obtain offspring population; S25, perform loop iteration: get The offspring population returns to the updated population, and The populations are merged and the process returns to step S22 to repeat the loop operation until a parameter combination that meets the requirements is screened out.
7. The GNSS coordinate time series prediction method based on genetic algorithm optimized SVM according to claim 6, characterized in that: The specific steps of S22 are: 1) Use the fitness function to calculate the individual fitness. The fitness function is expressed as: ; in, is the fitness function, is the number of samples included in the training set, is the predicted value, is a real value; 2) Adaptively adjust the crossover probability, mutation probability, and the number of elite parent populations and the number of offspring populations; Based on the adjustment of individual fitness, the maximum and minimum fitness of individuals in the population are determined, and the crossover probability and mutation probability are expressed as: ; ; in, is the crossover probability, is the mutation probability, is the maximum fitness of individuals in the population, is the average fitness of the population; , , , are preset constants, is the individual fitness; The number of elite parent populations and the number of offspring populations are determined according to the population size and algorithm settings.
8. The GNSS coordinate time series prediction method based on genetic algorithm optimized SVM according to claim 7, characterized in that: The specific steps of S24 are: 1) Select operation: The probability of being selected is calculated according to the fitness ratio of a population, expressed as: ; in, represents the number of individuals, is the probability of an individual being selected; 2) Mutation operation: The original gene value plus the random number is used to obtain the mutated gene value, which is expressed as: ; in, is the gene value after mutation, is the original gene value, is a random number; 3) Crossover operation, in A population randomly selects two populations, selects the intersection point for crossover to obtain a new population, which is expressed as: ; ; Among them, there are two populations, A and B. , ; is the intersection point, , Represents the population after crossover.
9. The GNSS coordinate time series prediction method based on genetic algorithm optimized SVM according to claim 1, characterized in that: In S3, the test results are cross-validated using root mean square error, mean absolute error and goodness of fit; The root mean square error is expressed as: ; The mean absolute error is expressed as: ; The goodness of fit is expressed as: ; in, Represents the average of the actual predicted values.
Citation Information
Patent Citations
Method for stability analysis of reference point for GNSS automatic deformation monitoring
CN106066901A
Prediction method of foundation pit displacement based on GA-BP neural network
CN109543237A
UWB positioning method based on particle swarm optimization support vector machine and autonomous integrity monitoring
CN115469267A
Continuous casting slab surface longitudinal crack process optimization method based on support vector machine and genetic algorithm
CN118296476A
Earth crust deformation time sequence simulation method based on particle swarm optimization random forest
CN118332521A
Cited By
Composite structure design method based on support vector machine and NSGA-II algorithm
CN121683440A
GNSS vertical displacement prediction method, storage medium and electronic equipment
CN122131339A