GNSS complex mountainous area landslide monitoring multipath signal unsupervised learning identification method based on characteristic decomposition
Through feature decomposition and K-means++ clustering algorithm, multi-path errors in GNSS signals are identified, satellite data affected by multi-path are eliminated, and the problem of reduced accuracy of GNSS signals in complex mountainous landslide monitoring is solved, and high-precision GNSS positioning is achieved.
Patent Information
- Application Number
- CN202510438253.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In complex mountainous landslide monitoring, GNSS signals are often affected by multipath errors, resulting in reduced monitoring accuracy and reliability. The existing unsupervised classification algorithms are poorly robust and difficult to meet the data processing requirements of long-term series and multiple interference sources.
The data processing method of feature decomposition is adopted, combined with the K-means++ clustering algorithm, and the signal-to-noise ratio, height angle, azimuth angle and double-difference residual are extracted as eigenvalues, feature decomposition and clustering are performed, and satellite data affected by multipath is eliminated using the average of the contour coefficient, recognition accuracy and absolute value of multipath error.
It significantly improves the GNSS positioning accuracy and ambiguity fixation rate, improves the multi-path signal recognition efficiency, solves the accuracy reduction problem caused by the multi-path effect of GNSS signals in complex environments, and provides more reliable technical support for landslide monitoring.
Smart Images

Figure CN120387041A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of satellite navigation and signal recognition, and particularly relates to an unsupervised learning recognition method for multipath signals in GNSS complex mountain landslide monitoring based on eigen - decomposition. Background Art
[0002] Landslides are one of the main types of geological disasters globally. Different landslide types, geological features, surrounding environments, and inducing factors determine the selection of monitoring methods and instruments. The Global Navigation Satellite System (GNSS), with advantages such as wide coverage, all - weather operation, continuous monitoring, and real - time feedback, has become one of the most widely used technologies in the three - dimensional surface deformation monitoring of landslides. In particular, the real - time kinematic relative positioning technology based on short baselines, due to its fast, efficient, and real - time high - precision positioning ability, is widely used in three - dimensional surface deformation monitoring. However, since the observation environment of GNSS receivers installed on mountains with potential landslide risks is usually harsh, including the presence of surrounding mountains and various vegetation, GNSS satellite signals often undergo severe reflection, refraction, and diffraction effects, resulting in large multipath errors, which seriously affect the accuracy and reliability of monitoring results.
[0003] In recent years, unsupervised classification algorithms have become one of the main methods for dealing with GNSS multipath error recognition due to their high automation and adaptability to the environment and data. In unsupervised classification algorithms, the selection of eigenvalues is the key to determining the clustering effect. Since the dimensions and distributions of each eigenvalue are different, directly classifying will affect the classification accuracy. Therefore, it is necessary to unify the dimensions and distributions of each eigenvalue. The conventional processing method is to normalize each eigenvalue, but simple normalization ignores the correlation information between samples. In addition, this method is vulnerable to sample gross errors and has poor robustness. Such linear transformations only achieve the consistency of the numerical range of each dimension, but ignore the potential coupling relationship between features. Simple normalization is highly sensitive to outliers. When gross error samples cause the extreme values of a certain feature dimension to shift, the normalization parameters will drift significantly, thus distorting the overall feature distribution and causing a large clustering center offset error. Such defects make it difficult for traditional methods to meet the data - processing requirements of long - time series and multi - interference sources in landslide monitoring scenarios. Therefore, there is an urgent need to study a technology for identifying multipath signals in complex mountain landslide monitoring. Summary of the Invention
[0004] Aiming at the problem of poor robustness in traditional multi-path signal recognition and processing methods, the purpose of the present invention is to propose an unsupervised learning recognition method for multi-path signals in GNSS complex mountain landslide monitoring. By selecting eigenvalues, introducing a data processing method of eigenvalue decomposition, and combining with the unsupervised learning K-means++ clustering algorithm, the multi-path errors in GNSS signals are accurately identified, satellite data affected by multi-path is effectively removed, and the ambiguity fixation rate and GNSS positioning accuracy are significantly improved.
[0005] Based on the above purpose, the technical solution adopted by the present invention is as follows:
[0006] In the first aspect, the present invention provides an unsupervised learning recognition method for multi-path signals in GNSS complex mountain landslide monitoring based on eigenvalue decomposition, including the following steps:
[0007] S1. Extract the signal-to-noise ratio, elevation angle, azimuth angle, and double-difference residual of the research station as eigenvalues;
[0008] S2. Perform eigenvalue decomposition on the eigenvalues extracted in step S1 to obtain a scaled data set;
[0009] S3. Use the K-means++ algorithm to cluster the scaled data set in step S2;
[0010] S4. Use the silhouette coefficient, recognition accuracy rate, and average value of the absolute value of the multi-path error of each cluster obtained after clustering in step S3 to remove satellite data greatly affected by multi-path signals, and then use the removed satellite data for GNSS positioning calculation.
[0011] The present invention selects the four GNSS data of signal-to-noise ratio, elevation angle, azimuth angle, and double-difference residual as eigenvalues. At the same time, considering the information loss and poor robustness that may occur in the normalized eigenvalues, the present invention introduces a method of eigenvalue decomposition to perform eigenvalue decomposition, and more referenceable and usable feature information can be obtained.
[0012] The present invention performs clustering analysis of eigenvalues with the K-means++ unsupervised clustering algorithm. This method introduces a probability weight mechanism to select the initial clustering centers, and its initial clustering centers are more representative and dispersed, effectively avoiding the problem of too long algorithm convergence time. At the same time, the silhouette coefficient, recognition accuracy rate, and average value of the absolute value of the multi-path error of each cluster are used to evaluate the clustering results, providing objective, reliable, and reasonable original data for GNSS positioning by screening data.
[0013] The present invention introduces a data processing method of eigenvalue decomposition, combines with the unsupervised learning K-means++ clustering algorithm to accurately identify the multipath errors in GNSS signals. The method of the present invention can effectively eliminate the satellite data affected by multipath errors, improve the efficiency of multipath signal recognition, and significantly improve the GNSS positioning accuracy. Thus, it solves the problem of reduced accuracy caused by the multipath effect of GNSS signals in complex geological environments, and provides more reliable technical support for landslide monitoring.
[0014] Preferably, the eigenvalue decomposition includes the following steps:
[0015] Input the signal-to-noise ratio, elevation angle, azimuth angle, and double-difference residuals into the dataset matrix X, calculate the correlation coefficient matrix R of matrix X, solve the fundamental solution system of matrix R and perform standard orthogonalization on it to obtain the orthogonal matrix P, multiply the dataset matrix X by the orthogonal matrix P to obtain the decorrelated matrix Y, and calculate the standard deviation σ for each dimension in the dataset matrix Y respectively i Form a diagonal matrix Λ, multiply the dataset matrix Y by the diagonal matrix Λ to obtain the scaled dataset Z.
[0016] Preferably, step 3 uses the K-means++ algorithm to cluster the scaled dataset Z, including the following steps:
[0017] Taking the CH value of the internal evaluation index of the K-means++ algorithm as an index to obtain the optimal number of clustering centers of the dataset Z; according to the optimal number of clustering centers, divide each sample point into the cluster closest to the centroid, calculate through the Euclidean distance, and perform iterative update of the clustering centers based on the algorithm until the objective function meets the requirements.
[0018] Preferably, the CH value is calculated based on the variance within the clusters and the variance between the clusters. The calculation formula of the CH value is:
[0019]
[0020] In the formula, CH(k) is the CH index; B k is the covariance between the cluster data; W k is the covariance of the data within the cluster; tr(·) is the trace of the matrix; n is the total number of samples; k is the number of clustering centers;
[0021] The calculation formula of the said Euclidean distance is:
[0022]
[0023] Perform iterative calculation based on the following algorithm and update the clustering centers, that is:
[0024]
[0025] When the objective function of the algorithm meets the requirements, that is, the change in the sum of squared errors between the samples within each cluster and the centroid is negligible, the iteration can be stopped. The expression of the objective function is as follows:
[0026]
[0027] Preferably, the silhouette coefficient SC after clustering is as follows:
[0028]
[0029] In the formula, S(i) represents the silhouette coefficient of a single sample, and its calculation formula is:
[0030]
[0031] Among them, the smaller a(i) is, the higher the compactness of sample i with its own cluster; the larger b(i) is, the greater the separation of sample i from other clusters; the closer the value of the silhouette coefficient s(i) is to 1, the better the clustering effect of sample i; the closer the value is to -1, the more suitable the sample is to be assigned to other clusters; the value close to 0 indicates that the sample is on the boundary of two clusters.
[0032] Preferably, the recognition accuracy formula is:
[0033]
[0034] Among them, S TP is the result of the model correctly classifying good satellites; S TN is the result of the model correctly classifying bad satellites; S FP is the wrong result of the model classifying good satellites; S FN is the wrong result of the model classifying bad satellites.
[0035] Preferably, the calculation formula for the average value of the absolute values of multipath errors of each cluster is as follows:
[0036]
[0037] In the formula, is the average value of the absolute values of multipath errors of the k-th cluster, and M(i) is the multipath error of cluster k.
[0038] In the second aspect, the present invention provides an application of the above unsupervised learning recognition method for multipath signals in GNSS complex mountain landslide monitoring.
[0039] The correct fixation of ambiguity is one of the key technologies for high-precision carrier-phase positioning, and the fixation rate is an important factor restricting high-precision positioning applications in complex environments. The higher the fixation rate, the more accurate the landslide displacement monitoring results. The present invention uses the fixation rate as an evaluation index for the application effect of the above-mentioned unsupervised learning recognition method for multipath signals in GNSS complex mountainous landslide monitoring. Through experiments, it is proved that the unsupervised learning recognition method for multipath signals provided by the present invention can increase the fixation rate to more than 99%.
[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0041] The present invention accurately identifies multipath errors in GNSS signals by introducing a data processing method of eigen-decomposition and combining with the unsupervised learning K-means++ clustering algorithm. This method effectively eliminates satellite data affected by multipath errors, significantly improves the GNSS positioning accuracy, and the fixation rate is not less than 99%. It solves the problem of reduced accuracy caused by the multipath effect of GNSS signals in complex geological environments, and provides more reliable technical support for landslide monitoring. Description of the Drawings
[0042] Figure 1 It is a flowchart for identifying multipath signals in complex mountainous landslide monitoring in Example 1;
[0043] Figure 2 It is a specific location map of the survey area;
[0044] Figure 3 It is a topographic and environmental map of the survey station in Example 1;
[0045] Figure 4 It is a sky map of the survey station drawn based on the observed satellites;
[0046] Figure 5 It is a distribution map of each cluster after K-means++ multipath recognition after normalization and eigen-decomposition;
[0047] Figure 6 It is a PDOP value map of the monitoring station;
[0048] Figure 7 It is a relationship map between the signal-to-noise ratio and the elevation angle of the monitoring station;
[0049] Figure 8 It is a CH value map for different numbers of clusters;
[0050] Figure 9 It is a histogram of the ambiguity fixation rate of three schemes. Detailed Embodiments
[0051] To better illustrate the purpose, technical solution, and advantages of the present invention, the present invention will be further described below in conjunction with specific embodiments. Those skilled in the art should understand that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0052] Embodiment 1
[0053] This embodiment provides an unsupervised learning recognition method for multipath signals in GNSS complex mountain landslide monitoring based on feature decomposition, as Figure 1 shown, including the following steps:
[0054] S1. Extract eigenvalue such as signal-to-noise ratio, elevation angle, azimuth angle, double-difference residual, etc. of the research station.
[0055] Among them, the signal-to-noise ratio data can be directly obtained from the observation file of the station. To obtain the azimuth angle and elevation angle, it is necessary to first perform iterative calculation of the satellite position to obtain the precise position of the satellite, then perform the correction of the earth's rotation, then convert the satellite coordinates from ECEF coordinates to ENU coordinates at the station center, and finally calculate the azimuth angle and elevation angle of the satellite.
[0056] To extract the double-difference residual, it is necessary to first perform single-difference between stations to eliminate the satellite clock error and satellite orbit error, and then perform inter-satellite single-difference to eliminate the receiver clock error, receiver hardware delay, and receiver antenna phase center error. The double-difference formula is obtained as follows:
[0057]
[0058] In the formula, assume there are stations p, q, and satellites j, k. In the formula, Δ▽ is the double-difference operator, is the pseudorange observation value, is the phase observation value, is the geometric distance after double-difference, λ is the wavelength, is the integer ambiguity, is the ionospheric error, is the tropospheric error, is the residual error (including multipath effect, etc.).
[0059] Due to the short baseline, the ionospheric error and tropospheric error are negligible, and the double-difference residual can be directly extracted from the above formula.
[0060] S2. Perform feature decomposition on the eigenvalue extracted in step S1.
[0061] Due to the differences in dimension and distribution among the eigenvalues, directly performing classification may affect the classification accuracy. Therefore, it is necessary to perform unified dimension and distribution processing on the eigenvalues. The common processing method is to normalize the eigenvalues, but this method has problems such as information loss and poor robustness. The present invention introduces the method of eigen - decomposition, decomposes the eigenvalues, and inputs the signal - to - noise ratio S, elevation angle α, double - difference residual υ, and azimuth angle AZ of the eigenvalues into the dataset matrix X, calculates the correlation coefficient matrix R of the matrix, solves the fundamental solution system of the matrix R and performs standard orthogonalization on it to obtain the orthogonal matrix P, multiplies the dataset matrix X by the orthogonal matrix P to obtain the decorrelated matrix Y, and calculates the standard deviation σ for each dimension in the dataset matrix Y i Form a diagonal matrix Λ, multiply the dataset matrix Y by the diagonal matrix Λ, that is, Z = Y·Λ to obtain the scaled dataset Z.
[0062] The specific process is as follows:
[0063] Input the signal - to - noise ratio S, elevation angle α, double - difference residual υ, and azimuth angle AZ of the eigenvalues into the dataset matrix X = [x1, x2, …, x m T =(x ij ) n×m , where n is the number of satellites and m is the number of eigenvalues. In this article, m = 4, then the expression of matrix X is:
[0064]
[0065] Then the correlation coefficient matrix R of matrix X can be obtained:
[0066]
[0067] In the formula, r ij is the correlation coefficient between the i - th dimension index and the j - th dimension index of the dataset, r ij = r ji ; Solve the characteristic equation |R - λI m | = 0 for the correlation coefficient matrix R, and the eigenvalues can be obtained: λ i (i = 1, 2, …, m), and find the fundamental solution system of matrix R: α = [α1, α2, …, α m , and arrange the fundamental solution system in the order of λ i size. Then perform standard orthogonalization on matrix α to obtain the orthogonal matrix P = [η1, η2, …, η m .
[0068] Multiply the dataset matrix X by the orthogonal matrix P, that is, Y = X·P to obtain the dataset matrix Y after eliminating the correlation; Calculate the standard deviation σ for each dimension in the dataset matrix Y i (i = 1, 2, …, m):
[0069]
[0070] Form a diagonal matrix Λ:
[0071]
[0072] Multiply the dataset matrix Y point - by - point by the diagonal matrix A to obtain the scaled dataset Z, i.e., Z = Y·Λ. Thus, the eigenvalue decomposition is completed.
[0073] S3. Use the K - means++ algorithm to cluster the dataset obtained from the feature decomposition in step S2.
[0074] In the face of a complex environment, to ensure the optimality of the signal classification quantity (i.e., the best number of clusters), this paper uses the CH value index, an internal evaluation index of the clustering algorithm, to obtain the best number of clusters k. After determining the number of cluster centers, each sample point is assigned to the cluster closest to the centroid, calculated by the Euclidean distance. Subsequently, the cluster centers are continuously iteratively updated until the objective function meets the requirements, so as to improve the accuracy and reliability of the classification results.
[0075] The CH value is calculated based on the variance within the clusters and the variance between the clusters. It takes into account both the dispersion degree of the data points within each cluster and the distance between different clusters. The calculation formula of the CH value is:
[0076]
[0077] In the formula, CH(k) is the CH index; B k is the covariance between the cluster - class data; W k is the covariance of the data within the cluster - class; tr(·) is the trace of the matrix; n is the total number of samples; k is the number of cluster centers. When the total number of samples and the number of cluster centers are fixed, if the data points within the category are more closely aggregated, then the clustering effect will be better. In the case of sufficient samples, as long as we reasonably select the feature vectors in the GNSS data, using the CH value for evaluation will be able to obtain accurate results.
[0078] k - means++ selects the initial cluster centers by introducing a probability - weight mechanism, making the selected cluster centers more representative and dispersed. Its basic principle in the initialization process of the cluster centers is to make the mutual distance between the initial cluster centers as far as possible. Then each sample point is assigned to the cluster closest to the centroid, calculated by the Euclidean distance. The formula is:
[0079]
[0080] Due to the large errors that may exist in the centroid parameters, resulting in poor clustering results, iterative calculations are required and the cluster centers are updated, that is:
[0081]
[0082] When the objective function of the algorithm meets the requirements, that is, the change in the sum of squared errors between the samples within each cluster and the centroid is negligible, the iteration can stop. The expression of the objective function is:
[0083]
[0084] S4. Use the results after clustering in step S3 to eliminate the satellite data with large multi-path signal influence.
[0085] After performing K-means++ clustering, the clustering effect is evaluated using the silhouette coefficient, recognition accuracy, and the average value of the absolute multi-path errors of each cluster. It can clearly distinguish the influence of multi-path errors on each cluster. After eliminating the cluster with the worst effect, positioning is performed using the optimized observation data, and its fixation rate and positioning accuracy have been significantly improved.
[0086] The specific evaluation indicators are as follows:
[0087] ① Silhouette coefficient:
[0088] The calculation of the silhouette coefficient is based on the following two factors:
[0089] a(i): The average distance from sample i to other samples in its affiliated cluster, used to measure the tightness of the sample with its own cluster.
[0090] b(i): The average distance from sample i to samples in the nearest neighboring different clusters, used to measure the separation degree of the sample from other clusters.
[0091] The calculation formula of the silhouette coefficient is:
[0092]
[0093] Among them, the smaller a(i) is, the higher the tightness of sample i with its own cluster; the larger b(i) is, the greater the separation degree of sample i from other clusters. The closer the value of the silhouette coefficient s(i) is to 1, the better the clustering effect of sample i; the closer the value is to -1, the more suitable the sample is to be assigned to other clusters; the value close to 0 indicates that the sample is on the boundary of two clusters.
[0094] By calculating the silhouette coefficients of all samples and taking the average value, the silhouette coefficient of the entire clustering result can be obtained; the overall silhouette coefficient SC of the clustering is:
[0095]
[0096] ② Recognition accuracy rate:
[0097] In the present invention, satellites with a median error more than three times the pseudorange multipath error are used as the correct labeling results, serving as an external indicator for evaluating the recognition accuracy rate of multipath. The formula for the recognition accuracy rate is:
[0098]
[0099] Among them, S TP is the result of the model correctly classifying good satellites. S TN is the result of the model correctly classifying poor satellites. S FP is the incorrect result of the model classifying good satellites. S FN is the incorrect result of the model classifying poor satellites.
[0100] ③ Average value of the absolute value of multipath error for each cluster:
[0101] To more intuitively reflect the recognition quality of multipath error, in the present invention, the average value of the absolute value of multipath error for each cluster is used as a quality evaluation indicator for reflecting the recognition result of multipath error. The formula is:
[0102]
[0103] In the formula, is the average value of the absolute value of multipath error for the k-th cluster, and M(i) is the multipath error of cluster k.
[0104] Example 2
[0105] In this example, the Tengqing Coal Mine in Huale Town, Shuicheng District, Liupanshui City is used as the test area. The method described in Example 1 is used to monitor the landslide. A reference station TQ01 is set up in an open area with sparse vegetation and wide visibility for easy observation; the monitoring stations are TQ02, TQ03, and TQ04, all located on the upper edge of the landslide body with dense vegetation and serious satellite signal occlusion. The dynamic monitoring distance is 1.8 km, which is a short baseline. All monitoring points use dual-system, multi-frequency GNSS receivers to collect landslide monitoring data in a complex landslide environment. The experiment is implemented using the self-developed Monitor software. Taking the TQ04 point as an example for data processing and analysis, the positioning processing strategy is shown in Table 1. Figure 3 This is the topographic and station environment map of this measurement area, Figure 4 and this is the sky map drawn based on the observed satellites. It can be seen from the topographic map and station environment map of the measurement area that there are a large number of tree occlusions around the monitoring stations. It can be seen from the topographic line sky map that the east side of the station is almost completely occluded by trees, and there is also partial occlusion in the northwest. Therefore, the data availability and continuity in the northwest and east are relatively poor.
[0106] Table 1 Positioning Processing Strategy
[0107] Processing solution Processing strategy Usage system GPS, BDS Frequency point GPS L1 & L2, BDS B1 & B2 Ambiguity fixing strategy Lambda Ratio value 3.0
[0108] Taking the signal-to-noise ratio, elevation angle, double-difference residual, and azimuth angle of the eigenvalue obtained from the monitoring data of the 259th day of the year 2021 as samples for machine learning, the total sample size is 102,287 (the sum of the number of visible satellites in each epoch). Use the CH value to obtain the optimal k value, and the CH value index is as Figure 5 shown. It can be seen from Figure 5 that when k = 5, the CH value is the largest, which is the optimal clustering number. Then, the results obtained after performing eigen-decomposition processing on the sample data and the results after using conventional normalization processing are respectively used for multipath identification using K-means++. Figure 6 Figure 1 is the distribution diagram of each cluster after K-means++ multipath identification after normalization processing and eigen-decomposition processing. It can be seen from Figure 6 the distribution diagram of each cluster after K-means++ multipath identification after normalization processing that the azimuth angle plays its anchoring effect well and successfully identifies the satellites that are severely blocked in the continuous interval on the east side of the monitoring station. However, introducing the azimuth angle as a feature variable, although it can effectively identify the severely blocked satellites, it will also cause a problem: for two satellites that are equally affected by severe obstacles, if their azimuth angles are too different. In this case, when calculating the Euclidean distance between these two satellites, the distance difference will also become larger accordingly. Eventually, they are divided into different clusters during the multipath identification process, resulting in the failure of multipath identification. And it can be seen from Figure 6 the distribution diagram of each cluster after K-means++ multipath identification after eigen-decomposition processing that this problem can be overcome to a certain extent after the sample is processed by eigen-decomposition. After eigen-decomposition, not only the satellites that are severely blocked in the continuous interval on the east side of the monitoring station are successfully identified, but also the satellites that are severely affected by obstacles in the northwest direction are identified. The identification result of eigen-decomposition is more in line with the real environment.
[0109] Next, use the two multipath identification results, through the formula The total silhouette coefficients and recognition accuracies of the two multi-path recognition results are calculated respectively. After feature decomposition processing, the total silhouette coefficient of the multi-path recognition result is 0.539 and the recognition accuracy is 93.6%. After normalization, the total silhouette coefficient of the multi-path recognition result is 0.317 and the recognition accuracy is 86.1%. Then, for each cluster, the average value of the absolute multi-path errors of the two data processing schemes after multi-path recognition is calculated through a formula. The results are shown in Table 2. It can be seen from Table 2 that the average value of the absolute multi-path errors of each cluster in the multi-path recognition result after feature decomposition is better than that after normalization. Especially for Cluster 2, which is the satellite seriously affected by multi-path errors, it can be seen that the average absolute value after feature decomposition is greater than that after normalization, indicating that the multi-path recognition effect after feature decomposition is better than the conventional processing method.
[0110] Average value of absolute multi-path errors of each cluster in multi-path recognition of the two data processing schemes in Table 2
[0111]
[0112] Next, to verify the feasibility of the K-mean++ GNSS signal classification and recognition method based on feature decomposition, TQ04 measured data is used for recognition verification. The data acquisition frequency is 15 s, the observation duration is 24 h, and the observation period is from 00:00:00 to 23:59:45 on the 259th day of the year 2021. Three recognition schemes are designed: Scheme 1, without classification processing; Scheme 2, after the sample data is conventionally normalized and then classified by K-means++ to eliminate the satellites seriously affected by multi-path errors; Scheme 3, after the sample data is processed by feature decomposition and then classified by K-means++ to eliminate the satellites seriously affected by multi-path errors.
[0113] The fixed rate, floating-point solution residual standard deviation, and fixed solution residual standard deviation are used to evaluate the accuracy of the three schemes for landslide monitoring. Among them, the fixed rate in GNSS (Global Navigation Satellite System) positioning is an index to measure the success rate of ambiguity resolution and is used to evaluate the reliability of the positioning algorithm. Its core meaning is: within a certain period of time, the proportion of the number of times the ambiguity is successfully fixed as an integer solution to the total number of solutions. The floating-point solution residual is the difference between the GNSS observation value and the model prediction value based on the floating-point solution when the integer ambiguity is not fixed. The floating-point solution residual reflects the fitting degree of the model to the observation value. The smaller the residual, the closer the model is to the true signal. The fixed solution residual is the difference between the observation value and the model prediction value based on the fixed solution (integer solution) after the ambiguity is fixed as an integer. The smaller the fixed solution residual, the higher the consistency between the integer solution and the actual observation value, and the more reliable the positioning accuracy.
[0114] The results are as Figure 7, it can be seen from Table 3 that the fixation rate and floating-point solution accuracy of Scheme 3 are significantly better than those of Scheme 1 and Scheme 2. Compared with Scheme 1 in the E, N, and U directions, the floating-point solution accuracy is improved by 55.3%, 30%, and 15% respectively, and the fixation rate is increased by 8.8%. Compared with Scheme 2, the floating-point solution accuracy is improved by 15.2%, 11.5%, and 3.8% respectively, and the fixation rate is increased by 5.4%. This is mainly because Scheme 3 can better identify satellite signals severely affected by environmental interference, significantly reducing the influence of multipath diffraction errors at the edges of obstacles around the monitoring station. And from Figure 6 It can be seen from the distribution maps of each cluster class that after the samples are processed by feature decomposition, multipath signals in the environment can be better identified. From Figure 4 it can be seen that the obstacles are distributed on the east side and the northwest side of the measuring station, exactly corresponding to the recognition area of Cluster 5 in Scheme 3. It can be seen that Scheme 3 can better reflect the true terrain of the measuring station. Therefore, compared with Scheme 2, the accuracy and fixation rate in the three directions of E, N, and U are both improved. The correct fixation of ambiguity is one of the key technologies for high-precision carrier-phase positioning, and the fixation rate is an important factor restricting high-precision positioning applications in complex environments. The higher the fixation rate, the more accurate the landslide displacement monitoring results. The above results show that Scheme 3 has higher accuracy for landslide monitoring.
[0115] Table 3 Statistical table of fixation rate and standard deviation of three schemes
[0116]
[0117] To further verify the recognition accuracy of the proposed algorithm, the observation data for 4 consecutive days from the 259th to the 262nd day of the year of 2021 at the TQ04 monitoring point are processed, and a correlation analysis is carried out on the signal clusters identified as being severely affected by multipath errors. The results are as Figure 8 shown. It can be seen that the recognition results for 4 consecutive days are basically similar, and on each day, satellites that are continuously and severely blocked on the east side and the northwest side of the monitoring point can be successfully identified, which is consistent with Cluster 2 identified above, indicating that the satellite signals in these areas are severely affected by multipath errors. From the number of signals identified as being severely affected by multipath errors, the number reaches 12,088 satellites, accounting for 11.8% of the total.
[0118] Finally, the fixation rate of the ambiguity of the calculation results of the observation data for 7 consecutive days from the 260th to the 266th day of the year of 2021 at the TQ04 monitoring station is further statistically analyzed. The results are as Figure 9 shown. From Figure 9It can be seen that: the fixed solution ratio of Solution 1 for 7 consecutive days is the lowest, with an average fixation rate of 89.3%; the average fixation rate of Solution 2 for 7 consecutive days is 93.4%, which is an improvement compared to Solution 1; Solution 3 has the highest ambiguity fixation rate, and the average fixation rate for 7 consecutive days is 99.4%, higher than that of Solution 1 and Solution 2.
[0119] The above are only several specific embodiments of the present invention disclosed. Those skilled in the art can make various changes and modifications to the embodiments of the present invention without departing from the spirit and scope of the present invention. However, the embodiments of the present invention are not limited thereto, and any changes that can be conceived by those skilled in the art should fall within the protection scope of the present invention.
Claims
1. An unsupervised learning identification method for multipath signals in GNSS complex mountain landslide monitoring based on eigenvalue decomposition, characterized in that, It includes the following steps: S1. Extract the signal-to-noise ratio, elevation angle, azimuth angle, and double-difference residual of the research station as eigenvalues; S2. Perform eigenvalue decomposition on the eigenvalues extracted in step S1 to obtain a scaled data set; S3. Use the K-means++ algorithm to cluster the data set scaled in step S2; S4. Use the silhouette coefficient, recognition accuracy, and average value of the absolute value of the multipath error of each cluster obtained after clustering in step S3 to eliminate satellite data greatly affected by multipath signals, and then use the eliminated satellite data for GNSS positioning calculation.
2. The method according to claim 1, wherein The eigenvalue decomposition includes the following steps: Input the signal-to-noise ratio, elevation angle, azimuth angle, and double-difference residuals into the dataset matrix X. Calculate the correlation coefficient matrix R of matrix X, solve the fundamental solution system of matrix R and perform standard orthogonalization on it to obtain the orthogonal matrix P. Multiply the dataset matrix X by the orthogonal matrix P to obtain the decorrelated matrix Y, and calculate the standard deviation σ for each dimension in the dataset matrix Y respectively. i Form the diagonal matrix Λ, and multiply the dataset matrix Y by the diagonal matrix Λ to obtain the scaled dataset Z.
3. The method according to claim 1, wherein Step 3 uses the K-means++ algorithm to cluster the scaled data set Z, including the following steps: Taking the CH value, an internal evaluation index of the K-means++ algorithm, as an index to obtain the optimal number of cluster centers of the data set Z; according to the optimal number of cluster centers, divide each sample point into the cluster closest to the centroid, calculate through the Euclidean distance, and perform iterative update of the cluster center based on the algorithm until the objective function meets the requirements.
4. The method according to claim 3, characterized in that The CH value is calculated based on the variance within the cluster and the variance between clusters. The calculation formula of the CH value is: where CH(k) is the CH index; B k is the covariance between cluster data; W k is the covariance of the data within the cluster; tr(·) is the trace of the matrix; n is the total number of samples; k is the number of cluster centers; The calculation formula of the Euclidean distance is: Perform iterative calculation and update the cluster center based on the following algorithm, that is: When the objective function of the algorithm meets the requirements, that is, the change in the sum of squared errors between the samples in each cluster and the centroid can be ignored, the iteration can stop. The expression of the objective function is:
5. The method according to claim 1, wherein The silhouette coefficient SC after clustering is: In the formula, S(i) represents the silhouette coefficient of a single sample, and its calculation formula is: Among them, the smaller a(i) is, the higher the tightness of sample i with its own cluster; the larger b(i) is, the greater the separation of sample i from other clusters; the closer the value of the silhouette coefficient s(i) is to 1, the better the clustering effect of sample i; the closer the value is to -1, the more suitable the sample is to be assigned to other clusters; the value close to 0 indicates that the sample is on the boundary of two clusters.
6. The method according to claim 1, wherein The formula for the recognition accuracy is: Among them, S TP is the result of the model correctly classifying good satellites; S TN is the result of the model correctly classifying bad satellites; S FP is the incorrect result of the model classifying good satellites; S FN is the incorrect result of the model classifying bad satellites.
7. The method according to claim 1, wherein The calculation formula for the average value of the absolute value of the multipath error of each cluster is as follows: Wherein, is the average value of the absolute values of the multipath errors of the k-th cluster, and M(i) is the multipath error of the k-th cluster.
8. Application of the multipath signal unsupervised learning recognition method according to claim 1 in GNSS complex mountain landslide monitoring.
Citation Information
Patent Citations
A satellite NLOS signal detection method based on unsupervised learning
CN109101902A
Area division method based on improved orthogonal decomposition
CN109116391A
GNSS (Global Navigation Satellite System) data processing method, system and equipment based on clustering analysis
CN117214929A