Data processing method, device, electronic device and storage medium
Through the mean low-rank autoregressive tensor completion (MLATC) method, missing data completion is performed on landslide monitoring data, solving the problem of missing data in landslide monitoring data, and improving data processing efficiency and prediction accuracy.
Patent Information
- Application Number
- CN202310421003.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-04-19
AI Technical Summary
There are serious data missing problems in landslide monitoring data, which leads to the inability to fully utilize the effectiveness of the monitoring system, affecting the accuracy of landslide displacement prediction.
The time series matrix of landslide monitoring data was completed by using mean low-rank autoregressive tensor completion (MLATC) method. This method obtains the missing data by obtaining the time series matrix, determining the mean of the reference data, constructing a tensor completion model and objective function, and solving the objective function.
In the absence of complete original data, effectively completing the missing parts of landslide monitoring data improves the efficiency and accuracy of data processing and enhances the reliability of landslide displacement prediction.
Smart Images

Figure CN116776089B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of geological disaster monitoring, and in particular to a data processing method, device, electronic equipment and storage medium. Background Art
[0002] The rapid development of informatization has brought about the continuous accumulation of massive data. Massive time series data has been applied to various industries, such as smart cities, environmental monitoring, and engineering technology. Time series data helps researchers from all walks of life to analyze data to obtain the best interpretation results. However, in the process of data collection and analysis, various problems will inevitably be encountered, and data missing is one of the typical problems. Data missing will increase the difficulty of data analysis and reduce the accuracy of data analysis. Therefore, how to achieve effective data completion and prediction for time series with large data volume and a certain scale of missing data is a challenging research task.
[0003] Landslide is one of the common geological disasters in nature. The frequent occurrence of extreme climate has increased the probability of landslides. The rapid development of the economy and society has further aggravated the losses caused by landslide disasters, highlighting the importance of landslide displacement prediction. Landslide monitoring data collection is the basis of landslide prediction, but various data missing seriously inhibit the effectiveness of the monitoring system. In the process of automatic landslide monitoring, data collection and transmission generally use different types of sensors and other electronic equipment. The automatic monitoring equipment is in an open-air environment all year round, and most of them inevitably suffer from wear, aging, power loss and other phenomena, which will lead to missing monitoring data. In addition, the geological environment in which most landslide disasters occur is relatively harsh. Phenomena such as heavy rain, hail, dense fog, and electromagnetic interference are extremely common. The geological disaster monitoring equipment installed and deployed in the field will inevitably be affected by the above harsh environment. During the operation of the monitoring equipment, there will be short-term, intermittent or long-term interruptions, which will cause the monitoring equipment to fail to send monitoring data to the server normally, which will lead to missing or lost monitoring time series in the server. Summary of the invention
[0004] The embodiments of the present invention provide a data processing method, device, electronic device and storage medium, which can complete missing data of landslide monitoring data.
[0005] In a first aspect, an embodiment of the present invention discloses a data processing method, the method comprising:
[0006] Obtain the time series matrix of landslide monitoring data;
[0007] Determine the mean value of each reference data in the landslide monitoring data; the reference data is the monitoring data adjacent to the missing data in the landslide monitoring data;
[0008] Determine a first matrix corresponding to the time series matrix according to the mean; the first matrix is a similarity matrix of the second matrix, and the second matrix is an orthogonal matrix of the time series matrix;
[0009] Determining an objective function based on the time series matrix, the first matrix and a pre-constructed tensor completion model; the objective function is used to indicate a second tensor obtained after low-rank tensor completion of the first tensor, the first tensor being a missing tensor corresponding to the time series matrix;
[0010] Solving the objective function to obtain the second tensor;
[0011] Missing data in the landslide monitoring data is determined according to the second tensor.
[0012] Optionally, determining a first matrix corresponding to the time series matrix according to the mean value includes:
[0013] Determine a first projection function based on the orthogonal projection of the second matrix on the spatial domain of the landslide monitoring data;
[0014] Determine a first parameter according to the mean, where the first parameter is used to indicate a time series length of the mean;
[0015] Determine a second projection function according to the first parameter and the first projection function; the second projection function is used to constrain the projection of the first matrix on the spatial domain of the landslide monitoring data;
[0016] A first matrix corresponding to the time series matrix is determined based on the second projection function.
[0017] Optionally, the objective function is a linear combination of the minimum nuclear norm of the second tensor and the autoregressive norm of the first matrix, and the objective function is expressed as:
[0018]
[0019] in, represents the nuclear norm of the second tensor X, A is the variable to be estimated, H represents the set of lagged time series, H={h1,h2,...,hi,...,hd}, is the weight parameter, represents the autoregressive norm of the first matrix Z, Indicates that the first matrix Z is converted into a third-order tensor; Represents the first matrix Z in the spatial domain of landslide monitoring data The second projection function of Represents the time series matrix Y in the spatial domain Orthogonal projection of .
[0020] Optionally, the second projection function is expressed as:
[0021]
[0022] in, t is the first parameter; m is the row subscript of the matrix Z, indicating the number of the monitoring sensor; n is the column subscript of the matrix Z, indicating the time series number; is the first projection function, expressed as:
[0023]
[0024] in, Represents the values of the second matrix in the interval (m,n).
[0025] Optionally, the autoregressive norm of the first matrix is expressed as:
[0026] ;
[0027] Wherein, t is the first parameter; A is the variable to be estimated, H represents the set of lagged time series, H={h1,h2,...,hi,...,hd}; represents the value of the second matrix in the interval (m, t); is the weighting factor.
[0028] Optionally, the nuclear norm of the second tensor is expressed as:
[0029] ;
[0030] in, , is the weighting factor, is the matrix corresponding to the second tensor X The i-th singular value of , M is the number of rows of the first matrix Z, indicating the number of monitoring sensors; N is the number of columns of the first matrix Z; The matrix is the matrix of the second tensor X expanded along the k-th mode, including:
[0031] .
[0032] Optionally, solving the objective function to obtain the second tensor includes:
[0033] Constructing a Lagrangian function in matrix form for the objective function;
[0034] Solving the Lagrangian function to obtain the second tensor;
[0035] Wherein, the Lagrangian function is:
[0036]
[0037] in, is the weighting factor, is the weight parameter, are the parameters to be learned, is the penalty parameter; It means expanding the tensor X into a matrix along mode-k, and F means finding the matrix norm; the Lagrangian function satisfies: .
[0038] In a second aspect, an embodiment of the present invention discloses a data processing device, the device comprising:
[0039] A data acquisition module is used to obtain the time series matrix of landslide monitoring data;
[0040] A first determination module is used to determine the mean of each reference data in the landslide monitoring data; the reference data is the monitoring data adjacent to the missing data in the landslide monitoring data;
[0041] A second determination module is used to determine a first matrix corresponding to the time series matrix according to the mean value; the first matrix is a similarity matrix of the second matrix, and the second matrix is an orthogonal matrix of the time series matrix;
[0042] A third determination module is used to determine an objective function based on the time series matrix, the first matrix and a pre-constructed tensor completion model; the objective function is used to indicate a second tensor obtained after low-rank tensor completion of the first tensor, the first tensor being a missing tensor corresponding to the time series matrix;
[0043] a fourth determining module, configured to solve the objective function to obtain the second tensor;
[0044] A fifth determination module is used to determine missing data in the landslide monitoring data according to the second tensor.
[0045] In a third aspect, an embodiment of the present invention discloses an electronic device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the aforementioned data processing method when executing the computer program.
[0046] In a fourth aspect, an embodiment of the present invention discloses a machine-readable medium having instructions stored thereon, which, when executed by one or more processors of a device, enables the device to execute the data processing method as described above.
[0047] The embodiments of the present invention include the following advantages:
[0048] The data processing method provided by the embodiment of the present invention converts the data completion problem of the time series matrix of landslide monitoring data into tensor completion. Without the need for complete original data, it can effectively complete the missing data and predict the displacement trend in the time series matrix of landslide monitoring data, thereby improving the processing efficiency and accuracy of missing data of landslide monitoring data. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0050] Figure 1 is a flow chart of steps of an embodiment of a data processing method of the present invention;
[0051] Figure 2 is a structural block diagram of an embodiment of a data processing device of the present invention;
[0052] Figure 3 It is a structural schematic diagram of the electronic device of the present invention. DETAILED DESCRIPTION
[0053] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0054] The terms "first", "second", etc. in the specification and claims of the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable when appropriate, so that the embodiments of the present invention can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, the term "and / or" in the specification and claims is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. In the embodiments of the present invention, the term "multiple" refers to two or more, and other quantifiers are similar.
[0055] In order to complete and predict the missing data of landslides, traditional time series models mostly focus on parameter models based on small-scale problems such as regression and exponential smoothing. For the incompleteness of missing long time series data, the filling methods can be roughly divided into deletion or filling. Data deletion is to remove some objectively abnormal monitoring data as abnormal data, which is mainly used in anomaly detection and feature analysis, while filling is to find the long-term time series change law of monitoring data and supplement the missing monitoring data. Common filling methods include statistical filling and machine learning filling.
[0056] The main methods of machine learning include missing value filling algorithms based on the nearest neighbor method, recurrent neural network and matrix decomposition. In the data filling method, for long time series problems of different dimensions, matrix decomposition can effectively explore the correlation between different time series, learn the overall characteristics of the time series matrix through the matrix decomposition method, and then make a low-rank approximation of the matrix with time characteristics, and then fill in the missing data.
[0057] Large-scale time series data are always accompanied by missing data. Tensor completion methods have been introduced into this field to complement traditional data completion methods based on probability statistics. Data completion schemes based on simple quantitative statistics have a relatively simple and efficient processing effect for small data sets and simple regression models, but lack feasibility for massive data in the era of data explosion. Modern research not only requires a wide variety of data variables and a long time period, but also requires fast processing speed, high versatility and portability. Multivariate variables and long time series can better describe the complex causal relationships between each other. Neural network technology is now often used to deal with the above situation, deleting missing data in order to form a complex intelligent model, but it is not the optimal solution for missing data, because deletion may strengthen or weaken the degree of connection of a certain causal relationship. Based on this, scholars are exploring the application of tensor decomposition technology to data completion in multivariate long time series, trying to improve the parsing speed and data missing problems.
[0058] Recent studies have found that low-rank matrices have certain advantages in analyzing multivariate long-term series data, including sequence tensor completion methods. Sequence tensor completion recovers the potential tensor from the sampling structure of the time series, assigns the position of the missing items as needed, seamlessly integrates the future values of the time series into the framework of missing data, and improves the accuracy of data completion. The low-rank matrix completion method performs Singular Spectrum Analysis and Singular Value Decomposition on the time series in sequence to complete the low-rank completion of the missing data of the time series, although its computational complexity is large. Therefore, by adding a time dimension to convert the positional time series into a high-dimensional tensor, the cost of complex calculations is better solved. This is also in line with the law of human activities, and there are both long-term and short-term activities. Therefore, some studies use tensors (sensors × 1day × 24hours) to represent the above activities. Tensor representation not only retains the dependencies between sensors, but also provides a new feasible solution for capturing local and global time patterns.
[0059] However, these models cannot effectively handle missing values, and most of them require inference as a preprocessing step, which may lead to potential bias. How to incorporate long-term / short-term patterns and sensor correlations in the presence of missing data remains an important research issue. To address the incompleteness of data, matrix / tensor decomposition methods have been applied to model multivariate and high-order (matrix / tensor variables) time series data. However, how to efficiently handle the large-scale, high-dimensional time series mentioned above has not been solved.
[0060] Some recent studies have used low-rank models for multivariate and high-order time series data, but only capture the global consistency / similarity between different time series, which does not conform to the global trend in the time dimension. Singular Spectrum Analysis (SSA) Hankel structured low-rank completion is a powerful method for time series analysis. Singular Spectrum Analysis (SSA) is a powerful method for time series analysis. SSA performs singular value decomposition (SVD) on the Hankel matrix obtained from the original univariate time series, and then uses these principal components to analyze the time series. This method can complete the inference and prediction tasks of missing data by low-rank completion of the Hankel matrix. However, this model is computationally expensive for the Hankel matrix / tensor model, and it may not be feasible to handle large and long multivariate time series.
[0061] In view of the data missing problem in the landslide monitoring process, the embodiment of the present invention proposes a landslide displacement time series prediction model based on mean-based low-rank autoregressive tensor completion (MLATC). It should be noted that the tensor completion model pre-constructed in the present invention is a mean-based low-rank autoregressive tensor completion model.
[0062] The low-rank autoregressive tensor completion model mainly includes completing low-rank matrix decomposition / tensor completion and constructing a time series autoregressive model. The low-rank matrix completion model uses the underlying low-rank structure to restore the incomplete matrix (assuming that the long-term data series of the landslide is an incomplete sequence). Considering that the deformation displacement of the landslide has a large correlation with the previous deformation, an autoregressive model is constructed to represent the deformation law of the landslide displacement over time. By introducing an autoregressive regularizer in the low-rank matrix decomposition to characterize the temporal dynamics of the landslide displacement deformation, the learned autoregressive regularizer is used to predict the time factor matrix, thereby realizing the completion of landslide displacement monitoring data and predictive modeling.
[0063] Reference Figure 1 , shows a flow chart of steps of an embodiment of a data processing method of the present invention, the method may include the following steps:
[0064] Step 101: Obtain a time series matrix of landslide monitoring data.
[0065] Step 102: determine the mean value of each reference data in the landslide monitoring data; the reference data is the monitoring data adjacent to the missing data in the landslide monitoring data.
[0066] Step 103: determine a first matrix corresponding to the time series matrix according to the mean; the first matrix is a similarity matrix of the second matrix, and the second matrix is an orthogonal matrix of the time series matrix.
[0067] Step 104, determine the objective function based on the time series matrix, the first matrix and a pre-constructed tensor completion model; the objective function is used to indicate a second tensor obtained after low-rank tensor completion of the first tensor, and the first tensor is the missing tensor corresponding to the time series matrix.
[0068] Step 105: Solve the objective function to obtain the second tensor.
[0069] Step 106: Determine missing data in the landslide monitoring data according to the second tensor.
[0070] The landslide monitoring data may be data collected by various monitoring sensors deployed at the landslide monitoring site. In the landslide monitoring system, the monitoring data of each on-site monitoring sensor may be collected by an intelligent terminal, and then the on-site monitoring data may be transmitted to a back-end data monitoring center.
[0071] It is understandable that the reasons for the missing data of landslide monitoring are relatively complex, and data may be missing during data collection, transmission, storage, and analysis. The characteristics of time series with missing data are roughly the same, such as the presence of noise, incompleteness, and abnormal mutation data in the time series. If there is missing data in the time series, it will have a huge impact on the subsequent landslide data analysis and monitoring and early warning. Therefore, in the process of data analysis, the data must be preprocessed, that is, data cleaning, and the processing of missing data is one of the key contents in data cleaning.
[0072] The missing data in the time series of landslide monitoring data generally have the following characteristics: (1) Long time span and large data volume. Landslide monitoring requires data collection of different factors that affect landslide displacement and deformation. The continuous extension of monitoring time leads to an increasing amount of data. The deformation of a landslide is not only related to itself, but may also be affected by other factors. There is a very high correlation between sensors. (2) Randomness. Data collection and transmission in the landslide monitoring system are completed automatically by on-site equipment. The components and modules used for collection and transmission are all electronic components. Considering that most landslide disasters are located in remote mountainous areas with relatively harsh environments, they are easily affected by other natural factors such as rainfall, hail, and electromagnetic interference. The above natural phenomena are random. Therefore, the data collected by the sensor will change with the changes in the landslide disaster body. (3) Spatial correlation. The cause mechanism of landslide disasters is complex. Corresponding monitoring equipment will be deployed in different areas of the landslide, and there will be a certain correlation between different sensors. Therefore, the spatial correlation between the sensors in the time series needs to be considered when filling in missing data.
[0073] Due to the different characteristics of time series, the data types in time series are also different. There are many reasons for missing data in time series, which can be roughly summarized as follows: (1) Data cannot be obtained. The landslide monitoring process is affected by the harsh natural environment. Problems with the power supply of the equipment and the temporary failure of the sensor will lead to the failure to complete data collection for a period of time or at a certain moment, resulting in data loss. (2) Data transmission failure. The transmission of landslide monitoring data mainly relies on wireless communication networks. When the wireless signal is interrupted, the data for a period of time or at a certain moment cannot be transmitted to the background monitoring center. (3) Interference from human factors. When technicians perform data preprocessing operations, relevant data may be transmitted incorrectly due to various reasons, which may also lead to data loss.
[0074] According to existing research, different time series have different data missing patterns. Considering the actual situation of landslide monitoring, the data missing types of landslide monitoring time series can be roughly divided into the following two types:
[0075] (1) Random missing data pattern: This is a random situation that may occur at any time during the monitoring process. There is no pattern or sign for this type of data missing, and the time of data missing is random and accidental.
[0076] (2) Non-random missing pattern. During the monitoring process, rainy weather and signal interruption may occur regularly or regularly. These phenomena may occur with the same probability in a specific time period. This type of data missing has a certain regularity.
[0077] The data processing method provided by the embodiment of the present invention can perform data completion and data prediction processing on landslide monitoring data in random missing mode and non-random missing mode.
[0078] Considering that the collected monitoring data includes information such as date, time, displacement, etc., in the embodiment of the present invention, it can be considered that the time series of the acquired landslide monitoring data is high-dimensional. , we need to first convert the multivariate time series matrix into a tensor, that is, Convert to a size of The third-order tensor, that is, the first tensor in the present invention , thus transforming the matrix completion problem into a tensor completion problem. The traditional tensor completion method is to assign the missing data to zero. Considering that the landslide displacement deformation has time accumulation and continuity, the average of the data before and after the missing data is taken instead of assigning it to zero when completing the landslide displacement data.
[0079] Specifically, in an embodiment of the present invention, the mean of each reference data in the landslide monitoring data is first determined, wherein the reference data is the monitoring data adjacent to the missing data in the landslide monitoring data, that is, the data located before or after the missing data and adjacent to the missing data in the time series matrix. Next, the first matrix Z corresponding to the time series matrix Y is determined based on the mean. It should be noted that, in an embodiment of the present invention, the first matrix Z is a similarity matrix of the orthogonal matrix Z0 of the time series matrix Y (that is, the "second matrix" in the present invention). Then, the objective function is determined based on the time series matrix Y, the first matrix Z and the pre-constructed tensor completion model. Among them, the objective function is used to indicate the first tensor The second tensor X is obtained after low-rank tensor completion, wherein the first tensor is the missing tensor corresponding to the time series matrix. The second tensor can be obtained by solving the objective function, and the missing data in the landslide monitoring data can be determined according to the second tensor.
[0080] It should be noted that in the embodiment of the present invention, the pre-built tensor completion model is a mean low-rank matrix autoregression tensor completion model. Of course, other models may also be used, such as the Low Rank Matrix Completion (LRMC) model, the Vector Auto-Regressive (VAR) model, and the like.
[0081] As an example, the low-rank matrix completion model uses the underlying low-rank structure to recover the incomplete matrix, defining the tensor The rank of ,in The number of in-situ sensors representing the landslide, Represents the number of days of collection, Represents the frequency of sampling per day. The objective function of the low-rank tensor completion model can be expressed as:
[0082] (1)
[0083] in, is the tensor to be found; is the missing tensor; for The set of observable landslide data locations in . Then, the tensor nuclear norm is replaced by the minimum rank and the tensor nuclear norm is defined as:
[0084] (2)
[0085] in, is the weighting factor; is the tensor X along the The matrix of the expanded pattern. Complete the data tensor The three mode expansion matrices are as follows (3):
[0086] (3)
[0087] In the formula, and They are the modular expansion matrices for interleaved sampling of three types of expansion modes.
[0088] The different order features between the data are integrated with each other through three mode expansion matrices, which effectively ensures the high-precision completion of the data. The singular value decomposition of each mode expansion matrix is:
[0089] (4)
[0090] in U , V , W is a left singular matrix; M , L , K is a right singular matrix; is a diagonal matrix, defined as:
[0091] (5)
[0092] In the formula is the total number of singular values of the qth modular expansion matrix, is the matrix X singular values. Based on the definition of tensor nuclear norm, the objective function is:
[0093] (6)
[0094] Among them, the matrix nuclear norm is defined as:
[0095] (7)
[0096] in for The nuclear norm of For the matrix rank; To transform the matrix The singular values of Singular values.
[0097] Because of the existence of interdependent matrix nuclear norm terms, the objective function of low-rank tensor completion is difficult to solve using ordinary methods. The Alternating Direction Method Of Multipliers (ADMM) is a widely used constraint optimization method in machine learning. It is an extension of the augmented Lagrangian method. The ADMM algorithm provides a framework for solving optimization problems with linear equality constraints, which facilitates the decomposition of the original optimization problem into several relatively easy-to-solve sub-optimization problems for iterative solution. With the further development of research, the low-rank tensor completion method has been further developed. By using QR decomposition to replace singular value decomposition to achieve the purpose of reducing computational complexity, the decomposed low-rank tensor completion reduces the time complexity of the low-rank tensor completion algorithm; the nonlinear set CG algorithm based on Riemann manifold reduces the decomposition of large-scale singular values in low-rank tensor completion; the tensor completion algorithm based on Douglas-Rachford separation technology further considers the existence of completion source data noise, greatly improving the robustness of the model. In addition, low-rank tensor completion also has a series of characteristics such as fast convergence speed and high computational accuracy.
[0098] In an optional embodiment of the present invention, a mean-based low-rank matrix autoregressive tensor completion model is used as the tensor completion model. Optionally, determining the first matrix corresponding to the time series matrix according to the mean includes:
[0099] Step S11, determining a first projection function based on the orthogonal projection of the second matrix on the spatial domain of the landslide monitoring data;
[0100] Step S12: determining a first parameter according to the mean, where the first parameter is used to indicate a time series length of the mean;
[0101] Step S13, determining a second projection function according to the first parameter and the first projection function; the second projection function is used to constrain the projection of the first matrix on the spatial domain of the landslide monitoring data;
[0102] Step S14: determining a first matrix corresponding to the time series matrix based on the second projection function.
[0103] The traditional tensor completion method is to assign missing data to zero. Considering that landslide displacement deformation has time accumulation and continuity, the embodiment of the present invention takes the average of the missing data before and after the landslide displacement data is completed instead of assigning it to zero. The specific difference is that the first matrix Z in the domain Further requirements are made on the projection to obtain a useful non-orthogonal projection.
[0104] Specifically, in the embodiment of the present invention, the orthogonal matrix Z0 based on the time series matrix, that is, the second matrix in the spatial domain of the landslide monitoring data The orthogonal projection on determines the first projection function, which can be expressed as:
[0105] (8)
[0106] It means that the matrix Z (the second matrix Z0 in the present invention) is in the domain The orthogonal projection of , m is the row subscript of the matrix Z, indicating the number of the monitoring sensor; n is the column subscript of the matrix Z, indicating the time series number.
[0107] Then, a first parameter is determined according to the mean of the reference data before and after the missing data, and a second projection function is determined according to the first parameter and the first projection function. The first parameter is used to indicate the time series length of the mean; and the second projection function is used to constrain the projection of the first matrix on the spatial domain of the landslide monitoring data.
[0108] Finally, the first matrix corresponding to the time series matrix Y can be determined according to the second projection function.
[0109] Optionally, the second projection function is expressed as:
[0110] (9)
[0111] in, t is the first parameter, and here t=10 can be taken.
[0112] is the first projection function, expressed as:
[0113] (10)
[0114] in, Represents the values of the second matrix in the interval (m,n).
[0115] Optionally, the objective function is a linear combination of the minimum nuclear norm of the second tensor and the autoregressive norm of the first matrix, and the objective function is expressed as:
[0116] (11)
[0117] in, represents the nuclear norm of the second tensor X, A is the variable to be estimated, H represents the set of lagged time series, H={h1,h2,...,hi,...,hd}, is the weight parameter, represents the autoregressive norm of the first matrix Z. The autoregressive norm of the first matrix Z is expressed as:
[0118] (12)
[0119] Wherein, t is the first parameter; A is the variable to be estimated, H represents the set of lagged time series, H={h1,h2,...,hi,...,hd}; represents the value of the second matrix in the interval (m, t); is the weighting factor.
[0120] Indicates that the first matrix Z is converted into a third-order tensor; Represents the first matrix Z in the spatial domain of landslide monitoring data The second projection function of Represents the time series matrix Y in the spatial domain Orthogonal projection of .
[0121] Optionally, the nuclear norm of the second tensor is expressed as:
[0122] (13)
[0123] in, , is the weighting factor, is the matrix corresponding to the second tensor X The i-th singular value of , M is the number of rows of the first matrix Z, indicating the number of monitoring sensors; N is the number of columns of the first matrix Z; The matrix is the matrix of the second tensor X expanded along the k-th mode, including:
[0124] (3).
[0125] Optionally, solving the objective function to obtain the second tensor includes:
[0126] Step S21, constructing a Lagrangian function in matrix form for the objective function;
[0127] Step S22, solving the Lagrangian function to obtain the second tensor;
[0128] Wherein, the Lagrangian function is:
[0129] (14)
[0130] in, is the weighting factor, is the weight parameter, are the parameters to be learned, is the penalty parameter; It means to expand the tensor X into a matrix along mode-k, and F means to find the matrix norm. The Lagrangian function satisfies: .
[0131] Formula (14) is transformed into is a variable, are the parameters to be learned.
[0132] Solving for Tensors hour, Should be established, then ,therefore for:
[0133] (15)
[0134] in, Represents the number of algorithm iterations. The VAR model is required to solve the matrix Z.
[0135] Next, the processing effect of the data processing method provided by the embodiment of the present invention will be described in combination with specific experimental data.
[0136] The displacement monitoring data of the Shuizhuyuan landslide in the Three Gorges Reservoir area was selected as the time series experimental data set for this study. The time series was collected from 7 GNSS monitoring sensors of the Shuizhuyuan landslide in the Three Gorges Reservoir area. The time interval was from July 15, 2017 to December 1, 2021, a total of 1600 days. The data monitoring cycle was to obtain 1 monitoring data on site every day, so the tensor size of the constructed data set was The data obtained by each sensor is a separate time series data. At this time, the original time series data matrix is .
[0137] In order to objectively describe the effectiveness of the prediction model, combined with the characteristics and main types of missing data in landslide monitoring, two different data missing processing methods, random and non-random, were used to process the original time series data set of the landslide.
[0138] Random missing (RM): Random missing refers to the scattered and random data loss phenomenon of landslide monitoring equipment during operation. The entire time series is divided according to the random data missing ratio, and 5%, 10%, 20%, and 40% are set respectively, representing data loss phenomena of different scales and levels in monitoring equipment.
[0139] Non-random missing (NM). Non-random missing means that the landslide monitoring equipment has regular equipment failures or regular network interruptions during operation. The data missing is also simulated at a ratio of 5%, 10%, 20%, and 40%, and compared with random missing data.
[0140] Different missing data sets are set for the time series, and the original landslide spatiotemporal displacement data set is artificially processed for missing data. Through two different data missing processing methods, random and non-random, the missing data completion and prediction effects of different models can be objectively evaluated. For landslide disasters with relatively complex causal mechanisms, there must be certain trend characteristics and correlations between different sensors on the landslide body. Therefore, for the missing data set after non-random missing processing, whether good data recovery and prediction effects can be achieved is an important indicator to measure this model. Therefore, for the missing data set after non-random missing processing, good data recovery and prediction effects are important indicators to measure this model.
[0141] Considering that the Shuizhuyuan landslide is a typical Three Gorges Reservoir landslide, it can also be seen from the actual monitoring data. The Shuizhuyuan landslide is currently in a state of slow deformation at a uniform speed. Therefore, in terms of dividing the training set and the test set, the first 1300 days are selected as the training set of the data, and the last 300 days are selected as the test set of the data, in order to test the accuracy and precision of the algorithm model in prediction, and use MAPE and RMSE to evaluate the prediction performance. The Shuizhuyuan landslide is currently in a state of slow deformation at a uniform speed. The first 1300 days are selected as the training set of the time series, and the last 300 days are selected as the test set. MAPE and RMSE are used to test the accuracy and precision of the algorithm model in prediction. In addition, the MLATC model is compared and analyzed with common tensor completion methods such as High-Accuracy Low-Rank Tensor Completion (HALRTC) and Temporal Regularized Matrix Factorization (TRMF).
[0142] The prediction results show that the MLATC model still achieves a very good completion effect. It is not only consistent with the original displacement data, but also effectively restores the missing data in some days with completely missing data. This is very helpful for analyzing the deformation law and deformation trend of the landslide, and also provides a data reference for understanding the correlation between different deformation areas of the landslide.
[0143]
[0144] Refer to Table 1, which shows the data recovery performance indicators under non-random missing conditions. From the results in Table 1, it can be seen that in the four different cases of NM5%, NM10%, NM20%, and NM40%, the MAPE value and RMSE value of the MLATC model are the best, indicating that the MLATC model has a significant improvement in data recovery performance compared with the HALRTC and TRMF models.
[0145] From the results in Table 1, we can see that the MLATC model has better completion effect than the HALRTC and TRMF models, which shows that the tensor autoregressive kernel norm can effectively replace the rank function, so that the low-rank tensor completion model can obtain a more accurate completion effect. In addition, the MLATC model introduces the autoregressive norm on the basis of the HALRTC model, which can make full use of the structural correlation and local trend feature information between high-dimensional multivariate time series data, and further clarify the correlation between multivariate time series data. This also proves that in the low-rank tensor completion structure, adding the autoregressive norm can further improve the completion effect and accuracy of the model. Since the tensor structure maintains the structural information of spatiotemporal data in the time dimension well, and makes full use of the correlation between the daily displacements of different sensors in the time series, the tensor completion result of the MLATC model is better than the completion effect of TRMF in the matrix form, which is also within the expected results. The experimental results under different missing ratios also verify this inference, and also confirm that the global time series data set can be used to achieve effective completion and rolling prediction of displacement monitoring data.
[0146] The reason is that the TRMF model only performs analysis and calculation in the matrix structure, and does not use the time domain smoothness of multivariate time series and the potential correlation information between series in the time and space dimensions. MLATC and HALRTC both use tensor structures and complete time series prediction through the method of quantity completion. In particular, the MLATC model introduces the autoregressive norm, which not only retains the structural information of the original time series through the tensor structure, but also makes full use of the correlation and trend feature information between the time series.
[0147] Similarly, taking the experimental results of four sensors numbered SZY-03, SZY-06, SZY-07, and SZY-08 as an example, after the random data missing processing with a missing ratio of 40% was performed on the data set of the Shuizhuyuan landslide, the prediction model not only had a very good completion effect on the missing data, but also showed a very good data fitting effect in terms of time series trend prediction.
[0148]
[0149] In the non-random case, the time series of the Shuizhuyuan landslide was processed with a missing ratio of 40% and then a 300-day time series prediction was performed. When the data missing ratio was 40%, the prediction results of non-completely random missing were basically consistent with the original displacement data. The MLATC model well realized the deformation trend characteristic fitting of the displacement and effectively predicted the displacement data.
[0150] Similarly, after random data missing processing with a missing ratio of 40% was performed on the Shuizhuyuan landslide data set, the prediction results of the MLATC model were basically close to the original displacement data. Although there was a certain range of fluctuations in the predicted data, the displacement was still effectively fitted overall.
[0151] The prediction performance of the MLATC model in the non-random missing and random missing cases is shown in Tables 3 and 4.
[0152]
[0153] By analyzing the prediction effects of each model, it can be concluded that in the embodiment of the present invention, after converting the original time series into a tensor structure, the prediction accuracy of the missing data in the time series matrix is effectively improved. In addition, under the framework of the tensor completion structure, the embodiment of the present invention utilizes the spatiotemporal monitoring data between displacement sensors in different areas of the landslide, and fully considers the deformation law characteristics of short-term correlation and long-term deformation consistency in the intrinsic deformation process of the landslide.
[0154] In addition, it can be understood that the methods used for data prediction and data completion of the time series of landslide monitoring data are the same. The data prediction of the time series is to treat the time series data to be predicted as the missing values of the data set, so in principle, for data prediction, it is also to complete the data completion task. Therefore, data completion is to perform basic iterative calculations on the original data on the whole time series, while time series prediction treats the monitoring data to be predicted as missing values.
[0155] In summary, the data processing method provided by the embodiment of the present invention converts the data completion problem of the time series matrix of landslide monitoring data into tensor completion. Without the need for complete original data, it can achieve effective completion of missing data and displacement trend prediction in the time series matrix of landslide monitoring data, thereby improving the processing efficiency and accuracy of missing data of landslide monitoring data.
[0156] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0157] Reference Figure 2 , shows a structural block diagram of a data processing device embodiment of the present invention, such as Figure 2 As shown, the device may specifically include:
[0158] A data acquisition module 201 is used to acquire a time series matrix of landslide monitoring data;
[0159] The first determination module 202 is used to determine the mean value of each reference data in the landslide monitoring data; the reference data is the monitoring data adjacent to the missing data in the landslide monitoring data;
[0160] A second determination module 203 is used to determine a first matrix corresponding to the time series matrix according to the mean value; the first matrix is a similarity matrix of the second matrix, and the second matrix is an orthogonal matrix of the time series matrix;
[0161] A third determination module 204 is used to determine an objective function based on the time series matrix, the first matrix and a pre-constructed tensor completion model; the objective function is used to indicate a second tensor obtained after low-rank tensor completion of the first tensor, the first tensor being a missing tensor corresponding to the time series matrix;
[0162] A fourth determination module 205 is used to solve the objective function to obtain the second tensor;
[0163] The fifth determining module 206 is configured to determine missing data in the landslide monitoring data according to the second tensor.
[0164] Optionally, the second determining module includes:
[0165] A first determination submodule is used to determine a first projection function based on the orthogonal projection of the second matrix on the spatial domain of the landslide monitoring data;
[0166] A second determining submodule, configured to determine a first parameter according to the mean, wherein the first parameter is used to indicate a time series length of the mean;
[0167] A third determination submodule is used to determine a second projection function according to the first parameter and the first projection function; the second projection function is used to constrain the projection of the first matrix on the spatial domain of the landslide monitoring data;
[0168] The fourth determination submodule is used to determine the first matrix corresponding to the time series matrix based on the second projection function.
[0169] Optionally, the objective function is a linear combination of the minimum nuclear norm of the second tensor and the autoregressive norm of the first matrix, and the objective function is expressed as:
[0170]
[0171] in, represents the nuclear norm of the second tensor X, A is the variable to be estimated, H represents the set of lagged time series, H={h1,h2,...,hi,...,hd}, is the weight parameter, represents the autoregressive norm of the first matrix Z, Indicates that the first matrix Z is converted into a third-order tensor; Represents the first matrix Z in the spatial domain of landslide monitoring data The second projection function of Represents the time series matrix Y in the spatial domain Orthogonal projection of .
[0172] Optionally, the second projection function is expressed as:
[0173]
[0174] in, t is the first parameter; m is the row subscript of the matrix Z, indicating the number of the monitoring sensor; n is the column subscript of the matrix Z, indicating the time series number; is the first projection function, expressed as:
[0175]
[0176] in, Represents the values of the second matrix in the interval (m,n).
[0177] Optionally, the Z autoregressive norm of the first matrix is expressed as:
[0178] ;
[0179] Wherein, t is the first parameter; A is the variable to be estimated, H represents the set of lagged time series, H={h1,h2,...,hi,...,hd}; represents the value of the second matrix in the interval (m, t); is the weighting factor.
[0180] Optionally, the nuclear norm of the second tensor X is expressed as:
[0181] ;
[0182] in, , is the weighting factor, is the matrix corresponding to the second tensor X The i-th singular value of , M is the number of rows of the first matrix Z, indicating the number of monitoring sensors; N is the number of columns of the first matrix Z; The matrix is the matrix of the second tensor X expanded along the k-th mode, including:
[0183] .
[0184] Optionally, the fourth determining module includes:
[0185] A construction submodule, used for constructing a Lagrangian function in matrix form for the objective function;
[0186] A computing submodule, used for solving the Lagrangian function to obtain the second tensor;
[0187] Wherein, the Lagrangian function is:
[0188]
[0189] in, is the weighting factor, is the weight parameter, are the parameters to be learned, is the penalty parameter; It means expanding the tensor X into a matrix along mode-k, and F means finding the matrix norm; the Lagrangian function satisfies: .
[0190] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0191] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0192] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0193] The embodiment of the present invention further provides an electronic device, referring to Figure 3 , including: a processor 401, a memory 402, and a computer program 4021 stored in the memory 402 and executable on the processor, wherein the processor 401 implements the data processing method of the aforementioned embodiment when executing the program.
[0194] The embodiment of the present invention also provides a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a device (server or terminal), the device can execute the above Figure 1The description of the data processing method in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated. For technical details not disclosed in the computer program product or computer program embodiment involved in this application, please refer to the description of the method embodiment of this application.
[0195] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed by the present invention. The description and examples are to be regarded as exemplary only, and the true scope and spirit of the present invention is indicated by the following claims.
[0196] It should be understood that the present invention is not limited to the exact construction that has been described above and shown in the drawings and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
[0197] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
[0198] The data processing method, device, electronic device and storage medium provided by the present invention are introduced in detail above. The principle and implementation mode of the present invention are explained in detail using specific examples. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation mode and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A data processing method, characterized in that: The method comprises: Obtain the time series matrix of landslide monitoring data; Determine the mean value of each reference data in the landslide monitoring data; the reference data is the monitoring data adjacent to the missing data in the landslide monitoring data; Determine a first matrix corresponding to the time series matrix according to the mean; the first matrix is a similarity matrix of the second matrix, and the second matrix is an orthogonal matrix of the time series matrix; Determining an objective function based on the time series matrix, the first matrix and a pre-constructed tensor completion model; the objective function is used to indicate a second tensor obtained after low-rank tensor completion of the first tensor, the first tensor being a missing tensor corresponding to the time series matrix; Solving the objective function to obtain the second tensor; determining missing data in the landslide monitoring data according to the second tensor; The reference data is monitoring data adjacent to the missing data in the landslide monitoring data, including: the reference data is data located before or after the missing data and adjacent to the missing data in the time series matrix; The step of determining a first matrix corresponding to the time series matrix according to the mean value includes: Determine a first projection function based on the orthogonal projection of the second matrix on the spatial domain of the landslide monitoring data; Determine a first parameter according to the mean, where the first parameter is used to indicate a time series length of the mean; Determine a second projection function according to the first parameter and the first projection function; the second projection function is used to constrain the projection of the first matrix on the spatial domain of the landslide monitoring data; A first matrix corresponding to the time series matrix is determined based on the second projection function.
2. The method according to claim 1, characterized in that The objective function is a linear combination of the minimum nuclear norm of the second tensor and the autoregressive norm of the first matrix, and the objective function is expressed as: ; in, Represents the second tensor The nuclear norm of , A is the variable to be estimated, H represents the set of lagged time series, H={h1,h2,...,hi,...,hd}, is the weight parameter, Represents the first matrix The autoregressive norm of Indicates that the first matrix Z is converted into a third-order tensor; Represents the first matrix Z in the spatial domain of landslide monitoring data The second projection function of Represents the time series matrix Y in the spatial domain Orthogonal projection of .
3. The method according to claim 2, characterized in that The second projection function is expressed as: ; Wherein, t is the first parameter; m is the row subscript of the matrix Z, indicating the number of the monitoring sensor; n is the column subscript of the matrix Z, indicating the time series number; is the first projection function, expressed as: ; in, Represents the values of the second matrix in the interval (m,n).
4. The method according to claim 2, characterized in that: First Matrix The autoregressive norm of is expressed as: ; Wherein, t is the first parameter; A is the variable to be estimated, H represents the set of lagged time series, H={h1,h2,...,hi,...,hd}; represents the value of the second matrix in the interval (m, t); is the weighting factor.
5. The method according to claim 2, characterized in that: The second tensor The nuclear norm of is expressed as: ; in, , is the weighting factor, is the second tensor The corresponding matrix The i-th singular value of , M is the number of rows of the first matrix Z, indicating the number of monitoring sensors; N is the number of columns of the first matrix Z; The matrix is the second tensor The matrix expanded along the kth mode includes: 。 6. The method according to claim 2, characterized in that The step of solving the objective function to obtain the second tensor includes: Constructing a Lagrangian function in matrix form for the objective function; Solving the Lagrangian function to obtain the second tensor; Wherein, the Lagrangian function is: ; in, is the weighting factor, is the weight parameter, are the parameters to be learned, is the penalty parameter; Refer to the tensor Expand into a matrix along mode-k, F represents the matrix norm; the Lagrangian function satisfies: .
7. A data processing device, characterized in that: The device comprises: A data acquisition module is used to obtain the time series matrix of landslide monitoring data; A first determination module is used to determine the mean of each reference data in the landslide monitoring data; the reference data is the monitoring data adjacent to the missing data in the landslide monitoring data; A second determination module is used to determine a first matrix corresponding to the time series matrix according to the mean value; the first matrix is a similarity matrix of the second matrix, and the second matrix is an orthogonal matrix of the time series matrix; A third determination module is used to determine an objective function based on the time series matrix, the first matrix and a pre-constructed tensor completion model; the objective function is used to indicate a second tensor obtained after low-rank tensor completion of the first tensor, the first tensor being a missing tensor corresponding to the time series matrix; a fourth determining module, configured to solve the objective function to obtain the second tensor; a fifth determining module, configured to determine missing data in the landslide monitoring data according to the second tensor; The reference data is monitoring data adjacent to the missing data in the landslide monitoring data, including: the reference data is data located before or after the missing data and adjacent to the missing data in the time series matrix; The second determining module is specifically used to: Determine a first projection function based on the orthogonal projection of the second matrix on the spatial domain of the landslide monitoring data; Determine a first parameter according to the mean, where the first parameter is used to indicate a time series length of the mean; Determine a second projection function according to the first parameter and the first projection function; the second projection function is used to constrain the projection of the first matrix on the spatial domain of the landslide monitoring data; A first matrix corresponding to the time series matrix is determined based on the second projection function.
8. An electronic device, characterized in that: The electronic device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, and the processor implements the data processing method according to any one of claims 1 to 6 when executing the computer program.
9. A machine-readable storage medium having instructions stored thereon, which, when the instructions are executed by one or more processors of a device, cause the device to execute the data processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Visual data tensor completion method based on smooth constraint and matrix decomposition
CN113222834A
Acoustic vector sensor DOA estimation method based on matrix decomposition under data loss
CN113687297A