A data recovery method based on multi-shift operators and matrix completion theory
By combining multi-shift operators and matrix fill theory, the data recovery method is optimized, and the row-rank correlation and low-rank characteristics are used to solve the problem of poor data recovery effect in the existing technology, achieving a more efficient data recovery effect.
Patent Information
- Application Number
- CN202210071574.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-01-21
AI Technical Summary
The existing data recovery methods fail to effectively utilize the correlation between row vectors and column vectors in the data matrix, resulting in poor recovery results. Especially when data is lost in low-cost sensor networks, existing methods cannot achieve the best recovery effect.
Multi-shift operator and matrix fill theory are used to characterize the correlation of each row and each column in the data matrix, and optimize the data recovery method model using non-smooth regular terms. The objective function is solved by combining generalized iteration distributed methods, alternating iterations to approach the optimal solution, and using matrix fill technology to restore the entire matrix.
It significantly improves the accuracy and effectiveness of data recovery, especially in the absence of data, which can better restore the low-rank characteristics of the data matrix and improves the data recovery performance.
Smart Images

Figure CN114610531B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of compressive sensing, and particularly to a data recovery method based on multi-shift operators and matrix completion theory. Background Art
[0002] In many practical problems, such as image and video processing, recommendation systems, text analysis, etc., the data to be recovered is usually represented by a matrix, which makes it more convenient to understand, model, process, and analyze the problem. However, these data often face problems such as missing, damaged, and noise pollution. How to obtain accurate data in these situations is the problem to be solved.
[0003] At the same time, in low-cost commodity sensor networks, such as air temperature sensor networks, it is very common to lose data due to sensor failures or communication failures. This requires data filling and recovery. In the data recovery method, the existing methods only use the low-rank property of the data matrix for matrix completion and recovery, ignoring the correlation between row vectors and column vectors in the data matrix, thus unable to achieve the best data recovery effect. Summary of the Invention
[0004] The purpose of the present invention is to provide a data recovery method based on multi-shift operators and matrix completion theory, aiming to solve the technical problem of poor recovery effect caused by the failure to consider the correlation between row vectors and column vectors in the matrix in the existing data recovery methods.
[0005] To achieve the above purpose, the present invention provides a data recovery method based on multi-shift operators and matrix completion theory, including the following steps:
[0006] Establish a data recovery method model based on multi-shift operators and matrix completion theory;
[0007] Optimize the data recovery method model using a non-smooth regularization term to obtain an objective function;
[0008] Design a generalized iterative distributed method to solve the objective function and iteratively approach the optimal solution alternately.
[0009] Among them, the data recovery method model is modeled as a graph G, and the corresponding collected data is a graph signal matrix.
[0010] Among them, in the graph signal matrix, the graph G can be composed of a row graph G r =(V r , E r , W r ) and a column graph G c =(V c , E c , W cis composed of, where the row graph consists of vertices V r = {1,...., M}, edges and an M×M order adjacency matrix W r and is composed in the same way for the column graph.
[0011] Among them, in the process of optimizing the data recovery method model using a non-smooth regularization term to obtain the objective function, the data recovery method model is equivalently rewritten using the non-smooth regularization terms of the quadratic total variation row vector and column vector, and the Kronecker product operation is used to describe the mutual relationship between each row and each column in the matrix.
[0012] Among them, in the process of designing a generalized iterative distributed method to solve the objective function and approaching the optimal solution by alternating iteration, the Newton method is used for updating, and the differentiable part and non-differentiable part of the objective function are solved in each iteration.
[0013] Among them, the data recovery method based on the multi-shift operator and matrix completion theory further includes a verification step:
[0014] Calculate the root mean square error parameter between the solved recovered signal matrix and the original signal matrix. The smaller the obtained root mean square error, the better the data recovery effect.
[0015] The present invention provides a data recovery method based on the multi-shift operator and matrix completion theory. The multi-shift operator is used to characterize the correlation between each row vector and each column vector in the data matrix. By regularizing the total variation of non-smoothness in each row vector and column vector, more accurate data recovery performance can be obtained. Moreover, the collected data has the low-rank property, and the matrix completion technology can be used to recover the entire matrix from the known partial data matrix. The combination of the multi-shift operator and matrix completion theory can characterize the low-rank property of the matrix while characterizing the correlation between each row and each column in the data matrix, thereby greatly improving the data recovery performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 is a schematic flowchart of a data recovery method based on the multi-shift operator and matrix completion theory of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0019] Please refer to Figure 1 , the present invention proposes a data recovery method based on multi-shift operators and matrix completion theory, including the following steps:
[0020] S1: Establish a data recovery method model based on multi-shift operators and matrix completion theory;
[0021] S2: Optimize the data recovery method model using non-smooth regularization terms to obtain an objective function;
[0022] S3: Design a generalized iterative distributed method to solve the objective function and iteratively approach the optimal solution alternately.
[0023] In the process of optimizing the data recovery method model using non-smooth regularization terms to obtain an objective function, the data recovery method model is equivalently rewritten using non-smooth regularization terms of quadratic total variation row vectors and column vectors, and the mutual relationship between each row and each column in the matrix is described through Kronecker product operations.
[0024] In the process of designing a generalized iterative distributed method to solve the objective function and iteratively approaching the optimal solution alternately, the Newton method is used for updating, and the differentiable part and the non-differentiable part of the objective function are solved in each iteration.
[0025] The data recovery method based on multi-shift operators and matrix completion theory further includes a verification step:
[0026] Calculate the root mean square error parameter between the solved recovered signal matrix and the original signal matrix. The smaller the obtained root mean square error, the better the data recovery effect.
[0027] Specifically, the network is modeled as a graph G, and the collected data can be regarded as a graph signal matrix, where each row represents all the data collected at a node i ∈ V. Due to the complex environment in reality, there may be missing data matrices, so the work of complementing and recovering is required. The data recovery of the graph signal matrix aims to recover the missing terms in the prior information of the low rank and minimum total variation of the graph signal matrix X ∈ R M×N of.
[0028] The graph G can be composed of a row graph G r =(V r , E r , W r ) and a column graph Gc =(V c , E c , W c ), where the row graph consists of vertices V r = {1,...., M}, edges and the M×M order adjacency matrix W r . The column graph is defined in the same way. Let X be a set of M-dimensional column vectors, denoted by subscripts as X = [x1,..., x N ; or let X be a set of N-dimensional row vectors, denoted by subscripts as X = [(x 1 ) T ,...,(x M ) T . T ; The column vectors x1,..., x N are vectors defined on the vertices V c .
[0029] The following is a further description from specific steps:
[0030] Step 1: Establish a data recovery method model based on multi-shift operators and matrix completion theory as:
[0031]
[0032] where X is the low-rank signal matrix to be recovered; Y is the partially observable signal matrix; M is the set composed of the subscripts of the observable samples; is as close to zero as possible, so that the true low-rank matrix can be almost completely recovered; is the total variation of the non-smoothness in each column vector, A n is the normalized adjacency matrix associated with the column graph; ||·|| * is the nuclear norm, which refers to the sum of the singular values of the matrix, used to convexly approximate the rank constraint and characterize the low-rank property of the graph signal matrix; α and β are both regularization parameter terms.
[0033] Step 2: According to the model established in Step 1, optimize and solve the model:
[0034] Use ρ r (X) and ρ c (X) to represent the non-smooth regularization terms of the row vectors and column vectors of the quadratic total variation S2(X) respectively. The above problem is equivalently rewritten as:
[0035]
[0036] where ρ r (X) controls the smoothness within columns, that is, the penalty along rows, and ρ r (X) is as follows:
[0037]
[0038] where is the Laplacian operator of the row graph G r i.e., and {(j,j′)∈E r}; represents the Kronecker operator, represents performing the Kronecker product operation with the M - order identity matrix.
[0039]
[0040] where is the Laplacian operator of the column graph G c i.e., and (j,j′)∈E c ; represents performing the Kronecker product operation with the N - order identity matrix. Further, the Kronecker product operations of and can be regarded as two shift operators respectively.
[0041] The objective function of the above problem can be composed of the differentiable sub - function and the non - differentiable part g(X) = β||X|| * . Let x = vec(X) denote the vectorized form of the matrix X, and h(X) is equivalent to:
[0042]
[0043] where Q M ∈R M×MN is the sample matrix, and Q M x = x M . The gradient and Hessian matrix of h(X) can be derived as:
[0044]
[0045]
[0046] Step 3: Design a generalized iterative distributed method to solve the above problem. In each iteration, solve the differentiable part and the non - differentiable part of the objective function, that is, use Newton's method to update x and z. The iteration of the proposed algorithm can be expressed as:
[0047]
[0048] x m+1 = SVT(z m+1 )
[0049] where m ≥ 0, P m is a series of local matrices approximating the inverse of the Hessian matrix. SVT is the Singular Value Thresholding operator. By using the singular value thresholding algorithm to solve for x, the obtained x is then assigned to Equation to solve for z. During the alternating iteration process of x and z, they are continuously assigned values and gradually approach the optimal solution.
[0050] Step 4: Calculate the Root Mean Squared Error (RMSE) parameter between the solved recovered signal matrix X and the original signal matrix to evaluate and compare the quality of the data recovery effect. That is, the smaller the root mean squared error, the better the data recovery effect.
[0051] Furthermore, the present invention proposes a specific simulation example for verification and illustration:
[0052] The input of the embodiment is sea surface temperature network data. The first 100 node data are taken as the original data X. And 10%, 20%, 30%, and 40% of the data are artificially damaged respectively as the observable data Y. And for different regularization parameters α, γ, β, all take 0.01.
[0053] Through the model of this technical solution, the data after matrix completion and recovery is solved. To evaluate the accuracy of the data recovery, the present invention uses the root mean squared error as the evaluation index of the recovery result. The RMSE is as follows:
[0054]
[0055] where, is the vectorization of the recovered data matrix, is the vectorization of the original data matrix, and N x is the vector length.
[0056] Table 1 Root mean squared error results under different data damage percentages in the simulation example experiment
[0057]
[0058] The comparison method is the standard data matrix recovery method, that is s.t.P M (X) = P M(Y). It can be found from the above table that the root mean square error of recovery decreases as the data destruction rate decreases, and the proposed method is always superior to the comparative method. It can be seen that the present method is superior to the above comparative methods. The combination of the multi-shift operator and the matrix completion theory can characterize the low-rank property of the matrix while characterizing the correlation of each row and each column in the data matrix, thereby greatly improving the data recovery performance.
[0059] The above-disclosed is only a preferred embodiment of the present invention. Of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. A data recovery method based on multi-shift operators and matrix completion theory, characterized in that It includes the following steps: Establish a data recovery method model based on multi-shift operators and matrix completion theory; The expression of the data recovery method model is: Where X is the low-rank signal matrix to be recovered; Y is the partially observable signal matrix; M is the set composed of the subscripts of observable samples; is as close to zero as possible, so that the true low-rank matrix can be almost completely recovered; is the total variation of the non-smoothness in each column vector, A n is the normalized adjacency matrix associated with the column graph; ||·|| * is the nuclear norm, which refers to the sum of the singular values of the matrix, used to convexly approximate the rank constraint and characterize the low-rank property of the graph signal matrix; both α and β are regularization parameter; Optimize the data recovery method model using non-smooth regularization terms to obtain an objective function; In the process of optimizing the data recovery method model using non-smooth regularization terms to obtain an objective function, use the non-smooth regularization terms of quadratic total variation row vectors and column vectors to equivalently rewrite the data recovery method model, and describe the mutual relationship between each row and each column in the matrix through Kronecker product operation; Design a generalized iterative distributed method to solve the objective function, and alternately iterate to approach the optimal solution.
2. The data recovery method based on multi-shift operators and matrix completion theory according to claim 1, wherein The data recovery method model is modeled as a graph G, and the corresponding collected data is a graph signal matrix.
3. The data recovery method based on multi-shift operators and matrix completion theory according to claim 2, wherein In the graph signal matrix, the graph G can be composed of a row graph G r =(V r , E r , W r ) and a column graph G c =(V c , E c , W c ), where the row graph consists of vertices V r ={1,....,M}, edges and an M×M order adjacency matrix W r . The column graph is defined in the same way.
4. The data recovery method based on multi-shift operators and matrix completion theory according to claim 1, wherein In the process of designing a generalized iterative distributed method to solve the objective function and alternately iterating to approach the optimal solution, use the Newton method for updating, and solve the differentiable part and the non-differentiable part of the objective function at each iteration.
5. The data recovery method based on multi-shift operators and matrix completion theory according to claim 1, wherein The data recovery method based on multi-shift operators and matrix completion theory further includes a verification step: Calculate the root mean square error parameter between the solved recovery signal matrix and the original signal matrix. The smaller the obtained root mean square error, the better the data recovery effect.
Citation Information
Patent Citations
Ubiquitous power Internet of Things perception data missing restoration method based on matrix filling
CN110705762A
Time-varying graph signal reconstruction method based on multiple shift operators
CN113190790A