Time series data denoising method and device, computer equipment and storage medium
By introducing approximate L0 norms and weighted kernel norms into the RPCA model, an improved RPCA model is constructed, which solves the problems of slow solution speed and large recovery error in high-dimensional timing data denoising, and realizes efficient and accurate denoising of timing data and improving data quality.
Patent Information
- Application Number
- CN202510196667.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
When processing high-dimensional timing data, the existing robust principal component analysis (RPCA) method has slow solution speed and large recovery errors, resulting in an increase in the deviation between the denoised data and the original data, and the data denoising efficiency and quality are reduced.
An approximately L0 norm based on fractional function structure is introduced in the RPCA model, and the weighted kernel norm and control terms are added to build an improved RPCA model. Through this model, the timing data is divided into low-rank matrix and sparse matrix, and efficient and accurate denoising processing is achieved through iterative solution of collaborative optimization problems.
The calculation and solution efficiency of the RPCA model in time-series data denoising and the accuracy of low-rank matrix recovery are improved, efficient and accurate denoising of time-series data is achieved, and data quality is improved.
Smart Images

Figure CN120123646A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a method, device, computer device, and storage medium for denoising time series data. Background Art
[0002] In the field of time series data processing and analysis, especially for typical time series data types such as spatio-temporal trajectory data and industrial sensing data, the time series data is often affected by various noises during the acquisition and transmission processes. These noises may mask the key information in the time series data, thereby affecting the accuracy of subsequent time series data processing and analysis results. Therefore, in the preprocessing stage of time series data, performing effective denoising processing on the noisy data to recover high-precision time series data is of crucial significance for improving the quality of time series data and the effect of subsequent applications.
[0003] The core idea of Robust Principal Component Analysis (RPCA) is that the data matrix contains low-rank structure information and sparse noise information. RPCA transforms the original data matrix into the sum of a low-rank matrix and a sparse matrix. The low-rank matrix part reflects the principal component information and low-frequency information of the original data matrix, while the sparse matrix contains the sparse information of the noise because the noise itself is sparse. The proposed RPCA method provides a reliable solution to the data denoising problem in the big data era. The RPCA method does not rely on prior assumptions about the noise type and distribution, has strong adaptability, and has been widely used in time series data denoising.
[0004] However, in actual applications of RPCA for time series data denoising, since the basic strategy of the RPCA solution algorithm is to solve the norm optimization problem in the RPCA model, a main problem of such algorithms is that as the dimension of the input matrix increases, the speed of algorithm solution slows down, and the recovery error increases, resulting in an increase in the deviation between the denoised data and the original data, a decrease in the efficiency of data denoising, and a decrease in the quality of the denoised data. Therefore, how to perform efficient and accurate denoising processing on noisy time series data and restore high-precision time series data has become an urgent problem to be solved in this field. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, and storage medium for denoising time series data, which introduce an approximate L0 norm based on a fractional function structure into the Robust Principal Component Analysis model, and add a weighted nuclear norm and a control term, improving the computational solution efficiency and the accuracy of low-rank matrix recovery when the model performs data denoising, thereby enabling efficient and accurate denoising of time series data.
[0006] A method for denoising time series data, the method includes:
[0007] Preprocess the time series data;
[0008] Based on the original RPCA model, introduce an approximate L0 norm based on the fractional function structure, and add a weighted nuclear norm and a control term to construct an improved RPCA model;
[0009] Input the preprocessed time series data into the improved RPCA model. Based on this model, divide the valid data in the input time series data into a low-rank matrix, divide the noise data in the input time series data into a sparse matrix, and construct a collaborative optimization problem that balances the low-rank characteristics of the low-rank matrix and the sparsity constraint of the sparse matrix. Iteratively solve this collaborative optimization problem by constructing an augmented Lagrangian function until the iteration convergence condition is satisfied and then stop the iteration, and output the denoised time series data and noise data.
[0010] In one embodiment, the preprocessing includes: normalizing the time series data.
[0011] In one embodiment, based on the original RPCA model, introduce an approximate L0 norm based on the fractional function structure, and add a weighted nuclear norm and a control term to construct an improved RPCA model, including:
[0012] Adopt a smoothing function ρ d (|t|) to replace the discontinuous norm in the original RPCA model, and construct an RPCA model introducing an approximate L0 norm, expressed as
[0013]
[0014] where, A is the low-rank matrix, E is the sparse matrix, D represents the original data matrix, s.t. is the constraint condition, ‖A‖ * represents the nuclear norm of the low-rank matrix, λ represents the regularization parameter, represents the approximate L0 norm of the variable x, x i is the i-th element in the variable x, N is the number of elements, is used for the approximate L0 norm, t represents the independent variable of the smoothing function, d represents the selected value, and 0.1 ≤ d ≤ 1;
[0015] Based on the RPCA model introducing an approximate L0 norm, perform nuclear norm weighting on the low-rank matrix A to obtain an RPCA model with a weighted nuclear norm, expressed as
[0016]
[0017] where, w A ={w A,j} represents the weights of the singular values in the low-rank matrix A, σj denotes the j-th singular value in the low-rank matrix A, w A,j denotes σ j 's weight; wherein, the weight of the weighted nuclear norm is used to control the sparsity of the sparse matrix;
[0018] Based on the RPCA model with weighted nuclear norm, a control term is further added to construct an improved RPCA model, expressed as
[0019]
[0020] wherein, λ 1 and λ 2 respectively represent the first regularization parameter and the second regularization parameter, represents the Frobenius norm control term; wherein, the control term is used to control the robustness of the low-rank matrix solution.
[0021] In one embodiment, a co-optimization problem that balances the low-rank property of the low-rank matrix and the sparsity constraint of the sparse matrix is constructed, and the co-optimization problem is iteratively solved by constructing an augmented Lagrangian function until the iterative convergence condition is satisfied and the iteration stops, and the denoised time-series data and noise data are output, including:
[0022] Based on the improved RPCA model, a co-optimization problem that balances the low-rank property of the low-rank matrix A and the sparsity constraint of the sparse matrix E is constructed;
[0023] An augmented Lagrangian function is constructed for the co-optimization problem, expressed as
[0024]
[0025] wherein, Y represents the Lagrange multiplier, μ represents the parameter, and B represents the intermediate variable;
[0026] The expression of the augmented Lagrangian function is iteratively solved by using the alternating iteration method until the iterative convergence condition is satisfied and the iteration stops, and the denoised time-series data and noise data are output.
[0027] In one embodiment, the expression of the augmented Lagrangian function is iteratively solved by using the alternating iteration method until the iterative convergence condition is satisfied and the iteration stops, and the denoised time-series data and noise data are output, including:
[0028] Let Y = (Y 1 , Y 2 ), μ = (μ 1 , μ 2 ), and the new expression of the augmented Lagrangian function is
[0029]
[0030] Among them, Y 1 and Y 2 respectively represent the first Lagrange multiplier and the second Lagrange multiplier, μ 1 and μ 2 respectively represent the first parameter and the second parameter, D represents the original data matrix, w A,j represents the weight of the j-th singular value σ j in the low-rank matrix A, n is the number of singular values, λ 1 and λ 2 respectively represent the first regularization parameter and the second regularization parameter, P d (E) represents the approximate L0 norm of the sparse matrix E;
[0031] The new expression of the augmented Lagrangian function is solved by an alternating iteration method to obtain the intermediate variable B, which is expressed as
[0032]
[0033] Then, first fix the sparse matrix E and update the low-rank matrix A, and then fix the low-rank matrix A and introduce the DC programming idea to update the sparse matrix E;
[0034] When the low-rank matrix A and the sparse matrix E converge to the corresponding and respectively, for Y 1 and Y 2 , μ 1 and μ 2 are iteratively updated until the preset iteration output accuracy is reached or the preset iteration number threshold is reached, and the denoised time series data and noise data are output; where k is the number of iterations, and respectively represent the low-rank matrix and the sparse matrix obtained by the convergence of the (k + 1)-th iteration.
[0035] In one embodiment, first fix the sparse matrix E and update the low-rank matrix A, and then fix the low-rank matrix A and introduce the DC programming idea to update the sparse matrix E, including:
[0036] Fix the sparse matrix E, and the expression for updating the low-rank matrix A is
[0037]
[0038] Among them, A k+1 represents the low-rank matrix of the (k + 1)-th iteration, U represents the left singular value vector, represents the soft threshold operator when updating the low-rank matrix A, Σ represents the singular value vector, V represents the right singular value vector, represents the singular value operator, wA = {w A,j} represents the weight of the singular value in the low-rank matrix A, and the superscript T represents the matrix transpose;
[0039] Fix the low-rank matrix A and introduce the DC programming idea to update the expression of the sparse matrix E as
[0040]
[0041] where E k+1 represents the sparse matrix at the (k + 1)-th iteration, represents the soft-thresholding operator when updating the sparse matrix E, λ represents the regularization parameter, and V k+1 represents V at the (k + 1)-th iteration, and V k+1 is defined as
[0042]
[0043] where sign(E k ) represents the soft-thresholding function, d represents the selected value, and 0.1 ≤ d ≤ 1, and E k represents the sparse matrix at the k-th iteration.
[0044] In one embodiment, the update methods of Y 1 and Y 2 are as follows:
[0045]
[0046] where Y 1 k+1 and respectively represent Y 1 and Y 2 at the (k + 1)-th iteration, Y 1 k and respectively represent Y 1 and Y 2 at the k-th iteration, and respectively represent μ 1 and μ 2 at the k-th iteration, and A k+1 、E k+1 and B k+1 respectively represent the low-rank matrix A, the sparse matrix E, and the intermediate variable B at the (k + 1)-th iteration;
[0047] The update methods of μ 1 and μ 2 are as follows:
[0048]
[0049] where ρ is a constant and ρ > 1, μ k represents the parameter μ in the k-th iteration, ε is a relatively small positive number, and respectively represent the sparse matrices obtained by the convergence of the (k + 1)-th and k-th iterations, ||D|| F represents the Frobenius norm of the original data matrix D.
[0050] A time-series data denoising device, the device includes:
[0051] A preprocessing module for preprocessing time-series data;
[0052] A model construction module for introducing an approximate L0 norm based on a fractional function structure and adding a weighted nuclear norm and a control term on the basis of the original RPCA model to construct an improved RPCA model;
[0053] A data denoising module for inputting the preprocessed time-series data into the improved RPCA model, dividing the valid data in the input time-series data into a low-rank matrix based on this model, dividing the noise data in the input time-series data into a sparse matrix, and constructing a collaborative optimization problem that balances the low-rank characteristics of the low-rank matrix and the sparsity constraint of the sparse matrix, and iteratively solving this collaborative optimization problem by constructing an augmented Lagrangian function until the iteration convergence condition is satisfied and the iteration stops, and outputting the denoised time-series data and noise data.
[0054] A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0055] Preprocess the time-series data;
[0056] On the basis of the original RPCA model, introduce an approximate L0 norm based on a fractional function structure and add a weighted nuclear norm and a control term to construct an improved RPCA model;
[0057] Input the preprocessed time-series data into the improved RPCA model, divide the valid data in the input time-series data into a low-rank matrix based on this model, divide the noise data in the input time-series data into a sparse matrix, and construct a collaborative optimization problem that balances the low-rank characteristics of the low-rank matrix and the sparsity constraint of the sparse matrix, and iteratively solve this collaborative optimization problem by constructing an augmented Lagrangian function until the iteration convergence condition is satisfied and the iteration stops, and output the denoised time-series data and noise data.
[0058] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0059] Preprocess the time series data;
[0060] Based on the original RPCA model, introduce an approximate L0 norm based on the fractional function structure, and add a weighted nuclear norm and a control term to construct an improved RPCA model;
[0061] Input the preprocessed time series data into the improved RPCA model. Based on this model, divide the valid data in the input time series data into a low-rank matrix, divide the noise data in the input time series data into a sparse matrix, and construct a collaborative optimization problem that balances the low-rank characteristics of the low-rank matrix and the sparsity constraint of the sparse matrix. Iteratively solve this collaborative optimization problem by constructing an augmented Lagrangian function until the iterative convergence condition is met and stop the iteration, and output the denoised time series data and noise data.
[0062] The above method, device, computer device and storage medium for denoising time series data construct an improved RPCA model for denoising time series data. By introducing an approximate L0 norm based on the fractional function structure, this model can accurately identify and separate the sparse noise in the input time series data using the L0 norm, achieving effective capture of noise. By adding a weighted nuclear norm for controlling the sparsity of the sparse matrix and a control term for controlling the robustness of the solution of the low-rank matrix in the model, it can effectively balance the robustness and sparsity in the data denoising solution process of the RPCA model, improve the computational solution efficiency when the model denoises time series data and the accuracy of the recovery of the low-rank matrix in the time series data, and can effectively balance the collaborative optimization of the low-rank characteristics of the low-rank matrix and the sparsity constraint of the sparse matrix, enhancing the stability of the solution space, and ultimately achieving efficient and accurate denoising of time series data. Description of the Drawings
[0063] Figure 1 It is a schematic flow chart of the time series data denoising method in an embodiment;
[0064] Figure 2 It is the schematic diagram of the smoothing function ρ d (|t|) in an embodiment;
[0065] Figure 3 It is a schematic flow chart of using the improved RPCA model to denoise time series data in an embodiment;
[0066] Figure 4 It is a schematic comparison diagram of the data after denoising the attribute K in the logging data and the noise-free original data in an embodiment;
[0067] Figure 5 It is the internal structure diagram of a computer device in an embodiment. Detailed Implementation Manner
[0068] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0069] In one embodiment, as Figure 1 shown, a method for denoising time series data is provided, including the following steps:
[0070] Step S1: Preprocess the time series data.
[0071] Among them, the preprocessing includes normalizing the time series data.
[0072] Step S2: On the basis of the original RPCA model, introduce an approximate L0 norm based on a fractional function structure, and add a weighted nuclear norm and a control term to construct an improved RPCA model.
[0073] Among them, the construction process of the improved RPCA model includes:
[0074] Adopt a smoothing function ρ d (|t|) to replace the discontinuous norm in the original RPCA model, and construct an RPCA model introducing an approximate L0 norm, expressed as
[0075]
[0076] Among them, A is a low-rank matrix, E is a sparse matrix, D represents the original data matrix, s.t. is a constraint condition, ||A|| * represents the nuclear norm of the low-rank matrix, λ represents the regularization parameter, represents the approximate L0 norm of the variable x, x i is the i-th element in the variable x, N is the number of elements, is used for the approximate L0 norm, t represents the independent variable of the smoothing function, d represents the selected value, and 0.1 ≤ d ≤ 1.
[0077] The function image of the smoothing function ρ d (|t|) is as Figure 2 shown. By introducing an approximate L0 norm into the original RPCA model, the sparse noise in the original data can be accurately identified and separated, so as to effectively capture and separate the noise.
[0078] On the basis of the RPCA model introducing an approximate L0 norm, the nuclear norm of the low-rank matrix A is weighted to obtain an RPCA model with a weighted nuclear norm, expressed as
[0079]
[0080] Among them, w A = {w A,j} represents the weight of the singular value in the low-rank matrix A, σ j represents the j-th singular value in the low-rank matrix A, and w A,j represents the weight of σ j ; among them, the weight of the weighted nuclear norm is used to control the sparsity of the sparse matrix.
[0081] Further, based on the RPCA model of the weighted nuclear norm, a control term is further added to construct an improved RPCA model, expressed as
[0082]
[0083] Among them, λ 1 and λ 2 respectively represent the first regularization parameter and the second regularization parameter, represents the F-norm control term; among them, the control term is used to control the robustness of the solution of the low-rank matrix.
[0084] By adding the weighted nuclear norm and the control term, the robustness and sparsity in the solution process of the RPCA model can be effectively balanced, so that the recovery accuracy of the RPCA model for the low-rank matrix can be effectively improved.
[0085] Step S3, input the preprocessed time-series data into the improved RPCA model. Based on this model, the valid data in the input time-series data is divided into a low-rank matrix, and the noise data in the input time-series data is divided into a sparse matrix. And a collaborative optimization problem that balances the low-rank characteristics of the low-rank matrix and the sparsity constraint of the sparse matrix is constructed, and the collaborative optimization problem is iteratively solved by constructing an augmented Lagrangian function until the iteration convergence condition is met and the iteration stops, and the denoised time-series data and noise data are output.
[0086] Among them, the process of the improved RPCA model for time-series data denoising is as Figure 3 shown, including the following steps:
[0087] First, take the preprocessed time-series data as the original data matrix D and input it into the improved RPCA model, and initialize the parameters, that is, divide the valid data in the input time-series data into a low-rank matrix A, divide the noise data in the input time-series data into a sparse matrix E, and construct a collaborative optimization problem that balances the low-rank characteristics of the low-rank matrix A and the sparsity constraint of the sparse matrix E.
[0088] Construct an augmented Lagrangian function for this collaborative optimization problem, expressed as
[0089]
[0090] Among them, \(Y\) represents the Lagrange multiplier, \(\mu\) represents the parameter, and \(B\) represents the intermediate variable.
[0091] Let \(Y=(Y 1 , Y 2 ), \mu = (\mu 1 , \mu 2 ), and a new expression of the augmented Lagrangian function is obtained as
[0092]
[0093] Among them, \(Y 1 and \(Y 2 represent the first Lagrange multiplier and the second Lagrange multiplier respectively, \(\mu 1 and \(\mu 2 represent the first parameter and the second parameter respectively, and \(P d (E)\) represents the approximate \(L_0\)-norm of the sparse matrix \(E\).
[0094] The new expression of the augmented Lagrangian function is solved by an alternating iteration method, and the intermediate variable \(B\) is obtained, expressed as
[0095]
[0096] Then, first fix the sparse matrix \(E\) and update the low-rank matrix \(A\), and the expression is
[0097]
[0098] Among them, \(A k+1 represents the low-rank matrix of the \((k + 1)\)th iteration, \(U\) represents the left singular vector, represents the soft threshold operator when updating the low-rank matrix \(A\), \(\Sigma\) represents the singular value vector, \(V\) represents the right singular vector, represents the singular value operator, and the superscript \(T\) represents the matrix transpose.
[0099] Then, fix the low-rank matrix \(A\) and introduce the DC (difference of convex functions) programming idea to update the sparse matrix \(E\), and the expression is
[0100]
[0101] Among them, \(E k+1 represents the sparse matrix of the \((k + 1)\)th iteration, represents the soft threshold operator when updating the sparse matrix \(E\), \(V k+1 represents \(V\) of the \((k + 1)\)th iteration, and \(V k+1 is defined as
[0102]
[0103] Among them, \(sign(E k) represents the soft threshold function, d represents the selected value, and 0.1 ≤ d ≤ 1, E k represents the sparse matrix of the k-th iteration.
[0104] When the low-rank matrix A and the sparse matrix E converge to the corresponding and respectively, for Y 1 and Y 2 , μ 1 and μ 2 are iteratively updated. Among them, and represent the low-rank matrix and the sparse matrix obtained by the (k + 1)-th iteration convergence respectively; Y 1 and Y 2 are updated as follows:
[0105]
[0106] Among them, Y 1 k+1 and represent Y 1 and Y 2 of the (k + 1)-th iteration respectively, Y 1 k and represent Y 1 and Y 2 of the k-th iteration respectively, and represent μ 1 and μ 2 of the k-th iteration respectively, A k+1 , E k+1 and B k+1 represent the low-rank matrix A, the sparse matrix E, and the intermediate variable B of the (k + 1)-th iteration respectively.
[0107] μ 1 and μ 2 are updated as follows:
[0108]
[0109] Among them, ρ is a constant and ρ > 1, μ k represents the parameter μ of the k-th iteration, ε is a relatively small positive number, and represent the sparse matrices obtained by the (k + 1)-th and k-th iteration convergences respectively, ||D|| F represents the Frobenius norm of the original data matrix D.
[0110] After reaching the preset iteration output accuracy or reaching the preset iteration number threshold, the denoised time series data and noise data are output.
[0111] The above method performs denoising on time series data by constructing an improved RPCA model. By introducing an approximate L0 norm based on a fractional function structure, this model can accurately capture sparse noise in time series data. By adding a weighted nuclear norm and a control term, it can effectively balance the robustness and sparsity in the process of solving data denoising using the RPCA model, improving the computational efficiency of the model for time series data denoising and the accuracy of low-rank matrix recovery in time series data. Based on this improved RPCA model, efficient and accurate denoising of time series data can be achieved.
[0112] To further verify the beneficial effects of the time series data denoising method proposed in this application, well logging data in industrial sensing data is specifically used as an example for data denoising. The denoising process includes:
[0113] The first step: Normalize the well logging data. Specifically, in this application, 10% sparse noise is added to the well logging data. The dimension of this data matrix is 8×256, and each row in the matrix represents an attribute data in the well logging data;
[0114] The second step: Based on the original RPCA model, introduce an approximate L0 norm based on a fractional function structure, and add a weighted nuclear norm and a control term to construct an improved RPCA model.
[0115] The third step: Input the normalized well logging data into the improved RPCA model for data denoising, and output the denoised well logging data and noise data.
[0116] Furthermore, three traditional data denoising algorithms, namely ILAM (inexact augmented Lagrangian multiplier method), ELAM (exact augmented Lagrangian multiplier method), and Mo-ST0 (inertial optimization smoothing algorithm), are selected to conduct a denoising performance comparison experiment with the method proposed in this application. The experimental results are shown in Table 1, and the performance evaluation indicators are relative error and iteration time.
[0117] Table 1 Denoising performance comparison
[0118] Algorithm Iteration Time (S) Relative Error ILAM 0.5549 0.07739 ELAM 2.1213 0.07214 Mo-ST0 0.3446 0.01083 The method proposed in this application 0.2932 0.01035
[0119] It can be seen from the experimental results shown in Table 1 that the method proposed in this application has higher computational efficiency, better recovery accuracy, and better denoising effect on time series data than other algorithms. Moreover, the effect of denoising attribute K in well logging data using the method proposed in this application is as Figure 4 shown. As Figure 4 can be seen, the data after denoising attribute K in the well logging data basically coincides with the original noise-free data, indicating that the method proposed in this application can achieve accurate data denoising.
[0120] In one embodiment, a device for denoising time-series data is provided, including:
[0121] A preprocessing module for preprocessing the time-series data;
[0122] A model construction module for introducing an approximate L0 norm based on a fractional function structure and adding a weighted nuclear norm and a control term on the basis of the original RPCA model to construct an improved RPCA model;
[0123] A data denoising module for inputting the preprocessed time-series data into the improved RPCA model, dividing the valid data in the input time-series data into a low-rank matrix, dividing the noise data in the input time-series data into a sparse matrix, and constructing a collaborative optimization problem that balances the low-rank characteristics of the low-rank matrix and the sparsity constraint of the sparse matrix, and iteratively solving the collaborative optimization problem by constructing an augmented Lagrangian function until the iteration converges, and then stopping the iteration, and outputting the denoised time-series data and noise data.
[0124] For the specific limitations of the device for denoising time-series data, reference may be made to the limitations on the method for denoising time-series data in the foregoing text, which will not be elaborated herein. Each module in the above device for denoising time-series data can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0125] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 5 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for denoising time-series data. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0126] Those skilled in the art can understand, Figure 5The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0127] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0128] Preprocess the time series data;
[0129] Based on the original RPCA model, introduce an approximate L0 norm based on the fractional function structure, and add a weighted nuclear norm and a control term to construct an improved RPCA model;
[0130] Input the preprocessed time series data into the improved RPCA model. Based on this model, divide the valid data in the input time series data into a low-rank matrix, divide the noise data in the input time series data into a sparse matrix, and construct a collaborative optimization problem that balances the low-rank characteristics of the low-rank matrix and the sparsity constraint of the sparse matrix. Iteratively solve this collaborative optimization problem by constructing an augmented Lagrangian function until the iteration convergence condition is met and then stop the iteration, and output the denoised time series data and noise data.
[0131] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0132] Preprocess the time series data;
[0133] Based on the original RPCA model, introduce an approximate L0 norm based on the fractional function structure, and add a weighted nuclear norm and a control term to construct an improved RPCA model;
[0134] Input the preprocessed time series data into the improved RPCA model. Based on this model, divide the valid data in the input time series data into a low-rank matrix, divide the noise data in the input time series data into a sparse matrix, and construct a collaborative optimization problem that balances the low-rank characteristics of the low-rank matrix and the sparsity constraint of the sparse matrix. Iteratively solve this collaborative optimization problem by constructing an augmented Lagrangian function until the iteration convergence condition is met and then stop the iteration, and output the denoised time series data and noise data.
[0135] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0136] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0137] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.
Claims
1. A time series data denoising method, characterized in that: The method comprises: Preprocess time series data; On the basis of the original RPCA model, the approximate L0 norm based on the fractional function structure is introduced, and the weighted nuclear norm and control terms are added to construct an improved RPCA model. The preprocessed time series data is input into the improved RPCA model. Based on the model, the valid data in the input time series data is divided into a low-rank matrix, and the noise data in the input time series data is divided into a sparse matrix. A collaborative optimization problem that balances the low-rank characteristics of the low-rank matrix and the sparsity constraints of the sparse matrix is constructed. The collaborative optimization problem is iteratively solved by constructing an augmented Lagrangian function until the iteration convergence condition is met, and the iteration is stopped, and the denoised time series data and noise data are output.
2. The method according to claim 1, characterized in that The preprocessing includes: normalizing the time series data.
3. The method according to claim 1, characterized in that On the basis of the original RPCA model, an approximate L0 norm based on the fractional function structure is introduced, and the weighted nuclear norm and control terms are added to construct an improved RPCA model, including: Using a smoothing function ρ based on the fractional function structure d (t) replaces the discontinuous norm in the original RPCA model and constructs an RPCA model that introduces the approximate L0 norm, which is expressed as Among them, A is a low-rank matrix, E is a sparse matrix, D represents the original data matrix, st is the constraint condition, ||A|| * represents the nuclear norm of the low-rank matrix, λ represents the regularization parameter, represents the approximate L0 norm of variable x, x i is the i-th element in variable x, N is the number of elements, Used to approximate the L0 norm, t represents the independent variable of the smoothing function, d represents the selected value, and 0.1≤d≤1; Based on the introduction of the RPCA model with approximate L0 norm, the low-rank matrix A is weighted by the nuclear norm to obtain the RPCA model with weighted nuclear norm, which is expressed as Among them, w A ={w A,j } represents the weight of the singular value in the low-rank matrix A, σ j represents the jth singular value in the low-rank matrix A, w A,j Represents σ j The weight of the weighted nuclear norm is used to control the sparsity of the sparse matrix; On the basis of the RPCA model of weighted nuclear norm, the control term is further added to construct an improved RPCA model, which is expressed as Among them, λ1 and λ2 represent the first regularization parameter and the second regularization parameter respectively. represents the F-norm control term, where the control term is used to control the robustness of the low-rank matrix solution.
4. The method according to claim 3, characterized in that A collaborative optimization problem is constructed to balance the low-rank characteristics of the low-rank matrix and the sparsity constraints of the sparse matrix. The collaborative optimization problem is iteratively solved by constructing an augmented Lagrangian function until the iteration convergence condition is met. The denoised time series data and noise data are output, including: Based on the improved RPCA model, a collaborative optimization problem of balancing the low-rank characteristics of the low-rank matrix A and the sparsity constraints of the sparse matrix E is constructed; The augmented Lagrangian function is constructed for this collaborative optimization problem and is expressed as Among them, Y represents the Lagrange multiplier, μ represents the parameter, and B represents the intermediate variable; The expression of the augmented Lagrangian function is iteratively solved by an alternating iteration method until the iteration convergence condition is met, and the iteration is stopped, and the denoised time series data and noise data are output.
5. The method according to claim 4, characterized in that The expression of the augmented Lagrangian function is iteratively solved by alternating iteration until the iteration convergence condition is met, and the denoised time series data and noise data are output, including: Let Y = (Y1, Y2), μ = (μ1, μ2), and the new expression of the augmented Lagrangian function is: Among them, Y1 and Y2 represent the first Lagrange multiplier and the second Lagrange multiplier respectively, μ1 and μ2 represent the first parameter and the second parameter respectively, D represents the original data matrix, and w A,j represents the jth singular value σ in the low-rank matrix A j The weight of n is the number of singular values, λ1 and λ2 represent the first regularization parameter and the second regularization parameter respectively, P d (E) represents the approximate L0 norm of the sparse matrix E; The new expression of the augmented Lagrangian function is solved by alternating iteration to obtain the intermediate variable B, which is expressed as Then, the sparse matrix E is fixed and the low-rank matrix A is updated. Then, the low-rank matrix A is fixed and the DC planning idea is introduced to update the sparse matrix E. When the low-rank matrix A and the sparse matrix E converge to the corresponding and When Y1 and Y2, μ1 and μ2 are iteratively updated until the preset iterative output accuracy is reached or the preset threshold of iteration times is reached, the denoised time series data and noise data are output; where k is the number of iterations, and They represent the low-rank matrix and sparse matrix obtained by convergence of the k+1th iteration respectively.
6. The method according to claim 5, characterized in that First, fix the sparse matrix E and update the low-rank matrix A. Then fix the low-rank matrix A and introduce the DC planning idea to update the sparse matrix E, including: Fix the sparse matrix E and update the expression of the low-rank matrix A as Among them, A k+1 represents the low-rank matrix of the k+1th iteration, U represents the left singular value vector, represents the soft threshold operator when updating the low-rank matrix A, Σ represents the singular value vector, V represents the right singular value vector, represents the singular value operator, w A ={w A,j } represents the weights of the singular values in the low-rank matrix A, and the superscript T represents the matrix transpose; Fix the low-rank matrix A and introduce the DC planning idea to update the expression of the sparse matrix E: Among them, E k+1 represents the sparse matrix of the k+1th iteration, represents the soft threshold operator when updating the sparse matrix E, λ represents the regularization parameter, V k+1 represents V of the k+1th iteration, and V k+1 Defined as Among them, sign(E k ) represents the soft threshold function, d represents the selected value, and 0.1≤d≤1, E k Represents the sparse matrix of the k-th iteration.
7. The method according to claim 5, characterized in that The update method of Y1 and Y2 is: in, and denote Y1 and Y2 of the k+1th iteration respectively, and denote Y1 and Y2 of the kth iteration respectively, and They represent μ1 and μ2 of the kth iteration, A k+1 、E k+1 and B k+1 Respectively represent the low-rank matrix A, sparse matrix E and intermediate variable B of the k+1th iteration; The update method of μ1 and μ2 is: Where ρ is a constant and ρ>1, μ k Represents the parameter μ of the kth iteration, ε is a relatively small positive number, and They represent the sparse matrices obtained by the k+1th and kth iterations, respectively, ‖D‖ F Represents the F-norm of the original data matrix D.
8. A time series data denoising device, characterized in that: The device comprises: Preprocessing module, used to preprocess time series data; A model building module is used to introduce an approximate L0 norm based on a fractional function structure and add a weighted nuclear norm and a control term on the basis of the original RPCA model to build an improved RPCA model; The data denoising module is used to input the preprocessed time series data into the improved RPCA model, divide the valid data in the input time series data into a low-rank matrix based on the model, divide the noise data in the input time series data into a sparse matrix, and construct a collaborative optimization problem that balances the low-rank characteristics of the low-rank matrix and the sparsity constraints of the sparse matrix. The collaborative optimization problem is iteratively solved by constructing an augmented Lagrangian function until the iteration is stopped when the iterative convergence condition is met, and the denoised time series data and noise data are output.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.