Color image reconstruction method, terminal equipment and storage medium
By introducing TR decomposition and FCTN decomposition with overlapping group sparse regularization constraints in the gradient domain, and combining them with nonlocal similarity block matching, the problem of global low rank and nonlocal self-similarity being ignored in existing methods is solved, and high-quality color image reconstruction is achieved.
Patent Information
- Application Number
- CN202510100356.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2026-01-13
AI Technical Summary
Existing color image inpainting methods ignore the global low-rank property and non-local self-similarity of images, resulting in poor reconstruction quality when dealing with large-area missing or irregular regions of missing data.
The TR decomposition model with overlapping group sparse regularization constraint is used to process images in the gradient domain. Combined with the FCTN decomposition model, the global low rank and local smoothness of the image are mined through non-local similarity block matching and hybrid tensor network, and the non-local similarity is used for image reconstruction.
It effectively reconstructs lost pixel information in color images, improving image infill quality, especially showing better results when large or irregular areas are missing.
Smart Images

Figure CN121329818A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing, and in particular to a color image reconstruction method, a terminal device and a storage medium. BACKGROUND
[0002] Color images contain rich spatial and color information, compared with hyperspectral images, remote sensing images, satellite images and other images, color images are more easily collected at low cost and understood by humans, and therefore have the most extensive application in various computer vision tasks. However, in the image acquisition, storage, transmission and other links, due to inevitable factors (such as sensor failure, storage unit damage, transmission signal loss, etc.), color images often have pixel loss phenomenon. In order to obtain more rich and complete image information, it has high theoretical and engineering value to estimate the missing pixel information (i.e. image filling) using the observed pixel information.
[0003] Tensor is a high-order extension of vector and matrix. Compared with matrix, it can express more rich information, so it is naturally used to represent and process multi-dimensional data in the real world (such as color images, video sequences, hyperspectral images, remote sensing images, etc.). In tensor data processing, tensor completion estimates missing elements from observed data, and has wide application in real life, such as video foreground and background separation, image cloud removal and seismic data reconstruction.
[0004] The color image filling problem is a typical ill-posed problem, and the core is to mine the intrinsic prior knowledge of the image. For a long time, traditional color image filling methods often use the way of processing matrix by channel to mine the image prior, such as median filtering method, total variation method, sparse representation method, etc. Although the above methods can achieve a certain effect of image restoration, but these methods all ignore the global low-rank prior of color image. Color image can be regarded as a third-order tensor containing one color dimension and two spatial dimensions, which has global low-rank property. This low-rank property can be captured by tensor decomposition. Unlike the rank of matrix, the rank of tensor has different definitions according to different tensor decompositions, such as Candecomp-Parafac (CP) rank, Tucker rank, tensor tube rank, tensor train (TT) rank, tensor ring (TR) rank, fiber rank, full connected tensor network (FCTN) rank, etc.
[0005] Tensor network decomposition can effectively mine the redundancy of high-dimensional tensors, which is a method to decompose high-dimensional tensors into low-dimensional factors and ensure the low-rank structure of high-dimensional tensors through the low-rank constraint of factors. Common tensor decompositions include CP decomposition, Tucker decomposition, tensor singular value decomposition, three-dimensional tensor singular value decomposition, tensor chain decomposition, tensor ring decomposition, FCTN decomposition, etc. These tensor decomposition methods have been gradually applied to image reconstruction tasks and have achieved good results. Among them, CP decomposition is defined as the sum of the smallest number of rank 1 tensor outer products. In practical calculations, the calculation of CP rank is an NP-hard problem. Most methods based on Tucker decomposition expand the tensor into a matrix, which inevitably destroys the global structure of high-dimensional images and only focuses on the correlation between one modality and other modalities. As the third-order counterpart of matrix singular value decomposition, Kilmer et al. proposed a new tensor singular value decomposition and defined the tube rank based on this decomposition (Kilmer M E, Braman K, Hao N, et al. Third-order tensors as operators on matrices: A theoretical and computational framework with applications in imaging [J]. SIAM Journal on Matrix Analysis and Applications, 2013, 34(1): 148-172). Subsequently, Lu et al. proposed a convex substitute for the tube rank (tensor core norm) and used it for tensor robust principal component analysis (Lu C, Feng J, Chen Y, et al. Tensor robust principal component analysis with a new tensor nuclear norm [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019, 42(4): 925-938). Methods based on tensor singular value decomposition treat each modality of the tensor unequally, so they cannot flexibly describe the correlation between different modalities. In recent years, more and more tensor decompositions based on tensor networks have emerged and shown strong ability to handle high-order tensors (especially tensors higher than three orders), such as tensor chain decomposition and TR decomposition. Due to the excellent data compression ability and image reconstruction ability of TT decomposition and TR decomposition, they have been successfully applied in many fields such as signal recovery, lightweight neural network design, etc.
[0006] Among the numerous tensor decomposition models, TR decomposition has the following advantages: (1) The number of variables in traditional tensor decomposition increases exponentially with respect to the tensor dimension, while the number of storage variables in TR decomposition is linearly related to the original tensor dimension, so it is much smaller than the number of variables in traditional tensor decomposition; (2) TR decomposition has the cyclic shift invariance property that most traditional tensor decompositions do not have, effectively improving the stability and flexibility of image reconstruction; (3) The decomposition form of TR is tensor to tensor, that is, the original tensor can be represented as a multilinear product of a series of 3-order tensor factors. This decomposition method avoids the information structure destruction problem caused by the unfolding operation in Tucker decomposition. Due to the above advantages, TR decomposition has been applied to classic image reconstruction problems such as image completion and image denoising since its inception, and has achieved good image reconstruction performance. However, the choice of TR rank is sensitive to the TR model. To solve the above problem, Yuan et al. imposed a nuclear norm minimization constraint on the unfolding matrix of the TR factor tensor and proposed a TR factor low-rank constraint (Tensor Ring with Low Rank Factors, TRLRF) model (Yuan L, Li C, Mandic D, et al. Tensor ring decomposition with rank minimization on latent space: An efficient approach for tensor completion [C]. AAAI Conference on Artificial Intelligence, 2019: 9151-9158). Note that when the image missing ratio is very high, considering only the global low-rank property cannot provide satisfactory image details. Fortunately, natural images often exhibit local continuity. To exploit this prior of images, Wu et al. introduced a sparse regularization constraint in the gradient domain of the modulo-2 unfolding matrix of the TR factor and proposed a TR with Gradient Factors Regularization (TRGFR) model (Wu P, Zhao X, Ding M, et al. Tensor ring decomposition-based model with interpretable gradient factors regularization for tensor completion [J]. Knowledge-Based Systems, 2023, 259: 110094), which provides robustness for the choice of TR rank through factor regularization constraints and takes into account both the global low-rank property and the local smoothness of images.It should be noted that there are two limitations in the TR-based tensor decomposition network: (1) only the connection between adjacent factors in the form of tensor contraction is established, rather than the connection of any factor, which leads to the limitation of tensor factor correlation representation; (2) the TR decomposition is highly sensitive to the arrangement order of the tensor modal, leading to the inflexibility of decomposition and application. To solve the above problems, Zheng et al. proposed a fully-connected tensor network decomposition network (Zheng Y, Huang T, Zhao X, et al. Fully-connected tensor network decomposition and its application to higher-order tensor completion [C]. AAAI Conference on Artificial Intelligence, 2021: 11071-11078), which effectively solves the above problems. It is worth noting that when dealing with tensors with a dimension of 3, the FCTN decomposition model is equivalent to the TR decomposition model. Therefore, for color images, the TR decomposition model can be selected to complete the image filling task.
[0007] The above-mentioned tensor decomposition-based method only focuses on the global correlation of the tensor, ignoring the non-local self-similarity (NSS) of the tensor. The non-local self-similarity of the image indicates that the similarity within the image is not limited to the adjacent area. In fact, some areas far apart also have a certain probability of high similarity. This feature is common in various images and is more obvious in images with repeated structures. It should be noted that this similarity cannot be captured in the original color image data space. Therefore, it is necessary to match and stack the image blocks with similarity to further mine the low-rank prior of the non-local similar blocks. Scholars have introduced this prior knowledge in image reconstruction problems. For example: Danielyan et al. proposed an image deblurring algorithm based on three-dimensional non-local patch matching (Danielyan A, Katkovnik V, Egiazarian K. Bm3d frames and variational image deblurring [J]. IEEE Transactions on Image Processing, 2012, 21(4): 1715-1728). Similarly, Gu et al. described the low-rank nature of non-local similar blocks with weighted kernel norm minimization (Gu S, Xie Q, Meng D, et al. Weighted nuclear norm minimization and its applications to low level vision [J]. International Journal of Computer Vision, 2016, 121(2): 183-208). Liu et al. combined the low-rank constraint of the non-local similar block matching matrix with the image total variation constraint to propose an effective image denoising method (Liu J, Osher S. Blockmatching local svd operator based sparsity and tv regularization for image denoising [J]. Journal of Scientific Computing, 2019, 78(1): 607-624). The above-mentioned methods all use the channel-independent image processing method to reconstruct the NSS block, ignoring the correlation of the image channel direction, and cannot fully utilize the NSS characteristics of the three-dimensional image. In order to overcome this problem, many methods use three-dimensional tensor groups as the basic unit of denoising.For example, Chen et al. proposed a hyperspectral image denoising model based on non-local TR decomposition (Chen Y, He W, Yokoya N, et al. Nonlocal tensor-ring decomposition for hyperspectral image denoising [J]. IEEE Transactions on Geoscience and Remote Sensing, 2019, 58(2): 1348-1362). Similarly, Zheng et al. proposed a hyperspectral image inpainting method based on non-local patch matching and fully connected tensor network (Zheng W-J, Zhao X-L, Zheng Y-B, et al. Nonlocal patch-based fully connected tensor network decomposition for multispectral image inpainting [J]. IEEE Geoscience Remote Sensing Letters, 2021, 19: 1-5). To further improve the quality of image inpainting, Zhao et al. (Zhao X-L, Yang J-H, Ma T-H, et al. Tensor completion via complementary global, local, and nonlocal priors [J]. IEEE Transactions on Image Processing, 2021, 31: 984-999) used TT rank to represent the global correlation of the image, and introduced a convolutional neural network denoiser and a block matching three-dimensional filter denoiser to preserve the local details and non-local similarity of the image, respectively. SUMMARY
[0008] To solve the above problems, the present application provides a color image reconstruction method, a terminal device and a storage medium.
[0009] The specific scheme is as follows:
[0010] A color image reconstruction method, comprising the following steps:
[0011] S1: After adding an overlapping group sparse regularization constraint in the gradient domain of the TR factor, a TR decomposition reconstruction model with factor regularization constraint is obtained, and then the TR decomposition reconstruction model is used to process the image to be filled to obtain a preliminary estimated image;
[0012] The TR decomposition reconstruction model is expressed as:
[0013]
[0014] Wherein, Ω represents the set of observation element indicators corresponding to the image to be filled; This represents the projection operator, which is used to retain elements in the observation element index set and map elements outside the observation element index set to 0; This represents the initial estimated image output by the model; TR(.) denotes TR decomposition. Represents the TR factor; ||.||2 represents the l2 norm; λ represents the k-th factor; k represents the index of the TR factor; K represents the total number of TR factors; λ k This represents the regularization balance parameter corresponding to the k-th factor; The superscript T indicates the sparse regularization term of the overlapping group; the superscript T indicates the transpose of the matrix; vec(.) indicates the matrix columnization operator; G indicates the magnitude of the combination value; D represents the tensor corresponding to the image to be filled. (k) This represents the cyclic difference matrix corresponding to the k-th factor; U represents the modulo-2 expansion of the k-th factor; (k) V represents the result of the semi-orthogonal transformation of the k-th factor; (k) Let I represent the transformation matrix of the k-th factor; I represents the identity matrix.
[0015] S2: After dividing the initially estimated image into small blocks, stack similar blocks into NSS groups;
[0016] S3: Based on all the obtained non-local self-similar groups, construct a set of (K+1) order NSS groups;
[0017] S4: Process the (K+1) order NSS group using FCTN decomposition to obtain the reconstructed NSS group;
[0018] S5: Aggregate the reconstructed NSS groups back to their original locations to obtain the filled image.
[0019] Furthermore, when solving the TR decomposition and reconstruction model, the PAM algorithm is used. U (k) V (k) and The four factors are updated alternately until the stopping condition is met.
[0020] A color image reconstruction terminal device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method described above in the embodiments of the present invention.
[0021] A computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the steps of the method of the embodiments of the application.
[0022] The application adopts the above technical solution, and compared with the existing image filling method, the method has superior performance, can reconstruct the lost pixel information in the color image with high quality, and achieves satisfactory results. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 Figures (a) and (b) show the difference between the TR decomposition model and the FCTN decomposition model in the first embodiment of the application, wherein, Figure 1 (a) is the TR decomposition model (K=3); Figure 1 (b) is the FCTN decomposition model (K=3); Figure 1 (c) is the TR decomposition model (K≥4); Figure 1 (d) is the FCTN decomposition model (K≥4).
[0024] Figure 2 Figures (a) and (b) show the flowchart of the method in the embodiment.
[0025] Figure 3 Figures (a) and (b) show the advantage of the calculation method of the overlapping group and the gradient in the embodiment, wherein, Figure 3 (a) is the image smoothing area extraction vector x1=(0,0,1,0,0) with outliers; T schematic diagram; Figure 3 (b) is the gradient of x1 under the periodic boundary condition; schematic diagram; Figure 3 (c) is the combined gradient schematic diagram of x1; Figure 3 (d) is the image edge area extraction vector x2=(1,2,3,4,5) without outliers; T schematic diagram; Figure 3 (e) is the gradient of x2 under the periodic boundary condition; schematic diagram; Figure 3 (f) is the combined gradient schematic diagram of x2. DETAILED DESCRIPTION
[0026] To further illustrate the embodiments, the application provides accompanying drawings. These drawings are part of the disclosure of the application, mainly used to illustrate the embodiments, and can be used to explain the operating principle of the embodiments in conjunction with the related description of the specification. Those skilled in the art should understand other possible implementation manners and advantages of the application by referring to these contents.
[0027] The application will be further described in conjunction with the accompanying drawings and specific embodiments.
[0028] Embodiment one:
[0029] Brief introduction of TR decomposition and FCTN decomposition.
[0030] (1) Notations and preliminary explanations.
[0031] In this paper, we use x, x, X and to denote scalar, vector, matrix and tensor, respectively. For a K-order tensor we use to denote its (i1, i2, …, i K )th element. The l2-norm of a vector is defined as
[0032] (2) TR decomposition model
[0033] Tensor ring decomposition (TRD) aims to represent a high-order (or high-dimensional) tensor by a sequence of third-order tensors that are cyclically multiplied. Specifically, let be a K-order tensor with size I1×I2×...×I K , denoted as TR decomposition decomposes a high-dimensional tensor into a sequence of latent tensors. In TR decomposition, the element of is defined as:
[0034]
[0035] where denotes the (i1, i2, …, i K )th element of the tensor ; Tr denotes the trace of a matrix; forms a circular sequence, is connected with , and the last tensor factor has size R K ×I K ×R1, i.e., R K+1 = R1; thus ensuring that is a square matrix; denotes the i k th side slice matrix of the latent tensor ; r TR = [R1, R2, …, R K ] T is called the TR rank. From equation (1), we have which is equivalent to the trace of the sequential product of the matrices {G (k) (i k )}. If we let R1 = R2 = … = R K = R, the number of parameters in TR decomposition is It is linearly related to the tensor order K.
[0036] Definition 1 (modulo k expansion): A tensor The modulo n expansion is represented as Tensor elements Mapping to matrix elements, i.e.:
[0037]
[0038] in The definition is as follows:
[0039]
[0040] To simplify the expression, equation (3) will be uniformly described as follows: in Let {k} be a vector formed by rearranging any elements in the subset after removing k from the set {1,2,…,K}.
[0041] Definition 2 (Inverse Modular Expansion): The inverse modulus k expansion matrix is of size I. k ×Π s≠k I s R [k] It indicates that its elements are defined as:
[0042]
[0043] Definition 3 (merging adjacent kernel factor subchains): Let It is the TR representation of a K-order tensor, where yes TR nuclei. Due to adjacent TR nuclei and Having the same mold size R k+1 Therefore, they can be combined into a single core through multilinear product. Its horizontal slice matrix is:
[0044]
[0045] Subchain tensors are merged by dividing the k-th factor. All nuclei outside, i.e. Represented as Its slice matrix is defined as:
[0046]
[0047] (3) Least squares algorithm for TR decomposition.
[0048] This section introduces a TR decomposition algorithm using Alternating Least Squares (ALS). In the ALS algorithm, when optimizing one TR factor, other TR factors are fixed, and this process will be repeated until the convergence condition is met. Given a K-order tensor The goal is to optimize the TR factors of a given TR rank r TR , that is,
[0049]
[0050] Solving equation (7) requires the use of Theorem 1.
[0051] Theorem 1: Given a TR decomposition Its inverse order k unfolding matrix can be written as:
[0052]
[0053] Proof:
[0054] According to the definition of TR decomposition in equation (1), we can get:
[0055]
[0056] Therefore, The inverse order k unfolding matrix of
[0057]
[0058] Then equation (8) holds, so we have:
[0059]
[0060] where rank denotes the rank of the matrix.
[0061] (4) FCTN decomposition model
[0062] Compared with traditional tensor decomposition networks (such as TT and TR decomposition), FCTN can fully exploit the correlation between decomposition factors and has axis permutation invariance, so it has better image reconstruction quality in practical applications. Specifically, by decomposing a K-order tensor into a series of K-order factor tensors with less parameters and establishing the relationship between any two factors, FCTN decomposition has the ability to fully represent the global correlation of the tensor. Mathematically, given a K-order tensor FCTN decomposition representation Each element of
[0063]
[0064] where FCTN factor and vectors (R 1,2 , R 1,3 ,..., R 1,K , R 2,3 ,..., R 2,K ,..., R K-1,K ) are called FCTN ranks, i.e., r FCTN . For simplicity, the above formula can be written as If R 1,2 = R 1,3 =... = R 1,K = R 2,3 =... = R 2,K =... = R K-1,K = R, the number of parameters in the FCTN decomposition is which is linearly related to the order K of the tensor.
[0065] Figure 1 Reflecting the difference between the TR decomposition model and the FCTN decomposition model, it can be found from the figure that the following rules: (1) as shown in Figure 1 (a) and Figure 1 (b), when the order K of the tensor is 3, the TR decomposition is equivalent to the FCTN decomposition; (2) as shown in Figure 1 (c) and Figure 1 (d), when the order K of the tensor is greater than or equal to 4, the TR decomposition cannot consider the correlation of non-adjacent factors. In contrast, the FCTN decomposition not only considers the correlation of adjacent factors, but also considers the correlation of non-adjacent factors, which can more effectively mine the mutual constraint relationship between factors.
[0066] The embodiment proposes a three-stage color image reconstruction method. As shown in Figure 2 , it includes the following steps:
[0067] S1: After adding an overlapping group sparse regularization constraint in the gradient domain of the TR factor, a TR decomposition reconstruction model with factor regularization constraint is obtained, and then the TR decomposition reconstruction model is used to process the image to be filled to obtain a preliminary estimated image;
[0068] S2: After dividing the preliminary estimated image into small blocks, similar blocks are stacked into NSS groups;
[0069] S3: According to all the obtained non-local self-similar groups, a group of (K+1) order NSS groups is constructed;
[0070] S4: The (K+1) order NSS group is processed by using the FCTN decomposition to obtain a reconstructed NSS group;
[0071] S5: The reconstructed NSS group is gathered to their original position to obtain a filled image.
[0072] The above steps are mainly divided into three stages. In the first stage, a TR decomposition network with factor regularization constraint is used to mine the local smoothness and global low-rankness of the three-dimensional color image. In the second stage, the preliminary estimated image in the first stage is used as the basis to find non-local similar blocks, and the similar blocks are stacked into a fourth-order tensor. Then, the FCTN in the third stage is used to further strengthen the low-rankness of the non-local similar blocks. Finally, the reconstructed non-local similar blocks are stacked into the reconstructed image.
[0073] Stage 1: Third-order TR reconstruction stage of color image.
[0074] The TR model effectively describes the relationship between different TR factors, thereby better describing the global low-rankness of the tensor. However, the traditional TR model is sensitive to the selection of the TR rank. Therefore, the embodiment introduces an overlapping group sparsity constraint in the gradient domain of the TR factor, thereby effectively characterizing the global low-rankness and local smoothness of the image by using the low-rankness of the TR factor and the overlapping group sparsity of its gradient. By increasing the overlapping group sparsity regularization constraint in the gradient domain of the TR factor, it is beneficial to improve the reconstruction performance of the estimated image and reduce the sensitivity of the TR rank selection. The model proposed in the embodiment further improves the robustness of the TR decomposition model to the TR rank selection from the two priors (overlapping group sparsity prior and low-rank prior) in the gradient domain of the TR factor. The motivation of the method proposed in the embodiment is explained from two aspects as follows:
[0075] (1) Overlapping group sparsity in the gradient domain of the TR factor. It is noted that formula (8) is established, and then ( denotes the Moore-Penrose generalized inverse), which indicates that each column of is a linear combination of the columns of R [k] An interesting phenomenon is that natural images usually have local smoothness, so should have overlapping group sparsity statistical characteristics.
[0076] (2) The rank of the gradient domain of the TR factor can be used to constrain the rank of the decomposed tensor.
[0077] Proof:
[0078] It is noted that the following inequality is established:
[0079]
[0080] where D is a circulant difference matrix, and rank(D (k) ) = I k -1, then:
[0081]
[0082] Comparing equation (11) and equation (14), we have According to the above analysis, we only need to calculate the gradient of TR factor By imposing low-rank constraint, we can guarantee the low-rank property of the original tensor.
[0083] Inspired by the above theoretical analysis, we adopt the data-driven semi-orthogonal transform V (k) Promote The combined sparsity, that is: where Therefore, the TR decomposition reconstruction model of the first stage is as follows:
[0084]
[0085] where Ω represents the observed element index set (clean positions in the broken graph). represents the projection operator, which retains the elements in the element index set and maps the elements outside the index set to 0. I represents the unit matrix. represents the overlapping group sparsity regular term, assuming that the vector then n represents the serial number of elements in the vector, N represents the total number of elements contained in the vector, G is the group value size, G l represents the number of left sampling points, G r represents the number of right sampling points, represents the maximum integer value less than or equal to x. Compared with the traditional sparse regular term ||x||1, the overlapping group sparse regular term has more advantages in protecting image local smoothness and image edges.
[0086] Figure 3 The calculation method of the proposed overlapping group and gradient and its advantages in image edge protection are shown. As can be seen from the figure, Figure 3 (a) shows the smooth region extraction vector x1 with outliers, where the outliers are represented by a dashed box, and the corresponding gradient is Figure 3 (b). Figure 3 (d) shows the image edge region extraction vector x2 in the image without outliers, and the corresponding gradient is Figure 3 (e). If the pixel-by-pixel soft threshold shrinkage operator in the baseline algorithm is used (v(i) represents the element of vector v, and τ represents the soft threshold value), then Figure 3 the second and third pixels of v1 in (b) and Figure 3 the first to fourth pixels of v2 in (e) cannot be effectively distinguished. This means that when selecting the threshold to remove the outliers in the smooth region, the gradient information of the image will be damaged, which will cause the over-smoothing phenomenon of the recovered image. If the combined gradient G=3 is used, as shown inFigure 3 (c) and Figure 3 (f) are shown, then the smooth region and the edge region are effectively distinguished, and a suitable threshold can be selected to eliminate outliers in the smooth region and protect the edge region from being over-smoothed.
[0087] Equation (15) can be rewritten as:
[0088]
[0089] To solve equation (16), the PAM algorithm is introduced in this embodiment to update U (k) ,V (k) , The four factors are updated alternately, that is:
[0090]
[0091] wherein, is the objective function in equation (16), and ρ>0 is a forcing parameter.
[0092] The following four sub-problems are solved respectively:
[0093] (1) The sub-problem of
[0094] The objective function of
[0095]
[0096] Equation (18) is rewritten in matrix form, and then there is:
[0097]
[0098] Let , then:
[0099]
[0100] The solution is:
[0101]
[0102] wherein, and ⊙ represent element-wise division and multiplication respectively, F is a normalized one-dimensional Fourier transform matrix, wherein repvec(v, n) represents repeating the column vector v n times, P is an orthogonal matrix, and satisfies P T P = I1( is a unit matrix); (D (k) ) T D (k) = FH A, F is a unitary matrix, satisfying FF H = I2( is the identity matrix); fold2 is a matrix-to-tensor modulo-2 reshaping operator.
[0103] Proof:
[0104] In equation (20), D (k) is a circulant matrix, so (D (k) ) T D (k) can be diagonalized by a normalized one-dimensional Fourier transform matrix F, i.e.,
[0105] (D (k) ) T D (k) = F H ΛF. (22)
[0106] In addition, since the matrix is a real symmetric matrix, it can be diagonalized by an orthogonal matrix P into a diagonal matrix, i.e.,
[0107]
[0108] Let then we have:
[0109]
[0110] Left-multiply F and right-multiply P on both sides of equation (24), we get:
[0111]
[0112] After simplification, we can get:
[0113]
[0114] Let then it can be rewritten as:
[0115] β k ΛY+ρY+YΣ=FCP. (27)
[0116] For β k ΛY+YΣ=β k Y⊙repvec(diag(Λ),R k R k+1 )+ρY+Y⊙[repvec(diag(Σ),I k )] T , the above equation can be transformed into:
[0117] β kY⊙repvec(diag(Λ),R k R k+1 )+ρY+Y⊙[repvec(diag(Σ),I k )] T = FCP, (28)
[0118] where diag denotes the operator that extracts the diagonal elements of a diagonal matrix and vectorizes them.
[0119] Extracting Y in the above equation, we have
[0120]
[0121] Let we have
[0122] Y⊙E = FCP. (30)
[0123] Then we have
[0124]
[0125] Since we have
[0126]
[0127] (2) (U (k) ) (s+1) is a subproblem.
[0128] (U (k) ) (s+1) The objective function of (U
[0129]
[0130] Since (V (k) ) T V (k) = I, we can rewrite it as
[0131]
[0132] The above equation can be computed by the overlapped group sparse shrinkage, i.e.
[0133]
[0134] where L OGS (x, G, τ) denotes the overlapped group sparse regularization iterative operator, and the specific iterative process is shown in Algorithm 1 shown in Table 1.
[0135] Table 1
[0136]
[0137]
[0138] where I denotes the identity matrix, M is a diagonal matrix whose mth diagonal element is:
[0139]
[0140] (3)(V (k) ) (s+1) subproblem
[0141] (V (k) ) (s+1) The solution of the subproblem requires the use of Theorem 2.
[0142] Theorem 2: The optimal solution of the following problem is A = VU T where U and V are the singular value decomposition matrices of BC T + ρD T .
[0143] Proof:
[0144]
[0145] According to the properties of the trace, Tr[A(BC T + ρD T )] reaches its upper bound if and only if A = VU T where BC T + ρD T = USV T .
[0146] (V (k) ) (s+1) The objective function of the problem
[0147]
[0148] According to Theorem 2, the closed-form solution of (37) is given by (38):
[0149] (V (k) ) (s+1) = ZW T , (38) where
[0150] (4) subproblem
[0151] The objective function of the problem
[0152]
[0153] The update is shown as formula (40):
[0154]
[0155] Ω c is the complement set of Ω.
[0156] In summary, the proposed TR with overlapping group sparsity (TR-OGS) decomposition model can achieve the following goals:
[0157] (1) Data dimension reduction and compression. Through tensor ring decomposition, the original high-order tensor data is converted into the form of low-order tensor, effectively reducing the data dimension and reducing the storage space occupation.
[0158] (2) Use overlapping group sparse regularization technology to accurately describe the local continuity of the image and effectively protect the gradient edge information of the image. Compared with the sparse regularization technology used in the baseline algorithm, the overlapping group sparse regularization constraint used in this embodiment replaces the pixel-by-pixel gradient with a combined gradient sparsity. On the one hand, it can effectively suppress the stair effect caused by the pixel-by-pixel gradient sparse constraint in the smooth area, and on the other hand, it can effectively distinguish between smooth areas and image edge areas by using combined gradients, while recovering the smooth area of the image and avoiding excessive smoothing of the image edge.
[0159] (3) Enhance the rank selection robustness of the TR model. The traditional TR model does not impose regularization constraints on the TR factors, resulting in a relatively sensitive selection of the rank. Through the gradient factor regularization technology, the local smooth data prior of the tensor factor is accurately expressed, thereby effectively improving the generalization ability of the TR tensor reconstruction model.
[0160] The TR-OGS algorithm proposed in this embodiment is shown in Algorithm 2 in Table 2.
[0161] Table 2
[0162]
[0163]
[0164] Stage 2: Non-local similar block matching
[0165] The reconstructed image obtained in the first stage is used as the input image of the second stage to ensure the accuracy of the matching. In the second stage, the initial image is first divided into small blocks, and then similar blocks are stacked into NSS (non-local self-similar) groups. Each original NSS group is a four-order tensor, including two spatial modes, one spectral mode and one similar group mode. Take a for example, which is divided into overlapping cubes The patch size is p and the overlap size is o. It can be seen that Then the key patches are selected from all patches at a fixed interval v. Since the non-local self-similar group is composed of the key patch and its similar patch stacked cubes. Therefore is the number of key patches. Secondly, since the non-local self-similar group has strong global correlation, it is taken as a basic reconstruction unit, and FCTN decomposition is introduced for each group to further mine the NSS group correlation.
[0166] Stage 3: Image reconstruction based on hybrid low-rank tensor network.
[0167] In the third stage, a (K+1) order non-local self-similar group is set, where H represents the number of block matching tensors, and FCTN decomposition is used to obtain the non-local self-similar group which can be expressed as:
[0168]
[0169] where is the projection operator, which is used to retain the entries in Ω j and set other entries to zero, Ω j is the index of known elements in . Using PAM (clustering algorithm) to solve and that is:
[0170]
[0171] where and are the iteration results of the s-th time, Similarly ρ>0 is the proximal parameter. According to the (43) in the literature (Zheng Y, Huang T, Zhao X, et al. Fully-connected tensor network decomposition and its application to higher-order tensor completion [C]. AAAI Conference on Artificial Intelligence, 2021:11071-11078), where represents the FCTN combination after removing the factor . represents the generalized tensor expansion form of .
[0172] Therefore The sub-problem is given by the formula:
[0173]
[0174] The analytical solution of formula (42) is:
[0175]
[0176] Where GenFold represents the operator that reshapes the matrix into a tensor in the literature (Zheng Y, Huang T, Zhao X, et al. Fully-connected tensor network decomposition and its application to higher-order tensor completion [C]. AAAI Conference on Artificial Intelligence, 2021: 11071-11078).
[0177] Finally, the reconstructed non-local self-similar groups are aggregated to their original positions to obtain the filled color image The embodiment proposes an image filling method based on a non-local hybrid low rank tensor network (Non Local Hybrid Low Rank Tensor Network) as shown in algorithm 3 shown in table 3.
[0178] Table 3
[0179]
[0180] The embodiment of the application proposes a non-local similar block matching and hybrid tensor network color image reconstruction method. In the method, first, a TR decomposition network with overlapping group sparse constraints is used to mine the global low rank priori and local smooth priori of the color image. On this basis, we send the reconstructed tensor estimated by the TR decomposition network to the second stage of the non-local block matching stage to construct a high-order tensor with potential non-local similarity. Finally, we introduce a fully connected tensor network to mine the non-local similarity priori of the tensor. The method exhibits excellent performance in the color image filling task. In particular, the proposed method exhibits more outstanding advantages compared to other image filling methods when dealing with complex scenarios of large area missing or irregular region data missing. This outstanding performance is due to the following reasons:
[0181] (1) The proposed method uses a TR decomposition model with overlapping group sparse regularization constraints to obtain preliminary estimates in the first stage, combining the global low-rank property of tensors and local smoothness priors, and addressing the sensitivity of the TR model to rank selection. At the same time, it provides more reliable data for the second stage of non-local similar block matching than the observed tensor with missing data.
[0182] (2) The proposed method uses non-local block matching to find similar blocks in the image through the algorithm, and uses the information of these similar blocks to fill in the missing areas. This non-local method can capture more global structural information of the image than traditional local methods, so it can effectively deal with complex scenarios of large-area missing or irregular area data missing and generate more natural and accurate filling results.
[0183] (3) The proposed method uses a hybrid fully connected tensor net to further enhance the filling ability of the algorithm. By constructing a tensor net, the algorithm can fully utilize the multi-dimensional information of the image, including color, texture and shape, etc. This makes the algorithm perform better when filling complex textures and areas with rich details.
[0184] Embodiment two:
[0185] The application also provides a color image reconstruction terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in the above method embodiments of Embodiment One of the application.
[0186] Further, as an executable solution, the color image reconstruction terminal device can be a desktop computer, a notebook, a palm computer, and a cloud server, etc. The color image reconstruction terminal device can include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above-mentioned composition structure of the color image reconstruction terminal device is only an example of the color image reconstruction terminal device, and does not constitute a limitation on the color image reconstruction terminal device, and can include more or fewer components than the above, or combine certain components, or different components, for example, the color image reconstruction terminal device can also include an input / output device, a network access device, a bus, etc., and the embodiments of the application do not limit this.
[0187] Further, as an executable solution, the processor can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor or the like, and the processor is a control center of the color image reconstruction terminal device, and connects all parts of the color image reconstruction terminal device through various interfaces and lines.
[0188] The memory can be used to store the computer program and / or modules, and the processor realizes various functions of the color image reconstruction terminal device by running or executing the computer program and / or modules stored in the memory, and calling data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required by a function; and the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a nonvolatile memory, for example, a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory devices.
[0189] The application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the steps of the method provided in the embodiments of the application.
[0190] The module / unit of the color image reconstruction terminal equipment integration, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on such an understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium can include any entity or device capable of carrying the computer program code, a recording medium, a U disk, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), and a software distribution medium, etc.
[0191] Although the present application has been particularly shown and described with reference to preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details can be made therein without departing from the spirit and scope of the application as defined by the appended claims.
Claims
1. A method of color image reconstruction, characterized by, The method comprises the following steps: S1: obtaining a TR decomposition reconstruction model with factor regularization constraint by adding overlapping group sparse regularization constraint in the gradient domain of TR factor, and then processing the image to be filled by the TR decomposition reconstruction model to obtain a preliminary estimation image; The TR decomposition reconstruction model is expressed as: Ω denotes the set of observation element indices corresponding to the image to be inpainted; denotes the projection operator that keeps the elements in the set of observation element indices and maps the elements out of the set of observation element indices to 0; denotes the preliminary estimation image output by the model; TR(.) denotes the TR decomposition; denotes the TR factor; ||.||2 denotes the l2 norm; denotes the k-th factor; k denotes the serial number of the TR factor; K denotes the total number of TR factors; λ k denotes the regularization balance parameter corresponding to the k-th factor; denotes the overlapping group sparse regularization term; the superscript T denotes the transpose of a matrix; vec(.) denotes the matrix vectorization operator; G denotes the combined value size; denotes the tensor corresponding to the image to be inpainted; D (k) denotes the cyclic difference matrix corresponding to the k-th factor; denotes the modulo 2 expansion form of the k-th factor; U (k) denotes the semi-orthogonal transformation result of the k-th factor; V (k) denotes the transformation matrix of the k-th factor; I denotes the identity matrix; S2: dividing the preliminary estimation image into small blocks, and stacking similar blocks into NSS groups; S3: constructing a group of (K+1) order NSS groups according to all the obtained non-local self-similar groups; S4: processing the (K+1) order NSS groups by FCTN decomposition to obtain reconstructed NSS groups; S5: gathering the reconstructed NSS groups to their original positions to obtain a filled image.
2. The color image reconstruction method of claim 1, characterized in that: In solving the TR decomposition reconstruction model, four factors of P, A, M and U (k) , V (k) and are alternately updated by PAM algorithm until the stop condition is reached and stopped.
3. The color image reconstruction method of claim 2, characterized in that: The stop condition is less than or equal to a set variation rate threshold, denotes the image obtained after the s-th iteration, denotes the image obtained after the s+1-th iteration.
4. A color image reconstruction terminal device, characterized by: The computer program is executed by the processor to implement the steps of the method in any one of claims 1-3.
5. A computer readable storage medium storing a computer program, characterized in that: The computer program is executed by the processor to implement the steps of the method in any one of claims 1-3.