Image reconstruction method and device, electronic equipment and storage medium
The high-frequency texture of the image is divided into multiple topics through the probabilistic latent semantic analysis topic learning model. Combining high-resolution and low-resolution atomic sets solves the problem of insufficient accuracy of high-frequency texture reconstruction and achieves detail recovery and structure preservation of high-resolution images.
Patent Information
- Application Number
- CN202510653583.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-19
AI Technical Summary
The reconstruction and restoration of high-frequency textures in the existing technology has the problem of insufficient accuracy, especially in the super-resolution reconstruction method based on sparse coding, which fails to effectively perform multi-dictionary reconstruction for different textures.
The probabilistic latent semantic analysis topic learning model is used to divide the high-frequency texture part of the training image into K topics, which are trained to form high-resolution and low-resolution atomic sets respectively. The low-resolution image is decomposed into high-frequency and low-frequency parts, and the high-frequency texture part is divided into overlapping documents and image blocks. The high-resolution image is reconstructed using a dictionary set associated with the topics.
The reconstruction accuracy of high-frequency textures is improved, the integrity and natural transition of image details are ensured, and the image quality is enhanced, so that high-resolution images retain the overall structure and enhance the details.
Smart Images

Figure CN120672574A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of communication technology, and specifically relates to an image reconstruction method, device, electronic device and storage medium. Background Art
[0002] Image resolution represents the density of information per unit size. High-resolution images contain more detailed features due to their rich pixel count. While high-resolution images can be obtained by improving hardware quality, sensor upgrades are expensive, and the actual acquisition process is susceptible to external interference, resulting in image resolution that cannot meet application requirements. Therefore, image processing technology is needed to address this issue.
[0003] Currently, this technology is mainly divided into three categories: interpolation-based, reconstruction-based, and learning-based methods. Interpolation-based methods introduce checkerboard and sawtooth effects due to the lack of prior knowledge; reconstruction-based methods are limited by fixed degradation models.
[0004] Although the existing super-resolution reconstruction method based on sparse coding can predict images using high- and low-resolution dictionaries, it does not perform multi-dictionary reconstruction for high-frequency images with different textures, resulting in insufficient accuracy in the reconstruction and restoration of high-frequency textures. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to provide an image reconstruction method, device, electronic device and storage medium, which can solve the problem of insufficient accuracy in the reconstruction and restoration of high-frequency textures.
[0006] In a first aspect, an embodiment of the present application provides an image reconstruction method, the method comprising:
[0007] Obtain a low-resolution image to be restored and a dictionary set; the dictionary set is a set of high-resolution and low-resolution atoms formed by dividing the high-frequency texture part of the training image into K topics through a probabilistic latent semantic analysis topic learning model, and then training each topic separately;
[0008] Decompose the low-resolution image to be restored into a high-frequency texture part and a low-frequency texture part;
[0009] Dividing the high-frequency texture portion into overlapping documents, and dividing each document into overlapping image blocks;
[0010] Use the probabilistic latent semantic analysis topic learning model to infer the topic to which the document belongs;
[0011] The image blocks are reconstructed using a dictionary set associated with the topic to which the document belongs, to obtain a high-resolution and high-frequency texture image;
[0012] A reconstructed high-resolution image is generated according to the high-resolution high-frequency texture image and the low-frequency texture part.
[0013] In a second aspect, an embodiment of the present application provides an image reconstruction device, the device comprising:
[0014] The acquisition module is used to obtain the low-resolution image to be restored and the dictionary set; the dictionary set is a set of high-resolution and low-resolution atoms formed by dividing the high-frequency texture part of the training image into K topics through the probabilistic latent semantic analysis topic learning model, and then training each topic separately;
[0015] A decomposition module, used for decomposing the low-resolution image to be restored into a high-frequency texture part and a low-frequency texture part;
[0016] A partitioning module, for partitioning the high-frequency texture portion into overlapping documents, and partitioning each document into overlapping image blocks;
[0017] The inference module is used to infer the topic of the document using the probabilistic latent semantic analysis topic learning model;
[0018] A reconstruction module is used to reconstruct the image block using a dictionary set associated with the topic to which the document belongs, to obtain a high-resolution and high-frequency texture image;
[0019] The generation module is used to generate a reconstructed high-resolution image according to the high-resolution high-frequency texture image and the low-frequency texture part.
[0020] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
[0021] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0022] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method described in the first aspect.
[0023] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the method described in the first aspect.
[0024] In an embodiment of the present application, a low-resolution image to be repaired and a dictionary set are obtained. The dictionary set is a high-resolution and low-resolution atomic set formed by training each theme separately after dividing the high-frequency texture part of the training image into K themes through a probabilistic latent semantic analysis topic learning model. The dictionary of topic classification can provide more accurate reconstruction rules for different texture features, decompose the low-resolution image to be repaired into a high-frequency texture part and a low-frequency texture part, and can process different frequency components in a targeted manner. The low-frequency part can be quickly amplified by simple interpolation, and the high-frequency part can be finely reconstructed through subsequent steps. The high-frequency texture part is divided into overlapping documents, and each document is divided into overlapping image blocks; by overlapping blocking, it is ensured that each image detail is The image blocks are completely covered, and the adjacent areas have information redundancy, which reduces the boundary discontinuity in subsequent reconstruction and makes the final image smoother and more natural. The probabilistic latent semantic analysis topic learning model is used to infer the topic of the document, thereby improving the accuracy of feature matching. The image blocks are reconstructed through a set of dictionaries associated with the topic to which the document belongs to obtain a high-resolution and high-frequency texture image. Since each dictionary focuses on a specific texture type, it can more accurately restore complex high-frequency details. According to the high-resolution and high-frequency texture image and the low-frequency texture part, a reconstructed high-resolution image is generated. The reconstructed high-resolution and high-frequency texture is combined with the interpolated low-frequency structure to generate a complete high-resolution image, which not only retains the overall structure of the image, but also enhances the details, significantly improving the image quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a flowchart of an image reconstruction method provided in an embodiment of the present application;
[0026] Figure 2 is a schematic diagram of an image reconstruction method provided in an embodiment of the present application;
[0027] Figure 3 is a structural diagram of an image reconstruction device provided in an embodiment of the present application;
[0028] Figure 4 It is a schematic diagram of the hardware structure of the electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the accompanying drawings of the embodiments of the present application to clearly describe the technical solutions of the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0030] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0031] In response to the problems arising in the related art, the embodiments of the present application provide an image reconstruction method, device, electronic device and storage medium, which can solve the problem of insufficient accuracy in the reconstruction and restoration of high-frequency textures in the related art.
[0032] The image reconstruction method provided in the embodiment of the present application is described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.
[0033] Figure 1 A flowchart of an image reconstruction method provided in an embodiment of the present application.
[0034] like Figure 1 As shown, the image reconstruction method may include steps 110 to 160, and the method is applied to an image reconstruction device, as shown below:
[0035] Step 110: Obtain a low-resolution image to be restored and a dictionary set. The dictionary set is a set of high-resolution and low-resolution atoms formed by dividing the high-frequency texture portion of the training image into K topics using a probabilistic latent semantic analysis topic learning model, and then training each topic separately to establish a mapping relationship between low-resolution image features and high-resolution image features.
[0036] Dictionary Pair: After classifying the high-frequency textures of training images by topic using probabilistic latent semantic analysis (pLSA), a pair of high-resolution and low-resolution atom sets is trained for each topic. The low-resolution dictionary captures local features, while the high-resolution dictionary stores the corresponding detail enhancement patterns. This establishes a nonlinear mapping from low-resolution blocks to high-resolution blocks, avoiding the limitations of a single dictionary in traditional sparse coding.
[0037] The dictionary set is an atomic set with specific semantics and feature expression capabilities formed by feature extraction, topic division and model training of training image data. Specifically, the dictionary set includes a high-resolution topic dictionary set and a low-resolution topic dictionary set:
[0038] High-resolution topic dictionary set: It consists of high-resolution topic dictionaries obtained by sparse representation training of high-resolution image blocks within each topic through the K-SVD algorithm. Each high-resolution topic dictionary is composed of multiple atoms that can represent high-frequency detail information such as specific texture patterns and edge features in high-resolution images, and is used to provide a feature expression basis for high-resolution reconstruction of image blocks.
[0039] Low-resolution topic dictionary set: The low-resolution image blocks within each topic are also trained and generated through sparse representation using the K-SVD algorithm. It contains multiple atoms that correspond to the corresponding texture, edge and other feature patterns in the low-resolution image and are used for feature matching with the low-resolution image blocks to be reconstructed.
[0040] The atoms in the dictionary set are the basic units for sparse representation of image blocks. The probabilistic latent semantic analysis topic learning model is used to mine topic relevance, enabling the dictionary set to classify image blocks with similar features at the semantic level, thereby providing support for establishing a mapping relationship between high-resolution and low-resolution features during the reconstruction process of low-resolution images to high-resolution images.
[0041] Probabilistic Latent Semantic Analysis (pLSA): An unsupervised topic model that assumes documents are composed of multiple topics, each of which corresponds to a distribution of words. It optimizes latent variables using the EM algorithm. It considers high-frequency textures as "documents" and image patches as "words," learning the latent topic distribution of texture structures, such as directional textures and periodic patterns.
[0042] High-frequency components are extracted from the training set images, which are then divided into K themes. For each theme, high- and low-resolution dictionaries are trained separately using KSVD or online dictionary learning to ensure inter-atom correlation. Thematic dictionaries improve the targeted nature of reconstruction.
[0043] Step 120, decomposing the low-resolution image to be restored into a high-frequency texture portion and a low-frequency texture portion;
[0044] Separating high-frequency and low-frequency textures: High-pass filtering is used to extract high-frequency components such as edges and noise; low-frequency components are obtained through Gaussian blurring or wavelet transforms, preserving information in smooth areas. High-frequency components dominate visual details, and separate processing prevents low-frequency information from interfering with reconstruction accuracy.
[0045] Step 130 , dividing the high-frequency texture portion into overlapping documents, and dividing each document into overlapping image blocks;
[0046] High-frequency images are divided into overlapping local regions, each of which is considered a "document" to preserve context. Documents are further divided into smaller overlapping blocks, which serve as the basic unit for sparse coding. Overlapping blocks suppress reconstruction artifacts, and multi-scale processing captures features at different granularities.
[0047] Step 140, using a probabilistic latent semantic analysis topic learning model to infer the topic to which the document belongs;
[0048] For all image patches in the current document, the joint probability of each topic is calculated. The dominant topic is determined by maximizing the likelihood function, and the most matching dictionary is adaptively selected to avoid the over-smoothing problem of the global dictionary.
[0049] Step 150 , reconstructing the image block using a dictionary set associated with the topic to which the document belongs, to obtain a high-resolution and high-frequency texture image;
[0050] Using a low-resolution dictionary corresponding to the selected topic, low frequencies provide basic structure and high frequencies provide details.
[0051] Step 160 : generating a reconstructed high-resolution image based on the high-resolution high-frequency texture image and the low-frequency texture portion.
[0052] In a possible embodiment, step 110 may specifically include the following steps:
[0053] Acquire multiple high-resolution training images, perform training image degradation processing on each high-resolution training image, and generate corresponding low-resolution training images;
[0054] Extracting high-frequency texture parts of each high-resolution training image and low-resolution training image respectively, forming the high-frequency texture parts of all high-resolution training images into a first training set, and forming the high-frequency texture parts of all low-resolution training images into a second training set;
[0055] The probabilistic latent semantic analysis topic learning model is used to perform topic mining on all documents in the first and second training sets, and K topics are obtained, where K is a positive integer.
[0056] Use K-SVD algorithm to train high-resolution topic dictionary and low-resolution topic dictionary for each topic, and obtain high-resolution topic dictionary set and low-resolution topic dictionary set;
[0057] For each atom in the low-resolution topic dictionary set, its nearest neighbors are found in the corresponding high-resolution topic dictionary set and low-resolution topic dictionary set, and a nearest neighbor regression matrix is calculated based on the nearest neighbors.
[0058] By simulating the degradation process of a real imaging system, high-resolution images are artificially downgraded to low-resolution images. For example, after blurring with a 7×7 Gaussian kernel, they are scaled down to the target size. This ensures that the high- and low-resolution image pairs are strictly aligned in terms of content, providing an accurate mapping foundation for subsequent dictionary training and avoiding reconstruction bias caused by data mismatch.
[0059] A high-pass filter is used to separate image details such as edges and noise. Low-frequency components are typically obtained through Gaussian smoothing or by retaining low-frequency wavelet coefficients. High-frequency components contain key information for super-resolution reconstruction, and separation can enhance texture details and avoid inefficient computation in smooth areas.
[0060] High-frequency regions in each image are considered "documents," and local image patches are considered "words." Using the Expectation Maximization (EM) algorithm, the algorithm iteratively learns two types of implicit relationships, estimating the probability that each document belongs to each topic. For example, a region with a 60% probability of belonging to an "edge topic" and a 40% probability of belonging to a "spot topic" is automatically discovered. This automatically discovers latent structural patterns in textures, such as directional edges and periodic textures, providing a basis for subsequent topic-based dictionary training. This approach is more robust than hard clustering.
[0061] The training of high-resolution and low-resolution image patches for each subject can be divided into two stages:
[0062] Sparse coding stage: The dictionary is fixed and the orthogonal matching pursuit (OMP) is used to solve the sparse representation of each image block.
[0063] Dictionary update stage: Atom-by-atom optimization is performed to update dictionary atoms through singular value decomposition (SVD) to minimize the reconstruction error.
[0064] A compact over-complete dictionary is obtained, where each atom captures a typical local feature, such as an edge segment in a specific direction, and the structural correspondence between high-resolution and low-resolution dictionaries is maintained.
[0065] For each atom in the low-resolution dictionary, the high-resolution dictionary is searched for the most similar atoms, such as those with the closest Euclidean distance. A linear mapping relationship between these neighbors is then fitted using the least-squares method. This local linear mapping from low-resolution features to high-resolution features allows for rapid recovery of detail during reconstruction, achieving greater accuracy than a global mapping.
[0066] Artificial degradation ensures pixel-level correspondence between high- and low-resolution image pairs, a prerequisite for dictionary effectiveness. Using unpaired data can lead to confusion in the mapping relationship. pLSA automatically classifies image patches into different themes based on texture features.
[0067] The probabilistic latent semantic analysis topic learning model (pLSA) is used to group documents and words with similar semantics into the same topic, and then a pair of high-resolution and low-resolution dictionaries are trained for each topic.
[0068] The specific calculation formula of the probabilistic latent semantic analysis topic learning model is:
[0069]
[0070] Among them, P(w i |d j ) represents the probability of the i-th word appearing in the j-th document; P(w i |z k ) represents the probability of the i-th word appearing under the k-th topic; P(z k |d j ) represents the probability of the kth topic appearing in the jth document;
[0071] Through the classic EM algorithm (maximum expectation algorithm), we can finally deduce:
[0072]
[0073] That is, the mixing coefficient P(z k |d j ), the maximum mixing coefficient can be used to determine the corresponding topic for the document, that is, Determine the subject of the document and finally convert P(w i |z k ) and P(z k ) is saved and used for estimation of test documents.
[0074] K-SVD optimizes two dictionaries separately while maintaining consistency in sparse coding between high- and low-resolution blocks. For example, an atom in the low-resolution dictionary corresponds to a "short horizontal line," while its high-resolution version corresponds to a "finer horizontal line with anti-aliasing."
[0075] The discreteness error of sparse coding is compensated by local linear mapping. For example, when a low-resolution block matches a "short horizontal line" atom during reconstruction, its high-resolution version is generated not only based on the corresponding atom, but also by combining the atoms of its neighbors, making the transition smoother.
[0076] In a possible embodiment, the step of performing topic mining on all documents in the first training set and the second training set using a probabilistic latent semantic analysis topic learning model to obtain K topics may specifically include the following steps:
[0077] Input the documents in the first training set and the second training set into the probabilistic latent semantic analysis topic learning model;
[0078] Through the probabilistic latent semantic analysis topic learning model, semantically similar documents and image blocks in the documents are classified into the same topic, and K topics are obtained.
[0079] Training set: A collection of image data used to train machine learning models, containing a large number of high-resolution images and their corresponding low-resolution images.
[0080] The first training set and the second training set may be used to distinguish training data of different types or sources, or to divide data. The specific understanding needs to be combined with the actual scenario.
[0081] In image restoration, a "document" is a structured partition of the high-frequency texture portion of an image. It's typically a large image region containing locally consistent texture patterns. As the fundamental unit of topic analysis, it facilitates the extraction of global semantic features from local textures.
[0082] Probabilistic Latent Semantic Analysis Topic Learning Model: An unsupervised learning model based on probabilistic statistics that aims to discover implicit "topic" structures in data. Here, "topic" refers to a collection of documents and image patches with similar semantic or visual features.
[0083] Assuming that each document is generated by a mixture of multiple potential topics, and each image patch corresponds to a typical feature under a certain topic, the model infers the potential topic structure by analyzing the distribution pattern of image patches in the document.
[0084] Specifically, the training set documents are fed into the topic learning model. The high-frequency texture components of all images in the first and second training sets are extracted, for example by filtering to remove low-frequency structures. The high-frequency texture components are then divided into overlapping documents, each containing several image patches. Each document can be considered a "text," and the image patches within it are considered "words." A "document-image patch" association matrix is then constructed.
[0085] All documents and their image patch features are fed into a probabilistic latent semantic analysis model as raw data for topic mining. The model assumes that each document is composed of a mixture of K potential topics with a certain probability, and each image patch is generated by one of these topics.
[0086] The model parameters are iteratively optimized using the Expectation Maximization (EM) algorithm to calculate the following probabilities:
[0087] The probability of each document belonging to each topic, such as document A has a 70% probability of belonging to topic 1 and a 30% probability of belonging to topic 2.
[0088] The probability of each topic generating various types of image patches. For example, in topic 1, 60% of the image patches are “slender edges” and 30% are “green spots”.
[0089] Based on the probabilistic association between documents and topics, the topic with the highest probability is taken as the "dominant topic" of the document, so that semantically similar documents are classified into the same topic, and the image blocks in these documents are also included in the feature set of the topic.
[0090] The documents and image patches in the training set are divided into K topics with clear semantics, each of which corresponds to a class of visual patterns. Image patches under the same topic have similar high-frequency texture features, while the features of different topics vary significantly.
[0091] Each theme corresponds to a separate dictionary set, which contains mapping pairs between low-resolution image patches and high-resolution image patches. Because image patches within the same theme have similar semantics, the mapping relationships in the dictionary are more targeted.
[0092] During image restoration, by inferring the subject of the document to be restored, the dictionary of that subject can be directly called for matching, reducing the search scope and improving the matching accuracy, thereby generating more semantically accurate high-resolution details. For example, grass texture will not be mistakenly matched to the dictionary of a brick wall.
[0093] Compressing the features of massive image patches into K topics reduces the computational effort required for subsequent restoration. Instead of searching the entire dictionary, searches can be performed within topic-specific dictionaries. Topics, as "semantic units," can cover new image patches that are not fully represented in the training set but have similar semantics.
[0094] Unlike clustering based solely on pixel similarity, the probabilistic latent semantic analysis model clusters by mining "semantic relevance," making topics more meaningful and the restoration results more consistent with human visual cognition. It allows documents to belong to a mixture of multiple topics, rather than rigidly categorizing them, more closely resembling the complexity of textures in real images. Removing the need for manual texture feature design, the model automatically learns topic-related feature representations from the data, making it applicable to complex and diverse image texture types.
[0095] In a possible embodiment, for each atom in the low-resolution topic dictionary set, the step of finding its nearest neighbor in the corresponding high-resolution topic dictionary set and low-resolution topic dictionary set may specifically include the following steps:
[0096] For atoms in the high-resolution topic dictionary set and the low-resolution topic dictionary set, a similarity matrix is calculated based on Euclidean distance or cosine similarity;
[0097] For each atom in the low-resolution topic dictionary, the top N high-resolution atoms and the top M low-resolution atoms with the highest similarity are selected from the similarity matrix to form the nearest neighbors of the atom;
[0098] The unique high-resolution atom with the highest similarity is selected from the nearest neighbors, and a nearest neighbor mapping relationship is established between the low-resolution atom and the selected high-resolution atom.
[0099] Atom: The basic unit in the dictionary, representing the feature vector of the image block, and achieving low-resolution to high-resolution mapping through matching between atoms.
[0100] Low-resolution atoms: features extracted from low-resolution image patches, such as vectors of pixel values.
[0101] High-resolution atoms: High-resolution image patch features corresponding to low-resolution atoms.
[0102] Euclidean distance: Suitable for comparing the absolute numerical differences of vectors and for direct comparison of pixel values.
[0103] Cosine similarity: Suitable for comparing relative patterns of vectors, such as texture direction and frequency, while ignoring overall brightness differences.
[0104] For all low-resolution atoms and high-resolution atoms under the same theme, the similarity between them is calculated. For each low-resolution atom, they are sorted from high to low according to the similarity, and the first N high-resolution atoms and the first M low-resolution atoms are selected to form their "nearest neighbors".
[0105] High-resolution neighborhood: Contains the high-resolution candidate atoms that are most similar to the current low-resolution atom.
[0106] Low-resolution neighborhood: contains similar low-resolution atoms under the same theme, which can be used to assist in judging the consistency of texture patterns.
[0107] Nearest neighbors: This refers to the set of atoms in both the high-resolution and low-resolution topic dictionary sets that are closest to a specific atom in the low-resolution topic dictionary. By calculating the similarity between atoms, the atoms with the highest similarity are selected, forming the "nearest neighbors" of the low-resolution atom. This defines the range within which to search for the corresponding high-resolution atom and is essentially a candidate set of atoms sorted by similarity.
[0108] Nearest neighbor mapping: This is a one-to-one correspondence between low-resolution and high-resolution atoms, established based on the "nearest neighbor" approach. After determining the nearest neighbor of a low-resolution atom, the high-resolution atom with the highest similarity to the low-resolution atom is selected from that neighborhood, and a mapping is established between the two. This mapping is used for subsequent image patch reconstruction, identifying high-resolution features corresponding to low-resolution image patch features through this mapping, and then using a regression matrix to achieve the conversion from low-resolution to high-resolution.
[0109] Select the most similar atom from the high-resolution neighborhood as the mapping target for the current low-resolution atom. This direct selection of the most similar atom establishes a one-to-one hard mapping relationship, simplifying the computational logic during the repair process.
[0110] By using numerical distance or similarity metrics, the visual similarity of image patches is converted into a computable metric, eliminating subjective judgment. Multiple candidate atoms are considered to reduce mismatches caused by noise or feature deviations. For example, if a low-resolution atom accidentally resembles a high-resolution atom due to local noise, other atoms in the neighborhood can assist in verifying its authenticity.
[0111] Using a trained dictionary to directly find matching high-resolution atoms avoids complex optimization processes and is suitable for real-time or large-scale image restoration. Atoms under the same theme share similar semantics, and the mapped high-resolution details can maintain the texture style of the input image.
[0112] In a possible embodiment, step 130 may specifically include the following steps:
[0113] The high-frequency texture part is divided into overlapping regions of size m×m, and each region is regarded as a document;
[0114] Each document is divided into overlapping image blocks of size n×n; m and n are both positive integers.
[0115] High-frequency texture: This portion of an image contains rapidly changing information such as details, edges, texture, and noise, typically extracted through filtering techniques. It is the core component of image restoration, requiring algorithms to reconstruct its high-resolution details.
[0116] Document: The first region division of the high-frequency texture part is a large image region with a fixed size. The texture within each document has local consistency, which facilitates subsequent topic analysis and feature matching.
[0117] Image block: The second subdivision of the document is a smaller image unit, such as the basic texture unit that makes up the document. It serves as the basic unit for dictionary matching and reconstruction, and restores the overall high-frequency texture by repairing each small block.
[0118] A preset square window is slid across the high-frequency textures, with each sliding step smaller than the window side length, to ensure overlap between adjacent documents. This avoids discontinuous textures at the edges of documents due to hard segmentation and ensures consistency between textures across adjacent documents during topic analysis. Textures within the same document are assumed to belong to the same or similar themes, facilitating subsequent clustering using topic models.
[0119] Within each document, a pre-set square window is used for further sliding partitioning, again using a step size smaller than the window side length to allow for overlap between image patches. Smaller patch sizes facilitate capturing finer texture features, while overlapping patches ensure that each pixel is covered by multiple patches, avoiding blocky boundaries after restoration. Each patch is fed into the dictionary matching process as a separate feature vector, whose features are used to find the most similar high-resolution corresponding patch in the dictionary.
[0120] High-frequency textures are divided into locally consistent regions, allowing subsequent topic models to quickly identify the dominant texture type in each region. Analysis of the entire image is transformed into analysis of multiple local documents, reducing computational effort while improving the accuracy of topic inference. Small image patches capture more specific texture patterns, enabling more accurate dictionary matching and reconstructing high-resolution details that are closer to the real texture.
[0121] Since each pixel belongs to multiple overlapping image blocks, the restoration results of different blocks are fused by weighted averaging during reconstruction, which can effectively eliminate the discontinuity of the block boundaries and make the restored texture transition natural.
[0122] Larger document sizes are suitable for capturing a wide range of texture patterns, but may contain a mixture of multiple topics and rely on the soft clustering ability of the topic model. Smaller document sizes are more likely to ensure the purity of a single topic, but may lose the context of texture associations.
[0123] Larger image patch sizes are suitable for reconstructing textures with global structures, but may contain too much detail, resulting in high feature vector dimensions and low matching efficiency. Smaller image patch sizes are suitable for capturing subtle textures, but may lead to matching ambiguity due to insufficient information.
[0124] Through a two-layer "document-image block" division, a progressive analysis from regional semantics to local features is achieved, ensuring both overall consistency of texture themes and precise matching of detailed features. The design of overlapping areas is a key method for eliminating artifacts from block-based restoration. By overlapping features of adjacent blocks and fusing the results, the restored image is visually continuous and uninterrupted. The size of the document and image blocks can be adjusted based on the specific image content to accommodate different restoration needs.
[0125] In a possible embodiment, step 150 may specifically include the following steps:
[0126] For each image patch, find the most relevant atom in the low-resolution topic dictionary set corresponding to the document topic;
[0127] The image blocks are reconstructed using the nearest neighbor regression matrix corresponding to the atoms with the highest correlation, and the reconstructed image blocks are overlapped and combined to obtain a high-resolution and high-frequency texture image.
[0128] Most relevant atom: The dictionary atom in the low-resolution topic dictionary that is most similar to the current image patch's features. The similarity between the image patch and all atoms in the dictionary is calculated using metrics such as Euclidean distance and cosine similarity. The atom with the highest similarity is the most relevant atom.
[0129] Nearest Neighbor Regression Matrix: A pre-trained mapping matrix that establishes relationships between low-resolution and high-resolution atoms. Using machine learning or statistical methods, feature transformation rules learned from a large number of low-resolution and high-resolution atom pairs are used to "translate" low-resolution features into high-resolution features. Once the nearest neighbors of a low-resolution atom are found, the corresponding high-resolution atoms are directly calculated using this matrix, achieving low-to-high-resolution reconstruction.
[0130] Overlapping: Overlapping high-resolution image blocks are weighted and fused to eliminate discontinuities at the block boundaries, generating a complete high-frequency texture image. The same pixel is reconstructed multiple times in multiple overlapping blocks, and the results of each block are combined through weighted distribution to ensure a smooth transition of pixel values.
[0131] Expand the pixel values of the current image block into a one-dimensional vector. Iterate through all atoms in the low-resolution dictionary corresponding to the document's topic, calculate the similarity between the image block vector and each atom, and select the atoms with the highest similarity. Search only within the dictionary for the document's topic to avoid semantic confusion caused by cross-topic matching.
[0132] During the training phase, each low-resolution atom is associated with a high-resolution atom, represented by a regression matrix. For each low-resolution atom found, the high-resolution atom's eigenvector is calculated using its corresponding regression matrix, and this eigenvector is then converted back to the pixel matrix of the image block. Because the image blocks overlap, each pixel is covered by multiple reconstructed blocks. A weight is assigned to each pixel in the image block, and the pixel values in the overlapping areas are summed according to the weights of each block to obtain the final pixel value. This prevents sudden changes in brightness or texture at the block boundaries, ensuring visual continuity of the restored high-frequency texture.
[0133] Searching for atoms only in a dictionary of subject-specific features ensures that the reconstructed high-resolution details are consistent with the texture type of the input image. This narrowing of the dictionary reduces computational effort, particularly in scenarios with a large number of subjects, K. The regression matrix can learn complex mappings between low- and high-resolution features, such as nonlinear enhancement from blurred edges to sharp edges, providing greater flexibility than simple nearest-neighbor atom replacement. Even if image patches do not perfectly match dictionary atoms, the regression matrix can generate reasonable high-resolution details based on the statistical patterns of the training data.
[0134] Weighted fusion creates a smooth transition between pixel values at the boundaries of adjacent blocks, making block separation imperceptible to the naked eye and improving the visual quality of the restored image. Pixels in overlapping areas are synthesized from the reconstruction results of multiple blocks, suppressing outliers caused by noise or mismatches in a single block while preserving true detail.
[0135] Combined with document topic analysis, this approach enables semantic understanding of inpainting, avoiding the texture distortion common in traditional methods. Regression matrices can be pre-trained and stored, requiring only matrix multiplication during inpainting. This significantly outperforms end-to-end deep learning-based models, making it suitable for real-time applications. Pixels in overlapping areas are determined by multiple blocks, minimizing the impact of a single block mismatch on the overall result.
[0136] In a possible embodiment, step 160 may specifically include the following steps:
[0137] Interpolate and reconstruct the low-frequency texture part to obtain a high-resolution low-frequency texture image;
[0138] The high-resolution high-frequency texture image and the high-resolution low-frequency texture image are combined to obtain a reconstructed high-resolution image.
[0139] Low-frequency texture: This refers to the slowly changing areas of an image, encompassing overall structure, outlines, and large-scale color blocks, but not detailed textures or edges. Pixel values vary smoothly in space, corresponding to low-frequency components in the frequency domain, and forming the image's "skeleton" structure.
[0140] Interpolation reconstruction: The process of increasing image resolution by inserting new pixels between pixels of a low-resolution image using mathematical methods.
[0141] Interpolation reconstruction can specifically include:
[0142] Nearest neighbor interpolation: directly copy the adjacent pixel values, which is fast but prone to aliasing.
[0143] Bilinear interpolation: New pixels are calculated by weighted average of 4 adjacent pixels. The result is smoother but the details are blurred.
[0144] Bicubic interpolation: uses weighted polynomial fitting of a larger neighborhood of pixels to balance smoothing and detail preservation.
[0145] High-Low Frequency Combination: The reconstructed high-resolution high-frequency texture is superimposed on the high-resolution low-frequency structure to restore the complete high-resolution image. In the frequency domain, low frequencies correspond to structure and high frequencies correspond to details, and the combination of the two constitutes the complete image frequency components.
[0146] The low-frequency portion has already been decomposed to remove high-frequency details, retaining only structural information. Its pixel values vary smoothly, making it suitable for direct amplification using interpolation methods. A computationally simple interpolation method is selected to upsample the low-frequency image to generate high-resolution low-frequency structure. The low-frequency structure does not rely on training data; its amplified pixel values can be estimated through smooth transitions between adjacent pixels. The high-resolution low-frequency image represents the overall structure of the image, while the high-resolution high-frequency image represents detailed texture.
[0147] High-resolution image = high-resolution low-frequency structure + high-resolution high-frequency texture;
[0148] Direct addition at the pixel level superimposes high-frequency details on the outline of low-frequency structures, forming a complete image with both clear structure and rich details.
[0149] The interpolated low-frequency image preserves the overall layout and outline of the original image, avoiding structural distortion. The low-complexity interpolation algorithm quickly generates high-resolution structures, providing the foundation for subsequent overlay of high-frequency details. The low-frequency portion provides clear large-scale structure, while the high-frequency portion complements delicate textures, resulting in an image that combines both "macro-contours" and "micro-details."
[0150] Restoring only high frequencies without addressing low frequencies can result in blurred structures. Using interpolation alone to amplify the entire image can result in lost high-frequency details due to the smoothing effect of interpolation. The fused image not only matches the human eye's perception of object structure but also meets the demand for detail clarity, approaching the look and feel of a true high-resolution image.
[0151] Low-frequency and high-frequency images are processed separately. Low-frequency images are rapidly reconstructed using interpolation, while high-frequency images rely on data-driven dictionary matching to restore details, achieving a balance between efficiency and accuracy. The low-frequency interpolation method can be combined with any high-frequency restoration algorithm, providing a flexible combination of techniques.
[0152] In an embodiment of the present application, a low-resolution image to be repaired and a dictionary set are obtained. The dictionary set is a high-resolution and low-resolution atomic set formed by training each theme separately after dividing the high-frequency texture part of the training image into K themes through a probabilistic latent semantic analysis topic learning model. The dictionary is used to establish a mapping relationship between low-resolution image features and high-resolution image features. The dictionary of topic classification can provide more accurate reconstruction rules for different texture features, decompose the low-resolution image to be repaired into a high-frequency texture part and a low-frequency texture part, and can process different frequency components in a targeted manner. The low-frequency part can be quickly amplified by simple interpolation, and the high-frequency part can be finely reconstructed through subsequent steps. The high-frequency texture part is divided into overlapping documents, and each document is divided into overlapping image blocks; by Overlapping blocks ensure that each image detail is fully covered and that adjacent areas have information redundancy, reducing boundary discontinuities in subsequent reconstruction and making the final image smoother and more natural. The probabilistic latent semantic analysis topic learning model is used to infer the topic of the document, thereby improving the accuracy of feature matching. The image blocks are reconstructed using a set of dictionaries associated with the topic to which the document belongs to obtain a high-resolution, high-frequency texture image. Since each dictionary focuses on a specific texture type, it can more accurately restore complex high-frequency details. Based on the high-resolution, high-frequency texture image and the low-frequency texture part, a reconstructed high-resolution image is generated. The reconstructed high-resolution, high-frequency texture is combined with the interpolated low-frequency structure to generate a complete high-resolution image, which not only retains the overall structure of the image but also enhances the details, significantly improving the image quality.
[0153] The following combination Figure 2 Describe the embodiments of the present application;
[0154] Step 1: Obtain multiple high-resolution images, perform image degradation on each high-resolution image to generate a low-resolution image, and extract the high-frequency texture parts of each high-resolution image and low-resolution image respectively. The high-frequency texture parts extracted from all high-resolution images constitute the first training set, and the high-frequency texture parts extracted from all low-resolution images constitute the second training set.
[0155] In the embodiment of the present application, a bilateral filter is used to extract high-frequency texture parts of high-resolution images and low-resolution images;
[0156] Step 2: Divide each training sample in the first training set and the second training set into m×m overlapping documents, and then divide the documents into n×n overlapping image blocks;
[0157] Each training sample of the first training set and the second training set is divided into m overlapping local regions, and each local region is regarded as a document. Then the document set consisting of m documents is expressed as D = (d1, d2, ..., d m); There are also word concepts in the document, continue to decompose m documents into n overlapping words, and form a word set W = (w1, w2, ..., w n ), where words are so-called image patches;
[0158] Step 3: Use the probabilistic latent semantic analysis topic learning model to perform topic mining on all documents to obtain K topics, and use the K-SVD algorithm to train the dictionary for each topic to obtain a high-resolution topic dictionary set and a low-resolution topic dictionary set respectively; in the embodiment of this application, the high-resolution topic dictionary set is recorded as Let the low-resolution theme dictionary set be
[0159] Assume that there are K potential topics in the document and word set in step 2, and topic z∈Z=(z1,z2,…,z K ) is a higher-level concept, that is, a topic is a collection of different documents grouped according to the co-occurrence semantic relationship of "words" within and between documents;
[0160] The specific steps for performing topic mining on all documents in the embodiment of this application are:
[0161] Documents with similar semantics and words in the documents are grouped into the same topic through a probabilistic latent semantic analysis topic learning model (pLSA), and then a pair of high-resolution and low-resolution dictionaries are trained for each topic. The probabilistic latent semantic analysis topic learning model (pLSA) in the embodiments of the present application is prior art and will not be described in detail here.
[0162] The specific calculation formula of the probabilistic latent semantic analysis topic learning model is:
[0163]
[0164] Among them, P(w i |d j ) represents the probability of the i-th word appearing in the j-th document; P(w i |z k ) represents the probability of the i-th word appearing under the k-th topic; P(z k |d j ) represents the probability of the kth topic appearing in the jth document;
[0165] Through the classic EM algorithm (maximum expectation algorithm), we can finally deduce:
[0166]
[0167] That is, the mixing coefficient P(z k |d j), the maximum mixing coefficient can be used to determine the corresponding topic for the document, that is, Determine the subject of the document and finally convert P(w i |z k ) and P(z k ) is saved and used for estimation of test documents;
[0168] In the embodiment of the present application, the specific steps of using the K-SVD algorithm to train the dictionary for each topic are:
[0169] Use the K-SVD algorithm and OMP algorithm to solve the following formula for x:
[0170]
[0171] st||α s ||0≤Q
[0172] Among them, Q represents sparsity, is a low-resolution image patch training set An image block in s yes The corresponding sparse representation coefficient;
[0173] The sparse representation theory states that high-resolution image patches have the same coefficient representation coefficients as the corresponding low-resolution image patches. Therefore, the high-resolution topic dictionary set The learning process is as follows:
[0174]
[0175] in, Represents the high-resolution image patch training set An image block in The calculation formula is solved using the pseudo-inverse of the matrix, as shown below:
[0176]
[0177] Thus, we can obtain high and low resolution dictionary pairs for each topic.
[0178] Step 4: For each atom in each low-resolution topic dictionary set, find the nearest neighbors of the atoms in the corresponding high-resolution topic dictionary set and the low-resolution topic dictionary set, and calculate the nearest neighbor regression matrix based on the nearest neighbors of the atoms in the high-resolution topic dictionary set and the low-resolution topic dictionary set;
[0179] The specific steps for calculating the nearest neighbor regression matrix in the embodiment of the present application are:
[0180] Assume that the input low-resolution image block y is judged to correspond to the topic z after the probabilistic latent semantic analysis topic learning model (pLSA) k , in the corresponding low-resolution theme dictionary The atom with the highest correlation with y is found in Then the solution process of the nearest neighbor regression matrix is as follows:
[0181]
[0182] Where β is the value of y in Low-resolution nearest neighbors corresponding to atoms The coefficient vector under , λ is the weight coefficient, which is used to eliminate singular problems and ensure the stability of the solution. When λ>0, the nearest neighbor regression algorithm is used to solve the calculation formula of β above, and the approximate solution is as follows:
[0183]
[0184] The sparse representation theory believes that the low-resolution image block y and its corresponding high-resolution image block x have the same coefficient vector β, so the high-resolution image block x can be represented by Corresponding high-resolution dictionary atomic neighborhood Combined with β, it is shown as follows:
[0185]
[0186] Through the calculation formula of the above β approximate solution and the calculation formula of x, it is found that the dictionary atom The nearest neighbor regression matrix can be expressed as follows:
[0187]
[0188] As can be seen from the above formula, as long as the subject dictionary pairs are trained in advance, the nearest neighbor sets of high and low resolution dictionary atoms corresponding to all low-resolution subject dictionary atoms can be found according to the correlation of dictionary atoms.
[0189] The nearest neighbor regression matrix corresponding to all atoms can be calculated in advance;
[0190] Step 5: Input the low-resolution image to be reconstructed, and decompose the low-resolution image to be reconstructed into a high-frequency texture part and a low-frequency texture part. In the same manner as in step 2, the high-frequency texture part of the low-resolution image to be reconstructed is decomposed into a document, and then the document is divided into image blocks.
[0191] Step 6: Use the probabilistic latent semantic analysis topic learning model to infer the topic of the document in step 5, and reconstruct the image blocks in the document to obtain a reconstructed high-resolution image;
[0192] The specific process of reconstruction includes:
[0193] For each image block in the document, find the atom with the highest correlation in the low-resolution topic dictionary set, reconstruct it using the nearest neighbor regression matrix corresponding to the atom with the highest correlation, and overlap and combine all the reconstructed high-resolution high-frequency texture image blocks to obtain a high-resolution high-frequency texture image;
[0194] The low-frequency texture portion decomposed from the low-resolution image to be reconstructed is interpolated and reconstructed to obtain a high-resolution low-frequency texture image; in the embodiment of the present application, a bicubic interpolation method is used to interpolate and reconstruct the low-frequency texture portion decomposed from the low-resolution image to be reconstructed;
[0195] The high-resolution high-frequency texture image and the high-resolution low-frequency texture image are combined to obtain a reconstructed high-resolution image.
[0196] In an embodiment of the present application, a low-resolution image to be repaired and a dictionary set are obtained. The dictionary set is a high-resolution and low-resolution atomic set formed by training each theme separately after dividing the high-frequency texture part of the training image into K themes through a probabilistic latent semantic analysis topic learning model. The dictionary is used to establish a mapping relationship between low-resolution image features and high-resolution image features. The dictionary of topic classification can provide more accurate reconstruction rules for different texture features, decompose the low-resolution image to be repaired into a high-frequency texture part and a low-frequency texture part, and can process different frequency components in a targeted manner. The low-frequency part can be quickly amplified by simple interpolation, and the high-frequency part can be finely reconstructed through subsequent steps. The high-frequency texture part is divided into overlapping documents, and each document is divided into overlapping image blocks; by Overlapping blocks ensure that each image detail is fully covered and that adjacent areas have information redundancy, reducing boundary discontinuities in subsequent reconstruction and making the final image smoother and more natural. The probabilistic latent semantic analysis topic learning model is used to infer the topic of the document, thereby improving the accuracy of feature matching. The image blocks are reconstructed using a set of dictionaries associated with the topic to which the document belongs to obtain a high-resolution, high-frequency texture image. Since each dictionary focuses on a specific texture type, it can more accurately restore complex high-frequency details. Based on the high-resolution, high-frequency texture image and the low-frequency texture part, a reconstructed high-resolution image is generated. The reconstructed high-resolution, high-frequency texture is combined with the interpolated low-frequency structure to generate a complete high-resolution image, which not only retains the overall structure of the image but also enhances the details, significantly improving the image quality.
[0197] The image reconstruction method provided in the embodiment of the present application can be executed by an image reconstruction device. In the embodiment of the present application, the image reconstruction device provided in the embodiment of the present application is described by taking the image reconstruction method executed by the image reconstruction device as an example.
[0198] Figure 3 is a block diagram of an image reconstruction device provided in an embodiment of the present application, the device 300 includes:
[0199] Acquisition module 310 is used to acquire a low-resolution image to be restored and a dictionary set; the dictionary set is a set of high-resolution and low-resolution atoms formed by dividing the high-frequency texture portion of the training image into K topics using a probabilistic latent semantic analysis topic learning model, and then training each topic separately to establish a mapping relationship between low-resolution image features and high-resolution image features;
[0200] a decomposition module 320 for decomposing the low-resolution image to be restored into a high-frequency texture portion and a low-frequency texture portion;
[0201] a division module 330 for dividing the high-frequency texture portion into overlapping documents, and dividing each document into overlapping image blocks;
[0202] An inference module 330 is configured to infer the topic of a document using a probabilistic latent semantic analysis topic learning model;
[0203] A reconstruction module 350 is configured to reconstruct the image block using a dictionary set associated with the topic to which the document belongs, to obtain a high-resolution and high-frequency texture image;
[0204] The generating module 360 is configured to generate a reconstructed high-resolution image based on the high-resolution high-frequency texture image and the low-frequency texture portion.
[0205] In a possible embodiment, the reconstruction module 350 is specifically configured to:
[0206] For each image patch, find the most relevant atom in the low-resolution topic dictionary set corresponding to the document topic;
[0207] The image blocks are reconstructed using the nearest neighbor regression matrix corresponding to the atoms with the highest correlation, and the reconstructed image blocks are overlapped and combined to obtain a high-resolution and high-frequency texture image.
[0208] In a possible embodiment, the generating module 360 is specifically configured to:
[0209] Interpolate and reconstruct the low-frequency texture part to obtain a high-resolution low-frequency texture image;
[0210] The high-resolution high-frequency texture image and the high-resolution low-frequency texture image are combined to obtain a reconstructed high-resolution image.
[0211] In a possible embodiment, the partitioning module 330 is specifically configured to:
[0212] The high-frequency texture part is divided into overlapping regions of size m×m, and each region is regarded as a document;
[0213] Each document is divided into overlapping image blocks of size n×n; m and n are both positive integers.
[0214] In a possible embodiment, the acquisition module 310 is specifically configured to:
[0215] Acquire multiple high-resolution training images, perform training image degradation processing on each high-resolution training image, and generate corresponding low-resolution training images;
[0216] Extracting high-frequency texture parts of each high-resolution training image and low-resolution training image respectively, forming the high-frequency texture parts of all high-resolution training images into a first training set, and forming the high-frequency texture parts of all low-resolution training images into a second training set;
[0217] The probabilistic latent semantic analysis topic learning model is used to perform topic mining on all documents in the first and second training sets, and K topics are obtained, where K is a positive integer.
[0218] Use K-SVD algorithm to train high-resolution topic dictionary and low-resolution topic dictionary for each topic, and obtain high-resolution topic dictionary set and low-resolution topic dictionary set;
[0219] For each atom in the low-resolution topic dictionary set, its nearest neighbors are found in the corresponding high-resolution topic dictionary set and low-resolution topic dictionary set, and a nearest neighbor regression matrix is calculated based on the nearest neighbors.
[0220] In a possible embodiment, the acquisition module 310 is specifically configured to:
[0221] Input the documents in the first training set and the second training set into the probabilistic latent semantic analysis topic learning model;
[0222] Through the probabilistic latent semantic analysis topic learning model, semantically similar documents and image blocks in the documents are classified into the same topic, and K topics are obtained.
[0223] In a possible embodiment, the acquisition module 310 is specifically configured to:
[0224] For atoms in the high-resolution topic dictionary set and the low-resolution topic dictionary set, a similarity matrix is calculated based on Euclidean distance or cosine similarity;
[0225] For each atom in the low-resolution topic dictionary, the top N high-resolution atoms and the top M low-resolution atoms with the highest similarity are selected from the similarity matrix to form the nearest neighbors of the atom;
[0226] The unique high-resolution atom with the highest similarity is selected from the nearest neighbors, and a nearest neighbor mapping relationship is established between the low-resolution atom and the selected high-resolution atom.
[0227] In an embodiment of the present application, a low-resolution image to be repaired and a dictionary set are obtained. The dictionary set is a high-resolution and low-resolution atomic set formed by training each theme separately after dividing the high-frequency texture part of the training image into K themes through a probabilistic latent semantic analysis topic learning model. The dictionary is used to establish a mapping relationship between low-resolution image features and high-resolution image features. The dictionary of topic classification can provide more accurate reconstruction rules for different texture features, decompose the low-resolution image to be repaired into a high-frequency texture part and a low-frequency texture part, and can process different frequency components in a targeted manner. The low-frequency part can be quickly amplified by simple interpolation, and the high-frequency part can be finely reconstructed through subsequent steps. The high-frequency texture part is divided into overlapping documents, and each document is divided into overlapping image blocks; by Overlapping blocks ensure that each image detail is fully covered and that adjacent areas have information redundancy, reducing boundary discontinuities in subsequent reconstruction and making the final image smoother and more natural. The probabilistic latent semantic analysis topic learning model is used to infer the topic of the document, thereby improving the accuracy of feature matching. The image blocks are reconstructed using a set of dictionaries associated with the topic to which the document belongs to obtain a high-resolution, high-frequency texture image. Since each dictionary focuses on a specific texture type, it can more accurately restore complex high-frequency details. Based on the high-resolution, high-frequency texture image and the low-frequency texture part, a reconstructed high-resolution image is generated. The reconstructed high-resolution, high-frequency texture is combined with the interpolated low-frequency structure to generate a complete high-resolution image, which not only retains the overall structure of the image but also enhances the details, significantly improving the image quality.
[0228] The image reconstruction device provided in the embodiment of the present application can implement each process implemented in the above method embodiment. To avoid repetition, it will not be described here.
[0229] Figure 4 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.
[0230] The electronic device may include a processor 401 and a memory 402 storing computer program instructions.
[0231] Specifically, the processor 401 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0232] The memory 402 may include a large-capacity memory for data or instructions. By way of example and not limitation, the memory 402 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 402 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 402 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 402 is a non-volatile solid-state memory. In a specific embodiment, the memory 402 includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.
[0233] The processor 401 reads and executes computer program instructions stored in the memory 402 to implement any one of the image reconstruction methods in the embodiments shown in the figures.
[0234] In one example, the electronic device may further include a communication interface 404 and a bus 410. Figure 4 As shown, the processor 401 , the memory 402 , and the communication interface 404 are connected via a bus 410 and communicate with each other.
[0235] The communication interface 404 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0236] Bus 410 comprises hardware, software or both, couples the parts of electronic equipment to each other.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations.In suitable cases, bus 410 can comprise one or more buses.Although the present application embodiment describes and shows specific bus, the application considers any suitable bus or interconnection.
[0237] The electronic device can execute the image reconstruction method in the embodiment of the present application, thereby realizing the combination Figure 2 Describe the image reconstruction method.
[0238] In addition, in combination with the image reconstruction method in the above embodiment, the embodiment of the present application can provide a computer readable storage medium for implementation. The computer readable storage medium stores computer program instructions; when the computer program instructions are executed by the processor, the Figure 1 image reconstruction method.
[0239] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.
[0240] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0241] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0242] The above description is only a specific embodiment of the present application. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the scope of protection of the present application.
Claims
1. An image reconstruction method, characterized in that: The method comprises: Obtain a low-resolution image to be restored and a dictionary set; the dictionary set is a set of high-resolution and low-resolution atoms formed by dividing the high-frequency texture part of the training image into K topics through a probabilistic latent semantic analysis topic learning model, and then training each topic separately; Decompose the low-resolution image to be restored into a high-frequency texture part and a low-frequency texture part; Dividing the high-frequency texture portion into overlapping documents, and dividing each document into overlapping image blocks; Use the probabilistic latent semantic analysis topic learning model to infer the topic to which the document belongs; The image blocks are reconstructed using a dictionary set associated with the topic to which the document belongs, to obtain a high-resolution and high-frequency texture image; A reconstructed high-resolution image is generated according to the high-resolution high-frequency texture image and the low-frequency texture part.
2. The method according to claim 1, characterized in that The image block is reconstructed by using a dictionary set associated with the topic to which the document belongs to obtain a high-resolution high-frequency texture image, including: For each image patch, find the most relevant atom in the low-resolution topic dictionary set corresponding to the document topic; The image blocks are reconstructed using the nearest neighbor regression matrix corresponding to the atoms with the highest correlation, and the reconstructed image blocks are overlapped and combined to obtain a high-resolution and high-frequency texture image.
3. The method according to claim 1, characterized in that The step of generating a reconstructed high-resolution image based on the high-resolution high-frequency texture image and the low-frequency texture portion includes: Interpolate and reconstruct the low-frequency texture part to obtain a high-resolution low-frequency texture image; The high-resolution high-frequency texture image and the high-resolution low-frequency texture image are combined to obtain a reconstructed high-resolution image.
4. The method according to claim 1, wherein The step of dividing the high-frequency texture portion into overlapping documents and dividing each document into overlapping image blocks comprises: The high-frequency texture part is divided into overlapping regions of size m×m, and each region is regarded as a document; Each document is divided into overlapping image blocks of size n×n; m and n are both positive integers.
5. The method according to claim 1, wherein The obtaining of the dictionary set includes: Acquire multiple high-resolution training images, perform training image degradation processing on each high-resolution training image, and generate corresponding low-resolution training images; Extracting high-frequency texture parts of each high-resolution training image and low-resolution training image respectively, forming the high-frequency texture parts of all high-resolution training images into a first training set, and forming the high-frequency texture parts of all low-resolution training images into a second training set; The probabilistic latent semantic analysis topic learning model is used to perform topic mining on all documents in the first and second training sets, and K topics are obtained, where K is a positive integer. Use K-SVD algorithm to train high-resolution topic dictionary and low-resolution topic dictionary for each topic, and obtain high-resolution topic dictionary set and low-resolution topic dictionary set; For each atom in the low-resolution topic dictionary set, its nearest neighbors are found in the corresponding high-resolution topic dictionary set and low-resolution topic dictionary set, and a nearest neighbor regression matrix is calculated based on the nearest neighbors.
6. The method according to claim 5, characterized in that The probabilistic latent semantic analysis topic learning model is used to perform topic mining on all documents in the first training set and the second training set to obtain K topics, including: Input the documents in the first training set and the second training set into the probabilistic latent semantic analysis topic learning model; Through the probabilistic latent semantic analysis topic learning model, semantically similar documents and image blocks in the documents are classified into the same topic, and K topics are obtained.
7. The method according to claim 5, characterized in that For each atom in the low-resolution topic dictionary set, finding its nearest neighbor in the corresponding high-resolution topic dictionary set and low-resolution topic dictionary set specifically includes: For atoms in the high-resolution topic dictionary set and the low-resolution topic dictionary set, a similarity matrix is calculated based on Euclidean distance or cosine similarity; For each atom in the low-resolution topic dictionary, the top N high-resolution atoms and the top M low-resolution atoms with the highest similarity are selected from the similarity matrix to form the nearest neighbors of the atom; The unique high-resolution atom with the highest similarity is selected from the nearest neighbors, and a nearest neighbor mapping relationship is established between the low-resolution atom and the selected high-resolution atom.
8. An image reconstruction device, characterized in that: The device comprises: The acquisition module is used to obtain the low-resolution image to be restored and the dictionary set; the dictionary set is a set of high-resolution and low-resolution atoms formed by dividing the high-frequency texture part of the training image into K topics through the probabilistic latent semantic analysis topic learning model, and then training each topic separately; A decomposition module, used for decomposing the low-resolution image to be restored into a high-frequency texture part and a low-frequency texture part; A partitioning module, for partitioning the high-frequency texture portion into overlapping documents, and partitioning each document into overlapping image blocks; The inference module is used to infer the topic of the document using the probabilistic latent semantic analysis topic learning model; A reconstruction module is used to reconstruct the image block using a dictionary set associated with the topic to which the document belongs, to obtain a high-resolution and high-frequency texture image; The generation module is used to generate a reconstructed high-resolution image according to the high-resolution high-frequency texture image and the low-frequency texture part.
9. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the image reconstruction method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the image reconstruction method according to any one of claims 1 to 7.