Omnidirectional image super-resolution reconstruction method based on latitude perception potential diffusion model
Through the omnidirectional image super-resolution reconstruction method based on latitude-aware potential diffusion model, the problems of incomplete texture recovery and structural misalignment of omnidirectional images in the polar region are solved, adaptive adjustment and global consistency reconstruction are realized, and the quality and stability of omnidirectional images are improved.
Patent Information
- Application Number
- CN202510993920.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-07-18
AI Technical Summary
The prior art ignores spatial differences and insufficient adaptive adjustment when processing omnidirectional images, resulting in incomplete recovery of polar regions and obvious deformation, lack of adaptive control capabilities for regional differences, and traditional training mechanisms lack fine constraints on structural consistency and information coordination.
The omnidirectional image super-resolution reconstruction method based on the latitude-aware latent diffusion model is adopted. By constructing a latitude-aware map and image features, the latent representation and residual representation are extracted using an image encoder, a multi-stage latent diffusion sequence is constructed, and a reverse iterative sampling is performed through the denoising neural network, combined with the joint loss function for training and optimization.
Adaptive adjustment of texture restoration intensity in different latitude areas is achieved, the problems of blurred polar structure recovery and insufficient enhancement of equatorial regions are solved, the reconstruction quality and robustness of omnidirectional images are improved, and the consistency of the global structure and the full transmission of channel information are ensured.
Smart Images

Figure CN120510040A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and image processing, and in particular to an omnidirectional image super-resolution reconstruction method based on a latitude-aware latent diffusion model. Background Art
[0002] In real life and practical applications, high-definition, structurally stable panoramic images are gradually becoming a key requirement in many scenarios. For example, in virtual reality systems, high-precision omnidirectional images not only affect the user's immersive experience but also directly determine the effectiveness of subsequent 3D modeling and spatial perception tasks. In fields such as geographic remote sensing, medical imaging, and digital city construction, raw images often have low resolution limitations and cannot meet the needs of high-precision analysis and automated recognition. Therefore, how to achieve high-quality reconstruction and detail enhancement of omnidirectional images while maintaining image semantic consistency has become a technical problem that needs to be solved urgently.
[0003] The current mainstream image super-resolution methods are mainly based on convolutional neural networks to build end-to-end mapping models, directly learning to reconstruct the relationship between low-resolution images and high-resolution images. Such methods have made great progress in natural image datasets, especially in visual perception. Some methods introduce attention mechanisms to enhance the expressiveness of local features, which can improve the local texture restoration effect. Some models use perceptual loss to optimize the similarity of the image semantic layer, which helps to improve the high-level visual consistency of the image.
[0004] Although existing methods perform well in general image processing tasks, they still have obvious limitations when processing omnidirectional images, a special type of data. First, mainstream models generally equate spherical projection images with ordinary image processing, and do not consider the fact that there is geometric distortion in different latitudes. This leads to incomplete texture recovery in polar regions, obvious deformation, and a lack of adaptive control capabilities for regional differences. Second, most methods ignore the evolutionary information in the latent space and only perform direct regression in the image space. They lack the ability to gradually reconstruct and are prone to texture jumps or context misalignment problems. In addition, traditional training mechanisms often use a single loss function for optimization. When faced with complex network structures with both backbone paths and control paths, there is a lack of fine constraints on structural consistency and information coordination. To this end, those skilled in the art have proposed an omnidirectional image super-resolution reconstruction method based on a latitude-aware latent diffusion model to solve the above problems. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides an omnidirectional image super-resolution reconstruction method based on a latitude-aware latent diffusion model, which solves the problems in the existing technology such as ignoring spatial differences, insufficient adaptive adjustment, and difficulty in structural coordination.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: an omnidirectional image super-resolution reconstruction method based on a latitude-aware latent diffusion model, comprising the following steps: S1, obtain a low-resolution omnidirectional image and upsample it to obtain a guidance image through interpolation method; S2, constructing a latitude perception map and fusing it with image features, then inputting it into the latitude perception network to extract multi-scale image perception features; S3. Encode the high-resolution image and the upsampled guide image using an image encoder to obtain an original latent representation and a guide latent representation, and calculate the difference between the original latent representation and the guide latent representation to obtain a latent residual representation; S4. According to the set diffusion scheduling function and Markov transition rule, the original latent representation and the residual representation are superimposed and Gaussian noise is added to construct a multi-stage latent diffusion sequence; S5. Starting from the final diffusion representation, reverse iterative sampling is performed based on the denoising neural network to restore the original potential representation. S6, input the reconstructed potential representation into the decoder to decode it into a high-resolution image; S7. Training and optimizing the model through a joint loss function, where the loss function includes potential reconstruction loss, image reconstruction loss, and perceptual loss.
[0007] The present invention provides an omnidirectional image super-resolution reconstruction method based on a latitude-aware latent diffusion model. It has the following beneficial effects: 1. This invention introduces a latitude map as an auxiliary control signal and integrates spatial semantic features through a multi-scale control module. This achieves the effect of adaptively adjusting the texture restoration strength in different latitude regions. Compared with the existing strategy of uniformly processing omnidirectional image areas, it solves the problems of blurred structure restoration in the polar regions and insufficient enhancement in the equatorial region.
[0008] 2. By constructing a latent space reverse sampling mechanism based on a diffusion model, the present invention can effectively capture cross-scale semantic distribution features in low-dimensional space, achieving a robust process of gradually evolving high-score features rather than a one-time reconstruction. Previous technologies mostly relied on discriminative networks to directly predict high-score images, which resulted in unstable training and easily distorted reconstruction results. This solution solves this optimization bottleneck.
[0009] 3. The present invention adopts a structurally consistent decoding network combined with a residual and skip connection design method, which ensures the consistency of the global structure and the sufficient transmission of channel information while restoring the image texture details. Compared with traditional decoders, it is significantly superior to the hierarchical isolated decoding structure in multi-level feature fusion, alleviating the problem of detail loss in the reconstruction process.
[0010] 4. The joint loss function training mechanism proposed in this invention, including latent consistency, pixel reconstruction, perceptual difference and controlled alignment loss, realizes end-to-end consistency supervision from semantic space to image space. Compared with the previous methods that only relied on a single MSE or perceptual loss, this multi-dimensional joint constraint method is more adaptable to complex image patterns and solves the defects of slow training convergence and weak generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION
[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0013] Please see the attached Figure 1 The embodiment of the present invention provides an omnidirectional image super-resolution reconstruction method based on a latitude-aware latent diffusion model, comprising the following steps: S1, obtain a low-resolution omnidirectional image and upsample it to obtain a guidance image through interpolation method; Specifically, in the proposed omnidirectional image super-resolution reconstruction method, step S1, as the key initial stage, is responsible for constructing a guidance image from the input low-resolution omnidirectional image. This guidance image serves as a structural guidance condition in the subsequent latent diffusion modeling process, throughout the entire information restoration chain. Therefore, to ensure the stability and consistency of the model's initial understanding of image content, it is necessary to quickly align the input image size to the target high-resolution scale in a low-complexity manner.
[0014] Generally speaking, the generation of guided images does not directly improve image detail quality, but rather serves as a lightweight prior structure that provides a coordinated interface between modules. This step not only provides initial pixel features for the subsequent latent encoder but also preserves the spatially perceptual structure to a certain extent, thereby helping the model avoid structural misclassification.
[0015] Specifically, in this embodiment, the input image is a low-resolution omnidirectional image acquired by a panoramic camera, and its image size is set to The goal is to obtain a size of The corresponding high-resolution reconstructed image is 、 ,and is the magnification factor, usually 2, 4 or 8 are selected.
[0016] In one possible implementation, the original image is upsampled using the bilinear interpolation method to obtain a preliminary guidance image. , which satisfies: ; in: Indicates that the input image is factored Perform bilinear interpolation upsampling operation.
[0017] Alternatively, bicubic interpolation, Laplacian Pyramid Upsampling, or other parameter-free reconstruction methods can be used, provided that the output image size is strictly consistent with the target resolution and does not introduce model learnable parameters.
[0018] In some embodiments, in order to enhance the structural fidelity of the interpolated guidance image, preprocessing methods such as illumination equalization, brightness normalization or edge enhancement can also be used to make the structural features more prominent, which facilitates the subsequent potential encoder to extract corresponding features.
[0019] Furthermore, the generated It will be input into the downstream potential representation modeling module as a structural guidance graph, and the subsequent extraction of the guidance potential representation linkage.
[0020] It should be noted that although While not inherently containing additional texture detail, it plays a crucial role in model design. Its purpose is not to generate content, but rather to provide positioning, providing a positional structural reference for subsequent model recovery of potential residuals. This process is similar to providing a "weak structural skeleton" in a fuzzy space.
[0021] It is worth noting that the construction of this guidance map is a one-time static process that does not require training. Its design goal is to minimize computational overhead while retaining sufficient spatial semantic features to assist the model in accurately reconstructing the target image details.
[0022] In addition, in some other embodiments, to match the omnidirectional image characteristics, the input image With guided images The number of channels is usually set to 3 (i.e., RGB channels), but for other sensor images such as thermal infrared and depth maps, the number of channels can be extended to 1 or more.
[0023] The above interpolation strategy does not explicitly learn the image content, but only relies on the local weight estimation of space, so the generated There will be blurring in the edge areas, which is an important basis for correction by potential diffusion modeling in subsequent steps.
[0024] S2, constructing a latitude perception map and fusing it with image features, then inputting it into the latitude perception network to extract multi-scale image perception features; Specifically, after generating the guidance image, the present invention proceeds to the next stage, step S2: latitude map construction and feature fusion. This step follows the structure guidance map generated in step S1 and further combines the geometric deformation differences in the vertical dimension of the omnidirectional image to construct a latitude map for perceiving the distortion characteristics of the longitude and latitude structure. This latitude map is then integrated into the image feature representation, improving the model's responsiveness to spatial location information.
[0025] In general, due to its spherical projection characteristics, omnidirectional images exhibit uneven distribution of spatial distortion in the longitudinal direction (i.e., height dimension) of the image. Pixels near the top or bottom of the image correspond to the polar regions of the sphere, while the middle portion approximately corresponds to the equator of the sphere. This type of nonlinear spatial deformation is often ignored in traditional planar image processing, resulting in structural dislocation or blurring in the polar regions of the reconstructed model. Therefore, in the present invention, a latitude map is introduced as a geometric compensation mechanism for spatial position to construct a spatial perception structure prior.
[0026] Specifically, in this embodiment, the height dimension of the input image is first Perform position indexing and set the image size to ,in Indicates the image height, Indicates the image width.
[0027] Constructing a 2D latitude map , which is defined as follows: ; in: Indicates the image Liedi The latitude weighted value of the row pixel; Indicates the height of the image in pixels; Represents the horizontal coordinate index of the image; Indicates the vertical coordinate index of the image; is the pi constant, which controls the amplitude range of the mapping. Smooth changes in the interval [-1,1]; introduced in the function form As a normalized position term, it ensures that the latitude map is symmetrically distributed around the center of the image.
[0028] As an option, you can also Further normalization is performed to compress its range to the interval [0,1]. The specific form is: ; However, in this embodiment, in order to enhance the contrast sensitivity between the positive and negative regions, the original cosine mapping value is directly used.
[0029] In one possible implementation, the latitude map It will be expanded and copied to multiple channels and concatenated with the backbone features. For example, for the backbone network features , latitude features can be constructed , and broadcast the copy as , and then with The cascade is: ; In this embodiment, the latitude map can be combined with the structure guide map Or its feature representation is input into the subsequent feature extraction network.
[0030] As an implementation method, the latitude map can also be input into the control branch to assist in constructing a modulation factor with spatial bias for controlling the backbone convolution path.
[0031] In some embodiments, the latitude coefficient can be further applied to the attention weight generation of each feature channel to form a weighted feature map to achieve a regional adaptation mechanism of "polar suppression and equatorial enhancement", thereby improving the model's recognition ability in edge areas.
[0032] In addition, in this embodiment, the constructed latitude map is considered to be a non-learnable structural prior and does not participate in the back-propagation weight update process. It is mainly used to guide the model to automatically adapt to the spatial perception response capabilities of each region in the image.
[0033] In general, to maintain structural consistency, the construction process of the latitude map directly corresponds to the input image resolution. In other words, if the input image is , then the latitude map should also be , to avoid dimension misalignment problems in subsequent splicing operations.
[0034] In order to enhance the feature alignment capability, in some implementations, the latitude map can also be preliminarily encoded through a convolutional module, so that it is transformed from the original position map into an implicit semantic map before participating in the fusion. However, this method is not the preferred path in this method because it introduces learnable parameters and destroys the stability of the spatial prior.
[0035] Furthermore, in the fusion structure, the latitude map is not only combined with the encoding features Splicing, also with the potential guide features in step S3 They are input into the multi-layer convolution module together to form a joint feature set that integrates contextual information, spatial location, and structural guidance.
[0036] This feature set will be used as conditional information to input into the noise prediction network in the subsequent potential diffusion modeling process. In the process, it plays the role of global spatial structure modulation.
[0037] S3. Use an image encoder to encode the high-resolution image and the upsampled guide image respectively to obtain the original latent representation and the guided latent representation, and calculate the difference between the original latent representation and the guided latent representation to obtain the latent residual representation; Specifically, after completing the generation of the guidance map and the construction of the latitudinal map, the present invention enters the latent representation construction phase, i.e., step S3. The primary goal of this phase is to extract the multi-layered representation of the image in the latent space and construct latent residual features by comparing the features of the guidance image with the high-quality image. These residual features play a key role in information compensation and restoration in the subsequent latent diffusion process and are the fundamental source of the model's detail restoration performance.
[0038] Typically, the explicit pixel information of an image contains a significant amount of redundancy, while the latent space can express the image's semantics and structure in a compressed manner. This method exploits this property by encoding the high-resolution image and the guidance image separately to obtain a latent representation, and characterizes detail loss through the residual relationship between the two. Compared to direct pixel modeling, this method boasts greater representation sparsity and modeling stability, facilitating the effective implementation of diffusion-based prediction mechanisms.
[0039] Specifically, in this embodiment, the input high-resolution image is first set to , the guide image is , both are the same size, .
[0040] Build two latent encoder networks separately: main encoder and boot encoder , used to extract the potential features of high-resolution images and guidance images. Among them, ; ; in: Represents the encoded representation of the target high-resolution image in the latent space, which is the original latent representation; Represents the encoded representation of the guidance image in the latent space, which is the guidance latent representation; , ,in, The encoding downsampling factor is usually 4 or 8; is the number of latent space channels, usually set to 64, 128 or other suitable values; The encoder structure generally uses lightweight convolution stacking, such as multi-layer residual modules or Conv-BN-ReLU structure, to avoid information overfitting.
[0041] In order to obtain the difference in the expression of high-frequency details, the latent residual tensor is calculated: ; in: Represents the encoded representation of the high-resolution image in the latent space, which is the original latent representation; Represents the encoded representation of the upsampled image in the latent space, which is the guided latent representation; Represents the potential residual feature map between the two, which is the potential residual representation.
[0042] The residual tensor reflects the detailed features missing in the guidance image, and is particularly responsive to texture, edge, and local structure information.
[0043] In some embodiments, to prevent training instability caused by excessive feature scale, the latent features can be normalized before calculating the residual. For example, the mean and standard deviation of each channel can be normalized: , ; ; in: , Represents the mean and standard deviation calculated by channel respectively.
[0044] As an option, attention mechanisms such as channel attention or spatial attention modules can be added to the residual features to further improve local feature sensitivity.
[0045] In one possible implementation, the latent feature extraction network not only receives the image input, but also integrates the latitude map features generated in step S2. For example, at the encoder input, with latitude map Splicing, the form is: ; ; This design enables the encoder to have spatial bias capabilities, which helps the model focus on the geometric characteristics of the polar regions and improves the robustness of the potential representation to structural distortions.
[0046] In addition, multi-scale structures can be added during feature encoding to capture the local and global feature levels of the image. Common structures include: Dilated Conv to expand the receptive field; Pyramid pooling (SPP module) is used to extract scale-invariant semantics; The deformable convolution kernel mechanism (Deformable Conv) is used to enhance the adaptability to polar distorted areas.
[0047] In this embodiment, the potential residual It will be used as an information perturbation term in the subsequent potential diffusion process, and its participation level will be dynamically adjusted according to the number of steps. Its construction method will directly affect the stability and prediction accuracy of the subsequent diffusion sequence generation.
[0048] In order to enhance the robustness of feature representation, the encoder structure is recommended to share some parameters with the denoising network, or adopt a symmetric structure to achieve collaborative constraints on the information path.
[0049] S4. According to the set diffusion scheduling function and Markov transition rule, the original latent representation and the residual representation are superimposed and Gaussian noise is added to construct a multi-stage latent diffusion sequence; Specifically, after constructing the latent representation and latent residual, the method proceeds to the diffusion sequence generation stage in the latent space, i.e., step S4. The core of this stage is to transform the target latent representation into a controlled perturbation sequence. Through a stepwise diffusion operation, the information degradation path is simulated, providing noise conditions and dynamic guidance cues for the subsequent reverse reconstruction process.
[0050] In general, traditional diffusion models often perform diffusion operations by adding Gaussian noise in pixel space. However, the diffusion strategy proposed in this paper does not act directly on image pixels, but is built in the latent feature space, which has stronger structural consistency and modeling efficiency. At this stage, the model introduces the latent residual as the perturbation benchmark and combines it with Gaussian randomness to iteratively construct the latent variable sequence at multiple step sizes. , thereby completing the gradual mapping of high-dimensional semantics to low-quality representations.
[0051] Specifically, in this embodiment, the potential representation obtained and potential residuals Used to construct diffusion sequences.
[0052] Define step index ,in Indicates the total number of diffusion steps, which is usually a fixed value, such as 50, 100 or 200. In order to control the disturbance amplitude of each step, the time scheduling factor is introduced , which is defined as: ; in: For the Diffusion intensity of the step; is the initial disturbance amplitude; is the maximum disturbance amplitude; To schedule the curvature control parameters and determine the nonlinear trend of the diffusion process; is the current time step; is the total number of diffusion steps.
[0053] Based on the above scheduling, construct the potential diffusion state , which is defined as follows: ; in: For the The diffusion potential characteristics of the step; is the original target latent representation; is the potential residual constructed in the previous stage; Control the degree of participation of residual information; represents standard normal Gaussian noise; is the noise amplitude coefficient, which controls the intensity of the noise disturbance; Indicates the standard deviation scaling term bound to the scheduling factor.
[0054] It should be noted that the above formula will explicitly express the residual With random Gaussian terms The fusion process enables the potential state of each step to have both structural guidance and noise generalization capabilities to a certain extent, providing sufficient conditions for subsequent counter-diffusion.
[0055] As an implementation method, the state transition between each diffusion step can be modeled in the form of a Markov chain. The forward transition probability of the potential diffusion state is defined as: ; in: represents the conditional probability distribution of the underlying states; , represents the increment of diffusion intensity between adjacent steps; represents the covariance matrix of this step; is the unit covariance matrix; Indicates the mean , the covariance is Gaussian distribution of random variables The probability density function of .
[0056] The Markov diffusion mechanism shows that the current state The previous state The shift is applied in the residual direction and noise samples are added. This transfer rule provides a theoretical basis for the denoising prediction model in the subsequent reverse sampling.
[0057] In some embodiments, in order to improve the stability and continuity of the sequence, an autoregressive adjustment mechanism can be introduced, that is, When introducing multiple historical status information such as , , constructing second-order or third-order Markov chains to further strengthen sequence consistency.
[0058] Alternatively, the diffusion scheduling factor and It can be fixed by pre-setting a function form, or it can be predicted and generated as a learnable parameter by an external control module (such as a time encoder or attention network) to achieve a more flexible scheduling mechanism.
[0059] Furthermore, in actual implementation, the sequence of all diffusion steps can be generated at once through the sampler module , a step-by-step iterative generation mode can also be adopted, expanding layer by layer according to time steps, which is beneficial to saving video memory and optimizing computing resource allocation.
[0060] After the diffusion sequence is generated, the final state It will be used as the initial input for the reverse denoising sampling in step S5 to recover the potential representation approximation and reconstruct the clear image content.
[0061] S5. Starting from the final diffusion representation, reverse iterative sampling is performed based on the denoising neural network to restore the original potential representation. Specifically, after completing the forward generation of the diffusion sequence, the present invention enters step S5, which is the reverse denoising sampling stage. This stage is based on the last state in the potential diffusion space. Taking the trained conditional denoising prediction model as the starting point, reverse reconstruction is performed to gradually eliminate the noise and residual signal through step-by-step reasoning to restore the original state of the potential representation.
[0062] In general, the reverse process of the diffusion model is driven by a parameterized denoising network, which gradually guides the potential state from the noisy state to a clean representation with a clear structure. In the present invention, the reverse denoising path not only relies on the diffusion state , while also combining the guided latent representation , original low-resolution image , Latitude map and the current step size , in order to form a multi-condition control mechanism and realize the coordinated control of information such as spatial structure, semantic prior and dynamic disturbance.
[0063] Specifically, in this embodiment, the parameterized conditional denoising network is defined as , whose input includes the current potential state , guide potential features , step length , low-resolution images and its corresponding latitude map , the output is the residual prediction under the current step size.
[0064] Reverse sampling The sampling mean of the step is defined as: ; in: Indicates the The reconstructed mean of the step; , is the diffusion scheduling factor defined in step S4; represents the increment of diffusion intensity between adjacent steps; represents the conditional noise estimation network, whose structure includes a multi-scale residual block, a time modulation module, and a control condition fusion module; All inputs were kept at a spatial scale consistent with the diffusion regime.
[0065] The covariance introduced during sampling is: ; in: For the The Gaussian noise covariance matrix of the step sampling; is the noise coefficient, taken from the forward process definition; is the unit covariance matrix, which is used to control the uniform distribution of noise.
[0066] Based on the above, the sampling process can be written as: ; The process starts from the current potential state Restore forward to , and finally The estimated potential original state is obtained when .
[0067] In one possible implementation, the time step It can be fed into the model through positional encoding, such as using sine and cosine embedding or a learnable method based on Learnable Embedding, so that the model can dynamically adapt to the recovery strategies at different stages.
[0068] To enhance the ability to express control information in the network, this embodiment introduces a pixel-aware control block (PACB), whose core structure is as follows: ; in: Represents the final fused output feature map; Intermediate feature graph representing control flow branches; Represents the fused feature map after concatenating the context feature and the control feature; represents a convolution operation with zero-initialized weights; Represents the corresponding element-wise multiplication operation.
[0069] This structure realizes the dynamic adjustment of backbone features by the control path and integrates residual information to achieve controllable sampling.
[0070] Generally, the PACB structure is placed after the multi-layer residual module, and control guidance is applied at the middle and high-level stages of feature expression to ensure the effectiveness of information injection.
[0071] In some embodiments, the PACB module can work in conjunction with time embedding to form a time-aware PACB to further improve timing consistency and control response stability.
[0072] It is worth noting that the above reverse sampling process is a conditional generation form. The model can accept different structural guidance conditions, maintain the diversity of the generated space, and at the same time have a certain degree of reconstruction fidelity.
[0073] After sampling is completed, the final estimated potential representation The input is sent to the decoder defined in step S6 for restoration output in the image domain.
[0074] S6, input the reconstructed potential representation into the decoder to decode it into a high-resolution image; Specifically, after completing the reverse denoising sampling operation on the diffusion potential sequence, the final potential representation estimate is obtained The method of the present invention enters step S6, which is the process of restoring the latent features to the image space. This step upsamples and restores the latent features through a structured decoding network, and ultimately outputs a super-resolution image with a size consistent with the high-resolution target image.
[0075] Generally speaking, the latent space and image space have significant differences in dimensionality distribution and information organization, so direct dimensionality mapping cannot be performed. A decoding network with good structural consistency and clear feature response is needed to reconstruct abstract semantics into specific texture content.
[0076] In the present invention, the decoder design needs to meet two goals: one is to maintain the correspondence between the reconstructed image and the original high-resolution image in structure and details, and the other is to achieve end-to-end mapping without introducing too much inference computation burden.
[0077] Specifically, in this embodiment, the input is the potential representation obtained in step S5 ,in represents the number of feature channels, , , is the feature scaling factor.
[0078] Set the final target image to ,in 、 is the original high-resolution target size.
[0079] To complete this mapping operation, construct a decoder , its output expression is: ; In one possible implementation, the decoder consists of multiple upsampling modules, each of which includes an upsampling operation, a convolution layer, a normalization layer, and an activation function.
[0080] In general, upsampling can take one of the following forms: Transposed Convolution; Sub-pixel convolution (PixelShuffle structure); Nearest neighbor interpolation or bilinear interpolation + convolution combination.
[0081] In this embodiment, the PixelShuffle structure is preferably used to reduce the grid artifact phenomenon. Specifically, assuming the upsampling multiple is , then the decoding network contains two layers of continuous pixel rearrangement modules, each layer performs Space enlargement.
[0082] The operation of each decoding block can be expressed as: ; in: Indicates the Layer input feature map; Realize the rearrangement of channel dimension to space dimension; It is a standard convolution operation with a kernel size of 3; It is an activation function used to enhance nonlinear expression capabilities.
[0083] In some embodiments, a 1×1 convolution can be added to the end of the decoder for channel mapping to convert the final feature map into RGB image format: ; In order to further improve the reconstruction quality, this embodiment also introduces a residual connection structure in the decoding stage, so that low-level features and high-level features form a fusion path: ; in: To ensure the size compatibility of skip connections, the spatial alignment operation is consistent with the decoding stride.
[0084] As an option, an attention module, such as channel attention or spatial attention, can be embedded in the decoder structure to improve the ability to restore texture details, especially for recovering detailed structures in polar regions.
[0085] In some embodiments, in order to enhance the stability and generalization of the decoder, a symmetrical structure design may be used to mirror the decoder structure to the encoder structure, thereby achieving structural consistency of the information path.
[0086] Furthermore, considering the spatially uneven distribution characteristics of the omnidirectional image, in this embodiment, the input features received by the decoder The modulation information of the latitude map in step S2 is already included, so there is no need to repeatedly add geometric priors in the decoding stage. Instead, polar weight compensation and equatorial region detail enhancement are completed through feature expression conduction.
[0087] The final generated will be with high resolution images During the training phase, supervised comparison is performed to construct the loss function. After training, the decoder has the ability to directly restore clear images.
[0088] S7. The model is trained and optimized through a joint loss function, where the loss function includes potential reconstruction loss, image reconstruction loss, and perceptual loss.
[0089] Specifically, after completing latent space decoding and obtaining the final reconstructed image, the present invention proceeds to step S7, the joint loss training phase, to further optimize the synergy between the model modules and improve the accuracy and stability of the reconstruction results. This step, by constructing a multi-level, cross-space loss function system, achieves simultaneous improvements in latent space modeling accuracy and image space restoration quality, and is a key part of the entire network training process.
[0090] Traditional image super-resolution methods typically use pixel-space mean squared error as the sole supervisory signal, making it difficult to effectively constrain the feature evolution process in the latent space. In this paper, by simultaneously introducing latent space alignment loss, image space reconstruction error, perceptual consistency constraints, and structural control module supervisory signals, we achieve multi-dimensional joint control of the modeling process.
[0091] Specifically, in this embodiment, four main types of loss items are constructed, namely: Latent Consistency Loss; Image space mean square error loss (MSE Loss); Perceptual Loss Control Alignment Loss.
[0092] First, define the latent space reconstruction loss as follows: ; in: is the original target latent representation; For the Step diffusion state; Sampling prediction network for reverse denoising; is the total number of diffusion steps; and They are low-resolution images and latitude maps respectively; represents the L2 norm.
[0093] This loss term is used to constrain the approximation accuracy of the target potential state at each step in the inverse denoising process.
[0094] Secondly, define the image space mean square error loss as follows: ; in: It is a real high-resolution image; is the reconstructed image output by the decoder; this item is mainly used to evaluate the accuracy of the reconstructed image at the pixel level.
[0095] In order to enhance the expressive power of structure perception, perceptual loss is further introduced: ; in: It is a fixed perceptual feature extraction network, such as the intermediate layer output of VGG-19; this loss focuses on the consistency of the image in the high-level semantic feature space and is used to supplement the limitations of the pixel space.
[0096] In addition, in order to maintain consistent response between the trunk path and the control path, a control consistency loss is designed in this embodiment, and the specific form is: ; in: Represents the output of the pixel perception control module; is the corresponding intermediate feature map in the trunk path; Indicates a gradient stop operation, used to prevent the main gradient from flowing back into the control path; Represents the consistency constraint loss for the pixel-aware control block output; this loss guides the control module output to align with the backbone semantic features, ensuring that the control signal is stable and reliable.
[0097] Finally, the above losses are combined to form the overall training objective function: ; in: represents the consistency constraint loss for the pixel-aware control block output; represents the joint loss function used for overall model training; represents the reconstruction loss of the latent space; Represents the pixel reconstruction error loss in the image space; Represents the perceptual loss based on the difference of deep image features; is the weighted coefficient of each loss term; each coefficient can be adjusted through cross-validation or empirical setting to ensure a balance between network convergence and reconstruction quality.
[0098] In some embodiments, a step-by-step scheduling strategy can be adopted, that is, weakening the weights of perception loss and control loss in the early stage of training, relying only on latent space and pixel space supervision; and gradually increasing the proportion of perception items and control items after the model has stably converged to enhance the model's structural understanding ability and response flexibility.
[0099] As an option, the loss composition structure can also be dynamically adjusted based on the task scenario, for example, emphasizing pixel accuracy in high-precision industrial imaging scenarios, while paying more attention to structural perception characteristics in visual perception tasks.
[0100] During the training process, this embodiment uses the Adam optimizer, and the initial learning rate is usually set to , and dynamically decreases when there is no improvement in the preset number of training rounds or validation set indicators.
[0101] All training processes are carried out under a multi-GPU parallel architecture and combined with a mixed precision training mechanism to improve training efficiency and stability.
[0102] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. An omnidirectional image super-resolution reconstruction method based on a latitude-aware latent diffusion model, characterized by: The following steps are involved: Obtain a low-resolution omnidirectional image and upsample it to obtain a guidance image through interpolation; Construct a latitude perception map and fuse it with image features before inputting it into a latitude perception network to extract multi-scale image perception features. Encode the high-resolution image and the upsampled guide image using an image encoder to obtain an original latent representation and a guide latent representation, and calculate the difference between the original latent representation and the guide latent representation to obtain a latent residual representation; According to the set diffusion scheduling function and Markov transition rule, the original latent representation and the residual representation are superimposed and Gaussian noise is added to construct a multi-stage latent diffusion sequence; Starting from the final diffusion representation, reverse iterative sampling is performed based on the denoising neural network to restore the original potential representation; The reconstructed latent representation is input into the decoder to decode into a high-resolution image; The model is trained and optimized through a joint loss function, which includes potential reconstruction loss, image reconstruction loss and perceptual loss.
2. The omnidirectional image super-resolution reconstruction method based on the latitude-aware latent diffusion model according to claim 1 is characterized in that: The constructing of the latitude perception map includes: Construct a cosine function weighted graph for the vertical coordinates of the input image to express the latitude distortion information; The latitude map is used as a weight map and is concatenated with the image features in the channel dimension to input the latitude perception network; The latitude-aware network extracts the structural and geometric features of the image to assist in the reconstruction of the latent representation; Among them, the latitude map The weight of the position is: ; in: Indicates the image Liedi The latitude weighted value of the row pixel; Indicates the height of the image in pixels; Represents the horizontal coordinate index of the image; Indicates the vertical coordinate index of the image; is the circumference constant of pi.
3. The omnidirectional image super-resolution reconstruction method based on the latitude-aware latent diffusion model according to claim 1, characterized in that: The acquisition of the original potential representation and the potential residual includes: Use the encoder to train real high-resolution images Encode, get representation ; Upsample the image Encoding is guided representation ; and The residual is expressed as: ; in: Represents the encoded representation of the high-resolution image in the latent space, which is the original latent representation; Represents the encoded representation of the upsampled image in the latent space, which is the guided latent representation; Represents the potential residual feature map between the two, which is the potential residual representation.
4. The omnidirectional image super-resolution reconstruction method based on the latitude-aware latent diffusion model according to claim 1, characterized in that: The multi-stage potential diffusion sequence is constructed as follows: Set the number of diffusion steps , minimum diffusion intensity , maximum diffusion intensity , scheduling index ; Construct a diffusion intensity scheduling function: ; In the Step 2 generates latent representation: ; in: For the Diffusion intensity of the step; is the minimum value of diffusion intensity; is the maximum value of diffusion intensity; is the exponential control factor of diffusion scheduling; is the total number of diffusion steps; is the current diffusion step index, and its value range is ; is the Gaussian noise scaling factor; is a Gaussian random noise tensor with zero mean and unit variance; is the original target latent representation; is the potential residual constructed in the previous stage; Indicates the standard deviation scaling term bound to the scheduling factor.
5. The omnidirectional image super-resolution reconstruction method based on the latitude-aware latent diffusion model according to claim 1, characterized in that: The Markov transition rule is defined as: ; in: The transition probability density function representing the potential diffusion; represents the potential representation of the previous diffusion step; Represents the encoded representation of the upsampled image in the latent space; represents the potential residual feature map; represents the increment of diffusion intensity between adjacent steps; Indicates the Diffusion intensity of the step; Indicates the Diffusion intensity of the step; represents the covariance matrix of this step; is the identity matrix with the same dimension as the latent representation; Indicates the mean , the covariance is Gaussian distribution of random variables The probability density function of .
6. The omnidirectional image super-resolution reconstruction method based on the latitude-aware latent diffusion model according to claim 1, characterized in that: The reverse iterative sampling process includes: At every step The denoising network predicts the mean and covariance of the next potential state, and the prediction form is: ; ; in: represents the mean vector used for sampling; represents the covariance matrix used for sampling; Indicates the Diffusion intensity of the step; Indicates the Diffusion intensity of the step; represents the intensity increment between adjacent diffusion steps; Indicates the The potential diffusion representation of the step; Representations guide latent representations; is the noise scaling factor; is the unit covariance matrix; For parameters The denoising neural network function is used to estimate the noise residual; is the original low-resolution omnidirectional image; is the latitude weighted graph corresponding to the image space; Indicates the current diffusion step index.
7. The omnidirectional image super-resolution reconstruction method based on the latitude-aware latent diffusion model according to claim 1, characterized in that: The denoising neural network comprises: The input contains the current diffusion potential representation , guide the latent representation , low-resolution images , Latitude map , current step count ; The network consists of a temporal modulation module and a pixel-aware control module, the feature enhancement of which is expressed as: ; in: Represents the final fused output feature map; Intermediate feature graph representing control flow branches; Represents the fused feature map after concatenating the context feature and the control feature; represents a convolution operation with zero-initialized weights; Represents the corresponding element-wise multiplication operation.
8. The omnidirectional image super-resolution reconstruction method based on the latitude-aware latent diffusion model according to claim 1, characterized in that: The decoder comprises: Input is the potential representation after reverse sampling ; It includes a decoding module composed of a multi-layer deconvolution structure to recover the latent representation into image space representation; The output is a reconstructed high-resolution image .
9. The omnidirectional image super-resolution reconstruction method based on the latitude-aware latent diffusion model according to claim 1, characterized in that: The loss function includes: Latent space reconstruction loss: ; Image space reconstruction loss: ; Perceptual Loss: ; in: represents the L2 norm, i.e. the Euclidean distance; Denoising neural network for predicting latent representations; represents the denoising reconstruction loss in the latent space; is the total number of diffusion steps; Represents the original representation of high-resolution images in the latent space; Indicates spread to potential representation of the step; represents the guided latent representation obtained by bilinear interpolation image encoding; is the original low-resolution omnidirectional image; is the latitude weighted graph corresponding to the image space; Indicates the current diffusion step index; represents the mean square error loss in pixel space; Represents a true high-resolution image; represents the reconstructed image generated by the decoder; represents the perceptual loss, which is used to measure the difference of images in the deep feature space; Represents the deep feature representation of an image extracted by a pre-trained image perception network.
10. The omnidirectional image super-resolution reconstruction method based on the latitude-aware latent diffusion model according to claim 9, characterized in that: The total loss function of the loss function is defined as: ; in: represents the consistency constraint loss for the pixel-aware control block output; represents the joint loss function used for overall model training; represents the reconstruction loss of the latent space; Represents the pixel reconstruction error loss in the image space; Represents the perceptual loss based on the difference of deep image features; is the weighting coefficient of each loss item.
Citation Information
Patent Citations
Conditional diffusion probability model training method and omnidirectional image generation method
CN118096532A
Method for generating high-resolution picture, computer device, and storage medium
US20200258197A1
Small image multi-object detection method based on super-resolution
WO2023060746A1
Cited By
Arbitrary-scale super-resolution method, device and equipment based on compact Gaussian splashing
CN121147020A
Unmanned aerial vehicle semantic perception domain generalization method and system based on probability diffusion modeling and storage medium
CN122156827A