High-precision panorama generation system and method
Through technical means such as semantic consistency and multi-perspective optimization, the geometric and lighting consistency of panoramas are improved, which solves the image quality and accuracy problems existing in existing technologies and expands the application of panoramas in multiple fields.
Patent Information
- Application Number
- CN202510691260.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing panorama generation methods have deficiencies in geometric consistency and illumination consistency, which affect image quality and accuracy and limit their application in fields such as virtual reality, augmented reality, and cultural heritage digitization.
By adopting semantic consistency, multi-view optimization, diffusion model restoration, multi-scale optimization, super-resolution processing and illumination consistency technologies, the geometric and illumination consistency of the panorama is improved through the image projection module, deformation restoration module, illumination consistency module and semantic consistency module.
It achieves high-quality and high-precision panoramic image generation, expands its application potential in cultural heritage protection, architectural design, autonomous driving and other fields, and improves the realism and reliability of images.
Smart Images

Figure CN120599098A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and computer graphics, and in particular to a high-precision panoramic image generation system and method. Background Art
[0002] High-precision panoramas, with their ultra-high-definition resolution, wide viewing angle (360° or 720°), and precise spatial positioning, provide unprecedented immersive and high-fidelity visual experiences for a wide range of fields, including architecture and interior design, cultural heritage preservation and digitization, autonomous driving and robotic navigation, industrial inspection and maintenance, real estate, and tourism. In architecture and interior design, designers can use high-precision panoramas to simulate spatial layouts and lighting effects in advance, reducing design errors. Clients can also use these panoramas to intuitively experience design solutions and accelerate their decision-making process. For cultural heritage preservation and digitization, millimeter-level scanning technology enables the creation of extremely detailed panoramas, preserving every detail of cultural relics and supporting long-term research and educational displays. In autonomous driving and robotic navigation, the ability to match panoramas with real-time scene information helps vehicles perceive complex road conditions, enabling precise navigation and obstacle avoidance. For industrial inspection and maintenance, engineers can use panoramas to view equipment details, reducing on-site repair costs. Furthermore, in the real estate market, homebuyers can view 360-degree details of properties, improving their decision-making efficiency. In the tourism industry, panoramic images of scenic spots attract potential visitors and promote tourism development.
[0003] However, existing panorama generation methods have significant deficiencies in geometric consistency and illumination consistency, which directly affect the quality and accuracy of the generated images. The geometric consistency problem mainly occurs in the process of stitching images from different perspectives. Due to the lack of accurate modeling of scene depth information and camera parameters, the stitched panorama may have objects misplaced, deformed, or chaotic occlusion relationships. The illumination consistency problem stems from the inadequate modeling of dynamic changes in ambient lighting conditions by existing methods. The light intensity, direction, and color temperature of the same scene will change significantly at different times or under different weather conditions, but traditional methods usually rely on a single illumination model or global brightness adjustment, which makes it difficult to accurately simulate local detail changes in complex lighting scenes. These problems jointly limit the practical application and development potential of panoramas in fields such as virtual reality, augmented reality, and cultural heritage digitization. Summary of the Invention
[0004] To overcome the above challenges, the present invention proposes an innovative high-precision panoramic image generation method and system. By adopting a variety of advanced technical means such as semantic consistency, multi-perspective optimization, diffusion model restoration, multi-scale optimization, super-resolution processing, illumination consistency and quality assessment, it effectively solves the problems of geometric consistency and illumination consistency, and realizes the generation of high-quality and high-precision panoramic images.
[0005] To achieve the above-mentioned object, the present invention provides a high-precision panoramic image generation system, comprising: an image projection module, a deformation repair module, an illumination consistency module, and a semantic consistency module;
[0006] The image projection module is used to generate an initial perspective image based on the initial image;
[0007] The deformation repair module is used to generate a repaired perspective image based on the initial perspective image and input text;
[0008] The lighting consistency module is used to generate a lighting consistency score and an initial panoramic image based on the initial image and the restored perspective image;
[0009] The semantic consistency module is used to output a panoramic image based on the initial panoramic image and input semantics.
[0010] Preferably, the workflow of the image projection module includes: projecting the initial image onto a unit sphere, generating the initial perspective image by defining the vertices of each image pixel and creating edges between adjacent pixels:
[0011]
[0012] Among them, (X, Y, Z) represents a point on the unit sphere; (x proj ,y proj ) represents a point on the perspective plane; f length Indicates focal length.
[0013] Preferably, the deformation repair module uses a rotation matrix to ensure that the repair is performed only on non-overlapping areas between adjacent views; the rotation matrix R includes:
[0014] R = RotationMatrix(r) = R x (θ x )·R y (θ y )·R z (θ z )
[0015] Among them, RotationMatrix represents the function of converting Euler angles into rotation matrices; the Euler angle of rotation around the x-axis is θ x, and its corresponding matrix The Euler angle of rotation around the y-axis is θ y , and its corresponding matrix The Euler angle of rotation around the z-axis is θ z , and its corresponding matrix
[0016] Preferably, the deformation repair module introduces a multi-scale diffusion model to gradually repair the image at different resolutions; the process includes:
[0017] Initial low-resolution image generation:
[0018] l low =Downsample(I original )
[0019] Among them, I original Represents the original image; I low Represents the low-resolution image after downsampling;
[0020] Stepwise Diffusion Inpainting: Apply a diffusion model at each scale to inpaint, gradually improving the resolution; for each size s, from low to high, perform a diffusion process and upsampling:
[0021]
[0022] in, Represents the image of step t; ∈ noise represents noise; T represents the total number of diffusion steps;
[0023] Combine multi-scale inpainted images into the final perspective image:
[0024]
[0025] Lap s =LaplacianPyramid(I scale )
[0026] Among them, w s represents the weight assigned to the Laplacian pyramid of each scale s, Lap s Represents a Laplacian pyramid containing multi-level detail information; I scale Indicates the repaired image; Indicates the use of CNN to extract features of each layer in the Laplacian pyramid.
[0027] Preferably, the multi-scale diffusion model adopts a Laplacian pyramid model; the Laplacian pyramid model is constructed by using a Gaussian pyramid and a Laplacian pyramid to integrate image details at different scales; and the construction process includes:
[0028] Gaussian pyramid construction: Gaussian pyramid generates a series of images of different resolutions through gradual downsampling and Gaussian filtering;
[0029] Construction of Laplacian Pyramid: Laplacian Pyramid obtains detail information by subtracting Gaussian pyramid images at adjacent scales;
[0030] Laplacian pyramid representation: Laplacian pyramid Lap is an image sequence containing detail information from high to low resolution:
[0031]
[0032] Among them, N s Represents the total number of scales s.
[0033] Preferably, the workflow of the lighting consistency module includes:
[0034] Feature extraction: Use CNN to extract the feature vector F of each image i :F i =CNN(I perspective,i ), where I perspective,i is the perspective image of the i-th viewing angle. There are multiple perspective images, which are represented as follows:
[0035] {I perspective,i , C i |=1,2,...,N v}
[0036] Among them, C i Indicates the corresponding camera parameters; N v Indicates the number of different perspectives;
[0037] Volume rendering: For each view angle i, sample multiple points t along the ray r j , calculate the volume density σ and color c of each point:
[0038] σ j ,c j =NeRF(r(t j ))
[0039] Among them, r(t j ) represents the j-th sampling point on the ray r; NeRF represents the rendering method using neural radiance field;
[0040] Lighting parameter estimation: Use volume rendering results to estimate the lighting parameters Light for each view i i :
[0041] Light i =LightEstimation(Fi ,σ j , c j ).
[0042] Preferably, before generating the initial panoramic image, the lighting consistency module uses neural radiance field technology to estimate the lighting conditions of each perspective and ensure lighting consistency of all perspectives.
[0043] Preferably, the workflow of the semantic consistency module includes:
[0044] First, a large language model is used to extract key semantic information about the input semantics from the text prompt:
[0045] t key,out =LLM(T key,in )
[0046] Among them, T key,in Indicates the text prompt for input; t key,out Represents the extracted text feature vector;
[0047] At the same time, feature vectors are extracted from the initial image:
[0048] f i =CNN(I i )
[0049] Among them, I i represents the i-th initial image, f i represents the extracted image feature vector;
[0050] Then perform feature fusion and transform the text feature t key,out With image feature f i Fusion:
[0051] f′ i =βf i , +(1–β)t key,out
[0052] Among them, β represents a weight parameter.
[0053] Preferably, a classifier or discriminator is used to check whether the generated image is semantically consistent with the text prompt:
[0054] Score(I′ i )=Classifier(I′ i )
[0055] Among them, I′ i Represents the fused image; Score(I′ i ) represents the consistency score between the image and the text hint.
[0056] The present invention also provides a high-precision panoramic image generation method, which is applied to the above system and includes the following steps:
[0057] Generate an initial perspective image based on the input initial image;
[0058] generating a restored perspective image based on the initial perspective image and input text;
[0059] generating a lighting consistency score and an initial panorama based on the initial image and the restored perspective image;
[0060] A panoramic image is output based on the initial panoramic image and input semantics.
[0061] Compared with the prior art, the present invention has the following beneficial effects:
[0062] This invention not only enhances the realism and reliability of panoramic images but also significantly expands their application scenarios. For example, in cultural heritage preservation, this method can more realistically reproduce the details of cultural relics, supporting deeper research and broader public education. In architectural design, it can provide a more precise design verification tool, helping designers better understand the actual effects of spatial layouts. In the field of autonomous driving, improved panoramic image generation technology can help improve the vehicle's understanding of the surrounding environment, thereby enhancing driving safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0064] Figure 1 Schematic diagram of the system structure of an embodiment of the present invention;
[0065] Figure 2 Schematic diagram of the system structure of the expansion solution of an embodiment of the present invention. DETAILED DESCRIPTION
[0066] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0067] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0068] Example 1
[0069] like Figure 1 As shown, this is a schematic diagram of the system structure of this embodiment, including: an image projection module, a deformation repair module, an illumination consistency module and a semantic consistency module; the image projection module is used to generate an initial perspective image based on the initial image; the deformation repair module is used to generate a repaired perspective image based on the initial perspective image and input text; the illumination consistency module is used to generate an illumination consistency score and an initial panoramic image based on the initial image and the repaired perspective image; the semantic consistency module is used to generate a semantically consistent output panoramic image based on the initial panoramic image and input semantics.
[0070] The following will describe in detail how the present invention solves technical problems in practical work in conjunction with this embodiment.
[0071] The inputs to the system in this embodiment are an initial image, input text, and input semantics. The initial image is a sequence of images, such as a video, not a single image. The input text is text related to the rotation matrix, and the input semantics is text related to the image content. The overall process is as follows:
[0072] Using spherical projection, image data from all viewpoints is mapped onto the unit sphere, and overlapping areas are fused to ensure a seamless panorama. Image inpainting techniques are then used to fill in missing areas in the panorama. Finally, lighting consistency optimization and other post-processing steps are applied to enhance the panorama's quality.
[0073] First, the image projection module projects the initial image onto the unit sphere, generating an initial perspective image by defining the vertices of each image pixel and creating edges between adjacent pixels. The specific process is as follows:
[0074] (1) Image pixel coordinates are converted to spherical coordinates
[0075] Image pixel coordinates: Assume the initial image resolution is W in ×H in , the pixel coordinates are (x in ,y in ) where x in ∈[0,W in -1], y in ∈[0,H in -1]
[0076] Normalize pixel coordinates: Normalize pixel coordinates to the range [0, 1]:
[0077]
[0078] in, Represents the normalized pixel coordinates.
[0079] Convert to spherical coordinate system: convert the normalized pixel coordinates to a point (X, Y, Z) on the unit sphere
[0080]
[0081] Among them, θ v Indicates the viewing angle, usually set to 90 degrees (i.e. radian).
[0082] (2) Generate initial perspective view
[0083] Define vertices and edges: For each pixel (x in ,y in ), define its vertex V on the unit sphere; then create edges between adjacent pixels to form a grid structure.
[0084] Project to perspective: Use perspective projection to project points on the unit sphere onto the perspective plane
[0085]
[0086] Among them, (x proj ,y proj ) represents a point on the perspective plane; f length Indicates focal length.
[0087] The input of the Deformation Repair module is the initial perspective image and the input text, and the output is the repaired perspective image. The purpose of the Deformation Repair module is to ensure that the repair is only performed on the non-overlapping areas between adjacent views using the rotation matrix R. This module ensures that the generated image maintains geometric and appearance consistency between different viewpoints, while repairing any missing areas, thereby improving the quality and detail of the final generated image. The workflow of the Deformation Repair module is as follows:
[0088] (1) Dynamically adjust the rotation matrix
[0089] Text feature extraction: Use LLM to extract key semantic information from the input text and generate text feature vectors:
[0090] t R,out =LLM(T R,in )
[0091] Among them, T R,in is the text prompt of the input; t R,out is the extracted text feature vector.
[0092] Rotation matrix parameterization: Parameterize the rotation matrix R into a learnable vector r:
[0093] r=[θ x ,θ y ,θ z ]
[0094] Among them, θ x ,θ y ,θ z Represents the rotation angles (Euler angles) around the x, y, and z axes respectively.
[0095] Parameter mapping: Use a multi-layer perceptron (MLP) to transform the text features t R,out Mapping to the rotation parameter r:
[0096] r=MLP(t R,out )
[0097] Among them, MLP is a multi-layer perceptron used to convert text features into rotation parameters.
[0098] Rotation matrix generation: Generate the rotation matrix R according to the rotation parameter r:
[0099] R = RotationMatrix(r) = R x (θ x )·R y (θ y )·R z (θz)
[0100] Among them, RotationMatrix is a function that converts Euler angles into rotation matrices. The Euler angle for rotation around the x-axis is θ x , and its corresponding matrix The Euler angle of rotation around the y-axis is θ y , and its corresponding matrix The Euler angle of rotation around the z-axis is θ z , and its corresponding matrix
[0101] Loss function:
[0102] View consistency loss By comparing images from adjacent perspectives, we can ensure a smooth transition when switching perspectives and reduce artifacts and inconsistencies. The expression is as follows:
[0103]
[0104] Among them, N vRepresents the number of different perspectives; i represents the index of the current perspective; j represents the index of the perspective adjacent to perspective i; neighbors(i) means returning the set of all perspectives adjacent to perspective i.
[0105] Dynamically adjust the rotation matrix loss function:
[0106]
[0107] in, represents semantic consistency loss; λ view and λ light represents the weight coefficient; Indicates lighting consistency. By comparing lighting parameters at different viewing angles, it ensures lighting consistency and avoids lighting inconsistency issues.
[0108] (2) Pixel position calculation: Use the rotation matrix R from one viewpoint A to another viewpoint B, that is, for the pixel P0 in viewpoint A, calculate its corresponding position P in viewpoint B. rot :
[0109] P rot =R·P0
[0110] Among them, P0 represents the pixel coordinates in viewpoint A; P rot Represents the pixel coordinates in viewpoint B; R represents the rotation matrix.
[0111] Binary mask: Non-overlapping areas may contain missing areas. A binary mask M is used to ensure that restoration is performed only on non-overlapping areas between adjacent views:
[0112] M=1-CheckOverlap(P rot )
[0113]
[0114] Among them, CheckOverlap(P rot ) checks whether the pixel is in the overlapping area and generates a mask M of the overlapping area. W and H represent the width and height of the image respectively. rot,x represents the x coordinate of the pixel in viewpoint B, P rot,y Represents the y coordinate of the pixel in viewpoint B.
[0115] Through the binary mask, the possible actual areas are screened out and then input into the diffusion model to fill the actual area content. This embodiment also introduces a "multi-scale diffusion model" (this embodiment uses the Lablas pyramid method) to gradually repair the image at different resolutions. The gradual repair from low resolution to high resolution can reduce artifacts and ensure the consistency of details in the generated image. The process includes:
[0116] (1) Initial low-resolution image generation: usually 64×64 or 128×128;
[0117] I low =Downsample(I original )
[0118] Among them, I original Represents the original image; I low Represents the low-resolution image after downsampling.
[0119] (2) Stepwise Diffusion Inpainting: Apply the diffusion model to inpaint at each scale, gradually improving the resolution. For each size s, from low to high, perform the diffusion process and upsampling:
[0120]
[0121] in, Represents the image of step t; ∈ noise represents noise; T represents the total number of diffusion steps.
[0122] (3) Merge the multi-scale inpainted images into the final perspective image:
[0123]
[0124] Lap s =LaplacianPyramid(I scale )
[0125] Among them, w s represents the weight assigned to the Laplacian pyramid of each scale s, Lap s Represents a Laplacian pyramid containing multi-level detail information; I scale Indicates the repaired image; in addition, w s Make corrections, such as scoring the semantic consistency of each scale s and obtaining the score result Score s (I′ i ), modify the weights and get If the quality evaluation is performed on the restored image of each size, the evaluation results Q are obtained respectively. s , then adjust the weight again: ReconstructCNN represents the method of using CNN for reconstruction, fusing the features of all layers. Indicates the use of CNN to extract features of each layer in the Laplacian pyramid.
[0126] The Laplacian pyramid is a multi-resolution image representation method used to decompose an image into details at different scales. It effectively integrates image details at different scales by constructing a Gaussian pyramid and a Laplacian pyramid.
[0127] (1) Gaussian pyramid construction: Gaussian pyramid generates a series of images of different resolutions through gradual downsampling and Gaussian filtering. The specific steps are:
[0128] 1) Input the initial image I;
[0129] 2) Gaussian filtering and downsampling: Gaussian filtering is applied to the current initial image I, followed by downsampling (usually to half of the original size), and this process is repeated until the required minimum resolution is reached.
[0130] G s =GaussianBlur(I s )
[0131] l s+1 =Downsample(G s )
[0132] Among them, G s Represents the Gaussian filtered image of the current scale; I s+1 Represents the downsampled image; I s Represents the initial image of the s-th scale.
[0133] (2) Construction of Laplacian Pyramid: The Laplacian Pyramid obtains detail information by subtracting Gaussian pyramid images at adjacent scales. The specific steps are as follows:
[0134] Sampling and Gaussian filtering: For each image G in the Gaussian pyramid s Upsample and then apply Gaussian filtering:
[0135]
[0136] in, represents the upsampled Gaussian image; G s+1 Represents the image at the s+1th level in the Gaussian pyramid.
[0137] Detail calculation: calculate the Gaussian image G of the current scale s and the upsampled Gaussian image The difference between the two is to get the Labrador pyramid layer Lap s :
[0138]
[0139] Final layer processing: The lowest resolution Gaussian image is directly used as the top layer of the Laplacian pyramid:
[0140] Lap top =G top
[0141] Among them, Lap top represents the top layer of the Laplace pyramid; G top Represents the bottom layer of the Gaussian pyramid.
[0142] (3) Laplacian pyramid representation: The Laplacian pyramid Lap is an image sequence that contains detailed information from high to low resolution:
[0143]
[0144] The loss function of the multi-scale diffusion model is as follows:
[0145] Diffusion loss Measure the restoration effect of the diffusion model at each scale to ensure that the diffusion model can effectively remove noise and restore the image at each scale:
[0146]
[0147] Among them, N s represents the total number of scales s; Represents the image of scale s at step t; ∈ noise represents noise; T represents the total number of diffusion steps.
[0148] Multi-scale consistency loss Ensure that the image features at each scale are semantically consistent with the text features:
[0149]
[0150] Among them, N s Represents the total number of scales s.
[0151] Multiscale diffusion loss
[0152]
[0153] in, Indicates lighting consistency. By comparing lighting parameters at different viewing angles, it ensures lighting consistency and avoids lighting inconsistency issues. Indicates loss of quality; Represents semantic consistency, ensuring that the image features at each scale are semantically consistent with the text features.
[0154] The input of the illumination consistency module is the initial image and the restored perspective image, and the output is the illumination consistency score and the initial panorama. It is important to note that the fully connected output is a perspective image with illumination consistency, and the perspective images from multiple perspectives need to be stitched together to form a complete panorama.
[0155] Before generating a panorama, the Lighting Consistency module uses Neural Radiance Field (NeRF) technology to estimate the lighting conditions for each viewpoint and ensure lighting consistency across all viewpoints. This step prevents inconsistent lighting in different areas of the generated panorama.
[0156] The specific process of the lighting consistency module includes:
[0157] (1) Feature extraction: Use CNN to extract the feature vector F of each image i :F i =CNN(I perspective,i ), where I perspective,i is the perspective image of the i-th viewing angle. There are multiple perspective images, which are represented as follows:
[0158] {I perspectivei , C i |i=1,2,...,N v}
[0159] Among them, C i Indicates the corresponding camera parameters; N v Indicates the number of different viewing angles.
[0160] (2) Volume rendering: For each view angle i, sample multiple points t along the ray r j , calculate the volume density σ and color c of each point:
[0161] σ j ,c j =NeRF(r(t j ))
[0162] Among them, r(t j ) represents the j-th sampling point on the ray r; NeRF represents the rendering method using neural radiance field.
[0163] (3) Lighting parameter estimation: Use volume rendering results to estimate the lighting parameters of each view (such as ambient light, directional light, etc.):
[0164] Light i =LightEstimation(F i ,σ j , c j )
[0165] Define a loss function to measure the difference between the volume rendering result and the lighting model:
[0166]
[0167] Among them, Model(Light i , r(t j )) represents the illumination model at point r(t j ) is the predicted color at .
[0168] Optimize lighting parameters using gradient descent:
[0169]
[0170] Among them, η light Represents the learning rate; Light parameter Light i It can be expressed as:
[0171] Light i ={Light ambient ,Light directional ,...}.
[0172] (4) Perspective consistency loss Light i : Ensure that lighting parameters are consistent across all viewing angles:
[0173]
[0174] Among them, Light i and Light j Represents lighting parameters at different viewing angles.
[0175] (5) Joint lighting optimization Jointly optimize the illumination consistency loss and NeRF's volume rendering loss
[0176]
[0177] in, represents the volume rendering loss of NeRF; λ c-light Represents the weight coefficient.
[0178] The input of the semantic consistency module is the initial panorama and input semantics, and the output is semantic consistency and output panorama. The specific process is as follows:
[0179] (1) First, a large language model (LLM) is used to extract key semantic information from the text prompt. This step can be expressed as:
[0180] t key,out =LLM(T key,in )
[0181] Among them, T key,in Indicates the text prompt for input; t key,out Represents the extracted text feature vector, which contains key information such as object category, color, and location.
[0182] (2) At the same time, extract the feature vector from the initial image. This step can be expressed as:
[0183] f i =CNN(I i )
[0184] Among them, I i represents the i-th initial image, f i Represents the extracted image feature vector.
[0185] (3) Then perform feature fusion, that is, text feature t key,out With image feature f i Fusion, ensuring that the generated image is semantically consistent with the textual prompt:
[0186] f′ i =βf i , +(1–β)t key,out
[0187] Among them, β represents a weight parameter used to balance the contribution of image features and text features.
[0188] (4) Semantic consistency check: Before generating the panorama, a classifier or discriminator is used to check whether the generated image is semantically consistent with the text prompt:
[0189] Score(I′ i )=Classifier(I′ i )
[0190] Among them, I′ i Represents the fused image; Score(I′ i ) represents the consistency score between the image and the text prompt. If the consistency score is low, it means that the generated image is not semantically consistent with the text prompt. If the consistency score is lower than the threshold ∈ score , you can adjust the fusion weight; if it is lower than the threshold ε score , then regenerate the image:
[0191] β′=β+λ adjust (1-Score(I′ i ))
[0192] Among them, β′ represents the updated weight parameter; β is the weight parameter before the update; λ adjustRepresents the adjustment factor, which is used to control the magnitude of the adjustment (usually a small integer, such as 0.1).
[0193] (5) Semantic consistency loss function:
[0194] Feature Difference Loss Ensure that image features and text features are semantically consistent, and reduce the semantic differences between generated images and text prompts.
[0195]
[0196] Among them, f i represents the extracted image feature vector; t key,out Represents the extracted text feature vector, which contains key information such as object category, color, and location.
[0197] Classifier consistency loss The semantic consistency between the generated image and the text prompt is measured by the classifier's score to ensure that the generated image conforms to the text prompt.
[0198]
[0199] Among them, Score(I′ i ) represents the consistency score between the image and the text hint.
[0200] Semantic Segmentation Consistency Loss Use the semantic segmentation model to evaluate the semantic consistency of different regions in the generated image to ensure that the semantics of different regions in the generated image are consistent with the description in the text prompt, and avoid generating irrelevant objects or regions.
[0201]
[0202] in, represents the image of the kth region; represents the semantics corresponding to the kth region image; K represents the total number of categories in the semantic segmentation region.
[0203] Semantic consistency loss
[0204]
[0205] Among them, λ cls and λ seg Represents the weight parameter.
[0206] Loss function optimization: by minimizing the comprehensive semantic consistency loss function This ensures that the generated image is consistent with the text prompt in multiple levels such as feature space, classifier score, and semantic segmentation. The optimization process can use gradient descent:
[0207]
[0208] Among them, θ semantic η represents the parameter of the semantic consistency model; semantic Represents the learning rate.
[0209] After the final initial image passes through this system, a semantically consistent output panorama is generated.
[0210] Example 2
[0211] This embodiment uses a no-reference image quality assessment indicator to assess image quality. If the assessment result does not meet the requirements, a regeneration or restoration process is automatically triggered to ensure that the final generated image meets expectations in terms of quality and semantics.
[0212] (1) NIQE: NIQE is a no-reference image quality assessment metric based on natural scene statistics. It evaluates image quality by comparing the statistical features of the input image with those of high-quality images.
[0213] 1) Image preprocessing: Convert the image into a grayscale image and then normalize the image.
[0214] 2) Feature extraction: Extracting natural scene statistical features of the image, such as Gaussian model parameters.
[0215] 3) Quality score: Use the pre-trained NIQE model to calculate the difference between the statistical features of the input image and the high-quality image to obtain the quality score.
[0216] (2) BRISQUE: BRISQUE is a no-reference image quality assessment metric based on image spatial features. It evaluates image quality by analyzing the statistical characteristics of the image.
[0217] 1) Image preprocessing: Convert the image into a grayscale image and then normalize the image.
[0218] 2) Feature extraction: Extract statistical features of the image, such as mean, variance, skewness, kurtosis, etc.
[0219] 3) Quality score: Use the pre-trained BRISQUE model to calculate the quality score of the input image.
[0220] (3) Image quality loss: Use no-reference image quality assessment metrics (such as NIQE and BRISQUE) to evaluate image quality and ensure that the generated images meet expectations in terms of visual quality:
[0221]
[0222] Example 3
[0223] This embodiment provides two training methods to train the system of the present invention.
[0224] (1) Each part can calculate the loss function separately and combine them together through pre-training of different modules, such as perspective consistency loss Dynamically adjust rotation matrix loss Multiscale diffusion loss Joint lighting optimization Semantic consistency loss Loss of image quality The advantage of this approach is that when a module is unavailable, it can be directly removed without retraining the entire network.
[0225] (2) Another method is to train the entire network as a whole, and its loss function is as follows:
[0226]
[0227] in, represents the reconstruction loss, which measures the pixel-level difference between the generated image and the real panorama; represents the perceptual loss, which measures the difference between the generated image and the real panorama in the high-level feature space; Represents style loss, ensuring that the style (such as texture) of the generated image is consistent with the real panorama Figure 1 To; Represents the total variation loss, smoothing the generated image and reducing noise; represents the temporal consistency loss (optional), which is applied to the video input to ensure that the panorama parts generated by adjacent frames transition smoothly in time; recon ,λ perc ,λ style ,λ tv ,λ temp Represents the weight of each sub-loss, which is used to balance the impact of different losses.
[0228] Reconstruction loss Measures the pixel-level difference between the generated image and the real panorama.
[0229]
[0230] Among them, I gen Represents the generated panorama, the shape of which is ((H, W, C), where H is the height, W is the width, and C is the number of channels (such as C = 3 for RGB images); I gt represents the real panorama; o, p, q represent the row, column, and channel indices of the pixel respectively.
[0231] (3) Perceptual loss The high-level features of the generated image and the real panorama are extracted through a pre-trained convolutional neural network (such as VGG), and the difference between the features is calculated:
[0232]
[0233] Among them, φ l N represents the feature extractor of the first layer of the pre-trained network (such as VGG), l represents the total number of elements in the feature map of layer l, represents the L2 norm (i.e., Euclidean distance).
[0234] Style Loss The Gram matrix is used to measure the texture style difference between the generated image and the real panorama.
[0235]
[0236] Among them, G l Represents the Gram matrix of the l-th layer feature map, Gram matrix G l Defined as G l (I) = φ l (I) T φ l (I), N gram Represents the total number of elements in the Gram matrix; represents the Frobenius norm.
[0237] Total variational loss Smooth the generated image and reduce noise.
[0238]
[0239] Among them, I gen (δ, η) is the value of the generated image at pixel (δ, η).
[0240] Timing consistency loss Ensure that the panorama parts generated by adjacent frames in the video input have a smooth transition in time.
[0241]
[0242] Among them, I gen (t) represents the panorama generated at the t-th frame.
[0243] Example 4
[0244] After outputting the panorama, this embodiment connects a diffusion model and connects the "input text" to the diffusion model through Transformer (as a condition), and adds a super-resolution network to the output to improve the quality of the final panorama. The structure is as follows Figure 2 shown.
[0245] Example 5
[0246] This embodiment also provides a high-precision panoramic image generation method, which includes the following steps: generating an initial perspective image based on an input initial image; generating a restored perspective image based on the initial perspective image and input text; generating an illumination consistency score and an initial panoramic image based on the initial image and the restored perspective image; and outputting a panoramic image based on the initial panoramic image and input semantics.
[0247] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A high-precision panoramic image generation system, characterized in that: include: Image projection module, deformation repair module, lighting consistency module and semantic consistency module; The image projection module is used to generate an initial perspective image based on the initial image; The deformation repair module is used to generate a repaired perspective image based on the initial perspective image and input text; The lighting consistency module is used to generate a lighting consistency score and an initial panoramic image based on the initial image and the restored perspective image; The semantic consistency module is used to output a panoramic image based on the initial panoramic image and input semantics.
2. The high-precision panoramic image generation system according to claim 1, characterized in that: The workflow of the image projection module includes: projecting the initial image onto a unit sphere, generating the initial perspective image by defining the vertices of each image pixel and creating edges between adjacent pixels: Among them, (X, Y, Z) represents a point on the unit sphere; (x proj ,y proj ) represents a point on the perspective plane; f length Indicates focal length.
3. The high-precision panoramic image generation system according to claim 1, characterized in that: The deformation repair module uses a rotation matrix to ensure that repair is performed only on non-overlapping areas between adjacent views; the rotation matrix R includes: R=RotationMatrix(r)=R x (i x )·R y (i y )·R z (i z ) Among them, RotationMatrix represents the function of converting Euler angles into rotation matrices; the Euler angle of rotation around the x-axis is θ x , the corresponding matrix The Euler angle of rotation around the y-axis is θ y , the corresponding matrix The Euler angle of rotation around the z-axis is θ z , the corresponding matrix 4. The high-precision panoramic image generation system according to claim 1, characterized in that: The deformation repair module introduces a multi-scale diffusion model to gradually repair the image at different resolutions. The process includes: Initial low-resolution image generation: I low =Downsample(I original ) Among them, I original Represents the original image; I low Represents the low-resolution image after downsampling; Stepwise Diffusion Inpainting: Apply a diffusion model at each scale to inpaint, gradually improving the resolution; for each size s, from low to high, perform a diffusion process and upsampling: in, Represents the image of step t; ∈ noise represents noise; T represents the total number of diffusion steps; Combine multi-scale inpainted images into the final perspective image: Lap s =LaplacianPyramid(I scale ) Among them, w s represents the weight assigned to the Laplacian pyramid of each scale s, Lap s Represents a Laplacian pyramid containing multi-level detail information; I scale Indicates the repaired image; Indicates the use of CNN to extract features of each layer in the Laplacian pyramid.
5. The high-precision panoramic image generation system according to claim 4, characterized in that: The multi-scale diffusion model adopts a Laplace pyramid model; the Laplace pyramid model is constructed by Gaussian pyramids. Pyramids and Laplacian pyramids are constructed to integrate image details at different scales; The construction process includes: Gaussian pyramid construction: Gaussian pyramid generates images of different resolutions through gradual downsampling and Gaussian filtering; Construction of Laplacian Pyramid: Laplacian Pyramid obtains detail information by subtracting Gaussian pyramid images at adjacent scales; Laplacian pyramid representation: Laplacian pyramid Lap is an image sequence containing detail information from high to low resolution: Among them, N s Represents the total number of scales s.
6. The high-precision panoramic image generation system according to claim 1, characterized in that: The workflow of the lighting consistency module includes: Feature extraction: Use CNN to extract the feature vector F of each image i :F i =CNN(I perspective,i ), where I perspective,i is the perspective image of the i-th view, expressed as follows: {I perspective,i ,C i |i=1,2,...,N v } Among them, C i Indicates the corresponding camera parameters; N v Indicates the number of different perspectives; Volume rendering: For each view angle i, sample several points t along the ray r j , calculate the volume density σ and color c of each point: σ j ,c j =NeRF(r(t j )) Among them, r(t j ) represents the j-th sampling point on the ray r; NeRF represents the rendering method using neural radiance field; Lighting parameter estimation: Use volume rendering results to estimate the lighting parameters Light for each view i i : Light i =LightEstimation(F i ,σ j ,c j )。 7. The high-precision panoramic image generation system according to claim 6, characterized in that: Before generating the initial panoramic image, the lighting consistency module uses neural radiance field technology to estimate the lighting conditions of each perspective and ensure lighting consistency across all perspectives.
8. The high-precision panoramic image generation system according to claim 1, characterized in that: The workflow of the semantic consistency module includes: First, a large language model is used to extract key semantic information about the input semantics from the text prompt: t key,out =LIM(T key,in ) Among them, T key,in Indicates the text prompt for input; t key,out Represents the extracted text feature vector; At the same time, feature vectors are extracted from the initial image: f i =CNN(I i ) Among them, I i represents the i-th initial image, f i represents the extracted image feature vector; Then perform feature fusion and transform the text feature t key,out With image feature f i Fusion: f′ i =βf i ,+(1-β)t key,out Among them, β represents a weight parameter.
9. The high-precision panoramic image generation system according to claim 8, characterized in that: Before generating the panorama, a classifier or discriminator is used to check whether the generated image is semantically consistent with the textual hint: Score(I′ i )=Classifier(I′ i ) Among them, I′ i Represents the fused image; Score(I′ i ) represents the consistency score between the image and the text hint.
10. A high-precision panoramic image generation method, the method being applied to the system according to any one of claims 1 to 9, characterized in that the steps include: Generate an initial perspective image based on the input initial image; generating a restored perspective image based on the initial perspective image and input text; generating a lighting consistency score and an initial panorama based on the initial image and the restored perspective image; A panoramic image is output based on the initial panoramic image and input semantics.