Method, device and computer device for generating stereoscopic panoramic image

CN116012432BActive Publication Date: 2026-09-11GUANGZHOU NANTIAN COMP SYST +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310087696.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-18
Publication Date
2026-09-11
Estimated Expiration
2043-01-18

AI Technical Summary

Technical Problem

[0004]因此,传统技术中存在对立体全景图像的生成效率不高的问题

Benefits of technology

[0045] The aforementioned method, apparatus, computer equipment, storage medium, and computer program product for generating stereoscopic panoramic images acquire an image sequence obtained by performing panoramic shooting operations on a target shooting scene; fuse at least two consecutive images to be processed in the image sequence to obtain a fused image; input the fused image into a pre-trained monocular depth estimation model to obtain a depth image corresponding to the fused image; perform edge sharpening processing on the depth image based on the depth information data corresponding to the depth image to obtain a sharpened depth image; determine the edge contour data corresponding to the fused image based on the sharpened depth image; segment the fused image into a foreground region and a background region based on the edge contour data; and generate a stereoscopic panoramic image of the target shooting scene based on the background edge pixel information corresponding to the background region. Thus, it achieves the fusion of consecutive images to be processed in an image sequence into a single image, enabling panoramic image synthesis of a target shooting scene. Based on obtaining a two-dimensional fused image, it converts the two-dimensional fused image into a stereoscopic panoramic image, generating a stereoscopic panoramic image without the need for other tools, thereby improving the generation efficiency of stereoscopic panoramic images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116012432B_ABST
    Figure CN116012432B_ABST
Patent Text Reader

Abstract

The application relates to a method and device for generating a stereoscopic panoramic image, computer equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring an image sequence obtained by performing a panoramic shooting operation on a target shooting scene, fusing at least two to-be-processed images connected in the image sequence to obtain a fused image; inputting the fused image into a pre-trained monocular depth estimation model to obtain a depth image corresponding to the fused image; performing edge sharpening processing on the depth image according to depth information data corresponding to the depth image to obtain a sharpened depth image; determining edge contour data corresponding to the fused image according to the sharpened depth image; segmenting the fused image into a foreground region and a background region according to the edge contour data; and generating a stereoscopic panoramic image for the target shooting scene according to the foreground region and the background region. The method can improve the generation efficiency of the stereoscopic panoramic image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, computer device, storage medium, and computer program product for generating stereoscopic panoramic images. Background Technology

[0002] With the innovative development of science and technology, people are obtaining social information in their daily lives from text to images and videos, and visual materials are derived from real-world photography or created using computer software.

[0003] Currently, when filming certain scenes, it's often impossible to display a panoramic view of the target scene from a single image. Furthermore, when shooting with a camera or photo editing software, the resulting two-dimensional images cannot show the three-dimensional features of the objects within the target scene. Obtaining a stereoscopic panoramic image of the target scene through video recording requires specialized software tools, demanding a high level of technical skill from the user.

[0004] Therefore, traditional technologies suffer from low efficiency in generating stereoscopic panoramic images. Summary of the Invention

[0005] Therefore, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for generating stereoscopic panoramic images that can improve the generation efficiency of stereoscopic panoramic images, in order to address the above-mentioned technical problems.

[0006] A method for generating a stereoscopic panoramic image, characterized in that the method includes:

[0007] Obtain an image sequence obtained by performing panoramic shooting on the target shooting scene, and fuse at least two consecutive images to be processed in the image sequence to obtain the fused image;

[0008] The fused image is input into a pre-trained monocular depth estimation model to obtain the depth image corresponding to the fused image;

[0009] Based on the depth information data corresponding to the depth image, the edge sharpening process is performed on the depth image to obtain the sharpened depth image;

[0010] Based on the sharpened depth image, determine the edge contour data corresponding to the fused image;

[0011] Based on the edge contour data, the fused image is segmented into foreground and background regions;

[0012] Based on the background edge pixel information corresponding to the background area, a stereoscopic panoramic image of the target shooting scene is generated.

[0013] In one embodiment, at least two consecutive images to be processed in an image sequence are fused to obtain a fused image, including:

[0014] Identify at least two adjacent images to be processed in the image sequence;

[0015] Extract image features from at least two images to be processed to obtain image feature matching points corresponding to at least two images to be processed;

[0016] Based on image feature matching points, determine the registration structure between at least two images to be processed;

[0017] Based on the registration structure, at least two images to be processed are fused to obtain the fused image.

[0018] In one embodiment, determining the registration structure between at least two images to be processed based on image feature matching points includes:

[0019] The homography matrix is ​​calculated for the image feature matching points to obtain the target image feature matching points; the target image feature matching points are the image feature matching points other than the abnormal image feature matching points.

[0020] The perspective transformation matrix is ​​calculated for the feature matching points of the target image to determine the registration structure between at least two images to be processed.

[0021] In one embodiment, at least two images to be processed are fused according to a registration structure to obtain a fused image, including:

[0022] Based on the registration structure, at least two images to be processed are registered to obtain the registered image;

[0023] The registered image is segmented into at least one block image;

[0024] The boundaries of each image block are repaired to obtain the repaired image;

[0025] Feature fusion is performed on the stitched areas of the repaired image to obtain the fused image.

[0026] In one embodiment, determining the edge contour data corresponding to the fused image based on the sharpened depth image includes:

[0027] Thresholding is applied to the sharpened depth image to obtain continuous edges and spots;

[0028] Mark continuous edges and blobs as binary graphs;

[0029] Based on the binary image, continuous edges and spots with fewer than a preset number of pixels are removed to obtain the removed image data.

[0030] Image similarity measurement is performed on the removed image data to obtain edge contour data.

[0031] In one embodiment, the background edge pixel information includes background edge color information and background edge depth information. Based on the background edge pixel information corresponding to the background region, a stereoscopic panoramic image of the target shooting scene is generated, including:

[0032] The background edge color information is input into a pre-trained color restoration network model to obtain the restored edge color information, and the background edge depth information is input into a pre-trained depth restoration network model to obtain the restored edge depth information.

[0033] The repair edge is generated based on the repair edge color information and repair edge depth information;

[0034] Based on the repaired edges, a stereoscopic panoramic image of the target shooting scene is synthesized.

[0035] A device for generating stereoscopic panoramic images, characterized in that the device comprises:

[0036] The fusion module is used to acquire the image sequence obtained by performing panoramic shooting operation on the target shooting scene, and fuse at least two consecutive images to be processed in the image sequence to obtain the fused image;

[0037] The input module is used to input the fused image into a pre-trained monocular depth estimation model to obtain the depth image corresponding to the fused image;

[0038] The processing module is used to perform edge sharpening processing on the depth image based on the depth information data corresponding to the depth image, so as to obtain a sharpened depth image.

[0039] The determination module is used to determine the edge contour data corresponding to the fused image based on the sharpened depth image;

[0040] The segmentation module is used to segment the fused image into foreground and background regions based on edge contour data;

[0041] The generation module is used to generate a stereoscopic panoramic image of the target shooting scene based on the background edge pixel information corresponding to the background area.

[0042] A computer device includes a memory and a processor, the memory storing a computer program, characterized in that the processor executes the computer program to implement the steps of the method described above.

[0043] A computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the above-described method.

[0044] A computer program product includes a computer program, characterized in that the computer program, when executed by a processor, implements the steps of the above-described method.

[0045] The aforementioned method, apparatus, computer equipment, storage medium, and computer program product for generating stereoscopic panoramic images acquire an image sequence obtained by performing panoramic shooting operations on a target shooting scene; fuse at least two consecutive images to be processed in the image sequence to obtain a fused image; input the fused image into a pre-trained monocular depth estimation model to obtain a depth image corresponding to the fused image; perform edge sharpening processing on the depth image based on the depth information data corresponding to the depth image to obtain a sharpened depth image; determine the edge contour data corresponding to the fused image based on the sharpened depth image; segment the fused image into a foreground region and a background region based on the edge contour data; and generate a stereoscopic panoramic image of the target shooting scene based on the background edge pixel information corresponding to the background region. Thus, it achieves the fusion of consecutive images to be processed in an image sequence into a single image, enabling panoramic image synthesis of a target shooting scene. Based on obtaining a two-dimensional fused image, it converts the two-dimensional fused image into a stereoscopic panoramic image, generating a stereoscopic panoramic image without the need for other tools, thereby improving the generation efficiency of stereoscopic panoramic images. Attached Figure Description

[0046] Figure 1 This is an application environment diagram of a method for generating stereoscopic panoramic images in one embodiment;

[0047] Figure 2 This is a flowchart illustrating a method for generating a stereoscopic panoramic image in one embodiment;

[0048] Figure 3 This is a flowchart illustrating a stereoscopic panoramic image generation process in one embodiment;

[0049] Figure 4 This is a screenshot of a stereoscopic panoramic image from one embodiment.

[0050] Figure 5 This is a schematic diagram showing the position of a continuous edge and a spot in an image in one embodiment;

[0051] Figure 6 This is a flowchart illustrating a method for generating a stereoscopic panoramic image in another embodiment;

[0052] Figure 7 This is a structural block diagram of a device for generating stereoscopic panoramic images in one embodiment;

[0053] Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0055] The method for generating stereoscopic panoramic images provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Server 104 acquires an image sequence obtained by performing panoramic shooting on the target scene, fuses at least two consecutive images in the image sequence to obtain a fused image; server 104 inputs the fused image into a pre-trained monocular depth estimation model to obtain a depth image corresponding to the fused image; server 104 performs edge sharpening processing on the depth image based on the depth information data corresponding to the depth image to obtain a sharpened depth image; server 104 determines the edge contour data corresponding to the fused image based on the sharpened depth image; server 104 segments the fused image into a foreground region and a background region based on the edge contour data; server 104 generates a stereoscopic panoramic image of the target scene based on the background edge pixel information corresponding to the background region. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle systems. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0056] In one embodiment, such as Figure 2 As shown, a method for generating stereoscopic panoramic images is provided, which can be applied to... Figure 1 Taking server 104 as an example, the following steps are included:

[0057] Step S202: Obtain the image sequence obtained by performing panoramic shooting operation on the target shooting scene, and fuse at least two connected images to be processed in the image sequence to obtain the fused image.

[0058] The target shooting scene can be the specific scene to be shot when performing panoramic shooting.

[0059] An image sequence can be a sequence of images captured consecutively when photographing a specific scene. For example, when photographing the facade of a coffee shop, multiple consecutive images of the coffee shop facade constitute an image sequence corresponding to the coffee shop facade. Adjacent images in an image sequence can have overlapping areas, and adjacent images can be two consecutive images captured by the camera device in a fixed shooting direction.

[0060] The image to be processed can be any image in an image sequence.

[0061] The fused image can be the image obtained by fusing the images in the image sequence.

[0062] In practice, the server acquires an image sequence of a specific scene captured by an image acquisition device, and then merges the images in the image sequence into a single image, i.e., generates the merged image.

[0063] Step S204: Input the fused image into the pre-trained monocular depth estimation model to obtain the depth image corresponding to the fused image.

[0064] The monocular depth estimation model can be a neural network model used to extract image depth values. For example, the monocular depth estimation model can be a monocular depth estimation model trained based on the ResNet-50 (a convolutional neural network model) depth residual network model.

[0065] Among them, a depth image can be an image in which the distance values ​​(depth values) of each point in a specific scene are captured by an image acquisition device as pixel values.

[0066] In practice, the server inputs the fused image into a pre-trained monocular depth estimation model to obtain the depth image corresponding to the fused image, that is, to obtain the depth value corresponding to each pixel in the fused image.

[0067] Step S206: Based on the depth information data corresponding to the depth image, perform edge sharpening processing on the depth image to obtain a sharpened depth image.

[0068] The depth information data can be the depth value corresponding to each pixel in the depth image.

[0069] The sharpened depth image can be a depth image obtained by sharpening a depth image.

[0070] In practice, the server performs bidirectional median filtering calculations based on the depth image information data corresponding to the fused image to obtain the sharpened depth image.

[0071] Step S208: Determine the edge contour data corresponding to the fused image based on the sharpened depth image.

[0072] Among them, edge contour data can refer to the positional data of the edge contour of an object in the image.

[0073] In practice, the server determines the edge contour data of the fused image based on the sharpened image.

[0074] Step S210: Based on the edge contour data, the fused image is segmented into a foreground region and a background region.

[0075] In practice, the server divides the fused image into two regions, a foreground region and a background region, based on the depth values ​​corresponding to the obtained edge contour data.

[0076] Step S212: Generate a stereoscopic panoramic image of the target shooting scene based on the background edge pixel information corresponding to the background area.

[0077] In practice, the server generates repaired edge information corresponding to the fused image based on the background edge pixel information corresponding to the background area, and then synthesizes a stereoscopic panoramic image of the target shooting scene based on the repaired edge information.

[0078] In practical applications, the process of generating stereoscopic panoramic images includes two parts: the fusion of two-dimensional images and the generation of three-dimensional images. Figure 3 An exemplary flowchart of the process for generating a stereoscopic panoramic image is provided.

[0079] In the process of generating this stereoscopic panoramic image, the fusion of two-dimensional images is first completed. By extracting the image feature points of each image in the image sequence, matching the feature points between each image, and then calculating the registration structure between the images based on the feature points, the images are fused according to the registration structure between the images, and the boundary cracks generated after fusion are repaired. The above process is repeated to complete the fusion of continuous images in the image sequence, and the fused image is obtained. In this way, the fusion of two-dimensional images is completed.

[0080] After fusing the 2D images, the generation of the 3D image is required. By calculating the depth information of the fused image, a layered depth image corresponding to the fused image is generated. Then, based on the layered depth image, the edge contours of the fused image are calculated. Further calculations are made of the composite color and depth values ​​for edge-occluded areas. These composite color and depth values ​​are then merged into the generated layered depth image. This process is repeated to synthesize pixels at all boundaries in the fused image. Finally, the final synthesized pixels are converted into a 3D mesh and rendered using a rendering model to create a new image, i.e., a stereoscopic panoramic image. The stereoscopic panoramic image is a dynamic effect image; see [link to documentation]. Figure 4 , Figure 4 It includes the fused two-dimensional image 402, and also includes a dynamic screenshot 404 of the stereoscopic panoramic image generated from the two-dimensional image 402.

[0081] The aforementioned method, apparatus, computer equipment, storage medium, and computer program product for generating stereoscopic panoramic images acquire an image sequence obtained by performing panoramic shooting operations on a target shooting scene; fuse at least two consecutive images to be processed in the image sequence to obtain a fused image; input the fused image into a pre-trained monocular depth estimation model to obtain a depth image corresponding to the fused image; perform edge sharpening processing on the depth image based on the depth information data corresponding to the depth image to obtain a sharpened depth image; determine the edge contour data corresponding to the fused image based on the sharpened depth image; segment the fused image into a foreground region and a background region based on the edge contour data; and generate a stereoscopic panoramic image of the target shooting scene based on the background edge pixel information corresponding to the background region. Thus, it achieves the fusion of consecutive images to be processed in an image sequence into a single image, enabling panoramic image synthesis of a target shooting scene. Based on obtaining a two-dimensional fused image, it converts the two-dimensional fused image into a stereoscopic panoramic image, generating a stereoscopic panoramic image without the need for other tools, thereby improving the generation efficiency of stereoscopic panoramic images.

[0082] In another embodiment, fusing at least two connected images to be processed in an image sequence to obtain a fused image includes: identifying at least two connected images to be processed in the image sequence; extracting image features from the at least two images to be processed to obtain image feature matching points corresponding to the at least two images to be processed; determining a registration structure between the at least two images to be processed based on the image feature matching points; and fusing the at least two images to be processed based on the registration structure to obtain a fused image.

[0083] Image features can refer to the color features, texture features, shape features, and spatial relationship features of an image.

[0084] Among them, image feature matching points can be points with the same name between two or more images, and the feature information of each feature point in the image feature matching points is similar.

[0085] The registration structure can refer to the correspondence between image feature matching points.

[0086] In practice, after the server identifies two connected images to be processed in the image sequence, it extracts the image features of the two images to obtain the image feature matching points corresponding to the two images. Based on the image feature matching points, the server determines the registration structure between the two images to be processed. Based on the registration structure, the server fuses the two images to be processed.

[0087] In practical applications, after identifying two adjacent images to be processed in an image sequence, the server extracts feature points from the two images using the SURF algorithm (Speeded-Up Robust Features, a robust local feature point detection and description algorithm). Then, it calculates matching feature points between the two images using the K-Nearest Neighbor algorithm (a classification algorithm). The server obtains image space coordinate transformation parameters from these matching feature points and performs image registration based on these parameters, thus completing the fusion of the two images. After fusing two adjacent images in the image sequence, the fused image needs to be fused with other images in the sequence until all images in the sequence are merged into a single image.

[0088] The technical solution of this embodiment extracts image features from connected images to be processed in an image sequence to obtain corresponding image feature matching points between the images to be processed, thereby determining the registration structure between the images to be processed. Based on the registration structure between the images to be processed, the images to be processed are fused to obtain a fused image. This can fuse the images in the image sequence into a single image, realizing the acquisition of a panoramic image of the target scene, which is beneficial for the generation of stereoscopic panoramic images and improves the generation efficiency of stereoscopic panoramic images.

[0089] In another embodiment, determining the registration structure between at least two images to be processed based on image feature matching points includes: performing homography matrix calculation on the image feature matching points to obtain target image feature matching points; the target image feature matching points are other image feature matching points besides abnormal image feature matching points; and performing perspective transformation matrix calculation on the target image feature matching points to determine the registration structure between at least two images to be processed.

[0090] The homography matrix can refer to the projection matrix from one plane to another. In practical applications, the homography matrix can characterize the mapping relationship between corresponding pixel positions in two images.

[0091] Among them, the target image feature matching point can be the image feature matching point determined by the homography matrix calculation.

[0092] Among them, abnormal image feature matching points can be image feature matching points that have been eliminated after calculation using the homography matrix.

[0093] In practical applications, image feature matching points are input into the RANSAC algorithm (Random Sample Consensus, an algorithm for detecting outliers in data) to calculate the homography matrix. Based on the calculation results, the calculated inliers are used as target image feature matching points, and the calculated outliers are used as outlier image feature matching points. Target image feature matching points refer to correct image feature matching points, while outlier image feature matching points refer to invalid or noisy image feature matching points.

[0094] The perspective transformation matrix represents the transformation relationship between the image before and after perspective. In practical applications, the perspective transformation matrix can refer to the pixel transformation relationship between two images.

[0095] In the specific implementation, the server uses the RANSAC algorithm to calculate the homography matrix, obtains the inlier data, i.e., the target feature matching points, and removes the outlier data, i.e., removes abnormal image feature matching points. The server inputs the obtained inlier data into the DLT algorithm (Direct Linear Transform, an algorithm that establishes a direct linear relationship between the image point coordinate system and the corresponding object point object space coordinates) to calculate the perspective transformation matrix, thereby determining the registration structure between the two images to be processed.

[0096] The technical solution of this embodiment calculates the homography matrix of image feature matching points, filters among each image feature matching point pair to obtain the target image feature matching point pair, calculates the perspective transformation matrix based on the target image feature matching point pair, and determines the registration structure between the images to be processed. This can accurately fuse the images to be processed in the image sequence and improve the image fusion efficiency, thereby improving the generation efficiency of stereoscopic panoramic images.

[0097] In another embodiment, at least two images to be processed are fused according to the registration structure to obtain a fused image, including: registering at least two images to be processed according to the registration structure to obtain a registered image; segmenting the registered image into at least one block image; repairing the boundaries of each block image to obtain a repaired image; and performing feature fusion on the stitching region of the repaired image to obtain a fused image.

[0098] The registered image can be formed by matching and superimposing multiple images. Image registration refers to the process of matching and superimposing two or more images acquired at different times, with different image acquisition devices, or under different conditions (weather, illumination, camera position and angle, etc.).

[0099] Among them, a segmented image can be a small image obtained by dividing an image into multiple small blocks.

[0100] The stitching area refers to the overlapping region between two images to be stitched together. By processing the pixel data of the overlapping region between the two images, a better image stitching effect can be achieved.

[0101] In the specific implementation, the server registers the image to be processed according to the registration structure to obtain the registered image. The server uses the APAP algorithm (As-Projective-As-Possible Image Stitching, an image stitching algorithm) to segment the registered image into small images. The server repairs the seams of each small image to eliminate ghosting and cracks at the stitching point of the registered image, thus obtaining the repaired image. The server then performs feature fusion on the stitching area of ​​the repaired image to obtain the fused image.

[0102] In practical applications, after the server obtains the repaired image, in order to eliminate the differences in lighting, noise, and exposure between the two images to be processed, the server adopts the Multiband Blending strategy (an image fusion algorithm) to perform Laplacian pyramid decomposition. The server divides the image into small blocks of different sizes, performs weighted average calculations to obtain the result corresponding to each small block, and then the server performs pyramid inverse reconstruction to obtain the fused image data.

[0103] The technical solution of this embodiment registers the image to be processed according to the registration structure to obtain a registered image. The registered image is then divided into at least one block image. Boundary repair is performed based on each block image to obtain a repaired image corresponding to the fused image. Feature fusion is then performed on the stitching area of ​​the repaired image to obtain a fused image. In this way, the display of the stitching area corresponding to the generated fused image is more natural, and the fused image can be generated more accurately, which is beneficial to improving the generation quality of stereoscopic panoramic images and thus improving the generation efficiency of stereoscopic panoramic images.

[0104] In another embodiment, determining the edge contour data corresponding to the fused image based on the sharpened depth image includes: thresholding the sharpened depth image to obtain continuous edges and spots; marking the continuous edges and spots as binary images; removing continuous edges and spots with fewer than a preset number of pixels based on the binary images to obtain removed image data; and performing image similarity measurement on the removed image data to obtain edge contour data.

[0105] The regions corresponding to continuous edges and spots can be as follows: Figure 5 The image area shown.

[0106] Here, a binary image can refer to an image represented as a binary file.

[0107] In the specific implementation, the server performs thresholding on the sharpened depth image, compares the disparity of adjacent pixels in the depth image to obtain continuous edges and blobs, and marks these continuous edges and blobs as binary images. Based on the binary images, the server removes fewer than a preset number of continuous edges and blobs, obtaining the removed image data. The server then uses the LPIPS (Learned Perceptual Image Patch Similarity) algorithm to measure image similarity in the removed image data, obtaining edge contour data. In practical applications, the server can remove less than 10 pixels from the binary images to obtain the removed image data.

[0108] In practical applications, thresholding refers to removing pixels in an image whose values ​​are higher or lower than a threshold. For example, setting the threshold to 127 means setting the value of all pixels with a value greater than 127 to 255, and setting the value of all pixels with a value less than 127 to 0.

[0109] The technical solution of this embodiment obtains continuous edges and spots by thresholding the sharpened depth image. By marking the continuous edges and spots as binary images, continuous edges and spots with fewer than a preset number of pixels are removed according to the binary images to obtain the removed image data. This eliminates noise data, which is beneficial for identifying edge contours in the image. Image similarity measurement is performed on the removed image data to obtain edge contour data, which is beneficial for accurately segmenting the fused image into foreground and background regions. This improves the accuracy of edge repair in the fused image and increases the generation efficiency of stereoscopic panoramic images.

[0110] In another embodiment, the background edge pixel information includes background edge color information and background edge depth information. Generating a stereoscopic panoramic image of the target shooting scene based on the background edge pixel information corresponding to the background region includes: inputting the background edge color information into a pre-trained color inpainting network model to obtain inpainted edge color information; and inputting the background edge depth information into a pre-trained depth inpainting network model to obtain inpainted edge depth information; generating inpainted edges based on the inpainted edge color information and inpainted edge depth information; and synthesizing a stereoscopic panoramic image of the target shooting scene based on the inpainted edges.

[0111] Among them, the background edge color information can refer to the color of the pixels in the background edge area.

[0112] Among them, background edge depth information can refer to the depth value of pixels in the background edge region.

[0113] Among them, repairing edge color information can be done by repairing the color of edge pixels.

[0114] Among them, the edge depth information can be the depth value of the edge pixel.

[0115] In the specific implementation, the server inputs the background edge color information into a pre-trained color inpainting network model to generate the repaired edge color. The server inputs the background edge depth information into a pre-trained depth inpainting network model to generate the repaired edge depth value. The server performs edge inpainting based on the repaired edge color and repaired edge depth value to obtain the repaired edge. The server synthesizes a stereoscopic panoramic image of the target shooting scene based on the pixel information corresponding to the repaired edge.

[0116] To facilitate understanding by those skilled in the art, the following provides a method for generating a stereoscopic panoramic image of a target shooting scene based on the background edge pixel information corresponding to the background region.

[0117] First, based on the flooding algorithm, the edge region corresponding to the fused image is expanded by 5 pixels and filled. The context data of the edge region of the background region is synthesized to obtain the pixel data of the background edge region.

[0118] Then, the pixel data of the background edge region is input into the edge inpainting network model to generate repaired edges. This edge inpainting network model includes a color inpainting network model and a depth inpainting network model. The color of the edge region corresponding to the fused image is input into the color inpainting network model to generate the repaired color. The server inputs the depth value of the edge region corresponding to the fused image into the depth inpainting network model to generate the repaired depth value. The server repeatedly applies the color inpainting network model and the depth inpainting network model until no more repaired edges are generated. Finally, the server converts the pixels corresponding to the repaired edges into a mesh and renders the interpolated areas, thus synthesizing a stereoscopic panoramic image.

[0119] The technical solution of this embodiment obtains repair edge color information by inputting background edge color information into a pre-trained color restoration network model, which can constrain the color synthesis of the repair edge. It also obtains repair edge depth information by inputting background edge depth information into a pre-trained depth restoration network model, which can constrain the depth value synthesis of the repair edge. By generating the repair edge based on the repair edge color information and repair edge depth information, the color information and depth information corresponding to the repair edge are merged into the layered depth image corresponding to the original fused image, which can accurately generate stereoscopic panoramic images and improve the generation efficiency of stereoscopic panoramic images.

[0120] In another embodiment, such as Figure 6 As shown, a method for generating stereoscopic panoramic images is provided, which can be applied to... Figure 1Taking server 104 as an example, the following steps are included:

[0121] Step S602: Obtain the image sequence obtained by performing panoramic shooting operation on the target shooting scene, and determine at least two connected images to be processed in the image sequence.

[0122] Step S604: Extract image features from at least two images to be processed to obtain image feature matching points corresponding to at least two images to be processed.

[0123] Step S606: Determine the registration structure between at least two images to be processed based on image feature matching points.

[0124] Step S608: According to the registration structure, at least two images to be processed are fused to obtain a fused image.

[0125] Step S610: Input the fused image into the pre-trained monocular depth estimation model to obtain the depth image corresponding to the fused image.

[0126] Step S612: Based on the depth information data corresponding to the depth image, perform edge sharpening processing on the depth image to obtain a sharpened depth image.

[0127] Step S614: Determine the edge contour data corresponding to the fused image based on the sharpened depth image.

[0128] Step S616: Based on the edge contour data, the fused image is segmented into foreground and background regions.

[0129] Step S618: Generate a stereoscopic panoramic image of the target shooting scene based on the background edge pixel information corresponding to the background area.

[0130] It should be noted that the specific limitations of the above steps can be found in the specific limitations of a method for generating stereoscopic panoramic images described above.

[0131] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0132] Based on the same inventive concept, this application also provides a stereoscopic panoramic image generation apparatus for implementing the stereoscopic panoramic image generation method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations of one or more stereoscopic panoramic image generation apparatus embodiments provided below can be found in the limitations of the stereoscopic panoramic image generation method described above, and will not be repeated here.

[0133] In one embodiment, such as Figure 7 As shown, a device for generating stereoscopic panoramic images is provided, comprising:

[0134] The fusion module 702 is used to acquire an image sequence obtained by performing a panoramic shooting operation on the target shooting scene, and fuse at least two consecutive images to be processed in the image sequence to obtain a fused image;

[0135] The input module 704 is used to input the fused image into a pre-trained monocular depth estimation model to obtain the depth image corresponding to the fused image;

[0136] The processing module 706 is used to perform edge sharpening processing on the depth image based on the depth information data corresponding to the depth image, so as to obtain a sharpened depth image.

[0137] The determination module 708 is used to determine the edge contour data corresponding to the fused image based on the sharpened depth image;

[0138] The segmentation module 710 is used to segment the fused image into foreground and background regions based on edge contour data;

[0139] The generation module 712 is used to generate a stereoscopic panoramic image of the target shooting scene based on the background edge pixel information corresponding to the background area.

[0140] In one embodiment, the fusion module 702 is specifically used to determine at least two connected images to be processed in the image sequence; extract image features from the at least two images to be processed to obtain image feature matching points corresponding to the at least two images to be processed; determine the registration structure between the at least two images to be processed based on the image feature matching points; and fuse the at least two images to be processed based on the registration structure to obtain a fused image.

[0141] In one embodiment, the fusion module 702 is specifically used to perform homography matrix calculation on the image feature matching points to obtain target image feature matching points; the target image feature matching points are other image feature matching points besides abnormal image feature matching points; the target image feature matching points are used to perform perspective transformation matrix calculation on the target image feature matching points to determine the registration structure between at least two images to be processed.

[0142] In one embodiment, the fusion module 702 is specifically used to register at least two images to be processed according to the registration structure to obtain a registered image; to segment the registered image into at least one block image; to repair the boundaries of each block image to obtain a repaired image; and to perform feature fusion on the stitching region of the repaired image to obtain a fused image.

[0143] In one embodiment, the determining module 708 is specifically used to perform thresholding processing on the sharpened depth image to obtain continuous edges and spots; mark the continuous edges and spots as binary images; based on the binary images, remove continuous edges and spots with fewer than a preset number of pixels to obtain removed image data; and perform image similarity measurement on the removed image data to obtain edge contour data.

[0144] In one embodiment, the background edge pixel information includes background edge color information and background edge depth information. The generation module 712 is specifically used to input the background edge color information into a pre-trained color restoration network model to obtain restored edge color information, and to input the background edge depth information into a pre-trained depth restoration network model to obtain restored edge depth information; generate restored edges based on the restored edge color information and restored edge depth information; and synthesize a stereoscopic panoramic image of the target shooting scene based on the restored edges.

[0145] Each module in the aforementioned stereoscopic panoramic image generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0146] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores XX data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a method for generating stereoscopic panoramic images.

[0147] Those skilled in the art will understand that Figure 8The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0148] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the method for generating a stereoscopic panoramic image described above. The steps of the method for generating a stereoscopic panoramic image here can be steps from the methods for generating stereoscopic panoramic images described in the various embodiments above.

[0149] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, causes the processor to perform the steps of the method for generating a stereoscopic panoramic image described above. The steps of the method for generating a stereoscopic panoramic image here may be steps from the methods for generating stereoscopic panoramic images described in the various embodiments above.

[0150] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, causes the processor to perform the steps of the stereoscopic panoramic image generation method described above. The steps of the stereoscopic panoramic image generation method described here can be steps from the stereoscopic panoramic image generation method of the various embodiments described above.

[0151] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0152] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0153] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0154] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for generating a stereoscopic panoramic image, characterized in that, The method includes: Obtain an image sequence obtained by performing panoramic shooting on a target shooting scene, and identify at least two consecutive images to be processed in the image sequence; Extract image features from the at least two images to be processed to obtain image feature matching points corresponding to the at least two images to be processed; Based on the image feature matching points, determine the registration structure between the at least two images to be processed; According to the registration structure, the at least two images to be processed are fused to obtain a fused image; The fused image is input into a pre-trained monocular depth estimation model to obtain the depth image corresponding to the fused image; Based on the depth information data corresponding to the depth image, the depth image is subjected to edge sharpening processing to obtain a sharpened depth image; The sharpened depth image is thresholded, and the disparity of adjacent pixels in the thresholded depth image is compared to obtain continuous edges and spots. The continuous edges and the spots are marked as binary images; Based on the binary image, fewer than a preset number of continuous edges and spots are removed to obtain the removed image data. Image similarity measurement is performed on the removed image data to obtain edge contour data; based on the edge contour data, the fused image is segmented into foreground and background regions, and a layered depth image corresponding to the fused image is generated; The background area is expanded according to a preset number of pixels, and the background edge pixel information corresponding to the expanded background area is extracted. The background edge pixel information includes background edge color information and background edge depth information. The background edge color information is input into a pre-trained color restoration network model to obtain restored edge color information, and the background edge depth information is input into a pre-trained depth restoration network model to obtain restored edge depth information. Based on the repaired edge color information and the repaired edge depth information, the synthesized color and depth values ​​are determined; The synthesized color and depth values ​​are merged into the layered depth image to obtain an updated layered depth image, and the updated layered depth image is converted into a 3D mesh; Based on the three-dimensional grid, a stereoscopic panoramic image of the target shooting scene is generated.

2. The method according to claim 1, characterized in that, Determining the registration structure between the at least two images to be processed based on the image feature matching points includes: The homography matrix is ​​calculated on the image feature matching points to obtain the target image feature matching points; the target image feature matching points are the image feature matching points other than abnormal image feature matching points. The perspective transformation matrix is ​​calculated by matching the feature points of the target image to determine the registration structure between the at least two images to be processed.

3. The method according to claim 1, characterized in that, The step of fusing the at least two images to be processed according to the registration structure to obtain a fused image includes: According to the registration structure, the at least two images to be processed are registered to obtain the registered images; The registered image is segmented into at least one block image; The boundaries of each of the image blocks are repaired to obtain the repaired image; Feature fusion is performed on the stitched area of ​​the repaired image to obtain the fused image.

4. A device for generating stereoscopic panoramic images, characterized in that, The device includes: The fusion module is used to acquire an image sequence obtained by performing panoramic shooting on a target shooting scene; identify at least two connected images to be processed in the image sequence; extract image features from the at least two images to be processed to obtain image feature matching points corresponding to the at least two images to be processed; determine the registration structure between the at least two images to be processed based on the image feature matching points; and fuse the at least two images to be processed based on the registration structure to obtain a fused image. The input module is used to input the fused image into a pre-trained monocular depth estimation model to obtain the depth image corresponding to the fused image; The processing module is used to perform edge sharpening processing on the depth image based on the depth information data corresponding to the depth image, so as to obtain a sharpened depth image; The determination module is used to perform thresholding processing on the sharpened depth image, and to perform disparity comparison of adjacent pixels on the thresholded depth image to obtain continuous edges and spots; to mark the continuous edges and spots as binary images; to remove fewer than a preset number of continuous edges and spots according to the binary images to obtain removed image data; and to perform image similarity measurement on the removed image data to obtain edge contour data. The segmentation module is used to segment the fused image into a foreground region and a background region based on the edge contour data, and generate a layered depth image corresponding to the fused image; The generation module is used to expand the background region according to a preset number of pixels, extract the background edge pixel information corresponding to the expanded background region, the background edge pixel information including background edge color information and background edge depth information; input the background edge color information into a pre-trained color restoration network model to obtain restored edge color information, and input the background edge depth information into a pre-trained depth restoration network model to obtain restored edge depth information; determine the synthesized color and depth values ​​according to the restored edge color information and the restored edge depth information; merge the synthesized color and depth values ​​into the layered depth image to obtain an updated layered depth image, and convert the updated layered depth image into a three-dimensional mesh; and generate a stereoscopic panoramic image of the target shooting scene based on the three-dimensional mesh.

5. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 3.

7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Image processing method and device

    CN114359123A

  • Binocular stereo panoramic image generation method and device, equipment and storage medium

    CN114742703A

  • Image synthesis apparatus and method, and computer-readable storage medium

    WO2019200807A1