Panoramic image processing method and apparatus
By using a neural network model to perform feature transformation and distortion correction in panoramic image processing, the problem of high-cost 3D model reconstruction is solved, achieving low-cost, high-precision depth image acquisition and 3D model construction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA DAMO (HANGZHOU) TECH CO LTD
- Filing Date
- 2022-08-25
- Publication Date
- 2026-07-17
AI Technical Summary
In existing technologies, remote property or goods inspection requires high-cost depth sensors for 3D model reconstruction, resulting in high costs. Furthermore, 2D videos lack the 3D experience and cannot meet the needs of online inspection.
A neural network model is used for panoramic image processing. The panoramic image features and depth images are converted through encoding and decoding layers. Deep learning is used to extract features and correct distortion on the sphere, avoiding the use of high-cost equipment.
It enables efficient acquisition of depth data from panoramic images without the aid of specialized instruments, reducing costs and improving the accuracy of depth images. It also supports the rapid construction of 3D models and enhances the user experience.
Smart Images

Figure CN115526923B_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of computer technology, and in particular to two panoramic image processing methods. Background Technology
[0002] In industries such as real estate and foreign trade, customers often prefer to "see for themselves" and need to physically inspect the property or goods before leasing or purchasing them. However, if the customer is far away from the property or goods they wish to buy, such as in different cities or countries, and due to special factors (such as weather or unavoidable natural disasters), the customer is unable to visit the property or goods in person for on-site inspection, then there is a widespread demand for online "inspection."
[0003] Traditionally, goods inspection is typically conducted via video or live streaming. Customers can see relatively clear 2D videos, but these are only two-dimensional and do not provide a 3D experience. For the reconstruction of 3D models of real estate or goods, depth data is needed to form point clouds for modeling. Simultaneously, point clouds from different locations obtained by sensors are stitched together based on distance information from these point clouds. However, acquiring depth data for 3D model reconstruction usually requires expensive dedicated depth sensors such as infrared cameras and LiDAR, which are not only cumbersome to use but also costly, making them prohibitively expensive for large-scale industrial applications. Summary of the Invention
[0004] In view of this, embodiments of this specification provide two panoramic image processing methods. One or more embodiments of this specification also relate to a panoramic image processing apparatus, a 3D model building terminal, an augmented reality (AR) device, an augmented reality (VR) device, a computing device, a computer-readable storage medium, and a computer program, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a panoramic image processing method is provided, comprising:
[0006] The target panoramic image is input into the image processing model, and the panoramic image features of the target panoramic image are obtained through the encoding layer of the image processing model.
[0007] According to the panoramic image spherical conversion algorithm, the panoramic image features are converted into spherical image coding features;
[0008] The spherical image encoding features are input into the decoding layer of the image processing model, and the spherical depth image of the target panoramic image is obtained through the decoding layer;
[0009] According to the panoramic image spherical conversion algorithm, the spherical depth image is converted into a panoramic depth image of the target panoramic image, wherein the image processing model is a neural network model.
[0010] According to a second aspect of the embodiments of this specification, a panoramic image processing apparatus is provided, comprising:
[0011] The panoramic feature determination module is configured to input the target panoramic image into the image processing model and obtain the panoramic image features of the target panoramic image through the encoding layer of the image processing model.
[0012] The encoding feature determination module is configured to convert the panoramic image features into spherical image encoding features according to the panoramic image spherical conversion algorithm;
[0013] The spherical depth image determination module is configured to input the spherical image encoding features into the decoding layer of the image processing model, and obtain the spherical depth image of the target panoramic image through the decoding layer;
[0014] The panoramic depth image determination module is configured to convert the spherical depth image into a panoramic depth image of the target panoramic image according to the panoramic image spherical conversion algorithm, wherein the image processing model is a neural network model.
[0015] According to a third aspect of the embodiments of this specification, a panoramic image processing method is provided, comprising:
[0016] In response to a user's panoramic image processing request, an image input interface is displayed to the user;
[0017] Receive the target panoramic image input by the user through the image input interface;
[0018] The target panoramic image is input into the image processing model, and the panoramic image features of the target panoramic image are obtained through the encoding layer of the image processing model.
[0019] According to the panoramic image spherical conversion algorithm, the panoramic image features are converted into spherical image coding features;
[0020] The spherical image encoding features are input into the decoding layer of the image processing model, and the spherical depth image of the target panoramic image is obtained through the decoding layer;
[0021] According to the panoramic image spherical conversion algorithm, the spherical depth image is converted into a panoramic depth image of the target panoramic image, and the panoramic depth image of the target panoramic image is displayed to the user through the image input interface. The image processing model is a neural network model.
[0022] According to a fourth aspect of the embodiments of this specification, a three-dimensional model building terminal is provided, comprising:
[0023] Memory and processor;
[0024] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, perform the following steps:
[0025] The target panoramic image of the target object is input into the image processing model, and the panoramic image features of the target panoramic image are obtained through the encoding layer of the image processing model.
[0026] According to the panoramic image spherical conversion algorithm, the panoramic image features are converted into spherical image coding features;
[0027] The spherical image encoding features are input into the decoding layer of the image processing model, and the spherical depth image of the target panoramic image is obtained through the decoding layer;
[0028] According to the panoramic image spherical conversion algorithm, the spherical depth image is converted into a panoramic depth image of the target panoramic image;
[0029] Based on the panoramic depth image of the target panoramic image, a three-dimensional model of the target object is constructed, wherein the image processing model is a neural network model.
[0030] According to a fifth aspect of the embodiments of this specification, an augmented reality (AR) device is provided, comprising:
[0031] Memory and processor;
[0032] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the above-described panoramic image processing method.
[0033] According to a sixth aspect of the embodiments of this specification, an augmented reality (VR) device is provided, comprising:
[0034] Memory and processor;
[0035] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the above-described panoramic image processing method.
[0036] According to a seventh aspect of the embodiments of this specification, a computing device is provided, comprising:
[0037] Memory and processor;
[0038] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, they implement the steps of the above-described panoramic image processing method.
[0039] According to an eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions that, when executed by a processor, implement the steps of the panoramic image processing method described above.
[0040] According to a ninth aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the panoramic image processing method described above.
[0041] One embodiment of this specification implements a panoramic image processing method and apparatus. The method includes inputting a target panoramic image into an image processing model; obtaining panoramic image features of the target panoramic image through the encoding layer of the image processing model; converting the panoramic image features into spherical image encoding features according to a panoramic image spherical conversion algorithm; inputting the spherical image encoding features into the decoding layer of the image processing model; obtaining a spherical depth image of the target panoramic image through the decoding layer; and converting the spherical depth image into a panoramic depth image of the target panoramic image according to the panoramic image spherical conversion algorithm. The image processing model is a neural network model.
[0042] Specifically, this method obtains the depth image of a panoramic image through a deep learning image processing model, without the need for additional instruments and equipment, thus saving costs. Furthermore, in practical applications, the deep learning process on planar images is extended to a sphere, and an image processing model specifically designed for spheres is run on the sphere, which solves the distortion caused by the ERP projection process of panoramic images, thereby improving the accuracy of obtaining the depth image of the panoramic image. Attached Figure Description
[0043] Figure 1 This is a schematic diagram of a 360-degree video recording, panoramic image, and cube bounding box image provided in one embodiment of this specification, as well as a schematic diagram of the mapping process from 360-degree video recording to panoramic image and cube bounding box image;
[0044] Figure 2 This is a schematic diagram of a depth image acquisition process for a panoramic image provided in one embodiment of this specification;
[0045] Figure 3 This is a schematic diagram illustrating a specific application scenario of a panoramic image processing method provided in one embodiment of this specification;
[0046] Figure 4 This is a flowchart illustrating a panoramic image processing method provided in one embodiment of this specification;
[0047] Figure 5 This is a flowchart illustrating the processing procedure of a panoramic image processing method provided in one embodiment of this specification.
[0048] Figure 6 This is a flowchart illustrating the processing procedure of another panoramic image processing method provided in one embodiment of this specification.
[0049] Figure 7 This is a schematic diagram of feature fusion in a panoramic image processing method provided in one embodiment of this specification;
[0050] Figure 8 This is a schematic diagram of the structure of a panoramic image processing device provided in one embodiment of this specification;
[0051] Figure 9 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0052] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0053] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0054] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0055] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0056] A 3D model is a polygonal representation of an object, typically displayed using a computer or other video equipment. The displayed object can be a real-world entity or a fictional object.
[0057] Panoramic images: Images that provide a 360-degree view of the surrounding environment. They can be captured using a panoramic camera or generated by fusing multiple perspective images of a single object from different angles.
[0058] Depth estimation: Given an image, a deep neural network model is used to estimate the distance from the actual position of each pixel in the scene to the optical center of the camera.
[0059] Image processing model: This image processing model is a neural network model based on the encoder-decoder framework. The encoder is the encoding layer of the neural network model, which is used for data compression (dimensionality reduction); the decoder is the decoding layer of the neural network model, which is used for parameter reconstruction (dimensionality increase).
[0060] HEALPixel algorithm: Hierarchical Equal Area isoLatitude Pixelation of asphere, is a method of pixelating a sphere that produces a subdivision of the sphere, where each pixel covers the same surface area as every other pixel.
[0061] Skip-connection: For depth estimation tasks, high-resolution images / feature maps are required. However, the feature extraction part of the network, through continuous layer-by-layer computation, eventually reduces the resolution of the feature maps to a very small level, which is detrimental to accurate depth estimation results. Skip-connection, on the other hand, can bring in feature maps from shallower layers and add them directly to the original feature maps. These shallower feature maps have higher resolution and contain richer local information, which is more conducive to the accuracy of the depth estimation results.
[0062] ERP: Equirectangular Projection.
[0063] Backbone: The term translates to "main network," meaning it's part of the overall network. This backbone network often refers to the feature extraction network, which extracts information from images for use by subsequent networks.
[0064] See Figure 1 , Figure 1The diagram illustrates 360-degree video recordings, panoramic images, and cube bounding box images, as well as the mapping process from 360-degree video recordings to panoramic and cube bounding box images.
[0065] Figure 1 Part a is a schematic diagram of 360-degree video recording, panoramic image, and cube bounding box image; part b is a schematic diagram of the mapping process from 360-degree video recording to panoramic image and cube bounding box image.
[0066] In practical applications, all panoramic cameras have an optically significant sphere inside the camera, such as... Figure 1 The 360-degree video in part b is used to generate the ERP panoramic image by dividing the network on the sphere and mapping it onto a 2:1 image; and by wrapping six images around the outside of the sphere and unfolding them, a cube bounding box image (i.e., cubemap) can be generated.
[0067] See Figure 2 , Figure 2 A schematic diagram of the panoramic depth image acquisition process for a panoramic image is shown.
[0068] Figure 2 It includes a panoramic image of the target building, as well as a cube map of the target building.
[0069] In practice, this method uses two backbones in the image processing model to extract features from the undistorted cube map and the distorted panoramic image, respectively. After feature fusion through the fusion module, the feature is input into the decoder of the image processing model for decoding to obtain the final panoramic depth image corresponding to the panoramic image. Subsequently, the 3D model of the target building can be constructed based on the panoramic depth image.
[0070] In the embodiments of this specification, this depth image acquisition process using the panoramic image can also acquire depth data of the target building image without the aid of professional instruments and equipment, so that a 3D model of the target building can be constructed based on the acquired depth data. However, this method requires two backbones to extract features from the cubemap and the panoramic image, resulting in a large number of parameters. Furthermore, although the left and right ends of the panoramic image are continuous in space, they are not continuous in the image, leading to inconsistent depth estimates at the left and right ends of the panoramic image and a large error. At the same time, this method fails to completely eliminate distortion and does not perform special processing on the distorted features extracted from the panoramic image, thus greatly affecting the depth estimation effect.
[0071] Based on this, two panoramic image processing methods are provided in this specification. One or more embodiments of this specification also relate to a panoramic image processing apparatus, a 3D model building terminal, an augmented reality (AR) device, an augmented reality (VR) device, a computing device, a computer-readable storage medium, and a computer program, which will be described in detail in the following embodiments.
[0072] See Figure 3 , Figure 3 A schematic diagram illustrating a specific application scenario of a panoramic image processing method provided according to an embodiment of this specification is shown.
[0073] Figure 3 The system includes a terminal 302 and a server 304. The terminal 302 includes, but is not limited to, mobile phones, tablets, and desktop computers. The server 304 can be understood as a physical server or a cloud server. For ease of understanding, this specification describes the server 304 as a cloud server in detail.
[0074] The panoramic image processing method provided in the embodiments of this specification is applied to a 3D model construction scene of a room as an example for detailed explanation.
[0075] In practice, the terminal 302 uploads the panoramic image of the room to be 3D modeled to the server 304. The panoramic image of the room can be taken by the panoramic camera embedded in the terminal 302, or by an external panoramic camera and then uploaded to the terminal 302. Alternatively, it can be uploaded by the user to the terminal 302 through other storage devices (such as hard drives).
[0076] The server 304 processes the panoramic image of the room using a pre-trained image processing model to obtain the panoramic depth image corresponding to the panoramic image of the room. Finally, the server 304 performs a 3D modeling of the room based on the panoramic depth image corresponding to the panoramic image of the room to obtain the 3D model of the room. The constructed 3D model of the room is then returned to the terminal 302 and displayed to the user through the display interface of the terminal 302.
[0077] Based on the panoramic depth image corresponding to the panoramic image of the room, there are multiple ways to perform 3D modeling of the room. For example, point cloud processing can be used: the panoramic depth image is converted into point cloud data, the relative relationship between point clouds is calculated using the normal distribution transformation method, and the point clouds are stitched together using the relative relationships between all panoramic images; mesh construction can be performed: for the stitched point cloud, the Poisson reconstruction method is used to obtain a triangular mesh; texture calculation can be performed: the correspondence between points on the triangular mesh and pixels on the image is calculated. Incorrect correspondence will increase the energy of the Markov random field between the image and the triangular mesh. Therefore, the correspondence between the triangular mesh points and image pixels with the lowest Markov random field energy can be found to obtain the texture of the triangular mesh; finally, the defects and redundancies in the mesh are repaired manually to obtain the final 3D model.
[0078] Alternatively, the room can be modeled in 3D using a different approach: point cloud processing. This involves converting the panoramic depth image into point cloud data; calculating stitching relationships by finding matching pixel relationships between two panoramic images, identifying two sets of matching 3D points, calculating the relative relationships between these sets using the Umeyama algorithm, and stitching the point clouds together based on these relative relationships across all panoramic images; mesh construction using the Dilaun triangulation method to obtain triangular meshes; mesh processing using the Lindström method for mesh degeneracy, the Surazsky algorithm for smoothing the triangular meshes, and the Liepa method for detecting and filling holes in the meshes; texture calculation by determining the correspondence between points on the triangular meshes and pixels in the image. Incorrect correspondences increase the Markov random field energy between the image and the triangular mesh. Therefore, finding the triangular mesh point-image pixel correspondence with the lowest Markov random field energy yields the triangular mesh texture; finally, manual repair is used to patch defects and redundancies in the mesh, resulting in the final 3D model.
[0079] Furthermore, the specific processing of the panoramic image can be implemented on either the terminal 302 or the server 304, and can be set according to the actual application. This specification embodiment does not impose any limitations on this. If the terminal 302 has sufficient computing resources, the processing of the panoramic depth image of the room's panoramic image and the construction of the room's 3D model can also be implemented on the terminal 302; this specification embodiment only illustrates the example of the processing of the panoramic depth image of the room's panoramic image and the construction of the room's 3D model being implemented on the server 304.
[0080] The panoramic image processing method provided in the embodiments of this specification obtains a precise panoramic depth image of the room to be 3D modeled through a deep learning image processing model, without the need for additional instruments and equipment, thus saving costs; and can quickly realize the 3D modeling of the room based on the panoramic depth image, improving the user experience.
[0081] See Figure 4 , Figure 4 A flowchart of a panoramic image processing method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0082] Step 402: Input the target panoramic image into the image processing model, and obtain the panoramic image features of the target panoramic image through the encoding layer of the image processing model.
[0083] The target panoramic image can be understood as a panoramic image captured by a standalone panoramic camera, or a panoramic camera embedded in a terminal such as a mobile phone or mobile computer; alternatively, it can be a target panoramic image generated by fusing multiple planar images. When the target panoramic image is generated by fusing multiple planar images, these multiple planar images all target the same object. For example, if the target object is a room, then all the multiple planar images are planar images of that room. The specific implementation method is as follows:
[0084] Before inputting the target panoramic image into the image processing model, the process further includes:
[0085] Acquire a panoramic image of the target object captured by a panoramic imaging device; or
[0086] At least two initial planar images of the target object are acquired, and the at least two initial planar images are fused according to a preset image fusion algorithm to obtain a target panoramic image of the target object.
[0087] Among them, panoramic shooting equipment includes stand-alone panoramic cameras or terminal devices with embedded panoramic cameras; and the target object can be any type and any size object, such as buildings, goods for sale, etc.
[0088] For ease of understanding, the embodiments in this specification all use residential buildings as examples for detailed description.
[0089] Taking a house for rent as an example, acquiring a panoramic image of the target object taken by a panoramic shooting device can be understood as acquiring a panoramic image of the house for rent taken by a mobile phone's panoramic camera; or
[0090] At least two initial planar images of the target object are acquired, and the two initial planar images are fused according to a preset image fusion algorithm to obtain a target panoramic image of the target object. This can be understood as acquiring multiple initial planar images of the house to be rented. These multiple initial planar images can be taken in real time by a mobile phone camera or obtained from historical photos taken in a database.
[0091] After acquiring multiple initial planar images of the house to be rented, these images can be fused using a preset image fusion algorithm to generate and obtain a target panoramic image of the house. The preset image fusion algorithm can be any image fusion algorithm that can fuse multiple planar images into a single panoramic image. For example, images taken from different angles using a mobile phone or camera (with some overlap between them) can be used to extract features using the SIFT (Scale Invariant Feature Transform) operator. Through feature matching, image rotation, and image fusion operations, the images can be stitched together to form a large-scene image. This specification does not limit the specific implementation of this method.
[0092] In the embodiments of this specification, a target panoramic image of the target object can be quickly obtained by using a panoramic shooting device or by fusing multiple initial planar images. Subsequently, a panoramic depth image corresponding to the target panoramic image of the target object can be obtained by using an image processing model.
[0093] Furthermore, the image processing model in the embodiments of this specification is a neural network model, which can be understood as running a neural network model specifically designed for image processing on a sphere.
[0094] Taking the target object as the house to be rented and the target panoramic image as the target panoramic image of the house to be rented, the target panoramic image is input into the image processing model. Through the encoding layer of the image processing model, the panoramic image features of the target panoramic image are obtained. It can be understood that the target panoramic image of the house to be rented is input into the image processing model, and the panoramic image features of the target panoramic image of the house to be rented are obtained through the encoding layer of the image processing model. In this embodiment, the features can be understood as feature maps.
[0095] Step 404: Convert the panoramic image features into spherical image coding features according to the panoramic image spherical conversion algorithm.
[0096] The panoramic image to spherical image conversion algorithm can be understood as the HEALPixel algorithm. Of course, the embodiments in this specification are not limited to using only this algorithm. Any algorithm that can convert panoramic images to spherical images can be used.
[0097] In practical applications, after obtaining the panoramic image features of the target panoramic image, the panoramic image features are converted into spherical image coding features through a panoramic image spherical conversion algorithm.
[0098] Step 406: Input the spherical image encoding features into the decoding layer of the image processing model, and obtain the spherical depth image of the target panoramic image through the decoding layer.
[0099] Specifically, after obtaining the spherical image encoding features, these features are input into the decoding layer of the image processing model. The decoding layer then decodes the features to obtain the spherical depth image of the panoramic image of the target.
[0100] Step 408: According to the panoramic image spherical conversion algorithm, convert the spherical depth image into a panoramic depth image of the target panoramic image.
[0101] After obtaining the spherical depth image of the target panoramic image, an inverse transformation can be performed according to the panoramic image spherical conversion algorithm to convert the spherical depth image into the panoramic depth image of the target panoramic image, thereby realizing the acquisition of the panoramic depth image of the target panoramic image.
[0102] See Figure 5 , Figure 5 A flowchart illustrating the processing procedure of a panoramic image processing method provided in one embodiment of this specification is shown.
[0103] Step 502: Input the target panoramic image into the encoding layer of the image processing model.
[0104] Step 504: Obtain the panoramic image feature map of the target panoramic image extracted by the coding layer.
[0105] Step 506: Convert the panoramic image feature map of the target panoramic image using the HEALPixel algorithm.
[0106] Step 508: Convert the panoramic image feature map into a distortion-free spherical image encoding feature map using the HEALPixel algorithm.
[0107] Step 510: Input the encoded feature map of the spherical image into the decoding layer of the image processing model.
[0108] Step 512: Obtain the spherical depth image of the panoramic image of the target obtained through the decoding layer.
[0109] Step 514: Perform an inverse transformation on the spherical depth image of the panoramic image of the target using the HEALPixel algorithm.
[0110] Step 516: Convert the spherical depth image into a panoramic depth image of the target panoramic image by using the inverse transform of the HEALPixel algorithm.
[0111] The method provided in the embodiments of this specification obtains the depth image of a panoramic image through a deep learning image processing model, without the need for additional instruments and equipment, thus saving costs. Furthermore, in practical use, the deep learning process on planar images is extended to a sphere, and an image processing model specifically designed for spheres is run on the sphere, which solves the distortion caused by the ERP projection process of panoramic images, thereby improving the accuracy of obtaining the depth image of the panoramic image.
[0112] In practical applications, the target panoramic image passes through multiple encoder layers of the image processing model. Through continuous layer-by-layer calculations to extract features, the resolution of the final extracted panoramic image feature map becomes very low, which is detrimental to accurate depth estimation results. Therefore, in this embodiment, to avoid this problem, after obtaining the spherical image decoding features of the target panoramic image through the decoding layer, the spherical image encoding features obtained through the encoding layer are superimposed with the spherical image decoding features obtained through the decoding layer to obtain the spherical depth image of the target panoramic image. This increases the accuracy of the subsequent depth estimation result obtained from the panoramic depth image of the target panoramic image based on the spherical depth image. The specific implementation method is as follows:
[0113] The step of inputting the spherical image encoding features into the decoding layer of the image processing model, and obtaining the spherical depth image of the target panoramic image through the decoding layer, includes:
[0114] The spherical image encoding features are input into the decoding layer of the image processing model, and the spherical image decoding features of the target panoramic image are obtained through the decoding layer.
[0115] Based on the spherical image encoding features and the spherical image decoding features, a spherical depth image of the target panoramic image is obtained.
[0116] Specifically, after obtaining the spherical image encoding features of the target panoramic image, the spherical image encoding features are input into the decoding layer of the image processing model, and the spherical image decoding features of the target panoramic image are obtained through the decoding layer; then the spherical image encoding features and the spherical image decoding features are superimposed to obtain the spherical depth image of the target panoramic image.
[0117] In practical applications, to ensure the effective superposition of spherical image encoded features and spherical image decoded features, a skip-connection approach can be used to superimpose features between two features. This involves using a skip-connection to bring over the spherical image encoded features from the shallower encoder layer and superimpose them with the spherical image decoded features from the decoder layer. Since the feature images from the shallower layers have higher resolution and contain richer local information, this improves the accuracy of depth estimation results. The specific implementation is described below:
[0118] The step of obtaining the spherical depth image of the target panoramic image based on the spherical image encoding features and the spherical image decoding features includes:
[0119] The spherical image encoding features and the spherical image decoding features are superimposed using a skip connection method;
[0120] Based on the superimposed spherical image encoding features and the spherical image decoding features, a spherical depth image of the target panoramic image is obtained.
[0121] Specifically, spherical image encoding features and spherical image decoding features are superimposed using a skip connection method. By superimposing the spherical image encoding features of the shallower encoder layer with the spherical image decoding features of the deeper decoder layer, the higher resolution of the feature image of the shallower encoder layer and the rich local information are used to improve the accuracy of depth estimation of the spherical depth image of the target panoramic image.
[0122] In practical implementation, taking an example where both the encoding and decoding layers are i-layers, the feature superposition of spherical image encoding features and spherical image decoding features will be described in detail:
[0123] The step of superimposing the encoded features and decoded features of the spherical image through skip connections includes:
[0124] S2. The spherical image encoded features obtained through the i-th encoding layer and the spherical image decoded features obtained through the j-th decoding layer are superimposed using a skip connection method.
[0125] Wherein, the initial layer of the i-th encoding layer is the first layer, and the initial layer of the j-th decoding layer is the last layer;
[0126] S4. Determine whether the i-th encoding layer is the last encoding layer and whether the j-th decoding layer is the first decoding layer.
[0127] If not, increment i by 1, decrement j by 1, and continue with step S2.
[0128] Where i and j are both positive integers.
[0129] Taking an example where i and j both belong to the range [1, n], and n is 4, the following steps are taken: The spherical image encoded features obtained through the first encoding layer are superimposed with the spherical image decoded features obtained through the fourth decoding layer using a skip connection. Then, the spherical image encoded features obtained through the second encoding layer are superimposed with the spherical image decoded features obtained through the third decoding layer using a skip connection, and so on, to complete the entire feature superposition process. Of course, in practical applications, the number of layers for i and j can be the same or different, depending on the specific requirements.
[0130] In the embodiments of this specification, by superimposing the spherical image coding features of the shallower encoder layer with the spherical image decoding features of the deeper decoder layer, the accuracy of depth estimation of the spherical depth image of the target panoramic image is improved based on the higher resolution of the feature image of the spherical image coding features of the shallower encoder layer and the rich local information.
[0131] Furthermore, since the embodiments in this specification pertain to the processing of panoramic images, to ensure the integrity and accuracy of the obtained panoramic depth image, it is necessary to consider the global receptive field, i.e., the context acquisition capability. Therefore, the embodiments in this specification set up a cross-attention mechanism fusion module in the decoder layer. Through the self-attention mechanism of the cross-attention mechanism fusion module (CAF), a better superposition between the spherical image encoded features and the spherical image decoded features is achieved. The specific implementation method is as follows:
[0132] The step of obtaining the spherical depth image of the target panoramic image based on the spherical image encoding features and the spherical image decoding features includes:
[0133] The spherical image encoding features and the spherical image decoding features are fused by the cross-attention mechanism fusion module of the decoding layer to perform attention calculation, thereby obtaining the correction amount of the spherical image encoding features and the correction amount of the spherical image decoding features;
[0134] Based on the correction amount of the spherical image encoding features, the correction amount of the spherical image decoding features, the spherical image encoding features, and the spherical image decoding features, a spherical depth image of the target panoramic image is obtained.
[0135] Specifically, for global features (spherical image decoding features), the correction amount Att0 is calculated using global Q information; for local features, the correction amount Att1 is calculated using local Q information, as described in Formula 1 below:
[0136]
[0137] In the embodiments of this specification, the spherical image encoded features and spherical image decoded features are self-attention calculated through the cross-attention mechanism fusion module in the decoding layer. The explanation of self-attention calculation is as follows:
[0138] The image is divided into many image patches, and the correlation between each image patch and other image patches is calculated using the formula mentioned above. Q is the query (image patch), and k is the key (other image patches). For a given image patch (query), all other image patches (keys) are used to calculate the correlation between the two image patches. The higher the correlation between two image patches, the larger the result of Q*K. V can be regarded as a "feature" in the neural network. It is equivalent to calculating the importance of each feature V using Q and K.
[0139] After obtaining the correction values for the spherical image coding features and the spherical image decoding features, the spherical depth image of the target panoramic image can be obtained based on these values. The specific implementation of Formula 2 is as follows:
[0140] X CA =FFN(LN(X0+Att0+X1+Att1))Formula 2
[0141] Furthermore, as described in the above embodiments, the specific superposition method of spherical image encoded features and spherical image decoded features can employ skip-connection. Therefore, the image processing model of the encoder-decoder layer with skip-connection in this embodiment can obtain a relatively accurate panoramic depth image. However, simply superimposing spherical image encoded features and spherical image decoded features through skip-connection results in a relatively simple fusion process, which cannot effectively balance local (features from skip-connection are biased towards local features) and global (decoder features are biased towards global features) features. Therefore, in this embodiment, to achieve better feature superposition, a CAF module (Cross-Attention Fusion Module) is used to fuse features from different dimensions of the decoder and skip-connection. A correction value is learned for each of the two different dimensions of features, thereby compensating for the defects caused by directly adding features from different dimensions. The specific implementation method is as follows:
[0142] The step of obtaining the spherical depth image of the target panoramic image based on the correction amount of the spherical image encoding features, the correction amount of the spherical image decoding features, the spherical image encoding features, and the spherical image decoding features includes:
[0143] Based on the correction amount of the spherical image encoding features and the correction amount of the spherical image decoding features, the spherical image encoding features and the spherical image decoding features are superimposed by a skip connection method;
[0144] Based on the superimposed spherical image encoding features and the spherical image decoding features, a spherical depth image of the target panoramic image is obtained.
[0145] Specifically, after obtaining the correction values of the spherical image encoding features and the spherical image decoding features, the spherical image encoding features and the spherical image decoding features are superimposed in the CAF module through a skip connection method based on the correction values of the spherical image encoding features and the spherical image decoding features. Finally, based on the feature image of the superimposed spherical image encoding features and spherical image decoding features, the spherical depth image of the target panoramic image is obtained. That is, the superposition of features in the CAF module is performed through a self-attention mechanism.
[0146] In practical applications, after obtaining the panoramic depth image of the target panoramic image, accurate 3D modeling can be performed based on this panoramic depth image. The specific implementation method is as follows:
[0147] After converting the spherical depth image into a panoramic depth image of the target panoramic image according to the panoramic image spherical conversion algorithm, the method further includes:
[0148] A three-dimensional model of the target object is constructed based on the panoramic depth image of the target panoramic image.
[0149] In practice, the method of constructing a 3D model based on panoramic depth images can be found in the above embodiments, and will not be repeated here.
[0150] Furthermore, with the image processing model pre-trained, after acquiring the target panoramic image, it can be directly input into the image processing model. The model can then perform the aforementioned processing steps internally and directly output the panoramic depth image of the target panoramic image, achieving a seamless user experience and enhancing the overall user experience. The specific implementation method is as follows:
[0151] After inputting the target panoramic image into the image processing model, the process further includes:
[0152] Obtain the panoramic depth image of the target panoramic image output by the image processing model;
[0153] Accordingly, the training steps of the image processing model are as follows:
[0154] Determine the sample panoramic image and the sample panoramic depth image corresponding to the sample panoramic image;
[0155] The sample panoramic image is input into the image processing model, and the sample panoramic image features are obtained through the encoding layer of the image processing model.
[0156] According to the panoramic image spherical conversion algorithm, the features of the sample panoramic image are converted into the coded features of the sample spherical image;
[0157] The encoded features of the sample spherical image are input into the decoding layer of the image processing model, and the sample spherical depth image of the sample panoramic image is obtained through the decoding layer.
[0158] According to the panoramic image spherical conversion algorithm, the sample spherical depth image is converted into a predicted panoramic depth image of the sample panoramic image;
[0159] The loss function of the image processing model is adjusted based on the sample panoramic depth image and the predicted panoramic depth image to train the image processing model.
[0160] Specifically, the training steps of the image processing model are the same as the specific implementation process of the image processing model in the above embodiments for processing the target panoramic image to obtain the panoramic depth image of the target panoramic image. Details not described in detail in the training steps of the image processing model can be found in the above embodiments. The specific implementation steps of the image processing model for processing the target panoramic image to obtain the panoramic depth image of the target panoramic image are as follows.
[0161] In the embodiments of this specification, the deep learning process (feature overlay, etc.) performed on a planar image is extended to a sphere, and a deep learning model specifically designed for a sphere (i.e., the image processing model of the embodiments of this specification) is run on the sphere. This solves the problem that the left and right ends of a panoramic image are continuous in space but discontinuous in the image. Furthermore, a panoramic depth image can be obtained by processing the panoramic image alone, reducing the overall number of parameters and greatly improving the processing efficiency of the image processing model.
[0162] The following is in conjunction with the appendix Figure 6 Taking the panoramic image processing method provided in this specification as an example of acquiring panoramic depth images from panoramic images of buildings, the panoramic image processing method will be further explained. Among other things, Figure 6 The diagram illustrates the processing steps of another panoramic image processing method provided in one embodiment of this specification, specifically including the following steps.
[0163] Step 602: Input the panoramic image of the house into the image processing model.
[0164] Step 604: Perform feature extraction and downsampling in the first, second, third, and fourth coding layers of the image processing model to obtain panoramic image feature maps output by the first, second, third, and fourth coding layers, respectively.
[0165] Step 606: Convert the panoramic image feature maps output from the first, second, third, and fourth coding layers into distortion-free spherical image coding feature maps using the HEALPixel algorithm in the panoramic image spherical conversion module.
[0166] Step 608: Input the spherical image encoding feature map output by the fourth encoding layer into the first decoding layer; after upsampling in the first decoding layer, obtain the spherical image decoding feature map; and through the cross self-attention mechanism fusion module of the first decoding layer, fuse the spherical image encoding feature map output by the third encoding layer after downsampling with the spherical image decoding feature map obtained by upsampling in the first decoding layer to obtain the first fused feature image.
[0167] See Figure 7 , Figure 7This diagram illustrates feature fusion via a self-attention mechanism in a panoramic image processing method provided in one embodiment of this specification.
[0168] Figure 7 The diagram illustrates the specific implementation of feature fusion through a cross-self-attention mechanism fusion module in each decoding layer; Figure 7 It can be seen that the spherical image encoded feature map output by the encoding layer of each layer, after being transformed by the HEALPixel algorithm in the panoramic image spherical transformation module, is processed by residual convolution, combined with spherical position encoding, and then input into the cross-attention mechanism module after layer normalization. At the same time, the spherical image encoded feature map output by the decoding layer corresponding to each encoding layer is also processed by residual convolution, combined with spherical position encoding, and then input into the cross-attention mechanism module after layer normalization. The spherical image encoded feature map and the spherical image decoded feature map are fused by the self-attention mechanism in the cross-self-attention mechanism fusion module. After obtaining the fused feature image, it is input into the feedforward network for processing, and then upsampled for further processing.
[0169] Step 610: Input the first fused feature image into the second decoding layer, and obtain the spherical image decoding feature map after upsampling in the second decoding layer; and perform feature fusion with the spherical image encoding feature map output by the second encoding layer after downsampling and the spherical image decoding feature map obtained by upsampling in the second decoding layer through the cross self-attention mechanism fusion module of the second decoding layer to obtain the second fused feature image.
[0170] Step 612: Input the second fused feature image into the third decoding layer, and obtain the spherical image decoding feature map after upsampling in the third decoding layer; and perform feature fusion with the spherical image encoding feature map output after downsampling in the first encoding layer and the spherical image decoding feature map obtained after upsampling in the third decoding layer through the cross self-attention mechanism fusion module of the third decoding layer to obtain the third fused feature image.
[0171] Step 614: Regress the third fused feature image through a depth regressor to obtain a spherical depth image, and then perform an inverse transformation of the spherical depth image through the HEALPixel algorithm to obtain the panoramic depth image corresponding to the panoramic image of the house.
[0172] The panoramic image processing method provided in this specification uses a single-backbone deep neural network, which results in fewer parameters and significantly higher accuracy in actual image processing. Furthermore, the panoramic image processing method uses the HEALPixel algorithm to convert the panoramic image into a distortion-free spherical image. Self-attention and convolution calculations are performed on the sphere. Since the sphere itself is a continuous surface, the left and right ends of the panoramic image are connected on the sphere, eliminating inconsistencies. Moreover, because the HEALPixel algorithm projects the extracted features from the panoramic image onto the distortion-free sphere, it specially processes the distortion in the features, greatly reducing the impact of distortion on the depth estimation task and improving the accuracy of the panoramic depth image. Simultaneously, the use of the CAF module significantly enhances the decoding layer's ability to perceive context, effectively utilizing scene information and further improving the accuracy of the acquired panoramic depth image.
[0173] Corresponding to the above method embodiments, this specification also provides embodiments of panoramic image processing apparatus. Figure 8 A schematic diagram of the structure of a panoramic image processing apparatus provided in one embodiment of this specification is shown. Figure 8 As shown, the device includes:
[0174] The panoramic feature determination module 802 is configured to input the target panoramic image into the image processing model and obtain the panoramic image features of the target panoramic image through the encoding layer of the image processing model.
[0175] The encoding feature determination module 804 is configured to convert the panoramic image features into spherical image encoding features according to the panoramic image spherical conversion algorithm;
[0176] The spherical depth image determination module 806 is configured to input the spherical image encoding features into the decoding layer of the image processing model, and obtain the spherical depth image of the target panoramic image through the decoding layer;
[0177] The panoramic depth image determination module 808 is configured to convert the spherical depth image into a panoramic depth image of the target panoramic image according to the panoramic image spherical conversion algorithm, wherein the image processing model is a neural network model.
[0178] Optionally, the spherical depth image determination module 806 is further configured to:
[0179] The spherical image encoding features are input into the decoding layer of the image processing model, and the spherical image decoding features of the target panoramic image are obtained through the decoding layer.
[0180] Based on the spherical image encoding features and the spherical image decoding features, a spherical depth image of the target panoramic image is obtained.
[0181] Optionally, the spherical depth image determination module 806 is further configured to:
[0182] The spherical image encoding features and the spherical image decoding features are superimposed using a skip connection method;
[0183] Based on the superimposed spherical image encoding features and the spherical image decoding features, a spherical depth image of the target panoramic image is obtained.
[0184] Optionally, the spherical depth image determination module 806 is further configured to:
[0185] The spherical image encoding features and the spherical image decoding features are fused by the cross-attention mechanism fusion module of the decoding layer to perform attention calculation, thereby obtaining the correction amount of the spherical image encoding features and the correction amount of the spherical image decoding features;
[0186] Based on the correction amount of the spherical image encoding features, the correction amount of the spherical image decoding features, the spherical image encoding features, and the spherical image decoding features, a spherical depth image of the target panoramic image is obtained.
[0187] Optionally, the spherical depth image determination module 806 is further configured to:
[0188] Based on the correction amount of the spherical image encoding features and the correction amount of the spherical image decoding features, the spherical image encoding features and the spherical image decoding features are superimposed by a skip connection method;
[0189] Based on the superimposed spherical image encoding features and the spherical image decoding features, a spherical depth image of the target panoramic image is obtained.
[0190] Optionally, the spherical depth image determination module 806 is further configured to:
[0191] S2. The spherical image encoded features obtained through the i-th encoding layer and the spherical image decoded features obtained through the j-th decoding layer are superimposed using a skip connection method.
[0192] Wherein, the initial layer of the i-th encoding layer is the first layer, and the initial layer of the j-th decoding layer is the last layer;
[0193] S4. Determine whether the i-th encoding layer is the last encoding layer and whether the j-th decoding layer is the first decoding layer.
[0194] If not, increment i by 1, decrement j by 1, and continue with step S2.
[0195] Optionally, the device further includes:
[0196] The panoramic image acquisition module is configured as follows:
[0197] Acquire a panoramic image of the target object captured by a panoramic imaging device; or
[0198] At least two initial planar images of the target object are acquired, and the at least two initial planar images are fused according to a preset image fusion algorithm to obtain a target panoramic image of the target object.
[0199] Optionally, the device further includes:
[0200] The 3D model building module is configured as follows:
[0201] A three-dimensional model of the target object is constructed based on the panoramic depth image of the target panoramic image.
[0202] Optionally, the device further includes:
[0203] The model processing module is configured as follows:
[0204] Obtain the panoramic depth image of the target panoramic image output by the image processing model;
[0205] Accordingly, the training steps of the image processing model are as follows:
[0206] Determine the sample panoramic image and the sample panoramic depth image corresponding to the sample panoramic image;
[0207] The sample panoramic image is input into the image processing model, and the sample panoramic image features are obtained through the encoding layer of the image processing model.
[0208] According to the panoramic image spherical conversion algorithm, the features of the sample panoramic image are converted into the coded features of the sample spherical image;
[0209] The encoded features of the sample spherical image are input into the decoding layer of the image processing model, and the sample spherical depth image of the sample panoramic image is obtained through the decoding layer.
[0210] According to the panoramic image spherical conversion algorithm, the sample spherical depth image is converted into a predicted panoramic depth image of the sample panoramic image;
[0211] The loss function of the image processing model is adjusted based on the sample panoramic depth image and the predicted panoramic depth image to train the image processing model.
[0212] The panoramic image processing device provided in this specification uses a deep learning image processing model to obtain the depth image of a panoramic image without the need for additional instruments and equipment, thus saving costs. Furthermore, in practical use, the deep learning process performed on planar images is extended to a sphere, and an image processing model specifically designed for spheres is run on the sphere, which solves the distortion caused by the ERP projection process of panoramic images, thereby improving the accuracy of obtaining the depth image of the panoramic image.
[0213] The above is a schematic scheme of a panoramic image processing device according to this embodiment. It should be noted that the technical solution of this panoramic image processing device and the technical solution of the panoramic image processing method described above belong to the same concept. For details not described in detail in the technical solution of the panoramic image processing device, please refer to the description of the technical solution of the panoramic image processing method described above.
[0214] In addition, this specification also provides another user-oriented panoramic image processing method, including:
[0215] In response to a user's panoramic image processing request, an image input interface is displayed to the user;
[0216] Receive the target panoramic image input by the user through the image input interface;
[0217] The target panoramic image is input into the image processing model, and the panoramic image features of the target panoramic image are obtained through the encoding layer of the image processing model.
[0218] According to the panoramic image spherical conversion algorithm, the panoramic image features are converted into spherical image coding features;
[0219] The spherical image encoding features are input into the decoding layer of the image processing model, and the spherical depth image of the target panoramic image is obtained through the decoding layer;
[0220] According to the panoramic image spherical conversion algorithm, the spherical depth image is converted into a panoramic depth image of the target panoramic image, and the panoramic depth image of the target panoramic image is displayed to the user through the image input interface. The image processing model is a neural network model.
[0221] The panoramic image processing method provided in this specification can provide users with a calling interface or a visual image input interface when facing users. When a panoramic image processing request is received from a user, the method can obtain the target panoramic image to be processed uploaded by the user through the calling interface, or receive the target panoramic image input by the user through the image input interface, process it, and display the panoramic depth image of the target panoramic image to the user through the calling interface or the image input interface, thereby improving the user experience.
[0222] Figure 9 A structural block diagram of a computing device 900 according to one embodiment of this specification is shown. The components of the computing device 900 include, but are not limited to, a memory 910 and a processor 920. The processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data.
[0223] The computing device 900 also includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interface (e.g., a Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) interface, a Wi-MAX interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0224] In one embodiment of this specification, the above-described components of the computing device 900 and Figure 9 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 9 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0225] The computing device 900 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 900 can also be a mobile or stationary server.
[0226] The processor 920 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the panoramic image processing method described above.
[0227] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-described panoramic image processing method belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-described panoramic image processing method.
[0228] One embodiment of this specification also provides a three-dimensional model building terminal, including:
[0229] Memory and processor;
[0230] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, perform the following steps:
[0231] The target panoramic image of the target object is input into the image processing model, and the panoramic image features of the target panoramic image are obtained through the encoding layer of the image processing model.
[0232] According to the panoramic image spherical conversion algorithm, the panoramic image features are converted into spherical image coding features;
[0233] The spherical image encoding features are input into the decoding layer of the image processing model, and the spherical depth image of the target panoramic image is obtained through the decoding layer;
[0234] According to the panoramic image spherical conversion algorithm, the spherical depth image is converted into a panoramic depth image of the target panoramic image;
[0235] Based on the panoramic depth image of the target panoramic image, a three-dimensional model of the target object is constructed, wherein the image processing model is a neural network model.
[0236] The above is an illustrative scheme of a 3D model building terminal according to this embodiment. It should be noted that the panoramic image processing method in the technical solution of this 3D model building terminal belongs to the same concept as the panoramic image processing method described above. For details not described in detail in the technical solution of the 3D model building terminal, please refer to the description of the technical solution of the panoramic image processing method described above.
[0237] This specification also provides an embodiment of an augmented reality (AR) device, comprising:
[0238] Memory and processor;
[0239] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the above-described panoramic image processing method.
[0240] The above is an illustrative scheme of an augmented reality (AR) device according to this embodiment. It should be noted that the technical solution of this AR device and the technical solution of the above-described panoramic image processing method belong to the same concept. Details not described in detail in the technical solution of the AR device can be found in the description of the technical solution of the above-described panoramic image processing method.
[0241] One embodiment of this specification also provides a virtual reality (VR) device, including:
[0242] Memory and processor;
[0243] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the above-described panoramic image processing method.
[0244] The above is an illustrative scheme of a virtual reality (VR) device according to this embodiment. It should be noted that the technical solution of this VR device and the technical solution of the above-described panoramic image processing method belong to the same concept. Details not described in detail in the technical solution of the VR device can be found in the description of the technical solution of the above-described panoramic image processing method.
[0245] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the panoramic image processing method described above.
[0246] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the panoramic image processing method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the panoramic image processing method described above.
[0247] An embodiment of this specification also provides a computer program, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the above-described panoramic image processing method.
[0248] The above is an illustrative example of a computer program according to this embodiment. It should be noted that the technical solution of this computer program belongs to the same concept as the technical solution of the panoramic image processing method described above. Details not described in detail in the computer program's technical solution can be found in the description of the technical solution of the panoramic image processing method described above.
[0249] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0250] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0251] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0252] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0253] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A panoramic image processing method, comprising: The target panoramic image is input into the image processing model, and the panoramic image features of the target panoramic image are obtained through the encoding layer of the image processing model. According to the panoramic image spherical conversion algorithm, the panoramic image features are converted into spherical image coding features; The spherical image encoding features are input into the decoding layer of the image processing model, and the spherical depth image of the target panoramic image is obtained through the decoding layer. The decoding layer includes a cross-attention mechanism fusion module, which is used to fuse the spherical image encoding features and the spherical image decoding features obtained by decoding through the decoding layer. According to the panoramic image spherical conversion algorithm, the spherical depth image is converted into a panoramic depth image of the target panoramic image, wherein the image processing model is a neural network model.
2. The panoramic image processing method according to claim 1, wherein inputting the spherical image encoding features into the decoding layer of the image processing model, and obtaining the spherical depth image of the target panoramic image through the decoding layer, comprises: The spherical image encoding features are input into the decoding layer of the image processing model, and the spherical image decoding features of the target panoramic image are obtained through the decoding layer. Based on the spherical image encoding features and the spherical image decoding features, a spherical depth image of the target panoramic image is obtained.
3. The panoramic image processing method according to claim 2, wherein obtaining the spherical depth image of the target panoramic image based on the spherical image encoding features and the spherical image decoding features comprises: The spherical image encoding features and the spherical image decoding features are superimposed using a skip connection method; Based on the superimposed spherical image encoding features and the spherical image decoding features, a spherical depth image of the target panoramic image is obtained.
4. The panoramic image processing method according to claim 2, wherein obtaining the spherical depth image of the target panoramic image based on the spherical image encoding features and the spherical image decoding features comprises: The spherical image encoding features and the spherical image decoding features are fused by the cross-attention mechanism fusion module of the decoding layer to perform attention calculation, thereby obtaining the correction amount of the spherical image encoding features and the correction amount of the spherical image decoding features; Based on the correction amount of the spherical image encoding features, the correction amount of the spherical image decoding features, the spherical image encoding features, and the spherical image decoding features, a spherical depth image of the target panoramic image is obtained.
5. The panoramic image processing method according to claim 4, wherein obtaining the spherical depth image of the target panoramic image based on the correction amount of the spherical image encoding features, the correction amount of the spherical image decoding features, the spherical image encoding features, and the spherical image decoding features comprises: Based on the correction amount of the spherical image encoding features and the correction amount of the spherical image decoding features, the spherical image encoding features and the spherical image decoding features are superimposed by a skip connection method; Based on the superimposed spherical image encoding features and the spherical image decoding features, a spherical depth image of the target panoramic image is obtained.
6. The panoramic image processing method according to claim 3 or 5, wherein the step of superimposing the spherical image encoded features and the spherical image decoded features by means of skip connections includes: S2. The spherical image encoded features obtained through the i-th encoding layer and the spherical image decoded features obtained through the j-th decoding layer are superimposed using a skip connection method. Wherein, the initial layer of the i-th encoding layer is the first layer, and the initial layer of the j-th decoding layer is the last layer; S4. Determine whether the i-th encoding layer is the last encoding layer and whether the j-th decoding layer is the first decoding layer. If not, increment i by 1, decrement j by 1, and continue with step S2.
7. The panoramic image processing method according to claim 1, further comprising, before inputting the target panoramic image into the image processing model: Acquire a panoramic image of the target object captured by a panoramic imaging device; or At least two initial planar images of the target object are acquired, and the at least two initial planar images are fused according to a preset image fusion algorithm to obtain a target panoramic image of the target object.
8. The panoramic image processing method according to claim 7, further comprising, after converting the spherical depth image into a panoramic depth image of the target panoramic image according to the panoramic image spherical conversion algorithm: A three-dimensional model of the target object is constructed based on the panoramic depth image of the target panoramic image.
9. The panoramic image processing method according to claim 1, further comprising, after inputting the target panoramic image into the image processing model: Obtain the panoramic depth image of the target panoramic image output by the image processing model; Accordingly, the training steps of the image processing model are as follows: Determine the sample panoramic image and the sample panoramic depth image corresponding to the sample panoramic image; The sample panoramic image is input into the image processing model, and the sample panoramic image features are obtained through the encoding layer of the image processing model. According to the panoramic image spherical conversion algorithm, the features of the sample panoramic image are converted into the coded features of the sample spherical image; The encoded features of the sample spherical image are input into the decoding layer of the image processing model, and the sample spherical depth image of the sample panoramic image is obtained through the decoding layer. According to the panoramic image spherical conversion algorithm, the sample spherical depth image is converted into a predicted panoramic depth image of the sample panoramic image; The loss function of the image processing model is adjusted based on the sample panoramic depth image and the predicted panoramic depth image to train the image processing model.
10. A panoramic image processing apparatus, comprising: The panoramic feature determination module is configured to input the target panoramic image into the image processing model and obtain the panoramic image features of the target panoramic image through the encoding layer of the image processing model. The encoding feature determination module is configured to convert the panoramic image features into spherical image encoding features according to the panoramic image spherical conversion algorithm; A spherical depth image determination module is configured to input the spherical image encoding features into the decoding layer of the image processing model, and obtain the spherical depth image of the target panoramic image through the decoding layer. The decoding layer includes a cross-attention mechanism fusion module, which is used to perform feature fusion on the spherical image encoding features and the spherical image decoding features obtained by decoding through the decoding layer. The panoramic depth image determination module is configured to convert the spherical depth image into a panoramic depth image of the target panoramic image according to the panoramic image spherical conversion algorithm, wherein the image processing model is a neural network model.
11. A panoramic image processing method, comprising: In response to a user's panoramic image processing request, an image input interface is displayed to the user; Receive the target panoramic image input by the user through the image input interface; The target panoramic image is input into the image processing model, and the panoramic image features of the target panoramic image are obtained through the encoding layer of the image processing model. According to the panoramic image spherical conversion algorithm, the panoramic image features are converted into spherical image coding features; The spherical image encoding features are input into the decoding layer of the image processing model, and the spherical depth image of the target panoramic image is obtained through the decoding layer. The decoding layer includes a cross-attention mechanism fusion module, which is used to fuse the spherical image encoding features and the spherical image decoding features obtained by decoding through the decoding layer. According to the panoramic image spherical conversion algorithm, the spherical depth image is converted into a panoramic depth image of the target panoramic image, and the panoramic depth image of the target panoramic image is displayed to the user through the image input interface. The image processing model is a neural network model.
12. A three-dimensional model building terminal, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, perform the following steps: The target panoramic image of the target object is input into the image processing model, and the panoramic image features of the target panoramic image are obtained through the encoding layer of the image processing model. According to the panoramic image spherical conversion algorithm, the panoramic image features are converted into spherical image coding features; The spherical image encoding features are input into the decoding layer of the image processing model, and the spherical depth image of the target panoramic image is obtained through the decoding layer. The decoding layer includes a cross-attention mechanism fusion module, which is used to fuse the spherical image encoding features and the spherical image decoding features obtained by decoding through the decoding layer. According to the panoramic image spherical conversion algorithm, the spherical depth image is converted into a panoramic depth image of the target panoramic image; Based on the panoramic depth image of the target panoramic image, a three-dimensional model of the target object is constructed, wherein the image processing model is a neural network model.
13. An augmented reality (AR) device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the panoramic image processing method according to any one of claims 1 to 9.
14. A virtual reality (VR) device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the panoramic image processing method according to any one of claims 1 to 9.