Methods and electronic device for capturing 360-degree image of a scene
Patent Information
- Application Number
- PCT/KR2026/002780
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-18
- Filing Date
- 2026-02-13
- Publication Date
- 2026-08-27
Smart Images

Figure KR2026002780_27082026_PF_FP_ABST
Abstract
Description
METHODS AND ELECTRONIC DEVICE FOR CAPTURING 360-DEGREE IMAGE OF A SCENE
[0001] Embodiments disclosed herein relate to an image processing method and system, and more particularly to methods and an electronic device for synthesizing 360-degree images.
[0002] With the rising popularity of extended reality (XR) devices, especially head-mounted displays (HMDs), a wide range of use cases and solutions are emerging in this space. However, a key challenge remains, the lack of seamless connectivity between augmented reality / virtual reality (AR / VR) headsets and existing devices like smartphones, which limits the ability to deliver a richer and more integrated user experience.
[0003] In an embodiment of the disclosure, 360-degree videos offer an immersive viewing experience by allowing users to look in all directions, simulating the sensation of standing at a center of a sphere and experiencing the scene firsthand. This feature enables consumers to capture their precious moments in a way that others relive the experience from every angle. While the XR devices are commonly used to enjoy such content, the availability of 360-degree videos remains limited due to the challenges involved in capturing them.
[0004] The 360-degree videos are widely used across various domains such as virtual tours on a real estate, travel, education and gaming, where more immersive experiences are created for live events like concerts and sports. By enabling viewers to explore scenes in all directions, the 360-degree videos significantly enhance storytelling and audience engagement through a more immersive and interactive viewing experience.
[0005] Currently, the user of the electronic device requires specialized 360-degree devices to create 360-degree content, and it may not be feasible for the user to carry / buy 360-degree devices.
[0006] Hence, there is a need in the art for solutions which will overcome the above-mentioned drawback(s), among others.
[0007] Accordingly, the embodiments herein provide a method for capturing a 360-degree image of a scene by an electronic device. In an embodiment of the disclosure, method includes capturing a first image of the scene using a first imaging device, where the first imaging device is positioned on a first surface of the electronic device. In an embodiment of the disclosure, the method includes capturing a second image of the scene using a second imaging device, where the second imaging device is positioned on the first surface of the electronic device. In an embodiment of the disclosure, the method includes capturing a third image of the scene using a third imaging device, where the third imaging device is positioned on a second surface of the electronic device, wherein the second surface faces in a direction opposite to the first surface. In an embodiment of the disclosure, the method includes generating a three-dimensional (3D) representation of the scene using the first image and the second image. In an embodiment of the disclosure, the method includes generating a fourth image by re-projecting the first image to align an optical axis of the first imaging device with the third imaging device based on the generated 3D representation and a pre-defined calibration matrix. In an embodiment of the disclosure, the method includes creating a two-dimensional (2D) rectilinear image representative of a 360-degree view of the scene using the third image and the fourth image.
[0008] Accordingly, the embodiments herein provide a method for capturing a 360-degree content of a scene by an electronic device. In an embodiment of the disclosure, the method includes receiving a user input to activate a 360 degree content capture mode on the electronic device. In an embodiment of the disclosure, the method includes performing a stereo depth estimation using a plurality of image streams from a first imaging device, and a second imaging device to generate a depth map. In an embodiment of the disclosure, the method includes reconstructing a three-dimensional (3D) volume of a scene using the depth map. In an embodiment of the disclosure, the method includes reprojecting texture and color information from an ultra-wide image stream received from the first imaging device into a viewpoint determined by extrinsic parameters. In an embodiment of the disclosure, the method includes converting both the reprojected image and the image from the third imaging device into an equirectangular projection format.
[0009] Accordingly, the embodiments herein provide an electronic device including a 360-degree scene capturing controller coupled with a memory storing a program or at least one instruction and at least one processor configure to individually or collectively execute the program or the at least one instruction. In an embodiment of the disclosure, wherein the program or the at least one instruction, when executed individually or collectively by the at least one processor, cause the electronic device to capture a first image of the scene using a first imaging device. In an embodiment of the disclosure, wherein the program or the at least one instruction, when executed individually or collectively by the at least one processor, cause the electronic device to capture a second image of the scene using the second imaging device. In an embodiment of the disclosure, wherein the program or the at least one instruction, when executed individually or collectively by the at least one processor, cause the electronic device to capture a third image of the scene using a third imaging device. In an embodiment of the disclosure, wherein the program or the at least one instruction, when executed individually or collectively by the at least one processor, cause the electronic device to generate a 3D representation of the scene using the first image and the second image. In an embodiment of the disclosure, wherein the program or the at least one instruction, when executed individually or collectively by the at least one processor, cause the electronic device to generate a fourth image by re-projecting the first image using the generated 3D representation and a pre-defined calibration matrix to align an optical axis of the first imaging device with the third imaging device. In an embodiment of the disclosure, wherein the program or the at least one instruction, when executed individually or collectively by the at least one processor, cause the electronic device to create a 2D rectilinear image representative of a 360-degree view of the scene using the third image and the fourth image.
[0010] Accordingly, the embodiments herein provide an electronic device including a 360-degree scene capturing controller coupled with memory storing a program or at least one instruction and at least one processor configure to individually or collectively execute the program or the at least one instruction. In an embodiment of the disclosure, wherein the program or the at least one instruction, when executed individually or collectively by the at least one processor, cause the electronic device to activate a 360- degree content capture mode on the electronic device. In an embodiment of the disclosure, wherein the program or the at least one instruction, when executed individually or collectively by the at least one processor, cause the electronic device to perform a stereo depth estimation using a plurality of image streams from a first imaging device and a second imaging device to generate a depth map. In an embodiment of the disclosure, wherein the program or the at least one instruction, when executed individually or collectively by the at least one processor, cause the electronic device to reconstruct a three-dimensional (3D) volume of a scene using the depth map. In an embodiment of the disclosure, wherein the program or the at least one instruction, when executed individually or collectively by the at least one processor, cause the electronic device to reproject texture and color information from an ultra-wide image stream received from the first imaging device into a viewpoint determined by extrinsic parameters. In an embodiment of the disclosure, wherein the program or the at least one instruction, when executed individually or collectively by the at least one processor, cause the electronic device to convert both the reprojected image and the image from the third imaging device into an equirectangular projection format.
[0011] According to an embodiment of the present disclosure, there may be provided a computer-readable recording medium having recorded thereon a program for causing a computer to execute at least one of the disclosed methods of controlling an electronic device.
[0012] These and other aspects of the embodiments herein will be better appreciated and understood when considered in conjunction with the following description and the accompanying drawings. It should be understood, however, that the following descriptions, while indicating at least one embodiment and numerous specific details thereof, are given by way of illustration and not of limitation. Many changes and modifications may be made within the scope of the embodiments herein without departing from the scope thereof, and the embodiments herein include all such modifications.
[0013] Embodiments herein are illustrated in the accompanying drawings, throughout which like reference letters indicate corresponding parts in the various figures. The embodiments herein will be better understood from the following description with reference to the following illustratory drawings. Embodiments herein are illustrated by way of examples in the accompanying drawings, and in which:
[0014] FIG. 1 depicts an example scenario illustrating limitation with an electronic device depicting panoramic image capture, according to existing arts;
[0015] FIG. 2 depicts example scenario illustrating limitation with the electronic device depicting a top view, a FOV and a missing FOV, according to existing arts;
[0016] FIG. 3 depicts example scenario illustrating limitation with the electronic device depicting a top view, a FOV and a missing FOV, according to existing arts;
[0017] FIG. 4 depicts various hardware components of an electronic device for capturing a 360-degree image of a scene, according to embodiments as disclosed herein;
[0018] FIG. 5 is a flowchart illustrating a method for capturing the 360-degree image of the scene by the electronic device, according to embodiments as disclosed herein;
[0019] FIG. 6 is another flowchart illustrating a method for capturing the 360-degree image of the scene by the electronic device, according to embodiments as disclosed herein;
[0020] FIG. 7 depicts a process of capturing a 360 degree image of the real-world scene using the electronic device, according to embodiments as disclosed herein;
[0021] FIG. 8 is an example method for capturing a 360-degree image of the real-world scene using the electronic device, according to embodiments as disclosed herein;
[0022] FIG. 9 depicts an example scenario in which the electronic device captures the 360-degree image of the real-world scene, according to embodiments as disclosed herein;
[0023] FIG. 10 depicts example scenario, wherein images are captured, and restored while capturing the 360-degree image of the real-world scene, according to embodiments as disclosed herein;
[0024] FIG. 11 depicts example scenario, wherein images are captured, and restored while capturing the 360-degree image of the real-world scene, according to embodiments as disclosed herein;
[0025] FIG. 12 depicts example scenario, wherein images are captured, and restored while capturing the 360-degree image of the real-world scene, according to embodiments as disclosed herein;
[0026] FIG. 13 depicts an example process of performing non overlapping camera calibration while capturing the 360 degree image of the real-world scene, according to embodiments as disclosed herein;
[0027] FIG. 14 depicts example scenario in which the proposed method performs a non-overlapping camera calibration and detecting calibration rig details, according to embodiments as disclosed herein;
[0028] FIG. 15 depicts example scenario in which the proposed method performs a non-overlapping camera calibration and detecting calibration rig details, according to embodiments as disclosed herein;
[0029] FIG. 16 depicts an example scenario of the proposed calibration process while capturing the 360 degree image of the real-world scene, according to embodiments as disclosed herein;
[0030] FIG. 17 depicts example scenario in which depth map based reprojection is performed while capturing the 360 degree image of the real-world scene, according to embodiments as disclosed herein;
[0031] FIG. 18 depicts example scenario in which depth map based reprojection is performed while capturing the 360 degree image of the real-world scene, according to embodiments as disclosed herein;
[0032] FIG. 19 depicts example scenario in which depth map based reprojection is performed while capturing the 360 degree image of the real-world scene, according to embodiments as disclosed herein;
[0033] FIG. 20 is a flow chart illustrating a method for generating Equirectangular images using conditional grids, according to embodiments as disclosed herein;
[0034] FIG. 21 depicts an equirectangular projection and diffusion operation while capturing the 360 degree image of the real-world scene, according to embodiments as disclosed herein;
[0035] FIG. 22 depicts an example scenario in which an equirectangular projection and diffusion is performed while capturing the 360 degree image of the real-world scene, according to embodiments as disclosed herein;
[0036] FIG. 23 depicts an example scenario of a SLAM diffusion, according to embodiments as disclosed herein;
[0037] FIG. 24 depicts an example scenario, wherein SLAM information is used for generating missing portions, according to embodiments as disclosed herein;
[0038] FIG. 25 depicts example scenario, in which a generative artificial intelligence (AI) is used for generating the missing FOV, according to embodiments as disclosed herein;
[0039] FIG. 26 depicts example scenario, in which a generative artificial intelligence (AI) is used for generating the missing FOV, according to embodiments as disclosed herein;
[0040] FIG. 27 depicts an example scenario, in which the proposed method generates the missing portions using a SLAM based generator, according to embodiments as disclosed herein;
[0041] FIG. 28 depicts an example scenario where the electronic device is used to capture 360-degree content, according to embodiments as disclosed herein; and
[0042] FIG. 29 depicts example scenario for capturing 360-degree content in one shot in different scenarios, according to embodiments as disclosed herein;
[0043] FIG. 30 depicts example scenario for capturing 360-degree content in one shot in different scenarios, according to embodiments as disclosed herein.
[0044] The principal object of embodiments herein is to disclose methods and systems (or electronic device) for synthesizing a 360-degree image, where the electronic device has at least one front facing camera (e.g., primary camera or the like), and at least one rear facing camera (e.g., secondary camera or the like).
[0045] Another object of embodiments herein is to capture a first image of a scene using a first imaging device, where the first imaging device is positioned on a rear side of the electronic device.
[0046] Another object of embodiments herein is to capture a second image of the scene using a second imaging device, where the second imaging device is positioned on the rear side of the electronic device.
[0047] Another object of embodiments herein is to capture a third image of the scene using a third imaging device, where the third imaging device is positioned on a front side of the electronic device.
[0048] Another object of embodiments herein is to generate a three-dimensional (3D) representation of the scene using the first image and the second image.
[0049] Another object of embodiments herein is to generate a fourth image by re-projecting the first image using the generated 3D representation and a pre-defined calibration matrix to align an optical axis of the first imaging device with the third imaging device.
[0050] Another object of embodiments herein is to create a two-dimensional (2D) rectilinear image representative of a 360-degree view of the scene using the third image and the fourth image.
[0051] Another object of embodiments herein is to generate a missing portion in the 2D rectilinear image using simultaneous localization and mapping (SLAM) information by receiving a plurality of images from the first imaging device and the second imaging device, and device's location and orientation using inertial measurement unit (IMU) and visual features.
[0052] Another object of embodiments herein is to generate an equi-rectangular projection of the scene by passing the 2D rectilinear image to a data driven model.
[0053] Another object of embodiments herein is to generate a 360-degree content of the scene based on the generated equi-rectangular projection of the scene.
[0054] The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted so as to not unnecessarily obscure the embodiments herein. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein may be practiced and to further enable those of skill in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein.
[0055] For the purposes of interpreting this specification, the definitions (as defined herein) will apply and whenever appropriate the terms used in singular will also include the plural and vice versa. It is to be understood that the terminology used herein is for the purposes of describing particular embodiments only and is not intended to be limiting. The terms "comprising", "having" and "including" are to be construed as open-ended terms unless otherwise noted.
[0056] The words / phrases "exemplary", "example", "illustration", "in an instance", "and the like", "and so on", "etc.", "etcetera", "e.g.,", "i.e.," are merely used herein to mean "serving as an example, instance, or illustration. Any embodiment or implementation of the present subject matter described herein using the words / phrases "exemplary", "example", "illustration", "in an instance", "and the like", "and so on", "etc.", "etcetera", "e.g.," , "i.e.," is not necessarily to be construed as preferred or advantageous over other embodiments.
[0057] Embodiments herein may be described and illustrated in terms of blocks which carry out a described function or functions. These blocks, which may be referred to herein as managers, units, modules, hardware components or the like, are physically implemented by analog and / or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits and the like, and may optionally be driven by a firmware. The circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like. The circuits constituting a block may be implemented by dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block of the embodiments may be physically separated into two or more interacting and discrete blocks without departing from the scope of the disclosure. Likewise, the blocks of the embodiments may be physically combined into more complex blocks without departing from the scope of the disclosure. Furthermore, any component referred to as a 'controller', 'manager', 'unit', or 'module' in the present disclosure should be understood as a hardware component, such as the processor 110 executing software instructions, and not as a nonce word devoid of structure.
[0058] It should be noted that elements in the drawings are illustrated for the purposes of this description and ease of understanding and may not have necessarily been drawn to scale. For example, the flowcharts / sequence diagrams illustrate the method in terms of the steps required for understanding of aspects of the embodiments as disclosed herein. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the present embodiments so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein. Furthermore, in terms of the system, one or more components / modules which comprise the system may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the present embodiments so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.
[0059] The accompanying drawings are used to help easily understand various technical features and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. As such, the present disclosure should be construed to extend to any modifications, equivalents, and substitutes in addition to those which are particularly set out in the accompanying drawings and the corresponding description. Usage of words such as first, second, third etc., to describe components / elements / steps is for the purposes of this description and should not be construed as sequential ordering / placement / occurrence unless specified otherwise.
[0060] It should be appreciated that the blocks in each flowchart and combinations of the flowcharts may be performed by one or more computer programs which include computer-executable instructions. The entirety of the one or more computer programs may be stored in a single memory or the one or more computer programs may be divided with different portions stored in different multiple memories.
[0061] Any of the functions or operations described herein can be processed by one processor or a combination of processors. The one processor or the combination of processors is circuitry performing processing and includes circuitry like an application processor (AP), a communication processor (CP), a graphical processing unit (GPU), a neural processing unit (NPU), a microprocessor unit (MPU), a system on chip (SoC), an IC, or the like.
[0062] The processor may include various processing circuitry and / or multiple processors. For example, as used herein, including the claims, the term “processor” may include various processing circuitry, including at least one processor, wherein one or more of at least one processor, individually and / or collectively in a distributed manner, may be configured to perform various functions described herein. As used herein, when “a processor”, “at least one processor”, and “one or more processors” are described as being configured to perform numerous functions, these terms cover situations, for example and without limitation, in which one processor performs some of recited functions and another processor(s) performs other of recited functions, and also situations in which a single processor may perform all recited functions. Additionally, the at least one processor may include a combination of processors performing various of the recited / disclosed functions, e.g., in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions.
[0063] Below are the explanations for the following terms used in the patent disclosure:
[0064] Spatial video: Spatial videos are stereoscopic 3D videos that offer a more immersive viewing experience.
[0065] 360-degree image: 360-degreeor full / spherical panoramic, denotes an image that captures a complete spherical view of the surroundings, from the sky directly above the user to the ground below the user.
[0066] Rectilinear: The rectilinear projection keeps all straight lines straight. As a result, the image stretches increasingly stronger towards a far edge.
[0067] Equirectangular projection: A common example of an equirectangular image or projection is a map of the earth. The image is not completely accurate because it takes a spherical object, and converts the spherical object into a flat object.
[0068] Sphere map: A sphere map, also known as a spherical environment map, is a 2D texture that shows the 360-degreeview of the scene around an object.
[0069] Primary camera: A back or rear camera in an electronic device with higher resolution, and higher field of view.
[0070] Secondary camera: A front camera in an electronic device with a lower resolution, and a lower Field of View.
[0071] Image Signal Processing (ISP): ISP is a dedicated processor that converts raw image data from a sensor into a more visually appealing image for the user.
[0072] Camera calibration: The process of determining one or more intrinsic parameters, such as, but not limited to, focal length, principal point and extrinsic (relation between camera / sensors).
[0073] FIG. 1 illustrates an example scenario of capturing a panoramic image using a conventional electronic device.
[0074] As shown in FIG. 1, the electronic device 10 captures a plurality of images while the electronic device 10 is moved along a horizontal direction relative to a scene. In this example, a first image corresponding to a first spatial region is captured as a first frame 21, a second image corresponding to a second spatial region adjacent to the first spatial region is captured as a second frame 22, and a third image corresponding to a third spatial region adjacent to the second spatial region is captured as a third frame 23. During the capture of the first frame 110, the second frame 21, and the third frame 23, the electronic device 10 is moved relative to the scene, and objects within the scene may change their positions over time. Accordingly, the captured frames 21, 22, and 23 may exhibit temporal inconsistencies.
[0075] For example, a moving object may appear at different positions in the first frame 21 and the second frame 22, resulting in duplicated objects when the frames are combined. Additionally, certain spatial regions between adjacent frames may not be captured due to timing differences or device motion, leading to missing regions in a panoramic image generated using conventional image stitching techniques. Accordingly, when the frames 21, 22, and 23 are stitched together using conventional panoramic image generation methods, visual artifacts such as duplicated objects, missing regions, or geometric distortions may occur, thereby degrading the quality of the resulting panoramic image.
[0076] FIG. 2 illustrates an example configuration of a front camera and a rear camera of an electronic device, along with their respective fields of view (FOVs).
[0077] As shown in the left portion of FIG. 2, a front camera 210 and a rear camera of the electronic device 100 are disposed on opposite surfaces of the electronic device 100 and have different optical axes. Accordingly, a field of view captured by the rear camera is not aligned with a field of view captured by the front camera with respect to a reference point. In this state, even when images captured by the front camera 210 and the rear camera are combined, the captured regions are unevenly distributed around the reference point, which may result in inaccurate spatial coverage when generating 360-degree image content.
[0078] As shown in the right portion of FIG. 2, an image corresponding to the rear camera may be transformed such that an optical axis of the rear camera is aligned with an optical axis of the front camera. In this example, a transformed rear camera field of view 222 is obtained based on re-projecting an image captured by the rear camera, such that the transformed rear camera field of view 222 is aligned with the front camera field of view 210 with respect to the reference point. By aligning the rear camera image with the front camera image in this manner, the electronic device 100 can obtain images that more uniformly cover regions around the reference point. Accordingly, 360-degree image content can be generated more accurately using the aligned images captured by the front camera 210 and the transformed rear camera image.
[0079] FIG. 3 depicts example scenarios illustrating limitation with the electronic device depicting a top view, a FOV and a missing FOV, according to existing arts.
[0080] As shown in FIG. 3, telephoto lenses, such as a 3Х telephoto lens or a 1.5Х telephoto lens, provide relatively narrow fields of view, for example approximately 38 degrees or 48 degrees. A normal lens provides a wider field of view, for example approximately 72 degrees, while wide-angle and super wide-angle lenses provide fields of view of approximately 113 degrees and 130 degrees, respectively. Even a fisheye lens provides a field of view of approximately 160 degrees, which is significantly less than a full 360-degree field of view. Accordingly, a single camera equipped with any of the lenses illustrated in FIG. 3 is inherently incapable of capturing a complete 360-degree surrounding environment. In an embodiment of the disclosure, even when multiple cameras are used in an electronic device, the combined field of view of the cameras remains substantially less than 360 degrees and is further affected by misalignment between camera optical axes, as described with reference to FIG. 2.
[0081] The embodiments herein achieve methods and systems (or electronic device) for synthesizing 360-degreeimages using an electronic device. In an embodiment of the disclosure, the method can be for capturing a 360-degree image of a scene by an electronic device. In an embodiment of the disclosure, the method includes capturing, by the electronic device, a first image of the scene using a first imaging device, where the first imaging device is positioned on a first surface of the electronic device. In an embodiment of the disclosure, the method includes capturing, by the electronic device, a second image of the scene using a second imaging device, where the second imaging device is positioned on the first surface of the electronic device. In an embodiment of the disclosure, the method includes capturing, by the electronic device, a third image of the scene using a third imaging device, where the third imaging device is positioned on a second surface of the electronic device, wherein the second surface faces in a direction opposite to the first surface. In an embodiment of the disclosure, the method includes generating, by the electronic device, a 3D representation of the scene using the first image and the second image. In an embodiment of the disclosure, the method includes generating, by the electronic device, a fourth image to align an optical axis of the first imaging device with the third imaging device based on the generated 3D representation and a pre-defined calibration matrix. In an embodiment of the disclosure, the method includes creating, by the electronic device, a two-dimensional (2D) rectilinear image representative of a 360-degree view of the scene using the third image and the fourth image.
[0082] In an embodiment of the present disclosure, a first surface and a second surface may correspond to surfaces of an electronic device 100 that face in opposite or substantially opposite directions. For example, when the electronic device 100 is a smartphone, the first surface may correspond to a rear surface of the electronic device. In this example, a first imaging device and a second imaging device positioned on the rear surface may correspond to a first primary camera and a second primary camera disposed on the rear surface of the electronic device. Further, the second surface may correspond to a front surface of the electronic device, and a third imaging device positioned on the front surface may correspond to a secondary camera. However, the present disclosure is not limited to the above-described example. The first surface and the second surface may not necessarily face completely opposite directions, and may instead be interpreted as different surfaces of the electronic device that face different directions. In this specification, references to a "rear" or a "front" of the electronic device are provided merely for ease of explanation, and may be understood as being replaced with, or corresponding to, the first surface and the second surface, respectively.
[0083] In this specification, the "3D representation of the scene" may refer to any three-dimensional structural representation generated based on image data, depth information, calibration parameters, or a combination thereof. In an embodiment of the disclosure, the 3D representation may comprise a depth map, a point cloud, a voxel-based representation, a mesh-based representation, or any other spatial representation that associates image pixels with corresponding three-dimensional coordinates. In an embodiment of the disclosure, the 3D representation is generated using stereo depth estimation based on the first image and the second image, and may further incorporate calibration information, camera pose information, or SLAM-based mapping information. The 3D representation is not limited to a complete or watertight geometric model of the scene, and may instead represent a partial, sparse, or approximate spatial structure sufficient for reprojection, alignment, or view synthesis purposes.
[0084] In this specification, the "2D rectilinear image representative of a 360-degree view of the scene" refers to a planar image representation formed by combining image content obtained from multiple imaging devices of the electronic device. In an embodiment of the disclosure, the 2D rectilinear image may include image regions derived from the third image and the fourth image, which are spatially aligned using calibration information and reprojection based on a 3D representation of the scene. The 2D rectilinear image does not correspond to a final spherical or equirectangular projection, but serves as an intermediate representation used for subsequent processing, including projection, inpainting, diffusion-based generation, or spherical mapping, to generate a 360-degree image of the scene. Accordingly, the term "representative of a 360-degree view" indicates that the 2D rectilinear image is used as an intermediate image containing visual information utilized in generating a 360-degree image, rather than requiring that the 2D rectilinear image itself directly represent or display a complete 360-degree view.
[0085] The proposed method will establish connect between the smartphone and XR device (e.g., VST) creating an ecosystem, so that the immersive XR experience would be much more richer than 2D. hence, the user need not invest any extra price to get the 360 experience. Based on the proposed method, the user man can create 360 degree content just using his phone for use in AR / VR headset experience without buying a specialized 360 degree content generation device. The proposed method uses a conditional grid which can significantly increase the image quality in spherical space.
[0086] As embodiments herein use the calibration information and depth reprojected image to fill the primary and secondary image in the rectilinear space, hence no overlapping FOV or feature matching is required, thereby making panorama video possible.
[0087] Embodiments herein make it possible for 360-degree media (e.g., 360-degree video, or the like) to be captured using the regular electronic device, wherein the capture media can be used in scenarios, such as, but not limited to, providing immersive video tours (for locations, such as, but not limited to, cities, monuments, museums, historical locations, and so on), immersive adventure experiences, and so on.
[0088] Referring now to the drawings, and more particularly to FIGS. 4 through 31, where similar reference characters denote corresponding features consistently throughout the figures, there are shown embodiments.
[0089] FIG. 4 depicts various hardware components of an electronic device 100 for capturing a 360-degree image of a scene, according to embodiments as disclosed herein. The electronic device 100 can be, for example, but not limited to a smart phone, a foldable phone, a laptop, a desktop computer, a notebook, a smart watch, a smart TV, a tablet, an immersive device, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, a Head-mounted display (HMD), a visual see-through (VST) device, or the like.
[0090] The electronic device 100 includes a processor 110, a communicator 120, a memory 130, a real-world 360-degree scene capturing controller 140, a data driven controller 150, a sensor 160, a first imaging device 170, a second imaging device 180 and a third imaging device 190. The processor 110 is coupled with the communicator 120, the memory 130, the real-world 360-degree scene capturing controller 140, the data driven controller 150, the sensor 160, the first imaging device 170, the second imaging device 180 and the third imaging device 190.
[0091] The real-world 360-degree scene capturing controller 140 captures a first image of the scene using the first imaging device 170 (e.g., first primary camera), where the first imaging device 170 is positioned on a first surface of the electronic device 100. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 receives an estimated transformation parameter associated with a plurality of features (e.g., edges, positions, field lines or the like) by matching the plurality of features in an actual second image frame and an expected second image frame of the secondary camera 180. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 calculates a time delta between actual second image frame and the expected second image frame. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 checks (or determines) if the time delta is more than a determined threshold. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 adjusts an electronic device parameter to reduce a camera synchronisation error.
[0092] For example, in a real-world scenario involving 360-degree scene capture, such as live VR streaming of a sports event, the real-world 360-degree scene capturing controller 140 receives estimated transformation parameters by matching the plurality of visual features such as player outlines, field markings, or other identifiable elements, between the actual second image frame captured by secondary camera 180 and an expected second image frame that the system predicts based on prior motion and synchronization data. After obtaining the transformation parameters, the real-world 360-degree scene capturing controller 140 calculates the time delta, which is the difference in timing between the actual and the expected second image frames. If this time delta exceeds a predetermined threshold, indicating a synchronization error, the real-world 360-degree scene capturing controller 140 takes corrective action. Specifically, it adjusts one or more electronic device parameters such as modifying the camera's internal clock, applying a timing offset, or fine-tuning the frame capture delay, to bring the camera back into synchronization. This process ensures that all cameras in the rig are temporally aligned, which is critical for producing seamless 360-degree video without visual artifacts, jitter, or misalignment when the video is stitched and viewed in a virtual reality environment.
[0093] In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 captures a second image of the scene using the second imaging device (e.g., second primary camera) 180, where the second imaging device 180 is positioned on the rear of the electronic device 100. In an embodiment of the disclosure, a relationship between the first imaging device 170, and the second imaging device 180 is retrieved from a pre-defined calibration matrix. In an embodiment of the disclosure, the pre-defined calibration matrix is determined by positing the electronic device 100 at a center of a calibration chamber (not shown) such that the first imaging device 170 and the second imaging device 180 are aligned to a center of the calibration chamber, receiving a sets of images from the first imaging device 170 and the second imaging device 180 at the same time, and determining the pre-defined calibration matrix by reprojecting key points associated with the sets of images. The calibration chamber comprises identical patterns against both the first imaging device 170 and the second imaging device 180 placed in a setup.
[0094] In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 captures a third image of the scene using the third imaging device 190 (e.g., secondary camera), where the third imaging device 190 is positioned on a front side of the electronic device 100. In an embodiment of the disclosure, the third image of the scene is time synchronized with the first image and the second image.
[0095] In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 generates a 3D representation of the scene using the first image and the second image. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 generates a fourth image by re-projecting the first image using the generated 3D representation and the pre-defined calibration matrix to align an optical axis of the first imaging device 170 with the third imaging device 190.
[0096] In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 estimates a depth map of the scene using the first image received from the first imaging device 170 using a stereo depth. The real-world 360-degree scene capturing controller 140 generates a mesh based on the depth map. The real-world 360-degree scene capturing controller 140 reprojects a red-green-blue (RGB) image from a first imaging device view to the centre by using calibration information and 3D points from the mesh. The real-world 360-degree scene capturing controller 140 generates the fourth image by reprojecting the RGB image from the first imaging device view to the centre by using calibration information and 3D points from the mesh.
[0097] In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 creates a two-dimensional (2D) rectilinear image representative of a 360-degree view of the scene using the third image and the fourth image. Each pixel in the rectilinear image from image coordinates are mapped to spherical coordinates based on a horizontal angular component (i.e., angle in the horizontal direction from the center of the image) and a vertical angular component (i.e., angle in the vertical direction from the center of the image). The horizontal and vertical angular components refer to angles that describe a direction a pixel is pointing in the 3D space, relative to the center of the camera's FOV.
[0098] In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 generates an equi-rectangular projection of the scene by passing the 2D rectilinear image to a data driven model (e.g., machine learning (ML) model, Artificial intelligence (AI) model, or the like). The data driven model is pre-trained to generate a missing region using a pre-defined region-specific condition. The pre-defined region-specific condition refers to how the missing region will be generated. While generating, the real-world 360-degree scene capturing controller 140 use conditions like pixels closer to the other in the spherical space, pixels close to equator, pixels close to poles, etc.
[0099] In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 stitches the first image received from the first imaging device 170 and the second image received from the second imaging device, based on the pre-defined calibration matrix. The real-world 360-degree scene capturing controller 140 computes a conditional grid for a part of the stitched image. The real-world 360-degree scene capturing controller 140 uses the conditional grid as an input to a diffusion based model to generate pixels. The diffusion-based model in the 360 image generation refers to a generative AI approach that creates 360-degree panoramic images (often in equirectangular format) using diffusion models (a class of probabilistic generative models known for producing high-quality and photorealistic images). In other words, the diffusion-based model can be a generative framework that learns to create 360-degree panoramic images by iteratively denoising random noise through a trained reverse diffusion process, conditioned or unconditioned on input signals such as text, depth, or partial images, while preserving the spatial continuity and spherical geometry specific to panoramic scenes. Values in the conditional grid indicate at least one of: border pixels, pixels near to equator, and pixels near poles. The real-world 360-degree scene capturing controller 140 generates the equi-rectangular projection of the scene based on the generated pixels.
[0100] In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 receives the rectilinear image having a defined focal length and optical center coordinates, where the defined focal length is dynamically calculated based on the field of view and resolution of the rectilinear image. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 maps each pixel in the rectilinear image from image coordinates to spherical coordinates. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 generates an equirectangular projection image by mapping the spherical coordinates to two-dimensional coordinates. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 renders the pixel values from the rectilinear image into the equirectangular projection image based on the computed coordinates. In an embodiment of the disclosure, the equirectangular projection image has an aspect ratio of 2:1, corresponding to a 360° horizontal and 180° vertical spherical field of view.
[0101] It should be understood that the real-world 360-degree scene capturing controller 140 described in the embodiments may be implemented by, or correspond to, the one or more processors 110 executing instructions stored in the memory 130. Therefore, operations described herein as being performed by the real-world 360-degree scene capturing controller 140 may be equivalently understood as being performed by the processor 110. Although the embodiments describe that operations are performed by the real-world 360-degree scene capturing controller 140 for clarity of functional description, it is to be understood that the controller 140 is physically implemented by the processor 110, and thus, the described operations are executed by the processor 110.
[0102] In an embodiment of the disclosure, the at least one missing region in the 2D rectilinear image is generated by receiving the plurality of images from the first imaging device 170 and the second image from the second imaging device 180, location of the electronic device 100, orientation using at least one sensor 160 and a visual feature, creating a map of an environment, estimating a position of the electronic device 100 with respect to the map, and generating the at least one missing region based on the estimated position of the device with respect to the map during the 360 images generation.
[0103] In an embodiment of the disclosure, the at least one missing region in the 2D rectilinear image is generated by identifying at least one background object and missing region from a set of frames using map and a pose data, classifying landmarks and noise using redundant information based on the identification, extracting at least one feature from the at least one background object and the missing region using the landmark, the noise, the map and the pose data, and generating at least one missing region based on the at least one extracted feature.
[0104] In an embodiment of the disclosure, the at least one missing portion in the 2D rectilinear image is generated using a Simultaneous localization and mapping (SLAM) information by receiving the plurality of images from the first imaging device 170 and the second imaging device 180, and device's location and orientation using IMU and visual features.
[0105] In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 generates a 360-degree content of the scene based on the generated equi-rectangular projection of the scene. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 displays the generated 360-degree image on a display (e.g., LCD, LED display or the like) for at least one of: a real-time preview, and a post-capture editing.
[0106] In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 receives a user input to activate a 360- degree content capture mode on the electronic device 100. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 performs a stereo depth estimation using a plurality of image streams from the first imaging device 170 and the second imaging device 180 to generate a depth map. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 reconstructs a 3D volume of the scene using the depth map. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 reprojects texture and color information from an ultra-wide image stream received from the first imaging device 170 into a viewpoint determined by extrinsic parameters.
[0107] In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 calibrates the first imaging device 170 and the third imaging device 190 using captured images to estimate extrinsic parameters comprising a pose matrix representing rotation and translation. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 reprojects texture and color information from the ultra-wide image stream into a viewpoint determined by the extrinsic parameters. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 converts both the reprojected image and the image from the third imaging device 190 into an equirectangular projection format.
[0108] In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 applies the data driven model trained for inpainting to fill in missing or occluded regions of the equirectangular projection. In an embodiment of the disclosure, the real-world 360-degree scene capturing controller 140 generates a complete 360-degree spherical image using the inpainted equirectangular projection.
[0109] The real-world 360-degree scene capturing controller 140 is implemented by analog and / or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits and the like, and may optionally be driven by firmware.
[0110] The processor 110 may include one or a plurality of processors. The one or the plurality of processors may be a general-purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU). The processor 110 may include multiple cores and is configured to execute the instructions stored in the memory 130.
[0111] In an embodiment of the disclosure, the processor 110 is configured to execute instructions stored in the memory 130 and to perform various processes. The memory 130 also stores instructions to be executed by the processor 110. The memory 130 may include non-volatile storage elements. Examples of such non-volatile storage elements may include magnetic hard discs, optical discs, floppy discs, flash memories, or forms of electrically programmable memories (EPROM) or electrically erasable and programmable (EEPROM) memories. In addition, the memory 130 may, in some examples, be considered a non-transitory storage medium. The term "non-transitory" may indicate that the storage medium is not embodied in a carrier wave or a propagated signal. However, the term "non-transitory" should not be interpreted that the memory 130 is non-movable. In certain examples, a non-transitory storage medium may store data that can, over time, change (e.g., in Random Access Memory (RAM) or cache).
[0112] In an embodiment of the disclosure, at least one of the plurality of modules / controller may be implemented through the AI model / the ML model using the data driven controller 150. The data driven controller 150 can be an ML model based controller and an AI model based controller. A function associated with the AI model may be performed through the non-volatile memory, the volatile memory, and the processor 110. the data driven controller 150 may be implemented by, or correspond to, the one or more processors 110, which may include specialized processing units such as a Neural Processing Unit (NPU), a Graphics Processing Unit (GPU), or a Tensor Processing Unit (TPU). Accordingly, operations described as being performed by the data driven controller 150, such as inference, training, or data processing using AI models, are executed by the processor 110 utilizing the memory 130. The distinction between the real-world 360-degree scene capturing controller 140 and the data driven controller 150 is logical for describing different functions, and they may be integrated into a single physical processor or distributed across multiple processors.
[0113] The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or AI model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.
[0114] Here, being provided through learning means that a predefined operating rule or AI model of a desired characteristic is made by applying a learning algorithm to a plurality of learning data. The learning may be performed in a device itself in which AI according to an embodiment is performed, and / o may be implemented through a separate server / system.
[0115] The AI model may include of a plurality of neural network layers. Each layer has a plurality of weight values, and performs a layer operation through calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.
[0116] The learning algorithm is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0117] Although FIG. 4 shows various hardware components of the electronic device 100 but it is to be understood that other embodiments are not limited thereon. In other embodiments, the electronic device 100 may include less or more number of components. In an embodiment of the disclosure, the labels or names of the components are used only for illustrative purposes and does not limit the scope of the invention. One or more components can be combined together to perform the same or substantially similar function in the electronic device 100.
[0118] FIG. 5 is a flowchart (S500) illustrating a method for capturing the 360-degree image of the scene by the electronic device 100, according to embodiments as disclosed herein. The operations (S502-S512) are handled by the real-world 360-degree scene capturing controller 140.
[0119] In step S502, the method includes capturing the first image of the scene using the first imaging device 170, where the first imaging device 170 is positioned on the rear side of the electronic device 100. In step S504, the method includes capturing the second image of the scene using the second imaging device 180, where the second imaging device 180 is positioned on the rear side of the electronic device 100.
[0120] In step S506, the method includes capturing the third image of the scene using the third imaging device 190, where the third imaging device 190 is positioned on the front side of the electronic device 100. In step S508, the method includes generating the 3D representation of the scene using the first image and the second image.
[0121] In step S510, the method includes generating the fourth image by re-projecting the first image using the generated 3D representation and a pre-defined calibration matrix to align the optical axis of the first imaging device 170 with the third imaging device 190. In step S512, the method includes creating the 2D rectilinear image representative of the 360-degree view of the scene using the third image and the fourth image.
[0122] FIG. 6 is another flowchart (S600) illustrating a method for capturing the 360-degree image of the scene by the electronic device 100, according to embodiments as disclosed herein. The operations (S602-S610) are handled by the real-world 360-degree scene capturing controller 140.
[0123] In step S602, the method includes receiving the user input to activate the 360-degree content capture mode on the electronic device 100. In step S604, the method includes performing the stereo depth estimation using the plurality of image streams from the first imaging device 170, and the second imaging device 180 to generate the depth map.
[0124] In step S606, the method includes reconstructing the 3D volume of the scene using the depth map. In step S608, the method includes reprojecting the texture and color information from the ultra-wide image stream received from the first imaging device 170 into the viewpoint determined by extrinsic parameters. In step S610, the method includes converting both the reprojected image and the image from the third imaging device 190 into an equirectangular projection format.
[0125] FIG. 7 depicts an example process of capturing the 360 degree image of the real-world scene using the electronic device 100, according to embodiments as disclosed herein.
[0126] The process begins with the synchronized capture of image data from the primary cameras 711, 712 and secondary rear camera 720. Due to differences in sensor characteristics, a style transfer 730 technique is applied to harmonize image quality, particularly across HDR, noise, and brightness levels. The spatial relationship between the primary cameras and secondary camera is retrieved using extrinsic calibration data, which defines their relative position and orientation.
[0127] The depth map 740 is then generated from the stereo image pair. This depth information is used to reproject the image from the primary cameras 711, 712 effectively transforming its viewpoint to align with the optical center of the secondary camera 720. The reprojected image is then merged with the actual image captured by the secondary camera 720 to create a rectilinear representation 750 of the scene, using the calibration parameters to ensure accurate geometric alignment.
[0128] The relationship between the primary cameras 711, 712 and secondary camera 720 is obtained from extrinsic calibration. Using the computed depth map 740, the primary camera image is reprojected to match the viewpoint of the secondary camera 720. The resulting reprojected image 750 is then fused with the real image from the secondary camera and mapped into rectilinear space 760, guided by the camera calibration parameters.
[0129] Next, an equirectangular image 770 is generated from the merged rectilinear content. However, due to occlusions and limited camera FOV, this equirectangular image 770 often contains holes (i.e., missing pixel regions). The conditional grid 780 is generated alongside, encoding spatial semantics such as pole regions, the equator, and areas of missing data. Using this equirectangular image 770 and the conditional grid 780 as inputs, the trained diffusion model is used to inpaint the missing regions (e.g., holes), inferring plausible content based on learned panoramic context. Finally, the completed equirectangular image 770 is projected into spherical space 790, making it suitable for immersive 360-degree viewing.
[0130] For instance, in a scenario where a user captures an outdoor park scene using the smartphone equipped with primary and rear cameras, the smartphone first aligns and merges images from both cameras into a partial 360-degree view. Due to occlusions behind the user or beyond the camera's FOV, certain areas remain unfilled. The diffusion model then realistically fills in the missing sky, foliage, or walking paths based on the existing context, resulting in a seamless, viewable 360-degree spherical image.
[0131] FIG. 8 is an example method for capturing the 360-degree image of the real-world scene using the electronic device 100, according to embodiments as disclosed herein.
[0132] In step S802, the method includes capturing and restoring the frames, where the primary and secondary cameras are switched on simultaneously in a time synchronized manner and frames are captured. Since the sensors 160 are different, the differences in the image quality factors such as HDR, Noise, Brightness, and so on, are compensated, so that both the images can be blended seamlessly.
[0133] In step S804, the method includes calibrating the secondary and the primary camera, where the secondary camera is aligned to the centre of the device horizontally, whereas the primary camera is aligned away from the centre of the device 100 to accommodate the sensors 160 physically. The calibration is important to understand the spatial relationship between the secondary and primary camera.
[0134] In step S806, the method includes reprojecting the primary cameras with the secondary camera. As the cameras are not aligned to the same axis, so that the 360-degree view can converge. The depth based reprojection of the primary camera is reprojected with the secondary camera. The reprojection is done by estimating stereo depth using the primary camera images, estimating 3D points, reprojecting from current view to the view of primary camera.
[0135] In step S808, the method include performing the equirectangular projection and SLAM diffusion. The equirectangular projection can be used for generated 360-degree images. The primary reprojected and secondary camera images are filled in the equirectangular space. Inpainting model based on diffusions are used to fill the remaining space. But since the transformation to map from equirectangular to sphere is known, it can be used as a condition to generate the missing pixels.
[0136] FIG. 9 depicts an example scenario in which the electronic device 100 captures the 360-degree image of the real-world scene, according to embodiments as disclosed herein. For example, a user 900 named Bob travels to Jim Corbett National Park and plans to capture his experience to share with friends and family. Bob opens the Galaxy Smartphone Camera 360 application (for example), which triggers the 360 engine 910 to begin generating content in the background. The method enables the synthesis of 360-degree images using a regular smartphone. In existing systems, the front and rear cameras of the smartphone have not been used together to generate true 360-degree content 920. However, the 360-degree scene capturing controller 140 (e.g., 360 engine or the like), running locally on the electronic device 100, generates a 360-degree image, encodes it, and stores it directly on the smartphone.
[0137] The method starts by performing axis alignment of newly detected features, along with calibration of challenging features in images captured simultaneously by the front and rear cameras. When these images are stitched, visual gaps may appear at the stitching boundaries due to misalignment or missing data. In the next step, the generative AI-based inpainting is applied to fill these gaps, creating a natural-looking equirectangular projection and delivering a seamless 360-degree viewing experience-all from a single smartphone.
[0138] Additionally, the 360 engine allows the user to seamlessly share the encoded 360-degree content 920 with other smartphones or VSTs. As a result, Bob's friends receive the 360-degree content on their devices and can experience Jim Corbett National Park as if they were there themselves.
[0139] FIGS. 10, 11, and 12 depict example scenarios, wherein images are captured, and restored while capturing the 360-degree image of the real-world scene, according to embodiments as disclosed herein.
[0140] As shown in FIG. 10-FIG. 12, two rear-facing cameras positioned opposite to each other, facing 180-degree apart, are used to simultaneously capture different sections of the scene. However, due to differences in the hardware (e.g., sensor type, lens quality, image processing pipeline), the captured images often vary in characteristics such as brightness, noise level, resolution, and dynamic range. To ensure a seamless blend between these images, it becomes essential to normalize their appearance. This is achieved using a style transfer-based deep learning approach, where one image is adjusted to match the style of the other, effectively harmonizing their visual qualities.
[0141] The image from the primary camera is used as the "style reference," while the image from the secondary camera serves as the "content input." A deep learning model takes these inputs and generates a corrected version of the secondary camera image that stylistically resembles the primary image. The result is a harmonized pair of images that can be merged seamlessly in further stages of 360-degree image generation. Further, the visual alignment ensures that the stitched output does not exhibit abrupt transitions or artifacts between the two views.
[0142] FIG. 10 illustrates an example scenario in which a first image captured by a first rear-facing camera and a second image captured by a second rear-facing camera are provided as input images prior to appearance normalization. In FIG. 10, the two images represent different sections of the same real-world scene captured simultaneously by cameras facing approximately 180 degrees apart, and the images exhibit visually inconsistent characteristics due to differences in camera hardware and image processing pipelines.
[0143] As shown in FIG. 11, an image captured by a primary camera is provided as a style reference image 1110, and an image captured by a secondary camera is provided as a content image 1120.
[0144] Feature representations are extracted from the style reference image 1110 and the content image 1120 through a plurality of feature layers of a neural network, and a style loss and a content loss are computed based on the extracted feature representations. The style loss and the content loss are jointly minimized within a style transfer module 1130 so as to transfer visual characteristics of the style reference image 1110 to the content image 1120 while preserving scene structure of the content image 1120.
[0145] As a result of the loss minimization performed by the style transfer module 1130, a style-normalized image 1140 is generated, wherein the generated image maintains content information derived from the secondary camera while exhibiting visual characteristics substantially similar to those of the primary camera image.
[0146] As shown in FIG. 12, a primary camera image 1210 is used as a style reference image, and a secondary camera image 1220 is used as a content image.
[0147] Through the style transfer operation described with reference to FIG. 11, a style-transferred secondary camera image 1230 is generated, wherein visual characteristics of the primary camera image 1210, including at least one of color tone, brightness, and contrast, are reflected in the secondary camera image 1220.
[0148] The style-transferred secondary camera image 1230 is visually harmonized with the primary camera image 1210, thereby reducing visual discontinuities between images captured by different cameras and enabling seamless subsequent processing, including image merging and 360-degree image generation.
[0149] FIG. 13 depicts an example process (S1300) of performing non overlapping camera calibration while capturing the 360 degree image of the real-world scene, according to embodiments as disclosed herein.
[0150] In Step S1302, the method includes determining the intrinsic parameters of the cameras by positioning the smartphone at the center of the calibration chamber, so as to ensure that both the primary and secondary cameras are aligned with the central axis of the chamber. The method further involves configuring the calibration chamber with identical reference patterns visible to both cameras, using a highly precise and symmetrical setup.
[0151] In Step S1304, the method includes capturing synchronized image sets from the primary and secondary cameras simultaneously, recording the reference patterns. In Step S1306, the method includes computing the extrinsic parameters between the primary and secondary cameras by reprojecting detected key points from the captured images as part of a stereo calibration process. In Step S1308, the method includes iteratively adjusting the camera parameters and optimizing the setup until the extrinsic parameters result in a minimal reprojection error, thereby achieving high-accuracy calibration.
[0152] FIG. 14 and FIG. 15 depict example scenarios in which the proposed method performs a non-overlapping camera calibration and detecting calibration rig details, according to embodiments as disclosed herein. A two-camera calibration problem is formulated as a set of rigidity constraints imposed by transformations between the cameras and calibration patterns. The multi-camera calibration problem is solved by iteratively solving for variables using closed-form techniques, minimizing algebraic error over the set of rigidity constraints, and minimizing reprojection errors. The extrinsic calibration of non-overlapping field of view cameras is non-trivial. Embodiments herein disclose a novel non-overlapping camera calibration procedure to do the same. As depicted in FIG. 14, two calibration targets (A, B) rigidly attached opposite to each other. Both targets (A, B) are exactly same and placed with laser-to-laser alignment. Images of partial targets can be used to estimate camera pose, targets consist of unique fiducial markers or caliburation chart. As shown in FIG. 14, an electronic device 100 is mounted on a stabilizing apparatus and positioned within a calibration environment, such that cameras of the electronic device 100 are configured to capture images of the calibration targets (A, B) arranged in opposite directions. The electronic device 100 can be in a stabilizing apparatus (tripod), configured to position the rear-facing camera to capture images from the primary target (i.e., target A) and simultaneously align the secondary-facing camera to capture images of a secondary target (i.e., target B), identical to the target A, thereby establishing a stereoscopic configuration. The two cameras of the electronic device 100 are facing in the opposite directions (i.e., looking 180-degreeaway from each other). The size of the rig and the sequence of patterns for each calibration target are identical and known beforehand. The images from both the secondary-facing and rear-facing cameras can be considered to be looking at the same calibration target (secondary-facing camera target in this case). The intrinsics of each camera in the rig to be calibrated are determined. Then, multiple sets of images are captured with the electronic device 100 mounted in the calibration rig at different positions, using both the secondary-facing and rear-facing cameras for calibration.
[0153] Since it is assumed that both the calibration targets are identical, the images captured by both the secondary-facing and rear-facing camera images can be assumed to be looking at the same calibration target (secondary-facing camera target in this case, without loss of generalization), and then the stereo calibration can be used to determine camera extrinsics. FIG. 15 illustrates an example image captured in an actual calibration environment corresponding to the conceptual configuration described with reference to FIG. 14. As shown in FIG. 15, the electronic device 100 is placed within a space surrounded by calibration patterns, and images are captured while the electronic device 100 is positioned to face calibration targets arranged in opposite directions.
[0154] The image shown in FIG. 15 represents an example implementation of the non-overlapping camera calibration setup, in which cameras of the electronic device 100 observe calibration patterns from different directions in a controlled environment. The orientation of one of the rear cameras is changed with the help of extrinsics to get its original pose as suggested as depicted in figure 15.
[0155] Building a 360-degree view of a scene with a device camera system requires the following: intrinsic camera calibration parameters (which are used to map the 3D world to a 2D image), extrinsic camera calibration parameters (which can be used to determine the camera's position and orientation in 3D, and help establish camera locations relative to each other), and field of view of each camera (which can be used to identify what part of the 360-degree scene is covered, can be inferred from intrinsics)
[0156] FIG. 16 depicts an example scenario of the proposed calibration process while capturing the 360 degree image of the real-world scene, according to embodiments as disclosed herein.
[0157] In an existing methods, in typical panorama image creation, the image stitching is performed based on feature matchings between consecutive images, which require overlapping field of view. Due to this, the panorama video is not possible. Based on the proposed method, the calibration information and depth reprojection image is used to fill the primary and the secondary image in the rectilinear space. Hence, no overlapping of field of view or feature matching is required, to create the panorama video.
[0158] As illustrated in FIG. 16, a first image 1610 and a second image 1620 are captured from different directions using cameras having non-overlapping fields of view. Calibration information associated with the cameras is used to align viewpoints of the first image 1610 and the second image 1620, and depth-based reprojection is applied to map image contents into a common rectilinear space. Based on the calibration information and the depth reprojection, image contents of the first image 1610 and the second image 1620 are combined to generate a rectilinear panorama image 1630 representing an expanded field of view.
[0159] FIG. 17 to FIG. 19 depict example scenarios (S1700, ) in which depth map based reprojection is performed while capturing the 360 degree image of the real-world scene, according to embodiments as disclosed herein.
[0160] In step S1702, the method includes receiving both the camera images for stereo rectification. In an embodiment of the disclosure, the method includes receiving the estimated transformation parameters associated with significant features by matching the significant features between the actual second image frame and the expected second image frame for block matching as depicted in step S1704.
[0161] In step S1706, the method includes performing the depth estimation using stereo depth estimation. As the primary and the secondary cameras are not placed in the same line of axis. Hence, directly using both the camera images will lead to different viewpoints. Additionally, calculating the time delta between the actual second image frame and the expected second image frame and reproject the primary camera image to align to the centre of the device using depth estimated from stereo camera.
[0162] In step S1708, the method includes checking whether the calculated time delta exceeds the predefined threshold. In an embodiment of the disclosure, the 3D points are used for depth based reprojection, wherein the images from primary camera and secondary camera are aligned to the same line of axis. In step S1710, the method includes adjusting camera parameters, such as exposure, to reduce the camera synchronization error based on the detected time delta, so as to obtain a centre aligned primary camera view.
[0163] In this phase, the image view is reprojected from the perspective of the primary camera to the centre of the device by estimating the depth map of the scene from the primary camera image using stereo depth, generating mesh(es) from the depth map, and reprojecting RGB images from the primary camera view to the centre by using calibration information and 3D points from the mesh.
[0164] As depicted in FIG. 18, the primary and secondary cameras are not placed in the same line of axis. Hence directly using them will lead to different views points. Hence, the primary camera image is reprojected to align to the centre of the device using depth estimated from the stereo cameras. For example, an electronic device 100 includes a primary camera and a secondary camera arranged along different axes, such that viewing rays corresponding to images captured by the primary camera and the secondary camera originate from different optical centers and are misaligned with respect to a centre viewpoint of the electronic device 100. Accordingly, directly using the images captured by the primary camera and the secondary camera results in different viewpoints.
[0165] As shown in FIG. 19, in this phase, the proposed method can be used to reproject the image view from the perspective of the primary camera to the centre of the device. In order to do this, the method includes estimating the depth map of the scene from the primary camera image using the stereo depth. In an embodiment of the disclosure, the method includes generating the mesh from the depth map. In an embodiment of the disclosure, the method includes reprojecting the RGB images from the primary camera view to the centre by using the calibration information and the 3D points from the mesh. As shown in FIG. 19, the proposed method reprojects image content based on depth information to align the viewpoint of the image to a centre viewpoint of the electronic device 100. In this phase, a depth map 1910 is estimated from an image captured by the primary camera, and a mesh 1920 is generated based on the depth map 1910. Using calibration information and three-dimensional points from the mesh 1920, image content is reprojected along reprojection rays toward a centre viewpoint, thereby generating a reprojected image aligned to the centre of the electronic device 100.
[0166] FIG. 20 is a flow chart illustrating a method for generating Equirectangular images using conditional grids, according to embodiments as disclosed herein.
[0167] In step S2002, the method includes receiving the first image from the primary camera and the second image from the secondary camera, and stitching the images based on determined calibration information. In step S2004, the method includes calculating the conditional grid for various parts of the stitched images. In step S2006, the method includes using the conditional grid as an input to the diffusion-based model to generate pixels. The values in the conditional grid indicating the border pixels, pixels near to equator, pixels near poles, etc.
[0168] FIG. 21 depicts an equirectangular projection and diffusion operation while capturing the 360 degree image of the real-world scene, according to embodiments as disclosed herein.
[0169] The projection obtained by a sphere mesh un-wrapped on a flat rectangular plane surface is known as an equirectangular projection. It is the direct mapping of the sphere latitude and longitude on the horizontal and vertical coordinate system, respectively. The diffusion models are generative models, meaning that they are used to generate data similar to the data on which they are trained. Fundamentally, the diffusion models work by destroying training data through the successive addition of Gaussian noise and then learning to recover. It can be trained to recover parts missing due to FOV (Field of View) limitation in electronic devices 100, as depicted in FIGS. 21, 22, 23.
[0170] As illustrated in FIG. 21, an input image 2110 is projected onto a spherical domain and converted into an equirectangular representation 2120.
[0171] Each pixel is mapped to a corresponding point on virtual sphere using latitude ( ) and longitude ( ). The image's focal length and optical center are used to calculate the angular displacement.
[0172] FIG. 22 depicts an example scenario in which an equirectangular projection and diffusion is performed while capturing the 360 degree image of the real-world scene, according to embodiments as disclosed herein.
[0173] As illustrated in FIG. 22, a rectilinear image 2210 is mapped to a spherical domain, wherein pixel coordinates (x, y) of the rectilinear image 2210 are transformed into angular coordinates ( , ).
[0174] The rectilinear image is mapped to a sphere. (x, y) is mapped to ( , ) for each pixel using arctangent function (the angle of the pixel relative to the camera viewpoint is calculated)
[0175]
[0176]
[0177] where x, y are the horizontal and vertical coordinates in the input image. cx, cy are the horizontal and vertical coordinates of the optical centre(centre of the image) in pixels. f is the focal length of the virtual camera. This is the measure of how zoomed in or out the image is. An arctangent function computes the angle formed between the pixels displacement from the center and the focal length.
[0178] This is converted to an equirectangular projection 2120. A 2D representation of spherical surface with 2:1 aspect ratio is considered, wherein the horizontal position (u) corresponds to the longitude ( ):
[0179]
[0180] The vertical position (v) corresponds to the latitude ( ):
[0181]
[0182] where We and Wh are the width and height equirectangular output image.
[0183] Based on the angular coordinates ( , ), an equirectangular projection image 2220 is generated.
[0184] Embodiments herein use the conditional grid to make the pixel generation more accurate in the spherical domain. When the generated equirectangular image is mapped to sphere, there will be more concentrated holes near the poles (many-to-one pixel mapping) and scattered holes near the equatorial region (due to scarce mapping). Considering this, along with the latent input here, embodiments herein will add conditional weight for each pixel. The weighting mechanism will act as a prior in the network which include the following example conditions: pixel around the pole having many-to-one pixel mapping, will be given lesser weightage, whereas pixel having one-to many near equator will be given higher weights, and holes (pixels) appearing near the vertical edges should take into account the left and right edge pixels since in spherical domain, and should be seamless. Thus, the diffusion model will be tailor made for the use-case, where it will focus more on pixels with higher weight and less on lesser weight pixels. (weighted fusion block).
[0185] FIG. 23 depicts an example scenario of a SLAM diffusion, according to embodiments as disclosed herein. FIG. 23 depicts the example scenario of using SLAM diffusion based generative process. The SLAM or device motion can be used to identify an area (for example a missing object) in a given frame but was visible in other frames. In an embodiment of the disclosure, the same can generate the missing portion with utmost fidelity.
[0186] As illustrated in FIG. 23, a plurality of image frames 2310 captured from a real-world scene are associated with SLAM information 2320 including device motion information of an electronic device 100. Based on the identified area and the SLAM information 2320, a SLAM diffusion-based generative model 2330 is used to generate image content corresponding to the missing portion, thereby producing a reconstructed image frame 2340.
[0187] FIG. 24 depicts an example scenario, wherein SLAM information is used for generating missing portions, according to embodiments as disclosed herein.
[0188] In step S2402, the method includes receiving a plurality of images (for example, Frame-1, Frame-2, Frame-3, Frame-4 ..., etc.) from the primary camera and the secondary camera. In an embodiment of the disclosure, the method include obtaining a location of the electronic device 100 and an orientation of the electronic device 100 and a visual feature.
[0189] In an embodiment of the disclosure, the electronic device 100 may receive a plurality of images from at least one of a first imaging device, a second imaging device, and a third imaging device. In this context, the first imaging device and the second imaging device may be understood as corresponding to the primary camera, and the third imaging device may be understood as corresponding to the secondary camera.
[0190] In an embodiment of the disclosure of the disclosure, extracting the visual feature may include analyzing the plurality of images to detect distinctive key points, such as corners, edges, or texture patterns, which serve as visual landmarks. The electronic device 100 may track a displacement of these key points across consecutive frames among the plurality of images to determine a relative motion. This extraction process allows the electronic device 100 to associate two-dimensional (2D) image data with three-dimensional (3D) spatial coordinates for mapping.
[0191] In an embodiment of the disclosure of the disclosure, the electronic device 100 may obtain pose information including a location and an orientation of the electronic device 100 using the at least one sensor 160. The at least one sensor 160 may include, but is not limited to, an Inertial Measurement Unit (IMU), an accelerometer, a gyroscope, or a magnetometer. The pose information may be obtained using well-known techniques in the art, such as by integrating data from these sensors to calculate the device's relative movement and rotation, or by performing sensor fusion algorithms that combine sensor data with visual odometry derived from the extracted visual features. In step S2404, the method includes creating a map of the environment and estimating the position of the device with respect to this map, based on the received image data and sensor inputs.
[0192] In an embodiment of the disclosure of the disclosure, the electronic device 100 may create a map of an environment based on the plurality of images. In an embodiment of the disclosure of the disclosure, the electronic device 100 estimate a position of the electronic device 100 with respect to the map based on the location, the orientation and the visual features.
[0193] In step S2406, the method includes using the obtained mapping and positional information for generating missing pixels during the formation of a 360-degree image, thereby enhancing spatial completeness and visual continuity.
[0194] In an embodiment of the disclosure of the disclosure, the electronic device 100 may generate the at least one missing region based on the estimated position of the device with respect to the map.
[0195] In an embodiment of the disclosure, the at least one missing region may be identified based on a comparison between an expected visible region of the environment, determined from the map and the estimated position and orientation of the electronic device 100, and an actually observed region in a current image frame. For example, as illustrated in FIG. 27, when a portion of the environment corresponding to a specific spatial location in the map is determined to be within a field of view of the electronic device 100 based on the estimated position of the electronic device 100, but image data corresponding to the spatial location is absent or occluded in the current frame, the corresponding area may be identified as a missing region. In such a case, image content corresponding to the identified missing region may be obtained from one or more other image frames, captured at different positions or orientations of the electronic device 100, and used as a prior for generating the missing pixels.
[0196] FIG. 25 and FIG. 26 depict example scenarios, in which a generative AI is used for generating the Missing FOV, according to embodiments as disclosed herein.
[0197] For every frame, the Gen AI Model is identifying the background objects or missing portions from other frames (for example Frame-1, ..., Frame-N, Frame-N+1) using the maps, the pose data.
[0198] In an embodiment of the disclosure, different landmarks (for example tables) and noise (for example moving human) are classified using redundant information and the features are extracted from the background or missing portions using the location and pose information. Additionally, the missing FOV is generated based on the extracted features, the objects and the textures. Moreover, the generation will not involve image stitching, rather an intelligent way identify missing portions and use that as a prior to generating, as depicted in FIG. 25-26.
[0199] As illustrated in FIG. 25, a plurality of image frames 2511, 2512, 2513, 2514, 2515, corresponding to Frame-1 through Frame-N, are provided together with map information 2520, and at least one region of interest or missing region 2530 is indicated in an example frame. As illustrated in FIG. 26, image frames 2610, 2621, 2622 corresponding to different time instances are provided as example inputs, and the image information is supplied to a generative AI model 2640 via an input representation 2630, such that an output image 2650 including a generated missing field-of-view region is obtained. FIG. 25 and FIG. 26 illustrate example input-output relationships of the generative AI model and do not limit the manner in which the missing field-of-view region is identified or generated.
[0200] FIG. 27 depicts an example scenario in which the proposed method generates the missing portions using a SLAM based generator, according to embodiments as disclosed herein.
[0201] In an embodiment of the disclosure of the disclosure, the at least one missing region in the 2D rectilinear image may be generated using Simultaneous Localization and Mapping (SLAM) information. In an embodiment of the disclosure of the disclosure, a SLAM-based generator may generate the at least one missing region based on a SLAM algorithm. In an embodiment of the disclosure of the disclosure, the SLAM information may include data used for performing the SLAM algorithm.
[0202] In an embodiment of the disclosure of the disclosure, the SLAM information may comprise a plurality of images from at least one of the first imaging device 170 and the second imaging device 180, and the third imaging device 190, a location and an orientation of electronic device, and visual features extracted from the plurality of images. For example, the proposal network 2700 (i.e., the SLAM-based generator) receives the plurality of images from the primary camera (for example, Frame-1 at 120-degree) and the second image from the secondary camera (frames chosen from SLAM), and device's location and orientation using IMU and visual features. In an embodiment of the disclosure, the SLAM based generator creates the map of the environment and estimates the position of the device 100 with respect to the map. The SLAM based generator uses the same information while generating the missing pixels during the 360-degree images generation (the frame-1 is generated at 180-degree).
[0203] FIG. 28 depicts an example scenario where the electronic device 100 is used to capture 360-degree content, according to embodiments as disclosed herein.
[0204] Consider that the user enables the camera with 360-degree option on their electronic device (e.g., smart phone or the like) 100. Using the received primary camera frame, the secondary camera frame, and one or more camera calibration parameters, the electronic device 100 can convert the style of the secondary camera image to suit the primary camera. The electronic device 100 further generates the required view to be generated for the primary camera frame into the centre aligned primary camera frame using the calibration information. The electronic device 100 creates the rectilinear image using both secondary and primary camera images. The electronic device 100 converts the rectilinear to Equirectangular using transformation. In an embodiment of the disclosure, the electronic device 100 generates the grid for the equirectangular image and use it as prior. In an embodiment of the disclosure, the electronic device 100 uses diffusion based generative model for inpainting to generate missing portion in the equirectangular space and reproject it to the sphere.
[0205] In other ways, once the user clicks on the 360-degree mode in the smartphone camera, three camera clients will be opened, one primary camera for ultra-wide angle field of view capture, one stereo camera to capture wide angle field of view, and one secondary camera. Ultra-Wide and Wide Camera is used to estimate depth using the stereo depth estimation, and further using the obtained depth map, to reconstruct the 3D volume. Moreover, from the ultra-wide images of the primary camera and secondary camera, the non-overlapping cameras are calibrated to estimate Extrinsics-Pose Matrix (R, T). In an embodiment of the disclosure, the electronic device 100 reprojects the image to the position provided from the calibration step with the 3D volume and colour, and texture information from the Ultra-Wide stream. Additionally, In an embodiment of the disclosure, the electronic device 100 converts the reprojected image and the secondary camera image to the equirectangular projection, captures the stored frames for inference using a trained generative AI (inpainting) model and converts the In-painted Equirectangular image to the 360-degree spherical image.
[0206] FIG. 29 and FIG. 30 depict example scenario for capturing 360-degree content in one shot in different scenarios, according to embodiments as disclosed herein.
[0207] 360-degree videos are immersive videos that allow viewers to look in all directions, as if they are standing at the centre of a sphere. This type of video provides a more engaging and interactive experience for viewers by allowing them to control the viewing angle, but difficult to capture due to specialized cameras. To create a 360-degree video, people use specialized cameras equipped with multiple lenses to capture a full 360-degree view of their surroundings. These cameras record footage simultaneously from all directions, which is then stitched together to create a seamless panoramic video. Viewers can also watch 360-degree videos using VR headsets for an even more immersive experience. When viewed through a VR headset, viewers can feel like they are actually present in the scene, as they can look around in all directions just as they would in real life. Overall, 360-degree videos provide a unique and engaging way for creators to tell stories and for viewers to experience content in a more immersive way.
[0208] In an embodiment of the disclosure, the 360-degree videos can be used for real-estate showcasing as depicted in FIG. 29, and FIG. 29 illustrates an example 360-degree image 2900. For example, a real estate agent, owner or visitor of a property can easily capture the 360-degree view without any specialized equipment. The content can be viewed by the customers and friends using the smartphone and / or the VST. The 360-degree one shot videos are image that is captured using the smartphone which can change the social media content generation and consumption because of the flexibility of not buying expensive equipment to capture such scenes as depicted in FIG. 30, and FIG. 30 illustrates another example 360-degree image 3000. In an embodiment of the disclosure, the users can get unique view every time when the user watches the video. This will increase user engagement towards social media content, when the smartphone fails to capture 360-degree content where the user has limited movement as depicted in FIG. 30.
[0209] The proposed method utilizes the current smartphones for creating 360-degree images using a proposed pipeline and calibrates secondary and back cameras of the smartphone to map the camera locations. In an embodiment of the disclosure, the method can be used to align the secondary and back cameras, using the depth based reprojection (DBR), so that the back camera image is moved to the line of axis of the secondary camera. Additionally, for generative AI for hole filling, approximately 100-degree of data needs to be synthesized using Gen AI and used to map 2D images to spherical mesh for 360-degree view.
[0210] Accordingly, the embodiments herein provide a method for capturing a 360-degree image of a scene by an electronic device. The method includes capturing, by the electronic device, a first image of the scene using a first imaging device, where the first imaging device is positioned on a rear of the electronic device. In an embodiment of the disclosure, the method includes capturing, by the electronic device, a second image of the scene using a second imaging device, where the second imaging device is positioned on the rear of the electronic device. In an embodiment of the disclosure, the method includes capturing, by the electronic device, a third image of the scene using a third imaging device, where the third imaging device is positioned on a front of the electronic device. In an embodiment of the disclosure, the method includes generating, by the electronic device, a three-dimensional (3D) representation of the scene using the first image and the second image. In an embodiment of the disclosure, the method includes generating, by the electronic device, a fourth image by re-projecting the first image using the generated 3D representation and a pre-defined calibration matrix to align an optical axis of the first imaging device with the third imaging device. In an embodiment of the disclosure, the method includes creating, by the electronic device, a two-dimensional (2D) rectilinear image representative of a 360-degree view of the scene using the third image and the fourth image.
[0211] In an embodiment of the disclosure, the method further includes generating, by the electronic device, an equi-rectangular projection of the scene by passing the 2D rectilinear image to a data driven model. The data driven model is pre-trained to generate at least one missing region using at least one pre-defined region-specific condition; The region specific conditions refer to how the missing region will be generated. While generating, the electronic device uses conditions like pixels closer to the other in the spherical space, pixels close to equator, pixels close to poles, etc. In an embodiment of the disclosure, the method includes generating, by the electronic device, a 360-degree content of the scene based on the generated equi-rectangular projection of the scene.
[0212] In an embodiment of the disclosure, the method includes displaying, by the electronic device, the generated 360-degree image on a display for at least one of: a real-time preview, and a post-capture editing.
[0213] In an embodiment of the disclosure, generating, by the electronic device, the equi-rectangular projection of the scene includes stitching the first image received from the first imaging device and the second image received from the second imaging device, based on the pre-defined calibration matrix, computing a conditional grid for at least one part of the stitched image, and using the conditional grid as an input to a diffusion based model to generate pixels, and generating the equi-rectangular projection of the scene based on the generated pixels. Values in the conditional grid indicate at least one of: border pixels, pixels near to equator, and pixels near poles
[0214] In an embodiment of the disclosure, at least one missing region in the 2D rectilinear image is generated by receiving a plurality of images from the first imaging device and the second image from the first imaging device, location of the electronic device, orientation using at least one sensor and a visual feature, creating a map of an environment, estimating a position of the electronic device with respect to the map, and generating the at least one missing region based on the estimated position of the device with respect to the map during the 360 images generation.
[0215] In an embodiment of the disclosure, at least one missing region in the 2D rectilinear image is generated by identifying at least one background object and missing region from a set of frames using map and a pose data, classifying landmarks and noise using redundant information based on the identification, extracting at least one feature from the at least one background object and the missing region using the landmark, the noise, the map and the pose data, and generating at least one missing region based on the at least one extracted feature.
[0216] In an embodiment of the disclosure, the at least one missing portion in the 2D rectilinear image is generated using a simultaneous localization and mapping (SLAM) information by receiving a plurality of images from the first imaging device and the second imaging device, and device's location and orientation using IMU and visual features.
[0217] In an embodiment of the disclosure, a relationship between the first imaging device, and the second imaging device is retrieved from the pre-defined calibration matrix, and the third image of the scene is time synchronized with the first image and the second image.
[0218] In an embodiment of the disclosure, the pre-defined calibration matrix is determined by positing the electronic device at a center of a calibration chamber such that the first imaging device and the second imaging device is aligned to a center of the calibration chamber, receiving a sets of images from the first imaging device and the second imaging device at the same time, and determining the pre-defined calibration matrix by reprojecting key points associated with the sets of images. The calibration chamber comprises identical patterns against both the first imaging device and the second imaging device placed in a setup.
[0219] In an embodiment of the disclosure, generating the fourth image by re-projecting the first image using the generated 3D representation and the pre-defined calibration matrix to align the optical axis of the first imaging device with the third imaging device camera includes estimating a depth map of the scene using the first image received from the first imaging device using a stereo depth, generating a mesh based on the depth map, reprojecting an red-green-blue (RGB) image from a first imaging device view to the centre by using calibration information and 3D points from the mesh, and generating the fourth image by reprojecting the RGB image from a first imaging device view to the centre by using calibration information and 3D points from the mesh.
[0220] In an embodiment of the disclosure, generating the equi-rectangular projection of the scene based on the 2D rectilinear image includes receiving the rectilinear image having a defined focal length and optical center coordinates, mapping each pixel in the rectilinear image from image coordinates to spherical coordinates, generating an equirectangular projection image by mapping the spherical coordinates to two-dimensional coordinates, and rendering the pixel values from the rectilinear image into the equirectangular projection image based on the computed coordinates. The focal length is dynamically calculated based on the field of view and resolution of the rectilinear image.
[0221] In an embodiment of the disclosure, each pixel in the rectilinear image from image coordinates are mapped to spherical coordinates based on a horizontal angular component and a vertical angular component.
[0222] In an embodiment of the disclosure, the equirectangular projection image has an aspect ratio of 2:1, corresponding to a 360° horizontal and 180° vertical spherical field of view.
[0223] In an embodiment of the disclosure, aligning the image from the first imaging device to a center view includes receiving an estimated transformation parameter associated with a plurality of features by matching the plurality of features in the actual second image frame and the expected second image frame of the secondary camera, calculating a time delta between actual second image frame and the expected second image frame, checking if the time delta is more than a determined threshold, and adjusting the electronic device parameter to reduce the camera synchronisation error.
[0224] Accordingly, the embodiments herein provide a method for capturing a 360-degree content of a scene by an electronic device. The method includes receiving, by the electronic device, a user input to activate a 360 degree content capture mode on the electronic device. In an embodiment of the disclosure, the method includes performing, by the electronic device, a stereo depth estimation using a plurality of image streams from a first imaging device, and a second imaging device to generate a depth map. In an embodiment of the disclosure, the method includes reconstructing, by the electronic device, a three-dimensional (3D) volume of a scene using the depth map. In an embodiment of the disclosure, the method includes reprojecting, by the electronic device, texture and color information from an ultra-wide image stream received from the first imaging device into a viewpoint determined by extrinsic parameters. In an embodiment of the disclosure, the method includes converting, by the electronic device, both the reprojected image and the image from the third imaging device into an equirectangular projection format.
[0225] In an embodiment of the disclosure, the method includes applying a data driven model trained for inpainting to fill in missing or occluded regions of the equirectangular projection. In an embodiment of the disclosure, the method includes generating a complete 360-degree spherical image using the inpainted equirectangular projection.
[0226] In an embodiment of the disclosure, reprojecting texture and color information from the ultra-wide image stream into a viewpoint determined by the extrinsic parameters includes calibrating the first imaging device and a third imaging device using captured images to estimate extrinsic parameters comprising a pose matrix representing rotation and translation, and reprojecting texture and color information from the ultra-wide image stream into a viewpoint determined by the extrinsic parameters.
[0227] Accordingly, the embodiments herein provide an electronic device including a 360-degree scene capturing controller coupled with a processor and a memory. The 360-degree scene capturing controller is configured to capture a first image of the scene using a first imaging device. In an embodiment of the disclosure, the 360-degree scene capturing controller is configured to capture a second image of the scene using the second imaging device. In an embodiment of the disclosure, the 360-degree scene capturing controller is configured to capture a third image of the scene using a third imaging device. In an embodiment of the disclosure, the 360-degree scene capturing controller is configured to generate a 3D representation of the scene using the first image and the second image. In an embodiment of the disclosure, the 360-degree scene capturing controller is configured to generate a fourth image by re-projecting the first image using the generated 3D representation and a pre-defined calibration matrix to align an optical axis of the first imaging device with the third imaging device. In an embodiment of the disclosure, the 360-degree scene capturing controller is configured to create a 2D rectilinear image representative of a 360-degree view of the scene using the third image and the fourth image.
[0228] Accordingly, the embodiments herein provide an electronic device including a 360-degree scene capturing controller coupled with a processor and a memory. In an embodiment of the disclosure, the 360-degree scene capturing controller is configured to receive a user input to activate a 360- degree content capture mode on the electronic device. In an embodiment of the disclosure, the 360-degree scene capturing controller is configured to perform a stereo depth estimation using a plurality of image streams from a first imaging device and a second imaging device to generate a depth map. In an embodiment of the disclosure, the 360-degree scene capturing controller is configured to reconstruct a three-dimensional (3D) volume of a scene using the depth map. In an embodiment of the disclosure, the 360-degree scene capturing controller is configured to reproject texture and color information from an ultra-wide image stream received from the first imaging device into a viewpoint determined by extrinsic parameters. In an embodiment of the disclosure, the 360-degree scene capturing controller is configured to convert both the reprojected image and the image from the third imaging device into an equirectangular projection format.
[0229] These and other aspects of the embodiments herein will be better appreciated and understood when considered in conjunction with the following description and the accompanying drawings. It should be understood, however, that the following descriptions, while indicating at least one embodiment and numerous specific details thereof, are given by way of illustration and not of limitation. Many changes and modifications may be made within the scope of the embodiments herein without departing from the scope thereof, and the embodiments herein include all such modifications.
[0230] Moreover, the propped method uses calibration information and depth based reprojection image to fill the primary and secondary image in the rectilinear space. Hence no overlapping field of view or feature matching is required.
[0231] Therefore, it is understood that the scope of the protection is extended to such a program and in addition to a computer readable means having a message therein, such computer readable storage means contain program code means for implementation of one or more steps of the method, when the program runs on a server or mobile device or any suitable programmable device. The method is implemented in at least one embodiment through or together with a software program written in e.g., Very high speed integrated circuit Hardware Description Language (VHDL) another programming language, or implemented by one or more VHDL or several software modules being executed on at least one hardware device. The hardware device can be any kind of portable device that can be programmed. The device may also include means which could be e.g., hardware means like e.g., an ASIC, or a combination of hardware and software means, e.g. an ASIC and an FPGA, or at least one microprocessor and at least one memory with software modules located therein. The method embodiments described herein could be implemented partly in hardware and partly in software. Alternatively, the invention may be implemented on different hardware devices, e.g., using a plurality of CPUs.
[0232] The foregoing description of the specific embodiments will so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify and / or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of embodiments, those skilled in the art will recognize that the embodiments herein can be practiced with modification within the scope of the embodiments as described herein.
Claims
1.A method for capturing a 360-degree image of a scene by an electronic device (100), the method comprising:capturing a first image of the scene using a first imaging device (170), wherein the first imaging device (170) is positioned on a first surface of the electronic device (100);capturing a second image of the scene using a second imaging device (180), wherein the second imaging device (180) is positioned on the first surface of the electronic device (100);capturing a third image of the scene using a third imaging device (190), wherein the third imaging device (190) is positioned on a second surface of the electronic device (100), wherein the second surface faces in a direction opposite to the first surface;generating a three-dimensional (3D) representation of the scene using the first image and the second image;generating a fourth image by re-projecting the first image to align an optical axis of the first imaging device (170) with the third imaging device (190) based on the generated 3D representation and a pre-defined calibration matrix; andcreating a two-dimensional (2D) rectilinear image representative of a 360-degree view of the scene using the third image and the fourth image.2.The method as claimed in claim 1, wherein the method comprises:generating an equi-rectangular projection of the scene by passing the 2D rectilinear image to a data driven model, wherein the data driven model is pre-trained to generate at least one missing region using at least one pre-defined region-specific condition; andgenerating a 360-degree content of the scene based on the generated equi-rectangular projection of the scene.3.The method as claimed in claim 2, wherein generating the equi-rectangular projection of the scene comprises:stitching the third image and the fourth image based on the pre-defined calibration matrix;computing a conditional grid for at least one part of the stitched image;using the conditional grid as an input to a diffusion based model to generate pixels, wherein values in the conditional grid indicate at least one of: border pixels, pixels near to equator, and pixels near poles; andgenerating the equi-rectangular projection of the scene based on the generated pixels.4.The method as claimed in claim 2, wherein at least one missing region in the 2D rectilinear image is generated by:receiving a plurality of images from at least one of the first imaging device (170) and the second imaging device (180), and the third imaging device (190);obtaining a location of the electronic device (100) and an orientation of the electronic device (100) and visual features;creating a map of an environment;estimating a position of the electronic device (100) with respect to the map; andgenerating the at least one missing region based on the estimated position of the device with respect to the mapduring the 360 images generation.5.The method as claimed in claim 2, wherein at least one missing region in the 2D rectilinear image is generated by:identifying at least one background object and missing region from a set of frames using map and a pose data;classifying landmarks and noise using redundant information based on the identification;extracting at least one feature from the at least one background object and the missing region using the landmark, the noise, the map and the pose data; andgenerate the at least one missing region based on the at least one extracted feature.6.The method as claimed in claim 2, wherein the at least one missing region in the 2D rectilinear image is generated using a Simultaneous localization and mapping (SLAM) information, the SLAM information comprising: a plurality of images from at least one of the first imaging device (170) and the second imaging device (180), and the third imaging device (190), and a location and an orientation of the electronic device, and visual features.7.The method as claimed in any one of claims 1 to 6, wherein a relationship between the first imaging device (170), and the second imaging device (180) is retrieved from the pre-defined calibration matrix, and the third image of the scene is time synchronized with the first image and the second image.8.The method as claimed in any one of claims 1 to 7, wherein the pre-defined calibration matrix is determined by:positing the electronic device (100) at a center of a calibration chamber such that the first imaging device (170) and the second imaging device (180) is aligned to a center of the calibration chamber, wherein the calibration chamber comprises identical patterns against both the first imaging device (170) and the second imaging device (180) placed in a setup;receiving a sets of images from the first imaging device (170) and the second imaging device (180) at the same time; anddetermining the pre-defined calibration matrix by reprojecting key points associated with the sets of images.9.The method as claimed in any one of claims 1 to 8, wherein generating the fourth image by re-projecting the first image using the generated 3D representation and the pre-defined calibration matrix to align the optical axis of the first imaging device (170) with the third imaging device (190) comprises:estimating a depth map of the scene using the first image received from the first imaging device (170) using a stereo depth;generating a mesh based on the depth map;reprojecting a red-green-blue (RGB) image from a first imaging device view to the centre by using calibration information and 3D points from the mesh; andgenerating the fourth image by reprojecting the RGB image from the first imaging device view to the centre by using calibration information and 3D points from the mesh.10.The method as claimed in any one of claims 1 to 9, wherein generating the equi-rectangular projection of the scene based on the 2D rectilinear image comprises:receiving the rectilinear image having a defined focal length and optical center coordinates, wherein the focal length is dynamically calculated based on the field of view and resolution of the rectilinear image;mapping each pixel in the rectilinear image from image coordinates to spherical coordinates;generating an equirectangular projection image by mapping the spherical coordinates to two-dimensional coordinates; andrendering the pixel values from the rectilinear image into the equirectangular projection image based on the computed coordinates.11.The method as claimed in claim 10, wherein each pixel in the rectilinear image from image coordinates are mapped to spherical coordinates based on a horizontal angular component and a vertical angular component.12.The method as claimed in any one of claim 10 to 11, wherein the equirectangular projection image has an aspect ratio of 2:1, corresponding to a 360° horizontal and 180° vertical spherical field of view.13.The method as claimed in any one of claim 1 to 12, wherein aligning the image from the first imaging device (170) to a center view comprises:receiving an estimated transformation parameter associated with a plurality of features by matching the plurality of features in the actual second image frame and the expected second image frame of the secondary camera;calculating a time delta between actual second image frame and the expected second image frame;checking if the time delta is more than a determined threshold; andadjusting the electronic device parameter to reduce the camera synchronisation error.14.An electronic device (100), comprising:a memory (130) storing a program or at least one instruction; andat least one processor (110) configure to individually or collectively execute the program or the at least one instruction;wherein the program or the at least one instruction, when executed individually or collectively by the at least one processor (110), cause the electronic device (100) to:capture a first image of the scene using a first imaging device (170);capture a second image of the scene using the second imaging device (180);capture a third image of the scene using a third imaging device (190);generate a three-dimensional (3D) representation of the scene using the first image and the second image;generate a fourth image by re-projecting the first image to align an optical axis of the first imaging device (170) with the third imaging device (190) based on the generated 3D representation and a pre-defined calibration matrix; andcreate a two-dimensional (2D) rectilinear image representative of a 360-degree view of the scene using the third image and the fourth image.15.A computer-readable recording medium having recorded thereon a program for causing a computer to execute the method of any one of claims 1 to 13.