System and method for creating 3D representation of images

WO2024210429A3PCT designated stage expired Publication Date: 2025-09-11SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/004161
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-04
Filing Date
2024-04-01
Publication Date
2025-09-11

AI Technical Summary

Technical Problem

Conventional methods for creating 3D representations using a single camera are inefficient due to low accuracy and high computational processing, and the use of multiple laterally separated cameras is ineffective for capturing small physical dimensions and determining accurate depth perception.

Method used

A method and system that capture two longitudinally separated images using a single camera, where objects are detected and paired based on similarity parameters, and their distances are estimated using the size ratio and movement distance to generate an accurate 3D representation of the field of view.

Benefits of technology

This approach provides high accuracy and low processing requirements, enabling effective 3D representation generation without the need for multiple cameras, suitable for small devices like mini-robots and autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024004161_12092025_PF_FP_ABST
    Figure KR2024004161_12092025_PF_FP_ABST
Patent Text Reader

Abstract

The present subject matter refers to a method and system for generating three-dimensional (3D) representation of a field of view. The method includes capturing an initial image of the field of view from an initial position and a final image of the field of view from a final position, detecting one or more objects in the initial image and the final image, pairing of each object in the initial image with a corresponding object in the final image, estimating a first distance of each paired object from the initial position based on: a ratio of a size of corresponding each paired object in the initial image to a size of corresponding each object in the final image, determining a physical size of each paired object based on the first distance of each paired object and generating a 3D representation.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR CREATING 3D REPRESENTATION OF IMAGES

[0001] The present disclosure relates to the field of image processing, and more particularly relates to a method and a system for creating a 3D representation of longitudinally separated images captured by a single camera.

[0002] 3D imaging technology has become increasingly popular in various industries, including but not limited to entertainment, gaming, medical imaging, etc. A three-dimensional (3D) image refers to an image that appears to have depth, height, and width, making it appear as if it exists in a physical space. Unlike traditional two-dimensional (2D) images, 3D images can be viewed from different angles and perspectives, allowing for a more immersive and realistic experience.

[0003] To enhance the user experience and for various purposes, creating an accurate 3D representation of a field of view during imaging is very important. Imaging devices such as mobile phones, digital cameras, and camera-configured systems such as mini robots, micro-robots, nanorobots, and autonomous vehicles use cameras to create the 3D representation of the field of view of the image to be captured. In order to generate the 3D representation of the field of view, accurate depth perception of objects during imaging is vital.

[0004] Certain conventional techniques employ a heuristic methodology that incorporates machine learning techniques to determine depth perception. In such conventional approaches, a vast training dataset is utilized as input to determine the depth perception of objects within the field of view. Nonetheless, such conventional solutions have low accuracy or precision in creating or generating an accurate 3D representation of the field of view. Additionally, the conventional approach involves significant computational processing due to the utilization of the extensive training dataset.

[0005] Additionally, certain conventional techniques determine depth perception through the utilization of multiple laterally separated. The utilization of the multiple laterally separated cameras leverages a principle of parallax to determine the depth perception. However, the deployment of the laterally separated cameras proves inefficient in determining the accurate depth perception as it necessitates multiple cameras to capture the objects within the field of view during the imaging. Furthermore, the conventional techniques are ineffective in capturing small physical dimensions despite the availability of multiple cameras.

[0006] Therefore, there lies a need for a solution to mitigate each of the above-discussed problems.

[0007] This summary is provided to introduce a selection of concepts in a simplified format that are further described in the detailed description of the invention. This summary is not intended to identify key or essential inventive concepts of the invention, nor is it intended for determining the scope of the invention.

[0008] In an implementation, the present subject matter refers to a method and system for generating three-dimensional (3D) representation of a field of view associated with an imaging device. The method includes capturing an initial image and a final image of the field of view from an initial position and a final position, respectively, and thereafter, detecting one or more objects in the captured initial image and the captured final image. The method further includes pairing each object in the initial image with a corresponding object in the final image based on a plurality of similarity parameters among the one or more objects. Thereafter, estimates a first distance of each paired object among the one or more objects, from the initial position based on: a ratio of a size of corresponding each paired object in the initial image to a size of corresponding each object in the final image, and a second distance between the initial position and the final position. The method further includes determining a physical size of each paired object among the one or more objects based on the estimated first distance of each paired object and generating a 3D representation of the field of view based on the estimated first distance and determined physical size of each paired object among the one or more objects.

[0009] To further clarify the advantages and features of the present invention, a more particular description of the invention will be rendered by reference to specific embodiments thereof, which is illustrated in the appended drawing. It is appreciated that these drawings depict only typical embodiments of the invention and are therefore not to be considered limiting its scope. The invention will be described and explained with additional specificity and detail with the accompanying drawings.

[0010] These and other features, aspects, and advantages of the present invention will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:

[0011] Figure 1illustrates a general architecture of an imaging device, in accordance with an embodiment of the present disclosure.

[0012] Figure 2illustrates a general operational flow for generating three-dimensional (3D) representation of a field of view associated with the imaging device, according to an embodiment of the present disclosure.

[0013] Figure 3illustrates a detailed operational flow for generating three-dimensional (3D) representation of a field of view associated with the imaging device, according to an embodiment of the present disclosure.

[0014] Figure 4illustrates an example scenario of capturing longitudinally separated images, according to an embodiment of the present disclosure.

[0015] Figure 5Aillustrates distortion removal techniques, according to an embodiment of the present disclosure.

[0016] Figure 5Billustrates distortion removal techniques, according to an embodiment of the present disclosure.

[0017] Figure 5Cillustrates distortion removal techniques, according to an embodiment of the present disclosure.

[0018] Figure 5Dillustrates distortion removal techniques, according to an embodiment of the present disclosure.

[0019] Figure 6illustrates an example spherical surface for calculating the resize factor 'R', according to an embodiment of the present disclosure.

[0020] Figure 7illustrates an example scenario for estimating the distance of each paired object, according to an embodiment of the present disclosure.

[0021] Figure 8Aillustrates a comparison of conventional art with the disclosed methodology.

[0022] Figure 8Billustrates a comparison of conventional art with the disclosed methodology.

[0023] Figure 9illustrates an example scenario for distance and size estimation using a mobile phone.

[0024] Figure 10illustrates an example scenario for creating accurate 3D representation using a mobile phone.

[0025] Figure 11illustrates an example scenario fordistance estimation in autonomous (Self Driving) Cars.

[0026] Figure 12illustrates an example scenario fordistance estimation in robots.

[0027] Further, skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have been necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent steps involved to help to improve understanding of aspects of the present invention. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present invention so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.

[0028] It should be understood at the outset that although illustrative implementations of the embodiments of the present disclosure are illustrated below, the present invention may be implemented using any number of techniques, whether currently known or in existence. The present disclosure should in no way be limited to the illustrative implementations, drawings, and techniques illustrated below, including the exemplary design and implementation illustrated and described herein, but may be modified within the scope of the appended claims along with their full scope of equivalents.

[0029] The term "some" as used herein is defined as "none, or one, or more than one, or all." Accordingly, the terms "none," "one," "more than one," "more than one, but not all" or "all" would all fall under the definition of "some." The term "some embodiments" may refer to no embodiments or to one embodiment or to several embodiments or to all embodiments. Accordingly, the term "some embodiments" is defined as meaning "no embodiment, or one embodiment, or more than one embodiment, or all embodiments."

[0030] The terminology and structure employed herein is for describing, teaching, and illuminating some embodiments and their specific features and elements and does not limit, restrict, or reduce the spirit and scope of the claims or their equivalents.

[0031] More specifically, any terms used herein such as but not limited to "includes," "comprises," "has," "consists," and grammatical variants thereof do NOT specify an exact limitation or restriction and certainly do NOT exclude the possible addition of one or more features or elements, unless otherwise stated, and furthermore must NOT be taken to exclude the possible removal of one or more of the listed features and elements, unless otherwise stated with the limiting language "MUST comprise" or "NEEDS TO include."

[0032] Whether or not a certain feature or element was limited to being used only once, either way, it may still be referred to as "one or more features" or "one or more elements" or "at least one feature" or "at least one element." Furthermore, the use of the terms "one or more" or "at least one" feature or element do NOT preclude there being none of that feature or element, unless otherwise specified by limiting language such as "there NEEDS to be one or more . . . " or "one or more element is REQUIRED."

[0033] Unless otherwise defined, all terms, and especially any technical and / or scientific terms, used herein may be taken to have the same meaning as commonly understood by one having ordinary skill in the art.

[0034] Embodiments of the present invention will be described below in detail with reference to the accompanying drawings.

[0035] The present disclosure relates to a method and system for generating three-dimensional (3D) representation of a field of view associated with an imaging device. According to an embodiment, a method for generating the 3D representation using two longitudinally separated images captured by a single camera in the imaging device is disclosed. In particular, two images are captured that are separated longitudinally rather than spatially using a single camera. Thereafter, the depth of objects from the two longitudinally separated images based on a change in the size of the objects due to a longitudinal movement of the imaging device and based on their distance from the imaging device, is calculated. Based on the calculated depth an accurate 3D representation of the field of view is generated. A detailed explanation of the same will be explained in the forthcoming paragraphs.

[0036] Figure 1illustrates a general architecture of an imaging device, in accordance with an embodiment of the present disclosure. According to an embodiment, the imaging device 100 includes at least one or more processors 101, a memory 103, a module / unit 105, a database 107, and at least one camera 109, coupled with each other.

[0037] As an example, an imaging device 100 corresponds to various imaging devices such as a personal computer (PC), a tablet PC, a personal digital assistant (PDA), a smartphone, a head-mounted device, a palmtop computer, a laptop computer, a desktop computer, digital cameras, robots, humanoids, autonomous vehicles, autonomous machines, or any other machine capable of executing a set of instructions.

[0038] As an example, the processor 101 may be a single processing unit or a number of units, all of which could include multiple computing units. The processor 101 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, logical processors, virtual processors, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. Among other capabilities, the processor 101 is configured to fetch and execute computer-readable instructions and data stored in the memory 103. In an alternate embodiment, the modules, units, and components as shown in the figure 1 include one or more processors, and the function of each of the modules / units 105 may be performed by the one or more processors 101 according to an alternate embodiment.

[0039] The memory 103 may include any non-transitory computer-readable medium known in the art including, for example, volatile memory, such as static random access memory (SRAM) and dynamic random access memory (DRAM), and / or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes.

[0040] In an example, the module(s), and / or unit(s) 105 may include a program, a subroutine, a portion of a program, a software component or a hardware component capable of performing a stated task or function. As used herein, the module(s), and / or unit(s) 105 may be implemented on a hardware component such as a server independently of other modules, or a module can exist with other modules on the same server, or within the same program. The module (s), and / or unit(s) 105 may be implemented on a hardware component such as processor one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. The module (s), and / or unit(s) 105 when executed by the processor(s) may be configured to perform any of the described functionalities.

[0041] As a further example, the database 107 may be implemented with integrated hardware and software. The hardware may include a hardware disk controller with programmable search capabilities or a software system running on general-purpose hardware. Examples of the database 107 are but are not limited to, in-memory databases, cloud databases, distributed databases, embedded databases, and the like. The database 107 amongst other things serves as a repository for storing data processed, received, and generated by one or more of the processors 101, and the modules / units 105.

[0042] The modules / units 105 may be implemented with an AI module that may include a plurality of neural network layers. Examples of neural networks include but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), and Restricted Boltzmann Machine (RBM). The learning technique is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning techniques include but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. At least one of a plurality of CNN, DNN, RNN, RMB models and the like may be implemented to thereby achieve execution of the present subject matter's mechanism through an AI model. A function associated with AI may be performed through the non-volatile memory, the volatile memory, and the processor. The processor may include one or a plurality of processors. At this time, one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU). The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.

[0043] Figure 2illustrates a general operational flow for generating three-dimensional (3D) representation of a field of view associated with the imaging device, according to an embodiment of the present disclosure.Figure 3illustrates a detailed operational flow for generating three-dimensional (3D) representation of a field of view associated with the imaging device, according to an embodiment of the present disclosure. Figure 2 will be explained by referring to figure 3. According to an embodiment, method 200 and method 300 in figures 2 and 3 respectively are performed by processor(s) 101 of figure 1.

[0044] According to an embodiment, consider a scenario where the imaging device 100 performs imaging to capture an image of a field of the view. During an operation, initially atoperation 201, the processor 101 is configured to capture an image from a camera. As an example, the one or more cameras 109 are implemented in the imaging device 100. However, according to an embodiment, for performing the method 200 and 300 a single camera is utilized. Further, the capturing operation may be explained by alternatively referring to the camera 109 or the imaging device 100, throughout the disclosure without deviating from the scope of the invention. According to an embodiment, atoperation 301,the processor 101 is configured to capture an initial image and a final image of the field of view from an initial position and a final position, respectively. The operation 301 is explained with the help of an exemplary scenario as shown in figure 4.

[0045] Figure 4illustrates an example scenario of capturing longitudinally separated images, according to an embodiment of the present disclosure. During an imaging operation of an object of interest 405 that is in the field of view of the imaging device 100, the camera 109 may be configured to capture an initial image 401 from an initial position P1. Let us consider that the distance between the camera 109 at position P1 and object of interest 405 is referred to as 'd'. Thereafter, the processor 101 of the imaging device 100 may be configured to perform a longitudinal movement of the imaging device 100 so as to capture a final image 403 from the next position P2. According to an embodiment, for performing the longitudinal movement of the imaging device 100, the processor 101 is configured to move the imaging device 100 longitudinally by a distance 'D' from the position P1 (initial position) to the position P2 (final position). As an example, the distance 'D' may be set in prior by the user in the imaging device 100. Thus, the distance between the camera at position P2 and the object of interest becomes 'd-D'. According to the embodiment, the value of 'D' is known by the longitudinal movement of the imaging device 100. Further, the value of 'd' is to be estimated. The estimation of the distance 'd' between the camera 109 at the position P1 and object of interest 405 shall be explained in detail in the forthcoming paragraphs.

[0046] According to an embodiment, the longitudinal movement is in a direction perpendicular towards or away from the field of view to a plane of the imaging device 100. Thus, the initial position P1 and the final position P2 are two longitudinally separated positions. Further, the final image 403 is captured only after capturing the initial image 401. Thus, according to an embodiment, the initial image 401 and the final image 403 are indicative of longitudinally separated images that are captured at the initial position P1 and the final position P2 which are longitudinally separated positions.

[0047] After the capturing of the initial and final image, the processor 101 of the imaging device 100 is configured to perform distortion removal at operation 203. In particular, during the capture of the initial image and the final image distortion may occur in the captured images ie., the initial image and the final image. Thus, by removing the distortions, the images are converted to undistorted images. Due to the projection of camera 109 longitudinally, different parts in the captured images are scaled differently. An example of a hyperbolically distorted image 501 is shown infigure 5A. According to an embodiment, the processor 101 of the imaging device 100 is configured to initially split the captured distorted images into small rectangles of thin segments as shown in figure 5B. According to an embodiment, each small rectangle 503(503a, 503b, 503c, 503d) is equal to the field of view. Each of the small rectangles 503 is of a different size due to stepwise hyperbolic distortion introduced by converting a flat image into an immersive image. Thereafter, each of the small rectangles 503 is resized by a resize factor 'R'. In order to calculate the resize factor 'R', the flat image is wrapped on an imaginary spherical surface around a viewing person, and the scaling factor of a slice projected on the viewing spherical surface will be the resize factor of the slice of the image as depicted inFigure 6. The resize factor 'R' is calculated based on equation 1 below.

[0048] --- (1)

[0049] Where, : Projecting sphere

[0050] : Projection plane

[0051] : Projecting element

[0052] : Radius of projecting sphere

[0053] : Angle of the projecting element from the central axis

[0054] : Angular width of the projecting element

[0055] : Length of the projecting element

[0056] : Length of the projected element

[0057] : Image width

[0058] : Distance of the projected element from the central axis

[0059] : Angular width of the field of view

[0060] : Resize factor

[0061] According to an embodiment, processor 101 is then configured to resize the small rectangles 503 based on the calculated resized factor 'R'. The output of the image resizer is shown in figure 5C as a resized image 505. For example, processor 101 may acquires a resized image 505a by resizing a small rectangle 503a. The resized image 505 is then stitched by merging each of the resized image 505a, 505b, 505c and 505d to get an undistorted image 507 as depicted in figure 5D. The undistorted image 507 which includes the initial and the final images is then processed to perform the object detection and pairing operation.

[0062] Referring back to figures 2 and 3, atoperation 205, the processor 101 is configured to perform object detection and pairing. In particular, after removing the distortion in the initial and final images, the processor 101 is configured to detect one or more objects in the captured initial image and the captured final image atoperation 205. In particular, atoperation 303, the processor 101 is configured to detect one or more objects in the captured initial image and the captured final image. According to an embodiment, the one or more objects in the captured initial image and the final image may be detected based on various known techniques, for example, color gradient image processing, machine learning models, a combination of the color gradient image processing and machine learning models and the like. According to some example techniques, the user may select a pixel area of objects of interest when the captured images are displayed on the camera preview screen.

[0063] According to an example embodiment, detecting the objects based on the color gradient image processing includes detecting the objects based on the steepness of the color gradient at the object boundaries. Thereafter, a histogram of gradients (HOG) of the objects is created. The HOGs are then converted into HOG description vectors. The HOG description vectors are then classified into object boundaries using any of the known classification algorithms, for example, support vector machine (SVM).

[0064] According to a further example embodiment, detecting the objects based on the color gradient image processing includes calculating a gradient vector for each pixel using a suitable kernel matrix for example, a Sobel operator, a Prewitt operator, etc. in each of the color channels. The image gradient vector is defined as a metric for every individual pixel, containing the pixel color changes in both the x-axis and y-axis. The image gradient vector is aligned with the gradient of a continuous multi-variable function, which is a vector of partial derivatives of all the variables. For example, f(x, y) records the color of the pixel at location (x, y), the gradient vector of the pixel (x, y) is defined in table 1.

[0065]

[0066] Accordingly, the magnitude of the vector (g) is given in equation 2, and the direction of the vector ( ) is given in equation 3.

[0067] ---- (2)

[0068] ---- (3)

[0069] After calculating the gradient vector, the color gradient image processing further includes identifying edges of pixels with high magnitudes and similar directions. Thus, the different objects are identified based on the quadrilaterals formed by the intersection of the edges.

[0070] Referring to the example scenario of figure 4, objects O1, O2, and O3 are detected in the initial image 401. Further, Objects O4, O5, and O6 are detected in the final image 403. According to an embodiment, after detecting the objects, the processor 101 is configured to extract the images of the detected objects in both the images i.e. from initial image 401 and the final image 403.

[0071] After detecting the objects, the processor 101, atoperation 305, is configured to pair each object in the initial image with a corresponding object in the final image based on a plurality of similarity parameters among the one or more objects. The plurality of similarity parameters includes a color, a size, and a position in the initial and final images of one or more objects. Thereafter, the processor 101 is configured to create a seven-dimensional similarity vector for each of the extracted object images in both images by calculating the dimensions (D') based on equation 4.

[0072] ---- (4)

[0073] Where,

[0074] : Average of red color in all pixels of the detected object

[0075] : Average of green color in all pixels of the detected object

[0076] : Average of blue color in all pixels of the detected object

[0077] : Horizontal pixel size of the detected object

[0078] : Vertical pixel size of the detected object

[0079] : Horizontal pixel position of the centroid of the detected object

[0080] : vertical pixel position of the centroid of the detected object

[0081] According to some embodiment, the similarity parameters corresponding to the position may be determined based on the shortest Euclidean distance D' between the vectors, for the objects to be paired. The technique of pairing may be explained by referring to the example scenario in figure 4. Consider that for pairing the object O1 with any of the objects in the final image 403, the object with similar color, size, and position in the final image 403 is considered. Accordingly, for the object O1 in the initial image, the object O4 is paired as the object O1 is determined to be similar to the object O4. Table 2 shows an example of pairing of the objects in the initial and final image for example scenario of figure 4. Further, operation 205 includes operations 303 and 305.

[0082] Object O1 paired with Object O4Object O2 is paired with Object O5Object O3 is paired with object O7

[0083] Table 2After the pairing of the objects in the initial and the final images, the method further proceeds to performoperation 307, the processor 101 is configured to estimate a distance (d)of each paired object among the one or more objects, from the initial position based on a ratio of a size of corresponding each paired object in the initial image to a size of corresponding each object in the final image and the distance (D) between the initial position and the final position. The distance of each paired object in the initial image 401, helps in determining the size of the objects that are to be represented in the 3D representation. The method for estimating will be explained with the help offigure 7. According to an embodiment, for estimating the distance (d), the processor 101 is configured to determine the size of each paired object in the initial image 401 based on the number of pixels of the each paired object in the initial image. Similarly, the processor 101 is configured to determine the size of each paired object in the final image 403 based on the number of pixels of each paired object in the final image 403. Referring to the figure 7, in the initial image 401, the size of each object i.e., objects O1, O2, and O3 in the initial image 401 is determined based on the number of pixels in objects O1, O2, and O3 respectively. Similarly, the size of each object i.e., objects O4, O5, and O6 in the final image 403 is determined based on the number of pixels in objects O4, O5, and O6 respectively. After determining the sizes of each paired object in the initial image 401 and the final image 403, the processor 101 is further configured to calculate the ratio (r) of the determined size of each paired object in the initial image to the determined size of the corresponding object in the final image.

[0084] According to an embodiment, the ratio (r) is given by the equation (5),

[0085] --- (5)

[0086] Where, : Size of the object of interest in the initial photo in the number of pixels,

[0087] : Size of the object of interest in the final photo in the number of pixels.

[0088] According to an embodiment, the the distance 'd' of an object from the initial position, in a case when the camera 109 is moved towards the field of view is given by equation 6, and the distance 'd' of an object from the initial position, in a case when the camera 109 is moved away from the field of view is given by equation 7.

[0089] --- (6)

[0090] --- (7)

[0091] Where D is the distance of the camera position between the initial and the final photo.

[0092] According to some embodiments, the distance D as explained above may be determined based on the longitudinal movement of the camera. Thus, the distance (D) between the initial position P1 and the final position P2 may be determined by using motion sensors like an accelerometer, gyroscope, and GPS sensor. Referring to the figure 7, a distance d1, distance d2, and distance d3 for the objects O1, O2, and O3 with respect to their corresponding paired objects O4, O5, and O6 respectively is being determined. Tables 3 and 4 illustrate detailed mathematical steps in a case when the camera is moving toward the field of view and a case when the camera is moving away from the field of view.

[0093] Step 1: Step 2: Step 3: Step 4: Step 5: Step 6:

[0094] Step 1: Step 2: Step 3: Step 4: Step 5: Step 6:

[0095] The steps as shown in the table 3 and 4 calculate the distance of the paired object from the camera by capturing two images of the object separated by a distance D, and then equating the ratio of change in the square root of an angular area of the paired object to a ratio of the distance of the paired object in the two images according to the step 2. Further, it can be seen that, at step 2, as the distance D increases the ratio of the size of each paired object in the initial image to the determined size of the corresponding object in the final image increases. Further, as the distance D decreases the ratio of the size of each paired object in the initial image to the determined size of the corresponding object in the final image decreases. Thus, the calculated ratio is indicative of a change in the size of each of the corresponding paired objects in the initial image in proportion to the size of each of the corresponding paired objects in the final image. Further, the estimated distance (d), of each paired object in the initial image, is indicative of the depth value of each of the paired objects in the initial image. This ensures an optimal field of view of the final 3D representation while reducing the resolution slightly.

[0096] After estimating the distance,at operation 309, the processor 101 is further configured to determine the physical size of each paired object among the one or more objects based on the estimated first distance of each paired object. In particular, for determining the physical size of each paired object, at first, a horizontal field of view Ahand vertical field of view Avof the camera 109 is determined based on equations 8 and 9 respectively.

[0097] --- (8)

[0098] - (9)

[0099] Thereafter, at the next step horizontal angular size and vertical angular size of the object is determined based on equations 10 and 11 respectively.

[0100] -- (10)

[0101] - (11)

[0102] Thereafter, at the next step horizontal physical size and vertical physical size of the object is determined based on equations 12 and 13 respectively.

[0103] --- (12)

[0104] --- (13)

[0105] Where, :Horizontal size of the camera sensor

[0106] : Vertical size of the camera sensor,

[0107] :Focal length of camera lens,

[0108] : Horizontal field of view of camera,

[0109] : Vertical field of view of camera,

[0110] : Horizontal size of the image in pixels,

[0111] : Vertical size of the image in pixels,

[0112] : Horizontal size of the image in pixels,

[0113] : Vertical pixel size of the image in pixels,

[0114] : Horizontal angular size of the object,

[0115] : Vertical angular size of the object,

[0116] : Horizontal physical size of the object,

[0117] : Vertical physical size of the object,

[0118] : Distance of the object.

[0119] Referring to figure 7, for the objects O1, O2, and O3 the corresponding physical size (horizontal and vertical size) of the aforesaid objects is determined as . Based on the estimated distance (d), the physical size of each of the paired objects, the processor 101 atoperations 311generates a 3D representation of the field of view. In particular, the processor 101 is configured to place the images of the objects with sizes at a distance from the initial viewer position (P1) in a 3D world. In a 3D world, the viewer can move within the 3D world. The resulting 3D map is an accurate 3D representation of the field of view. According to an embodiment, the output may be a 3D metafile containing images and actual real-world size and distance of the objects in the field of view as shown in the table 5.

[0120] Detected ObjectDetermined physical SizeDistanceObject 1 Object 2 Object 3

[0121] Figure 8A and Figure 8Billustrates a comparison of conventional art with the disclosed methodology. As can be seen at 801 due to the conventional lateral depth estimation the 3D representation of the image is not generated accurately. On the other hand at 803, due to the implementation of the longitudinal depth estimation, the 3D representation of the image is being generated accurately.

[0122] Figure 9illustrates an example scenario for distance and size estimation using a mobile phone. According to an example scenario, the user first clicks a first photo, then the user moves a distance D and clicks another photo. According to an example embodiment, the area of the object of interest will be selected by the user. In another embodiment, the object of interest will be selected by the user from a list of objects identified by: Color contrast image processing or Machine Learning. Thus, based on the user-selected object of interest, APP is configured to perform the disclosed methods 200 and 300 and provide the distance and size of the object. The distance and size may be used to generate the 3D representation of the selected object.

[0123] Figure 10illustrates an example scenario for creating accurate 3D representation using a mobile phone. According to an example scenario, the user first clicks a first photo, then the user moves a distance D and clicks another photo. According to an example embodiment, APP is configured to perform the disclosed methods 200 and 300, detect objects in the image and create a 3D representation of the image.

[0124] Figure 11illustrates an example scenario fordistance estimation in autonomous (Self Driving) Cars. According to an example scenario, the car first clicks a first photo, then the car moves a distance D, and then clicks another second photo. According to an example embodiment where in the car system, the method 200 and 300 is implemented. The car detects objects in the scene and thereafter, the car creates a distance model for the objects of interest.

[0125] Figure 12illustrates an example scenario fordistance estimation in robots. According to an example scenario, the robot first clicks a first photo, then the robot moves a distance D, and then clicks another second photo. According to an example embodiment where in the robots, the method 200 and 300 is implemented. The robots detect objects in the image. Then the robots create a distance model for the objects of interest. Traditional multiple-camera stereo vision cannot be used in small robots because of the lack of lateral separation.

[0126] Accordingly, the disclosed method provides high accuracy, and low processing due to avoid to huge training data sets. Further, the disclosed method may be implemented in small devices like mini-robots, microrobots, drones, and the like. Further, as the method utilizes a single camera, the power consumption and cost are low.

[0127] Some example embodiments disclosed herein may be implemented using processing circuitry. For example, some example embodiments disclosed herein may be implemented using at least one software program running on at least one hardware device and performing network management functions to control the elements.

[0128] While specific language has been used to describe the disclosure, any limitations arising on account of the same are not intended. As would be apparent to a person in the art, various working modifications may be made to the method in order to implement the inventive concept as taught herein.

[0129] The drawings and the forgoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein.

[0130] Moreover, the actions of any flow diagram need not be implemented in the order shown; nor do all of the acts necessarily need to be performed. Also, those acts that are not dependent on other acts may be performed in parallel with the other acts. The scope of embodiments is by no means limited by these specific examples. Numerous variations, whether explicitly given in the specification or not, such as differences in structure, dimension, and use of material, are possible. The scope of embodiments is at least as broad as given by the following claims.

[0131] Benefits, other advantages, and solutions to problems have been described above with regard to specific embodiments. However, the benefits, advantages, solutions to problems, and any component(s) that may cause any benefit, advantage, or solution to occur or become more pronounced are not to be construed as a critical, required, or essential feature or component of any or all the claims.

Claims

1.A method for generating three-dimensional (3D) representation of a field of view associated with an imaging device, the method comprising:capturing an initial image of the field of view from an initial position and a final image of the field of view from a final position, respectively;detecting one or more objects in the captured initial image and the captured final image;pairing, each object in the initial image with a corresponding object in the final image based on a plurality of similarity parameters among the one or more objects;estimating a first distance of each paired object among the one or more objects, from the initial position based on:a ratio of a size of corresponding each paired object in the initial image to a size of corresponding each object in the final image, anda second distance (D) between the initial position and the final position;determining a physical size of each paired object among the one or more objects based on the estimated first distance of each paired object; andgenerating a 3D representation of the field of view based on the estimated first distance and determined physical size of each paired object among the one or more objects.2.The method as claimed in claim 1, further comprising:determining the size of each paired object in the initial image based on a number of pixels of the each paired object in the initial image;determining the size of each paired object in the final image based on a number of pixels of the each paired object in the final image; andcalculating the ratio of the determined size of each paired object in the initial image to the determined size of corresponding object in the final image.3.The method as claimed in claim 2, wherein the calculated ratio is indicative of a change in the size of each of the corresponding paired object in the initial image in proportion to the size of each of the corresponding paired object in the final image.4.The method as claimed in claim 1, wherein the estimated first distance, of each paired object in the initial image, is indicative of a depth value of each of the paired object in the initial image.5.The method as claimed in claim 1, wherein the plurality of similarity parameters comprises a color, a size, and a position in the initial and final images of one or more objects.6.The method as claimed in claim 1, wherein the capturing the final image of the field of view comprises:performing a longitudinal movement of the imaging device by the second distance (D) from the initial position to the final position,wherein the initial image and the final image are indicative of longitudinally separated images that are captured at the initial position and the final position which are longitudinally separated positions, andwherein the final image is captured based on the longitudinal movement of the imaging device and after capturing of the initial image, andwherein the longitudinal movement is in the direction perpendicular towards or away from the field of view to a plane of the imaging device.7.An imaging device for generating 3D representation of field of view, the imaging device comprising:one or more processors configured to:capture an initial image of the field of view from an initial position and a final image of the field of view from a final position, respectively;detect one or more objects in the captured initial image and the captured final image;pair, each object in the initial image with a corresponding object in the final image based on a plurality of similarity parameters among the one or more objects;estimate a first distance, of each paired object among the one or more objects, from the initial position based on:a ratio of a size of corresponding each paired object in the initial image to a size of corresponding each object in the final image, anda second distance between the initial position and the final position;determine a physical size of each paired object among the one or more objects based on the estimated first distance of each paired object; andgenerate a 3D representation of the field of view based on the estimated first distance and determined physical size of each paired object among the one or more objects.8.The imaging device as claimed in claim 7, wherein the one or more processors is configured to:determine the size of each paired object in the initial image based on a number of pixels of the each paired object in the initial image;determine the size of each paired object in the final image based on a number of pixels of the each paired object in the final image; andcalculate the ratio of the determined size of each paired object in the initial image to the determined size of corresponding object in the final image.9.The imaging device as claimed in claim 8, wherein the calculated ratio is indicative of a change in the size of each of the corresponding paired object in the initial image in proportion to the size of each of the corresponding paired object in the final image.10.The imaging device as claimed in claim 7, wherein the estimated first distance, of each paired object in the initial image, is indicative of a depth value of each of the paired object in the initial image.11.The imaging device as claimed in claim 7, wherein the plurality of similarity parameters comprises a color, a size, and a position in the initial and final images of one or more objects.12.The imaging device as claimed in claim 7, wherein for capturing the final image of the field of view, the one or more processors is configured to:perform a longitudinal movement of the imaging device by the second distance from the initial position to reach the final position,wherein the initial image and the final image are indicative of longitudinally separated images that are captured at the initial position and the final position which are longitudinally separated positions, andwherein the second image is captured based on the longitudinal movement of the imaging device and after capturing of the initial image, andwherein the longitudinal movement is in the direction perpendicular towards or away from the field of view to a plane of the imaging device.13.A computer-readable storage medium, having a computer program stored thereon that performs, when executed by a processor, the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Systems, devices, and methods for generating posture estimate of object

    JP2020173795A

  • Information processing apparatus, information processing method, and program

    JP2022160233A

  • Information processing apparatus, control method, and program

    US20200349719A1

  • Information processing device and information processing method

    WO2021205772A1

  • KR20220015056A