System and method for detecting 3D association of objects
Through image sensors and inertial measurement sensors to detect objects and combine head attitude data, the problem of determining the proportion and depth of objects in 3D space in augmented reality applications is solved, and the accurate and continuous update effects of augmented reality applications are achieved.
Patent Information
- Application Number
- CN201980047687.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-07-17
- Filing Date
- 2019-07-17
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2039-07-17
AI Technical Summary
In augmented reality applications, it is difficult to accurately determine the proportion and depth of objects in 3D space, affecting the accuracy and continuity of object detection.
By using an image sensor and an inertial measurement sensor, the object is detected and bounded areas in the image are defined, and the position of the object in 3D space is determined in combination with head attitude data.
It enables the Augmented Reality Markup Application to run effectively with fewer system resources, providing accurate and continuous updates of objects in the environment.
Smart Images

Figure CN112424832B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to augmented reality systems and methods. More specifically, the present disclosure relates to systems and methods for detecting 3D association of objects. Background Art
[0002] Object detection in 3D space is an important aspect of augmented reality applications. However, augmented reality applications pose challenges regarding determining the scale and depth of detected objects in 3D space. Summary of the invention
[0003] Solution to the problem
[0004] The present disclosure provides systems and methods for detecting 3D association of objects.
[0005] In one embodiment, an electronic device for detecting 3D association of an object is provided. The electronic device includes at least one image sensor, an inertial measurement sensor, a memory, and at least one processor coupled to the at least one image sensor, the inertial measurement sensor, and the memory. The at least one processor is configured to capture an image of an environment using the at least one image sensor; detect an object in the captured image; define a bounded area around the detected object in the image; receive head pose data from the inertial measurement sensor or from another processor configured to calculate head pose using the inertial measurement sensor and the image sensor data; and determine the position of the detected object in 3D space using the head pose data and the bounded area in the captured image.
[0006] In another embodiment, a method for detecting a 3D association of an object is provided. The method includes: capturing an image of an environment using at least one image sensor; detecting an object in the captured image; defining a bounded area in the image around the detected object; receiving head pose data from an inertial measurement sensor or another processor configured to calculate head pose using inertial measurement sensor and image sensor data; and determining a position of the detected object in 3D space using the head pose data and the bounded area in the captured image.
[0007] In yet another embodiment, a non-transitory medium embodying a computer program for operating an electronic device is provided, wherein the electronic device is used to detect a 3D association of an object. The electronic device includes at least one image sensor, an inertial measurement sensor, a memory, and at least one processor. When the program code is executed by the at least one processor, the program code causes the electronic device to capture an image of an environment using the at least one image sensor; detect an object in the captured image; define a bounded area in the image around the detected object; receive head pose data from the inertial measurement sensor or from another processor configured to calculate head pose using the inertial measurement sensor and image sensor data; and determine a position of the detected object in 3D space using the head pose data and the bounded area in the captured image.
[0008] Other technical features may be apparent to those skilled in the art from the following drawings, description and claims.
[0009] Advantageous Effects of the Invention
[0010] The present disclosure provides a triangulation process that allows augmented reality tagging applications to continue to run efficiently even when fewer system resources are available. This is important for augmented reality applications where the user utilizes a camera to provide a continuous view of a scene or environment, as the user expects the application to display accurate and continuously updated information about objects in the environment on the user's device. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] For a more complete understanding of the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, wherein like reference numerals represent like components:
[0012] Figure 1 shows an example network configuration according to an embodiment of the present disclosure;
[0013] Figure 2 shows an example captured image for use with an augmented reality marking application according to an embodiment of the present disclosure;
[0014] Figure 3 An example of an augmented reality object detection and 3D association architecture according to an embodiment of the present disclosure is shown;
[0015] Figure 4 A flow chart illustrating an example object detection and 3D association process according to an embodiment of the present disclosure is shown;
[0016] Figure 5 A diagrammatic view of a 3D feature point reprojection process for detecting 3D association of objects according to an embodiment of the present disclosure is shown;
[0017] Figure 6A diagram showing another illustrative example of a 3D feature point reprojection process according to an embodiment of the present disclosure;
[0018] Figure 7 A flowchart of an example 3D feature point reprojection process according to an embodiment of the present disclosure is shown;
[0019] Figure 8 A diagram illustrating an example feature point triangulation process according to an embodiment of the present disclosure; and
[0020] Fig. 9 A flow chart of an example detected object triangulation process according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0021] In one embodiment, an electronic device for detecting 3D association of an object is provided. The electronic device includes at least one image sensor, an inertial measurement sensor, a memory, and at least one processor coupled to the at least one image sensor, the inertial measurement sensor, and the memory. The at least one processor is configured to capture an image of an environment using the at least one image sensor; detect an object in the captured image; define a bounded area around the detected object in the image; receive head pose data from the inertial measurement sensor or from another processor configured to calculate head pose using the inertial measurement sensor and the image sensor data; and determine the position of the detected object in 3D space using the head pose data and the bounded area in the captured image.
[0022] In an embodiment, in order to determine the position of the detected object, at least one processor is also configured to generate a 3D point cloud of at least a portion of the environment using at least one image sensor and an inertial measurement sensor; project feature points of the 3D point cloud into the captured image, wherein the captured image includes a bounded area; determine which of the feature points are located on the surface of the object, and remove feature points that are not located on the surface of the object.
[0023] In an embodiment, at least one processor is further configured to determine that a sparse graph is used to generate a 3D point cloud, generate a dense Figure 3 D point cloud, and the dense Figure 3 The feature points of the D point cloud are projected into a bounded area of the captured image.
[0024] In an embodiment, in order to determine the position of the detected object, at least one processor is also configured to capture multiple images using at least one image sensor; detect the object in each of the multiple images, define bounded areas in the multiple images, wherein each bounded area in the bounded areas is defined as surrounding the detected object in one of the multiple images, and triangulate the position of the object in one of the multiple images using the position of the object in at least another image of the multiple images.
[0025] In an embodiment, in order to triangulate the position of the object, at least one processor is also configured to generate 2D feature points in each image of a plurality of images, determine the 2D feature points within each bounded area of the bounded areas in the plurality of images, and triangulate the 2D feature points within a bounded area of a bounded area of one image of the plurality of images using the 2D feature points within another bounded area of another image of the plurality of images.
[0026] In an embodiment, to triangulate the feature points, the at least one processor is further configured to perform pixel matching between each bounded region of at least two of the plurality of images, and to perform triangulation between matching pixels of at least two of the plurality of images.
[0027] In an embodiment, the at least one processor is further configured to place a virtual object adjacent to the detected object in the 3D space, wherein the detected object and the virtual object are related to an object class.
[0028] In another embodiment, a method for detecting a 3D association of an object is provided. The method includes: capturing an image of an environment using at least one image sensor; detecting an object in the captured image; defining a bounded region in the image surrounding the detected object; receiving head pose data from an inertial measurement sensor or from another processor configured to calculate head pose using inertial measurement sensor and image sensor data; and determining a position of the detected object in 3D space using the head pose data and the bounded region in the captured image.
[0029] In yet another embodiment, a non-transitory medium embodying a computer program for operating an electronic device is provided, wherein the electronic device is used to detect a 3D association of an object. The electronic device includes at least one image sensor, an inertial measurement sensor, a memory, and at least one processor. When the program code is executed by the at least one processor, the program code causes the electronic device to capture an image of an environment using the at least one image sensor, detect an object in the captured image, define a bounded area in the image around the detected object, receive head pose data from the inertial measurement sensor or from another processor configured to calculate head pose using the inertial measurement sensor and image sensor data, and determine a position of the detected object in 3D space using the head pose data and the bounded area in the captured image.
[0030] Invention Mode
[0031] The following discussion is described with reference to the accompanying drawings. Figures 1 to 9 As well as various embodiments of the present disclosure. However, it should be understood that the present disclosure is not limited to the embodiments, and all changes and / or equivalents or substitutions thereof also belong to the scope of the present disclosure. Throughout the specification and the drawings, the same or similar reference numerals may be used to refer to the same or similar elements.
[0032] Before proceeding to the following detailed description, it may be advantageous to set forth definitions of certain words and phrases used throughout this patent document. The terms "transmit," "receive," and "communicate," and their derivatives, include direct and indirect communications. The terms "include" and "comprises," and their derivatives, are meant to include without limitation. The term "or" is inclusive, meaning and / or. The phrase "related to," and its derivatives, means to include, be included within, be interconnected with, include, be contained within, be connected to or connected with, be connected to or connected with, be communicable with, collaborate with, be intertwined, be in parallel, be close to, be bound to or bound with, have, have the characteristics of, have a relationship with, and the like.
[0033] In addition, the various functions described below may be implemented or supported by one or more computer programs, each of which is formed by a computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, processes, functions, objects, classes, instances, related data, or parts thereof suitable for implementation in appropriate computer-readable program codes. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as a read-only memory (ROM), a random access memory (RAM), a hard disk drive, a compact disk (CD), a digital video disc (DVD), or any other type of memory. "Non-transitory" computer-readable media excludes wired, wireless, optical or other communication links that transport temporary electrical signals or other signals. Non-transitory computer-readable media include media that can permanently store data and media that can store data and then rewrite data, such as a rewritable optical disk or an erasable storage device.
[0034] As used herein, the terms “having”, “may have”, “include” or “may include” a feature (eg, a number, a function, an operation, or a component such as a part) indicate the presence of the feature and do not exclude the presence of other features.
[0035] As used herein, the terms "A or B", "at least one of A and / or B", or "one or more of A and / or B" may include all possible combinations of A and B. For example, "A or B", "at least one of A and B", or "at least one of A or B" may indicate (1) including at least one A, (2) including at least one B, or (3) including all of at least one A and at least one B.
[0036] As used herein, the terms "first" and "second" may modify various components, regardless of importance, and do not limit these components. These terms are only used to distinguish one component from another component. For example, regardless of the order or importance of the devices, a first user device and a second user device may indicate user devices that are different from each other. For example, a first component may be represented as a second component, and vice versa without departing from the scope of the present disclosure.
[0037] It should be understood that when an element (e.g., a first element) is referred to as being (operably or communicatively) "coupled to or connected to another element (e.g., a second element)" or "connected to or connected to another element (e.g., a second element)", it may be coupled to or connected to or connected to another element directly or via a third element. Conversely, it should be understood that when an element (e.g., a first element) is referred to as being "directly coupled to / connected to another element (e.g., a second element)" or "directly connected to / connected to (e.g., a second element)", no other element (e.g., a third element) is interposed between the element and the other element.
[0038] As used herein, the term "configured (or configured) to" may be used interchangeably with the terms "suitable for", "capable of", "designed to", "adapted to", "enable to do", or "capable of" depending on the situation. The term "configured (or configured) to" does not basically mean "specially designed hardware to". Instead, the term "configured to" may mean that a device can implement an operation together with another device or component.
[0039] For example, the term "a processor configured (or set) to implement A, B, and C" may mean a general-purpose processor (e.g., a CPU or an application processor) that can perform operations by implementing one or more software programs stored in a memory device, or a special-purpose processor (e.g., an embedded processor) for implementing the operations.
[0040] The terms used herein are provided only to describe some embodiments thereof, rather than to limit the scope of other embodiments of the present disclosure. It should be understood that the singular forms "a", "a kind of" and "the" include plural references unless the context clearly specifies otherwise. All terms used herein including technical and scientific terms have the same meanings as those generally understood by those of ordinary skill in the art to which the embodiments of the present disclosure belong. It should also be understood that those terms such as those defined in common dictionaries should be interpreted as having meanings consistent with their meanings in the context of the relevant technology, and will not be interpreted as idealized or overly formal meanings, unless explicitly defined as such herein. In some cases, the terms defined herein may be interpreted as excluding embodiments of the present disclosure.
[0041] For example, examples of electronic devices according to embodiments of the present disclosure may include at least one of a smart phone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a net-book computer, a workstation, a PDA (personal digital assistant), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (e.g., smart glasses, a head-mounted device (HMD), electronic clothing, an electronic bracelet, an electronic necklace, an electronic accessory, an electronic tattoo, a smart mirror, or a smart watch).
[0042] Definitions for certain other words and phrases are provided throughout this patent document. Those of skill in the art should understand that in many, if not most, instances, such definitions apply to prior as well as future uses of such defined words and phrases.
[0043] Any description in this application that is to be construed as implying that any particular element, step, or function is not an essential element that must be included within the scope of the claims. The scope of patentable subject matter is limited solely by the claims. In addition, none of the claims is intended to invoke 35 U.S.C. 112(f) unless the exact words "means for..." are followed by a participle. Any other terms used in the claims, including but not limited to "mechanism", "module", "device", "unit", "component", "element", "member", "device", "machine", "system", "processor", or "controller", are understood by applicant to refer to structures known to persons skilled in the relevant art and are not intended to invoke 35 U.S.C. 112(f).
[0044] Although the present disclosure has been described with exemplary embodiments, various changes and modifications may be suggested to one skilled in the art. The present disclosure is intended to include such changes and modifications as fall within the scope of the appended claims.
[0045] Object detection in 3D space is an important aspect of augmented reality applications. Augmented reality applications present challenges in determining the scale and depth of detected objects in 3D space. For example, the scale of the detected object is used to determine the properties of the detected object. For example, if a car is detected in an image, the scale of the car can determine whether the car is a toy or a real car. Although 3D object detection (such as by using convolutional neural networks and geometric structures) is possible, such methods may be computationally expensive. Therefore, object detection for augmented reality can be implemented in the 2D domain, which provides less computation, but the object still needs to be associated with the 3D space. Although 3D space understanding can be determined using depth or stereo cameras, 3D space understanding can also be achieved using simultaneous localization and mapping (SLAM).
[0046] Figure 1 An example network configuration 100 is shown in accordance with various embodiments of the present disclosure. Figure 1 The embodiment of the network configuration 100 shown is for illustration only. Other embodiments of the network configuration 100 may be used without departing from the scope of the present disclosure.
[0047] According to an embodiment of the present disclosure, an electronic device 101 is included in a network configuration 100. The electronic device 101 may include at least one of a bus 110, a processor 120, a memory 130, an input / output (IO) interface 150, a display 160, a communication interface 170, or a sensor 180. In some embodiments, the electronic device 101 may exclude at least one component or may add another component.
[0048] The bus 110 includes circuits for connecting components 120 - 170 to one another and transferring communications (eg, control messages and / or data) between the components.
[0049] The processor 120 includes one or more of a central processing unit (CPU), an application processor (AP), or a communication processor (CP). The processor 120 can control at least one of the other components of the electronic device 101, and / or perform operations or data processing involving communication. In some embodiments, the processor may be a graphics processor unit (GPU).
[0050] For example, the processor 120 may receive a plurality of frames captured by a camera during a capture event. The processor 120 may detect an object located within one or more of the plurality of frames captured. The processor 120 may define a bounded area around the detected object. The processor 120 may also receive head posture information from an inertial measurement unit. The processor 120 may use the head posture information and the bounded area in the captured frames to determine the position of the detected object in the 3D space. The processor 120 may place a virtual object in the 3D space near the determined position of the detected object. The processor 120 may operate a display to display a 3D space including the detected object and the virtual object.
[0051] The memory 130 may include a volatile memory and / or a non-volatile memory. For example, the memory 130 may store commands or data related to at least one other component of the electronic device 101. According to an embodiment of the present disclosure, the memory 130 may store software and / or programs 140. The programs 140 include, for example: a kernel 141, middleware 143, an application programming interface (API) 145, and / or an application program (or "application") 147. At least a portion of the kernel 141, the middleware 143, or the API 145 may be represented as an operating system (OS).
[0052] For example, the kernel 141 can control or manage system resources (e.g., bus 110, processor 120, or memory 130) used to implement operations or functions implemented in other programs (e.g., middleware 143, API 145, or application 147). The kernel 141 provides an interface that allows the middleware 143, API 145, or application 147 to access various components of the electronic device 101 to control or manage system resources. The application 147 includes one or more applications for image capture, object detection, head or camera posture, simultaneous localization and mapping (SLAM) processing, and object classification and labeling. These functions can be implemented by a single application or multiple applications, wherein each application in the multiple applications implements one or more of these functions.
[0053] For example, the middleware 143 may function as a relay that allows the API 145 or the application 147 to communicate data with the kernel 141. A plurality of applications 147 may be provided. The middleware 143 is capable of controlling a work request received from the application 147, for example, by assigning a priority of using a system resource (e.g., the bus 110, the processor 120, or the memory 130) of the electronic device 101 to at least one of the plurality of applications 147.
[0054] The API 145 is an interface that allows the application 147 to control functions provided from the kernel 141 or the middleware 143. For example, the API 145 includes at least one interface or function (eg, command) for file control, window control, image processing, or text control.
[0055] The IO interface 150 serves as an interface that can, for example, transmit commands or data input from a user or other external devices to other components of the electronic device 101. In addition, the IO interface 150 can output commands or data received from other components of the electronic device 101 to the user or other external devices.
[0056] Display 160 includes, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, or a micro-electromechanical system (MEMS) display, or an electronic paper display. Display 160 may also be a depth perception display, such as a multi-focus display. Display 160 is capable of displaying, for example, various contents (e.g., text, images, videos, icons, or symbols) to a user. Display 160 may include a touch screen and may receive, for example, touch, gesture, proximity, or hover input using an electronic pen or a body part of a user.
[0057] For example, the communication interface 170 can establish communication between the electronic device 101 and an external electronic device (e.g., the first electronic device 102, the second electronic device 104, or the server 106). For example, the communication interface 170 can be connected to the network 162 or the network 164 through wireless communication or wired communication to communicate with the external electronic device. The communication interface 170 can be a wired transceiver or a wireless transceiver or any other component for transmitting and receiving signals (such as images, head or camera posture data or other information).
[0058] The electronic device 101 also includes one or more sensors 180, which can measure physical quantities or detect the activation state of the electronic device 101, and convert the measured or detected information into electrical signals. For example, the sensor 180 may include one or more buttons for touch input, one or more cameras, a gesture sensor, a gyroscope or a gyroscope sensor, an air pressure sensor, a magnetic sensor or a magnetometer, an acceleration sensor or an accelerometer, a grip sensor, a proximity sensor, a color sensor (e.g., a red, green, and blue (RGB) sensor), a biophysical sensor, a temperature sensor, a humidity sensor, an illumination sensor, an ultraviolet (UV) sensor, an electromyogram (EMG) sensor, an electroencephalogram (EEG) sensor, an electrocardiogram (ECG) sensor, an IR sensor, an ultrasonic sensor, an iris sensor, a fingerprint sensor, etc. The sensor 180 may also include an inertial measurement unit, which may include one or more accelerometers, gyroscopes, and other components. The sensor 180 may also include a control circuit for controlling at least one sensor included therein. Any of these sensors 180 may be located within the electronic device 101. One or more cameras may capture images for object detection and 3D environment mapping. An inertial measurement unit may track head pose and / or other movements of the electronic device 101 .
[0059] The first external electronic device 102 or the second external electronic device 104 may be a wearable device or a wearable device (e.g., a head mounted display (HMD)) in which the electronic device 101 may be mounted. When the electronic device 101 is mounted in the electronic device 102 (e.g., HMD), the electronic device 101 may communicate with the electronic device 102 through the communication interface 170. The electronic device 101 may be directly connected to the electronic device 102 to communicate with the electronic device 102 without involving a separate network. The electronic device 101 may also be an augmented reality wearable device, such as glasses, which include one or more cameras for SLAM and object detection for detecting 3D association of an object.
[0060] The wireless communication can use, for example, at least one of Long Term Evolution (LTE), Long Term Advanced Evolution (LTE-A), fifth generation wireless system (5G), millimeter wave or 60 GHz wireless communication, wireless USB, code division multiple access (CDMA), wideband code division multiple access (WCDMA), universal mobile telecommunications system (UMTS), wireless broadband (WiBro) or global system for mobile communications (GSM) as a cellular communication protocol. The wired connection may include at least one of universal serial bus (USB), high-definition multimedia interface (HDMI), recommended standard 232 (RS-232) or plain old telephone service (POTS).
[0061] The network 162 includes at least one of a communication network, for example, a computer network (eg, a local area network (LAN) or a wide area network (WAN)), the Internet, or a telephone network.
[0062] The first external electronic device 102 and the second external electronic device 104 and the server 106 may all be devices of the same type or different types as the electronic device 101. According to certain embodiments of the present disclosure, the server 106 includes a group of one or more servers. According to certain embodiments of the present disclosure, all operations or some operations implemented on the electronic device 101 may be implemented on another one or more other electronic devices (e.g., the electronic device 102 and the electronic device 104 or the server 106). According to certain embodiments of the present disclosure, when the electronic device 101 should automatically or upon request perform some functions or services, the electronic device 101 may request another device (e.g., the electronic device 102 and the electronic device 104 or the server 106) to perform at least some functions associated with it, rather than independently or additionally performing the function or service. Other electronic devices (e.g., the electronic device 102 and the electronic device 104 or the server 106) are capable of performing the requested function or additional function and transmitting the result of the execution to the electronic device 101. The electronic device 101 may provide the requested function or service by processing the received result as is or additionally. To this end, for example, cloud computing, distributed computing, or client-server computing techniques may be used.
[0063] although Figure 1 It is shown that the electronic device 101 includes the communication interface 170 to communicate with the external electronic device 104 or the server 106 via the network 162 , but according to an embodiment of the present disclosure, the electronic device 101 may independently operate without a separate communication function.
[0064] The server 106 may support driving the electronic device 101 by implementing at least one of the operations (or functions) implemented on the electronic device 101. For example, the server 106 may include a processing module or a processor that may support the processor 120 implemented in the electronic device 101.
[0065] although Figure 1 An example of a network configuration 100 is shown, but may be Figure 1 Various changes may be made. For example, network configuration 100 may include any number of each component in any suitable arrangement. In general, computing and communication systems have a variety of configurations, and Figure 1 The scope of the present disclosure is not limited to any particular configuration. Figure 1 One operating environment is shown in which the various features disclosed in this patent document may be used, but these features may be used in any other suitable system.
[0066] Figure 2 An example captured image 200 is shown for use with an augmented reality tagging application executed by a processor (such as the processor 120 of the electronic device 101) according to an embodiment of the present disclosure. Figure 2 As shown in the example captured image 200 of , the captured image 200 includes various detectable objects located within the frame of the captured image 200. When the processor detects the position of the objects within the captured image 200, a bounded area 202 is created and displayed around each detected object. The processor defines the bounded area 202 by the location or position of the associated detected object within the captured image 200. For example, the processor may determine the center (x, y) coordinates of the detected object within the frame and define the width (w) and height (h) of the bounded area so that the entire detected object is within the bounded area. When the bounded area is a square or rectangular area, and because the detected object can be of various shapes, there may be a portion of the bounded area 202 that does not include the surface of the detected object, such as a portion of the surrounding environment in the bounded area 202.
[0067] The application executed by the processor also classifies the detected objects and marks the detected objects with virtual objects. Image recognition can be implemented to classify each object in a predetermined object class. For example, Figure 2 As shown, a cup detected on a table in captured image 200 is classified into the "cup" class, a clock detected is classified into the "clock" class, a banana detected is classified into the "banana" class, and a cell phone detected is classified into the "cell phone" class. The processor may then mark each detected item with a graphic 204, displaying the graphic 204 near the bounded area 202 around the detected object. For example, Figure 2 As shown, the displayed graphic 204 adjacent to each bounded region 202 is a ball with the object class of the detected object displayed textually near the ball. In some embodiments, the displayed graphic can also be colored based on the class. For example, Figure 2The graphics 204 of the cup, the cell phone, the clock, and the banana detected in FIG. 2 may be displayed as a red ball, a blue ball, a green ball, and a yellow ball, respectively.
[0068] The detected objects can also be labeled with virtual objects that are visually similar to the associated object class. Figure 2 As shown, a virtual cup object 206 may be placed near the detected / cup. The virtual cup object 206 may have a different size than the detected cup. In order for the placement of the virtual object to be attractive to the user, the virtual object will be proportional and well placed in 3D space relative to the associated detected object of the virtual object in the captured image 200. Thus, the application determines the scale of the detected object (such as to determine whether it is a real car or a toy car) and presents the virtual object in the correct scale near the detected object. For example, if a clock is detected on a table, such as on Figure 2 In the example of FIG. 2 , a virtual clock similar to the one detected but of a different size may be placed near the detected clock for comparison. The processor also determines the depth of the detected object so that the virtual object can be placed at the correct depth in the captured image. This is particularly important for depth-aware displays such as multi-focal displays. In order to accurately display virtual objects, according to various embodiments of the present disclosure, SLAM is used to map the 3D space of the environment in which the object is detected, and the 2D bounded area 202 is associated with the 3D space.
[0069] Figure 3 An example of an augmented reality object detection and 3D association architecture 300 according to an embodiment of the present disclosure is shown. The architecture 300 includes a 3D association block 302. The 3D association block 302 defines various functions performed by a processor (such as processor 120). The processor receives data from two distinct data pipelines to implement 3D association of detected objects. One pipeline provides object detection information, and the other pipeline provides 3D head or camera pose and 3D environment information (SLAM data). For the object detection pipeline, the object detection block 304 receives data from an RGB camera 306 through an associated image signal processor (ISP) 308. The data received at the object detection block 304 includes RGB frames or images. Each RGB frame or image may include a timestamp.
[0070] The connection between the RGB camera 306 and the ISP 308 to the object detection block 304 can be an on-the-fly (OTF) connection so that the connection operates independently and in parallel with other OTF connections described with respect to the architecture 300. The object detection block 304 executed by the processor detects objects in the RGB image captured by the RGB camera 306. The object detection block 304 defines a bounded area around each detected object in the RGB image and assigns an object ID, object class, and timestamp to the detected object and / or bounded area. The object detection block 304 also defines the position and size of the bounded area around the detected object in accordance with various embodiments of the present disclosure, such as by determining the (x, y) position of the detected object and determining the size (h, w) of the bounded area around the detected object in the captured image. The detected object and bounded area information (such as object ID, object class, timestamp, position, and size) is stored in the memory 310. The memory 310 may be a temporary memory location or a permanent storage location local to the processor executing the object detection block 304 and / or the 3D association block 302 , a remote storage location such as a server, or other storage location.
[0071] For the 3D environment pipeline, for example, the 3D head pose block 312 receives data from the monochrome camera 314 through the associated ISP 316. The data received at the 3D head pose block 312 includes monochrome frames or images, and each frame or image may include a timestamp. The connection between the monochrome camera 314 and the ISP 316 to the 3D head pose block 312 can be an OTF connection so that the connection operates independently of and in parallel with other OTF connections described with respect to the architecture 300. The 3D head pose block 312 also receives motion and orientation data including timestamps (such as timestamps for each head pose or camera pose) from the inertial measurement unit (IMU) 318 through an interface (IF) 320. The connection between the IMU 318 and the IF 320 to the 3D head pose block 312 can be an OTF connection so that the connection operates independently of and in parallel with other OTF connections described with respect to the architecture 300. It should be understood that in some embodiments, a single camera such as the RGB camera 306 can be used for both data pipelines. For example, a single camera may capture images for object detection and used in conjunction with IMU data for SLAM processing. In some embodiments, RGB camera 306 may be a monochrome camera configured to capture images for object detection, or may be another type of image sensing device. In some embodiments, monochrome camera 314 may be a second RGB camera to provide data to 3D head pose block 312, or may be another type of image sensing device. In other embodiments, other devices such as laser scanners or sonar devices may be used to provide images, 3D spatial environment data for generating 3D point clouds, and other environment and SLAM data.
[0072] The 3D head pose block 312 executed by the processor uses the images captured by the monochrome camera 314 and the motion and orientation data such as the head or camera pose received from the IMU 318 to generate a 3D point cloud of the environment. Data from the 3D head pose block 312 (such as head pose, 3D point cloud, image with 2D feature points, or other information) can be stored in the memory 322. The memory 322 can be a temporary memory location or a permanent storage location local to the processor executing the 3D head pose block 312, the object detection block 304 and / or the 3D association block 302, a remote storage location such as a server, or other storage location. The memory 322 can also be the same memory as the memory 310, or a separate memory.
[0073] The 3D association block 302 executed by the processor receives data from the two data pipes and determines the position of the detected object in 3D space. According to various embodiments described herein, the position of the detected object can be determined by using at least the bounded area associated with the detected object and the head posture data received from the data pipe. Once the 3D position of the detected object is determined, the 3D position is stored in the memory 324 together with other object information such as object ID and timestamp. The memory 324 can be a temporary memory location or a permanent storage location local to the processor executing the 3D association block 302, the object detection block 304 and / or the 3D head posture block 312, a remote storage location such as a server, or other storage location. The memory 324 can also be the same memory as the memory 310 or the memory 322, or a separate memory.
[0074] The object information stored in the memory 324 is used by the rendering block 326 executed by the processor to attach graphics and virtual objects to the detected objects in 3D space. The rendering block 326 retrieves the object information from the memory 324, retrieves the 3D model from the memory 328, and retrieves the head posture data from the 3D head posture block 312. The 3D model can be retrieved according to the category of the detected object. The rendering block 326 scales the 3D model based on the determined scale of the detected object, and determines a 3D layout including depth for each 3D model based on the determined 3D position of the detected object. The memory 328 can be a temporary memory location or a permanent storage location local to the processor executing the rendering block 326, the object detection block 304, the 3D head posture block 312 and / or the 3D association block 302, can be a remote storage location such as a server, or other storage location. The memory 328 can also be the same memory as the memory 310, the memory 322 or the memory 324, or a separate memory. In some embodiments, the 3D association block 302 may have a fast communication / data channel to the object detection pipeline and the SLAM pipeline, such as by accessing a shared memory between the two pipelines. The rendering block 326 provides rendering information to the display processing unit (DPU) or display processor 330 to display an image on a display 332, wherein the image includes graphics and virtual objects attached to the detected objects.
[0075] although Figure 3 One example of an augmented reality object detection and 3D association architecture 300 is shown, but may be used for Figure 3Various changes may be made. For example, architecture 300 may include any number of each component in any suitable arrangement. It should be understood that the functions implemented by the various blocks of architecture 300 may be implemented by a single processor, or distributed between two or more processors, or within the same electronic device or arranged in separate electronic devices. It should also be understood that RGB camera 306, monochrome camera 314, and IMU 318 may be relative to Figure 1 The functions described may be implemented in at least one processor, a graphics processing unit (GPU), or other components. In general, computing architectures have a variety of configurations, and Figure 3 The scope of the present disclosure is not limited to any particular configuration.
[0076] Figure 4 A flow chart of an example object detection and 3D association process 400 according to an embodiment of the present disclosure is shown. Although the flow chart depicts a series of sequential steps, no inference should be drawn from this sequence regarding a particular order of implementation unless explicitly stated, implementation of steps or portions of steps sequentially rather than simultaneously or in an overlapping manner, or implementation of the specifically depicted steps without intervening or intermediate steps occurring. Figure 4 The process described in Figure 1 The method is implemented in the electronic device 101 and can be executed by the processor 120.
[0077] At box 402, the processor controls one or more cameras to capture images of the environment. At box 404, the processor detects an object in the captured image. At box 406, the processor defines a bounded area around the detected object. The bounded area includes a calculation to enclose the (x, y) position and size (w, h) of the detected object. At box 408, the processor generates head posture data. The head posture data can be a head or camera posture provided to the processor by the IMU 318, and may include other information such as a 3D point cloud, an image with 2D feature points, a timestamp, or other information. At box 410, according to various embodiments of the present disclosure, the processor uses the bounded area and the head posture data to determine the position of the detected object in a 3D space associated with the environment.
[0078] although Figure 4 An example process is shown, but the Figure 4 Various changes may be made. For example, although shown as a series of steps, various steps in each figure may overlap, occur in parallel, occur in a different order, or occur multiple times.
[0079] Figure 5A diagrammatic illustration of a 3D feature point reprojection process for 3D association of detected objects according to an embodiment of the present disclosure is shown. The process may be implemented by the electronic device 101 and may be executed by the processor 120. The reprojection process includes capturing an RGB frame 502 from a specific head pose (P_trgb) by a sensor such as an RGB camera 306. The processor detects an object 504 within the RGB frame 502, such as via the object detection block 304, and defines a bounded region 506 around the detected object 504. The bounded region 506 has (x, y) coordinates within the RGB frame 502, and the processor determines the size of the bounded region 506 to surround the detected object 504 based on the determined width (w) and height (h). The detected object and bounded region information (such as object ID, object class, timestamp, location, and size) may be stored in a memory (such as in the memory 310) for use during 3D association. Using SLAM, the processor generates a 3D point cloud 508 (such as via 3D head pose block 312) based on one or more parameters (such as head pose data received from an IMU such as IMU 318) and by determining a plurality of 3D feature points of the environment from generated image data (e.g., data from monochrome camera 314). SLAM data (such as head pose, 3D point cloud, image with 2D feature points, or other information) may be stored in a memory (such as in memory 322) for use during 3D association.
[0080] It should be understood that various SLAM algorithms can be used to generate a 3D point cloud of an environment, and the present disclosure is not limited to any particular SLAM algorithm. According to an embodiment of the present disclosure, the same camera that captures the RGB frame 502 or a separate camera can be used to create a 3D point cloud 508 in combination with motion and orientation data from the IMU. Multiple feature points of the 3D point cloud 508 correspond to detected features (objects, walls, etc.) in the environment and provide a mapping of the features within the environment.
[0081] The reprojection process also includes: the processor maps or reprojects the 3D point cloud through the captured RGB frame 502 and the bounded region 506, such as via the 3D association block 302, so that the 3D feature points are located within the field of view (FOV) 510 of the RGB camera of the pose P_trgb. RGB camera parameters 512 for the RGB frame 502 can be used to provide the FOV 510. The processor maps the bounded region 506 including the detected object 504 to the FOV 510. Once the bounded region 506 is mapped within the FOV 510 and the 3D point cloud 508 is reprojected into the FOV 510, the bounded region 506 will include a subset of the feature points from the 3D point cloud 508 projected through the bounded region 506, such as Figure 5As shown. In some embodiments, the processor may specifically generate feature points and reproject the feature points into the bounded area 506. For example, the processor may use a sparse graph instead of a dense graph to generate the 3D point cloud 508. The sparse graph may reduce the calculations used to generate the 3D point cloud 508, but feature points on the detected object 504 in the bounded area 506 may be lost. If the sparse graph is used to generate the 3D point cloud 508, the 3D association block 302 executed by the processor may request the SLAM block 312, which may also be executed by the same processor, to specifically generate 3D feature points within the detected bounded area 506 and keep the 3D feature points of the bounded area 506 in the 3D point cloud.
[0082] Once the 3D feature points are reprojected into the FOV 510, the processor determines which 3D feature points within the bounded area 506 are located on the surface of the detected object 504. The processor removes the 3D feature points that are not located on the surface of the detected object 504 from the bounded area 506. The processor determines the 3D position of the detected object 504 in the 3D space 514 including (x, y, z) coordinates based on the remaining 3D feature points on the surface of the detected object 504. The 3D position is stored in a memory (such as in the memory 324) along with other object information such as an object ID and a timestamp for use by the augmented reality marking application. During operation of the augmented reality marking application, the processor can use the determined 3D position of the detected object 504 to place the detected object 504 in the 3D space 514 at the rendering block 326 and for display on the display 332.
[0083] Figure 6 A diagram 600 of another illustrative example of a 3D feature point reprojection process according to an embodiment of the present invention is shown. The FOV 510 shown includes a bounded area 506. The processor reprojects 3D feature points 602 from the 3D point cloud 508 into the bounded area 506. The processor determines which 3D feature points 602 in the bounded area are not arranged on the surface of the detected object 504 and removes those feature points. The processor can then determine the position of the detected object 504 within the 3D space 514 based on the 3D feature points 602 arranged on one or more surfaces of the detected object 504.
[0084] Figure 7 A flow chart of an example 3D feature point reprojection process 700 according to an embodiment of the present disclosure is shown. Although the flow chart depicts a series of sequential steps, no inference should be drawn from this sequence regarding a particular order of implementation, implementation of steps or portions of steps sequentially rather than simultaneously or in an overlapping manner, or implementation of specifically depicted steps without intervening or intermediate steps occurring unless explicitly stated. Figure 7 The process described in Figure 1 The method is implemented in the electronic device 101 and can be executed by the processor 120.
[0085] At box 702, the processor captures an image using one or more cameras and defines a bounded area in the image surrounding the detected object. The process may also associate head pose data with an image corresponding to the head pose when the image was captured. At box 704, the processor generates a 3D point cloud. According to an embodiment of the present disclosure, a 3D point cloud may be generated based on one or more captured images and data received by the processor from an inertial measurement unit. At box 706, the processor reprojects 3D feature points of the 3D points through the field of view of the image captured in box 702. At decision box 708, the processor determines whether a sparse graph was used to generate the 3D point cloud at box 704. If not, process 700 moves to decision box 710.
[0086] At decision box 710, the processor determines whether there are feature points projected into the bounded area and not on the surface of the detected object. If so, process 700 moves to box 712. At box 712, the processor removes 3D feature points within the bounded area that are not located on the surface of the detected object. Then, the process moves to box 714. If the processor determines at decision box 710 that there are no feature points in the bounded area that are not located on the surface of the detected object, the process moves from decision box 710 to box 714. At box 714, the processor stores the position of the object in 3D space based on the 3D feature points arranged on the detected object. Based on box 714, other operations of the electronic device 101 may occur, such as presenting and displaying virtual objects adjacent to the 3D position of the detected object.
[0087] If the processor determines at decision box 708 that a 3D point cloud is generated using a sparse graph at box 704, the process moves from decision box 708 to decision box 716. At decision box 716, the processor determines whether SLAM is available. In some scenarios, such as if there are limited available system resources due to other processes running on the electronic device 101, SLAM may be offline. In some scenarios, SLAM may be online, but may run too far before or after the object detection pipeline. If the processor determines at decision box 716 that the SLAM pipeline is not available, the process moves to box 718. At box 718, the processor uses the pre-configured settings of the object class assigned to the detected object to place the object in 3D space. For example, the processor may store a default scale of the object class in a memory and use the default scale to determine the distance or depth at which the virtual object is placed in the captured image. The processor may also use a pre-configured distance of the object class. In this case, the processor places the detected object in 3D space according to the 2D (x, y) coordinates of the detected object and the pre-configured distance.
[0088] If the processor determines at decision block 716 that SLAM is available, process 700 moves to block 720. At block 720, the processor generates a dense map of 3D feature points within the bounded area and maintains the dense map for the bounded area in the 3D point cloud. Creating a separate dense map of the bounded area provides for generating sufficient feature points on the surface of the objects in the bounded area without using system resources to create a dense map of the entire environment. Process 700 moves from block 720 to block 710, where the processor determines whether any feature points from the dense map generated at block 720 are not located on the surface of the detected object, removes these feature points at block 712, and stores the 3D position of the detected object at block 714.
[0089] The 3D feature points generated by SLAM provide data about the 3D environment, including the depth and scale of objects and the characteristics of the environment. By reprojecting the 3D feature points onto the surface of the detected object, the 3D position of the detected object in the environment can be determined, including the depth and scale of the detected object. This allows the properties of the detected object to be determined. For example, by determining the scale of the detected object, a toy car can be distinguished from a real car, which can then be identified by relative Figure 2 The described augmented reality marking application correctly classifies the detected object. The depth of the object can also be used to provide information to the user (such as the distance from the user to the object). In addition, the depth and scale of the object detected in the image can be used to accurately attach the virtual object to the object detected in the image, so that the application places the virtual object near the object detected in the image and scales the virtual object as defined by the application (the virtual object may intentionally have a different scale than the detected object), and so that the virtual object remains near the object detected in the image even when the user moves the camera between different poses.
[0090] Using SLAM for 3D spatial understanding while performing object detection using 2D images requires less computation and system resources than attempting full 3D object detection, which provides faster object detection and object labeling. This is important for augmented reality applications, in which users utilize cameras to provide a continuous view of a scene or environment, because users expect applications to display accurate and continuously updated information about objects in the environment on the user's device.
[0091] although Figure 7 An example process is shown, but the Figure 7 Various changes may be made. For example, although shown as a series of steps, various steps in each figure may overlap, occur in parallel, occur in a different order, or occur multiple times.
[0092] In some scenarios, the system resources of the electronic device 101, or other devices that implement one or more operations described in the present disclosure, may be limited by other processes being performed by the electronic device 101. In such scenarios, the available system resources may prevent SLAM from generating a 3D point cloud. In this case, the bounded area of the captured image that can be processed by the SLAM pipeline is used to provide 2D feature points on the captured frame. If the available system resources do not allow the generation of 2D feature points, the object detection pipeline may implement pixel matching between bounded areas of the separately captured images to provide feature points via pixel matching. Once feature points are provided (such as by generating 2D feature points or by implementing pixel matching), the electronic device 101 may implement triangulation between matching feature points to associate the detected object with a 3D space. Therefore, the electronic device 101 may monitor the currently available system resources, and may also monitor the quality of the 3D point cloud, and switch the 3D association method (reprojection, 2D feature point triangulation, pixel matching triangulation) in real time according to the currently available system resources.
[0093] Figure 8 A diagram 800 of an example feature point triangulation process according to an embodiment of the present disclosure is shown. The process may be implemented by the electronic device 101 and may be executed by the processor 120. The camera 802 captures a first frame 804 of an object 806 included in the environment at a time T and a first pose. The camera 802 or another camera captures a second frame 808 of the object 806 at a time T+t and a second pose. The processor processes the first frame 804 and the second frame 808 and the pose data of the first pose and the second pose to detect the object 806 in each image and generate feature points on the surface of the object 806. If the available system resources allow the SLAM pipeline to provide 2D feature points in the image, the feature points may be 2D feature points. If the SLAM pipeline does not provide 2D feature points and only provides poses due to the available system resources, the object detection pipeline may perform pixel matching between bounded areas of the first frame 804 and the second frame 808 to determine where the same pixels and thus the same object features are located within the first frame 804 and the second frame 808. For example, the processor may determine that pixels in separate frames that include the same RGB values and possibly other pixel values are matching pixels.
[0094] The processor triangulates the feature points in the first frame 804 and the second frame 808 using the pose information provided by the SLAM pipeline. Figure 8As shown, when the first frame 804 and the second frame 808 are captured respectively, the feature points at the corners of the cube-shaped object can be triangulated using the pose data to determine the orientation of the camera 802 at time T and T+t. The processor triangulates the feature points of the object 806 from each frame 804 and frame 808 to determine the position of the same feature points in 3D space. Triangulation can be performed on multiple feature points (2D feature points or matching pixels) to place the triangulated object 810 including (x, y, z) coordinates in 3D space. The processor can triangulate multiple feature points to determine the 3D position and scale of the object. For example, the processor can triangulate the feature points at Figure 8 The feature points at each corner of the cube-shaped object shown are triangulated, which can provide enough information to accurately represent the entire shape in 3D space. It should be understood that although Figure 8 Two poses and two captured frames are shown, but more than two poses and frames may be used to provide multiple triangulation points.
[0095] Fig. 9 A flow chart of an example detected object triangulation process 900 according to an embodiment of the present disclosure is shown. Although the flow chart depicts a series of sequential steps, no inference should be drawn from this sequence regarding a particular order of implementation unless explicitly stated, implementation of steps or portions of steps sequentially rather than simultaneously or in an overlapping manner, or implementation of the specifically depicted steps without intervening or intermediate steps occurring. Fig. 9 The process described in Figure 1 The method is implemented in the electronic device 101 and can be executed by the processor 120.
[0096] At block 902, the processor determines that 3D feature points are not provided due to available system resources. At block 904, a plurality of frames are captured by the electronic device, wherein each frame is associated with a gesture. The processor performs object detection on each of the plurality of frames and defines a bounded region around at least one detected object in each of the plurality of frames. At decision block 908, the processor determines whether available system resources allow for generation of 2D feature points in the captured image. If so, process 900 moves to block 910.
[0097] At box 910, the processor generates 2D feature points within at least a bounded area of each frame in the plurality of frames. At box 912, the processor sets one or more 2D feature points within the bounded area as triangulation parameters. At box 914, the processor triangulates the 3D position of the detected object in two or more frames in the plurality of frames using the set triangulation parameters. In the case where the processor sets the triangulation parameters to the generated 2D feature points in box 912, the processor triangulates the one or more feature points using head pose data of the two or more images.
[0098] If the processor determines at decision block 908 that the available system resources do not allow for the generation of 2D feature points, the process 900 moves to block 916. At block 916, the processor performs pixel matching between pixels within the bounded region of a particular object in two or more of the plurality of frames. At block 918, the processor sets the one or more matched pixels as triangulation parameters or feature points to be used in triangulation. From block 918, the process 900 moves to block 914, where the processor triangulates the position of the detected object using the set triangulation parameters (in this case, the matched pixels) to determine the 3D position and scale of the detected object.
[0099] As described herein, triangulation using 2D feature points and pixel matching provides depth and scale of the detected object. This allows the nature of the detected object to be determined. For example, a toy car can be distinguished from an actual car by determining the scale of the detected object, which can then be determined by comparing the depth and scale of the detected object to the real car. Figure 2 The described augmented reality marking application correctly classifies the detected object. The depth of the object can also be used to provide information to the user (such as the distance from the user to the object). In addition, the depth and scale of the object detected in the image can be used to accurately attach the virtual object to the object detected in the image, so that the application places the virtual object near the object detected in the image, and scales the virtual object as defined by the application (the virtual object may intentionally have a different scale than the detected object), and so that the virtual object remains near the object detected in the image even when the user moves the camera between different poses.
[0100] The triangulation process described herein allows augmented reality tagging applications to continue to run efficiently even when fewer system resources are available. This is important for augmented reality applications (in which a user provides a continuous view of a scene or environment using a camera) because the user expects the application to display accurate and continuously updated information about objects in the environment on the user's device.
[0101] although Fig. 9 An example process is shown, but the Fig. 9 Various changes may be made. For example, although shown as a series of steps, various steps in each figure may overlap, occur in parallel, occur in a different order, or occur multiple times.
[0102] Nothing in this application should be construed as implying that any particular element, step, or function is an essential element that must be included in the claims scope. The scope of patentable subject matter is limited solely by the claims. Furthermore, none of the claims is intended to invoke 35 U.S.C. 112(f) unless the exact words "means for..." are followed by a participle.
Claims
1. Electronic devices, including: at least one image sensor; Inertial measurement sensors; Memory; as well as at least one processor coupled to the at least one image sensor, the inertial measurement sensor, and the memory, Wherein, the at least one processor is configured as: capturing an image of the environment using the at least one image sensor; Detecting objects in captured images; defining a bounded region in the image surrounding the detected object; receiving head pose data from the inertial measurement sensor; Determine available system resources; Switching to a designated 3D association method among a plurality of 3D association methods based on the determined available system resources; and determining the position of the detected object in 3D space using the head pose data and the bounded area in the captured image according to the specified 3D association method, Wherein, the at least one processor is further configured to: Based on determining that the available system resources are insufficient to allow generation of 2D feature points in the captured image, switching to a pixel matching triangulation method among the plurality of 3D association methods, Wherein, in the pixel matching triangulation method, in order to determine the position of the detected object, the at least one processor is further configured to: performing pixel matching between pixels located within the bounded region, Set the matched pixels as feature points to be used in triangulation, and The position of the detected object is determined by triangulation using the set feature points.
2. The electronic device according to claim 1, wherein: The at least one processor is further configured to: switch to a reprojection method among the plurality of 3D association methods based on determining that the available system resources are sufficient to provide 3D feature points, Wherein, in the reprojection method, in order to determine the position of the detected object, the at least one processor is further configured to: generating a 3D point cloud of at least a portion of the environment using the at least one image sensor and the inertial measurement sensor; Projecting feature points of the 3D point cloud into a captured image, wherein the captured image includes the bounded area; determining feature points located on the surface of the object among the feature points; and Feature points that are not located on the surface of the object are removed.
3. The electronic device according to claim 2, wherein: The at least one processor is further configured to: Determining that a sparse graph is used to generate the 3D point cloud; Generating a dense 3D point cloud of the bounded area; and The feature points of the dense 3D point cloud are projected into the bounded area of the captured image.
4. The electronic device according to claim 1, wherein: in, The at least one processor is further configured to: switching to a 2D feature point triangulation method among the plurality of 3D association methods based on determining that the available system resources are sufficient to allow generation of 2D feature points in the captured image but insufficient to provide 3D feature points, Wherein, in the 2D feature point triangulation method, in order to determine the position of the detected object, the at least one processor is further configured to: capturing a plurality of images using the at least one image sensor; detecting the object in each of the plurality of images; defining bounded regions in the plurality of images, wherein each of the bounded regions is defined around a detected object in one of the plurality of images; and The location of the object in one of the plurality of images is triangulated using the location of the object in at least another image of the plurality of images.
5. The electronic device according to claim 4, wherein: To triangulate the position of the object, the at least one processor is further configured to: generating 2D feature points in each of the plurality of images; Determining the 2D feature points located within each of the bounded regions in the plurality of images; as well as The 2D feature points located in one of the bounded regions of one of the images are triangulated using the 2D feature points located in another of the bounded regions of another of the images.
6. The electronic device according to claim 4, wherein: To triangulate the position of the object, the at least one processor is further configured to: performing pixel matching between each of the bounded regions of at least two images of the plurality of images; as well as Triangulation is performed between matching pixels of the at least two images of the plurality of images.
7. The electronic device according to claim 1, wherein: The at least one processor is further configured to place a virtual object adjacent to the detected object in the 3D space, wherein the detected object and the virtual object are associated with an object class.
8. A method for detecting 3D association of objects, the method comprising: capturing an image of the environment using at least one image sensor; Detecting objects in captured images; defining a bounded region in the image surrounding the detected object; receiving head pose data from an inertial measurement sensor; Determine available system resources; Switching to a specified 3D association method among a plurality of 3D association methods based on the determined available system resources; as well as determining the position of the detected object in 3D space using a processor using the head pose data and the bounded area in the captured image according to the specified 3D association method, Wherein, switching to the specified 3D association method comprises: based on determining that the available system resources are insufficient to allow generation of 2D feature points in the captured image, switching to a pixel matching triangulation method among the multiple 3D association methods, Wherein, in the pixel matching triangulation method, determining the position of the detected object includes: performing pixel matching between pixels located within the bounded region, Set the matched pixels as feature points to be used in triangulation, and The set feature points are used to perform triangulation to determine the position of the detected object in 3D space.
9. The method according to claim 8, wherein: Switching to the specified 3D association method includes: based on determining that the available system resources are sufficient to provide 3D feature points, switching to a reprojection method among the multiple 3D association methods, Wherein, in the reprojection method, determining the position of the detected object includes: generating a 3D point cloud of at least a portion of the environment using the at least one image sensor and the inertial measurement sensor; Projecting feature points of the 3D point cloud into a captured image, wherein the captured image includes the bounded area; determining feature points located on the surface of the object among the feature points; and Feature points that are not located on the surface of the object are removed.
10. The method of claim 9, determining the position of the detected object further comprising: Determining that a sparse graph is used to generate the 3D point cloud; Generating a dense 3D point cloud of the bounded area; as well as The feature points of the dense 3D point cloud are projected into the bounded area of the captured image.
11. The method according to claim 8, wherein: Switching to the specified 3D association method includes: based on determining that available system resources are sufficient to allow generation of 2D feature points in the captured image but insufficient to provide 3D feature points, switching to a 2D feature point triangulation method among the plurality of 3D association methods, Wherein, in the 2D feature point triangulation method, determining the position of the detected object includes: capturing a plurality of images using the at least one image sensor; detecting the object in each of the plurality of images; defining bounded regions in the plurality of images, wherein each of the bounded regions is defined around a detected object in one of the plurality of images; and The location of the object in one of the plurality of images is triangulated using the location of the object in at least another image of the plurality of images.
12. The method according to claim 11, wherein: Triangulating the position of the object includes: generating 2D feature points in each of the plurality of images; determining the 2D feature points located within each of the bounded regions in the plurality of images; and The 2D feature points located in one of the bounded regions of one of the images are triangulated using the 2D feature points located in another of the bounded regions of another of the images.
13. The method according to claim 11, wherein: Triangulating the position of the object includes: matching pixels between each of the bounded regions of at least two images of the plurality of images; and Matching pixels of the at least two images of the plurality of images are triangulated.
14. The method of claim 8, further comprising placing a virtual object adjacent to the detected object in the 3D space, wherein: The detected objects and the virtual objects are associated with object classes.
Citation Information
Patent Citations
3-dimensional scene analysis for augmented reality operations
US20170243352A1