Relative attitude estimation in multi-agent systems using a generated global reference frame map

WO2025184903A8PCT designated stage Publication Date: 2025-10-02QUALCOMM INC +5
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/080780
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-08
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing methods for inter-agent attitude estimation in multi-agent systems, such as those using GPS or motion capture, are ineffective in GPS-denied environments like indoor spaces due to the lack of sufficient point features for conventional camera-based techniques.

Method used

A method involving vanishing point estimation and re-identification techniques to generate a global reference frame map using parallel line segments, allowing for attitude estimation among agents by constructing reference coordinate frames and rotation matrices, even in environments with limited point features.

Benefits of technology

Enables accurate attitude estimation among agents in multi-agent systems, including indoor environments, facilitating better scene understanding and content sharing by determining the relative orientation of agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024080780_02102025_PF_FP_ABST
    Figure CN2024080780_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Techniques for relative attitude estimation in multi-agent system includes: obtaining a first reference coordinate frame based on line segments extracted from a first 2D image, of a 3D space, captured by a first device, obtaining a first rotation matrix of a first image sensor frame of the first device to the first reference coordinate frame, obtaining a second reference coordinate frame based on line segments extracted from a second 2D image, of the 3D space, captured by a second device, obtaining a second rotation matrix of a second image sensor frame of the second device to the second reference coordinate frame, and estimating an attitude of the first device with respect to the second device based on the first and second rotation matrices and a rotation chain.
Need to check novelty before this filing date? Find Prior Art

Description

RELATIVE ATTITUDE ESTIMATION IN MULTI-AGENT SYSTEMS USING A GENERATED GLOBAL REFERENCE FRAME MAP

[0001] INTRODUCTION

[0002] Field of the Disclosure

[0003] Aspects of the present disclosure relate to techniques for relative attitude estimation in multi-agent systems.

[0004] Description of Related Art

[0005] Multi-agent systems may include a number of agents. In some aspects, the agents may collaborate, such as to achieve a common goal, or perform one or more tasks collaboratively. As used herein, an agent is an entity or device. Example agents that make up a multi-agent system may include robots, augmented reality (AR) devices, virtual reality (VR) devices, wearable devices, cameras, satellites, unmanned aerial vehicles (UAVs) , and / or the like.

[0006] In certain aspects, agents of a multi-agent system may include image sensors (e.g., such as cameras) to perform computer vision tasks. For instance, these agents may collaboratively work together to perceive an environment for efficient and accurate situation awareness, which may be beneficial for performing tasks such as search and rescue, wide-area surveillance, environmental monitoring, collaborative mapping, and / or AR gaming, to name a few. For example, in the multi-agent system, image sensors of multiple agents may observe the same three dimensional (3D) space and generate two-dimensional (2D) images of the 3D space from different positions and / or at different angles, often with overlapping views, such that the fusion of all of their perceptions may lead to better scene understanding.

[0007] The ability to accurately estimate the attitude of agents, relative to one another in a multi-agent system, may be a key requirement to performing tasks, such as those listed above. As used herein, inter-agent attitude estimation refers to the estimation of the orientation of an agent relative to the orientation of another agent in the multi-agent system. For example, inter-agent attitude estimation may estimate a rotation between agent pairs in the multi-agent system. In practice, this may be achieved by using an external measurement system and / or technology, such as a global positioning system (GPS) or motion capture (mocap) . In some cases, however, agents may be deployed in locations where these systems and / or technologies are unavailable or ineffective, such as  GPS-denied environments. For example, agents deployed in indoor environments (e.g., such as inside a building) may not be able to use GPS signals for inter-agent attitude estimation due to GPS signals being blocked or reflected by the walls of the building. In such cases, other approaches may be needed for attitude estimation.SUMMARY

[0008] One aspect provides a method by an apparatus. The method includes obtaining a first reference coordinate frame based on a plurality of first line segments extracted from a first two-dimensional (2D) image, of a three-dimensional (3D) space, captured by a first device; obtaining a first rotation matrix of a frame of a first image sensor of the first device to the first reference coordinate frame; obtain a second reference coordinate frame based on a plurality of second line segments extracted from a second 2D image, of the 3D space, captured by a second device; obtaining a second rotation matrix of a frame of a second image sensor of the second device to the second reference coordinate frame; and estimating an attitude of the first device with respect to the second device based on the first rotation matrix, the second rotation matrix, and a rotation chain.

[0009] Other aspects provide: one or more apparatuses operable, configured, or otherwise adapted to perform any portion of any method described herein (e.g., such that performance may be by only one apparatus or in a distributed fashion across multiple apparatuses) ; one or more non-transitory, computer-readable media comprising instructions that, when executed by one or more processors of one or more apparatuses, cause the one or more apparatuses to perform any portion of any method described herein (e.g., such that instructions may be included in only one computer-readable medium or in a distributed fashion across multiple computer-readable media, such that instructions may be executed by only one processor or by multiple processors in a distributed fashion, such that each apparatus of the one or more apparatuses may include one processor or multiple processors, and / or such that performance may be by only one apparatus or in a distributed fashion across multiple apparatuses) ; one or more computer program products embodied on one or more computer-readable storage media comprising code for performing any portion of any method described herein (e.g., such that code may be stored in only one computer-readable medium or across computer-readable media in a distributed fashion) ; and / or one or more apparatuses comprising one or more means for performing any portion of any method described herein (e.g., such that performance would be by only one apparatus or by multiple apparatuses in a distributed fashion) . By way of example, an  apparatus may comprise a processing system, a device with a processing system, or processing systems cooperating over one or more networks. An apparatus may comprise one or more memories; and one or more processors configured to cause the apparatus to perform any portion of any method described herein. In some examples, one or more of the processors may be preconfigured to perform various functions or operations described herein without requiring configuration by software.

[0010] The following description and the appended figures set forth certain features for purposes of illustration.BRIEF DESCRIPTION OF DRAWINGS

[0011] The appended figures depict certain features of the various aspects described herein and are not to be considered limiting of the scope of this disclosure.

[0012] FIG. 1 depicts an example multi-agent system.

[0013] FIG. 2 depicts aspects of an example agent in a multi-agent system.

[0014] FIGS. 3A-3C depict an example workflow for the construction of a global reference frame map to enable attitude estimation among agents in a multi-agent system.

[0015] FIG. 4A depicts an example process flow for the construction of a reference coordinate frame, which is a constituent of a global reference frame map.

[0016] FIG. 4B depicts an example vanishing point representing an intersection of two-dimensional (2D) projections of parallel three-dimensional (3D) lines.

[0017] FIG. 5 depicts a method for inter-agent attitude estimation.DETAILED DESCRIPTION

[0018] Aspects of the present disclosure provide apparatuses, methods, processing systems, and computer-readable mediums for constructing a global reference frame map to realize the rotation among agents in a multi-agent system. The rotation may represent the estimated attitude of one agent relative to another agent within the multi-agent system.

[0019] Certain aspects described herein propose using (1) vanishing point estimation and (2) re-identification methods to generate 3D reference coordinate frames (simply referred to herein as “reference coordinate frames” ) , which make up the global reference frame map used for the relative attitude estimation in multi-agent systems. For example, vanishing point estimation may be used to estimate three orthogonal vanishing points  based on parallel line segments identified in and extracted from a 2D image of a 3D space. The 2D image may be an image captured by an agent in the system. Each vanishing point may represent a point where 2D projections of parallel line segments in the 3D space converge. A direction of parallel lines associated with each of the three vanishing points may represent a direction in a world coordinate system, and thus may be used to generate a reference coordinate frame.

[0020] After an initial reference coordinate frame is generated using the aforementioned techniques, line segments extracted from another 2D image and used to perform vanishing point estimation for generating another reference coordinate frame may be limited. For example, re-identification techniques may be used to identify extracted lines that were used to create an existing reference coordinate frame (e.g., the initial reference coordinate frame) such that only those extracted lines, which have not been previously used, are utilized to perform vanishing point estimation and generate another reference coordinate frame. As such, a subsequently created reference coordinate frame may be different than reference coordinate frame (s) previously created. This process of (1) re-identification and (2) vanishing point estimation may repeat for each 2D image of the 3D space generated by agents in the multi-agent system to generate multiple reference coordinate frames.

[0021] In addition to generating multiple reference coordinate frames, rotation matrices, each representing a rotation of a frame of an image sensor (e.g., used to capture a 2D image of a 3D space, such as a camera frame) , and associated with an agent in the multi-agent system relative to a reference coordinate frame, may be computed. For example, a rotation matrix may represent the rotation of an image sensor of the agent with respect to a particular reference coordinate frame. In some cases, a rotation matrix may be computed based on direction vectors corresponding to the vanishing points of a reference coordinate frame. In some cases, a rotation matrix may be computed based on direction vectors corresponding to the vanishing points of a reference coordinate frame and a pre-calibrated intrinsic matrix (K) of an image sensor associated with an agent that captured the 2D image used to generate the particular reference coordinate frame. As described herein, a rotation matrix may be computed between each agent capturing 2D images of the 3D space and each reference coordinate frame generated (or less than all reference coordinate frames generated) . Further, a rotation chain may be constructed  between each pair of reference coordinate frames generated using the aforementioned techniques.

[0022] Together the reference coordinate frames, the rotation matrices, and the rotation chains may make up a global reference frame map. In certain aspects, an agent may use the global reference frame map to determine an attitude of itself with respect to another agent in the multi-agent system (or vice versa) . In certain aspects, an agent may use the global reference frame map to determine an attitude of one agent in the multi-agent system with respect to another agent in the multi-agent system (e.g., where the agents are different from the agent performing the attitude estimation) .

[0023] As an illustrative example, a first reference coordinate frame may be generated based on first line segments extracted from a 2D image, of a 3D space, captured by a first camera associated with a first agent in the multi-agent system. A first rotation matrix may be computed between a frame of the first camera and the first reference coordinate frame. Further, a second reference coordinate frame may be generated based on second line segments extracted from a 2D image, of the same 3D space, captured by a second camera associated with a second agent in the multi-agent system. A second rotation matrix may be computed between a frame of the second camera and the second reference coordinate frame. A rotation chain may also be constructed between the first and second reference coordinate frames. In certain aspects, the first agent may estimate an attitude of itself with respect to the second agent based on the first rotation matrix, the second rotation matrix, and the rotation chain. In certain aspects, the second agent may estimate an attitude of itself with respect to the first agent based on the first rotation matrix, the second rotation matrix, and the rotation chain. In certain aspects, a third agent in the multi-agent system may estimate an attitude of the first agent with respect to the second agent, and vice versa.

[0024] Accordingly, the global reference frame map has the beneficial technical effect of enabling attitude estimation among any pair of agents in a multi-agent system. In certain aspects, this attitude estimation between agents may be useful for sharing content between agents. For example, a first agent may render content obtained at the first agent based on the attitude of the first agent determined with respect to a second agent, and send the rendered content to the second agent such that the content is rendered properly at the second agent even though the second agent is oriented differently than the first agent. In certain aspects, this attitude estimation between agents may be useful for  determining how to fuse together information about an environment perceived by multiple agents in the multi-agent system for better scene understanding of a 3D space.

[0025] Further, the use of parallel line classification and vanishing point estimation techniques allows for the creation of reference coordinate frames in the global reference frame map, even in the absence of a sufficient number of point features in a 3D space for which the reference coordinate frames are created. For example, conventional camera-based attitude estimation methods, such as visual Simultaneous Location and Mapping (vSLAM) and / or visual inertial odometer (VIO) are highly dependent on the existence of point features in 2D images of a 3D space for attitude estimation. For example, these conventional methods may use point features in 2D images to perform point feature matching to identify matching point features in two images and then estimate the relative geometric transformation between matched point features for attitude estimation. As such, a technical problem associated with these conventional methods involves the inability to use such methods in environments where point features are limited, such as indoor environments (e.g., an indoor office space may not include a sufficient number of point features for performing point feature matching to estimate relative attitude among agents) . The techniques described herein overcome this technical problem and improve upon the state of the art by enabling attitude estimation of agents deployed in any environment, including indoor environments having a limited number of point features.

[0026] Example Multi-Agent System

[0027] FIG. 1 depicts an example multi-agent system 100 including two independent entities, for example, a first agent 102 (1) and a second agent 102 (2) . As described herein, first agent 102 (1) and second agent 102 (2) may each be an entity configured to perceive and / or interact with an environment, such as 3D space 108.

[0028] For example, first agent 102 (1) may include a first image sensor 104 (1) (e.g., such as a first camera) and second agent 102 (2) may include a second image sensor 104 (2) (e.g., such as a second camera) . First agent 102 (1) may be situated at a different position than second agent 102 (2) in 3D space 108 and / or an orientation of first image sensor 104 (1) at first agent 102 (1) may be different than an orientation of second image sensor 104 (2) at second agent 102 (2) . First image sensor 104 (1) and second image sensor 104 (2) may observe the same 3D space 108 and generate 2D images, of 3D space 108, at different angles based on the differing locations of first agent 102 (1) and second agent 102 (2)  and / or differing orientations of first image sensor 104 (1) and second image sensor 104 (2) . For example, aspects of3D space 108 captured in a first 2D image plane 106 (1) associated with first image sensor 104 (1) may be different than aspects of 3D space 108 captured in a second 2D image plane 106 (2) associated with second image sensor 104 (2) . Different aspects may include different physical objects detected in 3D space 108, different perceptions (e.g., orientations, locations, etc. ) of the same physical objects detected in 3D space 108, different lines detected in 3D space 108, different perceptions of lines detected in 3D space 108, and / or others.

[0029] 3D space 108 may be any environment that can be perceived by first agent 102 (1) and second agent 102 (2) . In certain aspects, 3D space is an indoor environment such as an office space and / or a hallway. In certain aspects, 3D space 108 is an indoor environment having a limited number of point features that may be extracted for point feature matching.

[0030] In certain aspects, first agent 102 (1) and second agent 102 (2) may collaborate to perform one or more tasks such as, search and rescue, wide-area surveillance, environmental monitoring, collaborative mapping, and / or AR gaming, as described herein. In certain aspects, it may useful and / or necessary to understand the rotation between first agent 102 (1) and second agent 102 (2) , and more specifically, the attitude of first agent 102 (1) with respect to second agent 102 (2) and / or the attitude of second agent 102 (2) with respect to first agent 102 (1) . Techniques for performing attitude estimation among first agent 102 (1) and second agent 102 (2) are provided herein and are described with respect to workflow 300 depicted in FIGS. 3A-3C.

[0031] While FIG. 1 describes a multi-agent system having only two agents 102, other multi-agent systems may include more agents, and thus attitude estimation between different pairs of agents may be performed according to the techniques described herein.

[0032] FIG. 2 depicts aspects of an example agent 202 in a multi-agent system. Agent 202 in FIG. 2 may be an example of first agent 102 (1) and / or second agent 102 (2) in multi-agent system 100 in FIG. 1. In the depicted example, agent 202 may include one or more processors 206, one or more memories 208, and one or more image sensor (s) 204. Processor (s) 206, memory (ies) 208, and image sensor (s) 204 may be coupled by a bus 210, which may generally be configured for data exchange amongst the components. More generally, bus 210 may be configured to transmit programming instructions and / or  data among the processor (s) 206, memory (ies) 208, and image sensor (s) 204. Bus 210 may be representative of multiple buses while only one is depicted for simplicity.

[0033] Processor (s) 206 may be configured to retrieve and execute instructions stored in one or more memories, including local memory (ies) 208, as well as remote memory (ies) and data store (s) . Similarly, processor (s) 206 may be configured to store application data residing in local memory (ies) 208, as well as remote memory (ies) and data store (s) . In certain aspects, processor (s) 206 are representative of one or more central processing units (CPUs) , graphics processing units (GPUs) , tensor processing units (TPUs) , accelerators, and / or other processing devices.

[0034] Memory (ies) 208 may be configured as a volatile and / or a non-volatile computer-readable medium and, as such, may include one or more programming instructions thereon that, when executed by processor (s) 206, cause processor (s) 206 to complete various processes, such as the processes described herein with respect to FIGS. 3A-4A and 5. The programming instructions stored on memory (ies) 208 may be embodied as a plurality of software logic modules, where each logic module provides programming instructions for completing one or more tasks.

[0035] The one or more image sensors 204 of agent 202 may include, but are not limited to, optical sensor (s) (e.g., camera (s) , laser sensor (s) , etc. ) , thermal sensors, infrared sensors, and / or the like. In certain aspects, image sensor (s) 204 may be configured to produce, at least, 2D image data capturing a 3D space, such as 3D space 108 depicted in FIG. 1.

[0036] Aspects Related to the Construction of a Global Reference Frame Map for Attitude Estimation

[0037] FIGS. 3A-3C depict an example workflow 300 for the construction of a global reference frame map to enable attitude estimation among agents in a multi-agent system. For example, workflow 300 may be used to construct a global reference frame map to enable attitude estimation among first agent 102 (1) and second agent 102 (2) in example multi-agent system described and depicted with respect to FIG. 1.

[0038] As shown in FIG. 3A, workflow 300 begins at 302 with capturing a 2D image. In particular, at 302, a 2D image of a 3D space may be captured by an image sensor of an agent in the multi-agent system. The agent may be deployed to perceive the 3D space. A 2D image of 3D space 108 captured by first image sensor 104 (1) (e.g., such as a first  camera (c1) ) of first agent 102 (1) in FIG. 1 may be an example of the 2D image captured at 302. For illustration, the 2D image captured at 302 may be an image of a hallway inside an office building (e.g., as shown at 402 in FIG. 4A) .

[0039] Workflow 300 proceeds, at 304, with (1) extracting line segments and (2) determining a line descriptor per extracted line segment in the 2D image. For example, line segment detection techniques may be used to detect straight lines in the 2D image. The detected line segments may include Ln line segments where Ln = {l1, l2, l3, ... ln} and n is the total number of line segments detected in the 2D image. Further, a line descriptor may be generated per detected line segment to represent each individual line. Each line descriptor may be unique to the line segment it is describing. In certain aspects, the line descriptor may capture information about an appearance of a region around a corresponding line. In certain aspects, the line descriptor generated per line segment may be computed using information such as pixel intensity and / or other local image features. In certain aspects, a line band descriptor (LBD) method is used to generate the line descriptors for the detected line segments. In certain aspects, learning-based methods may be used to extract the line segments from the 2D image and generate a line descriptor per line.

[0040] Workflow 300 then proceeds, at 305, with determining whether at least one reference coordinate frame has been generated based on 2D image (s) captured by agent (s) in the multi-agent system. In this example, because the 2D image captured at 302 is the first image to be captured by any agent in the multi-agent system, no reference coordinate frames have been previously generated. Thus, workflow 300 proceeds to 306 to generate a first reference coordinate frame. The first reference coordinate frame, generated at 306, may be generated based on the line segments extracted at 304. Generation of the first reference coordinate frame, at 306, may involve using parallel line classification and vanishing point estimation techniques, which are described in detail with respect to FIG. 4A.

[0041] Specifically, FIG. 4A depicts an example process flow 400 for the construction of a reference coordinate frame using these techniques. As shown in FIG. 4A, process flow 400 may begin with capturing a 2D image, at 402, similar to FIG. 3A at 302. Process flow 400 may then proceed, at 404, with (1) extracting line segments and (2) determining a line descriptor per extracted line segment in the 2D image, similar to FIG. 3A at 302.

[0042] Process flow 400 proceeds, at 406, with generating a reference coordinate frame. Generating a reference coordinate frame, at 406, may involve performing steps 408-412.

[0043] For example, at 408, parallel line segments extracted from the 2D image (e.g., line segments parallel in 3D space) may be identified, and, at 410, vanishing points may be estimated. Each vanishing point detected may be a respective point where a subset of the line segments in the 2D image (e.g., parallel line segments) converge.

[0044] For example, as shown in FIG. 4B, a vanishing point 434 is an intersection point of 2D projections (e.g., in a 2D image 432) of parallel line segments 436 in a real-world (3D) environment. A vanishing point 434 may be used to calculate the direction vector of a set a parallel lines. Thus, by using parallel lines identified in a 2D image, an x-axis, a y-axis, and a z-axis (e.g., a 3D coordinate system) of a reference coordinate frame may be created.

[0045] Accordingly, at 410, a process to estimate vanishing points may include, as a first step, randomly selecting two lines, li and lj from detected line segments Ln = {l1, l2, l3, ... ln} (e.g., in the 2D image) . With line segments li and lj, a vanishing point may be estimated as: vij = li x lj

[0046] A score for this vanishing point (vij) may be defined as:

[0047] where

[0048] and

[0049] Further, lp may represent a line segment and δp may represent a length of the line segment lp. Variable, T, may refer to the transpose of a vector or a matrix. Variable,  may represent the normalized vector of v and may represent the  minimum angle between a current line segment lp and the selected line segments li and lj. Variable, ρth, may represent a user-defined threshold, which may be defined as: ρth = cos (70°)

[0050] This process may be repeated to create and score multiple vanishing points based on the detected line segments Ln = {l1, l2, l3, ... ln} (e.g., in the 2D image) . After multiple vanishing points (vij) have been created (e.g., based on selections from a subset of detected line segments Ln) , a vanishing point (vij) with a highest score may be selected as a vanishing point for an x-axis (vx) of the reference coordinate system.

[0051] The remaining lines in Ln = {l1, l2, l3, ... ln} may be used to repeat the above process to further identify a vanishing point for a y-axis (vy) of the reference coordinate system and a vanishing point for a z-axis (vz) of the reference coordinate system. Thus, after performing vanishing point estimation, at 410, three vanishing points may be identified from the detected line segments Ln, e.g., vx, vy, and vz.

[0052] At 412, process flow 400 proceeds with generating a reference coordinate frame based on a respective direction vector of the lines (e.g., parallel lines) (e.g., a subset of lines from Ln = {l1, l2, l3, ... ln} ) associated with each vanishing point (vx, vy, and vz) .

[0053] For example, the first reference coordinate frame (r1) may be a 3D coordinate system having three axes (e.g., an x-axis, a y-axis, and a z-axis) . The direction vector of the x-axis may be based on the direction vector of the y-axis may be based on and the direction vector of the z-axis may be based on or for example:

[0054] where K is the pre-calibrated intrinsic matrix of an image sensor associated with an agent that captured the 2D image used to generate the particular reference coordinate frame.

[0055] Based on process flow 400, the first reference coordinate frame may be generated at 306 in FIG. 3A. In FIG. 3A, workflow 300 proceeds, at 308, with recording line descriptors and vanishing points for the first reference coordinate frame. For example, line descriptors computed for line segments associated with each vanishing  point in the first reference coordinate frame may be added to a global reference frame map. Further, the vanishing points may also be added to the global reference frame map. In certain aspects, the global reference frame map is stored in memory, such as in memory (ies) 208 of an agent 202 depicted and described with respect to FIG. 2. As shown in FIG. 3A, descriptors (des) of different line segments associated with each vanishing point (vp) of the first reference coordinate frame may be stored in data structures associated with their respective vanishing point.

[0056] Workflow 300 proceeds, at 312, with estimating a rotation matrix (r1Rc1) of a frame of an image sensor, such as a camera (c1) of the agent (also referred to herein as “first camera (c1) frame” ) that captured the 2D image (e.g., at 302) to the first reference coordinate frame. Estimating this rotation matrix (r1Rc1) may estimate an attitude of the agent with respect to the first reference coordinate frame. In certain aspects, the rotation matrix r1Rc1, also referred to in this example as the first rotation matrix, is defined as:

[0057] Workflow 300 then proceeds back to 302 to generate another reference coordinate frame for another 2D image. Thus, at 302, another 2D image of the same 3D space (e.g., a hallway inside an office building) may be captured by an image sensor of another agent (e.g., an image sensor, such as a camera (c2) , of a second agent) in the multi-agent system. In this example, the first agent that produced, at 302, the first 2D image may be located and / or oriented different than the second agent that produced, at 302, the second 2D image. For example, the first agent and the second agent may perceive the 3D space differently. A 2D image of 3D space 108 captured by second image sensor 104(2) of second agent 102 (2) in FIG. 1 may be an example of the 2D image captured at 302 (e.g., in the second iteration ofworkflow 300) .

[0058] Workflow 300 proceeds, at 304, with (1) extracting line segments and (2) determining a line descriptor per extracted line segment in the second 2D image. Further, at 305, workflow 300 proceeds with determining whether at least one reference coordinate frame has been generated based on 2D image (s) captured by agent (s) in the multi-agent system. In this example, because the 2D image captured at 302 is the second image to be captured by an agent in the multi-agent system, a reference coordinate frame may already exist, which in this example may be the first reference coordinate frame generated at 306.  Thus, workflow 300 proceeds to 314 to perform reference coordinate frame re-identification.

[0059] Reference coordinate frame re-identification, at 314, may include matching line descriptors of lines segments for the second 2D image to line descriptor (s) recorded for line segments associated with existing reference coordinate frame (s) , such as the first reference coordinate frame in this example, to identify (1) existing line segment (s) in the second 2D image (e.g., line segment (s) in the second 2D image that match previously recorded line segments associated with existing reference coordinate frame (s) , such as the first reference coordinate frame) and / or (2) new line segment (s) in the second 2D image (e.g., line segment (s) in the second 2D image that do not match any of the previously recorded line segment (s) associated with existing reference coordinate frame (s) , such as the first reference coordinate frame) . In certain aspects, all line segments in the second 2D image, with their corresponding line descriptors, are checked against the recorded line descriptors. In certain other aspects, only a subset (e.g., a maximum threshold percentage, 70%, 60%, etc. ) of the line segments in the second 2D image, with their corresponding line descriptors, are checked against the recorded line descriptors to save processing resources. In certain aspects where only a threshold amount of the line segments of the 2D image are checked, line descriptors of line segments with longer lengths in the 2D image may be checked first. For example, there may be a higher chance that these longer length line segments appear in multiple images (e.g., a higher chance that these long length line segments also appeared in the first 2D image used to create the first reference coordinate frame at 306) . In certain aspects where less than all of the line segments are checked, instead of determining line descriptors for all extracted line segments in the second 2D image at 304, line descriptors for only the subset of the line segments that are checked at 314 (e.g., less than all line segments extracted from the second 2D image) may be determined.

[0060] In this example, it may be determined that (1) a first subset of the line segments in the second 2D image match the line segments previously recorded for the first reference coordinate frame and (2) a second subset of the line segments in the second 2D image do not match the line segments previously recorded for the first reference coordinate frame. Because the first subset of the line segments in the second 2D image match the line segments previously recorded for the first reference coordinate frame, workflow 300 proceeds, at 316, with estimating a rotation matrix (r1Rc2) of a frame of an image sensor,  such as a camera (c2) , associated with the agent (also referred to herein as “second camera (c2) frame” ) that captured the second 2D image (e.g., at 302) to the first reference coordinate frame. Estimating this rotation matrix (r1Rc2) may estimate an attitude of the agent with respect to the first reference coordinate frame. In certain aspects, the rotation matrix r1Rc2, also referred to in this example as the second rotation matrix, is defined as:

[0061] Workflow 300 then proceeds, at 318 in FIG. 3B, with using the new line segments in the second 2D image (e.g., determined at 314 in FIG. 3A) to generate a second reference coordinate frame (r2) (and does not use the existing line segments determined at 314 in FIG. 3A) . Similar to the first reference coordinate frame, generation of the second reference coordination frame, at 318, may also involve using parallel line classification and vanishing point estimation techniques.

[0062] For example, generating the second reference coordinate frame, at 318, may include steps 320, 322, and 324. At 320, the new line segments in the second 2D image (e.g., determined at 314 in FIG. 3A) are used to estimate vanishing points. Each vanishing point may be a point where a subset of the new line segments (e.g., a subset of the second subset of the line segments) in the second 2D image converge. For example, as shown at 320 in FIG. 3B, three vanishing points,  may be estimated based on the new line segments. The first vanishing point,  may correspond to a first direction vector of a first subset of parallel line segments among the new line segments with respect to the second camera frame. The second vanishing point,  may correspond to a second direction vector for a second subset of parallel line segments among the new line segments with respect to the second camera frame. Further, the third vanishing point,  may correspond to a third direction vector for a third subset of parallel line segments among the new line segments with respect to the third camera frame.

[0063] At 322, vanishing points,  may be transformed from the second camera frame to the existing first reference coordinate frame. For example,  may be transformed to the existing first reference coordinate frame by:

[0064] where r1Rc2 is the second rotation matrix between the second camera (c2) frame to the first reference coordinate frame (r1) (e.g., determined at 316 in FIG. 3A) . Further,  may be transformed to the existing first reference coordinate frame by:

[0065] Additionally,  may be transformed to the existing first reference coordinate frame by:

[0066] Here,  may correspond to a first direction vector for the first subset of parallel line segments among the new line segments with respect to the first reference coordinate frame.  may correspond to a second direction vector for the second subset of parallel line segments among the new line segments with respect to the first reference coordinate frame. Further,  may correspond to a third direction vector for the third subset of parallel line segments among the new line segments with respect to the first reference coordinate frame.

[0067] At 324,  and may be used to form the second reference coordinate frame (r2) . For example, the second reference coordinate frame may be a 3D coordinate system having three axes (e.g., an x-axis, a y-axis, and a z-axis) . The direction of the x-axis may be based on the direction of the y-axis may be based on  and the direction of the z-axis may be based on or for example:

[0068] Workflow 300 proceeds, at 326 in FIG. 3C, with recording line descriptors and vanishing points for the second reference coordinate frame. For example, line descriptors computed for line segments associated with each vanishing point in the second reference coordinate frame may be added to the global reference frame map (e.g., already containing line descriptors and vanishing points information for the first reference coordinate frame generated at 306) . Further, the vanishing points associated with the second reference coordinate frame may also be added to the global reference frame map.

[0069] Workflow 300 proceeds, at 328 in FIG. 3C, with estimating a rotation matrix (r2Rc2) of the second camera frame (c2) to the second reference coordinate frame (r2) (e.g., generated at 318 in FIG. 3B, and more specifically, at 324 in FIG. 3B) . Estimating this  rotation matrix (r2Rc2) may estimate an attitude of the second agent with respect to the second reference coordinate frame. In certain aspects, the rotation matrix r2Rc2, also referred to in this example as the third rotation matrix, is defined as:

[0070] The above equation may be used to determine the rotation matrix given line segments used to create the second reference coordinate frame are line segments observed by the second agent (e.g., via a camera associated with the second agent) . In some cases where the line segments associated with a reference coordinate frame are not observed by an agent however, yet a rotation matrix between a camera frame of a camera associated with the agent and the rotation coordinate frame is to be determined, the rotation matrix may instead be determined based on a rotation chain between the reference coordinate frame and another reference coordinate frame in a global reference frame map. For example, the rotation matrix r2Rc2, may be defined as:r2Rc2 = r2Rr1r1Rc2

[0071] where r1Rc2 is the second rotation matrix previously determined at 316 in FIG. 3A and r2Rr1 represents the rotation chain between the second reference coordinate frame and the first reference coordinate frame, which may be determined according to the equations provided below for step 330.

[0072] Workflow 300 then proceeds, at 330 in FIG. 3C, with constructing a rotation chain (r2Rr1) between the first reference coordinate frame (e.g., generated at 306 in FIG. 3A) and the second reference coordinate frame (e.g., generated at 318 in FIG. 3B, and more specifically, at 324 in FIG. 3B) . For example, the rotation chain (r1Rr2) may be determined as:

[0073] where is the three column vector of the rotation chain (r1Rr2) . Further, the rotation chain r2Rr1 may be defined as:

[0074] In certain aspects, after 330, workflow 300 may again return to 302 in FIG. 3A to generate another reference coordinate frame for another 2D image (e.g., which may be captured by a third agent in the multi-agent system) , which may be added to the global reference frame map for attitude estimation.

[0075] However, in certain other aspects, after 330, workflow 300 may proceed, at 332, with attitude estimation. In this example, attitude estimation at 326 may include estimating an attitude (c2Rc1) of the first agent with respect to the second agent (or vice versa) based on the first rotation matrix r1Rc1 (e.g., determined at 312 in FIG. 3A) , the third rotation matrix r2Rc2 (e.g., determined at 328 in FIG. 3C) , and the rotation chain r2Rr1 between the first reference coordinate frame and the second reference coordinate frame. For example, an attitude of the first agent with respect to the second agent may be based on a rotation chain defined as:c2Rc1 = r2Rc2r2Rr1r1Rc1

[0076] As more reference coordinate frames are generated (along with the computation of rotation matrices and rotation chains, per workflow 300) , the global reference frame map is updated to allow for the realization of rotation among multiple agents in the multi-agent system. The rotation may represent the estimated attitude of one agent relative to another agent within the multi-agent system.

[0077] Example Method for Relative Attitude Estimation

[0078] FIG. 5 shows a method 500 for attitude estimation by an apparatus. For example, the apparatus may estimate the attitude between a first device (e.g., an example of first agent 102 (1) depicted and described with respect to FIG. 1) and a second device (e.g., an example of second agent 102 (2) depicted and described with respect to FIG. 1) in a multi-agent system. In certain aspects, the apparatus is the first device and thus, performs method 500 to estimate an attitude of the apparatus with respect to the second device. In certain aspects, the apparatus is a third device in the multi-agent system and thus, performs method 500 to estimate an attitude of other independent devices in the multi-agent system.

[0079] Method 500 begins at block 502 with obtaining a first reference coordinate frame based on a plurality of first line segments extracted from a first 2D image, of a 3D space, captured by a first device.

[0080] Method 500 then proceeds to block 504 with obtaining a first rotation matrix of a frame of a first image sensors of the first device to the first reference coordinate frame.

[0081] Method 500 then proceeds to block 506 with obtaining a second reference coordinate frame based on a plurality of second line segments extracted from a second 2D image, of the 3D space, captured by a second device.

[0082] Method 500 then proceeds to block 508 with obtaining a second rotation matrix of a frame of a second image sensor of the second device to the second reference coordinate frame.

[0083] Method 500 then proceeds to block 510 with estimating an attitude of a first device with respect to a second device based on the first rotation matrix, the second rotation matrix, and a rotation chain.

[0084] In one aspect, the rotation chain is constructed between the first reference coordinate frame and the second reference coordinate frame.

[0085] In one aspect, the first device is the apparatus. The apparatus may include the first image sensor configured to capture the first 2D image. In one aspect where the first device is the apparatus, method 500 further includes rendering content based on the estimated attitude of the first device with respect to the second device and sending the content to the second device.

[0086] In one aspect, method 500 further includes rendering content based on the estimated attitude of the first device; and sending the content to the first device.

[0087] In one aspect, obtaining the first reference coordinate frame at block 502 includes extracting the plurality of first line segments from the first 2D image; estimating a plurality of first vanishing points based on the plurality of first line segments, wherein each of the plurality of first vanishing points comprises a respective point where a respective subset of the plurality of first line segments converge; and generating the first reference coordinate frame based on, for each of the plurality of first vanishing points, a respective first direction vector for the respective subset of the plurality of first line segments.

[0088] In one aspect, obtaining the first rotation matrix at block 504 includes computing the first rotation matrix based on the plurality of first vanishing points and a first pre-calibrated intrinsic matrix of the first image sensor, at the first device, used to capture the first 2D image.

[0089] In one aspect, method 500 further includes computing a first descriptor for each first line segment in each subset of the plurality of first line segments; adding the first descriptor computed for each first line segment to a global reference frame map for the first reference coordinate frame; and adding the plurality of first vanishing points to the global reference frame map for the first reference coordinate frame.

[0090] In one aspect, obtaining the second reference coordinate frame at block 506 includes extracting the plurality of second line segments from the second 2D image; determining a first subset of the plurality of second line segments to use to generate the second reference coordinate frame; estimating a plurality of second vanishing points based on the first subset of the plurality of second line segments, wherein each of the plurality of second vanishing points comprises a respective point where a respective subset of the first subset of the plurality of second line segments converge; for each second vanishing point of the plurality of second vanishing points: determine a second direction vector for the respective subset of the first subset of the plurality of second line segments associated with the second vanishing point, wherein the second direction vector is determined for the frame of the second image sensor; and transform the second direction vector determined for the frame of the second image sensor to a third direction vector in the first reference coordinate frame; and generating the second reference coordinate frame based on the third direction vector associated with each of the plurality of second vanishing points.

[0091] In one aspect, determining the first subset of the plurality of second line segments to use to generate the second reference coordinate frame includes, for each second line segment of one or more of the plurality of second line segments: computing a second descriptor for the second line segment; determining whether the second descriptor matches one of the first descriptors added to the global reference frame map; adding the second line segment to the first subset of the plurality of second line segments when the second descriptor does not match any of the first descriptors added to the global reference frame map; and adding the second line segment to a second subset of the plurality of second line segments when the second descriptor does match at least one of the first descriptors added to the global reference frame map.

[0092] In one aspect, method 500 further includes computing a third rotation matrix of the frame of the second image sensor of the second device to the first reference coordinate frame, wherein transforming the second direction vector to the third direction  vector includes transforming the second direction vector to the third direction vector based on the third rotation matrix.

[0093] In one aspect, computing the third rotation matrix of the frame of the second image sensor of the second device to the first reference coordinate frame is based on the second descriptor for each of the second subset of the plurality of second line segments matching at least one of the first descriptors.

[0094] In one aspect, method 500 further includes computing a second descriptor for each second line segment in each subset of the first subset of the plurality of second line segments; adding the second descriptor computed for each second line segment in each subset of the first subset of the plurality of second line segments to the global reference frame map; and adding the plurality of second vanishing points to the global reference frame map for the second reference coordinate frame.

[0095] In one aspect, obtaining the second rotation matrix at block 508 includes computing the second rotation matrix based on the third direction vector associated with each of the plurality of second vanishing points.

[0096] In one aspect, obtaining the rotation chain constructed between the first reference coordinate frame and the second reference coordinate frame includes computing the rotation chain based on the third direction vector associated with each of the plurality of second vanishing points.

[0097] In one aspect, method 500, or any aspect related to it, may be performed by an agent, such as agent 202 of FIG. 2, which includes various components operable, configured, or adapted to perform the method 500.

[0098] Note that FIG. 5 is just one example of a method, and other methods including fewer, additional, or alternative operations are possible consistent with this disclosure.

[0099] Example Clauses

[0100] Implementation examples are described in the following numbered clauses:

[0101] Clause 1: A method of attitude estimation performed by an apparatus, comprising: obtaining a first reference coordinate frame based on a plurality of first line segments extracted from a first two-dimensional (2D) image, of a three-dimensional (3D) space, captured by a first device; obtaining a first rotation matrix of a frame of a first image sensor of the first device to the first reference coordinate frame; obtaining a second  reference coordinate frame based on a plurality of second line segments extracted from a second 2D image, of the 3D space, captured by a second device; obtaining a second rotation matrix of a frame of a second image sensor of the second device to the second reference coordinate frame; obtaining a rotation chain constructed between the first reference coordinate frame and the second reference coordinate frame; and estimating an attitude of the first device with respect to the second device based on the first rotation matrix, the second rotation matrix, and the rotation chain.

[0102] Clause 2: The method of clause 1, wherein the first device is the apparatus, and further comprising the first image sensor configured to capture the first 2D image.

[0103] Clause 3: The method of clause 2, further comprising: rendering content based on the estimated attitude of the first device with respect to the second device; and sending the content to the second device.

[0104] Clause 4: The method of any clause 1, further comprising: rendering content based on the estimated attitude of the first device; and sending the content to the first device.

[0105] Clause 5: The method of any one of clauses 1-4, wherein obtaining the first reference coordinate frame comprises: extracting the plurality of first line segments from the first 2D image; estimating a plurality of first vanishing points based on the plurality of first line segments, wherein each of the plurality of first vanishing points comprises a respective point where a respective subset of the plurality of first line segments converge; and generating the first reference coordinate frame based on, for each of the plurality of first vanishing points, a respective first direction vector for the respective subset of the plurality of first line segments.

[0106] Clause 6: The method of clause 5, wherein obtaining the first rotation matrix, comprises: computing the first rotation matrix based on the plurality of first vanishing points and a first pre-calibrated intrinsic matrix of the first image sensor, at the first device, used to capture the first 2D image.

[0107] Clause 7: The method of any one of clauses 5-6, further comprising: computing a first descriptor for each first line segment in each subset of the plurality of first line segments; adding the first descriptor computed for each first line segment to a global reference frame map for the first reference coordinate frame; and adding the  plurality of first vanishing points to the global reference frame map for the first reference coordinate frame.

[0108] Clause 8: The method of clause 7, wherein obtaining the second reference coordinate frame comprises: extracting the plurality of second line segments from the second 2D image; determining a first subset of the plurality of second line segments to use to generate the second reference coordinate frame; estimating a plurality of second vanishing points based on the first subset of the plurality of second line segments, wherein each of the plurality of second vanishing points comprises a respective point where a respective subset of the first subset of the plurality of second line segments converge; for each second vanishing point of the plurality of second vanishing points: determining a second direction vector for the respective subset of the first subset of the plurality of second line segments associated with the second vanishing point, wherein the second direction vector is determined for the frame of the second image sensor; and transforming the second direction vector determined for the frame of the second image sensor to a third direction vector in the first reference coordinate frame; and generating the second reference coordinate frame based on the third direction vector associated with each of the plurality of second vanishing points.

[0109] Clause 9: The method of clause 8, wherein determining the first subset of the plurality of second line segments to use to generate the second reference coordinate frame comprises: for each second line segment of one or more of the plurality of second line segments: computing a second descriptor for the second line segment; determining whether the second descriptor matches one of the first descriptors added to the global reference frame map; adding the second line segment to the first subset of the plurality of second line segments when the second descriptor does not match any of the first descriptors added to the global reference frame map; and adding the second line segment to a second subset of the plurality of second line segments when the second descriptor does match at least one of the first descriptors added to the global reference frame map.

[0110] Clause 10: The method of clause 9, further comprising computing a third rotation matrix of the frame of the second image sensor of the second device to the first reference coordinate frame, wherein transforming the second direction vector to the third direction vector comprises transforming the second direction vector to the third direction vector based on the third rotation matrix.

[0111] Clause 11: The method of clause 10, wherein computing the third rotation matrix comprises computing the third rotation matrix of the frame of the second image sensor of the second device to the first reference coordinate frame based on the second descriptor for each of the second subset of the plurality of second line segments matching at least one of the first descriptors.

[0112] Clause 12: The method of any one of clauses 9-11, further comprising: computing a second descriptor for each second line segment in each subset of the first subset of the plurality of second line segments; adding the second descriptor computed for each second line segment in each subset of the first subset of the plurality of second line segments to the global reference frame map; and adding the plurality of second vanishing points to the global reference frame map for the second reference coordinate frame.

[0113] Clause 13: The method of any one of clauses 8-12, wherein obtaining the second rotation matrix comprises computing the second rotation matrix based on the third direction vector associated with each of the plurality of second vanishing points.

[0114] Clause 14: The method of any one of clauses 8-13, wherein obtaining the rotation chain constructed between the first reference coordinate frame and the second reference coordinate frame comprises computing the rotation chain based on the third direction vector associated with each of the plurality of second vanishing points.

[0115] Clause 15: The method of any one of clauses 1-14, wherein the rotation chain is constructed between the first reference coordinate frame and the second reference coordinate frame.

[0116] Clause 16: One or more apparatuses, comprising: one or more memories comprising executable instructions; and one or more processors configured to execute the executable instructions and cause the one or more apparatuses to perform a method in accordance with any one of clauses 1-15.

[0117] Clause 17: One or more apparatuses, comprising: one or more memories; and one or more processors, coupled to the one or more memories, configured to cause the one or more apparatuses to perform a method in accordance with any one of Clauses 1-15.

[0118] Clause 18: One or more apparatuses, comprising: one or more memories; and one or more processors, coupled to the one or more memories, configured to perform a method in accordance with any one of Clauses 1-15.

[0119] Clause 19: One or more apparatuses, comprising means for performing a method in accordance with any one of Clauses 1-15.

[0120] Clause 20: One or more non-transitory computer-readable media comprising executable instructions that, when executed by one or more processors of one or more apparatuses, cause the one or more apparatuses to perform a method in accordance with any one of Clauses 1-15.

[0121] Clause 21: One or more computer program products embodied on one or more computer-readable storage media comprising code for performing a method in accordance with any one of Clauses 1-15.

[0122] Additional Considerations

[0123] The preceding description is provided to enable any person skilled in the art to practice the various aspects described herein. The examples discussed herein are not limiting of the scope, applicability, or aspects set forth in the claims. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. For example, changes may be made in the function and arrangement of elements discussed without departing from the scope of the disclosure. Various examples may omit, substitute, or add various procedures or components as appropriate. For instance, the methods described may be performed in an order different from that described, and various actions may be added, omitted, or combined. Also, features described with respect to some examples may be combined in some other examples. For example, an apparatus may be implemented or a method may be practiced using any number of the aspects set forth herein. In addition, the scope of the disclosure is intended to cover such an apparatus or method that is practiced using other structure, functionality, or structure and functionality in addition to, or other than, the various aspects of the disclosure set forth herein. It should be understood that any aspect of the disclosure disclosed herein may be embodied by one or more elements of a claim.

[0124] The various illustrative logical blocks, modules and circuits described in connection with the present disclosure may be implemented or performed with a general  purpose processor, an AI processor, a digital signal processor (DSP) , an ASIC, a field programmable gate array (FPGA) or other programmable logic device (PLD) , discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, a system on a chip (SoC) , or any other such configuration.

[0125] As used herein, a phrase referring to “at least one of” a list of items refers to any combination of those items, including single members. As an example, “at least one of: a, b, or c” is intended to cover a, b, c, a-b, a-c, b-c, and a-b-c, as well as any combination with multiples of the same element (e.g., a-a, a-a-a, a-a-b, a-a-c, a-b-b, a-c-c, b-b, b-b-b, b-b-c, c-c, and c-c-c or any other ordering of a, b, and c) .

[0126] As used herein, the term “determining” encompasses a wide variety of actions. For example, “determining” may include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, a database or another data structure) , ascertaining and the like. Also, “determining” may include receiving (e.g., receiving information) , accessing (e.g., accessing data in a memory) and the like. Also, “determining” may include resolving, selecting, choosing, establishing and the like.

[0127] As used herein, “coupled to” and “coupled with” generally encompass direct coupling and indirect coupling (e.g., including intermediary coupled aspects) unless stated otherwise. For example, stating that a processor is coupled to a memory allows for a direct coupling or a coupling via an intermediary aspect, such as a bus.

[0128] The methods disclosed herein comprise one or more actions for achieving the methods. The method actions may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of actions is specified, the order and / or use of specific actions may be modified without departing from the scope of the claims. Further, the various operations of methods described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software component (s) and / or module (s) ,  including, but not limited to a circuit, an application specific integrated circuit (ASIC) , or processor.

[0129] The following claims are not intended to be limited to the aspects shown herein, but are to be accorded the full scope consistent with the language of the claims. Reference to an element in the singular is not intended to mean only one unless specifically so stated, but rather “one or more. ” The subsequent use of a definite article (e.g., “the” or “said” ) with an element (e.g., “the processor” ) is not intended to invoke a singular meaning (e.g., “only one” ) on the element unless otherwise specifically stated. For example, reference to an element (e.g., “a processor, ” “a controller, ” “a memory, ” “a transceiver, ” “an antenna, ” “the processor, ” “the controller, ” “the memory, ” “the transceiver, ” “the antenna, ” etc. ) , unless otherwise specifically stated, should be understood to refer to one or more elements (e.g., “one or more processors, ” “one or more controllers, ” “one or more memories, ” “one more transceivers, ” etc. ) . The terms “set” and “group” are intended to include one or more elements, and may be used interchangeably with “one or more. ” Where reference is made to one or more elements performing functions (e.g., steps of a method) , one element may perform all functions, or more than one element may collectively perform the functions. When more than one element collectively performs the functions, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and / or each function need not be performed in whole by only one element (e.g., different elements may perform different sub-functions of a function) . Similarly, where reference is made to one or more elements configured to cause another element (e.g., an apparatus) to perform functions, one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions. Unless specifically stated otherwise, the term “some” refers to one or more. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims.

Claims

1.An apparatus comprising:one or more memories; andone or more processors, coupled to the one or more memories, configured to cause the apparatus to:obtain a first reference coordinate frame based on a plurality of first line segments extracted from a first two-dimensional (2D) image, of a three-dimensional (3D) space, captured by a first device;obtain a first rotation matrix of a frame of a first image sensor of the first device to the first reference coordinate frame;obtain a second reference coordinate frame based on a plurality of second line segments extracted from a second 2D image, of the 3D space, captured by a second device;obtain a second rotation matrix of a frame of a second image sensor of the second device to the second reference coordinate frame; andestimate an attitude of the first device with respect to the second device based on the first rotation matrix, the second rotation matrix, and a rotation chain.2.The apparatus of claim 1, wherein the rotation chain is constructed between the first reference coordinate frame and the second reference coordinate frame.3.The apparatus of claim 1, wherein the first device is the apparatus, and further comprising the first image sensor configured to capture the first 2D image.4.The apparatus of claim 3, wherein the one or more processors are configured to cause the apparatus to:render content based on the estimated attitude of the first device with respect to the second device; andsend the content to the second device.5.The apparatus of claim 1, wherein the one or more processors are configured to cause the apparatus to:render content based on the estimated attitude of the first device; andsend the content to the first device.6.The apparatus of claim 1, wherein to obtain the first reference coordinate frame, the one or more processors are configured to cause the apparatus to:extract the plurality of first line segments from the first 2D image;estimate a plurality of first vanishing points based on the plurality of first line segments, wherein each of the plurality of first vanishing points comprises a respective point where a respective subset of the plurality of first line segments converge; andgenerate the first reference coordinate frame based on, for each of the plurality of first vanishing points, a respective first direction vector for the respective subset of the plurality of first line segments.7.The apparatus of claim 6, wherein to obtain the first rotation matrix, the one or more processors are configured to cause the apparatus to:compute the first rotation matrix based on the plurality of first vanishing points and a first pre-calibrated intrinsic matrix of the first image sensor, at the first device, used to capture the first 2D image.8.The apparatus of claim 6, wherein the one or more processors are configured to cause the apparatus to:compute a first descriptor for each first line segment in each subset of the plurality of first line segments;add the first descriptor computed for each first line segment to a global reference frame map for the first reference coordinate frame; andadd the plurality of first vanishing points to the global reference frame map for the first reference coordinate frame.9.The apparatus of claim 8, wherein to obtain the second reference coordinate frame, the one or more processors are configured to cause the apparatus to:extract the plurality of second line segments from the second 2D image;determine a first subset of the plurality of second line segments to use to generate the second reference coordinate frame;estimate a plurality of second vanishing points based on the first subset of the plurality of second line segments, wherein each of the plurality of second vanishing points comprises a respective point where a respective subset of the first subset of the plurality of second line segments converge;for each second vanishing point of the plurality of second vanishing points:determine a second direction vector for the respective subset of the first subset of the plurality of second line segments associated with the second vanishing point, wherein the second direction vector is determined for the frame of the second image sensor; andtransform the second direction vector determined for the frame of the second image sensor to a third direction vector in the first reference coordinate frame; andgenerate the second reference coordinate frame based on the third direction vector associated with each of the plurality of second vanishing points.10.The apparatus of claim 9, wherein to determine the first subset of the plurality of second line segments to use to generate the second reference coordinate frame, the one or more processors are configured to cause the apparatus to:for each second line segment of one or more of the plurality of second line segments:compute a second descriptor for the second line segment;determine whether the second descriptor matches one of the first descriptors added to the global reference frame map;add the second line segment to the first subset of the plurality of second line segments when the second descriptor does not match any of the first descriptors added to the global reference frame map; andadd the second line segment to a second subset of the plurality of second line segments when the second descriptor does match at least one of the first descriptors added to the global reference frame map.11.The apparatus of claim 10, wherein:the one or more processors are configured to cause the apparatus to compute a third rotation matrix of the frame of the second image sensor of the second device to the first reference coordinate frame, andto transform the second direction vector to the third direction vector, the one or more processors are configured cause the apparatus to transform the second direction vector to the third direction vector based on the third rotation matrix.12.The apparatus of claim 11, wherein to compute the third rotation matrix, the one or more processors are configured to cause the apparatus to compute the third rotation matrix of the frame of the second image sensor of the second device to the first reference coordinate frame based on the second descriptor for each of the second subset of the plurality of second line segments matching at least one of the first descriptors.13.The apparatus of claim 10, wherein the one or more processors are configured to cause the apparatus to:compute a second descriptor for each second line segment in each subset of the first subset of the plurality of second line segments;add the second descriptor computed for each second line segment in each subset of the first subset of the plurality of second line segments to the global reference frame map; andadd the plurality of second vanishing points to the global reference frame map for the second reference coordinate frame.14.The apparatus of claim 9, wherein to obtain the second rotation matrix, the one or more processors are configured to cause the apparatus to compute the second rotation matrix based on the third direction vector associated with each of the plurality of second vanishing points.15.The apparatus of claim 9, wherein to obtain the rotation chain constructed between the first reference coordinate frame and the second reference coordinate frame, the one or more processors are configured to cause the apparatus to compute the rotation chain based on the third direction vector associated with each of the plurality of second vanishing points.16.A method by an apparatus, comprising:obtaining a first reference coordinate frame based on a plurality of first line segments extracted from a first two-dimensional (2D) image, of a three-dimensional (3D) space, captured by a first device;obtaining a first rotation matrix of a frame of a first image sensor of the first device to the first reference coordinate frame;obtaining a second reference coordinate frame based on a plurality of second line segments extracted from a second 2D image, of the 3D space, captured by a second device;obtaining a second rotation matrix of a frame of a second image sensor of the second device to the second reference coordinate frame; andestimating an attitude of the first device with respect to the second device based on the first rotation matrix, the second rotation matrix, and a rotation chain.17.The method of claim 16, wherein the rotation chain is constructed between the first reference coordinate frame and the second reference coordinate frame.18.The method of claim 16, wherein the first device is the apparatus, and further comprising the first image sensor configured to capture the first 2D image.19.One or more non-transitory computer-readable media comprising executable instructions that, when executed by one or more processors of an apparatus, cause the apparatus to perform operations comprising:obtaining a first reference coordinate frame based on a plurality of first line segments extracted from a first two-dimensional (2D) image, of a three-dimensional (3D) space, captured by a first device;obtaining a first rotation matrix of a frame of a first image sensor of the first device to the first reference coordinate frame;obtaining a second reference coordinate frame based on a plurality of second line segments extracted from a second 2D image, of the 3D space, captured by a second device;obtaining a second rotation matrix of a frame of a second image sensor of the second device to the second reference coordinate frame; andestimating an attitude of the first device with respect to the second device based on the first rotation matrix, the second rotation matrix, and a rotation chain.20.The one or more non-transitory computer-readable media of claim 19, wherein the rotation chain is constructed between the first reference coordinate frame and the second reference coordinate frame.