Method, data processing system, and non-transitory machine-readable medium for determining a pose of a first wide-angle image
By dedistorting and feature matching of wide-angle images, the unreliable pose positioning caused by wide-angle image distortion is solved, and more accurate and stable AR device pose positioning and virtual object display are achieved.
Patent Information
- Application Number
- CN202410328157.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-10
- Filing Date
- 2021-05-27
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2041-05-27
AI Technical Summary
When using wide-angle images, the image feature matching caused by distortion is unreliable, which affects the device's posture positioning and the accurate display of virtual objects.
By dedistorting the area of the wide-angle image, image features are detected and matched, triangulation is performed with point cloud features, estimated poses are determined and image poses are optimized, and the initial pose positioning and tracking of the equipment is improved.
Improves the pose positioning accuracy and stability of wide-angle images, ensuring that virtual objects are displayed more accurately and stably on display devices.
Smart Images

Figure CN118247348B_ABST
Abstract
Description
[0001] This application is a divisional application of the application with national application number 202180049099.2, international application date May 27, 2021, national entry date January 9, 2023, and invention name “3D reconstruction using wide-angle imaging equipment”.
[0002] Priority claim
[0003] This application claims the benefit of priority to U.S. patent application No. 16 / 897,749, filed on June 10, 2020, which is incorporated herein by reference in its entirety. Technical Field
[0004] Implementations of the present disclosure relate generally to mobile computing technology, and more particularly, but not by way of limitation, to a system for presenting augmented-reality (AR) content at a client device. Background Art
[0005] Augmented reality is an interactive experience in which objects present in the real world are enhanced by computer-generated sensory information, sometimes across multiple sensory modalities, including vision, hearing, touch, body sensation, and smell. The primary value of augmented reality is the way in which components of the digital world are integrated into a person's perception of the real world, not as a simple display of data, but through the integration of immersion where they are perceived as a natural part of the environment.
[0006] AR systems can utilize virtual 3D models to locate and track the user's AR device. 3D reconstruction is a technique for inferring the geometry of a scene captured by a collection of images. Summary of the Invention
[0007] The pose of the wide-angle image is determined by dedistorting a region of the wide-angle image (602), determining an estimated pose of the dedistorted region of the wide-angle image (604, 614), and deriving the pose of the wide-angle image from the estimated pose of the dedistorted region (616, 618). The estimated pose of the dedistorted region can be determined by comparing features in the dedistorted region with features in previous dedistorted regions from one or more previous wide-angle images and by comparing features in the dedistorted region with features in the point cloud.
[0008] The present disclosure provides a method for determining the pose of a first wide-angle image using one or more processors, comprising: dedistorting an area of the first wide-angle image; dedistorting an area of the second wide-angle image; detecting image features in the dedistorted area of the first wide-angle image and image features in the dedistorted area of the second wide-angle image; identifying a pair of matching image features between the dedistorted area of the first wide-angle image and the dedistorted area of the second wide-angle image; determining a relative pose between the dedistorted area of the first wide-angle image and the dedistorted area of the second wide-angle image having the pair of matching image features; triangulating the pair of matching image features to generate 3D point positions of the matching image features; determining an estimated pose of the dedistorted area of the first wide-angle image based on the 3D point positions of the matching image features; and deriving the pose of the first wide-angle image from the estimated pose of the dedistorted area of the first wide-angle image.
[0009] The present disclosure provides a data processing system comprising: one or more processors; a wide-angle image capture device, and one or more machine-readable media storing instructions that, when executed by the one or more processors, cause the system to perform operations, the operations comprising: dedistorting a region of a first wide-angle image; dedistorting a region of a second wide-angle image; detecting image features in the dedistorted region of the first wide-angle image and image features in the dedistorted region of the second wide-angle image; identifying pairs of matching image features between the dedistorted region of the first wide-angle image and the dedistorted region of the second wide-angle image; determining a relative pose between the dedistorted region of the first wide-angle image and the dedistorted region of the second wide-angle image having the pairs of matching image features; triangulating the pairs of matching image features to generate 3D point positions of the matching image features; determining an estimated pose of the dedistorted region of the first wide-angle image based on the 3D point positions of the matching image features; and deriving the pose of the first wide-angle image from the estimated pose of the dedistorted region of the first wide-angle image.
[0010] The present disclosure provides a non-transitory machine-readable medium comprising instructions, which, when read by a machine, causes the machine to perform operations for determining a pose of a first wide-angle image, the operations comprising: dedistorting a region of the first wide-angle image; dedistorting a region of the second wide-angle image; detecting image features in the dedistorted region of the first wide-angle image and image features in the dedistorted region of the second wide-angle image; identifying a pair of matching image features between the dedistorted region of the first wide-angle image and the dedistorted region of the second wide-angle image; determining a relative pose between the dedistorted region of the first wide-angle image and the dedistorted region of the second wide-angle image having the pair of matching image features; triangulating the pair of matching image features to generate 3D point positions of the matching image features; determining an estimated pose of the dedistorted region of the first wide-angle image based on the 3D point positions of the matching image features; and deriving the pose of the first wide-angle image from the estimated pose of the dedistorted region of the first wide-angle image. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To easily identify the discussion of any particular element or act, the highest digit or digits in a reference number refer to the figure number in which the element is first introduced.
[0012] Figure 1 is a block diagram illustrating an example messaging system for exchanging data (eg, messages and associated content) over a network, wherein the messaging system includes an augmented reality system, in accordance with some implementations.
[0013] Figure 2 is a block diagram showing additional details regarding a messaging system according to an example embodiment.
[0014] Figure 3 is a block diagram illustrating various modules of an augmented reality system according to certain example embodiments.
[0015] Figure 4 Dewarping a wide-angle image into multiple linearly corrected images is shown.
[0016] Figure 5A 、 Figure 5B and Figure 5C The processing of a viewpoint sequence with reference to a 3D point cloud is schematically shown.
[0017] Figure 6 is a flowchart illustrating a 3D reconstruction method according to one example.
[0018] Figure 7 An interface flow chart according to an example is shown.
[0019] Figure 8is a block diagram illustrating a representative software architecture that may be used in conjunction with the various hardware architectures described herein and that may be used to implement the various embodiments.
[0020] Figure 9 is a block diagram illustrating components of a machine capable of reading instructions from a machine-readable medium (eg, a machine-readable storage medium) and performing any one or more of the methodologies discussed herein, according to some example embodiments. DETAILED DESCRIPTION
[0021] As mentioned above, augmented reality (AR) is an interactive experience in which objects present in the real world are enhanced by computer-generated sensory information. Some AR systems use point clouds to generate and present AR content, where a point cloud is a group of data points corresponding to the features and / or external surfaces of objects in the real world.
[0022] Structure from Motion (SfM) is a technique for estimating 3D structure from a sequence of 2D images. The output is the pose (i.e., the 6-DOF position and orientation of the device that captured the image) of each image and a 3D point cloud, where each point in the point cloud has a description of its appearance in two or more images. The general scheme is:
[0023] - Detect image features in each image, such as corner points.
[0024] - Match features between images.
[0025] -Estimating relative pose of image pairs with feature matching.
[0026] Starting from an initial image pair, image features are triangulated to estimate their 3D positions.
[0027] - Iteratively register new images by matching image features against the current 3D point cloud.
[0028] -Triangulate the new image to add new points to the point cloud and optimize the image pose and 3D point positions.
[0029] When using a wide-angle image source such as a fisheye or 360-degree camera, this processing is unreliable due to the distortion of image features. The disclosure herein seeks to mitigate this unreliability by dewarping a region of a wide-angle image received from such a wide-angle source, and in some examples estimating the pose of the dedistorted region using SfM techniques. The estimated pose of the dedistorted region can then be used to derive the pose of the wide-angle image, and thereby the pose of the device that includes the wide-angle image source. As used herein, the term wide-angle is intended to cover the output of any imaging device that includes intentional distortion.
[0030] Improvements to the initial pose for positioning a device, as well as improvements to tracking of the device after positioning, allow for more precise and / or stable positioning of virtual objects (or other augmented information) within an image or image stream to be displayed on a display device. Thus, the methods and systems described herein improve the functionality of devices or systems that include augmented reality functionality or otherwise utilize 3D reconstruction.
[0031] Thus, in certain example embodiments, a method for determining a pose of a wide-angle image using one or more processors is provided, the method comprising: dedistorting a region of the wide-angle image; determining an estimated pose of the dedistorted region of the wide-angle image; and deriving the pose of the wide-angle image from the estimated pose of the dedistorted region. Determining the estimated pose of the dedistorted region may include comparing features in the dedistorted region with features in other regions from one or more other wide-angle images that have already been dedistorted. Determining the estimated pose may also include comparing features in the dedistorted region with features in a point cloud. Furthermore, the pose of the wide-angle image may be derived from an average of at least some of the estimated poses.
[0032] In some example embodiments, determining the estimated pose includes determining a group of dedistorted regions with consistency between the estimated poses, and deriving the pose of the wide-angle image from the estimated poses of the dedistorted regions with such consistency. The consistency may be determined with reference to sensor data selected from the group consisting of motion sensor data and position sensor data. The poses of the dedistorted regions that are not in the group of dedistorted regions with consistency between the estimated poses may then be derived from the pose of the wide-angle image.
[0033] In some example embodiments, the method further comprises optimizing the estimated pose of the dedistorted region and then deriving an updated pose of the wide-angle image from at least some of the optimized estimated poses. The pose of the wide-angle image may then be optimized after optimizing the estimated pose of the dedistorted region.
[0034] In some example embodiments, a data processing system is provided that includes one or more processors, a wide-angle image capture device, and one or more machine-readable media storing instructions that, when executed by the one or more processors, cause the system to perform the operations described above in paragraphs
[0018] to
[0020] , including but not limited to: receiving a wide-angle image from the wide-angle image capture device, dedistorting a region of the wide-angle image, determining an estimated pose of the dedistorted region of the wide-angle image, and deriving a pose of the wide-angle image from the estimated pose of the dedistorted region.
[0035] In some example embodiments, a non-transitory machine-readable medium is provided, comprising instructions that, when read by a machine, cause the machine to perform the operations described above in paragraphs
[0018] to
[0020] , including but not limited to: receiving a wide-angle image from an image capture device, dedistorting a region of the wide-angle image, determining an estimated pose of the dedistorted region of the wide-angle image, and deriving a pose of the wide-angle image from the estimated pose of the dedistorted region.
[0036] Figure 1 1 is a block diagram illustrating an example messaging system 100 for exchanging data (e.g., messages and associated content) over a network. Messaging system 100 includes one or more client devices 102 hosting multiple applications including messaging client applications 104. Each messaging client application 104 is communicatively coupled to other instances of messaging client applications 104 and a messaging server system 108 via a network 106 (e.g., the Internet).
[0037] Thus, each messaging client application 104 is able to communicate and exchange data with another messaging client application 104 and with the messaging server system 108 via the network 106. The data exchanged between the messaging client applications 104 and between the messaging client applications 104 and the messaging server system 108 includes functions (e.g., commands to invoke functions) and payload data (e.g., text, audio, video, or other multimedia data).
[0038] The messaging server system 108 provides server-side functionality to specific messaging client applications 104 via the network 106. Although certain functionality of the messaging system 100 is described herein as being performed by either the messaging client application 104 or the messaging server system 108, it should be understood that the location of certain functionality within the messaging client application 104 or within the messaging server system 108 is a design choice. For example, it may be technically preferable to initially deploy certain technologies and functionality within the messaging server system 108 but later migrate such technologies and functionality to the messaging client application 104 where the client device 102 has sufficient processing power.
[0039] The messaging server system 108 supports various services and operations provided to the messaging client applications 104. Such operations include sending data to the messaging client applications 104, receiving data from the messaging client applications 104, and processing data generated by the messaging client applications 104. In some embodiments, this data includes, by way of example, message content, client device information, geolocation information, media annotations and overlays, message content persistence conditions, social network information, and live event information. In other embodiments, other data is used. Data exchange within the messaging system 100 is invoked and controlled through functionality available via the GUI of the messaging client applications 104.
[0040] Specifically, turning now to messaging server system 108, an application program interface (API) server 110 is coupled to and provides a programming interface to application server 112. Application server 112 is communicatively coupled to database server 118, which facilitates access to a database 120 having stored therein data associated with messages processed by application server 112.
[0041] Specific discussion is given of an application program interface (API) server 110 that receives and sends message data (e.g., commands and message payloads) between a client device 102 and an application server 112. In particular, the application program interface (API) server 110 provides a set of interfaces (e.g., routines and protocols) that the messaging client application 104 can call or query to invoke functionality of the application server 112. The application program interface (API) server 110 exposes various functions supported by the application server 112, including: account registration, login functionality, sending messages from a particular messaging client application 104 to another messaging client application 104 via the application server 112, sending media files (e.g., images or videos) from a messaging client application 104 to a messaging server application 114, and setting up a collection of media data (e.g., a story) for possible access by another messaging client application 104, retrieving a friend list of a user of a client device 102, retrieving such a collection, retrieving messages and content, adding friends to and removing friends from a social graph, locating friends within a social graph, and opening and applying events (e.g., related to a messaging client application 104).
[0042] The application server 112 hosts a number of applications and subsystems, including a messaging server application 114, an image processing system 116, a social networking system 122, and an AR system 124. According to certain example embodiments, the AR system 124 is configured to generate and / or host a point cloud based on image data. In some embodiments, the AR system 124 is shown as being located in the application server 112, but the AR system can also be partially or fully hosted on the client device 102. Further details of the AR system 124 can be found below. Figure 3 Found in.
[0043] The messaging server application 114 implements a number of message processing techniques and functions, particularly relating to the aggregation and other processing of content (e.g., text and multimedia content) included in messages received from multiple instances of the messaging client application 104. As will be described in further detail, text and media content from multiple sources can be aggregated into content collections (e.g., referred to as stories or galleries). The messaging server application 114 then makes these collections available to the messaging client application 104. Given the hardware requirements for other processor- and memory-intensive processing of data, such processing can also be performed on the server side by the messaging server application 114.
[0044] The application server 112 also includes an image processing system 116 that is dedicated to performing various image processing operations, typically on images or videos received within the payload of messages at the messaging server application 114 .
[0045] The social networking system 122 supports various social networking functionalities and services and makes these functionalities and services available to the messaging server application 114. Examples of functionalities and services supported by the social networking system 122 include identifying other users of the messaging system 100 who have relationships with or are "following" a particular user, and also identifying interests and other entities of a particular user.
[0046] The application server 112 is communicatively coupled to a database server 118 , which facilitates access to a database 120 in which data associated with messages processed by the messaging server application 114 is stored.
[0047] Figure 2 is a block diagram illustrating additional details regarding messaging system 100 according to an example embodiment. Specifically, messaging system 100 is shown as including messaging client application 104 and application server 112, which in turn contain several subsystems, namely, transient timer system 202, collection management system 204, and annotation system 206.
[0048] The transient timer system 202 is responsible for implementing temporary access to content granted by the messaging client application 104 and the messaging server application 114. To this end, the transient timer system 202 includes a plurality of timers that, based on duration and display parameters associated with the message, message collection (e.g., media collection), or graphical element, selectively display and enable access to messages and associated content via the messaging client application 104. Additional details regarding the operation of the transient timer system 202 are provided below.
[0049] The collection management system 204 is responsible for managing collections of media (e.g., collections of text, image video, and audio data). In some examples, collections of content (e.g., messages, including images, video, text, and audio) can be organized into "event galleries" or "event stories." Such collections can be made available for a specified time period, such as the duration of the event to which the content relates. For example, content related to a concert can be made available as a "story" for the duration of the concert. The collection management system 204 can also be responsible for publishing an icon to the user interface of the messaging client application 104 that provides notification of the existence of a particular collection.
[0050] The collection management system 204 also includes a curation interface 208 that enables collection managers to manage and curate specific content collections. For example, the curation interface 208 enables event organizers to curate content collections related to a specific event (e.g., removing inappropriate content or redundant messages). In addition, the collection management system 204 uses machine vision (or image recognition technology) and content rules to automatically curate content collections. In some embodiments, users can be paid compensation to include user-generated content in a collection. In such cases, the curation interface 208 operates to automatically pay such users for the use of their content.
[0051] The annotation system 206 provides various functions that enable users to annotate media content associated with a message or otherwise modify or edit the media content associated with the message. For example, the annotation system 206 provides functions related to generating and publishing media overlays for messages processed by the messaging system 100. The annotation system 206 can operatively provide media overlays (e.g., filters, lenses) to the messaging client application 104 based on the geographic location of the client device 102. In another example, the annotation system 206 can operatively provide media overlays to the messaging client application 104 based on other information, such as social network information of the user of the client device 102. Media overlays can include audio and visual content and visual effects. Examples of audio and visual content include pictures, text, logos, animations and sound effects, and animated facial models. Examples of visual effects include color overlays. Audio and visual content or visual effects can be applied to media content items (e.g., photos or videos) at the client device 102. For example, a media overlay includes text that can be superimposed on a photo generated by the client device 102. In another example, the media overlay includes a location identification overlay (e.g., Venice Beach), the name of a live event, or a business name overlay (e.g., Beach Cafe). In another example, the annotation system 206 uses the geographic location of the client device 102 to identify a media overlay that includes the name of a business at the geographic location of the client device 102. The media overlay may include other tags associated with the business. The media overlay may be stored in the database 120 and accessed through the database server 118.
[0052] In one example embodiment, the annotation system 206 provides a user-based publishing platform that enables a user to select a geographic location on a map and upload content associated with the selected geographic location. The user can also specify the circumstances under which a particular media overlay should be provided to other users. The annotation system 206 generates a media overlay including the uploaded content and associates the uploaded content with the selected geographic location.
[0053] In another example embodiment, the annotation system 206 provides a merchant-based publishing platform that enables merchants to select specific media overlays associated with a geographic location through a bidding process. For example, the annotation system 206 associates the highest-bidding merchant's media overlay with the corresponding geographic location for a predefined amount of time.
[0054] Figure 3 is a block diagram illustrating components of the AR system 124 according to certain example embodiments, wherein the components configure the AR system 124 to perform operations to generate AR parameters and perform AR functions.
[0055] In one example, the AR system 124 is shown as including an image module 302, a 3D reconstruction module 304, and a point cloud module 306, all of which are configured to communicate with each other (e.g., via a bus, shared memory, or switch). Any one or more of these modules may be implemented using one or more processors 308 (e.g., by configuring such one or more processors to perform the functions described for that module) and may therefore include one or more processors 308. The image module 302 is used to dedistort and segment images received from a wide-angle lens (e.g., from a fisheye lens, a 360-degree panoramic lens, or other lens that includes intentional distortion). The 3D reconstruction module 304 is used to generate pose and 3D point data as described in more detail below. The point cloud module 306 is used to store the 3D point cloud data generated by the 3D reconstruction module 304, but may also download and store a portion of an existing 3D cloud module at another time based on the GPS coordinates of the client device 102.
[0056] Any one or more of the modules described may be implemented using hardware alone (e.g., one or more processors 308 of a machine) or using a combination of hardware and software. For example, any module of the AR system 124 described may physically include an arrangement of one or more processors 308 (e.g., a subset of one or more processors of a machine or one or more processors in a machine) configured to perform the operations described herein for that module. As another example, any module of the AR system 124 may include software, hardware, or both software and hardware that configures an arrangement of one or more processors 308 (e.g., one or more processors of a machine) to perform the operations described herein for that module. Thus, different modules of the AR system 124 may include and configure different arrangements of such processors 308 or a single arrangement of such processors 308 at different points in time. Furthermore, any two or more modules of the AR system 124 may be combined into a single module, and the functionality described herein for a single module may be subdivided between multiple modules. Furthermore, according to various example embodiments, modules described herein as being implemented within a single machine, database, or device may be distributed across multiple machines, databases, or devices.
[0057] Figure 4 De-distorting 400 a wide-angle image 402 (eg, from a fisheye, 360-degree panoramic lens, or other lens including intentional distortion) into one or more linearly corrected images corresponding to at least a portion of the wide-angle image 402 is shown.
[0058] As can be seen, the wide angle image 402 includes a plurality of regions 404, 406, 408 that are distorted due to being generated using a wide angle lens. In the illustrated embodiment, the wide angle image 402 includes central regions of interest 406, 408, etc., as indicated by letters "a" through "h." Using known dedistortion techniques, and depending on the nature of the wide angle lens, these regions can be dedistorted as shown in FIG. Figure 4 4 and 5. As shown in FIG, regions 406 and 408 are dedistorted and linearized by image module 302. For example, region 406 and region 408 are transformed into dedistorted region 410 and dedistorted region 412, respectively.
[0059] like Figure 4 As shown in the example, the dedistorted regions have known relative poses. That is, the pose of each dedistorted region 410, 412, etc. relative to the camera frame is known or defined. These relative poses are configurable and can be specified by the user.
[0060] This paper will use the convention Nx to refer to the dedistorted image area, where N is the number of wide-angle images that have been captured and x is the letter corresponding to the dedistorted area. Figure 4 The dedistorted regions in are shown as square, aligned and directly adjacent, but this is not required. In some cases, it may be desirable for the dedistorted regions to overlap, as this can provide a better understanding of the following references. Figure 6 Better frame-to-frame region matching is discussed.
[0061] FIG5 schematically illustrates the processing of multiple viewpoints by a reference point cloud 502 (e.g., Figure 6 ), point cloud 502 is iteratively created as described in more detail below. (In the example shown), point cloud 502 represents a building. A plurality of dedistorted regions 504-512 have been previously registered to a 3D model of the environment, as described in more detail below. Each region 504-512 is represented as a rectangular pyramid, with the apex of the pyramid representing the position and orientation (i.e., pose) of the virtual camera and the base of the pyramid representing a view of point cloud 502 corresponding to dedistorted region Nx.
[0062] Each of regions 504 to 514 and its corresponding pose are determined by performing the method described below, treating each dedistorted region as an independent image. However, matching between dedistorted regions extracted from the same original wide-angle image is prevented because these matches cannot provide useful information.
[0063] For the purpose of illustration, Figure 5A As shown, region 504 corresponds to the dedistorted region 1a, ie, the dedistorted region a in frame 1 of the input image ( Figure 4 412 in FIG, region 506 corresponds to the dedistorted region 1b, region 508 corresponds to the dedistorted region 2d, region 510 corresponds to the dedistorted region 2a, region 512 corresponds to the dedistorted region 2c, and region 514 corresponds to the dedistorted region 2b. Region 514 (2c) is the registered region for frame 2 from wide-angle image 402.
[0064] like Figure 5B As shown, the new registration region 514 (2b) is shown to fit between regions 510 (2a) and 512 (2c), thereby providing consensus with regions 510 (2a) and 512 (2c) in the pose of frame 2 of the wide-angle image. Figure 5C As shown, it can now be concluded that region 508 (2d) is misregistered and this can be corrected as shown by arrow 516 so that region 508 is located next to region 512 (2c).
[0065] Figure 6A flowchart 600 illustrating a 3D reconstruction method according to one example performed by one or more processors or modules of the AR system 124 is shown.
[0066] The method begins at block 602, where the method is described above with reference to Figure 4 As described, a plurality of dewarped regions with known relative poses are generated from a plurality of wide-angle input images.
[0067] The following SfM steps are then performed on the dedistorted region at block 604:
[0068] - Detect image features, such as corner points, in each dedistorted region.
[0069] Matching image features between the dedistorted regions to identify matching feature pairs. The matching is performed between the dedistorted regions of the current frame of the wide-angle image and the dedistorted regions of one or more previous or subsequent frames. For the purpose of the matching, each dedistorted region is treated as an independent image.
[0070] - Estimating the relative poses of pairs of dedistorted regions that have feature matches. Note that in this case, these relative poses are not known relative poses between regions in a single image frame, but rather relative poses between pairs of dedistorted regions determined based on the identification and comparison of 2D features common to image regions in different frames. For example, dedistorted region 506 (1b) may include 2D features that are also found in dedistorted region 510 (2a), thereby allowing the relative pose between dedistorted region 506 and dedistorted region 510 to be determined.
[0071] Execution of these steps produces for each dedistorted region a set of 2D features and their matches to 2D features in other dedistorted regions, and an output of estimated relative poses between dedistorted regions with 2D feature matches.
[0072] Then, at block 606, two dedistorted regions are initially selected as registered regions that have a high number of matching features between them (thus providing a certain degree of confidence in their estimated relative pose). At this point in this example, a 3D point cloud 502 does not yet exist, although this method can also be used to supplement an existing point cloud. Therefore, the method now proceeds to block 608 where 3D points of the point cloud 502 are generated by triangulating the matching 2D features between the initial registered regions, as is known in the art.
[0073] The method then proceeds to block 610 where steps are now taken to optimize the region pose and 3D point positions to minimize the reprojection error, i.e., the error between the positions of features in the region and the positions of the corresponding 3D points when projected back into the region in which they were observed. This optimization is a well-known technique in SfM and is referred to as "Bundle Adjustment." After the bundle adjustment, a determination is made in decision block 612 as to whether there are more regions to register. If not, the method ends. If so, the method continues at block 614.
[0074] Then at block 614, another dedistorted region is selected and registered by matching 2D features in the selected region with 3D points in the current 3D point cloud 502. Using the matched 2D / 3D features, the pose of the selected region is determined.
[0075] Then, at decision block 616, a determination is made as to whether there is consistency between the newly registered region and other registered regions or sensor data from the same wide-angle image. Consistency can be, for example, that the N dedistorted regions are consistent with the 3D pose of the source wide-angle image, or that the newly registered region is consistent with Y dedistorted regions and wherein the Y dedistorted regions are consistent with the reported position sensor data and / or motion sensor data associated with the wide-angle image 402. After multiple images / regions with associated GPS data (obtained from position component 938) have been registered, the real-world position of the point cloud and other registered regions can be inferred. The amount of consistency required can be user-configurable.
[0076] In the first case, it may be specified that at least four regions from the same wide angle image be registered before inferring a pose of the wide angle image that can be used to assess consistency. In the second case, it may be specified that consistency with a new region may be determined once two registered regions agree within, for example, 10 cm of each other and within, for example, 5 meters of position sensor data (e.g., position reported by a GPS receiver and depending on the reported GPS accuracy) for the pose of the source wide angle image. Similarly, it may be specified that consistency with a new region may be determined once the poses of two regions agree within, for example, five degrees of a pose inferred from motion sensor data (e.g., orientation obtained or derived from an accelerometer, gyroscope, magnetometer, etc.).
[0077] The poses of different dedistorted regions from the same image frame can be compared to each other because their relative poses are known as described above, and the poses of the dedistorted regions can be compared to position or motion sensor data because the 3D model can be anchored to the real world and the pose of each region is known relative to the camera frame. If there is consistency in block 616, the method proceeds to block 618.
[0078] If consistency is not determined in decision block 616, the method proceeds to block 608, where the currently selected region is used to triangulate new 3D points based on the 2D feature matching performed in block 604 to further generate the 3D point cloud 502. The method then proceeds to block 610, where a bundle adjustment is now performed on the registered regions to optimize the image region pose and 3D point positions. After the bundle adjustment, a determination is made in decision block 612 as to whether there are any more dedistorted regions of any more image frames to be registered. If not, the method ends. If so, the method continues at block 614.
[0079] Returning now to decision block 616, if consistency has been determined for a particular set of regions of a wide-angle image, the pose of the wide-angle image is inferred at block 618. Once estimates of the poses of the consistent registration regions are known, and since the relative pose between each registration region and the wide-angle image in the same frame is known, the estimated pose of the wide-angle image can be inferred from the poses of the consistent registration regions of the same frame. These estimated poses for each registration region of the same frame for which consistency has been determined are then averaged to generate a best guess as to the true pose of the source wide-angle image. As used herein, the term "average" is defined as a number that measures the central tendency of a given set of numbers, including but not limited to a weighted or unweighted mean or median.
[0080] At block 620, the pose of each currently registered region for a particular wide-angle image is then corrected by setting the pose of each registered region to the pose of the image determined in block 618, making appropriate adjustments in each case to account for the relative positioning of each registered region (see Figure 5C ). In this way, the poses of the dedistorted regions that are not in the group of dedistorted regions with consistency between the estimated poses are derived from the pose of the wide-angle image.
[0081] At block 622, the currently unregistered regions of the particular wide-angle image are registered. This is a process that forces the registration of regions of the wide-angle image (with known pose) that have not yet been registered. After the pose of the wide-angle image is inferred in block 618, the pose of any regions that have not yet been registered can be assigned by setting the pose of each unregistered region to the image pose determined in block 618, making appropriate adjustments in each case to account for the relative positioning of each registered region. This is faster than waiting for the unregistered regions to be registered via blocks 614 through 620.
[0082] The method then proceeds to block 608 where new 3D points are triangulated using the currently selected region based on the 2D feature matching performed in block 604 to further generate the 3D point cloud 502 .
[0083] In this case, after inferring the image pose of the particular wide-angle image in block 618, correcting the image region poses in block 620, and registering the unregistered regions in block 622, the poses of the image regions of the particular wide-angle image are rigidly fixed relative to the inferred pose of the wide-angle image. In other words, the poses of the image regions of the particular wide-angle image can only be adjusted based on changes in the inferred pose of the particular wide-angle image.
[0084] The method then proceeds to box 610 where a bundle adjustment is now performed to optimize the image pose and 3D point positions in order to minimize the reprojection error. It is noteworthy that the bundle adjustment performed in box 610 is different depending on whether the method of arriving at box 610 is from "no consistency" in box 616 or from box 622. In the former case, the pose of the regions is optimized independently in box 610 before consistency is achieved for the pose of the wide-angle image. This is slower, but provides robustness - if one region is in the wrong position, it does not affect other regions. Once there is consistency on the pose of the wide-angle image and the regions are rigidly fixed relative to that pose, only the pose of the image is optimized, which is faster than optimizing each region independently. Since the relative pose of the dedistorted regions in the same frame is known, the reprojection error can be determined by reprojecting the 3D points into the dedistorted regions where the corresponding 2D features are found.
[0085] After bundling adjustments for the pose of a particular wide-angle image, a determination is made in decision block 612 as to whether there are any more dedistorted regions of any more wide-angle images to register. If not, the method ends. If so, the method returns to block 614.
[0086] The output of the optimization performed in 610 is the final pose of all regions and images that have been successfully registered and the 3D point cloud 502. The 3D point cloud 502 can be referenced as follows Figure 7 The 3D point cloud 502 may be stored on the client device 102 for use in ongoing SfM or augmented reality operations and / or may be uploaded to the application server 112 to create or supplement a 3D point cloud hosted on the application server 112 .
[0087] Figure 7 FIG. 7 is an exemplary interface flow diagram 700 illustrating the display of location-based AR content rendered by the client device 102 according to some example embodiments. Figure 7 As shown, the interface flow chart includes interface diagram 702 and interface diagram 704 .
[0088] In one example, client device 102 may cause a rendering of interface diagram 702 to be displayed on a display of client device 102. For example, client device 102 may capture image data via a camera and generate the interface depicted by interface diagram 702.
[0089] As shown in interface diagram 704, the client device 102 can access media content within a repository (e.g., database 120) based on the location of the client device 102. Media content (e.g., media content 706), including virtual objects or other augmented information or images, can be associated with a location within the media repository such that a reference to the location within the repository identifies the media content 706. Alternatively, the media content can be located in the memory of the client device 102. Media content can also be identified by user preference or selection.
[0090] The client device 102 may then cause a rendering of the media content 706 to be displayed at a location within the GUI based on the positioning or tracking pose generated from the image captured by the client device 102 and the 3D cloud generated in the flowchart 600 , as seen in the interface diagram 704 .
[0091] Software Architecture
[0092] Figure 8 is a block diagram illustrating an example software architecture 806 that can be used in conjunction with the various hardware architectures described herein. Figure 8 is a non-limiting example of a software architecture, and it should be understood that many other architectures may be implemented to facilitate the functionality described herein. The software architecture 806 may be implemented in a manner such as Figure 9 900, which includes, among other things, a processor 904, a memory 914, and I / O components 918. A representative hardware layer 852 is shown and may represent, for example, Figure 8 800. A representative hardware layer 852 includes a processing unit 854 with associated executable instructions 804. Executable instructions 804 represent executable instructions of a software architecture 806, including implementations of the methods, components, and the like described herein. Hardware layer 852 also includes a memory and / or storage module, i.e., memory / storage 856, also having executable instructions 804. Hardware layer 852 may also include other hardware 858.
[0093] exist Figure 8In the example architecture of , software architecture 806 can be conceptualized as a stack of layers, where each layer provides specific functionality. For example, software architecture 806 may include layers such as operating system 802, library 820, application 816, and presentation layer 814. In operation, applications 816 and / or other components within a layer may call application programming interfaces (APIs) API calls 808 through the software stack and receive responses in response to API calls 808. The layers shown are representative in nature, and not all software architectures have all layers. For example, some mobile operating systems or dedicated operating systems may not provide framework / middleware 818, while other operating systems may provide such a layer. Other software architectures may include additional layers or different layers.
[0094] The operating system 802 can manage hardware resources and provide public services. The operating system 802 may include, for example, a kernel 822, services 824, and drivers 826. The kernel 822 may act as an abstraction layer between the hardware and other software layers. For example, the kernel 822 may be responsible for memory management, processor management (e.g., scheduling), component management, networking, security settings, etc. The services 824 may provide other public services to other software layers. The drivers 826 are responsible for controlling or interfacing with the underlying hardware. For example, depending on the hardware configuration, the drivers 826 may include a display driver, a camera driver, drivers, flash memory drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers), Drivers, audio drivers, power management drivers, etc.
[0095] The libraries 820 provide common infrastructure used by the applications 816 and / or other components and / or layers. The libraries 820 provide functionality that allows other software components to perform tasks more easily than by directly interfacing with the functionality of the underlying operating system 802 (e.g., kernel 822, services 824, and / or drivers 826). The libraries 820 may include system libraries 844 (e.g., C standard libraries), which may provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. In addition, the libraries 820 may include API libraries 846 such as media libraries (e.g., libraries that support the rendering and manipulation of various media formats such as MPREG4, H.264, MP3, AAC, AMR, JPG, and PNG), graphics libraries (e.g., OpenGL frameworks that can be used to render 2D and 3D graphics content on a display), databases (e.g., SQLite that can provide various relational database functions), network libraries (e.g., WebKit that can provide web browsing functions), etc. The library 820 may also include various other libraries 848 to provide numerous other APIs to the applications 816 and other software components / modules.
[0096] The framework / middleware 818 (sometimes also referred to as middleware) provides a higher-level common infrastructure that can be used by applications 816 and / or other software components / modules. For example, the framework / middleware 818 can provide various graphical user interface (GUI) functions, advanced resource management, advanced location services, etc. The framework / middleware 818 can provide a wide range of other APIs that can be used by applications 816 and / or other software components / modules, some of which can be specific to a particular operating system 802 or platform.
[0097] Applications 816 include built-in applications 838 and / or third-party applications 840. Examples of representative built-in applications 838 may include, but are not limited to, contact applications, browser applications, book reader applications, location applications, media applications, messaging applications, and / or game applications. Third-party applications 840 may include applications that are run by entities other than the vendor of a particular platform using Android TM or IOS TM Software Development Kit (SDK) for developing applications and can be used on platforms such as IOS TM ANDROID TM 、 Mobile software running on the mobile operating system of the Phone or other mobile operating systems. The third-party application 840 can call API calls 808 provided by the mobile operating system (e.g., operating system 802) to facilitate the functions described herein.
[0098] Applications 816 may create user interfaces to interact with users of the system using built-in operating system functionality (e.g., kernel 822, services 824, and / or drivers 826), libraries 820, and framework / middleware 818. Alternatively or additionally, in some systems, interaction with the user may occur through a presentation layer, such as presentation layer 814. In these systems, application / component "logic" may be separated from aspects of the application / component that interact with the user.
[0099] Figure 9 is a block diagram illustrating components of a machine 900 capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium) and performing any one or more of the methodologies discussed herein, according to some example embodiments. Specifically, Figure 9A diagrammatic representation of a machine 900 in the example form of a computer system is shown within which instructions 910 (e.g., software, programs, applications, applet, apps, or other executable code) may be executed for causing the machine 900 to perform any one or more of the methodologies discussed herein. Thus, the instructions 910 may be used to implement the modules or components described herein. The instructions 910 transform a general-purpose, unprogrammed machine 900 into a specific machine 900 that is programmed to perform the functions described and illustrated in the manner described. In alternative embodiments, the machine 900 operates as a standalone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the machine 900 may operate in the capacity of a server or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 900 may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart home appliance), other smart devices, a network appliance, a network router, a network switch, a network bridge, or any machine capable of executing, sequentially or otherwise, the instructions 910 specifying actions to be taken by the machine 900. Furthermore, while only a single machine 900 is shown, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute the instructions 910 to perform any one or more of the methodologies discussed herein.
[0100] The machine 900 may include a processor 904, a memory / storage 906, and an I / O component 918, which may be configured to communicate with each other, for example, via a bus 902. The memory / storage 906 may include a memory 914, such as a main memory or other storage device, and a storage unit 916, both of which are accessible to the processor 904, for example, via the bus 902. The storage unit 916 and the memory 914 store instructions 910 that embody any one or more of the methods or functions described herein. The instructions 910 may also reside, completely or partially, within the memory 914, within the storage unit 916, within at least one of the processors 904 (e.g., within a cache memory of the processor), or within any suitable combination thereof during execution thereof by the machine 900. Thus, the memory 914, the storage unit 916, and the memory of the processor 904 are examples of machine-readable media.
[0101] The I / O components 918 may include various components for receiving input, providing output, generating output, transmitting information, exchanging information, capturing measurements, etc. The specific I / O components 918 included in a particular machine 900 will depend on the type of machine. For example, a portable machine such as a mobile phone will likely include a touch input device or other such input mechanism, while a headless server machine will likely not include such a touch input device. It will be understood that the I / O components 918 may include Figure 9 Many other components are not shown. The I / O components 918 are grouped according to function for the purpose of simplifying the following discussion only and are by no means limiting. In various example embodiments, the I / O components 918 may include output components 926 and input components 928. The output components 926 may include visual components (e.g., a display such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., a speaker), tactile components (e.g., a vibration motor, a resistance mechanism), other signal generators, and the like. The input components 928 may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input components), pointing-based input components (e.g., a mouse, touchpad, trackball, joystick, motion sensor, or other pointing instrument), tactile input components (e.g., a physical button, a touch screen or other tactile input component that provides location and / or force of a touch or touch gesture), audio input components (e.g., a microphone), and the like. The input component 928 may also include one or more image capture devices, such as a camera. The camera may include a wide-angle lens (e.g., a fisheye lens, a 360-degree panoramic lens, or other lens including intentional distortion), from which the methods and systems described herein receive wide-angle images for processing.
[0102] In other example embodiments, the I / O component 918 may include a biometric component 930, a motion sensor component 934, an environmental component 936, or a positioning component 938, as well as various other components. For example, the biometric component 930 may include components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), etc. The motion sensor component 934 may include an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), etc. The environment component 936 may include, for example, an illumination sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers that detect ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), a sound sensor component (e.g., one or more microphones that detect background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor (e.g., a gas detection sensor for detecting concentrations of hazardous gases to ensure safety or for measuring pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment. The positioning component 938 may include a position sensor component (e.g., a global positioning system (GPS) receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure from which altitude can be derived), an orientation sensor component (e.g., a magnetometer), etc. In this regard, it should be noted that a magnetometer can be considered an orientation sensor and a motion sensor because changes in the magnetometer output also indicate rotational motion.
[0103] Various technologies may be used to implement communications. I / O components 918 may include communications components 940 operable to couple machine 900 to network 932 or device 920 via coupling 922 and coupling 924, respectively. For example, communications components 940 may include a network interface component or other suitable device to interface with network 932. In other examples, communications components 940 may include wired communications components, wireless communications components, cellular communications components, near field communications (NFC) components, components (e.g., low power ), Device 920 may be another machine or any of a variety of peripheral devices (eg, a peripheral device coupled via a universal serial bus (USB)).
[0104] In addition, the communication component 940 can detect the identifier or can include a component that is operable to detect the identifier. For example, the communication component 940 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional bar codes such as Universal Product Code (UPC) bar codes, multi-dimensional bar codes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D bar codes, and other optical codes) or an acoustic detection component (e.g., a microphone for identifying an audio signal of a tag). In addition, various information can be obtained via the communication component 940, such as location via Internet Protocol (IP) geolocation, location information via Internet Protocol (IP), ... Signal triangulation to obtain location, location obtained by detecting NFC beacon signals that can indicate a specific location, etc.
[0105] Glossary
[0106] "Carrier signal" in this context refers to any intangible medium capable of storing, encoding, or carrying instructions for execution by a machine, and includes digital or analog communication signals or other intangible media that facilitate communication of such instructions. Instructions may be sent or received over a network via a network interface device using a transmission medium and using any of a number of well-known transmission protocols.
[0107] A "client device" in this context refers to any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. A client device may be, but is not limited to, a mobile phone, a desktop computer, a laptop computer, a portable digital assistant (PDA), a smartphone, a tablet computer, an ultrabook, a netbook, a laptop computer, a multiprocessor system, a microprocessor-based or programmable consumer electronics product, a game console, a set-top box, or any other communication device that a user can use to access a network.
[0108] A "communication network" in this context refers to one or more parts of a network, which may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a part of the Internet, a part of the public switched telephone network (PSTN), a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, The network, other types of networks, or a combination of two or more such networks. For example, the network or a portion of the network may include a wireless network or a cellular network, and the coupling may be a code division multiple access (CDMA) connection, a global system for mobile communications (GSM) connection, or other types of cellular or wireless coupling. In this example, the coupling may implement any of various types of data transmission technologies, such as single carrier radio transmission technology (1xRTT), evolution data optimized (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rates for GSM evolution (EDGE) technology, the third generation partnership project (3GPP) including 3G, fourth generation wireless (4G) networks, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), world wide interoperability for microwave access (WiMAX), long term evolution (LTE) standards, other standards defined by various standards setting organizations, other long range protocols, or other data transmission technologies.
[0109] In this context, a "transient message" is a message that can be accessed for a limited duration. Transient messages can be text, images, videos, and more. The access period for a transient message can be set by the sender of the message. Alternatively, the access period can be a default setting or a setting specified by the recipient. Regardless of the setting technique, the message is transient.
[0110] "Machine-readable medium" in this context refers to a component, device, or other tangible medium that is capable of storing instructions and data, either temporarily or permanently, and may include, but is not limited to, random access memory (RAM), read-only memory (ROM), buffer memory, flash memory, optical media, magnetic media, cache memory, other types of storage devices (e.g., erasable programmable read-only memory (EEPROM)), and / or any suitable combination thereof. The term "machine-readable medium" should be taken to include a single medium or multiple media (e.g., a centralized or distributed database or associated cache memory and servers) that can store instructions. The term "machine-readable medium" should also be taken to include any medium or combination of multiple media that can store instructions (e.g., code) executed by a machine, such that the instructions, when executed by one or more processors of the machine, cause the machine to perform any one or more of the methods described herein. Thus, "machine-readable medium" refers to a single storage device or device, as well as a "cloud-based" storage system or storage network that includes multiple storage devices or devices. The term "machine-readable medium" does not include the signal itself.
[0111] "Component" in this context refers to a device, physical entity or logic with boundaries defined by function or subroutine calls, branch points, application program interfaces (APIs) or other technologies that provide partitioning or modularization of specific processing or control functions. A component can be combined with other components via its interface to perform machine processing. A component can be a packaged functional hardware unit designed for use with other components and can be part of a program that generally performs a specific function among related functions. A component can constitute a software component (e.g., a code embodied on a machine-readable medium) or a hardware component. A "hardware component" is a tangible unit that can perform certain operations and can be configured or arranged in a certain physical manner. In various example embodiments, one or more computer systems (e.g., a stand-alone computer system, a client computer system or a server computer system) or one or more hardware components (e.g., a processor or a processor group) of a computer system can be configured as a hardware component for performing certain operations described herein by software (e.g., an application or an application portion). Hardware components can also be implemented mechanically, electronically or in any suitable combination thereof. For example, a hardware component can include a dedicated circuit or logic that is permanently configured to perform certain operations. The hardware component can be a special-purpose processor, such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). The hardware component can also include a programmable logic or circuit that is temporarily configured to perform certain operations by software. For example, the hardware component can include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a specific machine (or specific component of a machine) that is uniquely customized to perform the configured function, rather than a general-purpose processor. It will be understood that it can be decided to mechanically implement the hardware component in a circuit that is dedicated and permanently configured or in a circuit that is temporarily configured (for example, configured by software) for cost and time considerations. Accordingly, the phrase "hardware component" (or "hardware-implemented component") should be understood to include a tangible entity, that is, an entity that is physically constructed, permanently configured (for example, hardwired) or temporarily configured (for example, programmed) to operate in some way or perform certain operations described herein. Considering that the hardware component is temporarily configured (for example, programmed), it is not necessary to configure or instantiate each hardware component in the hardware component at any one time. For example, in the case where a hardware component includes a general-purpose processor that is configured by software to become a special-purpose processor, the general-purpose processor can be configured as different special-purpose processors (e.g., including different hardware components) at different times. The software configures one or more specific processors accordingly to constitute a specific hardware component at one time and to constitute different hardware components at different times. Hardware components can provide information to other hardware components and receive information from other hardware components. Therefore, the described hardware components can be considered to be communicatively coupled.When there are multiple hardware components at the same time, communication can be achieved by (for example, by appropriate circuits and buses) signal transmission between or among two or more hardware components. In the embodiment in which multiple hardware components are configured or instantiated at different times, communication between such hardware components can be achieved, for example, by storing information in a memory structure accessed by multiple hardware components and retrieving information in the memory structure. For example, a hardware component can perform an operation, and the output of the operation is stored in a memory device coupled to its communication ground. Then, other hardware components can access the memory device at a subsequent time to retrieve the stored output and process it. The hardware component can also initiate communication with an input device or an output device, and can operate on resources (for example, the collection of information). The various operations of the example methods described herein can be performed at least in part by one or more processors that are temporarily configured or permanently configured to perform related operations (for example, by software). Whether it is temporarily configured or permanently configured, such a processor can constitute a processor-implemented component that operates to perform one or more operations or functions described herein. As used herein, "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the method described herein can be implemented at least in part by a processor, wherein specific one or more processors are examples of hardware. For example, at least some of the operations of the method can be performed by one or more processors or the components implemented by the processor. In addition, one or more processors can also operate to support the execution of related operations in a "cloud computing" environment or as a "software as a service" (SaaS) operation. For example, at least some of the operations in the operation can be performed by a group of computers (as an example of a machine including a processor), wherein these operations can be accessed via a network (for example, the Internet) and via one or more appropriate interfaces (for example, application program interface (API)). The execution of certain operations in the operation can be distributed between processors, not only can reside in a single machine, but also can be deployed across multiple machines. In some example embodiments, a processor or the components implemented by the processor can be located in a single geographical location (for example, in a home environment, an office environment or a server cluster). In other example embodiments, a processor or the components implemented by the processor can be distributed across multiple geographical locations.
[0112] A "processor" in this context refers to any circuit or virtual circuit (a physical circuit simulated by logic executed on an actual processor) that manipulates data values according to control signals (e.g., "commands," "opcodes," "machine code," etc.) and produces corresponding output signals that are used to operate a machine. For example, a processor may be a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), or any combination thereof. A processor may also be a multi-core processor having two or more independent processors (sometimes called "cores") that can execute instructions simultaneously.
[0113] A "timestamp" in this context refers to a sequence of characters or encoded information that identifies when some event occurred, for example, giving a date and time of day, sometimes accurate to a fraction of a second.
[0114] "LIFT" in this context is a measure of the performance of a targeted model in predicting or classifying cases as having an enhanced response (relative to the entire population), measured against a random selection of targeted models.
[0115] In this context, phoneme alignment refers to the unit of speech that distinguishes one word from another. A phoneme can consist of a sequence of closures, bursts, and aspirations; or, a diphthong can transition from a back vowel to a front vowel. Therefore, a speech signal can be described not only by the phonemes it contains, but also by the positions of the phonemes. Therefore, phoneme alignment can be described as the "time alignment" of phonemes in a waveform to determine the proper sequence and position of each phoneme in the speech signal.
[0116] "Audio-to-visual conversion" in this context refers to converting an audible speech signal into visual speech, where the visual speech may include mouth shapes representing the audible speech signal.
[0117] In this context, a “Time Delay Neural Network (TDNN)” is an artificial neural network architecture whose primary purpose is to process sequential data. An example is converting continuous audio into a stream of categorical phoneme labels for speech recognition.
[0118] In this context, "Bidirectional Long Short-Term Memory (BLSTM)" refers to a recurrent neural network (RNN) architecture that memorizes values over arbitrary intervals. The stored values are not modified as learning progresses. RNNs allow for both forward and backward connections between neurons. BLSTM is well-suited for time series classification, processing, and forecasting, given unknown time lags and durations between events.
[0119] According to the above description, the embodiments of the present invention further disclose the following technical solutions, including but not limited to:
[0120] Solution 1. A method for determining the pose of a wide-angle image using one or more processors, comprising:
[0121] Dedistorting a region of the wide-angle image;
[0122] determining an estimated pose of the dedistorted region of the wide-angle image; and
[0123] The pose of the wide-angle image is derived from the estimated pose of the dedistorted region.
[0124] Option 2. A method according to Option 1, wherein determining the estimated pose of the dedistorted region includes: comparing features in the dedistorted region with features in other dedistorted regions from one or more other wide-angle images.
[0125] Option 3. A method according to Option 1, wherein determining the estimated pose includes comparing features in the dedistorted region with features in a point cloud.
[0126] Solution 4. The method of solution 1, wherein determining the estimated pose comprises:
[0127] determining a set of dedistorted regions having consistency between estimated poses; and
[0128] The pose of the wide-angle image is derived from the estimated pose of the dedistorted region having such consistency.
[0129] Solution 5. The method according to solution 4, further comprising:
[0130] The poses of the dedistorted regions that are not in the group of dedistorted regions having consistency between estimated poses are derived from the pose of the wide-angle image.
[0131] Solution 6. The method according to solution 4, further comprising:
[0132] Optimize the estimated pose of the dedistorted region.
[0133] Option 7. A method according to Option 6, wherein the updated pose of the wide-angle image is derived from at least some of the optimized estimated poses.
[0134] Option 8. The method according to Option 6 further includes: optimizing the pose of the wide-angle image after optimizing the estimated pose of the dedistorted area.
[0135] Option 9. The method according to Option 1, wherein deriving the pose of the wide-angle image comprises determining the pose of the wide-angle image as an average of at least some of the estimated poses.
[0136] Option 10. The method of Option 4, wherein consistency is determined with reference to sensor data selected from the group consisting of motion sensor data and position sensor data.
[0137] Solution 11. A data processing system comprising:
[0138] one or more processors;
[0139] wide-angle image capture device, and
[0140] One or more machine-readable media storing instructions that, when executed by the one or more processors, cause the system to perform operations comprising:
[0141] receiving a wide-angle image from the wide-angle image capture device;
[0142] Dedistorting a region of the wide-angle image;
[0143] determining an estimated pose of the dedistorted region of the wide-angle image; and
[0144] The pose of the wide-angle image is derived from the estimated pose of the dedistorted region.
[0145] Option 12. A data processing system according to Option 11, wherein determining the estimated pose of the dedistorted region includes: comparing features in the dedistorted region with features in other dedistorted regions from one or more other wide-angle images.
[0146] Solution 13. The data processing system of solution 11, wherein determining the estimated pose comprises:
[0147] determining a set of dedistorted regions having consistency between estimated poses; and
[0148] The pose of the wide-angle image is derived from the estimated pose of the dedistorted region having such consistency.
[0149] Option 14. The data processing system of Option 11, wherein consistency is determined with reference to position sensor data or motion sensor data.
[0150] Solution 15. The data processing system according to Solution 11, wherein the operations further comprise:
[0151] optimizing the estimated pose of the dedistorted region; and
[0152] The pose of the wide-angle image is optimized after optimizing the estimated pose of the dedistorted region.
[0153] Embodiment 16. A non-transitory machine-readable medium comprising instructions that, when read by a machine, cause the machine to perform operations for determining a pose of a wide-angle image, comprising:
[0154] receiving a wide-angle image from a wide-angle image capture device;
[0155] Dedistorting a region of the wide-angle image;
[0156] determining an estimated pose of the dedistorted region of the wide-angle image; and
[0157] The pose of the wide-angle image is derived from the estimated pose of the dedistorted region.
[0158] Option 17. A non-transitory machine-readable medium according to Option 16, wherein determining the estimated pose of the dedistorted region includes: comparing features in the dedistorted region with features in other regions that have been dedistorted from one or more other wide-angle images.
[0159] Option 18. The non-transitory machine-readable medium of option 16, wherein the operation of determining the estimated pose comprises:
[0160] determining a set of dedistorted regions having consistency between estimated poses; and
[0161] The pose of the wide-angle image is derived from the estimated pose of the dedistorted region having such consistency.
[0162] Option 19. The non-transitory machine-readable medium of Option 18, wherein consistency is determined with reference to position sensor data or motion sensor data.
[0163] Option 20. The non-transitory machine-readable medium of option 17, wherein the operations further comprise:
[0164] optimizing the estimated pose of the dedistorted region individually; and
[0165] The pose of the wide-angle image is optimized after separately optimizing the estimated pose of the dedistorted region.
Claims
1. A method for determining a pose of a first wide-angle image using one or more processors, comprising: Dedistorting a region of the first wide-angle image; dedistorting a region of the second wide-angle image; detecting image features in the dedistorted region of the first wide-angle image and image features in the dedistorted region of the second wide-angle image; identifying pairs of matching image features between the dedistorted region of the first wide-angle image and the dedistorted region of the second wide-angle image; determining a relative pose between a dedistorted region of the first wide-angle image and a dedistorted region of the second wide-angle image having a pair of matching image features; triangulating pairs of matching image features to generate 3D point locations of the matching image features; determining an estimated pose of the dedistorted region of the first wide-angle image based on the 3D point positions of the matching image features; as well as The pose of the first wide-angle image is derived from the estimated pose of the dedistorted region of the first wide-angle image.
2. The method according to claim 1, wherein Determining the estimated pose involves: determining a set of dedistorted regions of the first wide-angle image having consistency between estimated poses; and The pose of the first wide-angle image is derived from the estimated pose of the dedistorted region having such consistency.
3. The method according to claim 2, further comprising: The poses of the dedistorted regions of the first wide-angle image that are not in the group of dedistorted regions having consistency between estimated poses are derived from the pose of the first wide-angle image.
4. The method according to claim 2, further comprising: A consistency of the dedistorted region of the first wide-angle image is determined based on consistency of the estimated pose with position sensor data or motion sensor data.
5. The method according to claim 2, further comprising: The consistency of the specific dedistorted region of the first wide-angle image is determined based on the estimated pose of the specific dedistorted region being consistent with the estimated poses of a predetermined number of dedistorted regions of the first wide-angle image having consistency between the estimated poses.
6. The method according to claim 1, further comprising: Optimize the estimated pose of the dedistorted region. 7 . The method of claim 6 , further comprising optimizing the pose of the first wide-angle image after optimizing the estimated pose of the dedistorted region.
8. A data processing system comprising: one or more processors; wide-angle image capture device, and One or more machine-readable media storing instructions that, when executed by the one or more processors, cause the system to perform operations comprising: Dedistorting a region of the first wide-angle image; dedistorting a region of the second wide-angle image; detecting image features in the dedistorted region of the first wide-angle image and image features in the dedistorted region of the second wide-angle image; identifying pairs of matching image features between the dedistorted region of the first wide-angle image and the dedistorted region of the second wide-angle image; determining a relative pose between a dedistorted region of the first wide-angle image and a dedistorted region of the second wide-angle image having a pair of matching image features; triangulating pairs of matching image features to generate 3D point locations of the matching image features; determining an estimated pose of the dedistorted region of the first wide-angle image based on the 3D point positions of the matching image features; and The pose of the first wide-angle image is derived from the estimated pose of the dedistorted region of the first wide-angle image.
9. The data processing system according to claim 8, wherein: The operations to determine the estimated pose include: determining a set of dedistorted regions of the first wide-angle image having consistency between estimated poses; and The pose of the first wide-angle image is derived from the estimated pose of the dedistorted region having such consistency.
10. The data processing system according to claim 9, wherein: The operations further include: The poses of the dedistorted regions of the first wide-angle image that are not in the group of dedistorted regions having consistency between estimated poses are derived from the pose of the first wide-angle image.
11. The data processing system according to claim 9, wherein: The operations further include: A consistency of the dedistorted region of the first wide-angle image is determined based on consistency of the estimated pose with position sensor data or motion sensor data.
12. The data processing system according to claim 9, wherein: The operations further include: The consistency of the specific dedistorted region of the first wide-angle image is determined based on the estimated pose of the specific dedistorted region being consistent with the estimated poses of a predetermined number of dedistorted regions of the first wide-angle image having consistency between the estimated poses.
13. The data processing system according to claim 8, wherein: The operations further include: Optimize the estimated pose of the dedistorted region.
14. The data processing system according to claim 13, wherein: The operations further include: The pose of the first wide-angle image is optimized after optimizing the estimated pose of the dedistorted region.
15. A non-transitory machine-readable medium comprising instructions that, when read by a machine, cause the machine to perform operations for determining a pose of a first wide-angle image, the operations comprising: Dedistorting a region of the first wide-angle image; dedistorting a region of the second wide-angle image; detecting image features in the dedistorted region of the first wide-angle image and image features in the dedistorted region of the second wide-angle image; identifying pairs of matching image features between the dedistorted region of the first wide-angle image and the dedistorted region of the second wide-angle image; determining a relative pose between a dedistorted region of the first wide-angle image and a dedistorted region of the second wide-angle image having a pair of matching image features; triangulating pairs of matching image features to generate 3D point locations of the matching image features; determining an estimated pose of the dedistorted region of the first wide-angle image based on the 3D point positions of the matching image features; as well as The pose of the first wide-angle image is derived from the estimated pose of the dedistorted region of the first wide-angle image.
16. The non-transitory machine-readable medium of claim 15, wherein: The operations to determine the estimated pose include: determining a set of dedistorted regions of the first wide-angle image having consistency between estimated poses; and The pose of the first wide-angle image is derived from the estimated pose of the dedistorted region having such consistency.
17. The non-transitory machine-readable medium of claim 16, wherein: The operations further include: The poses of the dedistorted regions of the first wide-angle image that are not in the group of dedistorted regions having consistency between estimated poses are derived from the pose of the first wide-angle image.
18. The non-transitory machine-readable medium of claim 16, wherein: The operations further include: A consistency of the dedistorted region of the first wide-angle image is determined based on consistency of the estimated pose with position sensor data or motion sensor data.
19. The non-transitory machine-readable medium of claim 15, wherein: The operations further include: The consistency of the specific dedistorted region of the first wide-angle image is determined based on the estimated pose of the specific dedistorted region being consistent with the estimated poses of a predetermined number of dedistorted regions of the first wide-angle image having consistency between the estimated poses.
20. The non-transitory machine-readable medium of claim 17, wherein: The operations further include: optimizing the estimated pose of the dedistorted region individually; and The pose of the first wide-angle image is optimized after separately optimizing the estimated pose of the dedistorted region.
Citation Information
Patent Citations
Image processing device, image processing method, program thereof, recording medium containing the program
CN101305595A
Method and apparatus for processing wide angle image
CN107871308A