Enhanced three-dimensional structure generation
By using image dataset annotation and external spatial reference datasets, combined with machine learning and photogrammetry technology, the computational difficulty of generating real-world 3D maps is solved, and efficient and accurate 3D map generation is achieved.
Patent Information
- Application Number
- CN202380082784.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-01
- Filing Date
- 2023-11-30
- Publication Date
- 2025-07-11
AI Technical Summary
Generating three-dimensional maps of real-world environments is computationally difficult and computationally intensive, and prior art is difficult to efficiently generate accurate 3D maps.
By using the annotation of the image dataset and the external spatial reference dataset, combined with machine learning classifiers and photogrammetry technology, point cloud data is generated and corrected to achieve co-registration and splicing of point clouds, and a high-accuracy 3D map is generated.
It improves the accuracy and efficiency of 3D map generation, reduces the demand for computing resources, and can generate a highly realistic 3D map.
Smart Images

Figure CN120303699A_ABST
Abstract
Description
[0001] Priority Claim
[0002] This application claims the benefit of U.S. Patent Application No. 18 / 072,958, filed on December 1, 2022, the entire content of which is incorporated herein by reference. Technical Field
[0003] The present disclosure generally relates to data processing, and more particularly, to using image data generation models. Background Art
[0004] Generating three-dimensional (3D) maps of real-world environments is computationally difficult. Implementing robots and computer vision systems to map real-world environments involves significant expenditures on computer vision and robotic devices, as well as computationally intensive processing to generate accurate results. Brief Description of the Drawings
[0005] To easily identify the discussion of any particular element or action, one or more of the most significant digits in the reference numerals refer to the figure ("FIG") number in which that element or action is first introduced.
[0006] Figure 1 is a block diagram showing an example messaging system for exchanging data (e.g., messages and associated content) over a network.
[0007] Figure 2 is a block diagram showing further details of the messaging system with respect to Figure 1 of the messaging system.
[0008] Figure 3 is a schematic diagram showing data that can be stored in a database of a messaging server system according to certain example embodiments.
[0009] Figure 4 is a schematic diagram showing the structure of a message generated by a messaging client application for communication.
[0010] Figure 5 is a schematic diagram showing an example access restriction process according to some example embodiments, according to which access to content (e.g., multimedia payloads of ephemeral messages and associated data) or a collection of content (e.g., an ephemeral message story) can be time-limited (e.g., such that the access is ephemeral).
[0011] Figure 6 shows an example flowchart for generating a map from different image sources according to some example embodiments.
[0012] Figure 7 An apparatus for generating different image datasets from an orthogonal perspective according to some example embodiments is shown.
[0013] Figure 8 A 3D map generated from different data according to some example embodiments is shown.
[0014] Figure 9 An illustration of long - range drift in structure from motion reconstruction according to some example embodiments is shown.
[0015] Figure 10 An example flowchart of a method for implementing structure from motion to generate improved 3D data from image data according to some example embodiments is shown.
[0016] Figure 11 An illustration of minimizing errors in SfM - constructed data according to some example embodiments is shown.
[0017] Figure 12 An illustration of improved 3D data according to some example embodiments is shown.
[0018] Figure 13 A block diagram of components of a machine capable of reading instructions from a machine - readable medium (e.g., a machine - readable storage medium) and performing any one or more of the methods discussed herein according to some example embodiments is shown. Detailed Description
[0019] The following description includes systems, methods, techniques, instruction sequences, and computer program products that implement illustrative embodiments of the present disclosure. In the following description, for the purposes of illustration, numerous specific details are set forth in order to provide an understanding of the various embodiments of the inventive subject matter. However, it will be apparent to those skilled in the art that embodiments of the inventive subject matter may be practiced without these specific details. Generally, well - known instruction instances, protocols, structures, and techniques are not necessarily shown in detail.
[0020] One challenge in computer vision is to generate 3D maps using remote sensing data. Examples of remote sensing data include images or video data generated by drones and remote sensing data purchased from companies with aerial systems (e.g., airplanes, satellites according to some example embodiments). To address the foregoing problems, a 3D mapping system can be configured to generate a 3D map of the real-world environment using annotations of a larger image dataset (e.g., video provided by one or more end users of a website). The image dataset includes ground images that can be programmatically labeled with accurate tags (e.g., neural network image segmentation), where aerial image data from remote sensing (e.g., aerial LIDAR (Light Detection and Ranging)) and location data (e.g., GPS data) are used to improve and enhance the tag accuracy. In some example embodiments, the 3D mapping system implements photogrammetry (e.g., Alice Vision, Call Open) to create a point cloud from images (e.g., video clips of a city). Each pixel in the point cloud can be classified based on the consensus of each frame of a given video. The point cloud can be co-registered to a remote sensing reference dataset (e.g., data provided by an aerial device) to provide precise spatial coordinates for each pixel. Different patches of the point cloud can be stitched together to provide a complete 3D map of a given area, such as the downtown area of a city).
[0021] Figure 1 A block diagram of an example messaging system 100 for exchanging data (e.g., messages and associated content) over a network 106 is shown. The messaging system 100 includes a plurality of client devices 102, each client device 102 hosting a plurality of applications including a messaging client application 104. Each messaging client application 104 is communicatively coupled via the network 106 (e.g., the Internet) to other instances of the messaging client application 104 and a messaging server system 108.
[0022] Accordingly, each messaging client application 104 is capable of communicating and exchanging data with another messaging client application 104 and with the messaging server system 108 via the network 106. The data exchanged between the messaging client applications 104 and between the messaging client application 104 and the messaging server system 108 includes functions (e.g., commands to invoke functions) and payload data (e.g., text, audio, video, or other multimedia data).
[0023] The messaging server system 108 provides server - side functionality to a particular messaging client application 104 via the network 106. Although certain functions of the messaging system 100 are described herein as being performed by the messaging client application 104 or by the messaging server system 108, it should be understood that the location of certain functions within the messaging client application 104 or within the messaging server system 108 is a design choice. For example, it may be technically preferable to initially deploy certain technologies and functions within the messaging server system 108 and later migrate that technology and functionality to the messaging client application 104 where the client device 102 has sufficient processing power.
[0024] The messaging server system 108 supports various services and operations provided to the messaging client application 104. Such operations include: sending data to the messaging client application 104, receiving data from the messaging client application 104, and processing data generated by the messaging client application 104. As an example, the data may include message content, client device information, geographical location information, media annotations and overlays, message content persistence conditions, social network information, and live event information. Data exchange within the messaging system 100 is invoked and controlled through functions accessible via the user interface of the messaging client application 104.
[0025] Turning now specifically to the messaging server system 108, an application programming interface (API) server 110 is coupled to an application server 112 and provides a programming interface to the application server 112. The application server 112 is communicatively coupled to a database server 118, which facilitates access to a database 120 in which data associated with messages processed by the application server 112 is stored.
[0026] The API server 110 receives and sends message data (e.g., commands and message payloads) between the client device 102 and the application server 112. Specifically, the API server 110 provides a collection of interfaces (e.g., routines and protocols) that can be invoked or queried by the messaging client application 104 to activate the functions of the application server 112. The API server 110 exposes various functions supported by the application server 112, including: account registration; login functionality; sending a message from a particular messaging client application 104 to another messaging client application 104 via the application server 112; sending a media file (e.g., an image or video) from the messaging client application 104 to the messaging server application 114 for possible access by another messaging client application 104; setting a collection of media data (e.g., a story); retrieving such a collection; retrieving a friend list of a user of the client device 102; retrieving messages and content; adding friends to and deleting friends from the social graph; locating friends within the social graph; and opening application events (e.g., involving the messaging client application 104).
[0027] The application server 112 hosts multiple applications and subsystems, including the messaging server application 114, the image processing system 116, the social networking system 122, the mapping system 123, and the motion recovery architecture engine. The messaging server application 114 implements many message processing techniques and functions, particularly those related to the aggregation and other processing of content (e.g., text and multimedia content) included in messages received from multiple instances of the messaging client application 104. As will be described in more detail, text and media content from multiple sources can be aggregated into a collection of content (e.g., referred to as a story or a library). The messaging server application 114 then makes these collections available to the messaging client application 104. Given the hardware requirements for other processor- and memory-intensive data processing, the messaging server application 114 may also perform such processing on the server side.
[0028] The application server 112 also includes an image processing system 116 dedicated to performing various image processing operations, typically on images or videos received within the payload of a message at the messaging server application 114.
[0029] The social networking system 122 supports various social networking functions and services and makes these functions and services available to the messaging server application 114. To this end, the social networking system 122 maintains and accesses an entity graph within the database 120 (e.g., Figure 3The entity figure 304) in it. Examples of functions and services supported by the social networking system 122 include the identification of other users of the messaging system 100 with whom a particular user has a relationship or whom a particular user "follows", and the identification of the interests of other entities and a particular user. The application server 112 is communicatively coupled to the database server 118, which facilitates access to the database 120 that stores data associated with messages processed by the messaging server application 114. The mapping system 123 is configured to generate a 3D map from different image sources, as discussed in further detail below.
[0030] Figure 2 is a block diagram showing additional details regarding the messaging system 100 according to an example embodiment. Specifically, the messaging system 100 is shown to include a messaging client application 104 and an application server 112, which in turn includes a plurality of subsystems, namely a transient timer system 202, a collection management system 204, and an annotation system 206.
[0031] The transient timer system 202 is responsible for implementing temporary access to the content allowed by the messaging client application 104 and the messaging server application 114. To this end, the transient timer system 202 includes a plurality of timers that selectively display messages and associated content via the messaging client application 104 and enable access to messages and associated content based on the duration and display parameters associated with a message or a collection of messages (e.g., a story). Additional details regarding the operation of the transient timer system 202 are provided below.
[0032] The collection management system 204 is responsible for managing collections of media (e.g., collections of text, images, video, and audio data). In some examples, collections of content (e.g., messages, including images, video, text, and audio) can be organized into "event libraries" or "event stories". Such collections can be made available for a specified period of time (such as the duration of the event to which the content pertains). For example, content related to a concert can be made available as a "story" during the duration of that concert. The collection management system 204 may also be responsible for publishing an icon that provides notification of the existence of a particular collection to the user interface of the messaging client application 104.
[0033] The collection management system 204 also includes a curation interface 208 that allows the collection manager to manage and curate specific collections of content. For example, the curation interface 208 enables an event organizer to curate a collection of content related to a specific event (e.g., delete inappropriate content or redundant messages). Additionally, the collection management system 204 employs machine vision (or image recognition technology) and content rules to automatically curate content collections. In some embodiments, compensation may be paid to users to include user-generated content in the collection. In such cases, the curation interface 208 operates to automatically pay such users for the use of their content.
[0034] The annotation system 206 provides various functions that enable users to annotate or otherwise modify or edit media content associated with a message. For example, the annotation system 206 provides functions related to generating and publishing media overlays for messages to be processed by the messaging system 100. The annotation system 206 operably supplies media overlays (e.g., geo-filters or filters) to the messaging client application 104 based on the geographical location of the client device 102. In another example, the annotation system 206 operably supplies media overlays to the messaging client application 104 based on other information, such as the social network information of the user of the client device 102. Media overlays can include audio and visual content as well as visual effects. Examples of audio and visual content include pictures, text, logos, animations, and sound effects. Examples of visual effects include color overlays. The audio and visual content or visual effects can be applied to media content items (e.g., photos) at the client device 102. For example, the media overlay includes text that can be overlaid on a photo generated by the client device 102. In another example, the media overlay includes a location identifier (e.g., Venice Beach), the name of a live event, or the name of a merchant (e.g., Beach Café). In another example, the annotation system 206 uses the geographical location of the client device 102 to identify a media overlay that includes the name of a merchant located at the geographical location of the client device 102. The media overlay can include other markers associated with the merchant. The media overlay can be stored in the database 120 and accessed via the database server 118.
[0035] In one example embodiment, the annotation system 206 provides a user-based publishing platform that enables users to select a geographical location on a map and upload content associated with the selected geographical location. The user can also specify the context in which the specific content should be provided to other users. The annotation system 206 generates a media overlay that includes the uploaded content and associates the uploaded content with the selected geographical location.
[0036] In another example embodiment, the annotation system 206 provides a merchant-based publishing platform that enables a merchant to select a specific media coverage associated with a geographical location via an auction process. For example, the annotation system 206 associates the media coverage of the highest bidder with the corresponding geographical location within a predefined amount of time.
[0037] Figure 3 is a schematic diagram showing data 300 that can be stored in the database 120 of the messaging server system 108 according to some example embodiments. Although the contents of the database 120 are shown as including multiple tables, it should be understood that the data 300 can be stored in other types of data structures (e.g., as an object-oriented database). The database 120 includes message data stored within a message table 314. An entity table 302 stores entity data, and the entity data includes an entity graph 304. Entities whose records are maintained within the entity table 302 can include individuals, corporate entities, organizations, objects, locations, events, etc. Regardless of type, any entity for which data is stored with respect to the messaging server system 108 can be an identified entity. Each entity is set with a unique identifier and an entity type identifier (not shown).
[0038] The entity graph 304 also stores information about the relationships and associations between or among entities. For example, such relationships can be social relationships, professional relationships (e.g., working in the same company or organization), interest-based or activity-based relationships.
[0039] The database 120 also stores annotation data in the annotation table 312 in the form of an example of a filter. The filter whose data is stored in the annotation table 312 is associated with a video (whose data is stored in the video table 310) and / or an image (whose data is stored in the image table 308), and is applied to the video (whose data is stored in the video table 310) and / or the image (whose data is stored in the image table 308). In one example, the filter is an overlay that is displayed as being overlaid on an image or a video during presentation to a recipient user. Filters can be of various types, including user-selected filters from a library of filters presented to a sending user by the messaging client application 104 when the sending user is composing a message. Other types of filters include location-based filters (also known as geo-filters), which can be presented to the sending user based on a geographical location. For example, based on geographical location information determined by the global positioning system (GPS) unit of the client device 102, the messaging client application 104 can present location-based filters specific to a neighborhood or a particular location within the user interface. Another type of filter is a data filter, which can be selectively presented to the sending user by the messaging client application 104 based on other input or information collected by the client device 102 during the message creation process. Examples of data filters include the current temperature at a specific location, the current speed at which the sending user is traveling, the battery life of the client device 102, or the current time.
[0040] Other annotation data that can be stored within the image table 308 is so-called "lens" data. A "lens" can be a real-time special effect and sound that can be added to an image or a video.
[0041] As mentioned above, the video table 310 stores video data, which in one embodiment is associated with messages whose records are maintained in the message table 314. Similarly, the image table 308 stores image data associated with messages whose message data is stored in the message table 314. The entity table 302 can associate various annotations from the annotation table 312 with various images and videos stored in the image table 308 and the video table 310.
[0042] The story table 306 stores data regarding a collection of messages and associated image, video, or audio data, where the messages and associated image, video, or audio data are compiled into a collection (e.g., a story or a library). The creation of a particular collection can be initiated by a particular user (e.g., each user whose record is maintained in the entity table 302). A user can create a "personal story" in the form of a collection of content that has already been created and sent / broadcast by that user. To this end, the user interface of the messaging client application 104 can include an icon that a user can select to enable the sending user to add specific content to his or her personal story.
[0043] The collection can also form a "Live Story", which is a collection of content from multiple users created manually, automatically, or using a combination of manual and automatic techniques. For example, a "Live Story" can form a curated stream of content submitted by users from different locations and events. Users whose client devices 102 have location services enabled and are at a common location or event at a particular time can be presented, for example, via the user interface of the messaging client application 104, with options to contribute content to a particular Live Story. The messaging client application 104 can identify the Live Story to him or her based on the user's location. The end result is a "Live Story" told from a community perspective.
[0044] Another type of content collection is called a "Location Story", which enables users whose client devices 102 are located within a particular geographical location (e.g., on a college or university campus) to contribute to a particular collection. In some embodiments, contributing to a Location Story may require a second degree of authentication to verify that the end user belongs to a particular organization or other entity (e.g., is a student on a university campus).
[0045] Figure 4 is a schematic diagram showing the structure of a message 400 according to some embodiments, which is generated by the messaging client application 104 to be sent to another messaging client application 104 or a messaging server application 114. The content of a particular message 400 is used to populate a message table 314 stored in a database 120 accessible to the messaging server application 114. Similarly, the content of the message 400 is stored in memory as "in-transit" or "in-flight" data on the client device 102 or the application server 112. The message 400 is shown as including the following components:
[0046] · Message identifier 402: A unique identifier that identifies the message 400.
[0047] · Message text payload 404: The text to be generated by the user via the user interface of the client device 102 and included in the message 400.
[0048] · Message image payload 406: Image data captured by the camera device component of the client device 102 or retrieved from the memory of the client device 102 and included in the message 400.
[0049] · Message video payload 408: Video data captured by the camera device component or retrieved from the memory component of the client device 102 and included in the message 400.
[0050] · Message audio payload 410: Audio data that is captured by a microphone or retrieved from a memory component of the client device 102 and included in the message 400.
[0051] · Message annotation 412: Annotation data (e.g., filters, stickers, or other enhancements) representing an annotation to be applied to the message image payload 406, the message video payload 408, or the message audio payload 410 of the message 400.
[0052] · Message duration parameter 414: A parameter value indicating the amount of time, in seconds, for which the content of the message 400 (e.g., the message image payload 406, the message video payload 408, and the message audio payload 410) will be presented to a user or made accessible to the user via the messaging client application 104.
[0053] · Message geographic location parameter 416: Geographic location data (e.g., latitude coordinates and longitude coordinates) associated with the content payload of the message 400. Multiple message geographic location parameter 416 values may be included in the payload, where each of these parameter values is associated with a corresponding content item included in the content (e.g., a specific image in the message image payload 406 or a specific video in the message video payload 408).
[0054] · Message story identifier 418: A value that identifies one or more content collections (e.g., “stories”) associated with a specific content item in the message image payload 406 of the message 400. For example, multiple images within the message image payload 406 may each be associated with multiple content collections using identifier values.
[0055] · Message tag 420: One or more tags, where each of the one or more tags indicates a theme of the content included in the message payload. For example, in the case where a specific image included in the message image payload 406 depicts an animal (e.g., a lion), a tag value may be included within the message tag 420 indicating the relevant animal. Tag values may be generated manually based on user input or may be generated automatically using, for example, image recognition.
[0056] · Message sender identifier 422: An identifier (e.g., a messaging system identifier, an email address, or a device identifier) indicating the user of the client device 102 on which the message 400 was generated and from which the message 400 was sent.
[0057] · Message recipient identifier 424: An identifier (e.g., a messaging system identifier, an email address, or a device identifier) indicating the user of the client device 102 to which the message 400 is addressed.
[0058] The content (e.g., values) of the various components of message 400 can be pointers to locations in a table where content data values are stored. For example, the image values in message image payload 406 can be pointers to locations (or addresses) within image table 308. Similarly, the values within message video payload 408 can point to data stored within video table 310, the values stored within message annotation 412 can point to data stored within annotation table 312, the values stored within message story identifier 418 can point to data stored within story table 306, and the values stored within message sender identifier 422 and message receiver identifier 424 can point to user records stored within entity table 302.
[0059] Figure 5 is a schematic diagram showing an access restriction process 500 according to some example embodiments, according to which access to content (e.g., a transient message 502 and the multimedia payload of associated data) or a collection of content (e.g., a transient message story 504) can be time - restricted (e.g., such that the access is transient).
[0060] The transient message 502 is shown as being associated with a message duration parameter 506, the value of which determines the amount of time that the messaging client application 104 will display the transient message 502 to the receiving user of the transient message 502. In one embodiment, where the messaging client application 104 is an application client, the receiving user can view the transient message 502 for up to 10 seconds, according to the amount of time specified by the sending user using the message duration parameter 506.
[0061] The message duration parameter 506 and the message receiver identifier 424 are shown as inputs to a message timer 512, which is responsible for determining the amount of time to show the transient message 502 to a specific receiving user identified by the message receiver identifier 424. In particular, the transient message 502 is shown to the relevant receiving user only for the period of time determined by the value of the message duration parameter 506. The message timer 512 is shown as providing an output to a more generalized transient timer system 202, which is responsible for the overall timing of displaying content (e.g., the transient message 502) to the receiving user.
[0062] Figure 5The ephemeral message 502 shown in [Figure 0] is included within an ephemeral message story 504 (e.g., a personal story or an event story). The ephemeral message story 504 has an associated story duration parameter 508, and the value of the story duration parameter 508 determines the duration for which the ephemeral message story 504 is presented and accessible to the users of the messaging system 100. For example, the story duration parameter 508 could be the duration of a concert, where the ephemeral message story 504 is a collection of content belonging to that concert. Alternatively, when setting up and creating the ephemeral message story 504, the user (either the owning user or a curator user) can specify the value of the story duration parameter 508.
[0063] In addition, each ephemeral message 502 within the ephemeral message story 504 has an associated story participation parameter 510, and the value of the story participation parameter 510 determines the duration for which the ephemeral message 502 will be accessible within the context of the ephemeral message story 504. Thus, a particular ephemeral message 502 can "expire" and become inaccessible within the context of the ephemeral message story 504 before the ephemeral message story 504 itself expires according to the story duration parameter 508.
[0064] The ephemeral timer system 202 can also operatively remove a particular ephemeral message 502 from the ephemeral message story 504 based on a determination that the associated story participation parameter 510 has been exceeded. For example, when the sending user has established a story participation parameter 510 of 24 hours from posting, the ephemeral timer system 202 will remove the associated ephemeral message 502 from the ephemeral message story 504 after the specified 24 hours. The ephemeral timer system 202 also operates to remove the ephemeral message story 504 when the story participation parameter 510 for each and every ephemeral message 502 within the ephemeral message story 504 has expired or when the ephemeral message story 504 itself has expired according to the story duration parameter 508.
[0065] In response to the ephemeral timer system 202 determining that the ephemeral message story 504 has expired (e.g., is no longer accessible), the ephemeral timer system 202 communicates with the messaging system 100 (e.g., specifically, the messaging client application 104) such that a marker (e.g., an icon) associated with the relevant ephemeral message story 504 is no longer displayed within the user interface of the messaging client application 104.
[0066] The following is an example implementation of the mapping system 123 according to some example embodiments. First, the mapping system 123 performs semantic segmentation classifier training (e.g., an image segmentation neural network) on frames from a video (such as frames from a video social media post or a video generated by a client device). For example, the mapping system 123 uses a machine learning classifier to generate programming labels for each frame, and the machine learning classifier is trained to generate image segmentation labels for ground-based images (e.g., image data generated from a ground-based imaging device such as a client device). Second, the mapping system 123 places the semantically labeled image data into the geospatial space by converting the image frames into a point cloud using a photogrammetry computer vision solution such as Alice Vision. One problem with converting to a point cloud is the inaccuracy of the geographic information contained in ground-based sources (e.g., side-view video, GPS data). In some example embodiments, to address the problem of insufficient accuracy, the mapping system 123 uses an external spatial reference dataset (e.g., aerial data) to assist in the geographic correction of the point cloud. In some example embodiments, the external reference data is provided in different forms such as an aerial-generated visual dataset. That is, to address the problem of insufficient accuracy, the mapping system 123 implements an external spatial reference dataset to assist in the geographic correction of the ground-generated point cloud. The external reference can take various forms such as aerial LiDAR, a point cloud derived from aerial or satellite-tilted imaging techniques, high-resolution synthetic aperture radar (SAR), or other remote sensing sources. In some example embodiments, the mapping system 123 preprocesses the point cloud by dividing the video frames into segments (e.g., five-second segments) and constructing a 3D point cloud in a photogrammetry pipeline (e.g., implementing a photometric imaging scheme, SfM). According to some example embodiments, an example photogrammetry pipeline includes (1) camera device startup, (2) followed by image feature extraction, (3) followed by image matching, (4) followed by future matching, (5) followed by performing structure from motion (SfM). SfM involves estimating the 3D structure of a scene from a collection of two-dimensional images. SfM data can be generated in different ways based on different factors such as the number and type of camera devices used, whether the images are sorted, and whether the images are taken from different camera devices (e.g., camera devices of different user devices).
[0067] In some example embodiments, to co-register the point cloud to an external spatial reference, the mapping system 123 maximizes the number of possible point matches between the external reference data and the ground-based point cloud via densification. In some example embodiments, an additional densification process is performed since the two data sets are different data sets collected from different perspectives (such as orthogonal viewpoints (e.g., side view and top-down view)). In some example embodiments, image registration (e.g., co-registration) between multiple images is an image processing technique for aligning multiple scenes into a single integrated image.
[0068] In some example embodiments, to address these difficulties, the mapping system 123 increases the number of potential points for matching by identifying (e.g., interpolating) aerial data sources and ground data sources. In some example embodiments, for the aerial data source, the mapping system 123 densifies the point data corresponding to the facade (e.g., side) of a building to improve the data provided from the air since vertical surfaces (e.g., building walls, sides) typically contain few generated points in a given collection area due to the top-down view of the data collection device. In some example embodiments, the mapping system 123 then performs semantic segmentation on the LiDAR or 3D point cloud to determine building structures and ground structures. In some example embodiments, the mapping system 123 implements a machine learning neural network that is trained on the semantically segmented point cloud to perform densification (e.g., neural network-based point interpolation to densify points and generate interpolated points).
[0069] In some example embodiments, the mapping system 123 then uses a ground densification pipeline to densify the ground photogrammetry-derived point cloud, the ground densification pipeline including: (1) depth mapping, (2) followed by depth map filtering, (3) followed by meshing, and (4) followed by mesh filtering. The result of the pipeline generates an improved set of ground point cloud candidates that can be more easily co-registered to a reference external spatial data set (e.g., an enhanced reference LiDAR point cloud from one or more aerial devices).
[0070] In some example embodiments, co-registration of the two data sets includes: first using GPS data to obtain telemetry for approximating the point cloud location, then performing odometry-based alignment of the SfM point cloud, and then performing pose graph optimization of the SfM point cloud to the external spatial reference data. In this way, the mapping system 123 locks the SfM-derived point cloud to a robust spatial reference. In some example embodiments, the mapping system 123 then stitches each SfM point cloud in the SfM point cloud together to create a geographical location (such as Figure 8A seamless panoramic picture scroll of a mixed 3D point cloud of an urban area of the city shown, where each ground point cloud is differently colored or shaded (for example, one client device generates an image of a building from the right, another client generates an image of the building from the left, etc.).
[0071] One advantage of the segmentation and co-registration process is that each individual point cloud used to generate the 3D map is relatively small, so it does not require a large amount of computation to obtain. In this way, when different point clouds are stitched together to generate a 3D map, the smaller SfM models ensure that errors do not propagate far (for example, into adjacent point clouds where errors may spread). In this way, the local error of each ground point cloud is never correlated with its adjacent SfM model (for example, the ground point cloud obtained from other client devices). Additionally, in some example embodiments, image fragments (for example, image segments of an image or video frame) from the image frames of the client device used to generate the ground point cloud are subsequently stitched (for example, projected, applied as surface texture) to Figure 8 the 3D map, making the 3D map have a more realistic photographic appearance. In this way, the mapping system generates high-accuracy point clouds from commercial camera devices (such as client device cameras) via enhancement from an external spatial reference data set.
[0072] Figure 6 FIG. 6 shows an example flowchart of a method 600 for using a mapping system 123 to generate a 3D map by registering different image sets (such as ground-based point clouds, airborne point clouds) generated from orthogonal perspectives according to some example embodiments. At operation 605, the mapping system 123 identifies a ground point cloud data set. For example, multiple client user devices generate video data of different parts of a geographical location (such as the urban area of a city), point clouds are generated from the video data of each client device, and the multiple point clouds are identified by the mapping system 123 for further processing.
[0073] At operation 610, the mapping system 123 identifies remote sensing data. At operation 615, the mapping system 123 enhances the ground point cloud data (for example, densifies the point cloud, neural network-based densification to add additional points between sparse points).
[0074] At operation 620, the mapping system 123 enhances a set of remote sensing data, such as airborne device data of a geographical location obtained from a top-down perspective. In some example embodiments, the mapping system 123 enhances the remote sensing data via interpolation (for example, densification) to densify the facade of vertical surfaces (such as the outer walls of buildings) captured in the remote sensing data.
[0075] At operation 625, the mapping system 123 generates a 3D map of the physical environment. In some example embodiments, the mapping system 123 generates the 3D map by co-registering an enhanced ground point cloud dataset to enhanced remote sensing data. Co-registration of each enhanced ground point cloud dataset stitches the patches together with each other and with the remote sensing data to create an accurate 3D map.
[0076] Figure 7 An example of different datasets (such as different image data from orthogonal perspectives) according to some example embodiments is shown. In the example top-down view 700, the remote sensing data includes data generated from a top-down perspective (e.g., images, videos, range data, point clouds). Different remote devices can provide the remote sensing data, such as an aircraft or satellite that physically moves above a geographical location and uses a LIDAR or imaging device (e.g., a complementary metal oxide semiconductor (CMOS) camera device) to generate the remote sensing data. In the example side view 750, the local data includes data generated by a ground device that generates data from a side perspective (e.g., images, videos, range data, point clouds). Different devices can provide the local data, such as a user device (e.g., a smart phone, a camera device, a vehicle-based imaging system, a range system such as a lidar).
[0077] Figure 8 An example 3D map 800 of a geographical area (e.g., downtown Boulder, Colorado) generated by the mapping system 123 according to some example embodiments is shown. In Figure 8 the example shown, the 3D map 800 is shaded with different patterns to indicate regions of the 3D map 800 corresponding to different point clouds generated from different ground devices (e.g., different end-user devices). The different point clouds depict buildings, streets, and other physical features of the geographical area (such as downtown Boulder, Colorado). For example, according to some example embodiments, a first user device (not depicted) generates video data when the user of the device is stationary or moving around the geographical area, and the video data is then used to create a first SfM-based point cloud region 805 via a photometric pipeline, where processing (e.g., enhancement and co-registration with remote sensing data) is implemented as described above.
[0078] In addition, a second user device (not depicted) generates video data while a second user of the second user device is stationary or moving around in a geographic area, and the second video data set is then used to create a second SfM-based point cloud region 810 (e.g., which is processed and enhanced via remote sensing data as described above). In addition, a third user device (not depicted) generates a third video data set while a third user of the third user device is stationary or moving around in a geographic area, and the third video data set is then used to create a third SfM-based point cloud region 815 (e.g., which is processed and enhanced via remote sensing data as described above).
[0079] In addition, a fourth user device (not depicted) generates a fourth video data set while a fourth user of the fourth user device is stationary or walking around in a geographic area, and the fourth video data set is then used to create a fourth SfM-based point cloud region 820 (e.g., which is processed and enhanced via remote sensing data as described above). The resulting point clouds can then be stitched together to create a map 800 (e.g., a complete 3D panoramic map of the geographic area). In some example embodiments, the video data set is a clip of a social media post (e.g., a short message), while in other example embodiments, the video data set includes video data created from off-the-shelf consumer imaging solutions such as video recorders, digital single-lens reflex (DSLR) camera devices, or mirrorless camera devices.
[0080] As described above, one benefit of the independently derived point cloud regions 805, 810, 815, and 820 is that errors are confined to the respective point cloud regions and do not spread to adjacent regions. This can be beneficial, for example, in cases where one of the point clouds is inaccurate (e.g., due to poor video or aerial data quality), while still being able to create a highly useful 3D map 800.
[0081] In some example embodiments, the 3D map 800 does not display the different patterns of the different point clouds, but instead applies image fragments (e.g., image segments) of the video data as image textures to the 3D map 800, such that the 3D map 800 appears more photo-realistic.
[0082] In some example embodiments, the mapping system 123 implements an improved structure from motion engine 127 to generate accurate 3D map result data such as 3D point cloud data and 3D pose data (e.g., image pose data predicting the camera frustum position for each image in the images used in the SfM process). In some example embodiments, SfM involves detecting and matching features between images or video frames, tracking features across frames, triangulating 3D points and 3D geometries to generate 3D points that map to a point cloud of a given structure o and the position poses of the images (e.g., video frames of video social media posts) used to create the map 800.
[0083] In some example embodiments, the structure-from-motion involves bundle adjustment optimization processing. Bundle adjustment involves the combined optimization of the geometry of a 3D point cloud that makes up a structure (e.g., a geographical location or a scene) and the 3D poses of the images used to generate the SfM result data. In particular, for example, bundle adjustment optimizes the poses (e.g., including the position of the images used to generate the SfM data and the pose data of one or more structures of a building). One constraint for implementing bundle adjustment optimization involves minimizing the reprojection error. The reprojection error involves reprojection of the generated 3D position data into the current best estimate of the image position and pose (e.g., an image, a video frame) to determine the degree of alignment of the points and the data. Due to the inherent noise in the processing, errors may occur in the detected point cloud pixel positions, and due to the error accumulation when the bundle adjustment process is applied to many images, the final 3D structure begins to distort or deviate from what it should look like (e.g., a distorted 3D street while the real-world 3D street is straight).
[0084] To address the foregoing issues, the structure-from-motion engine 127 is configured to reduce the noise that causes the reprojection error, thereby creating a more accurate 3D model. In some example embodiments, the structure-from-motion engine 127 integrates external reference data (e.g., remote sensing data) into the bundle adjustment optimization process of SfM as a guidance in the adjustment process to ensure that the generated 3D point data matches the externally provided data (e.g., ensuring that the SfM-3D generated map looks the same as the external reference data). In some example embodiments, the external reference data is the remote sensing data discussed above (e.g., airborne lidar data, airborne point cloud) or other types of data, such as the 3D mesh of the structures modeled via the mapping system 123 that is stored. In this way, the structure-from-motion engine 127 ensures that in terms of the result data generated by the mapping system 123, the 3D reference data should be as close to each other as possible: the output from both should be 3D point cloud data of the same triangular geometry and the 3D pose data of the same all input images (e.g., the images used by the mapping system 123 to generate the map 800). According to some example embodiments, the resulting 3D map can then be used for further end-user experiences, such as an augmented reality social media post that can be published on a social media site as a transient message (e.g., transient message 502). For example, the map data can include one or more 3D building models of a city downtown, and the 3D building models can be used to generate (e.g., as viewed through a client device) a social media post depicting a cartoon 3D dinosaur walking along the street and climbing up one of the buildings, and the user can record a video clip of the dinosaur and then post it on a social media networking site.
[0085] At a high level, structure from motion is a computational technique for estimating three-dimensional structure from a sequence of two-dimensional images. The output may include a 3D pose for each image and a 3D point cloud, where each 3D point in the point cloud has a description of its appearance in two or more images (e.g., position data for the point in the two or more images). According to some example embodiments, the structure from motion engine 127 generates structure from motion data (e.g., 3D maps, point cloud data, and image pose data) by: (a) detecting image features (e.g., corners) in each image of an image sequence (e.g., a video sequence of social media posts), (b) matching features between different images, (c) generating relative pose data for an image pair using the matched features, (d) starting with an initial image pair, triangulating image features to estimate the 3D positions of the image features, and (e) iteratively registering new images and refining the reconstruction. In some example embodiments, the iterative registration and optimization process includes operations including: (e1) matching image features against a current 3D point cloud, (e2) triangulating a new image to add new points to the 3D point cloud, and (e3) performing bundle adjustment by simultaneously optimizing image pose and 3D point positions (e.g., concurrently in a single loss process).
[0086] Bundle adjustment can be implemented as an optimization task in the SfM process. The bundle adjustment process can refine a set of camera parameters (e.g., image pose and camera intrinsics) and structure parameters (e.g., 3D point positions) to find a set of parameters that most accurately predicts the 3D point positions observed in the available image set. In particular, for example, suppose that n 3D points are seen in m views, and let x ij is the projection of the i-th point on image j. Let v ij represents a binary variable, if point i is visible in image j, then v ij is equal to 1, otherwise it is equal to 0. It is also assumed that each camera device j is represented by a vector a j parameterized, and each 3D point i is represented by a vector b i Parameterization. The bundle adjustment minimizes the total reprojection error with respect to all 3D points and camera parameters:
[0087]
[0088] Among them, Q(a j , b i ) is the predicted projection of point i on image j, and d(x, y) represents the distance metric between the projected image point and the measured image point. The distance metric can be a pure Euclidean distance, or a more robust loss function can be utilized.
[0089] As described above, the Structure from Motion (SfM) engine 127 implements bundle adjustment to minimize the reprojection error (e.g., the distance between a 3D point reprojected into the image where it was observed and the position of that observation within the image). In some example embodiments, this optimization can highly depend on the measured positions of features within the image. Noise in these estimated feature positions and other unmodeled error sources can cause long - range drift in the estimated image pose, resulting in distorted and geometrically inaccurate 3D reconstructions.
[0090] Figure 9 An illustration 900 of long - range drift in Structure from Motion (SfM) reconstruction according to some example embodiments is shown. In particular, the ground truth scene geometry 905 corresponds to real - world structures (e.g., buildings, streets). The Structure from Motion (SfM) engine 127 can receive input images that serve as a camera frustum 910 representing the reconstructed image pose. Then, the Structure from Motion (SfM) engine 127 generates reconstructed 3D points (e.g., a point cloud) that should match the ground truth scene geometry 905. However, as shown by the drift 920, a portion of the reconstructed points includes drift due to the distorted reconstruction. As shown, the reconstructed 3D points no longer accurately represent the ground truth scene or the ground truth scene geometry.
[0091] In some example embodiments, the Structure from Motion (SfM) engine 127 is configured to perform bundle adjustment such that the bundle adjustment process minimizes both the reprojection error and the distance between the 3D points and their counterparts in a reference geometry (e.g., aerial lidar). In some example embodiments, the Structure from Motion (SfM) engine 127 is configured to implement bundle adjustment as follows:
[0092]
[0093] where, c i is the reference geometry counterpart of 3D point i, and l(x, y) represents a distance metric between the 3D point and its counterpart in the reference geometry. In some example embodiments, the distance metric is the Euclidean distance. In some example embodiments, the distance metric is configured to implement a robust loss function to mitigate the effect of outlier correspondences (e.g., outlier corresponding points). In some example embodiments, the distance metric is configured to implement orientation data, shape data, and color information data for the correspondence and minimization process.
[0094] Figure 10FIG. 0 shows an example flowchart of a method 1000 for implementing a structure from motion to generate improved 3D data from image data according to some example embodiments. At 1005, a structure from motion engine 127 matches image data. For example, the structure from motion engine 127 detects image features (e.g., such as corner points) in each image and performs matching of features between different images.
[0095] At operation 1010, the structure from motion engine 127 generates pose data. For example, the structure from motion engine 127 estimates the relative pose of an image pair using feature matching (e.g., estimating to view geometry, such as the relative frustum pose between an image pair). At operation 1015, the structure from motion engine 127 generates point position data. For example, starting from an initial image pair, the structure from motion engine 127 triangulates image features to estimate the 3D positions of the image features to triangulate 3D points.
[0096] At operation 1020, the structure from motion engine 127 iteratively updates the generated 3D data. For example, at operation 1020, the structure from motion engine 127 iteratively registers new images (e.g., new video posts from a social media site) for each image to be processed and optimizes the resulting reconstruction. In some example embodiments, to perform the iterative update at operation 1020, the structure from motion engine 127 iteratively performs the following operations for each image: (a) matches image features to the current 3D point cloud, (b) triangulates the new image to add new points to the 3D point cloud, (c) detects correspondences between the reconstructed point cloud and a reference geometry. In some example embodiments, a combination of position, orientation, shape, and color is used to find correspondences, and (d) performs an extended bundle adjustment optimization, where a reference geometry (e.g., aerial lidar data, mesh reconstruction of a building) is used to minimize the reprojection error and the distance between 3D points and their correspondences (e.g., corresponding points on a building).
[0097] Figure 11Illustrated is an illustration 1100 for minimizing errors in SfM construction data. In the illustrated example, the reference geometry data 1105 is reference data of real-world geometry and structures (e.g., accurate aerial lidar point cloud). The structure from motion engine 127 implements user images as camera device frustums 1110 to represent the reconstructed image poses, where the structure from motion is used to implement the images to generate a reconstructed 3D point cloud 115, and the reconstructed 3D point cloud 115 has an inaccurate distorted shape 1120 due to errors 1125 (e.g., errors between the reconstructed 3D points in the reference geometry), as discussed above. In some example embodiments, the structure from motion engine 127 is configured to minimize both the reprojection error and the distance between the reconstructed 3D points in the corresponding reference geometry data 1105.
[0098] Figure 12 An example architecture for improved 3D data is illustrated. In the illustrated example, as discussed above (e.g., method 1000), the structure from motion engine 127 implements the reference data such that the camera device frustums (e.g., of the images) are consistent and there is good alignment between the reconstructed and iteratively updated 3D point cloud 1210 in the reference geometry 1215, thereby correcting for the distortion due to long-distance drift.
[0099] The following are example embodiments:
[0100] Example 1. A method, comprising: identifying ground source image data generated using a plurality of client devices; identifying aerial-based image data generated from an orthogonal perspective relative to the ground source image data; generating enhanced ground source image data by correlating points of the ground source image data and the aerial-based image data; generating a three-dimensional map from the enhanced ground source image data, the 3D map including a stitched portion of enhanced ground source image data sets from different client devices among the plurality of client devices.
[0101] Example 2. The method according to Example 1, wherein the ground source image data includes a plurality of point clouds.
[0102] Example 3. The method according to any one of Examples 1 or 2, wherein the plurality of point clouds are generated by applying an imaging scheme to a video sequence generated by the plurality of client devices.
[0103] Example 4. The method according to any one of Examples 1 to 3, wherein the imaging scheme is a photometric imaging scheme for generating point cloud data from image data.
[0104] Example 5. The method according to any one of Examples 1 to 4, wherein the aerial-based image data is generated by an aircraft.
[0105] Example 6. The method according to any one of Examples 1 to 9, wherein the aerial image data includes aerial lidar data that images the ground from a top-down perspective.
[0106] Example 7. The method according to any one of Examples 1 to 9, further comprising: enhancing the ground source image data by performing densification using interpolation to add image details.
[0107] Example 8. The method according to any one of Examples 1 to 9, wherein a machine learning scheme is trained to perform densification on the ground source image data.
[0108] Example 9. The method according to any one of Examples 1 to 9, further comprising: enhancing the ground source image data by performing densification using interpolation to add image details.
[0109] Example 10. The method according to any one of Examples 1 to 9, wherein making the points correlated includes applying a point co-registration scheme to correlate the ground source image data and the aerial image data.
[0110] Example 11. A system, comprising: one or more processors of a machine; and at least one memory storing instructions that, when executed by the one or more processors, cause the machine to perform any method according to Example 10.
[0111] Example 12. A machine storage medium containing instructions that, when executed by a machine, cause the machine to perform any of the methods according to Examples 1 to 10.
[0112] Example 13. A method, comprising: identifying a plurality of images from a user device, one or more of the plurality of images depicting a physical structure; identifying reference data mapped to the physical structure; using a first image set from the plurality of images and the reference data to generate an initial 3D model of the physical structure, the reference data being implemented to minimize an error between points generated in the initial 3D model and corresponding points in the reference data; using a second image set from the plurality of images and the reference data to generate an updated 3D model of the physical structure; and storing the updated 3D model of the physical structure.
[0113] Example 14. The method according to Example 13, wherein the reference data includes point cloud data of points related to other points corresponding to the physical structure.
[0114] Example 15. The method according to any one of Examples 13 to 14, wherein the point cloud data is generated from an aerial vehicle.
[0115] Example mesh data describing the shape of the physical structure.
[0116] Example 17. The method according to any one of Examples 13 to 16, wherein generating the initial 3D model includes applying a motion recovery structure scheme to a pair of images among the plurality of images.
[0117] Example 18. The method according to any one of Examples 13 to 17, wherein generating the updated 3D model of the physical structure using the second image set includes registering the second image set to the initial 3D model.
[0118] Example 19. The method according to any one of Examples 13 to 18, wherein generating the updated 3D model of the physical structure using the second image set further includes: optimizing the image pose data while optimizing the three-dimensional point positions.
[0119] Example 20. The method according to any one of Examples 13 to 19, further includes: sending the updated 3D model of the physical structure to one or more client devices, and the client devices implement the updated 3D model of the physical structure to generate an augmented reality content item.
[0120] Example 21. The method according to any one of Examples 13 to 20, further includes: receiving a transient message including the augmented reality content item from one or more client devices.
[0121] Example 22. The method according to any one of Examples 13 to 21, further includes: posting the transient message on a social networking site.
[0122] Example 23. A system, including: one or more processors of a machine; and
[0123] at least one memory storing instructions that, when executed by the one or more processors, cause the machine to perform any of the methods according to Examples 13 to 22.
[0124] Example 24. A machine storage medium containing instructions that, when executed by the machine, cause the machine to perform any of the methods according to Examples 13 to 22.
[0125] Figure 13 is a block diagram showing components of a machine 1300 capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium) and performing any one or more of the methods discussed herein. Specifically, Figure 13FIG. 1300 shows a graphical representation of a machine 1300 in the example form of a computer system within which instructions 1316 (e.g., software, program, application, applet, app, or other executable code) can be executed to cause the machine 1300 to perform any one or more of the methods discussed herein. Similarly, the instructions 1316 can be used to implement the modules or components described herein. The instructions 1316 transform the general, unprogrammed machine 1300 into a particular machine 1300 programmed to perform the described and illustrated functions in the described manner. In alternative embodiments, the machine 1300 operates as a stand-alone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1300 can operate in a server-client network environment as a server machine or a client machine, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1300 can include but is not limited to: a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smartphone, a mobile device, a wearable device (e.g., a smartwatch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of sequentially or otherwise executing the instructions 1316 specifying the actions to be taken by the machine 1300. Moreover, although only a single machine 1300 is shown, the term "machine" should also be taken to include a collection of machines that individually or jointly execute the instructions 1316 to perform any one or more of the methods discussed herein.
[0126] The machine 1300 can include a processor 1310, a memory / storage 1330, and input / output (I / O) components 1350, which can be configured to communicate with each other, such as via a bus 1302. The memory / storage 1330 can include a main memory 1332, a static memory 1334, and a storage unit 1336, all of which can be accessed by the processor 1310, such as via the bus 1302. The storage unit 1336 and the memory 1332 store the instructions 1316 embodying any one or more of the methods or functions described herein. The instructions 1316 can also reside, completely or partially, within the memory 1332, within the storage unit 1336 (e.g., on the machine-readable medium 1338), within at least one of the processors 1310 (e.g., within a processor cache memory accessible by the processor 1312 or 1314), or within any suitable combination thereof during execution by the machine 1300. Accordingly, the memory 1332, the storage unit 1336, and the memory of the processor 1310 are examples of machine-readable media.
[0127] The I / O component 1350 can include a variety of components that receive input, provide output, generate output, send information, exchange information, capture measurement results, and so on. The specific I / O components 1350 included in a particular machine 1300 will depend on the type of the machine. For example, a portable machine such as a mobile phone will most likely include a touch input device or other such input mechanism, while a headless server machine will most likely not include such a touch input device. It will be understood that the I / O component 1350 can include Figure 13 many other components not shown. Merely for the sake of simplifying the following discussion, the I / O components 1350 are grouped according to function, and the grouping is in no way restrictive. In various example embodiments, the I / O component 1350 can include an output component 1352 and an input component 1354. The output component 1352 can include visual components (e.g., a display such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), auditory components (e.g., speakers), tactile components (e.g., vibration motors, resistance mechanisms), other signal generators, and so on. The input component 1354 can include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input components), pointing-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instruments), tactile input components (e.g., physical buttons, a touch screen that provides the position and / or force of a touch or touch gesture, or other tactile input components), audio input components (e.g., a microphone), and so on.
[0128] In other example embodiments, the I / O component 1350 may include a biometric component 1356, a motion component 1358, an environmental component 1360, or a positioning component 1362, as well as a variety of other components. For example, the biometric component 1356 may include components for detecting expressions (e.g., hand expressions, facial expressions, vocal expressions, body postures, or eye tracking), measuring biometric signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (e.g., voice recognition, retina recognition, face recognition, fingerprint recognition, or electroencephalogram-based recognition), and so on. The motion component 1358 may include an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotational sensor component (e.g., a gyroscope), and the like. The environmental component 1360 may include, for example, a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers that detect the ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an auditory sensor component (e.g., one or more microphones that detect background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor (e.g., a gas sensor that detects the concentration of hazardous gases or measures pollutants in the atmosphere for safety reasons), or other components that can provide an indication, measurement, or signal corresponding to the surrounding physical environment. The positioning component 1362 may include a position sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects the air pressure from which the altitude can be derived), an orientation sensor component (e.g., a magnetometer), and the like.
[0129] A variety of techniques may be used to implement communication. The I / O component 1350 may include a communication component 1364 that is operable to couple the machine 1300 to the network 1380 or the device 1370 via the coupling 1382 and the coupling 1372, respectively. For example, the communication component 1364 may include a network interface component or other suitable device to interface with the network 1380. In other examples, the communication component 1364 may include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, components (e.g., low power ), components, and other communication components that provide communication via other modalities. The device 1370 may be another machine or any of a variety of peripheral devices (e.g., a peripheral device coupled via USB).
[0130] In addition, the communication component 1364 can detect an identifier or include components operable to detect an identifier. For example, the communication component 1364 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting the following items: one-dimensional barcodes such as Universal Product Code (UPC) barcodes; multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF418, Hypercode, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying an audio signal of a tag). Additionally, various information can be obtained via the communication component 1364, such as location via Internet Protocol (IP) geolocation, location via signal triangulation, location via detecting an NFC beacon signal that can indicate a specific location, etc.
[0131] The "carrier signal" in this context refers to any intangible medium capable of storing, encoding, or carrying the instructions 1316 executed by the machine 1300, and includes digital or analog communication signals or other intangible media to facilitate the communication of such instructions 1316. The instructions 1316 can be sent or received over the network 1380 via a network interface device using a transmission medium and any one of a plurality of well-known transmission protocols.
[0132] The "client device" in this context refers to any machine 1300 that interfaces with the network 1380 to obtain resources from one or more server systems or other client devices 102. The client device 102 can be, but is not limited to, a mobile phone, a desktop computer, a laptop computer, a PDA, a smart phone, a tablet computer, an ultrabook, a netbook, a multi-processor system, a microprocessor-based or programmable consumer electronics system, a game console, an STB, or any other communication device that a user can use to access the network 1380.
[0133] The "communication network" in this context refers to one or more portions of the network 1380, which can be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the public switched telephone network (PSTN), a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, A network, another type of network, or a combination of two or more such networks. For example, a part of network or network 1380 may include a wireless network or a cellular network, and the coupling 1382 may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or another type of cellular or wireless coupling. In this example, the coupling may implement any of various types of data transmission technologies, such as Single-Carrier Radio Transmission Technology (1xRTT), Evolution-Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, Enhanced Data Rates for GSM Evolution (EDGE) technology, the 3rd Generation Partnership Project (3GPP) including 3G, Fourth Generation Wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High-Speed Packet Access (HSPA), Worldwide Interoperability for Microwave Access (WiMAX), Long-Term Evolution (LTE) standard, other standards defined by various standards-setting organizations, other long-distance protocols, or other data transmission technologies.
[0134] A "transient message" in this context refers to a message 400 that can be accessed within a limited duration of time. The transient message 502 can be text, an image, a video, etc. The access time for the transient message 502 can be set by the message sender. Alternatively, the access time can be a default setting or a setting specified by the recipient. Regardless of the setting technique, the message 400 is transient.
[0135] A "machine-readable medium" in this context refers to a component, device, or other tangible medium that can store instructions 1316 and data temporarily or permanently, and may include, but is not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), buffer memory, flash memory, optical media, magnetic media, cache memory, other types of storage devices (e.g., Erasable Programmable Read-Only Memory (EPROM)), and / or any suitable combination thereof. The term "machine-readable medium" should be considered to include a single medium or multiple media that can store the instructions 1316 (e.g., a centralized or distributed database or an associated cache memory and server). The term "machine-readable medium" can also be considered to include any medium or combination of media that can store instructions 1316 (e.g., code) executable by the machine 1300 such that the instructions 1316, when executed by one or more processors 1310 of the machine 1300, cause the machine 1300 to perform any one or more of the methods described herein. Thus, a "machine-readable medium" refers to a single storage device or apparatus and a "cloud"-based storage system or storage network including multiple storage devices or apparatuses. The term "machine-readable medium" does not include a signal per se.
[0136] A "component" in this context refers to a logic or device, physical entity having boundaries defined by functions or subroutine calls, branch points, APIs, or other technical definitions that provide partitioning or modularization of a particular processing or control function. Components can be combined with other components via their interfaces to perform machine processing. A component can be an encapsulated functional hardware unit designed to work with other components, as well as part of a program for a specific function that generally performs related functions. Components can constitute software components (e.g., code included on a machine-readable medium) or hardware components.
[0137] A "hardware component" is a tangible unit capable of performing certain operations and can be configured or arranged in some physical manner. In various example embodiments, one or more computer systems (e.g., a stand-alone computer system, a client computer system, or a server computer system) or one or more hardware components of a computer system (e.g., processor 1312 or a group of processors 1310) can be configured by software (e.g., an application or an application part) to operate to perform certain operations described herein as a hardware component. A hardware component can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component can include dedicated circuitry or logic permanently configured to perform certain operations. A hardware component can be a dedicated processor, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). A hardware component can also include programmable logic or circuitry temporarily configured to perform certain operations by software. For example, a hardware component can include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a particular machine (or a particular component of machine 1300) uniquely customized to perform the configured function and is no longer a general-purpose processor 1310.
[0138] It should be understood that a decision can be made, for cost and time considerations, as to whether to implement a hardware component mechanically in dedicated and permanently configured circuitry or in circuitry that is temporarily configured (e.g., configured by software). Accordingly, the phrase "hardware component" (or "hardware-implemented component") should be understood to include a tangible entity, i.e., an entity physically constructed, permanently configured (e.g., hard-wired) or temporarily configured (e.g., programmed) to operate in some manner or to perform certain operations described herein.
[0139] Consider implementations in which hardware components are temporarily configured (e.g., programmed) without having to configure or instantiate each hardware component in the hardware components at any given time. For example, in the case where a hardware component includes a general-purpose processor 1312 that is configured by software to be a dedicated processor, the general-purpose processor 1312 can be configured at different times to be different dedicated processors (e.g., including different hardware components). Software accordingly configures a particular processor 1312 or processor 1310 to form a particular hardware component at one moment and different hardware components at different moments.
[0140] Hardware components can provide information to and receive information from other hardware components. Thus, the described hardware components can be considered to be communicatively coupled. In the case where multiple hardware components are present, communication can be achieved by signal transmission between or among two or more of the hardware components (e.g., via appropriate circuitry and buses). In implementations in which multiple hardware components are configured or instantiated at different times, such communication between or among hardware components can be achieved, for example, by storing information in a memory structure accessed by the multiple hardware components and retrieving the information from the memory structure. For example, one hardware component can perform an operation and store the output of the operation in a memory device communicatively coupled thereto. Then, another hardware component can access the memory device at a subsequent time to retrieve the stored output and process it. Hardware components can also initiate communication with input devices or output devices and can operate on resources (e.g., collections of information).
[0141] The various operations of the example methods described herein can be performed, at least in part, by one or more processors 1310 that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors 1310 can constitute processor-implemented components that operate to perform one or more of the operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using one or more processors 1310. Similarly, the methods described herein can be at least in part processor-implemented, where a particular processor 1312 or processor 1310 is an example of hardware. For example, at least some of the operations of the method can be performed by one or more processors 1310 or processor-implemented components. Additionally, one or more processors 1310 can also be operative to support the execution of relevant operations in a "cloud computing" environment or to operate as "software as a service" (SaaS). For example, at least some of the operations can be performed by a group of computers (as an example of machines 1300 that include processors 1310), where the operations can be accessed via a network 1380 (e.g., the Internet) and via one or more appropriate interfaces (e.g., an API). The execution of certain operations can be distributed among the processors 1310 and can reside not only within a single machine 1300 but also be deployed across multiple machines 1300. In some example embodiments, the processors 1310 or processor-implemented components can be located in a single geographical location (e.g., within a home environment, an office environment, or a server farm). In other example embodiments, the processors 1310 or processor-implemented components can be distributed across multiple geographical locations.
[0142] A "processor" in this context refers to any circuit or virtual circuit (a physical circuit simulated by logic executed on an actual processor 1312) that manipulates data values according to control signals (e.g., "commands", "opcodes", "machine codes", etc.) and produces corresponding output signals that are applied to operate the machine 1300. A processor can be, for example, a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio frequency integrated circuit (RFIC), or any combination thereof. The processor 1310 can also be a multi-core processor 1310 having two or more independent processors 1312, 1313 (sometimes referred to as "cores") that can execute instructions 1316 simultaneously.
[0143] A "timestamp" in this context refers to a sequence of characters or encoded information that identifies when an event occurred, e.g., giving a date and time of day, sometimes down to a fraction of a second.
Claims
1. A method, comprising: Accessing a plurality of images generated by a user device, one or more of the plurality of images depicting a physical structure; Identifying reference data mapped to the physical structure; Using a first image set from the plurality of images and the reference data to generate an initial 3D model of the physical structure, the reference data being implemented to minimize an error between points generated in the initial 3D model and corresponding points in the reference data; Using the reference data to generate an updated 3D model of the physical structure; and Storing the updated 3D model of the physical structure.
2. The method according to claim 1, wherein The reference data includes point cloud data of points related to other points corresponding to the physical structure.
3. The method according to claim 2, wherein The point cloud data is generated from an aerial vehicle.
4. The method according to claim 1, wherein, The reference data is model mesh data describing the shape of the physical structure.
5. The method according to claim 1, wherein, Generating the initial 3D model includes applying a structure from motion scheme to a pair of images from the plurality of images.
6. The method according to claim 1, wherein Generating the updated 3D model of the physical structure includes: Registering a second image set for the initial 3D model.
7. The method according to claim 6, wherein, Generating the updated 3D model of the physical structure further includes: Optimizing image pose data while optimizing three-dimensional point positions.
8. The method according to claim 1 further comprises: Sending the updated 3D model of the physical structure to one or more client devices, the client devices implementing the updated 3D model of the physical structure to generate an augmented reality content item.
9. The method according to claim 8, further comprising: Receiving a transient message including the augmented reality content item from the one or more client devices.
10. The method according to claim 9, further comprising: Posting the transient message on a social networking site.
11. A system, comprising: One or more processors of a machine; And At least one memory storing instructions that, when executed by the one or more processors, cause the machine to perform operations, the operations including: Accessing a plurality of images generated by a user device, one or more of the plurality of images depicting a physical structure; Identifying reference data mapped to the physical structure; Using a first image set from the plurality of images and the reference data to generate an initial 3D model of the physical structure, the reference data being implemented to minimize an error between points generated in the initial 3D model and corresponding points in the reference data; Using the reference data to generate an updated 3D model of the physical structure; and Storing the updated 3D model of the physical structure.
12. The system according to claim 11, wherein The reference data includes point cloud data of points related to other points corresponding to the physical structure.
13. The system according to claim 12, wherein, The point cloud data is generated from an aerial vehicle.
14. The system according to claim 11, wherein, The reference data is model mesh data describing the shape of the physical structure.
15. The system according to claim 11, wherein, Generating the initial 3D model includes applying a structure from motion scheme to a pair of images from the plurality of images.
16. The system according to claim 11, wherein, Generating the updated 3D model of the physical structure includes: Registering a second image set for the initial 3D model.
17. The system according to claim 16, wherein, Generating the updated 3D model of the physical structure further includes: Optimizing image pose data while optimizing three-dimensional point positions.
18. The system according to claim 11, further comprising: Send the updated 3D model of the physical structure to one or more client devices, which implement the updated 3D model of the physical structure to generate an augmented reality content item.
19. The system according to claim 18, further comprising: Receiving, from the one or more client devices, a transient message including the augmented reality content item.
20. A machine storage medium containing instructions that, when executed by a machine, cause the machine to perform operations, the operations including: Accessing a plurality of images generated by a user device, one or more of the plurality of images depicting a physical structure; Identifying reference data mapped to the physical structure; Using a first set of images from the plurality of images and the reference data to generate an initial 3D model of the physical structure, the reference data being implemented to minimize an error between points generated in the initial 3D model and corresponding points in the reference data; Using the reference data to generate an updated 3D model of the physical structure; and Storing the updated 3D model of the physical structure.
Citation Information
Cited By
Augmented three-dimensional structure generation
US12412338B2