3D object modeling based on photometry

By combining photometric measurement methods with optimized structural and camera device parameters, a 3D object model is generated, solving the problems of inaccurate modeling and high resource consumption in existing technologies, and achieving efficient and accurate 3D object reconstruction.

CN115769260BActive Publication Date: 2026-05-26SNAP INC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SNAP INC
Filing Date
2021-04-28
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing methods for reconstructing 3D geometry from 2D images have limited applicability on a wide range of images, consume a lot of resources, and are limited by changes in camera parameters and lighting conditions, resulting in inaccurate or incomplete modeling.

Method used

By employing photometric measurement methods and jointly optimizing structural parameters, camera device parameters, and lens distortion parameters, a 3D object model is generated through pixel correspondence optimization, reducing memory requirements and improving accuracy.

Benefits of technology

It improves the accuracy and efficiency of 3D object modeling, reduces resource consumption, and is suitable for large-scale image processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115769260B_ABST
    Figure CN115769260B_ABST
Patent Text Reader

Abstract

Various aspects of this disclosure relate to a system and method for performing operations including: accessing a source image depicting a target structure; accessing one or more target images depicting at least a portion of the target structure; calculating a correspondence between a first set of pixels in the source image of a first portion of the target structure and a second set of pixels in one or more target images of the first portion of the target structure, the correspondence being calculated based on camera device parameters varying between the source image and one or more target images; and generating a three-dimensional (3D) model of the target structure based on the correspondence between the first set of pixels in the source image and the second set of pixels in one or more target images, based on joint optimization of the target structure and camera device parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority requirements

[0002] This application claims priority to U.S. Patent Application No. 16 / 861,034, filed April 28, 2020, which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure generally relates to three-dimensional (3D) geometric reconstruction, and more particularly to generating 3D models of objects based on two-dimensional images. Background Technology

[0004] For applications such as digital asset generation and cultural preservation, offline reconstruction of 3D geometry from 2D images is a key task in computer vision. Typically, dense geometries are constructed using sophisticated algorithms such as multi-view stereo (MVS). Obtaining and analyzing features present in various 2D images allows for the resolution of the 3D geometry of objects depicted in the images. Attached Figure Description

[0005] In the accompanying drawings, the same reference numerals may describe similar parts in different views, and the drawings are not necessarily drawn to scale. For ease of identification of any particular element or action being discussed, one or more of the most significant digits in the reference numerals refer to the drawing number at which that element is first introduced. Some embodiments are shown in the drawings by way of example, not limitation, in which:

[0006] Figure 1 It is a diagrammatic representation of a network environment in which the content of this disclosure can be arranged, based on some examples.

[0007] Figure 2 It is a graphical representation of a messaging system with both client-side and server-side functionality, based on some examples.

[0008] Figure 3 It is a graphical representation of the data structures maintained in the database based on some examples.

[0009] Figure 4 It is a graphical representation based on some example messages.

[0010] Figure 5 This is a flowchart illustrating example operation of an enhanced system according to an example implementation.

[0011] Figure 6 It is a graphical representation of a machine in the form of a computer system, based on some examples, within which a set of instructions can be executed to cause the machine to perform any or more of the methods discussed herein.

[0012] Figure 7It is a block diagram showing a software architecture in which examples can be implemented. Detailed Implementation

[0013] The following description includes systems, methods, techniques, instruction sequences, and computer program products embodying illustrative embodiments of this disclosure. In the following description, numerous specific details are set forth for illustrative purposes to provide an understanding of various embodiments. However, it will be apparent to those skilled in the art that embodiments can be practiced without these specific details. Generally, well-known examples of instructions, protocols, structures, and techniques are not necessarily shown in detail.

[0014] Many typical methods can be used to reconstruct the geometry of a 3D object from one or more 2D images. Such methods operate by matching features or pixels to structures depicted in the images. While such methods generally work well, their fundamental assumptions limit their general applicability to a wide range of images and can lead to poor or inaccurate object modeling.

[0015] For example, feature-based SfM (Structure from Motion) methods and algorithms assume fixed camera device parameters and minimize the geometric error of features matched between two images. Specifically, such algorithms operate offline and jointly optimize the structure and camera device parameters. These algorithms model the scene structure using densely triangulated meshes that are regularized for smoothing. A texture map is inferred using the texture-to-image error, and then the geometry of the 3D object is determined based on the texture map. Inferring the texture map significantly increases the number of variables due to texture, as well as the correlation between variables due to mesh and smoothing regularization. Therefore, optimization is either performed alternately on different sets of variables (texture, structure, and camera device parameters) or using a simple first-order gradient descent solver. Due to these complexities, such algorithms consume significant system resources, take long times to achieve results, and have limited applicability, not being widely applicable to a wide range of applications, including real-time image processing. Some MVS methods optimize based on both the depth and normal of the landmarks but do not use joint optimization with camera device parameters, leading to relatively inaccurate or incomplete object modeling.

[0016] Another known image processing algorithm, visual ranging, can be used to generate 3D models of objects depicted in images. These algorithms use structure to compute correspondences between images, minimizing errors between them. In doing so, these algorithms avoid the complexity of inferring the texture of objects. Such algorithms model the structure using sparse, ray-based landmarks, but assume known and fixed camera device parameters (camera device internal parameters) to avoid optimizing lens parameters. While this improves the efficiency of modeling 3D objects, the applicability of these algorithms is severely limited.

[0017] These typical methods also assume constant brightness at scene points across all images. This assumption is also severely limiting in applications and leads to poor object modeling because lighting conditions typically change due to the time of day or year, or the movement of light sources. These lighting conditions can cause objects to appear differently in different images depending on when, where, and how the image was captured.

[0018] The disclosed implementation improves the accuracy and efficiency of using electronic devices by generating 3D object models based on variable structural parameters, camera device parameters, and lens distortion parameters using photometric methods. The disclosed method reduces the memory requirements for generating 3D object models and is more feasible for larger problems. Specifically, the disclosed implementation provides an accurate surface model of the 3D object for an efficient optimization problem. The disclosed implementation jointly optimizes both structural parameters and camera device parameters, and optimizes the surface normals of each landmark in the joint structural and camera device optimization. The disclosed implementation also uses inter-image errors to optimize the lens distortion parameters to provide an accurate 3D object model with higher accuracy and lower memory requirements than previous algorithms (e.g., SfM and MVS methods). To generate the 3D object model, the disclosed implementation defines an optimization problem that takes into account structural parameters, camera device parameters, and lens distortion parameters. This optimization problem is solved based on a cost function that associates pixels of a portion of a structure in the source image with pixels in the target image. This significantly improves how to model objects in 3D based on 2D images. In particular, this significantly improves the user experience, reduces the amount of resources required to complete 3D object modeling tasks, and enhances the overall efficiency and accuracy of the device.

[0019] In one example, the 3D coordinates of a portion of the structure are determined by applying undistorted parameters or functions of the camera device used to capture the source image to the pixels of that portion of the structure in the source image. Then, the corresponding pixels of that portion of the structure in the target image are identified using the 3D coordinates by applying distortion parameters to a line drawn from the 3D coordinates to the position of the camera device used to capture the target image. Normalization is applied to the pixel values ​​to account for different lighting conditions, and the difference between the normalized pixels in the source and target images is then determined. This difference is reduced or minimized to solve an optimization problem to generate a 3D model of the object.

[0020] Networked computing environment

[0021] Figure 1 This is a block diagram illustrating an example messaging system 100 for exchanging data (e.g., messages and associated content) over a network. Messaging system 100 includes multiple instances of client devices 102, each hosting multiple applications including message clients 104. Each message client 104 is communicatively coupled to a message server system 108 and other instances of message client 104 via a network 106 (e.g., the Internet).

[0022] The messaging client 104 is able to communicate and exchange data with another messaging client 104 and the messaging server system 108 via the network 106. The data exchanged between messaging clients 104 and between messaging client 104 and messaging server system 108 includes functions (e.g., commands to activate functions) and payload data (e.g., text, audio, video, or other multimedia data).

[0023] Message server system 108 provides server-side functionality to a specific message client 104 via network 106. While some functions of message system 100 are described herein as being performed by message client 104 or message server system 108, the location of certain functions within message client 104 or message server system 108 can be a design choice. For example, it is technically preferred that certain technologies and functions be initially deployed within message server system 108, but later migrated to message client 104, which has sufficient processing power on client device 102.

[0024] The message server system 108 supports various services and operations provided to the message client 104. Such operations include sending data to and receiving data from the message client 104, and processing data generated by the message client 104. As an example, this data may include message content, client device information, geolocation information, media enhancements and overlays, message content persistence conditions, social network information, and live event information. Data exchange within the message system 100 is activated and controlled via functions available through the user interface (UI) of the message client 104.

[0025] Specifically, turning to message server system 108, application programming interface (API) server 110 is coupled to application server 112 and provides a programming interface to application server 112. Application server 112 is communicatively coupled to database server 118, which provides easy access to database 120, storing data associated with messages processed by application server 112. Similarly, web server 124 is coupled to application server 112 and provides a web-based interface to application server 112. To this end, web server 124 processes incoming network requests via Hypertext Transfer Protocol (HTTP) and several other related protocols.

[0026] Application Programming Interface (API) server 110 receives and sends message data (e.g., commands and message payloads) between client device 102 and application server 112. Specifically, API server 110 provides a set of interfaces (e.g., routines and protocols) that can be invoked or queried by messaging client 104 to activate the functionality of application server 112. API server 110 exposes various functionalities supported by application server 112, including: account registration; login functionality; sending messages from one messaging client 104 to another messaging client 104 via application server 112; sending media files (e.g., images or videos) from messaging client 104 to messaging server 114 and enabling possible access by another messaging client 104; setting up media data sets (e.g., stories); retrieving a user's friend list for client device 102; retrieving such a set; retrieving messages and content; adding and deleting entities (e.g., friends) in an entity graph (e.g., a social graph); locating friends in a social graph; and opening application events (e.g., related to messaging client 104).

[0027] Application server 112 hosts several server applications and subsystems, including, for example, message server 114, image processing server 116, and social networking server 122. Message server 114 implements several message processing technologies and functions, particularly those related to the aggregation and other processing of content (e.g., text and multimedia content) included in messages received from multiple instances of message client 104. As will be described in further detail, text and media content from multiple sources can be aggregated into content collections (e.g., referred to as stories or galleries). These collections are then made available to message client 104. Given the hardware requirements for such processing, additional processor- and memory-intensive data processing can also be performed on the server side by message server 114.

[0028] Application server 112 also includes image processing server 116, which is dedicated to performing various image processing operations, typically relative to the image or video within the payload of a message sent from or received from message server 114.

[0029] Social networking server 122 supports various social networking functions and services and makes these functions and services available to message server 114. To this end, social networking server 122 maintains and accesses entity graph 306 (such as...) within database 120. Figure 3 (As shown). Examples of functions and services supported by the social network server 122 include identifying other users of the messaging system 100 with whom a particular user has a relationship or who is “following” them, as well as identifying the interests and other entities of a particular user.

[0030] System Architecture

[0031] Figure 2 This is a block diagram illustrating further details of a messaging system 100 according to some examples. Specifically, the messaging system 100 is shown as including a messaging client 104 and an application server 112. The messaging system 100 contains multiple subsystems, which are supported on the client side by the messaging client 104 and on the server side by the application server 112. These subsystems include, for example, a short-lived timer system 202, a collection management system 204, an enhancement system 206, a map system 208, and a game system 210.

[0032] The short-lived timer system 202 is responsible for forcing temporary or time-limited access to content by the message client 104 and the message server 114. The short-lived timer system 202 includes several timers that selectively enable access (e.g., for rendering and displaying) of messages and associated content via the message client 104 based on duration and display parameters associated with the message or message set (e.g., a story). Further details regarding the operation of the short-lived timer system 202 are provided below.

[0033] The collection management system 204 is responsible for managing groups or collections of media (e.g., collections of text, images, video, and audio data). Collections of content (e.g., messages, including images, videos, text, and audio) can be organized into "event libraries" or "event stories." Such collections can be made available for a specified time period (e.g., the duration of an event related to the content). For example, content related to a concert can be made available as a "story" for the duration of the concert. The collection management system 204 can also be responsible for publishing icons that notify the user interface of the messaging client 104 of the existence of a specific collection.

[0034] Furthermore, the collection management system 204 includes a curation interface 212, which enables collection managers to manage and curate specific content collections. For example, the curation interface 212 allows event organizers to curate content collections related to a specific event (e.g., removing inappropriate content or redundant messages). Additionally, the collection management system 204 employs machine vision (or image recognition technology) and content rules to automatically curate content collections. In some examples, users may be paid compensation for including user-generated content in the collection. In such cases, the collection management system 204 operates to automatically pay such users for using their content.

[0035] Enhancement system 206 provides various functionalities that enable users to enhance (e.g., annotate or otherwise modify or edit) media content associated with a message. For example, enhancement system 206 provides functionalities related to generating and publishing media overlays for messages processed by messaging system 100. Enhancement system 206 operatively provides media overlays or enhancements (e.g., image filters) to messaging client 104 based on the geolocation of client device 102. In another example, enhancement system 206 operatively provides media overlays to messaging client 104 based on other information such as the social network information of the user of client device 102. Media overlays may include audio and visual content as well as visual effects. Examples of audio and visual content include pictures, text, logos, animations, and sound effects. Examples of visual effects include color overlays. Audio and visual content or visual effects may be applied to media content items (e.g., photos) at client device 102. For example, media overlays may include text or images that can be overlaid on top of a photograph taken by client device 102. In another example, media overlays include location identifier overlays (e.g., Venice Beach), names of live events, or business names overlays (e.g., beach cafes). In another example, enhancement system 206 uses the geolocation of client device 102 to identify media overlays that include the name of a merchant at the geolocation of client device 102. The media overlay may include additional tags associated with the merchant. The media overlay may be stored in database 120 and accessed through database server 118.

[0036] In some examples, enhancement system 206 provides a user-based publishing platform that allows users to select geolocations on a map and upload content associated with those geolocations. Users can also specify which media overlays should be provided to other users. Enhancement system 206 generates a media overlay that includes the uploaded content and associates it with the selected geolocation.

[0037] In other examples, enhancement system 206 provides a merchant-based publishing platform that enables merchants to select specific media overlays associated with geolocation via a bidding process. For example, enhancement system 206 associates the media overlay of the highest bidder with a corresponding geolocation for a predefined amount of time.

[0038] In other examples, enhancement system 206 uses photometric techniques described below to generate 3D models of one or more structures or landmarks in an image. Specifically, enhancement system 206 determines the correspondence between pixels in a source image and pixels in a target image associated with a portion of a target structure or landmark. Enhancement system 206 solves an optimization problem defined as a function of the determined correspondences to generate 3D models of the structures or landmarks. In one implementation, enhancement system 206 performs 3D model generation in real time based on a set of frames received from a camera feed received from a camera device of client device 102. In another implementation, enhancement system 206 performs 3D model generation offline based on a set of images previously captured by one or more client devices (e.g., images stored on the Internet).

[0039] Map system 208 provides various geolocation functions and supports the presentation of map-based media content and messages by messaging client 104. For example, map system 208 can display (e.g., stored in profile data 308) user icons or avatars on the map to indicate the current or past locations of the user's "friends," as well as media content (e.g., a collection of messages including photos and videos) generated by these friends within the map's context. For example, on the map interface of messaging client 104, messages posted by a user from a specific geolocation to messaging system 100 can be displayed to the specific user's "friends" within the context of that specific location on the map. Users can also share their location and status information with other users of messaging system 100 via messaging client 104 (e.g., using appropriate status avatars), where the location and status information is displayed to selected users within the context of the map interface of messaging client 104.

[0040] The gaming system 210 provides various gaming functions within the context of the messaging client 104. The messaging client 104 provides a game interface that displays a list of available games, which can be initiated by a user within the context of the messaging client 104 and played with other users of the messaging system 100. The messaging system 100 also enables specific users to invite other users to play specific games by sending invitations from the messaging client 104 to such other users. The messaging client 104 also supports both voice and text communication (e.g., chat) within the gaming context, provides leaderboards for the game, and also supports providing in-game rewards (e.g., coins and items).

[0041] Data Architecture

[0042] Figure 3This is a schematic diagram illustrating a data structure 300 that can be stored in a database 120 of a message server system 108, according to certain examples. Although the contents of the database 120 are shown as including several tables, it should be understood that the data can be stored in other types of data structures (e.g., as an object-oriented database).

[0043] Database 120 includes message data stored in message table 302. For any given message, this message data includes at least message sender data, message receiver (or recipient) data, and a payload. See below for reference. Figure 4 Further details are provided regarding information that can be included in the message and in the message data stored in message table 302.

[0044] Entity table 304 stores entity data and (for example, links to entity diagram 306 and profile data 308). Entities for which records are maintained within entity table 304 may include individuals, company entities, organizations, objects, locations, events, etc. Regardless of entity type, any entity whose data is stored in message server system 108 can be an identifiable entity. Each entity is provided with a unique identifier and an entity type identifier (not shown).

[0045] Entity graph 306 stores information about the relationships and associations between entities. As an example only, such relationships can be social relationships based on interests or activities, or professional relationships (e.g., working in a common company or organization).

[0046] Profile data 308 stores various types of profile data about a specific entity. Based on privacy settings specified by the specific entity, profile data 308 can be selectively used and presented to other users of messaging system 100. In the case of an individual, profile data 308 includes, for example, a username, phone number, address, settings (e.g., notification and privacy settings), and an avatar representation (or a set of such avatar representations) selected by the user. A specific user can then selectively include one or more of these avatar representations in the content of messages transmitted via messaging system 100 and in a map interface displayed to other users by messaging client 104. The set of avatar representations may include “status avatars,” which present a graphical representation of a user’s chosen state or activity at a specific time.

[0047] In the case that the entity is a group, in addition to the group name, members and various settings of the associated group (e.g., notifications), the group profile data 308 may similarly include one or more avatars associated with the group.

[0048] Database 120 also stores enhancement data, such as overlays or filters, in enhancement table 310. Enhancement data is associated with video (for which data is stored in video table 314) and images (for which data is stored in image table 316) and applied to video and images.

[0049] In one example, the filter is displayed as an overlay on an image or video during presentation to the receiving user. Filters can be of various types, including filters selected by the user from a set of filters presented to the sending user by the messaging client 104 when the sending user is composing a message. Other types of filters include geolocation filters (also known as geographic filters), which can be presented to the sending user based on geographic location. For example, a geolocation filter specific to a nearby or specific location can be presented by the messaging client 104 within the user interface based on geolocation information determined by the Global Positioning System (GPS) unit of the client device 102.

[0050] Another type of filter is a data filter, which can be selectively presented to the sending user by the messaging client 104 based on other inputs or information collected by the client device 102 during the message creation process. Examples of data filters include the current temperature at a specific location, the current speed of the sending user, the battery life of the client device 102, or the current time.

[0051] Other augmented data that can be stored in image table 316 includes augmented reality content items (e.g., corresponding to an application lens or augmented reality experience). Augmented reality content items can be real-time effects and sounds that can be added to images or videos.

[0052] As described above, augmented data includes augmented reality content items, overlays, image transformations, AR images, and similar items referring to modifications that can be applied to image data (e.g., video or images). This includes real-time modifications, which modify images as they are captured using the device sensors (e.g., one or more cameras) of client device 102 and then display the modified image on the screen of client device 102. This also includes modifications to stored content, such as modifications to video clips in a library that can be modified. For example, in client device 102, which has access to multiple augmented reality content items, a user can use a single video clip with multiple augmented reality content items to see how different augmented reality content items will modify the stored clip. For example, by selecting different augmented reality content items for the same content, multiple augmented reality content items with different pseudo-random motion models can be applied to that same content. Similarly, real-time video capture can be used with the modifications shown to demonstrate how the video image currently captured by the sensors of client device 102 will modify the captured data. Such data can be simply displayed on the screen without being stored in memory, or content captured by the device's sensors can be recorded and stored in memory with or without modification (or both). In some systems, preview functionality can show how different augmented reality content items will be displayed simultaneously in different windows of the display. For example, this allows viewing multiple windows with different pseudo-random animations on the display at the same time.

[0053] Therefore, using augmented reality content items and various systems, or other such transformation systems that modify content using that data, can involve the detection of objects (e.g., faces, hands, bodies, cats, dogs, surfaces, objects, etc.) in video frames, tracking these objects as they leave, enter, and move around within the field of view, and modifying or transforming them while tracking them. In various implementations, different methods can be used to achieve such transformations. Some examples may involve generating a 3D mesh model of one or more objects, and using transformations and animated textures of the model within the video to achieve the transformation. In other examples, tracking points on the objects can be used to place an image or texture (which can be 2D or 3D) at the tracked location. In further examples, neural network analysis of video frames can be used to place images, models, or textures within content (e.g., images or video frames). Thus, augmented reality content items refer both to the images, models, and textures used to create transformations within the content, and to the additional modeling and analysis information required to achieve such transformations through object detection, tracking, and placement.

[0054] Real-time video processing can be performed using any kind of video data (e.g., video streams, video files, etc.) stored in the memory of any type of computerized system. For example, a user can load a video file and store it in the device's memory, or a video stream can be generated using the device's sensors. Furthermore, computer-animated models can be used to process any object, such as a human face and parts of the human body, animals, or inanimate objects (e.g., chairs, cars, or other objects).

[0055] In some examples, when a specific modification is selected along with the content to be transformed, the elements to be transformed are identified by a computing device and then detected and tracked if they exist in the frames of the video. The elements of the object are modified according to the modification request, thereby transforming the frames of the video stream. For different types of transformations, the transformation of the video stream frames can be performed using different methods. For example, for frame transformations that primarily refer to changes in the form of the object's elements, feature points of each element of the object are calculated (e.g., using an Active Shape Model (ASM) or other known methods). Then, a feature point-based mesh is generated for each of at least one element of the object. This mesh is used for subsequent stages of tracking the elements of the object in the video stream. During tracking, the mesh mentioned for each element is aligned with the position of each element. Then, additional points are generated on the mesh. A first set of points is generated for each element based on the modification request, and a second set of points is generated for each element based on the first set of points and the modification request. The frames of the video stream can then be transformed by modifying the elements of the object based on the first and second set of points and the mesh. In this method, the background of the object being modified can also be altered or distorted by tracking and modifying the background.

[0056] In some examples, transformations of certain regions of an object using its elements can be performed by calculating feature points for each element of the object and generating a mesh based on those calculated feature points. Points are generated on the mesh, and various regions are then generated based on these points. The elements of the object are then tracked by aligning the regions of each element with the positions of at least one of the elements, and the properties of the regions can be modified based on modification requests, thereby transforming frames of the video stream. Depending on the specific modification request, the properties of the mentioned regions can be transformed in different ways. Such modifications can involve: changing the color of the region; removing at least a portion of the region from the frames of the video stream; including one or more new objects in the region based on the modification request; and modifying or distorting the elements of the region or object. In various implementations, any combination of such modifications or other similar modifications can be used. For certain models to be animated, some feature points can be selected as control points to determine the entire state space for options used in model animation.

[0057] In some examples of computer animation models that use face detection to transform image data, a specific face detection algorithm (e.g., Viola-Jones) is used to detect faces in the image. The Active Shape Model (ASM) algorithm is then applied to the facial regions of the image to detect facial feature reference points.

[0058] In other examples, other methods and algorithms suitable for face detection can be used. For example, in some implementations, landmarks are used to locate features that represent distinguishable points present in most of the images considered. For example, for a face landmark, the location of the left eye pupil could be used. If the initial landmark is not recognizable (e.g., if the person is wearing an eye patch), secondary landmarks can be used. Such a landmark recognition process can be used for any such object. In some examples, a set of landmarks forms a shape. The shape can be represented as a vector using the coordinates of points in the shape. One shape is aligned with another shape using a similarity transformation (allowing translation, scaling, and rotation) that minimizes the average Euclidean distance between the points of the shapes. The average shape is the mean of the aligned training shapes.

[0059] In some examples, a landmark search begins with an average shape aligned with the position and size of a face determined by a global face detector. This search then repeats the following steps: proposing provisional shapes by adjusting the localization of shape points through template matching of the image texture around each point, and then conforming the provisional shapes to a global shape model until convergence occurs. In some systems, individual template matching is unreliable, and the shape model pools the results of weak template matching to form a stronger overall classifier. The entire search is repeated at each level of the image pyramid, from coarse to fine resolution.

[0060] The transformation system can capture image or video streams on a client device (e.g., client device 102) and perform complex image manipulations locally on client device 102 while maintaining an appropriate user experience, computation time, and power consumption. Complex image manipulations can include size and shape changes, emotion shifts (e.g., changing a face from frowning to smiling), state shifts (e.g., aging an object, reducing its apparent age, or changing its gender), style shifts, application of graphical elements, and any other suitable image or video manipulations implemented by a convolutional neural network that has been configured to execute efficiently on client device 102.

[0061] In some examples, a computer animation model for transforming image data can be used by a system in which a user can use a client device 102 with a neural network to capture an image or video stream of the user (e.g., a selfie), the neural network operation being part of a messaging client 104 operating on the client device 102. A transformation system operating within the messaging client 104 determines the presence of a face in the image or video stream and provides a modification icon associated with the computer animation model to transform the image data, or the computer animation model can be presented as associated with the interface described herein. The modification icon includes changes that can be the basis for modifying the user's face in the image or video stream as part of a modification operation. Once a modification icon is selected, the transformation system initiates a process of transforming the user's image to reflect the selected modification icon (e.g., generating a smiley face on the user). Once the image or video stream is captured and the specified modification is selected, the modified image or video stream can be presented in a graphical user interface displayed on the client device 102. The transformation system can implement a complex convolutional neural network on a portion of the image or video stream to generate and apply the selected modification. In other words, once an edit icon is selected, the user can capture an image or video stream and see the changes in real-time or near real-time. Furthermore, while a video stream is being captured, the changes can be persistent, and the selected edit icon continues to be toggled. Machine learning neural networks can be used to achieve this type of modification.

[0062] A graphical user interface (GUI) presenting modifications performed by the transformation system can provide users with additional interactive options. Such options can be based on an interface used to initiate content capture and selection for a specific computer animation model (e.g., initiated from a content creator user interface). In various implementations, modifications can be persistent after an initial selection of the modification icon. Users can turn modifications on or off by tapping or otherwise selecting a face modified by the transformation system and save it for later viewing or browsing to other areas of the imaging application. In the case of multiple faces being modified by the transformation system, users can globally turn modifications on or off by tapping or selecting a single face modified and displayed within the GUI. In some implementations, individual faces within a set of multiple faces can be modified separately, or such modifications can be toggled individually by tapping or selecting a single face or a series of individual faces displayed within the GUI.

[0063] Story table 312 stores data about collections of messages and associated image, video, or audio data, compiled into collections (e.g., stories or libraries). The creation of a specific collection can be initiated by a specific user (e.g., each user whose records are maintained in entity table 304). A user can create a "personal story" in the form of a collection of content that has already been created and sent / broadcast by that user. For this purpose, the user interface of messaging client 104 may include user-selectable icons that allow the sending user to add specific content to his or her personal story.

[0064] The collection can also constitute a "live story" as a collection of content from multiple users, created manually, automatically, or using a combination of manual and automatic techniques. For example, a "live story" can constitute a curated flow of user-submitted content from various locations and events. Users whose client devices have location services enabled and are at a co-location event at a specific time can be presented with options, for example, via the user interface of messaging client 104, to contribute content to a specific live story. The live story can be identified to the user by messaging client 104 based on their location. The end result is a "live story" told from a community perspective.

[0065] Another type of content collection is called a "location story," which allows users whose client devices 102 are located in a specific geographic location (e.g., on a college or university campus) to contribute to a specific collection. In some examples, contributing to a location story may require secondary authentication to verify that the end user belongs to a specific organization or other entity (e.g., a student on a university campus).

[0066] As mentioned above, video table 314 stores video data, which, in one example, is associated with a message whose record is stored in message table 302. Similarly, image table 316 stores image data, which is associated with a message whose message data is stored in entity table 304. Entity table 304 allows various enhancements from enhancement table 310 to be associated with various images and videos stored in image table 316 and video table 314.

[0067] Data communication architecture

[0068] Figure 4This is a schematic diagram illustrating the structure of message 400 according to some examples, generated by message client 104 for transmission to another message client 104 or message server 114. The content of a particular message 400 is used to populate message table 302 stored in database 120, which is accessible by message server 114. Similarly, the content of message 400 is stored in memory as "in transit" or "in flight" data for client device 102 or application server 112. Message 400 is shown to include the following example components.

[0069] • Message Identifier 402: A unique identifier that identifies message 400.

[0070] • Message text payload 404: The text to be generated by the user via the user interface of the client device 102 and included in message 400.

[0071] • Message image payload 406: Image data captured by the camera component of the client device 102 or retrieved from the memory component of the client device 102 and included in the message 400. The image data of the message 400 used for sending or receiving can be stored in the image table 316.

[0072] • Message video payload 408: Video data captured by the camera device component or retrieved from the memory component of the client device 102 and included in the message 400. The video data used to send or receive the message 400 can be stored in the video table 314.

[0073] • Message audio payload 410: Audio data captured by the microphone or retrieved from the memory component of the client device 102 and included in message 400.

[0074] • Message enhancement data 412: Enhancement data (e.g., filters, stickers, or other annotations or enhancements) representing enhancements to be applied to the message image payload 406, message video payload 408, or message audio payload 410 of message 400. Enhancement data for sending or receiving message 400 can be stored in enhancement table 310.

[0075] • Message duration parameter 414: A parameter value that indicates the amount of time, in seconds, during which the content of the message (e.g., message image payload 406, message video payload 408, message audio payload 410) will be presented to the user or made accessible to the user via the message client 104.

[0076] • Message geolocation parameter 416: Geolocation data (e.g., latitude and longitude coordinates) associated with the content payload of the message. The payload may include multiple message geolocation parameter 416 values, each of which is associated with a content item included in the content (e.g., a specific image within the message image payload 406, or a specific video within the message video payload 408).

[0077] • Message Story Identifier 418: An identifier value that identifies one or more content sets (e.g., "Stories" identified in Story Table 312) by associating a specific content item in the message image payload 406 of message 400 with one or more content sets. For example, the identifier value can be used to associate multiple images within the message image payload 406 with multiple content sets, respectively.

[0078] • Message Tag 420: Each message 400 can be labeled with multiple tags, each of which indicates the subject of the content included in the message payload. For example, in the case where a specific image in the message image payload 406 depicts an animal (e.g., a lion), tag values ​​can be included in the message tag 420 indicating the relevant animal. Tag values ​​can be generated manually based on user input or automatically using, for example, image recognition.

[0079] • Message sender identifier 422: An identifier (e.g., a messaging system identifier, email address, or device identifier) ​​indicating the user of the client device 102 on which message 400 is generated and from which message 400 is sent.

[0080] • Message recipient identifier 424: An identifier (e.g., message system identifier, email address, or device identifier) ​​indicating the user of the client device 102 to which message 400 is addressed.

[0081] The contents (e.g., values) of the various components of message 400 can be pointers to locations in tables of their stored content data values. For example, image values ​​in message image payload 406 can be pointers to locations (or addresses of those locations) within image table 316. Similarly, values ​​in message video payload 408 can point to data stored in video table 314, values ​​stored in message enhancement data 412 can point to data stored in enhancement table 310, values ​​stored in message story identifier 418 can point to data stored in story table 312, and values ​​stored in message sender identifier 422 and message receiver identifier 424 can point to user records stored in entity table 304.

[0082] Figure 5This is a flowchart illustrating example operation of the enhanced system 206 during the execution of process 500 according to an exemplary embodiment. Process 500 can be implemented with computer-readable instructions executable by one or more processors, such that the operation of process 500 can be performed partially or entirely by functional components of message server system 108; therefore, process 500 is described below by way of example with reference to it. However, in other embodiments, at least some of the operations of process 500 can be arranged on various other hardware configurations. Therefore, process 500 is not intended to be limited to message server system 108 and can be implemented wholly or partially by any other component. Operations in process 500 can be executed in any order, in parallel, or can be completely skipped and omitted.

[0083] At operation 501, enhancement system 206 accesses source images depicting the target structure. For example, enhancement system 206 receives field or real-time camera feeds from the camera device of client device 102. As another example, enhancement system 206 retrieves previously captured images of structures or landmarks from the Internet or other local or remote sources. Enhancement system 206 selects the first image depicting the structure of interest from the received images as the source image. In some cases, enhancement system 206 selects the source image by processing a set of images and identifying which image has the highest visibility of the structure of interest.

[0084] As an example, enhancement system 206 generates a 3D coordinate frame for the target structure. That is, enhancement system 206 selects a region of the structure and generates a set of 3D coordinates (world coordinates) for that region. Enhancement system 206 calculates the visibility of the 3D coordinate frame over a set of images serving as depth maps and selects one of the images in the set as the source image based on the calculated visibility. In some implementations, enhancement system 206 calculates a pixel grid with specific intervals corresponding to the 3D coordinate frame. For example, enhancement system 206 projects the 3D coordinate frame onto each image in the image set to identify a set of pixels in each image that correspond to the 3D coordinate frame. That is, enhancement system 206 identifies pixel coordinates in each image that correspond to the 3D coordinates of the region of the structure. Enhancement system 206 samples the image set associated with the pixel grid to generate a matrix in which each column corresponds to a different image in the image set. Specifically, enhancement system 206 obtains the pixel value of the identified pixel coordinates for each image and stores the pixel value in the corresponding column of the matrix. Then, the enhancement system 206 calculates the solution for the mean, weighted mean, or robust sum of squares of each column of the matrix. The enhancement system 206 compares each column of the matrix with the calculated solution for the mean, weighted mean, or robust sum of squares, and selects the image from the image set whose corresponding column is closest in value to the calculated solution for the mean, weighted mean, or robust sum of squares as the source image.

[0085] In some implementations, the source frame index is determined by I k This indicates that the visibility of each landmark is determined by V. k These can be kept constant throughout the optimization. Poisson surface reconstruction is performed on the selected landmarks, and the resulting mesh is used to calculate visibility. The mesh is rendered into each view (each image in the image set) as a depth map, and the landmarks are projected into the views. Their depths are compared to the depth maps, and depths differing by less than 1% are determined to be visible. To avoid selecting source frames as photometric outliers (e.g., due to specular reflection), I k The frame whose tiles are closest to the robust average of the visible, normalized tiles is selected. The source frame is selected according to Equation 1 below:

[0086]

[0087] Among them, X k It is a 3xN world coordinate (3D coordinate) matrix, using a 4x4 point grid, spaced such that the average spacing of the visible view around the Kth landmark is 1 pixel. Starting with an unrobust average, μ is computed using iterative weighted least squares. To ensure that landmarks are initialized only in textured image regions, remove... (Assuming there are 256 gray levels) the boundary markers. The remaining functions of Equation 1 are described in more detail below. Specifically, The projection of the structure's 3D coordinates onto pixels in a given image j, taking into account camera, lens, and sensor parameters, is defined and described in more detail below. In some cases, to improve the convergence of photometric optimization to a good optimum, the optimization is initially run on a half-size image first, followed by a full-size image.

[0088] At operation 502, enhancement system 206 accesses one or more target images depicting at least a portion of the target structure. For example, after identifying a given source image, a target image depicting the same structure (e.g., a different view of the structure) is selected. The target image may be an image frame subsequently received after the source image in a real-time camera feed. The target frame may be a random frame selected from a set of images depicting the target structure.

[0089] In some implementations, enhancement system 206 upsamples or downsamples the target image based on the difference in the number of pixels corresponding to a portion of the target structure. For example, the source image may capture that portion of the target structure from a closer distance than the target image. In this case, the source image will have a greater number of pixels corresponding to that portion of the target structure than the target image. To ensure that the portion of the target structure represents the same number of pixels in both the source and target images, enhancement system 206 upsamples the target image to increase the number of pixels representing the target structure to match the number of pixels representing the target structure in the source image. In one implementation, enhancement system 206 identifies a first set of pixels in the source image corresponding to that portion of the target structure and a second set of pixels in the target image corresponding to that portion of the target structure. Enhancement system 206 may use metadata associated with the image to determine which pixels correspond to which portion of the structure. Enhancement system 206 calculates a first distance between each pixel in the first set of pixels and a second distance between each pixel in the second set of pixels. Enhancement system 206 selects sampling parameters based on the difference between the first and second distances. For example, if the first distance is greater than the second distance, enhancement system 206 selects a value for upsampling the target image. For example, if the first distance is less than the second distance, the enhancement system 206 selects a value for downsampling the target image. The enhancement system 206 applies the sampling parameters to the target image to upsample or downsample it.

[0090] At operation 503, enhancement system 206 calculates a correspondence between a first set of pixels in a source image of a first portion of the target structure and a second set of pixels in one or more target images of the first portion of the target structure, the correspondence being calculated based on camera device parameters that vary between the source image and the one or more target images. For example, as modeled by Equation 5 below, enhancement system 206 projects a view of a portion or the entire structure onto the source image. That is, enhancement system 206 draws rays or lines from the camera device position and pixels of a portion of the structure in the source image to identify the corresponding 3D world coordinates of that portion of the structure in the source image. When drawing rays or lines, enhancement system 206 bends the line based on the lens parameters of the camera device (e.g., applying the non-distortion parameters of the lens of the camera device used to capture the source image—pixels are non-distorted as they leave the lens of the camera device), so that the 3D world coordinates are accurately identified by the view from the camera device used to capture the source image. Specifically, the enhancement system 206 applies camera device parameters, lens parameters, and sensor parameters to the pixels in the source image to determine how light will reach the 3D world coordinates of the structure when viewed by the camera device used to capture the source image.

[0091] Next, using Equation 2 below, the enhancement system 206 draws rays or lines from the 3D world coordinates of that portion of the structure toward the camera device capturing the target image to identify pixels within the target image that represent the 3D world coordinates of that portion of the structure identified using the source image. When drawing the rays or lines, the enhancement system 206 bends the lines based on the lens parameters of the camera device (e.g., applying distortion parameters of the lens of the camera device used to capture the target image—distorting the 3D coordinates through the lens of the camera device to generate pixels in the target image), so that the 3D world coordinates are accurately represented by the view from the camera device used to capture the target image. Specifically, the enhancement system 206 applies camera device parameters, lens parameters, and sensor parameters to the 3D world coordinates to determine how light will reach the 3D world coordinates of the structure when viewed through the camera device used to capture the target image, in order to identify pixels in the target image corresponding to the 3D world coordinates of the structure.

[0092] A set of camera device parameters defines 3D points in world coordinates. The projection onto the image plane in pixel coordinates. The first set of camera device parameters includes camera device extrinsic parameters consisting of a pair of P rotations and translations for each image. The second set of camera device parameters includes the following intrinsic parameters of the camera device: a pair of C-sensor and lens calibration parameters for each camera device. Where C <= P. When C is less than P, some camera intrinsics are shared across the input images. In such cases, the index mapped from image i to camera j may need to be used as input. The projection of the 3D structure position onto the camera view is defined by Equation 2 below:

[0093]

[0094] in, It is the projection function:

[0095] It is a lens distortion function

[0096] Number, and The sensor calibration function is defined by Equation 3 below:

[0097]

[0098] For lens distortion to be applied to identify pixel coordinates corresponding to 3D structural coordinates (e.g., identifying how light bends by passing light from 3D coordinates through a lens to generate target image pixels), the standard polynomial radial distortion model defined as in Equation 4 below is used:

[0099] Where r = ||X|| 2 Formula 4

[0100] In some implementations, for the distortion from the world to the camera device, n = 2 and [l1, l2] = l j .

[0101] To identify the 3D coordinates of the structure (the transformation from the camera device to the world) based on the pixel coordinates of the source image, a representation of the inverse transformation can be used. The same model with a set of different polynomial coefficients. That is, the inverse of Equation 4 can be applied to the source image to determine the undistorted result produced from the image passing through the lens of the camera device, so as to map the source image pixel coordinates of the structure of interest to the corresponding 3D coordinates of the structure. The undistorted coefficients are calculated in closed form using a predetermined set of coefficients from the given lens distortion formula. The example lens distortion formula is discussed in more detail in Pierre Drap and Julien Lefèvre’s “An exact formula for calculating inverse radial lens distortions,” Sensors Journal, 16(6):807, 2016, the full contents of which are incorporated herein by reference.

[0102] A ray-based structural parameter setting was used, where each landmark is anchored to a pixel in the input image. Since image textures are compared around such points, assumptions about surface normals can be avoided; instead, surface normals are explicitly modeled. This makes the computation of 3D model normals invariant. Each landmark includes a given (fixed) pixel position x, a source frame index i, and variable surface plane parameter settings. In some cases, Martin Habbecke and Leif Kobbelt’s “Iterative multi-view plane fitting,” In Int. Autumn Seminar on Vision, Modeling and Visualization, pp. 73–80, 2006 (the entire contents of which are incorporated herein by reference), is used to calculate the 3D or world coordinates of the landmark (the structure of interest) according to Equation 5 below:

[0103] in,

[0104] Where X is the world coordinate, and x is the pixel coordinate in the image of this structure. The pixel-to-pixel correspondence X -> X' from frame i to frame j can be achieved by substituting Equation 5 into Equation 2. This can be represented by the function defined in Equation 6 below (for N image coordinates):

[0105]

[0106] i, j, and k represent the source frame index, target frame index, and landmark index, respectively, and This indicates the set of variables to be optimized and all problem parameters. L is the number of markers. Specifically, It is an optimization problem that is solved to generate a 3D model of the structure.

[0107] In each optimization iteration (stage) of solving the optimization problem, the parameter update δΘ is computed. Except for rotations, most parameters are parameterized to the minimum of the Euclidean space and are therefore updated additively (e.g., n). k ←n k +δn k For rotations, the update is parameterized (minimally) to R. i ←R i Ω(δr i ), where Ω(·) is the Rodrigues formula for converting a 3-vector into a rotation matrix. In some cases, the derivative of the rotation... This could mean the updated derivative. Generally, parameter updates are represented as

[0108] Enhancement system 206 determines or computes the difference between pixels in a source image and identified pixels in a target image. Based on this difference, enhancement system 206 solves an optimization problem for generating a 3D model of the target structure. That is, the optimization problem is solved based on a cost function defined as the difference between pixels in the source image and identified pixels in the target image. Specifically, parameterization provides a mapping from pixels in one image to pixels in another image via scene geometry and camera device position and intrinsic parameters. Cost measures the difference between the two sets of pixels. To ensure that the cost is invariant to changes in local lighting and accidental occlusion, a robust, locally normalized, least-squares NCC cost is used. NCC is discussed in more detail in co-owned U.S. Patent Application No. 16 / 457,607, filed June 28, 2019, the entire contents of which are incorporated herein by reference.

[0109] For anchored in the image Each landmark (indexed by k) in the image (using grayscale) defines a 4x4 pixel patch centered on the landmark, where i = I k It is the source image index of the k-th landmark. This is called part of the structure or landmark. A set of image coordinates is... Definition. This marker is considered visible within a subset of the input frames, the index set of which is defined by V. k Indication. The cost of all landmarks and images is defined by Equation 7 below:

[0110]

[0111]

[0112]

[0113] Among them, the mark The indicator is sampled via bilinear interpolation, and 1 represents its vector. The robust Giman-McLuhr kernel ρ robustly optimizes the cost at τ = 0.5, but any other suitable value or function can be used. Specifically, ρ can be any robustening function that downweights more extreme measurements. In V k In the middle, source frame I k It can be ignored because, by definition, it contributes no error. In Equation 7, A function is defined to normalize pixels to account for illumination differences, ensuring that the error or difference calculation remains illumination-invariant. Equation 7 accumulates the differences between the projections of pixels from the source image to corresponding pixels in a given set of target images over all specified structural portions. Specifically, P k It is a set of pixel coordinates in the source image i for a given part k of the structure. This provides a mapping of pixel coordinates from the source image to the target image j. A normalization function Ψ is applied to this set of pixel coordinates in the source image and the corresponding pixel coordinates in the target image, and the difference ε is calculated. jk The process continues iteratively (in stages) for each target image and each structural component k identified in the source image within a set of target images. The sum is output as a cost function E(Θ), which is used to update the parameters of the optimization problem at each iteration. The optimization problem parameters continue to be updated until the error is reduced or minimized to a specified threshold, or until a specified number of iterations is completed. At each update, a new sum of cost functions is calculated for the same or different sets of target images.

[0114] At operation 504, enhancement system 206 generates a three-dimensional (3D) model of the target structure based on the joint optimization of the target structure and camera device parameters, and based on the correspondence between the first group of pixels in the source image and the second group of pixels in one or more target images.

[0115] As mentioned above, Equation 7 defines the robust nonlinear least-squares cost for which numerous optimization problem solvers exist. These solvers typically involve calculating the partial derivatives of the residual error with respect to the optimization variables (known as the Jacobian determinant). Standard implementations of such solvers cache the entire Jacobian determinant, which can be infeasibly large, making optimization problems inefficient or requiring significant storage resources. In some large optimization problems, the Jacobian determinant may not even fit into memory, making such problems challenging to solve. According to some implementations, only one instance of the Jacobian determinant is stored and maintained per iteration or update, significantly reducing storage resources and improving the overall efficiency of the device. Specifically, the disclosed optimization problem has a specific structure common to the standard SfM formula. Without surface regularization, the landmarks are independent of each other. The VarPro algorithm (discussed below in conjunction with Algorithm 1) uses Schur complement to allow the disclosed implementation to construct and solve small-scale reduced-scale camera system (RCS) problems, and then uses embedded point iteration (EPI) to solve the structure. This decouples camera parameter updates from structural parameter updates, reducing the amount of data stored. Specifically, RCS involves the set of all problem variables excluding structural variables, which is represented as Based on Equation 8 below, the RCS is constructed and solved using the Levinberg scheme damping:

[0116]

[0117]

[0118]

[0119]

[0120] Among them, J + =(J T J) -1 J T H represents the pseudo-inverse of the matrix, and λ is the damping parameter. rcs and g rcs Both can accumulate one boundary marker at a time. H rcs It is a matrix of several camera device parameters (camera device motion (R and t), camera device sensor and focal length, lens parameters), and g rcs It is a vector or length of several camera device parameters. The structural parameters are represented by n. In some cases, by marginalizing all structural parameters, a linear system is constructed only among several camera device parameters. The algorithm models how the structure moves when the camera device moves.

[0121] This allows the disclosed technique to achieve low-memory computation of RCS by looping over the landmarks and marginalizing them within the loop. After the camera device parameters are updated, EPI is run until convergence is achieved independently on each landmark using Gaussian-Newton updates: Including the Jacobian determinant parameters in the RCS and EPI allows the optimization problem solver to operate on the Jacobian determinant of only one landmark at a time. This avoids the need to store the Jacobian determinant associated with more than one landmark at any given time or iteration, thus saving memory resources.

[0122] Algorithm 1 below describes the complete optimization based on the technique represented in Equation 8. As shown below, if the error or cost (S or E(Θ)) in the current iteration (stage) is lower than the error in the previous iteration (stage), the damping parameter in Equation 8 is decreased; otherwise, the damping parameter is increased, and the optimization problem parameters are set to the values ​​of the previous iterations. Then, the optimization update is attempted again for all images and landmarks to determine if the error or cost has decreased.

[0123]

[0124] The disclosed implementation provides a novel tool for 3D reconstruction tasks, used to generate a 3D model of a structure using a set of real-time or previously captured images. An improved method for generating a 3D model of the structure is provided by jointly refining the structural and camera device parameters using photometric measurement errors robust to local illumination variations. This results in a significant increase in the metric accuracy of the reconstruction.

[0125] According to the disclosed embodiments, the enhancement system 206 uses NCC least squares optimization to generate 3D models from images in a 3D model reconstruction technique. This 3D model reconstruction technique includes, for example, SfM (an offline method for jointly calculating camera positions and sparse geometry from a set of images based on moving structures), multi-view stereo image reconstruction (an offline method for generating dense geometry of a scene given a set of images and their camera positions), SLAM (Simultaneous Localization and Mapping, an online system that jointly calculates camera positions and sparse estimates of scene geometry for each consecutive frame of a video as each consecutive frame is captured in real time), and the like. In these example embodiments, the enhancement system 206 (operating on a server or user device) can implement the NCC least squares scheme using any construction method, and additionally, the output of the above methods can be used to track objects. For example, the enhancement system 206 can implement NCC least squares within a SLAM method to reconstruct an unknown scene / object on a user's mobile device (e.g., client device 102), and the enhancement system 206 can implement NCC least squares to track camera positions relative to the scene / object.

[0126] Buildings are examples of image features that can be uploaded to a database to create 3D models using augmentation system 206. For example, suppose the building is a typical urban building, and a 3D model of that building does not exist. Conventionally, creating a 3D model of this building might not be practical, as it would require careful measurement and real-world analysis. However, augmentation system 206 can implement a least-squares NCC scheme to correlate points of buildings in different user images to generate a 3D reconstructed model of the building. This 3D model data can then be sent to a client device for processing. For example, the 3D model of the building can be overlaid on the building in the image (e.g., in different user interface layers), and image effects can be applied to the 3D model to create an augmented reality experience. For example, an explosion effect can be applied to the 3D model of the building, or pixels of the image data can be remapped and enlarged on the 3D model so that the building appears significantly larger when viewed through live video. In some example implementations, the 3D model data is used for augmented reality effects in addition to identifying, tracking, or aligning buildings in the image. For example, a 2D representation can be generated based on a 3D model of a building, and this 2D representation can be compared with frames of real-time video to determine features in the real-time video depiction that are similar to or exactly match the building in the 2D representation. One benefit of the least-squares NCC method discussed above is that the augmented system 206 can efficiently apply least-squares NCC on the client device 102 to generate 3D model data on the client device 102 without server support.

[0127] Machine architecture

[0128] Figure 6 This is a schematic representation of machine 600, in which instructions 608 (e.g., software, programs, applications, applets, or other executable code) can be executed to cause machine 600 to perform any or more of the methods discussed herein. For example, instructions 608 can cause machine 600 to perform any or more of the methods described herein. Instructions 608 transform the general, non-programmed machine 600 into a specific machine 600 programmed to perform the described and illustrated functions in the described manner. Machine 600 can operate as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, machine 600 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 600 may include, but is not limited to: server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, web devices, network routers, network switches, network bridges, or any machine capable of sequentially or otherwise executing instructions 608 specifying actions to be taken by machine 600. Furthermore, although only a single machine 600 is shown, the term "machine" should also be considered as a collection of machines that individually or jointly execute instructions 608 to perform any one or more of the methods discussed herein. For example, machine 600 may include client device 102 or any of several server devices forming part of message server system 108. In some examples, machine 600 may also include both client and server systems, wherein certain operations of a particular method or algorithm are performed on the server side and certain operations of a particular method or algorithm are performed on the client side.

[0129] Machine 600 may include a processor 602, a memory 604, and an input / output (I / O) unit 638 that can be configured to communicate with each other via a bus 640. In the example, processor 602 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processor 606 and processor 610 that execute instructions 608. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") capable of executing instructions simultaneously. Although Figure 6 Multiprocessor 602 is shown, but machine 600 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.

[0130] Memory 604 includes main memory 612, static memory 614, and memory cells 616, all of which are accessible by processor 602 via bus 640. Main memory 604, static memory 614, and memory cells 616 store instructions 608 embodying any one or more of the methods or functions described herein. Instructions 608 may also reside wholly or partially in main memory 612, static memory 614, machine-readable medium 618 within memory cells 616, at least one of processors 602 (e.g., within the processor's cache memory), or any suitable combination thereof, during execution by machine 600.

[0131] I / O component 638 may include various components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurement results, etc. The specific I / O component 638 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine is unlikely to include such a touch input device. It will be understood that I / O component 638 may include... Figure 6Many other components are not shown. In various examples, I / O component 638 may include user output component 624 and user input component 626. User output component 624 may include visual components (e.g., displays such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tube (CRT) displays), acoustic components (e.g., speakers), haptic components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. User input component 626 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreens that provide position and force for touch or touch gestures, or other haptic input components), audio input components (e.g., microphones), etc.

[0132] In other examples, I / O component 638 may include biometric component 628, motion component 630, environmental component 632, or position component 634, as well as a wide range of other components. For example, biometric component 628 includes components for detecting expressions (e.g., hand expressions, facial expressions, vocal expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), and identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Motion component 630 includes accelerometer components (e.g., accelerometers), gravity sensor components, and rotation sensor components (e.g., gyroscopes).

[0133] Environmental component 632 includes, for example: one or more camera devices (with still image / photograph and video capabilities), lighting sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers for detecting ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones for detecting background noise), proximity sensor components (e.g., infrared sensors for detecting nearby objects), gas sensors (e.g., gas detection sensors for detecting the concentration of hazardous gases or measuring pollutants in the atmosphere for safety purposes), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment.

[0134] Regarding the camera device, client device 102 may have a camera device system including, for example, a front-facing camera on the front surface of client device 102 and a rear-facing camera on the rear surface of client device 102. The front-facing camera may be used, for example, to capture still images and videos (e.g., "selfies") of the user of client device 102, which can then be enhanced using the aforementioned enhancement data (e.g., filters). For example, the rear-facing camera may be used to capture still images and videos in a more conventional camera device mode, which are similarly enhanced using the enhancement data. In addition to the front and rear cameras, client device 102 may also include a 360° camera for capturing 360° photos and videos.

[0135] Furthermore, the camera system of client device 102 may include dual rear cameras (e.g., a main camera and a depth-sensing camera), or even include triple, quadruple, or quintuple rear camera configurations on the front and rear sides of client device 102. For example, these multi-camera systems may include wide-angle cameras, ultra-wide-angle cameras, telephoto cameras, macro cameras, and depth sensors.

[0136] The position component 634 includes a positioning sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure and from which altitude can be obtained), an orientation sensor component (e.g., a magnetometer), etc.

[0137] A wide variety of technologies can be used to implement communication. I / O component 638 also includes communication component 636, which is operable to couple machine 600 to network 620 or device 622 via a corresponding coupling or connection. For example, communication component 636 may include a network interface component or another suitable device to interface with network 620. In further examples, communication component 636 may include wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, etc. Components (e.g.) (low power consumption) Components and other communication components that provide communication via other modes. Device 622 can be any peripheral device from other machines or various peripheral devices (e.g., a peripheral device coupled via USB).

[0138] Furthermore, communication component 636 can detect identifiers or include components operable to detect identifiers. For example, communication component 636 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, UltraCode, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying audio signals from the tag). Additionally, various information can be derived via communication component 636, such as location via Internet Protocol (IP) geolocation, etc. The location of signal triangulation, the location of NFC beacon signals that can be detected to indicate a specific location, etc.

[0139] Various memories (e.g., main memory 612, static memory 614, and the memory of processor 602) and memory cell 616 may store one or more sets of instructions and data structures (e.g., software) implemented or used by any one or more of the methods or functions described herein. These instructions (e.g., instruction 608) cause various operations to implement the disclosed examples when executed by processor 602.

[0140] Instructions 608 can be sent or received over network 620 via a network interface device (e.g., the network interface component included in communication component 636), using a transmission medium and employing any of several known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instructions 608 can be sent or received via a transmission medium through a coupling (e.g., peer-to-peer coupling) to device 622.

[0141] Software Architecture

[0142] Figure 7 This is a block diagram 700 illustrating a software architecture 704, which can be installed on any or more devices described herein. The software architecture 704 is supported by hardware such as a machine 702 including a processor 720, memory 726, and I / O components 738. In this example, the software architecture 704 can be conceptualized as a stack of layers, where each layer provides specific functionality. The software architecture 704 includes layers such as an operating system 712, libraries 710, a framework 708, and an application 706. Operationally, the application 706 activates API calls 750 via the software stack and receives messages 752 in response to API calls 750.

[0143] Operating system 712 manages hardware resources and provides public services. Operating system 712 includes, for example, a kernel 714, services 716, and drivers 722. Kernel 714 acts as an abstraction layer between the hardware layer and other software layers. For example, kernel 714 provides memory management, processor management (e.g., scheduling), component management, network and security settings, and other functions. Services 716 can provide other public services to other software layers. Drivers 722 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 722 may include display drivers, camera drivers, etc. or Low-power drivers, flash drives, serial communication drivers (e.g., USB drives), Drivers, audio drivers, power management drivers, etc.

[0144] Library 710 provides common low-level infrastructure used by application 706. Library 710 may include system libraries 718 (e.g., the C standard library), which provide functions such as memory allocation, string manipulation, and mathematical functions. Furthermore, library 710 may include API libraries 724, such as media libraries (e.g., libraries for supporting the rendering and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codecs, Joint Picture Experts Group (JPEG or JPG), or Portable Web Graphics (PNG)), graphics libraries (e.g., the OpenGL framework for rendering graphic content on a display in two-dimensional (2D) and three-dimensional (3D) formats), database libraries (e.g., SQLite, which provides various relational database functions), web libraries (e.g., WebKit, which provides web browsing functionality), etc. Library 710 may also include various other libraries 728 to provide application 706 with many other APIs.

[0145] Framework 708 provides common high-level infrastructure for use by application 706. For example, framework 708 provides various graphical user interface (GUI) functions, advanced resource management, and advanced location services. Framework 708 can provide a wide range of other APIs that can be used by application 706, some of which may be specific to a particular operating system or platform.

[0146] In the example, application 706 may include a home application 736, a contacts application 730, a browser application 732, a book reader application 734, a location application 742, a media application 744, a messaging application 746, a game application 748, and a wide variety of other applications such as third-party application 740. Application 706 is a program that performs the functions defined in the program. One or more applications 706 can be created using various programming languages, such as object-oriented programming languages ​​(e.g., Objective-C, Java, or C++) or procedural programming languages ​​(e.g., C or assembly language). In a particular example, third-party application 740 (e.g., an entity other than the vendor of a particular platform using Android) TM or iOS TM Applications developed using a Software Development Kit (SDK) can be used on platforms such as iOS. TM ANDROID TM , Mobile software running on the phone's mobile operating system or another mobile operating system. In this example, a third-party application 740 can call API call 750 provided by the operating system 712 to facilitate the functions described herein.

[0147] Glossary

[0148] "Carrier signal" refers to any intangible medium capable of storing, encoding, or carrying instructions to be executed by a machine, and includes digital or analog communication signals or other intangible media to facilitate the communication of such instructions. Instructions can be sent or received over a network using a transmission medium via a network interface device.

[0149] "Client device" refers to any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. Client devices can be, but are not limited to, mobile phones, desktop computers, laptop computers, portable digital assistants (PDAs), smartphones, tablet computers, ultrabooks, netbooks, laptop computers, multiprocessor systems, microprocessor-based or programmable consumer electronics, game consoles, set-top boxes, or any other communication device that a user can use to access the network.

[0150] "Communication network" refers to one or more parts of a network, which can be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, a part of the Internet, a part of the Public Switched Telephone Network (PSTN), a POTS (Plain Old-Style Telephone Service) network, a cellular telephone network, a wireless network, etc. A network, other types of networks, or a combination of two or more such networks. For example, a network or part of a network may include a wireless network or a cellular network, and the coupling may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or other types of cellular or wireless coupling. In this example, the coupling can implement any data transmission technology of various types, such as Single Carrier Radio Transmission (1xRTT), Evolved Data Optimization (EVDO), General Packet Radio Service (GPRS), Enhanced Data Rate Evolution of GSM (EDGE), the 3rd Generation Partnership Project (3GPP) including 3G, fourth-generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High-Speed ​​Packet Access (HSPA), Global Microwave Access Interoperability (WiMAX), Long Term Evolution (LTE) standards, other data transmission technologies defined by various standards setting organizations, other long-distance protocols, or other data transmission technologies.

[0151] A "component" is a device, physical entity, or logic having boundaries defined by functional or subroutine calls, branch points, APIs, or other technologies that partition or modularize a particular processing or control function. A component can be combined with other components via its interface to perform machine processing. A component can be an encapsulated functional hardware unit designed for use with other components, and can be part of a program that typically performs a specific function within a related function.

[0152] Components can constitute software components (e.g., code embodied on a machine-readable medium) or hardware components. A “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in some physical manner. In various example implementations, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems) or one or more hardware components (e.g., processors or processor groups) of a computer system can be configured by software (e.g., an application or application portion) to perform certain operations described herein.

[0153] Hardware components can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component may include a dedicated circuit system or logic permanently configured to perform certain operations. A hardware component may be a dedicated processor, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). A hardware component may also include programmable logic or a circuit system temporarily configured by software to perform certain operations. For example, a hardware component may include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a specific machine (or a specific component of a machine) uniquely tailored to perform the configured function, and no longer a general-purpose processor. It will be understood that the implementation of a hardware component may be determined, for cost and time considerations, whether mechanically implemented in a dedicated and permanently configured circuit or in a temporarily configured (e.g., software-configured) circuit. Accordingly, the phrase “hardware component” (or “hardware-implemented component”) should be understood to include tangible entities, i.e., entities physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate or perform certain operations described herein.

[0154] Consider implementations where hardware components are temporarily configured (e.g., programmed), eliminating the need to configure or instantiate each hardware component at any given time. For example, where the hardware components include a general-purpose processor configured by software to become a dedicated processor, this general-purpose processor can be configured as different dedicated processors (e.g., including different hardware components) at different times. The software accordingly configures one or more specific processors to constitute a specific hardware component at one time and different hardware components at different times.

[0155] Hardware components can provide information to and receive information from other hardware components. Therefore, the described hardware components can be considered communicatively coupled. In the presence of multiple hardware components, communication can be achieved through signal transmission between or among two or more hardware components (e.g., via appropriate circuitry and buses). In embodiments where multiple hardware components are configured or instantiated at different times, such communication between hardware components can be achieved, for example, by storing information in a memory structure accessed by the multiple hardware components and retrieving information from that memory structure. For example, a hardware component can perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. Other hardware components can then access the memory device at a subsequent time to retrieve and process the stored output. Hardware components can also initiate communication with input or output devices and can operate on resources (e.g., information collection).

[0156] Various operations of the example methods described herein can be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute components of a processor implementation that operate to perform one or more operations or functions described herein. As used herein, a “processor-implemented component” refers to a hardware component implemented using one or more processors. Similarly, the methods described herein can be implemented, at least partially, by processors, where a particular one or more processors are examples of hardware. For example, at least some operations of the methods can be performed by one or more processors 602 or processor-implemented components. Furthermore, the one or more processors can also operate to support the execution of relevant operations in a “cloud computing” environment or operate as “Software as a Service” (SaaS). For example, at least some operations can be performed by a group of computers (as an example of machines including processors), where these operations are accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The execution of some operations can be distributed among processors, not residing only within a single machine, but deployed across multiple machines. In some example implementations, the processor or processor-implemented components may be located in a single geographic location (e.g., within a home environment, an office environment, or a server cluster). In other example implementations, the processor or processor-implemented components may be distributed across several geographic locations.

[0157] "Computer-readable storage medium" refers to both machine-readable storage media and transmission media. Therefore, these terms encompass both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" refer to the same thing and may be used interchangeably in this disclosure.

[0158] A "brief message" is a message that is accessible for a limited period of time. Brief messages can be text, images, videos, etc. The access time for a brief message can be set by the message sender. Alternatively, the access time can be a default setting or a setting specified by the recipient. Regardless of the setting method, the message is temporary.

[0159] "Machine storage medium" refers to one or more storage devices and media (e.g., centralized or distributed databases, and associated caches and servers) that store executable instructions, routines, and data. Therefore, this term should be considered to include, but is not limited to, solid-state memory and optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and device storage media include: non-volatile memory, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGAs, and flash memory devices; disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium," "device storage medium," and "computer storage medium" refer to the same thing and are used interchangeably in this disclosure. The terms "machine storage medium," "computer storage medium," and "device storage medium" expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are covered by the term "signal medium."

[0160] "Non-transitory computer-readable storage medium" refers to a tangible medium capable of storing, encoding, or carrying instructions that can be executed by a machine.

[0161] "Signal medium" means any intangible medium capable of storing, encoding, or carrying instructions executable by a machine, and includes digital or analog communication signals or other intangible media that facilitate the communication of software or data. The term "signal medium" should be considered to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal whose characteristics are set or altered in a manner that encodes information in the signal. The terms "transmission medium" and "signal medium" refer to the same thing and may be used interchangeably in this disclosure.

[0162] Changes and modifications may be made to the disclosed embodiments without departing from the scope of this disclosure. Such and other changes or modifications are intended to be included within the scope of this disclosure as set forth in the appended claims.

Claims

1. A method for modeling three-dimensional (3D) objects, comprising: Access the source image depicting the target structure; Access one or more target images depicting at least a portion of the target structure; Calculate the correspondence between a first group of pixels in the source image of the first part of the target structure and a second group of pixels in one or more target images of the first part of the target structure, the correspondence being calculated based on camera device parameters that vary between the source image and the one or more target images; as well as Based on the joint optimization of the target structure and camera device parameters, and based on the correspondence between the first group of pixels in the source image and the second group of pixels in one or more target images, a 3D model of the target structure is generated.

2. The method according to claim 1, wherein, The correspondence is also calculated based on the target structure, and the calculation of the correspondence includes: Identify the 3D coordinates of the first part of the target structure corresponding to the first group of pixels; By applying a first set of camera device parameters associated with the source image, a first line is projected from the first set of pixels to the 3D coordinates; By applying a second set of camera device parameters associated with the one or more target images, a second line is projected from the 3D coordinates to the second set of pixels; and Calculate the difference between each pixel in the first group of pixels and each pixel in the second group of pixels.

3. The method according to claim 2, further comprising: The second group of pixels is identified based on the second group of camera device parameters and the 3D coordinates; as well as Based on the aforementioned correspondence, photometric measurement errors are reduced to generate 3D models.

4. The method according to claim 2, further comprising: Based on the parameters of the first set of camera devices, the first set of pixels are dedistorted in order to project the first line; as well as Distortion is applied to the second line to project the second line onto the second group of pixels.

5. The method according to any one of claims 1 to 4, further comprising: Normalize the first group of pixels and the second group of pixels.

6. The method according to any one of claims 1 to 4, further comprising: Calculate the sum of squares of the differences.

7. The method according to any one of claims 1 to 4, further comprising: Calculate the pixel-to-3D coordinate correspondence between pixels in the source image and 3D points on the target structure; as well as Calculate the 3D coordinate-to-pixel correspondence between the 3D point on the target structure and the pixels in one or more target images.

8. The method according to any one of claims 1 to 4, wherein, The camera device parameters include rotation parameters, translation parameters, sensor parameters, lens dedistortion parameters, and lens distortion parameters.

9. The method according to any one of claims 1 to 4, further comprising: The definition includes an optimization problem involving multiple structural parameters and camera device parameters, wherein the optimization problem is invariant to illumination and surface normals; as well as The optimization problem is solved using a cost function based on the calculated correspondence to generate the 3D model.

10. The method according to claim 9, wherein, Solving the optimization problem involves decoupling the camera device parameter updates from the structural parameter updates to reduce the amount of data stored.

11. The method according to any one of claims 1 to 4, wherein, The source image and the one or more target images are received in real time from a camera feed from a camera device on a client device, and the method further includes applying one or more augmented reality elements to the camera feed based on the 3D model.

12. The method according to any one of claims 1 to 4, wherein, The source image and the one or more target images were previously captured and processed offline on the server.

13. The method according to any one of claims 1 to 4, further comprising: The resolution of the source image is matched with the resolution of one or more target images.

14. The method of claim 13, further comprising: Identify the first set of pixels in the source image that corresponds to the first part of the target structure; Identify a second set of pixels in one or more target images that corresponds to the first portion of the target structure; Calculate the first distance between each pixel in the first pixel set and the second distance between each pixel in the second pixel set; as well as The sampling parameters are selected based on the difference between the first distance and the second distance.

15. The method of claim 14, further comprising: Based on the sampling parameters, the one or more target images are upsampled or downsampled.

16. The method according to any one of claims 1 to 4, wherein, Accessing the source image includes: Generate the 3D coordinate frame of the target structure; Calculate the visibility of the 3D coordinate frame as a depth map for multiple images; and Based on the calculated visibility, one of the plurality of images is selected as the source image.

17. The method of claim 16, further comprising: Calculate the pixel grid with a specific interval corresponding to the 3D coordinate frame; The plurality of images associated with the pixel grid are sampled to generate a matrix, each column of which corresponds to a different image among the plurality of images; Calculate the solution for the average, weighted average, or robust sum of squares of each column of the matrix; as well as The image whose corresponding column is closest in value to the calculated average, the weighted average, or the robust sum of squares among the plurality of images is selected as the source image.

18. The method according to any one of claims 1 to 4, further comprising: In the initial optimization phase, the first set of images, which are reduced in size, are processed. as well as After the initial optimization phase, a second set of images, processed into full-size images, is used as one or more target images to improve convergence.

19. A three-dimensional object modeling system, comprising: The processor is configured to perform the following operations: Access the source image depicting the target structure; Access one or more target images depicting at least a portion of the target structure; Calculate the correspondence between a first group of pixels in the source image of the first part of the target structure and a second group of pixels in one or more target images of the first part of the target structure, the correspondence being calculated based on camera device parameters that vary between the source image and the one or more target images; as well as Based on the joint optimization of the target structure and camera device parameters, and based on the correspondence between the first group of pixels in the source image and the second group of pixels in one or more target images, a 3D model of the target structure is generated.

20. A non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations including: Access the source image depicting the target structure; Access one or more target images depicting at least a portion of the target structure; Calculate the correspondence between a first group of pixels in the source image of the first part of the target structure and a second group of pixels in one or more target images of the first part of the target structure, the correspondence being calculated based on camera device parameters that vary between the source image and the one or more target images; as well as Based on the joint optimization of the target structure and camera device parameters, and based on the correspondence between the first group of pixels in the source image and the second group of pixels in one or more target images, a three-dimensional 3D model of the target structure is generated.