Determining camera parameters using critical edge detection neural networks and geometric models
By combining critical edge detection neural network and geometric model, the accuracy, flexibility and efficiency of camera parameter determination in the prior art are solved, and efficient and accurate camera parameter determination and image enhancement effects are achieved.
Patent Information
- Application Number
- CN202510164577.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-01-25
- Filing Date
- 2019-11-06
- Publication Date
- 2025-05-23
AI Technical Summary
Existing digital image analysis systems have shortcomings in accuracy, flexibility and efficiency in determining camera parameters, such as high error rates, inaccurate image modifications and insufficient processing capabilities for complex images.
Critical edge detection neural network is used to combine geometric models to accurately, efficiently and flexibly determine camera parameters from a single digital image to generate enhanced digital images. Specific steps include weighting edges in digital images using a deep learning framework, determining camera parameters using vanishing lines in combination with geometric models, and improving network accuracy and efficiency through training data.
Improves the accuracy of camera parameters and the fidelity of image modification, enhances the flexibility and efficiency of the system, and can handle complex images without being affected by misleading shapes and lines.
Smart Images

Figure CN120031985A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application is a divisional application of the invention patent application with application number 201911077834.1, application date November 6, 2019, and invention name “Determining camera parameters using critical edge detection neural network and geometric model”. Technical Field
[0003] Embodiments of the present disclosure relate to determining camera parameters using a critical edge detection neural network and a geometric model. Background Art
[0004] Recent years have seen significant advances in the field of digital image analysis and editing. Due to advances in algorithms and hardware, traditional digital image analysis systems are now able to analyze and edit digital images in a variety of ways. For example, traditional digital image analysis systems can reproject digital images so that they appear visually aligned (e.g., vertically) and add or subtract objects depicted in the digital image. Indeed, traditional digital image analysis systems can (with camera parameters in control) resize and orient new digital objects so that they fit within the three-dimensional scene depicted in the digital image. In such three-dimensional synthesis, it is crucial to have an accurate estimate of the camera calibration to ensure that foreground and background elements have matching perspective distortion in the final rendered image.
[0005] Although conventional digital image analysis systems have evolved in recent years, they still suffer from several significant deficiencies in terms of accuracy, flexibility, and efficiency. For example, some conventional digital image analysis systems may utilize convolutional neural networks to determine camera parameters and modify digital images. Specifically, conventional digital analysis systems may train convolutional neural networks to identify camera parameters from digital images. However, such systems are not very precise and / or accurate. In fact, many digital image analysis systems that utilize convolutional neural networks to determine camera parameters have high error rates. As a result, such systems also generate inaccurate, realistic, or visually unappealing modified digital images.
[0006] Some conventional digital image analysis systems utilize geometric schemes to determine camera parameters from digital images. For example, such conventional systems may analyze geometric shapes in digital images to identify edges and determine camera parameters when capturing the digital images based on the identified edges. However, such systems are not robust or flexible. In fact, conventional systems utilizing geometric methods have significant problems with accuracy when analyzing digital images that contain misleading or confusing shapes and / or lines (e.g., lack of strong vanishing lines). For example, digital images containing various round objects, curved objects, or lines pointing in random directions may undermine the accuracy of the geometric model.
[0007] Some conventional systems can determine camera parameters by analyzing multiple digital images of the same subject (e.g., the same scene or object). Such systems are inefficient because they require a large amount of computer resources and require processing multiple images to extract a set of camera parameters. In addition, such systems provide little flexibility because they require digital images that meet specific standards, and in many cases, it is unlikely that the user has the images required for the system to work properly (i.e., a large number of digital images depicting the same subject). In addition, many systems require a large amount of computer processing resources and time to generate and utilize training data.
[0008] These, along with additional problems and controversies, exist regarding conventional digital image analysis systems. Summary of the invention
[0009] Embodiments of the present disclosure provide benefits and / or solve one or more of the foregoing or other problems in the art using systems, non-transitory computer-readable media, and computer-implemented methods, which are used to determine camera calibration and generate enhanced digital images using critical edge detection neural networks and geometric models. Specifically, the disclosed system can accurately, efficiently, and flexibly determine camera parameters such as focal length, pitch, roll, or yaw based on a single digital image. For example, in one or more embodiments, the disclosed system uses a deep learning-based framework to weight edges in a digital image (e.g., to identify vanishing lines associated with digital image perspective). The disclosed system can then use vanishing lines in combination with a geometric model to accurately identify camera parameters and generate a modified digital image. To further improve efficiency and accuracy, the disclosed system can also generate accurate training data from an existing digital image repository and use the training data to train a critical edge detection neural network. Specifically, the disclosed system can generate ground truth vanishing lines from a training digital image, and use the ground truth vanishing lines to train a critical edge detection neural network to identify critical edges. After training, the system can utilize a critical edge detection neural network in conjunction with a geometric model to more accurately determine camera parameters and generate a modified digital image.
[0010] Additional features and advantages of one or more embodiments of the present disclosure are summarized in the description which follows, and in part will be obvious from the description, or may be learned by practice of such exemplary embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] As briefly described below, the detailed description provides additional specificity and detail to one or more embodiments through use of the accompanying figures.
[0012] Figure 1A diagram illustrating an environment in which a camera parameter determination system may operate in accordance with one or more embodiments.
[0013] Figure 2 An overview of determining camera parameters for a digital image is shown in accordance with one or more embodiments.
[0014] Figures 3A-3C A process for generating ground truth vanishing lines from a training set of digital images in accordance with one or more embodiments is shown.
[0015] Figure 4 A flow chart for training a critical edge detection neural network is shown in accordance with one or more embodiments.
[0016] Figure 5 Utilizing a critical edge detection neural network to generate a vanishing edge map is shown in accordance with one or more embodiments.
[0017] Figure 6 Utilizing camera parameters of a digital image to generate an enhanced digital image in accordance with one or more embodiments is shown.
[0018] Figure 7 A block diagram of a camera parameter determination system is shown in accordance with one or more embodiments.
[0019] Figure 8 A flow chart illustrating a series of actions for determining camera parameters using a critical edge detection neural network in accordance with one or more embodiments.
[0020] Fig. 9 A flow chart illustrating a series of actions for training a critical edge detection neural network in accordance with one or more embodiments.
[0021] Fig.10 A block diagram of an example computing device is shown in accordance with one or more embodiments. DETAILED DESCRIPTION
[0022] The present disclosure describes one or more embodiments of a camera parameter determination system that can utilize a critical edge detection neural network in combination with a geometric model to identify camera parameters from a single digital image. The camera parameter determination system can train a critical edge detection neural network to identify vanishing lines in a digital image and generate an edge map (e.g., an edge map weighted for vanishing lines indicating perspective within the digital image). The camera parameter determination system can then utilize a geometric model to analyze the edge map generated by the critical edge detection neural network and identify a camera calibration. The camera parameter determination system can also further improve accuracy and efficiency by generating ground truth data for training the critical edge detection neural network.
[0023] To illustrate, a camera parameter determination system may generate ground truth vanishing lines from a set of training digital images. Specifically, in one or more embodiments, the camera parameter determination system identifies lines (e.g., edges) in the training digital images and determines vanishing points by analyzing the lines. Based on the distance between the vanishing points and the lines in the digital images, the camera parameter determination system may identify the ground truth vanishing lines. The camera parameter determination system may then train a critical edge detection neural network using the training digital images and the ground truth vanishing lines to identify vanishing lines in the digital images. While training the critical edge detection neural network, the camera parameter determination system may generate a vanishing edge map using the critical edge detection neural network, the vanishing edge map indicating vanishing lines for the digital images. In one or more embodiments, the camera parameter determination system utilizes the vanishing edge map (e.g., by applying a geometric model) to more accurately and efficiently determine camera parameters for the digital images.
[0024] As described above, in some embodiments, the camera parameter determination system identifies ground truth vanishing lines for a training set of images. More specifically, the camera parameter determination system may map the digital image onto a sphere and divide the sphere into regions or "bins". The camera parameter determination system may then identify each of the intersections of two or more lines on the sphere. In one or more embodiments, the camera parameter determination system utilizes a distance-based voting scheme between lines, intersections, and / or pixels of the digital image to determine a dominant vanishing point of the image. In some embodiments, the camera parameter determination system utilizes the dominant vanishing point to determine whether each line in the digital image is a ground truth vanishing line based on the distance between each line in the digital image and the dominant vanishing point.
[0025] In addition, as described above, the camera parameter determination system can train a critical edge detection neural network. Specifically, in one or more embodiments, the camera parameter determination system uses the ground truth vanishing lines of the training image to train the critical edge detection neural network to determine the vanishing lines in a supervised manner. Specifically, as described below with respect to Figure 4 Discussed in more detail, the camera parameter determination system can train a critical edge detection neural network to identify vanishing lines in a digital image using a loss function in conjunction with ground truth vanishing lines. In this manner, the camera parameter determination system can train a critical edge detection neural network to identify vanishing lines based on contextual information analyzed at various levels of abstraction within the digital image.
[0026] In addition, the camera parameter determination system can use the trained critical edge detection neural network to identify vanishing lines in the digital image and generate a vanishing edge map. In one or more embodiments, once the critical edge detection neural network is trained, the camera parameter determination system can use the critical edge detection neural network to generate a vanishing edge map including a plurality of edge weights. Specifically, the vanishing edge map can include a weight for each pixel indicating the probability that each pixel corresponds to a vanishing line in the digital image.
[0027] The camera parameter determination system can utilize the vanishing edge map to determine camera parameters for the corresponding digital image. More specifically, the camera parameter determination system can apply a geometric model to the vanishing edge map to accurately determine the focus, pitch, roll, or yaw. In fact, because the camera parameter determination system can apply the geometric model to the vanishing edge map (rather than the many misleading or erroneous lines contained in the digital image), the camera parameter determination system can improve the accuracy of the resulting camera parameters.
[0028] The camera parameter determination system can also utilize the determined camera parameters to perform various functions with the digital image. For example, the camera parameter determination system can generate an enhanced digital image. More specifically, the camera parameter determination system can accurately and seamlessly add objects to a three-dimensional scene depicted in the digital image. The camera parameter determination system can also re-project the digital image to align vertical and horizontal lines to a vanishing point. In addition, the camera parameter determination system can generate / estimate a three-dimensional model of a scene or object depicted in the digital image based on the camera parameters and the digital image.
[0029] The camera parameter determination system provides many advantages and benefits relative to conventional systems and methods. For example, the camera parameter determination system can improve accuracy relative to conventional systems. In fact, by applying a critical edge detection neural network to generate accurate vanishing lines and then utilizing a geometric model to analyze the vanishing lines, the camera parameter determination system can accurately determine camera parameters and generate more accurate and realistic enhanced digital images. Therefore, compared to conventional systems that utilize convolutional neural networks to predict camera parameters, the camera parameter determination system can utilize a geometric model that focuses on analyzing accurate vanishing lines generated by a critical edge detection neural network to generate accurate camera parameters.
[0030] In addition, the camera parameter determination system can improve flexibility relative to conventional systems. For example, in one or more embodiments, the camera parameter determination system utilizes a critical edge detection neural network to filter out inaccurate or misleading lines that fail to reflect the image perspective. The camera parameter determination system can then apply a geometric model to the remaining accurate vanishing lines. Therefore, the camera parameter determination system can robustly generate accurate camera parameters in a wide variety of digital images, even digital images that include circles or random lines. The camera parameter determination system can also improve flexibility by determining camera parameters based on a single digital image (rather than requiring the user to capture and provide multiple digital images of an object or scene).
[0031] In addition, the camera parameter determination system can improve the efficiency of implementing a computing system. Compared with traditional systems, the camera parameter determination system can determine camera parameters by analyzing a single digital image (rather than analyzing multiple digital images). In addition, the camera parameter determination system can efficiently generate and utilize training data (e.g., from an existing digital image repository) to train a critical edge detection neural network. This can greatly reduce the processing power and time required to generate and utilize traditional, labeled training data.
[0032] In summary, by utilizing a critical edge detection neural network, the camera parameter determination system can leverage contextual information in the image to determine the exact vanishing line (e.g., the first edge is from the floor, so that edge is likely to point to the vanishing point, or the second edge is part of a stair railing and may be distracting). Additionally, the camera parameter determination system can leverage the accuracy of the geometric model where such a model is likely to succeed, thereby yielding better overall performance relative to conventional systems.
[0033] As shown in the previous discussion, the present disclosure uses various terms to describe the features and benefits of a dynamic representation management system. Additional details about the meaning of these terms used in the present disclosure are provided below. For example, as used herein, the term "digital image" refers to any digital symbol, picture, icon or diagram. For example, the term "digital image" includes a digital file with the following or other file extensions: JPG, TIFF, BMP, PNG, RAW or PDF. The term "digital image" also includes one or more images (e.g., frames) in a digital video. In addition, the term "digital image" refers to a 3D object represented in a digital format. For example, the term "digital image" includes a digital file with the following or other file extensions: OBJ, DAE, 3DS, U3D and KMZ. Therefore, although many descriptions herein are expressed in terms of digital images, it should be understood that the present disclosure can also be applied to extracting attributes from digital videos and / or editing digital videos. In addition, as used herein, the term "training digital image" refers to a digital image used to train a neural network. Specifically, the term "training digital image" can include a digital image associated with ground truth data that can be used to train a neural network.
[0034] As used herein, the term "camera device" refers to any device that can be used to capture an image. Specifically, the term "camera device" may include a device that is capable of capturing any type of digital image as described above. For illustration, a camera device may include a digital camera or film camera, a mobile phone or other mobile device, a tablet computer, a computer, or any other device that can capture an image.
[0035] Additionally, as used herein, the term "camera parameters" refers to characteristics or properties of a camera device used to capture a digital image. Specifically, the term "camera parameters" may include characteristics of a camera device when capturing a digital image that affect the appearance of the digital image. For illustration, camera parameters may include focal length, field of view, pitch, roll, and / or yaw.
[0036] In addition, as used herein, the term "neural network" refers to a machine learning model that can be adjusted (e.g., trained) based on inputs to approximate unknown functions. Specifically, the term "neural network" may include a model of interconnected layers that transmit and analyze properties while changing the degree of abstraction to learn to approximate complex functions and generate outputs based on multiple inputs provided to the model. For example, the term "neural network" includes one or more machine learning algorithms. In other words, a neural network includes algorithms that implement deep learning techniques, i.e., machine learning that utilizes a collection of algorithms to attempt to model high-level abstractions in data. Additional details about exemplary neural networks and corresponding network architectures are provided below.
[0037] In addition, as used herein, the term "critical edge detection neural network" refers to a neural network for identifying vanishing lines from a digital image. For example, a critical edge detection neural network may include a neural network (e.g., for generating a vanishing edge map including the vanishing lines from the digital image). In one or more embodiments, the critical edge detection neural network includes a convolutional neural network, such as a stacked hourglass network. Additional details regarding the exemplary architectures and capabilities of the critical edge detection neural network are discussed in more detail below (e.g., regarding Figure 2 , 4 and 5-6).
[0038] In addition, as used herein, the term "vanishing point" refers to a region in an image that indicates perspective (e.g., a perspective-related point or direction where lines in the image appear to converge). Specifically, the term "vanishing point" may include a point, vector, line, or region in a digital image (e.g., a "viewpoint", a horizontal line, or the North Pole), where the two-dimensional perspective projections of lines that are parallel to each other in three-dimensional space appear to converge. The vanishing point may include a horizontal vanishing point (e.g., a common direction / line / point / region along a horizontal line, such as the Atlanta vanishing point), a vertical vanishing point (e.g., a common direction / line / point / region of the vertical lines in the digital image), or other vanishing points. The vanishing point may take various forms, such as a point where lines converge in two-dimensional space. In three-dimensional applications (e.g., when mapping a digital image to a three-dimensional panoramic sphere), the term vanishing point may also include a vanishing direction (e.g., a vector indicating the direction where the perspective lines in the image point or intersect). In fact, in some embodiments, the camera parameter determination system identifies a vanishing point that includes orthogonal vanishing directions in three-dimensional space (e.g., vanishing directions in the orthogonal x, y, and z directions). In fact, in one or more embodiments, the vanishing point corresponds to three Manhattan directions that respectively reflect the x, y, and z directions.
[0039] In addition, as used herein, the term "vanishing line" refers to a line in a digital image that corresponds to one or more vanishing points in the digital image. Specifically, the term "vanishing line" may include lines that converge at a vanishing point (e.g., at a horizontal vanishing point or a vertical vanishing point), which are close to or aligned with the vanishing point.
[0040] As used herein, the term "ground truth" refers to information of a known set of pixels that reflects a known set of attributes of a digital image used for training a neural network. For example, a ground truth vanishing line may include a known vanishing line identified from a digital image (e.g., a line pointing to a vanishing point). The camera parameter determination system may utilize the ground truth vanishing line to train the critical edge detection neural network.
[0041] Furthermore, as used herein, the term "training line" refers to a line identified in a training digital image. Specifically, the term "training line" may include a line identified by applying an edge detection model to a training digital image. Thus, a training line may include any line or edge depicted in a training digital image. As discussed in more detail below, a camera parameter determination system may determine a ground truth vanishing line from the training lines in the training digital image.
[0042] Additionally, as used herein, the term "vanishing edge map" refers to a representation of vanishing lines from a digital image. For illustration, a vanishing edge map may include an array, vector, database, or image (e.g., a black and white or grayscale image) where each entry represents a pixel of the digital image and indicates whether the pixel corresponds to a vanishing line. In one or more embodiments, the vanishing edge map includes weights (i.e., a weighted vanishing edge map) where each weight indicates a confidence value (e.g., a measure of confidence, such as a probability) that the corresponding pixel depicts a portion of a vanishing line within the digital image.
[0043] Similarly, as used herein, the term "weighted vanishing edge map" refers to a vanishing edge map of a digital image generated based on the weights of the lines in the image. For illustration, the camera parameter determination system may determine the weight of each line in the digital image based on various criteria, including the length of the line and / or the distance between the line and one or more vanishing points (i.e., the angular distance between the line and the vanishing point in the image, and / or the line distance between the vanishing point and the intersection of the line in the image). The camera parameter determination system may then generate a vanishing edge map by applying the weights. For example, the camera parameter determination system may determine the lines to be included in the weighted vanishing edge map based on the weights (e.g., whether the weights meet a threshold).
[0044] In addition, as used herein, the term "weight" refers to a value used to emphasize and / or de-emphasize one or more pixels. For example, a weight may include a confidence value (e.g., a measure of any confidence, such as a probability value). For example, the term "weight" may refer to a confidence value that a given line is a vanishing line or a confidence value that a given pixel is included in a vanishing line. In addition, as used herein, the term "training weight" refers to the weight of a line and / or pixel in a training image. More specifically, the training weights can be used as part of a training image to train a critical edge detection neural network.
[0045] Additionally, as used herein, the term "geometric model" refers to a model that analyzes lines and / or shapes in a digital image to estimate camera parameters. For example, the term "geometric model" may refer to a model that groups directional elements from a digital image and iteratively refines them to facilitate estimating the orientation of the image or camera parameters. To illustrate, the term "geometric model" may include a model that searches for converging lines and their intersections to determine vertical and / or horizontal orientation and / or camera parameters in a digital image.
[0046] Additional details regarding the camera parameter determination system will now be provided with respect to an illustrative diagram depicting an exemplary embodiment. Specifically, Figure 1 A camera parameter determination environment 100 is shown. Figure 1 As shown, camera parameter determination environment 100 includes client device 102, which includes client application 104 and is associated with user 106. Client device 102 communicates with server device 110 via network 108. Server device 110 may include digital media management system 112, which in turn may include camera parameter determination system 114.
[0047] although Figure 1 The camera parameter determination system 114 is shown as being implemented via the server device 110, but the dynamic representation management system 114 can be implemented via other components. For example, the camera parameter determination system 114 can be implemented in whole or in part by the client device 102. Similarly, the camera parameter determination system 114 can be implemented via both the client device 102 and the server device 110.
[0048] The client device 102 may include various types of computing devices. For example, the client device 102 may be a mobile device (e.g., a smart phone), a tablet computer, a laptop computer, a desktop computer, or a computer such as the following: Fig.10 Any other type of computing device further described. In addition, the client application 104 can include any of various types of client application programs. For example, the client application 104 can be an online application (e.g., a web browser), and the user 106 on the client device 102 can enter a uniform resource locator (URL) or other address that directs the web browser to the server device 110. Alternatively, the client application 104 can be a different native application developed for the client device 102.
[0049] Additionally, the one or more server devices 110 may include one or more computing devices including the following references: Fig.10In some embodiments, one or more server devices 110 include a content server. Server device 110 may also include an application server, a communication server, a web hosting server, a social networking server, or a digital content activity server.
[0050] Client device 102, server device 110, and network 108 may communicate using any communication platform and technology suitable for transmitting data and / or communication signals, including any known communication technology, devices, media, and protocols that support data communication, examples of which are provided in the accompanying drawings. Fig.10 Give a description.
[0051] Although not required, the camera parameter determination system 114 can be part of the digital media management system 112. The digital media management system 112 collects, monitors, manages, edits, distributes, and analyzes various media. For example, the digital media management system 112 can analyze and edit digital images and / or digital videos based on user input identified via one or more user interfaces at the client device 102. In one or more embodiments, the digital media management system 112 can utilize the camera parameter determination system 114 to determine camera parameters and / or modify the digital image based on the camera parameters. For example, the digital media management system 112 can provide a digital image to the camera parameter determination system 114, and the camera parameter determination system 114 can provide the digital media management system 112 with camera parameters for the provided image. In other embodiments, the server device 110 can include systems other than the digital media management system 112, and the camera parameter determination system 114 can receive the image via alternative means. For example, the server device 110 can receive the image from the client device 102 or from another source via the network 108. As described above, the camera parameter determination system 114 can efficiently, accurately, and flexibly determine the camera parameters of the digital image. Specifically, Figure 2 An overview of determining camera parameters for a digital image is shown in accordance with one or more embodiments.
[0052] Specifically, Figure 2 As shown, the camera parameter determination system 114 provides the digital image 202 to a critical edge detection neural network 204, which generates a vanishing edge map 206. The camera parameter determination system 114 can then apply a geometric model 208 to the vanishing edge map 206. In addition, the camera parameter determination system 114 can determine camera parameters 210 for the digital image 202 using the geometric model 208.
[0053] like Figure 2As shown, the camera parameter determination system 114 receives a digital image 202 and determines camera parameters for the image. The digital image 202 may be any of a variety of file types and may depict any of a variety of scene types. Furthermore, the camera parameter determination system 114 may determine camera parameters for a number of digital images and do so without regard to similarities or differences between file or scene types of the collection of digital images.
[0054] like Figure 2 As further shown, the camera parameter determination system 114 can utilize the critical edge detection neural network 204 to generate the vanishing edge map 206. As described in more detail with respect to FIGS. 3-4, the camera parameter determination system 114 can train the critical edge detection neural network to generate the vanishing edge map 206 including vanishing lines from the digital image. The critical edge detection neural network 204 can identify lines in the digital image 202 and can determine whether each line in the digital image is a vanishing line (e.g., a vanishing line corresponding to a vanishing point). As shown, the critical edge detection neural network 204 utilizes the vanishing lines identified from the digital image 202 to generate the vanishing edge map (e.g., it includes each vanishing line from the digital image 202, and no lines from the digital image that are not vanishing lines).
[0055] like Figure 2 As shown, the camera parameter determination system 114 generates camera parameters 210 from the vanishing edge map 206 using a geometric model 208. The geometric model 208 detects vanishing lines in the vanishing edge map 206 and uses the vanishing lines to estimate vanishing points and camera parameters for the digital image 202. This can be done by grouping the oriented lines from the vanishing edge map 206 into multiple vanishing points in the scene and performing iterative refinement. Applying the geometric model 208 to the vanishing edge map 206 (rather than the digital image 202) results in more accurate results because there are no misleading or "erroneous" line segments. Therefore, based on applying the geometric model 208 to the vanishing edge map 206, the camera parameter determination system 114 is able to more accurately and efficiently determine the camera parameters 210 for the digital image 202.
[0056] The geometric model 208 estimates the vanishing lines, vanishing points, and camera parameters of the digital image using camera calibration based on optimization. The geometric model can also perform vertical adjustments, where the model automatically modifies the lines in the image to straighten the digital image. These adjustments make the tilted lines consistent with the way human perception expects to see the image. In other words, the geometric model 208 can remove distortion relative to human observation. Such adjustments may be helpful in the context of images with strong geometric cues, such as images that include large man-made structures.
[0057] To determine the camera parameters for the image, the geometric model 208 performs an edge detection algorithm on the digital image to identify lines in the digital image. The geometric model 208 utilizes an energy function and iteratively optimizes the function to estimate various matrices (e.g., camera intrinsic parameter matrix, orientation matrix), which can then be used to estimate vanishing points, vanishing lines, and camera parameters for the image. However, as described above, for images without strong geometric cues or with many curved or closely parallel lines, the geometric model 208 may not be accurate in its estimation.
[0058] For example, in one or more embodiments, the camera parameter determination system 114 applies the geometric model 208 by utilizing the method described in Elya Shechtman, Jue Wang, Hyunjoon Lee, and Seungyong Lee in Camera Calibration and Automatic Adjustment Of Images, Patent No. 9098885B2, the entire contents of which are incorporated herein by reference. Similarly, the camera parameter determination system 114 can apply the geometric model 208 by utilizing the method described in Hyunjoon Lee, Eli Shechtman, Jue Wang, and Seungyong Lee in Automatic Upright Adjustment of Photographs, Journal of Latex Class Files, Vol. 6, No. 1 (January 2007), the entire contents of which are incorporated herein by reference.
[0059] In one or more embodiments, the camera parameter determination system 114 utilizes an alternative approach in response to determining that the vanishing edge map is insufficient for the geometric model 208 to generate accurate results. For example, in response to determining that the vanishing edge map has an insufficient number of vanishing lines for use in the geometric model 208 (e.g., less than a threshold number of vanishing lines that satisfies a confidence threshold), the camera parameter determination system 114 can utilize a direct convolutional neural network approach. Specifically, in one or more embodiments, the camera parameter determination system utilizes a convolutional neural network (CNN) based approach to directly determine camera parameters for a digital image.
[0060] As described above, the camera parameter determination system 114 may generate ground truth data for training a critical edge detection neural network. Figures 3A-3C The generation of ground truth data according to one or more embodiments is shown. More specifically, Figures 3A-3C A camera parameter determination system is shown that identifies lines in a digital image and determines whether each line is a vanishing line to generate a ground truth vanishing line set for the digital image. Specifically, Figure 3A The camera parameter determination system 114 is shown generating an unfiltered edge map. Then, Figure 3B The camera parameter determination system 114 is shown mapping the unfiltered edge map onto the spherical panorama and identifying the main vanishing points. Figure 3C The camera parameter determination system 114 is shown generating a ground truth edge map using the principal vanishing points and the unfiltered edge map.
[0061] like Figure 3A As shown, the camera parameter determination system 114 receives the digital image 302 and generates an unfiltered edge map 304 based on the digital image 302. The unfiltered edge map 304 includes lines detected from the digital image 302, (i.e., vanishing lines and non-vanishing lines from the digital image 302). To generate the unfiltered edge map 304, the camera parameter determination system 114 applies an edge detection algorithm to the digital image 302 to identify each of the lines from the digital image.
[0062] like Figure 3B As shown, the camera parameter determination system 114 may map the unfiltered edge map 304 onto a sphere (i.e., a spherical panorama that includes the sphere or a portion of the sphere), thereby generating an edge map sphere 306. Like the unfiltered edge map 304, the edge map sphere 306 includes lines detected from the digital image 302. After mapping the unfiltered edge map 304 onto the sphere, the camera parameter determination system 114 divides the edge map sphere 306 into various regions or “bins”. In addition, in one or more embodiments, the camera parameter determination system 114 identifies each intersection on the edge map sphere 306 and, based on the intersections, generates an intersection map sphere 308.
[0063] As described above, in one or more embodiments, the camera parameter determination system 114 may map the training lines from the panoramic digital image onto a spherical panorama (e.g., a sphere or ball). The camera parameter determination system 114 may then "sample" different portions of the panoramic digital image by dividing the image into several different, possibly overlapping sub-images. The camera parameter determination system 114 may prepare these sub-images as training images and perform the following steps with respect to Figure 3B-3C The steps listed above can be followed and these images can be used as training images when training a critical edge detection neural network.
[0064] like Figure 3BAs shown, the camera parameter determination system 114 can perform distance-based voting 310 for the vanishing points. As described above, in one or more embodiments, the camera parameter determination system 114 can determine a set of "bins" or areas on the spheres 306, 308. In one or more embodiments, the camera parameter determination system 114 divides the spheres uniformly or evenly. Additionally, in one or more embodiments, the camera parameter determination system 114 can determine distance-based "votes" for lines on the edge mapping sphere 306 and / or intersections on the intersection mapping sphere 308. The camera parameter determination system 114 can perform a Hough transform on the line segments detected on the edge mapping sphere 306, or any of a variety of similar transforms.
[0065] More specifically, the camera parameter determination system 114 may utilize any of a variety of voting schemes to determine the primary vanishing points of the digital image 302. The camera parameter determination system 114 may determine votes from pixels, lines, or intersections based on distance or orientation with respect to various potential vanishing points. For example, the camera parameter determination system 114 may initiate a pairwise voting scheme where the intersections between the lines are considered with respect to each of the lines involved in the intersection. In another embodiment, the camera parameter determination system 114 may determine the angular distance between each of the lines on the edge-mapped sphere 306 and each of the potential vanishing points to determine a "vote."
[0066] Finally, based on the votes, the camera parameter determination system 114 can determine the primary vanishing points 312 of the image. In one or more embodiments, the camera parameter determination system 114 identifies a predetermined number of vanishing points (e.g., the top three boxes as the top three vanishing points). In other embodiments, the camera parameter determination system 114 identifies an arbitrary number of vanishing points (e.g., any vanishing points that satisfy a threshold number or voting percentage). In some embodiments, the camera parameter determination system 114 identifies vanishing points corresponding to mutually orthogonal vanishing directions. For example, the horizontal vanishing directions 314, 316 and the vertical vanishing direction 318 are mutually orthogonal.
[0067] like Figure 3C As shown, the camera parameter determination system 114 performs an action 320 of weighting the lines using the principal vanishing points 312 and the unfiltered edge map 304 to generate a ground truth edge map 322. Specifically, the camera parameter determination system 114 may weight each line in the unfiltered edge map 304. The camera parameter determination system 114 may determine the weight (i.e., training weight) of each of the lines based on the distance between each of the lines of the image 302 and the principal vanishing points and / or based on the alignment of the line with each of the principal vanishing points of the image 302.
[0068] These training weights may reflect the probability or confidence value that each of the weighted lines is a vanishing line. In one or more embodiments, the weights are based on the distance from the identified vanishing point to the line, the intersection between the lines, or one or more of the pixels that make up the line. For example, the weights may be based on the angular distance between the vanishing point and the line. In another example, the weights may be based on the linear distance between the line intersection and the vanishing point. The camera parameter determination system 114 may then assign weights based on these measured distances, and may determine the weights based on a linear relationship between the distances and the weights themselves.
[0069] In one or more embodiments, the camera parameter determination system 114 may then use the weights for each of the lines to generate a ground truth edge map 316. The camera parameter determination system 114 may determine which lines to include in the ground truth vanishing edge map 322 as vanishing lines. In one or more embodiments, the weights may include a threshold at which the weights are considered vanishing lines. That is, weights above a predetermined threshold will cause the camera parameter determination system 114 to include the weighted lines in the ground truth vanishing edge map 322, while weights below a predetermined threshold will cause the camera parameter determination system 114 to exclude the weighted lines from the ground truth vanishing edge map 322.
[0070] The camera parameter determination system 114 may utilize further classification of lines in the digital image. For example, the camera parameter determination system 114 may utilize multiple (e.g., two or more) predetermined thresholds, one threshold determining a high weight within the vanishing edge map 316 and one threshold determining a low weight within the vanishing edge map 316. These two thresholds are given as examples, and it should be understood that the camera parameter determination system 114 may utilize any number of weight thresholds that are helpful in the context of the image to be processed. In addition, as described above, the camera parameter determination system 114 may utilize continuous weights that reflect the distance from the vanishing point.
[0071] Therefore, if Figure 3C As shown, the camera parameter determination system 114 generates a set of ground truth vanishing lines as part of the ground truth vanishing edge map 322. The camera parameter determination system 114 may repeat Figures 3A-3C The illustrated actions are applied to a plurality of digital images (eg, from an existing digital image repository) and generate a plurality of ground truth vanishing edge maps reflecting the ground truth vanishing lines for each of the digital images.
[0072] As described above, the camera parameter determination system 114 can utilize the ground truth vanishing lines of the training digital image to train a critical edge detection neural network. Figure 4 The camera parameter determination system 114 is shown for training a critical edge detection neural network 204 to accurately determine vanishing lines according to one or more embodiments. Figure 4 , the camera parameter determination system 114 utilizes the training images 402 and the corresponding ground truth vanishing line data 404 to train the critical edge detection neural network 204 .
[0073] like Figure 4 3, the camera parameter determination system 114 can utilize the training images 402 and the associated ground truth vanishing line data 404 to train the untrained critical edge detection neural network 406. As described above with respect to FIG3, the camera parameter determination system 114 can utilize the training images 402 to generate the ground truth vanishing line data 404. In one or more embodiments, the untrained critical edge detection neural network 406 utilizes the training images 402 and the ground truth vanishing line data 404 to learn to accurately identify vanishing lines from digital images.
[0074] Specifically, if Figure 4 As shown, the camera parameter determination system 114 generates a predicted vanishing edge map 408 using a critical edge detection neural network 406. Specifically, the critical edge detection neural network analyzes the training image 402 to generate a plurality of predicted vanishing lines. For example, the critical edge detection neural network can predict for each pixel whether the pixel belongs to a vanishing line in the training image 402.
[0075] Then, if Figure 4 As shown, the camera parameter determination system 114 compares the training vanishing edge map 408 for the training image 402 with the ground truth vanishing line data 404. More specifically, the camera parameter determination system 114 compares the training vanishing edge map 408 with the ground truth vanishing lines of the corresponding training image 402 using a loss function 410 that generates a calculated loss. Specifically, the loss function 410 can determine a loss metric (e.g., a difference metric) between the ground truth vanishing line data 404 and the vanishing lines in the training vanishing edge map 408.
[0076] In addition, if Figure 4 As shown, the camera parameter determination system 114 then uses the calculated loss to learn to more accurately identify vanishing lines from the digital image. For example, the camera parameter determination system 114 modifies the neural network parameters 414 (e.g., the internal weights of the layers of the neural network) to reduce or minimize the loss. Specifically, the camera parameter determination system uses back-propagation techniques to modify the internal parameters of the critical edge detection neural network 406 to reduce the loss caused by the application of the loss function 410.
[0077] By repeatedly analyzing training images, generating predicted vanishing edge maps, comparing the predicted vanishing edge maps to ground truth vanishing lines, and modifying neural network parameters, the camera parameter determination system 114 can train the critical edge detection neural network 406 to accurately identify vanishing lines from digital images. In practice, in one or more embodiments, the camera parameter determination system 114 iteratively trains the critical edge detection neural network 406 for a threshold amount of time, a threshold number of iterations, or until a threshold loss is reached. As described above, in one or more embodiments, the camera parameter determination system 114 utilizes a critical edge detection neural network that includes a convolutional neural network architecture. For example, Figure 5 2 shows an exemplary architecture and application of the critical edge detection neural network 204 according to one or more embodiments. Figure 5 As shown, and as about Figure 2 As discussed, critical edge detection neural network 204 receives digital image 202 , identifies vanishing lines in digital image 202 , and generates vanishing edge map 206 from digital image 202 that includes vanishing lines.
[0078] The critical edge detection neural network 204 can perform pixel-by-pixel prediction on the digital image 202. Specifically, the critical edge detection neural network can utilize a variant of an hourglass network. Therefore, in one or more embodiments, the critical edge detection neural network performs a bottom-up process by subsampling a feature map corresponding to the digital image 202, and performs a top-down process by upsampling the feature map by combining high-resolution features from the bottom layer. In one or more embodiments, instead of using a standard residual unit (e.g., a convolutional block) as a basic building block of the critical edge detection neural network 204, the critical edge detection neural network 204 includes an initial-like pyramid of features to identify vanishing lines. In this way, the camera parameter determination system 114 can capture multi-scale visual patterns (or semantics) when analyzing the digital image 202.
[0079] Typically, the structure of an hourglass neural network includes approximately equal top-down and bottom-up processing, and many contain two or more hourglass structures so that data is processed alternately bottom-up and top-down. This architecture allows the neural network to capture information at every scale and combine information across various resolutions, and has seen great success in the identification of objects in images. The specific structure of the various layers can vary depending on the purpose of the particular neural network, but a convolution-deconvolution architecture is employed that facilitates pixel-level prediction. For example, in one or more embodiments, the camera parameter determination system 114 can utilize an hourglass neural network, as described by Alejandro Newell, Kaiyu Yang, and Jia Deng in European Conference on Computer Vision (2016), Stacked Hourglass Networks for Human Pose Estimation, the entire contents of which are incorporated herein by reference.
[0080] To illustrate, Figure 5 As shown, the critical edge detection neural network 204 includes a plurality of pyramid feature units 502. Specifically, the critical edge detection neural network 204 utilizes various pyramid features similar to a pyramid feature network (PFN). Within these pyramid features, the critical edge detection neural network 204 may include various parallel convolutional layers similar to those found in a convolutional neural network (CNN) or more specifically in an inception network.
[0081] like Figure 5 As shown, the critical edge detection neural network 204 includes a plurality of convolutional layers, which include various convolution operations. Figure 5 An upper convolution branch 504 having a 1x1 convolution operation and a lower convolution branch 506 having a 1x1 convolution operation, a 3x3 convolution operation, and another 1x1 convolution operation are shown. These convolution branches 504, 506 together constitute a pyramid feature unit. As shown, the upper convolution branch 504 is parallel to the lower convolution branch 506. The pyramid feature unit 502 receives an input and performs a convolution operation on the input according to each of the convolution branches 504, 506. Then, at the output 508, the critical edge detection neural network 204 cascades the results and sends them to the next module. This approach helps to retain more details in the disappearing line segments and avoids blurry or distorted heat maps that may result from applying more traditional convolution blocks.
[0082] In one or more embodiments, the critical edge detection neural network 204 generates the vanishing edge map 206 by determining a confidence value for each pixel. The confidence value indicates a measure of confidence that the pixel from the critical edge detection neural network corresponds to a vanishing line (e.g., is included in or is part of a vanishing line). In other embodiments, the critical edge detection neural network 204 may determine a confidence value for each line. The critical edge detection neural network 204 may utilize these confidence values to determine which pixels and / or lines from the digital image to include in the vanishing edge map 206 as vanishing lines. In one or more embodiments, this determination is based on a predetermined threshold that the confidence values must comply with in order to be included. As described above, the camera parameter determination system 114 may then "feed" the vanishing edge map 206 to a geometric model, which will then consider only the lines included in the vanishing edge map to determine various camera parameters for the digital image.
[0083] In addition, the critical edge detection neural network 204 can determine weights for pixels and / or lines in the digital image based on confidence values corresponding to those pixels and / or lines. Then, based on these weights, the critical edge detection neural network can generate a weighted vanishing edge map that reflects the weights assigned in the edge map itself. That is, in one or more embodiments, the critical edge detection neural network 204 can generate a weighted vanishing edge map that reflects a measure of the confidence of each of the vanishing lines included therein. To illustrate, the weighted vanishing edge map can reflect that the geometric model should give more consideration to pixels and / or lines with higher confidence values, while the geometric model should give less consideration to pixels and / or lines with lower confidence values. Therefore, the camera parameter determination system 114 can use the geometric model and the weighted vanishing edge map to determine camera parameters.
[0084] In addition to (or as an alternative to) the confidence value, the camera parameter determination system 114 may also determine the weight based on the line length. Specifically, the camera parameter determination system 114 may give greater weight to pixels / lines corresponding to longer vanishing lines. Similarly, the camera parameter determination system 114 may give reduced weight to pixels / lines corresponding to shorter vanishing lines. Thus, the camera parameter determination system 114 may emphasize and / or de-emphasize pixels / lines based on the confidence value and / or line length.
[0085] The vanishing lines in the vanishing edge map 206 may correspond to different vanishing points (e.g., vanishing directions). More specifically, the vanishing edge map 206 may include vertical vanishing lines (e.g., vanishing lines having intersections corresponding to the vertical vanishing directions) and horizontal vanishing lines (e.g., vanishing lines having intersections corresponding to the horizontal vanishing directions). It should be understood that the camera parameter determination system 114 may utilize both vertical and horizontal vanishing lines to determine camera parameters for the digital image. That is, the critical edge detection neural network 204 may determine both vertical vanishing lines and horizontal vanishing lines, and may generate a vanishing edge map 206 including both vanishing lines. The camera parameter determination system 114 may utilize the vanishing edge map 206 and the geometric model to determine the camera parameters for the digital image.
[0086] As mentioned above about Figure 2 As discussed, the camera parameter determination system 114 may utilize the vanishing edge map and the geometric model to determine the camera parameters of a single digital image. Figure 6 The camera parameter determination system 114 is shown utilizing camera parameters 602 for a digital image 604 to perform various actions with the digital image 604. For example, the camera parameter determination system 114 presents a graphical user interface 606 for modifying the digital image. Based on user interaction with the graphical user interface 606, the camera parameter determination system 114 may utilize the camera parameters 602 to generate an enhanced digital image.
[0087] For example Figure 6 As shown, the camera parameter determination system 114 uses the camera parameters 602 to add the chair to the digital image 604. The camera parameter determination system 114 uses the camera parameters 602 to properly position the chair in the three-dimensional space (e.g., scene) depicted in the digital image 604. Indeed, as shown, the camera parameters indicate the pitch, roll, and yaw of the camera, which allows the geometric model to properly orient the chair within the digital image. Although Figure 6 An object is shown being added to the digital image 604, but it should be understood that the camera parameter determination system 114 can perform any of a variety of photo editing functions, including removing objects, modifying lighting or textures, combining images, or any other editing function. In addition, the camera parameters can also be used to generate a three-dimensional model of the scene depicted in the digital image.
[0088] Furthermore, if Figure 6 As shown, the camera parameter determination system 114 provides a visual aid based on the camera parameters 602 to be displayed with the digital image 604 to facilitate photo editing (e.g., by enabling more precise interaction with the user interface). Specifically, the camera parameter determination system 114 provides a line of sight corresponding to a vanishing point within the digital image 604. This approach enables precision in editing, thereby requiring less interaction before achieving the desired result.
[0089] In addition, the camera parameter determination system 114 can utilize the camera parameters 602 in the context of image searching. For example, in one or more embodiments, the camera parameter determination system 114 utilizes the camera parameters 602 to identify digital images with similar or identical camera parameters from an image database. Specifically, the camera parameter determination system 114 can determine the camera parameters for all digital images in the image database. The camera parameter determination system 114 can then search based on the determined camera parameters. For example, the camera parameter determination system 114 can provide a user interface in which the user 106 can specify search parameters based on one or more camera parameters (e.g., images with the same pitch, roll, and yaw as the input digital image). The camera parameter determination system 114 can then identify images that meet the search parameters.
[0090] Reference now Figure 7 , additional details will be provided regarding the capabilities and components of the camera parameter determination system 114 according to one or more embodiments. Specifically, Figure 7 A schematic diagram of an example architecture of a camera parameter determination system 114 hosted on a computing device 701 is shown. The camera parameter determination system 114 may represent one or more embodiments of the camera parameter determination systems 114 described previously.
[0091] As shown, the camera parameter determination system 114 is located on a computing device 701 as part of the digital media management system 112, as described above. In general, the computing device 701 can represent various types of computing devices (e.g., server device 110 or client device 102). For example, in some embodiments, the computing device 701 is a non-mobile device, such as a desktop or server. In other embodiments, the computing device 701 is a mobile device, such as a mobile phone, a smart phone, a PDA, a tablet computer, a laptop computer, etc. Additional details about the computing device 701 are provided below. Fig.10 Have a discussion.
[0092] like Figure 7 As shown, the camera parameter determination system 114 includes various components for performing the processes and features described herein. For example, the camera parameter determination system 114 includes a ground truth vanishing line data engine 702, a critical edge detection neural network 704, a neural network training engine 706, a camera parameter engine 707, and a data storage device 708. Each of these components is described in turn below.
[0093] like Figure 7As shown, the camera parameter determination system 114 may include a ground truth vanishing line data engine 702. The ground truth vanishing line data engine 702 may create, generate, and / or provide ground truth data to the camera parameter determination system 114. Figures 3A-3C As discussed, the ground truth vanishing line data engine 702 can generate ground truth vanishing lines for training a set of digital images. More specifically, the ground truth vanishing line data engine 702 can map digital images onto a sphere, determine intersection points in the digital images, determine vanishing points in the digital images, measure various distances in the digital images, and determine ground truth vanishing lines for the digital images.
[0094] In addition, if Figure 7 As shown, the camera parameter determination system 114 also includes a critical edge detection neural network 704. As described above with respect to FIGS. 3-5, the critical edge detection neural network 704 can determine vanishing lines in the digital image and generate a vanishing edge map, which includes the vanishing lines from the digital image (but does not include other lines from the digital image). As described below, the critical edge detection neural network 704 can be trained by a neural network training engine 706.
[0095] In addition, if Figure 7 As shown, the camera parameter determination system 114 also includes a neural network training engine 706. The neural network training engine 706 can train a neural network to perform various tasks using ground truth data. Figure 4 As discussed in more detail, the neural network training engine 706 can minimize the loss function and utilize ground truth data. More specifically, the neural network training engine 706 can train the critical edge detection neural network 704 to identify vanishing lines and generate a vanishing edge map. The neural network training engine 706 can use the ground truth data from the ground truth vanishing line data engine 702.
[0096] In addition, if Figure 7 As shown, the camera parameter determination system 114 includes a camera parameter engine 707. As discussed in more detail above, the camera parameter engine 707 may determine camera parameters based on a vanishing edge map generated by the critical edge detection neural network 704. As described above, the camera parameter engine may utilize a geometric model.
[0097] Likewise, Figure 7 As shown, the camera parameter determination system 114 includes a storage manager 708. The storage manager 708 can store and / or manage data on behalf of the camera parameter determination system 114. The storage manager 708 can store any data related to the camera parameter determination system 114. For example, the storage manager 708 can store disappearing line data 710 and camera parameter data 712.
[0098] Figure 7 A schematic diagram of a computing device 701 on which at least a portion of a camera parameter determination system 114 may be implemented in accordance with one or more embodiments is shown. Each component 702-712 of the camera parameter determination system 114 may include software, hardware, or both. For example, the components 702-712 may include one or more instructions, one or more instructions stored on a computer-readable storage medium, and may be executed by a processor of one or more computing devices, such as a client device 102 or a server device 110. When executed by one or more processors, the computer-executable instructions of the camera parameter determination system 114 may cause the computing device to perform the methods described herein. Alternatively, the components 702-712 may include hardware, such as a dedicated processing device that performs a specific function or group of functions. Alternatively, the components 702-712 of the dynamic representation management system 114 may include a combination of computer-executable instructions and hardware.
[0099] In addition, the components 702-712 of the camera parameter determination system 114 can be implemented, for example, as one or more operating systems, as one or more independent applications, as modules of one or more applications of an application, as one or more plug-ins, as one or more library functions or functions that other applications can call, and / or as a cloud computing model. Therefore, the components 702-712 can be implemented as independent applications, such as desktop or mobile applications. In addition, the components 702-712 can be implemented as one or more network-based applications hosted on a remote server. The components 702-712 can also be implemented in a set of mobile device applications or "apps". For illustration, the components 702-712 can be implemented in applications, including but not limited to LUMETRI TM , or ADOBE ADOBE, ADOBE DIMENSION, ADOBESTOCK, PHOTOSHOP, LIGHTROOM, PAINTCAN, LUMETRI, and ADOBE PREMIERE are registered trademarks or trademarks of Adobe Inc in the United States and / or other countries.
[0100] Figure 1-7 , and the corresponding text and examples provide many different methods, systems, devices, and non-transitory computer-readable media for the camera parameter determination system 114. Figure 8-9 As shown, in addition to the above description, one or more embodiments may also be described in terms of flowcharts that include acts for achieving certain results. Figure 8-9Can be performed with more or less actions.In addition, actions can be performed in different orders.In addition, the actions described herein can be repeated or performed in parallel with each other or with different instances of the same or similar actions.
[0101] As mentioned above, Figure 8-9 800, 900 for training and utilizing a critical edge detection neural network in accordance with one or more embodiments. Figure 8-9 Actions according to one embodiment are shown, but alternative embodiments may omit, add, reorder, and / or modify Figure 8-9 Any action shown. Figure 8-9 The actions of may be performed as part of a method. Alternatively, the non-transitory computer readable medium may include instructions that, when executed by one or more processors, cause the computing device 701 to perform Figure 8-9 In some embodiments, the system may perform Figure 8-9 action.
[0102] like Figure 8 As shown, a series of actions 800 includes an action 802 of identifying an image captured by a camera having camera parameters. For example, action 802 may involve identifying a digital image captured via a camera device having one or more camera parameters. Action 802 may also involve determining one or more parameters by applying a geometric model to a first vanishing line set and a second vanishing line set. In addition, action 802 may involve identifying digital images using an image search database to identify images having certain features.
[0103] In addition, if Figure 8As shown, a series of actions 800 includes an action 804 of generating a vanishing edge map from a digital image using a critical edge detection neural network 204. For example, action 804 may involve generating a vanishing edge map from a digital image using a critical edge detection neural network 204, wherein the vanishing edge map includes a plurality of vanishing lines from the digital image corresponding to vanishing points in the digital image, and wherein the critical edge detection neural network 204 is trained to generate the vanishing edge map from a training digital image and ground truth vanishing lines corresponding to ground truth vanishing points of the training digital image. In addition, action 804 may involve using the critical edge detection neural network 204, wherein the vanishing edge map includes vanishing lines from the digital image. In addition, action 804 may involve using the critical edge detection neural network 204 to generate a weighted vanishing edge map that reflects the confidence and / or probability that each line is a vanishing line. Action 804 may also include generating a vanishing edge map by generating a first vanishing line set and a second vanishing line set, wherein the first vanishing line set has intersections corresponding to vanishing points corresponding to a horizontal vanishing direction, and the second vanishing line set has second intersections corresponding to vanishing points corresponding to a vertical vanishing direction.
[0104] The vanishing edge map of action 804 may include confidence values corresponding to pixels of the digital image, the confidence values including a measure of confidence that the pixels correspond to vanishing lines. Furthermore, action 804 may include: determining weights for a plurality of lines based on the confidence values; generating a weighted vanishing edge map based on the weights for the plurality of lines; and generating one or more camera parameters based on the weighted vanishing edge map. Furthermore, the critical edge detection neural network 204 of action 804 may include a convolutional neural network.
[0105] In addition, if Figure 8 As shown, a series of actions 800 includes an action 806 of determining camera parameters for an image using a vanishing edge map. For example, action 806 may involve determining one or more camera parameters corresponding to a digital image using the vanishing edge map. Action 806 may also involve determining one or more camera parameters for a digital image using a weighted vanishing edge map. Action 806 may also include determining the camera parameters by placing more consideration on pixels and / or lines given larger weights in the weighted vanishing edge map and placing less consideration on pixels and / or lines given smaller weights in the weighted vanishing edge map.
[0106] In addition, the camera parameters of act 806 may include at least one of focal length, pitch, roll, or yaw. Act 806 may also involve determining one or more camera parameters of the digital image using a geometric model. In addition, the geometric model may determine one or more camera parameters of the digital image based on vanishing lines included in the vanishing edge map and excluding other lines from the digital image. In addition, act 806 may involve determining the camera parameters of the image using a geometric model, wherein the geometric model determines the camera parameters of the image based at least in part on one or more confidence values and / or weights associated with the vanishing lines included in the vanishing edge map.
[0107] Go to Fig. 9 , a series of actions 900 includes an action 902 of determining a vanishing point of a training image using training lines in a training image. For example, action 902 may include determining a vanishing point of a training digital image using training lines in a training digital image in a plurality of training images. Action 902 may also involve subdividing the training digital image into regions. Furthermore, action 902 may include determining a predetermined number of primary vanishing points based on the regions that received the most votes in a voting scheme.
[0108] Action 902 may also involve mapping the training lines to the spherical panorama, analyzing the training lines to generate a plurality of votes for a plurality of candidate vanishing regions, and determining a vanishing point from the plurality of candidate vanishing regions based on the plurality of votes. Additionally, action 902 may involve a voting scheme in which each pixel in the image votes for one or more regions of the training digital image as vanishing points. Additionally, action 902 may include a voting scheme in which each line in the training digital image votes for one or more regions of the image as vanishing points. Additionally, action 902 may include applying a Hough transform (or any of a variety of similarity transforms) to the training lines on the spherical panorama. Additionally, action 902 may involve determining a vertical vanishing point, a first horizontal direction, and a second horizontal direction of the training digital image.
[0109] Likewise, if Fig. 9 As shown, a series of actions 900 includes an action 904 of generating a set of ground truth vanishing lines for a training image. For example, action 904 may involve using an edge-mapped sphere to generate a set of ground truth vanishing lines for training a digital image. In addition, action 904 may involve subdividing the edge-mapped sphere into a plurality of sub-images, and determining the vanishing lines of each of the sub-images to serve as a training image for a critical edge detection neural network.
[0110] In addition, if Fig. 9As shown, a series of actions 900 includes an action 906 of determining distances between vanishing points and training lines. For example, action 906 may involve determining distances between vanishing points in the vanishing points and the training lines. Action 906 may also involve determining a first angular distance between a first vanishing line in the vanishing lines and a first vanishing point in the vanishing points. Furthermore, action 906 may involve determining distances between one or more intersection points of one or more line segments and one or more of the vanishing points.
[0111] In addition, if Fig. 9 As shown, a series of actions 900 includes an action 908 of including a training line as a vanishing line based on the determined distance. For example, action 908 may involve including the training line in a set of ground truth vanishing lines based on a distance between a vanishing point and the training line. Action 908 may also involve determining a weight of the training line based on the distance between the vanishing point and the training line, and comparing the weight of the training line to a distance threshold. Action 908 may additionally involve including the training line in a set of ground truth vanishing lines based on comparing the weight of the training line to a distance threshold.
[0112] In addition, if Fig. 9 As shown, a series of actions 900 includes an action 910 of generating predicted vanishing lines from a training image using a critical edge neural network. For example, action 910 may involve generating predicted vanishing lines from a training digital image using a critical edge detection neural network 204. Action 910 may also include utilizing a pyramid feature unit of a layer of the critical edge detection neural network, wherein the pyramid feature unit includes a convolution operation in parallel with a series of multiple convolution operations. In addition, action 910 may involve concatenating the results of the convolution operation from each layer of the pyramid feature unit.
[0113] Likewise, if Fig. 9 As shown, a series of actions 900 includes an action 912 of modifying parameters of a critical edge neural network by comparing the predicted vanishing lines to the ground truth vanishing lines. For example, action 912 can involve modifying parameters of the critical edge detection neural network 204 by comparing the predicted vanishing lines to a set of ground truth vanishing lines. In addition, action 912 can involve modifying parameters of the critical edge detection neural network 204 using a loss function, wherein the critical edge detection neural network modifies the parameters based on minimization of a loss determined by the loss function.
[0114] In addition to (or as an alternative to) the above-described actions, in some embodiments, the series of actions 800, 900 includes steps for training a critical detection edge neural network to generate a vanishing edge map from a training digital image. Figure 3A-5The described methods and acts may include corresponding acts for training a critical detection edge neural network to generate a vanishing edge map from a training digital image.
[0115] In addition to (or as an alternative to) the above actions, in some embodiments, the series of actions 800, 900 includes steps for generating a vanishing edge map of the digital image using a critical edge detection neural network. Figure 2 and Figure 5 The described methods and acts may include corresponding acts for generating a vanishing edge map of a digital image using a critical edge detection neural network.
[0116] Embodiments of the present disclosure may include or utilize a special-purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in more detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Specifically, one or more of the processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). Typically, a processor (e.g., a microprocessor) receives instructions from a non-transitory computer-readable medium (e.g., a memory) and executes those instructions to perform one or more processes, including one or more of the processes described herein.
[0117] Computer-readable media can be any available media that can be accessed by a general or special-purpose computer system. A computer-readable medium that stores computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium that carries computer-executable instructions is a transmission medium. Therefore, by way of example and not limitation, embodiments of the present disclosure may include at least two distinctly different types of computer-readable media: a non-transitory computer-readable storage medium (device) and a transmission medium.
[0118] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid-state drives (“SSD”) (e.g., RAM-based), flash memory, phase-change memory (“PCM”), other types of memory, other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired program code components in the form of computer-executable instructions or data structures and that can be accessed by a general or special purpose computer.
[0119] "Network" is defined as one or more data links that enable the transmission of electronic data between computer systems and / or modules and / or other electronic devices. When information is transmitted or provided to a computer via a network or other communication connection (hardwired, wireless, or a combination of hardwired or wireless), the computer properly regards the connection as a transmission medium. The transmission medium may include a network and / or a data link that can be used to carry the desired program code components, which are in the form of computer executable instructions or data structures and can be accessed by general or special computers. The above combinations should also be included in the scope of computer-readable media.
[0120] In addition, upon reaching various computer system components, program code means in the form of computer executable instructions or data structures may be automatically transferred from a transmission medium to a non-transitory computer readable storage medium (device) (or vice versa). For example, computer executable instructions or data structures received over a network or data link may be buffered in RAM within a network interface module (e.g., a "NIC") and then ultimately transferred to the computer system RAM and / or a less volatile computer storage medium (device) at the computer system. Thus, it should be understood that a non-transitory computer readable storage medium (device) may be included in a computer system component that also (or even primarily) utilizes a transmission medium.
[0121] Computer executable instructions include, for example, instructions and data, which, when executed by a processor, cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a specific function or function group. In some embodiments, computer executable instructions are executed by a general-purpose computer to transform a general-purpose computer into a special-purpose computer that implements the elements of the present disclosure. Computer executable instructions can be, for example, binary, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in a language specific to structural features and / or method actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features or actions described above. On the contrary, the described features and actions are disclosed as example forms for implementing the claims.
[0122] Those skilled in the art will appreciate that the present disclosure can be practiced in a network computing environment with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, etc. The present disclosure can also be practiced in a distributed system environment, in which local and remote computer systems linked by a network (by a hardwired data link, a wireless data link, or by a combination of a hardwired and wireless data link) all perform tasks. In a distributed system environment, program modules can be located in local and remote memory storage devices.
[0123] Embodiments of the present disclosure may also be implemented in a cloud computing environment. As used herein, the term "cloud computing" refers to a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing may be employed in the marketplace to provide universal and convenient on-demand access to a shared pool of configurable computing resources. A shared pool of configurable computing resources may be rapidly provisioned through virtualization and released with less management effort or service provider interaction, and then expanded accordingly.
[0124] The cloud computing model may consist of various features, such as, for example, on-demand self-service, broad network access, resource pools, rapid elasticity, measured services, and the like. The cloud computing model may also expose various service models, such as software as a service ("SaaS"), platform as a service ("PaaS"), and infrastructure as a service ("IaaS"). The cloud computing model may also be deployed using different deployment models, such as private cloud, community cloud, public cloud, hybrid cloud, and the like. Additionally, as used herein, the term "cloud computing environment" refers to an environment in which cloud computing is employed.
[0125] Fig.10 A block diagram of an example computing device 1000 is shown, which can be configured to perform one or more of the above-described processes. It will be understood that one or more computing devices such as computing device 1000 can represent the above-described computing devices (e.g., computing device 701, server device 110, and client device 102). In one or more embodiments, computing device 1000 can be a mobile device (e.g., a mobile phone, a smart phone, a PDA, a tablet computer, a laptop computer, a camera, a tracker, a watch, a wearable device, etc.). In some embodiments, computing device 1000 can be a non-mobile device (e.g., a desktop computer or another type of client device 102). In addition, computing device 1000 can be a server device that includes cloud-based processing and storage capabilities.
[0126] like Fig.10 As shown, computing device 1000 may include one or more processors 1002, memory 1004, storage device 1006, input / output interface 1008 (or "I / O interface 1008"), and communication interface 1010, which may be communicatively coupled by way of a communication infrastructure (e.g., bus 1012). Fig.10 The computing device 1000 is shown in FIG. Fig.10 The components shown in FIG. 1 are not intended to be limiting. Additional or alternative components may be used in other embodiments. Furthermore, in some embodiments, computing device 1000 includes more than Fig.10 The components shown are fewer than the components shown. Fig.10 Components of computing device 1000 are shown.
[0127] In certain embodiments, processor 1002 includes hardware for executing instructions, such as those constituting a computer program. By way of example and not limitation, to execute instructions, processor 1002 may retrieve (or fetch) instructions from internal registers, internal cache, memory 1004, or storage device 1006, and decode and execute them.
[0128] The computing device 1000 includes a memory 1004 coupled to the processor 1002. The memory 1004 may be used to store data, metadata, and programs executed by the processor. The memory 1004 may include one or more of volatile and non-volatile memories, such as random access memory ("RAM"), read-only memory ("ROM"), solid-state disk ("SSD"), flash memory, phase change memory ("PCM"), or other types of data storage devices. The memory 1004 may be internal or distributed memory.
[0129] The computing device 1000 includes a storage device 1006 for storing data or instructions. As an example, but not by way of limitation, the storage device 1006 may include the non-transitory storage medium described above. The storage device 1006 may include a hard disk drive (HDD), a flash memory, a universal serial bus (USB) drive, or a combination of these or other storage devices.
[0130] As shown, computing device 1000 includes one or more I / O interfaces 1008, which are provided to allow a user to provide input (e.g., user strokes) to computing device 1000, receive output from computing device 1000, and otherwise transfer data to and from computing device 1000. These I / O interfaces 1008 may include a mouse, keypad or keyboard, touch screen, camera, optical scanner, network interface, modem, other known I / O devices, or a combination of these I / O interfaces 1008. The touch screen may be activated with a stylus or finger.
[0131] The I / O interface 1008 may include one or more devices for presenting output to a user, including but not limited to a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In some embodiments, the I / O interface 1008 is configured to provide graphical data to the display for presentation to the user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may serve a particular implementation.
[0132] The computing device 1000 may further include a communication interface 1010. The communication interface 1010 may include hardware, software, or both. The communication interface 1010 provides one or more interfaces for communication (such as, for example, packet-based communication) between the computing device and one or more other computing devices or one or more networks. As an example and not limitation, the communication interface 1010 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wired-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network such as WI-FI. The computing device 1000 may further include a bus 1012. The bus 1012 may include hardware, software, or both that connect the components of the computing device 1000 to each other.
[0133] In the foregoing description, the present invention has been described with reference to specific example embodiments of the present invention. Various embodiments and aspects of the present invention are described with reference to the details discussed herein, and the accompanying drawings illustrate various embodiments. The above description and drawings are illustrative of the present invention and should not be construed as limiting the present invention. Many specific details are described to provide a thorough understanding of the various embodiments of the present invention.
[0134] Without departing from the spirit or essential features of the present invention, the present invention may be implemented in other specific forms. The described embodiments should be considered in all respects only as illustrative and non-restrictive. For example, the method described herein may be performed with fewer or more steps / actions, or the steps / actions may be performed in different orders. In addition, the steps / actions described herein may be repeated or performed in parallel with each other or with different instances of the same or similar steps / actions. Therefore, the scope of the present invention is indicated by the appended claims rather than the preceding description. All changes falling within the equivalent meanings and scopes of the claims should be included within their scope.
Claims
1. A method for determining camera parameters for a digital image, include: identifying a digital image captured via a camera device having camera parameters; generating an edge map by detecting edges within the digital image; generating a vanishing edge map by weighting the edges of the edge map based on whether the edges correspond to one or more vanishing lines of the digital image; and The camera parameters for the digital image are estimated using a geometric model and the vanishing edge map. 2 . The method of claim 1 , wherein estimating the camera parameters comprises estimating one or more of: focal length, pitch, roll, or yaw. 3 . The method of claim 1 , further comprising generating an enhanced digital image from the digital image based on the estimated camera parameters.
4. The method of claim 1 , wherein weighting the edges of the edge map based on whether the edges correspond to one or more vanishing lines of the digital image comprises determining a confidence value for pixels along the edges, the confidence value indicating a measure of confidence that the pixels correspond to vanishing lines. The method of claim 4 , further comprising excluding a plurality of edges of the edge map from the vanishing edge map based on the confidence values. 6 . The method of claim 5 , further comprising identifying that a subset of confidence values corresponding to the plurality of edges is below a predetermined threshold, and excluding the plurality of edges based on the subset of confidence values being below the predetermined threshold.
7. The method of claim 5, wherein estimating the camera parameters for the digital image using the geometric model and the vanishing edge map comprises considering the edges in the vanishing edge map using the geometric model without considering the plurality of edges excluded from the vanishing edge map.
8. The method of claim 1 , wherein weighting the edges of the edge map based on whether the edges correspond to one or more vanishing lines of the digital image comprises giving greater weights to edges corresponding to one or more longer vanishing lines, the one or more longer vanishing lines having a longer length than one or more shorter vanishing lines.
9. The method of claim 1, wherein generating the vanishing edge map comprises utilizing a critical edge detection neural network.
10. A non-transitory computer-readable storage medium comprising instructions that, when executed by at least one processor, cause a computing device to: identifying a digital image captured via a camera device having camera parameters; A vanishing edge map is generated from the digital image using a critical edge detection neural network by: generating an edge map by detecting edges within the digital image; and excluding a plurality of edges from the vanishing edge map based on a lack of a threshold correspondence of the plurality of edges of the edge map with one or more vanishing lines of the digital image; as well as The camera parameters for the digital image are estimated based on the vanishing edge map.
11. The non-transitory computer-readable storage medium of claim 10, further comprising instructions that, when executed by the at least one processor, cause the computing device to generate the vanishing edge map by generating one or more horizontal vanishing lines having intersection points corresponding to horizontal vanishing directions and one or more vertical vanishing lines having second intersection points corresponding to vertical vanishing directions.
12. The non-transitory computer-readable storage medium of claim 11, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine the camera parameters by applying the geometric model to the one or more horizontal vanishing lines and the one or more vertical vanishing lines.
13. The non-transitory computer-readable storage medium of claim 12, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine a confidence value for pixels along the edge, the confidence value indicating a measure of confidence that the pixel corresponds to a vanishing line.
14. The non-transitory computer-readable storage medium of claim 10, further comprising instructions that, when executed by the at least one processor, cause the computing device to: determining weights for the edges of the edge graph; generating a weighted vanishing edge graph based on the weights for the plurality of edges; and The camera parameters for the digital image are estimated based on the vanishing edge map by applying a geometric model to the weighted vanishing edge map. 15 . The non-transitory computer-readable storage medium of claim 14 , wherein applying the geometric model to the weighted vanishing edge map comprises relying more on edges having larger weights than edges having smaller weights when estimating the camera parameters.
16. A system, include: One or more storage devices, including: a digital image captured via a camera device having camera parameters; and Critical edge detection neural network; and One or more processors configured to cause the system to: generating a vanishing edge map from the digital image using the critical edge detection neural network by filtering out a plurality of edges detected in the digital image that lack a threshold correspondence with one or more vanishing lines of the digital image; and The camera parameters corresponding to the digital image are estimated using a geometric model and the vanishing edge map.
17. The system of claim 16, wherein the one or more processors are configured to cause the system to utilize the critical edge detection neural network to: Performing bottom-up processing on a feature map extracted from the digital image by downsampling; and The feature map extracted from the digital image is subjected to top-down processing by upsampling.
18. The system of claim 17, wherein the critical edge detection neural network comprises an hourglass neural network that allows for the bottom-up processing and the top-down processing.
19. The system of claim 16, wherein the one or more processors are configured to cause the system to utilize the critical edge detection neural network to identify vanishing lines for the digital image utilizing an initial-like pyramid of features extracted from the digital image.
20. The system of claim 19, wherein the one or more processors are configured to cause the system to utilize the critical edge detection neural network to capture multi-scale visual patterns in the digital image to assist in identifying the vanishing lines.
21. A system, include: One or more storage devices, including: a digital image captured via a camera device having camera parameters; and Critical edge detection neural network; and One or more computing devices configured to enable the system to: generating a vanishing edge map from the digital image using the critical edge detection neural network by generating a first set of vanishing lines having intersections corresponding to horizontal vanishing directions and a second set of vanishing lines having second intersections corresponding to vertical vanishing directions; and The camera parameters corresponding to the digital image are determined using the vanishing edge map by applying a geometric model to the first set of vanishing lines and the second set of vanishing lines.
22. The system of claim 21, wherein the one or more computing devices are further configured to cause the system to reproject the digital image so that the first set of vanishing lines and the second set of vanishing lines are aligned with a plurality of vanishing points.
23. The system of claim 21, wherein the vanishing edge map comprises confidence values corresponding to pixels of the digital image, the confidence values comprising a measure of confidence that the pixel corresponds to one or more vanishing lines in the first set of vanishing lines or the second set of vanishing lines.
24. The system of claim 23, wherein the one or more computing devices are further configured to cause the system to: determining a weight for the first set of vanishing lines based on the confidence value; generating a weighted vanishing edge map based on the weights for the first set of vanishing lines; and The camera parameters are generated based on the weighted vanishing edge map.
25. The system of claim 21, wherein the camera parameters include at least one of: focal length, pitch, roll, or yaw.
26. The system of claim 21, wherein the critical edge detection neural network comprises a convolutional neural network.
27. The system of claim 21, wherein the one or more computing devices are further configured to cause the system to generate an enhanced digital image from the digital image based on the determined camera parameters.
28. The system of claim 21, wherein the one or more computing devices are further configured to cause the system to generate one or more visual aids based on the determined camera parameters for display with the digital image during image editing.
29. The system of claim 21, wherein the one or more computing devices are further configured to cause the system to perform an image search based on the determined camera parameters.