Computer-implemented methods, systems, and non-transitory computer-readable media for automated analysis of visual data of images
By automating the analysis of image visual data and using neural networks to identify building spatial features, generate and match descriptors, the problem of capturing and using visual information inside buildings in existing technologies is solved, enabling rapid and highly accurate image acquisition location determination and improved navigation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-21
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies struggle to effectively capture, represent, and use visual information within buildings, including difficulties in constructing and maintaining floor plans, accurately scaling and filling room interior information, and visualizing and using floor plans in other ways.
The computer automatically analyzes the visual data of the images, uses a trained neural network to identify the spatial features of the buildings, generates circular descriptors for panoramic images, matches them with building location descriptors, determines the image acquisition location and orientation, and presents a two-dimensional floor plan of the building for navigation.
It enables the rapid and highly accurate determination of image acquisition locations without relying on depth sensors or other distance measurement devices, improving building navigation and information presentation while reducing computational resource requirements.
Smart Images

Figure CN116091914B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The following disclosure relates generally to techniques for automatically determining a capture location of an image on a floor plan of a building based on an analysis of visual data of the image relative to an analysis of the floor plan and subsequently using the determined capture location information in one or more ways (to improve navigation of the building using the capture location of the image). BACKGROUND
[0002] In various fields and situations (such as building analysis, property inventory, real estate acquisition and development, general contracting, cost estimates for renovations, automated navigation, etc.), it can be desirable to understand the interior of a house, office, or other building without having to physically go to and enter the building. However, it can be difficult to efficiently capture, represent, and use such building interior information, including displaying visual information captured within the interior of a building to a user at a remote location (e.g., to a user of a mobile device or other remote device) in a manner that enables the user to fully understand the layout and other details of the interior, including controlling the display in a manner selected by the user. Additionally, while a floor plan of a building can provide information about the layout and other details of the interior of the building, such use of a floor plan has several drawbacks, including that a floor plan can be difficult to construct and maintain, difficult to accurately scale and populate with information about the interior of a room, difficult to visualize and otherwise use, etc. For example SUMMARY
[0003] The present application provides computer-implemented methods, systems, and non-transitory computer-readable media that automate analysis of visual data of an image.
[0004] In one embodiment of the present application, a computer-implemented method of automating analysis of visual data of an image includes:
[0005] obtaining, by a computing device and for a building, building location description information including a plurality of building location circular descriptors for a plurality of building locations in the building, wherein each building location circular descriptor is associated with one of the building locations and has first angular information about a first potential spatial feature identified for a structural element of the building from a designated angular direction of the associated building location, wherein the first potential spatial feature is identified by a first trained neural network using a two-dimensional floor plan of the building;
[0006] generating, by the computing device, an image circular descriptor of a panoramic image, the panoramic image being captured in a room of the building and including visual information pertaining to at least some walls of the room, wherein the image circular descriptor has second angular information pertaining to second latent spatial features identified by a second trained neural network from the visual information of the panoramic image in a specified direction;
[0007] comparing, by the computing device, the image circular descriptor with the building location circular descriptors to determine one of the building location circular descriptors that is in the room and has first angular information that best matches the second angular information of the image circular descriptor;
[0008] associating, by the computing device and based on the comparison, the panoramic image with a determined location in the room and a determined orientation, the determined location being based on the building location associated with the determined one building location circular descriptor and the determined orientation being identified from the building location corresponding to a specified portion of visible information in the panoramic image; and
[0009] presenting, by the computing device, information including the two-dimensional floor plan of the building and showing the room with a visual indication identifying at least the determined location of the panoramic image, thereby using the presented information to navigate the building.
[0010] In another embodiment of the present application, a non-transitory computer readable medium stores content that causes one or more computing devices to perform automated operations, the automated operations comprising at least:
[0011] obtaining, by the one or more computing devices and for an image captured in an area associated with a building and including visual information pertaining to at least some structural elements of the building, an image circular descriptor of the image, the image circular descriptor being generated via analysis of the image by a first trained neural network and including information identifying features associated with the at least some structural elements in a specified direction within the visual information;
[0012] obtaining, by the one or more computing devices, building location circular descriptors, the building location circular descriptors each being associated with a building location, being generated via analysis of information about the associated building location by a second trained neural network, and including angular information about features associated with points of structural elements of the building in a specified angular direction from the associated building location;
[0013] comparing, by the one or more computing devices, the image circular descriptor and the building location circular descriptors to determine one of the building location circular descriptors having angular information that best matches the information included in the image circular descriptor;
[0014] associating, by the one or more computing devices, the image with a determined location for the building, the determined location being based on the associated building location for the determined one building location circular descriptor; and
[0015] providing, by the one or more computing devices, information about the determined location of the building for the image.
[0016] In yet another embodiment of the present application, a system for automated analysis of visual data of an image includes:
[0017] one or more hardware processors of one or more computing devices; and
[0018] one or more memories storing instructions that, when executed by at least one of the one or more hardware processors, cause at least one of the one or more computing devices to perform automated operations including at least:
[0019] obtaining description information for a region of buildings, the description information including building location circular descriptors for a plurality of building locations in the region, the building location circular descriptors being generated via a first trained neural network from analysis of information about the plurality of building locations, wherein each building location circular descriptor is associated with one of the plurality of building locations and has angular information about features associated with structural elements of the building in a specified angular direction from the associated building location;
[0020] generating, via a second trained neural network, an additional circular descriptor from analysis of recorded information for a recorded location in the region, wherein the additional circular descriptor includes information identifying features associated with at least some of the structural elements identifiable from the recorded information in a specified direction from the recorded location;
[0021] comparing the additional circular descriptor and the building location circular descriptors to determine one of the building location circular descriptors having angular information that best matches the information included in the additional circular descriptor;
[0022] based on the comparison, correlating the recorded information and the locations in the area, determining the locations in the area for the recorded locations based on the building locations associated with the determined one building location circular descriptor based on the building locations associated with the determined one building location circular descriptor; and
[0023] providing information related to the determined locations in the area for the recorded information. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 include diagrams depicting example building interior environments and computing systems used in embodiments of the present disclosure, including generating and presenting information representing interiors of buildings.
[0025] Figures 2A to 2I illustrate examples of automatically analyzing information related to floor plans of buildings and automatically analyzing visual data taken inside of buildings in order to automatically determine locations of image capture and present them on floor plans.
[0026] Figure 3 is a block diagram illustrating a computing system suitable for executing embodiments of a system that performs at least some of the techniques described in the present disclosure.
[0027] Figures 4A to 4B illustrate example embodiments of flowcharts for image floor plan location mapping manager (IFPLMM) system routines according to embodiments of the present disclosure.
[0028] Figure 5 illustrate example embodiments of flowcharts for building map viewer system routines according to embodiments of the present disclosure. DETAILED DESCRIPTION
[0029] The present disclosure describes techniques for using computing devices to perform automated operations related to determining locations of image capture based at least in part on analysis of visual data of images and comparison to corresponding analyzed floor plan information, and subsequently using the determined image capture location information in one or more further automated ways. In at least some embodiments, the images to be analyzed include one or more panoramic images or other images taken at one or more capture locations inside of a multi-room building (e.g., a house, an office, etc.) For example For example , and the determined image capture location information includes at least a location on a floor plan of the building and in some cases also includes an orientation or other directional information for at least a portion of the image, in at least some such implementations the automated image capture location determination is further performed without having or using any captured depth data from any depth sensor or other distance measuring device regarding distances from the capture location of the image to walls or other objects in the surrounding building. In various implementations, the determined image capture location information can also be used in various ways, such as in conjunction with a corresponding building floor plan and / or other generated mapping related information ( For example , a three-dimensional model of the interior of the building) including for controlling navigation of a mobile device ( For example , an autonomous vehicle), for display or other presentation in a corresponding GUI (graphical user interface) on one or more client devices, etc. Additional details regarding automated capture and use of determined image capture location information are included below, and in at least some implementations some or all of the techniques described herein can be performed via automated operation of an Image Floor Plan Location Mapping Manager ("IFPLMM") system, as further discussed below.
[0030] In at least some implementations and scenarios, for some or all of the captured images of a building each is a panoramic image captured at a capture location in or around the building, so that the panoramic image at each of a plurality of such capture locations is generated from one or more of: a video ( For example , taken at the capture location from a 360° video spun at the capture location by a user holding a smartphone or other mobile device), or a plurality of images captured in multiple directions from the capture location ( For example , spun at the capture location by a user holding a smartphone or other mobile device), or a simultaneous capture of all image information For example(using one or more fisheye lenses), etc. It will be understood that in some cases, such panoramic images can be presented using an isorectangular projection (where vertical lines and other vertical information in the environment are shown as straight lines in the projection, and where horizontal lines and other horizontal information in the environment are shown in a curved manner in the projection if they are above or below the horizontal center of the image, and where the amount of curvature increases with increasing distance from the horizontal center line), and providing up to 360° coverage around the horizontal axis and / or vertical axis, allowing a user viewing the initial panoramic image to move the viewing direction to different orientations within the initial panoramic image, resulting in different images (or "views") being rendered within the initial panoramic image (including using a planar coordinate system to present images rendered as stereoscopic images). Furthermore, in some embodiments, acquisition metadata regarding the capture of such panoramic images can be obtained and used in various ways, such as data acquired from IMU (Inertial Measurement Unit) sensors or other sensors of the mobile device while the mobile device is carried by the user or otherwise moved between acquisition locations, while in other embodiments, such acquisition metadata is not acquired or used. For example This is to determine the acquisition location of images on a building's floor plan based solely on visual data from the images. Additional details related to the acquisition and use of panoramic or other images of a building are included below.
[0031] As described above, the automated operation of the IFPLMM system may include determining the location of buildings within a defined area based at least in part on the analysis of visual information included in the image content. For example The acquisition location of images captured (within a room of a house or other building). In at least some embodiments, this automated determination of the acquisition location of images may include some or all of the following operations: using a trained neural network ( For example Deep convolutional neural networks (DNNs) are used to analyze images and determine the visual data features of image content. For example Based on the color and / or estimated depth of each pixel; based on information from associated groups of pixels, such as color and / or texture and / or estimated depth, or more generally, determining one or more latent spatial visual data features generated by a trained neural network; creating circular descriptors for images with degrees corresponding to the field of view of the image (e.g., a 360-degree circular descriptor representing a panoramic image with a 360° horizontal visual coverage, a 180-degree circular descriptor representing a panoramic image with a 180° horizontal visual coverage, a 72-degree circular descriptor representing a non-panoramic stereo image with a 72° horizontal visual coverage, etc.) and wherein each degree of the image circular descriptor encodes information about the corresponding degree in the image visual data from the determined features of the image ( For exampleFor example with respect to a direction within the image designated as 0° or otherwise designated as a starting angle or direction); and comparing the image circular descriptor to one or more corresponding building location circular descriptors from the building floor plan each representing corresponding locations associated with the building, where one or more of the building location circular descriptors are then identified as the best match and used to determine the capture location of the image in a room or other area associated with the building.
[0032] For purposes of an illustrative example, consider a panoramic image captured in a room of a building, where the panoramic image includes a 360° horizontal coverage around a vertical axis (e.g., a 360° horizontal coverage around a vertical axis of the camera used to capture the panoramic image) For example , a full circle showing all walls of the room from the capture location of the panoramic image, unless a portion of a wall is occluded by an intervening object in the room and / or by another intervening wall, such as for a room shape that is not purely rectangular, and where the x-axis and y-axis of the visual content of the image align with corresponding horizontal and vertical information in the room (e.g., the x-axis of the visual content of the image aligns with the horizontal direction of the room, and the y-axis of the visual content of the image aligns with the vertical direction of the room) For example , the boundaries between two walls, the boundaries between a wall and the floor, the bottom and / or top of a window and door, etc.), such that the image is not skewed or otherwise misaligned with respect to the room. For purposes of this example, the image capture can begin with a camera orientation in a north direction corresponding to a relative starting horizontal direction of 0° for the panoramic image, and continue in a circle performing image capture in a sequence of directions from the capture location using varying camera orientations, where a relative 90° horizontal direction for the panoramic image then corresponds to an east direction, a relative 180° horizontal direction for the panoramic image corresponds to a south direction, a relative 270° horizontal direction for the panoramic image corresponds to a west direction, and a relative 360° horizontal direction for the panoramic image returns to the north direction. In at least some implementations, information about the identified feature locations of the visual data of the panoramic image is encoded in a manner specific to such angles of direction from the capture location (e.g., a relative 0° horizontal direction for the panoramic image corresponds to a north direction, a relative 90° horizontal direction for the panoramic image corresponds to an east direction, a relative 180° horizontal direction for the panoramic image corresponds to a south direction, a relative 270° horizontal direction for the panoramic image corresponds to a west direction, and a relative 360° horizontal direction for the panoramic image returns to the north direction) For example , relative to the starting location of the panoramic image) resulting in an image circular descriptor for the image that encodes information for about 360° of visual coverage in the image (e.g., a 360° horizontal coverage around a vertical axis of the camera used to capture the panoramic image) Figures 2D to 2Ito combine feature information for a given horizontal direction of a series of vertical directions in that horizontal direction). In various implementations, such information about identified feature locations can be encoded and stored in various ways, including in some implementations an array or vector having one or more values for each direction angle to identify one or more features present at a given angular direction. Additionally, in other implementations and scenarios, feature information for an image can be identified and presented in ways other than based on angular differences from a starting direction of the image, resulting in other types of image descriptors being used in similar ways. Additional details regarding construction and use of such image circular descriptors are included below, including details regarding For example examples thereof and their associated descriptions.
[0033] As noted above, automated determination of image capture locations in a room (or other defined building area) of a building by implementations of the IFPLMM system can include matching angular information encoded in an image circular descriptor generated for an image to corresponding angular information in a building location circular descriptor generated from a floor plan of the building in order to determine a particular image capture location (and optionally orientation) in a room or other area. To generate a building location circular descriptor from a floor plan of a building, automated operations of the IFPLMM system can further include: starting with an existing floor plan of a building, such as a two-dimensional rasterized floor plan, that uses a top-down view to illustrate at least wall structure of the interior of the building and optionally has additional associated information such as semantic information about structural wall elements (e.g., whether a wall is a load-bearing wall, a non-load-bearing wall, a door, a window, and other inter-room wall openings) and / or structural information for additional areas associated with the building (e.g., one or more additional exterior structures such as a detached garage or carport, an accessory dwelling unit, a shed, etc.; one or more other exterior areas such as a backyard, a garden, a porch, a deck, a balcony, a patio, a walkway or other path, etc.); and generating a point cloud (e.g., a two-dimensional or “2D” point cloud; a three-dimensional or “3D” point cloud if additional height information is available; etc.) corresponding to at least the structural information of the floor plan, each point in the point cloud can be further assigned one or more types of associated information such as a 2D location on the floor plan (e.g., an X-Y coordinate), a relative location to a specified location (e.g., taking the lowest and leftmost point of the X-Y axis as 0, 0), a normal direction (e.g., a Z coordinate), etc. For example For example For example For example For example The location of a building on a floor plan is determined by the planar surface formed by the point and its adjacent points, as well as associated semantic information (if any) related to the location of the building. In at least some embodiments, this automated determination of circular descriptors for the location of one or more buildings on a floor plan may include some or all of the following operations: using a trained neural network ( For example Deep convolutional neural networks, such as those similar to those used to analyze visual data of images but trained separately, are used to analyze floor plan point cloud data and any additional associated information to determine corresponding features. For example One or more latent spatial data features generated by a trained neural network; and for each room or other area associated with the building, one or more circular descriptors of the building location for that room or other area are created. For example Multiple building location circular descriptors for multiple locations within the room or other area, such as in a grid or other selected information), each building location circular descriptor has a selected degree ( For example (360°) and each degree of the circular descriptor of the building location corresponds to information about any definite feature from the floor plan and the corresponding degree direction existing in the floor plan from the building location. For example The associated information of any floor plan point cloud points (where the direction specified as 0° or otherwise specified as the starting angle or direction, such as corresponding to the north direction or the upward direction on a 2D floor plan) is encoded. This building location circular descriptor can be predetermined, for example, before generating or using any corresponding image circular descriptor, or in some cases, can be dynamically created when compared with an image circular descriptor taken in the associated area of the building. In some embodiments, the characteristic information of the building location circular descriptor at a specific degree can encode information about the angle of incidence from the building location to any point in the point cloud at that specific degree from the building location (…). For example Using a set of enumerated ranges of incident angles, such as the normal direction relative to the point), and information on the estimated distance between the building location and the point's location ( For example, using a set of enumerated distance ranges). Additionally, in at least some implementations, the additional information can be associated with a building location circular descriptor for the building room or other area that was generated and used to compare the image circular descriptor for the new image, so as to also use and encode visual data from building location circular descriptors of one or more other images that have previously been localized within that room or other area (by at least determining a capture location of such images within the room or other area and by optionally further determining an orientation of the images within the room or other area so as to collectively specify a "pose" of the images within the room or other area, and such as where the orientation is optionally identified by associating particular image angles or other directions with corresponding floor plan angles or other directions of the room or other area so as to align different rotational coordinate systems of the images and floor plan, such as where the orientation is optionally identified by determining a geographic direction corresponding to directions within the images and floor plan, such as north, and the like).
[0034] Once multiple building location circular descriptors are generated or otherwise obtained for multiple room locations within a room or other area of a building, they can be compared or otherwise matched to an image circular descriptor for an image taken in the room or other area, so as to determine at least one of the best matching building location circular descriptors, where the capture location of the image is then identified based at least in part on the room location that the best matching building location circular descriptor most matches. For example, in some implementations and scenarios, the determined capture location of the image can be selected as the room location that the building location circular descriptor most matches, or conversely, in other implementations and scenarios, the determined capture location of the image can be determined to be within a small distance from the room location that the building location circular descriptor most matches (e.g., based on a distance metric used to compare the image circular descriptor to the building location circular descriptors). For example , such as by using a trained neural network (e.g., a neural network trained to determine a capture location of an image based on a difference between an image circular descriptor and a building location circular descriptor, such as a difference between an image circular descriptor and a building location circular descriptor that is most similar to the image circular descriptor). For example For example to a location between building locations having defined building location circular descriptors, separately from one or more neural networks used to determine features of the image and building floor plan point clouds. The matching process of the image circular descriptor and the building location circular descriptors can include determining a distance and / or a quantity of similarity / dissimilarity between the two circular descriptors in one or more ways, including in a rotation independent way, such as by determining a probability that the two circular descriptors match (where a highest probability of match corresponds to a minimum dissimilarity and / or distance), by measuring a difference between other encoding formats of the vectors of the circular descriptors being compared, and the like, as one non-exclusive example, a distance metric of a circular bulldozer can be used to compare vectors of two such circular descriptors in a rotation independent way.For example Regardless of whether the two circular descriptors use the same orientation as their corresponding relative starting points in the room, in other embodiments, the rotational differences between the two descriptors can be handled in other ways. Additionally, in some embodiments, the matching process may include matching the image circular descriptor with each possible building location circular descriptor ( Figures 2A to 2I The comparison is performed on all building location circular descriptors generated for one or more candidate buildings that may have had their images captured. In other implementations, only a subset of the building location circular descriptors for a specific building may be considered. For example This involves performing nearest neighbor gradient ascent or descent search using a defined similarity or dissimilarity metric. Additional details regarding the construction and use of such circular descriptors for building locations (including comparisons with one or more image circular descriptors) are included below, such as... For example Details of the examples and their associated descriptions.
[0035] In various embodiments, the described technology offers various benefits, including allowing the automatic enhancement of floor plans of multi-room buildings and other structures with information about the acquisition location of images taken in or around the building or other structure, including without having or using information from depth sensors or other distance measuring devices about distances from the image acquisition location to walls or other objects in the surrounding buildings or other structures. In at least some of these embodiments, the determination of image acquisition locations in areas associated with one or more buildings is further performed without having or using predicted room layouts from images and / or without any other images previously registered with determined acquisition locations on the building floor plan. Furthermore, this automated technology allows for the determination of such image acquisition location information faster than prior art and has higher accuracy in at least some embodiments, including by using information acquired from the actual building environment (rather than from a floor plan about how the building should theoretically be constructed), and if the corresponding building floor plan reflects the actual building environment and / or such changes, enabling the capture of structural element changes that occur after the initial construction of the building. This described technology further provides, at least in part, the ability to improve the performance of mobile devices based on the determined image acquisition locations. For example The benefits of semi-autonomous or fully autonomous vehicles for automated building navigation include a significant reduction in computational power and time spent attempting to learn the building's layout in other ways. Additionally, in some implementations, the described techniques can be used to provide an improved GUI where users can obtain information about the building's interior more accurately and quickly. For example(For use during internal navigation), including in response to search requests, as part of providing users with personalized information, as part of providing users with valuation estimates and / or other information about buildings, etc. The described technology also provides a variety of other benefits, some of which are further described elsewhere in this document.
[0036] As described above, the automated operation of the IFPLMM system may include: determining, at least in part, within a defined region based on the analysis of visual information included in the image content ( For example The acquisition location of images (within rooms of a house or other building). In at least some embodiments, this IFPLMM system can operate in conjunction with one or more separate ICA (Image Capture and Analysis) systems and / or one or more separate MIGM (Map Information and Generation Manager) systems to obtain and use building floor plans and other associated information from the MIGM system and / or obtain images of building location from the ICA system. In other embodiments, this IFPLMM system can incorporate some or all of the functionality of the ICA and / or MIGM systems as part of the IFPLMM system. In still other embodiments, the IFPLMM system can operate without using some or all of the functionality of the ICA and / or MIGM systems, such as if the IFPLMM system is from other sources ( For example Information about building floor plans and / or other related information is obtained from manual creation by one or more users, from the provision of such building floor plans and / or related information by one or more external systems or other sources, and / or if the IFPLMM system obtains information from other sources (e.g., from manual creation by one or more users, from the provision of such building floor plans and / or related information by one or more external systems or other sources). For example Information about the image to be located is obtained from the end user, such as through crowdsourcing. Additionally, the building floor plan used in the manner described herein can be in various formats (whether initially obtained and / or after initial automated analysis by the IFPLMM system), and in at least some embodiments includes a vectorized form containing specified information about the location of structural elements, such as one or more of the following: walls, windows, doorways and openings between other rooms, corners, etc. For example After initially receiving the building floor plans, analysis was performed to produce a vectorized form of the non-vectorized image.
[0037] Regarding the functionality of this ICA system, in at least some implementations, it can perform automated operations to collect data at one or more locations associated with a building. For example Acquire one or more images (inside multiple rooms of a building, at one or more external locations, etc.). For example, and optionally further gathering metadata related to the image gathering process and / or movement of the capture device between multiple gathering locations. For example, in at least some such implementations, such techniques can include capturing visual data from one or more gathering locations using one or more mobile devices (e.g., smartphones, tablets, etc.) held by a user and moved by the user around the gathering locations, and not gathering information about distances between the gathering locations and objects in the environment surrounding the gathering locations from any depth sensors or other distance measuring devices, such as capturing visual data from a series of multiple gathering locations within multiple rooms of a house (or other building). For example
[0038] For example For example The system can utilize 3D models of the building's interior and / or exterior, such as determining the relative positions of associated shapes of rooms or other areas by using inter-room passageway information and other information, and optionally adding distance scaling information and / or various other types of information to the generated floor plan. Additionally, in at least some embodiments, the MIGM system can perform further automated operations to determine additional information and correlate that information with specific rooms, areas, or locations within the building's floor plan and / or floor plan to analyze images and / or other environmental information captured inside the building. For example (audio) to determine specific properties ( For example The color and / or material type and / or other characteristics of a specific element, such as a floor, wall, ceiling, countertop, furniture, lighting fixture, appliance, etc.; the presence and / or absence of a specific element, such as an island in a kitchen; etc.), or other attributes that determine the relevant properties ( For example The orientation of building elements; building elements such as windows; views from a particular window or other location; etc. Additional details are included below regarding the operation of the computing device implementing the MIGM system to perform this automated operation, and in some cases, to further interact with one or more MIGM system operator users in one or more ways to provide further functionality.
[0039] In various implementations, the described techniques offer a variety of benefits, including improved autonomous operation of excavator engineering vehicles and / or other engineering vehicles. For example Fully autonomous operation) control, such as at least in part based on training one or more such engineering vehicles ( For example This technology involves developing one or more machine learning behavior models (of one or more types of construction vehicles) and using the trained machine learning behavior models to control the corresponding autonomous operation of one or more corresponding construction vehicles. Furthermore, this automation technology allows for faster and more accurate training and use than prior art, including a significant reduction in the computing power and time used. Additionally, in some embodiments, the described technology can be used to provide an improved GUI where users can obtain more accurate and faster information about the operation of excavator construction vehicles and / or other construction vehicles, including in response to search requests or other instructions, as part of providing personalized information to the user. The described technology also provides various other benefits, some of which are further described elsewhere.
[0040] For illustrative purposes, some embodiments are described below in which specific types of information are acquired, used, and / or presented in a specific manner and by using specific types of means for a particular type of structure. However, it will be understood that the described techniques may be used in other ways in other embodiments, and therefore the invention is not limited to the exemplary details provided. As a non-exclusive example, although in some embodiments specific types of circular descriptors are generated for images and room locations and these circular descriptors are compared or otherwise matched in a specific manner, it will be understood that other types of information for describing image content and room locations may be similarly generated and used in other embodiments, including for buildings (or other structures or layouts) separate from houses, and the determined image acquisition location information may be used in other ways in other embodiments. Additionally, the term "building" herein refers to any partially or completely enclosed structure that typically, but not necessarily, encompasses one or more rooms that visually or otherwise divide the interior space of the structure. Non-limiting examples of such buildings include houses, apartment buildings, or individual apartments, condominiums, office buildings, commercial buildings, or other wholesale and retail structures ( Figures 2A to 2I Shopping malls, department stores, warehouses, etc. When used herein with reference to the interior of a building, the acquisition location, or other locations (unless the context explicitly indicates otherwise), the term “acquisition” or “capture” can refer to any recording, storage, or input of media, sensor data, and / or other information relating to the spatial and / or visual and / or otherwise perceptible characteristics of the interior of a building, or a subset thereof, such as by a recording device or by another device receiving information from a recording device. As used herein, the term “panoramic image” can refer to a visual representation based on, including, or divisible into multiple discrete component images originating from substantially similar physical locations in different directions and depicting a larger field of view than any single discrete component image depicts, including (but not limited to) images from physical locations with a sufficiently wide viewing angle to include angles beyond what a person can perceive from a single direction of gaze. As used herein, the term "series" of acquisition locations generally refers to two or more acquisition locations, each of which is accessed at least once in a corresponding order, regardless of whether other non-acquisition locations have been accessed in between, and regardless of whether the access to said acquisition locations occurs during a single consecutive time period or at multiple different times, or by a single user and / or device or by multiple different users and / or devices. Additionally, various details are provided in the figures and text for illustrative purposes, but these details are not intended to limit the scope of the invention. For example, the dimensions and relative positions of elements in the figures are not necessarily drawn to scale, and some details are omitted and / or provided more prominently (…). For example(Through size and positioning) to enhance readability and / or clarity. Furthermore, the same reference numerals may be used in the accompanying drawings to identify the same or similar elements or actions.
[0041] Figure 1 This illustrates an example of automatically analyzing information about a building's floor plans and automatically analyzing visual data from images captured inside the building in order to automatically determine the image acquisition location and display it on the floor plan. Figure 2A For reference For example Further discussion of Building 198).
[0042] In particular, For example Example 2D floor plan information 230a is shown. Figure 1 (presented in vectorized format), along with elements including descriptions of floor plans of rooms or other areas ( Figures 2F to 2G Legend 269a shows semantic information about doors, windows, platforms, backyards and / or terraces, attached residential units, or other external structures, as well as optional geographical direction indicators 209. Although the example building is a multi-story residential building (with staircases for the higher floors shown near the bottom right of the floor plan), for simplicity, example 2D floor plan information 230a only shows sub-floor plans of the main floors. It will be understood that such floor plan information will typically include multiple floors or other separate sections, optionally with visual indications of how they are linked or otherwise associated, and the described type of processing can be performed on all such sections of the floor plan information. Regarding For example And in For example The text shows more details of its parts, individual rooms 260 (including the living room 260a as the leftmost room), and further discusses examples of the house 198 and its associated surrounding areas corresponding to the floor plan information 230a, which include an external platform or terrace or balcony 186, a larger external backyard or terrace 187, and external attachment structures 188. Figure 2B (Garages, sheds, attached living units, greenhouses, etc.). The automated operation of the IFPLMM system (not shown) includes analyzing 2D floor plan information 230a to generate a multi-point... P 232 related 2D point cloud 231 ( Figure 2A Based on a defined sampling rate or size, such as X points per Y distance. In this example implementation, the analysis of 2D floor plan information 230a includes generating each point... P Information for each point P The information includes 2D XY position vectors x Normal direction vector n and semantic informations , among which specific points P i Shown as having corresponding associated information x i , n i and s i .
[0043] For example continue lsf Examples are provided, and it is illustrated how the 2D point cloud 231 is supplied to the IFPLMM system by a neural network 274b that has been trained to assign features to the 2D point cloud of the floor plan. lsf (using training examples with positive and negative labels) to generate vectors with their respective features from the latent space. Figure 2C Multiple related points P 234 Enhanced 2D point cloud 233, where specific points P i Shown as having corresponding associated information Figures 2A to 2B i .
[0044] For example continue For example The example illustrates how the enhanced 2D point cloud 233 is supplied to the rendering component 273 of the IFPLMM system, which generates various circular descriptors 278 for building locations in rooms or other areas of a building, wherein, in this example, the circular descriptors for building locations are generated in a mesh layout. For example, room 260a is shown in more detail within the enhanced 2D point cloud 233, illustrating a mesh 268 for which example building locations for which circular descriptors for building locations are generated. The illustrated building locations 268 are only a subset of the building locations for which circular descriptors for building locations 278 are generated ( Figure 2C (This is part of a larger grid of building locations extending throughout the house, not shown. For the sake of brevity, only the building locations of individual rooms are described, but the total building locations will also include some or all of the other rooms in the building, and may further optionally include some or all of the illustrative area outside the main house 198.) For example Some or all of the following: an external platform, terrace, balcony, backyard, or additional external structure. For example Further information is provided regarding the circular descriptor 278c5 for the specific building location corresponding to the example grid location 268c5. It will be understood that such locations within a grid can be determined in various ways ( Figure 2Cand in other implementations, a set of room locations can have a form other than a grid (in some cases, including randomly or otherwise in an irregular manner). In at least some implementations, a building location circular descriptor will be generated for each of the room locations, such as for later use in comparing those building location circular descriptors to image circular descriptors for images in order to determine which of the building location circular descriptors best matches an image circular descriptor.
[0045] In this example, each building location circular descriptor uses a north direction to correspond to 0°, continuing in a clockwise manner through 360°. In various implementations, a building location circular descriptor for a given room location can be generated in various ways, including by using geometric techniques to determine an angular amount from a given room location and a starting direction to a given location of a point cloud point (e.g., on a wall). In some cases, the angular amount can be determined by using a line from the given room location and the starting direction to the given location of the point cloud point, and then determining an angle between the line and a line from the given room location and the starting direction to a point on the wall (e.g., a closest point on the wall). Figure 2C In this example, various point cloud points 233a-233q of room 260a are illustrated, with corresponding feature information illustrated in example building location circular descriptor 278c5. Although illustrated in a linear manner in example building location circular descriptor 278c5, it will be appreciated that in other implementations, a building location circular descriptor can instead be represented as a circle or ring (or as an arc if using less than 360°). Figure 2D In this example, various point cloud points 233a-233q of room 260a are illustrated, with corresponding feature information illustrated in example building location circular descriptor 278c5. Although illustrated in a linear manner in example building location circular descriptor 278c5, it will be appreciated that in other implementations, a building location circular descriptor can instead be represented as a circle or ring (or as an arc if using less than 360°). Figures 2A to 2C Further illustrated is information 270 regarding attributes that can be used as part of rendered feature information for a particular point cloud point, with information 272 illustrating an example of feature information that can be generated for point cloud points 234g-234i of building location circular descriptor 278c5 of building location 268c5. In this example, attributes include angular information in a set of enumerated ranges of incidence angles, distance information in a set of enumerated ranges of distances, and optionally other attribute information, in order to provide a sense of the surrounding geometry of the building location. In other implementations, attributes can be encoded in other ways (e.g., with exact values of incidence angles and / or distances rather than ranges; with attribute information other than incidence angles and / or distances, whether instead of or in addition to incidence angles and / or distances; etc.). In particular, in this example, with respect to point 234g, it has an incidence angle g g0 (corresponding to an incidence angle between 0° and 10°) and a distance value h g1 (corresponding to a distance between 0.5 meters and 1 meter), with corresponding feature f gThe vector is stored at the corresponding location of the building location circle descriptor 278c5 (labeled in this example for this point using "234g" reference numeral). Similar information is encoded and stored for each of the points that are visible from the building location 268c5, and similar processing is performed for each of the other building locations and their associated generated building location circle descriptors. In at least some points, when generating a building location circle descriptor, a first application of an occlusion test is applied to find a set of points that are visible to the building location, and the points are projected to the building location circle descriptor based on the angle at which the points are observed.
[0046] Figure 2C Continuing For example of the example, and illustrates alternative example panoramic images 255d and 250d that can be determined using the building location circle descriptors 278 generated in For example , and illustrates additional information that can be gathered for the location 265 in the room 260a from which the example panoramic images are actually captured. In this example, the panoramic image 255d represents a 180° panoramic image captured from a capture location 265 in the living room 260a of the house 198, as indicated using the information 265 and 267 on the floor plan excerpt for the house 260a, and the panoramic image 250d represents a 360° panoramic image taken from the capture location 265. Using such a panoramic image 255d or 250d, various subsets of the panoramic image can be displayed to an end user (not shown), with an example subset 251d shown as a portion of the panoramic image 255d. Alternatively, the subset 251d can instead represent a single stand-alone image captured from the capture location 265 in the living room 260a of the house 198, with the capture location determination performed for the stand-alone image instead (whether in addition to or in lieu of the capture location determination for the panoramic images 255d and / or 250d). In this example, one or more of the panoramic images 255d and 250d and the stand-alone image 251d are supplied to a neural network 274d of the IFPLMM system that is trained to generate a corresponding image circle descriptor 279 for each supplied image. Since in this example the panoramic image 255d and the stand-alone image 251d do not extend to a full 360° of horizontal extent, the corresponding image circle descriptor for either will encode information for less than 360°, as discussed elsewhere herein. For purposes of illustration, this example continues with the panoramic image 250d and its corresponding image circle descriptor 279d. In particular, the panoramic image 250d is captured in the living room of the house, and in this example includes a 360° horizontal coverage around a vertical axis, with the image displayed in an equi-rectangular format, and with the x and y axes of the visual content of the image corresponding to the horizontal and vertical information in the room For exampleAlignment includes the boundary between two walls, the boundary between a wall and the floor, the bottom and / or top of windows and doors, etc. In this example, image capture may begin, for example, with the camera orientation in a westward direction corresponding to the 0° relative starting horizontal direction of the panoramic image 250d and continue in a complete circle, where the relative 90° horizontal direction of the panoramic image then corresponds to the northward direction, the relative 180° horizontal direction of the panoramic image corresponds to the eastward direction, the relative 270° horizontal direction of the 360° panoramic image corresponds to the southward direction, and the relative 360° ending horizontal direction of the 360° panoramic image returns to the westward direction. Thus, for each degree from 0° to approximately 20° of the image circle descriptor 279d, the image circle descriptor will encode information about the features in the visual data of the image corresponding to the west wall of room 260a, at approximately the horizontal midpoint of the west-facing window at 0° ( For example The visual information transmitted through or reflected from the window begins with and includes vertical information in that horizontal direction within the visual data of the image. For example The image extends along the walls above and below the vertical slice of the window, and continues to the visible portions of the ceiling and floor in that horizontal direction, where example features 236a and 236b are shown within this range. Since the horizontal direction within the visual data of the image reaches the north end of the window and continues towards the northwest corner 195-1, the image circular descriptor 279d continues to encode information about features in the visual data of the image corresponding to the west wall of room 260a, where feature 236c corresponds to the northwest corner 195-1 shown at approximately 35°. Since the horizontal direction within the visual data of the image continues past the northwest corner 195-1 and continues towards the northeast corner 195-2, the image circular descriptor 279d continues to encode information about features in the visual data of the image corresponding to the north wall of room 260a, ending with feature 236d corresponding to the northeast corner 195-2 shown at approximately 165°. This encoding of visual data features continues for the entire 360°. As previously described, in various implementations, information regarding the determined location of identified features in a circular descriptor can be encoded and stored in various ways, including in the form of a vector or array having one or more values for each directional angle, to identify features present in a given angular direction. In other implementations, angular information other than a single degree of horizontality can be represented in the circular descriptor, such as less than a single degree or instead of multiple degrees, and / or verticality (either replacing or supplementing the degree of horizontality). Furthermore, while the panoramic image in the above example was captured with a westward starting direction, it will be understood that panoramic images can be captured in other ways in other cases. For example, other panoramic images may have different starting directions, or the panoramic image may be modified to capture its entire horizontal coverage simultaneously (…). Figure 2E(via one or more fisheye lenses), a specific orientation can be selected to treat it as 0° relative to the panoramic image. Figures 2A to 2D (Any choice; by using predefined directions, such as north; etc.).
[0047] Figure 2D continue Figure 2C Examples include one or more components 276 of the IFPLMM system, which will come from... Figure 2E Example image circular descriptor 279d and Figure 2C The building location circular descriptor 278 is taken as input, and the determined acquisition position 277 of the panoramic image 250d is generated on the floor plan 230a. Component 276 may include, for example, a rotation-based matcher component that selects the best-matching building location circular descriptor, and a neural network optionally trained to refine the acquisition positions of the image to include positions between the building locations corresponding to the building location circular descriptor 278. Additionally, Figure 2C With Figure 2E The method used is similar to further illustrate the excerpt of the floor plan of living room 260a, including showing the source For example The room location grid, but For example The grid 288 includes additional information about the degree of matching between the circular building location descriptor associated with each room location and the circular image descriptor 279d. For example (In a manner similar to a heatmap). In this example, the similarity / dissimilarity information 288 indicates that the grid room location 268c5 has the highest match with the image circular descriptor 279d ( For example The highest similarity, lowest dissimilarity, lowest distance, etc.), while the room locations in the 3rd and 4th rows of column d and the 4th row of column e have the second highest matching degree, and various other room locations generally decrease in their matching degree as their distance from room location 268c5 increases. In at least some embodiments, the comparison between the image circular descriptor and the building location circular descriptor of the room may include: at the selected room location ( Figure 1 Starting at a location such as room location 268g4 in this example, randomly selected, at the center of the room, or nearby, etc.; and repeatedly moving in the direction of adjacent room locations with higher matching scores using nearest neighbor search until the best match is identified, as described in excerpt 260f of room location 268c5, but in other implementations, other matching techniques may be used ( Figures 2A to 2EThis involves a detailed comparison with all building location circular descriptors, but without using this incremental search. After determining the room location with the best match, the corresponding location within the room can be assigned to the determined acquisition location 289 of the 360° panoramic image, such as the room location of the best-matching building location circular descriptor in the example illustrated here, or, in some implementations, within a short distance of that room location based on additional refinement performed by components of the IFPLMM system. For example The distance is calculated based on the amount and / or type of difference between the image circular descriptor and the best-matching building location circular descriptor. Once the acquisition location of the 360° panoramic image is determined (whether at the nearest building location 289 or the actual location 265), it can be compared with the floor plan. For example It may be stored and / or used in one or more ways in other respects, as discussed in more detail elsewhere in this document.
[0048] Although For example Not illustrated in the examples, but in some implementations and situations, acquisition location determination can be performed for images that may have been captured in any of a plurality of candidate areas of one or more buildings. In various implementations, this acquisition location determination activity can be performed in various ways to consider each possible area and / or building and find the best-matching building location across all these areas and / or buildings, narrowing the set of possible candidate areas before performing matching. For example This involves attempting to identify one or more features and / or visual elements present only in one or a subset of possible candidate regions from an image, using GPS or other location information associated with image acquisition, etc. In this implementation, the grid of circular descriptors for the building's location can extend across some or all of the building's rooms and / or other areas, and the best-match search for the circular descriptors of the image with the circular descriptors of the building's location can extend across multiple rooms or other areas of the circular descriptors of the building's location. Figure 2E This can include all building location circular descriptors considered for generating a building.
[0049] Regarding finding the best-matching circular building location descriptor for image circular descriptor 279d from multiple possible building locations within a room, in various embodiments, some or all of the circular building location descriptors for the room locations in the grid can be compared with image circular descriptor 279d in various ways to determine the degree of match. For example, in some embodiments, a distance metric for a circular bulldozer can be used to compare two such descriptors in a manner consistent with independent rotations, such that the two descriptors may have relative 0° starting directions pointing in different directions. In other embodiments, other measures of distance or similarity / dissimilarity can be used, such as by measuring the distance for each angle and aggregating this information across all angles.
[0050] Additionally, to facilitate comparison of two such circular descriptors in cases where distance or similarity / dissimilarity metrics are not independent rotations, in some implementations, additional automated operations can be performed to ensure that the information encoded in a given relative angular direction in one circular descriptor is being compared with the relative angular direction pointing to the same real-world direction in the other circular descriptor. For example, in some implementations, a brute-force method might be used, which compares each angular direction in one circular descriptor with a specific angular direction in the other circular descriptor (…). For example The two circular descriptors to be compared are compared (starting direction), thus ensuring that one of the comparisons uses the same direction. Alternatively, in other embodiments, automated operations can be performed to synchronize the two circular descriptors to be compared, such as by identifying which relative angular directions in one circular descriptor correspond to which relative angular directions in the other circular descriptor (starting direction). That is (This involves identifying the corresponding angular direction in another circular descriptor based on the relative 0° starting angle direction of one circular descriptor). Regarding... For example For example, this determination can identify that the 0° starting direction (corresponding to a westward direction) of the image circular descriptor 279d is the same as the 270° direction (or -90° direction) in each building location circular descriptor 278. Alternatively, in some embodiments, a finite number of instances of environmental characteristics can be identified, wherein the angular direction of each such instance of a circular descriptor is compared with a corresponding instance in another circular descriptor; such instances of characteristics may be directions orthogonal to or perpendicular to the wall plane. For example (identified by vanishing point analysis using lines in the image) so that there are four such instances for a 360° panoramic image in a typical rectangular room. Figure 2Fone or a limited number of instances of a wall element type in the environment, and can compare the angular direction in one circular descriptor for each such instance to the angular direction in another circular descriptor for an instance of the same wall element type, examples of such features can include doors Figures 2A to 2E , door start or end edges, inter-wall boundaries Figure 2G , where there are typically four such instances in a rectangular room, etc. In other implementations, other distance metrics and / or similarity / dissimilarity metrics can be used, and other techniques can be used to synchronize corresponding information in two or more circular descriptors being compared.
[0051] Figure 2D Continuing For example the example, and illustrating one example of a 2D floor plan 230f for the house 198, such as can be presented to an end user in a GUI 255f, where the living room 260a is the most westward room of the house. It will be appreciated that in some implementations, 3D or 2.5D floor plans showing wall height information can similarly be generated and displayed, either in addition to or in place of such 2D floor plans, where appropriate. For example One example is discussed. In this example, information has been added to the floor plan 255f to indicate the location of the determined capture location 265 of the 360° panoramic image 279d. In other implementations and scenarios, the location and / or orientation information for an image can be displayed in other ways, such as for the example stereoscopic image of the living room south side, including a visual indicator of the direction covered in the stereoscopic image, and / or for the additional panoramic image of the living room northwest corner, showing the capture location of the panoramic image and example starting orientation / direction information. When displayed as part of a GUI, the added information for the 360° panoramic image 279d on the displayed floor plan can be a user-selectable control (or associated with such a control) that allows an end user to select and display some or all of the associated 360° panoramic images (e.g., in a manner similar to For example ).
[0052] In this example, various other types of information are also illustrated on the 2D floor plan 255f. For example, such other types of information can include one or more of the following: room labels added to some or all of the rooms For example That is"Living room" for the living room; room dimensions added for some or all rooms; visual indications of furniture, appliances, or other built-in features added for some or all rooms; visual indications of the location of additional types of associations and links added for some or all rooms. Figure 2I The end user can choose to further display additional panoramic and / or stereoscopic images; the end user can choose to further present audio annotations and / or sound recordings, etc.; visual indicators for doors and windows added for some or all rooms; etc. Additionally, in this example, a user-selectable control 228 is added to indicate the current floor displayed on the floor plan and allows the end user to select a different floor to display. In some implementations, changes to floors or other floors can also be made directly from the floor plan, such as via selecting the corresponding connecting channel in the illustrated floor plan (…). For example (The staircase leading to floor 2). It will be understood that in some embodiments various other types of information may be added, in some embodiments some of the information of the described types may not be provided, and in other embodiments visual indicators of links and related information and user selections thereof may be displayed and selected in other ways.
[0053] For example continue For example Examples are provided, and an example of model 265g showing a floor plan of house 198 including height information is also provided. For example This model 265g, as part of a 2.5D or 3D model floor plan of the house, can be presented to the end user in a GUI. Such a model 265g can, for example, be additional mapping-related information generated by the MIGM system based on floor plans 230a and / or 230f, which shows additional information about height to illustrate the visual location of features such as windows and doors within the walls. In this example, information has been added to model 265g to represent the location of the determined acquisition position 265 of the 360° panoramic image 279d. Although... For example Not described herein, but in some implementations, additional information may be added to the displayed wall, such as images taken during video capture. For example (Rendering and illustrating actual paintings, wallpaper, or other surfaces from the house on the rendered model 265), and / or can be used in other ways to add specified colors, textures, or other visual information to walls or other surfaces.
[0054] For example continue θ Examples are provided, and information 290h is illustrated, which shows example information processing flows during automated operation of the IFPLMM system in at least some embodiments. Specifically, in SEIn the example of FIG. 1, the implementation of the IFPLMM system 140 is executing on one or more computing devices 180, and performs automated operations 281-285 to determine image capture locations for a building, and performs operation 287 to display and / or provide corresponding information, and optionally performs operation 289 to further use the determined image capture locations to improve subsequent automated operations of the IFPLMM system. In particular, the IFPLMM system receives 281 a rasterized building floor plan (e.g., from a database or other storage 294), and proceeds to generate 282 building location circle descriptors 288a for the building. Additionally, the IFPLMM system receives 283 an image captured at a location of the building, and proceeds to generate 284 an image circle descriptor 288b for the image. In step 285, the IFPLMM system performs a rotation-based matching of the circle descriptor 288b to the circle descriptors 288a, optionally performs a location refinement, to determine a capture location 286 of the image that is associated with the building, and proceeds in step 287 to display or otherwise provide (e.g., store, such as on one or more remote storage systems 180 over the network 170) the determined capture location. Additional details regarding such operations 281-285 are discussed in relation to FIGS. 2-7, and elsewhere herein. θ θ For example Additional details regarding such operations 289 are provided. Although only a single building and panoramic image example process is shown in relation to FIG. 1, it will be appreciated that similar processes can be performed for multiple buildings and / or multiple images (whether panoramic images or stereoscopic images) in order to determine capture location information for each of the multiple images relative to one or more candidate buildings in which the images were captured. SE θ θ
[0055] For example Continuing with the example of FIG. 1, the implementation of the IFPLMM system 140 is executing on one or more computing devices 180, and performs automated operations 281-285 to determine image capture locations for a building, and performs operation 287 to display and / or provide corresponding information, and optionally performs operation 289 to further use the determined image capture locations to improve subsequent automated operations of the IFPLMM system. In particular, the IFPLMM system receives 281 a rasterized building floor plan (e.g., from a database or other storage 294), and proceeds to generate 282 building location circle descriptors 288a for the building. Additionally, the IFPLMM system receives 283 an image captured at a location of the building, and proceeds to generate 284 an image circle descriptor 288b for the image. In step 285, the IFPLMM system performs a rotation-based matching of the circle descriptor 288b to the circle descriptors 288a, optionally performs a location refinement, to determine a capture location 286 of the image that is associated with the building, and proceeds in step 287 to display or otherwise provide (e.g., store, such as on one or more remote storage systems 180 over the network 170) the determined capture location. Additional details regarding such operations 281-285 are discussed in relation to FIGS. 2-7, and elsewhere herein. For example Examples are provided, and information 230i is illustrated, which includes a floor plan of the building and further illustrates information regarding example figure 245 having nodes 245a to 245m for each of some or all of the building's rooms or other areas. Specifically, each node may represent a room or other building area and optionally have inter-node edges corresponding to inter-room connections and / or inter-room adjacencies (whether or not connections exist), wherein the edges in this example correspond to inter-room connections (…). For example This allows node 245b for living room 260a to be connected to node 245a for corridor, but not to node 245f for the adjacent room on the north side of the living room. Once the acquisition location 265 of panoramic image 250d is determined, the panoramic image and / or its image circle descriptor can be associated with node 245b for the room or other area containing that acquisition location, where the visual data of the panoramic image ( For example Specific aspects of visual data, such as color data or other pixel values, texture map data, etc.; latent spatial features generated from visual data; etc.) are then used to update some or all of the circular descriptors of the building location in the living room. For example This is to supplement existing potential spatial features identified from floor map data, or to replace some or all of the existing potential spatial features identified from floor map data, as discussed in more detail elsewhere herein. In other embodiments, this determined acquisition location of the image can be used in conjunction with circular descriptors of building locations in ways other than via such a map, or such a map can be constructed in other ways ( γ It has multiple nodes for some or all of the rooms or other areas; has nodes grouped hierarchically or otherwise, such as by floor or other grouping; etc. In other implementations, some or all of the circular descriptor of building location may be updated in other ways (either supplementing visual data with images of the determined acquisition location, or replacing such visual data), one non-exclusive example involving adding information to such a circular descriptor of building location to include explicit indications of the building's wall elements and / or other structural elements (e.g., windows, doorways and non-doorway openings, walls, boundaries between walls and / or other boundaries, etc.).
[0056] Furthermore, although at least some rooms of a house are represented by associated nodes in an adjacency graph, in at least some implementations, some spaces within the house may not be designated as rooms for the purposes of the adjacency graph. ψIn the adjacency graph, spaces may not have individual nodes, such as cloakrooms, small areas (such as pantry rooms or cabinets), connecting areas (such as stairs and / or corridors), etc. In this example embodiment, the stairs have a corresponding node 245h and the walk-in cloakroom may optionally have a node 245l, while the pantry room has no node. However, in other embodiments, these spaces may not have nodes, or all or any combination of these spaces may have nodes. Additionally, in this example embodiment, areas adjacent to building entrances / exits outside the building also have nodes representing them, such as node 245j corresponding to the front yard (accessible from the building via the entrance door) and node 245i corresponding to the platform (accessible via the back door), as well as node 245l for attaching external structures and node 245m for larger backyards and / or terraces. In other embodiments, such external areas may not be represented as nodes (and in some embodiments, may instead be represented as attributes associated with adjacent outdoor doors or other openings and / or rooms). Similarly, in this example embodiment, information about areas visible from windows or other building locations can also be represented by nodes, such as optional node 245k corresponding to a view visible from a west-facing window in the living room. However, in other embodiments, such views may not be represented as nodes (and in some embodiments, may instead be represented as attributes associated with the corresponding window or other building location and / or the room with which they are located). It will be noted that although in ω Some edges are shown as passing through walls (such as the edge between node 245a for the corridor and node 245f for the master bedroom in the north center of the house), but the actual connection between the rooms corresponding to the nodes connected to this edge is based on a door or other non-door opening connection. atan (Based on the interior door near the northeast end of the corridor between the corridor and the master bedroom). Additionally, although not described in information 230i, the adjacency diagram of the house may further continue in other areas of the house (such as the second floor) not shown. In some embodiments, each node may further have associated attributes, such as those relating to the room or other area represented by the node, wherein non-exclusive examples include one or more of the following: room type; room size; location of windows and doors in the room and other openings between rooms; information about the shape of the room (whether 2D or 3D); view type of each window and optionally, orientation information of each window; optional orientation information of doors and other openings between rooms; information about other characteristics of the room, such as information from analysis of associated images and / or provided by end users viewing the floor plan and optionally its associated images. For example For examplesuch as carpet or other flooring material and color and / or texture and / or type of wall coverings and ceilings, etc.; light fixtures or other built-in elements; furniture or other items within the room; etc.; information about images taken in the room and / or copies thereof (optionally with associated location information within the room for each of the images); information about audio or other data captured in the room and / or copies thereof (optionally with associated location information within the room for each of the audio clips or other data segments); etc. Similarly, in some implementations, each edge can further have associated attributes such as relating to the inter-room connection represented by the edge, with non-exclusive examples including one or more of: inter-room connection type (e.g. δ , doorway, non-doorway opening, etc.); inter-room connection size (e.g. δθ , width; height and / or depth), etc.
[0057] Additionally, in at least some implementations, further automated operations can be performed as part of the automated determination of the capture location for images captured in a room or other building area, such as if corresponding information about the rooms or other areas of the building is available. For example, in at least some implementations, a geometric localization technique can be used to test the association between wall elements visible in an image and wall elements present in a room to confirm the degree of match of a building location circular descriptor that has been determined to be the best match for an image circular descriptor and / or as part of identifying such a best match building location circular descriptor. The geometric localization technique can include, for example, using a 2-point solver and / or a 3-point solver to determine one or more possible room shapes for the room and / or the location of elements within the room (or optionally receiving such information as a starting point), and then localizing the wall elements on the possible room shapes. In other implementations, the wall element locations can be determined in other ways, such as via use of depth sensing devices or other room mapping sensors in the room, via machine learning methods for analyzing images to identify room shapes and wall element locations, via input specified by one or more human operators, etc. Further, in some implementations, given a room location and information about the room shape and locations of wall elements, a new composite image can be generated as a projection / visualization of a view of the room from that room location with the wall elements shown in their locations, and the visual information of this composite image can be directly compared to actual images from the room to determine a degree of similarity / dissimilarity or other degree of match between the two images, with the inter-image comparison used to determine whether that room location is a match for the capture location of the actual images. In a similar manner, in some implementations, some or all of the building location circular descriptors for room locations in a room can be used as a basis for determining a degree of similarity / dissimilarity between images taken at those room locations (e.g. δImage circular descriptors (360° panoramic images) are generated, and then those room / image circular descriptors can be combined with new images taken in the room ( δθ The circular descriptors of images (with a horizontal coverage of less than 360°) are compared to determine the best-matching circular descriptor for building locations in a manner similar to that discussed above.
[0058] In some implementations, the automated determination of the acquisition location of images taken in a room may also include additional operations. For example, in at least some implementations, machine learning techniques may be used to learn the optimal encoding to allow the image to match the room location, such as candidate encodings from a variety of definitions, or by considering a variety of possible image elements and identifying the subset of image features that provide the best match for the corresponding room location. Additional details are included below regarding the various automated operations that may be performed by the IFPLMM system in at least some implementations.
[0059] As a non-exclusive example implementation, determining the acquisition location of a target image within a room (also referred to as "location" for the purposes of this example implementation) may include relative to a reference structural layout (or "map"). M To estimate the target image I Related 2D camera poses Camera pose in a 2D plane p = [ t , θ ] θ (2) Modeling is performed to achieve rotation on the deflection axis. For example [0, 2] π ) and peace t = [ x , y ], where the posture parameters t R 2 and For example [0, 2] π Define the camera's planar displacement vector and deflection axis rotation, respectively. The target image can be a panoramic image ( For example Maps (in a rectangular format) or stereoscopic images with a known field of view (FoV). M It can have various formats and encodes structural layout information in the 2D plane (also referred to as "occupancy" for the purposes of this example implementation).
[0060] In this example implementation, determining the acquisition location of the target image includes using a Monte Carlo localization (MCL) architecture, which defines a measurement model.P I p M ), which is expressed on the map M in terms of the likelihood of observing an image p from a camera pose I , where for simplicity, non-stochastic parameters are excluded M below. The posterior distribution I after observing p is the solution of interest, and the MCL architecture estimates the posterior distribution P p I , as follows:
[0061]
[0062] where P I is a normalizing constant that can be ignored, and P p is a prior camera pose distribution on the map that is assumed to be uniformly distributed within the map region. Finally, the full posterior can be approximated by extracting particles from P p , the likelihood of which will be estimated using the measurement model as defined in equation (5) below. P p
[0063] According to the metric learning architecture, instead of using the flat descriptors used in metric learning, the spatial visibility is encoded using a circular feature, resulting in a geometrically interpretable metric learning. The circular feature is defined as an ordered set of feature vectors, as follows:
[0064] F = { f α | α = 1 … V-1} (2)
[0065] where V is the number of feature segments. Each feature segment f α R D For 2D planes, the 2D πα / V ) to ( (2 π (α +1 ) ) / V ) range π / V local orientation FoV of rad. In this example implementation, the ordered set F is referred to as a circular feature because the first and last feature segments correspond to adjacent FoVs. The ordered feature segments correspond to 360° 2D spatial information, and V is the number of segments. In this design, the full 360° 2D spatial information is implicitly encoded in the order of the feature segments.
[0066] Two circular features F i = { f i α | α = 0 … V-1} and F j = { f j α | α = 0 … V-1} are defined as follows:
[0067]
[0068] where cos(.,.) computes the vector cosine similarity and normalizes the function output to [0, 1]. Define the rotation operator R(F, θ) that rotates the underlying spatial information of a circular feature F with a given angle θ
[0069]
[0070] where the 1D feature space is linearly interpolated when the index produces non-integer values. Finally, the measure model is defined as
[0071]
[0072] where A is the PDF normalization constant and F I and F t are the circular features encoded from the target image and rendered on the map at locations t , respectively.
[0073] To systematically reduce the rotational dimension from the MCL sampling step (because in For example (2), the MCL uses a large number of samples to approximate the camera pose posterior), and for sample locations with typical orientations t , its circular feature F t can be found by the following: I
[0074]
[0075] Substituting into equation 5 provides the solution to eliminate... Figures 2A to 2I And only by translation t Simplified measurement model for the given conditions:
[0076]
[0077] To solve equation 6, we need to use [0, 2] π Uniform sampling in) Figure 1 t The F t The rotation is performed, and the best one is retained. This discretization search initializes the rotation with coarse values, which will be refined later, as discussed below. The rotation matching process is highly efficient because it reuses the same circular features and does not introduce new assumptions.
[0078] Render circular features from a 2D floor plan with a given camera pose as follows. This is based on a general 2D map representation that encodes area occupancy information. M ( Figure 1 Floor plan, occupancy grid, etc., will be at the occupancy boundary ( For example Uniform sampling of points on the wall to extract 2D point clouds M = { m i | i = 0 … N-1}. Each point m i = [t i n i s i ] its position t i Normal vector n i and optional semantic information s i Encoding is performed. The normal vector is normalized to point within the room. When semantic information is available ( Figure 2A When dealing with labels and / or locations of doors, windows, etc., the semantic information is encoded as a binary mask appended to the point representation.
[0079] To avoid inefficient second-order rendering and encoding processes, latent space rendering is used to directly render circular feature vectors by aggregating features from visible map points. However, the visibility of the static environment is locally constant at most sampling locations, thus providing limited spatial context. Therefore, detailed rendering dynamics (such as the length and angle of incidence of the view line between features and sampling locations) are analyzed to mitigate the potential homogenization of the representation, and an adaptive rendering mechanism is defined. Therefore, a rendering codebook corresponding to the overdetermined latent feature space is used to assign view-dependent adaptive features to map points, thereby encoding the rendering dynamics. Figures 2D to 2IDuring rendering, features of map points are used, such as those determined by the codebook and rendering dynamics, where features from multiple codebooks are blended by addition. This process handles 2D point cloud maps. Figures 4A to 4B (via PointNet) to each map point m i Two sets G of potential features to be assigned i = { g i β | β = 0 … G-1} and H i = {h i γ | Figure 1 = 0 … H- 1}, which are represented as the distance and angle of incidence codebooks, respectively. The features in the codebook have the same characteristics as the circular feature segment g. β h γ R D Same size. During rendering, map point features are selected from the codebook based on the distance and angle of incidence relative to the rendering location. Assume the rendering location tˆ and map point m... i = [t i n i s i ], let d i = t i - tˆ, its rendering dynamics are calculated as
[0080]
[0081] in d i and For example i These are the distance and the angle of incidence, respectively. Clockwise, the angle of incidence is [0, 2]. π Distinguish between the four quadrants. (Using m) i The associated codebook G i and H i Then, its characteristic f is determined by the following formula. i
[0082]
[0083] in d max This is the predefined maximum distance from the codebook. Similar to Equation 4, for non-integer indices, linear interpolation between their two nearest codes in the codebook can be used. Finally, if m i If the location tˆ is reached through the visibility test, then f will be determined by the following formula. iProjected onto circular features F tˆ = { f i α | α = 0 … V-1}
[0084]
[0085] in For example i = Figure 1 2(d) i ( ) is the viewing angle. Finally, the features of the projected map points are averaged across each segment.
[0086] Circular features are extracted from panoramic and / or stereo images as follows. For panoramic images in an equal rectangular projection, each image column corresponds to a fixed amount of horizontal FoV. This capture configuration facilitates a bidirectional mapping between adjacent input image column groups and segments in the rendered circular representation. The query panorama is fed into the encoder ( Figure 2F The feature map is obtained from the ResNet50 encoder, which calculates the number of feature segments in each circular feature based on a pre-configured number of segments. This feature map is then compressed by average pooling in the vertical dimension to fit the feature size of the circular segments and then average pooled again in the horizontal direction. V Each element.
[0087] Stereo images are the most common type of images taken from cameras of pinhole camera models, possessing additional camera rotational degrees of freedom compared to panoramic images. If we assume the images have a known FoV and zero pitch / roll angle relative to the ground plane, each image column in a stereo image will correspond to a non-fixed but known horizontal FoV. It should be noted that for images of indoor targets, the pitch / roll angle can often be adjusted by estimating the room layout. The target image can be processed into a feature map (…). Figure 2G Using a ResNet50 encoder, and average pooling can be used to compress the vertical dimension, where a stereo-to-equirectangular transformation is performed on the feature map to obtain the final circular features. Since stereo images have a FoV significantly smaller than 360°, their circular features will have segments without assigned values, which will be masked in the computation. Equation 3 can also be renormalized to have a range [0, 1].
[0088] The accuracy of Monte Carlo localization architectures can depend on sampling density, but this can be inefficient for achieving high accuracy using sampling. A lightweight, continuous refinement branch is used to address the discretized pose sampling nature of MCL processing, thus improving the current estimate. This is achieved through the current best estimate. and , the refinement branch takes as input two circular features F I and R (F t , ) as input. The refinement network uses two ID convolutional layers with circular padding followed by a fully connected layer to predict two offsets For example t, Figure 1 for translation and rotation respectively + For example t, + Figure 1 . The updated map circular features F t are then rendered and their similarity computed to F I using equation 3. If the similarity score improves the original camera pose, it is accepted and the iteration continues (e.g., where convergence typically occurs within 3 iterations, and where in at least some cases the first refinement is accepted to quantize the estimate), otherwise the refinement is considered to have converged.
[0089] R The triplet loss can be used as supervision to learn the metric space. To form a triplet, the image circular feature F I is used as the anchor, the map circular feature F + = For example (F tgt , R gt ) at the ground truth camera pose is used as the positive feature, and the negative map circular feature F - = Figure 1 (F trnd , rnd ) is used under the outlier camera pose. The triplet is then defined by
[0090]
[0091] The similarity function S and the triplet loss L triplet uses an aggregated element-wise comparison and thus effectively discards any intra-feature context. An additional context loss can be used to provide the feature segments with the global extent of their circular features for learning contextual information (F Figure 3 , properties of the room / map). The circular feature context F is defined as the mean of its normalized feature segments
[0092]
[0093] which is applied to the training triplets in a similar fashion to equation 12.
[0094]
[0095] By the context loss, the circular features achieve a better coarse level representation, improving the recall of target images with limited FoV. This loss also acts as a regularization matrix by reducing the segments with large variance, resulting in smoother posterior estimates.
[0096] To train the refinement branch, the circular features within 0.5m radius and 30 degrees angle from the ground truth camera pose are sampled, and a regression loss is used to supervise the refinement branch as follows
[0097]
[0098] For the triplet and context loss, 100 negative samples are sampled and a single ground truth sample is broadcast in each training iteration. For the refinement loss, 20 hard negatives close to the ground truth camera pose are sampled, where the perturbation sampled from a uniform distribution is limited within 30 degrees and 0.5m radius. The mean of all losses is combined with equal weights, and the hyperparameters are set to G = H = 32, V = 16, D = 128 and d max = 10m. The map is sampled into a 2D point cloud at 10cm intervals at the occupancy boundary. 0 . 1 m × 0 . 1 m The circular features are rendered on a uniform grid. To estimate the relative rotation, as described in Equation 6, 16 uniformly sampled angles are evaluated and the best one is retained. Finally, the posterior distribution can be estimated using Equations 1 and 7. To extract the final estimate from the posterior grid map, 3x3 non-maximum suppression is applied to extract the maximum value. For the maximum values with scores greater than a threshold ( Figure 1 , 0.8), they are sent to the refinement branch to get the final estimate, where their probabilities are estimated as the uncertainty. Sorted by their probabilities, the top k estimates can be obtained. Although in some scenarios the example implementation discussed above can use a binary occlusion state with respect to the viewing line from the map location, other implementations and / or scenarios can further support partial occlusion, such as by modeling the map object (e.g., a wall) as a 3D mesh and computing the line of sight from the map location to the target object. For exampleModeling (furniture). Additionally, while the example implementation discussed above may not use information about whether a door is open in some cases, other implementations and / or situations may utilize this information. Furthermore, while the example implementation discussed above may use a 2D floor plan in some cases, other implementations and / or situations may extend to 3D space and / or complex rendering time dynamics of the model, such as non-Lambertian reflections (furniture). For example (reflector).
[0099] Already about Figure 1 Various details are provided regarding the non-exclusive example implementations discussed above, but it will be understood that the details provided are non-exclusive examples included for illustrative purposes, and other implementations may be performed in other ways without some or all of these details.
[0100] Figure 1 These are example block diagrams of various computing devices and systems that can participate in the described techniques in some implementations. In particular, For example This describes an IFPLMM (Image Location Mapping Manager) system 140, which executes on one or more server computing systems 180 to determine images 145 acquired in one or more building rooms or other building areas. For example For example The acquisition location of the panoramic image (such as the acquisition location specified with respect to the corresponding building floor plan 155). In at least some embodiments and situations, one or more users of the IFPLMM system client computing device 105 can further interact with the IFPLMM system 140 via network 170 to assist in some automated operations of the IFPLMM system and / or to subsequently begin using the determined image acquisition location information in one or more further automated ways. Additional details relating to the automated operation of the IFPLMM system are included elsewhere herein, including information regarding... For example , For example and Figure 1 Additional details.
[0101] in addition, For example Further illustration is given of the optional Internal Capture and Analysis (“ICA”) and / or MIGM (Map Information Generation Manager) system 160, which in this example executes on one or more server computing systems 180 (whether the same server computing system executing the IFPLMM system 140 or a different server computing system) to capture images relative to one or more buildings or other structures respectively. For example, one or more 360° panoramic images 165 optionally linked with relative position information between the panoramic images and / or generate and provide building floor plans 154 and / or other mapping related information (e.g., to one or more users of one or more client computing devices 175 via one or more computer networks 170). Figure 1 , based on the use of panoramic images 165 and optionally associated metadata regarding their capture and linking). For example One example of such a panoramic image of a particular house 198 is shown, as further discussed below, and For example and For example An example of an enhanced building floor plan is shown, with additional details included elsewhere herein related to the automated operation of the ICA and MIGM systems. In some embodiments, the ICA system 160 and / or the MIGM system 160 and / or the IFPLMM system 140 can be executed on the same server computing system, and the IFPLMM system can optionally obtain some or all of the images 145 and / or floor plans 155 from the ICA and MIGM systems, respectively, such as if multiple or all of those systems are operated by a single entity or otherwise executed in coordination with one another (e.g., if some or all of the functionality of those systems are integrated together into a larger system). For example In other embodiments, the IFPLMM system can instead obtain floor plan information and / or images from one or more other external sources without involving any such ICA or MIGM systems, and store them locally with the IFPLMM system for further analysis and use.
[0102] One or more users (not shown) of one or more client computing devices 175 can further interact with the IFPLMM system 140 and optionally the ICA system 160 and / or the MIGM system 160 via one or more computer networks 170 in order to obtain and use the determined capture location information for images (e.g., to obtain and view one or more such images and / or a generated floor plan on which one or more images have been located and optionally interact with it in order to optionally perform one or more of the following: change between a floor plan view and a view of a particular image at the capture location within or near the floor plan; change the horizontal and / or vertical viewing direction from which a corresponding view of a panoramic image is displayed in order to determine a portion of the panoramic image that the current user viewing direction is pointing to, etc.). Additionally, while For example not illustrated in FIG. 1, a floor plan (or a portion thereof) can be linked to or otherwise associated with one or more other types of information, including floor plans of multi-story or other multi-level buildings having interconnections (e.g., hallways, stairways, elevators, etc.) between the floors. For example For exampleMultiple associated sub-floor plans of different floors or levels (via connecting stairwells), two-dimensional (“2D”) floor plans of the building linked to three-dimensional (“3D”) renderings of the building, or otherwise associated with three-dimensional (“3D”) renderings of the building. Additionally, although in Figure 1 Not described herein, but in some embodiments, the client computing device 175 (or other device, not shown) may additionally receive and use the determined image acquisition location information (optionally combined with the generated floor plan and / or other generated mapping-related information) so that it can be transmitted through these devices. Figure 3 (Autonomous vehicles or other devices) control or assist automated navigation activities, whether to replace or supplement the display of generated information.
[0103] exist For example In the depicted computing environment, network 170 may be one or more publicly accessible linked networks, such as the Internet, that may be operated by various different parties. In other embodiments, network 170 may have other forms. For example, network 170 may instead be a private network, such as a corporate or university network that is completely or partially inaccessible to non-privileged users. In other embodiments, network 170 may include both private and public networks, wherein one or more private networks are accessible from one or more public networks and / or from one or more public networks. Furthermore, in various cases, network 170 may include various types of wired and / or wireless networks. Additionally, client computing device 175 and server computing system 180 may include various hardware components and stored information, as described below regarding For example Let's discuss this in more detail.
[0104] exist For example In the example, if an ICA system 160 exists, it can perform automated operations involving multiple associated acquisition locations ( For example Multiple panoramic images are generated in multiple rooms or other locations within a building or other structure, and optionally around some or all of the exterior of the building or other structure. For example Each of these is a 360° panorama around a vertical axis, such as those used to generate and provide a representation of the interior of a building or other structure. The technology may also include: analyzing information to determine the relative position / orientation between each of two or more acquisition locations; creating inter-panorama position / orientation links in the panorama to each of one or more other panoramas based on this determined position / orientation; and then providing information to display or otherwise present the multiple linked panoramic images of the various acquisition locations within the building.
[0105] Figure 1A block diagram depicts an exemplary building interior environment in which panoramic images have been generated and optionally linked, and are ready to generate and provide corresponding building floor plans and to present the panoramic images to a user. Specifically, For example Including building 198, which has an interior at least partially captured via multiple panoramic images, such as by a user (not shown) carrying a mobile device 185 with image acquisition capabilities traversing the interior of the building to reach a series of multiple acquisition locations 210. Implementation of the ICA system ( For example The ICA system 160 on the server computing system 180; some or all copies of the ICA system executing on the user's mobile device, such as the ICA application system 155 executing in the memory 152 of the device 185; etc.) can automatically execute or assist in capturing data representing the interior of a building, and optionally further analyze the captured data to generate linked panoramic images, thereby providing a visual representation of the interior of the building. Although the user's mobile device may include various hardware components, such as a camera or other imaging system 135, one or more sensors 148 ( For example The device may include a gyroscope 148a, accelerometer 148b, compass 148c, and other components such as one or more IMUs (or inertial measurement units) of the mobile device; an altimeter; a light detector, etc.); a GPS receiver; one or more hardware processors 132; a memory 152; a display 142; a microphone, etc. However, in at least some embodiments, the mobile device may not have access to or use equipment to measure the depth of objects in a building relative to the mobile device, making it possible to determine the relationship between different panoramic images and their acquisition locations, either partially or entirely, based on matching elements in different images and / or by using information from other hardware components listed (but without using any data from any such depth sensors). In other embodiments, the mobile device may have one or more sensors for measuring the depth of surrounding walls and other surrounding objects. Additionally, while a direction indicator 109 is provided for the viewer's reference, in at least some embodiments, the mobile device and / or the ICA system may not use this absolute direction information in order to determine the relative directions and distances between panoramic images 210 without considering actual geographic location or orientation.
[0106] During operation, the user associated with the mobile device arrives at the first acquisition position 210A in the first room inside the building (in this example, the entrance from the outer door 190-1 to the living room), and as the mobile device rotates about the vertical axis at the first acquisition position ( For example(While the user rotates his or her body, the mobile device remains fixed relative to the user's body) captures a view of a portion of the building's interior visible from the acquisition location 210A. For example The first room may contain some or all of it, and optionally small portions of one or more other adjacent or nearby rooms, such as through a doorway, hallway, staircase, or other connecting passageway originating from the first room. Actions of the user and / or the mobile device may be controlled or facilitated by using one or more programs executed on the mobile device, such as ICA application system 154, optional browser 162, control system 147, etc., and view capture may be performed by recording video and / or taking a series of one or more images, including capturing images depicting scenes that can be captured from the acquisition location. Figure 3 Several objects or other elements visible in a video frame For example Visual information (structural details). For example In the examples, such objects or other elements include various elements (or "wall elements") that structurally serve as part of a wall, such as doorways 190 and 197 and their doors ( For example (with revolving doors and / or sliding doors), windows 196, wall boundaries ( For example (Corner or edge) 195 (including corner 195-1 at the northwest corner of building 198 and corner 195-2 at the northeast corner of the first room). Additionally, For example Such objects or other elements in the examples may also include other elements within the room, such as furniture 191 to 193 ( For example The panoramic image includes, for example, sofas 191; chairs 192; tables 193; etc.), pictures or paintings hanging on the wall, or televisions or other objects 194 (such as 194-1 and 194-2), lamps, etc. Such panoramic images can be provided as input to an IFPLMM system for automatically determining the acquisition location of the panoramic image relative to a floor plan of the house 198 (assuming that the IFPLMM system can obtain such a floor plan).
[0107] Regarding such captured panoramic images, in some embodiments, the user may optionally provide textual or auditory identifiers associated with the panoramic image and / or its acquisition location, such as “entrance” for acquisition location 210A or “living room” for acquisition location 210B, while in other embodiments, the ICA system may automatically generate such identifiers. For example This can be achieved by automatically analyzing information from building video and / or other recorded data to perform corresponding automated determinations, such as through the use of machine learning, or without using identifiers. After the first acquisition location 210A has been appropriately captured ( For exampleAfter a complete rotation of the mobile device, the user can optionally move to the next acquisition position (e.g., acquisition position 210B), thereby optionally recording movement data, such as video and / or data from hardware components, during the movement between acquisition positions. For example Other data (from one or more IMUs, from cameras, etc.). At the next acquisition location, the user can similarly use a mobile device to capture one or more images from that acquisition location. This process can be repeated from some or all of the rooms of the building and optionally outside the building, as illustrated for acquisition locations 210C to 210L, in this example including capturing one or more panoramic images on an outer platform or terrace or balcony 186, capturing one or more panoramic images on a larger outer yard or terrace 187, and capturing one or more panoramic images on an external attached structure 188 ( For example One or more panoramic images are captured near or in the vicinity of (e.g., garages, sheds, attached living units, greenhouses, etc.). Further analysis of the video and / or other images captured for each acquisition location is performed to generate a panoramic image for each of acquisition locations 210A to 210L, including, in some embodiments, matching objects and other elements in different images. In addition to generating such panoramic images, further analysis may be performed to “link” at least some of the panoramic images together (for illustration, some corresponding lines 215 are shown between them) to determine the relative positional information between pairs of acquisition locations that are visible to each other, and to store the corresponding inter-panel links (…). Figure 4A to 4B These are links 215-AB between acquisition locations A and B, 215-BC between B and C, and 215-AC between A and C, respectively. In some implementations and situations, at least some acquisition locations that are not visible to each other are further linked. Figure 1 (Link 215-BE, not shown, between acquisition locations 210B and 210E).
[0108] about Figure 3 Various details are provided, but it will be understood that the details provided are non-exclusive examples included for illustrative purposes, and other implementations may be carried out in other ways without some or all of these details.
[0109] Figures 2A to 2Iis a block diagram illustrating an embodiment of one or more server computing systems 380 that execute embodiments of the IFPLMM system 389, and optionally one or more server computing systems 300 that execute embodiments of the ICA system 340 and the MIGM system 345. The server computing systems and the IFPLMM system can be implemented using a number of hardware components that form electronic circuitry adapted and configured to, when operating in conjunction, perform at least some of the techniques described herein. In the illustrated embodiment, each server computing system 300 includes one or more hardware central processing units ("CPUs") or other hardware processors 305, various input / output ("I / O") components 310, including a display 311, a network connection 312, a computer-readable media drive 313, and other I / O devices 315 (e.g., a keyboard, a mouse, or other pointing device, a microphone, a speaker, a GPS receiver, etc.), storage 320, and memory 330. For example Each server computing system 380 can include similar hardware components as the server computing systems 340, including one or more hardware CPU processors 382, various I / O components 382, storage 385, and memory 387, although some details of the server computing systems 300 are omitted from the server computing systems 380 for brevity.
[0110] The server computing systems 380 and the executing IFPLMM system 389 can communicate with other computing systems and devices, such as the following, via one or more networks 399 (e.g., the Internet, one or more cellular telephone networks, etc.): user client computing devices 390 (e.g., for viewing floor plans, associated images, and / or other related information); the ICA and MIGM server computing systems 300; one or more image capture mobile devices 360; optionally other navigable devices 395 that receive and use floor plans and determined image capture locations and optionally other generated information for navigation purposes (e.g., for use by semi-autonomous or fully autonomous vehicles or other devices); and optionally other computing systems (not shown) for storing and providing additional information related to buildings; for capturing interior data of buildings; for storing and providing information to client computing devices, such as additional supplemental information associated with images and their encompassed buildings or other surrounding environments, etc. For example For example For example
[0111] In the illustrated implementation, an implementation of the IFPLMM system 389 is executing in the memory 387 to perform at least some of the described techniques, such as using the processor 381 to execute software instructions of the system 389 in a manner that configures the processor 381 and the computing system 380 to perform automated operations that implement those described techniques. The illustrated implementation of the IFPLMM system can include one or more components (not shown) to each perform a portion of the functionality of the IFPLMM system, and the memory can further optionally execute one or more other programs 388. As one specific example, in at least some implementations, a copy of the ICA and / or MIGM system can be executed as one of the other programs 388, such as instead of or in addition to the ICA system 340 and the MIGM system 345 on the server computing system 300. The IFPLMM system 389 can further store and / or retrieve various types of data during its operation on the storage 385 (in one or more databases or other data structures), such as: various types of floor plan information and other building survey information 391 (similar to or the same as the information 391), 2D floor plans generated and saved along with semantic information about the locations of wall elements and other elements on those floor plans, 2.5D and / or 3D models generated and saved, building and room dimensions for use with associated floor plans, additional image and / or annotation information, and so on; information 393 about images whose capture locations are to be determined and associated information 392 about such determined capture locations; information 394 about generated building location circular descriptors and image circular descriptors; and optionally various other types of information. If present, the ICA system 340 and / or the MIGM system 345 can similarly store and / or retrieve various types of data during their operation on the storage 320 (in one or more databases or other data structures), and provide some or all of such information to the IFPLMM system 389 for use in its implementation (whether in a push and / or pull manner), such as various types of floor plan information and other building survey information 326 (similar to or the same as the information 391), various types of user information 322, captured 360° panoramic image information 324 (for analysis to generate floor plans; provision to users of client computing devices 390 for display, and so on, optionally including information about inter-panorama links that reflect relative location information of the panoramic images), and / or various types of optional additional information 329. For example For example For example For example For example For example For example (Various analytical information captured by the ICA system relating to the presentation or other uses of the interior of one or more buildings or other environments).
[0112] User client computing device 390 ( For example Mobile devices, image acquisition mobile devices 360, other navigable devices 395, and other computing systems may similarly include some or all of the same type of components described for server computing systems 300 and 380. As a non-limiting example, image acquisition mobile devices 360 are each shown as including one or more hardware CPUs 361, I / O components 362, storage devices 365, imaging systems 364, IMU hardware sensors 369, and memory 367, wherein a browser 368 and one or more client applications 369 (… Figure 5 One or both of the applications (specific to the IFPLMM system and / or ICA system) execute within memory 367 to participate in communication with the IFPLMM system 389, the ICA system 340, and / or other computing systems. Although specific components are not described for other navigable devices 395 or client computing systems 390, it will be understood that they may include similar and / or additional components.
[0113] We will also learn about the computing systems 300 and 380, as well as Figure 1 Other systems and devices included herein are merely illustrative and are not intended to limit the scope of the invention. Systems and / or devices may be modified to each comprise multiple interacting computing systems or devices and may be connected to other devices not specifically described, including via Bluetooth communication or other direct communication, through one or more networks (such as the Internet), via the Web, or via one or more private networks. Figure 3 Connecting to mobile communication networks, etc. More generally, the device or other computing system may include any combination of hardware that, when programmed or otherwise configured with specific software instructions and / or data structures, can interact and perform functions of the type described, including but not limited to desktop computers or other computers (…). For example This includes tablet computers, tablet PCs, database servers, network storage devices and other network devices, smartphones and other cellular phones, consumer electronic devices, wearable devices, digital music player devices, handheld gaming devices, PDAs, cordless phones, internet-connected appliances, and various other consumer products including appropriate communication capabilities. Additionally, in some embodiments, the functionality provided by the described IFPLMM system 389 may be distributed across various components, some of the functions described in the IFPLMM system 389 may not be provided, and / or other additional functions may be provided.
[0114] It will also be appreciated that, although various items are illustrated as being stored in memory or on storage while being used, these items or portions of them can be transferred between memory and other storage devices for purposes of memory management and data integrity. Alternatively, in other embodiments, some or all of the software components and / or systems can execute in memory on another device and communicate with the illustrated computing systems via inter-computer communication. Thus, some or all of the described techniques can be implemented in some embodiments by one or more software programs running on one or more computing systems (e.g., the IFPLMM system 389, executed on the server computing system 380) and / or data structures configured by the software programs. For example Some or all of the described techniques can be performed by hardware devices including one or more processors and / or memory and / or storage, such as by executing software instructions of the one or more software programs and / or by storing such software instructions and / or data structures, and in order to perform algorithms as described in the flowcharts and other disclosure herein, when configured by the one or more software programs and / or data structures. Moreover, in some embodiments, some or all of the systems and / or components can be implemented or provided in other ways, such as by being composed of one or more devices that are partially or wholly implemented in firmware and / or hardware (as opposed to being devices that are wholly or partially implemented by software instructions configuring a particular CPU or other processor), including but not limited to one or more application-specific integrated circuits (ASICs), standard integrated circuits, controllers (by executing appropriate instructions, and including microcontrollers and / or embedded controllers), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), etc. Some or all of the components, systems, and data structures can also be stored (as software instructions or structured data) on non-transitory computer-readable storage media, such as hard or flash drives or other non-volatile storage, volatile or non-volatile memory (RAM or flash RAM), network storage, or portable media articles (DVD disks, CD disks, optical disks, flash memory devices, etc.), for reading by appropriate drives or via appropriate connections. In some embodiments, the systems, components, and data structures can also be transmitted via generated data signals (as part of carrier waves or other analog or digital propagated signals) on a variety of computer-readable transmission media, including wireless and wired / cable-based media, and can take a variety of forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). Figure 5 For example For example For example For example For example For example For example This can be done as part of a single or multiple analog signal, or as multiple discrete digital packets or frames. In other embodiments, such a computer program product may also take other forms. Therefore, embodiments of this disclosure can be practiced using other computer system configurations.
[0115] For example A flowchart illustrating an example implementation of a system routine 400 for an Image Floor Plan Location Map Manager (IFPLMM) system is provided. This routine can be implemented, for example, by executing... For example The IFPLMM system 140, For example The IFPLMM system 389 and / or as per relevant regulations For example And it is implemented using the IFPLMM system described elsewhere in this document to perform automated operations related to determining the acquisition location of an image based at least in part on the analysis of the image visual data and subsequently using the determined acquisition location information in one or more automated ways. In the example of Figure 4, the acquisition location is determined relative to a floor plan of a building (such as a house), but in other embodiments, other types of mapping information may be used for other types of structures or for non-structural locations, and the determined acquisition location information may be used in ways other than those discussed with respect to routine 400, as discussed elsewhere in this document.
[0116] The described implementation of the routine begins at box 405, where information or instructions are received. The routine continues to box 410 to determine whether the instructions or other information received in box 405 instructs the determination of a circular descriptor for the building location of the indicated building, and if not, continues to box 440. Otherwise, the routine continues to execute boxes 415 through 430 to determine the circular descriptor for the building location, including obtaining a rasterized floor plan of the building in box 415. For example (for retrieving from storage devices, receiving in box 405, etc.), the rasterized floor plan optionally has associated semantic information about structural wall elements ( (Doors and openings between other rooms or areas, windows, boundaries between rooms or other areas, etc.). In box 420, the routine then generates a point cloud corresponding to the building structural elements shown on the floor plan. 2D point cloud), and determine the associated information of each point ( The routine then generates latent spatial features for each point using 2D XY position, normal direction, semantic data, etc., and uses a trained neural network to generate these features. In box 425, the routine then generates and stores a circular descriptor for each of the multiple building locations, which identifies the features of the point in the direction from that location. For each of the 360 degrees of horizontality). In block 430, the routine then optionally generates and stores a graph having nodes representing the rooms or other building areas and having inter-node edges corresponding to inter-area connections or other inter-area adjacencies, and associates each node with a generated building location circular descriptor corresponding to a building location in that room or other building area.
[0117] After block 430, or if instead it is determined in block 410 that the instruction or other information received in block 405 is not a determination of a building location circular descriptor for the indicated building, the routine continues to block 440 to determine whether the instruction or other information received in block 405 is an instruction to determine a capture location for an indicated image for the indicated building or within any known building), and if so, the routine continues to execute blocks 445-485 to do so, and otherwise continues to block 490.
[0118] In block 445, the routine obtains information about the image for which a capture location is to be determined, such as by receiving the image in block 405 or by otherwise retrieving a stored copy of the image. In block 450, the routine then proceeds to generate an image circular descriptor for the image, the image circular descriptor including information about features of the visual content of the image in each of a plurality of angular directions if the image is a 360° panoramic image, at each of the 360 degrees of horizontality of angular directions, such as relative to an angular direction determined to be the starting direction of the image).
[0119] In block 460, the routine then compares the image circular descriptor to some or all of the building location circular descriptors previously generated for one or more buildings for the indicated building, such as with respect to all rooms and / or all non-room areas; for one or more rooms and / or non-room areas with which the image can correspond; etc.) to determine a best matching building location circular descriptor, such as a building location circular descriptor having a minimum dissimilarity distance from the image circular descriptor. The routine further identifies a room location to be used as the determined capture location for the image based on a room location associated with the best matching building location circular descriptor, such as to use that associated room location as the determined capture location, or optionally instead to a further refinement of the determined capture location. In some implementations and scenarios, the routine can further refine the determined capture location according to one or more portions of the image the determined capture location corresponding to the image (e.g., the start direction of the image and / or the end direction of the image). In block 485, the routine then optionally adds information to the graph node corresponding to the determined capture location for the image, and updates some or all of the building location circular descriptors in the room or other area associated with that graph node using the visual data of the image and its determined capture location.
[0120] After block 485, the routine continues to block 488 to store the information determined and generated in blocks 415-485, and optionally display the determined image capture location information for the image in the enclosed room or other area on a floor map (or floor map excerpt) on the floor map, but in other embodiments, the determined information can be used in other ways (e.g., for automated navigation of one or more devices).
[0121] If instead it is determined in block 440 that the information or instructions received in block 405 are not a determination of a capture location for an image, then the routine instead continues to block 490 to perform one or more other indicated operations as appropriate. Such other operations can include, for example, receiving and responding to a request for previously determined image capture location information and / or for associated images , a request for such information to be displayed on one or more client devices, a request for such information to be provided to one or more other devices for use in automated navigation, etc.), obtaining and storing information about the building for use in later operations , information about floor plans and associated wall element locations in the rooms in the floor plans, etc.), performing geometric positioning techniques to test for an association of wall elements visible in the image to wall elements present in the room (whether to confirm a degree of match of a building location circular descriptor that has already been determined to be the best match for the image circular descriptor, and / or as part of identifying such a best matching building location circular descriptor), using machine learning techniques to learn a best encoding to allow matching of the image to room locations, etc.
[0122] After block 488 or 490, the routine continues to block 495 to determine whether to continue, such as until an explicit termination indication is received, or instead only if an explicit continuation indication is received. If it is determined to continue, then the routine returns to block 405 to wait for and receive additional instructions or information, otherwise it continues to block 499 and ends.
[0123] An example implementation of a flowchart for a building map viewer system routine 500 is illustrated. The routine can be performed, for example, by a building map viewer system 100, or by a component thereof, such as a building map viewer system 100A, or by a component thereof, such as a building map viewer system 100B. The building viewer client computing device 175 and its software system (not shown) The client computing device 390 and / or a building information viewer or presentation system as described elsewhere herein shall perform this function in order to receive and display mapping information of a defined area. 3D computer models, 2.5D computer models, 2D floor plans, etc.), this mapping information includes visual indications of one or more determined image acquisition locations; and optionally, additional information associated with a specific location ( The image is displayed in the mapping information. In the example, the survey information presented is for the building ( (For the interior of a building), but in other implementations, other types of mapping information may be presented and used in other ways for other types of buildings or environments, as discussed elsewhere in this document.
[0124] The described implementation of the routine begins at block 505, where an instruction or message is received. At block 510, the routine determines whether the received instruction or message instructs the display or other presentation of one or more building areas (…). The routine retrieves information about the building's interior (and other structures), and if not, proceeds to box 590. Otherwise, the routine advances to box 512 to retrieve floor plans and / or other generated survey information about the building (and other structures). (3D computer model) and optional instructions for associated links to the building's interior and / or surrounding locations, and selection of the initial view of the retrieved information ( (A view of the floor plan, at least some 3D computer models, etc.). In box 515, the routine then displays or otherwise presents the current view of the retrieved information, and waits for user selection in box 517. After user selection in box 517, if it is determined in box 520 that the user selection corresponds to the current location ( If the user selects a value to change the current view, the routine continues to box 522 to update the current view based on the user's selection, and then returns to box 515 to update the displayed or otherwise presented information accordingly. The user selection and corresponding update of the current view may include, for example, displaying or otherwise presenting an associated link information of the user's selection (…). , a specific image associated with the visual indication displayed at the determined acquisition location, and changing the display mode of the current view ( Zoom in or out; rotate information as appropriate; select new portions of the floor plan and / or 3D computer model to display or otherwise present, such as where some or all of the new portion was previously not visible, or where the new portion is a subset of the previously visible information; etc.
[0125] If instead it is determined in block 510 that the instruction or other information received in block 505 will not present information representing the interior of a building, then the routine instead proceeds to block 590 to perform any other indicated operations (such as any housekeeping tasks), configure parameters for use in various operations of the system (at least in part based on information specified by a user of the system, such as a user of a mobile device that captured one or more interiors of a building, an operator user of the IFPLMM system, etc.), obtain and store other information about users of the system, respond to requests for generated and stored information, etc., as appropriate.
[0126] After block 590, or if it is determined in block 520 that the user selected does not correspond to the current location, then the routine proceeds to block 595 to determine whether to continue, such as until an explicit termination indication is received, or instead only if an explicit continuation indication is received. If it is determined to continue (e.g., if the user made a selection in block 517 relating to a new location to be presented), then the routine returns to block 505 to await additional instructions or information (or continues to block 512 if the user made a selection in block 517 relating to a new location to be presented), and if not, then proceeds to step 599 and ends.
[0127] The non-exclusive example implementations described herein are further described in the following clauses.
[0128] A01. A computer-implemented method comprising:
[0129] obtaining, by one or more computing devices and for a house having a plurality of rooms, a rasterized two-dimensional floor plan of the house having associated semantic information about locations of doors and windows and inter-wall boundaries of the plurality of rooms;
[0130] generating, by the one or more computing devices, building location description information for the house including:
[0131] generating, by sampling a structural location of the house shown on the rasterized two-dimensional floor plan, a two-dimensional point cloud having a plurality of points representing a structure of the house, including associating with each point information including a two-dimensional location of the point on the two-dimensional floor plan and including normal direction information for a set of neighboring points of the point and including semantic information for the point about any locations of doors and windows and inter-wall boundaries corresponding to the point;
[0132] determining, by the one or more computing devices, first latent space features associated with points of the two-dimensional point cloud by supplying the two-dimensional point cloud to a first trained neural network; and
[0133] generating architectural location circular descriptors for a plurality of architectural locations in a designated mesh image traversing the plurality of rooms of the house, including for each of the architectural locations, determining angular directions in 360 horizontal degrees from the architectural location to at least some points of the point cloud, and encoding in one of the architectural location circular descriptors associated with the architectural location, information about some of the first latent space features associated with the at least some points;
[0134] generating, by the one or more computing devices, an image circular descriptor for a panoramic image taken in one of the plurality of rooms and having 360 horizontal degrees of visual information, including determining, by supplying the panoramic image to a second trained neural network, second latent space features associated with visual data of the panoramic image, and wherein the image circular descriptor encodes information identifying a designated direction within the visual data to the second latent space features;
[0135] comparing, by the one or more computing devices, the image circular descriptor to the architectural location circular descriptors to determine one of the architectural location circular descriptors whose encoded information best matches the encoded information of the image circular descriptor;
[0136] associating, by the one or more computing devices and based on the comparison, the panoramic image with a determined location on the two-dimensional floor plan, wherein the determined location includes the architectural location in the one room associated with the determined one architectural location circular descriptor, and further includes orientation information correlating the determined angular direction of the architectural location with the identified designated location of the panoramic image; and
[0137] using, by the one or more computing devices, the determined location of the panoramic image on the two-dimensional floor plan of the house for navigation of at least the one room of the house.
[0138] A02. The computer-implemented method of clause A01, wherein generating the building location circular descriptor further comprises: obtaining a first enumerated set of ranges of angles of incidence; obtaining a second enumerated set of ranges of distances; and performing the encoding of each of the building location circular descriptors of information about some of the first potential spatial features by, for each of at least some points of the building location for which the building location circular descriptor is used: encoding in the building location circular descriptor information of one of 360 horizontal degrees from the building location to a point comprising one from the first enumerated set of ranges of angles of incidence and one from the second enumerated set of ranges of distances.
[0139] A03. The computer-implemented method of any of clauses A01-A02, further comprising using, by the one or more computing devices, the two-dimensional floor plan to further control navigation activities of an autonomous vehicle, including providing the two-dimensional floor plan for use by the autonomous vehicle in moving between the plurality of rooms of the house.
[0140] A04. The computer-implemented method of any of clauses A01-A03, wherein using the determined location further comprises displaying, by the one or more computing devices, the two-dimensional floor plan, the two-dimensional floor plan showing the plurality of rooms and including one or more visual indications on the displayed two-dimensional floor plan of the determined location and the orientation information of the panoramic image in the one room.
[0141] A05. A computer-implemented method comprising:
[0142] obtaining, by a computing device and for a building, building location description information comprising a plurality of building location circular descriptors for a plurality of building locations in the building, wherein each building location circular descriptor is associated with one of the building locations and has first angular information about first potential spatial features identified by a first trained neural network using a two-dimensional floor plan of the building for structural elements in a specified angular direction from the associated building location;
[0143] generating, by the computing device, an image circular descriptor for a panoramic image, the panoramic image captured in a room of the building and comprising visual information about at least some walls of the room, wherein the image circular descriptor has second angular information about second potential spatial features identified by a second trained neural network from the visual information of the panoramic image in a specified direction;
[0144] comparing, by the computing device, the image circular descriptor with the building location circular descriptors to determine one of the building location circular descriptors that is in the room and has first angular information that best matches the second angular information of the image circular descriptor;
[0145] associating, by the computing device and based on the comparison, the panoramic image with a determined location in the room and a determined orientation, the determined location being based on the building location associated with the determined one building location circular descriptor, and the determined orientation identifying at least one direction from the building location corresponding to a specified portion of visible information in the panoramic image; and
[0146] presenting, by the computing device, the two-dimensional floor plan of the building and showing information of the room with a visual indication identifying at least the determined location of the panoramic image, thereby using the presented information to navigate the building.
[0147] A06. The computer-implemented method of any of clauses A01-A05, wherein presenting a floor plan further comprises visually indicating the determined orientation, and wherein the method further comprises: presenting, by the computing device and in response to a user selection of the visual indication on the presented floor plan, at least a portion of the panoramic image corresponding to the determined orientation.
[0148] A07. The computer-implemented method of any of clauses A01-A06,
[0149] wherein the visual information of the panoramic image comprises a 360- degree horizontal visual coverage from a capture location of the panoramic image,
[0150] wherein, for each of the 360-degree horizontal visual coverage from the capture location, the image circular descriptor comprises information about at least some of the second latent space features, the at least some of the second latent space features being associated with any structural elements of the room that are visible from the capture location in a direction corresponding to the horizontal visual coverage, and
[0151] wherein, for each of the 360-degree horizontal from the building location associated with the building location circular descriptor, each of the building location circular descriptors comprises information about at least some of the first latent space features, the at least some of the first latent space features being associated with any structural elements of the surrounding room that are visible from the building location in a direction corresponding to the horizontal visual coverage.
[0152] A08. The computer-implemented method of clause A07, wherein the structural elements of the building include at least one door, at least one window, and at least one wall boundary, and wherein obtaining the building location description information includes generating the building location circular descriptor that includes generating a two-dimensional point cloud having a plurality of points from the two-dimensional floor plan, that includes associating information with each of the points, the information including two-dimensional location information of the point and normal direction information of the point and semantic information about any structural element associated with the point, and that includes analyzing the points and associated information to generate the first latent space features, wherein each of the points is associated with at least one of the first latent space features.
[0153] A09. The computer-implemented method of any one of clauses A07-A08, further comprising determining the one building location circular descriptor having the most matching angular information to the information included in the image circular descriptor without using any depth information gathered from any depth sensor about a depth of any surrounding element from the gathering location to the room by performing the generating and the comparing.
[0154] A10. The computer-implemented method of any one of clauses A07-A09, further comprising selecting the plurality of building locations in the building by specifying a grid of building locations that covers a floor of at least some of the plurality of rooms of the building.
[0155] A11. The computer-implemented method of clause A10, wherein comparing the image circular descriptor and the building location circular descriptors includes performing a nearest neighbor search of the building locations of the grid that includes identifying the determined one building location circular descriptor by repeatedly moving from at least one current building location in the grid to at least one neighbor building location in the grid if the dissimilarity of the at least one neighbor building location in the grid to the image circular descriptor is less than the dissimilarity of the at least one current building location in the grid to the image circular descriptor.
[0156] A12. The computer-implemented method of any one of clauses A07-A11,
[0157] wherein comparing the image circular descriptor and the building location circular descriptors further includes:
[0158] for a specified type of feature, analyzing the visual information to identify at least one of the 360 horizontal degree visual coverage from the gathering location in which the feature is present;
[0159] For each of at least some of the building location circular descriptors, comparing the image circular descriptor and the building location circular descriptors by:
[0160] identifying one or more of the 360 horizontal degrees from the building location associated with the building location circular descriptor that present the characteristic; and
[0161] synchronizing a location of each of the identified at least one of the 360 horizontal degree visual coverage from the capture location to a location of each of the identified one or more 360 horizontal degrees from the building location to determine whether information in the image circular descriptor at other horizontal degree coverage relative to the synchronized locations matches information in the building location circular descriptor at other horizontal degree coverage; and
[0162] selecting one of at least some of the building location circular descriptors as a determined one building location circular descriptor based on the selected one building location circular descriptor having the identified synchronized location, where the information in the building location circular descriptor at other horizontal degree coverage most matches the information in the image circular descriptor at other horizontal degree coverage, and using the identified synchronized location to determine an orientation of the panoramic image in the room.
[0163] A13. The computer-implemented method of clause A12, wherein the specified type of characteristic is one of: a visible wall normal to a line along the identified horizontal degree visual coverage, or a specified type of wall element visible at the identified horizontal degree visual coverage.
[0164] A14. The computer-implemented method of any one of clauses A01 to A13, wherein comparing the image circular descriptor and the building location circular descriptors comprises for each of at least some of the building location circular descriptors: determining a probability that the image circular descriptor and the building location circular descriptor are a match due to a difference being less than a specified threshold; and selecting one of the at least some building location circular descriptors having a highest probability of matching the image circular descriptor as a determined one building location circular descriptor.
[0165] A15. The computer-implemented method of any of clauses A01-A14, wherein comparing the image circular descriptor and the building location circular descriptors comprises, for each of at least some of the building location circular descriptors: using a circular bulldozer to distance measure a distance between the image circular descriptor and the building location circular descriptor; and selecting one of the at least some building location circular descriptors having a smallest measured distance to the image circular descriptor as the determined one building location circular descriptor.
[0166] A16. The computer-implemented method of any of clauses A01-A15,
[0167] further comprising: obtaining an angle range of a first enumerated set; obtaining a distance range of a second enumerated set; and generating each of the building location circular descriptors by: for each of at least some points of the structural elements visible from the building location of the building location circular descriptor, encoding information in the building location circular descriptor about some of the first latent spatial features; for one of 360 horizontal degrees from the building location to a point comprising one from the angle range of the first enumerated set and one from the distance range of the second enumerated set, encoding information in the building location circular descriptor.
[0168] A17. The computer-implemented method of any of clauses A01-A16, further comprising: determining a location of the panoramic image in the room by supplying the panoramic image and a building location associated with the determined one building location circular descriptor to a refinement neural network; and receiving an adjusted location based on the building location and adjusted to reflect the visual information of the panoramic image.
[0169] A18. The computer-implemented method of any of clauses A01-A17,
[0170] wherein associating the panoramic image with the determined location and the determined orientation further comprises operations performed by the computing device of:
[0171] for each of a plurality of building location circular descriptors associated with one of a plurality of building locations in the room, generating additional visual information for the building location circular descriptor, the additional visual information representing a view from the building location associated with the building location circular descriptor and comprising at least some of the second latent spatial features visible in a specified angular direction of the building location circular descriptor; and
[0172] determining a capture location of the additional image captured in the room by comparing an additional image circular descriptor generated from the additional image captured in the room to the plurality of building location circular descriptors, including using the generated additional visual information of the plurality of building location circular descriptors.
[0173] A19. The computer-implemented method of clause A18, further comprising generating a graph having a plurality of nodes, wherein at least one node represents each of a plurality of rooms of the building; associating the plurality of building location circular descriptors with one of the plurality of nodes representing the room; and after determining the location of the panoramic image, further associating the panoramic image with one of the nodes representing the room.
[0174] A20. The computer-implemented method of any of clauses A01-A19, wherein comparing the image circular descriptor to the building location circular descriptors includes using machine learning to identify one of the determined building location circular descriptors as most similar to the image circular descriptor.
[0175] A21. A computer-implemented method comprising a plurality of steps to perform an automated operation, the automated operation implementing the described techniques substantially as disclosed herein.
[0176] B01. A non-transitory computer-readable medium storing executable software instructions and / or other stored content that cause one or more computing systems to perform an automated operation, the automated operation implementing the method of any of clauses A01-A21.
[0177] B02. A non-transitory computer-readable medium storing executable software instructions and / or other stored content that cause one or more computing systems to perform an automated operation, the automated operation implementing the described techniques substantially as disclosed herein.
[0178] B03. A non-transitory computer-readable medium,
[0179] storing content that cause one or more computing devices to perform an automated operation, the automated operation comprising at least:
[0180] obtaining, by the one or more computing devices and for an image captured in an area associated with a building and including visual information related to at least some structural elements of the building, an image circular descriptor for the image, the image circular descriptor including information identifying features associated with the at least some structural elements in a specified direction within the visual information;
[0181] obtaining, by the one or more computing devices, building location circular descriptors, the building location circular descriptors each being associated with a building location and including angular information about features associated with points of structural elements of the building in a specified angular direction from the associated building location;
[0182] comparing, by the one or more computing devices, the image circular descriptor and the building location circular descriptors to determine one of the building location circular descriptors having angular information that best matches the information included in the image circular descriptor;
[0183] associating, by the one or more computing devices, the image with a determined location for the building, the determined location being based on the associated building location for the determined one building location circular descriptor; and
[0184] providing, by the one or more computing devices, information for the image related to the determined location of the building.
[0185] B04. The non-transitory computer-readable medium of clause B03, wherein the image is a panoramic image having 360 degrees of horizontal visual information, wherein obtaining the image circular descriptor includes generating the image circular descriptor by the one or more computing devices via analysis of the image via a trained neural network, and wherein providing the information related to the determined location of the image includes presenting a floor plan of the building including a visual indication of the determined location of the image.
[0186] B05. The non-transitory computer-readable medium of any of clauses B03-B04, wherein the area associated with the building includes at least one of a plurality of rooms of the building, and wherein the structural elements of the building include a plurality of doors or windows or inter-wall boundaries.
[0187] B06. The non-transitory computer-readable medium of any of clauses B03-B05, wherein the area associated with the building includes at least one exterior area proximate to the building, and wherein the structural elements of the building include a plurality of doors or windows or inter-wall boundaries.
[0188] B07. The non-transitory computer-readable medium of any one of clauses B03-B06, wherein the visual information of the image has less than 360 horizontal degrees of coverage, wherein the determined one additional angular descriptor is for a panoramic image taken at the determined location and having 360 horizontal degrees of coverage, and wherein comparing the angular descriptor of the image and the additional angular descriptor comprises matching the angular descriptor of the image and a subset of the determined one additional angular descriptor of the panoramic image.
[0189] C01. One or more computing systems comprising: one or more hardware processors and one or more memories storing instructions that, when executed by at least one of the one or more hardware processors, cause the one or more computing systems to perform automated operations that implement the method of any one of clauses A01-A21.
[0190] C02. One or more computing systems comprising one or more hardware processors and one or more memories storing instructions that, when executed by at least one of the one or more hardware processors, cause the one or more computing systems to perform automated operations that implement the described technology substantially as disclosed herein.
[0191] C03. A system comprising:
[0192] one or more hardware processors of one or more computing devices; and
[0193] one or more memories storing instructions that, when executed by at least one of the one or more hardware processors, cause at least one of the one or more computing devices to perform automated operations comprising at least:
[0194] obtaining descriptive information for a region of a building, the descriptive information comprising building location circular descriptors for a plurality of building locations in the region, wherein each building location circular descriptor is associated with one of the building locations and has angular information about features associated with structural elements of the building in a specified angular direction from the associated building location;
[0195] generating an additional circular descriptor for information recorded at a recorded location in the region, wherein the additional circular descriptor comprises information identifying features associated with at least some of the structural elements identifiable from the recorded information in a specified direction from the recorded location;
[0196] comparing the additional circular descriptor and the building location circular descriptors to determine one of the building location circular descriptors having angular information that best matches the information included in the additional circular descriptor;
[0197] based on the comparison, associating the recorded information and a location in the area, the location in the area being determined for the recorded information based on the building location associated with the determined one building location circular descriptor; and
[0198] providing information related to the determined location in the area for the recorded information.
[0199] C04. The system of clause C03, wherein the recorded information comprises a panoramic image having visual information, wherein the structural element comprises a wall element having at least one of a door or a window or an inter-wall boundary, wherein providing the information related to the determined location in the room comprises presenting a floor plan of a floor of the building comprising the area, and wherein the presented floor plan comprises a visual indication of the determined location in the area.
[0200] C05. The system of any one of clauses C03-C04, wherein the area of the building is one of a plurality of rooms of the building.
[0201] C06. The system of any one of clauses C03-C05, wherein the area of the building is an exterior area adjacent to the building.
[0202] D01. A computer program, which computer program is adapted to perform the method of any one of clauses A01-A21 when the computer program is run on a computer.
[0203] The aspects of the disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions. It will further be understood that in some embodiments, the functions provided by one or more of the illustrated routines can be provided in alternative ways, such as being divided among more routines or consolidated into fewer routines. Similarly, in some embodiments, illustrated routines can provide more or less functionality than is described, such as when other illustrated routines instead lack or include such functionality, or when the amount of functionality that is provided is altered. In some embodiments, the routines are implemented using computer program instructions that are loaded from or into a machine and executed by a processor of the machine. More specifically, the routines can be implemented using hardware description language (HDL) instructions in combination with a processor, a general purpose computer, a dedicated computer, or other similar machine. In some embodiments, the routines are implemented using instructions (e.g., machine language instructions) in combination with a processor, a general purpose computer, a dedicated computer, or other similar machine. In some embodiments, the routines are implemented using a combination of instructions and hardware logic (e.g., integrated circuits) having the functionality described herein. Finally, it will be appreciated that the routines can be implemented by a human being or a group of humans acting in conjunction with each other, such as in the case of a human operator manually performing the functionality described herein. The operations described above may, in some embodiments, be performed in a different order, and / or at different times - including substantially concurrently or in parallel - and the order of the actions may not be material to the described embodiments. In other words, as used herein the term "and / or" includes any and all combinations of one or more of the associated listed items. The operations described above may, in some embodiments, be performed in series or in parallel, and / or synchronously or asynchronously, and / or in a particular order, but in other embodiments, the operations can be performed in other orders and other manners. Any data structures discussed above can also be structured differently, such as by dividing a single data structure into multiple data structures and / or by merging multiple data structures into a single data structure. Similarly, in some embodiments, the illustrated data structures can store more or less information than described, such as when other illustrated data structures correspondingly lack or include such information, or when the amount or type of information stored varies.
[0204] In light of the above, it will be appreciated that, although specific embodiments have been described herein for purposes of illustration, various modifications can be made without departing from the spirit and scope of the application. Accordingly, the application is not limited except as by the appended claims and the elements recited therein. In addition, while certain aspects of the application can be presented in certain embodiments by way of an example, it is contemplated that various aspects of the application can be implemented in any number of different specific contexts. For example, while only some aspects of the application can be presented as being embodied in a computer-readable storage medium at a particular time, other aspects can similarly be embodied in a computer-readable storage medium at the same or different time.
Claims
1. A computer-implemented method for automating the analysis of visual data of images, comprising: The computing device obtains building location description information for the building, the building location description information including multiple circular descriptors of multiple building locations in the building, wherein each circular descriptor of building location is associated with one of the building locations and has first angular information related to a first latent spatial feature identified for a structural element of the building in a specified angular direction from the associated building location, wherein the first latent spatial feature is identified by a first trained neural network using a two-dimensional floor plan of the building; The computing device generates an image circular descriptor for a panoramic image captured in a room of the building and including visual information relating to at least some walls of the room, wherein the image circular descriptor has second angular information relating to a second latent spatial feature identified by a second trained neural network from the visual information of the panoramic image in a specified direction; The computing device compares the image circular descriptor with the building location circular descriptor to determine the building location circular descriptor that is in the room and has first corner information that best matches the second corner information of the image circular descriptor; The computing device, based on the comparison, associates the panoramic image with a determined location and a determined orientation within the room. The determined location is based on the building location associated with a determined circular descriptor of a building location, and the determined orientation identifies at least one direction from the building location corresponding to a specified portion of the visible information in the panoramic image. The computing device presents the two-dimensional floor plan including the building and indicates information about the rooms with visual indications that identify at least the determined locations of the panoramic image, thereby using the presented information to navigate the building.
2. The computer-implemented method according to claim 1, wherein, Presenting the floor plan also includes visually indicating the determined orientation, and wherein the method further includes: by the computing device and in response to a user selection based on the visual indication on the presented floor plan, presenting at least a portion of the panoramic image corresponding to the determined orientation.
3. The computer-implemented method according to claim 1, wherein, The visual information of the panoramic image includes a 360-degree horizontal visual coverage area from the acquisition position of the panoramic image. Wherein, for each of the 360-degree horizontal visual coverage area from the acquisition position, the image circular descriptor includes information relating to at least some of the second latent spatial features, which are associated with any structural elements of the room visible from the acquisition position in the direction corresponding to the horizontal visual coverage area, and Wherein, for each of the 360 horizontal degrees from the building location associated with the circular descriptor of the building location, each of the circular descriptors of the building location includes information relating to at least some of the first potential spatial features, which are associated with any structural elements of the surrounding rooms visible from the building location in a direction corresponding to the visual coverage of the horizontal degree.
4. The computer-implemented method according to claim 3, wherein, The structural elements of the building include at least one door, at least one window, and at least one boundary between walls, and wherein obtaining the building location description information includes generating a circular descriptor of the building location, which includes generating a two-dimensional point cloud having a plurality of points from the two-dimensional floor plan, which includes associating information with each of the points, the information including two-dimensional location information of the point and normal direction information of the point and semantic information about any structural elements associated with the point, and includes analyzing the points and the associated information to generate a first latent spatial feature, wherein each of the points is associated with at least one of the first latent spatial features.
5. The computer-implemented method of claim 3, further comprising determining a circular descriptor of a building location having angular information that best matches the information included in the circular descriptor of the image by performing the generation and the comparison, without using any depth information acquired from any depth sensor relating to the depth from the acquisition location to any surrounding elements of the room.
6. The computer-implemented method of claim 3, further comprising selecting the plurality of building locations in the building by specifying a grid of building locations covering floors of at least some of the rooms in the plurality of rooms of the building.
7. The computer-implemented method according to claim 6, wherein, Comparing the image circular descriptor and the building location circular descriptor includes performing a nearest neighbor search on the building location of the grid, which includes identifying a determined building location circular descriptor by repeatedly moving from at least one current building location in the grid to at least one nearest neighbor building location in the grid if the dissimilarity between at least one nearest neighbor building location in the grid and the image circular descriptor is less than the dissimilarity between at least one current building location in the grid and the image circular descriptor.
8. The computer-implemented method according to claim 3, wherein, The comparison between the image circular descriptor and the building location circular descriptor further includes: For a specified type of feature, the visual information is analyzed to identify the presence of at least one of the features within the 360-degree horizontal visual coverage area from the acquisition location; For each of at least some of the circular descriptors for building locations, the image circular descriptor and the building location circular descriptor are compared by the following operation: Identify one or more of the characteristics present in the 360-degree horizontal dimension from the building location associated with the circular descriptor of the building location; and The position of each of at least one of the identified 360-degree horizontal visual coverage areas from the acquisition location is synchronized to the position of each of one or more of the identified 360-degree horizontal coverage areas from the building location, to determine whether information in the image circular descriptor relative to the synchronized position that is in other horizontal coverage areas matches information in the building location circular descriptor that is in other horizontal coverage areas; and Based on a selected building location circular descriptor with an identified synchronization location, one of at least some of the building location circular descriptors is selected as a determined building location circular descriptor, wherein information in other horizontal coverage areas of the building location circular descriptor best matches information in other horizontal coverage areas of the image circular descriptor, and the identified synchronization location is used to determine the orientation of the panoramic image in the room.
9. The computer-implemented method according to claim 8, wherein, The specified type of characteristic is one of the following: a visible wall orthogonal to a line along the identified horizontal visual coverage area, or a wall element of the specified type visible within the identified horizontal visual coverage area.
10. The computer-implemented method according to claim 1, wherein, Comparing the image circular descriptor and the building location circular descriptor includes, for each of at least some of the building location circular descriptors: determining the probability that the image circular descriptor and the building location circular descriptor are a match because the difference is less than a specified threshold; and selecting the building location circular descriptor with the highest probability of matching the image circular descriptor as the determined building location circular descriptor.
11. The computer-implemented method according to claim 1, wherein, Comparing the image circular descriptor and the building location circular descriptor includes, for each of at least some of the building location circular descriptors, using bulldozer distance to measure the difference between the image circular descriptor and the building location circular descriptor; And select one of the at least some building location circular descriptors that has the minimum measured distance to the image circular descriptor as the determined building location circular descriptor.
12. The computer-implemented method according to claim 1, further comprising: Obtain the angle range of the first enumeration group; obtain the distance range of the second enumeration group; Each of the circular descriptors for building locations is generated by: encoding information in the circular descriptor for building locations related to some of the first potential spatial features for each of at least some points of the structural element visible from the building location in the circular descriptor for building locations; and encoding information in the circular descriptor for building locations from the building location to one of the 360-degree horizontal distances from the building location to a point including one of the angular ranges from the first enumeration group and one of the distance ranges from the second enumeration group.
13. The computer-implemented method according to claim 1, further comprising: The location of the panoramic image in the room is determined by feeding the panoramic image and the building location associated with a circular descriptor of a determined building location to the refined neural network; And the adjusted position that receives the visual information based on the location of the building and adjusted to reflect the panoramic image.
14. The computer-implemented method according to claim 1, wherein, Associating the panoramic image with the determined location and the determined orientation also includes the following operations performed by the computing device: For each of a plurality of circular descriptors of building locations associated with one of a plurality of building locations in the room, additional visual information is generated for the circular descriptor of the building location, the additional visual information representing a view of the building location associated with the circular descriptor of the building location and including at least some of the second potential spatial features visible in a specified angular direction of the circular descriptor of the building location; as well as The acquisition location of the additional images captured in the room is determined by comparing the additional image circular descriptor generated from the additional images captured in the room with the plurality of building location circular descriptors, including additional visual information generated using the plurality of building location circular descriptors.
15. The computer-implemented method according to claim 14, further comprising: Generate a graph with multiple nodes, wherein at least one node represents each of the multiple rooms in the building; Associating the circular descriptors of the multiple building locations with one of the multiple nodes representing the room; and after determining the location of the panoramic image, also associating the panoramic image with a node representing the room.
16. The computer-implemented method according to claim 1, wherein, Comparing the image circular descriptor and the building location circular descriptor includes using machine learning to identify a determined building location circular descriptor as the most similar to the image circular descriptor.
17. A non-transitory computer-readable medium storing content that causes one or more computing devices to perform automated operations, the automated operations comprising at least: The one or more computing devices obtain an image circular descriptor for an image captured in an area associated with a building and including visual information related to at least some structural elements of the building. The image circular descriptor is generated by analyzing the image via a first trained neural network and includes information identifying features associated with the at least some structural elements in a specified direction within the visual information. A circular descriptor of a building location is obtained by the one or more computing devices, each of the circular descriptors of a building location being associated with a building location, generated by analysis of information about the building location via a second trained neural network, and including angular information about features associated with points of structural elements of the building in a specified angular direction from the associated building location; The one or more computing devices compare the image circular descriptor and the building location circular descriptor to determine the building location circular descriptor that has the corner information that best matches the information included in the image circular descriptor; The image is associated with a determined location for the building by the one or more computing devices, the determined location being based on the associated building location for a circular descriptor of the determined building location; as well as The one or more computing devices provide information about the image relating to the determined location of the building.
18. The non-transitory computer-readable medium according to claim 17, wherein, The image is a panoramic image with 360-degree horizontal visual information, wherein the information provided in relation to the determined location of the image includes: a floor plan of the building that presents a visual indication of the determined location of the image, wherein the area associated with the building includes one of a plurality of rooms of the building or at least one of the external areas of the building, and wherein the structural elements of the building include a plurality of doors or windows or boundaries between walls.
19. A system for automatically analyzing visual data of images, comprising: One or more hardware processors of one or more computing devices; as well as One or more memories storing instructions that, when executed by at least one of the one or more hardware processors, cause at least one of the one or more computing devices to perform an automated operation, the automated operation comprising at least: Obtain descriptive information about a region of buildings, the descriptive information including circular descriptors of building locations for multiple building locations in the region, the circular descriptors of building locations being generated by analyzing information about the multiple building locations via a first trained neural network, wherein each circular descriptor of building location is associated with one of the multiple building locations and has angular information relating to features associated with structural elements of the building in a specified angular direction from the associated building location; For information recorded at a recording location in the region, an additional circular descriptor is generated by analyzing the recorded information via a second trained neural network, wherein the additional circular descriptor includes information that identifies at least some of the associated features among the structural elements that can be identified from the recorded information in a specified direction from the recording location; Compare the additional circular descriptor and the building location circular descriptor to determine the one of the building location circular descriptors that has the corner information that best matches the information included in the additional circular descriptor; Based on the comparison, the recorded information and the location within the region are associated, and the location within the region is determined for the recorded location based on the building location associated with a determined circular descriptor of a building location; and The recorded information provides information relating to the determined location within the area.
20. The system according to claim 19, wherein, The area of the building is one of a plurality of rooms in the building, and the recorded information includes a panoramic image with visual information, wherein the structural elements include a wall element having at least one of a door or a window or a boundary between walls, and wherein providing the information relating to a determined location in the room includes presenting a floor plan of the building including the area, wherein the presented floor plan includes visual indications of the determined location in the area.
Citation Information
Patent Citations
Generating floor maps for buildings from automated analysis of visual data from the buildings' interiors
CA3097164A1
Generating 3D models representing buildings
CN110033513A