Automatically generating and using building video based on building plan information analysis
By analyzing building images and floor plans and automatically generating building videos, the problem of difficulty in effectively capturing and using building internal information in the prior art is solved, and efficient building identification and navigation is achieved.
Patent Information
- Application Number
- CN202410213023.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-14
- Filing Date
- 2024-02-27
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively capture, represent and use building internal information, including identifying buildings that meet the criteria of concern and displaying visual information captured inside the building to users of remote locations.
By analyzing building images and floor plans, building videos are automatically generated using computing devices and used in an automated manner, such as for improved identification and navigation of buildings.
It realizes efficient capture and use of internal information of the building, allowing remote users to understand internal layout and details, and improves the efficiency of building identification and navigation.
Smart Images

Figure CN120010725A_ABST
Abstract
Description
Technical Field
[0001] The following disclosure generally relates to techniques for automatically generating building videos based on automatic analysis of acquired building information including building images and floor plans, and techniques for automatically using such generated building videos in further ways, such as for improved identification and navigation of buildings. Background Art
[0002] In various situations, such as building analysis, property inspections, real estate acquisition and development, general contracting, improvement cost estimation, etc., it may be desirable to understand the interior of a house or other building without physically traveling to and entering the building. However, it may be difficult to effectively capture, represent, and use such building interior information, including difficulty in identifying buildings that meet criteria of interest, and displaying visual information captured inside the building interior to a user at a remote location (e.g., enabling the user to understand the layout and other details of the interior, including controlling the display in a manner selected by the user). In addition, although a floor plan of a building may provide some information about the layout and other details of the building interior, such use of the floor plan has some disadvantages, including that the floor plan is difficult to construct and maintain, difficult to accurately scale and populate information about the interior, difficult to display and otherwise use, etc. Although a textual description of a building may sometimes exist, the textual description is often inaccurate and / or incomplete (e.g., lacks details about various attributes of the building, includes incorrect or misleading information, etc.). BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Figure 1A Included are diagrams depicting exemplary building interior environments and one or more computing systems used in embodiments of the present disclosure, including generating and using information representing a building, such as a video based on acquired building images and floor plans and determined properties of the building.
[0004] Figure 1B Examples of types of building description information are shown.
[0005] Figures 2A to 2V Examples are shown of automatically generating a video using information about a building from automatic analysis of building images and building floor plans for subsequent use in one or more automated ways.
[0006] Figure 3 is a block diagram illustrating a computing system suitable for executing an embodiment of a system that implements at least some of the techniques described in this disclosure.
[0007] FIG. 4A to FIG. 4DAn exemplary implementation of a flow chart for a Building Video Generation and Usage Manager (BVGUM) system routine is shown in accordance with an embodiment of the present disclosure.
[0008] Figure 5 An exemplary implementation of a flow chart for an Image Capture and Analysis (ICA) system routine according to an embodiment of the present disclosure is shown.
[0009] FIG. 6A to FIG. 6B An exemplary implementation of a flow chart for a Mapping Information Generation Manager (MIGM) system routine according to an embodiment of the present disclosure is shown.
[0010] FIG. 7A to FIG. 7B An exemplary implementation of a flow chart for a building information access system routine is shown in accordance with an implementation of the present disclosure. DETAILED DESCRIPTION
[0011] The present disclosure describes techniques for using a computing device to perform automated operations involving generating a video about a building by analyzing acquired building images and building floor plans and optionally other building information, such as for generating a video of the interior of a building based on images acquired within the building and a floor plan showing structural elements of the interior of the building, and subsequently using the generated video in one or more other automated manners, such as for improved building recognition and navigation. The automated techniques may include using information about objects (e.g., structural elements) and other properties of a building as part of video generation (e.g., selecting aspects of one or more rooms or other areas to highlight), for example, from automated analysis of information about the building (e.g., acquired images of the building, floor plans, etc.), and in some cases, automatically generating text descriptions about the determined building properties, such as by using one or more trained machine learning models (e.g., trained neural networks) and / or one or more trained language models (e.g., large language models), in at least some embodiments, such building information may be used for constructed multi-room buildings (e.g., houses, office buildings, etc.), and include panoramic images (e.g., with 360° horizontal video overlap) and / or other images (e.g., linear perspective images) acquired at acquisition locations in and around the building (e.g., without or using information from any depth sensor or other distance measurement device about the distance from the image's acquisition location to walls or other objects in the surrounding building). In some cases, the automated techniques may also include using the generated video in various ways, such as to assist in determining buildings that match specified criteria, for controlling navigation of a mobile device (e.g., an autonomous vehicle), for display or other presentation in a corresponding GUI (graphical user interface) on one or more client devices to enable virtual navigation of a building, etc. The following includes additional details regarding the automatic generation and use of video information about a building from the automated analysis of the building information, and in at least some embodiments, some or all of the techniques described herein may be performed via automated operation of a Building Video Generation and Usage Manager (“BVGUM”) system, as further discussed below.
[0012] Automatic operation of the BVGUM system in at least some embodiments may include obtaining one or more existing videos, each of the existing videos having visual data that overlaps at least some of a building (e.g., multiple rooms of a building and / or both the interior and exterior of the building); and automatically generating one or more additional videos for a portion of the building from at least one of the existing videos, for example by identifying one or more subset segments of at least one of the existing videos that meet one or more defined segment criteria and using the one or more subset segments as part or all of a newly generated video. Identifying one or more subset segments of at least one existing video that meets one or more defined segment criteria may include, for example, analyzing some or all frames of the existing video to identify one or more corresponding types of building information, such as a room or other area for which the frame includes visual data (e.g., a room or other area in which a camera capturing the video was located when the frame was captured), and / or one or more objects or building structural elements or other building attributes shown in the visual data of the frame, and / or the location in the video of transitions between rooms and / or other areas (e.g., through doorways or non-doorway wall openings), and / or locating some or all of the existing video to a building floor plan by determining some or all of the path through the building followed by the camera while the video was acquired (e.g., using SFM (Structure from Motion) or SLAM (Simultaneous Location And Mapping) or MVS (Multi-View Detection and Mapping)). Stereo, multi-view stereo) analysis, and / or detecting movement patterns within a path that meet one or more defined movement criteria (e.g., detecting a room or other area of interest based on a path that enters the room or other area, rotates in a circle, and then exits in substantially the same direction as entry). Such objects or building structural elements or other building attributes may have various forms and may be identified by analyzing the visual data of the frames in various ways, as discussed further below, non-exclusive examples of such building attributes include windows, doorways, non-doorway wall openings, walls and ceilings and floors and boundaries between at least two of them, built-in or movable objects, and the like.Additionally, in various embodiments, identification of a room or other area in which a frame includes visual data may be performed in various ways, with non-exclusive examples including the following: comparing the visual data of the frame with additional visual data of one or more images captured at a building to identify at least one image having a known location in the room or other area that matches the visual data (e.g., determining an inter-image pose between the image acquisition location and the frame acquisition location) to locate the frame within the room or other area (e.g., at the inter-image pose location of the frame); comparing one or more building attributes identified in the visual data of the frame with other building attributes determined from a floor plan of the building; Matching building attributes (e.g., analyzing a floor plan to identify structural elements that match visible structural elements in visual data of a frame) to locate a frame within a room or other area of the floor plan having matching building attributes (e.g., based on matching a specific location within the room or other area); determining that building objects or other building attributes identified from analysis of visual data of a frame are associated with a room type of a room in the building (e.g., associating a stove or refrigerator with a kitchen room type, associating a toilet or shower with a bathroom type, associating a bed or nightstand with a bedroom type, etc.), using motion patterns of the detected video path in the manner described above, etc.
[0013] Additionally, in some embodiments, the generation of the additional video from one or more existing videos may be based at least in part on user input and circumstances regarding segment criteria, including in response to information presented to one or more users regarding a floor plan of the building and / or regarding media blocks captured at the building, while in other embodiments, some or all of the segment criteria (e.g., one or more rooms or other areas that include visual data in the additional video, building attributes that include visual data in the additional video, etc.) may be automatically determined, such as using one or more machine learning models trained to determine such information. As a non-exclusive example, some or all of a floor plan of a building may be presented to the user, optionally with information indicated by the floor plan regarding the location where the building image was acquired and / or the floor plan of specific building attributes, and optionally with a visual representation of the path of the existing video superimposed on the floor plan, and in other embodiments and circumstances, the information regarding media blocks captured at the building may be grouped and / or presented in other ways, such as providing a list or other grouping of one or more media types (e.g., specific images, videos, audio recordings, etc.) that were partially or fully acquired within each of one or more rooms or other areas. In at least some embodiments, a user may specify some or all of the segmentation criteria for use in generating an additional video from one or more existing videos, such as by selecting one or more rooms or other areas (e.g., on a rendered floor plan) for visual data to be included in the additional video, by selecting a portion of an overlay visual representation of a path of an existing video (e.g., on a rendered floor plan) to include in the additional video (e.g., by drawing a box or other shape around the portion), by selecting one or more building attributes to include in the visual data of the additional video, etc. Additionally, once such an additional video is generated, information about it may similarly be overlaid on or otherwise associated with a floor plan of the building (e.g., overlaying the path of the additional video, including the additional video on a media piece shown for one or more rooms or other areas, etc.), and the generated additional video may similarly be presented to one or more users (e.g., sending the corresponding segmentation criteria to one or more users). The following includes additional details regarding: analyzing existing videos, identifying various types of information associated with existing videos, generating additional videos from one or more segments of one or more existing videos, presenting videos, presenting information about one or more videos on a floor plan, and presenting other types of building information, including information about Figure 2Q to Figure 2S Examples and their associated descriptions.
[0014] In at least some embodiments, automatic operation of the BVGUM system may include: using visual data from images acquired at the building and additional information from a floor plan of the building, such as by using one or more specified generation criteria for new videos (e.g., a path having a continuous sequence of positions or other positions from which visual data is provided, such as positions on a floor plan of the building; one or more orientations for each such position, such as the direction in three directions from a specified height at the position; the speed of movement between positions in the sequence; the type of visual transitions used between visual data for non-contiguous or non-adjacent positions in the sequence, etc.), automatically generating one or more new videos of a portion of the building, in at least some embodiments and scenarios, the generation of such new videos may include: using a NeRF (Neural Radiance Field) neural network and associated NeRF processing techniques with images at known locations to generate additional images (e.g., video frames) at other specified locations and orientations, and in at least some embodiments and scenarios, the generation of such new videos may include using a Gaussian Splatting processing technique with images at known locations to generate additional images (e.g., video frames) at other specified locations and orientations. Additional information from the floor plan used in the generation of the new video may include, for example: structural elements and other building attributes of rooms and / or other areas identified from the floor plan and optionally selected to include their visual data in the video, acquisition locations of images determined on the floor plan, etc. The existing images may be of various types (e.g., panoramic images, such as in equirectangular format; perspective images, such as in rectilinear and / or orthographic projection formats; frames of one or more existing videos, such as a single frame or a sequence of multiple consecutive frames, etc.), and in some embodiments and scenarios, the generation of the new video using visual data from one or more existing images may be based at least in part on user input, including in response to information presented to one or more users about the floor plan of the building and / or about media blocks captured at the building in order to specify some or all of the generation criteria for the new video, while in other embodiments, some or all of the generation criteria (e.g., a path for the new video to follow, building attributes to include in the visual data of the new video, using portions of the existing video as transitions between adjacent rooms, etc.) may be determined automatically, such as using one or more machine learning models trained to determine such information.As a non-exclusive example, some or all of a floor plan of a building may be presented to a user, optionally with information indicated on the floor plan or otherwise provided about acquisition locations of building images and / or building properties of the building, and in at least some embodiments, the user may specify some or all of the generation criteria based on the presented information, such as by selecting one or more rooms or other areas (e.g., on the presented floor plan) whose visual data is included in the new video, by selecting one or more building properties to include in the visual data of the new video, etc. Additionally, once such a new video is generated, information about it may similarly be overlaid on or otherwise associated with the floor plan of the building (e.g., overlaying a path for the new video, including the new video in a media piece shown for one or more rooms or other areas, etc.), and the generated new video may similarly be presented to one or more users (e.g., receiving corresponding segment criteria from one or more users). The following includes information about generating a new video using visual data of an image, presenting a video, presenting information about one or more videos on a floor plan, and presenting other types of building information (including information about a building). Figure 2T to Figure 2V Examples and their associated descriptions) for additional details.
[0015] In some embodiments and situations, the automatic operation of the BVGUM system may further include: automatically generating and adding additional visual data overlaid on one or more generated videos of the building, whether the additional video generated from at least a subset of the existing videos acquired at the building, and / or a new video generated using the visual data of the images acquired at the building and the additional information from the floor plan of the building. Such additional overlaid visual data may, for example, include the geographic shape and / or outline and / or other visual representations of objects at the building (e.g., objects that are partially or completely blocked or otherwise obscured from the current position and orientation of the visual data of the generated video, such as objects that are in the same room as the current position but blocked by one or more other objects and / or structural building elements, objects that are in different rooms or other building areas, such as exterior areas, objects that are blocked from the current position by one or more walls and / or other objects, etc.), while in other embodiments and situations, some or all of the additional overlaid visual data may include other types of visual data (e.g., visual representations of virtual objects that are not physically present at the building). In some embodiments and scenarios, the additional visual data may be generated and included in the generated video when the video is generated, and in some embodiments and scenarios, the additional visual data may be generated and included in the generated video after the video is generated.
[0016] In at least some embodiments, automated operation of the BVGUM system may include automatically analyzing visual data of images acquired in and around a building, and optionally associated image acquisition metadata (e.g., orientation information of the images, such as using orientation information from a compass sensor, location information from a GPS sensor, etc.), to generate one or more videos depicting the building. The automated technique may also include selecting one or more groups of images for the building, and for each such group of images, generating a video that includes visual overlays corresponding to selected building attributes of interest, and also includes an audible narrative based on automatically generated textual descriptions of the building attributes and optionally based on additional information (e.g., about the building as a whole, about transitions between multiple images of the group, etc.). In at least some such embodiments, automated operation includes selecting one or more building images for use in generating a video of the building, including determining a sequence of images if multiple images are selected, such image selection may include, for example, selecting images corresponding to particular rooms or other areas, that highlight particular types of building attributes, that have particular types of features, etc. Given a set of one or more selected images, the visual portion of the resulting video may be based on various types of manipulations of the visual data of such images, with non-exclusive examples including zooming, panning (e.g., within a panoramic image), tilting, and the like, including highlighting or emphasizing particular attributes of a building of interest for description, and using various types of transitions between the visual data of different images in the sequence. A narrative accompanying the video may further be automatically generated and synchronized with the video, including providing a narrative description of the selected building attributes, as further discussed below. In at least some such embodiments, the BVGUM system may use one or more machine learning models (e.g., one or more neural networks) to perform such image selection and sequence determination, and may be trained via supervised learning (e.g., using labeled versions of user-generated videos, such as video house tours generated by professional videographers or photographers), while in other embodiments, such machine learning models may be trained in an unsupervised manner (e.g., using unsupervised clustering). For building images used in video generation, in at least some embodiments and circumstances, some or all of the images acquired for the building and used in video generation may be panoramic images, each of which is acquired at one of a plurality of acquisition locations within or around the building, so that a panoramic image at each such acquisition location is generated from one or more videos at the acquisition location (e.g., a 360° video shot from a smart phone or other mobile device, which is held by a user rotating at the acquisition location), or multiple images acquired from the acquisition location in multiple directions (e.g., from a smart phone or other mobile device, which is held by a user rotating at the acquisition location), or all image information is captured simultaneously (e.g., using one or more fisheye lenses), etc.It should be understood that such panoramic images may in some cases be represented in a spherical coordinate system and provide up to 360° overlap around the horizontal and / or vertical axes, so that a user viewing the starting panoramic image can move the viewing direction within the starting panoramic image to different directions so that different images (or "views") are presented within the starting panoramic image (if the panoramic image is represented in a spherical coordinate system, this includes converting the image being rendered to a planar coordinate system). In addition, acquisition metadata about capturing such panoramic images can be obtained and used in various ways, such as data obtained from an IMU (inertial measurement unit) sensor or other sensor of a mobile device when the mobile device is carried by a user or moved between acquisition locations. Additional details about automatically generating building videos from building images are included below, including about. Figure 2D to Figure 2P Examples and descriptions of .
[0017] As described above, in at least some embodiments, automatic operation of the BVGUM system may include automatically determining attributes of a building of interest based at least in part on analyzing visual data of images acquired in and around the building and optionally associated image acquisition metadata, including in at least some cases by using one or more trained machine learning models (whether the same or different machine learning models are used to select images for use in video generation and / or for determining segment criteria and / or for determining generation criteria and / or for performing video generation), and in other embodiments, information about some or all of the building attributes may be determined in other ways, such as based in part on existing textual building descriptions. Such determined attributes may reflect characteristics of individual rooms or other areas of a building, such as corresponding to structural elements and other objects identified in the rooms and / or visible features or other attributes of objects and rooms. In particular, in at least some embodiments and cases, automatic analysis of the BVGUM system of building images may include identifying structural elements or various types of other objects in rooms of a building or in areas associated with the building (e.g., exterior areas, attached outbuildings or other structures, etc.), where non-exclusive examples of such objects include floors, walls, ceilings, windows, doorways, non-doorway wall openings, stair sets, fixtures (e.g., lighting or plumbing), appliances, cabinets, islands, fireplaces, countertops, other built-in structural elements, furniture, etc. Automatic analysis of acquired building images by the BVGUM system may further include determining specific attributes of each of some or all such identified objects, such as color, material type (e.g., surface material), estimated age, etc., and in some embodiments additional types of attributes, such as the direction facing the building object (e.g., for windows, doorways, etc.), natural lighting at a particular location (e.g., based on the geographic location and orientation of the building and the position of the sun at a specified time, such as time of day, time of month, time of month within a month, time of year, etc., and optionally corresponding to a particular object), views from a particular window or other location, etc. Attributes determined for a particular room based on one or more images acquired in the room (or based on one or more images acquired at a location having a view of at least some of the room) may include, for example, one or more of the following non-exclusive examples: room type, room dimensions, room shape (e.g., two-dimensional, or "2D," such as relative positions of walls; three-dimensional, or "3D," such as 3D point clouds and / or planar surfaces of walls, floors, and ceilings, etc.), type of room use (e.g., public space versus private space) and / or function (e.g., entertainment), location of windows and doorways in the room and other openings between rooms, types of connections between rooms, dimensions of connections between rooms, etc.In at least some such embodiments, for such automatic image analysis, the BVGUM system may use one or more machine learning models (e.g., classification neural network models) that are trained by supervised learning (e.g., using labeled data that identifies images with every possible object and attribute), while in other embodiments, such machine learning models may be trained in an unsupervised manner (e.g., using unsupervised clustering). The following includes additional details regarding automatically analyzing acquired images and / or other environmental data associated with a building to determine attributes of the building and its rooms, including regarding. Figure 2D to Figure 2P Examples and their associated descriptions.
[0018] As described above, in at least some embodiments, automatic operation of the BVGUM system may also include: automatically analyzing types of building information other than the acquired building images to determine additional attributes of the building, including, in at least some cases, determining attributes reflecting some or all (e.g., two or more rooms of the building) characteristics of the building by using one or more trained machine learning models (e.g., one or more trained neural networks, and whether the same or different as the machine learning models used to analyze images and / or select images for videos and / or generate videos from selected images and / or determine segment criteria and / or determine generation criteria), such as some or all layouts corresponding to some or all of the rooms of the building (e.g., based at least in part on interconnections between rooms and / or other inter-room adjacencies), such other types of building information may include, for example, one or more of the following: floor plans; a set of interlinked images, such as for a virtual tour; an existing text description of the building (e.g., information listing the building, such as included on a multiple listing service MLS, etc.). Such a floor plan of a building may include a 2D (two-dimensional) representation of various information about the building (e.g., rooms, doorways between rooms and other room-to-room connections, exterior doorways, windows, etc.), and may be further associated with various types of supplemental or additional information about the building (e.g., data for a plurality of other building-related attributes), such additional building information may, for example, include one or more of the following: a 3D or three-dimensional model of the building including height information (e.g., for building walls and room-to-room openings and other vertical areas); a 2.5D or two- and half-dimensional model of the building that, when reproduced, includes a visual representation of walls and / or other vertical surfaces without explicitly modeling the measured heights of those walls and / or other vertical surfaces; images and / or other types of data captured in the rooms of the building, including panoramic images (e.g., 360° panoramic images), etc., as discussed in more detail below. In some embodiments and circumstances, the floor plan and / or its associated information may further represent at least some information regarding the exterior of the building (e.g., for some or all of the real estate on which the building is located), such as exterior areas adjacent to doorways or other wall openings between the building and the exterior, or more generally, some or all of the exterior areas of a real estate that includes one or more buildings or other structures (e.g., a house and one or more outdoor buildings or other accessory structures, such as a garage, shed, pool house, separate guest room, mother-in-law unit or other accessory living unit, pool, patio, deck, walkway, etc.).
[0019] In at least some embodiments and situations, automatic analysis of building floor plans and / or other building information by the BVGUM system may include determining building attributes based on information about the building as a whole, such as objective attributes that can be independently verified and / or replicated (e.g., number of bedrooms, number of bathrooms, square footage, connectivity between rooms, etc., etc.), and / or subjective and objective attributes with associated uncertainty (e.g., whether the building has an open floor plan; has a typical / normal layout versus an atypical / odd / unusual layout; a standard versus a non-standard floor plan; an accessibility-friendly floor plan, such as by being accessible to one or more features, such as wheelchairs or other disabled persons and / or senior age, etc.). In at least some embodiments and situations, automatic analysis of the BVGUM system of building floor plans may further include determining building attributes based at least in part on information about inter-room adjacencies (e.g., inter-room connections between two or more rooms or other areas), such as based at least in part on the layout of some or all of the rooms of the building (e.g., all rooms on the same floor or all rooms that are part of a room grouping), including some or all of such subjective attributes, as well as other types of attributes, such as movement flow patterns of people through rooms. At least some of the building attributes so determined may be further based on information about the location and / or orientation of the building (e.g., about views available from windows or other exterior openings of the building, about the orientation of windows or other structural elements or other objects of the building, about natural lighting information available at a specified date and / or season and / or time, etc.). In at least some such embodiments, the BVGUM system may use one or more machine learning models (e.g., classification neural network models) for such automatic analysis of building floor plans, the machine learning models being trained by supervised learning (e.g., using labeled data identifying floor plans or other groups of rooms or other areas having each possible feature or other attribute), while in other embodiments such machine learning models may be trained in an unsupervised manner (e.g., using unsupervised clustering). The following includes additional details about the automatic analysis of a building's floor plan to determine attributes of the building, including information about Figure 2D to Figure 2V Examples and their associated descriptions.
[0020] As described above, in at least some embodiments, the automatic operation of the BVGUM system may also include: automatically generating a description of the building based on the automatically determined characteristics and other attributes, including, in at least some embodiments and situations, using one or more trained language models to generate descriptions of each of some or all of such determined attributes. In various embodiments, the generated descriptions of individual attributes may be further combined in various ways, such as by grouping the attributes and their associated descriptions in various ways (e.g., by room or other area; by attribute type, such as by object type and / or color and / or surface material; by degree of specificity or generality, such as grouping building-wide attributes and including their generated descriptions, then generating descriptions of the attributes grouped by room, then generating descriptions of the attributes corresponding to individual structural elements and other objects, etc.). After the attributes and / or building description of the building are generated or otherwise obtained, such as based on analysis of information about the building (e.g., images of the building, floor plans, and optionally other associated information), the generated building information may be used by the BVGUM system in various ways, including in some embodiments as part of a narrative that generates visual data accompanying the generated video and describes the information shown in the visual data. The generation of such a video narrative may include, for example, using one or more trained language models that take inputs such as objects and / or other attributes, associated location information (e.g., one or more rooms, one or more floors or other groups of rooms, etc.), timing and / or sequence information (e.g., a series of objects and / or other attributes to be highlighted or otherwise displayed in the video), etc., and generate corresponding textual descriptions. The following includes additional details regarding the automatic generation of descriptions of determined building attributes and the use of such generated descriptions as part of the video narrative, including information regarding Figure 2D to Figure 2P Examples and their associated descriptions.
[0021] After automatically generating a video of a building based on analysis of an image of the building and optionally other associated information, in some embodiments, the generated building information may also be used by the BVGUM system to automatically determine in various embodiments that the building matches one or more specified criteria (e.g., search criteria) in various ways, including identifying a building as similar to or otherwise matching one or more other buildings based on the building's corresponding video or other building information. Such criteria may include any one or more attributes or specified combinations thereof, and / or may more generally match the content of the building video narrative, with examples including based on specific objects and / or other attributes, based on adjacency information about which rooms are connected to each other and related inter-room relationship information (e.g., with respect to the overall building layout), based on specific rooms or other areas and / or attributes of those rooms or other areas, etc. Non-exclusive and non-limiting illustrative examples of criteria may include: kitchen with brick overlap island and dark wood flooring and a view facing north; building with bathroom adjacent to bedroom (i.e., no intervening hall or other room); deck adjacent to family room (optionally with a certain type of connection between them, such as French doors); 2 bedrooms facing south; master bedroom on second floor with a view of the ocean or more generally water; any combination of such specified criteria, etc. The following includes additional details regarding the use of the generated building information to help further identify buildings as matching specified criteria or otherwise used, including regarding Figure 2D to Figure 2V Examples and their associated descriptions.
[0022] The described techniques provide various benefits in various embodiments, including allowing information about multi-room buildings and other structures to be identified and used more efficiently and quickly and in ways that were not previously available, including generating one or more videos of a building having at least visual data for selected or automatically determined types of building data to provide improved navigation of the building and / or to help identify other related buildings. In addition, the described techniques may help automatically identify buildings that match specified criteria based at least in part on automatic analysis of various types of building information (e.g., images, floor plans, etc.), such criteria may be based, for example, on one or more of the following: attributes of specific objects within the building (e.g., in specific rooms or other areas, or more generally, attributes of those rooms or other areas), such as determined from analysis of one or more images acquired at the building; similarity to one or more other buildings; adjacency information about which rooms are interconnected and related inter-room relationship information, such as adjacency information about the overall building layout; similarity to specific building or other area features or other attributes; similarity to subjective attributes about features of the floor plan, etc. In addition, such automatic techniques allow such identification of matching buildings to be determined by using information obtained from the actual building environment (rather than from plans about how the building should be theoretically constructed), and enable the capture of changes in structural elements and / or visual appearance elements that occur after the building is initially constructed. Such described techniques also provide the benefit of allowing improved automatic navigation of buildings by mobile devices (e.g., semi-autonomous or fully autonomous vehicles) based at least in part on the identification of buildings that match specified criteria, including significantly reducing the computing power and time used to attempt to otherwise learn the layout of the building. In addition, in some embodiments, the described techniques can be used to provide an improved GUI in which a user can more accurately and quickly identify one or more buildings that match the specified criteria, and obtain information about the indicated buildings (e.g., for navigating the interior of the one or more buildings), including in response to a search request, as part of providing personalized information to a user, as part of providing a user with value estimates and / or other information about a building (e.g., after analyzing information about one or more target building floor plans, the floor plans are similar to one or more initial floor plans, or match the specified criteria), etc. Various other benefits are also provided by the described techniques, some of which are further described elsewhere herein.
[0023] Additionally, in some embodiments, one or more target buildings that are similar to specified criteria associated with a particular end-user are identified (e.g., identified as previously of interest to the end-user based on one or more initial buildings selected by the end-user and / or based on explicit and / or implicit activity by the end-user to specify such buildings; based on one or more search criteria specified by the end-user, whether explicitly and / or implicitly, etc.), and used in further automated activities to personalize interactions with the end-user. In various embodiments, such further automated personalized interactions may be of various types, and in some embodiments may include displaying or otherwise presenting to the end-user information about the one or more target buildings and / or additional information associated with those buildings. Additionally, in at least some embodiments, the video generated or otherwise presented to the end-user can be personalized for the end-user in various ways, such as based on the length of the video, the type of room shown, the type of property shown, etc., including in some embodiments dynamically generating a new building video for the end-user recipient based on information specific to the recipient, selecting one of multiple available videos of the building to present to the end-user recipient based on information specific to the recipient, customizing an existing building video for the end-user recipient terminal based on information specific to the recipient (e.g., removing portions of an existing video), etc. The following includes additional details regarding end-user personalization and / or presentation of the indicated building, including information regarding Figure 2D to Figure 2V Examples and their associated descriptions.
[0024] As described above, the automatic operation of the BVGUM system may include using the acquired building images and / or other building information, such as floor plans. In at least some embodiments, such a BVGUM system may operate in conjunction with one or more independent ICA (Image Capture and Analysis) systems and / or one or more independent MIGM (Mapping Information and Generation Manager) systems to obtain and use images and floor plans and other associated information for buildings from the ICA and / or MIGM systems, while in other embodiments, such a BVGUM system may be combined with some or all of the functions of such ICA and / or MIGM as part of the BVGUM system. In other embodiments, the BVGUM system may operate without using some or all of the functions of the ICA and / or MIGM systems, for example, if the BVGUM system obtains building images, floor plans and / or other relevant information from other sources (e.g., such building images, floor plans and / or related information are manually created or provided by one or more users).
[0025] With respect to the functionality of such an ICA system, it may, in at least some embodiments, perform automated operations to acquire images (e.g., panoramic images) at various acquisition locations associated with a building (e.g., inside multiple rooms of a building), and optionally further acquire metadata related to the image acquisition process (e.g., image pose information, such as using compass heading and / or GPS-based location) and / or movement of the acquisition device between acquisition locations, and in some embodiments, such acquisition and subsequent use of the acquired information may be performed without having or using information from a depth sensor or other distance measuring device regarding the distance from the image acquisition location to walls or other objects in the surrounding building or other structure. For example, in at least some such embodiments, such techniques may include using one or more mobile devices (e.g., a camera having one or more fisheye lenses and mounted on a rotatable tripod or having an automatic rotation mechanism; a camera having one or more fisheye lenses sufficient to capture 360° horizontally without rotation; a smartphone held and moved by a user, such as rotating the user's body and held smartphone in a 360° circle around a vertical axis; a camera held by or mounted on a user or user's clothing; a camera mounted on an air-based and / or ground-based drone or other robotic device, etc.) to capture visual data from a sequence of multiple acquisition locations within multiple rooms of a house (or other building). Additional details regarding the operation of devices implementing the ICA system are included elsewhere herein, such as performing such automatic operations and in some cases further interacting with one or more ICA system operator users in one or more ways to provide further functionality.
[0026] With respect to the functionality of such a MIGM system, it may, in at least some embodiments, perform automated operations to analyze a plurality of 360° panoramic images (and optionally other images) that have been acquired for the interior of a building (and optionally the exterior of a building), and generate a corresponding floor plan for the building, such as by determining the location of room shapes and passageways connecting the rooms for some or all of these panoramic images, and in at least some embodiments and circumstances by determining structural wall elements and optionally other objects in some or all of the rooms of the building. The types of structural wall elements corresponding to connecting passageways between two or more rooms may include one or more doorway openings and non-doorway wall openings in other rooms, windows, stairways, non-room corridors, etc., and the automated analysis of the images may identify such elements based at least in part on identifying the outlines of passageways, identifying content within passageways that is different from that outside of them (e.g., different colors or shading), etc. The automated operations may also include generating a floor plan of the building using the determined information, and optionally generating other mapping information for the building, such as by using the inter-room passageway information and other information to determine the relative locations of related room shapes to each other, and optionally adding distance scaling information and / or various other types of information to the generated floor plan. Additionally, in at least some embodiments, the MIGM system may perform further automated operations to determine and associate additional information with a building floor plan and / or a particular room or location within a floor plan, to analyze images and / or other environmental information (e.g., audio) captured within a building interior to determine specific objects and attributes (e.g., color and / or material type and / or other characteristics of specific structural elements or other objects, such as floors, walls, ceilings, countertops, furniture, fixtures, appliances, cabinets, islands, fireplaces, etc.; the presence and / or absence of specific objects or other elements, etc.), or otherwise determine relevant attributes (e.g., the direction facing a building object, such as a window; a view from a particular window or other location, etc.). The following includes additional details regarding the operation of one or more computing devices implementing the MIGM system to perform such automated operations, and in some cases further interact with one or more MIGM system operator users in one or more ways to provide further functionality.
[0027] For purposes of illustration, some embodiments are described below in which particular types of information are obtained, used, and / or presented in particular ways for particular types of structures and using particular types of devices, however, it will be understood that the described techniques may be used in other ways in other embodiments and thus the invention is not limited to the exemplary details provided. As a non-exclusive example, although in some embodiments particular types of data structures (e.g., videos, floor plans, virtual tours of interconnected images, generated building descriptions, etc.) are generated and used in particular ways, it should be understood that other types of information for describing buildings, including for buildings (or other structures or layouts) that are separate from houses, may be similarly generated and used in other embodiments, and that buildings identified as matching specified criteria may be used in other ways in other embodiments. Additionally, the term "building" refers herein to any partially or fully enclosed structure, typically but not necessarily including one or more rooms that visually or otherwise separate the interior space of the structure, non-limiting examples of such buildings include a house, an apartment building or individual apartments therein, a dormitory, an office building, a commercial building or other wholesale and retail structure (e.g., a shopping mall, a department store, a warehouse, etc.), a supplemental structure on a house together with another main structure (e.g., a garage separate from the house or a house on the house), etc. The terms "acquire" or "capture" as used herein with respect to the interior of a building, an acquisition location, or other location (unless the context clearly indicates otherwise) may refer to any preservation, storage, or recording of any media, sensor data, and / or other information related to spatial characteristics and / or visual characteristics and / or other perceptible characteristics of the interior of a building or a subset thereof, such as by a recording device or by another device that receives information from a recording device. As used herein, the term "panoramic image" may refer to a visual representation that is based on, includes, or is separable into multiple discrete component images that originate from substantially similar physical locations in different directions and depict a larger field of view than any discrete component image depicted alone, including images with sufficiently wide-angle viewing angles from the physical location to include angles in a single direction that exceed the angles perceptible from a person's gaze. As used herein, the term "sequence" of acquisition locations generally refers to two or more acquisition locations, each of which is visited at least once in a corresponding order, regardless of whether other non-acquisition locations are visited between them, and regardless of whether the visits to the acquisition locations occur during a single continuous time period, or at multiple different times, or are performed by a single user and / or device, or by multiple different users and / or devices. In addition, various details are provided in the drawings and text for illustrative purposes, but are not intended to limit the scope of the invention.For example, the sizes and relative positions of elements in the drawings are not necessarily drawn to scale, with some details omitted and / or greater prominence provided (e.g., by size and positioning) to enhance readability. In addition, the same reference numerals may be used in the drawings to identify the same or similar elements or actions.
[0028] Figure 1AIncluded are example block diagrams of various computing devices and systems that may participate in the described techniques in some embodiments, such as with respect to an example building 198 (a house in this example) shown and an example Building Video Generation and Usage Manager (“BVGUM”) system 140 executing on one or more server computing systems 180 in this example embodiment. In the illustrated embodiment, the BVGUM system 140 analyzes building information 142 obtained for one or more buildings (e.g., images, such as images 165 acquired by an ICA system; floor plans, such as floor plans 155 generated by a MIGM system; existing building videos, etc.), and uses the building information and information generated from its analysis to generate one or more building videos 141 for the one or more buildings, the one or more building videos 141 having visual data for the one or more buildings and, in some cases, including accompanying narration, optionally using supporting information provided by a system operator user via a computing device 105 through an intervening one or more computer networks 170, and in some embodiments and cases by using one or more trained machine learning and / or language models 144 as part of analyzing the generation of the building information 142 and / or videos 141. In some embodiments, the building information 142 analyzed by the BVGUM system may be obtained in a manner other than via an ICA and / or MIGM system (e.g., if such ICA and / or MIGM system is not part of the BVGUM system), so as to receive building images and / or floor plans and / or existing building videos from other sources. The BVGUM system may also use the building information so generated in one or more other automated manners, including in some embodiments as part of criteria for identifying buildings or other indications that match one another, such criteria in some embodiments may be provided or otherwise associated by particular users (e.g., objects or other attributes specified by users, floor plans or other building information indicated by those users, other existing building videos previously identified as being of interest to the users, etc.), and corresponding information 143 about various users may further optionally be stored and used to identify buildings that meet such criteria and subsequently use the information of the identified buildings in one or more further automated manners (e.g., using generated videos of the buildings). Additionally, in at least some embodiments and scenarios, one or more users of the client computing devices 105 may further interact with the BVGUM system 140 via one or more networks 170 to assist in some automated operations of the BVGUM system for identifying buildings that meet the criteria and / or subsequently using the identified floor plans in one or more other automated manners. Other details related to the automated operations of the BVGUM system are included elsewhere herein, including with respect to Figure 2D to Figure 2V and FIG. 4A to FIG. 4D .
[0029] Additionally, in this example, an internal capture and analysis (“ICA”) system (e.g., an ICA system 160 such as part of a BVGUM system executed on one or more server computing systems 180; an ICA system application 154 executed on a mobile image acquisition device 185, etc.) captures information 165 about one or more buildings or other structures (e.g., by capturing one or more 360° panoramic images and / or other images of multiple acquisition locations 210 in an exemplary house 198), and a MIGM (Mapping Information Generation Manager) system 160 executed on one or more server computing systems 180 (e.g., as part of a BVGUM system) also uses the captured building information and optional additional supporting information (e.g., provided by a system operator user via a computing device 105 through an intervening one or more computer networks 170) to generate and provide building floor plans 155 and / or other mapping-related information (not shown) for one or more buildings or one or more other structures. Although the ICA and MIGM systems 160 are shown in this exemplary embodiment as being executed on the same server computing system(s) 180 as the BVGUM system(s) (e.g., all systems are operated by a single entity or are otherwise executed in cooperation with one another, e.g., some or all functionality of all systems are integrated together), in other embodiments, the ICA system 160 and / or the MIGM system 160 and / or the BVGUM system 140 may be separately operated on one or more other systems separate from the one or more systems 180 (e.g., on a mobile device 185; one or more other computing systems not shown, etc.), whether instead of or in addition to being executed on the one or more systems 180. In some embodiments, the BVGUM may be configured to include a copy of those systems executing on device 180 (e.g., having a copy of the MIGM system 160 executing on device 185 to incrementally generate at least a partial building plan as building images are needed by the ICA system 160 executing on device 185 and / or by that copy of the MIGM system, with another copy of the MIGM system optionally executing on one or more server computing systems to generate a final complete building plan after all images have been acquired), and in other embodiments, the BVGUM may alternatively operate without the ICA system and / or the MIGM system, and instead obtain the panoramic images (or other images) and / or building plan from one or more external sources. Additional details related to the automated operation of the ICA and MIGM systems are included elsewhere herein, including with respect to FIG. 2A to FIG. 2D and about Figure 5 and FIG. 6A to FIG. 6B .
[0030] The various components of the mobile image acquisition computing device 185 are also Figure 1A, includes one or more hardware processors 132 (e.g., CPU, GPU, etc.) that execute software (e.g., ICA application 154, optional browser 162, etc.) using executable instructions stored and / or loaded on one or more memory / storage components 152 of the device 185, and optionally includes one or more imaging systems 135 of one or more types to acquire visual data of one or more panoramic images 165 and / or other images (not shown, such as linear perspective images), in some embodiments, some or all of such images 165 may be provided by one or more separate associated camera devices 184 (e.g., via wired / wired connections, via Bluetooth or other inter-device wireless communication, etc.), whether in addition to or in place of images captured by the mobile device 185. The illustrated embodiment of the mobile device 185 also includes: one or more sensor modules 148, which in this example include a gyroscope 148a, an accelerometer 148b, and a compass 148c (for example, as part of one or more IMU units on the mobile device that are not separately shown); one or more control systems 147 that manage I / O (input / output) and / or communications and / or networking of the device 185 (for example, in order to receive instructions from a user and present information to a user), such as for other device I / O and communication components 143 (for example, a network interface or other connection, a keyboard, a mouse or other pointing device, a microphone, a speaker, a GPS receiver, etc.); a display system 149 (for example, with a touch-sensitive screen); optionally, one or more depth sensing sensors or one or more types of other ranging components 136; optionally, a GPS (or global positioning system) sensor 134 or other position determination sensor (not shown in this example); optional other components (for example, one or more lighting components), etc. The other computing devices / systems 105, 175, and 180 and / or camera device 184 may include various hardware components and information stored in a manner similar to mobile device 185, which are not shown in this example for the sake of brevity and are referenced below. Figure 3 Discuss in more detail.
[0031] One or more users (e.g., end users, not shown) of one or more client computing devices 175 may further interact with the BVGUM system 140 (and optionally the ICA system 160 and / or the MIGM system 160) via one or more computer networks 170 to participate in providing input for generating a video and / or receiving information presented about the generated video and corresponding buildings and / or identifying buildings that meet target criteria and / or identifying building videos that meet target criteria, and subsequently using the identified and / or generated information (e.g., generated building videos) in one or more other automated manners, such client computing devices may each execute a building information access system (not shown) used by the user in the interaction, as discussed in more detail elsewhere herein, including with respect to FIG. 7A to FIG. 7B . Such interaction by one or more users may include, for example, specifying criteria to be used in generating a video, or specifying criteria to be used in searching for corresponding buildings or building videos, or otherwise providing building information regarding user focus criteria, or obtaining and optionally requesting information about one or more indicated buildings (e.g., playing or otherwise presenting a sequence of multiple videos of one or more buildings, such as in a playlist and / or by using an autoplay function), and interacting with the corresponding provided building information (e.g., changing between a floor plan and a view of a particular image at an acquisition location within or near the floor plan; changing the horizontal and / or vertical viewing direction of a corresponding view of a displayed panoramic image to determine the portion of the panoramic image to which the current user viewing direction is directed; viewing generated textual building information or other generated building information, etc.). Additionally, the floor plan (or portion thereof) may be linked to or associated with one or more other types of information, including a floor plan for a multiple-story or multi-story building, a floor plan with multiple associated sub-story floor plans for interconnecting different floors or levels (e.g., by connecting stairways), a two-dimensional ("2D") floor plan for a building to be linked to or associated with a three-dimensional ("3D") rendering of a building, etc. Additionally, although in Figure 1A Not shown, but in some embodiments, client computing device 175 (or other device, not shown) may receive and use information about a building (e.g., an identified floor plan and / or other mapping-related information) in an additional manner to control or assist in the automated navigation activities of those devices (e.g., by an autonomous vehicle or other device), rather than or in addition to displaying the identified information.
[0032] exist Figure 1AIn the computing environment shown, network 170 may be one or more publicly accessible linked networks, possibly operated by various different parties, such as the Internet. In other implementations, network 170 may have other forms. For example, network 170 may instead be a private network, such as a company or university network that is completely or partially inaccessible to non-privileged users. In other implementations, network 170 may include private networks and public networks, wherein one or more private networks may access and / or access one or more public networks. In addition, network 170 may include various types of wired and / or wireless networks in various circumstances. In addition, client computing device 175 and server computing system 180 may include various hardware components and stored information, as described below with reference to Figure 3 Discuss in more detail.
[0033] exist Figure 1A In an example, the ICA system may perform automated operations involved in generating multiple 360° panoramic images at multiple relevant acquisition locations (e.g., in multiple rooms or other locations within a building or other structure, and optionally around some or all of the exterior of the building or other structure), such as using visual data acquired via one or more mobile devices 185 and / or associated camera devices 184, and for generating and providing a representation of the interior of a building or other structure. For example, in at least some such embodiments, such techniques may include using one or more mobile devices (e.g., a camera having one or more fisheye lenses and mounted on a rotatable tripod or having an automatic rotation mechanism, a camera having a fisheye lens that is sufficiently horizontal to capture 360° without rotation, a smartphone held and moved by a user, a camera held by a user or mounted on a user or on a user's clothing, etc.) to capture data from a sequence of multiple acquisition locations within multiple rooms of a house (or other building), and optionally further capture movement involving the acquisition device (e.g., movement between acquisition locations, such as rotation; movement between some or all acquisition locations, such as for linking multiple acquisition locations together, etc.), in at least some cases without measuring distances between acquisition locations or having other measured depth information from objects in the environment surrounding the acquisition locations (e.g., without using any depth sensing sensors). After information about the acquisition location is captured, the technology may include generating a 360° panoramic image from the acquisition location having 360° horizontal information about a vertical axis (e.g., a 360° panoramic image showing the surrounding room in an equirectangular format), and then providing the panoramic image for subsequent use by the MIGM and / or BVGUM system.
[0034] In addition, although Figure 1A, but a floor plan (or portion thereof) may be linked to or otherwise associated with one or more additional types of information, such as one or more associated and linked images or other associated and linked information, including a two-dimensional ("2D") floor plan for a building being linked to or otherwise associated with a separate 2.5D model floor plan rendering of the building and / or a 3D model floor plan rendering of the building, etc., and including a floor plan for a multi-story or other multi-story building having multiple associated sub-floor floor plans for different floors or levels that are interconnected (e.g., via connecting stairways), or are portions of a common 2.5D and / or 3D model. Thus, non-exclusive examples of end-user interaction with a displayed or otherwise generated 2D floor plan of a building may include one or more of the following: changing between a view of a particular image and a floor plan view at an acquisition location within or near the floor plan; changing between a 2D floor plan view and a 2.5D or 3D model view, the 2.5D or 3D model view optionally including an image of a wall mapped to the displayed model; changing the horizontal and / or vertical viewing direction of a corresponding subset view of a displayed panoramic image (or an entrance into the panoramic image) in order to determine the portion of the panoramic image in the 3D coordinate system to which the current user viewing direction points, and rendering a corresponding plan image showing the portion of the panoramic image without curvature or other deformation present in the original panoramic image, etc. Additionally, although in Figure 1A Not shown, but in some embodiments, client computing device 175 (or other device, not shown) may receive and use the generated floor plan and / or other generated mapping-related information in additional ways to control or assist in the automated navigation activities of those devices (e.g., by an autonomous vehicle or other device), whether to display the generated information, or in addition to displaying the generated information.
[0035] Figure 1A Also described are exemplary building interior environments in which 360° panoramic images and / or other images are acquired, for example, by an ICA system, and used by a MIGM system (e.g., under control of a BVGUM system) to generate and provide one or more corresponding building floor plans (e.g., multiple incremental partial building floor plans), and / or such building information is used by a BVGUM system as part of an automatic building information generation operation. Specifically, Figure 1AOne floor of a multi-story house (or other building) 198 is shown having an interior that is captured at least in part by a plurality of panoramic images, such as by a mobile image capture device 185 having image capture capabilities and / or one or more associated camera devices 184 as they move through the interior of the building to a sequence of a plurality of capture locations 210 (e.g., starting at capture location 210A, moving along a travel path 115, etc. to acquisition location 210B, and ending at acquisition location 210-O or 210P on the exterior of the building). An embodiment of an ICA system may automatically perform or assist in capturing data representing the interior of the building (and further analyzing the captured data to generate a 360° panoramic image to provide a visual representation of the interior of the building), and an embodiment of a MIGM system may analyze the visual data of the captured images to generate one or more building floor plans (e.g., a plurality of incremental building floor plans) for the house 198. Although such a mobile image acquisition device may include various hardware components, such as a camera, one or more sensors (e.g., gyroscopes, accelerometers, compasses, etc., such as part of one or more IMUs of a mobile device, or inertial measurement units; altimeters; light detectors, etc.), a GPS receiver, one or more hardware processors, memory, a display, a microphone, etc., in at least some embodiments, the mobile device may not have access to or use a device to measure the depth of objects in a building relative to the position of the mobile device, such that in such embodiments, the relationship between different panoramic images and their acquisition locations may be determined based in part or in whole on elements in the different images, but without using any data from any such depth sensor, while in other embodiments, such depth data may be used. In addition, although in Figure 1A A direction indicator 109 is provided for indicating the reader relative to the exemplary house 198, but in at least some embodiments, the mobile device and / or the ICA system may not use such absolute direction information and / or absolute position, such that in such embodiments, the relative direction and distance between the acquisition locations 210 are determined without regard to the actual geographic location or direction, while in other embodiments, such absolute direction information and / or absolute position may be obtained and used.
[0036] In operation, the mobile device 185 and / or one or more camera devices 184 arrive at a first acquisition location 210A within a first room within a building interior (in this example, in a living room accessible via an exterior door 190-1) and capture or acquire a view of a portion of the building interior visible from the acquisition location 210A (e.g., some or all of the first room, and optionally, small portions of one or more other adjacent or neighboring rooms, such as through a doorway wall opening, a non-doorway wall opening, a hallway, a stairway, or other connecting passageway from the first room). The view capture may be performed in various ways as discussed herein, and may include a plurality of structural elements or other objects visible in the image captured from the acquisition location, in which the view capture is performed. Figure 1A In the example of FIG. 1 , such objects within a building 198 include walls, floors, ceilings, doorways 190 (including 190-1 through 190-6, such as having swinging and / or sliding doors), windows 196 (including 196-1 through 196-8), boundaries between walls and other walls / ceilings / floors, such as for corners or edges 195 between walls (including corner 195-1 in the northwest corner of the building 198, corner 195-2 in the northeast corner of the first room, corner 195-3 in the southwest corner of the first room, corner 195-4 in the southeast corner of the first room, corner 195-5 at the north edge of a room passage between the first room and the hallway, furniture 191-193 (e.g., recliner 191; chair 192; table 193, etc.), pictures or paintings hung on the walls or televisions or other hanging objects 194 (e.g., 194-1 and 194-2), light fixtures ( Figure 1A Not shown), various built-in appliances or other fixtures or other structural elements ( Figure 1A ), etc. The user may also optionally provide a text or auditory tag identifier to be associated with the acquisition location and / or surrounding rooms, such as "Living Room" for one of the acquisition locations 210A or 210B or for the room including the acquisition locations 210A and / or 210B, and / or a descriptive annotation with one or more phrases or sentences about the room and / or one or more objects in the room, while in other embodiments, the ICA and / or MIGM system may automatically generate such identifiers and / or annotations (e.g., by automatically analyzing images and / or videos and / or other recorded information for the building to perform corresponding automatic determinations, such as by using machine learning; based at least in part on input from ICA and / or MIGM system operator users, etc.) or may not use an identifier.
[0037] After the first acquisition location 210A has been captured, the movable mobile device 185 and / or one or more camera devices 184 or the mobile mobile device 185 and / or one or more camera devices 184 may move to the next acquisition location (such as acquisition location 210B) under their own power, optionally recording images and / or video and / or other data from hardware components (e.g., from one or more IMUs, from cameras, etc.) during the movement between acquisition locations. At the next acquisition location, the mobile device 185 and / or one or more camera devices 184 may similarly capture 360° panoramic images and / or other types of images from the acquisition location. The process may be repeated for some or all rooms of the building and in some cases outside the building, as shown in this example for additional acquisition locations 210C-210P, in which images from acquisition locations 210A to 210-O are acquired in a single image acquisition session (e.g., in a substantially continuous manner, such as within a total of 5 minutes or 15 minutes), and images from acquisition location 210P are optionally acquired at a different time (e.g., so as to be adjacent to the street of the building or the building's foreground). In this example, multiple acquisition locations 210K-210P are located outside of, but associated with, a building 198 on surrounding real estate 241, including acquisition locations 210L and 210M in one or more additional structures 189 on the same real estate (e.g., an ADU or accessory dwelling unit; a garage; a shed, etc.), acquisition location 210K on an exterior deck or patio 186, and acquisition locations 210N-210P at multiple yard locations on real estate 241 (e.g., a back yard 187, a side yard 188, a front yard including acquisition location 210P, etc.). The acquired images at each acquisition location may be further analyzed, including in some embodiments presenting or otherwise placing each panoramic image in an equirectangular format, either at the time of image acquisition or later, and further analyzed by the MIgm and / or BVGUM system in the manner described herein.
[0038] Figure 1B 1 shows an example of the type of building description information 110b that may be used in some embodiments, such as existing building information that is then analyzed and used by the BVGUM system. Figure 1BIn the example of , the building description information 110b includes an overview text description, and various property data, such as may be used in part or in whole as listing information for an MLS system. In this example, the property data is grouped into sections (e.g., overview properties, further interior detail properties, further property detail properties, etc.), but in other embodiments, the property data may not be grouped or may be grouped in other ways, or more generally, the building description information may not be separated into a property list and a separate text overview description. In this example, the separate text overview description emphasizes features that may be of interest to the viewer, such as the house style type, interesting information about the rooms and other building features (e.g., has been recently updated or has other interesting features), interesting information about the property and surrounding or other environment, etc. Additionally, in this example, the attribute data includes various types of objective attributes about the rooms and buildings and limited information about facilities, but various types of details shown in italics (e.g., about subjective attributes, about connectivity and other adjacencies between rooms, about other specific structural elements or objects and about attributes of such objects, etc.) may be missing in this example, as may be determined by the BVGUM system by analyzing building images and / or other building information (e.g., floor plans).
[0039] about Figure 1A and Figure 1B Various details are provided, but it is to be understood that the details provided are non-exclusive examples included for purposes of illustration and that other implementations may be performed in other ways without some or all of such details.
[0040] Figures 2A to 2V An example of automatically generating a video using information about a building from automatic analysis of building images and other building information for subsequent use in one or more automated ways (eg, for building 198) is shown.
[0041] In particular, Figure 2A An exemplary image 250a is shown, for example, in Figure 1A1, a non-panoramic perspective image taken in a northeast direction starting from acquisition location 210B in a living room of a house 198 (or a northeast-facing subset view of a 360° panoramic image starting from the acquisition location and formatted in a rectilinear manner), in which example a direction indicator 109a is further displayed to show that the image is taken in a direction in the northeast direction. In the example shown, the displayed image includes built-in elements (e.g., light fixture 130a, two windows 196-1, etc.), furniture (e.g., chair 192-1), and a picture 194-1 hanging on the north wall of the living room. No inter-room passages (e.g., doorways or other wall openings) entering or leaving the living room are visible in the image. However, multiple room boundaries are visible in image 250a, including horizontal wall-ceiling and wall-floor boundaries between the visible portion of the living room's north wall and the living room's ceiling and floor, horizontal wall-ceiling and wall-floor boundaries between the visible portion of the living room's east wall and the living room's ceiling and floor, and vertical wall-to-wall boundary 195-2 between the north wall and the east wall.
[0042] Figure 2B Continue to show Figure 2A of examples, and shows that Figure 1A 198 in the living room of the house 198, further displaying the direction indicator 109b to show the northwest direction in which the image was taken. In this example image, a small portion of one of the windows 196-1 and a portion of the window 196-2 and the new lighting fixture 130b continue to be visible. In addition, in the image 250b, a portion similar to Figure 2A Horizontal and vertical room boundaries are visible in a .
[0043] Figure 2C Continue to show FIG. 2A to FIG. 2B of examples, and shows that Figure 1A The third perspective image 250c captured in a southwest direction in the living room of the house 198, such as the third perspective image 250c captured from the acquisition location 210B, further displays the direction indicator 109c to show the southwest direction from which the image was captured. In this example image, a portion of the window 196-2 continues to be visible, as does the recliner 191 and the visual horizontal and vertical room boundaries in a manner similar to Figure 2A and Figure 2B The example image also shows two inter-room passages for the living room, which in this example include doorways 190-1 ( Figure 1A Identified as a door to the exterior of the house, such as a front yard), and a doorway 190-6 having a sliding door for moving between the living room and the side yard 188, such as Figure 1AAs shown in the information in, additional non-doorway wall opening 263a exists in the east wall of the living room to move between the living room and the hallway, but is not visible in images 250a-250c. It should be understood that various other perspective images can be acquired from acquisition position 210B and / or other acquisition positions and displayed in a similar manner.
[0044] Figure 2D Continue to show FIG. 2A to FIG. 2C , and shows a 360° panoramic image 255d (eg, acquired from acquisition location 210B) that displays the entire living room in an equirectangular format. Since the panoramic image does not have the same FIG. 2A to FIG. 2C The same direction as the perspective image of Figure 2D No direction indicator 109 is displayed in the panoramic image, although the pose information of the panoramic image may include one or more associated directions (e.g., the starting and / or ending directions of the panoramic image, such as if acquired by rotation). Portions of the visual data of panoramic image 255d correspond to first perspective image 250a (shown approximately in the center portion of image 250d), while the left portion of image 255d and the far right portion of image 255d contain visual data corresponding to those of perspective images 250b and 250c, so that, for example, starting with image 255d, a series of perspective images may be presented (e.g., for use in a video) that includes some or all of images 250a-250c (and optionally a large number of intermediate images so that if the resulting video uses 30 frames per second, the presented perspective images correspond to 5 seconds of pan and / or tilt within image 255d). This exemplary panoramic image 255d includes windows 196-1, 196-2, and 196-3, furniture 191-193, doorways 190-1 and 190-6, and a non-doorway wall opening 263a leading to a hallway room (which shows a portion of doorway 190-3 visible in an adjacent hallway). Image 255d also shows various room boundaries in a manner similar to the perspective image, but with the horizontal boundaries being displayed in an increasingly curved manner as one moves farther from the horizontal centerline of the image, with visible boundaries including vertical wall-to-wall boundaries 195-1 to 195-4, a vertical boundary 195-5 on the left / north side of the hallway opening, a vertical boundary on the south / right side of the hallway opening, and horizontal boundaries between walls and floors and between walls and ceilings.
[0045] Figure 2DAlso shown is information including one example 230d of a portion of a 2D floor plan for house 198 (e.g., corresponding to floor 1 of the house), such as may be presented to an end user in GUI 260d, wherein the living room is the westernmost room of the house (as reflected by directional indicator 209), it being understood that in some embodiments, a 3D or 2.5D floor plan with wall height information presented may be similarly generated and displayed, either in addition to or in lieu of such a 2D floor plan. In this example, various types of information are shown on 2D floor plan 230d. For example, this type of information may include one or more of the following: room labels added to some or all rooms (e.g., "Living Room" for a living room); room dimensions added for some or all rooms; visual indications of objects, such as installed fixtures or appliances (e.g., kitchen appliances, bathroom items, etc.) or other built-in elements (e.g., kitchen island) added for some or all rooms, optionally with associated labels and / or descriptive annotations (e.g., double steel kitchen sink, kitchen island with red Coria surface, LED track lighting, white tile flooring, etc.); visual indications added to some or all rooms of the location of additional types of associated and linked information (e.g., additional panoramic images and / or perspective images selectable by the end user for further display; audio or non-audio annotations selectable by the end user for further presentation, such as Such as "Kitchen includes Brand X refrigerator with feature Y, Brand Z built-in stove / oven, etc."; sound recording that the end user can select for further presentation to listen to the street noise level from bedroom 1, etc.); visual indications added to some or all rooms of structural elements such as doors and windows; visual indications of visual appearance information (e.g., color and / or material type and / or texture of installed items, such as floor overlays or wall overlays or surface overlays); views from particular windows or other building locations and / or visual indications of other information about the exterior of the building (e.g., type of exterior space; items present in the exterior space; other related buildings or structures such as sheds, garages, pools, decks, patios, walkways, gardens, etc.); keywords or legends 269 identifying visual indicators for one or more types of information, etc. When displayed as part of a GUI such as 260d, some or all of such displayed information may be user-selectable controls (or associated with such controls) allowing an end-user to select and display some or all of the associated information (e.g., selecting the 360° panoramic image indicator for acquisition location 210B to view some or all of the panoramic image (e.g., in a manner similar to FIG. 2A to FIG. 2DIn addition, in this example, a user-selectable control 228 is added to indicate the current floor displayed for the floor plan and allow the end user to select a different floor to be displayed, and in some embodiments, changes to floors or other levels can also be made directly from the floor plan, such as via selection of corresponding connecting pathways in the shown floor plan (e.g., stairs to floor 2). It should be understood that various other types of information may be added in some embodiments, some of the types of information shown may not be provided in some embodiments, and visual indications and user selections of linked and associated information may be displayed and selected in other ways in other embodiments.
[0046] Figure 2E and Figure 2F Continue to show FIG. 2A to FIG. 2D An example of Figure 2E Information 255e is shown, which includes an image 250e1 of the southwest portion of the living room (in a similar manner to Figure 2C Information 255e also shows a list 248p of objects and additional attributes of interest that were identified based at least in part on the visual data of image 250e1, indicating that visual characteristics of the west window include its type (e.g., a picture window), the type of latch hardware, information about the view through the window, and optionally various other attributes (e.g., size, the orientation / direction it faces, etc.). Image 250e1 also indicates that doorway 190-1 has been identified as an object in the room, with a "front door" label 246p1 (whether automatically determined or based at least in part on information provided by one or more associated users) and an automatically determined bounding box location 199a shown. Additionally, information 248p indicates that further attributes include determined visual characteristics of the door, such as the type of door and information about the door's door knobs and door hinges, which are further visually indicated 131p on image 250p. Figure 2F Shows that the Figure 2EAdditional visual data extracted from images 250e1 and / or 250e2 as part of determining objects and other attributes of the room, and specifically including close-up example images 250f1, 250f2, and 250f3 corresponding to the front door of doorway 190-1 and its hardware 131p, such as for determining the corresponding attributes of the front door. In some embodiments and scenarios, if image 250e1 is selected for use in a video of house 198 and it is determined that there is interest in the front door and its visual characteristics, then visual data of image 250e1 (or another image in the living room) may be selected for display, which includes such images 250f1, 250f2, and 250f3 (e.g., via zoom, pan, tilt, etc.) to highlight those objects and other attributes in the video, while corresponding narration provides a description of those objects and other attributes. Other objects, such as one or more ceiling light fixtures, furniture, walls and other surfaces, etc., may be similarly identified (e.g., based at least in part on a list of defined types of objects that are expected or typical for a room of type "living room"), and optionally described via associated narration in the generated video, and the visual data of the selected one or more corresponding images showing these objects. Optionally, in a manner that highlights or emphasizes them (e.g., by zooming in to highlight the object). Additionally, a "living room" tag 246p3 for the room is also determined and shown (automatically or based at least in part on information provided by one or more relevant users). Figure 2E An alternative or additional image 250e2 is also provided, which in this example is a panoramic image with a 360° visual overlay of the living room (in a similar manner to Figure 2D Such panoramic images may be used in place of or in addition to perspective images such as image 250e1 to determine objects and other attributes and additional related information (e.g., location, labels, annotations, etc.), and to assess the overall layout of items in a room and / or the expected traffic flow of a room, wherein exemplary panoramic image 250e2 similarly shows position bounding boxes 199a and 199b for the front door and west window objects, and an additional position bounding box 199c for tables 193, 199d for ceiling lights 130b on the east wall, it being understood that in other embodiments a variety of other types of objects and / or other attributes may be determined, including other walls and surfaces (e.g., ceilings and floors) and other structural elements (e.g., windows 196-1 and 193, doorways 190-6, non-doorway wall openings 263a, etc.), other furniture (e.g., bed 191, chair 192, etc.), and the like.
[0047] Figure 2G Continue to show FIG. 2A to FIG. 2Fand provide examples of additional data that can be used to determine objects and other properties about other rooms based at least in part on analysis of one or more initial room-level images of other rooms of a building. Specifically, Figure 2G Information 255g including an image 250g1 is shown, for example for bathroom 1. Figure 2E In the manner of an image of a bathroom, image 250g1 includes indications 131v of objects in a bathroom that are identified and for which corresponding attribute data (e.g., corresponding to visual characteristics of the objects) is determined, in this example, the objects include a tile floor, a sink countertop, a sink faucet and / or other sink hardware, a bathtub faucet and / or other bathtub hardware, a toilet, etc., however, the location information, labels, and instructions provided are not shown in this example. In a similar manner, image 250g2 of a kitchen includes indications 131w of objects that are identified in the kitchen and for which corresponding attribute data (e.g., corresponding to visual characteristics of the objects) is determined, in this example, the objects include a refrigerator, a stove on a kitchen island, a sink faucet and / or other sink hardware, a countertop and / or a backsplash next to the sink, etc., however, the location information, labels, and instructions provided are not shown in this example. It should be understood that various other types of objects and other attributes may be determined in these and other rooms and further used to generate corresponding building description information, and in the manner of an image of a bathroom, image 250g1 includes indications 131v of objects that are identified and for which corresponding attribute data (e.g., corresponding to visual characteristics of the objects) is determined, in this example, the objects include a refrigerator, a stove on a kitchen island, a sink faucet and / or other sink hardware, a countertop and / or a backsplash next to the sink, etc., however, the location information, labels, and instructions provided are not shown in this example. Figure 2E to Figure 2G These types of data shown are non-exclusive examples and are used for illustration purposes.
[0048] Figure 2H to Figure 2K Continue to show Figures 2A to 2G and provide additional information related to analyzing the floor plan information to determine additional properties of the building. Specifically, Figure 2H Information 290h is shown including an example 2D floor plan 230h of a building, including information 222h determined regarding expected movement flow pattern properties through the building, as indicated using corresponding labels 221h, and such information 222h is optionally displayed on the floor plan (e.g., overlaid on the floor plan). In a similar manner, Fig.2I Additional information 290i is provided in connection with analyzing the floor plan 230i to determine information regarding various types of subjective attributes of the building (e.g., wheelchair accessibility, accessibility for persons with limited walking mobility, open floor plan, typical layout, modern style, etc.), as indicated using corresponding markers 221i, but in this example no corresponding locations or other indications are shown on the floor plan, although such corresponding locations may be determined and indicated in other embodiments and situations. Figure 2JAdditional information 290j is similarly provided in connection with analyzing the floor plan 230j to determine information regarding areas of the building corresponding to public and private space attributes 222j, as indicated using corresponding labels 221j, and such information 222j is optionally displayed on (e.g., overlaid on) the floor plan. Figure 2K Additional information 290k is provided in connection with analyzing the floor plan 230k to determine information regarding room types and / or functional attributes 222k (e.g., bedroom, bathroom, kitchen, dining, family room, wall cabinet, etc.) of the building, as indicated using corresponding labels 221k, and such information 222k is optionally displayed on (e.g., overlaid on) the floor plan. It should be understood that specific attributes regarding the rooms and / or the building as a whole may be determined by analyzing such floor plans in a variety of ways, and Figure 2H to Figure 2K The types of information shown in are non-exclusive examples provided for purposes of illustration, such that similar and / or other types of information may be determined in other manners in other embodiments.
[0049] Figure 2L Continue to show Figures 2A to 2K , and information 2901 corresponding to building 198 is shown. In particular, while path 115 shows a sequence of images acquired at locations 210A through 210-O of a building, the BVGUM system may determine a different sequence of images 2251 whose visual data will be included in the video to be generated in a corresponding order, optionally for a selected subset of the acquired images. In this example, a sequence of panoramic images acquired at acquisition locations 210A, 210C, 210G, 210J, 210K, etc. is selected to correspond to starting at the entrance of the house, but quickly proceeding to illustrate information about the kitchen (e.g., because the kitchen is determined to be typically of high interest and therefore of high interest to a particular recipient for whom the video is being generated, to correspond to a configuration setting or other instruction regarding video generation based on the order in which non-panoramic images and / or other information in the rooms of the building are acquired by one or more users, etc.). While in Figure 2L 198, so as to display other types of information (e.g., one or more specific rooms, such as bathrooms and / or bedrooms; one or more groups of rooms, such as corresponding to different floors or public or private spaces, etc.). In at least some embodiments, one or more trained machine learning models (e.g., one or more neural network models) can be used to determine a set of one or more images selected for use in a video and / or a determined sequence of multiple such selected images.
[0050] Figure 2M and Figure 2N Continue to show Figures 2A to 2L An example of and information about selecting visual data of an image to include in the visual portion of a video being generated. Specifically, Figure 2M Image 255m is shown in a manner similar to image 255d, but with various objects indicated as being annotated in the video, including selecting visual data of image 255m to be included in the video displaying those objects. In this example, the selected objects include a vaulted ceiling as indicated at 299f, track lighting 130b as indicated at 299d, a front door 190-1 as indicated at 299a, a west picture window 196-2 as indicated at 299b, a south window 196-3 as indicated at 299g, a sliding door 190-6 as indicated at 299h, an east wall as indicated at 299e, a table 193 as indicated at 299c, etc., for example, and the flow of arrows between those selected objects may indicate that corresponding subsets of the visual data of image 255m are displayed successively in the video to highlight the order in which those objects are displayed (e.g., in the form of a DAG or directed acyclic graph). Figure 2M Further shown is exemplary text information 265m, which may be included as narration within the video regarding the room and the selected object (e.g., audible narration in the audio portion of the video; text narration, such as for closed captioning, etc.) to describe properties corresponding to objects such as a vaulted ceiling, track lighting 130b, front door 190-1 (e.g., door type, door hardware, description of location to which the door leads, etc.), west picture window 196-2 (e.g., window type, hardware, view, orientation, etc.), south window 196-3, sliding door 190-6, east wall (e.g., color, type of surface material, etc.), table 193 (e.g., material, size, etc.), etc. Figure 2N Also shown is a sequence 242 of visual data subsets (e.g., video frames) rendered from a panoramic image 255m that may be included in the video to be generated to correspond to pans and / or tilts within the panoramic image. In at least some embodiments, descriptions of objects and other attributes are generated using one or more trained language models.
[0051] As a non-exclusive example, the selected subset of visual data from the selected image may be from a single viewpoint and one or more viewing angles (e.g., the acquisition position of the image and in a determined direction with a determined level of zero or more zooms, e.g., each frame of the video corresponds to a perspective image subset of the selected image corresponding to the viewing angle), whether continuous or non-contiguous. If multiple images are selected for use in the video (e.g., in a determined order), the visual data selected from those multiple images may correspond to multiple viewpoints and one or more viewing angles for each such viewpoint (e.g., visual data is selected from the acquisition position of each image and in one or more viewing angles, e.g., each frame of the video corresponds to a perspective image portion of the selected image, and one or more perspective image portions for each selected image), whether continuous or non-contiguous. With respect to identifying objects to be depicted in a video, the BVGUM system may, in at least some embodiments and circumstances, create a graph (e.g., a GCN, or graph convolutional network, which is used to learn a DAG having edges with direction and order) that includes information about one or more of: which objects and / or other attributes to be depicted, and optionally for how long; a sequence of objects and / or other attributes depicted within an image; a determined sequence of use of multiple images; cinematic transitions or other types of transitions between visual data of adjacent images in the determined sequence, and the like.
[0052] With respect to generating text descriptions for video narration, the BVGUM system may use one or more trained language models for visual floor speech, image captioning, and text image retrieval in at least some embodiments and situations, where text generated from an image or image sequence is combined and / or aggregated in at least some embodiments and situations (e.g., to manipulate style, grammar, and / or modality in order to provide a rich and effective recipient experience; generate multiple generated texts, etc.), whether using multiple discrete models or a single end-to-end model. In at least some embodiments and situations, the one or more trained language models may include one or more trained vision and language (VLM, Vision and Language model) models, such as a large model trained to generate descriptions / captions for input images using a large corpus of training tuples (e.g., image, caption tuples), some benefits of VLM models include not needing to explicitly prompt the model about the entity you want to describe, which often results in more abstract and powerful descriptions. In at least some embodiments and scenarios, the one or more trained language models may also include at least one of a pre-trained language model, a knowledge-enhanced language model, a parsing and / or tagging and / or classification model (e.g., a dependency parser, a group parser, a sentiment classifier, a semantic role tagger, etc.), an algorithm for controlling language quality (e.g., a tagger, a word extractor, a regular expression matcher), and a multimodal vision and language model that can automatically regress or mask decoding, such tagging and / or classification models may include, for example, a semantic role tagger, a sentiment classifier, and a semantic classifier to identify semantic concepts associated with the entire sequence of words and tokens or any components thereof (e.g., identifying the semantic role of an entity in the sequence, such as a patient or agent, and classifying the overall sentiment or semantics of the sequence, such as how positive or negative it is about the subject, how fluent the sequence is, or how well it encourages the reader to take a certain action). The one or more trained language models may, for example, perform iterative generation (decoding) of words, subwords, and tokens based on representations of prompts, prefixes, control codes, and contextual information such as features derived from visual / sensor information, knowledge bases, and / or graphics, and the parsing model may further perform operations including analyzing the internal structure of word and token sequences to identify their components according to one or more grammars (e.g., dependencies, context-free grammars, head-driven phrase structure grammars, etc.) in order to identify modifications that can be made to word, subword, and / or token sequences to further develop desired language qualities.For example, the one or more trained language models may be organized into a directed acyclic graph providing a structure of inputs, outputs, data sources, and model interactions, where the structure is aligned with the data sources, for example with respect to one or more of: space, where the context of text generation is related to a specific point in a building, such as a location or room where a panorama was taken, so that the generated text will be aligned with that location; time, where the context of text generation is a temporal sequence of frames in a video sequence or slide show, so that the generated text will be aligned with that sequence of frames, etc. Inputs to the one or more trained language models may include, for example, one or more of the following: structured and / or unstructured data sources (e.g., publicly or privately available, such as real estate records, tax records, MLS records, Wikipedia articles, homeowners association and / or colleague documents, news articles, nearby or visible landmarks, etc.) that provide information about the building being analyzed and / or associated physical space and its environment, and / or provide general and common sense information about buildings and the real estate market (e.g., homes and related elements, overall housing market information, information related to fair housing practices, biases related to terms and phrases that assist in language generation, etc.); information about objects and / or other attributes (e.g., fixture type and location, surface material, surface color, surface texture, room dimensions, degree of natural light present in a room, walkability score, expected commute time, etc.); captured and / or synthesized visual and / or sensor information and any derived information, such as structured and unstructured sequences (including single items) of images, panoramas, videos, depth maps, point clouds, and segmentation maps, etc. The one or more trained language models may also be designed and / or configured to, for example, implement one or more of the following: modality, to reflect the way language is able to express relationships to reality and truth (e.g., things that are forbidden, such as "You shouldn't go to school"; advice provided through subject-assisted reversal, such as "Don't go to school?", etc.); fluency, to reflect measures of the natural quality of language with respect to a set of grammatical rules (e.g., "big stinky brown dog" rather than "stinky big brown dog"); style, to reflect patterns of word and grammatical structure selection (e.g., short descriptions using interesting and restrained language; informal style, such as for sending text messages; formal style, such as for English papers or conference submissions; voice, to reflect the way subjects and objects are organized relative to verbs (e.g., active and passive voice), etc.
[0053] As generated Figure 2MAs a non-exclusive example of text 265m, the BVGUM system may use a pipeline of language models and algorithms, including from what are called "knowledge enhanced natural language generation models" (KENLG). The KENLG model is provided with a knowledge source and hints about entities within the knowledge source that the model should attend to. The KENLG model absorbs these inputs and generates descriptions of these entities accordingly. The KENLG model may focus on several knowledge sources, including a knowledge base (KB), a knowledge graph (KG), and unstructured text, such as a Wikipedia page. The knowledge base represents predicate relationships between entities (e.g., Living Room, HasA, Fireplace; Fireplace, MadeOf, Stucco; Kitchen Counter, HasProperty, Spacious, etc.). The knowledge graph is a translation of the knowledge base into a graph structure, in which the nodes of the graph represent entities, while the edges of the graph represent predicate relationships, and the hints are sequences of entities for which descriptions are to be generated. For example, using the example tuple above, if the model was given the prompt (living room), the expected output of the model would be "The living room has a lovely stucco wall fireplace." In this way, information about houses or other buildings generated by any number of one or more upstream feature extraction models and data collection processes can be aggregated into a knowledge base and then into a knowledge graph in which the KENLG model can participate. The benefits of representing buildings in terms of a knowledge graph in this manner include representing buildings in a natural way, as the combination of interrelated spaces, viewpoints, and objects with attributes (when combined with the KENLG model) can be used to generate a large number of descriptions based on the provided prompts. For a multi-sentence description such as 265m, its generation may include (in part) determining a prompt sequence to provide the model with prompts for generating its constituent clauses based on those entities, and then using the KENLG model to generate text for those prompts. For example, a prompt sequence may be identified as a sequence of observation points (e.g., Figure 2M 299a-299h), and / or information about the layout of rooms and buildings, such as Figure 2-O 265o further discussed in the text. After generating a sequence of clauses corresponding to the prompt sequence (e.g., as derived from the points of views 299a-299h), one or more types of algorithms may be applied to perform language modifications to compose a convincing narrative, including constructing prepositional phrases to smoothly connect the constituent clauses (e.g., "When you enter the living room..."), to making changes to the style and modality of each clause (e.g., to express competence, such as "You will notice that..."; to express obligation, such as "You must know that...", etc.).
[0054] With respect to selecting objects and attributes to be discussed in the video being generated, such as using visual data of images showing those objects or other attributes, objects and attributes may be selected in various ways in various embodiments. As a non-exclusive example, a set of objects and other attributes may be predefined, such as based on input from a user (e.g., a visit to a building or to some or all of a building, such as a lease or rental; based on tracked activity of a user in requesting or viewing information about a building (e.g., viewing images and / or other information about a building); based on tracked activity of a user in purchasing or changing objects within a building, such as during a remodeling period, etc.), and information about such a set of objects and other attributes may be stored in various ways (e.g., in a database) and may be used to train one or more models (e.g., one or more machine learning models for selecting images and / or portions of images, one or more language models for generating text descriptions, etc.). In other embodiments and scenarios, the set of objects and other attributes may be determined in other ways, whether in lieu of or in addition to such predefinition, such as to be learned (e.g., based at least in part on analyzing professional photographs or other images of buildings to identify objects that are the focus of or otherwise included in such images).
[0055] Figure 2N Also shown is information about examples of synchronizing a generated narration of a video (e.g., in an audio portion of the video) with corresponding visual data in a visual portion of the video, e.g., causing the narration to occur simultaneously with the corresponding visual data or otherwise accompany the visual data (e.g., introducing the visual data prior to display). As previously described, a sequence 242 of visual data subsets (e.g., video frames) that may be included in a video to be generated is shown such that image 250b is displayed when the accompanying narration occurs around the vaulted ceiling (e.g., at approximately 4.2 seconds in the narration), image 250b and / or 250n1 are displayed when the accompanying narration occurs around the built-in track lighting (e.g., at approximately 4.2 seconds in the narration) (at around time 8.3 seconds to zoom in on the lighting in one or more images 250n1), panning counterclockwise toward the front door in one or more intermediate images 250n2 (e.g., between approximately 10 and 15 seconds), image 250n3 is displayed when the accompanying narration occurs near the front door (e.g., at around time 15.6 seconds, such as zooming in on some or all of the doors), image 250n4 is displayed when the accompanying narration occurs around the west image windows (e.g., at approximately 20.4 seconds, such as zooming in on some or all of the doors, etc.).
[0056] With respect to the narration generated for video synchronization, synchronization may be performed in various ways in various embodiments. As a non-exclusive example, as visual data is shown in the video, the narration may be presented for one or more objects and other attributes shown in the visual data. In other embodiments and scenarios, additional activities may be performed to generate a smoothly flowing narration over time, such as optimizing a combination of visual continuity and smooth changes in narrative themes (e.g., based at least in part on learning from a set of narrative home tour videos).
[0057] FIG. 2O (referred to herein as “2-O” for clarity) further illustrates information regarding an example of including visual data in a video corresponding to a transition between visual data of two selected images (in this example, image 255m acquired at acquisition location 210A, and an additional image acquired at acquisition location 210C, not shown). Figure 2-O In an example, the added visual data for the transition may include zooming from the visual data shown in frame 250o1 to the end of the visual data shown in frame 250o2, for example using one or more intermediate frames (not shown) to correspond to the zoom shown by arrow 227 (the arrow is included for the benefit of the reader, but is not shown in the generated video in at least some embodiments and cases), and then further transitioning (not shown) to the additional image acquired at acquisition location 210C, such as using one or more intermediate frames (not shown) to blend the visual data of frame 250o2 into a frame that includes only visual data from the additional image acquired at acquisition location 210C. Figure 2-O Further shown is an example of additional narration 265o (e.g., to be included in the audio portion of the video) synchronized with the visual data of one or more transitions in the video, such as being automatically generated to describe the transition between the acquisition positions of the image. In other embodiments, narration 265o may be included in the video without any additional visual data corresponding to the transition (e.g., displaying frame 250o1 followed by frame 250o2 of the same size), or visual data of one or more transitions may be displayed in the visual portion of the video without any accompanying narration.
[0058] With respect to the order in which the sequence of images to be included in the video is selected, the order may be determined in various ways in various embodiments. As a non-exclusive example, some or all of the images acquired in the building may be selected for the sequence in an order corresponding to the order in which the images were acquired, such as Figure 1A Some or all of the images along the path 115 shown may be determined based at least in part on timestamps, overlapping visual data, IMU data, etc. As another non-exclusive example, a different order 2251 may be determined, such as in terms of Figure 2LIn situations where different images are acquired at substantially different times (e.g., during multiple image acquisition sessions), images may be selected for some or all rooms on a room-by-room basis (e.g., after the images are located within the rooms, such as may be reflected on a floor plan or other building model), rather than or in addition to the order in which the images were acquired and / or using a different order as described above, the order of the rooms may be selected, for example, in the order in which the images were acquired in the rooms and / or in a different manner (e.g., as described with respect to FIG. Figure 2L When selecting multiple images within a room, in other embodiments the images may be selected in an order corresponding to, for example, one or more of the following: an order of priority of the types of features or other information shown in the images; to minimize visual jumps between images within the room (and improve visual transitions); etc.
[0059] Figure 2P Continue to show Figure 2A to Figure 2-O, and information 290P is shown, which shows an example data flow interaction for at least some automated operations of one embodiment of a BVGUM system (referred to as a BVGUM-A system in this example). Specifically, an embodiment of the BVGUM system 140 is shown as being executed on one or more computing systems 180, and in this example embodiment, information about a building to be analyzed is received, the information including stored images from a memory or database 295, a floor plan from a memory or database 296, and optionally other building description information 297p (e.g., a list) and / or other information 298 (e.g., other building information), such as tags, annotations, etc.; other configuration settings or instructions for controlling the generation of the video, such as length, type of information to be included, etc., information about the intended recipient of the video to be generated, such as for personalizing the generated video to the recipient, etc.). Input information is received in step 281p, wherein the received image information is optionally forwarded to a BVGUM image analyzer component 282 for analysis (e.g., identifying objects and optionally other attributes, such as local attributes specific to a particular image and / or room), the received floor plan information is optionally forwarded to a BVGUM floor plan analyzer component 283 for analysis (e.g., determining other building attributes, such as overall attributes corresponding to some or all of the building as a whole), and other received building information is optionally forwarded to a BVGUM other information analyzer component 284 for analysis (e.g., for determining other building attributes, such as based on a textual description of the building), and the output of components 282 and / or 283 and / or 284 form some or all of the determined building attributes 274 of the building, in other embodiments, information about some or all of such objects and / or other attributes may be received (e.g., as part of other information 298) and included directly in the building attribute information 274, and as discussed in more detail elsewhere herein, the operation of such components 282 and / or 283 and / or 284 may include or use one or more trained machine learning models.
[0060] In addition, the received images are forwarded to a BVGUM image selector and optional sequence determiner component 285p for analysis (e.g., to determine one or more image groups, each image group having one or more selected images for a corresponding video to be generated, optionally with a determined image sequence if multiple images are selected for the image group), the output of component 285p being one or more such image groups 275p, as discussed in more detail elsewhere herein, the operation of such component 285p may include or use one or trained machine learning models. The image groups 275p and the determined building attributes 274 are then provided to a BVGUM attribute selector and text description generator component 286p, wherein the output of component 286p is generated building text description information 276p, as discussed in more detail elsewhere herein, the operation of such component 286p may include or use one or trained language models as part of generating the text description information. In at least some embodiments, for each image group, component 286p may select some or all objects visible in one or more selected images of the image group (e.g., using information from BVGUM component 282 and / or information received in other information 298) for use as attributes of the building, and optionally further select other building attributes (e.g., attributes corresponding to visual characteristics of some or all of the selected objects, such as using information from BVGUM component 282 and / or information received in other information 298; attributes corresponding to information about a plurality of rooms, optionally for the building as a whole or for a floor or other subset of the building, including room layouts about the building or a subset of the building, such as using information from BVGUM component 283 and / or information received in other information 298; other attributes obtained from analyzing textual descriptions of the building or other building information, such as using information from BVGUM component 284 and / or information received in other information 298, etc.). For the selected attributes of the image group, BVGUM component 286p may then generate a textual description of each such attribute, and then combine the attribute descriptions to form an overall textual description of the building for use with the image group.
[0061] The image groups 275p and the generated building text description information 276p are then provided to a BVGUM building video generation component 287p, the output of which is a generated building video with an accompanying narrative description 277p for each image group. In at least some embodiments, for a group of images, component 287p may select visual data for one or more images in the group to include in the video (e.g., in accordance with an accompanying determined sequence, if any) so as to display visual data for an object in a selected attribute whose textual description is part of the information 276p for the group of images, and optionally display (e.g., highlight) visual data corresponding to other such selected attributes in other manners, as discussed in more detail elsewhere herein, the visual data selected for a selected panoramic or perspective image may be used in one or more frames of the video being generated, including optionally using techniques such as panning, tilting, zooming, etc., and corresponding series of visual data groups used in consecutive frames, and in some cases displaying a single visual data group from an image (e.g., some or all of a selected perspective image, a subset of a selected panoramic image, etc.) in multiple consecutive video frames (e.g., showing the same scene for one or more seconds). Component 287p may also select and use textual description information from information 276p for the group of images to generate a narration to accompany the selected visual data of the video in a synchronized manner, such as an audible narration for the audio portion of the video (e.g., using automatic text-to-speech generation, obtaining and using a manually provided transcript of information for the narration, etc.), and / or a visually displayed textual narration (e.g., in a manner similar to closed captioning). Additionally, for a group of images having multiple selected images, component 287p may add additional information corresponding to transitions between visual data of different images, such as additional visual data using one or more types of cinematic transitions and / or additional narration describing the transitions.
[0062] In some embodiments, the BVGUM system may also include a BVGUM building matcher component 288 to receive the determined building attributes 274 and / or the generated building description information 276P and / or information about the selected images of the image group 275P, and use the information to identify that the current building and / or one or more generated videos of the building match one or more specified criteria (e.g., at a later time after generating the information 274-276P and optionally 277P, such as when the corresponding criteria are received from one or more client computing systems 182 via one or more networks 170), and if so, the component 288 generates matching building information 279, which may include information about the building, such as one or more of the generated building videos 277P and optionally some or all of the building information 274 and / or 276P. After generating one or more of these types of information 277p, 274, and / or 276p, the BVGUM system may further perform step 289 to display or otherwise provide some or all of the generated and / or determined information for transmission over the network 170 to one or more client computing systems 182 for display (e.g., rendering the video 277p, such as by displaying a visual portion of the video and optionally playing a synchronized audio portion of the video), to one or more remote storage systems 181 for storage, or to one or more other recipients for further use. Additional details regarding the operation of the various BVGUM system components and the corresponding types of information that are analyzed and generated are included elsewhere herein.
[0063] Figure 2Q Continue to show Figures 2A to 2P , and with something like Figure 2LIn a manner similar to that shown in FIG. 2 (including an indication of the locations of acquisition locations 210A to 210-O of the building images, but without details of the image sequence 2251 or the path 115 traveled when acquiring the building images), information 290q is shown, which shows an example of a floor plan of a building 198 and surrounding courtyards 187 and 188 on the property on which the building is located, and an additional path 225q (starting at start location 225q-a and ending at end location 225q-b) is shown for an existing building video that has been captured by a camera device (not shown), along which the camera device travels through portions of the building and reaches the exterior area 186 / 187, it being understood that such a path may travel through zero or one or more acquisition locations. In this example, various information from the floor plan may be used to generate one or more additional videos, each corresponding to a subset of the existing building video. As a non-exclusive example, an additional video may be generated to include visual data of a kitchen / dining room so as to correspond to a subset 224q of path 225q. In some embodiments, such a subset 224q may be automatically determined by the BVGUM system in various ways, as discussed in more detail elsewhere herein, non-exclusive examples of which include the BVGUM system analyzing visual data of an existing building video to determine one or more types of result information and determining to generate additional video for each of one or more rooms or other areas through which the path passes, and non-exclusive examples of the types of result information determined include one or more of: the location of the path in the existing video (e.g., enabling generation of a visual representation of the path 225q); transitions between rooms (e.g., a transition 223q into and out of a kitchen / dining room); movement patterns 226q of the path that are associated with a room or otherwise satisfy one or more defined criteria; the presence of building attributes visible in the visual data of the existing video (e.g., building attributes 229q in the kitchen / dining room, which in this example include various objects and surfaces in the kitchen), etc. In some embodiments, such a subset 224q may be selected by one or more users in a variety of ways, whether in lieu of or in addition to automatic determination by the BVGUM system, and as discussed in more detail elsewhere herein, non-exclusive examples include a user (not shown) using a displayed GUI (not shown) that includes some or all of the information 290q for selecting or otherwise specifying a kitchen / dining great room, and / or for selecting or otherwise specifying an area of a building that includes the kitchen / dining great room, and / or for selecting or otherwise specifying a portion of a path 225q in the kitchen / dining great room, and / or for selecting or otherwise specifying one or more building attributes in the kitchen / dining great room (e.g., one or more of the building attributes 229q), etc.The generated additional videos corresponding to the subset of pathways 224q and the corresponding subset of the existing building videos may be further presented to one or more users once generated. Additionally, once the additional building video corresponding to the kitchen / dining room is generated, various types of information may be associated therewith, such as a visual representation and / or other information about the pathway associated therewith, one or more labels for one or more rooms or other areas for which the video includes visual data (e.g., "kitchen / dining room" in this example), one or more building attributes for which the video has visual data, and the like.
[0064] Figure 2R Continue to show Figures 2A to 2Q and shows an example of a GUI 260r that may be displayed to one or more users to include Figure 2Q at least some of the information 290q to enable a user to select or other segment criteria specification to identify a subset of existing building videos for use in generating additional building videos corresponding to the path subset 224q based on the floor plan information presented to the user, in which example the floor plan 230r shown in the GUI 260r includes a visual representation 225r of the path of the existing building videos, but does not include information overlaid on the floor plan regarding acquisition locations 210A to 210-O or building attributes 229q, but includes additional display information 231r and 232r associated with the floor plan to respectively list media clips of the building grouped by room and at least some of the determined building attributes grouped by room. In this example, for example, a user (not shown) may specify that additional videos corresponding to the kitchen / dining great room be generated by one or more of: selecting 233r the kitchen / dining great room, such as by clicking on the floor plan within the room if the rooms are individually selectable, or by drawing a shape (not shown) around the kitchen / dining great room, or selecting the kitchen / dining great room from a displayed list of rooms (not shown), etc.; selecting 232r1 or otherwise specifying one or more building attributes in the kitchen / dining great room (e.g., from a list 232r of building attributes grouped by room; by selecting a visual representation of the building attributes overlaid on the floor plan (not shown), or otherwise specifying the building attributes, etc.), etc. Additionally, in various embodiments, one or more existing and / or generated videos associated with the building floor plan may be displayed on the floor plan or otherwise associated with the floor plan in various ways, such as overlaying a visual representation of the video path on the floor plan, displaying icons (e.g., user selectable) representing some or all of the videos in each room traversed by the video path, displaying a list of videos (e.g., grouped by room or other area), etc.
[0065] Figure 2S Continue to show Figures 2A to 2R, and information 290s is shown, 290s showing example data flow interactions for at least some automated operations of one embodiment of a BVGUM system (referred to in this example as a BVGUM-B system). Specifically, an embodiment of the BVGUM system 140 is shown as being executed on one or more computing systems 180, and in this example embodiment, receiving information about a building to be analyzed, the information including stored images from a memory or database 295, floor plans from a memory or database 296, one or more existing building videos from a memory or database 297, and optionally other information 298 (e.g., other building information, such as tags, annotations, etc.; other configuration settings or instructions for controlling video generation, such as length, type of information to be included, etc.). In step 281s, input information is received, the received video information is forwarded to the BVGUM video analyzer component 291s for analysis (e.g., identifying associated paths or other positioning information for one or more video locations; objects visible in the visual data of the video and optionally other building attributes, etc.), the received floor plan information is forwarded to the BVGUM floor plan analyzer component 283 for analysis (e.g., determining building attributes from the floor plan, such as some or all global attributes corresponding to the building as a whole and / or to specific rooms or other areas), and the received image is optionally forwarded to the BVGUM image analyzer component 282 for analysis in a manner similar to Figure 2P , and other received building information is optionally forwarded to the BVGUM other information analyzer component 284 for use in a manner similar to Figure 2P , and the output of components 282 and / or 283 and / or 284 forms some or all of the determined building attributes 274 of the building, and the output of component 291s forms the determined video attributes 292, in other embodiments, information about some or all of such video attributes and / or information about objects and other building attributes may be received (e.g., as part of other information 298) and directly included in the video attribute information 292 and building attribute information 274, respectively, and as discussed in more detail elsewhere herein, the operation of such components 291s and / or 282 and / or 283 and / or 284 may include or use one or more trained machine learning models. The determined building attributes 274 are then optionally provided to a BVGUM attribute selector component 286s, wherein the optional output of component 286s is selected as the building attributes 276s in order to select some or all objects of one or more specified types for which visual data is to be included in the additional video to be generated.
[0066] The existing video and video attributes 292 and the determined building attributes 274 and optionally the selected building attributes are then provided to a BVGUM video subset determiner component 285s, the output of which is one or more subsets of one or more existing videos for use in each of one or more additional videos in its generation, such as based in part or in whole on user input 294s obtained by component 285s and / or received in other information 298, and / or based in part or in whole on automatic operation of the BVGUM system to determine one or more existing video subsets (e.g., a subset of visual data including the selected building attribute 276s), as discussed in more detail elsewhere herein. The correspondingly selected one or more existing videos and / or determined subsets thereof are then provided to a BVGUM additional video generation component 287s, wherein the output of component 287s is at least one generated additional video 277s having visual data for at least a subset of the existing videos, including, in some embodiments and cases, using a single existing building video subset as the generated additional video without further modification or manipulation, and in some embodiments and cases, sequentially using two or more existing building video subsets as the generated additional video without further modification or manipulation, and in some embodiments and cases, performing further modifications or other manipulations (e.g., cropping frames, adjusting lighting and other visual attributes, selecting specific poses and / or zoom levels to be used as subset portions of each of one or more frames, etc.) on one or more such existing video subsets as part of generating the additional video. In addition, for additional videos that include two or more existing video subsets, component 287s may add additional information corresponding to transitions between visual data of different subsets, such as additional visual data describing the transitions using one or more types of cinematic transitions and / or narration. The output from component 287s may also optionally include additional information 278s about one or more generated additional videos for presentation on or in association with the floor plan presentation, such as a visual representation of the path of the generated additional videos, an indication of one or more rooms or other areas, inclusion of the additional videos in a list or other grouping of associated media data in the one or more rooms or other areas, and the like.
[0067] In some embodiments, the BVGUM system may also include a BVGUM building matcher component 288 to receive the determined building attributes 274 and / or the determined video attributes 292 and use that information to identify that the current building and / or one or more generated videos of the building match one or more specified criteria (e.g., at a later time after generating the information 274 and 292 and 277s, such as after receiving the corresponding criteria from one or more client computing systems 182 via one or more networks 170), and if so, component 288 generates matching building information 279, which may include information about the building, such as one or more of the generated additional videos 277, and optionally also some or all of the building information 274 and / or the additional information 278S. After generating one or more of these types of information 277s, 278s, 274, and / or 292, the BVGUM system may further perform step 289 to display or otherwise provide some or all of the generated and / or determined information, such as by sending such information to one or more client computing systems 182 via network 170 for display (e.g., rendering video 277s, such as by displaying a visual portion of the video and optionally playing a synchronized audio portion of the video), to one or more remote storage systems 181 for storage, or to one or more other recipients for further use. Additional details regarding the operation of various BVGUM system components and the corresponding types of information that are analyzed and generated are included elsewhere herein.
[0068] Figure 2T Continue to show Figures 2A to 2S , and shows an example of a GUI 260t that can be similar to Figure 2R 225t is displayed to one or more users in a manner that is not disclosed to one or more users, but in this example enables a user to select or otherwise specify one or more generation criteria for a new video to be generated that is not based on one or more subsets of one or more existing building videos so as to instead be generated using visual data of a subset of building images. In this example, a path 225t is specified by the user and / or automatically determined by the BVGUM system for generating such a new video, wherein the path 225t is optionally not displayed to the user (e.g., until after the new video is generated, until after the generation criteria determination is completed, etc.). In this example, the floor plan 230t shown in the GUI 260t resembles Figure 2R260r, but does not include path 225r for the existing video (e.g., because that video does not exist, such as reflected in the updated grouping of media data for the kitchen / dining room in information 231t; because the existing video is not relevant to the generation of the new video, etc.). However, the GUI 260t does include one or more indications of generation criteria for the new video to be generated, such as user selection or automatic determination 233t of one or more rooms or other areas for which visual data is to be included in the new video (in this example, the kitchen / dining great room, and the family room, such as by selecting the rooms on the floor plan, drawing one or more shapes on the floor plan including the rooms, etc.), and / or user selection or automatic determination 232t1 of one or more building attributes for which visual data is to be included in the new video, and / or user selection or automatic determination of one or more building images 232t2 for use in generating the new video, and / or user specification or automatic determination of a path 225t for the new video, and / or user specification or automatic determination 236tr of one or more building locations for which visual data is to be included in the new video, and / or user specification or automatic determination of an orientation or pose (not shown) to be used from one or more locations for including visual data in the new video, etc.
[0069] Figure 2U Continue to show FIG. 2A to FIG. 2T , and shows information 290U, which is similar to Figure 2QAn example of a floor plan of a building 198 and surrounding courtyards 187 and 188 on the property on which the building is located is shown in a manner that is similar to that of FIG. 1 (but without indicating the locations of some building image acquisition locations), wherein an additional path 225U (starting at start location 225U-a and ending at end location 225U-a) is shown for a new video to be generated using visual data of building images at least at acquisition locations 210G and 210J and 250r, and optionally further at one or more of acquisition locations 210I and 210K, it being understood that such a path may travel through zero or one or more acquisition locations. In this example, various additional information is shown, such as automatically determined orientations or poses 234u at various locations along the path, such as may be automatically determined by the BVGUM system to highlight building attributes that have been selected and / or determined, or otherwise meet one or more defined criteria. In this example, the determined orientation or pose includes executing a 360 degree circle at the starting position 225u-a, then proceeding at a base tangent to the path for the first approximately half of the path, then changing direction to pose highlight a particular attribute in the kitchen (e.g., an island, stove, refrigerator, sink, etc.) before optionally returning to a base tangent orientation or pose near the end of the path (not shown) and / or by completing another 360 degree circle at the end of the path (not shown). In various embodiments, the path and / or orientation / pose may be determined in various ways, as discussed in more detail elsewhere herein.
[0070] Figure 2V Continue to show Figures 2A to 2U An example is provided, and information 290V is shown, which shows an example data flow interaction for at least some automated operations of an embodiment of a BVGUM system (referred to as a BVGUM-C system in this example). In particular, an embodiment of the BVGUM system 140 is shown as being executed on one or more computing systems 180, and in this example embodiment, information about a building to be analyzed is received, the information including stored images from a memory or database 295, a floor plan from a memory or database 296, and optionally other information 298 (e.g., one or more existing building videos; other building information, such as labels, annotations, etc.; other configuration settings or instructions for controlling video generation, such as length, type of information to be included, etc.). Input information is received in step 281v, and the received floor plan information is forwarded to a BVGUM floor plan analyzer component 283 for use in a manner similar to Figure 2S The received image is optionally forwarded to a BVGUM image analyzer component 282 for use in a manner similar to that of FIG. 284 (e.g., determining building properties from the floor plan, such as global properties corresponding to some or all of the building as a whole and / or to specific rooms or other areas). Figure 2P and Figure 2S , and other received building information is optionally forwarded to the BVGUM other information analyzer component 284 for use in a manner similar to Figure 2P and Figure 2S The determined building attributes 274 are then optionally provided to a BVGUM attribute selector component 286v, wherein the optional output of component 286v is selected as the building attributes 276v, so as to select some or all objects of one or more specified types for which visual data is to be included in the additional video to be generated.
[0071] The building image and determined building attributes 274 and optionally selected building attributes 276V are then provided to a BVGUM determiner component 285V which determines a path to use and optionally determines building attributes to highlight, wherein the output of component 285V is provided to a BVGUM video generation component 287v for use in generating one or more new videos and optionally to a BVGUM posture determiner component 291, such as based in part or in whole on user input 294v obtained by component 285v and / or received in other information 298, and / or based in part or in whole on automatic operation of the BVGUM system to determine such information, as discussed in more detail elsewhere herein. If a BVGUM pose determiner component 291 is used, the output of component 285v and / or user input 294v may be further used to determine the orientation / pose used at each of one or more locations along the determined path, and this information is further provided to component 287v for use in generating one or more new videos, wherein the output of component 287v is at least one generated new video 277v having at least a portion of visual data for one or more rooms or other areas of the building, including visual data from images along the path, or having visual data corresponding to the orientation / pose determined from the path. The output from component 287v may also optionally include additional information 278v about the one or more generated new videos for presentation on or in association with the floor plan presentation, such as a visual representation of the path of the generated new video, an indication of one or more rooms or other areas, inclusion of the new video in a list or other grouping of associated media data, and the like.
[0072] In some embodiments, the BVGUM system may also include a BVGUM building matcher component 288 to receive the determined building attributes 274 and use the information to identify that the current building and / or one or more generated videos of the building match one or more specified criteria (e.g., at a later time after generating the information 274 and 277v, such as after receiving the corresponding criteria from one or more client computing systems 182 via one or more networks 170), and if so, the component 288 generates matching building information 279, which may include information about the building, such as one or more generated new videos 277v and optionally some or all of the building information 274 and / or additional information 278v. After generating one or more of these types of information 277v, 278v, 274, and / or 292, the BVGUM system may further perform step 289 to display or otherwise provide some or all of the generated and / or determined information, such as by sending such information to one or more client computing systems 182 via network 170 for display (e.g., rendering video 277v, such as by displaying a visual portion of the video and optionally playing a synchronized audio portion of the video), to one or more remote storage systems 181 for storage, or to one or more other recipients for further use. Additional details regarding the operation of various BVGUM system components and the corresponding types of information that are analyzed and generated are included elsewhere herein.
[0073] Already referenced Figures 2A to 2V Various details are provided, but it is to be understood that the details provided are non-exclusive examples included for purposes of illustration and that other implementations may be performed in other ways without some or all of such details.
[0074] Figure 3 is a diagram showing execution of a BVGUM system 340 (e.g., in a manner similar to Figure 1ABlock diagram of an embodiment of one or more server computing systems 300 implementing a building information access system (e.g., one or more server computing systems 300 and BVGUM system 140) and one or more server computing systems 380 executing an implementation of an ICA system 388 and a MIGM system 389 (server computing systems and BVGUM and / or ICA and / or MIGM systems). One or more server computing systems and BVGUM and / or ICA and / or MIGM systems may be implemented using multiple hardware components that form electronic circuits suitable and configured to perform at least some of the techniques described herein when in combined operation. One or more computing systems and devices may also optionally execute a building information access system (e.g., one or more server computing systems 300) and / or optional other programs 335 and 383 (e.g., one or more server computing systems 300 and 380, respectively, in this example), although the building information access system is not shown in this example. In the illustrated embodiment, each server computing system 300 includes one or more hardware central processing units ("CPUs") or other hardware processors 305, various input / output ("I / O") components 310, storage devices 320, and memory 330, and the illustrated I / O components include a display 311, a network connection 312, a computer-readable media drive 313, and other I / O devices 315 (e.g., I / O devices, keyboards, mice or other pointing devices, microphones, speakers, GPS receivers, etc.). Each server computing system 380 can have similar components, although for simplicity, only one or more hardware processors 381, memory 387, storage devices 384, and I / O components 382 are shown in this example.
[0075] In the illustrated embodiment, one or more server computing systems 300 and executing BVGUM system 340, one or more server computing systems 380 and executing ICA system 388 and MIGM system 389, and optionally executing a building information access system (not shown) can communicate with each other and other computing systems and devices, such as via one or more networks 399 (e.g., the Internet, one or more cellular telephone networks, etc.), including interacting with user client computing devices 390 (e.g., for viewing building information, such as generated building videos, building descriptions, floor plans, and images and / or other related information, such as by interacting with the building information access system or executing a copy of the building information access system), and / or mobile image acquisition devices 360 (e.g., for acquiring images and / or other information of a building or other environment to be modeled, such as in a manner similar to that of a mobile device). Figure 1A185), and / or other navigable device 395 that optionally receives and uses the floor plan and optionally other generated information for navigation purposes (e.g., for use by a semi-autonomous or fully autonomous vehicle or other device). In other embodiments, some of the described functionality may be combined in fewer computing systems, such as combining the BVGUM system 340 and a building information access system in a single system or device, combining the BVGUM system 340 and image acquisition functionality of device 360 in a single system or device, combining the ICA system 388 and the MIGM system 389 and image acquisition functionality of device 360 in a single system or device, combining the BVGUM system 340 with one or two of the ICA system 388 and the MIGM system 389 in a single system or device, combining the BVGUM system 340 with the ICA system 388 and the MIGM system 389 and image acquisition functionality of one or more devices 360 in a single system or device, etc.
[0076] In the illustrated embodiment, an embodiment of the BVGUM system 340 is executed in the memory 330 of one or more server computing systems 300 to perform at least some of the described techniques, such as by using the processor 305 to execute software instructions of the system 340 in a manner that configures the processor 305 and the computing system 300 to perform automated operations that implement the described techniques. The illustrated embodiment of the BVGUM system may include one or more components not shown to each perform a portion of the functionality of the BVGUM system, such as in the manner discussed elsewhere herein, and the memory may further optionally execute one or more other programs 335. As a specific example, in at least some embodiments, a copy of the ICA and / or MIGM system may be executed as one of the other programs 335, for example, instead of or in addition to the ICA system 388 or MIGM system 389 on one or more server computing systems 380, and / or a copy of the building information access system may be executed as one of the other programs 335. The BVGUM system 340 may also store and / or retrieve various types of data on the memory 320 during its operation (e.g., in one or more databases or other data structures), such as various user information 322, floor plans and other associated information 324 (e.g., generated and saved 2.5D and / or 3D models, building and room dimensions for use with associated floor plans, additional image and / or annotation information, etc.), images and associated information 326, generated building videos 328 (optionally with narrative descriptions) and other generated building information (e.g., determined building properties, generated property descriptions, generated building descriptions, etc.), and / or various types of optional additional information 329 (e.g., various analytical information related to the presentation or other use of one or more building interiors or other environments).
[0077] In addition, embodiments of the ICA system 388 and MIGM system 389 in the illustrated embodiment are executed in the memory 387 of the server computing system 380 to perform techniques related to generating panoramic images and floor plans of buildings, such as by using the processor 381 to execute software instructions of the systems 388 and / or 389 in a manner that configures the processor 381 and the computing system 380 to perform automated operations that implement these techniques. The illustrated embodiments of the ICA and MIGM systems may include one or more components not shown to perform portions of the functionality of the ICA and MIGM systems, respectively, and the memory may further optionally execute one or more other programs 383. The ICA system 388 or MIGM system 389 may also store and / or retrieve various types of data on the storage device 384 (e.g., in one or more databases or other data structures) during operation, such as video and / or image information 386 acquired for one or more buildings (e.g., for analysis to generate floor plans, to provide displayed 360° videos or images to users of the client computing device 390, etc.), floor plan and / or other generated mapping information 387, and optional other information 385 (e.g., additional image and / or annotation information for use with the associated floor plan, building and room dimensions for use with the associated floor plan, various analytical information related to the presentation or other use of one or more building interiors or other environments, etc.). Although Figure 3 Not shown, the ICA and / or MIGM system may also store and use additional types of information, such as information about other types of buildings to be analyzed and / or provided to the BVGUM system, about ICA and / or MIGM system operator users and / or end users, etc.
[0078] Some or all of the user client computing device 390 (e.g., mobile device), mobile image acquisition device 360, optional other navigable device 395, and other computing systems (not shown) may similarly include some or all of the same types of components shown for the server computing system 300. As a non-limiting example, the mobile image acquisition device 360 is each shown to include one or more hardware CPUs 361, I / O components 362, memory and / or storage devices 367, one or more imaging systems 365, IMU hardware sensors 369 (e.g., for acquiring video and / or images, associated device movement data, etc.), and optional other components. In the illustrated example, one or both of a browser and one or more client applications 368 (e.g., applications dedicated to the BVGUM system and / or ICA system and / or MIGM system) are executed in the memory 367 to participate in communications with the BVGUM system 340, the ICA system 388, the MIGM system 389, and / or other computing systems. Although specific components are not illustrated for other navigable devices 395 or other computing devices / systems 390, it will be appreciated that they may include similar and / or additional components.
[0079] It should also be understood that Figure 3 The computing systems 300 and 380 and other systems and devices included in the are merely illustrative and are not intended to limit the scope of the invention. The system and / or device may instead each include multiple interactive computing systems or devices, and may be connected to other devices not specifically shown, including via Bluetooth communication or other direct communication, via one or more networks such as the Internet, via the Web, or via one or more dedicated networks (e.g., mobile communication networks, etc.). More generally, the device or other computing system may include any combination of hardware that can interact and perform the functions of the type described, optionally when programmed or otherwise configured with specific software instructions and / or data structures, including but not limited to desktop computers or other computers (e.g., input boards, tablet computers, etc.), database servers, network storage devices and other network devices, smart phones and other cellular phones, consumer electronic devices, wearable devices, digital music player devices, handheld gaming devices, PDAs, wireless phones, Internet devices, and various other consumer products including appropriate communication capabilities. Additionally, in some implementations, the functionality provided by the illustrated BVGUM system 340 may be distributed among various components, some of the described functionality of the BVGUM system 340 may not be provided, and / or other additional functionality may be provided.
[0080] It should also be understood that although various items are shown as being stored in memory or stored in storage devices when used, these items or portions thereof may be transferred between memory and other storage devices for the purposes of memory management and data integrity. Alternatively, in other embodiments, some or all of the software components and / or systems may be executed in memory on another device and communicate with the illustrated computing system via inter-computer communication. Thus, in some embodiments, when configured by one or more software programs (e.g., by the BVGUM system 340 executed on the server computing system 300, by the building information access system executed on the server computing system 300 or other computing systems / devices, etc.) and / or data structures, some or all of the described techniques may be performed by a hardware device including one or more processors and / or memory and / or storage devices. Such as by executing software instructions of one or more software programs and / or by storing such software instructions and / or data structures, and in order to perform algorithms as described in the flowcharts and other disclosures herein. Additionally, in some embodiments, some or all of the systems and / or components may be implemented or provided in other ways, such as by being comprised of one or more devices implemented in part or in whole in firmware and / or hardware (e.g., rather than devices implemented in whole or in part by software instructions that configure a particular CPU or other processor), including, but not limited to, one or more application specific integrated circuits (ASICs), standard integrated circuits, controllers (e.g., by executing appropriate instructions, and including microcontrollers and / or embedded controllers), field programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), etc. Some or all of the components, systems, and data structures may also be stored (e.g., as software instructions or structured data) on a non-transitory computer-readable storage medium, such as a hard disk or flash drive or other non-volatile storage device, volatile or non-volatile memory (e.g., RAM or flash RAM), network storage device, or portable media article (e.g., DVD disk, CD disk, optical disk, flash memory device, etc.) to be read by an appropriate drive or via an appropriate connection. In some embodiments, systems, components, and data structures may also be transmitted via generated data signals (e.g., as part of a carrier wave or other analog or digital propagation signal) on various computer-readable transmission media, including wireless-based and wire / cable-based media, and may take various forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). In other embodiments, such computer program products may also take other forms. Therefore, embodiments of the present disclosure may be implemented with other computer system configurations.
[0081] FIG. 4A to FIG. 4DAn exemplary embodiment of a flow chart for a Building Video Generation and Usage Manager (BVGUM) system routine 400 is shown. For example, the BVGUM system routine 400 may be executed by executing Figure 1A BVGUM system 140, Figure 3 BVGUM system 340, and / or as referenced Figure 2D to Figure 2V and the BVGUM system described elsewhere herein to perform automated operations related to automatically generating a building video having one or more indicated types of visual data (e.g., user-selected and / or automatically determined building objects or other building attributes, user-selected and / or automatically determined rooms or other areas, etc.), optionally accompanied by an automatically generated narrative, and subsequently using the generated building video in one or more automated ways. FIG. 4A to FIG. 4D In an exemplary embodiment of the invention, the indicated building may be a house or other building, and generating a video of the building includes generating descriptive information for selected building attributes and using them to narrate a video including visual data of images acquired for the building, but in other embodiments, other types of data structures and analyses may be used for other types of structures or for non-structural locations, and the generated building information may be used in other ways than discussed with respect to routine 400, as described elsewhere herein.
[0082] The illustrated implementation of the routine begins at block 405, where information or instructions are received. The routine continues to block 410 to determine whether the instructions or other information received in block 405 indicate that one or more videos are to be generated for the indicated building (e.g., based at least in part on existing images and / or videos of the indicated building), and if so, the routine continues to execute some or all of blocks 415-465 and / or 905-990 to do so, otherwise continuing to block 476. In block 415, the routine optionally obtains recipient-specific configuration settings and / or information (e.g., information received in block 405, stored information, etc.) used in video generation, such as corresponding to video length, type of information included in the video (e.g., room type, object type, other property type, etc.). In block 420, the routine then determines whether existing building information (e.g., images, floor plans with at least 2D room shapes positioned relative to each other, one or more existing videos, textual building descriptions, lists or other indications of building objects and / or other building attributes, labels and / or descriptive annotations associated with images and / or rooms and / or objects, etc.) is available for the building, and if so, proceeds to block 422 to retrieve such existing building information. If instead it is determined in block 420 that building information is not available, the routine instead executes blocks 425-440 to generate such images and floor plans and associated information, including optionally obtaining available information about the building in block 425 (e.g., building dimensions and / or other information about building dimensions and / or structure; exterior images of the building, such as from an overhead and / or from a nearby street, etc., such as from a public source), initiating execution of the ICA system routine in block 430 to obtain images and optional additional data for the building (for example, in the case of a building that is not available in the block 430). Figure 5 ), and initiating execution of the MIGM system routine in block 440 to generate building data associated with the floor plan and optional additional mapping using the image acquired from block 430 (for example, in Figure 6A-6B An example of such a routine is shown in ).
[0083] In blocks 441-465 and blocks 905-990, the routine performs several activities as part of using the building information from blocks 430 and 440 or from block 422 to generate one or more videos for the building. In particular, in block 441, the routine includes: if such information is not already available from blocks 422 and / or 430, analyzing each image using one or more trained machine learning models (e.g., one or more trained classification neural networks) to identify structural elements and other objects and determine other attributes associated with such objects (e.g., color, surface material, style, location, orientation, descriptive labels) or for use with the building. In box 442, the routine then optionally analyzes other building information (e.g., floor plans, text descriptions, etc.) to determine other properties of the building using one or more trained machine learning models (e.g., one or more trained classification neural networks), such as based at least in part on the layout information (e.g., connectivity and other adjacency information for groups of two or more rooms), the determined properties may, for example, include attributes that each classify the floor plan of the building according to one or more subjective factors (e.g., accessibility-friendly, open floor plan, atypical floor plan, etc.), room types for some or all rooms in the building, types of connections between rooms and other adjacencies between some or all rooms (e.g., connected by doors or other openings, adjacent to but not connected to intervening walls, non-adjacent, etc.), one or more objective attributes, etc.
[0084] In box 443, the routine determines whether the instructions or other information received in box 405 indicate the use of one or more existing building videos to generate at least one additional video, such as by providing corresponding segment criteria for use in generating the additional video or otherwise indicating the generation of one or more additional videos, and if so, proceeds to boxes 905-935 to do so. Specifically, in box 905, the routine retrieves one or more existing building videos (such as retrieved in box 422 or received in box 405), and in box 910, proceeds to optionally present information for the indicated building in the displayed GUI to obtain user-specified segment criteria from one or more end-user recipients and / or other users of the one or more additional videos to be generated (e.g., if not received from the building information access routine or other source in box 405, then in box 910, the routine proceeds to boxes 905-935 to optionally present information for the indicated building in the displayed GUI to obtain user-specified segment criteria from one or more end-user recipients and / or other users of the one or more additional videos to be generated). FIG. 7A to FIG. 7B). Such presentation of building information may include, for example, displaying a visual representation of the floor plan at one or more associated locations on the floor plan overlaid with a visual representation of one or more existing videos (e.g., displaying a path of the existing videos) or otherwise provided in association with the floor plan (e.g., to provide a list or other group of such existing building videos), and optionally, displaying information about building properties and building images at associated locations on the floor plan or otherwise provided in association with the floor plan (e.g., providing a list or other group of building images associated with rooms or other areas of the building, providing a list or other group of building properties associated with rooms or other areas of the building, etc.), and providing one or more user-selectable controls to enable one or more users to provide one or more segment criteria for generating one or more additional videos from the one or more existing videos. In box 915, the routine optionally receives user input from at least one user (e.g., an end-user recipient to whom the generated new video will be presented) to provide one or more segment criteria regarding various aspects of the building for which visual data is to be included in the one or more additional videos to be generated by, for example, indicating one or more portions of one or more existing videos to be used to generate the corresponding additional videos (e.g., by selecting portions of the presented visual representation of the existing videos, by including one or more rooms or other areas of the generated additional video to include visual data from one or more existing videos, by specifying one or more building properties to include visual data from one or more existing videos to include in the generated additional video, by providing a linguistic description of the analysis as one or more such types of data, etc.). In box 920, the routine then analyzes each existing building video to associate some or all frames with rooms or other areas of interest, including optionally identifying building attributes, such as building objects and / or structural elements (e.g., by matching visual data of one or more building images with known locations, and / or matching with data from a floor plan, and / or identifying building objects and / or structural elements associated with room types of rooms or other areas on the floor plan), optionally detecting transitions between rooms and / or other areas for determining where to begin or end additional video for that or the next room in the existing building, optionally detecting motion patterns along a video path that is associated with a room or other area of interest, or otherwise satisfies one or more defined criteria, etc. In addition, in some embodiments, the routine may further determine whether the frame of the video has a sufficiently wide and / or tall viewing angle that is the viewing angle for standard playback, using only a portion of the frame.In box 925, the routine then determines aspects of the building for which visual data is to be included in one or more additional videos to be generated, such as based in part or in whole using the segment criteria specified in box 915 or received in box 405, and / or based in part or in whole in an automatic manner (e.g., for an existing video having a path through multiple rooms or other areas of a building, to generate an additional video for each such room or other area that includes a portion of the existing video), and additionally, if the video frames are determined to have a sufficiently wide and / or high viewing angle that standard playback of the frames uses only a portion of the frames, the routine may further determine a particular orientation to be used from a particular location (e.g., to call attention to a particular building attribute) so as to correspond to particular portions of at least some of the video frames. In box 930, the routine then generates one or more additional videos having visual data for the architectural aspect determined in box 920, wherein each additional video includes one or more subset portions of one or more existing videos and includes at least visual data from the one or more subset portions, such as by using data from the analysis of the one or more existing videos in box 925, and if the video frame is determined to have a viewing angle that is wide and / or high enough that standard playback of the frame uses only a portion of the frame, the routine may further determine specific portions of at least some of the video frames of the subset portion of the existing video to be presented during playback of the generated additional video (whether by default, with the user being able to change the viewpoint during playback while viewing within the frame of the generated additional video, or as the only option for playback) so as to display the determined architectural attribute of interest. In box 935, the routine then optionally generates one or more visual representations associated with one or more of the generated additional videos to be overlaid on or otherwise included in the floor plan of the building (e.g., a path along which the visual data of the generated video moves, a group of multiple media blocks associated with a room or other area, a user-selectable icon within the room or other area to which the generated additional video corresponds, etc.), and optionally presents the generated additional videos and / or at least one of the generated visual representations along with the floor plan to one or more end-user recipients (e.g., one or more users from whom user input was received in box 915), such as by updating the information presented in box 910.
[0085] If it is determined in box 443 that one or more additional videos are not to be generated from one or more existing building videos, the routine continues to box 444 to determine whether the instructions or other information received in box 405 indicate that at least one new video is to be generated using information associated with the building floor plan, such as by providing corresponding generation criteria for use in generating the new video or otherwise indicating generation of one or more new videos, and if so, continues to execute boxes 955-990 to do so. Specifically, in box 955, the routine optionally obtains user input from at least one user (e.g., an end-user recipient to whom the generated new video will be presented) to provide one or more generation criteria regarding various aspects of the building for which visual data is to be included in the one or more new videos to be generated (e.g., if not received from the building information access routine or other source in box 405, where Figure 7A-7BAn example of such a routine is discussed in ), such as using a user to select aspects from a displayed GUI with building information and / or from verbal input. The displayed GUI may include, for example, a floor plan with a visual representation of one or more images superimposed at one or more associated locations on the floor plan or provided in association with the floor plan (e.g., to provide a list or other group of such images), and optionally, information about building properties displayed at associated locations on the floor plan or provided in association with the floor plan (e.g., to provide a list or other group of building properties associated with rooms or other areas of the building, etc.), and providing one or more user-selectable controls to enable the user to provide one or more generation criteria for generating one or more new videos using the visual data of the images and the structural information about the building from the floor plan. In addition, user input may be based on, for example, indicating one or more images having visual data for use in generating a corresponding new video (e.g., by directly selecting an image, by specifying one or more rooms or other areas for which visual data is to be included in the generated new video and wherein the indicated images have corresponding acquisition locations or otherwise include such visual data, by specifying one or more building attributes for which visual data is to be included in the generated new video and wherein the indicated images have corresponding visual data, by providing a verbal description that is analyzed to identify one or more such images, etc.), and / or by indicating one or more rooms or other areas for which the new video will have visual data, such as by directly selecting one or more rooms or other areas, specifying a path that the visual data of the new video follows through the one or more rooms or other areas, etc. In box 960, the routine then determines one or more areas of the building for which to include visual data in one or more new videos to be generated, the one or more new videos including one or more rooms or other areas, such as based in part or in whole using the generation criteria specified in box 955 or received in box 405, and / or based in part or in whole in an automatic manner (e.g., rooms or other areas through which an indicated path passes, rooms or other areas that include one or more indicated building attributes, etc.).In box 965, the routine then determines multiple aspects of the building in the determined one or more building areas for which visual data is to be included in the one or more new videos to be generated (e.g., one or more building attributes for which visual data is to be included), such as based in part or in whole using the generation criteria specified in box 955 or received in box 405, and / or based in part or in whole on an automatic manner (e.g., determining a path through the indicated rooms or other areas that would be likely or appear to be followed by a human videographer and / or predicted to provide results that meet one or more defined criteria; determining one or more orientations from one or more locations for which visual data is included so as to point to particular selected building attributes, etc., and optionally using one or more trained machine learning models), and in addition, if the video frame to be included in the one or more new videos has a sufficiently wide and / or tall viewing angle that standard playback of the frame uses only a portion of the frame, the routine may further determine a particular orientation to use from a particular location (e.g., to highlight a particular building attribute of interest). In box 970, the routine then determines a building image having an acquisition location in the determined building area, or otherwise includes visual data for at least some of the determined building areas, and in box 975, proceeds to generate one or more new videos having visual data for the determined areas and corresponding to the determined aspects, such as by using NeRf and / or SFM generation techniques based on the visual data of the determined building image, and optionally including a corresponding generated narrative about the determined building areas and / or building attributes and other determined aspects as discussed in more detail elsewhere herein, if some or all of the frames of the generated one or more new videos have a sufficiently wide and / or high viewing angle such that standard playback of the frames uses only a portion of the frames, the routine may further determine specific portions of at least some of the video frames of the new videos to be presented during playback of the generated one or more new videos (whether by default the user is able to change the viewpoint within the frames of the generated new videos during playback, or viewing the playback as the only option for playback) so as to display the determined building attributes of interest.In box 990, the routine then optionally generates one or more visual representations associated with the one or more generated new videos, which are to be overlaid on or otherwise included in the floor plan of the building (e.g., a path along which the visual data of the generated videos moves, a group of multiple media associated with a room or other area, a user-selectable icon within the room or other area to which the generated new videos correspond, etc.), and optionally presents some or all of the generated new videos and / or at least one of the generated visual representations along with the floor plan to one or more end-user recipients (e.g., one or more users from whom user input was received in box 955, such as by updating the information presented in box 955).
[0086] If it is determined in box 444 that one or more new videos are not to be generated using information associated with the building floor plan, the routine continues to box 446 to analyze the images and optionally other building information to determine one or more image groups, each image group comprising a determined sequence of one or more selected videos and optionally multiple selected images, such as using one or more trained machine learning models (e.g., one or more neural networks) and based on any configuration settings and / or recipient information from box 415. After box 446, the routine continues to box 450 to select objects and optionally other attributes corresponding to the selected images to be depicted in one or more videos being generated (e.g., for each of some or all of the selected images, determine the objects that are visible in the image and, optionally, determine the location of the objects within the image, and optionally select further building attributes determined in box 444 and / or obtained in box 422), such as based on any configuration settings and / or recipient information from box 415, and optionally using one or more trained machine learning models (e.g., one or more neural networks), whether the same or different trained machine learning models used in box 446, in some embodiments and circumstances, object and optionally other attribute selection may be performed as part of the image selection in box 446. The routine in box 450 also includes: generating a text description for each of the selected objects and other attributes, such as based on any configuration settings and / or recipient information from box 415, and optionally using one or more trained language models (e.g., one or more trained transformer-based machine learning models), and optionally combining the generated descriptions to generate an overall building text description.
[0087] After box 450, the routine continues to box 455 to generate a visual portion of a video for each image group using one or more selected images of the image group, the image group including visual data from each of the images (e.g., in an order corresponding to a determined sequence of images), according to any configuration settings and / or recipient information from box 415, and, for example, utilizing one or more frames of each image corresponding to one or more selected groups of visual data from the image (e.g., multiple visual data groups from images corresponding to one or more of pans, tilts, zooms, etc. within the image, and including visual data corresponding to one or more selected objects or other selected attributes), and optionally also including visual data corresponding to one or more transitions between visual data of different images, in other embodiments, the visual data included in the generated video may include selected visual data groups from selected images at different locations within the video, such as by inserting visual data groups from one or more other selected images. In block 460, then, for each image group, the routine generates a synchronized narration of the video of the image group (e.g., for the audible portion of the video), based at least in part on the textual description of the selected object or other selected attribute of the image group generated from block 450, and optionally includes further descriptive information (e.g., one or more transitions between visual data corresponding to different images to provide an introduction and / or summary, etc.), according to any configuration settings and / or recipient information from block 415, and such as using one or more trained language models, in other embodiments, generating the narration of the video before generating the visual portion of the video, wherein the visual data included in the visual portion is selected to be synchronized with the narration. After block 460, the routine in block 465 then optionally provides one or more of the generated videos for presentation or otherwise presents the one or more generated videos.
[0088] If it is determined in block 410 that the instructions or other information received in block 405 are not to generate one or more building videos, the routine continues to block 476 to determine whether the instructions or other information received in block 405 are to modify an existing building video. If so, the routine proceeds to block 478 to obtain modification instructions or other criteria regarding how to perform the modification and information to indicate the video to be modified (e.g., an indication of a particular building) received in block 405 to retrieve the video and generate a new video by modifying the retrieved video according to the modification instructions or other criteria. Such modification criteria may include, for example, one or more of the following: one or more indicated time lengths (e.g., minimum time, maximum time, start and end times of a subset of videos, etc.), wherein the retrieved video is modified accordingly (e.g., removing segments based on associated priorities; removing start and / or end portions, etc.); indicating one or more rooms (e.g., based on one or more room types), wherein the retrieved video is modified to exclude information about such one or more rooms; indicating one or more objects (e.g., based on one or more object types), wherein the retrieved video is modified to exclude information about such one or more objects; indicating one or more room groupings (e.g., floors, multi-room apartments or town houses of dormitories or larger buildings, multiplex units, etc.), wherein the retrieved video is modified to exclude information about such one or more room groupings, etc. In some embodiments and scenarios, the modification criteria may also be based on a specific recipient in order to personalize the modified video to that recipient, wherein corresponding criteria specific to that recipient are retrieved (e.g., from stored recipient preference information). After box 465 or 478 or 530 or 590, the routine continues to box 489 to store the generated videos and optionally some or all of the other generated building information from boxes 420-487 and / or 905-990, and optionally further provide one or more generated videos and / or at least some of the other generated building information to one or more corresponding recipients (e.g., in box 405, receiving information and / or instructions from a user or other entity recipient, or otherwise specified in such information and / or instructions).
[0089] If it is determined in box 476 that the instructions or other information received in box 405 is not to modify an existing building video, the routine continues to box 482 to determine whether the instructions or other information received in box 405 is to identify one or more generated building videos that meet the indicated criteria (e.g., based on information about rooms and / or objects and / or other attributes described in the video, such as based at least in part on the narration accompanying the video) and / or identify one or more target buildings with such generated building videos, and if not, continue to box 490. Otherwise, the routine continues to box 484 to retrieve candidate building videos (e.g., building videos previously generated for one or more indicated buildings in boxes 443-478 and / or 905-990) and compare information about such videos with the specified criteria. In box 486, then, for each candidate video and / or associated building, the routine determines the degree of match of the candidate video / building's information with the criteria, and if there are multiple indicated criteria, determining the degree of match may include combining information for multiple criteria in one or more ways (e.g., average, cumulative total, etc.). The routine also optionally sorts the multiple candidate videos / buildings based on their degree of match and selects one or more best matches to be used as the identified target video or building (e.g., all matches are above a defined threshold, a single best match, etc., and optionally based on instructions or other information received in box 405), wherein the selected one or more best matches have the highest degree of match to the specified criteria. In box 488, the routine then presents or otherwise provides information for the selected one or more candidate videos / one or more buildings (e.g., providing one or more selected candidate videos for presentation, such as in an order based on degree of match; providing information about one or more selected candidate buildings for presentation, etc.), such as through a building information access routine, wherein information about FIG. 7A to FIG. 7B An example of such a routine is discussed.
[0090] If it is determined in box 482 that the information or instructions received in box 405 are not to identify one or more other target generated videos and / or associated buildings using one or more specified criteria, the routine continues to box 490 to perform one or more other indicated operations as appropriate. Such other operations may include, for example, receiving and responding to requests for previously generated video and / or other building information (e.g., requests for such information for display or other presentation on one or more client devices, requests for such information to be provided to one or more other devices for use in automatic navigation, etc.), training one or more neural networks or other machine learning models (e.g., classification neural networks) to determine objects and associated attributes based on analysis of images and / or other acquired environmental data, training one or more neural networks (e.g., classification neural networks) or other machine learning models to determine building attributes by analyzing building floor plans (e.g., based on one or more subjective factors such as accessibility-friendliness, open floor plans, atypical floor plans, non-standard floor plans, etc.), training one or more machine learning models (e.g., language models) to generate attribute description information for determined objects and optionally other indicated building attributes and / or generating building description information for a building having multiple such objects and optionally other indicated building attributes, obtaining and storing information about users of the routine (e.g., the current user's search and / or selection preferences), etc.
[0091] After frame 488 or 489 or 490, routine continues to frame 495, to determine whether to continue, for example, until receiving explicit instruction to terminate, or only when receiving explicit instruction to continue. If it is determined to continue, routine returns to frame 405 to wait for additional instructions or information, and otherwise continues to frame 499 and ends.
[0092] Although not targeted FIG. 4A to FIG. 4B Exemplary embodiments and FIG. 4C to FIG. 4DSome of the automated operations of the BVGUM system are described herein, but in some embodiments, a human user may further help facilitate some operations of the BVGUM system, such as providing an operator user and / or end user of the BVGUM system with one or more types of input that are further used in subsequent automated operations. As non-exclusive examples, such a human user may provide one or more of the following types of input: providing input to help identify objects and / or other attributes from the analyzed images, floor plans, and / or other building information so as to provide input in boxes 441 and / or 442 for use as part of the automated operation of the one or more boxes; providing input in box 446 for use as part of a subsequent automated operation to help select images for a group of images and / or determine a sequence of selected images; providing input in box 450 for use as part of a subsequent automated operation to help select objects and / or other attributes to be depicted in the video, and / or to help generate textual descriptions of the objects and / or other attributes, and / or to help generate a textual description of the building based at least in part on the generated textual descriptions of the objects and / or other attributes; providing input in box 455 for use as part of a subsequent automated operation to help select one or more visual data groups from the selected images, and / or to help specify transitions to use between images; providing input in box 460 for use as part of a subsequent automated operation to help determine a narration for the video based on the generated textual description, and / or to generate an audible version of the narration, etc. Additional details regarding implementations in which one or more human users provide input used in otherwise automated operations of the BVGUM system are included elsewhere herein.
[0093] Figure 5 An example flow chart of an implementation of an ICA (image capture and analysis) system routine 500 is shown. The routine may be implemented by, for example, the ICA system 160 of FIG. Figure 3 ICA system 388, and / or as described in Figures 2A to 2Vand the ICA systems described elsewhere herein to acquire 360° panoramic images and / or other images at a capture location within a building or other structure, e.g., for subsequent generation of associated floor plans and / or other mapping information. Although portions of the example routine 500 are discussed with respect to acquiring particular types of images at particular capture locations, it will be appreciated that this routine or similar routines may be used to acquire video (with video frame images) and / or other data (e.g., audio), either in lieu of or in addition to such panoramic or other stereo images (e.g., on real estate on which a target building is located to display courtyards, decks, patios, accessory structures, etc.). Additionally, while the illustrated embodiments acquire and use information from the interior of a target building, it will be appreciated that other embodiments may perform similar techniques on other types of data, including information on non-building structures and / or on the exteriors of one or more target buildings of interest. Additionally, some or all of the routines may be performed on a mobile device used by a user to acquire the image information, and / or some or all of the routines may be performed by a system remote from such a mobile device. In at least some embodiments, the information may be acquired from a mobile device that is used by a user to acquire the image information. FIG. 4A to FIG. 4D Block 430 of routine 400 of the embodiment of the present invention calls routine 500, provides corresponding information from routine 500 to routine 400 as part of the implementation of block 430, and in such case, returns processing control to routine 400 after blocks 577 and / or 599. In other embodiments, routine 400 may perform additional operations in an asynchronous manner without waiting for such a return of processing control (e.g., performing other processing activities while waiting for corresponding information from routine 500 to be provided to routine 400).
[0094] The illustrated implementation of the routine begins at block 505, where instructions or information are received. At block 510, the routine determines whether the received instructions or information indicate acquisition of visual data and / or other data representative of a building interior (optionally based on information provided about one or more additional acquisition locations and / or other guiding acquisition instructions), and if not, continues to block 590. Otherwise, the routine proceeds to block 512 to receive an instruction to begin an image acquisition process at a first acquisition location (e.g., from a user of a mobile image acquisition device that will perform the acquisition process). After block 512, the routine proceeds to block 515 to perform an acquisition location image acquisition activity for acquiring a 360° panoramic image of an acquisition location of a target building interior of interest, such as via one or more fisheye lenses and / or non-fisheye rectilinear lenses on a mobile device, and providing at least 360° of horizontal overlap around a vertical axis, although other types of images and / or other types of data may be acquired in other embodiments. As a non-exclusive example, the mobile image acquisition device may be a rotating (scanning) panoramic camera equipped with a fisheye lens (e.g., with 180° horizontal overlap) and / or other lenses (e.g., with less than 180° horizontal overlap, such as a conventional lens or a wide-angle lens or an ultra-wide lens). The routine may also optionally obtain annotations and / or other information about the acquisition location and / or surroundings from the user, for example, for later use in presenting information about the acquisition location and / or surroundings.
[0095] After completion of block 515, the routine continues to block 520 to determine whether there are more acquisition locations to acquire images, for example based on corresponding information provided by the user of the mobile device and / or received in block 505. In some embodiments, the ICA routine will only acquire a single image and then proceed to block 577 to provide the image and corresponding information (e.g., return the image and corresponding information to the BVGUM system and / or MIGM system for further use before receiving additional instructions or information to acquire one or more next images at one or more next acquisition locations). If there are more acquisition locations to acquire additional images at the current time, the routine continues to block 522 to optionally initiate the capture of link information (e.g., acceleration data) during the movement of the mobile device along the path of travel away from the current acquisition location and toward the next acquisition location within the building interior. The captured link information may include additional sensor data recorded during such movement (e.g., from one or more IMUs or inertial measurement units, on the mobile device or otherwise carried by the user) and / or additional visual information (e.g., images, videos, etc.). Initiating the capture of such link information may be performed in response to an explicit instruction from a user of the mobile device or based on one or more automatic analyses of information recorded from the mobile device. Additionally, in some embodiments, the routine may also optionally monitor the motion of the mobile device during movement to the next acquisition location and provide one or more guidance prompts (e.g., to the user) regarding the motion of the mobile device, the quality of the sensor data and / or visual information being captured, the relevant lighting / environmental conditions, the desirability of capturing the next acquisition location, and any other appropriate aspects of capturing link information. Similarly, the routine may optionally obtain annotations and / or other information from the user regarding the path of travel, such as for later use in presenting information about the path of travel or the result of the panoramic image inter-connection link. At box 524, the routine determines that the mobile device has arrived at the next acquisition location (e.g., based on an instruction from the user, based on stopping the forward movement of the mobile device for at least a predetermined amount of time, etc.), serving as the new current acquisition location, and returns to box 515 to perform acquisition location image acquisition activities for the new current acquisition location.
[0096] If it is determined in block 520 that there are not any more acquisition locations at which to acquire image information of the current building or other structure at the current time, the routine proceeds to block 545 to optionally pre-process the acquired 360° panoramic images prior to subsequent use thereof (e.g., for generating associated mapping information, for providing information about structural elements or other objects in a room or other enclosed area, etc.) in order to generate images of a particular type and / or a particular format (e.g., performing an equirectangular projection on each such image, having straight vertical data, such as the sides of a typical rectangular door frame or a typical boundary between two adjacent walls that remains straight, and having straight horizontal data, such as the top of a typical rectangular door frame or a boundary between a wall and a floor that remains straight at the horizontal midline of the image, but gradually curves in a convex manner relative to the horizontal midline in the equirectangular projection image as the distance from the horizontal midline increases in the image. In block 577, the images and any associated generated or obtained information are stored for later use and optionally provided to one or more recipients (e.g., provided to block 430 of routine 400 if called from this block). FIG. 6A to FIG. 6B One example of a routine for generating a floor plan representation of a building interior from generated panoramic information is shown.
[0097] If it is determined in block 510 that the instruction or other information received in block 505 is not to acquire images and other data representing the interior of a building, the routine continues to block 590 to perform any other indicated operations, as appropriate, to configure parameters to be used in various operations of the system (e.g., based at least in part on information specified by a user of the system (e.g., based at least in part on information specified by a user of the system, such as a user of a mobile device capturing one or more building interiors, an operator user of the ICA system, etc.), in response to a request to generate and store information (e.g., identifying one or more sets of interconnected linked panoramic images, each of which represents a building or a portion of a building that matches one or more specified search criteria) , one or more panoramic images matching one or more specified search criteria, etc.), to generate and store panoramic image-to-panoramic connections between panoramic images of buildings or other structures (e.g., for each panoramic image, determining a direction within the panoramic image toward one or more other acquisition locations of one or more other panoramic images, so as to enable later display of an arrow or other visual representation with the panoramic image, and for each such determined direction from the panoramic image, enabling an end user to select one of the displayed visual representations to switch to display of another panoramic image at another acquisition location corresponding to the selected visual representation), to obtain and store other information about users of the system, to perform any housekeeping tasks, etc.
[0098] After frame 577 or frame 590, routine proceeds to frame 595 to determine whether to continue, for example, until receiving explicit instruction to terminate, or only when receiving explicit instruction to continue. If it is determined to continue, routine returns to frame 505 to wait for additional instructions or information, and if not, proceeds to step 599 and ends.
[0099] Although not targeted Figure 5 512 and 524, which are used as part of the automatic operation of the frame; performing activities related to image acquisition in frame 515 (e.g., participating in image acquisition, such as activating a shutter, implementing settings on a camera and / or associated sensors or components, rotating a camera as part of capturing a panoramic image, etc.; setting the position and / or orientation of one or more camera devices and / or associated sensors or components, etc.); providing input in frames 515 and / or 522, which are used as part of subsequent automatic operations, such as labels, annotations or other descriptive information about a particular image, surrounding room and / or objects in the room, etc. Additional details about embodiments in which one or more human users provide input are included elsewhere herein, and the input is further used for additional automatic operations of the ICA system.
[0100] FIG. 6A to FIG. 6B An exemplary embodiment of a flow chart of a MIGM (Mapping Information Generation Manager) system routine 600 is shown. The routine may be executed by, for example, executing the MIGM system 160 of FIG. 1, Figure 3 MIGM system 389, and / or as regards Figures 2A to 2V and the MIGM system described elsewhere herein to determine a room shape of a room (or other defined area) by analyzing information from one or more images captured in the room (e.g., one or more 360° panoramic images) to generate a partial or complete floor plan of a building or other defined area based at least in part on the one or more images of the area and optionally additional data captured by a mobile computing device, and use the determined room shape to generate other mapping information of a building or other defined area based at least in part on the one or more images of the area and optionally additional data captured by a mobile computing device. FIG. 6A to FIG. 6BIn the example of , the determined room shape of the room can be a 2D room shape representing the locations of the walls of the room or a 3D fully closed combination of planar surfaces representing the locations of the walls, ceiling, and floor of the room, and the generated mapping information for the building (e.g., a house) can include a 2D floor plan and / or a 3D computer model floor plan. However, in other embodiments, other types of room shapes and / or mapping information can be generated and used in other ways, including for other types of structures and defined areas, as discussed elsewhere herein. In at least some embodiments, the room shape can be generated from FIG. 4A to FIG. 4D Block 440 of routine 400 of the embodiment of the present invention calls routine 600, provides corresponding information from routine 600 to routine 400 as part of the implementation of such block 440, and in such case, returns processing control to routine 400 after blocks 688 and / or 699. In other embodiments, routine 400 may perform additional operations in an asynchronous manner without waiting for such a return of processing control (e.g., proceed to block 445 once corresponding information from routine 600 is provided to routine 400, perform other processing activities while waiting for corresponding information from routine 600 to be provided to routine 400, etc.).
[0101] The illustrated implementation of the routine begins at box 605, where information or instructions are received. The routine continues to box 610 to determine whether image information is already available for analysis of one or more rooms (e.g., for some or all of the indicated building, such as based on one or more such images received in box 605 as previously generated by the ICA routine), or whether such image information is currently to be acquired. If it is determined in box 610 that some or all of the image information is currently to be acquired, the routine continues to box 612 to acquire such information, optionally waiting for one or more users or devices to move throughout one or more rooms of the building and acquire panoramic images or other images at one or more acquisition locations in the one or more rooms (e.g., multiple acquisition locations in each room of the building), optionally together with metadata information about the acquisition and / or interconnect information related to movement between acquisition locations, as discussed in more detail elsewhere herein. Implementation of box 612 may, for example, include calling an ICA system routine to perform such an activity, wherein Figure 5One exemplary embodiment of an ICA system routine for performing such image acquisition is provided. If it is determined in block 610 that an image is not currently being acquired, the routine continues to block 615 to obtain one or more existing panoramic images or other images from one or more acquisition locations in one or more rooms (e.g., multiple images acquired at multiple acquisition locations including at least one image and acquisition location in each room of a building), optionally along with metadata information about the acquisition and / or interconnected information related to movement between acquisition locations, such as, in some cases, has been provided in block 605 along with corresponding instructions.
[0102] After either box 612 or 615, the routine continues to box 620 where it determines whether mapping information for a set of inter-links (sometimes referred to as a "virtual tour" to enable an end user to move from any one of the images in the linked set to one or more other images linked to the starting current image, including, in some embodiments, by selecting a user-selectable control for each such other linked image displayed with the current image, optionally by overlaying a visual representation of such user-selectable controls and corresponding inter-image directions over the visual data of the current image, and similarly moving from the next image to one or more additional images linked to the next image, and so on) is generated, and if so, continues to box 625. The routine in box 625 selects at least some pairs of images (e.g., based on images of the pair having overlapping visual content), and based on shared visual content and / or based on other captured linked interconnection information (e.g., movement information) associated with the images of the pair (whether moving directly from an acquisition location of one image in a pair to an acquisition location of another image in the pair, or moving between these starting and ending acquisition locations via one or more other intermediate acquisition locations of other images). The routine in box 625 may further optionally use relative orientation information of at least the paired images to determine the global relative positions of some or all images relative to each other in a common coordinate system, and / or generate inter-image links and corresponding user-selectable controls as described above. Additional details on creating such linked sets of images are included elsewhere herein.
[0103] Following block 625, or if it is determined in block 620 that the instructions or other information received in block 605 is not to determine a linked set of images, the routine continues to block 635 to determine whether the instructions received in block 605 indicate to generate other mapping information (e.g., a floor plan) for the indicated building, and if so, the routine continues to perform some or all of blocks 637-685 to do so, and otherwise continues to block 690. In block 637, the routine optionally obtains additional information about the building, such as from activities performed during acquisition and optionally analysis of the images, and / or from one or more external sources (e.g., online databases, information provided by one or more end users, etc.). Such additional information may include, for example, the exterior dimensions and / or shape of the building, additional imagery and / or annotated information acquired corresponding to specific locations on the exterior of the building (e.g., surrounding the building and / or for other structures on the same property, from one or more overhead locations, etc.), additional imagery and / or annotated information acquired corresponding to specific locations within the building (optionally, for locations different from the location at which the panoramic image or other image was acquired), etc.
[0104] After box 637, the routine continues to box 640 to select the next room (starting with the first one) for which one or more images (e.g., 360° panoramic images) acquired in the room are available, and analyze the visual data of the images for the room to determine the room shape (e.g., by determining at least the wall locations), optionally together with determining uncertainty information about the walls and / or other portions of the room shape, and optionally including identifying other wall and floor and ceiling elements (e.g., wall structural elements / objects such as windows, doorways and stairs and other inter-room wall openings and connecting passages, wall boundaries between a wall and another wall and / or ceiling and / or floor, etc.) and their locations within the determined room shape of the room. In some embodiments, room shape determination may include determining a 2D room shape using the boundaries of the walls to one another and at least one of the floor or ceiling (e.g., using one or more trained machine learning models), while in other embodiments, room shape determination may be performed in other ways (e.g., by generating a 3D point cloud of some or all of the room walls and optionally the ceiling and / or floor. For example, by analyzing at least visual data of the panoramic image and optionally additional data captured by the image acquisition device or an associated mobile computing device, optionally using one or more of SFM (Structure from Motion) or SLAM (Simultaneous Localization and Mapping) or MVS (Multi-View Stereo) analysis. In addition, the activity of block 645 Initial pose information for each of those panoramic images (e.g., provided with acquisition metadata for the panoramic images), and / or additional metadata for each panoramic image (e.g., acquisition height information for a camera device or other image acquisition device used to acquire the panoramic images relative to the floor and / or ceiling) may also be optionally determined and used. Additional details regarding determining room shapes and identifying additional information for rooms are included elsewhere herein. After box 640, the routine continues to box 645, where it determines whether there are more rooms for which room shapes have been determined based on the images acquired in those rooms, and if so, returns to box 640 to select the next such room for which room shape has been determined.
[0105] If it is determined in block 645 that there are no more rooms for which room shapes are generated, the routine continues to block 660 to determine whether to further generate at least a partial floor plan of the building (e.g., based at least in part on the determined room shapes from block 640, and optionally further information about how the determined room shapes are positioned relative to each other). If not, such as when only one or more room shapes are determined without generating further mapping information for the building (e.g., the room shape of a single room is determined based on one or more images acquired in the room by the ICA system), the routine continues to block 688. Otherwise, the routine continues to block 665 to retrieve one or more room shapes (e.g., the room shapes generated in block 645), or otherwise obtain one or more room shapes for the rooms of the building (e.g., based on manually provided input), whether 2D or 3D room shapes, and then continues to block 670. In block 670, the routine uses the one or more room shapes to create an initial floor plan (e.g., an initial 2D floor plan using 2D room shapes and / or an initial 3D floor plan using 3D room shapes), such as a partial floor plan that includes one or more room shapes but less than all room shapes of the building, or a complete floor plan that includes all room shapes of the building. If multiple room shapes are present, the routine further determines in block 670 the positioning of the room shapes relative to each other, such as by using visual overlap between images from multiple acquisition locations to determine the relative positions of those acquisition locations and room shapes surrounding those acquisition locations, and / or by using other types of information (e.g., using inter-room pathways connecting rooms, optionally applying one or more constraints or optimizations, etc.). In at least some embodiments, the routine in block 670 further refines some or all of the room shapes by generating a binary segmentation mask of overlapping relatively positioned room shapes, extracting polygons representing the outlines or contours of the segmentation mask, and separating the polygons into refined room shapes. Such a floor plan may include, for example, relative position and shape information for various rooms without providing any actual size information for the individual rooms or the building as a whole, and may also include multiple linked or associated sub-maps of the building (e.g., to reflect different floors, levels, sections, etc.). The routine also optionally associates the locations of doors, wall openings, and other identified wall elements on the floor plan.
[0106] After box 670, the routine optionally performs one or more steps 680 to 685 to determine and associate additional information with the floor plan. In box 680, the routine optionally estimates the dimensions of some or all rooms, such as from analysis of the images and / or their acquisition metadata or from general dimensional information obtained for the exterior of the building, and associates the estimated dimensions with the floor plan. It should be understood that if sufficiently detailed dimensional information is available, a building map, blueprint, etc. can be generated from the floor plan. After box 680, the routine continues to box 683 to optionally associate further information with the floor plan (e.g., with a specific room or other location within the building), such as additional existing images with specified locations and / or annotation information. In box 685, if the room shape from box 645 is not a 3D room shape, the routine also optionally estimates the height of the walls in some or all rooms, such as based on analysis of the image and optionally the size of known objects in the image, and height information about the camera when the image was acquired, and uses the height information to generate a 3D room shape for the room. The routine further optionally uses the 3D room shapes (whether from block 640 or block 685) to generate a 3D computer model floor plan of the building, wherein the 2D and 3D floor plans are associated with each other. In other embodiments, only the 3D computer model floor plan may be generated and used (including, if desired, by using a horizontal slice of the 3D computer model floor plan to provide a visual representation of the 2D floor plan).
[0107] Following box 685, or if it is determined in box 660 that a floor plan is not to be determined, the routine continues to box 688 to store the determined room shape and / or the generated mapping information and / or other generated information, optionally provide some or all of the information to one or more recipients (e.g., to box 440 of routine 400 if called from that box), and optionally further use some or all of the determined and generated information to provide a generated 2D floor plan and / or 3D computer model floor plan for display on one or more client devices and / or to one or more other devices, for use in automating navigation of such devices and / or associated vehicles or other entities, to similarly provide and use information about the determined room shape and / or linked collection of images and / or about determined additional information about room contents and / or pathways between rooms, etc.
[0108] If it is determined in box 635 that the information or instructions received in box 605 are not to generate mapping information for the indicated building, the routine continues to box 690 to perform one or more other indicated operations as appropriate. Such other operations may include, for example, receiving and responding to requests for previously generated floor plans and / or previously determined room shapes and / or other generated information (e.g., requests for such information to be displayed on one or more client devices, requests for such information to be provided to one or more other devices for use in automatic navigation, etc.), obtaining and storing information about the building for use in subsequent operations (e.g., information about the size, number or type of rooms, total square length, other adjacent or nearby buildings, adjacent or nearby vegetation, exterior images, etc.), etc.
[0109] After frame 688 or frame 690, routine continues to frame 695, to determine whether to continue, for example, until receiving explicit instruction to terminate, or only when receiving explicit instruction to continue. If it is determined to continue, routine returns to frame 605 to wait and receive additional instructions or information, and otherwise continues to frame 699 and ends.
[0110] Although not targeted FIG. 6A to FIG. 6BThe automated operations shown in the exemplary embodiments of the present invention are described, but in some embodiments, a human user can further help facilitate some operations of the MIGM system, such as providing one or more types of input to an operator user and / or end user of the MIGM system, which one or more types are further used in subsequent automated operations. As non-exclusive examples, such a human user can provide one or more types of input as follows: provide input to help link a collection of images so as to provide input in box 625, which is used as part of the automated operation of that box (e.g., specifying or adjusting an initially automatically determined orientation between one or more pairs of images, specifying or adjusting an initially automatically determined final global position of some or all images relative to each other, etc.); provide input in box 637, which is used as part of subsequent automated operations, such as information of one or more of the shown types about a building; provide input with respect to box 640 used as part of subsequent automated operations to specify or adjust initially automatically determined element locations and / or estimated room shapes, and / or manually combine information from multiple estimated room shapes of a room (e.g., providing input to box 670 used as part of subsequent operations to specify or adjust the initial automatically determined location of the room shape within the generated floor plan and / or to specify or adjust the initial automatically determined room shape itself within such floor plan; providing input to one or more of boxes 680 and 683 and 685 used as part of subsequent operations to specify or adjust initial automatically determined information of one or more types discussed with respect to those boxes; and / or specifying or adjusting initial automatically determined pose information (either initial pose information or subsequently updated pose information) for one or more of the panoramic images, etc. Additional details regarding embodiments in which one or more human users provide input that is further used in additional automatic operations of the MIUM system are included elsewhere herein.
[0111] FIG. 7A to FIG. 7B An exemplary embodiment of a flowchart for a building information access system routine 700 is shown. The routine may be executed by, for example, the building information access client computing device 175 of FIG. 1 and its software system (not shown), Figure 3The client computing device 390 of the present invention and / or a building information access viewer or presentation system as described elsewhere herein is executed to request and receive and present building information (e.g., video; individual images; floor plans and / or other mapping-related information, such as determined room structure layout / shape, virtual tours of linked images, etc.; generated building description information, etc.), provide user input related to the video to be generated, obtain and display information about images that match one or more indicated target images, obtain and display guided acquisition instructions (e.g., about other images acquired during the acquisition session and / or about the associated building, such as a portion of a displayed GUI), etc. FIG. 7A to FIG. 7B In the example of , the information presented is for one or more buildings (such as the interior of a house), but in other embodiments, other types of mapping information may be presented for other types of buildings or environments and used in other ways, as discussed elsewhere herein.
[0112] The illustrated implementation of the routine begins at block 705, where instructions or information are received. At block 705, the routine determines whether the instructions or information received at block 705 are to obtain user input related to generating one or more videos of the indicated building, and if so, continues to execute blocks 805-840. Otherwise, the routine continues to block 710 to determine whether the instructions or information received in block 705 are to present information for determination of one or more target buildings, and if so, continues to block 715 to determine whether the instructions or information received in block 705 are to select one or more target buildings using specified criteria (e.g., based at least in part on the indicated building), and if not, continues to block 720 to obtain an indication from the user of the target building to use (e.g., based on a current user selection, such as from a displayed list or other user selection mechanism; based on the information received in block 705, etc.). Otherwise, if it is determined in block 715 that one or more target buildings are to be selected from the specified criteria (e.g., based at least in part on the indicated building), the routine continues to block 725 where it obtains an indication of one or more search criteria to be used, such as from the current user selection or as indicated in information or instructions received in block 705, and then searches stored information about buildings (e.g., floor plans, videos, generated text descriptions, etc.) to determine one or more buildings that meet the search criteria or otherwise obtains an indication of one or more such matching target buildings, such as information currently or previously generated by the BVGUM system (an example of the operation of such a system will be described with reference to FIG. 1 ). FIG. 4A to FIG. 4D705 ), while in other embodiments, the routine may alternatively present information for multiple target buildings that meet the search criteria (e.g., in a ranked order based on degree of match; presenting one or more videos for each of multiple buildings in a sequence in a sequential manner, etc.), and receive a user selection of the best matching building from the multiple candidate target buildings.
[0113] After box 720 or 725, the routine continues to box 730 to determine whether the instruction or other information received in box 705 indicates to present one or more generated videos for each of the one or more target buildings, and if so, continues to box 732 to do so, including retrieving one or more existing generated videos for each target building (e.g., one or more existing generated videos that match the criteria specified in the information of box 705 or one or more existing generated videos determined in other ways, such as using preference information or other information specific to the recipient), or alternatively, in some embodiments and situations, requesting dynamically generated videos (e.g., by interacting with the BVGUM system to cause such generation, whether by newly generating a video or by modifying an existing video, and optionally providing one or more criteria used in such generation, such as using preference information or other information specific to the recipient), and initiating presentation of the retrieved and / or dynamically generated videos (e.g., sending the one or more videos to one or more client devices for presentation on those devices). After box 732, the routine continues to box 795.
[0114] If it is determined in box 730 that the instructions or other information received in box 705 do not indicate that one or more generated videos are to be presented, the routine continues to box 735 to retrieve information of the target building for display (e.g., a floor plan; other generated mapping information for the building, such as a set of linked images used as part of a virtual tour; generated building description information, etc.), and optionally associated link information indicating surrounding locations for the building interior and / or building exterior, and / or information about one or more generated interpretations or other descriptions of the target building, and select an initial view of the retrieved information (e.g., a floor plan, a view of a specific room shape, a specific image, some or all of the generated building description information, etc.). In box 740, the routine then displays or otherwise presents the current view of the retrieved information, and waits for a user selection in box 745. After the user selection in block 745, if it is determined in block 750 that the user selection corresponds to adjusting the current view of the current target building (e.g., changing one or more aspects of the current view), the routine continues to block 755 to update the current view in accordance with the user selection, and then returns to block 740 to update the displayed or otherwise presented information accordingly. The user selection and corresponding update of the current view may include, for example, displaying or otherwise presenting a piece of associated link information selected by the user (e.g., a particular image associated with the displayed visual indication of the determined acquisition location so that the associated link information is overlaid on at least some of the previously displayed; a particular other image that is linked to the current image and selected from the current image using a user-selectable control overlaid on the current image to represent the other image; etc.), and / or changing how the current view is displayed (e.g., zooming in or out; rotating the information if appropriate; selecting a new portion of the plan view to be displayed or otherwise presented, such as where some or all of the new portion was not previously visible, or where instead the new portion is a subset of the previously visible information; etc.). If it is determined in box 750 that the user chooses not to display further information about the current target building (e.g., display information for another building, end the current display operation, etc.), the routine continues to box 795, and if the user selection involves such further operation, returns to box 705 to perform the operation selected by the user.
[0115] If it is determined in box 710 that the instruction or other information received in box 705 is not to present information representing a building, the routine continues to box 760 to determine whether the instruction or other information received in box 705 indicates to identify other images (if any) corresponding to one or more indicated target images, and if so, continues to boxes 765-770 to perform such activities. In particular, in box 765, the routine receives an indication of one or more target images for matching (e.g., from the information received in box 705 or based on one or more current interactions with the user) and one or more matching criteria (e.g., the amount of visual overlap), and in box 770, identifies one or more other images (if any) that match the indicated target images, such as by interacting with the ICA and / or MIGM system to obtain the other images. Then, the routine displays or otherwise provides information about the identified other images in box 770, so as to provide information about them as part of the search results, to display one or more of the identified other images, etc. If it is determined in block 760 that the instructions or other information received in block 705 are not to identify other images corresponding to one or more indicated target images, the routine continues to block 775 to determine whether the instructions or other information received in block 705 correspond to guided acquisition instructions to obtain and provide information about one or more indicated target images (e.g., most recently acquired images) during the image acquisition period, if so, continue to block 780, and otherwise continue to block 790. In block 780, the routine obtains information about one or more types of guided acquisition instructions, such as by interacting with the ICA system, and displays or otherwise provides information about the guided acquisition instructions in block 780, such as by overlaying the guided acquisition instructions on the partial plan view and / or on the most recently acquired image in a manner discussed in more detail elsewhere herein.
[0116] If it is determined in box 707 that the instructions or information received in box 705 will obtain user input related to generating one or more videos of the indicated building, the routine continues to execute boxes 805-840 to do so. In particular, the routine in box 805 receives an indication of a building, such as based on the input provided in box 705 or based on information specified in box 805 by one or more users via a GUI presented to one or more users, so as to select the building from a list and / or from search results. In box 810, the routine then retrieves information about the building, which includes a floor plan of the building and information about building properties of the building, and optional additional building information (e.g., one or more existing videos of the building, images acquired at the building, etc.), and in box 815, if there are any existing videos and the retrieved information does not include a visual representation of such an existing video (e.g., the path traveled by the camera in the associated path through the building while capturing the visual data of the video), one or more such visual representations are generated for each such video. In box 820, the routine then presents information about the building to one or more users in the displayed GUI, such as a visual representation of the floor plan overlaid with a visual representation of the existing video (if any), and optionally information about building properties and building images that are displayed at associated locations on the floor plan or otherwise provided in association with the floor plan (e.g., to provide a list or other group of building images associated with rooms or other areas of the building, to provide a list or other group of building properties associated with rooms or other areas of the building, etc.). The routine also provides one or more user-selectable controls to enable one or more users to provide one or more segment criteria for generating one or more additional videos from one or more existing videos and / or provide one or more generation criteria for generating one or more new videos based at least in part on visual data of at least some of the building images, wherein the user input may include, for example, indicating one or more portions of one or more existing videos for generating corresponding additional videos and / or specifying the type of information to include in the new video to be generated (e.g., by selecting portions of the presented visual representation of the existing videos, by specifying one or more rooms or other areas for which visual data is to be included in additional videos generated from the existing videos or in a new video generated from the images, by specifying one or more building attributes for which visual data is to be included in additional videos generated from the one or more existing videos or in a new video generated from the images, etc.).In box 825, the routine then receives user input that uses one or more user-selectable controls to specify one or more segment criteria and / or one or more generation criteria for generating one or more additional videos and / or new videos, respectively, and in box 430, provides the user input to the BVGUM system routine to cause the generation of one or more additional videos and / or new videos, respectively, based on the specified segment criteria and / or generation criteria, wherein the user input is related to. FIG. 4A to FIG. 4D An example of such a routine is described. In block 840, the routine receives one or more generated videos and optionally presents one or more videos and / or information about the generated videos to the user (e.g., by updating the presentation of the floor plan in the GUI to add a visual representation of the path and / or other information about each such generated video).
[0117] In box 790, the routine continues to perform the other indicated operations as appropriate to configure parameters to be used in various operations of the system (e.g., based at least in part on information specified by users of the system, such as obtaining users of mobile devices within one or more buildings, operator users of BVGUM and / or MIGM systems, etc., including for personalizing the display of information for a particular user based on the preferences of the particular recipient user or other information specific to the recipient), to obtain and store other information about users of the system (e.g., preferences or other information specific to the recipient), to respond to requests for generated and stored information, to perform any housekeeping tasks, etc.
[0118] After box 732 or box 770 or box 780 or box 790 or box 840, or if it is determined in box 750 that the user selection does not correspond to the current building, the routine proceeds to box 795 to determine whether to continue, such as until an explicit indication to terminate is received, or only if an explicit indication to continue is received. If it is determined to continue (including whether the user makes a selection related to a new building to be presented in box 745), the routine returns to box 705 to wait for additional instructions or information (or if the user makes a selection related to a new building to be presented in box 745, then continue directly to box 735), and if not, proceed to step 799 and end.
[0119] Non-exclusive exemplary embodiments described herein are further described in the following clauses.
[0120] A01. A computer-implemented method for one or more computing devices to perform an automated operation, comprising:
[0121] obtaining, by the one or more computing devices, data for an indicative house having a plurality of rooms, the data comprising a plurality of images acquired at a plurality of acquisition locations of the indicative house, and further comprising a floor plan for the indicative house, the floor plan indicating a layout of the plurality of rooms by at least two-dimensional room shapes having structural elements of the plurality of rooms and placed at relative locations of the plurality of rooms and having associated locations of the plurality of acquisition locations, and the data further comprising a building video captured at the indicative house along a path passing through at least two of the plurality of rooms;
[0122] Generate, by the one or more computing devices and based on the obtained data, a first video for at least a first room of the at least two rooms, the first video having visual data of the first room, including:
[0123] analyzing, by the one or more computing devices, the building video to determine a subset of the building video corresponding to the first room, including determining that visual data of the subset of the building video matches additional visual data of at least one of the plurality of images whose respective acquisition location is in the first room;
[0124] as well as
[0125] using, by the one or more computing devices, the determined subset of the building video as at least a portion of the additional video of the first room;
[0126] generating, by the one or more computing devices and based on the obtained data, a second video along the path through the portion of the house, the second video having at least visual data of one or more rooms through which the path passes, comprising:
[0127] Analyzing, by the one or more computing devices, the floor plan to determine a plurality of the structural units in the one or more rooms, and determining a subset of the plurality of images having a plurality of images, each of the plurality of images being acquired at a location in the one or more rooms; and
[0128] combining the visual data of the plurality of images with visual data of a plurality of structural elements determined in the one or more rooms to generate the second video;
[0129] presenting, by the one or more computing devices, a visual representation of at least a portion of the floor plan including the first room and the one or more rooms, wherein in presenting the visual representation of the first video, first user-selectable information is overlaid on the first room, and in presenting the visual representation of the second video, second user-selectable information is overlaid on the one or more rooms;
[0130] presenting, by the one or more computing devices and in response to a first user selecting the first user-selectable information, the first video; and
[0131] The second video is presented by the one or more computing devices and in response to a second user selecting the second user-selectable information.
[0132] A02. A computer-implemented method for one or more computing devices to perform an automated operation, comprising:
[0133] obtaining, by the one or more computing devices, data for an indicative building having a plurality of rooms, the data comprising a plurality of images acquired at a plurality of acquisition locations of the indicative building, and further comprising a floor plan for the indicative building, the floor plan having at least two-dimensional room shapes of the plurality of rooms and having associated locations on the floor plan of the plurality of acquisition locations, and the data further comprising a video of the building captured at the indicative building along a path through at least two of the plurality of rooms;
[0134] Generate, by the one or more computing devices, an additional video having at least visual data of one of the at least two rooms based on the obtained data, the additional video comprising:
[0135] determining, by the one or more computing devices and based at least in part on the floor plan, at least one image of the plurality of images, wherein a corresponding acquisition location of the at least one image is in the one room;
[0136] Analyzing, by the one or more computing devices, the building video to determine a subset of the building video corresponding to the one room, including determining that visual data of the subset of the building video matches the determined additional visual data of the at least one image; and
[0137] Using, by the one or more computing devices, the determined subset of the building video as at least part of the additional video of the one room; and presenting, by the one or more computing devices, information about the generated additional video on a portion of the floor plan corresponding to the one room.
[0138] A03. A computer-implemented method for one or more computing devices to perform an automated operation, comprising:
[0139] acquiring data for an indicative building having a plurality of rooms, the data comprising a plurality of images acquired at a plurality of acquisition locations of the indicative building, and further comprising a floor plan of the indicative building having at least two-dimensional room shapes and associated positions on the floor plan of the plurality of acquisition locations, and the data further comprising a building video having at least visual data along a path through at least some of the indicative building;
[0140] Based on the obtained data, an additional video is generated, the additional video having at least visual data for the area indicating the building through which the path passes, including:
[0141] Analyzing the building video to determine a subset of the building video corresponding to the area includes at least one of: determining that visual data of the subset of the building video matches additional visual data of at least one of the plurality of images whose respective acquisition locations are in the area; or determining that visual data of the subset of the building video shows structural elements of buildings included in the subset of the floor plans corresponding to the area;
[0142] or determining that visual data of a subset of the building video includes one or more objects associated with a room type of a room associated with the area; and
[0143] using the determined subset of the building video as at least part of an additional video of the area; and
[0144] Information about the generated additional video is provided that is associated with the subset of the floor plan corresponding to the area.
[0145] A04. A computer-implemented method for one or more computing devices to perform an automated operation, comprising:
[0146] obtaining, by the one or more computing devices, data for an indicative building having a plurality of rooms, the data comprising a plurality of images acquired at a plurality of acquisition locations of the indicative building, and further comprising a floor plan of the indicative building, the floor plan having at least two-dimensional room shapes, the at least two-dimensional room shapes having structural elements of the plurality of rooms and having associated locations on the floor plan at the plurality of acquisition locations;
[0147] Generating, by the one or more computing devices, a new video along a path through the portion of the building based on the obtained data, the new video having at least visual data for at least one room through which the path passes, comprising:
[0148] analyzing, by the one or more computing devices, the floor plan to determine a plurality of structured elements in the at least one room, and determining a subset of the plurality of images having respective acquisition locations in the at least one room, the subset of images comprising a plurality of images; and
[0149] combining visual data of the plurality of images of the subset to generate a new video having visual data of the determined plurality of structural elements in the at least one room; and
[0150] Information about the generated new video is presented, by the one or more computing devices, on a portion of the floor plan corresponding to the at least one room.
[0151] A05. A computer-implemented method as described in any of clauses A01-A04, wherein, in the presented visual representation, presenting a visual representation of at least a portion of the floor plan having first user-selectable information superimposed on the first room includes: presenting a group of multiple media associated with the first room, the group including at least the first video and the at least one image, and the first user-selectable information includes one or more controls for playing the first video.
[0152] A06. A computer-implemented method as described in any of clauses A01-A05, wherein, in the presented visual representation, presenting a visual representation of at least a portion of the floor plan having second user-selectable information overlaid on the one or more rooms includes: presenting a path through the one or more rooms, and wherein presenting the second video includes: in response to user input including a location on the presented path, starting presentation at a point within the second video corresponding to the location on the presented path.
[0153] A07. The computer-implemented method of clause A06, further comprising: prior to presenting the visual representation of the at least portion of the floor plan having the overlapping first user-selectable information and the overlapping second user-selectable information:
[0154] presenting, by the one or more computing devices, an initial visual representation of some or all of the floor plan including at least the one or more rooms;
[0155] receiving, by the one or more computing devices, user input regarding the presented initial visual representation associated with the second video to be generated, the user input comprising at least one of: a first specification regarding the initial visual representation of the path; or a second specification regarding the initial visual representation of the one or more rooms; or a third specification regarding the initial visual representation of one or more objects in the one or more rooms,
[0156] Wherein, the generation of the second video is performed in response to the received user input.
[0157] A08. A computer-implemented method as described in any of clauses A01-A07, wherein analyzing the building video to determine the subset of the building video corresponding to the one room includes: obtaining information about structural elements visible in the one room from the floor plan, and analyzing visual data of the subset of the building video to identify one or more of the structural elements.
[0158] A09. A computer-implemented method as described in any of clauses A01-A08, wherein analyzing the building video to determine a subset of the building video corresponding to the one room includes: for each of at least some frames of the subset of the building video, comparing the frame with one or more of the at least one image to identify matching visual elements in the frame and in the one or more images.
[0159] A10. A computer-implemented method as described in any of clauses A01-A09, wherein the information about the generated additional video includes one or more user-selectable controls corresponding to the playback of the generated additional video, wherein the presenting comprises: sending, by the one or more computing devices and over one or more computer networks, the information about the additional video generated on the portion of the floor plan corresponding to the one room to one or more client devices so that a visual representation of the information and the one or more user-selectable controls is presented on the one or more client devices, and wherein the method further comprises:
[0160] receiving, by the one or more computing devices, user input corresponding to a selection of at least one of the one or more user-selectable controls; and
[0161] At least some of the generated additional videos are sent to the one or more client devices via the one or more computing devices and via the one or more computer networks in response to the user input, so that at least some of the generated additional videos are presented on the one or more client devices.
[0162] A11. A computer-implemented method as described in any of clauses A01-A10, wherein the one or more computing devices include a server computing device and a client computing device of a user, and wherein the method further comprises:
[0163] receiving, by the server computing device, a request from the client computing device for video information of the area indicating the building;
[0164] performing, by a server computing device, generating the additional video and providing the information in response to the request, including sending the provided information to the client computing device via one or more computer networks; and
[0165] The transmitted information is received by the client computing device, and the received transmitted information is presented on the client computing device.
[0166] A12. A computer-implemented method as described in any of clauses A01-A11, wherein the area indicating the building includes one of the multiple rooms, and wherein analyzing the building video to determine a subset of the building video corresponding to the area includes: using the floor plan to determine at least one image whose respective acquisition locations are in the area, and determining that the visual data of the subset of the building video matches the additional visual data of the determined at least one image.
[0167] A13. A computer-implemented method as described in any of clauses A01-A12, wherein the area indicative of the building includes one of the plurality of rooms, and wherein analyzing the building video to determine a subset of the building video corresponding to the area includes: using the floor plan to determine structural elements in the area, and determining that the visual data of the subset of the building video includes at least one of the determined structural elements.
[0168] A14. A computer-implemented method as described in any of clauses A01-A13, wherein providing the information includes presenting information about the generated additional video on a portion of the floor plan corresponding to the one room, the presentation comprising at least one of the following: presenting at least one of a set of multiple media associated with the one room on the presented portion of the floor plan, the set of multiple media including at least the generated additional video and the at least one image; or presenting a visual representation of a path overlaid on the presented portion of the floor plan.
[0169] A15. A computer-implemented method as described in clause A14, wherein the presenting includes: presenting a visual representation of the path superimposed on the presented portion of the plan view, the presented visual representation being user-selectable, and wherein the method further includes:
[0170] receiving user input comprising a selection of a location on the presented visual representation of the path; and
[0171] In response to the user input, at least some of the generated additional videos are presented, beginning at points in the generated additional videos that correspond to locations on the presented visual representation of the path.
[0172] A16. A computer-implemented method as described in any of clauses A14-A15, wherein the presenting comprises: presenting the set of multiple media associated with the one room on the presented portion of the floor plan, including providing one or more user-selectable controls to the presented set of multiple media, and wherein the method further comprises:
[0173] receiving user input comprising a selection of at least one of the user-selectable controls associated with the generated additional video; and
[0174] In response to the user input, at least some of the generated additional videos are presented.
[0175] A17. A computer-implemented method as described in any of clauses A01-A16, wherein providing the information includes presenting some of the generated additional videos, the presenting including selecting a subset of each frame of some of the generated additional videos to be displayed, and also including receiving user input indicating a target orientation, and also including presenting an additional portion of the generated additional video by using the target orientation to select a corresponding additional subset of each frame of the additional portion to be displayed.
[0176] A18. A computer-implemented method as described in any of clauses A01-A17, wherein the method further includes identifying one or more target attributes of the buildings in the area, and wherein providing the information includes presenting at least some of the generated additional videos in which the identified one or more target attributes are displayed, the presenting including selecting a subset of each frame of the at least some of the generated additional videos for display.
[0177] A19. A computer-implemented method as described in any of clauses A01-A18, wherein the area includes one of the plurality of rooms, and wherein analyzing the building video to determine the subset of the building video corresponding to the area includes: analyzing the visual data of the subset of the video to determine at least one of the beginning of the subset or the end of the subset based at least in part on identifying at least one room-to-room transition along the path.
[0178] A20. A computer-implemented method as described in any of clauses A01-A19, wherein the method further comprises: determining an area for the generated additional video based at least in part on an analysis of the building video to detect motion patterns along the path that meet one or more defined criteria; and selecting the area to include a portion of the path corresponding to the detected motion pattern.
[0179] A21. A computer-implemented method as described in any of clauses A01-A20, wherein the area includes one of a plurality of rooms of the room type, and wherein analyzing the building video to determine a subset of the building video corresponding to the area includes: analyzing visual data of the subset of the video to identify the one or more objects associated with the room type.
[0180] A22. A computer-implemented method as described in any of clauses A01-A21, wherein the area includes one of two or more rooms through which the path passes, and wherein analyzing the building video to determine a subset of the building video corresponding to the area includes: using at least one of a SLAM (simultaneous location and mapping) technique during capture of the building video or a SFM (structure from motion) technique after capture of the building video to associate each frame of the building video with one of the two or more rooms, and selecting frames of the building video for the subset associated with the one room.
[0181] A23. A computer-implemented method as described in any of clauses A01-A22, wherein the method further includes: before generating the additional video, receiving user input from a user and determining an area for the additional video based at least in part on the user input, the user input comprising at least one of: a portion of the floor plan corresponding to an area selected by the user on a display of the floor plan; or a group of one or more building properties of a building selected by the user and located within the area; or a portion of a path corresponding to an area selected by the user on a display of the floor plan with a visual representation of the path overlaid thereon; or one or more rooms within the area selected by the user; or one or more exterior portions within the area that are outside the building and on the real estate on which the building is located selected by the user.
[0182] A24. A computer-implemented method as described in any of clauses A01-A23, wherein generating the additional video further comprises: for each of one or more objects that are not visible in the visual data of the subset, generating one or more visual representations of the object; and superimposing the generated one or more visual representations in one or more frames of the additional video at one or more indicated locations in the building.
[0183] A25. A computer-implemented method as described in any of clauses A01-A24, wherein the method further comprises:
[0184] receiving, by the one or more computing devices, a request from a client computing device for the video information indicative of an area of the building including the at least one room; and
[0185] The one or more computing devices, in response to the request, perform generation of the new video and presentation of the information, including sending information about the generated new video to the client computing device through one or more computer networks so that the sent information is displayed on the client computing device.
[0186] A26. A computer-implemented method as described in any of clauses A01-A25, wherein the method further comprises: before generating the new video, receiving user input from a user, the user input comprising the user's designation of the path on the display of the plan view.
[0187] A27. The computer-implemented method of any of clauses A01-A26, wherein the method further comprises, before generating the new video:
[0188] receiving user input from a user, the user input comprising at least one of: a portion of a floor plan selected by the user on a display of the floor plan; or a group of one or more building properties of a building selected by the user; or a selection of one or more rooms by the user; or a selection by the user of an exterior of a building and an exterior area on the property where the building is located; and
[0189] A path is determined based at least in part on user input, including: passing through the portion of the floor plan if selected by the user; or passing through the one or more building properties if selected by the user; or passing through the one or more rooms if selected by the user; or passing through the exterior area if selected by the user.
[0190] A28. A computer-implemented method as described in clause A27, wherein determining the path further comprises: using a trained machine learning model to select the path to at least one of be plausible for human movement or provide a visually pleasing result.
[0191] A29. A computer-implemented method as described in any of clauses A27-A28, wherein determining the path further comprises: for each of one or more positions along the path, determining a position from which to include visual data in the generated new video.
[0192] A30. A computer-implemented method as described in clause A29, wherein determining the orientation of one of the locations comprises: selecting the orientation of the one location to be at least one of: substantially tangent to the path at the one location; or substantially perpendicular to a tangent to the path at the one location; or pointing to at least one selected building attribute.
[0193] A31. A computer-implemented method as described in any of clauses A01-A30, wherein presenting information about the generated new video on the portion of the floor plan includes at least one of: presenting a set of multiple media associated with one of the at least one rooms, the set of multiple media including at least the generated new video and at least one image of the multiple images; or presenting a visual representation of the path overlaid on the presented portion of the floor plan.
[0194] A32. A computer-implemented method as described in clause A31, wherein the presenting includes presenting a visual representation of the path superimposed on the presented portion of the plan view, the presented visual representation being user-selectable, and wherein the method further comprises:
[0195] receiving user input comprising a selection of a location on the presented visual representation of the path; and
[0196] In response to the user input, at least some of the generated new video is presented, beginning at a point in the generated new video that corresponds to a location on the presented visual representation of the path.
[0197] A33. A computer-implemented method as described in any of clauses A31-A32, wherein the presenting includes presenting the set of multiple media associated with the one room on the presented portion of the floor plan, including providing one or more user-selectable controls to the presented set, and wherein the method further includes:
[0198] receiving user input comprising a selection of at least one of the user-selectable controls associated with the generated new video; and
[0199] In response to the user input, at least some of the generated new videos are presented.
[0200] A34. A computer-implemented method as described in any of clauses A01-A33, wherein presenting the information includes presenting some of the generated new videos, the presenting includes selecting a subset of each frame of the some of the generated new videos to be displayed, and also includes receiving user input indicating a target orientation, and also includes presenting an additional portion of the generated new video by using the target orientation to select a corresponding additional subset of each frame of the additional portion to be displayed.
[0201] A35. A computer-implemented method as described in any of clauses A01-A34, wherein generating the new video further comprises: for each of one or more objects that are not visible in the combined visual data of the multiple images, generating one or more visual representations of the object; and superimposing the generated one or more visual representations in one or more frames of the new video at one or more indicated locations in the building.
[0202] A36. A computer-implemented method comprising the steps of performing a plurality of automated operations that implement techniques substantially as disclosed herein.
[0203] B01. A non-transitory computer-readable medium having stored executable software instructions and / or other storage contents, wherein the stored executable software instructions and / or other storage contents enable one or more computing systems to perform automatic operations to implement the method of any of clauses A01-A29.
[0204] B02. A non-transitory computer-readable medium having stored executable software instructions and / or other storage contents, wherein the stored executable software instructions and / or other storage contents enable one or more computing systems to perform automatic operations that implement the techniques substantially as disclosed herein.
[0205] C01. One or more computing systems comprising one or more hardware processors and one or more memories having stored instructions, which, when executed by at least one of the one or more hardware processors, cause the one or more computing systems to perform automatic operations of a method implementing any of clauses A01-A29.
[0206] C02. One or more computing systems comprising one or more hardware processors and one or more memories having stored instructions that, when executed by at least one of the one or more hardware processors, cause the one or more computing systems to perform automated operations that implement the techniques substantially as disclosed herein.
[0207] D01. A computer program adapted to perform the method of any one of clauses A01-A29 when said computer program is run on a computer.
[0208] Aspects of the present disclosure are described herein with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present disclosure.It should be understood that each frame of the flowchart and / or block diagram and the combination of frames in the flowchart and / or block diagram can be implemented by computer-readable program instructions.It will be further understood that in some implementations, the functions provided by the above routines can be provided in an alternative manner, such as splitting between more routines, or merging into fewer routines.Similarly, in some implementations, the routines shown can provide more or less functions than described, such as when other routines shown lack or include such functions respectively, or when the amount of functions provided changes.In addition, although various operations can be shown as being performed in a particular manner (e.g., serial or parallel, or synchronous or asynchronous) and / or in a particular order, in other implementations, operations can be performed in other orders and other ways.Any data structure discussed above can also be constructed in different ways, such as by dividing a single data structure into multiple data structures and / or by merging multiple data structures into a single data structure. Similarly, in some implementations, illustrated data structures may store more or less information than depicted, such as when other illustrated data structures lack or include such information, respectively, or when the amount or type of information stored varies.
[0209] It will be appreciated from the above that, although specific embodiments are described herein for purposes of illustration, various modifications may be made without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited except by the corresponding claims and the elements cited by those claims. In addition, although certain aspects of the present invention may be presented in the form of certain claims at certain times, the inventors contemplate various aspects of the present invention in the form of any available claims. For example, although only some aspects of the present invention may be described as being embodied in a computer-readable medium at a particular time, other aspects may also be so embodied.
Claims
1. A computer-implemented method comprising: obtaining, by one or more computing devices, data for an indicative building having a plurality of rooms, the data comprising a plurality of images acquired at a plurality of acquisition locations of the indicative building, and further comprising a floor plan for the indicative building, the floor plan having at least two-dimensional room shapes of the plurality of rooms and having associated locations on the floor plan at the plurality of acquisition locations, and the data further comprising a video of the building captured at the indicative building along a path through at least two of the plurality of rooms; Generating, by the one or more computing devices and based on the obtained data, an additional video having at least visual data of one of the at least two rooms, comprising: determining, by the one or more computing devices and based at least in part on the floor plan, at least one image of the plurality of images, wherein a corresponding acquisition location of the at least one image is in the one room; Analyzing, by the one or more computing devices, the building video to determine a subset of the building video corresponding to the one room, including determining that visual data of the subset of the building video matches the determined additional visual data of the at least one image; and using, by the one or more computing devices, the determined subset of the building video as at least a portion of the additional video of the one room; and Information about the generated additional video is presented, by the one or more computing devices, on a portion of the floor plan corresponding to the one room.
2. The computer-implemented method of claim 1, wherein: Analyzing the building video to determine a subset of the building video corresponding to the one room includes obtaining information about structural elements visible in the one room from the floor plan and analyzing visual data of the subset of the building video to identify one or more of the structural elements.
3. The computer-implemented method of claim 1, wherein: Analyzing the building video to determine a subset of the building video corresponding to the one room includes: for each of at least some frames of the subset of the building video, comparing the frame to one or more of the at least one image to identify matching visual elements in the frame and in the one or more images.
4. The computer-implemented method of claim 1, wherein: The information about the generated additional video includes one or more user-selectable controls corresponding to playback of the generated additional video, wherein the presenting includes: sending, by the one or more computing devices and over one or more computer networks, the information about the additional video generated on the portion of the floor plan corresponding to the one room to one or more client devices so that a visual representation of the information and the one or more user-selectable controls is presented on the one or more client devices, and wherein the method further includes: receiving, by the one or more computing devices, user input corresponding to a selection of at least one of the one or more user-selectable controls; and At least some of the generated additional videos are sent to the one or more client devices via the one or more computing devices and via the one or more computer networks in response to the user input so that at least some of the generated additional videos are presented on the one or more client devices.
5. A system comprising: one or more hardware processors of one or more computing devices; as well as One or more memories having stored instructions that, when executed by at least one of the one or more hardware processors, cause at least one of the one or more computing devices to perform automated operations, the automated operations comprising at least: acquiring data for an indicative building having a plurality of rooms, the data comprising a plurality of images acquired at a plurality of acquisition locations of the indicative building and further comprising a floor plan of the indicative building having at least two-dimensional room shapes and having associated locations on the floor plan at the plurality of acquisition locations, and the data further comprising a building video having at least visual data along a path through at least some of the indicative building; Based on the obtained data, an additional video is generated, the additional video having at least visual data for the area indicating the building through which the path passes, including: Analyzing the building video to determine a subset of the building video corresponding to the area includes at least one of: determining that visual data of the subset of the building video matches additional visual data of at least one of a plurality of images of their respective acquisition locations in the area; or determining that the visual data of the subset of the building video shows structural elements of buildings included in the subset of floor plans corresponding to the area; or determining that the visual data of the subset of the building video includes one or more objects associated with a room type of a room associated with the area; and using the determined subset of the building video as at least part of an additional video of the area; and Information about the generated additional video is provided that is associated with the subset of the floor plan corresponding to the area.
6. The system of claim 5, wherein: The at least one computing device comprises a server computing device, and wherein the one or more computing devices further comprises a client computing device of a user, and wherein the stored instructions comprise software instructions that, when executed by the one or more computing devices, cause the one or more computing devices to perform further automated operations comprising: receiving, by the server computing device, from the client computing device, a request for video information of the area indicative of the building; performing, by the server computing device, generating the additional video and providing the information in response to the request, including sending the provided information to the client computing device via one or more computer networks; and The transmitted information is received by the client computing device, and the received transmitted information is presented on the client computing device.
7. The system of claim 5, wherein: The area indicative of the building includes one of the multiple rooms, and wherein analyzing the building video to determine a subset of the building video corresponding to the area includes: using the floor plan to determine the at least one image whose respective acquisition location is in the area, and determining that the visual data of the subset of the building video matches the additional visual data of the determined at least one image.
8. The system of claim 5, wherein: The area indicative of the building includes one of the multiple rooms, and wherein analyzing the building video to determine a subset of the building video corresponding to the area includes: using the floor plan to determine structural elements in the area, and determining that the visual data of the subset of the building video includes at least one of the determined structural elements.
9. The system of claim 5, wherein: Providing the information includes presenting information about the generated additional video on a portion of the floor plan corresponding to the one room, the presenting including presenting a visual representation of the path overlaid on the presented portion of the floor plan, the presented visual representation being user-selectable, and wherein the automatic operation further includes: receiving user input comprising a selection of a location on the presented visual representation of the path; and In response to the user input, at least some of the generated additional videos are presented, beginning at points in the generated additional videos that correspond to locations on the presented visual representation of the path.
10. The system of claim 5, wherein: Providing the information includes presenting information about the generated additional video on a portion of the floor plan corresponding to the one room, the presenting including presenting a group of a plurality of media associated with the one room on the presented portion of the floor plan, and providing the presented group to one or more user-selectable controls, and wherein the automatic operation further includes: receiving user input comprising a selection of at least one of the user-selectable controls associated with the generated additional video; and In response to the user input, at least some of the generated additional videos are presented.
11. The system of claim 5, wherein: Providing the information includes presenting some of the generated additional videos, the presenting including selecting a subset of each frame of some of the generated additional videos to be displayed, and also including receiving user input indicating a target orientation, and also including presenting an additional portion of the generated additional video by using the target orientation to select a corresponding additional subset of each frame of the additional portion to be displayed.
12. The system of claim 5, wherein: The automatic operation also includes identifying one or more target attributes of the building in the area, and wherein providing the information includes presenting at least some of the generated additional videos in which the identified one or more target attributes are displayed, the presenting including selecting a subset of each frame of the at least some of the generated additional videos for display.
13. The system of claim 5, wherein: The area includes one of the multiple rooms, and wherein analyzing the building video to determine the subset of the building video corresponding to the area includes: analyzing the visual data of the subset of the video to determine at least one of the beginning of the subset or the end of the subset based at least in part on identifying at least one room-to-room transition along the path.
14. The system of claim 5, wherein: The automatic operation also includes: determining an area for the generated additional video based at least in part on an analysis of the building video to detect motion patterns along the path that meet one or more defined criteria; and selecting the area to include a portion of the path corresponding to the detected motion pattern.
15. The system of claim 5, wherein: The region includes one of a plurality of rooms of the room type, and wherein analyzing the building video to determine a subset of the building video corresponding to the region includes analyzing visual data of the subset of the video to identify the one or more objects associated with the room type.
16. The system of claim 5, wherein: The automatic operation also includes: before generating the additional video, receiving user input from a user and determining an area for the additional video based at least in part on the user input, the user input including at least one of: a portion of the floor plan corresponding to the area selected by the user on the display of the floor plan; or a group of one or more building properties of the building selected by the user and located within the area; or a portion of a path corresponding to the area selected by the user on the display of the floor plan with a visual representation of the path overlaid; or one or more rooms within the area selected by the user; or one or more exterior portions within the area that are outside the building and on the real estate on which the building is located selected by the user.
17. The system of claim 5, wherein: Generating the additional video also includes: for each of one or more objects that are not visible in the visual data of the subset, generating one or more visual representations of the object; and superimposing the generated one or more visual representations in one or more frames of the additional video at one or more indicated locations in the building.
18. A non-transitory computer-readable medium having stored content, the stored content causing one or more computing devices to perform an automated operation, the automated operation comprising at least: obtaining, by the one or more computing devices, data for an indicative building having a plurality of rooms, the data comprising a plurality of images acquired at a plurality of acquisition locations of the indicative building, and further comprising a floor plan of the indicative building, the floor plan having at least two-dimensional room shapes, the at least two-dimensional room shapes having structural elements of the plurality of rooms and having associated locations on the floor plan at the plurality of acquisition locations; generating, by the one or more computing devices and based on the obtained data, a new video along a path through the portion of the building, the new video having at least visual data for at least one room through which the path passes, comprising: analyzing, by the one or more computing devices, the floor plan to determine a plurality of structural elements in the at least one room and determining a subset of the plurality of images having respective acquisition locations in the at least one room, the subset of images comprising a plurality of images; and combining visual data of the plurality of images of the subset to generate a new video having visual data of the determined plurality of structural elements in the at least one room; and presenting, by the one or more computing devices, information about the generated new video on a portion of the floor plan corresponding to the at least one room.
19. The non-transitory computer readable medium of claim 18, wherein: The automatic operation also includes, before generating the new video: receiving a user input from a user, the user input comprising at least one of: a portion of a floor plan selected by a user on a display of the floor plan; or a group of one or more building properties of a building selected by the user; or a selection of one or more rooms by the user; or selection by the user of the exterior of a building and an exterior area on the real estate on which said building is located; as well as Determining a path based at least in part on the user input includes: passing through the portion of the floor plan if selected by the user; or passing through the one or more building properties if selected by the user; or passing through the one or more rooms if selected by the user; or passing through the exterior area if selected by the user.
20. A computer-implemented method comprising: obtaining, by one or more computing devices, data for an indicative house having a plurality of rooms, the data comprising a plurality of images acquired at a plurality of acquisition locations of the indicative house, and further comprising a floor plan for the indicative house, the floor plan indicating a layout of the plurality of rooms by at least two-dimensional room shapes having structural elements of the plurality of rooms and placed at relative locations of the plurality of rooms and having associated locations of the plurality of acquisition locations, and the data further comprising a building video captured at the indicative house along a path passing through at least two of the plurality of rooms; Generate, by the one or more computing devices and based on the obtained data, a first video for a first room of the at least two rooms, the first video having visual data of the first room, including: analyzing, by the one or more computing devices, the building video to determine a subset of the building video corresponding to the first room, including determining that visual data of the subset of the building video matches additional visual data of at least one of the plurality of images whose respective acquisition location is in the first room; and using, by the one or more computing devices, the determined subset of the building video as at least a portion of the additional video of the first room; generating, by the one or more computing devices and based on the obtained data, a second video along the path through the portion of the house, the second video having at least visual data of one or more rooms through which the path passes, comprising: analyzing, by the one or more computing devices, the floor plan to determine a plurality of structural elements in the one or more rooms and determining a subset of the plurality of images having a plurality of images, each of the plurality of images being acquired at a location in the one or more rooms; and combining the visual data of the plurality of images with visual data of a plurality of structural elements determined in the one or more rooms to generate the second video; presenting, by the one or more computing devices, a visual representation of at least a portion of the floor plan including the first room and the one or more rooms, wherein in the presented visual representation representing the first video, first user-selectable information is overlaid on the first room, and in the presented visual representation representing the second video, second user-selectable information is overlaid on the one or more rooms; presenting, by the one or more computing devices and in response to a first user selecting the first user-selectable information, the first video; and The second video is presented, by the one or more computing devices and in response to a second user selecting the second user-selectable information.
Citation Information
Cited By
House resource video processing method and device, and storage medium
CN121640343A