Room-specific geometric representation

CN121095486APending Publication Date: 2025-12-09APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510740952.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-06-02
Filing Date
2025-06-05
Publication Date
2025-12-09

Smart Images

  • Figure CN121095486A_ABST
    Figure CN121095486A_ABST
Patent Text Reader

Abstract

The invention relates to a room-specific geometric representation. Various implementations provide one or more geometric representations of a physical environment based on a room-specific subset of sensor data. For example, a method may include obtaining sensor data of a physical environment including a plurality of rooms, and the sensor data including an image of the physical environment. The method may also include obtaining room boundary information associated with the physical environment, wherein the room boundary information is determined based on the sensor data. The method may also include identifying a room-specific subset of the sensor data based on the room boundary information. The method may also include generating one or more geometric representations of the physical environment based on the room-specific subsets of sensor data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to electronic devices that use sensors to scan a physical environment to generate a geometric representation of the scanned physical environment. Background Technology

[0002] Existing active / passive scanning systems and technologies can be improved by evaluating and using sensor data acquired during the scanning process to generate three-dimensional (3D) representations, such as 3D planar maps representing the physical environment. Summary of the Invention

[0003] The various specific embodiments disclosed herein include devices, systems, and methods for providing room-specific use of sensor data for three-dimensional (3D) reconstruction and / or understanding of the physical environment surrounding a user (e.g., for mesh and / or plane generation for multi-room 3D floor plans). For example, room-specific meshes and / or planes can be formed by generating multiple geometric representations of a large environment (e.g., multiple meshes and / or multiple planes) and taking into account identified room boundaries. Given such room boundaries, captured data (e.g., keyframes of image and / or depth data) can be selectively used to generate geometric representations; for example, only those keyframes corresponding to a given room can be used to update the 3D mesh of that room. These techniques are applicable to any form of room scanning, such as passive scanning performed during Simultaneous Localization and Mapping (SLAM), to reconstruct / understand the physical environment (not just for floor plan generation). In some specific embodiments, portions of the SLAM map can be updated based on which room the user is in, and portions of the SLAM map associated only with the user's room can be passed to the application, etc.

[0004] In some implementations, these room-specific processes can be used in conjunction with the generation of other parametric primitives (e.g., spheres, cylinders, etc.) or implicit representations (such as NeRF / Gaussian sputtering). In some implementations, individual objects can be represented in the scene using appropriate representations (e.g., walls represented by planes, pillars by cylinders, more general objects by meshes or implicit representations, etc.). In other words, other representations can also be linked to and / or limited to specific rooms or bounded regions, similar to the mesh and / or plane generation used for multi-room 3D floor plans as described herein.

[0005] In some implementations, keyframes suitable for each room and clustered together can be selected and fed into the meshing process. Additionally, keyframe clustering can improve the efficiency and accuracy of meshing used to generate a 3D representation of the physical environment (e.g., a 3D floor plan). For example, keyframe clustering can improve wall alignment, avoid wall contention, prevent a room's mesh from being affected by details in adjacent rooms, assist when mirroring is present, handle Simultaneous Localization and Mapping (SLAM) drift, improve the appearance of openings between rooms, and / or address issues such as thin wall structures (e.g., walls sometimes disappear). Furthermore, instead of generating meshes for smaller segments that may span multiple rooms, meshes can be clustered by room, such that each room is associated with one or more meshes, where the meshes do not span multiple rooms. A plane may span multiple rooms (e.g., floor, ceiling, shared walls, etc.), but similarly, the plane will be generated using only captured data corresponding to one or more rooms in which the plane appears.

[0006] The various specific embodiments disclosed herein may also include devices, systems, and methods for providing associations between meshes and / or planes and rooms (e.g., for use in viewing extended reality (XR) environments) in a multi-room 3D floor plan. For example, the process for associating meshes and / or planes may include enabling functionality in XR that selectively uses a subset of the geometric representations (e.g., meshes and / or planes) of a large environment based on their association with a specific room. For example, a room may be associated with one or more meshes, and a plane may span multiple rooms, but will only be associated with the room in which the plane is found. The association of meshes and / or planes with rooms may be used by operating system-level functions that use room-specific meshes / planes, or it may be a third-party application that tells the system which room it is in and receives room-specific meshes / planes for its own functions (e.g., limiting the information associated with room-specific meshes / planes for a particular application).

[0007] The various specific embodiments disclosed herein may also include providing applications with limited access to sensor-based physical environment information (e.g., data regarding planes, grids, virtual object placement, etc.). Access may be restricted based on associating sensor-based environmental information with corresponding keyframes (e.g., a collection of images / other data obtained from a specific location / pose). A device may associate multiple pieces of sensor-based environmental information (e.g., planes, grids, virtual object placement) with corresponding keyframes from which each piece of information is determined. For example, based on determining that keyframe-1 is used to identify plane A, plane A is associated with keyframe-1; based on determining that keyframe-2 is used to identify plane B, plane B is associated with keyframe-2, and so on.

[0008] In some implementations, once a user grants permission for an application to access sensor-based environmental information, the application is granted access to specific information to protect the user's privacy. In some implementations, specification information may be associated with keyframes that relate to the application's location during its use (e.g., during the current application session and / or previous application sessions). Specifically, when the application is used, only a few of the keyframes known to the device are related to the device's location at the time the application is running (e.g., keyframes from nearby locations used for SLAM and / or evaluating the surrounding environment, such as identifying planes, grids, etc.). In the example above, if the application is not used in the location associated with keyframe-2, it will not have access to the plane B information. The application may only have access to sensor-based environmental information associated with these keyframes, effectively granting the application access only to information "visible" to the device during the application's use.

[0009] In some implementations, the system can provide centralized processing and accumulation of privacy data across multiple features (e.g., grids, planes, etc.). In some implementations, the system can provide application-specific persistence of privacy data (e.g., extending the visibility of privacy data beyond the current session, allowing the device to have access to all information from previous sessions to which it previously had access). In some implementations, the system can provide a unique way of using keyframes as privacy entities, thereby providing a novel way of grouping sensor-based datasets, where grouping facilitates limited distribution to applications in a privacy-preserving manner, for example, ensuring that applications are not given access to information associated with parts of the environment that are not visible during application usage.

[0010] Generally, an innovative aspect of the subject matter described in this specification can be embodied in a method at an electronic device having a processor, the method comprising the actions of: acquiring sensor data of a physical environment comprising multiple rooms, the sensor data including images of the physical environment; acquiring room boundary information associated with the physical environment, wherein the room boundary information is determined based on the sensor data; identifying room-specific subsets of the sensor data based on the room boundary information; and generating one or more geometric representations of the physical environment based on these room-specific subsets of the sensor data.

[0011] These and other implementation schemes may optionally include one or more of the following features.

[0012] In some aspects, identifying a room-specific subset of the sensor data based on the room boundary information includes identifying a room-specific keyframe set for each of the plurality of rooms. In some aspects, the sensor data includes first sensor data acquired during a first time period, and the method further includes: acquiring second sensor data during a second time period, wherein the second sensor data corresponds to the first room among the plurality of rooms; and updating the room-specific keyframe set for the first room among the plurality of rooms in response to acquiring the second sensor data.

[0013] In some aspects, these actions may also include updating the three-dimensional (3D) representation of the first room based on the updated set of room-specific keyframes. In some aspects, these actions may also include determining a set of planes associated with each of the plurality of rooms based on these room-specific subsets of the sensor data.

[0014] In some respects, identifying these sets of planes for each room includes: the floor plane, the ceiling plane, and the wall planes of at least one or more walls that identify each room.

[0015] In some aspects, these actions may also include: generating a three-dimensional (3D) representation of the physical environment based on the one or more geometric representations. In some aspects, these actions may also include: presenting a real-time view of the 3D representation on a display of the electronic device.

[0016] In some aspects, the real-time view of the room includes a floor plan generated when the sensor data is acquired. In some aspects, these images of the physical environment are based on at least one of the following: real-time view images, ultrawide view images including views different from these real-time view images, and semantically labeled images corresponding to these real-time view images or these ultrawide view images.

[0017] In some respects, the room boundary information is determined based on 3D semantic data, including 3D point clouds containing semantic labels associated with at least a portion of the 3D points within the point cloud. In some respects, these semantic labels identify the walls, wall structures, objects, and classifications of these objects in each room.

[0018] Generally, an innovative aspect of the subject matter described in this specification can be embodied in a method at an electronic device having a processor, the method comprising the actions of: identifying a room in a multi-room physical environment in which the electronic device is operating, wherein a geometric representation is associated with a room in the multi-room physical environment. These actions may include: obtaining a room-specific subset of the geometric representation, wherein the room-specific subset of the geometric representation is identified based on identifying which geometric representation is associated with the room. These actions may include: providing a view of an extended reality (XR) environment depicting virtual content and the room in the multi-room physical environment, wherein the virtual content is provided based on performing functions using the room-specific subset of the geometric representation.

[0019] These and other implementation schemes may optionally include one or more of the following features.

[0020] In some respects, room identification in a multi-room physical environment is based on the location of tracking electronic devices. In other respects, room identification in a multi-room physical environment is based on the positioning of electronic devices relative to a corresponding floor plan associated with the multi-room physical environment.

[0021] In some respects, these actions may also include selectively providing a room-specific subset of the geometric representation to an application on an electronic device. In some respects, virtual content may be invisible from the view of the XR environment when the electronic device moves to another room.

[0022] Generally, an innovative aspect of the subject matter described in this specification can be embodied in a method in an electronic device having a processor, the method comprising the actions of: associating a set of environmental information based on sensor data with corresponding keyframes in a set of keyframes associated with a physical environment. These actions may include: identifying a subset of the set of keyframes corresponding to the use of an application in the physical environment. These actions may include: identifying a subset of the set of environmental information based on the subset of the keyframe set. These actions may include: granting the application limited access to the set of environmental information based on sensor data, wherein the limited access is restricted to a subset of the set of environmental information based on sensor data.

[0023] These and other implementation schemes may optionally include one or more of the following features.

[0024] In some aspects, associating a set of environmental information based on sensor data with corresponding keyframes includes: associating a first keyframe with a first plane based on an identifier of a first plane using a first keyframe, and associating a second keyframe with a second plane based on an identifier of a second plane using a second keyframe. In some aspects, associating a set of environmental information based on sensor data with corresponding keyframes includes: associating a first keyframe with a first grid based on an identifier of a first grid using a first keyframe, and associating a second keyframe with a second grid based on an identifier of a second grid using a second keyframe.

[0025] In some aspects, the set of environmental information based on sensor data is associated with corresponding keyframes based on room-specific subsets of the identifiers associated with the physical environment. In other aspects, the subset of the set of environmental information based on sensor data is based on room boundary information associated with the physical environment.

[0026] In some aspects, identifying a subset of the keyframe set that corresponds to the application's use in the physical environment is based on identifying keyframes within the keyframe set based on the device's location. In other aspects, identifying a subset of the keyframe set that corresponds to the application's use in the physical environment is based on identifying keyframes within the keyframe set based on the device's pose.

[0027] According to some embodiments, an apparatus includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform any of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors of the apparatus, cause the apparatus to perform or cause to perform any of the methods described herein. According to some embodiments, an apparatus includes: one or more processors, non-transitory memory, and means for performing or causing to perform any of the methods described herein. Attached Figure Description

[0028] To enable those skilled in the art to understand this disclosure, more detailed descriptions can be made with reference to some exemplary embodiments, some of which are shown in the accompanying drawings.

[0029] Figures 1A to 1B Examples are given based on the physical environment of some specific implementations.

[0030] Figure 2 Examples are given based on some specific implementations. Figures 1A to 1B Part of the 3D point cloud of the first room.

[0031] Figure 3 Examples are given based on some specific implementations. Figures 1A to 1B Part of the first 3D floor plan of the first room.

[0032] Figure 4 yes Figure 3 A 3D plan view.

[0033] Figure 5 yes Figure 3 and Figure 4 Another view of the 3D plan view.

[0034] Figure 6 This is a view of the second 3D floor plan of the second room.

[0035] Figure 7 It is based on a combination of specific implementations. Figures 3 to 5 The first 3D plan view and Figure 6 The second 3D plan view is a combination of 3D plan views.

[0036] Figures 8A to 8I An exemplary multi-room 3D floor plan process is illustrated according to some specific implementations.

[0037] Figures 9A to 9D Exemplary processes for generating multiple geometric representations of a planar graph, according to some specific implementations, are illustrated.

[0038] Figures 10A to 10B An exemplary room-specific subset identification process for a geometric representation of a floor plan is illustrated according to some specific implementations.

[0039] Figure 11 This is a flowchart illustrating a pipeline for generating an exemplary 3D room floor plan based on keyframes in some specific implementation.

[0040] Figure 12 This is a flowchart illustrating a method for improving a 3D planar diagram according to some specific implementations.

[0041] Figure 13 This is a flowchart illustrating a method for generating a room-specific geometric representation according to some specific implementations.

[0042] Figure 14 This is a flowchart illustrating a method, according to some specific implementation, for providing a view of an extended reality (XR) environment based on a room-specific geometric representation.

[0043] Figure 15 This is a flowchart illustrating, according to some specific implementations, a method for restricting access to a collection of sensor-based environmental information.

[0044] Figure 16 It is a block diagram based on some specific implementations of electronic devices.

[0045] As is customary practice, various features illustrated in the accompanying drawings may not be drawn to scale. Therefore, for clarity, the dimensions of various features may be arbitrarily expanded or reduced. Furthermore, some drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used throughout the specification and drawings to denote similar features. Detailed Implementation

[0046] Numerous details have been described to provide a thorough understanding of the exemplary embodiments illustrated in the accompanying drawings. However, the drawings illustrate only some exemplary aspects of this disclosure and should not be considered limiting. Those skilled in the art will understand that other effective aspects and / or variations do not include all the specific details set forth herein. Furthermore, well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure further relevant aspects of the exemplary embodiments described herein.

[0047] Figures 1A to 1B An exemplary physical environment 100 is illustrated. In this example, Figure 1A An exemplary electronic device 110 is shown operating in a first room 190 of physical environment 100. Figure 1A In this example, the first room 190 includes a door 130 (providing an opening to a second room 195 in the physical environment 100), a door frame 140, and a window 150 (with a window frame 160) on the wall 120. The first room 190 also includes a table 170 and potted plants 180. Figure 1B As shown in the top view, the first room 190 and the second room 195 are adjacent to each other, that is, a portion of the wall 120 of the first room 190 is adjacent to the wall 196 of the second room 195 (e.g., on their opposite side).

[0048] Electronic device 110 includes one or more cameras, microphones, depth sensors, motion sensors, or other sensors that can be used to capture information about physical environment 100 and assess that physical environment. The acquired sensor data can be used to generate 3D representations, such as 3D point clouds, 3D meshes, or 3D planar maps.

[0049] In one example, user 102 moves around physical environment 100, and device 110 captures sensor data that generates one or more 3D planar views of physical environment 100. Device 110 can be moved to capture sensor data from different viewpoints, such as at various distances, angles, heights, etc. Device 110 can provide user 102 with information that is helpful to the environment scanning process. For example, device 110 can provide a view from a camera during a room scan showing the contents of a currently captured RGB image (e.g., a real-time camera feed). Alternatively, device 110 can provide a view of a real-time 3D point cloud or a real-time 3D planar view to facilitate the scanning process or otherwise provide feedback to user 102 informing them which parts of physical environment 100 have been captured in sensor data and which parts of physical environment 100 require more sensor data for accurate representation in 3D representation and / or 3D planar views.

[0050] Device 110 performs a scan of the first room 190 to capture a first 3D plan view 300 from which the first room 190 is generated. Figures 3 to 5 The data includes, for example, dense point-based representations such as 3D point clouds. Figure 2 This is used to represent a first room 190 and to generate a first 3D floor plan 300, which may represent the 3D positions of the walls, wall openings, windows, doors, and objects in the first room 190. In some embodiments, the 3D floor plan uses non-point cloud data, such as one or more parametric representations, to define the positions of such elements. For example, such parametric representations may define 2D / 3D shapes (e.g., primitives) that represent the positions and sizes of elements in the room in the 3D floor plan. In some embodiments, a 3D floor plan of a room such as the first room 190 is generated based on a 3D point cloud generated during a first scan of the first room 190, for example, a scan captured when a user 102 walks around the first room 190 and sensor data is captured.

[0051] Figure 2 Examples are given. Figures 1A to 1BThis is a portion of the 3D point cloud of the first room 190. In some embodiments, the 3D point cloud 200 is generated based on one or more images (e.g., grayscale, RGB, etc.), one or more depth images, and motion data regarding the movement of the device between different image captures. In some embodiments, an initial 3D point cloud is generated based on sensor data and subsequently densified via an algorithm, a machine learning model, or other processes that add additional points to the 3D point cloud. The 3D point cloud 200 may include information identifying the 3D coordinates of points in a 3D coordinate system. Each point may be associated with characteristic information, such as identifying the point's color based on the color of the corresponding portion of an object or surface in the physical environment 100, identifying the surface normal direction based on the surface normal direction of the corresponding portion of the object or surface in the physical environment 100, identifying a semantic label (which identifies the type of object associated with the point), identifying the estimated material type, etc.

[0052] In an alternative embodiment, a 3D mesh is generated in which the points of the 3D mesh have 3D coordinates, such that groups of mesh points are identified as surface portions, such as triangles, corresponding to the surface of the first room 190 of the physical environment 100. Such points and / or associated shapes (e.g., triangles) may be associated with color, surface normal direction, semantic labels, and / or estimated material.

[0053] exist Figure 2 In the example, the 3D point cloud 200 includes a point set 220 representing a wall 120, a point set 230 representing a door 130, a point set 240 representing a door frame 140, a point set 250 representing a window 150, a point set 260 representing a window frame 160, a point set 270 representing a table 170, and a point set 280 representing a potted plant 180. In this example, the points of the 3D point cloud 200 are depicted as having relatively uniformity and highlighting points on the edges of objects for easier understanding of the figures. However, it should be understood that the 3D point cloud 200 need not include uniformly distributed points, nor need it include points representing object edges that are highlighted or otherwise differ from the points in the 3D point cloud 200.

[0054] The 3D point cloud 200 can be used to identify one or more boundaries and / or regions (e.g., walls, floors, ceilings, etc.) within the first room 190 of the physical environment 100. The relative positions of these surfaces can be determined relative to the physical environment 100 and / or the 3D point-based representation 200. In some implementations, plane detection algorithms, machine learning models, or other techniques are performed using sensor data and / or 3D point-based representations (such as the 3D point cloud 200). The plane detection algorithm can detect the 3D position of one or more planes of the physical environment 100 in a 3D coordinate system. The detected planes can be defined by one or more boundaries, corners, or other 3D spatial parameters. The detected planes can be associated with one or more types of features, such as walls, ceilings, floors, tabletops, countertops, cabinet fronts, etc., and / or can be semantically labeled. The detected planes associated with certain features (e.g., walls, floors, ceilings, etc.) can be analyzed relative to whether such planes include windows, doors, and openings. Similarly, the 3D point cloud 200 can be used to identify any representation of an object. For example, 3D point cloud 200 can be used to identify one or more boundaries of an object, identify bounding boxes around the object (e.g., bounding boxes corresponding to table 170 and plant 180), etc.

[0055] Use 3D point cloud 200 to generate representation Figures 1A to 1B The first 3D representation of the physical environment of the first room 100 and the physical environment of the first room 190 (such as...) Figures 3 to 5 The first plan view 300 illustrated herein. For example, detected planes, boundaries, bounding boxes, etc., can be detected and used to generate the shapes (e.g., 2D / 3D primitives) of elements representing the first room 190 of the physical environment 100. Figures 3 to 5 In this example, wall representations 310a-310d represent the walls of the first room 190 (e.g., wall representation 310b represents wall 120), floor representation 320 represents floor 320 of the first room 190, door representations 350a-350b represent the doors of the first room 190 (e.g., door representation 350a represents door 130), window representations 360a-360d represent the windows of the first room 190 (e.g., window representation 360a represents window 150), table representation 380 is the bounding box representing table 170, and flower representation 290 is the bounding box representing potted plant 180. In this example, because the 3D plan view includes object representations for non-room boundaries (e.g., for 3D objects within the room, such as table 170 and potted plant 180), the 3D plan view can be considered a 3D room scan. In other specific implementations, the 3D plan view only represents room boundary features, such as walls, floors, doors, windows, etc.

[0056] Use a similar (but different) scanning process to generate Figures 1A to 1BThe physical environment of the second room of 100 and the second 3D representation of the physical environment of 195 (such as...) Figure 6 (See the second plan view 600 illustrated). This second scan can be based on sensor data obtained within the second room 195. For example, detected planes, boundaries, bounding boxes, etc., can be detected and used to generate shapes, such as 2D / 3D primitives representing elements of the second room 195 of the physical environment 100. Figure 6 In the diagram, walls 610a-610d represent the walls of the second room 195, floor 620 represents the floor of the second room 195, door 650a represents the door of the second room 195, and windows 660a-660b represent the windows of the second room 195.

[0057] As described, Figures 3 to 5 First 3D plan view 300 and Figure 6 The second 3D plan view 600 is generated by a first room scan and a second room scan, respectively. Such room scans can be different or discontinuous, making device motion tracking between scans unavailable, inaccurate, or otherwise unusable for accurately correlating the different 3D plan views in position. The specific embodiments disclosed herein address this lack of positional correlation by using various techniques to determine the positional correlation between multiple different 3D plan views.

[0058] Given a defined positional relationship between multiple distinct 3D planar diagrams, the 3D planar diagrams can be combined (e.g., stitched together into a single representation) to form a single combined 3D planar diagram. Figure 7 It is a combination Figures 3 to 5 The first 3D plan view 300 and Figure 6The second 3D plan view is a combined 3D plan view 700. As shown, 3D plan views 300 and 600 are positioned adjacent to each other and aligned with each other in a manner that precisely corresponds to the positional relationship between the first room 190 and the second room 195 represented by 3D plan views 300 and 600. For example, the floors of the first room 190 and the second room 195 may be flush with each other in physical environment 100, and therefore floor representations 320 and 620 may be flush with each other (on the same plane) in combined 3D plan view 700. As another example, wall representation 310b (corresponding to wall 120) and wall representation 610d (corresponding to wall 196) may be adjacent to each other based on the fact that walls 120 and 196 are on opposite sides of the same wall in physical environment 100. In some specific implementations, such adjacent walls are merged into a single wall (e.g., as described with respect to Figure 10). Similarly, walls, doors and door openings, windows and window openings, other openings, and other features can be accurately aligned or otherwise positioned based on the precise positional relationships between different 3D plan views 300 and 600. In the case of merged walls, doors, windows, etc., from the merged walls can be projected onto the merged walls to provide an accurate and aligned appearance.

[0059] Figures 8A to 8I An exemplary multi-room 3D floor plan process is illustrated. In this example, such as Figure 8A As shown, the user begins the first scan from the starting position 830 within the first room 810 of the physical environment 800. Figure 8B As shown, during the first scan, the user walks along path 832, capturing images of various parts of the first room 810. Figure 8B As shown, at some point after the first scan is completed, a first 3D planar image 835 is generated. At this time, the user does not need to continue scanning and the device does not need to track its position. The user may (or may not) rest or otherwise wait (e.g., wait for minutes, hours, days, weeks, months, etc.) before performing a second scan. Figures 8A to 8I An exemplary multi-room 3D floor plan process is illustrated, which includes a user scanning one room at a time. However, in some specific implementations, the scanning process may be a continuous process, in which the system may keep each room updated as the user traverses different rooms (e.g., updating multiple rooms at once).

[0060] like Figure 8D As shown, the user begins a second scan of the first room 810 of the physical environment 800 from location 850. Figure 8EAs shown, during this initial portion of the second scan, the device repositions itself within the first room 810, for example, by capturing sensor data as the user moves around within the first room (e.g., along path 860). This repositioning may involve matching features from the sensor data captured during the first scan with sensor data captured during the initial portion of the second scan, both sets of sensor data being based on data captured within the first room 810. For example, feature points detected in the 2D image of the second scan can be mapped with respect to the 3D positions of those feature points determined based on the positioning performed from the sensor data from the first scan. In some specific implementations, Simultaneous Localization and Mapping (SLAM) techniques are used for repositioning during the second scan based on data from the first scan.

[0061] In some implementations, the user interface guides the user to begin a second scan in a previously scanned room (e.g., in the first room 810), guides the user to obtain repositioning data (e.g., by moving around or using a mobile device to capture sensor data from the previously scanned room), and / or notifies the user once repositioning is complete, for example, by guiding the user to move to the second room to generate a second 3D floor plan of the second room.

[0062] like Figure 8F As shown, the device tracks a user moving (e.g., walking) along path 870 from a first room 810 of a physical environment 800 to a second room 820. Tracking this device movement can be based on motion sensor (e.g., accelerometer, gyroscope, etc.) data and / or visual inertial ranging (VIO) and / or other image-based motion tracking. Tracking via motion sensors, VIO, etc., can be continuous as the user moves the device from the first room 810 to the second room 820. This tracking may involve determining, based on sensor data, whether the user is changing floors / floors within the building, for example, by detecting stairs, whether the user is traversing stairs, whether the user is going up or down stairs, the height of the stairs, etc. For example, image and / or depth data can be captured and used to determine whether the device is moving from the ground floor to the basement or vice versa.

[0063] The repositioning and subsequent tracking of the movement from the first room 810 to the second room 820 provides data that enables the determination of the positional relationship between the 3D floor plans generated from the first and second scans.

[0064] like Figure 8G As shown, after the user has been relocated ( Figure 8E And moved to the second room 820. Figure 8FAfterwards, the device scans the second room 820. In some embodiments, the device automatically detects that it is capturing data from a new room (i.e., different from the first room 810) and automatically begins capturing scan data to generate a 3D floor plan of the second room 820. This can occur while moving along path 870 (e.g., while the user is still in the first room 810 or after the user has entered the second room 820). In some embodiments, the user provides input to begin capturing scan data to generate a 3D floor plan of the second room 820. During the scanning of the second room 820, for example, as the user moves along path 880, the user can capture sensor data of the second room 820. Figure 8H As shown, at a certain point after the second scan ends, a second 3D plan view 885 is generated. The repositioning and subsequent tracking of the movement from the first room 810 to the second room 820 provide positioning data that enables the determination of the positional relationship between the 3D plan views generated from the first and second scans. Figure 8I As shown, this positional relationship is used to generate a combined 3D plan view 890, wherein the first 3D plan view 835 and the second 3D plan view 885 are combined in a positionally accurate manner. An optimization process can be used to combine the rooms in a way that reduces or minimizes artifacts and / or errors, such as misaligned walls, corners, doors, windows, multiple doors, non-parallel corridor walls, and misaligned doors / door openings / windows.

[0065] Figures 9A to 9D Exemplary processes for generating multiple geometric representations of a planar diagram, according to some specific implementations, are illustrated. For example... Figure 9A As illustrated, during the scanning of the representation of the physical environment 900 (e.g., as... Figures 8A to 8I As illustrated, different geometric representations (e.g., 3D meshes) can be generated. For example, geometric representations 902a and 902b appear to include all or at least most of bedroom-1 910, geometric representation 902g covers a portion of bedroom-2 920, geometric representation 902c covers a portion of living room 940, geometric representation 902d covers a portion of dining room 950, geometric representation 902e covers a portion of kitchen 960, and geometric representation 910f covers a portion of lounge 930. Figure 9B This illustrates a large-sized 2D mesh used to detect specific physical structures, such as floors or ceilings, within a multi-room environment. For example, geometric representations 904a and 904b could represent two 2D planes representing the ceiling at 90°. During scanning of the environment, and / or after completing a full scan of the physical environment, a mesh like... Figure 9C The illustrated 2D plan view 980 and / or as shown Figure 9DThe illustrated 3D plan view 990. In some embodiments, the 2D plan view 980 and / or the 3D plan view 990 may be used by an application on the device. Additionally, in some embodiments, the 2D plan view 980 and / or the 3D plan view 990 may be provided as a preview to the user performing the scan, either during or after its generation.

[0066] Figures 10A to 10B An exemplary room-specific subset identification process for a geometric representation of a floor plan, according to some specific implementations, is illustrated. Specifically, Figure 10A , Figure 10B Examples 1000A and 1000B are 3D plan view representations used to obtain room boundary information from a scan of the physical environment (e.g., represented as 900) and selectively identify specific subsets of the geometric representations of a room. In other words, room plan view algorithms assign different geometric representations (e.g., 3D mesh regions) to specific rooms to identify each room or region. This algorithm can acquire sensor data and determine semantic data, planes, object detection / bounding boxes, identified room structures (walls, openings, windows, doors, etc.).

[0067] like Figure 10A As shown, the 3D floor plan representation 1000A includes geometric representation 1002a associated with bedroom-1 1010, geometric representation 1002d associated with bedroom-2 1020, geometric representation 1002b associated with living room 1040, dining room 1050, and kitchen 1060, and geometric representation 1002c associated with lounge 1030. Therefore, because living room 1040, dining room 1050, and kitchen 1060 are all a large open area, geometric representation 1002b is a large 3D mesh representation of that area / room. However, as... Figure 10B As shown, the 3D floor plan representation 1000B is based on the determination that the living room 1040, dining room 1050, and kitchen 1060 should be separated by different geometric representations 1004a, 1004b, and 1004c, respectively. For example, the geometric representation could be a segmented space based on a "function" of area. In other words, the system may not need to identify thin structures (e.g., walls) to technically define rooms separated from other rooms. The identifiers dividing an area into two rooms or sections can be automatically determined based on user feedback / interaction or on a semantic understanding of objects within the room (e.g., a kitchen area could be identified based on cabinets, an island with a stove, etc.).

[0068] Figure 11This is a flowchart illustrating an exemplary 3D room floor plan generation pipeline 1100 that can be performed at a device such as device 110 of Figure 1. In this pipeline 1100, sensor data is acquired at sensor data and tracking box 1104. Such sensor data may include captured images, depth sensor data, ambient light sensor data, motion sensor data, and / or any other type of sensor data useful in scanning, providing, feedback, and / or 3D room scan generation. At sensor data and tracking box 1104, the device may track its pose (i.e., position and / or orientation) as it captures sensor data. At 3D modeling box 1106, the data from sensor data and tracking box 1104 is used.

[0069] The 3D modeling frame 1106 can use sensor data (e.g., during scanning of the physical environment) to generate and update a 3D model (e.g., a 3D point cloud or a 3D mesh) representing the physical environment. As more sensor data is received and processed, the 3D model can be refined and updated. This update can occur in real time during and / or after the scanning process. The 3D modeling frame 1106 can provide a 3D model comprising points or mesh polygons corresponding to surface portions of the physical environment. Such points and / or mesh polygons may each have a 3D location and be associated with additional information, including but not limited to color information, surface normal information, semantic label information, and estimated material, such as identifying the type of object corresponding to each point or mesh polygon. Color, surface normal, semantic, and estimated material information can be determined based on evaluating sensor data (e.g., using an algorithm or machine learning model). The 3D modeling frame 1106 can provide the 3D model to the wall / opening detection and consistency frame 1110 and / or the 3D object detection frame 1120. The 3D models provided to boxes 1110 and 1120 can be updated over time, such as during the scanning process and / or during the capture of sensor data after the scanning process.

[0070] The wall / opening detection and consistency box 1110 uses a 3D model to detect walls and openings within a physical environment. This may include predicting planar surfaces corresponding to walls, floors, ceilings, etc., and / or the boundaries of such planar surfaces. In some implementations, a machine learning model evaluates the 3D model and / or sensor data to identify planar surfaces and / or detect walls, openings, etc. This may include using location information and additional information associated with point / mesh polygons of the 3D model. For example, this may include material and / or semantics estimated using location, color, surface normals associated with point / mesh polygons of a 3D point cloud or 3D mesh.

[0071] The wall / opening detection and consistency box 1110 uses the detected walls, openings, etc., and compares them with other data to ensure that the positioning, size, shape, etc., of the walls, openings, etc., are consistent with each other. The wall / opening detection and consistency box 1110 provides the adjusted walls, openings, etc., to the window / door detection box 1112 and the wall / opening height estimation box 1114.

[0072] Window / door detection frame 1112 detects windows and doors on walls. This detection may utilize a 3D model from frame 1106, sensor data from frame 1104, and / or data about walls, openings, etc., from frame 1110. In some implementations, window / door detection frame 1112 detects points / mesh polygons of the 3D model within a threshold distance of the detected walls, openings, etc., and associates those point / mesh polygon vertices with the walls. For example, this may include projecting some point cloud points onto the plane of the wall. Windows, doors, etc., can be detected based on the projected points, with or without semantic information. In some implementations, an algorithm or machine learning model interprets the 3D model and the detected walls, openings, etc., to predict the position and size of windows, doors, etc.

[0073] Wall / Opening Height Estimation Box 1114 estimates the height of walls and openings. This type of detection can utilize a 3D model from Box 1106, sensor data from Box 1104, and / or data about walls, openings, etc., from Box 1110. This type of detection may include the use of algorithms or machine learning models.

[0074] The outputs of boxes 1112 and 1114 are used to generate a plan view 1116 that specifies the location and size of elements of the physical environment of an approximate plan / building (e.g., walls, floors, ceilings, openings, windows, doors, etc.). This 3D plan view can parameterize the planar elements of the physical environment, for example, by specifying the locations of two or more points (e.g., relative corner points defining the shape and position of a rectangle in a 3D coordinate system) that provide sufficient information to form a rectangle, polygon, or other 2D shape. In some implementations, the approximate plan / building (e.g., walls, floors, ceilings, openings, windows, doors, etc.) has a certain thickness and is represented (e.g., using parameters specifying the 2D shape and thickness).

[0075] The 3D model from box 1106 is also output to 3D object detection box 1120. 3D object detection box 1120 can detect objects such as tables, televisions, screens, refrigerators, fireplaces, shelves, ovens, chairs, stairs, sofas, dishwashers, cabinets, stoves, beds, toilets, washing machines, dryers, sinks, bathtubs, etc. Such detection can utilize the 3D model from box 1106 and / or sensor data from box 1104. Such detection can include the use of algorithms or machine learning models. In some implementations, machine learning models evaluate 3D models and / or sensor data to identify bounding boxes or other primitive shapes surrounding 3D objects. This can include using positional information and additional information associated with point / mesh polygons of the 3D model. For example, this can include material and / or semantics estimated using position, color, surface normals associated with point / mesh polygons of a 3D point cloud or 3D mesh. As a specific example, a set of points corresponding to a table-type object can be identified based on semantic labels associated with points in the point cloud. The bounding boxes around these points can be determined based on their positions. Such bounding boxes can be oriented based on the surface normals of points, for example, making the bounding box orientation match the orientation of a table.

[0076] At object boundary refinement box 1122, the boundaries of the 3D object detected at box 1120 are refined. This refinement can utilize the 3D model from box 1106, sensor data from box 1104, and the 3D object detected at box 1120. Sensor information from box 1140 may include frame updates from box 1130, such as images associated with sensor data used to provide a real-time preview during scanning, semantically labeled images, etc. The frame updates from box 1130 can then be analyzed at box 1135 for keyframe tracking. For example, keyframes can be tracked and used for the resulting 3D mesh of one or more portions of a 3D plan view. In some implementations, keyframes suitable for a room and clustered together can be selected and fed into the meshing process. Additionally, keyframe clustering can improve the efficiency and accuracy of meshing for generating a 3D representation of the physical environment (e.g., a 3D plan view) by leveraging clustering techniques used for efficient feature extraction methods.

[0077] In some implementations, the block 1135 for keyframe tracking may include associating sensor-based environmental information with corresponding keyframes (e.g., a set of images / other data obtained from a specific location / pose). For example, the device may associate multiple pieces of sensor-based environmental information (e.g., planes, grids, virtual object placement) with corresponding keyframes from which each piece of information is determined. For example, based on determining that keyframe-1 is used to identify plane A, plane A is associated with keyframe-1; based on determining that keyframe-2 is used to identify plane B, plane B is associated with keyframe-2, and so on.

[0078] In some implementations, once a user grants permission for an application to access sensor-based environmental information, the application is only granted access to privacy information associated with keyframes that are relevant to the application's location during its use (e.g., during the current application session and / or previous application sessions). Specifically, when the application is used, only a few of the keyframes known to the device are relevant to the device's location at the time the application is running (e.g., keyframes from nearby locations used for SLAM and / or evaluating the surrounding environment, such as identifying planes, grids, etc.). In the example above, if the application is not used in the location associated with keyframe-2, it will not have access to the plane B information. The application may only have access to sensor-based environmental information associated with these keyframes, effectively granting the application access only to information "visible" to the device during the application's use.

[0079] Additionally, refinement may be used by the tutor box 1140, or otherwise to provide feedback to the user by adjusting the position of an object representation (e.g., bounding box edges) that may not be precisely aligned with the corresponding real-world edges depicted in the live image data. Therefore, in some embodiments, refinement may be used to display edge indications on the live view during scanning. In some embodiments, such refinement is used only for live view enhancement during scanning. In some embodiments, such refinement is used only to improve the 3D object representation for generating a 3D room floor plan. In some embodiments, such improvement is used for both. Unrefined and / or refined 3D objects may be provided to the wall / object alignment box 1124.

[0080] The wall / object alignment box 1124 adjusts the 3D object representation (e.g., 3D bounding box representation) based on the plan view 1116. For example, the 3D bounding box of a table positioned near a wall can be adjusted to be parallel to the wall, against the wall, etc. In some implementations, the 3D object representation within a threshold distance of the wall in the plan view 1116 is automatically adjusted to be aligned with the wall.

[0081] The output of the wall / object alignment box 1124 provides the bounding box of the 3D object or other 3D primitive representation for use in generating the 3D room floor plan 1150.

[0082] 3D primitive representations can parameterize 3D objects, for example, by specifying the positions of two or more vertices (e.g., relative corner points defining the shape and position of a 3D box within a 3D coordinate system) that provide sufficient information to form a 3D box, cone, cylinder, wedge, sphere, torus, pyramid, etc.

[0083] Sensor data and tracking box 1104 also provide data used by frame update box 1130. Frame update box 1130 includes 2D frame data captured in the physical environment during scanning. This frame update box can include frame-based data (e.g., 2D images, 2D depth images, semantically labeled 2D images, etc.) captured at a relatively fast rate during scanning. Frame data can be updated at a faster rate than updating the 3D model at 3D modeling box 1106. Frame update box 1130 provides 2D frame data to tutor box 1140, mirror detection box 1142, and planar boundary refinement box 1144.

[0084] The guidance frame 1140 can provide guidance or other information during the scanning process to facilitate the scanning process. For example, the guidance frame can provide a real-time view of the image data being captured (e.g., via video), determine how the user should move the device to capture data of the remaining portions of the physical environment, and guide the user to take actions to improve the quality of the image capture (e.g., move more slowly, rescan the area, move to scan a new area, increase ambient lighting, etc.).

[0085] Mirror detection frame 1142 uses 2D frame data from frame update frame 1130 to detect mirrors in the physical environment. Mirror detection may include algorithms or machine learning processes configured to detect reflective surfaces within the physical environment. Mirror detection frame 1142 can provide information about the detected mirrors, which is used to generate a 3D room floor plan 1150.

[0086] The floor plan room boundary refinement box 1144 uses 2D frame data from frame update box 1130 and floor plan 1116 to determine and / or refine the room boundaries of the floor plan. For example, after receiving a keyframe from box 1135, the floor plan room boundary refinement box 1144 may identify room-specific boundaries (e.g., planes associated with the room), generate coarse boundaries (e.g., low-level estimates), and use this information to generate a mesh / plane for each identified room before applying refinement. The information associated with the room-specific mesh / plane for each identified room can then be sent to keyframe update box 1160.

[0087] Keyframe update frame 1160 can select, track, and update keyframes suitable for each room and that can be clustered together and fed into the meshing process. Additionally, keyframe clustering can improve the efficiency and accuracy of meshing used to generate a 3D representation of the physical environment (e.g., a 3D floor plan). For example, keyframe clustering can improve wall alignment, avoid wall contention, prevent a room's mesh from being affected by details in adjacent rooms, assist when mirroring is present, handle Simultaneous Localization and Mapping (SLAM) drift, improve the appearance of openings between rooms, and / or address issues such as thin wall structures (e.g., sometimes walls disappear). Furthermore, instead of generating meshes for smaller segments that may span multiple rooms, meshes can be clustered by room, such that each room is associated with one or more meshes, where the meshes do not span multiple rooms. The 3D room floor plan frame 1150 can then utilize the room-specific mesh / planar information.

[0088] Keyframe update box 1160 can update a set of keyframes associated with the physical environment (e.g., all keyframes of a house that the device has). Keyframe update box 1160 can also update sensor-based environmental information associated with the corresponding keyframes. For example, the device can update multiple pieces of sensor-based environmental information (e.g., plane, grid, virtual object placement, etc.) associated with the corresponding keyframe from which each piece of information is determined.

[0089] In some embodiments, the refinement of the floor plan room boundary refinement box 1144 may be used by the tutor box 1140 or otherwise to provide the user with feedback on the location of wall edges and other boundaries determined for the floor plan 1116, which may not be precisely aligned with the corresponding real-world edges depicted in the live image data. Therefore, in some embodiments, the floor plan refinement may be used to display edge indications on the live view during scanning. Floor plan boundary refinement may include using 2D RGB images, 2D semantically labeled images, 2D depth data, or other 2D data obtained or generated therefrom to determine adjustments to the boundaries of walls, openings, windows, doors, etc., in the floor plan. In some embodiments, such refinement is used only to provide enhancement to the live view during scanning. In some embodiments, such refinement is used only to improve the floor plan 1116 used to generate the 3D room floor plan 1150. In some embodiments, such refinement is used for both.

[0090] The 3D room floor plan 1150 thus combines the 2D shape of the floor plan 1116, which represents walls, openings, doors, windows, and other planar / architectural elements, with a 3D object representation from frame 1124. The 3D room floor plan may also take into account information from tutor frame 1140 and mirror detection 1142. Due to the relatively high level / parameter representation, the resulting 3D room floor plan 1150 can be generated efficiently and accurately. In some specific implementations, the 3D room floor plan 1150 is generated relatively quickly, for example, during or shortly after scanning the physical environment, and does not require significant waiting (e.g., minutes, hours, days, etc.) for significant manual modifications or other post-processing.

[0091] Using parametric representations to define 3D room floor plans 1150 allows for the use of simple, compact datasets that can be efficiently stored, managed, presented, modified, shared, transmitted, or otherwise used. Such 3D room floor plans offer significant advantages over non-parametric ones, such as those utilizing dense point clouds or 3D meshes with hundreds or thousands of vertices representing hundreds or thousands of triangular faces. Parametric representations can utilize 3D boundary shapes as primitives to represent the shapes of tables, objects, appliances, etc. Such representations significantly simplify details while still providing 3D room floor plans that accurately model the essential aspects of a room.

[0092] Figure 12 This is a flowchart illustrating a process 1200 for improving a combined 3D floor plan. In this example, elements of the combined 3D floor plan are aggregated and grouped, as shown in aggregation box 1210 and grouping box 1220. Aggregation box 1210 identifies doors and door openings in box 1212, walls in box 1214, and windows in box 1216. Detected elements in multiple rooms that are spatially close together at mapping box 1222 and grouped together in grouping box 1224 are used to improve the combined 3D floor plan. For example, mapping and grouping may identify a wall connecting / separating two rooms and a door connecting the two rooms, and each of these groups may be identified and used as a constraint when improving / optimizing the 3D floor plan at merging / improving box 1230.

[0093] Merge / Improve box 1230 performs room updates at box 1232, which can move or rotate individual 3D floor plans relative to each other, for example, to make walls parallel and / or perpendicular, corridors straight, etc. At box 1234, the process merges merged elements (e.g., planes representing adjacent walls). This can be based on the grouping from box 1220. The merging of wall planes can assume wall thickness. For example, wall thickness can be based on the geographical location of the house, such as walls in a particular area being expected to have a thickness of X inches, etc.

[0094] At box 1234, the process also updates the corners of elements in adjacent rooms to align them with each other. One or more optimization processes can be used. For example, such optimization can use constraints from grouping and repeatedly rotate and / or move each figure in the 3D plan view to minimize error / distance measurements. During such optimization, each 3D plan view can (or can not) be treated as a rigid body. Such optimization improves alignment and reduces error gaps between adjacent 3D plan view elements in the combined 3D plan view.

[0095] At box 1236, wall elements (e.g., windows, doors, etc.) are reprojected onto any walls merged at box 1234. These elements may have been positioned on different planes before the walls were merged, and therefore the merge could make them appear to "float" in space. Reprojecting the wall elements (e.g., doors, windows, etc.) onto the merged wall locations resolves this "floating" disconnected appearance. Reprojection may involve calculating orthographic projections onto the new merged wall locations.

[0096] At box 1238, keyframes are selectively identified for each room or segment of a room in the floor plan (e.g., keyframe room association). For example, only those keyframes corresponding to a given room can be used to update the 3D mesh for that room. In some implementations, keyframes suitable for each room can be clustered together and fed into the meshing process. For example, a finite set of all collected keyframes can be used for each identified room. In some implementations, at box 1238, the system can associate sensor-based environmental information with corresponding keyframes (e.g., a set of images / other data obtained from a specific location / pose). For example, the device can associate multiple pieces of sensor-based environmental information (e.g., planes, meshes, virtual object placement) with the corresponding keyframe from which each piece of information is determined. For example, based on determining that keyframe-1 is used to identify plane A, plane A is associated with keyframe-1; based on determining that keyframe-2 is used to identify plane B, plane B is associated with keyframe-2, and so on. In some implementations, once a user grants permission for an application to access sensor-based environmental information, the application is only granted access to privacy information associated with keyframes that are relevant to the application's location during its use (e.g., during the current application session and / or previous application sessions). Specifically, when the application is used, only a few of the keyframes known to the device are relevant to the device's location at the time the application is running (e.g., keyframes from nearby locations used for SLAM and / or evaluating the surrounding environment, such as identifying planes, grids, etc.). In the example above, if the application is not used in the location associated with keyframe-2, it will not have access to the plane B information. The application may only have access to sensor-based environmental information associated with these keyframes, effectively granting the application access only to information "visible" to the device during the application's use.

[0097] At box 1240, a normalization process is performed to give the combined 3D planar graph a normalized appearance. The normalization process may include keyframe updates at box 1242. For example, keyframes may be tracked and used with normalized data to modify the resulting 3D mesh of the 3D planar graph. Additionally, corresponding keyframes associated with the corresponding use of the application in the physical environment (e.g., based on the device's current position and / or pose) may be updated.

[0098] Process 1200 can improve the combined 3D floor plan in various ways. Process 1200 can remove the appearance of double walls, align corners, and align other elements, so that the combined 3D floor plan has an accurate, easy-to-understand appearance and otherwise conforms to user expectations. Additionally, process 1200 can (e.g., via keyframes) utilize room-specific subsets of sensor data and restrict such data to specific applications (e.g., restricting the application to access a grid associated with the bedroom but not with the adjacent bathroom).

[0099] Figure 11 and Figure 12 Examples of how to generate 3D floor plans using room / grid associations are provided. However, the techniques described in this article (including, but not limited to, those described below for specific purposes) Figure 13 and Figure 14 The described methods 1300 and 1400 can be used for different types of purposes besides generating 3D floor plans (e.g., games, application placement, immersive experiences, etc.). Identifying and using room / mesh associations allows applications to use only the meshes / planes associated with the user's room (resulting in more efficient processing), allows the system to display application windows only in the user's room (thus updating the SLAM map), allows immersive environments to interact only with those parts of the mesh that are in the user's room, etc. In other words, exemplary embodiments provide a process for identifying meshes / planes associated with a room and then using that known association to perform a function / application, and... Figure 11 and Figure 12 The 3D room floor plan generation described herein is an example of generating and using mesh / room associations. For example, the association of meshes and / or planes with rooms can be used by operating system-level functions that use room-specific meshes / planes, or it can be a third-party application that tells the system which room it is in and receives room-specific meshes / planes for its own functions (e.g., restricting the information associated with room-specific meshes / planes for a particular application).

[0100] Figure 13 This is a flowchart illustrating a method 1300 for generating a 3D reconstruction of a physical environment, wherein portions of the reconstructed environment are associated with corresponding rooms. In some embodiments, a device such as electronic device 110 performs method 1300. In some embodiments, method 1300 is performed on a mobile device, desktop computer, laptop computer, HMD, or server device. Method 1300 is performed by processing logic components, including hardware, firmware, software, or combinations thereof. In some embodiments, method 1300 is performed on a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory).

[0101] In various specific embodiments disclosed herein, method 1300 provides room-specific use of sensor data for three-dimensional (3D) reconstruction and / or understanding of the physical environment surrounding a user (e.g., for mesh and / or plane generation for multi-room 3D floor plans). For example, room-specific meshes and / or planes can be formed by generating multiple geometric representations of a large environment (e.g., multiple meshes and / or multiple planes) and taking into account identified room boundaries. Given such room boundaries, captured data (e.g., keyframes of image and / or depth data) can be selectively used to generate geometric representations; for example, only those keyframes corresponding to a given room can be used to update the 3D mesh of that room. Mesh / plane association can be used for various applications such as generating 3D floor plans, updating SLAM maps, games, immersive experiences, etc.

[0102] At block 1302, method 1300 obtains sensor data of a physical environment comprising multiple rooms, the sensor data including images of the physical environment. The sensor data may include 3D point clouds and sequences of 2D images corresponding to views of the rooms captured during scanning of the rooms. In some specific implementations, the sensor data includes image data (e.g., from an RGB camera), depth data (e.g., depth images from a depth camera), ambient light sensor data (e.g., from an ambient light sensor), and / or motion data from one or more motion sensors (e.g., accelerometers, gyroscopes, etc.). The sensor data may include keyframes of the images.

[0103] In some implementations, sensor data includes visual inertial odometry (VIO) data determined based on image data. A 3D point cloud can provide semantic information about one or more elements of a room. A 3D point cloud can provide information about the location and appearance of surface portions within the physical environment. In some implementations, the 3D point cloud is acquired over time (e.g., during a scan of the room), and can be updated, with updated versions of the 3D point cloud obtained over time. For example, when the 3D representation is updated / adjusted over time (e.g., while a user is scanning the room, such as...). Figures 8A to 8I As shown, a 3D representation can be obtained (and analyzed / processed).

[0104] In some implementations, the 2D image corresponds to the view that the user can see during the live view. Alternatively, the 2D image captured during a room scan is an ultrawide image that differs from the live view shown to the user (e.g., capturing additional sensor data not shown to the user during the scan). Additionally, in some implementations, the 2D image captured during a room scan includes semantically labeled images corresponding to the live view and / or the ultrawide view. In some implementations, passive scanning of the room may occur (e.g., the user is using the device for other purposes, and the device is scanning a room behind the scene without active user involvement).

[0105] At box 1304, method 1300 obtains room boundary information associated with the physical environment, which is determined based on sensor data. For example, room-specific grids and / or planes can be formed by generating multiple geometric representations of a large environment (e.g., multiple grids and / or multiple planes) and taking into account the identified room boundaries.

[0106] In some implementations, the image of the physical environment is based on at least one of the following: a real-time view image, an ultrawide view image including a view different from the real-time view image, and a semantically labeled image corresponding to the real-time view image or the ultrawide view image. In some implementations, room boundary information is determined based on 3D semantic data including a 3D point cloud containing semantic labels associated with at least a portion of the 3D points within the 3D point cloud. In some implementations, the semantic labels identify the walls, wall structures, objects, and object categories of each room. In some implementations, the semantic labels may identify the type of room (e.g., living room, kitchen, bedroom, etc.).

[0107] At box 1306, method 1300 identifies room-specific subsets of sensor data based on room boundary information. In other words, a limited number of keyframes are selectively identified for each room / segment. Given such room boundaries, the captured data (e.g., keyframes of image and / or depth data) can be selectively used to generate a geometric representation; for example, only those keyframes corresponding to a given room can be used to update the 3D mesh of that room. Additionally or alternatively, in some implementations, room-specific subsets of sensor data can be identified by matching / clustering keyframes with semantic labels of the same room type (e.g., if the keyframes have room labels, such as semantically labeled rooms).

[0108] In some implementations, identifying room-specific subsets of sensor data based on room boundary information includes identifying a set of room-specific keyframes for each of multiple rooms. For example, a mesh for a given room is generated, using only the keyframes for that room. Keyframes can be captured for each room during live scanning or during device usage. The keyframe process can compute clusters based on sensor data (e.g., 3D point clouds), then associate the keyframes with rooms to generate coarse boundaries (e.g., low-level estimates) for the rooms, and then use this information to generate the mesh / plane. A refinement module can refine the boundaries of the coarse boundaries (the published room planes) to provide a two-level approach, e.g., coarse then refined.

[0109] In some implementations, if a keyframe is identified as being associated with two different rooms, keyframes can be selected for a specific room based on which keyframe is pointing at which room. For example, the system can view each keyframe by room, score the keyframes, and determine which keyframes to process. Additionally, a confidence level can be determined and scored for each keyframe associated with each room. In some implementations, the system generates a mesh for more than one room at a time, and then the system can trim portions of the mesh that are identified as not being part of a room.

[0110] In some implementations, the sensor data is first sensor data acquired within a first time period, and method 1300 further includes: acquiring second sensor data within a second time period, wherein the second sensor data corresponds to a first room among a plurality of rooms; and updating a room-specific keyframe set for the first room among the plurality of rooms in response to acquiring the second sensor data. In some implementations, method 1300 further includes updating a three-dimensional (3D) representation of the first room based on the updated room-specific keyframe set. For example, using a keyframe clustering algorithm, keyframes can be adapted to each room and clustered together and fed into the meshing process. Keyframe clustering can improve the efficiency and accuracy of meshing used to generate a 3D representation of the physical environment (e.g., a 3D floor plan). For example, keyframe clustering can be used to improve wall alignment, avoid wall contention, prevent the mesh of one room from being affected by details in adjacent rooms, assist when mirroring is present, address Simultaneous Localization and Mapping (SLAM) drift, improve the appearance of openings between rooms, and / or resolve issues such as thin wall structures (e.g., sometimes walls disappear).

[0111] At box 1308, method 1300 generates one or more geometric representations of the physical environment based on a room-specific subset of sensor data. For example, keyframes can be used to update the appropriate mesh corresponding to the appropriate room; for instance, a finite set of all collected keyframes is used for each room. Keyframe selection implements room-specific keyframe clustering for mesh generation (e.g., keyframes suitable for each room are clustered together and fed into the mesh generation process). Plane generation / detection can also be based on a room-specific subset of sensor data.

[0112] In some implementations, method 1300 further includes determining a set of planes associated with each of the plurality of rooms based on a room-specific subset of sensor data. For example, plane generation and / or plane detection may also be based on a room-specific subset of sensor data (e.g., keyframes). In some implementations, determining the set of planes for each room includes: floor planes, ceiling planes, and wall planes that identify at least one or more walls of each room.

[0113] In some implementations, method 1300 further includes generating a 3D representation of the physical environment based on one or more geometric representations. For example, a room might be defined as a 3D bounded volume, but it might not be desirable to restrict each geometric representation to a single definition. For example, the same open space could be two rooms, such as an open dining room leading to a kitchen. Thus, as discussed herein, the geometric representation could be a piecewise space based on a “function” of area. In other words, the system might not need to identify thin structures (e.g., walls) to technically define rooms separated from another room. For example, a house with an open living room and dining room leading to a kitchen, or an office building with corridors and several offices or cubicles. Therefore, if a large space exists, the system might need to define some areas for partitioning. The defined areas could be defined hierarchical partitioning, such as first physical objects (walls) and then logical ones (corridors, cubicles, etc.). In some implementations, the process may involve manual input on a user interface (e.g., interaction with the representation) to determine the room designations of the geometric representation (e.g., dividing a large grid into two rooms without walls between them, such as a kitchen and dining area).

[0114] In some embodiments, method 1300 further includes presenting a real-time view of a 3D representation on a display of an electronic device. In some embodiments, the real-time view of the room includes a floor plan generated when sensor data is acquired. For example, a real-time view of a 3D mesh is generated as a preview while scanning the environment, or for use in applications on the device.

[0115] Figure 14This is a flowchart illustrating method 1400 for providing a view of an extended reality (XR) environment based on a room-specific geometric representation. In some embodiments, a device such as electronic device 110 performs method 1400. In some embodiments, method 1400 is performed on a mobile device, desktop computer, laptop computer, HMD, or server device. Method 1400 is performed by processing logic components, including hardware, firmware, software, or combinations thereof. In some embodiments, method 1400 is performed on a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory).

[0116] In various specific embodiments disclosed herein, method 1400 provides the association of meshes and / or planes with rooms, for example, for use in a multi-room 3D floor plan to be utilized when viewing an XR environment. For example, the process for associating meshes and / or planes may include enabling functionality in the XR that selectively uses a subset of the geometric representations (e.g., meshes and / or planes) of a large environment based on the association of geometric representations with specific rooms. For example, a room may be associated with one or more meshes, and a plane may span multiple rooms, but will only be associated with the room in which the plane is found. The association of meshes and / or planes with rooms may be used by operating system-level functions that use room-specific meshes / planes, or it may be a third-party application that tells the system which room it is in and receives room-specific meshes / planes for its own functions (e.g., restricting information associated with room-specific meshes / planes for a particular application).

[0117] At box 1402, method 1400 identifies a room in a multi-room physical environment in which an electronic device is operating, wherein a geometric representation is associated with the room in the multi-room physical environment. The geometric representation may include 2D / 3D meshes and / or planes.

[0118] In some implementations, rooms in a multi-room physical environment are identified based on the location of a tracking electronic device. For example, the identification of a room in a multi-room physical environment (e.g., a bedroom in a house) could be the device's current room. In some implementations, rooms in a multi-room physical environment are identified based on the location of an electronic device relative to a corresponding floor plan associated with the multi-room physical environment. For example, room identification could be determined based on the tracking device's location, location relative to a known floor plan / mapping, understanding / detection of semantic information about objects in the room, or other techniques.

[0119] In some implementations, room boundary information is associated with the physical environment and determined based on acquired sensor data. For example, room-specific grids and / or planes can be formed by generating multiple geometric representations of a large environment (e.g., multiple grids and / or multiple planes) and taking into account the identified room boundaries. In some implementations, algorithms for 3D reconstruction of the physical environment acquire sensor data and determine semantic data, planes, object detection / bounding boxes, identified room structures (walls, openings, windows, doors, etc.), etc. In some implementations, room boundary information is determined based on 3D semantic data including a 3D point cloud containing semantic labels associated with at least a portion of the 3D points within the point cloud. In some implementations, semantic labels identify the walls, wall structures, objects, and object classifications for each room.

[0120] At box 1404, method 1400 obtains a room-specific subset of the geometric representation, which is identified based on which geometric representations are associated with a room. For example, only the meshes and planes associated with the current room are identified. In other words, a limited number of keyframes for each room / segment are selectively identified. Given such room boundaries, the captured data (e.g., keyframes of image and / or depth data) can be selectively used to generate the geometric representation; for example, only those keyframes corresponding to a given room can be used to update the 3D mesh of that room.

[0121] In some implementations, identifying room-specific subsets of sensor data based on room boundary information involves identifying a set of room-specific keyframes for each of multiple rooms. For example, a grid is generated for a given room, using only the keyframes for that room. Keyframes can be captured for each room during live scanning or during device usage. The keyframe process can compute clusters, then associate the keyframes with rooms to generate coarse boundaries (e.g., low-level estimates) for the rooms, and then use this information to generate the grid / plane. A refinement module can refine the boundaries of the coarse boundaries (the published room planes) to provide a two-level approach, e.g., coarse then refined.

[0122] In some implementations, if a keyframe is identified as being associated with two different rooms, keyframes can be selected for a specific room based on which keyframe is pointing at which room. For example, the system can view each keyframe by room, score the keyframes, and determine which keyframes to process. Additionally, a confidence level can be determined and scored for each keyframe associated with each room.

[0123] At box 1406, method 1400 provides a view of an XR environment depicting virtual content and a multi-room physical environment, the virtual content being provided based on room-specific subsets of geometric representations. For example, the process for associating grids and / or planes may include enabling functionality in the XR that selectively uses a subset of the geometric representations of a large environment (e.g., grids and / or planes) based on the association of geometric representations with specific rooms. For instance, a room may be associated with one or more grids, and a plane may span multiple rooms, but will only be associated with the room where the plane is found.

[0124] In some implementations, the association of grids and / or planes with rooms can be used by operating system-level functions that utilize room-specific grids / planes, or it can be a third-party application that tells the system which room it is in and receives room-specific grids / planes for its own functions (e.g., restricting information associated with room-specific grids / planes used for a particular application). Examples include a house with an open living room and dining room leading to a kitchen, or an office building with hallways and several offices or cubicles. Therefore, if a large space exists, the system may need to define some areas for partitioning. The defined areas can be defined hierarchical partitioning, such as first physical objects (walls) and then logical ones (hallways, cubicles, etc.). In some implementations, the process may involve manual input on a user interface (e.g., interaction with a representation) to determine the room designation in the geometric representation (e.g., dividing a large grid into two rooms without walls between them, such as a kitchen and dining area).

[0125] In some implementations, method 1400 further includes selectively providing a room-specific subset of the geometric representation to an application on an electronic device. For example, to provide better privacy management, the room-specific subset of the geometric representation (e.g., grid data, planar data, etc.) may be restricted to allowing a particular application access to certain rooms (e.g., not allowing it to see inside the bathroom).

[0126] In some implementations, virtual content may not be visible from the XR environment's view when an electronic device moves to another room. For example, a virtual photograph placed on a physical wall may only be visible to someone standing inside the room, but not through the wall, and / or may be invisible to someone standing outside the room but looking at the room from a doorway or another open space.

[0127] Figure 15This is a flowchart illustrating a method 1500 for restricting access to a set of environmental information based on sensor data. In some embodiments, a device such as electronic device 110 performs method 1500. In some embodiments, method 1500 is performed on a mobile device, desktop computer, laptop computer, HMD, or server device. Method 1500 is performed by processing logic components, including hardware, firmware, software, or combinations thereof. In some embodiments, method 1500 is performed on a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory).

[0128] In various embodiments disclosed herein, method 1500 grants an application limited access to sensor-based physical environment information (e.g., data regarding planes, grids, virtual object placement, etc.). Access may be restricted based on associating the sensor-based environment information with corresponding keyframes (e.g., a collection of images / other data obtained from a specific location / pose). The device may associate multiple pieces of sensor-based environment information (e.g., planes, grids, virtual object placement) with corresponding keyframes from which each piece of information is determined. For example, based on determining that keyframe-1 is used to identify plane A, plane A is associated with keyframe-1; based on determining that keyframe-2 is used to identify plane B, plane B is associated with keyframe-2, and so on. In some embodiments, once a user grants permission for the application to have access to the sensor-based environment information, the application is only granted access to privacy information associated with a keyframe that is related to the application's location during use of the application (e.g., during the current application session and / or previous application sessions). Specifically, when an application is used, only a few of the keyframes known to the device are related to the device's location at the time the application is running (e.g., keyframes from nearby locations used for SLAM and / or evaluating the surrounding environment, such as identifying planes, grids, etc.). In the example above, if the application is not used in the location associated with keyframe-2, it will not have access to the information about plane B. The application may only have access to sensor-based environmental information associated with these keyframes, effectively granting the application access only to information "visible" to the device during the application's use.

[0129] At box 1502, method 1500 associates a set of environmental information based on sensor data with corresponding keyframes in a set of keyframes associated with the physical environment. For example, the set of keyframes may include all keyframes that the device has acquired or created for a particular house.

[0130] In some specific implementations, associating a set of environmental information based on sensor data with corresponding keyframes includes: associating a first keyframe with a first plane based on the identification of a first plane using a first keyframe, and associating a second keyframe with a second plane based on the identification of a second plane using a second keyframe. For example, based on determining that keyframe-1 is used to identify plane A, plane A is associated with keyframe-1; based on determining that keyframe-2 is used to identify plane B, plane B is associated with keyframe-2, and so on. In some specific implementations, associating a set of environmental information based on sensor data with corresponding keyframes includes: associating a first keyframe with a first grid based on the identification of a first grid using a first keyframe, and associating a second keyframe with a second grid based on the identification of a second grid using a second keyframe.

[0131] In some embodiments, the set of environmental information based on sensor data is determined based on sensor data of the physical environment obtained via one or more sensors at the device or from another device. Sensor data may include images or depth data of the physical environment. Sensor data may include 3D point clouds and sequences of 2D images corresponding to views of the room captured during a scan of the room. In some embodiments, sensor data includes image data (e.g., from an RGB camera), depth data (e.g., depth images from a depth camera), ambient light sensor data (e.g., from an ambient light sensor), and / or motion data from one or more motion sensors (e.g., accelerometers, gyroscopes, etc.). Sensor data may include keyframes of images. In some embodiments, sensor data includes visual inertial odometry (VIO) data determined based on image data. 3D point clouds can provide semantic information about one or more elements of the room. 3D point clouds can provide information about the location and appearance of surface portions within the physical environment. In some embodiments, 3D point clouds are acquired over time (e.g., during a scan of the room), and can be updated, with updated versions of the 3D point clouds obtained over time. For example, when the 3D representation is updated / adjusted over time (e.g., when a user scans a room, such as...). Figures 8A to 8I As shown, a 3D representation can be obtained (and analyzed / processed).

[0132] At box 1504, method 1500 identifies a subset of the keyframe set that corresponds to the application's use in the physical environment. For example, when the application is used, it determines which keyframes among the device's known keyframes are relevant to the device's location when the application is running. For example, keyframes are identified from nearby locations used for SLAM and / or evaluating the surrounding environment, such as identifying planes, grids, etc.

[0133] At box 1506, method 1500 identifies a subset of the set of environmental information based on sensor data, based on a subset of the set of keyframes. For example, it identifies sensor data associated with a subset of keyframes that has been associated with an application used at a previous location in the environment.

[0134] At box 1508, method 1500 grants the application limited access to a set of environmental information based on sensor data. The limited access is restricted to a subset of the set of environmental information based on sensor data. For example, the application may only have access to information associated with keyframes corresponding to the use of the application, effectively granting the application access only to information based on content "visible" to the device during the use of the application.

[0135] In some implementations, the subset of keyframes in the keyframe set that corresponds to the application's use in the physical environment is identified based on the device's location. Alternatively, in some implementations, the subset of keyframes in the keyframe set that corresponds to the application's use in the physical environment is identified based on the device's pose.

[0136] In some implementations, a set of environmental information based on sensor data is associated with corresponding keyframes based on identifying a room-specific subset of the physical environment. In other words, a limited number of keyframes are selectively identified for each room / segment. Given such room boundaries, the captured data (e.g., keyframes of image and / or depth data) can be selectively used to generate a geometric representation; for example, only those keyframes corresponding to a given room can be used to update the 3D mesh of that room. Additionally or alternatively, in some implementations, room-specific subsets of sensor data can be identified by matching / clustering keyframes with semantic labels of the same room type (e.g., if the keyframes have room labels, such as semantically labeled rooms). Room-specific subsets provide representations of privacy entities as keyframes. Some implementations are chunk-based (e.g., visible or invisible 3x3x3m blocks), but keyframe identification using a subset of the keyframe set corresponding to the application's use in the physical environment provides a more "room-based" application for limiting the data provided to the application based on what is visible from the device's perspective.

[0137] In some implementations, identifying room-specific subsets of sensor data based on room boundary information includes identifying a set of room-specific keyframes for each of multiple rooms. For example, a mesh for a given room is generated, using only the keyframes for that room. Keyframes can be captured for each room during live scanning or during device usage. The keyframe process can compute clusters based on sensor data (e.g., 3D point clouds), then associate the keyframes with rooms to generate coarse boundaries (e.g., low-level estimates) for the rooms, and then use this information to generate the mesh / plane. A refinement module can refine the boundaries of the coarse boundaries (the published room planes) to provide a two-level approach, e.g., coarse then refined.

[0138] In some implementations, if a keyframe is identified as being associated with two different rooms, keyframes can be selected for a specific room based on which keyframe is pointing at which room. For example, the system can view each keyframe by room, score the keyframes, and determine which keyframes to process. Additionally, a confidence level can be determined and scored for each keyframe associated with each room. In some implementations, the system generates a mesh for more than one room at a time, and then the system can trim portions of the mesh that are identified as not being part of a room.

[0139] Figure 16 This is a block diagram of electronic device 1600. Device 1600 illustrates an exemplary device configuration of electronic device 110. Although certain specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features are not illustrated for the sake of brevity and so as not to obscure further relevant aspects of the specific embodiments disclosed herein. Therefore, as a non-limiting example, in some specific implementations, device 1600 includes one or more processing units 1602 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 1606, one or more communication interfaces 1608 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BlueTooth, ZigBee, SPI, I2C and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 1610, one or more output devices 1612, one or more internal and / or external image sensor systems 1614, memory 1620, and one or more communication buses 1604 for interconnecting these components and various other components.

[0140] In some embodiments, one or more communication buses 1604 include circuitry for interconnecting system components and controlling communication between system components. In some embodiments, the one or more I / O devices and sensors 1606 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, or one or more depth sensors (e.g., structured light, time-of-flight, etc.).

[0141] In some embodiments, one or more output devices 1612 include one or more displays configured to present a view of a 3D environment to a user. In some embodiments, one or more output devices 1612 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more displays correspond to waveguide displays such as diffraction, reflection, polarization, and holography. In one example, device 1600 includes a single display. In another example, device 1600 includes displays for each of the user's eyes.

[0142] In some embodiments, one or more output devices 1612 include one or more audio generating devices. In some embodiments, one or more output devices 1612 include one or more speakers, surround sound speakers, speaker arrays, or headphones for generating spatialized sound (e.g., 3D audio effects). Such devices can virtually place sound sources in a 3D environment, including behind, above, or below one or more listeners. One or more output devices 1612 may additionally or alternatively be configured to generate haptic feedback.

[0143] In some embodiments, the one or more image sensor systems 1614 are configured to acquire image data corresponding to at least a portion of the physical environment. For example, the one or more image sensor systems 1614 may include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), monochrome cameras, IR cameras, depth cameras, event-based cameras, etc. In various embodiments, the one or more image sensor systems 1614 also include an illumination source emitting light, such as a flash. In various embodiments, the one or more image sensor systems 1614 also include an on-camera image signal processor (ISP) configured to perform multiple processing operations on the image data.

[0144] Memory 1620 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 1620 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 1620 may optionally include one or more storage devices remotely located to one or more processing units 1602. Memory 1620 includes non-transitory computer-readable storage media.

[0145] In some embodiments, memory 1620 or a non-transitory computer-readable storage medium of memory 1620 stores an optional operating system 1630 and one or more instruction sets 1640. Operating system 1630 includes procedures for handling various basic system services and for performing hardware-related tasks. In some embodiments, instruction set 1640 includes executable software defined by binary information stored in charge. In some embodiments, instruction set 1640 is software executable by one or more processing units 1602 to implement one or more of the techniques described herein.

[0146] Instruction set 1640 includes a plan view instruction set 1642, which is configured to, upon execution: acquire sensor data; provide a view / representation; select a set of sensor data and / or generate a 3D point cloud, 3D mesh, 3D plan view, and / or other 3D representation of the physical environment as described herein. Instruction set 1640 also includes a keyframe instruction set 1644, which is configured to select a keyframe corresponding to a given room, which can be used to update the 3D mesh of that room, as described herein. Instruction set 1640 may be embodied in a single software executable or multiple software executables.

[0147] Although instruction set 1640 is shown residing on a single device, it should be understood that in other embodiments, any combination of elements may reside in separate computing devices. Furthermore, the accompanying drawings serve more as a functional description of the various features present in a particular embodiment, and differ from the structural schematics of the embodiments described herein. As will be appreciated by those skilled in the art, items shown individually may be combined, and some items may be separate. The actual number of instruction sets and how features are allocated therein will vary depending on the specific embodiment and may depend in part on the particular combination of hardware, software, and / or firmware chosen for that particular embodiment.

[0148] It should be understood that the specific embodiments described above are cited by way of example, and this disclosure is not limited to what has been specifically shown and described above. Rather, the scope includes both combinations and sub-combinations of the various features described above, as well as variations and modifications of the various features that would occur to those skilled in the art upon reading the foregoing description and that are not disclosed in the prior art.

[0149] As described above, one aspect of this technology involves collecting and using sensor data, which may include user data, to improve the user experience of electronic devices. This disclosure envisions that, in some cases, the collected data may include personal information data that uniquely identifies a particular person or can be used to identify the interests, characteristics, or preferences of a particular person. Such personal information data may include motion data, physiological data, demographic data, location-based data, telephone numbers, email addresses, home addresses, device characteristics of personal devices, or any other personal information.

[0150] This disclosure recognizes that the use of such personal information data in the present invention can benefit users. For example, personal information data can be used to improve the content viewing experience. Therefore, the use of such personal information data may enable planned control over electronic devices. Furthermore, this disclosure also anticipates other uses of personal information data that benefit users.

[0151] This disclosure further envisions that entities responsible for the collection, analysis, disclosure, transmission, storage, or other use of such personal information and / or physiological data will comply with established privacy policies and / or privacy practices. Specifically, such entities should implement and adhere to privacy policies and measures recognized as meeting or exceeding industry or governmental requirements for maintaining the privacy and security of personal information data. For example, personal information from users should be collected for legitimate and reasonable purposes of the entity and not shared or sold outside of these legitimate purposes. Furthermore, such collection should only be conducted after receiving informed consent from users. Additionally, such entities should take any necessary steps to safeguard and protect access to such personal information data and ensure that others with access to such personal information data comply with their privacy policies and procedures. Moreover, such entities may subject themselves to third-party assessments to demonstrate their compliance with widely accepted privacy policies and practices.

[0152] Regardless of the foregoing, this disclosure also contemplates specific implementations allowing users to selectively block the use or access to personal information data. That is, this disclosure contemplates providing hardware or software components to prevent or block access to such personal information data. For example, with regard to a content delivery service tailored to a user, the technology of this invention can be configured to allow a user to choose to "join" or "opt out" of the collection of personal information data during service registration. In another example, a user may choose not to provide personal information data for a target content delivery service. In yet another example, a user may choose not to provide personal information but allow the transmission of anonymous information for improving device functionality.

[0153] Therefore, while this disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, it is also contemplated that various embodiments can be implemented without access to such personal information data. That is, various embodiments of the present invention will not become inoperable due to the absence of all or part of such personal information data. For example, preferences or settings can be inferred based on non-personal information data or an absolute minimum amount of personal information, such as content requested by a device associated with a user, other non-personal information available to the content delivery service, or publicly available information, thereby selecting content and delivering it to the user.

[0154] In some implementations, data is stored using a public / private key system that allows only the data owner to decrypt the stored data. In other implementations, data may be stored anonymously (e.g., without identification and / or without personal information about the user, such as legal name, username, time, and location data). This prevents other users, hackers, or third parties from identifying the user associated with the stored data. In some implementations, users can access their stored data from a different user device than the one used to upload the stored data. In these cases, users may need to provide login credentials to access their stored data.

[0155] This document sets forth numerous specific details to provide a comprehensive understanding of the claimed subject matter. However, those skilled in the art will understand that the claimed subject matter can be practiced without these specific details. In other instances, methods, apparatus, or systems known to a person of ordinary skill have not been described in detail so as not to obscure the claimed subject matter.

[0156] Unless otherwise specifically stated, it should be understood that throughout this specification, discussions using terms such as “processing,” “calculating,” “calculating,” “determining,” and “identifying” refer to the actions or processes of computing devices, such as one or more computers or similar electronic computing devices, which manipulate or convert data representing physical electronic or magnetic quantities within the memory, registers, or other information storage, transmission, or display devices of a computing platform.

[0157] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include computer systems based on multi-purpose microprocessors that access stored software that programs or configures the computing system from a general-purpose computing device to a special-purpose computing device that implements one or more specific embodiments of the subject matter of this invention. The teachings contained herein may be implemented in the software used for programming or configuring the computing device using any suitable programming, scripting, or other type of language or combination of languages.

[0158] Specific implementations of the methods disclosed herein can be performed in the operation of such computing devices. The order of the boxes presented in the example above can be varied; for example, the boxes can be reordered, combined, and / or divided into sub-blocks. Some boxes or processes can be executed in parallel.

[0159] The use of "applies to" or "configured to" in this document implies open and inclusive language, which does not exclude applicability to or configuration for performing additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, as processes, steps, calculations, or other actions "based on" one or more of the stated conditions or values ​​may in practice be based on additional conditions or values ​​beyond those stated. The headings, lists, and numbering included herein are for illustrative purposes only and are not intended to be restrictive.

[0160] It will also be understood that while terms such as "first," "second," etc., may be used in this document to describe various elements, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first node may be called a second node, and similarly, a second node may be called a first node, changing the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. First nodes and second nodes are both nodes, but they are not the same node.

[0161] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of these embodiments and the appended claims, the singular forms “a,” “an,” and “the” are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will be further understood that the term “comprising,” as used in this specification, specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0162] As used herein, the term "if" can be interpreted as meaning "when the prerequisite is true" or "when the prerequisite is true" or "in response to determination" or "according to determination" or "in response to detection" that the prerequisite is true, depending on the context. Similarly, the phrases "if it is determined [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" are interpreted as meaning "when it is determined that the prerequisite is true" or "in response to determination" or "according to determination" that the prerequisite is true or "when it is detected that the prerequisite is true" or "in response to detection" that the prerequisite is true, depending on the context.

[0163] The foregoing description and summary of the present invention should be understood as illustrative and exemplary in every respect, and not restrictive, and the scope of the invention disclosed herein is determined not only by the detailed description of the illustrative specific embodiments, but also by the full extent permitted by patent law. It should be understood that the specific embodiments shown and described herein are merely illustrative of the principles of the invention, and various modifications can be made by those skilled in the art without departing from the scope and spirit of the invention.

Claims

1. A method comprising: In electronic devices with processors: Obtain sensor data of the physical environment of multiple rooms, the sensor data including images of the physical environment; Obtain room boundary information associated with the physical environment, wherein the room boundary information is determined based on the sensor data; Based on the room boundary information, identify a room-specific subset of the sensor data; as well as One or more geometric representations of the physical environment are generated based on a room-specific subset of the sensor data.

2. The method of claim 1, wherein identifying a room-specific subset of the sensor data based on the room boundary information comprises: Identify the set of room-specific keyframes for each of the multiple rooms.

3. The method according to claim 2, wherein the sensor data includes first sensor data obtained within a first time period, and the method further includes: Obtain second sensor data within a second time period, wherein the second sensor data corresponds to the first room among the plurality of rooms; as well as In response to obtaining the second sensor data, the room-specific keyframe set of the first room in the plurality of rooms is updated.

4. The method according to claim 3, further comprising: The three-dimensional (3D) representation of the first room is updated based on the updated set of room-specific keyframes.

5. The method according to claim 1, further comprising: Based on the room-specific subset of the sensor data, a set of planes associated with each of the plurality of rooms is determined.

6. The method of claim 5, wherein determining the set of planes for each room comprises: Identify the floor plane, ceiling plane, and wall plane of at least one or more walls for each room.

7. The method according to claim 1, further comprising: A three-dimensional (3D) representation of the physical environment is generated based on the one or more geometric representations.

8. The method according to claim 7, further comprising: A real-time view of the 3D representation is displayed on the screen of the electronic device.

9. The method of claim 8, wherein the real-time view of the room includes a floor plan generated when the sensor data is acquired.

10. The method of claim 1, wherein the image of the physical environment is based on at least one of: a real-time view image, an ultrawide view image including a view different from the real-time view image, and a semantically labeled image corresponding to the real-time view image or the ultrawide view image.

11. The method of claim 1, wherein the room boundary information is determined based on 3D semantic data including a 3D point cloud, the 3D point cloud including semantic labels associated with at least a portion of the 3D points within the 3D point cloud.

12. The method of claim 11, wherein the semantic tags identify the walls, wall structures, objects, and classifications of the objects in each room.

13. An apparatus comprising: Non-transitory computer-readable storage medium; and One or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions that, when executed on the one or more processors, cause the one or more processors to perform operations, the operations including: Obtain sensor data of the physical environment of multiple rooms, the sensor data including images of the physical environment; Obtain room boundary information associated with the physical environment, wherein the room boundary information is determined based on the sensor data; Based on the room boundary information, a room-specific subset of the sensor data is identified; and One or more geometric representations of the physical environment are generated based on a room-specific subset of the sensor data.

14. The device of claim 13, wherein identifying a room-specific subset of the sensor data based on the room boundary information comprises: Identify the set of room-specific keyframes for each of the multiple rooms.

15. The device of claim 14, wherein the sensor data includes first sensor data obtained within a first time period, and the method further comprises: Obtain second sensor data within a second time period, wherein the second sensor data corresponds to the first room among the plurality of rooms; as well as In response to obtaining the second sensor data, the room-specific keyframe set of the first room in the plurality of rooms is updated.

16. The apparatus of claim 15, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, further cause the one or more processors to perform operations including: The three-dimensional (3D) representation of the first room is updated based on the updated set of room-specific keyframes.

17. The apparatus of claim 13, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, further cause the one or more processors to perform operations including: Based on the room-specific subset of the sensor data, a set of planes associated with each of the plurality of rooms is determined.

18. The device of claim 17, wherein determining the set of planes for each room comprises: Identify the floor plane, ceiling plane, and wall plane of at least one or more walls for each room.

19. The apparatus of claim 13, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, further cause the one or more processors to perform operations including: To generate a three-dimensional (3D) representation of the physical environment based on the one or more geometric representations; and A real-time view of the 3D representation is presented on the display of the electronic device, wherein the real-time view of the room includes a floor plan generated when the sensor data is acquired.

20. A non-transitory computer-readable storage medium storing program instructions executable on a device to perform operations, the operations including: Obtain sensor data of the physical environment of multiple rooms, the sensor data including images of the physical environment; Obtain room boundary information associated with the physical environment, wherein the room boundary information is determined based on the sensor data; Based on the room boundary information, identify a room-specific subset of the sensor data; as well as One or more geometric representations of the physical environment are generated based on a room-specific subset of the sensor data.