Localization of a user in a physical environment

The method addresses localization challenges by scanning and labeling objects in a physical environment, constructing virtual scenes, and comparing pose data to achieve precise user positioning and orientation, enhancing applications like navigation and augmented reality.

WO2025178523A1PCT designated stage Publication Date: 2025-08-28INTER IKEA SYST +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/SE2024/051007
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-21
Filing Date
2024-11-27
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing localization techniques face challenges such as signal blockage, interference, and high computing resource requirements, particularly in indoor environments, making them impractical for applications like navigation and augmented reality.

Method used

A computer-implemented method that scans a physical environment, assigns semantic labels to objects, constructs a virtual scene, and compares pose data with predetermined virtual scenes to localize users based on partial matches, using computer vision and optimization algorithms.

Benefits of technology

Enables precise user localization with reduced resource utilization, overcoming limitations of existing methods by providing accurate positioning and orientation in various environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SE2024051007_28082025_PF_FP_ABST
    Figure SE2024051007_28082025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method (100) comprising scanning (110) a physical environment (10) for a plurality of objects (12); assigning (120) a semantic label to recognized objects (12); constructing (130) a virtual scene (20) involving virtual representations (22) of the objects (12) based on the semantic labels; comparing (140) pose data of the constructed virtual scene (20) to pose data a predetermined virtual scene (30), the predetermined virtual scene (30) involving predetermined object representations (32) being semantically labelled; and in response to the comparing (140) indicating an at least partial match between pose data of two or more of the virtual representations (22) and pose data of two or more of the predetermined object representations (32), localizing (150) a user (40) in relation to two or more objects (12) in the physical environment (10) whose virtual representations (22) caused the at least partial match.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] LOCALIZATION OF A USER IN A PHYSICAL ENVIRONMENT

[0002] TECHNICAL FIELD

[0003] The present invention relates to a computer-implemented method for localization of a user in a physical environment. The present invention also relates to an associated computerized system, storage medium and computer program product.

[0004] BACKGROUND

[0005] Localization techniques aim to determine where a user or a device is located within a physical environment. The purpose of localization may be to enable certain applications like navigation, augmented reality experiences, location-aware services, proximity-based interactions, resource allocation, security and access control, and personalization of services, to name a few examples. The present inventor has identified limitations with contemporary user localization techniques and is herein presenting improvements in this regard.

[0006] SUMMARY

[0007] Various established localization techniques are prevalent in the field, each offering distinct approaches. Satellite-based localization approaches involve receiving satellite signals that are processed using e.g. triangulation techniques to calculate locations based on signal arrival times. In radio-based localization, such as Wi-Fi or Bluetooth, signal strengths or time-of-flight are measured for providing locational insights. Other radio-based approaches, such as RFID, employ electromagnetic fields for object identification and tracking. Vision-based tracking techniques, that are also known in the art to some extent, either employ scanning techniques to detect and enumerate visual features or involve the use of visual markers.

[0008] Despite the utility of known localization techniques, each approach is associated with limitations that warrant continuous improvement. Satellite-based approaches often struggle with indoor positioning due to signal blockage and reduced accuracy in confined spaces, making it less effective for applications within buildings. Radio localization methods, while often suitable for indoor environments, can be affected by signal interference, multi-path effects, and limitations in dense urban areas. Visionbased tracking faces challenges related to variable lighting conditions, occlusions, or requires a complex setup including visual markers, thereby rendering the approaches impractical or a subject to substantial computing resource requirements. This is especially the case for complex visual processing or real-time analysis, such as user localization.

[0009] The present invention seeks to mitigate, alleviate eliminate or circumvent one or more of the above-mentioned deficiencies and / or disadvantages in the art, singly or in any combination, by providing in a first aspect a computer-implemented method for localizing a user in a physical environment. The method comprises scanning the physical environment for a plurality of objects, assigning a semantic label to each one of the plurality of objects recognized in the physical environment, constructing a virtual scene involving virtual representations of the plurality of objects based on the semantic labels, comparing pose data of the constructed virtual scene to pose data of at least one predetermined virtual scene, each predetermined virtual scene involving a plurality of predetermined object representations being respectively semantically labelled, and in response to the comparing indicating an at least partial match between pose data of two or more of the virtual representations and pose data of two or more of the predetermined object representations, localizing the user in relation to two or more objects in the physical environment whose virtual representations caused the at least partial match.

[0010] In one or more embodiments, the localizing comprises calculating average pose data of said two or more objects in the physical environment whose virtual representations caused the at least partial match and localizing the user with respect to the average pose data.

[0011] In one or more embodiments, the localizing comprises applying an optimization algorithm.

[0012] In one or more embodiments, the localizing of the user is performed in relation to all of the objects in the physical environment whose virtual representations caused the at least partial match in the comparing. In one or more embodiments, the comparing comprises calculating similarity metrics of the pose data of the constructed virtual scene and of the pose data of the at least one predetermined virtual scene.

[0013] In one or more embodiments, assigning the semantic labels comprises applying a computer vision algorithm to the plurality of objects being scanned.

[0014] In one or more embodiments, the computer vision algorithm ranks the plurality of objects based on object types, each rank indicating a level of pose confidence.

[0015] In one or more embodiments, substantially immovable objects are assigned relatively higher ranks compared to movable objects.

[0016] In one or more embodiments, the at least partial match is a match of orientation data and position data of the respective pose data.

[0017] In one or more embodiments the at least partial match is a match of an extent of the semantic labels.

[0018] In one or more embodiments, said at least one predetermined virtual scene is obtained as output from a software-based planner tool or output from a mobile computing device having scanned and semantically labelled a physical environment involving a plurality of objects.

[0019] In one or more embodiments, said at least one predetermined virtual scene is stored in a look-up table comprising a plurality of entries, each entry involving a combination of two or more predetermined object representations.

[0020] In one or more embodiments, the comparing further comprising applying one or more filters to limit the total size of the look-up table.

[0021] In one or more embodiments the filters comprise one or more of a geographical filter, IMU filter, object category filter, spatial surroundings filter, virtual scene filter and a numeric filter.

[0022] In one or more embodiments, based on the determined pose of the user, the method further comprising one or more of: displaying one or more graphical indicators, providing audio adjustments, providing haptic feedback, providing olfactory feedback, and providing ambient simulations.

[0023] In one or more embodiments, the graphical indicators comprise one or more of: cues of spatial interactions associated with the objects and / or the physical environment, cues of environmental adaptations associated with the objects and / or the physical environment, cues of user-safety features associated with the objects and / or the physical environment, cues of movement guiding indicators in the physical environment.

[0024] In one or more embodiments, the localizing comprises determining a position and / or orientation of the user.

[0025] In a second aspect, a computerized system is provided. The computerized system comprises a processor configured to perform the functionality of the method according to the first aspect.

[0026] In a third aspect, a non-transitory computer-readable storage medium is provided. The computer-readable storage medium comprises instructions, which when executed by one or more processors of a computerized system, cause the processor to perform the functionality of the method according to any one of the embodiments.

[0027] In a fourth aspect, a computer program product is provided. The computer program product comprises computer code for performing the method according to the first aspect or any of the embodiments being dependent thereon when the computer program code is executed by a processing device.

[0028] Further advantageous features of the invention are elaborated in embodiments disclosed herein. In addition, advantageous features of the invention are defined in the dependent claims. All references to "a / an / the [element, device, component, means, step, etc.]" are to be interpreted openly as referring to at least one instance of the element, device, component, means, step, etc., unless explicitly stated otherwise. The steps of any method disclosed herein do not have to be performed in the exact order disclosed, unless explicitly stated.

[0029] BRIEF DESCRIPTION OF THE DRAWINGS

[0030] These and other aspects, features and advantages of which the invention is capable of will be apparent and elucidated from the following description of embodiments of the present invention, reference being made to the accompanying drawings in which:

[0031] FIG. 1 is an exemplary illustration of user localization in a physical environment according to one example. FIG. 2 is an exemplary flowchart diagram of a method for localizing a user according to various examples.

[0032] FIGs 3A-D are various exemplary illustrations of localizing a user in physical environments.

[0033] FIG. 4 is an exemplary illustration of various graphical indicators that can be provided once the user has been localized in the physical environment.

[0034] FIG. 5 depicts an exemplary schematic view of a non-transitory computer- readable storage medium.

[0035] FIG. 6 depicts a block diagram of a computerized system according to one embodiment.

[0036] DETAILED DESCRIPTION

[0037] The present invention will now be described more fully hereinafter with reference to the accompanying drawings, in which currently preferred embodiments of the invention are shown. This invention can, however, be embodied in many different forms and should as such not be interpreted as limited to the embodiments set forth herein. These embodiments are provided for thoroughness and completeness, and to fully convey the scope of the invention to the skilled person.

[0038] With reference to FIG. 1, an exemplary physical environment 10 is shown where a user 40 is present. This particular physical environment 10 is an indoor room, more specifically a living room, but any alternative physical environment can be readily envisaged in which a user can be located. Other exemplary physical environments include indoor locations such as a home environment (bathroom, kitchen, dining room, bedroom, laundry room, etc.), office environment (cubicle, conference room, reception area, waiting area, etc.), shopping environment (department store, clothing store, electronics store, food court, bookstore, spa area, etc.), arena environment (sports venue, football stadium, ice hockey rink, concert venue, etc.), outdoor environment (park, playground, sports venue, garden, etc.), hybrid environment (balcony, patio, terrace, outdoor dining area, etc.), other environment (hotel room, hospital room, classroom, gym / fitness center, library, movie theater, airport lounge, music studio, art gallery, laboratory, etc.), or the like. These are just a few exemplary physical environments where the subject matter of the present disclosure may be applied and can vary widely based on specific designs and purposes of application areas.

[0039] The physical environment 10 comprises a plurality of objects 12. The type of objects generally depends on the type of environment, and it shall thus be understood that different physical environments include different types of objects. For the living room shown in FIG. 1, the objects 12 include, amongst other things, a computer screen 12-1, a ceiling 12-2, a wall 12-3 extending around the room, a couch 12-4, a mat 12-5, a magazine container 12-6 and a bookcase 12-7.

[0040] The user 40 is, in this example, carrying out a localization procedure. As discussed in the background section, localization aims to determine where the user 40 or a related mobile computing device 50 is located within the physical environment 10. More specifically, the localization may involve determining a position and / or orientation of the user 40. While the user 40 intuitively will understand its location in relation to the various objects 12 in the room, localization has to be performed in order for a computer system to know where the user 40 is located. This process serves various purposes, such as integrating computer-generated content into the perspective of the user 40, or conducting calculations tied to the location of the user 40 (e.g., for applications like navigation, augmented reality experiences, and location-aware services, proximity-based interactions, resource allocation, security and access control, and personalization of services), thereby enabling enhanced functionalities. Accordingly, the present inventor has developed a way of establishing a local coordinate system such that the location of the user 40 can be pinpointed with high precision with respect to the spatial surroundings of the physical environment 10.

[0041] FIG. 1 highlights some notable features of an exemplary user localization approach. In a first step shown in the top illustration of the room, the user 40 scans the physical environment 10 for purposes of recognizing the objects 12. In this example this is done using an exemplary mobile computing device 50. The mobile computing device 50 may be an image capturing device configured to capture images of the surroundings, such as a VR headset, smartphone, tablet, digital camera, smart glasses, body cameras, 360-degree cameras, wearable action cameras, VR headset, or the like. The mobile computing device 50 may further obtain one or more sets of sensor- retrieved data from various multimodal sources. The sets of sensor-retrieved data may include point cloud data other than those obtained through visual imagery. The additional sets of sensor-retrieved data may enable the creation of more detailed and dynamic models of the room, preferably processed in tandem with the visual point cloud data for improving the localization accuracy. To this end, the one or more sets of sensor-retrieved data may complement the visual scanning of the environment.

[0042] Point cloud data may be obtained from sensors arranged in the mobile computing device 50, or sensors arranged external to the mobile computing device 50 and operable to communicate said point cloud data to the mobile computing device 50. The point cloud data may include wave sensing data such as electromagnetic wave data for non-visible light, including active or passive configurations such as time of flight data, radar data, LIDAR data, or the like. The point cloud data may relate to inertial measurement data such as accelerometer data, gyroscope data, magnetometer data, or the like. The point cloud data may include static or dynamic air pressure data, which may be active and / or passive, such as ultrasonic transducer data, microphone data, broadband acoustic impulse generator data, altimeter data, piezoresistive pressure data, or the like. The point cloud data may include contact sensing data such as capacitive pressure data, microelectromechanical systems (MEMS) pressure data, strain gauge pressure transducer or manometer data, touch and collision sensor data, or the like. The point cloud data may include static electric or magnetic field data such as electrostatic or magnetostatic sensing data. The point cloud data may include vibration data that can deduce the size of resonant surfaces in a space.

[0043] Purely by way of example, one or more of the above-mentioned point cloud data can be obtain by inductive sensors, capacitive sensors, hall effect sensors, eddy current sensors, thermal sensors, ultrasonic sensors, RFID sensors (wherein objects or locations in the environment may be equipped with an RFID tag) vibration sensors, particle sensing devices, pressure sensors, or sound sensors. The sensors may be used to capture information relating to the location, size, or shape of the objects, or relative distance measurements during capture with other sensors to enhance the accuracy of other sensors readings. The one or more sensors may work independently or in conjunction with one another (e.g. sensor fusion).

[0044] While the user scans 40 the objects 12 (and optionally retrieves other point cloud data of the surroundings according to one or more of the examples described above), correctly recognized objects will be assigned respective labels. In this example the couch is recognized as “Couch”, the floor as “Floor”, the wall as “Wall”, the ceiling as “Ceiling”, the window as “Window”, the table as “Table”, and the chair as “Other”. The semantic labels of FIG. 1 are just examples, and it shall be understood that a plurality of the other objects 12 of the room may alternatively or additionally be assigned with respective semantic labels. Other appropriate semantic labels may be applied based on object type and the implemented scanning technique. The semantic label may be a tag or other identifier assigned to the recognized objects 12. The semantic label may be used to convey the meaning or semantics of the recognized objects 12. The meaning of semantics may refer to information about the type or characteristic of the objects 12. A type may alternatively be understood as a category or class, such as the labels of FIG. 1. The characteristic may be a position property, an orientation property, a dimension property, a weight property, a color property, a shape property, a texture property, a temperature property, a material property, a transparency property, a state property, a density property, an age property, or the like. The characteristic and / or type may be used for purposes of localizing the user, as will be discussed in further detail with reference to FIG. 2.

[0045] In some embodiments, the objects 12 are broken down into constituent objects, each of which are assigned a semantic label, and treated as an object 12 according to the above.

[0046] In some embodiments, the objects 12 are broken down into minimal constituent objects according to their semantic labels, i.e. the smallest possible object which can be associated with a relevant semantic label. For example, a “Couch” may be broken down into “Pillows”, “Chaise longue”, and “Seat module”.

[0047] Based on the semantic labels of the objects 12 that have been recognized by the scanning, a virtual scene 20 is constructed. This is shown in the bottom right illustration. The virtual scene 20 comprises a plurality of virtual representations 22 of the plurality of objects 12 that had respective semantic labels assigned to them. It shall be noted that the virtual scene 20 does not necessarily correspond to the physical environment 10 because not all objects 12 need to be scanned. It is sufficient that at least two objects 12 are recognized (a plurality of objects can include two or more), assigned with respective semantic labels, and serve as a basis for the virtual scene 20 to be constructed. In the example, the virtual scene 20 includes virtual representations 22 of five of the objects 12, namely virtual representations of a ceiling 22-1, a first part of a wall 22-2, a couch 22-3, a table 22-4 and a second part of a wall 22-5. In addition, it shall be understood that one or more of the objects 12 can be associated with different virtual placements in the virtual scene 20, such as the table in the shown example.

[0048] Since the virtual representations 22 in the virtual scene 20 are based on semantic labels, they include respective information related to the semantic labels, such as the type and characteristic as discussed above. In the bottom right illustration, the characteristics of position property (denoted al; a3) and orientation property (denoted xl, yl, zl; x3, y3, z3) are depicted. In this example three-dimensional coordinates are employed, but for other examples at least two-dimensional coordinates may be envisaged. Together these two properties correspond to pose data of the virtual representation 22 of the corresponding object 12 to which the semantic label has been assigned.

[0049] Pose data of the constructed virtual scene 20 is then compared to pose data of at least one predetermined virtual scene 30. This is depicted by the arrow from the bottom right illustration to the bottom left illustration. A predetermined virtual scene 30 shall be understood as a reference virtual scene being stored in a storage location, such as a cloud-storage unit, a local storage device, or other suitable storage means. The predetermined virtual scene may be obtained as output from a software-based planner tool, such as room planner, or previous outputs from a mobile computing device, such as from a previously scanned room. For example, in the embodiment using a predetermined scene obtained from a software-based planner tool, the virtual scene may be compared to a predetermined virtual scene previously created and rendered, in 2D or 3D, in programs such as CAD, Blender, D5 Render, Enscape 3D or Unity. The predetermined virtual scene may correspond to multiple similar physical environments, enabling accurate detection of the user in any such environments using only one predetermined scene. In a specific example, a company may build a plurality of similar stores, including similar objects. Regardless of which of the plurality of stores a user is in, they may still be accurately located therewithin using a single predetermined scene that matches two or more of plurality of stores. Yet alternatively, any hybrids between a previously created and rendered scene, and a real scene, may be envisaged as exemplary predetermined scenes.

[0050] Each predetermined virtual scene, such as the predetermined virtual scene 30, comprises a plurality of predetermined object representations 32 being respectively semantically labelled. This means that each predetermined object representation 32 included in the predetermined virtual scene 30, in this case predetermined object representations of a ceiling 32-1, a first part of a wall 32-2, a couch 32-3, a table 32-4 and a second part of a wall 32-5, are semantically labelled. Since the predetermined object representations 32 are respectively semantically labelled, they also include information pertaining to types and characteristics, including pose data. This pose data is thus compared to the pose data of the virtual representations 22 as discussed above, and in this example more specifically a position property (denoted a2; a4) and an orientation property (denoted x2, y2, z2; x4, y4, z4). By comparing pose data in this way, a local coordinate system can be established with respect to a position, and optionally also an orientation, to where the user 40 was located upon scanning the objects 12.

[0051] For the localization to be possible, the comparison requires an at least partial match between the pose data of two or more of the virtual representations 22 and pose data of two or more of the predetermined object representations 32. For example, the user 40 may be given an estimate origin position of (0, 0, 0) in the local coordinate system, and the two or more matching objects 12, 22, 32 be given respective positions in the same local coordinate system with respect to the user position of (0, 0, 0). To this end, the user 40 can be localized in relation to said two or more objects 12 in the physical environment 10 whose virtual representations caused the at least partial match. What this means is generally that the user 40 can be localized in the physical environment 10 merely by way of scanning, recognizing, and assigning semantic labels to at least two objects 12, provided that said at least partial match can be established for (pose data of) virtual representations 22 of said at least two objects with (pose data of) predetermined object representations 32 of at least two objects. This approach includes advantages of being able to employ vision-based scanning techniques, sensor-based scanning techniques, or sensor-fusion based scanning techniques, that rely on semantic labels, and using the information provided by said semantic labels for two or more objects 12 to localize the user 40. Thanks to the storage and comparison of the constructed virtual scene 20 with (one or more) predetermined virtual scenes, together with insights that pose data of two or more objects 12 are required, the approach is efficient in terms of both time and resource utilization. Moreover, the approach has been shown to be surprisingly accurate in terms of user localization preciseness.

[0052] An at least partial match may be a match of orientation data and position data of the respective pose data of the compared objects (i.e., the virtual representations 22 and the predetermined object representations 32). In some examples, the at least partial match may be a match of an extent of semantically labelled objects 12. In these examples, the semantically labelled objects are also assigned respective bounding boxes during the assigning step. The extent of a bounding box of an object 12 refers to a size of a bounding volume that is a closed volume which contains, preferably as tightly as possible (i.e., with an as small enclosure as possible), the virtual representation 22 of the object 12. As a mere exemplary form, the bounding box may be an oriented or axis- aligned bounding box, capsule, cylinder, ellipsoid, sphere, slab, triangle, convex or concave hull, eight-direction discrete orientation polytope, or any combination thereof. The extent of the bounding box may also involve salient data points, similar to the point clouds, for the employed algorithm. Considering the extent of the bounding box may thus improve the localization accuracy in certain cases where very similar objects in appearance are compared to one another, but with different sizes. An example could be two tables with identical designs but created with different dimensional attributes, such as 50 cm x 100 cm and 100 cm x 200 cm, respectively.

[0053] FIG. 2 is an exemplary flowchart of a computer-implemented method 100 for localizing a user 40 in a physical environment 10. The boxes depicted in uniform lines correspond to more general steps of the method 100, while the boxes depicted in dashed lines are optional embodiments with respective technical advantages. One or more of these optional embodiments may be considered in combination.

[0054] The method 100 involves a step of scanning 110 the physical environment 10 for a plurality of objects 12. As discussed in relation to FIG. 1, the scanning may in some examples be performed by a mobile computing device 50 applying 112 a near real-time scanning technique. The near real-time scanning technique may be based on algorithms known in the art, such as ARKit, which utilizes feature point tracking, simultaneous localization and mapping (SLAM), visual-inertial odometry (VIO) plane detection and motion and orientation tracking. Other known near real-time scanning technologies may include ARCore or OpenCV, or the like. The scanning 110 may be performed one or more times and can for example be done by the user 40 navigating its way across and / or around the physical environment 10 while using the mobile computing device 50 to capture images of the surroundings.

[0055] The method 100 further involves a step of assigning 120 a semantic label to each one of the plurality of objects 12 recognized in the physical environment 10. This may be done during the scanning 110 while each object 12 is recognized, or at a later stage once at least portions of the physical environment 10 have been scanned. The purpose of assigning semantic labels to the objects 12 is to be able to create a parametric room representation (which will later be referred to as a virtual scene). The assigning 120 of semantic labels may be done using various different techniques known in the art.

[0056] In some examples, the assigning 120 of the semantic labels may comprise applying 122 a computer vision algorithm to the plurality of objects 12 being scanned. By way of example, semantic labels may be assigned using a supervised learning approach, such as a convolutional neural network (CNN) or a single-shot detector (SSD). In examples where a convolutional neural network is used, the CNN may be trained on a labelled dataset where the objects 12 are annotated with their corresponding semantic labels. The CNN is configured to learn how to extract features and patterns that characterize each semantic label during the training process. The CNN may be region-based (e.g. R-CNN, optionally Faster R-CNN), which additionally provides bounding boxes around the objects which are assigned with semantic labels. This relates to the embodiment as discussed above of matches between extents of semantic labels. Other techniques may include image segmentation (e.g. using U-Net or DeepLab), transfer learning (e.g. using pretrained models such as ImageNet), or the like.

[0057] In some examples, the computer vision algorithm as discussed above further ranks the plurality of objects 12 based on respective object types, where each rank indicates a level of pose confidence. The pose confidence may be any number, for instance between 0-100, which tells us how confidently the pose of the corresponding object can be used for purposes of comparing pose data at a later stage (which will be described shortly). Generally, a higher pose confidence allows for a higher chance of success when comparing pose data of virtual representations 22 to pose data of the predetermined object representations 32. Hence, the pose confidence may be a value indicating how reliably it can be used for purposes of localizing the user 40.

[0058] In some examples, substantially immovable objects 12 are assigned relatively higher ranks compared to movable objects 12. Substantially immovable objects 12, such as walls, ceiling, windows, ducts, etc., are less likely to be moved around in e.g. a room. Therefore, the insight is made that substantially immovable objects 12 have a higher pose confidence with regards to an estimated position. This is in comparison with substantially movable objects, such as chairs, tables, mats, electronic devices, etc. It is therefore more likely that a subsequent comparison of virtual representations 22 of substantially immovable objects with corresponding predetermined object representations 32 of substantially immovable objects results in a more accurate match. This is due to the fact that these objects are less likely to be associated with a plurality of different positions in different scenes.

[0059] In some examples, the objects 12 are ranked based on whether they fit in a single image, wherein objects 12 that fit within a single image of the scene 20 are ranked higher than objects that do not fit within a single image, or frame, of the scan. To this end, the method may involve identifying, for each image, if the whole object fits, i.e., can be seen, completely in at least one image of the scan.

[0060] In some examples, it is determined that an object 12 fits in a single image based on predefined rules and an object’s semantic label. For example, the predefined rules may determine that certain objects are assumed to fit in a single image or assumed to not fit in a single image. Specifically, the predefined rules may say that objects that are often large, such as walls, are assumed to not fit in a single image and objects that are often smaller such as windows and doors will fit in a single image. Alternatively, the single frame determination may comprise determining if an object’s bounding box fits inside an image.

[0061] In some embodiments, objects 12 that have been assigned semantic labels indicating that the pose of the object 12 is difficult to determine may be ranked lower, or altogether omitted. For example, semantic labels indicating a spatial symmetry, such as a rotational symmetry (i.e. round, circular, cylindrical, etc) may be disregarded or ranked lower.

[0062] The method 100 further involves a step of constructing 130 a virtual scene 20 involving virtual representations 22 of the plurality of objects 12 based on the semantic labels. The virtual scene 20 shall be understood as a parametric room representation of the physical environment 10. The parametric representation is a mathematical description of curves, surfaces or objects in which coordinates are expressed as functions of one or more parameters. The parameters act as independent variables that determine the shape and position of geometric entities, in this case being the objects 12. This approach enables dynamic and flexible modeling, allowing for changes in shape and orientation by adjusting the parameter values. One way of constructing the virtual scene 20 is to utilize a room layout estimation (RLE) process. The RLE process includes creating semantic point clouds, performing semantic projection Z-slicing and end-to-end line detection, providing a 2D room layout, and optionally various postprocessing techniques (such as anti-aliasing, depth of field, motion blue, bloom, etc.).

[0063] The method 100 further involves a step of comparing 140 pose data of the constructed virtual scene 20 to pose data of at least one predetermined virtual scene 30. Each predetermined virtual scene 30 includes a plurality of predetermined object representations 32 being respectively semantically labelled. Pose data includes position data and orientation data. Position data refers to where the virtual representations 22 and the predetermined object representations 32 are located in the respective virtual scenes 20, 30. That is, how they are positioned in relation to one another, and ultimately to the user 40 (which will be described in more detail soon). The position data may comprise spatial coordinates representative of a space of the virtual scenes 20, 30. The position data may alternatively comprise coordinates in a plane in the space of the virtual scenes 20, 30. Orientation data refers to how the virtual representations 22 and the predetermined object representations 32 are oriented in the respective virtual scenes 20, 30 in relation to one or more axes, such as a horizontal or vertical axis extending through the scenes 20, 30.

[0064] In some examples, the predetermined virtual scene 30 comprises synthetic objects, wherein a synthetic object is an object wholly created in a virtual rendering of a scene. One or more synthetic objects may be input by a user through a room-planner tool. The pose and position of these objects may be pre-calculated, i.e. calculated before the comparing 140. There is thus no need for a dedicated scanning in order to generate the predetermined virtual scene 30.

[0065] In some examples, the comparing 140 involves calculating 142 similarity metrics of the pose data of the respective scenes 20, 30, and more specifically of the pose data of the virtual / object representations 22, 32 of the scenes 20, 30. The similarity metrics are based on the pose data, i.e., how similar the position data and the orientation data are to one another. The similarity metrics may be calculated using a point set registration (PSR) method. The PSR involves aligning a point cloud of the virtual representation 22 with a point cloud of the predetermined object representation 32. The aligning may involve estimating relative transformations between the point clouds. By way of employing a PSR method, multiple datasets can be amalgamated into a common coordinate system for purposes of user localization. Any PSR method known in the art may be employed for purposes of calculating similarity metrics, such as Iterative Closest Point (ICP), SoftAssign, Expectation Maximization (EM)-ICP, Levenberg- Marquardt (LM)-ICP, Kernel Correlation (KC), Trimmed Iterative (Tr)-ICP, Fast (F)- ICP, Generalized (G)-ICP, Voxelized Generalized (VG)-ICP, Coherent Point Drift (CPD), Expectation Conditional Maximization for Point Registration (ECMPR), Gaussian Mixture Model Registration (GMMReg), Normal Distribution Transform Normal Distribution Transform (NDT), Robust EM- Segmentation, Maximum Likelihood Mixture Decoupling (MLMD), Support Vector Regression (SVR), Joint Registration of Multiple Point Sets (JRMPC), or Hierarchical Gaussian Mixtures for Adaptive 3D Registration (HGMR). In some examples, the calculating 142 comprises calculating similarity metrics of the pose data for one pose for each virtual representation 22, 32. A single pose comparison calculation comprises constructing a 3D object through temporal orientation of poses of objects in a scan, taking into account the specific time of capture. This 'time' can refer to intervals between images in a scan, movement between points, or a globally synchronized timestamp applicable to the images.

[0066] In some examples, the comparing 140 may involve calculating 142 similarity metrics of the pose data for two poses for each virtual representation 22, 32. In such dual pose comparisons, the pose is first compared according to an image from a scan, followed by a comparison using a second image from the scan. Additionally, calculating similarity metrics may further comprise a statistical method that weighs multiple poses based on the time of obtaining the pose, i.e. weighting historical poses wherein the most recent pose is weighted higher than older poses. The weighting may be a logarithmic weighting.

[0067] In some examples, the at least one predetermined virtual scene 30 may be stored in a look-up table. The look-up table is a data structure, or a mathematical function, used to map input values to corresponding output values. Each entry of the look-up table involves a combination of two or more predetermined object representations 32. For example, the look-up table may be an array of entries where each entry stores two or more predetermined object representations 32. The look-up table may be predetermined for each virtual scene 30. Data entries of the look-up table may thus be retrieved and compared to the contents of the virtual scene 20, i.e., the virtual representations 22 of the objects 12. The system can quickly retrieve the precomputed results from the look-up table, which may make comparisons to the virtual representations 22 efficient.

[0068] In some examples, the look-up table comprises a denormalized list, the denormalized list indicating partial matches between virtual objects 22, 32 in the scenes 20, 30 as single-liaison entries referencing look-up tables including the two or more predetermined object representations 32.

[0069] In some examples, when a user scans the scene 20, the method may further entail determining if the scene 20 has been previously scanned. If the existence of a matching pre-scanned scene is determined, the already constructed virtual scene thereof may be utilized. The determination is conducted utilizing the pre-scanned scene as a predetermined virtual scene 30, and in response to the physical environment and the pre-scanned scene matching according to a similarity score being above a predetermined threshold, it is determined that the pre-scanned scene and the scanned scene are the same. The similarity score may be based on the similarity of objects found within the scanned scene 20 and the predetermined scene. Thereafter, information associated with the pre-scan may be utilized, such as associated look-up tables according to the above. Furthermore, the method may involve adding further virtual representations 22 after a matching pre-scanned environment has been determined, such as detected objects 12 that were not in the scene 20 during the pre-scan, to the predetermined and pre-scanned scene 30. In this manner, the predetermined list can evolve over time, adding in further entries to the look-up tables as necessary.

[0070] In examples where predetermined object representations 32 of a predetermined virtual scene 30 are stored in a look-up table, the comparing 140 may involve applying one or more filters to limit the total size of a look-up table. Computing relative distances for objects towards a large database of possible spaces may be computationally complex. Therefore, by making an initial estimate with regards to various different types of factors and / or data, at least partial matches between the virtual scene 20 and the predetermined virtual scene 30 may be found at a quicker rate. The computational complexity may thus be reduced, which improves performance of the user localization.

[0071] The filters may include one or more of a geographical filter, IMU filter, object category filter, spatial surroundings filter, virtual scene filter and numeric filter.

[0072] The geographical filter can provide an initial estimate of the user’s 40 position, for example by using other types of tracking technologies such as GPS / GNSS. This may be achieved by arranging a positioning receiver, e.g. a GPS receiver, in the mobile computing device 50. The receiver can alternatively be arranged anywhere in the vicinity of where the user is located, such as in the same house or apartment. By utilizing the geographical filter, the user localization method can disregard one or more entries in the look-up table which does not match the approximate geographical position of the user. The IMU filter can provide an initial estimate of the user’s 40 position, for example by utilizing gyro, accelerometer and / or magnetometer data. This may be achieved by arranging an IMU, e.g. a gyro, accelerometer and / or magnetometer, in the mobile computing device 50 operated by the user 40. Accordingly, a coarse pose estimate of the user 40 can be obtained, such that the user localization method can disregard one or more entries in the look-up table which do not match the approximate user pose determined by the IMU sensor units.

[0073] The object category filter may filter the look-up table based on object types, as determined by the semantic labels. For example, if the user 40 is located in a room which involves a specific type of dining table chairs and ceiling stucco patterns, the look-up table may be limited to entries having only this particular combination of object types, e.g. a ceiling with stucco patterns and the specific types of dining table chairs. Entries of ceilings being non-stucco patterned and entries with other types of dining table chairs can thus be filtered from the look-up table.

[0074] The spatial surroundings filter may filter the look-up table based on the spatial surroundings where the user 40 is located at the time the physical environment 10 is scanned. This may allow for a removal of e.g. rooms not being of interest, such as bathrooms when the user is located in a kitchen having typical spatial surroundings of a kitchen. The spatial surroundings filter may thus involve a room type or location type.

[0075] The virtual scene filter may filter the look-up table based on scenes that have been previously visited by the user 40. For example, if the user 40 has already been in e.g. a house and has scanned all of the rooms therein, the virtual scene filter may be applied such that only the previously scanned rooms are possible matches for the user localization. This can be beneficial to establish a user position in a quick way when, for example, furniture has been introduced, removed, or rearranged in a room, or renovations have taken place.

[0076] The numeric filter may be any type of number that can additionally filter the look-up table based on different factors, including but not limited to a number of walls, a number of furniture items, a number indicative of a size of the room, a number being indicative of when the house was constructed, or the like. The method 100 further involves a step of localizing 150 the user. This occurs in response to an at least partial match between the virtual scene 20 and the predetermined virtual scene 30 occurring, optionally using one or more of the filters as discussed above for providing initial estimates. The at least partial match refers to that the match is not necessarily perfect (perfect would be exactly the same objects at exactly corresponding pose data). In some examples it is sufficient if a match can be established at least to some extent. This may be determined by the calculated similarity metrics as discussed above. Purely by way of example, if pose data of a first virtual object representation differs from pose data of a predetermined object representation with an angle of 1 radian in relation to a horizontal axis of a room and a difference in the z-axis of 20 discrete points, this may be sufficient to establish a match therebetween. If, at the same time, a second virtual object representation differs from pose data of a predetermined object representation with a larger angle and / or larger discrete value, the addition of a second object comparison may provide sufficient confidence in the user localization. The skilled person will appreciate that a plurality of different scenarios like this can appear for any number of objects having any pose data. The notable insight here is that an at least partial match between two or more virtual representations of objects in the physical environment can serve as a sufficient condition for purposes of providing an accurate user localization.

[0077] The localizing 150 may be performed using at least two different approaches, each approach involving different considerations and advantages. In both of the approaches, the user 40 is localized based on pose data of the objects 12 whose virtual representations 22 caused an at least partial match, although the procedure differs slightly.

[0078] The first approach is exemplified in FIG. 3 A. Herein, the user 40 is localized using triangulation techniques based on pose data of each individual object 12 whose virtual representation 22 caused the at least partial match. Due to the match, the relative positions of the objects 12-1, 12-2 are known, depicted as the patterned line therebetween. Moreover, an estimate origin position of the user 40 performing the scanning is also known. Therefore, the localizing 150 in this approach involves a triangulation of the position of the user 40 with respect to the known pose data of the respective objects 12-1, 12-2.

[0079] Another scenario of the first approach is exemplified in FIG. 3B. This is a more advanced, although possibly more realistic scenario. The user 40 is located in a physical environment 10 involving a plurality of objects 12. The user has scanned the room and identified eight objects 12. As seen in the illustration, only four virtual representations of them, namely objects 12-1, 12-2, 12-3, 12-4 have resulted in an at least partial match. The other objects 12-5, 12-6, 12-7, 12-8 yielded no relevant matching during the comparison step. To this end, the localizing of the user 40 is made with respect to the objects 12-1, 12-2, 12-3, 12-4 whose virtual representations caused the at least partial match.

[0080] The second approach is exemplified in FIG. 3C. The difference compared to the first approach is that average pose data of the two or more objects 12 whose virtual representations caused the at least partial match is calculated. This corresponds to steps 152 and 154 of FIG. 2. The user 40 is then localized with respect to the average pose data. This can be done thanks to knowing the user’s 40 estimated origin position, and the calculated average pose data of the two or more objects 12 whose virtual representations caused the at least partial match.

[0081] Similar to FIG. 3B, FIG. 3D shows a more advanced example, but for the second approach. As seen in the illustration, the average pose data has been determined based on the four objects 12-1, 12-2, 12-3, 12-4, and the user 40 is localized based on this virtual center of gravity point in the physical environment 10 which corresponds to the average pose data.

[0082] Either one of the first or second approach may also involve a step of applying an optimization algorithm, which is shown at step 156 of FIG. 2. This shall be understood as a part of the localization procedure. The pose data of the matching objects is known, as is the estimated origin position of the user 40. Depending on whether the first or the second approach is considered, the optimization algorithm may be applied with respect to each one of the pose data of the respective objects 12 (first approach), or with respect to the average pose data (second approach). Any suitable optimization method known in the art may be applied for this purpose, such as a least squares method, gradient method, maximum likelihood estimation, quadratic function, or the like. A gradient method is preferably applied since the average error is not static but a function of current position and orientation of the user 40.

[0083] In some examples, the localizing 150 may be performed in relation to all of the objects 12 in the physical environment whose virtual representations 22 caused the at least partial match in the comparing 140. This may further increase the localization accuracy of the user 40.

[0084] As further seen in FIG. 2, once the user 40 has been successfully localized using the described approaches, the method 100 may further involve a plurality of additional steps. All of these are carried out based on the determined pose of the user 40. Hence, they are enabled thanks to the localization of the user 40. More specifically, these steps include displaying 160 one or more graphical indications 60, providing 161 audio adjustments, providing 162 haptic feedback, providing 163 olfactory feedback, providing 164 vibration alerts, and providing 165 ambient simulations. This will now be described in further detail with reference to FIG. 4.

[0085] In FIG. 4, an exemplary user scenario is shown. In this example, the user 40 has been localized using one or more of the approaches described herein. Consequently, additional types of applications can be enabled for the user 40, some of which are shown in FIG. 4. As seen in the illustration, a plurality of different graphical indicators 60 are shown. The graphical indicators 60 may be shown in a view of the mobile computing device 50, for example projected at a screen of a pair of AR glasses.

[0086] The graphical indicators 60 may comprise cues of spatial interactions, environmental adaptations or user-safety features associated with the objects 12 and / or the physical environment 10. The graphical indicators 60 may alternatively or additionally comprise cues of movement guiding indicators in the physical environment. Spatial interactions may include object interactions, gesture recognitions, or the like, and can relate to gaming, virtual product placement, education / training, remote assistance, interactive art and exhibits, collaborative design, virtual try-ons in retail, enhanced cultural experiences, etc. Environmental adaptations may relate to lighting features, sound, shadows, reflections, perspective for e.g. contextual overlays, translations, accessibility enhancements, environmental simulations, or the like. User- safety features may relate to hazard alerts, emergency evacuation guidance, heads-up displays for navigation, safety training simulations, collision avoidance assistance, personal safety alarms, safety information overlays, health monitoring and alerting, or the like. Movement guiding indicators may be wayfinding arrows, virtual footprints, path highlighting, animated guiding, dynamic route adjustments, interactive maps, proximity warnings, or the like.

[0087] Moreover, different types of feedback, alerts, adjustments, simulations, etc., can be provided in conjunction with or separately from the graphical indicators 60. Feedback may include audio adjustments, for example an increase or decrease of an audio based on user location. Feedback may include haptic feedback for e.g. warning a user upon being located at certain potentially hazardous locations or near potentially hazardous objects. Feedback may include providing certain scents (olfactory feedback) for certain locations to potentially enhance a certain feeling or expression. Alerts may include vibrations when dangerous locations are entered. Ambient simulations can include certain effects associated with ambient conditions, e.g. climate-related simulations for certain locations.

[0088] In FIG. 4 it is noted that the information referred to above can be associated with the user 40 or with the objects 12-1, 12-2, 12-3, 12-4 thanks to the user’s 40 location being established in relation to the objects 12-1, 12-2, 12-3, 12-4. The objects 12-5, 12- 6, 12-7, 12-8 are therefore not associated with any particular information in this example.

[0089] The object 12-1 is associated with four types of information. Firstly, a warning icon is shown at 60-1, indicating that the user 40 should take care when interacting with the object 12-1. Secondly, two types of vibration alerts are shown at 60-2, 60-3. Finally, a guiding path is created at 60-8 such that the user can obtain navigational support towards the object 12-1. The object 12-2 is associated with two types of information. Firstly, a speaker icon 60-4 is presented. Secondly, a touch icon 60-5 is presented. These icons illustrate potential actions that can be carried out with respect to the object 12-2. The object 12-2 may thus be a speaker, and its audio can be adjusted by way of interacting with the icon 60-5. The object 12-3 is associated with one type of information, namely a warning to move in the vicinity of the object 12-3. The user 40 may thus learn that movement within this area can pose a risk of danger. The object 12- 4 is associated with one type of information, namely an indicator that the object 12-4 is a lighting device, such as a lamp, that can be interacted with to adjust ambient surroundings.

[0090] With reference to FIG. 5, a schematic illustration of a (non-transitory) computer- readable (storage) medium 300 is shown according to one exemplary embodiment. The computer-readable medium 300 may be associated with or connected to the system 200 as described herein and is capable of storing a computer program product 310. The computer-readable medium 300 in the disclosed embodiment is a memory stick, such as a Universal Serial Bus (USB) stick. The USB stick 300 comprises a housing 330 having an interface, such as a connector 340, and a memory chip 320. In the disclosed embodiment, the memory chip 320 is a flash memory, i.e., a non-volatile data storage that can be electrically erased and re-programmed. The memory chip 320 stores the computer program product 310 which is programmed with computer program code (instructions) that when loaded into a processor device, will perform a method, for instance the method 100 explained with reference to FIG. 2. The USB stick 300 is arranged to be connected to and read by a reading device for loading the instructions into the processor device. It should be noted that a computer-readable medium can also be other mediums such as compact discs, digital video discs, hard drives or other memory technologies commonly used. The computer program code (instructions) can also be downloaded from the computer-readable medium via a wireless interface to be loaded into the processing device.

[0091] In FIG. 6, an exemplary computerized system 200 is shown. The computerized system 200 may be employed for implementing one or more of the functionalities as previously described in this disclosure. The computerized system 200 may include a number of units known to the skilled person for implementing the functionalities as described in the present disclosure. The computerized system 200 may comprise one or more computing units capable of including firmware, hardware, and / or executing software instructions to implement the functionality described herein. The computerized system 200 may comprise one or more processor devices (may also be referred to as a control unit) 230, one or more memories 235 and one or more buses 240. The processor devices 230 may be included in the computing devices 222a-e and the cloud-based computing resource 212, respectively. The computerized system 200 may include at least one computing device having the processor device 230. A system bus 240 may provide an interface for system components including, but not limited to, the memories 235 and the processor devices 230. The processor device 230 may include any number of hardware components for conducting data or signal processing or for executing computer code stored in the memories. The processor device 230 may, for example, include a general-purpose processor, an application specific processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), a circuit containing processing components, a group of distributed processing components, a group of distributed computers configured for processing, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The processor device 230 may further include computer executable code that controls operation of the programmable device.

[0092] The system bus 240 may be any of several types of bus structures that may further interconnect to a memory bus (with or without a memory controller), a peripheral bus, and / or a local bus using any of a variety of bus architectures. The memories 235 may be one or more devices for storing data and / or computer code for completing or facilitating methods described herein. The memories 235 may include database components, object code components, script components, or other types of information structure for supporting the various activities herein. Any distributed or local memory device may be utilized with the systems and methods of this description. The memories 235 may be communicably connected to the processor device 230 (e.g., via a circuit or any other wired, wireless, or network connection) and may include computer code for executing one or more processes described herein. The memories may include non-volatile memories (e.g., read-only memory (ROM), erasable programmable read-only memories (EPROM), electrically erasable programmable readonly memories (EEPROM), etc.), and volatile memories (e.g., random-access memory (RAM)), or any other medium which can be used to carry or store desired program code in the form of machine-executable instructions or data structures and which can be accessed by a computer or other machine with a processor device. A basic input / output system (BIOS) may be stored in the non-volatile memories and can include the basic routines that help to transfer information between elements within the computer system.

[0093] A storage 245 may be operably connected to the computerized system 200 via, for example, I / O interfaces (e.g., card, device) 250 and I / O ports 255. The storage 245 can include, but is not limited to, devices like a magnetic disk drive, a solid state drive, an optical drive, a flash memory card, a memory stick, etc. The storage 245 may also include a cloud-based server implemented using any commonly known cloudcomputing platform. The storage 245 or memory 235 can store an operating system that controls and allocates resources of the computerized system 200.

[0094] The computerized system 200 may interact with network devices 260 via the I / O interfaces 250, or the I / O ports 255. Through the network devices 260, the computerized system 200 may interact with a network. Through the network, the computerized system 200 may be logically connected to remote computers. Through the network, the serverside platform 210 may communicate with the client-side platform 220, as described above. The networks with which the computerized system 200 may interact include, but are not limited to, a local area network (LAN), a wide area network (WAN), and other networks.

[0095] The operational steps described in any of the exemplary aspects herein are described to provide examples and discussion. The steps may be performed by hardware components, may be embodied in machine-executable instructions to cause a processor to perform the steps, or may be performed by a combination of hardware and software. Although a specific order of method steps may be shown or described, the order of the steps may differ. In addition, two or more steps may be performed concurrently or with partial concurrence.

[0096] The terminology used herein is for the purpose of describing particular aspects only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. It will be further understood that the terms "comprises," "comprising," "includes," and / or "including" when used herein specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0097] It will be understood that, although the terms first, second, etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element without departing from the scope of the present disclosure.

[0098] Relative terms such as "below" or "above" or "upper" or "lower" or "horizontal" or "vertical" may be used herein to describe a relationship of one element to another element as illustrated in the Figures. It will be understood that these terms and those discussed above are intended to encompass different orientations of the device in addition to the orientation depicted in the Figures. It will be understood that when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly connected" or "directly coupled" to another element, there are no intervening elements present.

[0099] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms used herein should be interpreted as having a meaning consistent with their meaning in the context of this specification and the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0100] It is to be understood that the present disclosure is not limited to the aspects described above and illustrated in the drawings; rather, the skilled person will recognize that many changes and modifications may be made within the scope of the present disclosure and appended claims. In the drawings and specification, there have been disclosed aspects for purposes of illustration only and not for purposes of limitation, the scope of the inventive concepts being set forth in the following claims.

Claims

CLAIMS1. A computer-implemented method (100) for localizing a user (40) in a physical environment (10), the method (100) comprising: scanning (110) the physical environment (10) for a plurality of objects (12); assigning (120) a semantic label to each one of the plurality of objects (12) recognized in the physical environment (10); constructing (130) a virtual scene (20) involving virtual representations (22) of the plurality of objects (12) based on the semantic labels; comparing (1 0) pose data of the constructed virtual scene (20) to pose data of at least one predetermined virtual scene (30), each predetermined virtual scene (30) involving a plurality of predetermined object representations (32) being respectively semantically labelled; and in response to the comparing (140) indicating an at least partial match between pose data of two or more of the virtual representations (22) and pose data of two or more of the predetermined object representations (32), localizing (150) the user (40) in relation to two or more objects (12) in the physical environment (10) whose virtual representations (22) caused the at least partial match.

2. The computer-implemented method (100) of claim 1, wherein the localizing (150) comprises: calculating (152) average pose data of said two or more objects (12) in the physical environment (10) whose virtual representations (22) caused the at least partial match; and localizing (154) the user (40) with respect to the average pose data.

3. The computer-implemented method (100) of any preceding claim, wherein the localizing (150) comprises applying (156) an optimization algorithm.

4. The computer-implemented method (100) of any preceding claim, the localizing (150) of the user is performed in relation to all of the objects (12) in thephysical environment (10) whose virtual representations (22) caused the at least partial match in the comparing (140).

5. The computer-implemented method (100) of any preceding claim, wherein the comparing (140) comprises calculating (142) similarity metrics of the pose data of the constructed virtual scene (20) and of the pose data of the at least one predetermined virtual scene (30).

6. The computer-implemented method (100) of any preceding claim, wherein assigning (120) the semantic labels comprises applying (122) a computer vision algorithm to the plurality of objects (12) being scanned.

7. The computer-implemented method (100) of claim 6, wherein the computer vision algorithm ranks the plurality of objects (12) based on object types, each rank indicating a level of pose confidence.

8. The computer-implemented method (100) of claim 7, wherein substantially immovable objects (12) are assigned relatively higher ranks compared to movable objects (12).

9. The computer-implemented method (100) of any preceding claim, wherein the at least partial match is a match of orientation data and position data of the respective pose data.

10. The computer-implemented method (100) of any preceding claim, wherein the at least partial match is a match of an extent of the semantic labels.

11. The computer-implemented method (100) of any preceding claim, wherein said at least one predetermined virtual scene (30) is obtained as: output from a software-based planner tool; oroutput from a mobile computing device (50) having scanned and semantically labelled a physical environment involving a plurality of objects.

12. The computer-implemented method (100) of any preceding claim, wherein said at least one predetermined virtual scene (30) is stored in a look-up table comprising a plurality of entries, each entry involving a combination of two or more predetermined object representations (32).

13. The computer-implemented method (100) of claim 12, wherein the comparing (140) further comprising applying (148) one or more filters to limit the total size of the look-up table.

14. The computer-implemented method (100) of claim 13, wherein the filters comprise one or more of a geographical filter, IMU filter, object category filter, spatial surroundings filter, virtual scene filter and a numeric filter.

15. The computer-implemented method (100) of any preceding claim, wherein based on the determined pose of the user (40), the method (100) further comprising one or more of: displaying (160) one or more graphical indicators (60); providing (161) audio adjustments; providing (162) haptic feedback; providing (163) olfactory feedback; and providing (165) ambient simulations.

16. The computer-implemented method (100) of claim 15, wherein the graphical indicators (60) comprise one or more of: cues of spatial interactions associated with the objects (12) and / or the physical environment (10), cues of environmental adaptations associated with the objects (12) and / or the physical environment,cues of user-safety features associated with the objects (12) and / or the physical environment (10), and cues of movement guiding indicators in the physical environment (10).

17. The computer-implemented method (100) of any preceding claim, wherein the localizing (150) comprises determining a position and / or orientation of the user (40).

18. A computerized system (200) comprising a processor (230) configured to perform the functionality of the method (100) of any one of claims 1-17.

19. A non-transitory computer-readable storage medium (300) comprising instructions, which when executed by one or more processors (230) of a computerized system (200), cause the processor (230) to perform the functionality of the method (100) of any one of claims 1-17.

20. A computer program product comprising computer code for performing the functionality of the method (100) of any one of claims 1-17.

Citation Information

Patent Citations

  • Vehicle position and posture determination method and apparatus, and electronic device

    EP3989170A1

  • CBCT image denoise system and method

    KR1020240172483A

  • Pose determination with semantic segmentation

    US20190080467A1

  • Object-based localization

    US20190392212A1

  • Systems and methods for aiding a visual positioning system with indoor wayfinding

    US20220164985A1