Method for simultaneous location estimation and mapping

JP2025514171A5Pending Publication Date: 2026-04-27VOXELSENSORS SRL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
VOXELSENSORS SRL
Filing Date
2023-04-25
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Standard SLAM methods are inefficient and prone to delays due to the need to scan the entire scene for changes in object position or orientation, especially when objects move quickly, leading to power consumption issues and user discomfort such as motion sickness.

Method used

A low-latency SLAM method that identifies the locations of a second set of points of an object within a shorter time window, converts these points to 3D positions, and determines a transformation to fit these positions to a denser first set of 3D positions acquired in a previous time window, allowing for efficient and accurate mapping and localization.

Benefits of technology

The method achieves low latency and efficient SLAM, enabling smooth object movement perception with minimal delay, preventing motion sickness, and ensuring objects appear realistic and clear within the scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to a sensing method, the method comprising the step of determining a second set of locations (2) of points of at least one object (3) in a scene (4) within a second time window (T2) relative to a first reference point (28), preferably such as a first viewpoint. The method further comprises the step of transforming the second set of locations (2) into a second set of 3D positions (6). The method further comprises the step of determining a transformation for the second set of 3D positions (6) such that the transformed second set of 3D positions (6) best fits the first set of 3D positions (5). The first set of 3D positions (5) is denser than the second set of 3D positions (6), the first set of 3D positions (5) being within a first time window (T1) prior to the second time window (T2), the second time window (T2) being shorter than the first time window (T1), preferably at least half as short.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method for simultaneous localization and mapping (SLAM), and in particular to a low latency method for SLAM. [Background technology]

[0002] In the technical field, various techniques are described for obtaining 3D information of an environment and objects in said environment, for example for the purpose of creating an augmented reality of the environment. Simultaneous Localization and Mapping (SLAM) makes it possible to build a map of the environment and the objects therein, and to localize a sensor or camera within said map.

[0003] A problem with standard SLAM methods is that it takes time to process changes in the position or orientation of an object in a scene. For example, every time an object moves from one point to another, the entire scene needs to be scanned to find the new location and orientation of the object, and therefore all the points that define the object. This means that every time a movement occurs, the entire scene needs to be scanned, i.e., each point of the scene and the points of the object in the scene. This is a disadvantage because it is not an efficient process and consumes a lot of power. It also has the disadvantage that the processing speed is slow, which causes a lot of delays.

[0004] Another drawback is that if an object moves very fast in the scene, standard SLAM methods cannot quickly find the new position and orientation of the object. Instead, there is a delay between the actual movement of the object and the moment the method finds the new position and orientation of the object. This leads to confusion for users when using the method for augmented reality. This further leads to the fact that if the object moves too fast, 1) the user perceives the object with a delay from its actual movement, and 2) it may cause motion sickness, as the object seems to be rapidly changing position in an unrealistic way. Furthermore, the object cannot be displayed smoothly in the scene unless the object's movement is very slow.

[0005] Furthermore, for objects that move very quickly through a scene, scanning the scene for all the points that define said object may not be completed because the object may move before the scan is complete, causing the user to see only a portion of the object, or even worse, to see one object in two different locations in the scene, because the object has moved before the scan of the scene is complete, causing the user to see a distorted, blurry, and even confusing representation.

[0006] Therefore, there is a need for improved SLAM methods and systems that employ said methods that are fast and efficient.

[0007] The present invention aims to solve at least some of the above problems. Summary of the Invention

[0008] It is an object of embodiments of the present invention to provide low latency, high speed, accurate and efficient sensing, such as optical sensing, preferably for Simultaneous Localization and Mapping (SLAM). The above object is achieved by a method, a sensor, an apparatus and a system according to the present invention.

[0009] In a first aspect, the present invention relates to a method for sensing, such as optical sensing, preferably for Simultaneous Localization and Mapping (SLAM), comprising: a) determining a second set of locations of points of at least one object in the scene within a second time window relative to a first reference point, preferably a first viewpoint; b) converting the second set of configurations into a second set of 3D positions; The sensing method includes: c) determining a transformation to the second set of 3D positions such that the transformed second set of 3D positions best fits the first set of 3D positions, the first set of 3D positions being denser than the second set of 3D positions (6); the first set of 3D positions is acquired within a first time window prior to the second time window; The second time window is shorter than the first time window, preferably at least half as short.

[0010] Step (c) is preferably (but not necessarily) performed by calculating a third set of 3D positions by determining a transformation of the second set of 3D positions such that the third set of 3D positions best fit relative to the first set of 3D positions.

[0011] It is an advantage of embodiments of the present invention that a low latency and efficient method for SLAM is obtained. It is an advantage of embodiments of the present invention that smooth motion of objects is obtained. It is an advantage of embodiments of the present invention that a user perceives the motion of objects in a scene with minimal latency. It is an advantage of embodiments of the present invention that motion sickness is prevented. It is an advantage of embodiments of the present invention that a realistic motion of objects is perceived by a user. It is an advantage of embodiments of the present invention that objects in a scene are prevented from appearing blurry, distorted or incomplete to a user.

[0012] An advantage of embodiments of the present invention is that only a partial identification of the current location of points defining an object is required. An advantage of embodiments of the present invention is that only a partial identification of an object is required, allowing low latency and efficient mapping and localization of said object. It is an advantage of embodiments of the present invention that only a partial identification of the object and the location of points defining the object is required, yet an accurate identification is obtained. It is an advantage of embodiments of the present invention that the need to identify all points defining an object is eliminated. It is an advantage of embodiments of the present invention that it is not necessary to repeatedly identify all points of an object in every instance. It is an advantage of embodiments of the present invention that the points in the second set of 3D locations are much fewer than the points in the first set of 3D locations, preferably at least half fewer, more preferably at least 1 / 5 fewer, even more preferably at least 1 / 10 fewer, even more preferably at least 1 / 20 fewer, even more preferably at least 1 / 50 fewer, and most preferably at least 1 / 100 fewer, thus finding the transformed second or third set of 3D locations that are the best fit of the first set of 3D points is a fast, power-efficient and low latency step. The method is further advantageous for reducing power consumption since it is not necessary to identify all points defining the object in the second time window, i.e. the second or third set of 3D positions are sufficient for fitting to the first set of 3D positions.

[0013] It is an advantage of embodiments of the present invention that an unknown current configuration of points defining an object is obtained based on previous configurations of points defining said object. It is an advantage of embodiments of the present invention that it is possible to identify changes to the position of an object, i.e. changes to the positions of points defining said object, that occur within a very short time window (e.g. less than 100 microseconds).

[0014] An advantage of embodiments of the present invention is that a continuous localization estimate of objects in the scene is obtained.An advantage of embodiments of the present invention is that a quasi-continuous mapping of objects in the scene is obtained, where quasi means that the updates of the map and localization estimates are significantly faster than the movements of the objects and / or the observer.

[0015] It is an advantage of embodiments of the present invention that a rotation of an object and the points defining said object are calculated. It is an advantage of embodiments of the present invention that the points defining said object are adjusted based on a change in the object's orientation.

[0016] It is an advantage of embodiments of the present invention that both the translation and the rotation of an object between the first and second time windows are obtained.

[0017] An advantage of embodiments of the present invention is that it is possible to map a user's environment and identify the user's posture or head pose in the digital environment using sparse and dense data sets with low latency, which is particularly advantageous in head mounted display applications.

[0018] Preferred embodiments of the first aspect of the invention include one or suitable combinations of two or more of the following features.

[0019] The method preferably comprises the steps of: - determining a location of a first set of points of the at least one object in the scene within the first time window with respect to the first reference point; - transforming said first set of positions to said first set of three-dimensional positions; It further has:

[0020] The method preferably comprises the steps of: - supplementing said second set of three-dimensional locations to said first set of three-dimensional locations.

[0021] It is an advantage of embodiments of the present invention that the first set of 3D positions can be relied upon to complement the second set of 3D positions.It is an advantage of embodiments of the present invention that the second set of 3D positions can be built on the basis of the first set of 3D positions.

[0022] The method preferably comprises the steps of: - illuminating dots on said scene along an illumination trace, preferably in a Lissajous manner; - monitoring the dot; - determining said first and second sets of locations along said trace and relative to said first reference point; - identifying a third set of configurations along said trace, within said first time window, and relative to a second reference point, which is preferably a second viewpoint; - identifying a fourth set of locations along the trace, within the second time window, and relative to the second reference point; - triangulating said first and third sets of configurations to obtain 3D positions of said first set; - triangulating the second and fourth sets of configurations to obtain 3D positions of the second set; The present invention further comprises:

[0023] An advantage of an embodiment of the present invention is that the points of the second set of locations (or the fourth set of locations) along the illumination trace are much fewer than the points of the first set of locations (or the third set of locations) along the illumination trace. An advantage of an embodiment of the present invention is that only points along the illumination trace are identified, allowing low-latency, fast and efficient mapping and localization of the object in an accurate manner. An advantage of an embodiment of the present invention is that the illumination trace covers, for example, at most 20%, preferably at most 10%, more preferably at most 5%, even more preferably at most 2%, and most preferably at most 1% of the scene during the second time window.

[0024] An advantage of embodiments of the present invention is that it is not necessary to illuminate the entire scene and identify all points of objects in the scene. An advantage of embodiments of the present invention is that it is possible to obtain information about all objects in the scene by only partially monitoring the scene. An advantage of embodiments of the present invention is that different areas of the scene are monitored equally or approximately equally. An advantage of embodiments of the present invention is that different areas of the scene have an equal or approximately equal chance of being scanned by the dots, in particular in the Lissajous method.

[0025] This has the advantage that only one method (in this case only triangulation) needs to be used to obtain the first and second sets of 3D positions, i.e. a simpler, faster and more efficient method is obtained compared to other methods that use two different methods to obtain dense and sparse 3D data sets.

[0026] The method preferably comprises the steps of: - defining a first set of lines by successive points of said first set of 3D positions; - defining a second set of lines by successive points of said second set of 3D positions; - determining a transformation for the second set of lines such that the second set of lines best fits the first set of lines; It further has:

[0027] It is an advantage of embodiments of the present invention that an accurate fitting of the second set of 3D positions to the first set of 3D positions is obtained.

[0028] The method preferably comprises the steps of: - dividing said second time window into sub-windows in each of which a subset of 3D positions are identified; - calculating for each of said subset of 3D positions a corresponding transformed set of 3D positions; The present invention further comprises: Each of the 3D positions in the transformed set is obtained by transforming the 3D positions of the second set within the sub-window such that the 3D positions of the transformed set each best fit the 3D positions of the first set.

[0029] An advantage of embodiments of the present invention is that it is still possible to identify changes in the location of points defining an object or rotation of said object that occur within a very short time window.An advantage of embodiments of the present invention is that it is still possible to monitor objects that move or rotate very fast within the scene.An advantage of embodiments of the present invention is that the method allows the flexibility to consider the entire second time window or to consider one or more sub-windows depending on the need, depending on the movement or rotation of the object within the scene.

[0030] This method step preferably comprises the steps of: - extrapolating a prediction set of 3D positions for a next time instance based on the second set of 3D positions and the first set of 3D positions by identifying a trajectory of the object between the first time window and the second time window; - determining the 3D position of the real set; - validating said predicted set by at least partially comparing it with said actual set; The present invention further comprises:

[0031] An advantage of embodiments of the present invention is that a set of 3D positions of objects for subsequent detection is predicted. An advantage of embodiments of the present invention is that the movement of objects in a scene is predicted before it actually occurs. An advantage of embodiments of the present invention is that high speed and low latency mapping is obtained while obtaining high accuracy mapping. An advantage of embodiments of the present invention is that inaccurate identification of points (e.g. false positives) are filtered.

[0032] The method preferably further comprises the step of combining said first and second time windows and their points.An advantage of embodiments of the present invention is that denser data is obtained, consisting of a larger number of points defining said object within a longer time window.

[0033] In a second aspect, the present invention relates to an optical sensor for simultaneous localization and mapping (SLAM), comprising: a first plurality of pixel sensors, each comprising a photodetector; a processing unit, . the first plurality of photodetectors are adapted to output a second set of locations of at least one object point in the scene within a second time window relative to the first plurality of photodetectors; the processing unit is adapted to transform the second set of configurations into a second set of 3D positions; the processing unit having a processor that, in use, determines a transformation to the second set of 3D positions such that the second set of 3D positions best fits the first set of 3D positions; the first set of 3D locations is denser than the second set of 3D locations; the first set of 3D positions are within a first time window prior to the second time window; the second time window is shorter than the first time window, preferably at least half as short, Preferably, the first set of 3D positions are obtained from a storage location, preferably a local or cloud-based storage location.

[0034] A preferred embodiment of the second aspect of the invention comprises - the first plurality of photodetectors are further adapted to output the first set of locations of points of the at least one object in the scene within the first time window, and the processing unit is further adapted to convert the first set of locations into the first set of 3D locations; the sensor further comprises a primary optical system capable of generating an image of the scene onto the first plurality of pixel sensors.

[0035] In a third aspect, the present invention relates to an optical sensing device for simultaneous localization and mapping (SLAM), comprising the optical sensor according to the second aspect, further comprising at least one light source adapted to illuminate at least one dot on the scene along an illumination trace, preferably in a Lissajous manner, the optical sensing device further comprising a second plurality of pixel sensors, each pixel sensor comprising a photodetector; the first plurality of photodetectors are adapted to monitor the dots and output the first and second sets of locations along the illumination trace; the second plurality of photodetectors are adapted to monitor the dots relative to the second plurality of photodetectors and output a third set of location of points of the object along the illumination trace within the first time window and a fourth set of location of points of the object along the trace within the second time window; The processing unit is capable of triangulating the first and third sets of configurations to obtain the first set of 3D positions, and the processing unit is capable of triangulating the second and fourth sets of configurations to obtain the second set of 3D positions.

[0036] Preferred embodiments of the third aspect of the invention include one or suitable combinations of two or more of the following features. the optical sensing device further comprises at least one memory element, the memory element being adapted to store at least the first set of 3D positions and / or the second set of 3D positions. - the processing unit is capable of extrapolating a predicted set of 3D positions for a next instance based on the second set of 3D positions and the first set of 3D positions, the processing unit is capable of identifying a trajectory of the object between the first time window and the second time window, and the processing unit is further capable of validating the predicted set by comparing it with an actual set of 3D positions.

[0037] In a fourth aspect, the present invention relates to an optical sensing system for simultaneous localization and mapping (SLAM), comprising the optical sensing device according to the third aspect, further comprising a secondary optical system capable of generating an image of the scene on the second plurality of pixel sensors.

[0038] These and other characteristics, features and advantages of the present invention will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, which illustrate, by way of example, the principles of the invention, the description being given for the purposes of example only, without limiting the scope of the invention. [Brief description of the drawings]

[0039] The present disclosure is further explained by the following description and the accompanying drawings. [Figure 1] FIG. 1 shows in (a,b) a first set of arrangements (1) and a second set of arrangements (2) of points of an object (3) in 2D, and in (c,d) a first set of 3D positions (5) and a second set of 3D positions (6) in 3D, according to an embodiment of the invention. [Diagram 2]FIG. 2 relates to an embodiment of the invention and shows in (a, b) a first set arrangement (1) and a second set arrangement (2) of points of an object (3) in 2D along an illumination trace (9) and with reference to a first reference point (28), in (c, d) a third set arrangement (11) and a fourth set arrangement (12) of points of an object (3) in 2D with reference to a second reference point (29), and in (e, f) a comparison of the first and third sets of arrangements (1, 11). ) and a first set of 3D positions (5) and a second set of 3D positions (6) in 3D resulting from triangulating the second and third sets of configurations (2, 12), respectively, and in (g,h) showing matching or fitting a third set of lines (35), which are a transformation of the second set of lines (34), to the first set of lines (33) such that the third set of lines (35) is the best fit to the first set of lines (33). [Diagram 3] FIG. 3 shows an object (3) with different orientations (a, b, c) according to an embodiment of the invention. [Figure 4] FIG. 4, for an embodiment of the invention, shows points of an object (3) identified over four time windows of different widths (14, 15, 16, 17), with a dense set of data (20) obtained for the longest time window (17). [Diagram 5] FIG. 5 relates to an embodiment of the invention and shows an object (3) moving along a trajectory (19) between the first and second time windows (T1, T2), and a set of predictive 3D positions (18) are obtained. [Figure 6] FIG. 6 relates to an embodiment of the invention and shows in (a) identifying points along a Lissajous trajectory (10) that are used to define a second set of contour lines (30) in (b) and in (c, e) identifying points along a denser Lissajous trajectory (10) that are used to define a first set of contour lines (31) and a fourth set of contour lines (32) in (d, f). [Figure 7] FIG. 7 shows an optical sensor (21) according to an embodiment of the present invention. [Figure 8]8 shows an optical sensing device (26) comprising a first and a second plurality of pixel sensors (28, 29) and a light source (27) for illuminating dots (8) on a scene (4) in a Lissajous pattern (10) according to an embodiment of the present invention. Reference signs in the claims should not be construed as limiting the scope thereof. In different drawings, the same reference signs refer to the same or similar elements. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0040] The present invention relates to a sensing method and an optical sensor for sensing, e.g. optical sensing, preferably for Simultaneous Localization and Mapping (SLAM), with low latency.

[0041] The present invention will be described with respect to particular embodiments and with reference to certain drawings but the invention is not limited thereto but only by the claims. The drawings described are schematic and non-limiting. In the drawings, the size of some of the elements may be exaggerated and not drawn to scale for illustrative purposes. The dimensions and relative dimensions do not correspond to actual reductions to the practice of the invention.

[0042] The terms first, second, etc. in this specification and claims are used to distinguish between similar elements and are not necessarily used to describe an order in time, space, sequence, or otherwise. The terms so used are interchangeable under appropriate circumstances, with it being understood that the embodiments of the invention described herein are capable of operating in sequences other than those described or illustrated herein.

[0043] Moreover, terms such as top, under, and the like are used in this specification and claims for descriptive purposes and not necessarily to describe relative positions, and it is to be understood that the terms so used are interchangeable under appropriate circumstances, and that the embodiments of the invention described herein are capable of operation in orientations other than those described or illustrated herein.

[0044] In this specification, numerous specific details are described. However, it will be understood that embodiments of the present invention may be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this specification.

[0045] References throughout this specification to "one embodiment" or "one embodiment" mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. Thus, the appearances of the phrases "in one embodiment" or "in one embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment, but may. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments, as would be apparent to one of ordinary skill in the art from this disclosure.

[0046] Similarly, in describing exemplary embodiments of the invention, it should be understood that various features of the invention may be grouped together in a single embodiment, figure, or description for the purpose of streamlining the disclosure and aiding in understanding one or more various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single previously disclosed embodiment. Thus, the claims following the detailed description are expressly incorporated into this specification, with each claim standing on its own as a separate embodiment of the invention.

[0047] Furthermore, although some embodiments described herein include some features and not others included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the present invention and form different embodiments, as would be understood by one of ordinary skill in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0048] Unless otherwise defined, all terms used in disclosing the present invention, including technical and scientific terms, have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. As a further guide, definitions of terms are included to better understand the teachings of the present invention.

[0049] As used herein, the following terms have the following meanings: "A," "an," and "the" as used herein refer to both the singular and the plural, unless the context clearly indicates otherwise. By way of example, "a contaminant" refers to one or more contaminants.

[0050] The recitation of numerical ranges by endpoints includes not only the recited endpoints but also all values ​​and fractions subsumed within that range.

[0051] In a first aspect, the present invention relates to a sensing, eg optical sensing, preferably a Simultaneous Localization and Mapping (SLAM) method.

[0052] The method includes (a) identifying a location of a second set of points of at least one object in the scene within a second time window, said set including, for example, points that define the object, e.g., points that indicate that the object is present in a portion of the scene, e.g., points that define a contour of the object, such as identifying points.

[0053] The method further includes (b) transforming the second set of configurations into a second set of 3D positions, e.g., the second set of 3D positions having depth information, e.g., xyz information, for each point defining the at least one object.

[0054] The method further includes (c) determining a transformation to the second set of 3D positions such that the transformed second set of 3D positions best fit the first set of 3D positions. The first set of 3D positions is denser than the second set of 3D positions, i.e., the first set of 3D positions includes more points than the second set of 3D positions. The first set of 3D positions are acquired within a first time window prior to the second time window.

[0055] Step (c) is preferably performed by calculating a third set of 3D positions by determining a transformation to the second set of 3D positions such that the third set of 3D positions best fit the first set of 3D positions. For example, fitting or matching the third set of 3D positions or the second set of 3D positions to the first set of 3D positions. The third set of 3D positions are a transformed version of the second set of 3D positions. This is not a point to point mapping, but a matching step between the overall second or third set of 3D positions and the overall first set of 3D positions, e.g. matching a group of points to another group of points. For example allowing similar points of the objects to be matched to each other as a group. The matching or fitting can be performed with an accuracy of at least 80%, preferably at least 90%, more preferably at least 95%, most preferably at least 98%.

[0056] Matching or fitting in the present invention is preferably performed by matching or fitting a sparse set of data points to a dense set of data points, because it is cheaper, easier and faster to convert and fit a sparse set of points to a dense set of points than to convert and fit a dense set of points to a sparse set of points, also because a dense set of data often represents the whole world (e.g. the whole digital world), whereas a sparse set of data represents only a part of said world (e.g. where the user is looking). However, in principle both approaches are possible.

[0057] Such high accuracy of matching is possible because objects do not move very quickly between the first and second time windows. For example, the time between the end of the first time window and the beginning of the second time window is typically 0 milliseconds where the windows overlap, but the windows can overlap in time, e.g., an overlap time of 100 microseconds to 1 millisecond, or there can be a gap between the windows, e.g., between 100 microseconds and 1 millisecond.

[0058] The first set of 3D positions are preferably obtained from a storage location, such as a local or cloud-based storage location.

[0059] The second time window is shorter than the first time window, preferably at least half or one third shorter. The ratio of the second time window to the first time window may be between 1:2 and 1:100, preferably between 1:5 and 1:50, more preferably between 1:10 and 1:20. Those skilled in the art will understand that the ratio of the second time window to the first time window depends on the moving speed of the object depending on the application. For example, more points are identified in the first window than in the second window. For example, the sampling rate of both the first and second windows is the same, but the length of the first time window is longer, so that more points are identified in the first time window than in the second time window.

[0060] The width of the second time window is long enough to identify a small fraction of the number of points for the dense representation obtained from the first time window. The number of points in the second time window is at most equal to the number of points in the first time window, preferably at most 50%, more preferably at most 10%, and most preferably at most 5%. For example, the second time window may be at most 5 milliseconds, preferably 1 millisecond, and more preferably at most 100 microseconds.

[0061] The second set of locations is preferably identified with respect to a first reference point, such as a first viewpoint. The reference point may for example be the position of a system performing the method or for example the viewpoint of a sensor in the system. This allows identifying objects in the scene with respect to the reference point, thereby allowing all objects in the scene to be mapped. Thus, for example, a map of objects in the scene, preferably all objects in the scene, can be constructed and further a system performing the method in the scene can be localized. The at least one object is preferably a plurality of objects.

[0062] The method preferably further comprises the steps of identifying a first set of configurations of points of the at least one object in the scene within the first time window relative to the first reference point, and transforming the first set of configurations into the first set of 3D positions. The first and second sets of configurations are 2D (i.e. two-dimensional), e.g. in xy axes, e.g. each point having an x ​​coordinate and a y coordinate. The first set of configurations is a dense collection of data and the second set of configurations is a sparse collection of data.

[0063] As mentioned above, step (c) is preferably performed by determining a transformation of the second set of 3D positions to obtain the third set of 3D positions. For example, a translation (i.e. a movement in xyz) and / or a rotation (i.e. a change in orientation) of the second set of 3D positions corresponding to a translation and / or rotation of the at least one object is determined to obtain the third set of 3D positions, thereby providing a best fit to the first set of 3D positions, for example to obtain information about the movement and / or rotation of objects in the scene between the time window and the second time window, for example to determine a change in orientation of the at least one object with respect to a reference orientation or a reference axis system (e.g. an xyz axis system), for example if the orientation of the object is changed about at least one axis.

[0064] The translation and / or rotation determination is preferably done in stages. For example, an initial translation and / or rotation estimate is determined and then the transformed second or third set of 3D positions are fitted or matched to the first set of 3D positions. However, if the initial translation and / or rotation is not successful, the matching or fitting process continues by correcting or improving the translation and / or rotation estimate. This is repeated until the match is successful, or until a best match is obtained, or until a percentage error (e.g. a minimum error) is obtained. For example, a percentage error in matching the transformed second or third set of 3D positions to the first set of 3D positions may be tolerated.

[0065] Calculating the change in orientation, e.g., rotation, of an object is important because at least some points defining the object may not be identifiable, e.g., may disappear, due to the rotation and may be identifiable during the first time window but not during the second time window, while at least some other points defining the object may appear due to the rotation and may be identifiable, e.g., during the second time window but not during the first time window. Calculating the orientation thus facilitates matching in an accurate manner, e.g., by adjusting points in the first and / or second set of 3D positions based on the change in orientation to adjust points defining the object.

[0066] The method and transformation are preferably performed on rigid bodies, i.e. bodies that do not deform, such as boxes or cylinders, which may move and / or rotate in the scene but do not usually deform, i.e. do not change shape. However, the invention also works for non-rigid bodies, e.g. by identifying a new first set of configurations and a new first set of 3D positions more frequently than for rigid bodies. For example, in the case of an inflating balloon, the transformation allows for non-rigid deformations as well as translational and / or rotational changes of the object.

[0067] After determining the transformation, we know where the object has moved and how it has rotated. Thus, an object can be defined primarily based on the transformed second or third set of 3D positions, without the need to identify all points of the object during the second time window. This allows for fast, smooth, and low-latency mapping and localization of objects in a scene. In other words, using a sparse data set, e.g., the second set of 3D positions, the method can be used to map the user's environment and locate the user in the environment with very low latency. The sparse data set can also be used to locate the user's head pose in the environment, e.g., a digital environment, with very low latency, e.g., where the user is looking.

[0068] It should be noted that the sparse data set may represent a scene (e.g. in scene coordinates) and the dense data set may represent the world (e.g. in world coordinates). Considering for example a room or a given space, the dense data includes data representing the room, e.g. all objects and e.g. walls of the room, defined relative to a reference point in world coordinates. However, the sparse data set includes data representing a scene or part of the room, e.g. the scene that the user of the method is looking at, in scene or sensor coordinates. This is particularly advantageous for the following reasons: First, the sparse data set is transformed into a dense data set and fitted. For example, the method tries to find the best fit in world coordinates of the scene, e.g. where the scene is located in the room. Secondly, finding the best fit allows finding the position of the user of the method in world coordinates and the user's head pose, i.e. where the user is looking.

[0069] Preferably, the method further comprises the step of complementing the second set of 3D positions with the first set of 3D positions. Based on the accuracy of the best fit in step (c), the complementing step may also be accurate. This is advantageous because it avoids the need to identify missing points in the second set of 3D positions during the second time window. In other words, the complementation allows the second set of 3D positions to be constructed based on the first set of 3D positions, enabling low latency and efficient SLAM.

[0070] Preferably, the ratio of the number of points in said second set of 3D locations to the number of points in said first set of 3D locations is between 1:2 and 1:50, preferably between 1:10 and 1:20.

[0071] Preferably, the method is performed continuously, which has the advantage that it allows for continuous localization and mapping of the scene and objects therein.

[0072] Preferably, the method steps may further comprise illuminating dots on the scene along an illumination trace. For example, the dots are illuminated by a light source, for example a dot light source. For example, a mechanism (for example a MEMS mirror) moves the dots to illuminate different parts of the scene. The illumination trace is preferably continuous, but does not necessarily have to be, for example it can be switched on and off.

[0073] The method steps may further include monitoring the dots, for example by a sensor viewing the scene.

[0074] The method steps may further include identifying the first and second sets of locations along the illumination trace with respect to the first reference point. The method steps may further include identifying a third set of locations along the illumination trace with respect to the second reference point within the first time window. The second reference point is preferably a second viewpoint. The method steps may further include identifying a fourth set of locations along the trace with respect to the second reference point within the second time window. The third set of locations differs from the first set of locations in that they are identified from different reference points, e.g. the second and first reference points. For example, both the first and third sets of locations are identified within the first time window and have much more points than the second and fourth sets of locations because the first time window is longer than the second time window. The fourth set of locations and the second set of locations also differ in that they are identified from different reference points, e.g. the second and first reference points.

[0075] The method steps may further include triangulating the first and third sets of configurations to obtain 3D positions of the first set. The method steps may further include triangulating the second and fourth sets of configurations to obtain 3D positions of the second set. The first and third sets of configurations may be triangulated to obtain 3D positions of the first set. For example, by monitoring the dots and identifying points along the illumination trace, it is possible to associate points of one set of configurations at one reference point with corresponding points of another set of configurations at another reference point. This allows depth information to be obtained for objects in the scene.

[0076] The illumination trace is also advantageous in that only the points along the illumination trace are identified within the second time window, which allows for low-latency mapping. For example, not all points in the scene are identified. The illumination trace also covers only a portion of the scene during the second time window, for example covering only at most 20% or at most 10% of the scene, preferably at most 5%, more preferably at most 2%, most preferably at most 1%. Similarly, in future time windows, it is only necessary to identify points along the illumination trace, and only occasionally to identify points along a longer time window, such as the first time window.

[0077] Preferably, future time windows, e.g., a third time window, a fourth time window, a fifth time window, and the set of configurations identified over these time windows can be combined to obtain a new set of configurations of a number of points that define the object within a longer time window, which may then be converted into 3D positions and used to map the 3D positions in future time windows.

[0078] Preferably, the first set of 3D positions may be updated periodically. For example, the first set of 3D positions may be updated after a predefined time period, for example by identifying a new set of points within a third time window adapted to identify many points of the object, such as the first time window, long enough to identify a new dense set of 3D points. This allows matching or fitting a new sparse set of 3D points to the new dense set of 3D positions. In other words, every predefined time, a new dense data set may be obtained and used to match or fit the new sparse data set.

[0079] Preferably, the dots may be sparsely illuminated on the scene. For example, only a certain part of the scene is illuminated and the points are identified along the illumination trace caused by the dots within a given time window. For example, different parts of the scene are illuminated, for example such that points of an object in any part of the scene can be identified. However, the identification of the points is performed, for example in a coarse manner such that only a partial identification of the points of the object in the scene is performed. For example, if an object is defined by 100 points, only a part of these points are identified, for example at most 5 points or at most 10 points or at most 20 points. Identifying these points allows the overall shape or contour of the object to be roughly identified. In the previous time window, a denser scan is performed in which more points are identified, for example for an object with 100 points, at least 80 points are identified, for example at least 90 points, for example at least 95 points, for example at least 98 points. By first converting the points from 2D to 3D and mapping the former point cloud (e.g., including 5 or 10 points) defining the object to the latter point cloud (e.g., including 95 or 98 points), it is possible to obtain complementary points based on the latter point cloud (i.e., dense group) to complement the former point cloud (i.e., sparse group). Note that the complementary points are based on the latter point cloud (i.e., dense group), but are not necessarily the exact points taken from the latter point cloud (i.e., sparse group). For example, when an object rotates, some complementary points may need to be adjusted.

[0080] Preferably, the dots may be illuminated in a Lissajous fashion, since this allows for a coarse identification of points that define an object without identifying all the points of the object each time. The Lissajous fashion is an example, but other sparse illuminations such as raster scanning can achieve similar results. This is advantageous in obtaining a low-latency and efficient method for SLAM. For example, an object on the right side of a scene and the same object on the left side of a scene may have a similar number of identified points due to similar scanning patterns of the dots in different parts of the scene. Further advantages of the Lissajous fashion are shown below and in the figures.

[0081] Preferably, the method may further comprise the step of defining a first set of lines by successive points of the first set of 3D locations. The method may further comprise the step of defining a second set of lines by successive points of the second set of 3D locations. For example, the successive points may be successive points on a Lissajous trajectory, e.g., each line crossing the object once. This is also shown in Fig. 2(g,h). It is clear that, since the lines are three-dimensional lines, the lines may or may not be straight.

[0082] The method may further include determining a transformation for the second set of lines such that the second set of lines best fits the first set of lines. Finding a best fit or match between two sets of lines may be more accurate and / or easier than finding a best fit or match between two sets of points. This may be done by calculating a third set of lines, the third set of lines being transformed versions of the second set of lines.

[0083] This step may therefore include determining a transformation to the second set of lines to obtain the third set of lines such that the third set of lines best fit to the first set of lines, for example by determining a translation (i.e. movement in xyz) and / or rotation (i.e. change in orientation) of the second set of lines corresponding to a translation and / or rotation of the at least one object to obtain the third set of lines, thus obtaining a best fitting or matching to the first set of lines.

[0084] Preferably, the method may further comprise the step of dividing the second time window into sub-windows, each of the sub-windows comprising a subset of 3D positions. For example, the subset of 3D positions correspond to the arrangement of the subsets identified in each of the sub-windows. The method may further comprise the step of calculating, for each of the subset of 3D positions, a corresponding transformed set of 3D positions. Each of the transformed set of 3D positions is obtained by transforming the second set of 3D positions within the sub-window such that each transformed set of 3D positions best fits the first set of 3D positions.

[0085] By dividing the second time window into sub-windows, for example, information about an object moving or rotating very fast in a scene can be obtained by obtaining information about the object in each of the sub-windows. For example, each of the subsets of 3D positions in each of the sub-windows may be fitted to the first set of 3D positions. Depending on the speed of the object's movement or rotation, 3D positions in one or more of the sub-windows may be considered for the fitting.

[0086] Preferably, the number of points in the 3D positions of the subset within the very short sub-window may need to be larger, for example by identifying more points of the object within said sub-window, so that it is possible to fit the 3D positions of the subset within the very short sub-window to the first set of 3D positions, as long as the 3D positions of the subset have a sufficient number of points for such fitting.

[0087] Preferably, the method may further comprise combining the second set of 3D locations with a fourth set of 3D locations (e.g. the fourth set of 3D locations are within a third time window), for example if the second set of 3D locations does not have enough points to fit to a dense set of 3D locations, e.g. the first set of 3D locations. Alternatively, the method may further comprise combining the first and second time windows and their points, or combining any other time windows with each other.

[0088] Preferably, the method may further comprise storing a dense set of data, e.g. the first set of 3D positions and / or the first set of configurations, in at least one memory element, e.g. for use in fitting or matching thereto. For example, the dense set of data is an aggregation of a set of configurations in a series of time windows, the series of time windows being numbered such that the dense set of data comprises, e.g., at least 50%, preferably at least 80%, more preferably at least 90%, most preferably at least 95% of the points of at least one object in the scene. The step may also comprise storing the sparse set of data, or other data. Such storage may be done locally or in a cloud-based storage location.

[0089] Preferably, the method may further comprise a step of updating the dense set of data, e.g., in a time window much later than the first time window, a new dense set of data may obtain a new first set of configurations, e.g., especially if the object has moved and / or rotated significantly. The same steps for obtaining a new set of configurations are performed, after which the dense set of data is updated, and then further fitting or matching and triangulation is performed based on the new first set of configurations or a new set of 3D positions.

[0090] Preferably, the method steps include the steps of: defining a first set of contour lines by the outermost points of the first set of 3D locations; and defining a second set of contour lines by the outermost points of the second set of 3D locations. The contour lines are defined, for example, by joining the outermost points of the first or second set of 3D locations to each other to form the contour of the object. The second set of contour lines are then transformed into the third set of contour lines by determining a transformation such that the third set of contour lines or the transformed second set of contour lines match or fit the first set of contour lines. Contour lines are defined using a larger set of 3D locations, which may be more advantageous and accurate for fitting than point clouds. Therefore, the method steps preferably include the step of fitting the second set of contour lines to the first set of contour lines.

[0091] The use of Lissajous trajectories as described above is advantageous to define the second set of configurations in a better way, because after a few or several Lissajous trajectories many points of the second set of configurations are identified, based on which the 3D positions of the second set are defined, and then the contour of the second set is defined, the more Lissajous trajectories and the longer the identification window of the second set of configurations is the better defined. This embodiment is better shown in Fig. 6. Alternatively, the contour of the first and second sets may be defined in 2D by the first and second sets of configurations, and then the contour of the first and second sets may be transformed into 3D.

[0092] Contour lines in this embodiment refer to lines that define the outermost points of an object, e.g. the contour of an object, as shown in Figure 6. This allows for segmentation of objects, e.g. the location of each object can be known by drawing a set of contour lines around said object.

[0093] Preferably, the method steps may further comprise extrapolating a 3D position of a prediction set for a next time instance or time window based on the second set of 3D positions and the first set of 3D positions. The extrapolation comprises determining a trajectory of the object between the first and second time windows. For example, if an object trajectory is moving rightward between the first and second time windows, the object is expected to move further rightward in the next time instance. Furthermore, a moving speed of the object can be estimated between the first and second time windows, which is used to better extrapolate the 3D position of the prediction set for the next time instance or time window.

[0094] This is advantageous for predicting the movement of the object beyond the first and second time windows before it actually occurs. The method steps may further comprise a step of identifying a real set of 3D positions. The real set of 3D positions may consist of only a few points of the object, for example the real set of 3D positions correspond to a partial identification of the points of the object. The method steps may further comprise a step of validating the predicted set by at least partial comparison with the real set. The real set of 3D positions are used only for validation purposes, for example to verify whether the object has changed its trajectory or its movement speed.

[0095] This has the added advantage of filtering out false positives or inaccurate point identifications: for example, if an object is observed to have moved in a certain direction at a certain speed, then any point identification away from the predicted set of 3D positions is likely to be a false positive unless it relates to another object in the scene.

[0096] In a second aspect, the present invention relates to an optical sensor for SLAM. The system comprises a first plurality of pixel sensors, each pixel sensor comprising a photodetector. The first plurality of photodetectors are adapted to output a second set of locations of at least one object point in a scene within a second time window relative to the first plurality of photodetectors, e.g., a middle point of the first plurality is reference.

[0097] The optical sensor further comprises a processing unit adapted to convert the second set of configurations into a second set of 3D positions, which are then stored in a random access memory of the processing unit and retrieved whenever needed.

[0098] The processing unit has a processor that determines (e.g. is capable of determining) during use a transformation to the second set of 3D positions such that the second set of 3D positions best fit the first set of 3D positions. For example, determining a translation and / or rotation of the second set of 3D positions relative to the first plurality of pixel sensors, which corresponds to a translation and / or rotation of an object in the scene. This is preferably done continuously. As shown in the method of the first aspect, the transformation is advantageous in that it allows to obtain a low-latency and high-efficiency method and thus a low-latency and high-efficiency optical sensor. Thus, the optical sensor can map objects in the scene and localize the first plurality of pixel sensors relative to the objects. The optical sensor may include a set of instructions that, when executed by a processor, performs the method of the first aspect.

[0099] The first set of 3D positions is denser than the second set of 3D positions. The first set of 3D positions is within a first time window that precedes the second time window. The second time window is shorter than the first time window, preferably at least half as short. The first set of 3D positions are preferably obtained from a storage location, such as a local or cloud-based storage location.

[0100] Preferably, the first plurality of photodetectors are further adapted to output a first set of configurations of points of the at least one object in the scene within the first time window, and the processing unit is further adapted to convert the first set of configurations into a first set of 3D positions.

[0101] Preferably, the processing unit is capable of complementing said second set of 3D positions with said first set of 3D positions, thereby constructing said second set of 3D positions, after which the processing unit can be connected to a means for generating an image of said scene.

[0102] Preferably, the pixel sensors may be arranged in multiple rows and / or multiple columns. For example, the pixel sensors may be arranged in a matrix. Such an arrangement allows detection of a wider area. Preferably, each photodetector may be a single photon detector, for example a single photon avalanche diode (SPAD).

[0103] Preferably, each photodetector may be arranged in a reverse bias configuration. The photodetector may detect photons incident thereon. The photodetector has an output for outputting an electrical detection signal upon detection of a photon. For example, the detection signal may be represented by a signal constituted by a logic "1" (e.g., detection), and no detection signal may be represented by a signal constituted by a logic "0" (e.g., no detection). Alternatively, the detection signal may be represented by or as a result of a pulse signal, for example, a transition from logic "0" to logic "1" and then back from logic "1" to logic "0".

[0104] Preferably, there are at least twice as many points in the first set of 3D positions as there are points in the second set of 3D positions, preferably at least 5 times as many, more preferably at least 10 times as many, even more preferably at least 20 times as many, and most preferably at least 50 times as many, because the first time window is longer than the second time window, e.g. because more points of the object have been identified in the first time window.

[0105] Preferably, the optical sensor further comprises primary optics capable of generating an image of said scene onto said first plurality of pixel sensors.

[0106] In a third aspect, the present invention relates to an optical sensing device for SLAM, the optical sensing device comprising the optical sensor according to the second aspect.

[0107] The optical sensing device further comprises at least one light source, the light source adapted to illuminate at least one dot on the scene along an illumination trace. The optical sensing device further comprises a second plurality of pixel sensors, each pixel sensor comprising a photodetector. The number of the second plurality is preferably the same as the number of the first plurality. For example, the first and second plurality of pixel sensors are preferably the same number of pixel sensors of the same configuration. However, the first and second plurality are not located at the same position. For example, the first plurality is shifted along the x-axis from the second plurality. This allows triangulation of the outputs of the first and second plurality of pixel sensors to obtain depth information of the scene, as described below. Preferably, the light source is adapted to a wavelength detectable by the pixel sensors, for example between 100 nanometers and 10 micrometers, preferably between 100 nanometers and 1 micrometer.

[0108] A first plurality of the photodetectors are adapted to monitor the dots, the photodetectors being further adapted to identify and output the first and second set of locations along the illumination trace, and similarly, a second plurality of the photodetectors are adapted to monitor the dots relative to the second plurality and output a third set of locations of the object's points along the illumination trace within the first time window and a fourth set of locations of the object's points along the illumination trace within the second time window.

[0109] The processing unit is capable of triangulating the first and third sets of configurations to obtain the first set of 3D positions, and the processing unit is further capable of triangulating the second and fourth sets of configurations to obtain the second set of 3D positions.

[0110] A light source is advantageous in enabling triangulation. For example, using the light source illuminated on a scene (e.g., environment), and the first and second plurality of pixel sensors are arranged, for example, at different orientations relative to each other, and the first and second plurality of pixel sensors have a shared field of view of the scene, it is possible to convert the xy-time data of the two sensors into xyz-time data by triangulation. For example, the first and second plurality of xy-time data can be converted into xyz-time data by adapting the light source to illuminate a dot on the scene in an illumination trace, the dot scanning the scene (preferably continuously), the first and second plurality monitoring the dot, and outputting the position of a point of at least one object in the scene (e.g. a point on the surface of the object) along the trace in multiple instances. The light source can, for example, act as a reference point and synchronize the arrangement of the points of the first plurality of the objects with the arrangement of the points of the second plurality of the objects to create a z dimension.

[0111] Preferably, the set of 3D positions can be filtered, for example, to reduce the number of false positive pixels and triangulate only the data of pixels with true detections. For example, as described in PCT IB2021 054688, PCT EP2021 087594, PCT IB2022 000323, PCT IB2022 058609, by checking the detection of nearby detection units (e.g. pixels) or by checking the persistence of detections over time. This is an intra-pixel filtering mechanism. Another filtering mechanism may be used. For example, after projecting the detections to the periphery, it is possible to remove noisy detections because the projection pattern is known (e.g. Lissajous, etc.). For example, filtering can be done based on the expected pulse width, the expected dot projection size, past detection values ​​of the pixel and its neighboring pixels, and the current detection value. The detected signals are then output to row and column signal buses, and the position (i.e. x, y coordinates) of the laser dot at each timestamp can be known. For example, filtering may be performed based on past bus information to identify detection locations where photons resulting from active illumination are most likely to be detected.

[0112] Preferably, the light source advantageously illuminates and scans the scene to be imaged in a Lissajous manner or pattern. This illumination trajectory is advantageous in that after a few illumination cycles a significant part of the image is already illuminated, allowing efficient and fast image detection. Other illumination patterns, such as raster scanning, are also conceivable.

[0113] Preferably, the first and / or second plurality of pixel sensors may be comprised of more than 100 sensors, preferably more than 1,000 sensors, more preferably more than 10,000 sensors, even more preferably more than 100,000 sensors, and most preferably more than 1,000,000 sensors. For example, the first and / or second plurality of pixel sensors may be arranged in a matrix, the first and / or second plurality of pixel sensors being comprised of 1,000 pixel sensor rows and 1,000 pixel sensor columns.

[0114] Preferably, the optical sensing device further comprises at least one memory element, the memory element being adapted to store at least the first set of 3D positions and / or the second set of 3D positions, e.g., the first set of 3D positions has at least twice as many data points of the object as the second set of 3D positions, the memory element may also store other data, such as the first and / or second set of configurations and / or the third set of 3D positions.

[0115] Preferably, the processing unit is capable of extrapolating a predicted set of 3D positions for a next instance based on the second set of 3D positions and the first set of 3D positions. The processing unit is capable of determining a trajectory of the object between the first time window and the second time window. The processing unit is further capable of validating the predicted set by comparing with an actual set of 3D positions.

[0116] In a fourth aspect, the present invention relates to an optical sensing system for SLAM, comprising the optical sensing device according to the third aspect, the optical sensing system further comprising a secondary optical system capable of generating an image of the scene on the second plurality of pixel sensors.

[0117] The image is for example constructed of the at least one object in the scene. The optical sensing system is capable of constructing an image of the scene with low latency, as shown in the first and second aspects, since the optical sensor is capable of constructing a second set of 3D positions of the object, constituted by a smaller number of data points, based on a first set of 3D positions of the object, constituted by a larger number of data points, of the object.

[0118] Preferably, the optical sensing system is used for 3D vision applications, for example the optical sensing system is used to visualize an object in three dimensions, for example said at least one optical sensor may enable said 3D vision.

[0119] Any features of the second aspect (i.e. the sensor), the third aspect (i.e. the device) and the fourth aspect (i.e. the system) may be correspondingly described in the first aspect (i.e. the method).

[0120] In a fifth aspect, the present invention relates to the use of a method according to the first aspect and / or a sensor according to the second aspect and / or an apparatus according to the third aspect and / or a system according to the fourth aspect for sensing, such as optical sensing, preferably SLAM.

[0121] Further features and advantages of embodiments of the present invention will now be described with reference to the drawings, in which it should be noted that the present invention is not limited to the specific embodiments shown in these drawings or described in the examples, but only by the claims.

[0122] FIG. 1 shows in (a, b) a first set of configurations (1) and a second set of configurations (2) of points of an object (3). The set of configurations (1, 2) are determined in 2D. The first set of configurations (1) is determined within a first time window (T1). The first set of configurations (1) is relative to a first reference point (28), which is for example a system implementing the method of the invention or for example a camera of the system. The second set of configurations (2) is determined within a second time window (T2) that is later than the first time window (T1). FIG. 1 also shows that the first time window (T1) is much longer than the second time window (T2). This makes the first set of configurations (1) much larger than the second set of configurations (2). For example, more points are identified in the first set of configurations (1) than in the second set of configurations (2) because the specific time, i.e. the first time window (T1), is much longer than the second time window (T2). As shown in FIG. 1(c, d), the sets of configurations (1, 2) are transformed from 2D to 3D. For example, the first set of configurations (1) is transformed into a first set of 3D positions (5) and the second set of configurations (2) is transformed into a second set of 3D positions (6). The second set of 3D positions (6) are then fitted or matched to the first set of 3D positions (5). For example, the third set of 3D positions (7) is fitted or matched to the first set of 3D positions (5) by transforming the second set of 3D positions (6) into a third set of 3D positions (7) such that the third set of 3D positions (7) fits in an optimal manner to the first set of 3D positions (5). The fitting or matching is performed at the level of the whole group of points, rather than point-by-point mapping. Thus, the number of points in said second set of 3D positions (6) must be large enough to allow a correct and accurate mapping (e.g., by having a large enough second time window (T2)), but small enough to allow a low-latency and fast mapping process. After successful matching or fitting, the transformation required to obtain such a successful fitting is determined.For example, it may be determined how far the second set of 3D positions were relative to the first set of 3D positions (i.e., how much the object moved) and how much orientation changed between the second set of 3D positions relative to the first set of 3D positions (i.e., how much the object rotated).

[0123] After successful mapping, the second set of 3D positions (6) are complemented by the first set of 3D positions (5), e.g., the first set of 3D positions (5) allows the second set of 3D positions (6) to be as complete as the first set of 3D positions (5).

[0124] FIG. 2 shows at (a, b) a first set of arrangements (1) and a second set of arrangements (2) of points of an object (3). A moving dot (8) generated by a light source (27) is illuminated on a scene (4) along an illumination trace (9), which in this case is a Lissajous pattern (10), but can also be other patterns. The density of the lines of the illumination trace (9) on the scene (4) is related to the time window through which the moving dot (8) has moved on the scene (4). The further the moving dot (8) scans the scene (4), the higher the density of the lines and the greater the number of points of the object (3) that are identified along the illumination trace (9) of the moving dot (8). As shown, the number of points along the trace (9) within a first time window (T1) is greater than the number of points along the trace (9) within a second time window (T2), simply because the time window (T1) is longer than the second time window (T2). The first and second sets of configurations (1, 2) are identified with reference to a first reference point (28). The third and fourth sets of configurations (11) and (12) are identified with reference to a second reference point (29) that is different from the first reference point (28), for example, the first and second reference points (28, 29) are separated from each other. As shown in FIG. 2(c, d), the third and fourth sets of configurations (11, 12) show the object (3) with a slightly different orientation than in FIG. 2(a, b) due to the different reference points. By triangulating the first and third sets of configurations (1, 11), a first set of 3D positions (5) is obtained, as shown in FIG. 2(e). Similarly, by triangulating the second and fourth sets of configurations (2, 12), a second set of 3D positions (6) is obtained. By triangulation, the depth information of the set of configurations (1, 2, 11, 12) can be obtained to obtain the first and second sets of 3D positions (5, 6). The second set of 3D positions (6) are then matched to the first set of 3D positions (5) in the same manner as described in FIG. 1.

[0125] FIG. 2(g) shows another method of matching or fitting based on a first set of lines (33) of the first set of 3D positions (5) and a second set of lines (34) of the second set of 3D positions (6). The first set of lines (33) is a dense set of lines defined by a dense collection of data points of the object (3). The first lines (33) consist of lines connecting consecutive points of the first set of 3D positions (5). Similarly, the second set of lines (34) is obtained. Then, a third set of lines (35) is calculated by transforming the second set of lines (34) such that the third set of lines (35) is best fitted or matched to the first set of lines (33). Matching two sets of lines is easier, faster and more accurate than matching two groups of points. Furthermore, the second or third set of lines (34, 35) may be complemented by the first set of lines (33).

[0126] Fig. 3 shows an object (3) in 3D with different orientations. The orientations are calculated, for example, relative to a reference orientation for the system performing the method. Calculating the orientation is useful for matching a third set of 3D positions (7) to the first set of 3D positions (5). Calculating the orientation is important because at least some points of the second set of 3D positions (6) may be hidden due to a change in orientation, in which case the matching of said second or third set of 3D positions (6, 7) to said first set of 3D positions (5) should be adjusted accordingly.

[0127] FIG. 4 shows the points of an object (3) identified over four time windows (14, 15, 16, 17) of different widths. Time window (14) is the shortest and time window (17) is the longest. It is clear that the longer the time window, the more points are identified. For a sufficiently long time window, for example the longest time window (17), a dense set of data (20) is obtained. After converting the dense set of data (20) to 3D, for example by triangulation (as shown above), it becomes an ideal data set to match with a sparse set of data or to provide a complementary 3D position. However, 3D data obtained within a shorter time window can also be used for matching. For example, a 3D data set acquired within time window (15) can be matched with a 3D data set acquired within time window (16) as well as with (14) against time window (16) and with (14) against time window (15). In other words, it is not necessary to match a sparse data set to a dense data set; it is also possible to match a semi-sparse data set to a semi-dense data set, or a sparse data set to a denser data set.

[0128] FIG. 5 shows an object (3) moving a trajectory (19) between a first time window (T1) and a second time window (T2). As explained above, a third set of 3D positions (7) is mapped to the first set of three-dimensional positions (5), which are then used as complementary 3D positions (7) to complement the second set of 3D positions (6). It is then desired to determine the location of the object (3) and the points that define said object (3) at a next time instance (T3) in the near future, for example 1 millisecond later. A predictive set of 3D positions (18) is extrapolated based on a trajectory (19) of the past movement of said object (3), for example between the first time window (T1) and the second time window (T2). Not only the movement itself, but also the rotation and orientation changes of said object (3) are taken into account to obtain said predictive set of 3D positions (18). For validation, an actual set of 3D positions of the object (3) at the next time instance (T3) is determined. The predicted set is compared to the actual set for validation, for example by partial determination. It should be noted that the extrapolation of the predicted set of 3D positions (18) works better if it is based on a larger set of 3D positions, for example at least 2 sets of 3D positions, for example 5 sets of 3D positions, for example 10 sets of 3D positions. For example, if the object points are determined over a longer period, the trajectory (19) will be predicted more accurately.

[0129] Figure 6 shows in (a) identifying points along a Lissajous trajectory (10), which are converted to 3D and then used to define a second set of contours (30) in (b,g).

[0130] In (a) of FIG. 6, only a few points of the object (3) have been identified, but the second set of contours (30) shows the object. In (c) of FIG. 6, the same as (a) is shown, but with a denser Lissajous trajectory (10), for example, identifying the points along the trajectory (10) is done in (c) within a longer window than in (a). By identifying the points that define the object (3) with the denser Lissajous trajectory (10) in (c) and converting them to 3D, in (d, h) a first set of contours (31) is defined, which gives a better idea of ​​the object (3) than the second set of contours (30). Another Lissajous trajectory (10) is shown in (e) and a fourth set of contours (32) is defined in (f, i).

[0131] The second set of contour lines (30) is then matched or fitted to the first set of contour lines (31) or the fourth set of contour lines (32), such that a rough identification of only a few points of the object (3) in the second set of contour lines (30) is sufficient to identify the translation and / or rotation of the object (3) during different time windows, as shown in Figure 6(a,b).

[0132] Figure 7 shows an optical sensor (21) comprising a first plurality (28) of pixel sensors (22), each pixel sensor (22) comprising a photodetector (23). The optical sensor (21) further comprises a processing unit (24) and a memory element (25). Photons are incident on the photodetectors (23), which are adapted to output a first set of locations (1) and a second set of locations (2) of an object (3) in a scene (4). The optical sensor (21) is shown in Figure 8 as part of an optical sensing device (26), in which there are a first plurality (28) and a second plurality (29) of pixel sensors (22), as well as a light source (27). The optical sensing device (26) allows to triangulate a first set of configurations (1) acquired by the first plurality (28) with a third set of configurations (11) acquired by the second plurality (29) and a second set of configurations (2) acquired by the first plurality of sensors (28) with a fourth set of configurations (12) acquired by the second plurality (29) to obtain the first and second sets of 3D positions (5, 6). This is done by illuminating the light source (27), which illuminates dots (8) on the scene (4) in a Lissajous manner (10), and the first and second plurality (28, 29) monitor the dots (8) and identify the set of configurations (1, 2, 11, 12) along a trajectory of the dots (8), in this case along a Lissajous trajectory (10). It should be noted that to allow said triangulation, the first and second pluralities (28, 29) are not in the same position but are offset relative to one another in the x and / or y axis, for example as shown in FIG.

[0133] Other arrangements for achieving the objectives of the method and apparatus embodying the present invention will be apparent to those skilled in the art. Next, the details of a specific embodiment of the present invention will be described. However, no matter how detailed the above description is in text, it will be apparent that the present invention can be applied in many ways. It should be noted that the use of a specific term in describing a particular feature or aspect of the present invention should not be interpreted as meaning that the term in this specification is redefined to be limited to the particular feature or aspect of the invention with which the term is associated. [Explanation of symbols]

[0134] T1 First time window T2 Second time window 1 First set of configurations 2 Second set of configurations 3 Objects 4 Scenes 5 1st set 3D position 6 2nd set 3D position 7 3D position of the third set 8 Dots 9 Irradiation trace 10 Lissajous trajectory 11 3rd set layout 12 4th set layout 14, 15, 16, 17 Four time windows of different widths 18 3D position of the prediction set 19 Trajectories 20 Dense data sets 21 Optical sensors 22 pixel sensor 23 photodetector 24 processing unit 25 Memory element 26 Optical sensing device 27 Light source 28 first reference point or first plurality of pixel sensors 29 second reference point or second plurality of pixel sensors 30 second set of contour lines 31 First set of contour lines 32 Fourth set of contour lines 33 1st set of lines 34 2nd set of lines 35 3rd set of lines

Claims

1. a) Preferably, a step of identifying a second set of locations (2) in 2D of points of at least one object (3) in a scene (4) within a second time window (T2), with reference to a first reference point (28), such as a first viewpoint, b) The step of converting the arrangement (2) of the second set to the 3D positions (6) of the second set, In a sensing method having, c) The process further comprises determining the transformation to the 3D position (6) of the second set such that the transformed 3D position (6) of the second set best fits the 3D position (5) of the first set corresponding to the arrangement (1) of the first set in 2D, The 3D positions (5) of the first set are denser than the 3D positions (6) of the second set, the 3D positions (5) of the first set are acquired within a first time window (T1) prior to the second time window (T2), and the second time window (T2) is shorter than the first time window (T1), preferably at least half as short. Sensing method.

2. In the sensing method according to Claim 1, The 3D positions (5) of the first set and the 3D positions (6) of the second set are each obtained by triangulation of the 2D arrangement of the two sets, using two different reference points (28, 29) as a reference. Sensing method.

3. In the sensing method according to claim 1, A step of identifying the arrangement (1) of the first set of points of the at least one object (3) in the scene (4) within the first time window (T1) with reference to the first reference point (28), A step of converting the arrangement (1) of the first set to the 3D position (5) of the first set, Having, Sensing method.

4. In the sensing method described in claim 1, The method further comprises the step of complementing the 3D positions (6) of the second set with complementary points based on the 3D positions (5) of the first set. Sensing method.

5. In the sensing method according to any one of claims 1 to 4, The steps include irradiating a dot (8) onto the scene (4) along the irradiation trace (9), The steps include monitoring the dot (8), The steps include: determining the arrangement (1, 2) of the first and second sets along the irradiation trace (9) with reference to the first reference point (28); The steps include: identifying the arrangement (11) of the third set along the irradiation trace (9) within the first time window (T1), preferably with reference to a second reference point (29), which is a second viewpoint; The steps include: identifying the arrangement (12) of the fourth set within the second time window (T2) along the irradiation trace (9), with reference to the second reference point (29); The steps include: triangulating the arrangement (1, 11) of the first and third sets to obtain the 3D position (5) of the first set; The steps include: triangulating the arrangement (2, 12) of the second and fourth sets to obtain the 3D position (6) of the second set; It also has, Sensing method.

6. In the sensing method according to any one of claims 1 to 4, The steps include defining a line (33) of the first set by consecutive points of the 3D positions (5) of the first set, The steps include defining a line (34) of the second set by consecutive points of the 3D positions (6) of the second set, A step of determining a transformation for the second set of lines (34) such that the second set of lines (34) best fits the first set of lines (33), It also has, Sensing method.

7. In the sensing method according to any one of claims 1 to 4, The steps include dividing the second time window (T2) into sub-windows in which the 3D positions of subsets are identified, For each of the 3D positions in the subset, the step of calculating the corresponding 3D position in the transformed set, It also has, Each of the three-dimensional positions of the converted set is obtained by converting the three-dimensional positions of the second set in the sub-window such that each of the three-dimensional positions of the converted set best fits the three-dimensional position of the first set. Sensing method.

8. In the sensing method according to any one of claims 1 to 4, Based on the 3D positions (6) of the second set and the 3D positions (5) of the first set, the 3D positions (18) of the predicted set for the next time instance (T3) are extrapolated by identifying the trajectory (19) of the object (3) between the first time window (T1) and the second time window (T2). Steps include identifying the 3D position of the actual set, The steps include verifying the prediction set by comparing it with the actual set, It also has, Sensing method.

9. In the sensing method according to any one of claims 1 to 4, The method further comprises the steps of combining the first and second time windows (T1, T2) and their respective points. Sensing method.

10. An optical sensor (21) for SLAM that performs position estimation and map creation simultaneously, A first plurality of (28) pixel sensors (22) each equipped with a photodetector (23), Processing unit (24), It has, The first plurality (28) of photodetectors (23) are adapted to output a second set of arrangements (2) of points of at least one object (3) in a scene (4) within a second time window (T2), with reference to the first plurality (28), wherein the second set of arrangements (2) is 2D, and the processing unit (24) is adapted to convert the second set of arrangements (2) into a second set of 3D positions (6). The processing unit (24) has a processor that determines, during use, the conversion of the 3D position (6) of the second set to the 3D position (6) of the second set so that the 3D position (6) of the second set best fits the 3D position (5) of the first set corresponding to the arrangement (1) of the first set in 2D. The 3D positions (5) of the first set are denser than the 3D positions (6) of the second set, the 3D positions (5) of the first set are in the first time window (T1) prior to the second time window (T2), and the second time window (T2) is shorter than the first time window (T1), preferably less than half the length. The processing unit (24) is adapted to acquire the 3D positions (5) of the first set and the 3D positions (6) of the second set by triangulating the arrangement of two sets in 2D with respect to two different reference points (28, 29). Characterized by, Optical sensor (21).

11. In the optical sensor (21) according to claim 10, The first plurality (28) of photodetectors (23) are further adapted to output the first set of arrangements (1) of points of at least one object (3) in the scene (4) within the first time window (T1), and the processing unit is further adapted to convert the first set of arrangements (1) into the first set of 3D positions (5). Optical sensor (21).

12. In the optical sensor (21) according to claim 10 or claim 11, The optical sensor (21) further comprises a primary optical system capable of generating an image of the scene (4) on the first plurality of (28) pixel sensors (22). Optical sensor (21).

13. An optical sensing device (26) comprising the optical sensor (21) described in claim 10, The optical sensing device (26) further comprises at least one light source (27), the light source (27) adapted to illuminate at least one dot (8) on the scene (4) along an illumination trace (9), the optical sensing device (26) further comprises a second plurality (29) of pixel sensors (22), each pixel sensor (22) comprising a photodetector (23), the first plurality (28) of the photodetectors (23) adapted to monitor the dot (8) and output the first and second sets of arrangements (1, 2) along the illumination trace (9), the second plurality (29) of the photodetectors (23) based on the second plurality (28) The processing unit (24) is configured to monitor the dot (8) and output the arrangement (11) of a third set of points of the object (3) along the irradiation trace (9) within the first time window (T1), and to output the arrangement (12) of a fourth set of points of the object (3) along the irradiation trace (9) within the second time window (T2), wherein the processing unit (24) is capable of triangulating the arrangements (1, 11) of the first and third sets in order to obtain the 3D position (5) of the first set, and the processing unit (24) is capable of triangulating the arrangements (2, 12) of the second and fourth sets in order to obtain the 3D position (6) of the second set. Optical sensing device (26).

14. In the optical sensing device (26) according to claim 13, The optical sensing device (26) further comprises at least one memory element (25), the memory element (25) being configured to store at least the first set of 3D positions (5) and / or the second set of 3D positions (6). Optical sensing device (26).

15. In the optical sensing device (26) according to claim 13 or claim 14, The processing unit (24) is capable of extrapolating the 3D position of a predicted set (18) for the next instance (T3) based on the 3D position (6) of the second set and the 3D position (5) of the first set; the processing unit (24) is capable of identifying the trajectory (19) of the object (3) between the first time window (T1) and the second time window (T2); and the processing unit (24) is further capable of verifying the predicted set (18) by comparing it with the 3D position of the actual set. Optical sensing device (26).

16. An optical sensing system comprising an optical sensing device (26) according to claim 13 or claim 14, The optical sensing system further comprises a secondary optical system capable of generating an image of the scene (4) on the second plurality of (29) pixel sensors (22). Optical sensing system.