Feature Space Distribution Management Method and Electronic Device for Simultaneous Localization and Mapping

By collecting and optimizing feature points in three-dimensional space in the AR system, and making them evenly distributed in 2D images and 3D space, the problem of poor positioning and mapping effects in the prior art is solved, and more efficient feature management and positioning accuracy are achieved.

CN115004229BActive Publication Date: 2025-05-27GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180011166.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-11
Filing Date
2021-02-07
Publication Date
2025-05-27
Estimated Expiration
2041-02-07

AI Technical Summary

Technical Problem

The existing AR system fails to effectively manage the spatial distribution of feature in positioning and map building, resulting in poor positioning and map building results.

Method used

By collecting two-dimensional images in three-dimensional space, select a subset of feature points that are evenly distributed in the image, and optimize the coordinates of these feature points based on point cloud and camera posture to ensure the uniform distribution of feature points in 3D space.

Benefits of technology

It realizes more accurate and efficient positioning and mapping, avoids the positioning accuracy and frame rate problems caused by uneven distribution of feature points, and improves the processing speed and the utilization efficiency of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115004229B_ABST
    Figure CN115004229B_ABST
Patent Text Reader

Abstract

A method for managing the spatial distribution of features for localization and mapping may include acquiring a 2D image of a 3D space. The 2D image may include a plurality of tracked feature points. The method may further include: selecting a first subset of feature points that are evenly distributed in the 2D image from the plurality of feature points, and retrieving the coordinates of the first subset of feature points from the point cloud. These coordinates define the positions of the first subset of feature points in the global coordinate map of the 3D space. The method may further include selecting a second subset of feature points that are evenly distributed in the 3D space from the first subset of feature points, optimizing the coordinates of the second subset of feature points based on the camera pose, and updating the coordinates of the second subset of feature points in the point cloud.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to extended reality technologies, and more specifically but not limited thereto, to management of feature space distribution for localization and mapping. Background Art

[0002] Augmented reality (AR) superimposes virtual content on a user's view of the real world. With the development of AR software development kits (SDKs), the mobile industry has introduced smartphone AR into the mainstream. AR SDKs typically provide six degrees-of-freedom (6DoF) tracking capabilities. A user can use a smartphone's camera to scan the environment, and the smartphone performs visual inertial odometry (VIO) in real time. After continuously tracking the camera pose, virtual objects can be placed into the AR scene to create an illusion of real and virtual objects merging together.

[0003] Despite the progress of AR systems, there is still a need in the art for improved methods and systems related to management of feature space distribution for localization and mapping. Summary of the Invention

[0004] Embodiments may include a method for localization and mapping. In some embodiments, the method may include acquiring a two-dimensional (2D) image of a three-dimensional (3D) space. The 2D image includes a plurality of feature points. The method may further include selecting a first subset of feature points that are uniformly distributed in the 2D image from the plurality of feature points. The method may further include determining coordinates of the feature points in the first subset of feature points from a point cloud. The point cloud may include previously determined coordinates that define the positions of at least the feature points in the first subset of feature points in a global coordinate map of the 3D space. The method may further include selecting a second subset of feature points that are uniformly distributed in the 3D space from the first subset of feature points at least in part based on the determined coordinates. The method may further include optimizing the coordinates of the feature points in the second subset of feature points at least in part based on the pose of the camera that acquired the 2D image. The pose of the camera may be determined based on data received from an inertial measurement unit (IMU) that is used to track the pose of the camera in the 3D space. The method may further include updating the coordinates of the feature points in the second subset of feature points in the point cloud using the optimized coordinates of the feature points in the second subset of feature points.

[0005] In some embodiments, the above-mentioned multiple feature points may be a first set of feature points. The method may further include detecting a second set of feature points in the above-mentioned 2D image. The method may further include selecting a third subset of feature points from the above-mentioned second set of feature points such that the combination of the above-mentioned first subset of feature points and the above-mentioned third subset of feature points can be evenly distributed in the above-mentioned 2D image. In some embodiments, the above-mentioned third subset of feature points may be selected such that the number of feature points in the combination of the above-mentioned first subset of feature points and the above-mentioned third subset of feature points can be within a predetermined range.

[0006] In some embodiments, the method may further include tracking the above-mentioned third subset of feature points in one or more subsequently acquired 2D images. The method may further include determining the coordinates of the feature points in the above-mentioned third subset of feature points using triangulation. The method may further include adding the above-mentioned third subset of feature points and the above-mentioned coordinates of the feature points in the above-mentioned third subset of feature points to the above-mentioned point cloud. In some embodiments, the method may further include acquiring another 2D image of the above-mentioned 3D space. The above-mentioned another 2D image may include a part of the above-mentioned second subset of feature points and / or a part of the above-mentioned third subset of feature points. The method may further include selecting a fourth subset of feature points that are evenly distributed in the above-mentioned another 2D image from a part of the above-mentioned second subset of feature points and / or a part of the above-mentioned third subset of feature points. The method may further include determining the coordinates of the feature points in the above-mentioned fourth subset of feature points from the above-mentioned point cloud. The method may further include selecting a fifth subset of feature points from the above-mentioned fourth subset of feature points such that the above-mentioned fifth subset of feature points can be evenly distributed in the above-mentioned 3D space. The method may further include optimizing the coordinates of the feature points in the above-mentioned fifth subset of feature points at least partially based on the pose of the camera that acquired the above-mentioned another 2D image.

[0007] In some embodiments, the above-mentioned multiple feature points may be detected based on the colors and / or color intensities of one or more neighboring pixels and surrounding pixels representing each of the above-mentioned multiple feature points in one or more previously acquired 2D images. In some embodiments, the method may further include dividing the above-mentioned 2D image into a plurality of regions of equal size. The even distribution of the above-mentioned first subset of feature points in the above-mentioned 2D image may mean that the feature points in the above-mentioned first subset of feature points are distributed in different regions among the above-mentioned plurality of regions. In some embodiments, the method may further include dividing the space in the above-mentioned global coordinate map corresponding to the above-mentioned 3D space into a plurality of blocks of equal size. The even distribution of the above-mentioned second subset of feature points in the above-mentioned 3D space may mean that the feature points in the above-mentioned second subset of feature points are distributed in different volumes among the above-mentioned plurality of volumes.

[0008] The embodiment may also include an electronic device for managing the feature space distribution for positioning and mapping. The electronic device may include a camera, one or more processors, and a memory. The memory may have instructions that, when executed by the one or more processors, may cause the electronic device to perform: acquiring a two-dimensional (2D) image of a three-dimensional (3D) space using the camera. The 2D image may include a plurality of feature points. When executed by the one or more processors, the instructions may further cause the electronic device to perform: selecting a first subset of feature points that are evenly distributed in the 2D image from the plurality of feature points. When executed by the one or more processors, the instructions may further cause the electronic device to perform: determining the coordinates of the feature points in the first subset of feature points from a point cloud. The point cloud may include previously determined coordinates that define the positions of at least the feature points in the first subset of feature points in a global coordinate map of the 3D space. When executed by the one or more processors, the instructions may further cause the electronic device to perform: selecting a second subset of feature points that are evenly distributed in the 3D space from the first subset of feature points at least partially based on the determined coordinates.

[0009] In some embodiments, when executed by the one or more processors, the instructions may further cause the electronic device to perform: optimizing the coordinates of the feature points in the second subset of feature points at least partially based on the pose of the camera that acquired the 2D image. The pose of the camera may be determined based on data received from a sensor unit of the electronic device. The sensor unit may include an inertial measurement unit (IMU) for tracking the pose of the camera in the 3D space. When executed by the one or more processors, the instructions may further cause the electronic device to perform: updating the coordinates of the feature points in the second subset of feature points in the point cloud using the optimized coordinates of the feature points in the second subset of feature points.

[0010] In some embodiments, the plurality of feature points may be a first set of feature points. When executed by the one or more processors, the instructions may further cause the electronic device to perform: detecting a second set of feature points in the 2D image. When executed by the one or more processors, the instructions may further cause the electronic device to perform: selecting a third subset of feature points from the second set of feature points such that the combination of the first subset of feature points and the third subset of feature points may be evenly distributed in the 2D image.

[0011] In some embodiments, when executed by the one or more processors, the instructions may further cause the electronic device to perform: tracking the third subset of feature points in one or more subsequently acquired 2D images. When executed by the one or more processors, the instructions may further cause the electronic device to perform: using triangulation to determine the coordinates of the feature points in the third subset of feature points. When executed by the one or more processors, the instructions may further cause the electronic device to perform: adding the third subset of feature points and the coordinates of the feature points in the third subset of feature points to the point cloud.

[0012] In some embodiments, when executed by the one or more processors, the instructions may further cause the electronic device to perform: dividing the 2D image into a plurality of regions of equal size. The uniform distribution of the first subset of feature points in the 2D image may mean that the feature points in the first subset of feature points are distributed in different regions among the plurality of regions. In some embodiments, when executed by the one or more processors, the instructions may further cause the electronic device to perform: dividing the space in the global coordinate map corresponding to the 3D space into a plurality of blocks of equal size, wherein the uniform distribution of the second subset of feature points in the 3D space may mean that the feature points in the second subset of feature points are distributed in different blocks among the plurality of blocks.

[0013] The embodiment may further provide a non-transitory machine-readable medium having instructions for using an electronic device to manage the spatial distribution of features for positioning and mapping, the electronic device having a camera and one or more processors. The instructions may be executed by the one or more processors to cause the electronic device to perform: acquiring a two-dimensional (2D) image of a three-dimensional (3D) space using the camera. The 2D image may include a plurality of feature points. The instructions may be executed by the one or more processors to cause the electronic device to further perform: selecting a first subset of feature points that are uniformly distributed in the 2D image from the plurality of feature points. The instructions may be executed by the one or more processors to cause the electronic device to further perform: determining the coordinates of the feature points in the first subset of feature points from a point cloud. The point cloud may include previously determined coordinates that define the positions of at least the feature points in the first subset of feature points in the global coordinate map of the 3D space. The instructions may be executed by the one or more processors to cause the electronic device to further perform: selecting a second subset of feature points that are uniformly distributed in the 3D space from the first subset of feature points at least partially based on the determined coordinates.

[0014] In some embodiments, the above instructions may be executed by the above one or more processors to cause the above electronic device to further perform: optimizing the coordinates of the feature points in the above second subset of feature points at least partially based on the pose of the camera that captures the above 2D image. The pose of the above camera may be determined based on data received from the sensor unit of the above electronic device. The above sensor unit may include an inertial measurement unit (IMU) for tracking the pose of the above camera in the above 3D space. The above instructions may be executed by the above one or more processors to cause the above electronic device to further perform: updating the coordinates of the feature points in the above second subset of feature points in the above point cloud using the optimized coordinates of the feature points in the above second subset of feature points.

[0015] In some embodiments, the above multiple feature points may be a first set of feature points. The above instructions may be executed by the above one or more processors to cause the above electronic device to further perform: detecting a second set of feature points in the above 2D image. The above instructions may be executed by the above one or more processors to cause the above electronic device to further perform: selecting a third subset of feature points from the above second set of feature points such that the combination of the above first subset of feature points and the above third subset of feature points may be evenly distributed in the above 2D image.

[0016] In some embodiments, the above instructions may be executed by the above one or more processors to cause the above electronic device to further perform: tracking the above third subset of feature points in one or more subsequently captured 2D images. The above instructions may be executed by the above one or more processors to cause the above electronic device to further perform: determining the coordinates of the feature points in the above third subset of feature points using triangulation. The above instructions may be executed by the above one or more processors to cause the above electronic device to further perform: adding the above third subset of feature points and the above coordinates of the feature points in the above third subset of feature points to the above point cloud.

[0017] In some embodiments, the above third subset of feature points may be selected such that the number of feature points in the combination of the above first subset of feature points and the above third subset of feature points may be within a predetermined range. In some embodiments, the above instructions may be executed by the above one or more processors to cause the above electronic device to further perform: dividing the above 2D image into a plurality of regions of equal size. The above first subset of feature points being evenly distributed in the above 2D image may mean that the feature points in the above first subset of feature points are distributed in different ones of the above plurality of regions.

[0018] Compared with traditional technologies, the present invention has many advantages. For example, embodiments of the present disclosure relate to methods and systems for providing a feature management mechanism for managing the spatial distribution of feature points to implement a more accurate and efficient positioning and mapping system. In addition, some embodiments dynamically select feature points that are not only evenly distributed in a 2D image but also evenly distributed in the 3D space represented by the 2D image. These and other embodiments of the present invention, along with many of its advantages and features, will be described in more detail in conjunction with the following text and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a simplified schematic diagram of a 3D space in which some embodiments may be implemented.

[0020] Figure 2A Schematically shows feature points distributed in a 3D space according to an embodiment of the present invention.

[0021] Figure 2B Schematically shows according to an embodiment of the present invention Figure 2A the distribution of the feature points in a 2D image representing the front view of the 3D space.

[0022] Figure 2C Schematically shows according to an embodiment of the present invention Figure 2A the distribution of the feature points in a 2D image representing the right view of the 3D space.

[0023] Figure 3 is a block diagram showing the functions of feature management for positioning and mapping when implementing an extended reality session according to an embodiment of the present invention.

[0024] Figure 4 is a flowchart showing an embodiment of a feature management method for positioning and mapping when implementing an extended reality session according to an embodiment of the present invention.

[0025] Figure 5 Shows a block diagram of an embodiment of a computer system.

[0026] In the drawings, similar components and / or features may have the same reference numerals. In addition, each component of the same type can be distinguished by adding a dash and a second label for distinguishing similar components after the reference numeral. If only the first reference numeral is used in the specification, it means that it applies to any similar components with the same first reference numeral, regardless of the second reference numeral. DETAILED DESCRIPTION

[0027] In the following description, various embodiments will be described. For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments may be practiced without these specific details. In addition, well-known features may be omitted or simplified in order to make the described embodiments clear.

[0028] Various platforms have been developed to implement various extended reality technologies, such as VR, AR, MR, etc. Using different application programming interfaces (APIs), some platforms can enable a user's mobile phone to sense its environment, understand, and interact with the world. Some APIs can be applicable to different operating systems, such as Android and iOS, to achieve a shared extended reality experience. Some platforms can utilize three main functions, such as motion tracking, environment understanding, and light estimation, to combine virtual content with the real world seen through the camera of the user's mobile phone. For example, motion tracking can enable the user's mobile phone or camera to understand and track its position relative to the world. Environment understanding can enable the mobile phone to detect the size and position of a flat horizontal surface (such as the ground or a coffee table). Light estimation can enable the mobile phone to estimate the current lighting conditions of the environment.

[0029] Some platforms can allow developers to integrate the camera and motion functions of the user device to create an augmented reality experience in an application installed on the user device. The platform can combine device motion tracking, camera scene capture, advanced scene processing, and display convenience to simplify the task of establishing an extended reality experience. These technologies can be used to create different types of extended reality experiences using the rear camera and / or front camera of the user device.

[0030] Motion tracking systems used by existing platforms, such as simultaneous localization and mapping (SLAM), visual inertial simultaneous localization and mapping (VISLAM), or visual inertial odometry (VIO), can synchronously acquire an image stream and an inertial measurement unit (IMU) data stream, and then create a combined output through processing and feature point detection. One class of existing solutions to the motion tracking and / or localization and mapping problem mainly relies on feature points detected in two-dimensional (2D) images.

[0031] Generally speaking, the more evenly distributed the detected feature points are, the better the positioning and mapping effects that can be achieved. Existing positioning and mapping systems only rely on feature points in 2D images and only consider the distribution of feature points in 2D images. Existing systems do not consider the distribution of feature points in the 3D space represented by 2D images. However, the distributions of feature points in 2D images and 3D spaces both affect the results of positioning and mapping. This is because the distribution of feature points in 2D images depends on the perspective of the camera. Therefore, maintaining an even distribution of detected feature points in 2D images does not necessarily mean that the feature points have an even distribution in 3D space, which can have a great impact on the accuracy and frame rate of the positioning and mapping system.

[0032] The embodiments described in this paper provide a feature management mechanism for managing the spatial distribution of feature points to achieve a more accurate and efficient positioning and mapping system. The embodiments described here dynamically select feature points that are not only evenly distributed in 2D images but also evenly distributed in the 3D space represented by 2D images. By continuously monitoring the spatial distribution of feature points, more accurate positioning and mapping results can be obtained. In addition, by maintaining the even distribution of feature points in 2D images and 3D spaces, diverse feature points representing the entire 3D space rather than a local area can be selected to obtain better optimization results. Moreover, by selecting feature points that are evenly distributed in 2D images and 3D spaces, the selection of duplicate feature points (multiple feature points in a single square or cube) can be avoided, and the problem scale can be managed or reduced, thereby improving the processing speed, such as faster optimization speed, lower power consumption, and more efficient utilization of computing resources.

[0033] Figure 1 is a simplified schematic diagram of a 3D space in which some embodiments can be implemented. Although Figure 1 shows an indoor space, the various embodiments described in this paper can also be implemented in outdoor environments or other environments where extended reality sessions can be performed.

[0034] An extended reality session such as an AR session or an MR session can be implemented using the electronic device 110. The electronic device 110 can include a camera 112 for acquiring 2D images of the 3D space 100 for implementing the extended reality session, and a display 114 for displaying the 2D images acquired by the camera 112 and presenting one or more rendered virtual objects 116. The electronic device 110 can include additional software and / or hardware, including but not limited to the IMU 118, to implement the extended reality session and various other functions of the electronic device 110, such as allowing the electronic device 110 to estimate and track the pose (e.g., position and orientation) of the camera 112 relative to the 3D space for motion tracking.

[0035] Although Figure 1FIG. 0 shows a tablet computer as an example of an electronic device 110, but the various embodiments described herein may not be limited to implementation on a tablet computer. The various embodiments may be implemented using many other types of electronic devices, including but not limited to various mobile devices, handheld devices (such as smartphones or tablet computers), hands-free devices, wearable devices (such as head-mounted displays (HMDs), optical see-through head-mounted displays (OST HMDs), or smart glasses), or any device capable of implementing AR, MR, or other extended reality applications.

[0036] In Figure 1 the illustrated embodiment, although the electronic device 110 uses the display 114 to display or render a camera view with or without virtual objects superimposed thereon, the embodiments described herein may be implemented without rendering the camera view on the display. For example, some OST-HMDs and / or smart glasses may not include a display for displaying or rendering a camera view. The OST-HMDs and / or smart glasses may include a transparent display screen through which the user can see the real-world environment. Nevertheless, the embodiments described herein may be implemented in OST-HMDs and / or smart glasses to achieve a more accurate and efficient implementation of an extended reality session.

[0037] To implement an extended reality session, an extended reality application, such as an AR and / or MR application that may be stored in the memory of the electronic device 110, may be activated to start the extended reality session. Via the application, the electronic device 110 may scan the surrounding environment, such as the 3D space 100, in which the extended reality session may be conducted. The 3D space 100 may be scanned using the camera 112 of the electronic device 110 to create a global coordinate map or framework of the 3D space 100. Other available sensor inputs, such as inputs received from the lidar, radar, sonar, or various other sensors of the electronic device 110, may also be used to create the global coordinate map. Since new inputs may be received from the various sensors of the electronic device 110, the global coordinate map may be continuously updated.

[0038] The global coordinate map can be a 3D map and can be used to track the pose of the camera 112 when the user moves the electronic device 110 in the 3D space 100, such as the position and orientation of the camera 112 relative to the 3D space. The position can represent or indicate where the camera 112 is, and the orientation can represent or indicate the direction that the camera 112 faces or points to. The camera 112 can move in three degrees of freedom to change its position and can rotate in three degrees of freedom to change its orientation. The global coordinate map can also be used to continuously track the positions of individual feature points in the 3D space.

[0039] The construction of the global coordinate map of the 3D space 100 and / or the tracking of the position and orientation of the camera 112 can be performed using various positioning and mapping systems, such as SLAM for constructing and / or updating a map of an unknown environment while tracking the position and orientation of a moving object (such as the camera 112 of the electronic device 110 described herein).

[0040] As new inputs are received from the camera 112, it may be necessary to optimize the global coordinate map of the construction of the 3D space 100, the estimation of the pose of the camera 112, and / or the estimation of the feature point positions through the positioning and mapping system. It can be understood that maintaining the ability to accurately locate the electronic device 110 and / or the camera 112 will affect the accuracy and efficiency of the implemented extended reality session. As described above, existing positioning and mapping technologies mainly rely on feature points detected in 2D images of the 3D space (such as Figure 1 the corners 122a, 122b, 122c, 122d of the table 120 shown) to construct the global coordinate map and / or estimate the camera pose and the positions of the feature points. However, as also described above, the uniform distribution of the detected feature points in the 2D image does not necessarily mean that these feature points are uniformly distributed in the 3D space, which will be discussed in more detail with reference to Figures 2A - 2C The non-uniform distribution of feature points in the 3D space will have a great impact on the accuracy of the positioning and mapping system.

[0041] Figure 2A Schematically shows feature points distributed in the 3D space according to an embodiment of the present invention. Although a hexahedron is shown in Figure 2A for illustration, the 3D space in which an extended reality session can be implemented can have any form and / or size. Figure 2B Schematically shows the distribution of the same feature points in a 2D image representing the front view of the 3D space according to an embodiment of the present invention. Figure 2C Schematically shows the distribution of the same feature points in a 2D image representing the right view of the 3D space according to an embodiment of the present invention.

[0042] As shown in Figure 2AAs shown, the 3D space is divided into multiple volumes (or blocks), for example, multiple cubes, and these blocks have equal predefined sizes. It should be understood that the real or physical 3D space is not divided into multiple cubes; rather, the corresponding space in the global coordinate map is divided. Different cube sizes (dimensions) can be defined in multiple layers. For example, in the first layer, the 3D space can be divided into a first set of cubes having a first size, such as Figure 2A the eight cubes shown in Figure 2A . In the second layer, each of the first set of cubes having the first size can be further divided into a second set of cubes having equal second sizes, where the second size is smaller than the first size, and in the third layer, each of the second set of cubes in the second layer can be further divided into a third set of cubes having equal third sizes, where the third size is smaller than the second size, and so on. Depending on computing power, the required resolution, the availability of detected feature points, and other factors, feature points can be tracked in different layers. Although the division of the 3D space into cubes is described by way of example, in various embodiments, the 3D space can be divided into other forms of 3D volumes. Although Figure 2A shows eight cubes for illustration, the 3D space can be divided into more cubes, such as dozens, hundreds, or thousands of cubes, or even a greater number of cubes. Additionally, although only one layer of division is shown for clarity,

[0043] multiple layers of division can be implemented.

[0044] In some embodiments, to determine whether the cubes in which the feature points are distributed are distributed or scattered in the 3D space rather than concentrated in a local area, after determining that the feature points are distributed in different cubes, the distribution of the feature points or the cubes with feature points along each of the three dimensions of the 3D space can also be determined, which will be discussed in more detail below. By selecting and / or tracking the feature points evenly distributed in the 3D space (e.g., distributed among different cubes), more accurate positioning and mapping of the entire 3D space can be achieved.

[0045] Further referring to Figures 2A - 2C , four feature points 202, 204, 206, 208 can be detected and tracked in the 2D image of the 3D space for positioning and mapping. The 2D image is, for example, Figure 2B the front view of the 3D space shown in Figure 2C and the right view of the 3D space shown in Figure 2B . Similar to dividing the 3D space into multiple layers, each 2D image can also be divided into multiple layers. For example, in the first layer, each 2D image can be divided into a first set of regions, such as squares, and each region in the first set has an equal first size, such as Figure 2C the four squares shown in or

[0046] the four squares shown in . In the second layer, each of the first set of squares with the first size can be further divided into a second set of squares with an equal second size, and the second size is smaller than the first size. And in the third layer, each of the second set of squares with the second size can be further divided into a third set of squares with an equal third size, and the third size is smaller than the second size, and so on. Figure 2B Although it is exemplarily described that each 2D image is divided into squares, in various embodiments, each 2D image can be divided into other forms of 2D regions. Although Figure 2C and Figure 2B show four squares for illustration, each 2D image can be divided into more squares, such as dozens, hundreds or thousands of squares, or even more squares. In addition, although Figure 2C and show only one layer of division for clarity, multiple layers of division can be implemented.

[0047] When discussing the distribution of feature points in the context of a 2D image, in some embodiments, a uniform distribution of feature points means that the feature points are distributed in different squares among a plurality of squares in a layer, and the squares in which the feature points are distributed are dispersed throughout the 2D image rather than being concentrated in a local area of the 2D image. In some embodiments, to determine whether the squares in which the feature points are distributed are distributed or scattered in the 2D image rather than being concentrated in a local area, after determining that the feature points are distributed in different squares, the distribution along each of the two dimensions of the 2D image can also be determined, which will be discussed in more detail below.

[0048] In some embodiments, a uniform distribution of feature points in a 2D image can mean that each of the squares in which the feature points are distributed can include more than one feature point, and / or the squares in which the feature points are distributed are distributed throughout the 2D image rather than being concentrated in a local area of the 2D image. The total number of feature points included in each square can be less than a predetermined number. The number of feature points included in each square can be the same or different. In some embodiments, when the number of feature points included in each square is greater than 1, the square can be further divided into smaller squares such that each of the squares in which the feature points are distributed can include one feature point.

[0049] Continuing to refer to Figure 2B , when viewed only from the front view of the 3D space, the feature points 202, 204, 206, 208 appear to be uniformly distributed in the 2D image because each feature point is distributed in a different square. Additionally, the squares having feature points are distributed or scattered throughout the 2D image. Specifically, the squares having feature points are uniformly distributed along the vertical dimension of the two-dimensional image and are uniformly distributed along the horizontal dimension of the two-dimensional image. However, referring to Figure 2C , when viewed from the right view of the 3D space, it can be seen that the feature points 202, 204, 206, 208 are not evenly distributed because the feature points 204 and 206 are distributed in the same square, that is, Figure 2C the bottom-left square of the 2D image. Therefore, in many existing localization and mapping systems, when only the 2D positions of the feature points are examined, the uniform distribution of the feature points may not actually be maintained, and the localization and mapping results may be inaccurate because as Figure 2B shows, the tracked feature points may appear to be uniformly distributed in a 2D image, but in fact may not be uniformly distributed as shown in the 2D image of Figure 2C .

[0050] Embodiments of the present invention described herein avoid this problem associated with existing localization and mapping systems by introducing a feature management mechanism that examines the distribution of feature points in 2D images and 3D space. By examining the distribution of feature points in 2D images and 3D space, feature points that are not only evenly distributed in 2D images but also evenly distributed in 3D space can be selected for tracking, enabling more accurate and efficient localization and mapping. As an example, in the embodiments described herein, by examining the distribution of feature points in 3D space, feature point 210 and feature points 202, 204, 208 can be selected and tracked, while feature point 206 is not selected. This is because feature points 202, 204, 206, 208 are not evenly distributed in 3D space, while feature points 202, 204, 208, 210 can be considered to be evenly distributed in 3D space.

[0051] Specifically, in 3D space, feature point 202 is distributed in the front upper left cube, feature point 204 is distributed in the front lower left cube, feature point 206 is distributed in the front lower right cube, feature point 208 is distributed in the rear upper right cube, and feature point 210 is distributed in the rear lower right cube. When feature points 202, 204, 206, 208 are selected, although feature points 202, 204, 206, 208 are distributed in different cubes, these cubes are distributed in a clustered manner in 3D space because three of the four cubes with feature points are located in the front region of the 3D space, while only one cube with a feature point is located in the rear region of the 3D space. However, when feature points 202, 204, 208, 210 are selected, not only are the feature points distributed among different cubes, but the cubes with feature points are not distributed in a clustered manner, but are evenly distributed along each of the three dimensions of the three-dimensional space. Specifically, the two cubes with feature points 202, 204 among the four cubes are in the front region of the 3D space, and the two cubes with feature points 208, 210 among the four cubes are in the rear region of the 3D space; the two cubes with feature points 202, 204 among the four cubes are in the left region of the 3D space, and the two cubes with feature points 208, 210 among the four cubes are in the right region of the 3D space; the two cubes with feature points 202, 208 among the four cubes are in the top region of the 3D space, and the two cubes with feature points 204, 210 among the four cubes are in the bottom region of the 3D space.

[0052] Therefore, by examining the distribution of feature points in 3D space, feature points that are more evenly distributed, such as feature points 202, 204, 208, 210, can be selected instead of feature points that may be unevenly distributed, such as feature points 202, 204, 206, 208. By selecting feature points that are evenly distributed in 3D space, an even distribution in different 2D images or views of the 3D space can be maintained, such asFigure 2B and Figure 2C the uniform distribution of the feature points 202, 204, 208, and 210 shown, thereby achieving a more accurate 2D representation of the 3D space and more accurate and efficient localization and mapping.

[0053] Figure 3 is a block diagram 300 showing the functions of feature management for localization and mapping according to some embodiments. The functions of the various blocks within the flowchart can be executed by the hardware and / or software components of the electronic device 110. Note that alternative embodiments may vary the functions shown to add, remove, combine, separate, and / or rearrange the various functions shown. Such variations will be understood by those of ordinary skill in the art.

[0054] At block 305, the feature points in each 2D image received can be tracked, where the 2D images are from the image data 302 stream. These tracked feature points can be feature points that were previously detected in the received 2D images and continuously tracked, including those detected in the previous 2D image and continuously tracked in the currently received 2D image. Image processing techniques can be used to detect and track the feature points in the 2D images. In some embodiments, the feature points in each 2D image can be detected and tracked based on the color, color intensity, or brightness of one or more pixels and the surrounding pixels. For example, the 2D image can be analyzed to detect adjacent pixels having the same or similar color and / or color intensity. Then, the size of each group of adjacent pixels can be analyzed to determine whether each group of pixels can represent or can be classified as a point, edge, plane, surface, etc. For example, when a group of adjacent pixels is determined to be within a predetermined radius, the group of adjacent pixels can be classified as a point, and the corresponding feature represented by the group of adjacent pixels can be classified as a feature point. The predetermined radius can be measured by a predetermined number of pixels, such as 5 pixels, 10 pixels, 15 pixels, 20 pixels, 25 pixels, 50 pixels, 100 pixels, etc., and the number of pixels can be determined according to the required resolution, processing power, etc.

[0055] As an example, referring back to Figure 1, a 2D image of the 3D space 100 can be analyzed. The neighboring pixels representing each of the corners 122a, 122b, 122c, 122d of the table 120 in the 2D image can be determined to have the same or similar color and / or color intensity and can be further determined to be within a predetermined radius. Thus, the four corners 122a, 122b, 122c, 122d can be detected as feature points and can be tracked in the 2D image of the acquired 3D space. These feature points may help to define or construct geometric constraints for implementing an extended session, such as geometric constraints of the surface on which virtual objects (such as virtual object 116) can be presented. Although feature points are described as exemplary features that can be detected and tracked, the features that can be tracked are not limited to points. For example, depending on the specific application, edges, regions, planes, etc. can be detected and tracked.

[0056] At block 310, outliers among the detected and tracked feature points can be removed. Removing the outlier feature points can include pausing or stopping the tracking of the feature point. In some embodiments, the outlier feature points can include feature points that interfere with the uniform distribution of the above-mentioned feature points in the current 2D image. For example, referring to Figure 2B and Figure 2C , among the feature points 202, 204, 206, 208, 210, the feature point 206 can be considered an outlier feature point because the feature points 202, 204, 208, 210 are uniformly distributed in the Figure 2B and Figure 2C 2D image, while including the feature point 206 will result in a non-uniform distribution of the feature points. As described above, the uniform distribution can be determined by the following method: dividing the 2D image into multiple regions of square or other appropriate shapes, determining whether the feature points are distributed in different squares, and / or checking the distribution of the squares along each dimension of the 2D image. Thus, the outlier feature points can be removed by selecting the feature points that are uniformly distributed in the current 2D image.

[0057] In some embodiments, the outlier feature points can also include feature points whose speed or moving speed may be greater than a predetermined threshold. In some embodiments, the predetermined threshold can be a predefined speed value or moving speed. In some embodiments, the predetermined threshold can be the percentage difference in speed or moving speed between different feature points. For example, by tracking the feature points, the speed of the feature points can be determined. The feature points moving at a speed or moving speed higher than the predetermined threshold can be removed. The change in speed or moving speed between the feature points can also be determined. For example, the average speed or average moving speed of the feature points can be determined. The feature points whose speed or moving speed exceeds a predetermined percentage of the average value can be removed, for example, exceeding 30%, exceeding 40%, exceeding 50%, exceeding 60%, exceeding 70%, exceeding 80%, exceeding 90%, or more.

[0058] At block 315, new feature points in the current 2D image can be detected. Similar to detecting feature points in a previous 2D image, feature points in the current 2D image can be detected by analyzing the colors, color intensities, sizes, etc. of neighboring and surrounding pixels.

[0059] At block 320, the newly detected feature points can be selected such that a uniform distribution of feature points in the current 2D image can be maintained. In some embodiments, to maintain a uniform distribution of feature points in the current 2D image, the distribution of the newly detected feature points can be analyzed by checking whether the newly detected feature points are distributed in different squares of the 2D image, and / or by checking whether the newly detected feature points are distributed throughout the 2D image or are clustered in one or more local regions. Duplicate feature points, i.e., feature points that are distributed in the same square, can be removed so that only one feature point in each square can be selected. The clustered feature points can be reduced such that the distribution of the newly detected feature points in the entire 2D image can be uniform.

[0060] In some embodiments, in addition to analyzing the distribution of the newly detected feature points, the distribution of the previously detected and tracked feature points combined with the newly detected feature points can also be analyzed. If there are duplicate feature points, the newly detected feature points can be removed and the previously detected and tracked feature points can be retained because the coordinates of the previously detected and tracked feature points may have been determined and optimized multiple times (as will be discussed below) and are thus more accurate. In some embodiments, among the duplicate feature points, for other considerations, the newly detected feature points can be retained while the previously detected and tracked feature points are removed. Once the newly detected feature points are selected, these newly selected feature points can be tracked in subsequently acquired 2D images in a manner similar to tracking the previously detected feature points in the previous 2D image at block 305.

[0061] In some embodiments, when selecting the newly detected feature points at block 320, only the distribution of the newly detected feature points can be checked and the detected feature points can be selected accordingly such that the newly detected feature points selected are uniformly distributed. The combined distribution of the previously detected and tracked feature points and the newly detected feature points can be evaluated in another outlier rejection operation at block 310 to remove the feature points that interfere with the uniform distribution, so that the remaining feature points are uniformly distributed.

[0062] In some embodiments, when a newly detected feature point is selected at block 320, an appropriate number of newly detected feature points may be selected such that, in the current 2D image, the number of the selected newly detected feature points combined with the previously detected and continuously tracked feature points may be within a predetermined range. Accordingly, an excessive number of newly detected feature points may not be selected so as not to overload the computing power of the device implementing the extended reality session.

[0063] At block 325, a correspondence may be established between the feature points in the current 2D image and their coordinates (e.g., 3D coordinates) in 3D space. The 3D coordinates may be coordinates in a coordinate map that is established for mapping the 3D space and tracking the pose of the camera and / or the positions of the detected and tracked feature points in the 3D space. An exemplary global coordinate map may include a global coordinate map of the 3D space 100 established by the electronic device 110 for mapping the 3D space 100 and tracking the pose of the camera 112 and / or the positions of the detected and tracked feature points (e.g., feature points 122a, 122b, 122c, 122d).

[0064] The 3D coordinates of the feature points may be determined by triangulation. Specifically, when a certain feature point is detected and tracked in two or more 2D images, the feature point is observed by a camera (e.g., camera 112) from two or more different camera poses (i.e., different camera positions and / or different camera orientations). Accordingly, based on the different camera positions and / or orientations from which the feature point is observed, the 3D coordinates of the feature point in the 3D global coordinate map may be determined by triangulation. The determined positions or coordinates of the detected and tracked feature points may be stored in a point cloud (e.g., a 3D point cloud). Establishing a correspondence between the feature points in the 2D image and their 3D coordinates may include retrieving the 3D coordinates of the feature points in the current 2D image from the point cloud. It may be understood that not all of the feature points in the current 2D image may have available or known 3D coordinates stored in the point cloud. For example, the 3D coordinates of the newly detected feature points in the current 2D image are not yet known because these feature points are observed from only one camera pose.

[0065] At block 330, the coordinates of previously detected and tracked feature points can be optimized and / or updated. For example, based on a stream of IMU data 304 received from an IMU sensor (e.g., the IMU sensor 118 of the electronic device 110 for tracking the pose of the camera 112 in the 3D space 100) used to track the pose of the camera in the 3D space, the current camera pose can be determined. By receiving the image data 302 and the IMU data 304 in a synchronized manner, based on the visual correspondence of the feature points and the camera pose for acquiring the current 2D image, the positions of the previously detected and tracked feature points in the current 2D image can be optimized and / or updated. Since more image data and IMU data are available, the coordinates of the feature points can be calculated or updated with the new data, and the calculated or updated coordinates may be more accurate and / or optimized. As can be understood, since the feature points can be continuously detected from the new camera pose and the coordinates of the feature points can be tracked in the new 2D image, the coordinates of the feature points can be continuously optimized and / or corrected. In some embodiments, an optimization module available in an existing localization and mapping system can be used to perform the optimization, such as the optimization modules of SLAM, VISLAM, VIO, etc.

[0066] In some embodiments, at block 330, the coordinates of the feature points newly detected in the previous 2D image and continuously tracked in the current 2D image can also be determined by triangulation based on the camera pose for acquiring the previous 2D image and the camera pose for acquiring the current 2D image. When more image data and IMU data are available, the coordinates of these feature points can be continuously optimized. In a similar manner, if the positions or coordinates of the feature points newly detected from the current 2D image are continuously tracked in one or more subsequently acquired 2D images, these positions or coordinates can be determined and / or optimized. At block 335, the point cloud can be updated by storing the coordinates of the feature points that have been determined and / or optimized at block 330.

[0067] In some embodiments, the feature points evenly distributed in the 3D space can be selected at block 340 before the optimization at block 330. In these embodiments, only the feature points evenly distributed in the 3D space will be optimized at block 330 and updated in the point cloud at block 335. In other words, by selecting the feature points evenly distributed in the 3D space, only a subset of all the detected and tracked feature points in the 2D image can be optimized, which is particularly advantageous when the optimization module for localization and mapping may be limited by the system processing power.

[0068] Specifically, at block 340, based on the coordinates retrieved from the point cloud, feature points that have been previously detected and continuously tracked in the current 2D image and are evenly distributed in the 3D space can be selected. As described above, the even distribution of feature points in the 3D space can be determined by dividing the 3D space into a plurality of cubes (or other suitable forms of 3D volumes / blocks) of equal size in one or more layers and checking whether the feature points are distributed in different cubes. In some embodiments, the distribution of the cubes with feature points along each of the three dimensions of the 3D space can also be evaluated. For example, the 3D space can have an X dimension, a Y dimension, and a Z dimension. The distribution of the cubes with feature points along the X dimension, the Y dimension, and / or the Z dimension can be evaluated. In some embodiments, the distribution of the cubes along each of the X, Y, and / or Z dimensions can be evaluated based on the respective X, Y, and / or Z coordinates of the feature points within each cube. For example, the X coordinates of all the feature points can be checked to evaluate whether the cubes containing these feature points are evenly distributed in the X dimension, the Y coordinates of all the feature points can be checked to evaluate whether the cubes containing these feature points are evenly distributed in the Y dimension, and / or the Z coordinates of all the feature points can be checked to evaluate whether the cubes containing these feature points are evenly distributed in the Z dimension. By checking whether the feature points are distributed in different cubes and / or checking the distribution of the cubes with feature points along each of the three dimensions of the 3D space, feature points that are evenly distributed in the 3D space can be determined and selected. Then, the selected feature points can be used for subsequent processing.

[0069] The functions described above with reference to blocks 305 - 325 and 340 can be collectively referred to as the feature management mechanism 345, which has several advantages. By, for example, selecting newly detected feature points that are evenly distributed in at least the 2D image at block 320, duplicate feature points in the same square can be avoided, thus effectively utilizing computing power and preventing degradation of system performance. Additionally, by continuously monitoring the spatial distribution of feature points and dynamically selecting feature points that are evenly distributed not only in each 2D image but also in the 3D space, more accurate optimization and / or localization and mapping results can be obtained. Since the scale of the optimization problem (i.e., the number of feature points to be optimized and their corresponding 3D coordinates) may become smaller due to the removal of points that are unevenly distributed in the 2D image and / or 3D space, a faster optimization speed, lower power consumption, and less memory and CPU usage can be achieved.

[0070] It should be understood that Figure 3 A specific method is provided that shows the functions of feature management for localization and mapping according to an embodiment of the present invention. As described above, according to alternative embodiments, other series of steps can also be performed. For example, alternative embodiments of the present invention can perform the above steps in a different order. Additionally, Figure 3Each of the steps shown can include multiple sub - steps, which can be executed in various orders suitable for each step. Additionally, additional steps can be added or removed depending on the specific application. Those of ordinary skill in the art will recognize many variations, modifications, and alternatives.

[0071] Figure 4 is a flowchart showing an embodiment of a feature management method 400 for improving localization and mapping when implementing an extended reality session according to an embodiment of the present invention. Method 400 can be implemented, for example, by using an electronic device 110 as described herein. Thus, the means for performing one or more of the functions shown in each of the boxes Figure 4 can include hardware and / or software components of the electronic device 110. As with other Figure 1 herein, Figure 4 is provided as an example. The functions of other embodiments may be different from the functions shown. The differences can include performing additional functions, replacing and / or removing selected functions, performing functions in a different order or simultaneously, etc.

[0072] At block 405, a 2D image of a 3D space (e.g., 3D space 100) in which an extended reality session can be implemented can be captured by a camera of the electronic device (e.g., camera 112 of electronic device 110). The 2D image can include feature points that have been detected in one or more previously captured 2D images of the 3D space and continuously tracked in the current 2D image. As described above, feature points can be detected and / or tracked based on the color, color intensity, size, etc. of one or more neighboring pixels and surrounding pixels. The 2D image can also include new feature points to be detected in subsequent operations.

[0073] At block 410, a subset of feature points, e.g., a first subset of feature points, can be selected from the previously detected and continuously tracked feature points. The first subset of feature points can be selected such that these feature points are evenly distributed in the 2D image. As described above, to select feature points that are evenly distributed in the 2D image, the 2D image can be divided into multiple regions of the same size, e.g., squares. In some embodiments, multi - layer image partitioning can be implemented. In some embodiments, the first subset of feature points can be selected such that these feature points are distributed in different squares of the 2D image. Further, the first subset of feature points can be selected such that the squares in which the feature points are distributed can be evenly distributed along each of the two dimensions of the 2D image to ensure that the selected feature points are distributed or scattered in the 2D image rather than clustered in a local area. After the first subset of feature points has been selected from the previously detected and continuously tracked feature points, the remaining feature points can cease to be tracked.

[0074] In some embodiments, in addition to removing feature points with non-uniform distribution in the 2D image, method 400 may further include removing feature points whose speed or moving speed may be higher than a predetermined threshold. The predetermined threshold may be a predetermined speed or moving speed, or a predetermined percentage difference in speed or moving speed between different feature points.

[0075] At block 415, new feature points may be detected in the current 2D image. Similar to detecting feature points in previously acquired 2D images, new feature points may be detected based on the color, color intensity, size, etc. of one or more neighboring pixels and surrounding pixels in the current 2D image.

[0076] At block 420, a subset of feature points, such as a second subset of feature points, may be selected from the newly detected feature points. The second subset of feature points may be selected such that a uniform distribution of feature points in the current 2D image can be maintained. In some embodiments, in order to maintain a uniform distribution of feature points in the current 2D image, the second subset of feature points may be selected such that these feature points are evenly distributed in the current 2D image. In some embodiments, in order to maintain a uniform distribution of feature points in the current 2D image, the second subset of feature points may be selected such that the combined first subset of feature points and the second subset of feature points can be evenly distributed in the current 2D image. If there are duplicate feature points (i.e., multiple feature points in a single square) when combining the first and second subsets of feature points, the newly detected feature points may be removed and the previously detected feature points may be retained, or vice versa.

[0077] In some embodiments, in addition to maintaining a uniform distribution, when selecting the newly detected feature points, the total number after combining the first subset of feature points and the selected second subset of feature points may be maintained within a predetermined range so that the computing resources are not overburdened. The selected second subset of feature points may be added for tracking in subsequently acquired 2D images.

[0078] At block 425, the coordinates of the first subset of feature points (i.e., feature points that have been previously detected, continuously tracked, and evenly distributed in the current 2D image), such as 3D coordinates, may be retrieved from the point cloud, where the previously determined feature point coordinates may be stored in the point cloud. The coordinates may define the position of the feature point in a 3D coordinate system (e.g., a global coordinate map or frame) of the 3D space. As described above, when detecting and tracking feature points in two or more 2D images, the 3D coordinates of any feature point may be determined by triangulation. That is, when the camera (e.g., camera 112) observes the feature point from two or more different camera poses, the 3D coordinates of the feature point in the 3D global coordinate map may be determined based on the camera poses via triangulation. Based on the retrieved coordinates, a correspondence between the first subset of feature points in the 2D image and their 3D coordinates may be established.

[0079] At block 430, based on the retrieved coordinates, a third subset of feature points can be selected from the first subset of feature points such that the selected third subset of feature points is evenly distributed in the 3D space. As described above, to select feature points that are evenly distributed in the 3D space, the 3D space can be divided into a plurality of volumes / blocks of the same size in one or more layers, such as cubes. In some embodiments, a third subset of feature points can be selected such that these feature points are distributed in different cubes of the 3D space. Additionally, a third subset of feature points can be selected such that the cubes in which the feature points are distributed can be evenly distributed along each of the three dimensions of the 3D space to ensure that the selected feature points are distributed or scattered in the 3D space rather than being clustered in a local area. After the third subset of feature points has been selected from the first subset of feature points, the remaining feature points in the first subset of feature points can be not tracked.

[0080] At block 435, the coordinates of the third subset of feature points can be optimized based at least in part on the new camera pose and the current 2D image for acquiring the current 2D image. The camera pose can be determined based at least in part on data received from an IMU (such as the IMU 118 of the electronic device 110), which is used to track the camera pose as the camera moves in the 3D space. Since only the coordinates of the third subset of feature points that are evenly distributed in the 3D space can be optimized, rather than the entire set of previously detected and tracked feature points or the entire first subset of feature points, the scale or total number of the optimized feature points can be smaller, achieving a higher optimization rate.

[0081] At block 440, the coordinates of the feature points newly detected in the previous 2D image and continuously tracked in the current 2D image can be determined. Specifically, the coordinates can be determined by triangulation based on the camera pose for acquiring the previous 2D image and the camera pose for acquiring the current 2D image.

[0082] At block 445, the point cloud can be updated by storing the coordinates of the feature points that have been optimized at block 435 and / or the coordinates of the feature points that have been determined at block 440. After updating the point cloud, method 400 can return to block 405 and can acquire another 2D image of the 3D space. At least some points in the second subset of feature points and / or the third subset of feature points can be present in the newly acquired 2D image, and these points can be tracked in the newly acquired 2D image. Method 400 can continue with operations 410-445 to select feature points and optimize the coordinates of the selected feature points to achieve more accurate and efficient localization and mapping.

[0083] It should be understood that Figure 4The specific steps shown provide a particular method, showing an embodiment of a feature management method for improving localization and mapping in implementing an extended reality session according to an embodiment of the present invention. As described above, according to alternative embodiments, other series of steps may also be performed. For example, alternative embodiments of the present invention may perform the above steps in a different order. In addition, Figure 4 each of the steps shown may include multiple sub-steps, and these sub-steps may be performed in various orders according to each step. In addition, additional steps may be added or deleted according to a specific application. Those of ordinary skill in the art will recognize many variations, modifications, and alternatives.

[0084] Figure 5 is a simplified block diagram of a computing device 500. The computing device 500 may implement some or all of the functions, behaviors, and / or capabilities described above using electronic storage or processing, as well as other functions, behaviors, or capabilities not explicitly described. The computing device 500 includes a processing subsystem 502, a storage subsystem 504, a user interface 506, and / or a communication interface 508. The computing device 500 may also include other components (not explicitly shown), such as a battery, a power controller, and other components operable to provide various enhanced functions. In various embodiments, the computing device 500 may be implemented on a desktop or laptop computer, a mobile device (e.g., a tablet computer, a smart phone, a mobile phone), a wearable device, a media device, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a processor, a controller, a microcontroller, a microprocessor, or an electronic unit designed to perform the above functions or combinations of functions.

[0085] The storage subsystem 504 can be implemented using local storage and / or removable storage media, for example, using disks, flash memory (e.g., Secure Digital cards, Universal Serial Bus flash drives), or any other non-transitory storage medium or combination of media, and can include volatile and / or non-volatile storage media. Local storage can include random access memory (RAM), including dynamic RAM (DRAM), static RAM (SRAM), or battery-backed RAM. In some embodiments, the storage subsystem 504 can store one or more application programs and / or operating system programs to be executed by the processing subsystem 502, including programs for implementing some or all of the operations to be performed by a computer as described above. For example, the storage subsystem 504 can store one or more code modules 510 to implement one or more of the method steps described above.

[0086] Firmware and / or software implementations can be realized using modules (e.g., processes, functions, etc.). A machine-readable medium tangibly embodying instructions can be used to implement the methods described herein. The code modules 510 (e.g., instructions stored in a memory) can be implemented within or external to a processor. As used herein, the term "memory" refers to a class of long-term, short-term, volatile, non-volatile, or other storage media and is not limited to any particular type of memory or memories or the type of medium on which memory is stored.

[0087] In addition, the term "storage medium" or "storage device" can refer to one or more memories for storing data, including read only memory (ROM), RAM, magnetic RAM, core memory, disk storage media, optical storage media, flash devices, and / or other machine-readable media for storing information. The term "machine-readable medium" includes, but is not limited to, portable or fixed storage devices, optical storage devices, wireless channels, and / or various other storage media capable of storing instructions and / or data.

[0088] In addition, embodiments can be implemented by hardware, software, scripting languages, firmware, middleware, microcode, hardware description languages, and / or any combination thereof. When implemented in software, firmware, middleware, scripting languages, and / or microcode, the program code or code segments for performing tasks can be stored in a machine-readable medium such as a storage medium. A code segment (e.g., code module 510) or machine-executable instructions can represent a process, function, subroutine, program, routine, subroutine, module, software package, script, class, or a combination of instructions, data structures, and / or program statements. By passing and / or receiving information, data, arguments, parameters, and / or memory contents, a code segment can be coupled to another code segment or hardware circuit. Information, variables, parameters, data, etc. can be passed, forwarded, or transmitted in a suitable manner, including memory sharing, message passing, token passing, network transmission, etc.

[0089] The implementation of the above technologies, blocks, steps, and devices can be accomplished in various ways. For example, these technologies, blocks, steps, and devices can be implemented in hardware, software, or a combination thereof. For hardware implementation, the processing unit can be implemented within one or more ASICs, DSPs, DSPDs, PLDs, FPGAs, processors, controllers, microcontrollers, microprocessors, other electronic units designed to perform the above functions, and / or a combination of the above.

[0090] Each code module 510 can include a set of instructions (code) implemented on a computer-readable medium that directs the processor of the computing device 500 to perform corresponding actions. The instructions can be used to run sequentially, in parallel (e.g., under different processing threads), or in a combination thereof. After loading the code module 510 on a general-purpose computer system, the general-purpose computer is transformed into a special-purpose computer system.

[0091] A computer program incorporating various features described herein (e.g., in one or more code modules 510) can be encoded and stored on various computer-readable storage media. The computer-readable medium encoded with the program code can be packaged together with a compatible electronic device, or the program code can be provided separately from the electronic device (e.g., downloaded via the Internet or as a separately packaged computer-readable storage medium). The storage subsystem 504 can also store information useful for establishing a network connection using the communication interface 508.

[0092] The user interface 506 may include input devices (e.g., touchpad, touch screen, scroll wheel, click wheel, dial, button, switch, keypad, microphone, etc.), output devices (e.g., video screen, indicator light, speaker, headphone jack, virtual or augmented reality display, etc.), and support electronic devices (e.g., digital-to-analog or analog-to-digital converters, signal processors, etc.). A user may operate the input devices of the user interface 506 to invoke the functions of the computing device 500, and may view and / or hear the output from the computing device 500 through the output devices of the user interface 506. For some embodiments (e.g., for processes using ASICs), there may be no user interface 506.

[0093] The processing subsystem 502 may be implemented as one or more processors (e.g., integrated circuits, one or more single-core or multi-core microprocessors, microcontrollers, central processing units, graphics processing units, etc.). In operation, the processing subsystem 502 may control the operation of the computing device 500. In some embodiments, the processing subsystem 502 may execute various programs in response to program code and may maintain multiple simultaneously executing programs or processes. At a given time, some or all of the program code to be executed may reside in the processing subsystem 502 and / or a storage medium (e.g., the storage subsystem 504). By programming, the processing subsystem 502 may provide various functions for the computing device 500. The processing subsystem 502 may also execute other programs (including programs that may be stored in the storage subsystem 504) to control other functions of the computing device 500.

[0094] The communication interface 508 may provide voice and / or data communication capabilities for the computing device 500. In some embodiments, the communication interface 508 may include radio frequency (RF) transceiver components for accessing wireless data networks (e.g., Wi-Fi networks, 3G, 4G / LTE, etc. mobile communication technologies), components for short-range wireless communication (e.g., using Bluetooth communication standards, NFC, etc.), other components, or combinations of technologies. In some embodiments, in addition to or instead of a wireless interface, the communication interface 508 may provide a wired connection (e.g., universal serial bus, Ethernet, universal asynchronous receiver / transmitter, etc.). The communication interface 508 may be implemented using a combination of hardware (e.g., driver circuits, antennas, modulators / demodulators, encoders / decoders, and other analog and / or digital signal processing circuits) and software components. In some embodiments, the communication interface 508 may support multiple communication channels simultaneously. In some embodiments, the communication interface 508 is not used.

[0095] It should be understood that the computing device 500 is illustrative and can be subject to variations and modifications. The computing device can have various functions not specifically described (e.g., voice communication via a cellular telephone network) and can include components suitable for such functions.

[0096] In addition, although the computing device 500 is described with reference to specific boxes, it should be understood that these boxes are defined for convenience of description and are not intended to imply a particular physical arrangement of component parts. For example, the processing subsystem 502, the storage subsystem, the user interface 506, and / or the communication interface 508 can be in one device or distributed among multiple devices.

[0097] Furthermore, the boxes do not need to correspond to physically distinct components. For example, by programming a processor or providing appropriate control circuitry, the boxes can be configured to perform various operations, and depending on how the initial configuration is obtained, the various boxes can be reconfigurable or non-reconfigurable. Embodiments can be implemented in a variety of devices, including electronic devices implemented using a combination of circuitry and software. The electronic devices described herein can be implemented using the computing device 500.

[0098] The various features described herein, such as methods, apparatuses, computer-readable media, etc., can be implemented using a combination of dedicated components, programmable processors, and / or other programmable devices. The processes described herein can be implemented on the same processor or different processors. In cases where a component is described as being for performing certain operations, such a configuration can be implemented, for example, by designing electronic circuitry to perform the operations, by programming a programmable electronic circuit (such as a microprocessor) to perform the operations, or a combination of the above. In addition, although the above embodiments refer to specific hardware and software components, those skilled in the art will understand that different combinations of hardware and / or software components can also be used, and that specific operations described as being implemented in hardware can be implemented in software and vice versa.

[0099] Specific details are given in the above description to provide an understanding of the embodiments. However, it should be understood that the embodiments can be practiced without these specific details. In some cases, well-known circuits, processes, algorithms, structures, and techniques are shown without unnecessary details to avoid obscuring the embodiments.

[0100] Although the principles of the present disclosure have been described above in connection with specific apparatuses and methods, it should be understood that this description is only by way of example and not a limitation on the scope of the present disclosure. The embodiments are selected and described to explain the principles of the invention and its practical applications, so that other technicians in the art can utilize the invention in various embodiments and make various modifications to suit the specific purposes contemplated. It should be understood that this description is intended to cover modifications and equivalents.

[0101] In addition, it should be noted that embodiments may be described as processes, which are depicted as flowcharts, data flow diagrams, structural diagrams, or block diagrams. Although a flowchart may describe operations as a sequential process, many operations may be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process terminates when its operations are complete, but there may be other steps not included in the figure. A process may correspond to a method, function, procedure, subroutine, subprogram, etc.

[0102] References to "a", "an", or "the" are intended to mean "one or more" unless specifically stated to the contrary. Conditional language used herein, such as "may", "might", "for example", etc., unless otherwise expressly stated or otherwise understood in the context in which it is used, generally is intended to convey that certain examples include while other examples do not include certain features, elements, and / or steps. Thus, such conditional language generally does not imply that one or more examples in any way require the aforementioned features, elements, and / or steps, or that one or more examples must include means for deciding, with or without author input or prompting, whether or not these features, elements, and / or steps are included or will be performed in any particular example. The terms "comprising", "including", "having", etc. are synonyms and are used inclusively in an open-ended manner and do not exclude additional elements, features, acts, operations, etc. In addition, the term "or" is used in its inclusive (rather than exclusive) sense, e.g., when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. The use of "adapted to" or "configured to" herein refers to open and inclusive language and does not exclude devices adapted to or configured to perform additional tasks or steps. In addition, the use of "based on" is open and inclusive because a process, step, calculation, or other action "based on" one or more recited conditions or values may in fact be based on additional conditions or values other than those recited. Similarly, the use of "at least partially based on" is open and inclusive because a process, step, calculation, or other action "at least partially based on" one or more recited conditions or values may in practice be based on additional conditions or values other than those recited. Headings, lists, and numbers included herein are for ease of explanation only and do not imply limitation.

Claims

1. A method for positioning and mapping, the method comprises: acquiring a 2D image of a 3D space, wherein the 2D image includes a plurality of feature points; selecting a first subset of feature points that are evenly distributed in the 2D image from the plurality of feature points; determining the coordinates of the feature points in the first subset of feature points from a point cloud, wherein the point cloud includes previously determined coordinates that define the positions of at least the feature points in the first subset of feature points in a global coordinate map of the 3D space; selecting a second subset of feature points that are evenly distributed in the 3D space from the first subset of feature points, at least partially based on the determined coordinates; optimizing the coordinates of the feature points in the second subset of feature points, at least partially based on the pose of the camera that acquired the 2D image, wherein the pose of the camera is determined based on data received from an inertial measurement unit (IMU) that is used to track the pose of the camera in the 3D space; and updating the coordinates of the feature points in the second subset of feature points in the point cloud using the optimized coordinates of the feature points in the second subset of feature points.

2. The method according to claim 1, wherein, the plurality of feature points are a first set of feature points, and the method further comprises: detecting a second set of feature points in the 2D image; and selecting a third subset of feature points from the second set of feature points such that the combination of the first subset of feature points and the third subset of feature points is evenly distributed in the 2D image.

3. The method according to claim 2, wherein, the third subset of feature points is selected such that the number of feature points in the combination of the first subset of feature points and the third subset of feature points is within a predetermined range.

4. The method according to claim 2, further comprises: tracking the third subset of feature points in one or more subsequently acquired 2D images; determining the coordinates of the feature points in the third subset of feature points using triangulation; and adding the third subset of feature points and the coordinates of the feature points in the third subset of feature points to the point cloud.

5. The method according to claim 4, further comprises: acquiring another 2D image of the 3D space, wherein the another 2D image includes a part of the second subset of feature points and / or a part of the third subset of feature points; selecting a fourth subset of feature points that are evenly distributed in the another 2D image from the part of the second subset of feature points and / or the part of the third subset of feature points; determining the coordinates of the feature points in the fourth subset of feature points from the point cloud; selecting a fifth subset of feature points from the fourth subset of feature points such that the fifth subset of feature points is evenly distributed in the 3D space; and optimizing the coordinates of the feature points in the fifth subset of feature points, at least partially based on the pose of the camera that acquired the another 2D image.

6. The method according to claim 1, wherein, Detecting the plurality of feature points based on one or more neighboring pixels and surrounding pixels' color and / or color intensity representing each feature point among the plurality of feature points in one or more previously acquired 2D images.

7. The method according to claim 1, further comprising: Dividing the 2D image into a plurality of regions of the same size, wherein the first subset of feature points being evenly distributed in the 2D image means that the feature points in the first subset of feature points are distributed in different regions among the plurality of regions.

8. The method according to claim 1, further comprising: Dividing the space in the global coordinate map corresponding to the 3D space into a plurality of blocks of the same size, wherein the second subset of feature points being evenly distributed in the 3D space means that the feature points in the second subset of feature points are distributed in different blocks among the plurality of blocks.

9. The method according to claim 1, wherein selecting the first subset of feature points that are evenly distributed in the 2D image from the plurality of feature points comprises: removing feature points with a velocity or moving speed higher than a predetermined threshold from the plurality of feature points, wherein the predetermined threshold is a predetermined speed or a predetermined percentage difference in speed between different feature points.

10. An electronic device for managing the feature space distribution for positioning and mapping, the electronic device comprising: a camera; one or more processors; and a memory having a program that, when executed by the one or more processors, causes the electronic device to perform the following operations: acquiring a two-dimensional (2D) image of a three-dimensional (3D) space using the camera, wherein the 2D image includes a plurality of feature points; selecting a first subset of feature points that are evenly distributed in the 2D image from the plurality of feature points; determining coordinates of the feature points in the first subset of feature points from a point cloud, wherein the point cloud includes previously determined coordinates that define at least the positions of the feature points in the first subset of feature points in a global coordinate map of the 3D space; selecting a second subset of feature points that are evenly distributed in the 3D space from the first subset of feature points at least partially based on the determined coordinates; optimizing the coordinates of the feature points in the second subset of feature points at least partially based on the pose of the camera that acquired the 2D image, wherein the pose of the camera is determined based on data received from a sensor unit of the electronic device, the sensor unit including an inertial measurement unit (IMU) for tracking the pose of the camera in the 3D space; and updating the coordinates of the feature points in the second subset of feature points in the point cloud using the optimized coordinates of the feature points in the second subset of feature points.

11. The electronic device according to claim 10, wherein the plurality of feature points are a first set of feature points, and when executed by the one or more processors, the program further causes the electronic device to perform the following operations: detecting a second set of feature points in the 2D image; and Select a third subset of feature points from the second set of feature points such that the combination of the first subset of feature points and the third subset of feature points is evenly distributed in the 2D image.

12. The electronic device according to claim 11, wherein when executed by the one or more processors, the program further causes the electronic device to perform the following operations: track the third subset of feature points in one or more subsequently acquired 2D images; determine the coordinates of the feature points in the third subset of feature points using triangulation; and add the third subset of feature points and the coordinates of the feature points in the third subset of feature points to the point cloud.

13. The electronic device according to claim 10, wherein when executed by the one or more processors, the program further causes the electronic device to perform the following operations: divide the 2D image into a plurality of regions of the same size, wherein the fact that the first subset of feature points is evenly distributed in the 2D image means that the feature points in the first subset of feature points are distributed in different regions among the plurality of regions.

14. The electronic device according to claim 10, wherein when executed by the above one or more processors, the program further causes the electronic device to perform the following operations: divide the space in the global coordinate map corresponding to the 3D space into a plurality of blocks of the same size, wherein the fact that the second subset of feature points is evenly distributed in the 3D space means that the points in the second subset of feature points are distributed in different blocks among the plurality of blocks.

15. A non-transitory machine-readable medium having instructions, wherein the instructions are executable by one or more processors to cause an electronic device having the one or more processors to perform the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Refined three-dimensional reconstruction method of cultural relic sequence image based on triangulation net interpolation and restriction

    CN107767440A

  • Improved ICP point cloud rapid splicing method and device, electronic equipment and storage medium

    CN110175954A