Feature map generation method and device, storage medium and computer device

CN117576494BActive Publication Date: 2026-09-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210945938.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-08
Publication Date
2026-09-25
Estimated Expiration
2042-08-08

AI Technical Summary

Technical Problem

[0003]随着自动驾驶等应用越来越广泛,对定位的精度要求越来越高,然而相关技术中构建的特征地图在使用过程中经常存在定位精度低的问题

Benefits of technology

[0010]上述特征地图生成方法、装置、计算机设备、存储介质和计算机程序产品,通过获得针对目标场景拍摄得到的多帧图像,从各帧图像上分别提取图像特征点,基于提取的图像特征点在所属图像的位置确定对应的特征描述子,将各帧图像的图像特征点中具有匹配关系的图像特征点,组成特征点集合,从特征点集合中确定代表特征点,计算特征点集合中剩余的图像特征点对应的特征描述子与代表特征点对应的特征描述子之间的差异,基于计算得到的差异确定特征点集合的位置误差,基于位置误差迭代更新特征点集合中剩余的图像特征点,当满足迭代停止条件时,得到更新后的特征点集合;基于更新后的特征点集合中各个图像特征点在所属图像的位置,确定更新后的特征点集合对应的空间特征点,基于空间特征点生成特征地图,由于在生成特征地图的过程中,基于图像特征点的特征描述子对图像特征点进行了位置优化,可以使得生成的特征地图更加鲁棒,从而利用该特征地图在定位的过程中定位精度得到了大大提升。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117576494B_ABST
    Figure CN117576494B_ABST
Patent Text Reader

Abstract

The application relates to a feature map generation method and device, a storage medium and a computer device, which can be applied to the field of maps or the field of automatic driving, and comprises the following steps: obtaining multiple frames of images, extracting image feature points from each frame of image, determining corresponding feature descriptors based on the extracted image feature points; image feature points with a matching relationship in the image feature points of each frame of image are combined to form a feature point set; a representative feature point is determined from the feature point set, the difference between the feature descriptors corresponding to the remaining image feature points in the feature point set and the feature descriptor corresponding to the representative feature point is calculated; the position error is determined based on the calculated difference, the remaining image feature points in the feature point set are iteratively updated based on the position error, and the updated feature point set is obtained when the iteration stopping condition is met; the spatial feature points are determined based on the updated feature point set to generate a feature map for positioning. The method can improve the positioning accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, computer device, storage medium, and computer program product for generating feature maps. Background Technology

[0002] With the development of computer technology, visual positioning technology has emerged. In visual positioning technology, feature maps can be constructed. A feature map is a data structure that uses relevant geometric features (such as points, lines, and surfaces) to represent the observation environment, thereby assisting the moving device to be positioned. For example, in autonomous driving, feature maps can be constructed to locate autonomous vehicles.

[0003] As applications such as autonomous driving become more widespread, the requirements for positioning accuracy are becoming increasingly stringent. However, the feature maps constructed using related technologies often suffer from low positioning accuracy during use. Summary of the Invention

[0004] Therefore, it is necessary to provide a feature map generation method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve positioning accuracy in response to the above-mentioned technical problems.

[0005] On one hand, this application provides a feature map generation method. The method includes: obtaining multiple frames of images captured for a target scene; extracting image feature points from each frame of images; determining corresponding feature descriptors based on the positions of the extracted image feature points in their respective images; forming a feature point set by combining image feature points with matching relationships from the image feature points of each frame of images; determining a representative feature point from the feature point set; calculating the difference between the feature descriptors corresponding to the remaining image feature points in the feature point set and the feature descriptors corresponding to the representative feature point; determining the positional error of the feature point set based on the calculated difference; iteratively updating the remaining image feature points in the feature point set based on the positional error; obtaining an updated feature point set when an iteration stopping condition is met; determining spatial feature points corresponding to the updated feature point set based on the positions of each image feature point in the updated feature point set in its respective image; and generating a feature map based on the spatial feature points. The feature map is used to locate a moving device to be located in the target scene.

[0006] On the other hand, this application also provides a feature map generation apparatus. The apparatus includes: a feature extraction module, used to acquire multiple frames of images captured for a target scene, extract image feature points from each frame, and determine corresponding feature descriptors based on the positions of the extracted image feature points in their respective images; a feature point set determination module, used to form a feature point set by combining image feature points with matching relationships from the image feature points of each frame; a difference calculation module, used to determine a representative feature point from the feature point set, and calculate the difference between the feature descriptors corresponding to the remaining image feature points in the feature point set and the feature descriptors corresponding to the representative feature point; a position update module, used to determine the position error of the feature point set based on the calculated difference, iteratively update the remaining image feature points in the feature point set based on the position error, and obtain an updated feature point set when an iteration stopping condition is met; and a feature map generation module, used to determine spatial feature points corresponding to the updated feature point set based on the positions of each image feature point in the updated feature point set in its respective image, and generate a feature map based on the spatial feature points. The feature map is used to locate a moving device to be positioned in the target scene.

[0007] On the other hand, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the above-described feature map generation method.

[0008] On the other hand, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the above-described feature map generation method.

[0009] On the other hand, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the above-described feature map generation method.

[0010] The aforementioned feature map generation method, apparatus, computer equipment, storage medium, and computer program product acquire multiple frames of images captured for a target scene, extract image feature points from each frame, determine corresponding feature descriptors based on the positions of the extracted image feature points in their respective images, form a feature point set by combining matching image feature points from each frame, determine a representative feature point from the feature point set, calculate the difference between the feature descriptors corresponding to the remaining image feature points in the feature point set and the feature descriptors corresponding to the representative feature point, determine the positional error of the feature point set based on the calculated difference, iteratively update the remaining image feature points in the feature point set based on the positional error, and obtain an updated feature point set when the iteration stops. Based on the positions of each image feature point in the updated feature point set in its respective image, determine the spatial feature points corresponding to the updated feature point set, and generate a feature map based on the spatial feature points. Because the position of the image feature points is optimized based on the feature descriptors during the feature map generation process, the generated feature map is more robust, thereby greatly improving the positioning accuracy during the localization process. Attached Figure Description

[0011] Figure 1 This is an application environment diagram of the feature map generation method in one embodiment;

[0012] Figure 2 This is a flowchart illustrating a feature map generation method in one embodiment;

[0013] Figure 3 This is a schematic diagram illustrating the composition of a feature point set in one embodiment;

[0014] Figure 4 This is a schematic diagram of the process of generating a feature map based on spatial feature points in one embodiment;

[0015] Figure 5 This is a schematic diagram illustrating the determination of a corresponding position in an input image in one embodiment;

[0016] Figure 6 This is a schematic diagram of the feature extraction model in one embodiment;

[0017] Figure 7 This is a flowchart illustrating the location information determination steps in one embodiment;

[0018] Figure 8 This is a structural block diagram of a feature map generation device in one embodiment;

[0019] Figure 9 This is an internal structural diagram of a computer device in one embodiment;

[0020] Figure 10 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0022] The feature map generation method provided in this application can be applied to Intelligent Traffic Systems (ITS) and Intelligent Vehicle Infrastructure Cooperative Systems (IVICS). Wherein:

[0023] Intelligent Traffic Systems (ITS), also known as Intelligent Transportation Systems, effectively integrate advanced technologies (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) into transportation, service control, and vehicle manufacturing. This strengthens the connection between vehicles, roads, and users, thereby forming a comprehensive transportation system that ensures safety, improves efficiency, enhances the environment, and conserves energy.

[0024] Intelligent Vehicle Infrastructure Cooperative Systems (IVICS) are a development direction of Intelligent Transportation Systems (ITS). IVICS utilizes advanced wireless communication and next-generation Internet technologies to implement comprehensive, real-time dynamic information exchange between vehicles and infrastructure. Based on the collection and fusion of dynamic traffic information across all times and spaces, it conducts active vehicle safety control and cooperative road management, fully realizing effective collaboration between people, vehicles, and roads. This ensures traffic safety, improves traffic efficiency, and ultimately forms a safe, efficient, and environmentally friendly road traffic system.

[0025] The feature map generation method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, motion device 102 communicates with server 104 via a network. Motion device 102 refers to either an autonomously moving device or a passively moving device. Autonomously moving devices can be various vehicles, robots, etc., while passively moving devices can be, for example, terminals carried by the user and moving with the user, such as smartphones, tablets, and portable wearable devices. Motion device 102 is equipped with a camera. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Specifically, in the feature map generation stage, the camera on any motion device can capture multiple frames of images of the target scene and send them to the server. The server generates and saves a feature map based on each frame. In the positioning information determination stage, the motion device to be positioned can send inertial measurement data, velocity measurement data, and target images captured in the target scene to the server. The server can determine the positioning information of the motion device to be positioned based on this data and the saved feature map, and then send the positioning information to the motion device to be positioned.

[0026] It is understood that in other embodiments, when any moving device moves in the target scene, its imaging device can capture multiple frames of images of the target scene, then generate and save a feature map based on each frame. Thus, when the moving device moves again in the target scene, its positioning information can be determined based on the saved feature map. Simultaneously, the feature map generated by the moving device can also be sent to a server. When other moving devices to be positioned move in the target scene, they can download the feature map and determine their positioning information based on it. Alternatively, when other moving devices to be positioned move in the scene, they can send inertial measurement data, velocity measurement data, and target images captured in the scene to the server. The server can then determine the positioning information of the moving device to be positioned based on this data and the saved feature map, and return the positioning information to the moving device to be positioned.

[0027] In one embodiment, such as Figure 2 As shown, a feature map generation method is provided, which can be applied to... Figure 1 Taking the server in the example of this example, it can be understood that this feature map generation method can also be applied to... Figure 1 The method refers to a system consisting of motion equipment, or a system comprised of a server and motion equipment, where the interaction between the server and motion equipment is implemented. Specifically, the method includes the following steps:

[0028] Step 202: Obtain multiple frames of images captured for the target scene, extract image feature points from each frame, and determine the corresponding feature descriptor based on the position of the extracted image feature points in their respective images.

[0029] In this context, the target scene refers to the specific scene for which the generated feature map is intended. The target scene can be the environment in which the vehicle is located, for example, the scene determined by the vehicle's possible driving route. During the vehicle's journey through this scene, multiple frames of surrounding images are captured by a camera. Image feature points are pixels in the image that can be used to describe features of the scene, such as significant edge points, histogram of oriented gradients (HOR), and Haar features. Feature descriptors correspond one-to-one with image feature points. A feature descriptor is a representation of the statistical results of the Gaussian image gradient in the neighborhood of a feature point, and it can be used to describe the corresponding image feature point.

[0030] Specifically, the motion device can capture multiple frames of images and transmit them to the server for processing in real time, or the motion device can simply store the captured multiple frames of images and input them to the server for processing in some way after image acquisition is complete. After obtaining multiple frames of images captured for the target scene, the server can extract image feature points from each frame. Based on the location of the extracted image feature points in their respective images, it determines the corresponding feature descriptors, thereby obtaining the image feature points of each frame and the feature descriptors of each image feature point.

[0031] In one embodiment, image feature point extraction can be achieved using, but is not limited to, algorithms such as Good Features to Track, for which corresponding functions are provided in the computer vision library OpenCV. In other embodiments, image feature point extraction can also be performed by training a machine learning model. This machine learning model includes multiple convolutional layers, each processing the original image differently and outputting a feature image. This feature image represents the probability that each location in the original image is a feature point, and thus, the original feature points can be determined based on the feature image. It is understood that multiple image feature points can be extracted from each frame of the image. "Multiple" refers to at least two.

[0032] Step 204: Collect image feature points that have matching relationships from the image feature points of each frame image and form a feature point set.

[0033] In this context, matching image feature points refer to similar image feature points. In one embodiment, matching image feature points can be determined by their feature descriptors; when the feature descriptors of two image feature points reach a certain degree of similarity, they are considered to be matched.

[0034] Specifically, the server can group matching image feature points from each frame into feature point sets, thus obtaining multiple feature point sets. For example, ... Figure 3 As shown, assuming there are a total of 3 frames of images, the first frame includes image feature points A1, A2, and A3, the second frame includes image feature points B1, B2, B3, and B4, and the third frame includes image feature points C1, C2, and C3. Assuming that A1, B1, and C1 are image feature points that have a matching relationship with each other, A2, B2, and C2 are image feature points that have a matching relationship with each other, and A3, B3, and C3 are image feature points that have a matching relationship with each other, then A1, B1, and C1 can form feature point set 1, A2, B2, and C2 can form feature point set 2, and A3, B3, and C3 can form feature point set 3.

[0035] In one embodiment, assuming there are M frames in total, with i ranging from 1 to M, firstly, N image feature points are extracted from the i-th frame. The method for extracting image feature points can be referred to above. If i = 1, that is, for the first frame, a corresponding feature point set can be created for each image feature point. If i > 1, with j ranging from 1 to N, it is determined whether there is an image feature point in the (i-1)-th frame that matches the j-th image feature point in the i-th frame. If a matching image feature point exists, the j-th image feature point is added to the feature point set corresponding to the matching image feature point (since the (i-1)-th frame has been processed, this feature point set must already exist); if no matching image feature point exists, a feature point set corresponding to the j-th image feature point is created. Once no new image feature points are added to a feature point set, the feature point set can be considered to be completed. By processing the image frame by frame using the above method, after the M-th frame is processed, multiple feature point sets are obtained. Each feature point set includes at least one image feature point, or includes a sequence of image feature points that have a matching relationship with each other. It's understandable. However, in practical applications, if feature maps are being built in real time, M may not be known. But the specific steps are similar to those above. You just need to keep incrementing i until all images have been processed.

[0036] Step 206: Determine the representative feature point from the feature point set, and calculate the difference between the feature descriptors corresponding to the remaining image feature points in the feature point set and the feature descriptors corresponding to the representative feature point.

[0037] Here, the representative feature point set refers to the image feature points in the feature point set that can represent the entire feature point set. The remaining image feature points in the feature point set refer to the image feature points in the feature point set other than the representative feature point. For example, suppose a feature point set includes four image feature points: A1, B1, C1, and D1, where A1 is the representative feature point, then B1, C1, and D1 are the remaining image feature points. In one embodiment, the server can randomly select an image feature point from each feature point set as the representative feature point for each set. In other embodiments, the server can calculate the average feature point for each feature point set and determine the image feature point in each set that is closest to its average feature point as the representative feature point.

[0038] Specifically, to prevent overall shifts during iterative updates of image feature points in the feature point set, in this embodiment, a representative feature point can be determined in each feature point set. During the iterative update process, the position of the representative feature point is kept fixed, thereby calculating the difference between the feature descriptors corresponding to the remaining image feature points in each feature point set and the feature descriptors corresponding to the representative feature points of each feature point set, and obtaining the differences corresponding to each remaining image feature point.

[0039] In one embodiment, for each set of feature points, the server can calculate the absolute difference between the feature descriptor corresponding to each remaining image feature point in the set and the feature descriptor corresponding to the entire set of feature points, thus obtaining the difference for each remaining image feature point. In other embodiments, after calculating the absolute difference, the server can calculate the square of the absolute difference to obtain the difference for each remaining image feature point.

[0040] Step 208: Determine the position error of the feature point set based on the calculated difference, and iteratively update the remaining image feature points in the feature point set based on the position error. When the iteration stopping condition is met, the updated feature point set is obtained.

[0041] The iteration stopping condition can be, for example, one of the following: the position error reaches the minimum value, the number of iterations reaches the preset number, or the iteration time reaches the minimum duration.

[0042] Specifically, since each feature point set determines a spatial feature point when determining image feature points, to improve the accuracy of the determined spatial feature points, it is necessary to reduce the overall positional error of the feature point set. Therefore, in this embodiment, for each feature point set, the server can statistically analyze the differences between the remaining image feature points in that set. Based on the statistically obtained differences, the positional error of the feature point set is determined. The positions of all image feature points other than the representative feature point are iteratively updated in the direction that minimizes this positional error. Each update is equivalent to optimizing the position of the image feature points. Based on the descriptors corresponding to the optimized image feature points, the positional error is recalculated, and the next update is performed. This step is repeated multiple times to optimize the position of the image feature points. When the iteration stopping condition is met, the updated image feature points and the representative feature point belonging to the same feature point set are combined to form an updated feature point set. During the update process, a gradient descent algorithm can be used to update the positions of the image feature points.

[0043] In one embodiment, to prevent degradation during the optimization process, the server can calculate the singular values ​​of the Hessian matrix. If the largest singular value divided by the smallest singular value is greater than a preset threshold, the update is abandoned.

[0044] Step 210: Based on the position of each image feature point in the updated feature point set in its respective image, determine the spatial feature points corresponding to the updated feature point set, and generate a feature map based on the spatial feature points. The feature map is used to locate the moving device to be located in the target scene.

[0045] Spatial feature points refer to three-dimensional feature points, that is, the corresponding points of feature points on the image in three-dimensional space. In this embodiment, the feature map can be a data structure including multiple spatial feature points, and its specific form is not limited. The moving device to be located refers to the moving device that needs to be located. The moving device to be located and the moving device that sends multiple frames of images can be the same moving device or different moving devices. The pose of the image to which the image feature points belong refers to the pose of the camera when that frame of image is captured. This pose can be obtained through pose transformation based on the pose of the moving device at the same moment and the relative pose relationship between the camera and the moving device.

[0046] Specifically, for each updated feature point set, the server can perform triangulation calculations based on the position and pose of each image feature point in its respective image, obtaining the spatial feature points corresponding to each feature point set. Further, the server can generate a feature map based on these spatial feature points, store the feature map, and then use this feature map to assist in the positioning of the moving device in subsequent positioning processes. Triangulation calculation is an existing method for mapping two-dimensional image feature points to three-dimensional spatial feature points, which will not be elaborated on here. It is understood that the descriptor of a spatial feature point can be the average of the descriptors of all image feature points that generated that spatial feature point.

[0047] In one embodiment, the server determines the pose of the image to which the image feature points belong through the following steps: First, the relative pose between the moving device and the camera is obtained. This relative pose typically remains unchanged during the movement of the moving device and can be obtained through calibration. Then, the pose of the moving device at each moment is determined based on the inertial measurement data and velocity measurement data uploaded by the moving device. Next, the pose of the moving device at each moment is aligned with the acquisition time of multiple frames of images. This alignment refers to determining the pose of the moving device corresponding to each frame of image, where the data acquisition time (the time when inertial measurement data and velocity measurement data are acquired) is the same as (or within the allowable error range) the acquisition time of that frame of image. Finally, an attitude transformation is performed based on the pose of the moving device corresponding to each frame of image and the relative pose between the moving device and the camera to obtain the pose of that frame of image.

[0048] In the aforementioned feature map generation method, multiple frames of images captured for the target scene are obtained. Image feature points are extracted from each frame. Based on the location of the extracted image feature points in their respective images, corresponding feature descriptors are determined. Image feature points with matching relationships in each frame are grouped into a feature point set. A representative feature point is determined from the feature point set. The difference between the feature descriptors corresponding to the remaining image feature points in the feature point set and the feature descriptors corresponding to the representative feature point is calculated. The positional error of the feature point set is determined based on the calculated difference. The remaining image feature points in the feature point set are iteratively updated based on the positional error. When the iteration stops, the updated feature point set is obtained. Based on the location of each image feature point in the updated feature point set in its respective image, the spatial feature points corresponding to the updated feature point set are determined. A feature map is generated based on the spatial feature points. Since the image feature points are optimized for position based on the feature descriptors during the feature map generation process, the generated feature map is more robust, thereby greatly improving the positioning accuracy during the localization process.

[0049] In one embodiment, determining the position error of the feature point set based on the calculated difference includes: taking each remaining image feature point in the feature point set as a target feature point, calculating the matching confidence between each target feature point and the representative feature point; calculating the position error of each target feature point based on the matching confidence and difference corresponding to each target feature point; and statistically analyzing the position error of each target feature point to obtain the position error of the feature point set.

[0050] Among them, the matching confidence between the target feature point and the representative feature point is used to characterize the degree of matching between the target feature point and the representative feature point. The higher the degree of matching, the more similar the two feature points are.

[0051] Specifically, for each set of feature points, the server can use each remaining image feature point in the set as a target feature point. For each target feature point, the server can calculate the matching confidence between the target feature point and a representative feature point, and then multiply the matching confidence by the difference to obtain the position error corresponding to the target feature point. Finally, the server calculates the position error of each target feature point to obtain the position error of the entire set of feature points. The calculation can be one of the following: summation, averaging, or median calculation.

[0052] In one specific embodiment, the server can calculate the position error using the following formula (1), where, Let u and v represent image feature points, i(u) represent the u-th image feature point in the i-th frame, and k(v) represent the v-th feature point in the k-th frame. w uv To match the confidence level, p u p represents the position of image feature point u on the image. v F represents the position of image feature point v on the image. i(u) [p u ] represents p u The descriptor, F k(v) [p v ] represents p v Descriptors:

[0053]

[0054] In this embodiment, by calculating the matching confidence of image feature points, the positional error of image feature points is obtained based on the matching confidence and the difference. This makes the positional error of each image feature point more accurate. As a result, the positional error of the feature point set obtained by statistically analyzing the positional errors of each image feature point in the feature point set is more accurate, thus obtaining a more accurate feature map and further improving the positioning accuracy.

[0055] In one embodiment, calculating the matching confidence between each target feature point and the representative feature point includes: obtaining the feature descriptor of each target feature point and the feature descriptor of the representative feature point; calculating the vector similarity between the feature descriptor of each target feature point and the feature descriptor of the representative feature point, and using each vector similarity as the matching confidence between the corresponding target feature point and the representative feature point.

[0056] Vector similarity describes the degree of similarity between two vectors. Since feature descriptors are in vector form, vector similarity can be calculated. In one embodiment, vector similarity can be, for example, cosine similarity.

[0057] Specifically, the server can obtain the feature descriptors for each target feature point and the feature descriptor for the representative feature point. Then, it calculates the vector similarity between the feature descriptors of each target feature point and the feature descriptor of the representative feature point, using each vector similarity as the matching confidence score for its corresponding target feature point. For example, assuming a feature point set includes image feature points A1, B1, and C1, where C1 is the representative feature point, the server can obtain the feature descriptors for A1, B1, and C1 respectively, calculate the vector similarity between the feature descriptor of image feature point A1 and the representative feature point C1, and calculate the vector similarity between the feature descriptor of image feature point B1 and the representative feature point C1, using this as the matching confidence score between image feature point A1 and the representative feature point C1.

[0058] In the above embodiments, the matching confidence is obtained by calculating the vector similarity between feature descriptors. Since feature descriptors describe image feature points, the obtained matching confidence is more accurate.

[0059] In one embodiment, determining representative feature points from a set of feature points includes: calculating the average feature point position corresponding to the feature point set based on the position of each image feature point in the feature point set in its respective image; determining image feature points whose distance from the average feature point position satisfies a distance condition from the feature point set; and using the determined image feature points as representative feature points.

[0060] The distance condition includes one of the following: the distance between the point and the average feature point is less than or equal to a distance threshold, or the point is sorted before a sorting threshold when arranged in ascending order of distance from the average feature point.

[0061] Specifically, for each feature point set, the server can obtain the position of each image feature point in the set within its corresponding image. The position values ​​in the same dimension are summed and averaged to obtain the target value for that dimension. The target values ​​for each dimension determine the average feature point position corresponding to the feature point set. For example, suppose a feature point set includes image feature points A1, B1, and C1, where A1 is located at (x1, y1), B1 at (x2, y2), and C1 at (x3, y3). Then the average feature point position corresponding to this feature point set is ((x1+x2+x3) / 3, (y1+y2+y3) / 3).

[0062] For each feature point set, after calculating the average feature point position corresponding to the feature point set, the server can calculate the distance between the position of each image feature point in the feature point set and the average feature point position. Based on the calculated distance, image feature points that meet the distance condition are selected, and the selected image feature points are determined as representative feature points.

[0063] In a specific embodiment, the distance condition includes the distance between the image feature point and the average feature point position being less than or equal to a distance threshold. After calculating the distance between each image feature point and the average feature point position corresponding to its respective feature point set, the server compares each distance with the distance threshold. If only one image feature point has a distance less than the distance threshold between it and the average feature point position corresponding to its respective feature point set, then that image feature point is determined as the representative feature point. If multiple image feature points have a distance less than the distance threshold between them and the average feature point position corresponding to their respective feature point sets, then one of these image feature points can be selected as the representative feature point. For example, the image feature point with the smallest distance can be selected as the representative feature point.

[0064] In another specific embodiment, the distance condition includes sorting the image feature points in ascending order based on their distance from the average feature point position, prior to a sorting threshold. After calculating the distance between each image feature point and the average feature point position corresponding to its set, the server can sort the image feature points in ascending order based on their distance and select a representative feature point from the image feature points sorted before the sorting threshold. For example, if the sorting threshold is 2, the image feature point sorted first can be selected as the representative feature point.

[0065] In the above embodiments, based on the position of each image feature point in the feature point set in the corresponding image, the average feature point position corresponding to the feature point set is calculated, and image feature points whose distance from the average feature point position meets the distance condition are determined from the feature point set. The determined image feature points are used as representative feature points, and the determined representative feature points can better reflect the overall positional characteristics of the feature point set.

[0066] In one embodiment, the feature point set includes multiple features; determining representative feature points from the feature point set includes: for each feature point set, filtering the feature point set if the feature point set meets the filtering conditions; and proceeding to the step of determining representative feature points from the feature point set if the feature point set does not meet the filtering conditions.

[0067] In this embodiment, the filtering conditions include at least one of the following: the distance between the initial spatial feature points calculated based on the feature point set and the camera capturing the multi-frame images is greater than a first preset distance threshold; the distance between the initial spatial feature points calculated based on the feature point set and the camera capturing the multi-frame images is less than a second preset distance threshold, and the second preset distance threshold is less than the first preset distance threshold; the disparity calculated based on the feature point set is greater than a preset disparity threshold; and the average reprojection error calculated based on the feature point set is greater than a preset error threshold.

[0068] Here, the initial spatial feature points refer to the spatial feature points determined based on the positions of each image feature point in the unupdated feature point set within its respective image. Filtering the feature point set involves removing that feature point from multiple feature point sets.

[0069] Specifically, the server can calculate the distance between the initial spatial feature points and the shooting device of the multi-frame images. If the distance is greater than the first preset distance threshold, that is, the spatial feature points are too far from the shooting device, the set of feature points is filtered out. If the distance is less than the second preset distance threshold, that is, the spatial feature points are too close to the shooting device, the set of feature points is filtered out. The second preset distance threshold is less than the first preset distance threshold.

[0070] Furthermore, the server can also perform disparity calculation based on the feature point set. If the calculated disparity is greater than a preset disparity threshold, the feature point set will be filtered out.

[0071] Furthermore, the server can also project the initial spatial feature points calculated based on the feature point set onto the images to which each image feature point in the feature point set belongs, calculate the distance between each image feature point and the projected feature points projected onto their respective images, obtain each projection distance, and then calculate the average value of the projection distances to obtain the average reprojection error. If the average reprojection error is greater than a preset error threshold, the feature point set is filtered out.

[0072] For an unfiltered set of feature points, the server can proceed to the step "determine representative feature points from the feature point set" to identify representative feature points from these feature point sets. Then, the server can optimize the positions of the image feature points in these feature point sets using the method provided in the above embodiments to obtain each updated feature point set. Finally, based on the position of each image feature point in each updated feature point set in its respective image, the server can determine the spatial feature points corresponding to each updated feature point set, thereby obtaining multiple spatial feature points and generating a feature map.

[0073] In the above embodiments, by setting filtering conditions, the feature point set that meets the filtering conditions is filtered, which further improves the robustness of the feature map, thereby further improving the positioning accuracy when using the feature map for assisted positioning.

[0074] In one embodiment, such as Figure 4 As shown, generating a feature map based on spatial feature points includes:

[0075] Step 402: Based on the feature descriptors of each image feature point in the updated feature point set, determine the average descriptor corresponding to the updated feature point set.

[0076] Specifically, for each updated set of feature points, the server can calculate the average descriptor corresponding to that set of feature points using the following formula (2):

[0077]

[0078] Among them, u j Here, f is the average descriptor, j represents the j-th feature point set (the updated feature point set), and f is the descriptor of the image feature points in the j-th feature point set. This represents the feature descriptor subset corresponding to the j-th feature point set.

[0079] Step 404: From the feature descriptors of each image feature point in the updated feature point set, select the feature descriptors whose similarity to the average descriptor meets the similarity condition, and use the selected feature descriptors as reference descriptors.

[0080] The similarity condition can be one of the following: the similarity is greater than a preset similarity threshold, or the similarity is sorted before the sorting threshold when arranged in descending order of similarity.

[0081] In a specific embodiment, the similarity condition includes a similarity greater than a preset similarity threshold. Then, for each updated feature point set, after calculating the average descriptor corresponding to the feature point set, the server calculates the similarity between the feature descriptor of each image feature point in the feature point set and the average descriptor. Then, each similarity is compared with the similarity threshold. If only one image feature point has a similarity greater than the preset similarity threshold, then the feature descriptor of that image feature point is determined as the reference descriptor. If multiple image feature points have similarities greater than the preset similarity threshold, then one of the feature descriptors corresponding to these image feature points can be selected as the reference descriptor. For example, the feature descriptor with the highest similarity can be selected as the reference descriptor.

[0082] In another specific embodiment, the distance condition includes sorting the features before a sorting threshold when arranged in descending order of similarity. Then, for each updated set of feature points, after the server calculates the similarity between the feature descriptors of each image feature point in the set and the average descriptor, it can sort the feature descriptors of each image feature point in descending order of similarity and select a reference descriptor from the feature descriptors sorted before the sorting threshold. For example, if the sorting threshold is 2, the feature descriptor sorted first can be selected as the reference descriptor.

[0083] In another specific embodiment, the server can calculate the reference descriptor by referring to the following formula (3):

[0084]

[0085] Among them, f j Here, j represents the j-th feature point set (the updated feature point set), and f represents the descriptor of the image feature points in the j-th feature point set. This represents the feature descriptor subset corresponding to the j-th feature point set.

[0086] Step 406: Project the spatial feature points onto the images to which each image feature point belongs in the updated feature point set, to obtain multiple projected feature points, and determine the feature descriptor corresponding to the projected feature points based on the position of the projected feature points on their respective images.

[0087] Step 408: Based on the difference between the feature descriptor corresponding to the projected feature point and the reference descriptor, determine the reprojection error corresponding to the projected feature point.

[0088] Step 410: Calculate the reprojection error corresponding to each projection feature point to obtain the target error. Iterate and update the spatial feature points based on the target error. When the iteration stopping condition is met, obtain the target spatial feature points corresponding to the updated feature point set. Generate a feature map based on the target spatial feature points.

[0089] Specifically, for each updated feature point set, after determining its corresponding spatial feature point, the server can project the spatial feature point onto the image to which each image feature point in the feature point set belongs, obtaining multiple projected feature points corresponding to the spatial feature point. Further, based on the position of each projected feature point on its respective image, the server can determine the feature descriptor corresponding to each projected feature point. Then, the difference between each projected feature point and the reference descriptor corresponding to the updated feature point set calculated in step 404 is calculated to obtain the reprojection error corresponding to each projected feature point. Finally, the reprojection errors are statistically analyzed to obtain the target error corresponding to the updated feature point set. The server iteratively updates the spatial feature points corresponding to the updated feature point set in the direction of minimizing the target error, that is, the updated spatial feature point is used as the current spatial feature point, and the process returns to step 406. Steps 406-410 are repeated continuously until the iteration stopping condition is met. The obtained spatial feature point is the target spatial feature point, and a feature map can then be generated based on the target spatial feature point. The iteration stopping condition can be one of the following: the target error reaches its minimum value, the number of iterations reaches a preset number, or the iteration duration reaches a preset duration.

[0090] In a specific embodiment, when the server performs steps 406 to 410 above, it can refer to the following formula (4) to calculate the target error:

[0091]

[0092] in, Let Z(j) be the target error, j be the set of the j-th feature points (the updated set of feature points), Z(j) be the set of images to which each image feature point in the j-th feature point set belongs, i be the i-th frame image, and C be the target error. i P represents the camera intrinsic parameters corresponding to the i-th frame image. j It refers to the spatial feature points corresponding to the j-th feature point set, R i Let t be the rotation matrix corresponding to the i-th frame image. i Let f be the translation matrix corresponding to the i-th frame image. j This is the reference descriptor corresponding to the j-th feature point set.

[0093] In the above embodiments, by determining a reference descriptor, spatial feature points are projected onto the images to which each image feature point belongs in the updated feature point set, resulting in multiple projected feature points. Based on the position of the projected feature points on their respective images, the feature descriptor corresponding to each projected feature point is determined. Based on the difference between the feature descriptor corresponding to each projected feature point and the reference descriptor, the reprojection error corresponding to each projected feature point is determined. The reprojection error corresponding to each projected feature point is statistically analyzed to obtain the target error. Based on the target error, the spatial feature points are iteratively updated. When the iteration stopping condition is met, the target spatial feature point is obtained, thus achieving position optimization of the spatial feature points. The feature map generated based on the optimized target spatial feature points can further improve the positioning accuracy when used for localization.

[0094] In one embodiment, the multi-frame images are captured by a camera mounted on the target motion device; the feature map generation method further includes: acquiring inertial measurement data and velocity measurement data of the target motion device when capturing the multi-frame images; using the inertial measurement data and velocity measurement data to calculate the initial pose of the motion device to be located; determining pre-integration information based on the inertial measurement data; constructing a factor map based on the pre-integration information and velocity measurement data; adjusting the initial pose based on the factor map to obtain the target pose; and generating a feature map based on spatial feature points, including: establishing a correspondence between spatial feature points and the target pose; and generating a feature map based on the correspondence and spatial feature points.

[0095] In one embodiment, extracting image feature points from each frame of the image and determining the corresponding feature descriptor based on the position of the extracted image feature points in their respective images includes: inputting the image into a trained feature extraction model, and outputting a first tensor corresponding to the image feature points and a second tensor corresponding to the feature descriptor through the feature extraction model; the first tensor is used to describe the probability of feature points appearing in each region of the image; performing non-maximum suppression processing on the image based on the first tensor to determine the image feature points from the image; converting the second tensor into a third tensor with the same size as the image, and determining the vector in the third tensor that matches the position of the image feature point in its respective image as the descriptor corresponding to the image feature point.

[0096] Specifically, the server inputs the image into a trained feature extraction model. The model outputs a first tensor corresponding to the image feature points and a second tensor corresponding to the feature descriptors. Both the first and second tensors are multi-channel tensors, and the size of each channel is smaller than that of the original input image. The values ​​at each position in the first tensor describe the probability of feature points appearing in each corresponding region of the original input image. For example, assuming the image size input to the feature extraction model is H x W, the output first tensor could be H / N1 x W / N1 x X1, and the second tensor could be H / N2 x W / N2 x X2, where N1, N2, X1, and X2 are all positive integers greater than 1.

[0097] In one embodiment, when performing nonmaximum suppression processing on an image based on a first tensor, the server can first convert the first tensor into a probability map with the same size as the input image, search for local maxima in the probability map, and determine the location of the local maxima as the target location. Since the size of the probability map is the same as that of the input image, the pixel in the input image at the same location as the target location can be directly determined as the image feature point of the input image.

[0098] In another embodiment, considering that converting the first tensor into a probability map of the same size as the input image is time-consuming, the server can perform non-maximum suppression processing on the image based on the first tensor through the following steps:

[0099] 1. Obtain the maximum value of the first tensor at each position along the direction of multiple channels, and the channel index corresponding to each maximum value, to obtain the third and fourth tensors respectively.

[0100] Specifically, assuming the first tensor includes N (N is greater than or equal to 2) channels, for each pixel position in the first tensor, the server can search for the maximum value along the direction of the N channels, and use the maximum value found at each pixel position as the value at the corresponding position in the third tensor, thus obtaining the third tensor. At the same time, the channel index of the maximum value found at each pixel position is used as the value at the corresponding position in the third tensor, thus obtaining the fourth tensor.

[0101] 2. Determine the target value from the third tensor and search the neighborhood of the target value in the third tensor. The neighborhood of the target value includes multiple target locations, the corresponding location of the target location in the image, and the image distance between the target location and the corresponding location of the target value in the image is less than a preset distance threshold.

[0102] Specifically, the server can sort the values ​​in the third tensor from smallest to largest to obtain a set of values. Then, it iterates through each value in the set, checking if it's less than a preset threshold. If it is, it continues to the next value; if it's greater than the threshold, it identifies the value as the target value. This allows for a search of the neighborhood of the target value in the third tensor. Since the third tensor is smaller than the original input image, and image feature points refer to pixels in the input image, the neighborhood of the target value needs to be determined based on the corresponding pixel position of the target value in the third tensor in the original input image. In other words, if the neighborhood of the target value includes multiple target positions, the image distance between the corresponding position of each target position in the input image and the corresponding position of the target value in the image must be less than a preset distance threshold. That is, the corresponding position of each target position in the input image must fall within the neighborhood of the corresponding position of the target value in the image. For example... Figure 5 As shown, assuming the target value is located at point A, and the corresponding position of point A in the input image is point B, if... Figure 5 The dashed box in the image represents the neighborhood of point B. Then, the corresponding position of each target location of point A in the neighborhood of the third tensor falls within the dashed box in the input image.

[0103] In one embodiment, considering that the features extracted from different channels in the first tensor are different, the corresponding position of the pixel position in the third tensor in the original image is related to the channel where the pixel position is located. For the pixel position (i, j) in the third tensor, the index value is determined from the corresponding position in the fourth tensor as D[i, j], and its corresponding position in the original image is (N x i + D[i, j] / 8, N x j + D[i, j] % 8), where N is the scaling ratio of the third tensor relative to the original input image. For example, suppose the original input image is 640x480, the first tensor is 80x60x64, the second tensor is 80x60x256, the third tensor is 80x60 (each value represents the maximum value of the first tensor's 64 dimensions, decimal type), and D is 80x60 (each value represents the index corresponding to the maximum value of the first tensor's 64 dimensions, integer type). The first tensor's 64 dimensions correspond to each 8x8 region of the original image. Then, the coordinates (32, 53, 35) of the first tensor correspond to the coordinates (32x8 + 35 / 8, 53x8 + 35%8) = (260, 427) in the original image.

[0104] Therefore, the distance between corresponding positions of two pixel locations in the fourth tensor can be used as the distance between corresponding positions of those two pixel locations in the original input image. For example, for a pixel location (i, j) and another pixel location (i+n, j+n) in the third tensor, the distance between corresponding positions of these two pixel locations in the original image can be obtained by calculating the distance between pixel locations (i, j) and (i+n, j+n) in the fourth tensor.

[0105] 3. If the search result indicates that the target value is greater than the corresponding value at other locations in the neighborhood, the target pixel in the image corresponding to the location of the target value is determined as the image feature point.

[0106] In this context, the target pixel is determined from the image based on the location of the target value and its corresponding channel index. The channel index is determined from the fourth tensor based on the location of the target value. For example, suppose the pixel coordinates of a target value in the third tensor are (i, j). The corresponding position of this pixel in the fourth tensor is also (i, j). If the value at this position in the fourth tensor is D[i, j], then if the search result indicates that the target value is greater than the corresponding values ​​at other positions in the neighborhood, the pixel with coordinates (N x i + D[i, j] / 8, N x j + D[i, j] % 8) in the original input image is determined as the target pixel corresponding to the location of the target value, where N is the scaling factor of the third tensor relative to the original input image.

[0107] In a specific embodiment, the specific structure of the feature extraction model in the above embodiments can be as follows: Figure 6As shown, the first convolutional block is a 3x3 fully convolutional layer with a stride of 1 and 64 output channels; the first pooling block is a 3x3 max pooling layer with a stride of 1 and 64 output channels; the second convolutional block is a 3x3 fully convolutional layer with a stride of 2 and 64 output channels; the third convolutional block is a 3x3 fully convolutional layer with a stride of 1 and 64 output channels; and the fourth convolutional block is a 3x3 fully convolutional layer. The fifth convolutional block is a 3x3 fully convolutional layer with a stride of 2 and 64 output channels; the sixth convolutional block is a 3x3 fully convolutional layer with a stride of 1 and 64 output channels; the seventh convolutional block is a 3x3 fully convolutional layer with a stride of 2 and 128 output channels; the eighth convolutional block is a 3x3 fully convolutional layer with a stride of 1 and 128 output channels; the ninth convolutional block is a 3x3 fully convolutional layer with a stride of 1 and 128 output channels; the eleventh convolutional block is a 3x3 fully convolutional layer with a stride of 1 and 128 output channels; the twentieth convolutional block is a 3x3 fully convolutional layer with a stride of 1 and 128 output channels; the th... The ninth convolutional block is a 3x3 fully convolutional layer with a stride of 1 and 128 output channels; the tenth convolutional block is a 1x1 fully convolutional layer with a stride of 2 and 64 output channels; the eleventh convolutional block is a 1x1 fully convolutional layer with a stride of 2 and 64 output channels; the twelfth convolutional block is a 1x1 fully convolutional layer with a stride of 2 and 128 output channels; the thirteenth convolutional block is a 1x1 fully convolutional layer with a stride of st... The first convolutional block has a stride of 2 and 128 output channels; the fourteenth convolutional block is a 3x3 fully convolutional layer with a stride of 1 and 128 output channels; the fifteenth convolutional block is a 3x3 fully convolutional layer with a stride of 1 and 64 output channels; the sixteenth convolutional block is a 3x3 fully convolutional layer with a stride of 1 and 128 output channels; and the seventeenth convolutional block is a 3x3 fully convolutional layer with a stride of 1 and 256 output channels.

[0108] Assuming the input image has dimensions H x W, the output of the 15th convolutional block of the feature extraction model is a tensor A of feature points with dimensions H / 8 x W / 8 x 64, and the output on the right is a tensor B of descriptors with dimensions H / 8 x W / 8 x 256. The specific steps for extracting feature points and descriptors are as follows:

[0109] 1. Along the 64-channel dimension, obtain the maximum value and the index corresponding to the maximum value, thus obtaining two H / 8xW / 8 tensors C and D.

[0110] 2. Arrange the probability values ​​in tensor C from largest to smallest into a set E, and set a target set F to store the feature point indices and confidence scores.

[0111] 3. Traverse set E and obtain the indices i and j of the corresponding values ​​in tensor D.

[0112] 4. If C[i,j] is less than a certain threshold (e.g., 0.05), skip.

[0113] 5. Traverse the neighborhood n of C[i,j].

[0114] 6. Calculate the distance between D[i+n, j+n] (or D[in, jn]) and D[i, j], which is the distance between the coordinates (8x(i+n)+D[i+n, j+n] / 8, 8x(j+n)+D[i+n, j+n]%8) and (8x i+D[i, j] / 8, 8x j+D[i, j]%8) on the original graph. If it is greater than a certain threshold, skip it.

[0115] 7. If C[i+n, j+n] (or C[in, jn]) is greater than C[i, j], then exit the traversal in step 5; otherwise, continue executing step 5.

[0116] 8. If step 5 is completed and C[i,j] is greater than any C[i+n,j+n] (or C[in,jn]), then place C[i,j] and (ix8+D[i,j] / 8,jx8+D[i,j]%8) in the target set F.

[0117] 9. Continue with step 3.

[0118] 10. Perform bilinear interpolation on tensor B to obtain tensor G with dimensions HxWx256, and calculate the L2 norm along the channel direction.

[0119] 11. Based on the results of the target set F, find the corresponding descriptor in tensor G. That is, for the index of each image feature point in the target set F, find the position in tensor G that has the same index, and use the vector composed of the values ​​of each channel at that position as the feature descriptor of that image feature point. The feature descriptor is then a 256-dimensional vector. For example, for the index (10,13) of an image feature point in the target set F, find the position in tensor G that corresponds to (10,13), and use the vector composed of the values ​​of each channel at that position as the feature descriptor of that image feature point.

[0120] In the above embodiments, since it is not necessary to convert the first tensor into a probability map with the same size as the input image, the efficiency of image feature point extraction is improved.

[0121] In one embodiment, obtaining multiple frames of images captured for a target scene includes: obtaining multiple original frames of images of the target scene captured by a fisheye camera, performing distortion correction on the multiple original frames of images, and obtaining multiple frames of images captured for the target scene.

[0122] In this embodiment, the multi-frame images of the target scene obtained by the server are captured by a fisheye camera, and the fisheye camera imaging model is approximately a unit spherical projection model. Generally, the fisheye camera imaging process is decomposed into two steps: first, three-dimensional spatial points are linearly projected onto a virtual unit sphere; then, points on the unit sphere are projected onto the image plane. This process is non-linear. The design of the fisheye camera introduces distortion, therefore the images formed by the fisheye camera are distorted, with radial distortion being particularly severe. Therefore, its distortion model mainly considers radial distortion. The projection function of the fisheye camera is designed to project a large scene onto a finite image plane as much as possible. Based on different projection functions, the design model of the fisheye camera can be roughly divided into four types: equidistant projection model, equal solid angle projection model, orthographic projection model, and stereoscopic projection model. In this embodiment, any one of these four models can be used to correct the distortion of the multi-frame original images captured by the fisheye camera to obtain the multi-frame images of the target scene.

[0123] In the above embodiments, since the multiple frames of images are captured by a fisheye camera, which has a wider field of view than a pinhole camera, it can perceive more environmental information and extract more image feature points, thereby further improving the robustness of the generated feature map and thus improving the positioning accuracy.

[0124] In one embodiment, such as Figure 7 The diagram illustrates the process of determining location information using a feature map generated according to an embodiment of this application, including the following steps:

[0125] Step 702: Obtain the inertial measurement data, velocity measurement data, and target image captured by the motion device in the target scene, and use the inertial measurement data and velocity measurement data to determine the initial pose of the motion device to be located.

[0126] The inertial measurement data can be obtained through an Inertial Measurement Unit (IMU). The velocity measurement data can be obtained through a velocity sensor; for example, when the moving device to be located is a vehicle, the velocity measurement data can be obtained through a wheel speed sensor. Both the inertial measurement data and the velocity measurement data are measured while the moving device to be located is moving in the target scene.

[0127] Specifically, the server can receive inertial measurement data, velocity measurement data, and target images captured by the moving device to be located in the target scene, all sent by the device. Based on a preset kinematic model, the server calculates the initial pose of the moving device using the inertial measurement data and velocity measurement data. The preset kinematic model reflects the relationship between vehicle position, velocity, acceleration, and time. This embodiment does not limit the specific form of the model; in practical applications, it can be reasonably set according to requirements. For example, it can be improved from an existing single-vehicle model to obtain the desired model.

[0128] Step 704: Based on the initial pose, determine the spatial feature points that match the position from the generated feature map to obtain the target spatial feature points.

[0129] In one embodiment, the server can find spatial feature points that match the location represented by the initial pose from the feature map, and use them as target spatial feature points. In other embodiments, the feature map also includes storing the poses corresponding to each spatial feature point. The poses corresponding to the spatial feature points can be the poses of the moving device when capturing multiple frames of images during the generation of the feature map. Then, in the process of determining the positioning information, the server can compare the initial pose of the moving device to be located with the poses corresponding to each spatial feature point, and determine the spatial feature point corresponding to the pose with the highest matching degree as the target feature point.

[0130] Step 706: Determine image feature points that match the target spatial feature points from the target image, form matching pairs between the determined image feature points and the target spatial feature points, and determine the positioning information of the motion device based on the matching pairs.

[0131] Specifically, the server can compare the descriptor corresponding to the target spatial feature point with the feature descriptors corresponding to each image feature point in the target image. The image feature point corresponding to the feature descriptor with the highest similarity is determined as the image feature point matching the target spatial feature point. This determined image feature point and the target spatial feature point form a matching pair, and the positioning information of the moving device can then be determined based on the matching pair. The descriptor corresponding to the target spatial feature point can be the average value of the feature descriptors of each image feature point in the set of feature points corresponding to the target spatial feature point.

[0132] In one embodiment, determining positioning information based on matching pairs can specifically employ the PnP algorithm, an existing method that will not be elaborated upon here. In another embodiment, determining positioning information based on matching pairs specifically includes: projecting spatial feature points from the matching pair onto the target image to obtain projected feature points; calculating the reprojection error using the projected feature points and image feature points from the matching pair; determining the pose corresponding to the minimum value of the least squares function of the reprojection error as the corrected pose; and correcting the initial pose using the corrected pose to obtain positioning information. Further, the server can return the positioning information to the motion device to be positioned.

[0133] In the above embodiments, since the image feature points are optimized in position based on the feature descriptors of the image feature points during the generation of the feature map, the generated feature map is more robust, thereby greatly improving the positioning accuracy during the positioning process.

[0134] In a specific embodiment, the feature map generation method of this application can be applied to parking application scenarios, specifically including the following steps:

[0135] I. Server generates feature map

[0136] 1. Obtain multiple frames of original images of the target scene captured by a fisheye camera, perform distortion correction on the multiple frames of original images, and obtain multiple frames of images captured for the target scene.

[0137] Specifically, the target vehicle equipped with a fisheye camera can drive in the garage, and the fisheye camera can capture multiple frames of the environment in the garage. These frames are then sent to the server, which performs distortion correction on the multiple frames of the original images to obtain multiple frames of images captured for the target scene.

[0138] It should be noted that the target vehicle and the vehicle that needs to be parked can be the same vehicle or different vehicles.

[0139] 2. Extract image feature points from each frame of the image, and determine the corresponding feature descriptor based on the position of the extracted image feature points in the image to which they belong.

[0140] Specifically, for each frame of image, the server can input the image into a trained feature extraction model, and the feature extraction model outputs a first tensor corresponding to the image feature points and a second tensor corresponding to the feature descriptors. The first tensor is used to describe the probability of feature points appearing in each region of the image. Non-maximum suppression processing is performed on the image based on the first tensor to determine the image feature points from the image. The second tensor is converted into a third tensor with the same size as the image, and the vector in the third tensor that matches the position of the image feature point in its respective image is determined as the descriptor corresponding to the image feature point.

[0141] The first tensor includes multiple channels. Non-maximum suppression processing is performed on the image based on the first tensor to determine image feature points from the image. This includes: obtaining the maximum value of the first tensor at each position along the direction of the multiple channels and the channel index corresponding to each maximum value, respectively obtaining a third tensor and a fourth tensor; determining the target value from the third tensor and searching the neighborhood of the target value's location in the third tensor. The neighborhood of the target value's location includes multiple target locations, and the image distance between the corresponding position of the target location and the corresponding position of the target value's location in the image is less than a preset distance threshold; if the search result indicates that the target value is greater than the value corresponding to other positions in the neighborhood, the target pixel point in the image corresponding to the target value's location is determined as the image feature point of the image; the target pixel point is determined from the image based on the target value's location and the corresponding channel index value, and the channel index value is determined from the fourth tensor based on the target value's location.

[0142] 3. Form a feature point set by combining the image feature points that have matching relationships among the image feature points of each frame.

[0143] 4. For each set of feature points, if the set of feature points meets the filtering conditions, the set of feature points is filtered; if the set of feature points does not meet the filtering conditions, proceed to step 5. The filtering conditions include at least one of the following: the distance between the initial spatial feature points calculated based on the feature point set and the imaging device of the multi-frame images is greater than a first preset distance threshold; the distance between the initial spatial feature points calculated based on the feature point set and the imaging device of the multi-frame images is less than a second preset distance threshold, and the second preset distance threshold is less than the first preset distance threshold; the disparity calculated based on the feature point set is greater than a preset disparity threshold; and the average reprojection error calculated based on the feature point set is greater than a preset error threshold.

[0144] 5. Determine the representative feature point from the feature point set, and calculate the difference between the feature descriptors corresponding to the remaining image feature points in the feature point set and the feature descriptors corresponding to the representative feature point.

[0145] Specifically, the server determines representative feature points from the feature point set through the following steps: based on the position of each image feature point in the feature point set in its respective image, calculate the average feature point position corresponding to the feature point set; determine image feature points whose distance from the average feature point position meets the distance condition from the feature point set, and use the determined image feature points as representative feature points; wherein, the distance condition includes one of the following: the distance from the average feature point position is less than or equal to a distance threshold, or the feature point is sorted before a sorting threshold when sorted in ascending order of distance from the average feature point position.

[0146] 6. Take each remaining image feature point in the feature point set as a target feature point, calculate the matching confidence between each target feature point and the representative feature point, calculate the position error of each target feature point based on the matching confidence and difference of each target feature point, and statistically analyze the position error of each target feature point to obtain the position error of the feature point set.

[0147] The process of calculating the matching confidence between each target feature point and the representative feature point includes: obtaining the feature descriptor of each target feature point and the feature descriptor of the representative feature point; calculating the vector similarity between the feature descriptor of each target feature point and the feature descriptor of the representative feature point, and using each vector similarity as the matching confidence between the corresponding target feature point and the representative feature point.

[0148] 7. Iteratively update the remaining image feature points in the feature point set based on the position error. When the iteration stopping condition is met, the updated feature point set is obtained.

[0149] Specifically, the server can use the gradient descent algorithm to update the positions of the remaining image feature points in the feature point set in the direction of minimizing the position error, determine the descriptors corresponding to the obtained image feature points from the third tensor, and then recalculate the position error. This process is repeated until the iteration stopping condition is met.

[0150] Through the above steps, multiple updated feature point sets can be obtained. The server can then determine whether any of these feature point sets satisfy the filtering conditions. For the feature point sets that satisfy the filtering conditions, the filtering is performed. For the remaining feature point sets after filtering, subsequent steps are executed. The filtering conditions can be referred to the description in the above embodiments.

[0151] 8. Based on the position of each image feature point in the updated feature point set in its respective image, determine the spatial feature points corresponding to the updated feature point set, thereby obtaining multiple spatial feature points.

[0152] 9. Optimize the position of each spatial feature point, specifically including the following steps:

[0153] 9.1 For each spatial feature point, based on the feature descriptors of each image feature point in the updated feature point set corresponding to that spatial feature point, determine the average descriptor corresponding to the updated feature point set.

[0154] 9.2 From the feature descriptors of each image feature point in the updated feature point set, select the feature descriptors whose similarity to the average descriptor meets the similarity condition, and use the selected feature descriptors as reference descriptors.

[0155] 9.3 Project the spatial feature points onto the images to which each image feature point belongs in the updated feature point set to obtain multiple projected feature points. Determine the feature descriptor corresponding to the projected feature points based on their positions on their respective images.

[0156] 9.4. Based on the difference between the feature descriptor and the reference descriptor corresponding to the projected feature point, determine the reprojection error corresponding to the projected feature point.

[0157] 9.5. Calculate the reprojection error corresponding to each projection feature point to obtain the target error. Iterate and update the spatial feature points based on the target error. When the iteration stopping condition is met, the target spatial feature point is obtained. This target feature point is the spatial feature point after position optimization.

[0158] 10. Generate a feature map based on the optimized target space feature points and save the feature map.

[0159] II. Parking Based on Feature Maps

[0160] 1. When a vehicle that needs to park enters the garage entrance, it can download the feature map from the server. The user can input the location of the target parking space, and the vehicle can plan a parking route from the garage entrance to the target parking space based on the feature map.

[0161] 2. The vehicle automatically drives along the planned parking route. During the driving process, the vehicle is positioned in the following ways:

[0162] 2.1. Obtain current inertial measurement data through the IMU, current speed measurement data through the wheel speed sensor, and current target image by the camera installed on the vehicle.

[0163] 2.2 Determine the current initial pose using inertial measurement data and velocity measurement data.

[0164] 2.3. Based on the current initial pose, determine the spatial feature points for position matching from the saved feature map to obtain the target spatial feature points.

[0165] 3.4. Determine image feature points that match the target spatial feature points from the target image, form matching pairs between the determined image feature points and the target spatial feature points, and determine the current position based on the matching pairs.

[0166] 4. When the current location reaches the target parking space, the vehicle will automatically drive into the target parking space and complete the parking process.

[0167] In another specific embodiment, the feature map generation method of this application can be applied to the application scenario of automatic cleaning by a sweeping robot. In this application scenario, the sweeping robot first walks in the area that needs to be cleaned, collects multiple frames of images in the area, generates a feature map according to the feature map generation method provided in the embodiment of this application, and then, in the subsequent automatic cleaning process, the cleaning route can be planned through the feature map. During the automatic cleaning process, the robot automatically locates itself based on the feature map to perform the cleaning task according to the planned cleaning route.

[0168] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0169] Based on the same inventive concept, this application also provides a feature map generation apparatus for implementing the feature map generation method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations of one or more feature map generation apparatuses and positioning information determination apparatus embodiments provided below can be found in the limitations of the feature map generation method described above, and will not be repeated here.

[0170] In one embodiment, such as Figure 8 As shown, a feature map generation device 800 is provided, comprising:

[0171] The feature extraction module 802 is used to obtain multiple frames of images captured for the target scene, extract image feature points from each frame of images, and determine the corresponding feature descriptor based on the position of the extracted image feature points in the respective images.

[0172] The feature point set determination module 804 is used to form a feature point set by combining image feature points with matching relationships among the image feature points of each frame image;

[0173] The difference calculation module 806 determines the representative feature point from the feature point set and calculates the difference between the feature descriptors corresponding to the remaining image feature points in the feature point set and the feature descriptors corresponding to the representative feature point.

[0174] The position update module 808 is used to determine the position error of the feature point set based on the calculated difference, and iteratively update the remaining image feature points in the feature point set based on the position error. When the iteration stopping condition is met, the updated feature point set is obtained.

[0175] The feature map generation module 810 is used to determine the spatial feature points corresponding to the updated feature point set based on the position of each image feature point in its respective image, and to generate a feature map based on the spatial feature points. The feature map is used to locate the moving device to be located in the target scene.

[0176] The aforementioned feature map generation device acquires multiple frames of images captured for a target scene, extracts image feature points from each frame, determines corresponding feature descriptors based on the location of each extracted feature point in its respective image, forms a feature point set by combining matching image feature points from each frame, identifies a representative feature point from this set, calculates the difference between the feature descriptors corresponding to the remaining image feature points in the set and the feature descriptors corresponding to the representative feature point, determines the positional error of the feature point set based on the calculated difference, iteratively updates the remaining image feature points in the feature point set based on the positional error, and obtains an updated feature point set when the iteration stops. Based on the location of each image feature point in the updated feature point set in its respective image, it determines the spatial feature points corresponding to the updated feature point set, and generates a feature map based on the spatial feature points. Because the image feature points are optimized for position based on the feature descriptors during the feature map generation process, the generated feature map is more robust, thereby greatly improving the positioning accuracy during localization.

[0177] In one embodiment, the position update module 808 is used to take each remaining image feature point in the feature point set as a target feature point, calculate the matching confidence between each target feature point and the representative feature point, calculate the position error of each target feature point based on the matching confidence and difference of each target feature point, and calculate the position error of each target feature point by statistically analyzing the position error of each target feature point.

[0178] In one embodiment, the position update module 808 is further configured to obtain the feature descriptors of each target feature point and the feature descriptor of the representative feature point; calculate the vector similarity between the feature descriptors of each target feature point and the feature descriptor of the representative feature point, and use each vector similarity as the matching confidence between the corresponding target feature point and the representative feature point.

[0179] In one embodiment, the difference calculation module 806 is further configured to calculate the average feature point position corresponding to the feature point set based on the position of each image feature point in the feature point set in the image to which it belongs; determine the image feature points in the feature point set whose distance from the average feature point position meets the distance condition, and use the determined image feature points as representative feature points; wherein the distance condition includes one of the following: the distance from the average feature point position is less than or equal to a distance threshold, or the feature point is sorted before a sorting threshold when sorted in ascending order of distance from the average feature point position.

[0180] In one embodiment, the feature point set includes multiple features; the difference calculation module 806 is further configured to, for each feature point set, filter the feature point set if the feature point set meets the filtering conditions; and if the feature point set does not meet the filtering conditions, proceed to the step of determining representative feature points from the feature point set; wherein the filtering conditions include at least one of the following: the distance between the initial spatial feature points calculated based on the feature point set and the imaging device of the multi-frame images is greater than a first preset distance threshold; the distance between the initial spatial feature points calculated based on the feature point set and the imaging device of the multi-frame images is less than a second preset distance threshold, the second preset distance threshold is less than the first preset distance threshold; the disparity calculated based on the feature point set is greater than a preset disparity threshold; and the average reprojection error calculated based on the feature point set is greater than a preset error threshold.

[0181] In one embodiment, the feature map generation module is further configured to: determine an average descriptor corresponding to the updated feature point set based on the feature descriptors of each image feature point in the updated feature point set; select feature descriptors whose similarity to the average descriptor satisfies a similarity condition from the feature descriptors of each image feature point in the updated feature point set, and use the selected feature descriptors as reference descriptors; project spatial feature points onto the images to which each image feature point in the updated feature point set belongs, obtaining multiple projected feature points, and determine the feature descriptors corresponding to the projected feature points based on their positions on the images to which they belong; determine the reprojection error corresponding to the projected feature points based on the difference between the feature descriptors corresponding to the projected feature points and the reference descriptors; statistically analyze the reprojection errors corresponding to each projected feature point to obtain the target error; iteratively update the spatial feature points based on the target error; when the iteration stopping condition is met, obtain the target spatial feature points corresponding to the updated feature point set; and generate a feature map based on the target spatial feature points.

[0182] In one embodiment, the feature extraction module is further configured to input the image into a trained feature extraction model, and output a first tensor corresponding to the image feature points and a second tensor corresponding to the feature descriptors through the feature extraction model; the first tensor is used to describe the probability of feature points appearing in each region of the image; non-maximum suppression processing is performed on the image based on the first tensor to determine the image feature points from the image; the second tensor is converted into a third tensor with the same size as the image, and the vector in the third tensor that matches the position of the image feature point in its respective image is determined as the descriptor corresponding to the image feature point.

[0183] In one embodiment, the first tensor includes multiple channels. The feature extraction module is further configured to: obtain the maximum value of the first tensor at each position and the channel index corresponding to each maximum value along the direction of the multiple channels, thereby obtaining a third tensor and a fourth tensor; determine the target value from the third tensor and search the neighborhood of the target value in the third tensor, wherein the neighborhood of the target value includes multiple target positions, the corresponding position of the target position in the image, and the image distance between the target position and the corresponding position of the target value in the image is less than a preset distance threshold; if the search result indicates that the target value is greater than the value corresponding to other positions in the neighborhood, determine the target pixel in the image corresponding to the target value as the image feature point of the image; the target pixel is determined from the image based on the target value and the corresponding channel index value, and the channel index value is determined from the fourth tensor based on the target value.

[0184] In one embodiment, the feature extraction module is further configured to: obtain multiple frames of original images of the target scene captured by a fisheye camera, perform distortion correction on the multiple frames of original images, and obtain multiple frames of images captured for the target scene.

[0185] In one embodiment, the above-mentioned apparatus further includes: a positioning information determination module, configured to initially acquire inertial measurement data, velocity measurement data, and target images captured by the motion device to be positioned in a target scene; determine the initial pose of the motion device to be positioned using the inertial measurement data and velocity measurement data; determine spatial feature points matching the position from the generated feature map based on the initial pose to obtain target spatial feature points; determine image feature points matching the target spatial feature points from the target image; form matching pairs between the determined image feature points and the target spatial feature points; and determine the positioning information of the motion device based on the matching pairs.

[0186] Each module in the aforementioned feature map generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0187] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores feature map data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a feature map generation method.

[0188] In one embodiment, a computer device is provided, which may be a terminal installed within the aforementioned sports equipment, such as a vehicle-mounted terminal, and its internal structure diagram may be as follows: Figure 10As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a feature map generation method. The display unit of the computer device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0189] Those skilled in the art will understand that Figure 9 , Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0190] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the feature map generation method described above.

[0191] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the feature map generation method described above.

[0192] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the feature map generation method described above.

[0193] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0194] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0195] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0196] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for generating a feature map, characterized in that, The method includes: Obtain multiple frames of images captured for the target scene, extract image feature points from each frame, and determine the corresponding feature descriptor based on the position of the extracted image feature points in their respective images; The image feature points that have matching relationships among the image feature points of each frame are combined into a feature point set; Based on the position of each image feature point in the feature point set in its respective image, calculate the average feature point position corresponding to the feature point set. From the set of feature points, determine the image feature points whose distance from the average feature point position satisfies the distance condition, and use the determined image feature points as representative feature points; Calculate the difference between the feature descriptors corresponding to the remaining image feature points in the feature point set and the feature descriptors corresponding to the representative feature points; The positional error of the feature point set is determined based on the calculated difference. The remaining image feature points in the feature point set are iteratively updated based on the positional error. When the iteration stopping condition is met, the updated feature point set is obtained. Based on the position of each image feature point in the updated feature point set within its respective image, spatial feature points corresponding to the updated feature point set are determined, and a feature map is generated based on the spatial feature points. The feature map is used to locate the moving device to be located in the target scene.

2. The method according to claim 1, characterized in that, The determination of the positional error of the feature point set based on the calculated differences includes: Each remaining image feature point in the feature point set is taken as a target feature point, and the matching confidence between each target feature point and the representative feature point is calculated. Based on the matching confidence and difference of each target feature point, calculate the position error of each target feature point. The positional error of the feature point set is obtained by statistically analyzing the positional error of each of the target feature points.

3. The method according to claim 2, characterized in that, The step of calculating the matching confidence between each of the target feature points and the representative feature points includes: Each of the target feature points is obtained separately, and the feature descriptor of the representative feature point is also obtained. Calculate the vector similarity between the feature descriptor of each target feature point and the feature descriptor of the representative feature point, and use each vector similarity as the matching confidence between the corresponding target feature point and the representative feature point.

4. The method according to claim 1, characterized in that, The distance condition includes one of the following: the distance between the point and the average feature point is less than or equal to a distance threshold, or the point is sorted before a sorting threshold when arranged in ascending order of distance from the average feature point.

5. The method according to any one of claims 1 to 4, characterized in that, The feature point set includes multiple features; the method further includes: For each set of feature points, the set of feature points is filtered if it meets the filtering conditions. If the feature point set does not meet the filtering conditions, the process proceeds to the following steps: calculating the average feature point position corresponding to the feature point set based on the position of each image feature point in the feature point set in its respective image; and determining image feature points whose distance from the average feature point position meets the distance condition from the feature point set, and using the determined image feature points as representative feature points. The filtering conditions include at least one of the following: The distance between the initial spatial feature points calculated based on the feature point set and the imaging device of the multi-frame images is greater than a first preset distance threshold; The distance between the initial spatial feature points calculated based on the feature point set and the capturing device of the multi-frame images is less than a second preset distance threshold, and the second preset distance threshold is less than the first preset distance threshold. The disparity calculated based on the set of feature points is greater than a preset disparity threshold. The average reprojection error calculated based on the set of feature points is greater than a preset error threshold.

6. The method according to claim 1, characterized in that, The generation of a feature map based on the spatial feature points includes: Based on the feature descriptors of each image feature point in the updated feature point set, determine the average descriptor corresponding to the updated feature point set; From the feature descriptors of each image feature point in the updated feature point set, select the feature descriptors whose similarity to the average descriptor meets the similarity condition, and use the selected feature descriptors as reference descriptors; The spatial feature points are projected onto the images to which each image feature point in the updated feature point set belongs, resulting in multiple projected feature points. The feature descriptor corresponding to the projected feature point is determined based on the position of the projected feature point on its respective image. Based on the difference between the feature descriptor corresponding to the projected feature point and the reference descriptor, the reprojection error corresponding to the projected feature point is determined; The target error is obtained by statistically analyzing the reprojection error corresponding to each projection feature point. The spatial feature points are iteratively updated based on the target error. When the iteration stops, the target spatial feature points corresponding to the updated feature point set are obtained. A feature map is generated based on the target spatial feature points.

7. The method according to claim 1, characterized in that, The step of extracting image feature points from each frame of the image and determining the corresponding feature descriptor based on the position of the extracted image feature points in their respective images includes: The image is input into a trained feature extraction model, which outputs a first tensor corresponding to the image feature points and a second tensor corresponding to the feature descriptors. The first tensor is used to describe the probability of feature points appearing in each region of the image. Non-maximum suppression processing is performed on the image based on the first tensor to determine the image feature points of the image; The second tensor is converted into a third tensor with the same size as the image, and the vector in the third tensor that matches the position of the image feature point in the image is determined as the descriptor corresponding to the image feature point.

8. The method according to claim 7, characterized in that, The first tensor includes multiple channels, and the non-maximum suppression processing of the image based on the first tensor to determine image feature points from the image includes: The maximum value of the first tensor at each position and the channel index corresponding to each maximum value are obtained along the direction of the plurality of channels to obtain the third tensor and the fourth tensor respectively. The target value is determined from the third tensor, and the neighborhood of the target value in the third tensor is searched. The neighborhood of the target value includes multiple target positions, the corresponding position of the target position in the image, and the image distance between the target position and the corresponding position of the target value in the image is less than a preset distance threshold. If the search result indicates that the target value is greater than the value corresponding to other positions in the neighborhood, the target pixel in the image corresponding to the position of the target value is determined as the image feature point of the image. The target pixel is determined from the image based on the location of the target value and the corresponding channel index value, and the channel index value is determined from the fourth tensor based on the location of the target value.

9. The method according to any one of claims 1 to 8, characterized in that, The process of obtaining multiple frames of images captured for the target scene includes: A multi-frame original image of the target scene is obtained by capturing it with a fisheye camera, and distortion correction is performed on the multi-frame original image to obtain the multi-frame image captured for the target scene.

10. The method according to any one of claims 1 to 9, characterized in that, The method includes: Acquire inertial measurement data, velocity measurement data, and target images captured by the motion device in the target scene, and use the inertial measurement data and velocity measurement data to determine the initial pose of the motion device to be located; Based on the initial pose, spatial feature points matching the position are determined from the feature map to obtain the target spatial feature points; Image feature points that match the target spatial feature points are determined from the target image, and the determined image feature points and the target spatial feature points are combined to form a matching pair. The positioning information of the motion device is determined based on the matching pair.

11. A feature map generation apparatus, characterized in that, The device includes: The feature extraction module is used to obtain multiple frames of images captured for the target scene, extract image feature points from each frame of images, and determine the corresponding feature descriptor based on the position of the extracted image feature points in the image to which they belong. The feature point set determination module is used to form a feature point set by combining image feature points with matching relationships from the image feature points of each frame image; The difference calculation module is used to calculate the average feature point position corresponding to the feature point set based on the position of each image feature point in the feature point set in the corresponding image; determine the image feature points in the feature point set whose distance from the average feature point position meets the distance condition, and use the determined image feature points as representative feature points; and calculate the difference between the feature descriptors corresponding to the remaining image feature points in the feature point set and the feature descriptors corresponding to the representative feature points. The position update module is used to determine the position error of the feature point set based on the calculated difference, and iteratively update the remaining image feature points in the feature point set based on the position error. When the iteration stopping condition is met, the updated feature point set is obtained. The feature map generation module is used to determine the spatial feature points corresponding to the updated feature point set based on the position of each image feature point in its respective image, and to generate a feature map based on the spatial feature points. The feature map is used to locate the moving device to be located in the target scene.

12. The feature map generation apparatus according to claim 11, characterized in that, The position update module is further configured to take each remaining image feature point in the feature point set as a target feature point, calculate the matching confidence between each target feature point and the representative feature point, calculate the position error of each target feature point based on the matching confidence and difference of each target feature point, and statistically analyze the position error of each target feature point to obtain the position error of the feature point set.

13. The feature map generation apparatus according to claim 12, characterized in that, The location update module is further configured to obtain the feature descriptors of each of the target feature points and the feature descriptor of the representative feature point; calculate the vector similarity between the feature descriptors of each of the target feature points and the feature descriptor of the representative feature point, and use each vector similarity as the matching confidence between the corresponding target feature point and the representative feature point.

14. The feature map generation apparatus according to claim 11, characterized in that, The distance condition includes one of the following: the distance between the point and the average feature point is less than or equal to a distance threshold, or the point is sorted before a sorting threshold when arranged in ascending order of distance from the average feature point.

15. The feature map generation apparatus according to any one of claims 11 to 14, characterized in that, The set of feature points includes multiple points; The difference calculation module is further configured to, for each feature point set, filter the feature point set if the feature point set meets the filtering conditions; if the feature point set does not meet the filtering conditions, proceed to the step of calculating the average feature point position corresponding to the feature point set based on the position of each image feature point in the feature point set in its respective image; and determine the image feature points whose distance from the average feature point position meets the distance condition from the feature point set, and use the determined image feature points as representative feature points. The filtering conditions include at least one of the following: The distance between the initial spatial feature points calculated based on the feature point set and the imaging device of the multi-frame images is greater than a first preset distance threshold; The distance between the initial spatial feature points calculated based on the feature point set and the capturing device of the multi-frame images is less than a second preset distance threshold, and the second preset distance threshold is less than the first preset distance threshold. The disparity calculated based on the set of feature points is greater than a preset disparity threshold. The average reprojection error calculated based on the set of feature points is greater than a preset error threshold.

16. The feature map generation apparatus according to claim 11, characterized in that, The feature map generation module is further configured to: determine an average descriptor corresponding to the updated feature point set based on the feature descriptors of each image feature point in the updated feature point set; select feature descriptors whose similarity to the average descriptor satisfies a similarity condition from the feature descriptors of each image feature point in the updated feature point set, and use the selected feature descriptors as reference descriptors; project the spatial feature points onto the image to which each image feature point in the updated feature point set belongs, obtaining multiple projected feature points, and determine the feature descriptor corresponding to the projected feature points based on their positions on the images to which they belong; Based on the difference between the feature descriptor corresponding to the projected feature point and the reference descriptor, the reprojection error corresponding to the projected feature point is determined; the reprojection error corresponding to each projected feature point is statistically analyzed to obtain the target error; the spatial feature points are iteratively updated based on the target error; when the iteration stopping condition is met, the target spatial feature points corresponding to the updated feature point set are obtained; and a feature map is generated based on the target spatial feature points.

17. The feature map generation apparatus according to claim 11, characterized in that, The feature extraction module is further configured to input the image into a trained feature extraction model, and output a first tensor corresponding to the image feature points and a second tensor corresponding to the feature descriptors through the feature extraction model; the first tensor is used to describe the probability of feature points appearing in each region of the image; non-maximum suppression processing is performed on the image based on the first tensor to determine the image feature points of the image from the image; the second tensor is converted into a third tensor with the same size as the image, and the vector in the third tensor that matches the position of the image feature point in its respective image is determined as the descriptor corresponding to the image feature point.

18. The feature map generation apparatus according to claim 17, characterized in that, The first tensor includes multiple channels. The feature extraction module is further configured to obtain the maximum value of the first tensor at each position and the channel index corresponding to each maximum value along the direction of the multiple channels, thereby obtaining a third tensor and a fourth tensor respectively; determine the target value from the third tensor, and search the neighborhood of the position where the target value is located in the third tensor. The neighborhood of the position where the target value is located includes multiple target positions. The corresponding position of the target position in the image and the image distance between the target position and the corresponding position of the target value in the image are less than a preset distance threshold. If the search result indicates that the target value is greater than the value corresponding to other positions in the neighborhood, the target pixel in the image corresponding to the position of the target value is determined as the image feature point of the image. The target pixel is determined from the image based on the location of the target value and the corresponding channel index value, and the channel index value is determined from the fourth tensor based on the location of the target value.

19. The feature map generation apparatus according to any one of claims 11 to 18, characterized in that, The feature extraction module is also used to obtain multiple frames of original images of the target scene captured by a fisheye camera, and to perform distortion correction on the multiple frames of original images to obtain the multiple frames of images captured for the target scene.

20. The feature map generation apparatus according to any one of claims 11 to 19, characterized in that, The device further includes a positioning information determination module, which is used to acquire inertial measurement data, velocity measurement data and target image captured by the moving device in the target scene, and to determine the initial pose of the moving device to be positioned using the inertial measurement data and the velocity measurement data. Based on the initial pose, spatial feature points matching the position are determined from the feature map to obtain the target spatial feature points; Image feature points that match the target spatial feature points are determined from the target image, and the determined image feature points and the target spatial feature points are combined to form a matching pair. The positioning information of the motion device is determined based on the matching pair.

21. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 10.

22. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.

23. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • End to end framework for geometry-aware multi-scale keypoint detection and matching in fisheye images

    US20190156145A1

  • Method, electronic device and storage medium for vehicle localization

    US20220164595A1