Feature matching method, pose estimation method, equipment, medium and movable platform
By using the tensor core (TensorCore) of the graphics processing unit (GPU) to perform matrix multiplication and addition operations during the feature matching process, the problem of large computational complexity is solved, and more efficient feature matching and pose estimation are achieved.
Patent Information
- Application Number
- CN202510842641.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-23
AI Technical Summary
The feature matching process in existing technologies requires a large amount of computation, resulting in slow calculation speed and affecting efficiency, especially in the fields of autonomous driving, drones and robots.
The tensor core TensorCore in the graphics processing unit (GPU) is used for feature matching, and the cosine similarity calculation after normalization is converted into matrix multiplication and addition calculation, which reduces the calculation complexity and improves the calculation speed.
The calculation speed of feature matching is improved, GPU computing resources are fully utilized, and the utilization rate and data output of graphics processors are improved.
Smart Images

Figure CN120689638A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a feature matching method, a pose estimation method, a device, a medium and a movable platform. Background Art
[0002] Feature matching methods, as a core technology in computer vision and image processing, are widely used in fields such as autonomous driving, drones, and embodied robots.
[0003] In related technologies, when performing feature matching, similarity is mainly calculated through brute force matching (for example, L2 distance) to obtain feature matching results.
[0004] The brute force matching method mentioned above causes the feature matching process to include multiple types of calculations. For example, the L2 distance includes difference calculations, square calculations, and square root calculations. This results in a large amount of calculations in the feature matching process, which in turn affects the calculation speed and further affects the efficiency of feature matching. Summary of the Invention
[0005] The embodiments of the present application provide a feature matching method, a pose estimation method, a device, a medium and a mobile platform, which can fully utilize the hardware resources of the GPU to improve the computing speed and thus improve the speed of feature matching, thereby meeting the performance requirements in various application scenarios.
[0006] In a first aspect, an embodiment of the present application provides a feature matching method, comprising:
[0007] Determining first feature data and second feature data; wherein the first feature data and the second feature data indicate data for feature matching based on image data collected by a camera sensor;
[0008] Determine the feature matching method based on the normalized cosine similarity;
[0009] The first feature data and the second feature data are sent to at least one tensor computing core TensorCore in a graphics processing unit GPU, so that the at least one tensor computing core TensorCore performs feature matching processing on the first feature data and the second feature data according to the feature matching method to obtain a feature matching result.
[0010] In a possible implementation, before sending the first feature data and the second feature data to at least one tensor computing core (TensorCore) in a graphics processing unit (GPU), the method further includes:
[0011] Determining a third matrix size of a matching matrix block according to a first matrix size of the first feature data and a second matrix size of the second feature data; wherein the matrix block indicates a computing block corresponding to the tensor computing core TensorCore;
[0012] According to the third matrix size, data alignment processing is performed on the first feature data and the second feature data respectively.
[0013] In a possible implementation, the number of the first feature data is multiple, and the number of the second feature data is multiple; and the first feature data and the second feature data are respectively aligned according to the third matrix size, including:
[0014] Determine a feature matching mode; wherein the feature matching mode indicates the number of feature data to be matched at a single time; the feature data includes a single first feature data and a single second feature data;
[0015] Determining a matching data alignment method according to the feature matching pattern;
[0016] According to the third matrix size and the data alignment mode, data alignment processing is performed on the first feature data and the second feature data respectively.
[0017] In a possible implementation, determining a matching data alignment mode according to the feature matching pattern includes:
[0018] If the feature matching mode indicates that feature matching is performed on a single feature data at a time, determining that the matching data alignment mode is a zero-padding alignment mode;
[0019] If the feature matching mode indicates that feature matching is performed on a plurality of feature data at a time, the matching data alignment mode is determined to be a feature data alignment mode.
[0020] In one possible implementation, the method further includes:
[0021] The at least one tensor computing core TensorCore performs feature matching processing on the first feature data and the second feature data according to the feature matching method to obtain an initial matching result;
[0022] Performing cross-validation processing on the initial matching result to obtain a cross-validation result;
[0023] After determining that the verification is passed based on the cross-validation result, the initial matching result is determined as the feature matching result.
[0024] In a possible implementation, the initial matching result corresponds to a plurality of feature data; the feature data includes a single first feature data and a single second feature data; and cross-validation processing is performed on the initial matching result, including:
[0025] Determining, from the initial matching results, initial matching sub-results corresponding to each of the feature data according to the feature data splicing identifier;
[0026] A cross-validation process is performed on each of the initial matching sub-results to obtain the cross-validation result.
[0027] In a second aspect, an embodiment of the present application provides a pose estimation method, comprising:
[0028] Acquire multiple frames of image data of the target object;
[0029] Determining characteristic image data corresponding to each frame of image data;
[0030] According to any one of the feature matching methods of the first aspect above, feature matching processing is performed on the feature image data to obtain a feature matching result;
[0031] The position and posture of the target object are determined according to the feature matching result.
[0032] In a third aspect, an embodiment of the present application provides a feature matching device, comprising:
[0033] A first determining unit, configured to determine first feature data and second feature data; wherein the first feature data and the second feature data indicate data for feature matching based on image data collected by a camera sensor;
[0034] A second determining unit is used to determine a feature matching method according to the normalized cosine similarity;
[0035] A data processing unit is used to send the first feature data and the second feature data to at least one tensor computing core TensorCore in a graphics processing unit GPU, so that the at least one tensor computing core TensorCore performs feature matching processing on the first feature data and the second feature data according to the feature matching method to obtain a feature matching result.
[0036] In a possible implementation, the device further includes a data alignment unit configured to:
[0037] Before sending the first feature data and the second feature data to at least one tensor computing core TensorCore in a graphics processing unit (GPU), determining a third matrix size of a matching matrix block based on a first matrix size of the first feature data and a second matrix size of the second feature data; wherein the matrix block indicates a computing block corresponding to the tensor computing core TensorCore;
[0038] According to the third matrix size, data alignment processing is performed on the first feature data and the second feature data respectively.
[0039] In a possible implementation, the number of the first feature data is multiple, and the number of the second feature data is multiple; in this case, the data alignment unit is configured to:
[0040] Determine a feature matching mode; wherein the feature matching mode indicates the number of feature data to be matched at a single time; the feature data includes a single first feature data and a single second feature data;
[0041] Determining a matching data alignment method according to the feature matching pattern;
[0042] According to the third matrix size and the data alignment mode, data alignment processing is performed on the first feature data and the second feature data respectively.
[0043] In a possible implementation, the second determining unit is configured to:
[0044] If the feature matching mode indicates that feature matching is performed on a single feature data at a time, determining that the matching data alignment mode is a zero-padding alignment mode;
[0045] If the feature matching mode indicates that feature matching is performed on a plurality of feature data at a time, the matching data alignment mode is determined to be a feature data alignment mode.
[0046] In one possible embodiment, the device is further used for:
[0047] The at least one tensor computing core TensorCore performs feature matching processing on the first feature data and the second feature data according to the feature matching method to obtain an initial matching result;
[0048] Performing cross-validation processing on the initial matching result to obtain a cross-validation result;
[0049] After determining that the verification is passed based on the cross-validation result, the initial matching result is determined as the feature matching result.
[0050] In a possible implementation manner, the initial matching result corresponds to a plurality of feature data; the feature data includes a single first feature data and a single second feature data; in this case, the apparatus is further configured to:
[0051] Determining, from the initial matching results, initial matching sub-results corresponding to each of the feature data according to the feature data splicing identifier;
[0052] A cross-validation process is performed on each of the initial matching sub-results to obtain the cross-validation result.
[0053] In a fourth aspect, an embodiment of the present application provides a posture estimation device, comprising:
[0054] A data acquisition unit, used for acquiring multiple frames of image data of a target object;
[0055] A feature extraction unit, configured to determine feature image data corresponding to each frame of image data;
[0056] a feature matching unit, configured to perform feature matching processing on the feature image data according to the feature matching method described in any one of the first aspects above, to obtain a feature matching result;
[0057] A pose estimation unit is used to determine the pose of the target object according to the feature matching result.
[0058] In a fifth aspect, an embodiment of the present application provides a computer device, including: a memory, a processor;
[0059] The memory stores computer-executable instructions;
[0060] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the first aspect and / or various possible implementations of the first aspect, or executes the implementation of the second aspect.
[0061] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect and / or various possible implementations of the first aspect, or to implement the implementation of the second aspect.
[0062] In the seventh aspect, an embodiment of the present application provides a movable platform, comprising: a camera sensor and a processor; the processor is communicatively connected to the camera sensor, and the processor is used to execute the first aspect and / or various possible implementations of the first aspect above, or to implement the implementation of the second aspect above.
[0063] In an eighth aspect, an embodiment of the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect, or implements the implementation of the second aspect.
[0064] The feature matching method, pose estimation method, device, medium, and mobile platform provided in the embodiments of the present application can determine the feature matching method based on the cosine similarity after normalization after determining the first feature data and the second feature data, thereby converting the feature matching calculation process into a matrix multiplication and addition calculation process. Compared with the brute force matching method, it reduces the computational complexity and the amount of data calculation, and is more compatible with the high-computing-power computing unit (i.e., the tensor computing core TensorCore) in the graphics processing unit GPU. Afterwards, the first feature data and the second feature data can be sent to at least one tensor computing core TensorCore in the graphics processing unit GPU, so that at least one tensor computing core TensorCore performs feature matching processing on the first feature data and the second feature data according to the feature matching method to obtain a feature matching result. This implementation method can not only improve the calculation speed of feature matching, but also save the computing resources of the graphics processing unit GPU while giving full play to the computing performance of the graphics processing unit GPU, thereby improving the utilization rate of the graphics processing unit GPU, and thus helping to improve the overall data output of the data closed loop. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0066] Figure 1 A schematic diagram of the processing flow of a front-end visual odometry provided in an embodiment of the present application;
[0067] Figure 2 A schematic diagram of a flow chart of a feature matching method provided in an embodiment of the present application;
[0068] Figure 3 A flowchart of another feature matching method provided in an embodiment of the present application;
[0069] Figure 4 A schematic diagram of an implementation flow of a feature matching method provided in an embodiment of the present application;
[0070] Figure 5 A flowchart of a pose estimation method provided in an embodiment of the present application;
[0071] Figure 6 A schematic structural diagram of a feature matching device provided in this application;
[0072] Figure 7 A schematic structural diagram of a posture estimation device provided in this application;
[0073] Figure 8 A schematic diagram of the structure of a computer device provided in an embodiment of the present application;
[0074] Figure 9 A schematic structural diagram of a movable platform provided in an embodiment of the present application.
[0075] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0076] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0077] First, let’s explain the terms involved in this application:
[0078] SLAM algorithm: Simultaneous Localization and Mapping, a technology that uses sensor data to estimate the position of mobile devices such as robots in real time and build a map of the environment;
[0079] CUDA core: also known as CudaCore, which is a unit in the graphics processing unit GPU used to handle general parallel computing;
[0080] Tensor Core: also known as Tensor Core, is a dedicated computing unit in the graphics processing unit GPU, which improves computing efficiency through multiplication and addition fusion calculations.
[0081] In the fields of autonomous driving, drones, and embodied robots, positioning and navigation systems are indispensable, and SLAM algorithms are the core technology. SLAM algorithms are mainly divided into front-end visual odometry, back-end nonlinear optimization, mapping, and loop detection. The processing process of the front-end visual odometry can be found in Figure 1 .
[0082] Figure 1This is a schematic diagram of the processing flow of a front-end visual odometer provided in an embodiment of the present application. Figure 1 As shown, it includes: extracting features from the image data collected by the camera sensor, then performing feature matching processing on the current frame features and the previous frame features, and performing pose estimation based on the feature matching results.
[0083] In related technologies, when performing feature matching, similarity is mainly calculated through brute force matching (for example, L2 distance) to obtain feature matching results.
[0084] When calculating the similarity based on the L2 distance, please refer to the following formula (1).
[0085] (1)
[0086] The brute force matching method mentioned above causes the feature matching process to include multiple types of calculations, such as difference calculations, square calculations, and square root calculations, as shown in formula (1). This makes the feature matching process more complex and involves a large amount of calculations, which easily affects the calculation speed and further affects the efficiency of feature matching.
[0087] Research has found that currently, accelerated computing on graphics processing units (GPUs) typically relies on parallel acceleration via the CUDA cores within the GPU. This approach, which relies on CUDA cores with lower computing power to parallelize the values of each dimension, results in a low upper limit for accelerated computing and fails to fully utilize the performance of the GPU.
[0088] The feature matching method provided in the present application converts the similarity calculation method in the feature matching process from one involving multiple types of calculations to a process involving only multiplication and addition calculations, and performs matrix multiplication and addition operations based on the tensor computing core TensorCore with relatively strong actual computing power in the graphics processing unit GPU, thereby reducing the computational complexity of feature matching and improving the processing speed of feature matching. While giving full play to the computing performance of the graphics processing unit GPU, it also saves the computing resources of the graphics processing unit GPU.
[0089] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0090] Figure 2 A flow chart of a feature matching method provided in an embodiment of the present application is shown as follows: Figure 2As shown, the method includes:
[0091] S201: Determine first feature data and second feature data.
[0092] The first feature data and the second feature data indicate data for feature matching based on image data collected by the camera sensor.
[0093] In one example, the camera sensor may indicate any type of camera, such as a monocular camera, a binocular camera, etc.
[0094] In one example, after collecting image data using a camera sensor, feature extraction can be performed to obtain feature image data corresponding to each frame of image data. In this case, the features of the current frame can be determined as first feature data, and the features of the previous frame can be determined as second feature data.
[0095] In one example, when extracting features from image data, any feature extraction method may be used, which is not specifically limited here.
[0096] S202: Determine a feature matching method based on the normalized cosine similarity.
[0097] In one example, the feature matching method can be understood as a calculation method for performing feature matching on the first feature data and the second feature data, that is, a calculation formula.
[0098] Generally, the similarity between the first feature data and the second feature data can be calculated by feature matching, and the matching points in the first feature data and the second feature data can be determined based on the value of the similarity, thereby completing the feature matching of the first feature data and the second feature data.
[0099] In the embodiment of the present application, a calculation formula for feature matching can be determined based on the cosine similarity after normalization, and the calculation formula can be determined as the feature matching method.
[0100] In specific implementation, the similarity can be calculated first by using cosine similarity, wherein the calculation process of cosine similarity can be shown in the following formula (2).
[0101] (2)
[0102] By normalizing the cosine similarity shown in formula (2), the following formula (3) is satisfied.
[0103] (3)
[0104] Based on this, the calculation process of cosine similarity shown in formula (2) can be simplified as shown in the following formula (4).
[0105] (4)
[0106] At this time, the normalized cosine similarity shown in formula (4) can be determined as the feature matching method, and according to this formula, the complex feature matching calculation can be converted into a multiplication and addition operation that is more suitable for the tensor computing core TensorCore in the graphics processor GPU.
[0107] S203. Send the first feature data and the second feature data to at least one tensor computing core TensorCore in the graphics processing unit GPU, so that the at least one tensor computing core TensorCore performs feature matching processing on the first feature data and the second feature data according to the feature matching method to obtain a feature matching result.
[0108] From the above description, it can be seen that after determining the first feature data and the second feature data, the embodiment of the present application can determine the feature matching method based on the cosine similarity after normalization, thereby converting the feature matching calculation process into a matrix multiplication and addition calculation process. Compared with the brute force matching method, it reduces the computational complexity and the amount of data calculation, and is more compatible with the high-computing-power computing unit (i.e., the tensor computing core TensorCore) in the graphics processing unit GPU. Afterwards, the first feature data and the second feature data can be sent to at least one tensor computing core TensorCore in the graphics processing unit GPU, so that at least one tensor computing core TensorCore performs feature matching processing on the first feature data and the second feature data according to the feature matching method to obtain a feature matching result. This implementation method can not only improve the calculation speed of feature matching, but also save the computing resources of the graphics processing unit GPU while giving full play to the computing performance of the graphics processing unit GPU, thereby improving the utilization rate of the graphics processing unit GPU, and then improving the overall data output of the data closed loop.
[0109] Figure 3 A flow chart of another feature matching method provided in an embodiment of the present application is shown as follows: Figure 3 As shown, this embodiment Figure 2 Based on the embodiment, the feature matching method is described in detail, and the method includes:
[0110] S301: Determine first feature data and second feature data.
[0111] The first feature data and the second feature data indicate data for feature matching based on image data collected by the camera sensor.
[0112] In one example, this step can refer to the content described in S201 above, and will not be described in detail here.
[0113] In one possible implementation, after determining the first feature data and the second feature data, and before sending the first feature data and the second feature data to at least one tensor computing core TensorCore in a graphics processing unit GPU, the embodiment of the present application can adjust the matrix sizes corresponding to the first feature data and the second feature data, respectively, so that the first feature data and the second feature data can be adapted to the matrix calculation of the tensor computing core TensorCore. For details, see steps S302 to S303 described below.
[0114] S302 : Determine a third matrix size of the matching matrix block according to the first matrix size of the first feature data and the second matrix size of the second feature data.
[0115] Among them, the matrix block indicates the computing block corresponding to the tensor computing core TensorCore.
[0116] In an example, the matrix blocks of the tensor computing core TensorCore can be divided into at least the following types: 16*16*16 matrix blocks, 32*32*16 matrix blocks, 64*64*16 matrix blocks, 32*8*16 matrix blocks, 8*32*16 matrix blocks, 32*16*16 matrix blocks, 32*64*16 matrix blocks, etc. There is no limitation on the size of the third matrix corresponding to the matrix block, which is subject to the matrix block size supported by the specific graphics processor GPU.
[0117] In one example, the first matrix size of the first feature data (or the second matrix size of the second feature data) can be determined by the number of feature points included in the first feature data (or the second feature data) and the dimension of each feature point. For example, if the first feature data includes N feature points, each with a dimension of 128, then the first matrix size is N*128; if the second feature data includes M feature points, each with a dimension of 128, then the second matrix size is M*128.
[0118] At this time, if the values of N and M are close, then the size of the third matrix can be determined to be 16*16*16; if the difference between N and M is large, and N is greater than M, then the size of the third matrix can be determined to be 32*8*16; if the difference between N and M is large, and N is less than M, then the size of the third matrix can be determined to be 8*32*16.
[0119] S303 : Perform data alignment processing on the first characteristic data and the second characteristic data respectively according to the third matrix size.
[0120] This implementation can first determine the matrix blocks for calculating the first feature data and the second feature data, and perform data alignment processing according to the matrix sizes corresponding to the matrix blocks, the first feature data and the second feature data, so as to make full use of the computing resources of the tensor computing core TensorCore and thus improve computing performance.
[0121] In one possible implementation, since in actual application scenarios, there are multiple first feature data and second feature data that need to be matched, in step S303 above, data alignment processing is performed on the first feature data and the second feature data, respectively, based on the size of the third matrix. For details, please refer to the process described below.
[0122] First, a feature matching mode is determined, wherein the feature matching mode indicates the number of feature data to be matched in a single time; the feature data includes a single first feature data and a single second feature data.
[0123] In one example, the feature matching mode may include a single matching mode and a batch processing mode, wherein the single matching mode may indicate that feature matching is performed on a single feature data at a time, and the batch processing mode may indicate that feature matching is performed on multiple feature data at a time.
[0124] At this time, by supporting different feature matching modes, the feature matching method can be made flexible and diverse, thereby meeting the needs of various application scenarios and improving the scope of application of the feature matching method.
[0125] Then, according to the feature matching pattern, the matching data alignment method is determined.
[0126] Finally, according to the third matrix size and the data alignment mode, data alignment processing is performed on the first characteristic data and the second characteristic data respectively.
[0127] In one example, the first characteristic data and the second characteristic data may be cropped according to the third matrix size, and the portion that does not meet the third matrix size may be padded with data according to the data alignment method, thereby completing the data alignment process.
[0128] Optionally, if the feature matching mode indicates that feature matching is performed on a single feature data at a time, the matching data alignment mode is determined to be a zero-padding alignment mode; if the feature matching mode indicates that feature matching is performed on multiple feature data at a time, the matching data alignment mode is determined to be a feature data alignment mode.
[0129] Based on this, in the zero-padding alignment mode, after data cropping, the portion that does not meet the third matrix size requirement can be padded with zeros.
[0130] In the feature data alignment mode, after data cropping, new feature data can be added to the part that does not meet the third matrix size requirement, so that batch processing can be realized by adding new feature data, which can further improve data processing efficiency.
[0131] In a possible implementation, if some feature data that does not meet the third matrix size requirement still exists after all feature data are filled, zero padding can be used to align the data of the part that does not meet the third matrix size requirement.
[0132] In one possible implementation, in batch processing mode, zeros can be directly added when there is less data to be filled, and new feature data can be added when there is more data to be filled. This not only reduces the number of data cropping times and improves the calculation speed, but also makes the batch processing mode more flexible.
[0133] For example, if the matrix block size of 16*16*16 is used, then the first feature data and the second feature data are aligned as multiples of 16 before being sent to the tensor computing core TensorCore; if the matrix block size of 64*64*16 is used, then the first feature data and the second feature data are aligned as multiples of 64 before being sent to the tensor computing core TensorCore, so as to utilize the tensor computing core TensorCore for calculation and processing.
[0134] In one example, after determining the data that needs to be feature matched, feature matching processing can be performed according to the process described in S304 to S305 below.
[0135] S304: Determine a feature matching method based on the normalized cosine similarity.
[0136] In one example, this step can refer to the content described in S202 above, and will not be described in detail here.
[0137] S305. Send the first feature data and the second feature data to at least one tensor computing core TensorCore in the graphics processing unit GPU, so that the at least one tensor computing core TensorCore performs feature matching processing on the first feature data and the second feature data according to the feature matching method to obtain a feature matching result.
[0138] In a possible implementation, after feature matching processing is performed on the first feature data and the second feature data, the accuracy of the feature matching can be ensured by cross-verification consistency.
[0139] In a specific implementation, after the first feature data and the second feature data are sent to at least one tensor computing core TensorCore in a graphics processing unit (GPU), the at least one tensor computing core TensorCore performs feature matching processing on the first feature data and the second feature data according to a feature matching method to obtain an initial matching result. Then, a cross-validation process is performed on the initial matching result to obtain a cross-validation result. After determining that the verification is passed based on the cross-validation result, the initial matching result is determined as the feature matching result.
[0140] In one example, cross-validation can be performed by determining whether the points on the first feature data and the second feature data correspond to the points with the highest matching degree. For example, for point a on the first feature data, if it is determined that the point with the highest similarity to point a in the second feature data is point b, then consistency can be determined by determining whether the point in the first feature data with the highest similarity to point b in the second feature data is point a. If so, consistency is satisfied, indicating that verification has passed; if not, consistency is not satisfied, indicating that verification has failed.
[0141] In practice, the maximum value in column j can be found in row i of the initial matching result. If the maximum value in column j is also in row i, then consistency is satisfied, indicating that verification has passed.
[0142] In a possible implementation, the data filling can be performed based on the feature data, so that the initial matching result corresponds to multiple feature data, wherein the feature data includes a single first feature data and a single second feature data.
[0143] Based on this, when performing cross-validation on the initial matching results, the initial matching sub-results corresponding to each feature data can be determined from the initial matching results according to the feature data splicing identifier. Then, cross-validation is performed on each initial matching sub-result to obtain a cross-validation result.
[0144] In one example, the feature data splicing identifier can indicate a splicing position or a splicing character. In a specific implementation, the starting position and amount of filling can be recorded during the feature data filling process to obtain the splicing position. Alternatively, during the feature data filling process, a preset character can be filled between different feature data to obtain the splicing character. The feature data splicing identifier is not specifically limited here and is subject to implementation.
[0145] This implementation can perform cross-validation on each feature data separately when the initial matching result corresponds to multiple feature data, thereby avoiding interference between different feature data and improving the accuracy of cross-validation.
[0146] Figure 4 A schematic diagram of an implementation flow of a feature matching method provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, the feature matching method provided in the embodiment of the present application can add an L2 norm normalization layer after the output layer of the feature extraction network, so that after the features are extracted by the output layer of the feature extraction network, the feature image data can be obtained after being processed by the L2 norm normalization layer. At this time, the first feature data and the second feature data that need to be feature matched can be determined from the feature image data. At this time, the cosine similarity after normalization can be achieved by performing a dot product operation on the first feature data and the second feature data. At the same time, data alignment processing can be performed on the first feature data and the second feature data. Afterwards, the first feature data and the second feature data are sent to at least one tensor computing core TensorCore in the graphics processing unit GPU, so that the at least one tensor computing core TensorCore performs feature matching processing on the first feature data and the second feature data to obtain a feature matching result.
[0147] In a possible implementation, the feature matching method in the above embodiment can be applied to the pose estimation method, see Figure 5 . Figure 5 A flow chart of a pose estimation method provided in an embodiment of the present application is shown as follows: Figure 5 As shown, the method includes:
[0148] S501: Acquire multiple frames of image data of a target object.
[0149] In one example, the target object can be determined based on the application scenario of the pose estimation method. For example, in an autonomous driving scenario, the target object can be a vehicle; in a drone application scenario, the target object can be a drone; in an embodied robot scenario, the target object can be a robot, etc.
[0150] S502: Determine the characteristic image data corresponding to each frame of image data.
[0151] In one example, feature extraction can be performed on each frame of image data according to a feature extraction method to obtain feature image data. The feature extraction method is not limited here and is subject to practicability.
[0152] S503: Perform feature matching processing on the feature image data according to the above feature matching method to obtain a feature matching result.
[0153] Here, the feature image data that needs to be feature matched can be determined as the first feature data and the second feature data mentioned above, and feature matching processing is performed.
[0154] S504: Determine the position and posture of the target object based on the feature matching result.
[0155] In one example, the relative motion of the camera sensor from the previous frame to the current frame can be determined based on the feature matching results, and positioning can be performed based on the relative motion and the internal and external parameters of the camera sensor to determine the position and posture of the target object.
[0156] This implementation can perform pose estimation through the feature matching method provided in the above embodiment, which can improve the calculation speed of pose estimation and thus improve the performance of the pose estimation method.
[0157] In a possible implementation, the pose estimation method can be applied to the aforementioned SLAM algorithm, thereby improving the performance of the SLAM algorithm.
[0158] Figure 6 A structural diagram of a feature matching device provided in this application is shown as follows: Figure 6 As shown, the feature matching device 60 provided in this embodiment includes:
[0159] The first determining unit 601 is configured to determine first feature data and second feature data, wherein the first feature data and the second feature data indicate data for feature matching based on image data collected by a camera sensor.
[0160] The second determining unit 602 is configured to determine a feature matching method according to the normalized cosine similarity.
[0161] The data processing unit 603 is used to send the first feature data and the second feature data to at least one tensor computing core TensorCore in the graphics processing unit GPU, so that the at least one tensor computing core TensorCore performs feature matching processing on the first feature data and the second feature data according to the feature matching method to obtain a feature matching result.
[0162] In a possible implementation, the apparatus further includes a data alignment unit 604 configured to:
[0163] Before sending the first feature data and the second feature data to at least one tensor computing core TensorCore in a graphics processing unit (GPU), determining a third matrix size of a matching matrix block based on a first matrix size of the first feature data and a second matrix size of the second feature data; wherein the matrix block indicates a computing block corresponding to the tensor computing core TensorCore;
[0164] According to the third matrix size, data alignment processing is performed on the first characteristic data and the second characteristic data respectively.
[0165] In a possible implementation, there are multiple first feature data and multiple second feature data. In this case, the data alignment unit 604 is configured to:
[0166] Determine a feature matching mode; wherein the feature matching mode indicates the number of feature data to be matched at a single time; the feature data includes a single first feature data and a single second feature data;
[0167] Determine the matching data alignment method based on the feature matching pattern;
[0168] According to the third matrix size and the data alignment mode, data alignment processing is performed on the first characteristic data and the second characteristic data respectively.
[0169] In a possible implementation, the second determining unit 602 is configured to:
[0170] If the feature matching mode indicates that feature matching is performed on a single feature data at a time, determining that the matched data alignment mode is a zero-padding alignment mode;
[0171] If the feature matching mode indicates that feature matching is performed on multiple feature data at a time, the matched data alignment mode is determined to be the feature data alignment mode.
[0172] In one possible embodiment, the device is further used for:
[0173] At least one tensor computing core TensorCore performs feature matching processing on the first feature data and the second feature data according to the feature matching method to obtain an initial matching result;
[0174] Perform cross-validation processing on the initial matching results to obtain cross-validation results;
[0175] After the cross-validation result determines that the validation is passed, the initial matching result is determined as the feature matching result.
[0176] In a possible implementation, the initial matching result corresponds to multiple feature data; the feature data includes a single first feature data and a single second feature data; in this case, the apparatus is further configured to:
[0177] According to the feature data splicing identifier, the initial matching sub-result corresponding to each feature data is determined from the initial matching result;
[0178] A cross-validation process is performed on each initial matching sub-result to obtain a cross-validation result.
[0179] The feature matching device provided in this embodiment can execute the feature matching method provided in the above method embodiment. Its implementation principle and technical effects are similar, and are not described in detail in this embodiment.
[0180] Figure 7 A schematic diagram of the structure of a posture estimation device provided in this application is shown as follows: Figure 7 As shown, the posture estimation device 70 provided in this embodiment includes:
[0181] The data acquisition unit 701 is configured to acquire multiple frames of image data of a target object.
[0182] The feature extraction unit 702 is used to determine feature image data corresponding to each frame of image data.
[0183] The feature matching unit 703 is configured to perform feature matching processing on the feature image data according to any one of the feature matching methods in the first aspect to obtain a feature matching result.
[0184] The pose estimation unit 704 is used to determine the pose of the target object according to the feature matching result.
[0185] The posture estimation device provided in this embodiment can execute the posture estimation method provided in the above method embodiment. Its implementation principle and technical effects are similar, and are not described in detail in this embodiment.
[0186] Figure 8 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 8 As shown, the computer device 80 provided in this embodiment includes: at least one processor 801 and a memory 802. Optionally, the computer device 80 also includes a communication component 803. The processor 801, the memory 802 and the communication component 803 are connected via a bus 804.
[0187] During the specific implementation process, at least one processor 801 executes the computer execution instructions stored in the memory 802, so that at least one processor 801 executes the above-mentioned feature matching method or pose estimation method.
[0188] The specific implementation process of the processor 801 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.
[0189] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASICs), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.
[0190] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.
[0191] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.
[0192] Figure 9 A structural diagram of a movable platform provided in an embodiment of the present application is shown in FIG. Figure 9 As shown, the movable platform 90 provided in the embodiment of the present application includes: a camera sensor 901 and a processor 902. The processor 902 is in communication with the camera sensor 901, and the processor 902 is used to execute the above feature matching method or pose estimation method.
[0193] The present application also provides a computer program product, including a computer program, which implements the above-mentioned feature matching method or pose estimation method when executed by a processor.
[0194] The present application also provides a computer-readable storage medium, which stores computer-executable instructions. When a processor executes the computer-executable instructions, the above-mentioned feature matching method or posture estimation method is implemented.
[0195] The readable storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0196] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.
[0197] The division of units is merely a logical functional division; actual implementations may employ alternative divisions, such as combining or integrating multiple units or components into another system, or omitting or disabling certain features. Furthermore, any direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units, either through an interface, electrical, mechanical, or other means.
[0198] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0199] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0200] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0201] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0202] Finally, it should be noted that those skilled in the art will readily identify other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The present invention is not limited to the precise structure described above and illustrated in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof. The scope of the present invention is limited solely by the appended claims.
Claims
1. A feature matching method, characterized in that: include: Determining first feature data and second feature data; wherein the first feature data and the second feature data indicate data for feature matching based on image data collected by a camera sensor; Determine the feature matching method based on the normalized cosine similarity; The first feature data and the second feature data are sent to at least one tensor computing core TensorCore in a graphics processing unit GPU, so that the at least one tensor computing core TensorCore performs feature matching processing on the first feature data and the second feature data according to the feature matching method to obtain a feature matching result.
2. The method according to claim 1, characterized in that Before sending the first feature data and the second feature data to at least one tensor computing core (TensorCore) in a graphics processing unit (GPU), the method further includes: Determining a third matrix size of a matching matrix block according to a first matrix size of the first feature data and a second matrix size of the second feature data; wherein the matrix block indicates a computing block corresponding to the tensor computing core TensorCore; According to the third matrix size, data alignment processing is performed on the first feature data and the second feature data respectively.
3. The method according to claim 2, characterized in that There are multiple first feature data and multiple second feature data; and according to the third matrix size, performing data alignment processing on the first feature data and the second feature data, respectively, including: Determine a feature matching mode; wherein the feature matching mode indicates the number of feature data to be matched at a single time; the feature data includes a single first feature data and a single second feature data; Determining a matching data alignment method according to the feature matching pattern; According to the third matrix size and the data alignment mode, data alignment processing is performed on the first feature data and the second feature data respectively.
4. The method according to claim 3, characterized in that Determining a matching data alignment method based on the feature matching pattern includes: If the feature matching mode indicates that feature matching is performed on a single feature data at a time, determining that the matching data alignment mode is a zero-padding alignment mode; If the feature matching mode indicates that feature matching is performed on a plurality of feature data at a time, the matching data alignment mode is determined to be a feature data alignment mode.
5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: The at least one tensor computing core TensorCore performs feature matching processing on the first feature data and the second feature data according to the feature matching method to obtain an initial matching result; Performing cross-validation processing on the initial matching result to obtain a cross-validation result; After determining that the verification is passed based on the cross-validation result, the initial matching result is determined as the feature matching result.
6. The method according to claim 5, characterized in that The initial matching result corresponds to a plurality of feature data; the feature data includes a single first feature data and a single second feature data; Performing cross-validation processing on the initial matching results includes: Determining, from the initial matching results, initial matching sub-results corresponding to each of the feature data according to the feature data splicing identifier; A cross-validation process is performed on each of the initial matching sub-results to obtain the cross-validation result.
7. A pose estimation method, characterized in that: include: Acquire multiple frames of image data of the target object; Determining characteristic image data corresponding to each frame of image data; According to the feature matching method of any one of claims 1 to 6, feature matching processing is performed on the feature image data to obtain a feature matching result; The position and posture of the target object are determined according to the feature matching result.
8. A computer device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the feature matching method according to any one of claims 1 to 6, or executes the pose estimation method according to claim 7.
9. A computer storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the feature matching method according to any one of claims 1 to 6, or the pose estimation method according to claim 7.
10. A movable platform, characterized in that: include: A camera sensor and a processor; the processor is communicatively connected to the camera sensor, and the processor is used to execute the feature matching method as described in any one of claims 1 to 6, or the pose estimation method as described in claim 7.
Citation Information
Patent Citations
GA (genetic algorithm) optimized SVM (support vector machine) and normalization-combined palmprint and palm vein fusion identification method
CN106548134A
Progressive modification of generative adversarial neural networks
CN110059793A
Face picture feature comparison method and device, computer equipment and storage medium
CN111274996A
Tensorcore-based convolutional neural network operation method and device
CN112215345A
Medical endoscope video electronic image stabilization method
CN117934316A