Pose determination method, apparatus, device, and readable storage medium
By combining data from cameras and radar, and decoupling the calculation of pose transformation data, the pose error problem caused by a single sensor is solved, achieving higher accuracy and robustness.
Patent Information
- Application Number
- CN202610781681.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-08-25
AI Technical Summary
In existing technologies, errors are prone to occur when using a single sensor to determine pose, leading to inaccurate pose.
A combination of camera and radar is used to determine depth information and feature descriptors by acquiring image frames from the camera and point cloud data from the radar, calculate pose transformation data, and combine the results of the two in a decoupled manner to improve accuracy.
This technology enables accurate pose determination even when sensors malfunction, improving the accuracy of pose determination and enhancing the robustness and reliability of the system.
Smart Images

Figure CN122636726A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, specifically relating to a pose determination method, apparatus, device, and readable storage medium. Background Technology
[0002] With the rapid development of intelligent transportation, roadside intelligent devices are playing an increasingly prominent role in traffic perception systems. To achieve high-precision, all-around perception of the traffic environment, it is usually necessary to deploy sensors, such as cameras and radar, to determine the poses of people and objects in the current traffic environment.
[0003] However, in related technologies, pose determination is achieved using a single sensor. If this sensor malfunctions, pose errors can easily occur. Summary of the Invention
[0004] This application provides a pose determination method, apparatus, device, and readable storage medium to solve the problem of pose errors.
[0005] Firstly, a pose determination method is provided, including:
[0006] Acquire at least two consecutive image frames captured by the camera, and at least two consecutive point cloud data frames acquired by the radar;
[0007] Determine the depth information and feature descriptors corresponding to each image frame in the at least two consecutive image frames;
[0008] The first pose transformation data corresponding to the at least two consecutive image frames is determined based on the depth information and feature descriptors corresponding to each image frame;
[0009] Based on the at least two consecutive frames of point cloud data, determine the ground point cloud corresponding to each frame in the at least two consecutive frames;
[0010] Based on the ground point cloud corresponding to each frame, determine the second pose transformation data corresponding to the at least two consecutive frames of point cloud data;
[0011] The pose is determined based on the first pose transformation data and / or the second pose transformation data.
[0012] Optionally, determining the first pose transformation data corresponding to the at least two consecutive image frames based on the depth information and feature descriptors corresponding to each image frame includes:
[0013] Based on the depth information corresponding to each image frame, determine the three-dimensional coordinates of at least one feature point corresponding to each image frame;
[0014] Based on the feature descriptors corresponding to each image frame, determine the similarity between the three-dimensional coordinates of at least one feature point corresponding to at least two image frames;
[0015] Determine a three-dimensional feature point pair, wherein the three-dimensional feature point pair is a feature point pair between at least two image frames whose similarity is greater than or equal to a first preset value;
[0016] The first pose transformation data is determined based on the three-dimensional feature point pairs.
[0017] Optionally, determining the second pose transformation data corresponding to the at least two consecutive frames of point cloud data based on the ground point cloud corresponding to each frame includes:
[0018] Based on the ground point cloud corresponding to each frame, the third pose transformation data is determined, and the third pose transformation data is the initial pose transformation data.
[0019] Determine the static point cloud in the ground point cloud corresponding to each frame;
[0020] The second pose transformation data is determined based on the static point cloud corresponding to each frame and the third pose transformation data.
[0021] Optionally, the third pose transformation data includes: an initial rotation matrix and an initial translation vector; the static point cloud includes a first static point cloud of the first frame point cloud data and a second static point cloud of the second frame point cloud data; the first frame point cloud data is the point cloud data preceding the second frame point cloud data; and determining the second pose transformation data based on the static point clouds corresponding to each frame and the third pose transformation data includes:
[0022] Based on the initial rotation matrix and the initial translation vector, determine the first point cloud distance between the first static point cloud and the second static point cloud;
[0023] The initial rotation matrix and the initial translation vector are iterated at least once by minimizing the objective function to obtain the iterated rotation matrix and the iterated translation vector; wherein, the second point cloud distance is less than the first point cloud distance, and the second point cloud distance is the point cloud distance between the first static point cloud and the second static point cloud determined by the iterated rotation matrix and the iterated translation vector;
[0024] The iterated rotation matrix and the iterated translation vector are used as the second pose transformation data.
[0025] Optionally, determining the third pose transformation data based on the ground point cloud corresponding to each frame includes:
[0026] The ground point cloud corresponding to each frame is fitted to obtain the plane equation and normal vector corresponding to the ground point cloud of each frame.
[0027] Based on the plane equation and the normal vector, determine the pitch angle change, roll angle change, and altitude change;
[0028] The third pose transformation data is determined based on the changes in pitch angle, roll angle, and altitude.
[0029] Optionally, the method further includes:
[0030] Acquire the current image frame captured by the camera and the current point cloud data captured by the radar;
[0031] Obtain the reference image frame and reference point cloud data;
[0032] Based on the current image frame and the reference image frame, the first pose error of the camera is determined;
[0033] Based on the current point cloud data and the reference point cloud data, the second pose error of the radar is determined;
[0034] If the first pose error is greater than or equal to the fourth preset value, a pose error warning for the camera is generated.
[0035] If the second pose error is greater than or equal to the fifth preset value, a pose error warning for the radar is generated.
[0036] Secondly, a pose determination device is provided, comprising:
[0037] The first acquisition module is used to acquire at least two consecutive image frames captured by the camera and at least two consecutive point cloud data frames collected by the radar.
[0038] The first determining module is used to determine the depth information and feature descriptors corresponding to each image frame in the at least two consecutive image frames;
[0039] The second determining module is used to determine the first pose transformation data corresponding to the at least two consecutive image frames based on the depth information and feature descriptors corresponding to each image frame;
[0040] The third determining module is used to determine the ground point cloud corresponding to each frame in the at least two consecutive frames of point cloud data based on the at least two consecutive frames of point cloud data.
[0041] The fourth determining module is used to determine the second pose transformation data corresponding to the at least two consecutive frames of point cloud data based on the ground point cloud corresponding to each frame.
[0042] The fifth determining module is used to determine the pose based on the first pose transformation data and / or the second pose transformation data.
[0043] Thirdly, an electronic device is provided, comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the pose determination method as described in the first aspect.
[0044] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the pose determination method as described in the first aspect.
[0045] Fifthly, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps of the pose determination method as described in the first aspect.
[0046] In this embodiment, the electronic device can acquire at least two consecutive image frames captured by a camera and at least two consecutive frames of point cloud data acquired by a radar. The electronic device then determines the depth information and feature descriptors corresponding to each image frame in the at least two consecutive image frames, and then determines the first pose transformation data corresponding to the at least two consecutive image frames based on the depth information and feature descriptors. The electronic device then determines the ground point cloud corresponding to each frame in the at least two consecutive point cloud data, and then determines the second pose transformation data corresponding to the at least two consecutive point cloud data based on the ground point cloud corresponding to each frame. The electronic device then determines the pose using the first pose transformation data and / or the second pose transformation data. In this process, the electronic device can use both the camera and radar to determine the pose simultaneously, and the pose determination processes of the camera and radar are completely decoupled, without affecting or interacting with each other. Even if one of the camera or radar malfunctions, the other can be used to determine the pose, avoiding pose errors and improving the accuracy of pose determination. Attached Figure Description
[0047] Figure 1 This is a flowchart of a pose determination method provided in an embodiment of this application;
[0048] Figure 2 This is a schematic diagram illustrating the principle of a camera pose determination method provided in an embodiment of this application;
[0049] Figure 3 This is a flowchart of a pose error early warning method provided in an embodiment of this application;
[0050] Figure 4 This is a schematic diagram of the structure of a pose determination device provided in an embodiment of this application;
[0051] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0052] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0053] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and are not used to describe a specified order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, a first object can be one or more. Furthermore, in the specification and claims, "and" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0054] It is worth noting that the technologies described in this application are not limited to Long Term Evolution (LTE) / LTE-Advanced (LTE-A) systems, but can also be used in other wireless communication systems, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single-carrier Frequency-Division Multiple Access (SC-FDMA), and other systems. The terms "system" and "network" in this application are often used interchangeably, and the described technologies can be used with the systems and radio technologies mentioned above, as well as with other systems and radio technologies. However, the following description describes New Radio (NR) systems for illustrative purposes, and NR terminology is used in most of the following description. These technologies can also be applied to applications beyond NR systems, such as 6th generation (6G) radio systems. th Generation 6G communication system.
[0055] pass Figure 1 The process of the pose determination method of this application is described in detail below.
[0056] Step 101: Acquire at least two consecutive image frames captured by the camera and at least two consecutive point cloud data frames acquired by the radar.
[0057] In some embodiments, the electronic device can acquire consecutive image frames captured by the camera, preprocess each image frame sequentially, and sort and store the image frames according to their timestamps. A process lock can be added when the electronic device acquires image frames to ensure sequential processing and avoid mutual exclusion.
[0058] The process of preprocessing an image frame by an electronic device can be as follows: the electronic device normalizes the image so that the pixel value range of the image is scaled to [0, 1], and the channel format and size of the image are unified.
[0059] After obtaining at least two consecutive preprocessed image frames, the electronic device can determine the first pose transformation data based on two or more image frames. The following explanation uses the example of the electronic device determining the first pose transformation data based on two image frames.
[0060] The electronic device can also acquire at least two consecutive frames of point cloud data from radar and determine the second pose transformation data based on these two frames. The following explanation uses the example of the electronic device determining the second pose transformation data based on two frames of point cloud data.
[0061] Step 102: Determine the depth information and feature descriptors corresponding to each image frame in the at least two consecutive image frames.
[0062] like Figure 2 As shown, the electronic device can input two image frames into the label-free distillation (DINOv2) encoder, and after passing through the Transformer network, obtain depth information and feature descriptors from the depth output head and feature output head, respectively.
[0063] In this embodiment, by introducing an end-to-end visual Transformer encoding structure, direct prediction from continuous image pairs to relative pose transformations is achieved, improving the modeling capability for large viewpoint changes, illumination changes, and scene complexity at the overall network structure level. Furthermore, a global feature modeling approach based on a self-attention mechanism is adopted, enabling the model to simultaneously focus on local details and global spatial relationships, providing a more stable and consistent feature representation for subsequent geometric reconstruction and pose estimation.
[0064] Step 103: Determine the first pose transformation data corresponding to the at least two consecutive image frames based on the depth information and feature descriptors corresponding to each image frame.
[0065] In some embodiments, determining the first pose transformation data corresponding to the at least two consecutive image frames based on the depth information and feature descriptors corresponding to each image frame includes:
[0066] Based on the depth information corresponding to each image frame, determine the three-dimensional coordinates of at least one feature point corresponding to each image frame;
[0067] Based on the feature descriptors corresponding to each image frame, determine the similarity between the three-dimensional coordinates of at least one feature point corresponding to at least two image frames;
[0068] Determine a three-dimensional feature point pair, wherein the three-dimensional feature point pair is a feature point pair between at least two image frames whose similarity is greater than or equal to a first preset value;
[0069] The first pose transformation data is determined based on the three-dimensional feature point pairs.
[0070] For example, the electronic device inputs image frame 1 and image frame 2 into the DINOv2 encoder to obtain depth information depth_1 and feature descriptor features_1 corresponding to image frame 1, and depth information depth_2 and feature descriptor features_2 corresponding to image frame 2. The electronic device then calculates the depth information depth_1 using a pinhole camera model to obtain the three-dimensional coordinates of each feature point in space corresponding to image frame 1, and calculates the depth information depth_2 using the pinhole camera model to obtain the three-dimensional coordinates of each feature point in space corresponding to image frame 2. The electronic device then uses a nearest neighbor distance ratio algorithm and a bidirectional cross-validation algorithm, along with feature descriptor features_1 and feature descriptor features_2, to calculate the similarity between the three-dimensional coordinates. The electronic device uses feature point pairs with a similarity greater than or equal to a first preset value as three-dimensional feature point pairs. The electronic device then calculates the first pose transformation data corresponding to the three-dimensional feature point pairs using a least squares algorithm, etc., and this first pose transformation data includes a first rotation matrix and a first translation vector.
[0071] In this embodiment, depth information and geometric structural features are decoupled, modeled, and fused. Specifically, the network predicts pixel-level depth information and corresponding high-dimensional feature descriptors for the image, where the depth branch focuses on characterizing the scale and spatial hierarchy of the scene, while the feature branch expresses the discriminative matching relationships between pixels. Furthermore, by explicitly combining depth information and feature descriptors, feature points in the two-dimensional image are mapped to three-dimensional spatial points with real-scale constraints, thereby obtaining the three-dimensional coordinates of the feature points in each frame.
[0072] Furthermore, by independently predicting and fusing depth features and geometric structural features, three-dimensional spatial constraints are explicitly introduced, enabling feature matching and pose solving to be established under a consistent geometric reference framework, thereby enhancing the stability and interpretability of pose estimation from an algorithmic perspective.
[0073] Before the electronic device determines the first pose transformation data, the process of determining the first pose transformation data can be trained using a loss function, so that the first pose transformation data calculated after training is more accurate.
[0074] The loss function can be: Where N is the total number of feature points, and p is the three-dimensional coordinate. It is the predicted camera pose transformation matrix. It is the actual camera pose. This represents a projection function that projects 3D points onto a 2D image.
[0075] In this embodiment, the Viewpoint Consistency Reprojection Error (VCRE), which is highly relevant to practical engineering applications, is further introduced into the model optimization objective, subjecting the network to explicit constraints on geometric projection consistency during the training phase. This optimization method can directly constrain the consistency of predicted pose in the image projection space, making the model output more in line with the actual needs of roadside online calibration and pose monitoring.
[0076] By combining the above-mentioned technical means, this application achieves stable estimation and continuous online monitoring of camera pose changes, significantly improving the robustness, reliability and engineering practicality of the system in long-term operation scenarios.
[0077] Step 104: Based on the at least two consecutive frames of point cloud data, determine the ground point cloud corresponding to each frame in the at least two consecutive frames.
[0078] Step 105: Based on the ground point cloud corresponding to each frame, determine the second pose transformation data corresponding to the at least two consecutive frames of point cloud data.
[0079] In some embodiments, determining the second pose transformation data corresponding to the at least two consecutive frames of point cloud data based on the ground point cloud corresponding to each frame includes:
[0080] Based on the ground point cloud corresponding to each frame, the third pose transformation data is determined, and the third pose transformation data is the initial pose transformation data.
[0081] Determine the static point cloud in the ground point cloud corresponding to each frame;
[0082] The second pose transformation data is determined based on the static point cloud corresponding to each frame and the third pose transformation data.
[0083] Optionally, determining the third pose transformation data based on the ground point cloud corresponding to each frame includes:
[0084] The ground point cloud corresponding to each frame is fitted to obtain the plane equation and normal vector corresponding to the ground point cloud of each frame.
[0085] Based on the plane equation and the normal vector, determine the pitch angle change, roll angle change, and altitude change;
[0086] The third pose transformation data is determined based on the changes in pitch angle, roll angle, and altitude.
[0087] In this embodiment, the electronic device can use a Random Sample Consensus (RANSAC) algorithm to determine the third pose transformation data. For example, the electronic device can use the RANSAC algorithm to perform ground plane fitting on each frame of point cloud data, and then calculate the pitch angle change based on the plane equation and corresponding normal vector obtained from fitting two frames of ground point cloud data. Roll angle change And the height change z, thereby calculating the third pose transformation data.
[0088] The third pose transformation data T can be characterized as:
[0089] .
[0090] The electronic device then determines the static point cloud in the ground point cloud corresponding to each frame, and determines the second pose transformation data based on the static point cloud corresponding to each frame and the third pose transformation data.
[0091] Optionally, the third pose transformation data includes: an initial rotation matrix and an initial translation vector; the static point cloud includes a first static point cloud of the first frame point cloud data and a second static point cloud of the second frame point cloud data; the first frame point cloud data is the point cloud data preceding the second frame point cloud data; and determining the second pose transformation data based on the static point clouds corresponding to each frame and the third pose transformation data includes:
[0092] Based on the initial rotation matrix and the initial translation vector, determine the first point cloud distance between the first static point cloud and the second static point cloud;
[0093] The initial rotation matrix and the initial translation vector are iterated at least once by minimizing the objective function to obtain the iterated rotation matrix and the iterated translation vector; wherein, the second point cloud distance is less than the first point cloud distance, and the second point cloud distance is the point cloud distance between the first static point cloud and the second static point cloud determined by the iterated rotation matrix and the iterated translation vector;
[0094] The iterated rotation matrix and the iterated translation vector are used as the second pose transformation data.
[0095] In this embodiment, the objective function to be minimized can be: .
[0096] Where N is the number of point clouds, This is the first static point cloud. This is the second static point cloud. Let be the rotation matrix before iteration, and t be the translation vector before iteration.
[0097] The first point cloud distance between the first and second static point clouds is determined by using the first static point cloud, the second static point cloud, the initial rotation matrix, and the initial translation vector. Then, the initial rotation matrix and initial translation vector are iteratively applied by minimizing the objective function until the iterated rotation matrix, the iterated translation vector, and the second static point cloud distance are smaller than the first point cloud distance. The electronic device then uses the iterated rotation matrix and iterated translation vector as the second pose transformation data.
[0098] In this embodiment, dynamic point cloud is removed by segmenting dynamic and static points to eliminate objects that change significantly in a short period of time (such as moving vehicles and pedestrians), while static point clouds are used to improve the stability of sensor calibration. Specifically, by introducing ground plane structure constraints, separating dynamic and static point clouds, and using direction correction techniques based on long-term traffic flow statistical characteristics, the problems of point cloud registration being easily affected by dynamic targets such as vehicles and pedestrians in dynamic traffic environments, relying on external reference point clouds or high-precision maps, and being difficult to achieve long-term stable online calibration are solved.
[0099] By fully utilizing the above-mentioned technical means and taking into account the stable ground plane structure and the long-term consistency of traffic flow direction in roadside scenarios, the radar can achieve phased decoupled estimation and online correction of pitch angle, roll angle, altitude and yaw angle by relying solely on its own time-series point cloud data without the need for external reference sensors or human intervention. This significantly improves the stability and robustness of 3D radar pose estimation in complex traffic scenarios.
[0100] Step 106: Determine the pose based on the first pose transformation data and / or the second pose transformation data.
[0101] In some embodiments, the process of the camera determining pose transformation data is completely decoupled from the process of the radar determining pose transformation data. The electronic device can acquire the pose transformation data of the radar or camera based on the radar's or camera's operating status, whether a malfunction has occurred, or whether the pose data is accurate. For example, assuming the camera malfunctions, the pose transformation data is inaccurate, or the operating status is poor, the electronic device can acquire the corresponding pose transformation data from the radar. That is, the radar and camera do not interfere with each other, and the electronic device can determine the pose based on either of the pose transformation data.
[0102] In some embodiments, the electronic device may also determine the pose by combining the pose transformation data of radar and camera. For example, the final pose may be determined primarily by radar pose transformation data and secondarily by camera pose transformation data, or it may be determined primarily by camera pose transformation data and secondarily by radar pose transformation data.
[0103] As described in steps 101-106 above, the electronic device can acquire at least two consecutive image frames captured by the camera and at least two consecutive frames of point cloud data collected by the radar. The electronic device then determines the depth information and feature descriptors corresponding to each image frame in the at least two consecutive image frames, and then determines the first pose transformation data corresponding to the at least two consecutive image frames based on the depth information and feature descriptors. The electronic device then determines the ground point cloud corresponding to each frame in the at least two consecutive point cloud data, and then determines the second pose transformation data corresponding to the at least two consecutive point cloud data based on the ground point cloud corresponding to each frame. The electronic device then determines the pose using the first pose transformation data and / or the second pose transformation data. In this process, the electronic device can use both the camera and radar to determine the pose simultaneously, and the pose determination processes of the camera and radar are completely decoupled, without affecting or interacting with each other. Even if one of the camera or radar malfunctions, the other can be used to determine the pose, avoiding pose errors and improving the accuracy of pose determination.
[0104] In some embodiments, the electronic device can also determine whether the camera pose transformation data is accurate by using the camera pose error, and determine whether the radar pose transformation data is accurate by using the radar pose error.
[0105] Optionally, the method further includes:
[0106] Acquire the current image frame captured by the camera and the current point cloud data captured by the radar;
[0107] Obtain the reference image frame and reference point cloud data;
[0108] Based on the current image frame and the reference image frame, the first pose error of the camera is determined;
[0109] Based on the current point cloud data and the reference point cloud data, the second pose error of the radar is determined;
[0110] If the first pose error is greater than or equal to the fourth preset value, a pose error warning for the camera is generated.
[0111] If the second pose error is greater than or equal to the fifth preset value, a pose error warning for the radar is generated.
[0112] In this embodiment, the electronic device can be implemented using the formula: Determine the pose error.
[0113] in, This represents the pose error between the current image frame and the reference image frame, or the pose error between the current point cloud data and the reference point cloud data. In other words, the pose error of the camera can be obtained using this formula, and the pose error of the radar can also be obtained using this formula.
[0114] Due to rotational error Translation error It consists of two parts. The rotation matrix of the current frame relative to the reference frame is obtained from pose estimation; Representing the rotation matrix The trace is the sum of its diagonal elements.
[0115] Among them, rotational error By formula: The calculated rotation error represents the minimum rotation angle between two frames, which can intuitively reflect the magnitude of the attitude deviation.
[0116] Translation error By formula: The calculation shows that... These are the translation vector components of the current frame relative to the reference frame, representing the displacement changes in the three coordinate axes.
[0117] Electronic devices then pass through rotational error With translation error Add them together to get the error. This allows us to determine the degree of pose offset, providing a basis for subsequent reference frame calibration, early warning judgment, and reference frame update.
[0118] After the electronic device determines the first pose error, if the first pose error is greater than or equal to the fourth preset value, then the camera pose has an error, and a pose error warning is generated. Similarly, after the electronic device determines the second pose error, if the second pose error is greater than or equal to the fifth preset value, then the radar pose has an error, and a pose error warning is generated.
[0119] In some embodiments, if the first pose error is less than the fourth preset value, the current image frame is used as the new reference image frame. Similarly, if the second pose error is less than the fifth preset value, the point cloud data corresponding to the current frame is used as the new reference point cloud data.
[0120] In some embodiments, assuming that pose error is detected multiple times consecutively and the pose error is greater than or equal to a sixth preset value, the reference image frame or reference point cloud data is reset, and the parameters of the radar or camera are updated.
[0121] The following is through Figure 3The calculation process for the above pose error is explained.
[0122] Step 301: Acquire real-time data from the sensor.
[0123] In some embodiments, real-time data includes: the current image frame from the camera or the current point cloud data from the radar.
[0124] Step 302: Determine the pose error.
[0125] In some embodiments, the electronic device determines the first pose error corresponding to the current image frame and the reference image frame, or determines the second pose error corresponding to the current point cloud data and the reference point cloud data.
[0126] Step 303: Determine whether the pose error is greater than or equal to the warning value.
[0127] Step 304: If the pose error is greater than or equal to the warning value, generate a pose error warning.
[0128] In some embodiments, if the first pose error is greater than or equal to a fourth preset value, then a camera pose error has occurred, and a pose error warning is generated. If the second pose error is greater than or equal to a fifth preset value, then a radar pose error has occurred, and a pose error warning is generated.
[0129] Step 305: If the pose error is less than the warning value, update the reference frame.
[0130] In some embodiments, if the first pose error is less than the fourth preset value, the current image frame is used as the new reference image frame. Similarly, if the second pose error is less than the fifth preset value, the point cloud data corresponding to the current frame is used as the new reference point cloud data.
[0131] Step 306: If the pose error is continuously greater than or equal to the update value, then reset the reference frame and update the sensor parameters.
[0132] In some embodiments, assuming that pose error is detected multiple times consecutively and the pose error is greater than or equal to a sixth preset value, the reference image frame or reference point cloud data is reset, and the parameters of the radar or camera are updated.
[0133] Through the unified management and collaborative analysis of the aforementioned pose errors, an online self-calibration and monitoring system for multi-source sensors was constructed, enabling dynamic calibration and early warning of miscalibration of the extrinsic parameters of each sensor. Specifically, a pose error model was constructed by decoupling and fusing rotation and translation errors for modeling and evaluation. The rotation error is calculated based on the trace of the rotation matrix to determine the minimum rotation angle, accurately reflecting changes in sensor attitude. The translation error uses Euclidean distance to measure spatial displacement deviation. By uniformly measuring and summing these two errors, the algorithm simultaneously reflects attitude and position changes under a single metric, avoiding monitoring blind spots caused by relying on only a single error source. This error modeling method is particularly suitable for slow drift or local abrupt changes that occur in multi-sensor systems during long-term operation.
[0134] Building upon this foundation, a dual-threshold discrimination mechanism is introduced to serve both safety monitoring and the adaptive update needs of the electronic equipment. When the pose error exceeds a warning value (either the fifth or fourth preset value), the electronic equipment immediately triggers the warning mechanism to alert users to potential anomalies such as sensor loosening, collisions, or calibration failures, thereby ensuring the safety and reliability of the electronic equipment's operation. When the error does not exceed the warning threshold, the electronic equipment automatically performs external parameter calibration, achieving online parameter correction without manual intervention. This integrated "monitoring-calibration" design enables the electronic equipment to continuously maintain calibration accuracy without affecting normal operation.
[0135] Furthermore, to address the issue of environmental changes or the gradual invalidation of the initial reference frame in long-term operating scenarios, a time-consistent reference frame update strategy was implemented. By recording the historical sequence of errors and determining whether the error is stable and exceeds the reference frame update value over multiple consecutive frames, the electronic device can distinguish between "instantaneous noise" and "systematic changes," updating the reference frame and its corresponding extrinsic parameters only when the latter occurs. This strategy effectively avoids the instability caused by frequent reference frame updates while enhancing the electronic device's adaptability to scene changes and sensor aging.
[0136] This application is also particularly applicable to scenarios with extremely high requirements for external parameter accuracy and system stability, such as autonomous vehicles, mobile robots, and multi-sensor fusion systems. It can continuously monitor the status of external parameters during actual operation and achieve self-calibration and dynamic update of reference frames while ensuring safety, significantly reducing maintenance costs and improving the robustness and reliability of the system in long-term operation.
[0137] As can be seen from the above embodiments, this application addresses the problem of pose drift and difficulty in online calibration during the long-term stable operation of multiple roadside sensors. It constructs a unified technical solution for multi-sensor decoupling, phased implementation, online self-calibration, and monitoring. Specifically, the camera pose determination process and the radar pose determination process are independent in algorithm structure but work collaboratively at the system level, jointly serving sensor online calibration and out-of-calibration early warning. Furthermore, while sharing time-series data and early warning logic, the electronic equipment can be specifically improved based on the imaging mechanism and environmental sensitivity of different sensors, thereby achieving system-level robustness enhancement.
[0138] Furthermore, the roadside multi-source sensor online calibration method based on decoupled modeling adopts a decoupled and collaborative design paradigm. By constructing improved pose estimation algorithms for cameras and radars respectively, it achieves asynchronous parallel processing, fundamentally avoiding error coupling and mutual interference between sensors. The upper-level collaborative center introduces a fusion error evaluation model and a dual-threshold decision mechanism, which uniformly quantifies the rotation and translation errors estimated by decoupling into a comprehensive index. Based on preset warning and update values, it autonomously triggers external parameter calibration, system warnings, or reference frame replacement, realizing a leap from "passive maintenance" to "active perception and autonomous management," providing architectural guarantees for the long-term stable operation of electronic equipment.
[0139] Please refer to Figure 4 , Figure 4 This is a schematic diagram of a pose determination device according to an embodiment of this application. The pose determination device 400 includes:
[0140] The first acquisition module 401 is used to acquire at least two consecutive image frames captured by the camera and at least two consecutive point cloud data frames collected by the radar.
[0141] The first determining module 402 is used to determine the depth information and feature descriptors corresponding to each image frame in the at least two consecutive image frames;
[0142] The second determining module 403 is used to determine the first pose transformation data corresponding to the at least two consecutive image frames based on the depth information and feature descriptors corresponding to each image frame;
[0143] The third determining module 404 is used to determine the ground point cloud corresponding to each frame in the at least two consecutive frames of point cloud data based on the at least two consecutive frames of point cloud data.
[0144] The fourth determining module 405 is used to determine the second pose transformation data corresponding to the at least two consecutive frames of point cloud data based on the ground point cloud corresponding to each frame.
[0145] The fifth determining module 406 is used to determine the pose based on the first pose transformation data and / or the second pose transformation data.
[0146] Optionally, the second determining module 403 may further include:
[0147] The first determining unit is used to determine the three-dimensional coordinates of at least one feature point corresponding to each image frame based on the depth information corresponding to each image frame.
[0148] The second determining unit is used to determine the similarity between the three-dimensional coordinates of at least one feature point corresponding to at least two image frames based on the feature descriptors corresponding to each image frame;
[0149] The third determining unit is used to determine a three-dimensional feature point pair, wherein the three-dimensional feature point pair is a feature point pair between at least two image frames whose similarity is greater than or equal to a first preset value;
[0150] The fourth determining unit is used to determine the first pose transformation data based on the three-dimensional feature point pairs.
[0151] Optionally, the fourth determining module 405 may also include:
[0152] The fifth determining unit is used to determine the third pose transformation data based on the ground point cloud corresponding to each frame, wherein the third pose transformation data is the initial pose transformation data.
[0153] The sixth determining unit is used to determine the static point cloud in the ground point cloud corresponding to each frame;
[0154] The seventh determining unit is used to determine the second pose transformation data based on the static point cloud corresponding to each frame and the third pose transformation data.
[0155] Optionally, the third pose transformation data includes: an initial rotation matrix and an initial translation vector; the static point cloud includes a first static point cloud of the first frame point cloud data and a second static point cloud of the second frame point cloud data; the first frame point cloud data is the point cloud data preceding the second frame point cloud data; the seventh determining unit may further include:
[0156] The first determining subunit is used to determine the first point cloud distance between the first static point cloud and the second static point cloud based on the initial rotation matrix and the initial translation vector;
[0157] An iterative subunit is used to perform at least one iteration on the initial rotation matrix and the initial translation vector by minimizing the objective function to obtain the iterated rotation matrix and the iterated translation vector; wherein, the second point cloud distance is less than the first point cloud distance, and the second point cloud distance is the point cloud distance between the first static point cloud and the second static point cloud determined by the iterated rotation matrix and the iterated translation vector;
[0158] A sub-unit is generated to use the iterated rotation matrix and the iterated translation vector as the second pose transformation data.
[0159] Optionally, the fifth determining unit may also include:
[0160] The fitting subunit is used to fit the ground point cloud corresponding to each frame to obtain the plane equation and normal vector of the ground point cloud corresponding to each frame.
[0161] The second determining subunit is used to determine the pitch angle change, roll angle change, and altitude change based on the plane equation and the normal vector;
[0162] The third determining subunit is used to determine the third pose transformation data based on the pitch angle change, the roll angle change, and the height change.
[0163] Optionally, the pose determination device 400 also includes:
[0164] The second acquisition module is used to acquire the current image frame captured by the camera and the current point cloud data captured by the radar;
[0165] The third acquisition module is used to acquire reference image frames and reference point cloud data;
[0166] The sixth determining module is used to determine the first pose error of the camera based on the current image frame and the reference image frame;
[0167] The seventh determining module is used to determine the second pose error of the radar based on the current point cloud data and the reference point cloud data;
[0168] The first generation module is used to generate a camera pose error warning when the first pose error is greater than or equal to a fourth preset value.
[0169] The second generation module is used to generate a pose error warning for the radar when the second pose error is greater than or equal to a fifth preset value.
[0170] The pose determination device 400 provided in this application embodiment can perform the above-described... Figure 1 The method embodiments shown are similar in principle and technical effect, and will not be described again here.
[0171] This application also provides an electronic device. Since the principle by which this electronic device solves the problem is similar to the pose determination method in this application, the implementation of this electronic device can be found elsewhere. Figure 1 The implementation of the method shown will not be repeated here. Figure 5 As shown, the electronic device of this application embodiment includes: a processor 510, configured to read a program from a memory 520 and execute the following processes:
[0172] Acquire at least two consecutive image frames captured by the camera, and at least two consecutive point cloud data frames acquired by the radar;
[0173] Determine the depth information and feature descriptors corresponding to each image frame in the at least two consecutive image frames;
[0174] The first pose transformation data corresponding to the at least two consecutive image frames is determined based on the depth information and feature descriptors corresponding to each image frame;
[0175] Based on the at least two consecutive frames of point cloud data, determine the ground point cloud corresponding to each frame in the at least two consecutive frames;
[0176] Based on the ground point cloud corresponding to each frame, determine the second pose transformation data corresponding to the at least two consecutive frames of point cloud data;
[0177] The pose is determined based on the first pose transformation data and / or the second pose transformation data.
[0178] Optionally, the processor 510 is further configured to read the program in the memory 520 and perform the following steps: determining the first pose transformation data corresponding to the at least two consecutive image frames based on the depth information and feature descriptors corresponding to each image frame includes:
[0179] Based on the depth information corresponding to each image frame, determine the three-dimensional coordinates of at least one feature point corresponding to each image frame;
[0180] Based on the feature descriptors corresponding to each image frame, determine the similarity between the three-dimensional coordinates of at least one feature point corresponding to at least two image frames;
[0181] Determine a three-dimensional feature point pair, wherein the three-dimensional feature point pair is a feature point pair between at least two image frames whose similarity is greater than or equal to a first preset value;
[0182] The first pose transformation data is determined based on the three-dimensional feature point pairs.
[0183] Optionally, the processor 510 is further configured to read the program in the memory 520 and perform the following steps: determining the second pose transformation data corresponding to the at least two consecutive frames of point cloud data based on the ground point cloud corresponding to each frame, including:
[0184] Based on the ground point cloud corresponding to each frame, the third pose transformation data is determined, and the third pose transformation data is the initial pose transformation data.
[0185] Determine the static point cloud in the ground point cloud corresponding to each frame;
[0186] The second pose transformation data is determined based on the static point cloud corresponding to each frame and the third pose transformation data.
[0187] Optionally, the third pose transformation data includes: an initial rotation matrix and an initial translation vector; the static point cloud includes a first static point cloud of the first frame point cloud data and a second static point cloud of the second frame point cloud data; the first frame point cloud data is the point cloud data preceding the second frame point cloud data; the processor 510 is further configured to read the program in the memory 520 and execute the following steps: determining the second pose transformation data based on the static point clouds corresponding to each frame and the third pose transformation data includes:
[0188] Based on the initial rotation matrix and the initial translation vector, determine the first point cloud distance between the first static point cloud and the second static point cloud;
[0189] The initial rotation matrix and the initial translation vector are iterated at least once by minimizing the objective function to obtain the iterated rotation matrix and the iterated translation vector; wherein, the second point cloud distance is less than the first point cloud distance, and the second point cloud distance is the point cloud distance between the first static point cloud and the second static point cloud determined by the iterated rotation matrix and the iterated translation vector;
[0190] The iterated rotation matrix and the iterated translation vector are used as the second pose transformation data.
[0191] Optionally, the processor 510 is also used to read the program in the memory 520 and perform the following steps: determining the third pose transformation data based on the ground point cloud corresponding to each frame includes:
[0192] The ground point cloud corresponding to each frame is fitted to obtain the plane equation and normal vector corresponding to the ground point cloud of each frame.
[0193] Based on the plane equation and the normal vector, determine the pitch angle change, roll angle change, and altitude change;
[0194] The third pose transformation data is determined based on the changes in pitch angle, roll angle, and altitude.
[0195] Optionally, the processor 510 is also used to read the program from the memory 520 and perform the following steps:
[0196] Acquire the current image frame captured by the camera and the current point cloud data captured by the radar;
[0197] Obtain the reference image frame and reference point cloud data;
[0198] Based on the current image frame and the reference image frame, the first pose error of the camera is determined;
[0199] Based on the current point cloud data and the reference point cloud data, the second pose error of the radar is determined;
[0200] If the first pose error is greater than or equal to the fourth preset value, a pose error warning for the camera is generated.
[0201] If the second pose error is greater than or equal to the fifth preset value, a pose error warning for the radar is generated.
[0202] Among them, Figure 5 In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits of one or more processors represented by processor 510 and memory represented by memory 520 together. The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides the interface.
[0203] The electronic device provided in this application embodiment can perform the above-described functions. Figure 1 The method embodiments shown are similar in principle and technical effect, and will not be described again here.
[0204] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described pose determination method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0205] This application also provides a computer program product, including computer instructions. When executed by a processor, the computer instructions implement the various processes of the above-described pose determination method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0206] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0207] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0208] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A pose determination method, characterized in that, include: Acquire at least two consecutive image frames captured by the camera, and at least two consecutive point cloud data frames acquired by the radar; Determine the depth information and feature descriptors corresponding to each image frame in the at least two consecutive image frames; The first pose transformation data corresponding to the at least two consecutive image frames is determined based on the depth information and feature descriptors corresponding to each image frame; Based on the at least two consecutive frames of point cloud data, determine the ground point cloud corresponding to each frame in the at least two consecutive frames; Based on the ground point cloud corresponding to each frame, determine the second pose transformation data corresponding to the at least two consecutive frames of point cloud data; The pose is determined based on the first pose transformation data and / or the second pose transformation data.
2. The method according to claim 1, characterized in that, Determining the first pose transformation data corresponding to at least two consecutive image frames based on the depth information and feature descriptors corresponding to each image frame includes: Based on the depth information corresponding to each image frame, determine the three-dimensional coordinates of at least one feature point corresponding to each image frame; Based on the feature descriptors corresponding to each image frame, determine the similarity between the three-dimensional coordinates of at least one feature point corresponding to at least two image frames; Determine a three-dimensional feature point pair, wherein the three-dimensional feature point pair is a feature point pair between at least two image frames whose similarity is greater than or equal to a first preset value; The first pose transformation data is determined based on the three-dimensional feature point pairs.
3. The method according to claim 1, characterized in that, The step of determining the second pose transformation data corresponding to at least two consecutive frames of point cloud data based on the ground point cloud corresponding to each frame includes: Based on the ground point cloud corresponding to each frame, the third pose transformation data is determined, and the third pose transformation data is the initial pose transformation data. Determine the static point cloud in the ground point cloud corresponding to each frame; The second pose transformation data is determined based on the static point cloud corresponding to each frame and the third pose transformation data.
4. The method according to claim 3, characterized in that, The third pose transformation data includes: an initial rotation matrix and an initial translation vector. The static point cloud includes a first static point cloud of the first frame point cloud data and a second static point cloud of the second frame point cloud data. The first frame point cloud data is the point cloud data preceding the second frame point cloud data. Determining the second pose transformation data based on the static point clouds corresponding to each frame and the third pose transformation data includes: Based on the initial rotation matrix and the initial translation vector, determine the first point cloud distance between the first static point cloud and the second static point cloud; The initial rotation matrix and the initial translation vector are iterated at least once by minimizing the objective function to obtain the iterated rotation matrix and the iterated translation vector; wherein, the second point cloud distance is less than the first point cloud distance, and the second point cloud distance is the point cloud distance between the first static point cloud and the second static point cloud determined by the iterated rotation matrix and the iterated translation vector; The iterated rotation matrix and the iterated translation vector are used as the second pose transformation data.
5. The method according to claim 3, characterized in that, The step of determining the third pose transformation data based on the ground point cloud corresponding to each frame includes: The ground point cloud corresponding to each frame is fitted to obtain the plane equation and normal vector corresponding to the ground point cloud of each frame. Based on the plane equation and the normal vector, determine the pitch angle change, roll angle change, and altitude change; The third pose transformation data is determined based on the changes in pitch angle, roll angle, and altitude.
6. The method according to claim 1, characterized in that, The method further includes: Acquire the current image frame captured by the camera and the current point cloud data captured by the radar; Obtain the reference image frame and reference point cloud data; Based on the current image frame and the reference image frame, the first pose error of the camera is determined; Based on the current point cloud data and the reference point cloud data, the second pose error of the radar is determined; If the first pose error is greater than or equal to the fourth preset value, a pose error warning for the camera is generated. If the second pose error is greater than or equal to the fifth preset value, a pose error warning for the radar is generated.
7. A pose determination device, characterized in that, include: The first acquisition module is used to acquire at least two consecutive image frames captured by the camera and at least two consecutive point cloud data frames collected by the radar. The first determining module is used to determine the depth information and feature descriptors corresponding to each image frame in the at least two consecutive image frames; The second determining module is used to determine the first pose transformation data corresponding to the at least two consecutive image frames based on the depth information and feature descriptors corresponding to each image frame; The third determining module is used to determine the ground point cloud corresponding to each frame in the at least two consecutive frames of point cloud data based on the at least two consecutive frames of point cloud data. The fourth determining module is used to determine the second pose transformation data corresponding to the at least two consecutive frames of point cloud data based on the ground point cloud corresponding to each frame. The fifth determining module is used to determine the pose based on the first pose transformation data and / or the second pose transformation data.
8. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the pose determination method as described in any one of claims 1 to 6.
9. A computer-readable storage medium for storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps in the pose determination method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps in the pose determination method as described in any one of claims 1 to 6.