A method and device for estimating state information
By tracking and associating feature points of multiple image acquisition devices, combining IMU data, constructing state transfer and reprojection error equations, and using Kalman filtering theory to fuse sensor data, the accuracy and robustness problems of state information estimation in multi-sensor systems are solved, thereby improving the safety and accurate positioning capabilities of autonomous vehicles.
Patent Information
- Application Number
- CN202110217283.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-26
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2041-02-26
AI Technical Summary
Existing filtering schemes cannot effectively estimate vehicle state information of a multi-sensor system containing any number of cameras, affecting the safe driving of autonomous vehicles.
By tracking and associating feature points of multiple image acquisition devices and combining IMU data, the state transfer equation and reprojection error equation are constructed, and the sensor data are fused using Kalman filtering theory to estimate the state information of the target object.
It achieves high-precision and more robust state estimation of multi-sensor systems containing any number of image acquisition devices, improving the safety and accurate positioning capabilities of autonomous vehicles.
Smart Images

Figure CN114964217B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automation technology, and in particular to a state information estimation method and device. Background Art
[0002] In autonomous vehicle and robot mapping and positioning solutions, accurately and reliably estimating the motion state of a vehicle or robot using various sensors is a crucial issue. The system that estimates the motion state of the target vehicle or robot is called the robot's state estimator.
[0003] In related technologies, methods for estimating target motion state information are generally divided into two approaches: filtering and optimization. Filtering solutions offer better real-time performance and higher accuracy, making them easier to deploy in lightweight solutions for autonomous vehicles and robotic motion.
[0004] At present, in order to ensure the safe driving of autonomous vehicles, autonomous vehicles are often equipped with multiple cameras facing different directions to capture the surrounding environment of the vehicle, obtain surrounding environment information, and then combine the sensor data collected by other multiple sensors for positioning to ensure the accurate positioning of the vehicle and thus ensure the safe driving of the vehicle.
[0005] Current filtering solutions do not have the ability to estimate vehicle state information for a multi-sensor system containing an arbitrary number of cameras. Therefore, how to provide a method for estimating vehicle state information for a multi-sensor system containing an arbitrary number of cameras becomes an urgent problem to be solved. Summary of the Invention
[0006] The present invention provides a state information estimation method and apparatus to estimate the state information of an object in a multi-sensor system comprising any number of image acquisition devices. The specific technical solution is as follows:
[0007] In a first aspect, an embodiment of the present invention provides a state information estimation method, the method comprising:
[0008] Obtaining a current image captured by a multi-image acquisition device and current sensor data captured by other sensors set for the target object at the current moment, wherein the other sensors include an IMU;
[0009] Determining initial state information of the target object at the current moment by using IMU data corresponding to a moment before the current moment and previous state information of the target object at the previous moment;
[0010] Using the feature points detected in each current image and the relative position relationship between the image acquisition devices corresponding to each current image, a matching point pair between each current image is determined as a first matching point pair corresponding to the current image;
[0011] Using the feature points detected in each current image and the feature points detected in the previous image, a matching point pair between each current image and the previous image is determined as a second matching point pair corresponding to the current image;
[0012] Determine the three-dimensional position information corresponding to each feature point to be used based on the first matching point pair and the second matching point pair corresponding to the current image and the first matching point pair and the second matching point pair corresponding to the images at N moments before the current moment;
[0013] Based on the three-dimensional position information, the image position information of each feature point to be used, the initial state information and the current sensor data, the current state information of the target object at the current moment is determined.
[0014] Optionally, the initial state information includes: initial velocity information and initial posture information, wherein the initial posture information includes: initial posture information and initial position information;
[0015] The step of determining the initial state information of the target object at the current moment by using the IMU data corresponding to the moment before the current moment and the previous state information of the target object at the previous moment includes:
[0016] Determine the angular velocity and acceleration information of the target object at the previous moment using IMU data corresponding to the previous moment;
[0017] Constructing a first state transfer equation using the angular velocity information at the previous moment and the previous posture information in the previous state information; determining the initial posture information of the target object at the current moment using the first state transfer equation;
[0018] constructing a second state transition equation using the previous posture information, the acceleration information at the previous moment, and the previous velocity information in the previous state information; and determining the initial velocity information of the target object at the current moment using the second state transition equation;
[0019] A third state transfer equation is constructed using the initial velocity information, the previous velocity information, and the previous position information in the previous state information; and the initial position information of the target object at the current moment is determined using the third state transfer equation.
[0020] Optionally, the step of determining the three-dimensional position information corresponding to each feature point to be utilized based on the first matching point pair and the second matching point pair corresponding to the current image and the first matching point pair and the second matching point pair corresponding to the images at N moments before the current moment includes:
[0021] According to the triangulation algorithm, the three-dimensional position information corresponding to each feature point to be used is determined based on the first matching point pair and the second matching point pair corresponding to the current image, the first matching point pair and the second matching point pair corresponding to the images at the previous N moments before the current moment, the device posture information of the image acquisition device corresponding to each current image, the device posture information of the image acquisition device corresponding to each image at the previous N moments, and the posture information of the target object corresponding to the current image and each image at the previous N moments.
[0022] Optionally, the step of determining the current state information of the target object at the current moment based on the three-dimensional position information, the image position information of each feature point to be utilized, the initial state information, and the current sensor data includes:
[0023] Determining intermediate state information of the target object at a current moment based on the current sensor data and the initial state information;
[0024] Determine projection position information of a projection point of a spatial point corresponding to each feature point to be utilized in the image in which the feature point is located by using the three-dimensional position information, the intermediate pose information in the intermediate state information, the object pose information in the state information of the target object corresponding to each image at each of the previous N moments, and the device pose information and intrinsic parameter matrix of each image acquisition device;
[0025] Constructing a reprojection error equation based on the projection position information corresponding to each feature point to be utilized and the image position information of each feature point to be utilized in the current image;
[0026] Based on the reprojection error equation, current state information of the target object at a current moment is determined.
[0027] Optionally, the step of determining the current state information of the target object at the current moment based on the reprojection error equation includes:
[0028] Based on the reprojection error equation, construct a target measurement equation;
[0029] The target measurement equation and the filter update equation are used to determine the current state information of the target object at the current moment.
[0030] In a second aspect, an embodiment of the present invention provides a state information estimation device, the device comprising:
[0031] A first acquisition module is configured to obtain a current image captured by a multi-image acquisition device set on the target object at a current moment and current sensor data captured by other sensors, wherein the other sensors include an IMU;
[0032] A first determining module is configured to determine initial state information of the target object at the current moment by using IMU data corresponding to a moment before the current moment and previous state information of the target object at the previous moment;
[0033] A second determination module is configured to determine matching point pairs between the current images as first matching point pairs corresponding to the current images by using feature points detected in the current images and relative positional relationships between image acquisition devices corresponding to the current images;
[0034] a third determining module configured to determine, by using the feature points detected in each current image and the feature points detected in the previous image, a pair of matching points between each current image and its previous image as a second matching point pair corresponding to the current image;
[0035] A fourth determination module is configured to determine the three-dimensional position information corresponding to each feature point to be used based on the first matching point pair and the second matching point pair corresponding to the current image and the first matching point pair and the second matching point pair corresponding to the images at N moments before the current moment;
[0036] The fifth determination module is configured to determine the current state information of the target object at a current moment based on the three-dimensional position information, the image position information of each feature point to be used, the initial state information and the current sensor data.
[0037] Optionally, the initial state information includes: initial velocity information and initial posture information, wherein the initial posture information includes: initial posture information and initial position information;
[0038] The first determining module is specifically configured to determine the angular velocity information and acceleration information of the target object at the previous moment by using the IMU data corresponding to the previous moment;
[0039] Constructing a first state transfer equation using the angular velocity information at the previous moment and the previous posture information in the previous state information; determining the initial posture information of the target object at the current moment using the first state transfer equation;
[0040] constructing a second state transition equation using the previous posture information, the acceleration information at the previous moment, and the previous velocity information in the previous state information; and determining the initial velocity information of the target object at the current moment using the second state transition equation;
[0041] A third state transfer equation is constructed using the initial velocity information, the previous velocity information, and the previous position information in the previous state information; and the initial position information of the target object at the current moment is determined using the third state transfer equation.
[0042] Optionally, the fourth determination module is specifically configured to determine the three-dimensional position information corresponding to each feature point to be used based on the first matching point pair and the second matching point pair corresponding to the current image, the first matching point pair and the second matching point pair corresponding to the images at the previous N moments before the current moment, the device posture information of the image acquisition device corresponding to each current image, the device posture information of the image acquisition device corresponding to each image at the previous N moments, and the posture information of the target object corresponding to the current image and each image at the previous N moments in accordance with the triangulation algorithm.
[0043] Optionally, the fifth determining module includes:
[0044] a first determining unit configured to determine intermediate state information of the target object at a current moment based on the current sensor data and the initial state information;
[0045] The second determining unit is configured to determine projection position information of a projection point of a spatial point corresponding to each feature point to be utilized in the image in which the spatial point is located, using the three-dimensional position information, the intermediate pose information in the intermediate state information, the object pose information in the state information of the target object corresponding to each image at each of the previous N moments, and the device pose information and intrinsic parameter matrix of each image acquisition device;
[0046] a construction unit configured to construct a reprojection error equation based on projection position information corresponding to each feature point to be utilized and image position information of each feature point to be utilized in the current image;
[0047] The third determining unit is configured to determine the current state information of the target object at a current moment based on the reprojection error equation.
[0048] Optionally, the third determining unit is specifically configured to construct a target measurement equation based on the reprojection error equation;
[0049] The target measurement equation and the filter update equation are used to determine the current state information of the target object at the current moment.
[0050] As can be seen from the above content, an embodiment of the present invention provides a state information estimation method and device, which obtains a current image captured by a multi-image acquisition device set for a target object at the current moment and current sensor data captured by other sensors, wherein the other sensors include an IMU; uses the IMU data corresponding to the moment before the current moment and the previous state information of the target object at the previous moment to determine the initial state information of the target object at the current moment; uses the feature points detected by each current image and the relative posture relationship between the image acquisition devices corresponding to each current image to determine the matching point pairs between each current image as the first matching point pairs corresponding to the current image; uses the feature points detected by each current image and the feature points detected by the previous image to determine the matching point pairs between each current image and its previous image as the second matching point pairs corresponding to the current image; based on the first matching point pair and the second matching point pair corresponding to the current image and the first matching point pair and the second matching point pair corresponding to the images at the previous N moments before the current image, determines the three-dimensional position information corresponding to each feature point to be used; based on the three-dimensional position information, the image position information of each feature point to be used, the initial state information and the current sensor data, determines the current state information of the target object at the current moment.
[0051] By applying the embodiments of the present invention, the state can be expanded, that is, the tracking results of the feature points of each of the multiple image acquisition devices, i.e., the second matching point pair corresponding to the current image and its N previous time images, and the feature point association relationship between the multiple image acquisition devices, i.e., the first matching point pair corresponding to the current image and its N previous time images, can be used to construct the three-dimensional position information corresponding to each feature point to be used, so as to construct a multi-state constraint, fuse the current sensor data of other sensor data, and determine the current state information of the target object at the current moment, so as to obtain a state estimation result with higher accuracy and higher robustness, so as to realize the estimation of the state information of the object of the multi-sensor system including any number of image acquisition devices. Of course, it is not necessary to achieve all the advantages described above at the same time when implementing any product or method of the present invention.
[0052] The innovative features of the embodiments of the present invention include:
[0053] 1. By extending the state, that is, utilizing the tracking results of the feature points of each of the multiple image acquisition devices, i.e., the second matching point pair corresponding to the current image and its images N moments before, and the feature point association relationship between the multiple image acquisition devices, i.e., the first matching point pair corresponding to the current image and its images N moments before, the three-dimensional position information corresponding to each feature point to be used can be constructed to construct multi-state constraints, fuse the current sensor data of other sensor data, and determine the current state information of the target object at the current moment, so as to obtain a state estimation result with higher accuracy and higher robustness, so as to realize the estimation of the state information of the object of the multi-sensor system containing any number of image acquisition devices.
[0054] 2. Based on the error state Kalman filter, the IMU data at the previous moment and the previous state information of the target at the previous moment are used to construct the state transfer equation to obtain the initial state information of the target object at the current moment, providing a basis for the subsequent determination of the current state information of the target object at the current moment with higher accuracy and robustness.
[0055] 3. Considering that other sensor data arrive earlier than images, we first determine the intermediate state information of the target object at the current moment based on the current sensor data and initial state information that arrive earlier. Then, referring to the error state Kalman filter theory, we construct a reprojection error equation for the projection position information of the spatial points corresponding to each feature point to be used in the image and the image position information of each feature point to be used, and then construct a target measurement equation to optimize the intermediate state information and obtain current state information with high accuracy and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely some embodiments of the present invention. Those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0057] Figure 1 A schematic diagram of a flow chart of a state information estimation method provided by an embodiment of the present invention;
[0058] Figure 2 A schematic structural diagram of a state information estimation device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0059] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0060] It should be noted that the terms "including" and "having" and any variations thereof in the embodiments of the present invention and the accompanying drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.
[0061] The present invention provides a state information estimation method and apparatus to achieve state information estimation of an object in a multi-sensor system comprising any number of image acquisition devices.
[0062] Figure 1 A flow chart of a state information estimation method provided in an embodiment of the present invention. The method may include the following steps:
[0063] S101: Obtaining a current image captured by a multi-image acquisition device and current sensor data captured by other sensors set on a target object at a current moment.
[0064] Among them, other sensors include IMU.
[0065] The state information estimation method provided in the embodiments of the present invention can be applied to any electronic device with computing capabilities, which can be a terminal or a server. In one implementation, the functional software implementing the method can exist in the form of independent client software or as a plug-in for currently related client software, for example, in the form of a functional module of an autonomous driving system, which is possible. The electronic device can be a device installed on a target object or a device installed on a target object, which is possible.
[0066] In one case, the electronic device can be a multi-sensor state estimator, which inputs data from multiple sensors, such as data collected by multiple image acquisition devices, IMU (Inertial measurement unit), GNSS (Global Navigation Satellite System), wheel speed meter or wheel speed sensor, etc., and obtains various state information of the entire multi-sensor system, i.e., the target object for setting the multi-sensor system, under the assumption that all sensors are approximately rigid body transformations, mainly including position information, posture information, and speed information.
[0067] The target object can be an autonomous vehicle or an intelligent robot. The multi-image acquisition device provided for the target object can be a system of multiple image acquisition devices installed on the same rigid body of the target object in any number, position, and orientation.
[0068] During the movement of the target object, the multi-image acquisition device and other sensors can collect data in real time and send it to the electronic device, so that the electronic device obtains the image collected by the multi-image acquisition device set for the target object at the current moment as the current image; and obtains the sensor data collected by other sensors from the previous moment to the current moment as the current sensor data.
[0069] Other sensors may include, but are not limited to, IMU (Inertial Measurement Unit), wheel speed sensor, and GNSS (Global Navigation Satellite System), GPS (Global Positioning System), and radar.
[0070] In one scenario, the relative positions of the multiple image acquisition devices and the target object can be considered fixed, and the relative positions of the other sensors and the target object can be considered fixed. Accordingly, when the pose information of any one of the image acquisition devices, the target object, or the other sensors is determined, the pose information of the other objects is also determined.
[0071] S102: Determine the initial state information of the target object at the current moment by using the IMU data corresponding to the previous moment and the previous state information of the target object at the previous moment.
[0072] In this step, the electronic device can obtain the state information of the target object at the previous moment as the previous state information, and obtain the IMU data corresponding to the previous moment of the current moment. Based on the error-state Kalman filter (ESKF, error-state Kalman Filter), the IMU data corresponding to the previous moment and the previous state information are used to predict and determine the initial state information of the target object at the current moment. The state information may include but is not limited to: the position information and speed information of the target object, wherein the position information includes position information and attitude information. The IMU data corresponding to the previous moment is: the IMU data set by the target object from the previous moment before the current moment to the previous moment before the current moment.
[0073] The state information of the target vehicle can be predicted and determined by the following state transition equation. Specifically, it can be expressed by the following formula (1):
[0074]
[0075] Among them, t k-1 Indicates the moment before the current moment, t kIndicates the current moment, Indicates the previous state information, u k Indicates the IMU data corresponding to the moment before the current moment, The initial state information at the current moment.
[0076] In another embodiment of the present invention, the initial state information Includes: initial speed information And initial posture information, wherein the initial posture information includes: initial posture information And initial position information
[0077] The S102 may include the following steps 011-014:
[0078] 011: Use the IMU data from the moment before the current moment to determine the angular velocity information and acceleration information of the target object at the previous moment.
[0079] 012: Construct a first state transfer equation using the angular velocity information at the previous moment and the previous posture information in the previous state information; and determine the initial posture information of the target object at the current moment using the first state transfer equation.
[0080] 013: Use the previous posture information, the acceleration information at the previous moment, and the previous velocity information in the previous state information to construct a second state transfer equation; use the second state transfer equation to determine the initial velocity information of the target object at the current moment.
[0081] 014: Construct a third state transition equation using the initial velocity information, the previous velocity information, and the previous position information in the previous state information; and determine the initial position information of the target object at the current moment using the third state transition equation.
[0082] In this implementation, the electronic device debiases and removes gravity acceleration from the IMU data corresponding to the previous moment of the current moment, and obtains the angular velocity information and acceleration information of the target object at the previous moment, respectively using ω k-1 With α k-1 express.
[0083] Then, the angular velocity information at the previous moment and the previous posture information in the previous state information are used to construct a first state transition equation; the first state transition equation is used to determine the initial posture information of the target object at the current moment. Specifically, the first state transition equation can be expressed by the following formula (2):
[0084]
[0085] in, Indicates the previous posture information in the previous state information, Represents quaternion multiplication.
[0086] Furthermore, the electronic device constructs a second state transition equation using the previous posture information, the acceleration information at the previous moment, and the previous velocity information in the previous state information; and uses the second state transition equation to determine the initial velocity information of the target object at the current moment. Specifically, the second state transition equation can be expressed by the following formula (3):
[0087]
[0088] in, Indicates the previous speed information in the previous state information.
[0089] Furthermore, the electronic device constructs a third state transition equation using the initial velocity information, the previous velocity information, and the previous position information in the previous state information; and uses the third state transition equation to determine the initial position information of the target object at the current moment. Specifically, the third state transition equation can be expressed by the following formula (4):
[0090]
[0091] in, Indicates the previous position information in the previous state information.
[0092] S103: Using the feature points detected in each current image and the relative posture relationship between the image acquisition devices corresponding to each current image, determine the matching point pairs between each current image as the first matching point pairs corresponding to the current image.
[0093] The relative positional relationship between the image acquisition devices arranged on the target object is fixed. Accordingly, the electronic device may pre-store the relative positional relationship between the image acquisition devices.
[0094] After obtaining each current image, the electronic device may first use a preset feature point detection algorithm to detect feature points in each current image to obtain feature points in each current image. In one embodiment, the preset feature point detection algorithm may be a detection algorithm that can detect feature points in an image, such as a FAST feature point detection algorithm.
[0095] Furthermore, the electronic device determines the region of interest in each current image based on the relative posture relationship between the image acquisition devices corresponding to each current image, wherein the region of interest in the current image is the region that overlaps with other current images. For each current image, a preset feature descriptor extraction algorithm is used to extract feature descriptors for the feature points in the region of interest in the current image, and feature descriptors corresponding to each feature point in the region of interest in the current image are obtained. Based on the Fast Library for Approximate Nearest Neighbors algorithm and the feature descriptors corresponding to each feature point in the region of interest in each current image, the feature points in the region of interest in each current image are matched to obtain the feature point pairs, i.e., matching point pairs, between each current image as the first matching point pair corresponding to the current image. The first matching point pair corresponding to the current image, i.e., the k-th frame image, can be expressed as: c i ,c j represent the i-th image acquisition device and the j-th image acquisition device respectively, c i ,c j The value of is an integer between [1, n], where n represents the number of image acquisition devices set on the target vehicle.
[0096] The preset feature descriptor extraction algorithm may be a BRIER feature descriptor extraction algorithm.
[0097] S104: using the feature points detected in each current image and the feature points detected in the previous image, determining a matching point pair between each current image and its previous image as a second matching point pair corresponding to the current image.
[0098] In this step, the electronic device uses the feature points detected in the current image and the feature points detected in the previous image, and uses the sparse optical flow KLT algorithm to track the feature points on the current image and the previous image, and obtains a matching feature point pair between the current image and the previous image, that is, a matching point pair, as the second matching point pair corresponding to the current image. The second matching point pair corresponding to the current image can be expressed as: c i represents the i-th image acquisition device, c i The value range of is an integer between [1, n], where n represents the number of image acquisition devices set on the target vehicle.
[0099] The embodiment of the present invention does not limit the order of executing S104 and S103. The electronic device may execute S103 first and then S104, or execute S104 first and then S103, or execute S103 and S104 in parallel.
[0100] S105: Determine the three-dimensional position information corresponding to each feature point to be used based on the first matching point pair and the second matching point pair corresponding to the current image and the first matching point pair and the second matching point pair corresponding to the images at N moments before the current moment.
[0101] The images at N moments before the current moment refer to images captured by a multi-image capture device of the target object at each moment within N moments before the current moment.
[0102] After obtaining the first and second matching point pairs corresponding to the current image, the electronic device refers to a multi-state constrained Kalman filter (MSCKF). To establish multiple constraints, the electronic device expands the states in the MSCKF during the state information estimation process for the target object including multiple image acquisition devices. Specifically, the corresponding sliding window is first expanded, and the length of the sliding window is set to N+1. The sliding window contains the pose information in the state information of the target object, which is specifically expressed as follows:
[0103] x k =[π k ,π k-1 ,…π e …π k-N ], Integers in ;
[0104] Among them, x k Represents the state information in the sliding window corresponding to the current image; π k Indicates the pose information in the state information of the target corresponding to the current image. In one case, the state information of the IMU can be directly used to represent the state information of the target object. Accordingly, the state information of the target object contained in the sliding window is the pose information of the IMU set for the target object. In another case, the state information of the target object can be determined by using the state information of the IMU and the pre-stored pose conversion relationship between the IMU and the target object. Both are possible. e Right now Represents the pose information of the target object corresponding to the current image and the image at time e in the previous N time periods in the world coordinate system.
[0105] Accordingly, the first matching point pair and the second matching point pair corresponding to each image in the N images before the current moment are obtained; based on the first matching point pair and the second matching point pair corresponding to the current image, the first matching point pair and the second matching point pair corresponding to the images N before the current moment, the device pose information of the image acquisition device corresponding to the current image, and the device pose information of the image acquisition device corresponding to each image in the N images before the current moment, the three-dimensional position information corresponding to each feature point to be utilized is determined. The feature point to be utilized is: a feature point whose three-dimensional position information can be determined from the feature points detected in each current image and its N images before the current moment.
[0106] The process for determining the first matching point pair corresponding to each of the N images before the current moment is the same as the process for determining the first matching point pair corresponding to the current image. The process for determining the second matching point pair corresponding to each of the N images before the current moment is the same as the process for determining the second matching point pair corresponding to the current image, and is not further described here. N is a positive integer.
[0107] The device posture information of the image acquisition device corresponding to the image may refer to the device posture information when the image acquisition device acquires the image.
[0108] In one implementation of the present invention, S105 may include the following steps:
[0109] According to the triangulation algorithm, the three-dimensional position information corresponding to each feature point to be used is determined based on the first matching point pair and the second matching point pair corresponding to the current image, the first matching point pair and the second matching point pair corresponding to the images at the previous N moments before the current moment, the device posture information of the image acquisition device corresponding to each current image, the device posture information of the image acquisition device corresponding to each image at the previous N moments, and the posture information of the target object corresponding to the current image and each image at the previous N moments.
[0110] In one implementation, the relative positional relationship between the target object and the image acquisition device it is mounted on is determined, and the electronic device may pre-store the relative positional relationship between the target object and each of the image acquisition devices it is mounted on. When the positional information of the target object is determined, the positional information of each of the image acquisition devices is also determined.
[0111] The electronic device uses a triangulation algorithm to obtain a matching result for all feature points of the current image and its N-moment images corresponding to each position information in the sliding window corresponding to the current image, based on the image position information of each feature point in the first matching point pair on its image and the image position information of each feature point in the second matching point pair on its image, the image position information of each feature point in the first matching point pair on its image and the image position information of each feature point in the second matching point pair on its image corresponding to each image at the previous N moments of the current moment, the device posture information of the image acquisition device corresponding to each current image, the device posture information of the image acquisition device corresponding to each image at the previous N moments of the current moment, and the posture information of the target object corresponding to the current image and each image at the previous N moments of the current moment. Perform triangulation to determine the three-dimensional position information corresponding to each feature point to be used. The specific calculation process of the triangulation algorithm can refer to the calculation process of the triangulation algorithm in the related art, which will not be repeated here.
[0112] In one case, the three-dimensional position information corresponding to each feature point to be used, the image position information of each feature point to be used, and the device pose information and internal parameter matrix of the image acquisition device corresponding to the image where each feature point to be used is used to construct the minimum reprojection error. The Levenberg-Marquardt method is used to iteratively optimize the minimum reprojection error to obtain more accurate three-dimensional position information corresponding to each feature point to be used. Among them, the three-dimensional position information set corresponding to each feature point to be used can be expressed as in, Represents the three-dimensional position information of the spatial point corresponding to the qth feature point to be used in world coordinates.
[0113] The image position information of each feature point to be utilized is the image position information of each feature point to be utilized in the image in which it is located. The image in which each feature point to be utilized is located includes the current image and the images at the N moments before the current moment.
[0114] S106: Determine the current state information of the target object at the current moment based on the three-dimensional position information, the image position information of each feature point to be utilized in the current image, the initial state information and the current sensor data.
[0115] In this step, considering that the image acquisition device captures images at a lower frequency than other sensors capture other sensor data, the electronic device can first refer to the error state Kalman filter theory and, based on the current sensor data corresponding to the current moment and the initial state information first obtained by the electronic device, update the state information of the target object to obtain intermediate state information of the target object. Furthermore, referring to the error state Kalman filter theory, a corresponding measurement equation is constructed using the three-dimensional position information, the image position information of each feature point to be utilized in the current image, and the intermediate state information. Furthermore, based on this measurement equation, the current state information of the target object at the current moment is determined. This current state information can be state information in a world coordinate system.
[0116] By applying the embodiments of the present invention, the state can be expanded, that is, the tracking results of the feature points of each of the multiple image acquisition devices, that is, the second matching point pair corresponding to the current image and the images N moments before it, and the feature point association relationship between the multiple image acquisition devices, that is, the first matching point pair corresponding to the current image and the images N moments before it, are used to construct three-dimensional position information corresponding to each feature point to be used, so as to construct multi-state constraints, fuse the current sensor data of other sensor data, and determine the current state information of the target object at the current moment, so as to obtain a state estimation result with higher accuracy and higher robustness, so as to realize the estimation of the state information of an object in a multi-sensor system including any number of image acquisition devices.
[0117] In another embodiment of the present invention, the S106 may include the following steps 021-024:
[0118] 021: Based on the current sensor data and the initial state information, determine the intermediate state information of the target object at the current moment.
[0119] 022: Determine the projection position information of the projection point of the spatial point corresponding to each feature point to be used in the image in which it is located by using the three-dimensional position information, the intermediate posture information in the intermediate state information, the object posture information in the state information of the target object corresponding to each image at the previous N moments, and the device posture information and internal parameter matrix of each image acquisition device.
[0120] 023: Construct a reprojection error equation based on the projection position information corresponding to each feature point to be used and the image position information of each feature point to be used in the current image.
[0121] 024: Based on the reprojection error equation, determine the current state information of the target object at the current moment.
[0122] In one scenario, the current sensor data may include, but is not limited to, current IMU data collected by the IMU between the moment before the current moment and the current moment, current GNSS data collected by the GNSS between the moment before the current moment and the current moment, and current wheel speed data collected by the wheel speed sensor between the moment before the current moment and the current moment. Specifically, the electronic device iteratively updates the initial state information based on the current IMU data, current GNSS data, and current wheel speed data to obtain intermediate state information of the target object.
[0123] Here, referring to the error state Kalman filter theory, a measurement equation for the state update process of the target object using the above current sensor data is constructed. The measurement equation can be uniformly expressed by the following formula (5):
[0124]
[0125] Among them, z k Indicates the measurement value, that is, the current IMU data, current GNSS data or current wheel speed data, R k represents the measurement noise. h(·) represents the function that maps the system state to the measurement value, where the system state refers to the state information of the target object output by the system. When the state of the target object is updated using the current sensor data, the initial state information is the initial value of the system state for the state update.
[0126] Specifically, for GNSS measurements, z k represents the current GNSS data. Accordingly, the above formula (5) can be expressed as the following formula (5.1):
[0127]
[0128] in, Indicates a frame of GNSS data in the current GNSS data. Indicates that the frame GNSS data is substituted into the state information of the updated target object. The system state before the update system is the state information of the target object. GNSS This value represents the noise measured by the system when this frame of GNSS data is substituted into the status update system to update the status information of the target object.
[0129] For wheel speed measurement, z k represents the current wheel speed data. Accordingly, the above formula (5) can be expressed as the following formula (5.2):
[0130]
[0131] in, Indicates a frame of wheel speed data in the current wheel speed data. The wheel speed data of this frame is substituted into the state information of the updated target object, and the system state before the state update system is the state information of the target object. odo This represents the noise measured by the system when the wheel speed data for this frame is substituted into the state update system that updates the state information of the target object.
[0132] For IMU measurement, if the target object is stationary, that is, And if the IMU data measured by the IMU within the past preset time period, such as 1 second, indicates that the change in the angular velocity and acceleration of the target vehicle is less than a preset threshold, then the above formula (5) can be expressed as the following formula (5.3):
[0133]
[0134] Among them, 0 represents a frame of IMU data in the current IMU data, The system state before the state update system is replaced by the state information of the target object in the frame IMU data, that is, the state information of the target object, R static This value represents the noise measured by the system when this frame of IMU data is substituted into the state update system that updates the state information of the target object.
[0135] In one case, if the target object is not in a stationary state, the state information of the target object may not be updated using the current IMU data.
[0136] In this implementation, there is no limitation on the order of different quantities of updates when using current sensor data to update the state information of the target object. The state update system updates the state information of the target object using the type of current sensor data it obtains.
[0137] Subsequently, the filter update amount of the state update system can be calculated based on each of the above measurement equations. The update equation corresponding to the filter update amount is expressed by the following formula (6):
[0138]
[0139] Among them, P k|k-1 It represents the predicted state covariance, and the initial value is the initial state covariance. The initial state covariance can be expressed by the following formula (7):
[0140]
[0141] in, State transfer equation State quantity The derivative of Q kis the state transfer error, which is generally the noise parameter of IMU and is a constant.
[0142] K k Expressed as Kalman gain, it indicates the magnitude by which the current system state needs to be adjusted; Represents the current state information of the target object at the current moment, P k|k Represents the current state covariance at the current moment; That is, h(·) is the state quantity The derivative of . Represents generalized addition, including vector addition and rotational vector addition.
[0143] After the electronic device uses the current sensor data and initial state information to determine the intermediate state information of the target object at the current moment, it uses the posture information in the intermediate state information of the target object and the object posture information in the state information of the target object corresponding to the images at each moment of the previous N moments, the three-dimensional position information corresponding to each feature point to be used, the image position information of each feature point to be used, and the device posture information and internal parameter matrix of the image acquisition device corresponding to the image where each feature point to be used is located to construct a reprojection error equation, and then determine the current state information of the target object at the current moment based on the reprojection error equation.
[0144] In one case, the process of constructing the reprojection error equation can be:
[0145] For each spatial point corresponding to a feature point to be utilized, the spatial point corresponding to the feature point to be utilized is transformed from the world coordinate system to the target object coordinate system, and then to the device coordinate system of the image acquisition device corresponding to the image where the feature point to be utilized is located, using the three-dimensional position information of the spatial point corresponding to the feature point to be utilized, the object pose information in the state information of the target object corresponding to the image where the feature point to be utilized is located, and the device pose information of the image acquisition device corresponding to the image where the feature point to be utilized is located. Furthermore, in combination with the intrinsic parameter matrix of the image acquisition device corresponding to the image where the feature point to be utilized is used, the spatial point corresponding to the feature point to be utilized is projected from the device coordinate system of the image acquisition device corresponding to the image where the feature point to be utilized is located to the image coordinate system of the image where the feature point to be utilized is located, thereby obtaining the projection position information of the spatial point corresponding to the feature point to be utilized at the projection point of the image where the feature point to be utilized is located. Furthermore, for each feature point to be utilized, the image position information of the feature point to be utilized and the projection position information of the spatial point corresponding to the feature point to be utilized at the projection point of the image where the feature point to be utilized are used to construct a reprojection error equation.
[0146] Specifically, in theory, the projection position information of the spatial point corresponding to the feature point to be utilized on the projection point of the image where the feature point to be utilized is coincident with the image position of the feature point to be utilized on the image where the feature point to be utilized is located. Accordingly, the reprojection error equation can be expressed by the following formula (8):
[0147]
[0148] Indicates that the qth feature point to be used is in the image where it is located, that is, the cth feature point in the eth moment image in the current image and its previous N moments. i Image position information on images captured by an image acquisition device, i.e., visual measurement; Indicates the 3D position information corresponding to the qth feature point to be used. The value of q ranges from 1 to S and is an integer between them. S represents the total number of feature points to be used. Indicates the cth image in the eth moment image in the current image and its N previous moments. i The device pose information of each image acquisition device, i.e., the external parameter, Represents the pose information of the target object corresponding to the current image and the image at time e in the previous N images; Indicates the cth image in the eth moment image in the current image and its N previous moments. i The intrinsic parameter matrix of an image acquisition device; Represents the vertical axis coordinate value in the three-dimensional position information corresponding to the qth feature point to be used.
[0149] Then, based on the reprojection error equation, adjust value so that the above reprojection error equation holds true, and the current state information of the target object at the current moment is determined.
[0150] In another embodiment of the present invention, the step 024 may include the following steps 0241-0242:
[0151] 0241: Based on the reprojection error equation, construct the target measurement equation.
[0152] 0242: Use the target measurement equation and the filter update equation to determine the current state information of the target object at the current moment.
[0153] Referring to the error state Kalman filter theory, in order to update the current state information of the target object at the current moment, it is necessary to construct a measurement equation z that is the same as the processing method of the sensor data collected by other sensors. k =h(x k )+R k In this implementation, the target measurement equation is constructed based on the above reprojection error equation.
[0154] Among them, the left side of the equal sign in the above formula (8) That is, visual measurement means z k , the right side of the equal sign in the above formula (8) represents h(x k ). Accordingly, the visual measurement equation can be expressed as:
[0155]
[0156] Through the above formula, we can observe that the state quantity in the visual measurement equation contains The image position information of the feature point to be used, that is, the position of the feature point, is not used as the state quantity in the target measurement equation. According to the theory of multi-state constrained Kalman filter MSCKF, through the first-order approximation, multiplying both sides of the equation by Left null space elimination The residual effect of the final visual target measurement equation is obtained
[0157] Through the target measurement equation and the filter update equation, that is, the update equation corresponding to the filter update amount of the above formula (6), the pose information in the sliding window corresponding to the current image and the system state corresponding to the current moment are obtained. That is, the current state information of the target object at the current moment.
[0158] Corresponding to the above method embodiment, the embodiment of the present invention provides a state information estimation device, such as Figure 2 As shown, the device may include:
[0159] The first acquisition module 210 is configured to obtain a current image captured by a multi-image acquisition device set on the target object at a current moment and current sensor data captured by other sensors, wherein the other sensors include an IMU;
[0160] The first determining module 220 is configured to determine the initial state information of the target object at the current moment by using the IMU data corresponding to the previous moment and the previous state information of the target object at the previous moment;
[0161] The second determination module 230 is configured to use the feature points detected in each current image and the relative position relationship between the image acquisition devices corresponding to each current image to determine the matching point pairs between the current images as the first matching point pairs corresponding to the current images;
[0162] The third determining module 240 is configured to use the feature points detected in each current image and the feature points detected in the previous image to determine a matching point pair between each current image and its previous image as a second matching point pair corresponding to the current image;
[0163] The fourth determining module 250 is configured to determine the three-dimensional position information corresponding to each feature point to be used based on the first matching point pair and the second matching point pair corresponding to the current image and the first matching point pair and the second matching point pair corresponding to the images at N moments before the current moment;
[0164] The fifth determination module 260 is configured to determine the current state information of the target object at a current moment based on the three-dimensional position information, the image position information of each feature point to be utilized, the initial state information and the current sensor data.
[0165] By applying the embodiments of the present invention, the state can be expanded, that is, the tracking results of the feature points of each of the multiple image acquisition devices, that is, the second matching point pair corresponding to the current image and the images N moments before it, and the feature point association relationship between the multiple image acquisition devices, that is, the first matching point pair corresponding to the current image and the images N moments before it, are used to construct three-dimensional position information corresponding to each feature point to be used, so as to construct multi-state constraints, fuse the current sensor data of other sensor data, and determine the current state information of the target object at the current moment, so as to obtain a state estimation result with higher accuracy and higher robustness, so as to realize the estimation of the state information of an object in a multi-sensor system including any number of image acquisition devices.
[0166] In another embodiment of the present invention, the initial state information includes: initial velocity information and initial posture information, wherein the initial posture information includes: initial posture information and initial position information;
[0167] The first determining module 220 is specifically configured to determine the angular velocity information and acceleration information of the target object at the previous moment by using the IMU data corresponding to the previous moment;
[0168] Constructing a first state transfer equation using the angular velocity information at the previous moment and the previous posture information in the previous state information; determining the initial posture information of the target object at the current moment using the first state transfer equation;
[0169] constructing a second state transition equation using the previous posture information, the acceleration information at the previous moment, and the previous velocity information in the previous state information; and determining the initial velocity information of the target object at the current moment using the second state transition equation;
[0170] A third state transfer equation is constructed using the initial velocity information, the previous velocity information, and the previous position information in the previous state information; and the initial position information of the target object at the current moment is determined using the third state transfer equation.
[0171] In another embodiment of the present invention, the fourth determination module 250 is specifically configured to determine the three-dimensional position information corresponding to each feature point to be used based on the first matching point pair and the second matching point pair corresponding to the current image, the first matching point pair and the second matching point pair corresponding to the images at the previous N moments before the current moment, the device posture information of the image acquisition device corresponding to each current image, the device posture information of the image acquisition device corresponding to each image at the previous N moments, and the posture information of the target object corresponding to the current image and each image at the previous N moments according to the triangulation algorithm.
[0172] In another embodiment of the present invention, the fifth determining module 260 includes:
[0173] A first determining unit (not shown in the figure) is configured to determine the intermediate state information of the target object at a current moment based on the current sensor data and the initial state information;
[0174] A second determining unit (not shown in the figure) is configured to determine projection position information of a projection point of a spatial point corresponding to each feature point to be utilized in the image in which the spatial point is located, using the three-dimensional position information, the intermediate pose information in the intermediate state information, the object pose information in the state information of the target object corresponding to each image at each of the previous N moments, and the device pose information and intrinsic parameter matrix of each image acquisition device;
[0175] a construction unit (not shown in the figure), configured to construct a reprojection error equation based on the projection position information corresponding to each feature point to be utilized and the image position information of each feature point to be utilized in the current image;
[0176] A third determining unit (not shown in the figure) is configured to determine the current state information of the target object at the current moment based on the reprojection error equation.
[0177] In another embodiment of the present invention, the third determining unit is specifically configured to construct a target measurement equation based on the reprojection error equation;
[0178] The target measurement equation and the filter update equation are used to determine the current state information of the target object at the current moment.
[0179] The above-mentioned system and device embodiments correspond to the system embodiment and have the same technical effects as the method embodiment. For detailed descriptions, please refer to the method embodiment. The device embodiment is obtained based on the method embodiment. For detailed descriptions, please refer to the method embodiment section and will not be repeated here. It should be understood by those skilled in the art that the accompanying drawings are only schematic diagrams of one embodiment, and the modules or processes in the accompanying drawings are not necessarily required to implement the present invention.
[0180] Those skilled in the art will appreciate that the modules in the apparatuses of the embodiments may be distributed in the apparatuses of the embodiments as described in the embodiments, or may be located in one or more apparatuses different from the embodiments with corresponding changes. The modules in the above embodiments may be combined into one module or further divided into multiple sub-modules.
[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A state information estimation method, characterized in that: The method comprises: Obtaining a current image captured by a multi-image acquisition device and current sensor data captured by other sensors set on the target object at the current moment, wherein the other sensors include an IMU, and the multi-image acquisition device is a system of multiple image acquisition devices installed on the same rigid body in any number, position, and orientation; Determining initial state information of the target object at the current moment by using IMU data corresponding to a moment before the current moment and previous state information of the target object at the previous moment; Using the feature points detected in each current image and the relative position relationship between the image acquisition devices corresponding to each current image, a matching point pair between each current image is determined as a first matching point pair corresponding to the current image; Using the feature points detected in each current image and the feature points detected in the previous image, a matching point pair between each current image and the previous image is determined as a second matching point pair corresponding to the current image; Determine the three-dimensional position information corresponding to each feature point to be used based on the first matching point pair and the second matching point pair corresponding to the current image and the first matching point pair and the second matching point pair corresponding to the images at N moments before the current moment; Based on the three-dimensional position information, the image position information of each feature point to be used, the initial state information and the current sensor data, the current state information of the target object at the current moment is determined, including: based on the current sensor data and the initial state information, the intermediate state information of the target object at the current moment is determined; using the three-dimensional position information, the intermediate posture information in the intermediate state information, the object posture information in the state information of the target object corresponding to each image at the previous N moments, and the device posture information and internal parameter matrix of each image acquisition device, the projection position information of the projection point of the spatial point corresponding to each feature point to be used in the image in which it is located is determined; based on the projection position information corresponding to each feature point to be used and the image position information of each feature point to be used in the current image, a reprojection error equation is constructed; based on the reprojection error equation, the current state information of the target object at the current moment is determined.
2. The method according to claim 1, wherein The initial state information includes: initial velocity information and initial posture information, wherein the initial posture information includes: initial posture information and initial position information; The step of determining the initial state information of the target object at the current moment by using the IMU data corresponding to the moment before the current moment and the previous state information of the target object at the previous moment includes: Determine the angular velocity and acceleration information of the target object at the previous moment using IMU data corresponding to the previous moment; Constructing a first state transfer equation using the angular velocity information at the previous moment and the previous posture information in the previous state information; determining the initial posture information of the target object at the current moment using the first state transfer equation; constructing a second state transition equation using the previous posture information, the acceleration information at the previous moment, and the previous velocity information in the previous state information; and determining the initial velocity information of the target object at the current moment using the second state transition equation; A third state transfer equation is constructed using the initial velocity information, the previous velocity information, and the previous position information in the previous state information; and the initial position information of the target object at the current moment is determined using the third state transfer equation.
3. The method according to claim 1, wherein The step of determining the three-dimensional position information corresponding to each feature point to be used based on the first matching point pair and the second matching point pair corresponding to the current image and the first matching point pair and the second matching point pair corresponding to the images at N moments before the current moment includes: According to the triangulation algorithm, the three-dimensional position information corresponding to each feature point to be used is determined based on the first matching point pair and the second matching point pair corresponding to the current image, the first matching point pair and the second matching point pair corresponding to the images at the previous N moments before the current moment, the device posture information of the image acquisition device corresponding to each current image, the device posture information of the image acquisition device corresponding to each image at the previous N moments, and the posture information of the target object corresponding to the current image and each image at the previous N moments.
4. The method according to claim 1, wherein The step of determining the current state information of the target object at the current moment based on the reprojection error equation includes: Based on the reprojection error equation, construct a target measurement equation; The target measurement equation and the filter update equation are used to determine the current state information of the target object at the current moment.
5. A state information estimation device, characterized in that: The device comprises: A first acquisition module is configured to obtain a current image captured by a multi-image acquisition device set on the target object at a current moment and current sensor data captured by other sensors, wherein the other sensors include an IMU, and the multi-image acquisition device is a system of multiple image acquisition devices installed on the same rigid body in any number, position, and orientation; A first determining module is configured to determine initial state information of the target object at the current moment by using IMU data corresponding to a moment before the current moment and previous state information of the target object at the previous moment; A second determination module is configured to determine matching point pairs between the current images as first matching point pairs corresponding to the current images by using feature points detected in the current images and relative positional relationships between image acquisition devices corresponding to the current images; a third determining module configured to determine, by using the feature points detected in each current image and the feature points detected in the previous image, a pair of matching points between each current image and its previous image as a second matching point pair corresponding to the current image; A fourth determination module is configured to determine the three-dimensional position information corresponding to each feature point to be used based on the first matching point pair and the second matching point pair corresponding to the current image and the first matching point pair and the second matching point pair corresponding to the images at N moments before the current moment; a fifth determining module configured to determine current state information of the target object at a current moment based on the three-dimensional position information, the image position information of each feature point to be utilized, the initial state information, and the current sensor data; The fifth determining module includes: a first determining unit configured to determine intermediate state information of the target object at a current moment based on the current sensor data and the initial state information; The second determining unit is configured to determine projection position information of a projection point of a spatial point corresponding to each feature point to be utilized in the image in which the spatial point is located, using the three-dimensional position information, the intermediate pose information in the intermediate state information, the object pose information in the state information of the target object corresponding to each image at each of the previous N moments, and the device pose information and intrinsic parameter matrix of each image acquisition device; a construction unit configured to construct a reprojection error equation based on projection position information corresponding to each feature point to be utilized and image position information of each feature point to be utilized in the current image; The third determining unit is configured to determine the current state information of the target object at a current moment based on the reprojection error equation.
6. The device according to claim 5, characterized in that The initial state information includes: initial velocity information and initial posture information, wherein the initial posture information includes: initial posture information and initial position information; The first determining module is specifically configured to determine the angular velocity information and acceleration information of the target object at the previous moment by using the IMU data corresponding to the previous moment; Constructing a first state transfer equation using the angular velocity information at the previous moment and the previous posture information in the previous state information; determining the initial posture information of the target object at the current moment using the first state transfer equation; constructing a second state transition equation using the previous posture information, the acceleration information at the previous moment, and the previous velocity information in the previous state information; and determining the initial velocity information of the target object at the current moment using the second state transition equation; A third state transfer equation is constructed using the initial velocity information, the previous velocity information, and the previous position information in the previous state information; and the initial position information of the target object at the current moment is determined using the third state transfer equation.
7. The device according to claim 5, characterized in that The fourth determination module is specifically configured to determine the three-dimensional position information corresponding to each feature point to be used based on the first matching point pair and the second matching point pair corresponding to the current image, the first matching point pair and the second matching point pair corresponding to the images at the previous N moments before the current moment, the device posture information of the image acquisition device corresponding to each current image, the device posture information of the image acquisition device corresponding to each image at the previous N moments, and the posture information of the target object corresponding to the current image and each image at the previous N moments in accordance with the triangulation algorithm.
8. The device according to claim 5, wherein The third determining unit is specifically configured to construct a target measurement equation based on the reprojection error equation; The target measurement equation and the filter update equation are used to determine the current state information of the target object at the current moment.