Positioning method, device, electronic device and computer-readable storage medium
By acquiring and fusing the data of lidar and depth sensors, the problem of data in the prior art cannot be synchronously fusion, and the calculation accuracy of the position relationship between humans and robots is improved.
Patent Information
- Application Number
- CN202010651193.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-08
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-07-08
AI Technical Summary
In the prior art, the measurement time of the camera and the depth sensor cannot be synchronized, resulting in the inability to fuse the camera data and the depth sensor data, thereby reducing the calculation accuracy of the relative position relationship between the human body and the robot.
A positioning method is proposed to obtain a point cloud map measured by lidar and depth map measured by depth sensors, extract the target area from it, and fuse multiple time-outsync sensor measurement data based on the object's historical motion state to improve the accuracy of the target position.
Through data fusion, the accuracy of target position determination is improved and the calculation accuracy of position relationship between human body and robot is enhanced.
Smart Images

Figure CN113917475B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular, to a positioning method, device, electronic device, and computer-readable storage medium. Background Art
[0002] In the process of interaction between a robot and a human, an important application is to track the human to move to a specified position following the human. Therefore, it is necessary to detect the human body in real time and estimate the relative position relationship between the human body and the robot.
[0003] In the prior art, the distance between the human body and the robot is usually calculated based on the data of a camera and a depth sensor. However, the measurement times of the camera and the depth sensor are usually not synchronized, so that the camera data and the depth sensor data cannot be fused, and the accuracy of the calculated relative position relationship between the human body and the robot is poor. Summary of the Invention
[0004] This application aims to solve at least one of the technical problems in the related art to some extent.
[0005] To this end, the first object of this application is to propose a positioning method, which extracts a first target area from a point cloud map and a second target area from a depth map according to an object included in a region of interest in a visual image, and fuses multiple sensor measurement data with different time synchronizations according to the historical motion state of the object, thereby improving the accuracy of target position determination after fusion.
[0006] The second object of this application is to propose a positioning device.
[0007] The third object of this application is to propose an electronic device.
[0008] The fourth object of this application is to propose a non-transitory computer-readable storage medium.
[0009] To achieve the above object, an embodiment of the first aspect of this application proposes a positioning method, including:
[0010] Obtain a point cloud map measured by a lidar; obtain a depth map measured by a depth sensor; extract a first target area from the point cloud map and a second target area from the depth map; wherein, the first target area and the second target area detect the same object as the region of interest in the synchronously acquired visual image;
[0011] Fuse the first positioning information carried by the first target area and the second positioning information carried by the second target area according to the historical motion state of the object to obtain the target position of the object.
[0012] Optionally, as a first possible implementation manner of the first aspect, fusing the first positioning information carried by the first target area and the second positioning information carried by the second target area according to the historical motion state of the object to obtain the target position of the object includes:
[0013] Determining a first observation position where the object is located according to the first positioning information carried by the first target area;
[0014] Determining a second observation position where the object is located according to the second positioning information carried by the second target area;
[0015] Performing an iterative correction process on one of the first observation position and the second observation position according to the historical motion state of the object;
[0016] Updating the historical motion state according to the position obtained by the iterative correction process;
[0017] Performing an iterative correction process on the other of the first observation position and the second observation position according to the updated historical motion state to obtain the target position.
[0018] Optionally, as a second possible implementation manner of the first aspect, performing the iterative correction process includes:
[0019] Obtaining a predicted motion state according to the historical motion state adopted in the current iterative correction process; the historical motion state adopted in the current iterative correction process is generated according to the position obtained by the previous iterative correction process and the historical motion state adopted in the previous iterative correction process;
[0020] Obtaining a predicted observation position according to the predicted motion state;
[0021] Correcting the first observation position or the second observation position for the current iterative correction process according to the predicted observation position.
[0022] Optionally, as a third possible implementation manner of the first aspect, before performing the iterative correction process on one of the first observation position and the second observation position according to the historical motion state of the object, it further includes:
[0023] Comparing the first observation moment when the point cloud map is obtained and the second observation moment when the depth map is obtained;
[0024] If the first observation moment is earlier than the second observation moment, determining that the first observation position is iteratively corrected earlier than the second observation position;
[0025] If the first observation moment is later than the second observation moment, it is determined that the second observation position is iteratively corrected prior to the first observation position;
[0026] If the first observation moment is equal to the second observation moment, the order of iterative correction for the first observation position and the second observation position is randomly determined.
[0027] Optionally, as a fourth possible implementation manner of the first aspect, the historical motion state includes the historical position and historical velocity of the object;
[0028] Correspondingly, for each iterative correction process, the historical position in the historical motion state used in the current iterative correction process is generated based on the position obtained from the previous iterative correction process and the historical position used in the previous iterative correction process;
[0029] For each iterative correction process, the historical velocity in the historical motion state used in the current iterative correction process is determined based on the historical position used in the current iterative correction process and the historical position used in the previous iterative correction process.
[0030] Optionally, as a fifth possible implementation manner of the first aspect, the obtaining of the predicted observation position according to the predicted motion state includes:
[0031] Substitute the predicted motion state into the observation equation to obtain the predicted observation position;
[0032] Wherein, the observation equation is the product of the predicted motion state and the transformation matrix superimposed with the measurement noise term;
[0033] The transformation matrix is used to indicate the transformation relationship between the predicted motion state and the predicted observation position;
[0034] The measurement noise term conforms to a Gaussian white noise distribution with a set covariance; the set covariance is determined according to the device accuracy and measurement confidence.
[0035] Optionally, as a sixth possible implementation manner of the first aspect, the correcting of the first observation position or the second observation position for the current iterative correction process according to the predicted observation position includes:
[0036] Determine the measurement residual for the first observation position or the second observation position for the current iterative correction process;
[0037] If the measurement residual is less than the difference threshold, correct the first observation position or the second observation position for the current iterative correction process according to the predicted observation position.
[0038] Optionally, as the seventh possible implementation manner of the first aspect, the measurement residual is the difference between the predicted observation position and the first observation position or the second observation position in the current iterative correction process.
[0039] Optionally, as the eighth possible implementation manner of the first aspect, the extracting the first target area from the point cloud map includes:
[0040] Determine the rectangular coordinate position in the image coordinate system for the region of interest;
[0041] Map the rectangular coordinate position in the image coordinate system to the polar coordinate system of the lidar to obtain the polar coordinate position;
[0042] Extract the first target area from the point cloud map according to the polar coordinate position.
[0043] Optionally, as the ninth possible implementation manner of the first aspect, the determining the rectangular coordinate position in the image coordinate system for the region of interest includes:
[0044] Determine the rectangular coordinate position for the left and right boundaries of the region of interest.
[0045] Optionally, as the tenth possible implementation manner of the first aspect, the mapping the rectangular coordinate position in the image coordinate system to the polar coordinate system of the lidar to obtain the polar coordinate position includes:
[0046] Map the rectangular coordinate position in the image coordinate system to the camera coordinate system through the internal parameter matrix of the camera to obtain the rectangular coordinate position of the camera coordinate system; wherein, the camera is used to collect the visual image;
[0047] Map the rectangular coordinate position of the camera coordinate system to the polar coordinate system through the external parameter matrix between the camera and the lidar to obtain the polar coordinate position.
[0048] Optionally, as the eleventh possible implementation manner of the first aspect, the determining the first observation position where the object is located according to the first positioning information carried by the first target area includes:
[0049] Determine the depth of each pixel point in the first target area according to the first positioning information carried by the first target area;
[0050] According to the depth of each pixel point, count the pixel point number indication value corresponding to each set depth;
[0051] Determine the target depth from each set depth according to the peak value of the pixel point number indication value;
[0052] Locate the first observation position where the object is located according to the target depth.
[0053] Optionally, as the twelfth possible implementation of the first aspect, determining the target depth from the set depths according to the peak value of the pixel point number indication value includes:
[0054] Determine the foreground depth and the background depth from the set depths; wherein, the background depth has the maximum peak value of the pixel point number indication value; the foreground depth has the first peak value of the pixel point number indication value in ascending order of depth;
[0055] Select the target depth from the foreground depth and the background depth according to the pixel point number indication values corresponding to the foreground depth and the background depth.
[0056] Optionally, as the thirteenth possible implementation of the first aspect, selecting the target depth from the foreground depth and the background depth according to the pixel point number indication values corresponding to the foreground depth and the background depth includes:
[0057] If the ratio of the pixel point number indication values of the foreground depth and the background depth is greater than the ratio threshold, use the foreground depth as the target depth;
[0058] If the ratio of the pixel point number indication values of the foreground depth and the background depth is not greater than the ratio threshold, use the background depth as the target depth.
[0059] Optionally, as the fourteenth possible implementation of the first aspect, after locating the first observation position where the object is located according to the target depth, further includes:
[0060] Determine the measurement confidence of the first observation position according to the pixel point number indication value corresponding to the target depth;
[0061] If the foreground depth and the background depth are the same, increase the measurement confidence of the first observation position, wherein the measurement confidence is used to generate a measurement noise term for the observation equation adopted in the iterative correction process, and the observation equation is used to substitute the predicted motion state into the observation equation after obtaining the predicted motion state according to the historical motion state adopted in the iterative correction process to obtain the predicted observation position.
[0062] Optionally, as the fifteenth possible implementation of the first aspect, after statistically calculating the pixel point number indication values corresponding to the set depths according to the depths of the pixels, further includes:
[0063] Filter out the set depth with the number indication value of the pixel points less than the number threshold.
[0064] Optionally, as the sixteenth possible implementation manner of the first aspect, the counting the number indication value corresponding to each set depth according to the depths of the pixel points includes:
[0065] For each set depth, determine the depth statistical range;
[0066] According to the depths of the pixel points, count the number of pixel points whose depths match the corresponding depth statistical range to obtain the number indication value corresponding to the corresponding set depth.
[0067] Optionally, as the seventeenth possible implementation manner of the first aspect, the first observation position includes an observation distance and an observation angle; the first positioning information includes a depth and an angle;
[0068] The positioning the first observation position where the object is located according to the target depth includes:
[0069] According to the angles carried by the pixel points corresponding to the target depth, position the observation angle of the object;
[0070] According to the depths carried by the pixel points corresponding to the target depth, position the observation distance of the object.
[0071] To achieve the above object, an embodiment of the second aspect of the present application proposes a positioning device, including:
[0072] An acquisition module, configured to acquire a point cloud map measured by a lidar;
[0073] The acquisition module is further configured to acquire a depth map measured by a depth sensor;
[0074] An extraction module, configured to extract a first target area from the point cloud map and extract a second target area from the depth map; wherein, the first target area and the second target area detect the same object as the region of interest in the synchronously acquired visual image;
[0075] A fusion module, configured to fuse the first positioning information carried by the first target area and the second positioning information carried by the second target area according to the historical motion state of the object to obtain the target position of the object.
[0076] To achieve the above object, an embodiment of the third aspect of the present application proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the positioning method described in the first aspect is implemented.
[0077] To achieve the above object, an embodiment of the fourth aspect of the present application provides a computer program stored thereon, characterized in that when the program is executed by a processor, it implements the positioning method described in the first aspect.
[0078] The technical solution provided by the embodiment of the present application may include the following beneficial effects:
[0079] According to the object included in the region of interest in the visual image, the first target region is extracted from the point cloud map, and the second target region is extracted from the depth map. According to the historical motion state of the object, the sensor measurement data that are out of sync at multiple times are fused, and the accuracy of determining the target position is improved after the fusion.
[0080] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0082] Figure 1 is a schematic flowchart of a positioning method provided by an embodiment of the present application;
[0083] Figure 2 is a schematic flowchart of another positioning method provided by an embodiment of the present application;
[0084] Figure 3 is a schematic diagram of coordinate system conversion;
[0085] Figure 4 is a schematic flowchart of yet another positioning method provided by an embodiment of the present application;
[0086] Figure 5 is a schematic flowchart of an iterative correction process provided by an embodiment of the present application;
[0087] Figure 6 is a schematic diagram of the definitions of each state variable in the polar coordinate system;
[0088] Figure 7 is a schematic flowchart of another iterative correction process provided by an embodiment of the present application;
[0089] Figure 8 is a schematic flowchart of yet another positioning method provided by an embodiment of the present application;
[0090] Figure 9 is a schematic flowchart of yet another positioning method provided by an embodiment of the present application;
[0091] Figure 10A schematic diagram of a histogram provided for this application;
[0092] Figure 11 Another schematic diagram of a histogram provided for this application; and
[0093] Figure 12 A schematic structural diagram of a positioning device provided for an embodiment of this application. Detailed implementation manners
[0094] The embodiments of this application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain this application, and should not be construed as a limitation to this application.
[0095] The positioning method, device, electronic device, and computer-readable storage medium of the embodiments of this application will be described below with reference to the accompanying drawings.
[0096] Figure 1 A schematic flowchart of a positioning method provided for an embodiment of this application.
[0097] As Figure 1 shown, the method includes the following steps:
[0098] Step 101, obtain a point cloud map measured by a lidar.
[0099] Step 102, obtain a depth map measured by a depth sensor.
[0100] In this embodiment, both the point cloud map measured by the lidar and the depth map measured by the depth sensor carry corresponding positioning information.
[0101] Among them, the positioning information carried by the point cloud map includes the angular information and depth information of each object measured by the lidar, and the depth information indicates the relative distance information between the object and the lidar.
[0102] Among them, the depth sensor, for example, is a (Depth Red Green Blue, RGB-D) sensor. The positioning information carried by the depth map includes the depth information of each object measured by the depth sensor, and the depth information indicates the relative distance information between the object and the depth sensor.
[0103] Step 103, extract a first target area from the point cloud map and extract a second target area from the depth map, where the first target area and the second target area detect the same object as the region of interest in the synchronously acquired visual image.
[0104] Among them, the region of interest refers to the region in the visual image that contains an object. For example, if the object is a human body, the region of interest is the region that contains the human body.
[0105] In this embodiment, when a point cloud map is measured by a lidar and a depth map is measured by a depth sensor, a visual image is also synchronously collected by a camera, and the region of interest is determined by detecting the visual image. For example, the region of interest is determined from the image by a target detection algorithm (Single Shot MultiBox Detector, SSD) or a target detection algorithm (You Only Look Once, YOLO).
[0106] According to the determined region of interest, a first target region is extracted from the point cloud map. When extracting the first target region from the point cloud map, since the reference coordinates corresponding to the region of interest are in the image coordinate system, and the point cloud map measured by the lidar is in the polar coordinate system of the lidar, it is necessary to convert to the same coordinate system. To reduce the computational amount, the region of interest is converted from the image coordinate system to the polar coordinate system of the lidar, achieving the unification of the coordinate system, and thus realizing the extraction of the first target region from the point cloud map. Among them, the coordinate system conversion method will be specifically introduced in the next embodiment.
[0107] According to the determined region of interest, a second target region is also extracted from the depth map. That is to say, the first target region and the second target region detect the same object as the region of interest in the synchronously collected visual image, providing a basis for data fusion of different sensors to determine the position of the object.
[0108] Step 104, according to the historical motion state of the object, fuse the first positioning information carried by the first target region and the second positioning information carried by the second target region to obtain the target position of the object.
[0109] Specifically, according to the first positioning information carried by the first target region, determine the first observation position where the object is located. According to the second positioning information carried by the second target region, determine the second observation position where the object is located. According to the historical motion state of the object, perform an iterative correction process on one of the first observation position and the second observation position. Update the historical motion state according to the position obtained from the iterative correction process. Furthermore, according to the updated historical motion state, perform an iterative correction process on the other of the first observation position and the second observation position to obtain the target position. In the iterative correction process, the first observation position and the second observation position are fused to obtain the target position. That is to say, in the iterative correction process, the measurement data of the depth sensor and the measurement data of the lidar sensor are fused. After fusion, the amount of information for determining the target position is increased, and the accuracy of determining the target position is improved.
[0110] In the positioning method according to the embodiments of the present application, a point cloud map obtained by lidar measurement is acquired, a depth map obtained by depth sensor measurement is acquired, a first target area is extracted from the point cloud map, and a second target area is extracted from the depth map. Among them, the first target area and the second target area detect the same object as the region of interest in the synchronously acquired visual image. According to the historical motion state of the object, the first positioning information carried by the first target area and the second positioning information carried by the second target area are fused to obtain the target position of the object. According to the object included in the region of interest in the visual image, the first target area is extracted from the point cloud map, and the second target area is extracted from the depth map. According to the historical motion state of the object, the measurement data of multiple sensors with different time synchronization are fused, and the accuracy of target position determination is improved after fusion.
[0111] Based on the previous embodiment, this embodiment provides an implementation manner, which specifically illustrates how to extract the first target area from the point cloud image.
[0112] As Figure 2 shown, the steps of extracting the first target area from the point cloud map in step 103 may include the following steps:
[0113] Step 201, for the region of interest, determine the rectangular coordinate position in the image coordinate system.
[0114] Specifically, the image coordinate system is a plane coordinate system. Therefore, the left and right boundaries corresponding to the region of interest can be retained, and for the left and right boundaries of the region of interest, the rectangular coordinate positions are determined. As Figure 3 shown, the left and right boundaries corresponding to the region of interest correspond to four vertices. According to the four vertices corresponding to the left and right boundaries, the rectangular coordinate position in the image coordinate system is determined.
[0115] Step 202, map the rectangular coordinate position in the image coordinate system to the polar coordinate system of the lidar to obtain the polar coordinate position.
[0116] As Figure 3 shown, the rectangular coordinate position of the region of interest in the image coordinate system is mapped to the camera coordinate system through the internal parameter matrix of the camera to obtain the rectangular coordinate position in the camera coordinate system. Among them, the camera is used to collect the visual image. Furthermore, the rectangular coordinate position in the camera coordinate system is mapped to the polar coordinate system of the lidar through the external parameter matrix between the camera and the lidar to obtain the polar coordinate position corresponding to the region of interest in the polar coordinate system of the lidar, realizing the conversion of the same coordinate system. And converting the region of interest from the image coordinate system to the polar coordinate system of the lidar reduces the amount of calculation.
[0117] Step 203, according to the polar coordinate position, extract the first target area from the point cloud map.
[0118] Specifically, according to the polar coordinate position corresponding to the region of interest in the polar coordinate system of the lidar, the region corresponding to the polar coordinate position is determined from the point cloud map, and the region corresponding to the polar coordinate position is used as the first target region, where the first target region and the region of interest in the synchronously acquired visual image detect the same object.
[0119] In the positioning method of this embodiment, the coordinate system of the region of interest in the visual image collected by the camera is converted to obtain the corresponding polar coordinate position in the polar coordinate system of the lidar, realizing the conversion of the same coordinate system. And converting the region of interest from the image coordinate system to the polar coordinate system of the lidar reduces the computational amount. The first target region is extracted from the point cloud map according to the polar coordinate position, so that the first target region and the region of interest in the synchronously acquired visual image detect the same object, facilitating subsequent fusion with the data collected by the depth sensor.
[0120] In one embodiment of the present application, for extracting the second target region from the depth map in step 103, as a possible implementation, the obtained depth map is registered with the visual image collected by the camera. Furthermore, the depth map is cropped according to the region of interest in the visual image to obtain the second target region that detects the same object as the region of interest.
[0121] Based on the above embodiments, this embodiment provides an implementation manner, which details how to fuse the measurement data of multiple sensors with different time synchronizations to improve the accuracy of target position determination. Among them, the fusion algorithm can be implemented based on the Unscented Kalman Filter (UKF), or based on the classical Kalman filter, or the Extended Kalman Filter (EKF). In this embodiment, the implementation based on the Unscented Kalman Filter is taken as an example for illustration, but it does not limit this embodiment.
[0122] As Figure 4 shown, the above step 104 may include the following steps:
[0123] Step 401, determine the first observation position where the object is located according to the first positioning information carried by the first target region.
[0124] Among them, the first positioning information includes depth and angle, and the first observation position includes the observation angle and the observation distance.
[0125] Specifically, according to the first positioning information carried by each pixel point in the first target area, the pixel points whose first positioning information matches each set depth are counted to obtain the pixel point number indication value corresponding to each set depth. According to the peak value of the pixel point number indication value, the target depth is determined from each set depth, and based on the target depth, the first observation position where the object is located is determined. Among them, the method for determining the first observation position will be described in detail in the subsequent embodiments.
[0126] Step 402: Determine the second observation position where the object is located according to the second positioning information carried by the second target area.
[0127] Among them, the second positioning information includes depth, and the second observation position includes the observation distance.
[0128] Specifically, according to the second positioning information carried by each pixel point in the second target area, the pixel points whose second positioning information matches each set depth are counted to obtain the pixel point number indication value corresponding to each set depth. According to the peak value of the pixel point number indication value, the target depth is determined from each set depth, and based on the target depth, the second observation position where the object is located is determined. Among them, the method for determining the second observation position will be described in detail in the subsequent embodiments.
[0129] Step 403: Perform an iterative correction process on one of the first observation position and the second observation position according to the historical motion state of the object.
[0130] Step 404: Update the historical motion state according to the position obtained from the iterative correction process.
[0131] Step 405: Perform an iterative correction process on the other one of the first observation position and the second observation position according to the updated historical motion state to obtain the target position.
[0132] In this embodiment, the iterative correction process can be performed on the first observation position first, or on the second observation position first. Among them, how to determine the order of the iterative correction processes for the first observation position and the second observation position will be described in detail in the next embodiment.
[0133] In one embodiment, an iterative correction process is performed on the first observation position according to the historical motion state of the object, and the historical motion state is updated according to the position obtained from the iterative correction process to obtain an updated historical motion state. Furthermore, an iterative correction process is performed on the second observation position according to the updated historical motion state to obtain a target position. That is to say, the target position is obtained by performing iterative correction on the second observation position on the basis of the iterative correction of the first observation position, that is, obtained by fusing the first positioning information and the second positioning information. That is to say, the measurement data of the depth sensor and the measurement data of the lidar sensor are fused, and the amount of information used to determine the target position is increased through fusion, thereby improving the accuracy of target position determination.
[0134] Similarly, an iterative correction process is first performed on the second observation position, and then an iterative correction process is performed on the first observation position. The principle of fusing to determine the target position is the same as above and will not be elaborated here.
[0135] It should be noted that in this embodiment, the fusion is performed in the time sequence of data collection by two sensors, namely the lidar sensor and the depth sensor. In practical applications, three or more sensors can also be fused according to the time sequence of data collection, which is not limited in this embodiment.
[0136] In the positioning method of the embodiment of the present application, an iterative correction process is performed on one of the first observation position and the second observation position according to the historical motion state of the object. The historical motion state is updated according to the position obtained from the iterative correction process. According to the updated historical motion state, an iterative correction process is performed on the other of the first observation position and the second observation position to obtain a target position. The target position is obtained by fusing the first positioning information and the second positioning information. That is to say, the measurement data of the depth sensor and the measurement data of the lidar sensor are fused, and the amount of information used to determine the target position is increased through fusion, thereby improving the accuracy of target position determination.
[0137] Based on the above embodiment, Figure 5 It is a schematic flowchart of an iterative correction process provided by an embodiment of the present application.
[0138] As Figure 5 shown, in the above steps 403 and 405, when performing the iterative correction process, the following steps may be included:
[0139] Step 501, obtain a predicted motion state according to the historical motion state used in the current iterative correction process, where the historical motion state used in the current iterative correction process is generated according to the position obtained from the previous iterative correction process and the historical motion state used in the previous iterative correction process.
[0140] In one embodiment of the present application, the historical motion state used in the current iteration correction process is brought into the system state equation to obtain the predicted motion state.
[0141] In this embodiment, the system state equation is the product of the historical motion state and the state transition matrix superimposed with the process noise term.
[0142] For example, the system state equation is expressed as: X(t + 1) = g(t)X(t) + w(t), where X(t) represents the historical motion state, w(t) is the process noise term, which conforms to Gaussian white noise with an expectation of 0 and a covariance of Q, and its covariance Q can be initialized to a constant value. g(t) is the state transition matrix. For example, where X(t) is the historical motion state. The dist(t) included in the historical motion state indicates the relative distance of the object in the historical state; vel(t) indicates the relative linear velocity of the object; sita(t) indicates the relative angle of the object; omega(t) indicates the angular velocity of the object. X(t + 1) is the predicted motion state predicted based on the historical motion state X(t).
[0143] Taking the robot scenario as an example, the lidar sensor and depth sensor set in the robot collect the image information of people in real time. Figure 6 It is a schematic diagram of the definition of each state variable in the polar coordinate system, as Figure 6 shown. Then the definitions of the variables of the historical motion state in the polar coordinate system of the lidar are as follows:
[0144] Figure 6 In it, the triangle indicates the position of the robot, the pentagram indicates the position of the object, that is, the position of the person. The number 1 indicates dist(t), that is, the relative distance between the person and the robot, which is always positive.
[0145] The number 2 indicates sita(t): the relative angle of the person in the robot coordinate system, with the front being 0 degrees and the clockwise direction being positive.
[0146] The number 3 indicates vel(t), that is, the relative linear velocity between the person and the robot, with moving away being positive, that is, the direction indicated by the arrow is positive.
[0147] The number 4 indicates omega(t), that is, the angular velocity of the person moving tangentially in the robot coordinate system, with the clockwise direction being positive, that is, the direction indicated by the arrow is positive.
[0148] where X(t + 1) is the predicted motion state predicted based on the historical motion state X(t).
[0149] Step 502: Obtain the predicted observation position according to the predicted motion state.
[0150] In one embodiment of the present application, the predicted motion state is brought into the observation equation to obtain the predicted observation position. Among them, the observation equation is the product of the predicted motion state and the transformation matrix superimposed with the measurement noise term. Among them, the transformation matrix is used to indicate the transformation relationship between the predicted motion state and the predicted observation position; the measurement noise term conforms to the Gaussian white noise distribution with a set covariance, and the set covariance is determined according to the device accuracy and the measurement confidence level. Among them, the measurement confidence level includes the measurement confidence level of the first observation position or the measurement confidence level of the second observation position, and the method for determining the measurement confidence level will be described in subsequent embodiments..
[0151] In this embodiment, after substituting the predicted motion state into the observation equation, the predicted observation position can be obtained, and the predicted observation position includes the predicted observation angle and the predicted observation distance.
[0152] For example, the observation equation for the predicted observation distance is: dist(t + 1) = h1 * X(t + 1) + v1(t + 1);
[0153] The observation equation for the predicted observation angle is: sita(t + 1) = h2 * X(t + 1) + v2(t + 1).
[0154] Among them, dist(t + 1) represents the observation distance, sita(t + 1) represents the observation angle, h1 and h2 represent the transformation matrix, and v1(t + 1) and v2(t + 1) represent the measurement noise terms. For example, h1 = [1 0 0 0], h2 = [0 0 1 0]. Step 503, according to the predicted observation position, correct the first observation position or the second observation position in the current iteration correction process.
[0155] In one embodiment of the present application, if the current iteration correction process is to correct the first observation position, then according to the predicted observation position, correct the first observation position in the current iteration correction process to obtain the corrected position.
[0156] Among them, the predicted observation position includes the predicted angle and the predicted distance, and the first observation position includes the first observation angle and the second observation distance of the object. Therefore, the corrected position also includes the observation distance and the observation angle.
[0157] As a possible implementation, determine the weight of the predicted observation position and the weight of the first observation position, and perform weighted calculation according to the weight of the predicted observation position and the weight of the first observation position to obtain the position obtained in the current iteration correction process. Among them, the corrected position also includes the observation distance and the observation angle.
[0158] In another embodiment of the present application, if the current iterative correction process is to correct the second observation position, then the second observation position for the current iterative correction process is corrected according to the predicted observation position to obtain the corrected position.
[0159] Among them, the predicted observation position includes a predicted angle and a predicted distance, and the first observation position includes the first observation angle and the second observation distance of the object. Thus, the position obtained after correction also includes the observation distance and the observation angle.
[0160] As a possible implementation, the weight of the predicted observation position and the weight of the second observation position are determined, and weighted calculation is performed according to the weight of the predicted observation position and the weight of the second observation position to obtain the position obtained in the current iterative correction process. That is to say, the second observation position is corrected by the predicted observation position.
[0161] It should be noted that the historical motion state adopted in the current iterative correction process is generated according to the position obtained in the previous iterative correction process and the historical motion state adopted in the previous iterative correction process. For example, in this embodiment, the iterative correction process in round A1 corrects the first observation position or the second observation position through the predicted observation position to obtain the corrected distance S. The historical motion state adopted in the next iterative correction process in round A2 is to replace the distance in the historical motion state adopted in round A1 with the distance S corrected in round A1 to obtain the historical motion state that needs to be adopted in the next iterative correction process in round A2, realizing the iterative correction of the historical motion state and improving the accuracy of each iterative correction.
[0162] In a possible implementation of the embodiment of the present application, the historical motion state includes the historical position and historical speed of the object. Among them, the historical position includes the relative distance and relative angle of the object, that is to say, the distance is determined by the relative distance and relative angle between the object and the collected machine. The historical speed includes the relative linear speed and angular speed of the object.
[0163] Correspondingly, for each iterative correction process, the historical position in the historical motion state adopted in the current iterative correction process is generated according to the position obtained in the previous iterative correction process and the historical position adopted in the previous iterative correction process. And for each iterative correction process, the historical speed in the historical motion state adopted in the current iterative correction process is determined according to the historical position adopted in the current iterative correction process and the historical position adopted in the previous iterative correction process. By correcting the historical position and historical speed in the historical motion state in each iterative correction process, it is realized that each iterative correction process adopts the corrected historical motion state, improving the accuracy of the position obtained in each iterative correction process.
[0164] It should be noted that the speed of the object to be corrected in each iteration process can be used to predict the behavior of the object. For example, it can be determined whether the object is moving away from or approaching the acquisition machine, etc., so as to realize the tracking of the object.
[0165] In the positioning method of this embodiment, through iterative correction, the first observation position or the second observation position in the current iterative correction process is corrected according to the predicted observation position to obtain the corrected position. After the corrections for both the first observation position and the second observation position are completed, the corrected target position is obtained, so that the determined target position integrates the information of the first observation position and the second observation position, improving the accuracy of target position determination.
[0166] In practical applications, when the difference between the predicted observation position and the first observation position or the second observation position in the current iterative correction process is large, it indicates that the credibility of the predicted observation position is low. If the predicted observation position is used to correct the obtained first observation position or second observation position, the error of the corrected position will be larger. Therefore, in order to improve the reliability of iterative correction, in this embodiment, according to the measurement residual, the size of the difference between the predicted observation position and the first observation position or the second observation position in the current iterative correction process is judged to identify whether it is necessary to correct the first observation position or the second observation position in this iteration process. Thus, as Figure 7 shown, the above step 503 may include the following steps:
[0167] Step 701: Determine the corresponding measurement residual for the first observation position or the second observation position in the current iterative correction process.
[0168] Step 702: If the measurement residual is less than the difference threshold, correct the first observation position or the second observation position in the current iterative correction process according to the predicted observation position.
[0169] Among them, the measurement residual is the difference between the predicted observation position and the first observation position or the second observation position in the current iterative correction process.
[0170] In this embodiment, since the acquisition times of the detection data by the lidar sensor or the depth sensor are not synchronized, therefore, each time iterative correction is performed, the first observation position or the second observation position may be used for correction. Thus, the corresponding measurement residual is determined between the obtained predicted observation position and the first observation position or the second observation position.
[0171] Specifically, in a scenario where the first observation position performs the current iterative correction process, the measurement residual is determined based on the first observation position and the predicted observation position, and it is judged whether the measurement residual is less than the difference threshold. If the measurement residual is less than the difference threshold, the obtained predicted observation position has a high credibility and can be used to subsequently correct the first observation position to determine the corrected position; if the measurement residual is greater than the difference threshold, the obtained predicted observation position has a low credibility, that is, it is regarded as noise data and is not used to correct the first observation position to improve the accuracy of the corrected position.
[0172] In another scenario, where the second observation position performs the current iterative correction process, the measurement residual is determined based on the second observation position and the predicted observation position, and it is judged whether the measurement residual is less than the difference threshold. If the measurement residual is less than the difference threshold, the obtained predicted observation position has a high credibility and can be used to subsequently correct the second observation position to determine the corrected position; if the measurement residual is greater than the difference threshold, the obtained predicted observation position has a low credibility, that is, it is regarded as noise data and is not used to correct the second observation position to improve the accuracy of the corrected position.
[0173] In the positioning method of this embodiment, for the first observation position or the second observation position that performs the current iterative correction process, the corresponding measurement residual is determined. When the measurement residual is less than the difference threshold, the first observation position or the second observation position that performs the current iterative correction process is corrected according to the predicted observation position. Since the measurement residual indicates the credibility level of the obtained predicted observation position, the reliability and accuracy of correcting the first observation position or the second observation position in each iterative correction process are improved.
[0174] In the above embodiment, it is described that an iterative correction process is performed on one of the first observation position and the second observation position according to the historical motion state of the object. In this embodiment, a method for determining the order of performing the iterative correction on the first observation position and the second observation position is provided. As Figure 8 shown, before the above step 403, the following steps may be included:
[0175] Step 801, compare the first observation moment when the point cloud map is obtained and the second observation moment when the depth map is obtained.
[0176] In this embodiment, when the lidar measures the point cloud map and the depth sensor measures the depth map, the acquisition times are not synchronized. For the convenience of distinction, the moment when the point cloud map is obtained is called the first observation moment, and the moment when the depth map is obtained is called the second observation moment. That is to say, the first observation moment when the point cloud map is obtained is different from the second observation moment when the depth map is obtained. Therefore, it is necessary to compare the first observation moment and the second observation moment to determine the sequence of the first observation moment and the second observation moment.
[0177] Step 802, if the first observation moment is prior to the second observation moment, determine that the first observation position is prior to the second observation position for iterative correction.
[0178] Step 803, if the first observation moment is later than the second observation moment, determine that the second observation position is prior to the first observation position for iterative correction.
[0179] Step 804, if the first observation moment is equal to the second observation moment, randomly determine the order of iterative correction for the first observation position and the second observation position.
[0180] In the positioning method of this embodiment, if it is determined by comparison that the first observation moment is prior to the second observation moment, it is determined that the first observation position is prior to the second observation position for iterative correction. That is to say, iterative correction is first performed according to the first observation position at the first observation moment, and then iterative correction is performed according to the second observation position at the second observation moment. If it is determined by comparison that the first observation moment is later than the second observation moment, it is determined that the second observation position is prior to the first observation position for iterative correction. If the first observation moment is equal to the second observation moment, the order of iterative correction for the first observation position and the second observation position is randomly determined. That is to say, one of the first observation position and the second observation position is randomly determined to be iteratively corrected first, and then the other of the first observation position and the second observation position is iteratively corrected, realizing iterative correction of the first observation position and the second observation position alternately, thereby realizing effective fusion of the lidar measurement data and the depth sensor measurement data and improving the accuracy.
[0181] Based on the above embodiment, this embodiment provides an implementation manner, illustrating how to determine the first observation position of the object according to the first target area determined from the point cloud map measured by the lidar. As Figure 9 shown, the above step 401 may include the following steps:
[0182] Step 901, determine the depth of each pixel point in the first target area according to the first positioning information carried by each pixel point in the first target area.
[0183] Step 902, according to the depth of each pixel point, count the pixel point number indication values corresponding to each set depth.
[0184] In one embodiment of the present application, the pixels whose depths match each set depth are counted to obtain the pixel number indication value corresponding to each set depth. Specifically, for each pre-determined set depth, a depth statistical range is determined, and the number of pixels whose depths match the corresponding depth statistical range is counted to obtain the pixel number indication value corresponding to the corresponding set depth. By expanding the range of each set depth to obtain the depth statistical range and counting the number of pixels within the corresponding depth statistical range as the number of pixels corresponding to the corresponding set depth, the number of pixels corresponding to the corresponding set depth is increased, taking into account both accuracy and precision.
[0185] In this embodiment, a histogram can be used to display the number of pixels corresponding to each set depth. Figure 10 FIG. is a schematic diagram of a histogram provided for this embodiment, Figure 10 which shows the number of pixels corresponding to each set depth.
[0186] Step 903: Filter out the set depths whose pixel number indication values are less than the number threshold.
[0187] Specifically, the depth map measured by the depth sensor, such as the depth measured by an RGBD camera, Figure 1 generally has many randomly distributed noise points, which are likely to form outliers during the process of determining the target position and reduce the stability of the control system. Therefore, according to the pixel number indication value corresponding to the corresponding set depth determined by statistics, the set depths whose pixel number indication values are less than the number threshold are filtered out, so as to remove the set depths belonging to outliers and not use them for determining the target depth, improving the accuracy of target depth determination and also improving the stability of the control system.
[0188] Step 904: Determine the target depth from each set depth according to the peak value of the pixel number indication value.
[0189] Specifically, the foreground depth and the background depth are determined from each set depth. Among them, the background depth has the maximum peak value of the pixel number indication value, and the foreground depth has the first peak value of the pixel number indication value in ascending order of depth. According to the pixel number indication values corresponding to the foreground depth and the background depth, the target depth is selected from the foreground depth and the background depth.
[0190] As a possible implementation manner of the embodiment of the present application, if the ratio of the pixel number indication values of the foreground depth and the background depth is greater than the ratio threshold, the foreground depth is used as the target depth; if the ratio of the pixel number indication values of the foreground depth and the background depth is not greater than the ratio threshold, the background depth is used as the target depth.
[0191] In one scenario,Figure 11 Another histogram schematic diagram provided for this application. As Figure 11 shown, in ascending order of depth, the maximum peak of the pixel point number indication value is the peak indicated by c. Therefore, the depth corresponding to the peak indicated by c is the foreground depth. And the maximum peak of the pixel point number indication value is the peak indicated by b. Therefore, the depth corresponding to the peak indicated by b is the background depth. Determine the ratio of the pixel point number indication value corresponding to the foreground depth indicated by c to the pixel point number indication value corresponding to the background depth indicated by d, and compare it with the ratio threshold. If the determined ratio is less than the ratio threshold, it is considered that the foreground is an occluder and the background is the target. For example, if the target is a human face or a human body, then the background depth indicated by d is used as the target depth.
[0192] In another scenario, as Figure 10 shown, the peak indicated by a is the maximum peak of the pixel point number indication value in ascending order of depth, which is the foreground depth. The peak indicated by b is the maximum peak of the pixel point number indication value, which is the background depth. If the ratio of the pixel point number indication value corresponding to the foreground depth indicated by a to the pixel point number indication value corresponding to the background depth indicated by b is greater than the ratio threshold, it is considered that the background is an occluder and the foreground is the target, and then the foreground depth indicated by a is used as the target depth.
[0193] Step 905: Locate the first observation position where the object is located according to the target depth.
[0194] Among them, the first observation position includes the observation distance and the observation angle.
[0195] Specifically, according to the angles carried by the pixel points corresponding to the target depth, locate the observation angle of the object. As a possible implementation, the angles carried by the pixel points corresponding to the target depth can be weighted and averaged to calculate the observation angle of the object; according to the depths carried by the pixel points corresponding to the target depth, locate the observation distance of the object. As a possible implementation, the depths carried by the pixel points corresponding to the target depth can be weighted and averaged to calculate the depth information of the object, and the observation distance of the object is located according to the depth information of the object. Furthermore, according to the located observation distance and observation angle of the object, locate the first observation position where the object is located.
[0196] Step 906: Determine the measurement confidence of the first observation position according to the pixel point number indication value corresponding to the target depth.
[0197] Among them, the pixel point number indication value corresponding to the target depth is proportional to the measurement confidence of the first observation position. That is to say, the larger the pixel point number indication value corresponding to the target depth, the higher the measurement confidence of the first observation position.
[0198] In an actual scenario, the target depth can be the background depth or the foreground depth. Therefore, when determining the confidence level of the first observation position, it is also determined according to the corresponding background depth or foreground depth.
[0199] In one scenario, if the target depth is the background depth, the measurement confidence level of the target position can be determined according to the pixel count indication value corresponding to the background depth. As a possible implementation, the measurement confidence level of the first observation position is determined using the ratio of the pixel count indication value of the background depth to the pixel count indication value of the foreground depth. For example, the confidence level is C1, and C1 = (back_ratio / doubless_ratio) * 100, where back_ratio is the ratio of the pixel count indication value of the background depth to the pixel count indication value of the foreground depth, and doubless_ratio is the full score threshold of the confidence level.
[0200] In another scenario, if the target depth is the foreground depth, the measurement confidence level of the target position can be determined according to the pixel count indication value corresponding to the foreground depth. As a possible implementation, the measurement confidence level of the first observation position is determined using the ratio of the pixel count indication value of the foreground depth to the pixel count indication value of the background depth. For example, the confidence level is C2, and C2 = (fore_ratio / doubless_ratio) * 100, where fore_ratio is the ratio of the pixel count indication value of the foreground depth to the pixel count indication value of the background depth, and doubless_ratio is the full score threshold of the confidence level.
[0201] Step 907, if the foreground depth and the background depth are the same, increase the measurement confidence level of the first observation position.
[0202] Among them, the measurement confidence level is used to Figure 5 generate a measurement noise term for the observation equation used in the iterative correction process described in the embodiment. The observation equation is used to substitute the predicted motion state into the observation equation after obtaining the predicted motion state according to the historical motion state adopted in the iterative correction process, so as to obtain the predicted observation position.
[0203] Specifically, if the foreground depth and the background depth are the same, that is, the foreground and background depths coincide, which means that the accuracy of the currently determined first observation position is relatively high, then the measurement confidence level of the first observation position needs to be increased. As a possible implementation, the measurement confidence level of the first observation position determined using the foreground depth can be increased to 2 times, that is, 2C2, or the measurement confidence level of the first observation position determined using the foreground depth can be increased to 2 times, that is, 2C1.
[0204] It should be noted that the method for determining the second observation position of the positioning object, the method for determining the measurement confidence of the second observation position, and the method for determining the measurement confidence of the first observation position have the same principle, which will not be elaborated here.
[0205] In the positioning method of the present application, for the first positioning information carried by each pixel point in the first target area determined in the point cloud map collected by the lidar, the pixel points where the first positioning information matches each set depth are counted, and the pixel point number indication value corresponding to each set depth is obtained. According to the pixel point number indication value corresponding to each set depth, the set depths with pixel point number indication values less than the number threshold are filtered out to remove the noise depth information introduced in the depth map and avoid the impact on the system. And according to the peak value of the pixel point number indication value, the target depth is determined from each set depth to locate the first observation position of the positioning object and determine the measurement confidence of the first observation position, realizing the processing of the data collected by the lidar to improve the accuracy of the target position determined by subsequent data fusion based on multiple sensors.
[0206] Based on the above embodiments, in an embodiment of the present application, according to the second positioning information carried by the second target area, the second observation position of the object is determined. For the specific implementation method, reference can be made to the explanation of determining the first observation position of the object according to the first positioning information carried by the first target area in the previous embodiment. The implementation principles are the same and will not be elaborated here.
[0207] To implement the above embodiments, the present application also proposes a positioning device.
[0208] Figure 12 It is a schematic structural diagram of a positioning device provided by an embodiment of the present application.
[0209] As Figure 12 shown, the device includes: an acquisition module 81, an extraction module 82, and a fusion module 83.
[0210] The acquisition module 81 is used to acquire the point cloud map measured by the lidar.
[0211] The above acquisition module 81 is also used to acquire the depth map measured by the depth sensor.
[0212] The extraction module 82 is used to extract the first target area from the point cloud map and extract the second target area from the depth map, where the first target area and the second target area detect the same object as the region of interest in the synchronously acquired visual image.
[0213] The fusion module 83 is used to fuse the first positioning information carried by the first target area and the second positioning information carried by the second target area according to the historical motion state of the object to obtain the target position of the object.
[0214] Further, in a possible implementation manner of the embodiment of the present application, the above-mentioned fusion module 83 is specifically configured to:
[0215] Determine the first observation position where the object is located according to the first positioning information carried by the first target area, and determine the second observation position where the object is located according to the second positioning information carried by the second target area; according to the historical motion state of the object, perform an iterative correction process on one of the first observation position and the second observation position, update the historical motion state according to the position obtained by the iterative correction process, and according to the updated historical motion state, perform an iterative correction process on the other of the first observation position and the second observation position to obtain the target position.
[0216] As a possible implementation manner, the above-mentioned fusion module 83 is specifically configured to:
[0217] Obtain the predicted motion state according to the historical motion state adopted in the current iterative correction process; the historical motion state adopted in the current iterative correction process is generated according to the position obtained by the previous iterative correction process and the historical motion state adopted in the previous iterative correction process; obtain the predicted observation position according to the predicted motion state; correct the first observation position or the second observation position undergoing the current iterative correction process according to the predicted observation position.
[0218] As a possible implementation manner, the historical motion state includes the historical position and historical speed of the object;
[0219] Correspondingly, for each iterative correction process, the historical position in the historical motion state adopted in the current iterative correction process is generated according to the position obtained by the previous iterative correction process and the historical position adopted in the previous iterative correction process; for each iterative correction process, the historical speed in the historical motion state adopted in the current iterative correction process is determined according to the historical position adopted in the current iterative correction process and the historical position adopted in the previous iterative correction process.
[0220] As a possible implementation manner, the above-mentioned fusion module 83 is further configured to:
[0221] Compare the first observation moment when the point cloud map is obtained and the second observation moment when the depth map is obtained; if the first observation moment is earlier than the second observation moment, determine that the first observation position is iteratively corrected earlier than the second observation position; if the first observation moment is later than the second observation moment, determine that the second observation position is iteratively corrected earlier than the first observation position; if the first observation moment is equal to the second observation moment, randomly determine the order of iterative correction of the first observation position and the second observation position.
[0222] As a possible implementation manner, the above-mentioned fusion module 83 is specifically configured to:
[0223] Substitute the predicted motion state into the observation equation to obtain the predicted observation position;
[0224] Wherein, the observation equation is the product of the predicted motion state and the transformation matrix plus the measurement noise term; the transformation matrix is used to indicate the transformation relationship between the predicted motion state and the predicted observation position; the measurement noise term conforms to the Gaussian white noise distribution with a set covariance; the set covariance is determined according to the device accuracy and the measurement confidence level.
[0225] As a possible implementation manner, the above-mentioned fusion module 83 is specifically configured to:
[0226] Determine the measurement residual for the first observation position or the second observation position in this iterative correction process; if the measurement residual is less than the difference threshold, then correct the first observation position or the second observation position in this iterative correction process according to the predicted observation position.
[0227] As a possible implementation manner, the measurement residual is the difference between the predicted observation position and the first observation position or the second observation position in this iterative correction process.
[0228] Furthermore, in a possible implementation manner of the embodiment of the present application, the above-mentioned extraction module 82 is specifically configured to:
[0229] Determine the rectangular coordinate position in the image coordinate system for the region of interest; map the rectangular coordinate position in the image coordinate system to the polar coordinate system of the lidar to obtain the polar coordinate position; extract the first target region from the point cloud map according to the polar coordinate position.
[0230] As a possible implementation manner, the above-mentioned extraction module 82 is specifically configured to:
[0231] Determine the rectangular coordinate positions for the left and right boundaries of the region of interest.
[0232] As a possible implementation manner, the above-mentioned extraction module 82 is specifically configured to:
[0233] Map the rectangular coordinate position in the image coordinate system to the camera coordinate system through the internal parameter matrix of the camera to obtain the rectangular coordinate position in the camera coordinate system, where the camera is used to collect visual images, and map the rectangular coordinate position in the camera coordinate system to the polar coordinate system through the external parameter matrix between the camera and the lidar to obtain the polar coordinate position.
[0234] Furthermore, in a possible implementation manner of the embodiment of the present application, the above-mentioned fusion module 83 is specifically configured to:
[0235] Determine the depth of each pixel point in the first target area according to the first positioning information carried by the first target area, and count the pixel point number indication values corresponding to each set depth according to the depth of each pixel point; determine the target depth from each set depth according to the peak value of the pixel point number indication value; locate the first observation position where the object is located according to the target depth.
[0236] As a possible implementation manner, the above-mentioned fusion module 83 is specifically configured to:
[0237] Determine the foreground depth and the background depth from each set depth; wherein, the background depth has the maximum peak value of the pixel point number indication value; the foreground depth has the first peak value of the pixel point number indication value in the order of increasing depth from small to large, and select the target depth from the foreground depth and the background depth according to the pixel point number indication values corresponding to the foreground depth and the background depth.
[0238] As a possible implementation manner, the above-mentioned fusion module 83 is specifically configured to:
[0239] If the ratio of the pixel point number indication value of the foreground depth to that of the background depth is greater than the ratio threshold, use the foreground depth as the target depth; if the ratio of the pixel point number indication value of the foreground depth to that of the background depth is not greater than the ratio threshold, use the background depth as the target depth.
[0240] As a possible implementation manner, the above-mentioned fusion module 83 is further configured to:
[0241] Determine the measurement confidence of the first observation position according to the pixel point number indication value corresponding to the target depth; if the foreground depth and the background depth are the same, increase the measurement confidence of the first observation position, wherein the measurement confidence is used to generate a measurement noise term for the observation equation adopted in the iterative correction process, and the observation equation is used to substitute the predicted motion state into the observation equation after obtaining the predicted motion state according to the historical motion state adopted in the iterative correction process to obtain the predicted observation position.
[0242] As a possible implementation manner, the above-mentioned fusion module 83 is further configured to:
[0243] Filter out the set depths with pixel point number indication values less than the number threshold.
[0244] As a possible implementation manner, the above-mentioned fusion module 83 is specifically configured to:
[0245] For each set depth, determine the depth statistical range, and count the number of pixel points whose depth matches the corresponding depth statistical range according to the depth of each pixel point, so as to obtain the pixel point number indication value corresponding to the corresponding set depth.
[0246] As a possible implementation, the first observation position includes an observation distance and an observation angle, and the first positioning information includes a depth and an angle;
[0247] The above-mentioned fusion module 83 is specifically configured to locate the observation angle of the object according to the angles carried by the pixel points corresponding to the target depth, and locate the observation distance of the object according to the depths carried by the pixel points corresponding to the target depth.
[0248] It should be noted that the foregoing explanation of the positioning method embodiment also applies to the positioning device of this embodiment, and will not be elaborated here.
[0249] To implement the above embodiment, the present application also proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the positioning method described in the foregoing method embodiment is implemented.
[0250] The electronic device proposed by the present application may be, but is not limited to, a robot.
[0251] To implement the above embodiment, the present application also proposes a non-transitory computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the positioning method described in the foregoing method embodiment is implemented.
[0252] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0253] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0254] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0255] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing a logical function, and can be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0256] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0257] Those of ordinary skill in the art can understand that all or part of the steps carried out in the method of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0258] In addition, in each of the embodiments of the present application, the functional units can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0259] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A positioning method, characterized in that, The method includes: Obtaining a point cloud map measured by a lidar; Obtaining a depth map measured by a depth sensor; Extracting a first target area from the point cloud map and a second target area from the depth map; wherein, the first target area and the second target area detect the same object as the region of interest; the region of interest is obtained by detecting a visual image collected by a camera; Fusing the first positioning information carried by the first target area and the second positioning information carried by the second target area according to the historical motion state of the object to obtain the target position of the object; The fusing the first positioning information carried by the first target area and the second positioning information carried by the second target area according to the historical motion state of the object to obtain the target position of the object includes: Determining a first observation position where the object is located according to the first positioning information carried by the first target area; Determining a second observation position where the object is located according to the second positioning information carried by the second target area; Performing an iterative correction process on one of the first observation position and the second observation position according to the historical motion state of the object; Updating the historical motion state according to the position obtained by the iterative correction process; Performing an iterative correction process on the other of the first observation position and the second observation position according to the updated historical motion state to obtain the target position.
2. The positioning method according to claim 1, characterized in that, The performing the iterative correction process includes: Obtaining a predicted motion state according to the historical motion state adopted in the current iterative correction process; the historical motion state adopted in the current iterative correction process is generated according to the position obtained by the previous iterative correction process and the historical motion state adopted in the previous iterative correction process; Obtaining a predicted observation position according to the predicted motion state; Correcting the first observation position or the second observation position for the current iterative correction process according to the predicted observation position.
3. The positioning method according to claim 1, characterized in that, Before performing the iterative correction process on one of the first observation position and the second observation position according to the historical motion state of the object, it further includes: Comparing a first observation time when the point cloud map is obtained and a second observation time when the depth map is obtained; If the first observation time is earlier than the second observation time, determining that the first observation position is iteratively corrected earlier than the second observation position; If the first observation time is later than the second observation time, determining that the second observation position is iteratively corrected earlier than the first observation position; If the first observation time is equal to the second observation time, randomly determining the order of iterative correction of the first observation position and the second observation position.
4. The positioning method according to claim 2, wherein The historical motion state includes the historical position and historical speed of the object; Correspondingly, for each iterative correction process, the historical position in the historical motion state adopted in the current iterative correction process is generated according to the position obtained by the previous iterative correction process and the historical position adopted in the previous iterative correction process; For each iteration correction process, the historical speed in the historical motion state adopted in this iteration correction process is determined according to the historical position adopted in this iteration correction process and the historical position adopted in the previous iteration correction process.
5. The positioning method according to claim 2, wherein The obtaining of the predicted observation position according to the predicted motion state includes: Substituting the predicted motion state into the observation equation to obtain the predicted observation position; Wherein, the observation equation is the product of the predicted motion state and the transformation matrix superimposed with the measurement noise term; The transformation matrix is used to indicate the transformation relationship between the predicted motion state and the predicted observation position; The measurement noise term conforms to a Gaussian white noise distribution with a set covariance; the set covariance is determined according to the device accuracy and measurement confidence.
6. The positioning method according to claim 2, wherein The correction of the first observation position or the second observation position for this iteration correction process according to the predicted observation position includes: Determining a measurement residual for the first observation position or the second observation position for this iteration correction process; If the measurement residual is less than the difference threshold, the first observation position or the second observation position for this iteration correction process is corrected according to the predicted observation position.
7. The positioning method according to claim 6, characterized in that, The measurement residual is the difference between the predicted observation position and the first observation position or the second observation position for this iteration correction process.
8. The positioning method according to any one of claims 1-7, characterized in that, The extraction of the first target area from the point cloud map includes: Determining the rectangular coordinate position in the image coordinate system for the region of interest; Mapping the rectangular coordinate position in the image coordinate system to the polar coordinate system of the lidar to obtain a polar coordinate position; Extracting the first target area from the point cloud map according to the polar coordinate position.
9. The positioning method according to claim 8, wherein The determining of the rectangular coordinate position in the image coordinate system for the region of interest includes: Determining the rectangular coordinate position for the left and right boundaries of the region of interest.
10. The positioning method according to claim 8, characterized in that, The mapping of the rectangular coordinate position in the image coordinate system to the polar coordinate system of the lidar to obtain a polar coordinate position includes: Mapping the rectangular coordinate position in the image coordinate system to the camera coordinate system through the internal parameter matrix of the camera to obtain the rectangular coordinate position in the camera coordinate system; wherein, the camera is used to collect the visual image; Mapping the rectangular coordinate position in the camera coordinate system to the polar coordinate system through the external parameter matrix between the camera and the lidar to obtain the polar coordinate position.
11. The positioning method according to any one of claims 1-7, characterized in that, The determining of the first observation position where the object is located according to the first positioning information carried by the first target area includes: Determining the depth of each pixel point in the first target area according to the first positioning information carried by the first target area; Statistically calculating the pixel point number indication value corresponding to each set depth according to the depth of each pixel point; Determining the target depth from each set depth according to the peak value of the pixel point number indication value; Locating the first observation position where the object is located according to the target depth.
12. The positioning method according to claim 11, characterized in that, The determining of the target depth from each set depth according to the peak value of the pixel point number indication value includes: Determine a foreground depth and a background depth from each set depth; wherein, the background depth has the maximum peak of the pixel number indication value; the foreground depth has the first peak of the pixel number indication value in the order of increasing depth; Select the target depth from the foreground depth and the background depth according to the pixel number indication values corresponding to the foreground depth and the background depth.
13. The positioning method according to claim 12, wherein The selecting the target depth from the foreground depth and the background depth according to the pixel number indication values corresponding to the foreground depth and the background depth includes: If the ratio of the pixel number indication values of the foreground depth and the background depth is greater than a ratio threshold, use the foreground depth as the target depth; If the ratio of the pixel number indication values of the foreground depth and the background depth is not greater than the ratio threshold, use the background depth as the target depth.
14. The positioning method according to claim 12, wherein After positioning the first observation position where the object is located according to the target depth, further include: Determine the measurement confidence of the first observation position according to the pixel number indication value corresponding to the target depth; If the foreground depth and the background depth are the same, increase the measurement confidence of the first observation position, wherein the measurement confidence is used to generate a measurement noise term for the observation equation adopted in the iterative correction process, and the observation equation is used to substitute the predicted motion state into the observation equation after obtaining the predicted motion state according to the historical motion state adopted in the iterative correction process to obtain the predicted observation position.
15. The positioning method according to claim 11, characterized in that, After counting the pixel number indication values corresponding to each set depth according to the depths of the respective pixels, further include: Filter out the set depths with pixel number indication values less than a number threshold.
16. The positioning method according to claim 11, characterized in that, The counting the pixel number indication values corresponding to each set depth according to the depths of the respective pixels includes: For each set depth, determine a depth statistical range; According to the depths of the respective pixels, count the number of pixels whose depths match the corresponding depth statistical range to obtain the pixel number indication value corresponding to the corresponding set depth.
17. The positioning method according to claim 11, characterized in that, The first observation position includes an observation distance and an observation angle; the first positioning information includes a depth and an angle; The positioning the first observation position where the object is located according to the target depth includes: Position the observation angle of the object according to the angles carried by the respective pixels corresponding to the target depth; Position the observation distance of the object according to the depths carried by the respective pixels corresponding to the target depth.
18. A positioning device, characterized in that, The device includes: An acquisition module, configured to acquire a point cloud map measured by a lidar; The acquisition module is further configured to acquire a depth map measured by a depth sensor; An extraction module, configured to extract a first target area from the point cloud map and extract a second target area from the depth map; wherein, the first target area and the second target area detect the same object as the region of interest; the region of interest is obtained by detecting a visual image acquired by a camera; A fusion module, configured to fuse the first positioning information carried by the first target area and the second positioning information carried by the second target area according to the historical motion state of the object to obtain the target position of the object; Specifically, the fusion module is configured to determine a first observation position where the object is located according to the first positioning information carried by the first target area; determine a second observation position where the object is located according to the second positioning information carried by the second target area; perform an iterative correction process on one of the first observation position and the second observation position according to the historical motion state of the object; update the historical motion state according to the position obtained by the iterative correction process; and perform an iterative correction process on the other of the first observation position and the second observation position according to the updated historical motion state to obtain the target position.
19. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the positioning method according to any one of claims 1-17 is implemented.
20. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the positioning method according to any one of claims 1-17 is implemented.
Citation Information
Patent Citations
Unmanned aerial vehicle intelligent perception system and method based on multiple sensors
CN107450577A