Attitude determination method and device, driver attitude determination method and device and vehicle
By combining two-dimensional images and three-dimensional point cloud perspective multi-point algorithm and iterative close-point algorithm, the problem that the ICP algorithm is difficult to converge in driver head attitude determination is solved, achieving higher accuracy and stable attitude matching.
Patent Information
- Application Number
- CN202510185525.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-07-22
AI Technical Summary
The existing iterative nearest point algorithm (ICP) has high requirements for initial values when determining the driver's head attitude and is difficult to converge, resulting in low posture accuracy, especially when the driver's head attitude changes.
By obtaining the driver's two-dimensional captured images and three-dimensional point clouds, the reference pose is determined using the perspective multi-point algorithm (PNP), and point cloud matching is combined with the ICP algorithm to introduce joint optimization of two-dimensional matching degree and three-dimensional matching degree to ensure the accuracy and stability of the initial pose value.
It improves the accuracy and stability of driver attitude determination, avoids solution failure, reduces the amount of calculation, and improves the matching accuracy during attitude changes.
Smart Images

Figure CN120355783A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of vehicles, and more particularly, to a method for determining an attitude, a method for determining a driver's attitude, an electronic device, a vehicle, a computer program product, and a non-volatile computer-readable storage medium including a computer program. Background Art
[0002] With the development of in-vehicle camera technology and deep learning algorithm software, the content of vehicle safety has also been extended. The in-vehicle near-infrared camera hardware and the behavior recognition algorithm installed on the vehicle computer can judge the state of the driver during the driving process and determine whether the driver is in a distracted state. This requires obtaining an accurate head pose of the driver without introducing interference items. The depth camera can be used to achieve non-contact and non-interference assistance. The head point cloud at the initial moment and the head point cloud at the current moment are matched to solve the attitude change amount, and the initial attitude is added to obtain the head attitude.
[0003] In the related art, the head pose is usually obtained based on the Iterative Closest Point (ICP) algorithm. Since the existing ICP algorithm has high requirements for the initial value, it is difficult for the initial value of the attitude matrix without prior information to converge, resulting in low accuracy of the output attitude. Summary of the Invention
[0004] The embodiments of the present application provide a method for determining an attitude, a method for determining a driver's attitude, an electronic device, a vehicle, a program product, and a medium, which can ensure the accuracy of the output attitude.
[0005] An attitude determination method according to an embodiment of the present application includes: obtaining a captured image of a target object and a target point cloud; determining a reference attitude of the target object based on the captured image; and determining an initial attitude of the target object based on the reference attitude and the target point cloud.
[0006] In some embodiments, the determining the reference attitude of the target object based on the captured image includes: obtaining two-dimensional target key points of the captured image; determining three-dimensional target key points corresponding to each of the two-dimensional target key points based on the two-dimensional target key points and a preset mapping relationship; and determining the reference attitude based on the two-dimensional target key points and the corresponding three-dimensional target key points.
[0007] In some embodiments, determining the reference pose based on the two-dimensional target key points and the corresponding three-dimensional target key points includes: matching the two-dimensional target key points and the corresponding three-dimensional target key points based on the perspective multi-point algorithm to obtain m matching pairs, each matching pair including a two-dimensional target key point and a matching three-dimensional target key point; determining the reference pose based on the m matching pairs.
[0008] In some embodiments, it further includes: projecting the target point cloud model onto a two-dimensional plane based on the reference pose and a preset target point cloud model to obtain a two-dimensional image of the target point cloud model in the reference pose; determining two-dimensional feature points corresponding to the three-dimensional feature points of the target point cloud model based on the two-dimensional image; updating the mapping relationship based on the three-dimensional feature points of the target point cloud model and the two-dimensional feature points corresponding to the three-dimensional feature points.
[0009] In some embodiments, it further includes: determining a two-dimensional matching degree based on the two-dimensional target key points of the captured image and a preset mapping relationship; determining a three-dimensional matching degree based on n pairs of point clouds obtained by matching the target point cloud and a source point cloud, where the source point cloud is determined based on a preset target point cloud model; determining whether the two-dimensional matching degree and the three-dimensional matching degree satisfy a preset iterative convergence condition or whether the number of iterations reaches a preset number; if so, outputting a first target initial pose based on the initial poses in multiple iterative processes; if not, re-entering the step of acquiring the captured image and the target point cloud of the target object.
[0010] In some embodiments, determining the two-dimensional matching degree based on the two-dimensional target key points of the captured image and a preset mapping relationship includes: determining the three-dimensional target key points corresponding to each of the two-dimensional target key points based on the two-dimensional target key points of the captured image and the preset mapping relationship; matching the two-dimensional target key points and the corresponding three-dimensional target key points based on the perspective multi-point algorithm to obtain m matching pairs, each matching pair including a two-dimensional target key point and a matching three-dimensional target key point; determining the two-dimensional matching degree based on the m matching pairs.
[0011] In some embodiments, the two-dimensional matching degree Z1 is implemented based on the following formula: where w2d is 1 / m, K is the internal parameter of the camera that captures the image, pj is the point cloud coordinate of the j-th point of the source point cloud, r is the rotation angle, t is the offset, and l i is the image coordinate of the captured image.
[0012] In some embodiments, determining the three-dimensional matching degree based on the n pairs of point clouds obtained by matching the target point cloud and the source point cloud includes: matching the target point cloud and the source point cloud based on the iterative closest point algorithm to obtain the n pairs of point clouds; and determining the three-dimensional matching degree based on the n pairs of point clouds.
[0013] In some embodiments, the three-dimensional matching degree Z2 is implemented based on the following formula: where w3d = 1 / n, pi is the point cloud coordinate of the source point cloud, r is the rotation angle, t is the offset, q i is the point cloud coordinate of the target point cloud, T is the matrix transpose, and n qi is the normal vector of q i
[0014] In some embodiments, outputting the first target initial pose based on the initial poses in multiple iteration processes includes: when the two-dimensional matching degree and the three-dimensional matching degree meet the preset iteration convergence condition, determining the initial pose that meets the preset iteration convergence condition as the first target initial pose; when the number of iterations reaches the preset number, determining the initial pose with the minimum sum of the two-dimensional matching degree and the three-dimensional matching degree as the first target initial pose.
[0015] In some embodiments, it further includes: adjusting a preset target point cloud model based on the target point cloud and the initial pose.
[0016] In some embodiments, adjusting the preset target point cloud model based on the target point cloud and the initial pose includes: transforming the target point cloud model based on the initial pose so that the pose of the target point cloud model is consistent with the initial pose; and adjusting the transformed target point cloud model based on the target point cloud so that the geometric features of the target point cloud model are consistent with the target object.
[0017] In some embodiments, the target point cloud model includes a three-dimensional deformable model.
[0018] In some embodiments, determining the initial pose of the target object based on the reference pose and the target point cloud includes: transforming the target point cloud based on the reference pose so that the pose of the target point cloud is consistent with the pose of a preset source point cloud; and matching the transformed target point cloud with the source point cloud to determine the initial pose.
[0019] In some embodiments, it further includes: determining a second target initial pose of the target object based on the target point cloud and a preset source point cloud; outputting the one with a smaller three-dimensional matching degree between the first target initial pose and the second target initial pose as the initial pose of the target object.
[0020] In some embodiments, it further includes: performing distraction detection on the target object based on the initial pose of the target object to determine whether the target object is distracted.
[0021] In some embodiments, the performing distraction detection on the target object based on the initial pose of the target object to determine whether the target object is distracted includes: determining the pitch angle and yaw angle of the target object based on the initial pose of the target object; determining that the target object is distracted when the pitch angle or the yaw angle is outside the corresponding angle range; and determining that the target object is not distracted when both the pitch angle and the yaw angle are within the corresponding angle range.
[0022] In some embodiments, the determining the initial pose of the target object based on the reference pose and the target point cloud includes: transforming the target point cloud based on the reference pose so that the pose of the target point cloud is consistent with the pose of a preset source point cloud; and performing matching between the transformed target point cloud and the source point cloud to determine the initial pose.
[0023] This application also provides a method for determining a driver's pose, including: acquiring a captured image and a target point cloud of the driver; determining a reference pose of the driver based on the captured image; and determining an initial pose of the driver based on the reference pose and the target point cloud.
[0024] This application provides an electronic device, including a processor, where the processor is connected to a memory; a computer program is stored in the memory, and the processor executes the computer program to implement the pose determination method described in any of the above embodiments; and / or implement the driver pose determination method described in any of the above embodiments.
[0025] This application provides a vehicle, including the electronic device described in any of the above embodiments.
[0026] This application also provides a computer program product, including a computer program, where the computer program includes instructions for executing the pose determination method described in any of the above embodiments, or instructions for executing the driver pose determination method described in any of the above embodiments.
[0027] The present application also provides a non - volatile computer - readable storage medium containing a computer program. When the computer program is executed by a processor, the processor is caused to execute the posture determination method or the driver posture determination method described in any of the above - mentioned embodiments.
[0028] In the posture determination method, the driver posture determination method, the electronic device, the vehicle, the program product and the medium according to the embodiments of the present application, by acquiring a captured image of a target object (such as a driver) and a target point cloud (for example, acquiring a captured image of the target object through an in - vehicle image camera of the vehicle and acquiring a target point cloud of the target object through a depth camera), then based on the captured image, (by means of, such as, a perspective multi - point algorithm, etc.) determining a reference posture of the target object, and finally based on the reference posture and the target point cloud, determining an initial posture of the target object. Since a two - dimensional captured image of the target object is introduced for posture reference when determining the initial posture, it can not only improve the accuracy of the obtained initial posture of the target object, but also enable more nearest - neighbor matching point pairs to be obtained when determining the posture of the target object based on the target point cloud, providing a better initial posture value for subsequent solution, improving the stability during the solution of point cloud transformation, and avoiding the problem of solution failure.
[0029] Additional aspects and advantages of the embodiments of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The above - mentioned and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:
[0031] Figure 1 is a schematic diagram of an application scenario of the posture determination method according to some embodiments of the present application;
[0032] Figure 2 is a flowchart of the posture determination method according to some embodiments of the present application;
[0033] Figure 3 is a flowchart of the posture determination method according to some embodiments of the present application;
[0034] Figure 4 is a flowchart of the posture determination method according to some embodiments of the present application;
[0035] Figure 5 is a flowchart of the posture determination method according to some embodiments of the present application;
[0036] Figure 6 is a schematic diagram of a scenario of the posture determination method according to some embodiments of the present application;
[0037] Figure 7 It is a schematic flowchart of the attitude determination method according to some embodiments of the present application;
[0038] Figure 8 It is a schematic flowchart of the attitude determination method according to some embodiments of the present application;
[0039] Figure 9 It is a schematic flowchart of the attitude determination method according to some embodiments of the present application;
[0040] Figure 10 It is a schematic flowchart of the attitude determination method according to some embodiments of the present application;
[0041] Figure 11 It is a schematic diagram of the scenario of the attitude determination method according to some embodiments of the present application;
[0042] Figure 12 It is a schematic diagram of the scenario of the attitude determination method according to some embodiments of the present application;
[0043] Figure 13 It is a schematic diagram of the scenario of the attitude determination method according to some embodiments of the present application;
[0044] Figure 14 It is a schematic diagram of the scenario of the attitude determination method according to some embodiments of the present application;
[0045] Figure 15 It is a schematic diagram of the scenario of the attitude determination method according to some embodiments of the present application;
[0046] Figure 16 It is a schematic diagram of the scenario of the attitude determination method according to some embodiments of the present application;
[0047] Figure 17 It is a schematic diagram of the scenario of the attitude determination method according to some embodiments of the present application;
[0048] Figure 18 It is a schematic diagram of the scenario of the attitude determination method according to some embodiments of the present application;
[0049] Figure 19 It is a schematic diagram of the scenario of the attitude determination method according to some embodiments of the present application;
[0050] Figure 20 It is a schematic diagram of the scenario of the attitude determination method according to some embodiments of the present application;
[0051] Figure 21 It is a schematic flowchart of the attitude determination method according to some embodiments of the present application;
[0052] Figure 22 It is a schematic flowchart of the attitude determination method according to some embodiments of the present application;
[0053] Figure 23 is a schematic flow chart of a posture determination method according to some embodiments of the present application;
[0054] Figure 24 is a schematic flow chart of a posture determination method according to some embodiments of the present application;
[0055] Figure 25 is a schematic flow chart of a posture determination method according to some embodiments of the present application;
[0056] Figure 26 is a schematic flow chart of a posture determination method according to some embodiments of the present application;
[0057] Figure 27 is a schematic diagram of a scenario of a posture determination method according to some embodiments of the present application;
[0058] Figure 28 is a schematic flow chart of a posture determination method according to some embodiments of the present application;
[0059] Figure 29 is a schematic diagram of modules of a posture determination device according to some embodiments of the present application;
[0060] Figure 30 is a schematic diagram of modules of a driver posture determination device according to some embodiments of the present application;
[0061] Figure 31 is a schematic diagram of the connection state of a non - volatile computer - readable storage medium and a processor according to some embodiments of the present application. Detailed Embodiments
[0062] The following describes in detail the embodiments of the present application. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the embodiments of the present application and should not be construed as limiting the embodiments of the present application.
[0063] For the convenience of understanding the present application, the following explains the nouns and background technologies that appear in the present application:
[0064] With the development of in - vehicle camera technology and deep - learning algorithm software, the content of vehicle safety has also been expanded. The in - vehicle near - infrared camera hardware and the behavior recognition algorithm installed on the vehicle computer can judge the state of the driver during the driving process of the vehicle and determine whether the driver is in a distracted state. This requires obtaining the accurate head posture of the driver without introducing interference items. The depth camera can be used to achieve non - contact and non - interference assistance. The head point cloud at the initial moment and the head point cloud at the current moment are matched to solve the posture change amount, and the initial posture is added to obtain the head posture.
[0065] The existing classic point cloud matching algorithm is the Iterative Closest Point (ICP) algorithm (an algorithm for point cloud registration). The ICP algorithm is used to solve the registration problem of two point cloud data sets in three-dimensional space. Its purpose is to find a rigid body transformation (including rotation and translation) so that the two point clouds can be optimally aligned. The rotation and translation information contained in this rigid body transformation can be used to represent the posture.
[0066] The ICP algorithm is generally based on the target point cloud and the source point cloud. The target point cloud is the point cloud to which the source point cloud needs to be registered, and is the reference standard for registration. The source point cloud is the original point cloud, which needs to be transformed to achieve registration with the target point cloud. During the calculation process, the source point cloud is continuously iteratively transformed to align with the target point cloud as much as possible, so as to achieve the purpose of minimizing the distance error between the two and other objective functions. In each iteration, the transformation matrix is calculated according to the corresponding relationship with the target point cloud, and the transformation matrix is applied for update, gradually approaching the target point cloud, and finally achieving alignment with the target point cloud. Based on the alignment relationship, the posture to be calculated is determined.
[0067] However, when performing head point cloud matching based on the ICP algorithm, there are the following disadvantages:
[0068] (1) The depth camera usually obtains the point cloud information of the entire scene, not just the point cloud of the head. When calculating based on the ICP algorithm, the calculation amount of the entire scene or human body point cloud matching is large, and the local head effect (for example, the effect of the facial features, etc.) cannot be guaranteed;
[0069] (2) Perform point cloud matching on the designated area of the head in the scene. Since the number of point clouds at different times is different and the matching pairs are unknown, an iterative solution is required. The existing ICP algorithm has high requirements for the initial value, and the initial value of the pose matrix without prior information is difficult to converge, resulting in low output pose accuracy.
[0070] (3) When the driver raises or lowers his head or turns his head left or right, the head point cloud obtained by the depth camera is incomplete, that is, only a part of the face is presented. The existing ICP algorithm will not be able to calculate, that is, it will fail.
[0071] In order to solve the above technical problems, an embodiment of the present application provides a posture determination method.
[0072] The following is an introduction to the application scenarios of the technical solution of this application. The posture determination method provided by this application can be applied to Figure 1 In the vehicle shown.
[0073] Optionally, the vehicle includes a depth camera (3D camera or RGB-D camera) and / or an image camera. The image camera records the planar appearance of an object, including color, texture, etc., by acquiring two-dimensional image information. In addition to acquiring the planar image of an object, the depth camera can also acquire the depth information of the object, that is, the distance from each pixel point to the camera, forming three-dimensional data.
[0074] For example, a portrait photo taken by an image camera can only show the physical features of a person, while a portrait taken by a depth camera can simultaneously obtain the distances between various parts of the person's face and the camera, which can be used for more accurate face recognition or 3D modeling.
[0075] The pose determination method of the present application will be elaborated in detail below:
[0076] Please refer to Figure 1 and Figure 2 . An embodiment of the present application provides a pose determination method. For example, please refer to Figure 1 . Figure 1 An exemplary usage scenario diagram showing the application of the pose determination method to vehicle 100 is shown. The pose determination method includes:
[0077] Step 011: Acquire a captured image of the target object and the target point cloud.
[0078] Among them, the target object can be a user. For example, the pose determination method can be applied to a vehicle, and the target object can be a driver; for another example, when the pose determination method is applied to a motion analysis system, the target object can be an athlete, etc.
[0079] Among them, the captured image can be a two-dimensional image such as the head features of a user acquired by an image camera. For example, the captured image can be the head image of a driver during driving.
[0080] Among them, the target point cloud can be the point cloud of the head position of a user acquired by a depth camera. For example, the target point cloud can be the point cloud data of the head position of a driver acquired by the depth camera of the vehicle during driving. For another example, it can be to acquire the depth image of the current scene through a depth camera, obtain the point cloud data in the current scene, and then preprocess the acquired point cloud data, including denoising, filtering, segmentation, etc., to remove irrelevant background point clouds and noise points, so as to obtain the target point cloud that can represent the head position of the driver. For the convenience of narration, in the following narration, an example will be given where the pose determination method is applied to a vehicle, the target object includes a driver, the captured image includes a two-dimensional image of the driver's head, and the target point cloud includes the point cloud data of the driver's head position.
[0081] Specifically, the vehicle acquires a captured image of the target object through an image camera, and scans the target object through a depth camera to obtain point cloud data of the target object during driving. The point cloud data may include three-dimensional coordinates of the surface of the target object, etc.
[0082] Step 012: Based on the captured image, determine the reference pose of the target object.
[0083] Among them, the reference pose can be considered as the position and orientation of the head of the target object in space when the target object is concentrating on driving. For example, the reference pose can be represented by a rotation matrix, etc.
[0084] Specifically, for example, in the case of obtaining the captured image, the captured image can be preprocessed first, including denoising, enhancing contrast, edge detection, etc., to improve the accuracy of subsequent processing, and then feature points or feature regions, etc. (for example, corner points, edges, textures, etc. in the captured image) can be extracted from the preprocessed image. Based on the feature points or feature regions, through deep learning algorithms, etc., the reference pose corresponding to the target object in the captured image is determined.
[0085] Step 013: Based on the reference pose and the target point cloud, determine the initial pose of the target object.
[0086] Among them, please refer to Figure 3 , optionally, Step 013: Based on the reference pose and the target point cloud, determine the initial pose of the target object, including:
[0087] Step 0131: Transform the target point cloud based on the reference pose so that the pose of the target point cloud is consistent with the pose of a preset source point cloud;
[0088] Step 0132: Match the transformed target point cloud with the source point cloud to determine the initial pose.
[0089] Specifically, in the case of determining the reference pose and obtaining the target point cloud of the target object, the reference pose of the target object represented by the two-dimensional image and the target point cloud obtained based on the depth camera can be aligned through a registration algorithm (such as the ICP algorithm, etc.) (for example, the threshold between the point cloud data corresponding to the reference pose and the point cloud data of the target point cloud is less than a preset threshold, etc.), and then the target point cloud with the same pose is matched with the preset source point cloud, so as to determine the initial pose of the target object.
[0090] In this way, by obtaining the captured image of the target object (such as the driver) and the target point cloud (for example, obtaining the captured image of the target object through the in-vehicle image camera of the vehicle and obtaining the target point cloud of the target object through the depth camera), and then based on the captured image, (through algorithms such as the perspective multi-point algorithm) determining the reference pose of the target object, and finally based on the reference pose and the target point cloud, determining the initial pose of the target object. Since the two-dimensional captured image of the target object is introduced when determining the initial pose, it can not only improve the accuracy of the obtained initial pose of the target object, but also enable more nearest neighbor matching point pairs to be obtained when determining the pose of the target object based on the target point cloud, providing a better initial pose value for subsequent solution, improving the stability during the solution of the point cloud transformation, and avoiding the problem of solution failure.
[0091] Moreover, compared with the current algorithms that only solve the pose based on the ICP algorithm, the calculation amount of this application is smaller and the accuracy is higher.
[0092] Please refer to Figure 4 , in some embodiments, step 012: Based on the captured image, determine the reference pose of the target object, including:
[0093] Step 0121: Obtain the two-dimensional target key points of the captured image;
[0094] Step 0122: Based on the two-dimensional target key points and the preset mapping relationship, determine the three-dimensional target key points corresponding to each two-dimensional target key point;
[0095] Step 0123: Based on the two-dimensional target key points and the corresponding three-dimensional target key points, determine the reference pose.
[0096] Among them, the two-dimensional target key points can be multiple important landmark points used to represent the head and face of the target user. For example, the two-dimensional target key points can be two-dimensional head key points, and the three-dimensional target key points can be three-dimensional head key points, such as the eyes, nose, mouth, and chin of the face.
[0097] Among them, please refer to Figure 5 , optionally, step 0123: Based on the two-dimensional target key points and the corresponding three-dimensional target key points, determine the reference pose, including:
[0098] Step 01231: Based on the perspective multi-point algorithm, match the two-dimensional target key points and the corresponding three-dimensional target key points to obtain m matching pairs, and each matching pair includes a two-dimensional target key point and a matching three-dimensional target key point;
[0099] Step 01232: Based on the m matching pairs, determine the reference pose.
[0100] Among them, the Perspective-n-Point (PNP) algorithm is a method for solving the correspondence between 3D and 2D points. For example, please refer to Figure 6 , Figure 6 which exemplarily shows a model for solving 2D key points of the head based on the PNP algorithm.
[0101] Specifically, through a pre-trained face detection and face key point detection model, for example, through an image processing method: such as the Haar cascade classifier based on the cross-platform computer vision and machine learning software library (OpenCV) (an object detection algorithm based on Haar features and cascade structure); or, based on deep learning methods such as the Mediapipe algorithm (an algorithm for detecting facial, palm, and human pose features), OpenPose (a human pose estimation algorithm), etc., or, through a pre-trained deep learning model (such as a convolutional neural network model (CNN model), etc.) to process the captured image to obtain the target key points in the 2D captured image (for example, represented by the image coordinates in the 2D captured image). Then, through the preset mapping relationship between the 2D target key points and the 3D target key points, to determine the position of the 2D target key points in the 3D space. For example, taking R to represent the camera pose calibrated by the depth camera and T to represent the camera pose calibrated by the image camera as an example for illustration, the coordinate transformation relationship between the two can be solved by the following formula:
[0102]
[0103] where R r can be can be can be can be
[0104] Taking the coordinates of the world coordinate system as (x, y, z), the pose of the target object's head in the depth camera coordinate system is R1, T1, and the pose in the image camera coordinate system is R2, T2. The relationship is:
[0105] R2 = R * R1
[0106] T2 = R * T1 + T
[0107] Assume that the internal parameters of the depth camera include (fx, fy, cx, cy) (where fx represents the focal length of the depth camera in the x-axis direction (horizontal direction), which helps determine the scaling ratio of the object in the horizontal direction in the image; fy represents the focal length of the depth camera in the y-axis direction (vertical direction), which affects the scaling ratio of the object in the vertical direction in the image; cx represents the offset of the optical axis of the depth camera in the x-axis direction in the depth image coordinate system, which can be used to determine the horizontal principal point coordinates of the camera image; cy represents the offset of the optical axis of the camera in the y-axis direction in the image coordinate system, which can be used to determine the vertical principal point coordinates of the camera image).
[0108] At the p point in the three-dimensional world (p = (p x , p y , p z )), the conversion relationship between it and the depth value of the depth map collected by the depth camera includes:
[0109] p x = (u - cx) * p z / fx
[0110] p y = (v - cy) / p z / fy
[0111] Where u and v are the image coordinates after the p point is imaged.
[0112] In this way, in the case of obtaining the conversion relationship between the depth camera and the in-vehicle image camera, the reference pose of the target object under the image camera can be obtained. That is, based on the internal parameters, external parameters of the camera, and camera calibration, the two-dimensional target key points can be converted into three-dimensional key points in the depth camera coordinate system.
[0113] In this way, based on the two-dimensional target key points and the corresponding three-dimensional target key points, the two-dimensional target key points and the corresponding three-dimensional target key points are matched through methods such as the least squares method and the PNP algorithm to obtain m matching pairs. Based on the m matching pairs, the pose of the two-dimensional target key points in the three-dimensional space (i.e., the direction and position of the head relative to the depth camera coordinate system) can be determined, that is, the reference pose is determined.
[0114] Please refer to Figure 7 , in some embodiments, the pose determination method further includes:
[0115] Step 014: Project the target point cloud model onto two dimensions based on the reference pose to obtain a two-dimensional image of the target point cloud model in the reference pose;
[0116] Step 015: Determine the two-dimensional feature points corresponding to the three-dimensional feature points of the target point cloud model based on the two-dimensional image;
[0117] Step 016: Update the mapping relationship based on the 3D feature points of the target point cloud model and the 2D feature points corresponding to the 3D feature points.
[0118] Among them, the target point cloud model includes a 3D deformable model (3-Dimension Morphable Model, 3DMM model).
[0119] Specifically, in the case of obtaining the reference pose, by combining the reference pose with a preset target point cloud model (for example, an average head model in the 3DMM model library, etc.), and then projecting the target point cloud model combined with the reference pose into the image camera coordinate system to simulate the 2D image of the target point cloud model obtained based on the image camera, the 2D feature points corresponding to each 3D feature point in the target point cloud model can be determined. And based on the 2D feature points corresponding to each 3D feature point of the target point cloud model, the preset mapping relationship between the 3D key points and the 2D key points is updated, which can improve the authenticity of the mapping relationship between the 3D key points and the 2D key points and improve the accuracy of the output initial pose.
[0120] Please refer to Figure 8 , in some embodiments, the pose determination method further includes:
[0121] Step 017: Determine the 2D matching degree based on the 2D target key points of the captured image and the preset mapping relationship;
[0122] Step 018: Determine the 3D matching degree based on n pairs of point clouds where the target point cloud and the source point cloud are matched. The source point cloud is determined based on the preset target point cloud model;
[0123] Step 019: Whether the 2D matching degree and the 3D matching degree meet the preset iterative convergence condition or whether the number of iterations reaches the preset number;
[0124] Step 020: If so, output the first target initial pose based on the initial poses in multiple iterative processes;
[0125] If not, re-enter the step of obtaining the captured image and the target point cloud of the target object.
[0126] Please refer to Figure 9 , optionally, Step 017: Determine the 2D matching degree based on the 2D target key points of the captured image and the preset mapping relationship, including:
[0127] Step 0171: Determine the 3D target key points corresponding to each 2D target key point based on the 2D target key points of the captured image and the preset mapping relationship;
[0128] Step 0172: Based on the perspective multi-point algorithm, match the 2D target key points and the corresponding 3D target key points to obtain m matching pairs, where each matching pair includes a 2D target key point and a corresponding 3D target key point;
[0129] Step 0173: Determine the 2D matching degree based on the m matching pairs.
[0130] Optionally, the 2D matching degree Z1 can be implemented based on the following formula:
[0131]
[0132] where, w 2d is the loss function coefficient of the PNP algorithm, w 2d = 1 / m, r represents the rotation angle, t represents the offset, and li represents the image coordinates.
[0133] Among them, the iteration convergence condition can include the convergence condition that in the 2D matching degree, the pixel distance between the image coordinates of the captured image and the image coordinates of the corresponding 2D image of the preset target point cloud model is less than the preset pixel threshold; and in the 3D matching degree, the distance between the point cloud corresponding to the transformed reference pose and the target point cloud is less than the preset distance.
[0134] Please refer to Figure 10 , optionally, Step 018: Determine the 3D matching degree based on the n pairs of point clouds obtained by matching the target point cloud and the source point cloud, including:
[0135] Step 0181: Based on the iterative closest point algorithm, match the target point cloud and the source point cloud to obtain n pairs of point clouds;
[0136] Step 0182: Determine the 3D matching degree based on the n pairs of point clouds.
[0137] Among them, the 3D matching degree Z2 can be implemented based on the following formula:
[0138]
[0139] For example, taking the determination of the initial pose based on the ICP algorithm as an example, assuming the source point cloud includes {p i}, and the target point cloud includes {q i}, where i = 1, 2, 3... n is taken as an example to illustrate, the ICP pose matching algorithm finds the rigid body transformation R, T to obtain:
[0140]
[0141] The above formula includes six parameters and has a closed-form solution, which can be quickly solved by SVD decomposition. The depth estimation principle of depth cameras has three methods: binocular stereo vision, laser ranging, and structured light ranging. The obtained point clouds are disordered, cluttered, and relatively uneven in density. There are no clear point cloud matching pairs in the disordered point clouds. It is necessary to use the Approximate Nearest Neighbors (ANN) algorithm to match the point clouds, and the solved R and T are used as the initial iterative solution for the pose, which are substituted into ANN to update the matching point pairs.
[0142] Considering the problems of uneven sampling and slow convergence speed brought by the ANN algorithm, the usual solution method is to change the loss function to minimize the distance to the tangent space, denoted as n qi as q i the normal vector at the point cloud:
[0143]
[0144] The rotation matrix R can be expressed in terms of Euler angles as:
[0145]
[0146] The rotation matrix A can be linearized. When the head of the target object rotates, it can be calculated through the following relational expression:
[0147] Ap i ≈p i +r×p i
[0148] where p i is the coordinate of the source point cloud, and r represents the rotation angle.
[0149] By linearizing the rotation matrix A, denoted as n qi as the normal vector at qi, the tangent space solution loss function of ICP is given and substituted into the Rodriguez formula as:
[0150]
[0151] Then, when determining the initial pose, it is necessary to ensure the projection alignment relationship between the three-dimensional and two-dimensional features. Therefore, the optimization function can be:
[0152]
[0153] Then solve the Jacobi matrix of this equation: The derivative of the parameter at the minimum point is 0, corresponding to a 6x6 linear equation system Ax = b. Solving the solution of the equation system gives the rotation and translation parameters, and the results are as Figures 11 to 14 shown ( Figures 11 to 14(For an example diagram of finding matching pairs based on the ANN algorithm), assume the number of matching point clouds is M, then the pose matching degree is defined by the matching degree of the matching point clouds:
[0154]
[0155] Since the solution steps of the above formula depend on the initial pose matching value, when solving under the condition that the rotation angle between different poses is small, that is, directly inputting the point clouds under two poses for solution, it is often extremely easy to fall into a local minimum.
[0156] For another example, taking the acquisition of the side face point cloud of a human face as an example, Figure 15 is the point cloud of the side face of a person under the perspective of a depth camera, Figure 16 is the point cloud of the side face of a person under the front view. Figures 17 to 20 Are, in sequence, the initial matching value of the point cloud of the side face of a person and the target point cloud model, the matching value based on the ICP algorithm, and the side face point cloud and front view point cloud of the ICP matching based on the PNP constraint.
[0157] The PNP algorithm solves the projection relationship between the marked key points of the three-dimensional model and the key points of the two-dimensional image. Denote K as the camera internal parameter and X as all the points on the object, then the camera imaging principle is:
[0158] L = K(RX + t)
[0159] Denote the x and y coordinates of the i-th point of the 2D face key points in the image as p i are the x, y, and z coordinates of the three-dimensional face key points of the corresponding 2D points marked by the model. The number of face key points is n, and the loss function to be solved is:
[0160]
[0161] When the head rotation angle (Yaw) of the face key points in the image changes and the real key points of the left or right face are not visible, at this time, solve by means of the visible key points in the image. For example, a circle point set of face key points can be set (a circle point set includes the set of all point clouds at the same height). When the face angle changes, the real face key points are shown as points on the circle in the image (that is, the pixel coordinate positions of the face key points are different under different camera poses. When the face rotation angle (yaw) is large, the face key points of the side face are not visible, and the actual correspondence of the face key points in the image is the equal-height points of the three-dimensional key points).
[0162] When iteratively solving the pose change, the 3D face key points are integrated and imaged, and the 3D face key points are updated according to the coordinate positions of the point sets. In the pose scenarios of the driver looking up, looking down, and turning the side face, there is self-occlusion of the face, and the point cloud obtained by the depth camera lacks some head point clouds. If the set pose change value is small at this time, it is still easy to fall into a local minimum during the iterative solution, resulting in the failure of pose estimation.
[0163] Therefore, optimization can be jointly carried out by introducing the captured image and the target point cloud. For example, it can include the joint optimization function of the following two-dimensional matching degree Z1 and three-dimensional matching degree Z2:
[0164]
[0165] where w 3d and w 2d are the constraint weights of the 3D point cloud and the 2D image features respectively. Since w 3d = 1 / n and w 2d = 1 / m, when n is smaller, the constraint weight of the 3D point cloud is larger, and when m is larger (for example, the face size is larger and the number of 2D key points is larger), the constraint weight of the 2D image features is smaller.
[0166] Specifically, after obtaining the two-dimensional target key points of the captured image, the corresponding three-dimensional target key points in the target point cloud model can be determined based on the mapping relationship between the two-dimensional target key points and the three-dimensional target key points. Then, based on the PNP algorithm, the two-dimensional target key points and the corresponding three-dimensional target key points are matched (for example, when the distance between the position of the two-dimensional target key point in space and the corresponding three-dimensional target key point is less than a preset threshold, it is determined that the matching is completed), so as to obtain m matching pairs that match each other. Each matching pair includes a two-dimensional target key point and a matched three-dimensional target key point. Based on the m matching pairs, the two-dimensional matching degree Z1 is determined. Then, the source point cloud is transformed to obtain n pairs of point clouds that match the target point cloud and the source point cloud, so as to determine the three-dimensional matching degree. The three-dimensional matching degree can be used to characterize the matching situation between the target point cloud and the source point cloud. Among them, the source point cloud can include the point cloud data of each point of the preset target point cloud model. According to the joint optimization of the two-dimensional matching degree Z1 and the three-dimensional matching degree Z2, even when there are missing point clouds (such as in scenarios where the driver looks up or down), the initial pose of the target object can still be calculated based on the determination of the target point cloud model and the joint optimization function of the two-dimensional matching degree Z1 and the three-dimensional matching degree Z2. When both the two-dimensional matching degree and the three-dimensional matching degree meet their respective corresponding convergence conditions, or when the number of iterations reaches the preset number of times, the initial pose in the multiple iteration processes is determined as the first target initial pose and the first target initial pose is output. Otherwise, the captured image and the target point cloud of the target object are re-obtained and re-iterated. In the iteration process, the point cloud matching and the three-dimensional and two-dimensional matching relationship joint limit loss function are used, so that the face defective point cloud can also correctly fit the pose, and the interference that may be brought by the auxiliary information is specifically removed at the end stage.
[0167] Please refer to Figure 21 , in some embodiments, step 020: Output the first target initial pose based on the initial poses in the multiple iteration processes, including:
[0168] Step 0201: When the two-dimensional matching degree and the three-dimensional matching degree meet the preset iteration convergence conditions, determine the initial pose that meets the preset iteration convergence conditions as the first target initial pose;
[0169] Step 0202: When the number of iterations reaches the preset number of times, determine the initial pose with the minimum sum of the two-dimensional matching degree and the three-dimensional matching degree as the first target initial pose.
[0170] Among them, the preset number of times can be determined according to the actual iteration situation. For example, it can be 8 times, 9 times, 10 times, 11 times, etc.
[0171] Specifically, when the number of iterations reaches the preset number, the initial pose with the minimum sum between the two-dimensional matching degree Z1 and the three-dimensional matching degree Z2 in each iteration is used as the first target initial pose to ensure the effect of the output first target initial pose.
[0172] Please refer to Figure 22 , in some embodiments, the pose determination method further includes:
[0173] Step 021: Adjust the preset target point cloud model based on the target point cloud and the initial pose.
[0174] Specifically, since there are significant differences in the shapes and sizes of the faces and heads of different target users, in order to further improve the accuracy of the output initial pose, the preset target point cloud model can be adjusted according to the target point cloud and the initial pose, so that the adjusted target point cloud model fits better with the head shape of the target user, thereby improving the accuracy of the output initial pose.
[0175] Please refer to Figure 23 , in some embodiments, Step 021: Adjust the preset target point cloud model based on the target point cloud and the initial pose, including:
[0176] Step 0211: Transform the target point cloud model based on the initial pose so that the pose of the target point cloud model is consistent with the initial pose;
[0177] Step 0212: Adjust the transformed target point cloud model based on the target point cloud so that the geometric features of the target point cloud model are consistent with the target object.
[0178] Specifically, the target point cloud model can be subjected to corresponding multiple transformations according to the obtained multiple initial poses, so that the pose of the target point cloud model is consistent with the initial pose. When the pose of the target point cloud model is consistent with the initial pose, the target point cloud model is adjusted according to the difference between the target point cloud and the point cloud data of the transformed target point cloud model, so that the geometric features of the target point cloud model are consistent with the target object.
[0179] More specifically, for example, based on the 3DMM model, about 200 precise human heads and the expression point clouds of the set expressions of each human head are scanned by a fixed-position bracket and a high-precision hardware scanning device, and then the scanned human heads are decomposed by principal components analysis (PCA) into m shape bases and n expression bases. Assuming that S is the human face shape, X is the average human face shape, is the shape difference base between people, and is the expression shape difference base of people, then there is:
[0180] = + · + ·
[0181] = + · + ·
[0182] By solving for and, a target point cloud model that is the same shape and size as the head of the target object can be determined.
[0183] It can be understood that there is a topological structure between the point clouds of the 3DMM human head three-dimensional model, and the normal vectors of the point clouds can be solved. Assuming the number of 3D key points is N and the number of 3DMM shape bases is M, it is necessary to set the face key points such that 2*N > M. Otherwise, during the shape fitting process, the fitting equations of PNP degenerate into underdetermined equations, and the constraints become ineffective.
[0184] Please refer to Figure 24 , in some embodiments, the pose determination method further includes:
[0185] Step 022: Based on the target point cloud and a preset source point cloud, determine the second target initial pose of the target object;
[0186] Step 023: Output the one with the smaller three-dimensional matching degree among the first target initial pose and the second target initial pose as the initial pose of the target object.
[0187] Among them, the preset source point cloud can be a preset point cloud set that can represent the head of the user.
[0188] Specifically, through the ICP algorithm, transformations such as rotation and translation can be performed on the preset source point cloud so that the distance between each transformed source point cloud and the target point cloud is less than a preset threshold distance required for matching, thereby determining the second target initial pose of the target object. Then, calculate the average value of the distances between the transformed source point cloud and the target point cloud in the second target initial pose (i.e., the three-dimensional matching degree), and compare it with the three-dimensional matching degree of the first target initial pose. It is considered that the smaller the three-dimensional matching degree, the higher the matching degree between the initial pose and the pose of the target object. Therefore, output the one with the smaller three-dimensional matching degree among the first target initial pose and the second target initial pose as the initial pose of the target object to obtain a driver's head pose accuracy closer to the global optimum.
[0189] Please refer to Figure 25 , in some embodiments, the pose determination method further includes:
[0190] Step 024: Based on the initial pose of the target object, perform distraction detection on the target object to determine whether the target object is distracted.
[0191] Specifically, taking the posture determination method applied to a driving assistance system as an example, assuming the target object is the driver, by obtaining the initial posture of the driver and post-processing the obtained initial posture of the target object, converting the posture matrix of the initial posture into Euler angles, and respectively setting corresponding distraction thresholds for the Euler angles in the scenarios of the driver turning their head, looking up, and looking down during driving, based on the Euler angles and each distraction threshold, it is determined whether the driver is distracted; for another example, the posture determination method can also be applied to the virtual reality field. Assuming the target object is a user interacting with a virtual object during a game process of a virtual reality system, by obtaining the initial posture of the user and performing distraction detection to determine whether the user is still using the human-computer interaction of the virtual reality system, etc.
[0192] Please refer to Figure 26 Optionally, step 024: Based on the initial posture of the target object, perform distraction detection on the target object to determine whether the target object is distracted, including:
[0193] Step 0241: Based on the initial posture of the target object, determine the pitch angle and yaw angle of the target object;
[0194] Step 0242: In the case where the pitch angle or the yaw angle is outside the corresponding angle range, determine that the target object is distracted;
[0195] Step 0243: In the case where both the pitch angle and the yaw angle are within the corresponding angle ranges, determine that the target object is not distracted.
[0196] Among them, the pitch angle can be used to represent the looking up and looking down situations of the target object, and the yaw angle can be used to represent the left and right head-turning situations of the target object.
[0197] Specifically, the initial posture of the target object can be transformed to determine the pitch angle and yaw angle of the target object, and then the pitch angle and yaw angle are respectively determined. In the case where at least one of the pitch angle and the yaw angle is outside the corresponding angle range, it is determined that the target object is distracted.
[0198] Please refer to Figure 27 Taking the posture determination method applied to the distraction module of a vehicle as an example, as Figure 27As shown, in the business application layer (the Android APP layer in the figure, referring to the specific business applications in the Android system), the camera interface (Camera API) (a set of interface classes that provide the APP with camera functions) can be called to start or stop the camera (including the image camera and the depth camera). When the camera is started, the Android Camera HAL (the interaction layer in the Android system that undertakes the operation of the camera function between the Android Framework layer and the driver layer) can send the camera-captured image data (camera image data) to the Camera API, and the Camera API then sends the camera image data to the vehicle's Driver Monitoring System (DMS). Through the DMS, the camera image data is sent to the vehicle's distraction module, and the distraction module can determine whether the target object is distracted through the pose determination algorithm and / or other distraction algorithms (such as the gaze tracking algorithm in the figure). When it is determined that the target object is distracted, a distraction signal can be issued, and the DMS then sends the distraction signal to the Camera application business module (specific camera services in the APP, such as camera photo preview, video recording, taking pictures, and the data module of the camera screen required by the APP) for subsequent processing.
[0199] Please refer to Figure 1 and Figure 28 , this application also proposes a driver pose determination method, which includes:
[0200] Step 031: Obtain the captured image of the driver and the target point cloud;
[0201] Step 032: Based on the captured image, determine the reference pose of the driver;
[0202] Step 033: Based on the reference pose and the target point cloud, determine the initial pose of the driver.
[0203] Specifically, please refer to Figure 1 , the driver pose determination method can be applied to, for example, Figure 1In the vehicle 100 shown, during the process of the driver driving the vehicle 100, first, a two-dimensional captured image of the driver's head and the target point cloud of the driver's head are obtained. Then, based on the captured image, the reference pose corresponding to the target object in the captured image is determined through a deep learning algorithm or the like. In the case of determining the reference pose of the driver and obtaining the target point cloud of the driver, the initial pose of the driver is determined. For example, the reference pose of the driver represented by the two-dimensional image and the target point cloud obtained based on the depth camera can be aligned through a registration algorithm (such as the ICP algorithm, etc.) (for example, the threshold between the point cloud data corresponding to the reference pose and the point cloud data of the target point cloud is less than a preset threshold, etc.). Then, the target point cloud with consistent poses is matched with the preset source point cloud to determine the initial pose of the driver.
[0204] It can be understood that the driver pose determination method can be considered as a method of applying the pose determination method to a vehicle to determine the driver's pose. The driver pose determination method can also include the pose determination method described in any of the above embodiments and achieve the same technical effects. For the sake of brevity, it will not be elaborated here.
[0205] In this way, by obtaining the captured image and the target point cloud of the driver, and determining the reference pose of the driver based on the captured point cloud, and finally determining the initial pose of the target object according to the reference pose and the target point cloud. Since a two-dimensional captured image of the driver is introduced when determining the initial pose, it can not only improve the accuracy of the obtained initial pose of the driver, improve the accuracy of subsequent judgment processing (for example, determining the driver's pose based on the driver pose determination method for distraction judgment and other processing), thereby improving the safety of vehicle driving, but also enable more nearest neighbor matching point pairs to be obtained when determining the driver's pose based on the target point cloud, provide a better initial pose value for subsequent solution, improve the stability during the solution of point cloud transformation, and avoid the problem of solution failure, that is, further ensure the accuracy of the obtained driver's pose.
[0206] Please refer to Figure 29 For better implementation of the pose determination method of the embodiments of the present application, the embodiments of the present application also provide a pose determination device 300. The pose determination device 300 includes a first acquisition module 301, a first determination module 302, and a second determination module 303. The first acquisition module 301 is used to acquire the captured image and the target point cloud of the target object; the first determination module 302 is used to determine the reference pose of the target object based on the captured image; the second determination module 303 is used to determine the initial pose of the target object based on the reference pose and the target point cloud.
[0207] Please refer to Figure 30, To facilitate better implementation of the driver posture determination method according to the embodiments of the present application, the embodiments of the present application further provide a driver posture determination device 400. The driver posture determination device 400 includes a second acquisition module 401, a third determination module 402, and a fourth determination module 403. The second acquisition module 401 is configured to acquire a captured image of the driver and a target point cloud; the third determination module 402 is configured to determine a reference posture of the driver based on the captured image; the fourth determination module 403 is configured to determine an initial posture of the driver based on the reference posture and the target point cloud.
[0208] In the foregoing, the posture determination device 300 and the driver posture determination device 400 have been described from the perspective of functional modules in combination with the accompanying drawings. Each functional module can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the embodiments of the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware-encoded processor, or executed and completed by a combination of the hardware and software modules in the encoded processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.
[0209] The electronic device according to the embodiments of the present application includes a processor, and the processor is connected to a memory; a computer program is stored in the memory, and the processor executes the computer program to implement the posture determination method described in any one of the above embodiments; and / or the driver posture determination method described in any one of the above embodiments, and can achieve the same technical effects. For the sake of brevity, it will not be repeated here.
[0210] Among them, the electronic device can be used as the processor of the vehicle, and the electronic device can be installed in the vehicle so that the vehicle can implement the posture determination method or the driver posture determination method described in any one of the above embodiments through the electronic device.
[0211] Please refer to again Figure 1 , The vehicle of the present application may include the posture determination device 300 and the driver posture determination device 400 described in any one of the above embodiments. For example, the posture determination device 300, the driver posture determination device 400, or the electronic device is the processor of the vehicle 100 (for example, please refer to Figure 1, the vehicle 100 may include a processor 10 and a memory 20. The attitude determination device 300, the driver attitude determination device 400, or the electronic device may serve as the processor 10 of the vehicle 100). The vehicle 100 implements the attitude determination method or the driver attitude determination method described in any of the above embodiments through the attitude determination device 300, the driver attitude determination device 400, or the electronic device. For the sake of brevity, details are not repeated here.
[0212] An embodiment of the present application also provides a computer program product, including a computer program, which includes the attitude determination method or the driver attitude determination method of any of the above embodiments. For the sake of brevity, details are not repeated here.
[0213] Please refer to Figure 30 , an embodiment of the present application also provides a computer-readable storage medium 500, on which a computer program 510 is stored. When the computer program 510 is executed by a processor 520, the steps of the attitude determination method of any of the above embodiments are implemented. For the sake of brevity, details are not repeated here.
[0214] In the description of this specification, the descriptions with reference to terms such as "certain embodiments", "in an example", "exemplarily", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiments or examples. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0215] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed. This should be understood by those skilled in the art of the embodiments of the present application.
[0216] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A posture determination method, characterized in that, Including: Obtaining a captured image and a target point cloud of a target object; Determining a reference pose of the target object based on the captured image; Determining an initial pose of the target object based on the reference pose and the target point cloud.
2. The attitude determination method according to claim 1, characterized in that The determining the reference pose of the target object based on the captured image includes: Obtaining two-dimensional target key points of the captured image; Determining three-dimensional target key points corresponding to the two-dimensional target key points based on the two-dimensional target key points and a preset mapping relationship; Determining the reference pose based on the two-dimensional target key points and the corresponding three-dimensional target key points.
3. The attitude determination method according to claim 2, wherein The determining the reference pose based on the two-dimensional target key points and the corresponding three-dimensional target key points includes: Based on a perspective multi-point algorithm, matching the two-dimensional target key points and the corresponding three-dimensional target key points to obtain m matching pairs, each matching pair including one of the two-dimensional target key points and a matching three-dimensional target key point; Determining the reference pose based on the m matching pairs.
4. The attitude determination method according to claim 2, characterized in that It further includes: Projecting the target point cloud model onto a two-dimensional plane based on the reference pose and a preset target point cloud model to obtain a two-dimensional image of the target point cloud model in the reference pose; Determining two-dimensional feature points corresponding to three-dimensional feature points of the target point cloud model based on the two-dimensional image; Updating the mapping relationship based on the three-dimensional feature points of the target point cloud model and the two-dimensional feature points corresponding to the three-dimensional feature points.
5. The attitude determination method according to any one of claims 1-4, characterized in that, It further includes: Determining a two-dimensional matching degree based on the two-dimensional target key points of the captured image and a preset mapping relationship; Determining a three-dimensional matching degree based on n pairs of point clouds obtained by matching the target point cloud and a source point cloud, the source point cloud being determined based on a preset target point cloud model; Whether the two-dimensional matching degree and the three-dimensional matching degree meet a preset iteration convergence condition or whether the number of iterations reaches a preset number; If so, outputting a first target initial pose based on the initial poses in multiple iteration processes; If not, re-entering the step of obtaining a captured image and a target point cloud of the target object.
6. The attitude determination method according to claim 5, wherein The determining the two-dimensional matching degree based on the two-dimensional target key points of the captured image and a preset mapping relationship includes: Determining three-dimensional target key points corresponding to each of the two-dimensional target key points based on the two-dimensional target key points of the captured image and a preset mapping relationship; Based on a perspective multi-point algorithm, matching the two-dimensional target key points and the corresponding three-dimensional target key points to obtain m matching pairs, each matching pair including one of the two-dimensional target key points and a matching three-dimensional target key point; Determining the two-dimensional matching degree based on the m matching pairs.
7. The attitude determination method according to claim 6, wherein The two-dimensional matching degree Z1 is implemented based on the following formula: where w 2d is 1 / m, K is the internal parameter of the camera that generates the captured image, p j is the point cloud coordinate of the j-th point of the source point cloud, r is the rotation angle, t is the offset, and l i is the image coordinate of the captured image.
8. The attitude determination method according to claim 5, wherein The determining the three-dimensional matching degree based on n pairs of point clouds obtained by matching the target point cloud and a source point cloud includes: Based on an iterative closest point algorithm, matching the target point cloud and the source point cloud to obtain the n pairs of point clouds; Determining the three-dimensional matching degree based on the n pairs of point clouds.
9. The attitude determination method according to claim 8, wherein The three-dimensional matching degree Z2 is implemented based on the following formula: Among them, w 3d = 1 / n, p i is the point cloud coordinate of the source point cloud, r is the rotation angle, t is the offset, q i is the point cloud coordinate of the target point cloud, T is the matrix transpose, n qi is the normal vector of q i .
10. The attitude determination method according to claim 5, characterized in that, Outputting a first target initial pose based on the initial pose in multiple iteration processes includes: When the two-dimensional matching degree and the three-dimensional matching degree meet the preset iteration convergence condition, determining the initial pose that meets the preset iteration convergence condition as the first target initial pose; When the number of iterations reaches the preset number, determining the initial pose with the minimum sum of the two-dimensional matching degree and the three-dimensional matching degree as the first target initial pose.
11. The attitude determination method according to any one of claims 1-10, characterized in that, It further includes: Adjusting a preset target point cloud model based on the target point cloud and the initial pose.
12. The attitude determination method according to claim 11, wherein The adjusting the preset target point cloud model based on the target point cloud and the initial pose includes: Transforming the target point cloud model based on the initial pose so that the pose of the target point cloud model is consistent with the initial pose; Adjusting the transformed target point cloud model based on the target point cloud so that the geometric features of the target point cloud model are consistent with the target object.
13. The attitude determination method according to any one of claims 4, 5, and 11, characterized in that The target point cloud model includes a three-dimensional deformable model.
14. The attitude determination method according to claim 5, wherein It further includes: Determining a second target initial pose of the target object based on the target point cloud and a preset source point cloud; Outputting the one with the smaller three-dimensional matching degree among the first target initial pose and the second target initial pose as the initial pose of the target object.
15. The attitude determination method according to claim 1 or 14, characterized in that, It further includes: Performing distraction detection on the target object based on the initial pose of the target object to determine whether the target object is distracted.
16. The attitude determination method according to claim 15, wherein The performing distraction detection on the target object based on the initial pose of the target object to determine whether the target object is distracted includes: Determining the pitch angle and yaw angle of the target object based on the initial pose of the target object; When the pitch angle or the yaw angle is outside the corresponding angle range, determining that the target object is distracted; When both the pitch angle and the yaw angle are within the corresponding angle ranges, determining that the target object is not distracted.
17. The attitude determination method according to claim 1, characterized in that, The determining the initial pose of the target object based on the reference pose and the target point cloud includes: Transforming the target point cloud based on the reference pose so that the pose of the target point cloud is consistent with the pose of a preset source point cloud; Matching the transformed target point cloud with the source point cloud to determine the initial pose.
18. A method for determining a driver's posture, characterized in that, It includes: Obtaining a captured image and a target point cloud of the driver; Determining a reference pose of the driver based on the captured image; Determining an initial pose of the driver based on the reference pose and the target point cloud.
19. An electronic device, characterized in that, It includes: A processor, the processor is connected to a memory; a computer program is stored in the memory, and the processor executes the computer program to implement the pose determination method according to any one of claims 1-17; and / or the driver pose determination method according to claim 18.
20. A vehicle, characterized in that, It includes: The electronic device according to claim 19.
21. A computer program product, characterized in that, It includes a computer program, and the computer program includes instructions for executing the pose determination method according to any one of claims 1 to 17 or the driver pose determination method according to claim 18.
22. A computer-readable storage medium, characterized in that, When the computer program is executed by a processor, the processor is caused to execute the attitude determination method according to any one of claims 1 to 17 or the driver attitude determination method according to claim 18.