Vehicle face coordinate recognition method and device, electronic equipment, medium and vehicle
By identifying the image two-dimensional coordinates of key face points and calculating the camera's external parameter matrix, the problem of monocular cameras not requiring depth information and 3D face annotation to obtain the three-dimensional coordinates of the face is solved, and the effect of providing real three-dimensional coordinates for the passenger monitoring system is achieved.
Patent Information
- Application Number
- CN202311540545.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
The prior art lacks technical solutions that do not require depth information or 3D face annotation when obtaining three-dimensional coordinates of faces based on a monocular camera.
By acquiring face images, identifying the image two-dimensional coordinates of face key points, calculating the rotation matrix and translation matrix, converting the standard three-dimensional coordinates of face key points into camera three-dimensional coordinates, realizing the acquisition of three-dimensional coordinates that do not require depth information and 3D face annotations.
It realizes the acquisition of the real three-dimensional coordinates of a face on the camera coordinate system based on a monocular camera, and provides accurate three-dimensional coordinate information for passenger monitoring systems.
Smart Images

Figure CN120020913A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicles, and particularly to a vehicle face coordinate recognition method, device, electronic device, medium and vehicle. Background Art
[0002] In a vehicle monitoring system, such as a passenger monitoring system (OMS, Occupancy Monitoring System), it is necessary to obtain the three-dimensional face pose and also detect the facial behaviors of passengers.
[0003] However, it is relatively difficult to directly obtain the three-dimensional face pose through a monocular camera. The three-dimensional face pose is represented by three-dimensional face coordinates. There are two existing methods:
[0004] 1. Obtain the three-dimensional face coordinate information through depth. This method requires depth information and needs to use a depth (Time of flight, ToF) camera for shooting, resulting in too high a cost;
[0005] 2. Through a machine learning model, for example, use a CNN model to predict the 3D key points of the face. However, this method requires 3D face annotation and a large amount of annotated data for training, consuming a large amount of learning and training resources.
[0006] Therefore, when the prior art obtains the three-dimensional face coordinates based on a monocular camera, there is a lack of a technical solution that does not require depth information or 3D face annotation. Summary of the Invention
[0007] Based on this, in view of the technical problem that the prior art lacks a technical solution that does not require depth information or 3D face annotation when obtaining the three-dimensional face coordinates based on a monocular camera, it is necessary to provide a vehicle face coordinate recognition method, device, electronic device, medium and vehicle.
[0008] The present invention provides a vehicle face coordinate recognition method, including:
[0009] Obtain a face image captured by a camera;
[0010] Identify the image two-dimensional coordinates of the face key points in the image coordinate system from the face image, where the image coordinate system is a coordinate system established on the face image;
[0011] According to the camera internal parameter matrix, calculate the rotation matrix and translation matrix of the camera external parameters that convert the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the image two-dimensional coordinates;
[0012] According to the rotation matrix and the translation matrix, convert the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system.
[0013] Further, the calculating the rotation matrix and the translation matrix of the camera external parameters for converting the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the image two-dimensional coordinates according to the camera internal parameter matrix includes:
[0014] Substitute the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space and the image two-dimensional coordinates into the formula: s*p = F*[R|t]P, where s is the scale coefficient, F is the camera internal parameter matrix, p is the image two-dimensional coordinates, P is the standard three-dimensional coordinates, R is the rotation matrix of the camera external parameters, and t is the translation matrix of the camera external parameters;
[0015] Calculate the rotation matrix and the translation matrix.
[0016] Further, the converting the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system according to the rotation matrix and the translation matrix includes:
[0017] According to the rotation matrix and the translation matrix, convert the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system.
[0018] Even further, the converting the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system according to the rotation matrix and the translation matrix includes:
[0019] Calculate the camera three-dimensional coordinates of the face key points in the camera coordinate system as: P' = R*P + t, where P' is the camera three-dimensional coordinates, R is the rotation matrix, t is the translation matrix, and P is the standard three-dimensional coordinates.
[0020] Further, before identifying the image two-dimensional coordinates of the face key points in the image coordinate system from the face image, the method further includes:
[0021] Perform classification and recognition on the face image to obtain the recognition category of the face image. When the recognition category is a non-front face angle type, set the confidence level of the recognized camera three-dimensional coordinates to an untrusted confidence level.
[0022] Even further:
[0023] Performing classification and recognition on the face image to obtain the recognition category of the face image includes: identifying a head detection frame and a face detection frame from the face image, performing classification and recognition on the image within the head detection frame to obtain the recognition category of the image within the head detection frame, and the range of the head detection frame is larger than the range of the face detection frame.
[0024] Identifying the two-dimensional image coordinates of the face key points in the image coordinate system from the face image includes: identifying the two-dimensional image coordinates of the face key points in the image coordinate system within the face detection frame.
[0025] Furthermore, the method further includes:
[0026] Judging the area where the passenger is located in the vehicle according to the camera three-dimensional coordinates; or
[0027] Obtaining the three-dimensional camera coordinates of the passenger's lips from the camera three-dimensional coordinates, and judging whether the passenger is speaking according to the three-dimensional camera coordinates of the passenger's lips; or
[0028] Obtaining the three-dimensional camera coordinates of the passenger's lips from the camera three-dimensional coordinates, and judging the position of the voice generation area according to the three-dimensional camera coordinates of the passenger's lips.
[0029] The present invention provides a vehicle face coordinate recognition device, including:
[0030] An image acquisition module, configured to acquire a face image captured by a camera;
[0031] An image two-dimensional coordinate recognition module, configured to identify the two-dimensional image coordinates of the face key points in the image coordinate system from the face image, and the image coordinate system is a coordinate system established on the face image;
[0032] A conversion calculation module, configured to calculate a rotation matrix and a translation matrix of the camera external parameters for converting the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the two-dimensional image coordinates according to the camera internal parameter matrix;
[0033] A camera three-dimensional coordinate conversion module, configured to convert the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system according to the rotation matrix and the translation matrix.
[0034] The present invention provides an electronic device, including:
[0035] At least one processor; and,
[0036] A memory communicatively connected to at least one of the processors; wherein,
[0037] The memory stores instructions executable by at least one of the processors, and the instructions are executed by at least one of the processors to enable at least one of the processors to execute the vehicle face coordinate recognition method as described above.
[0038] The present invention provides a storage medium that stores computer instructions, which are used to execute all steps of the vehicle face coordinate recognition method as described above when the computer executes the computer instructions.
[0039] The present invention provides a vehicle, including the vehicle face coordinate recognition device as described above, or the electronic device as described above.
[0040] The present invention obtains a face image through a monocular camera, recognizes the image two-dimensional coordinates of face key points in the image coordinate system from the face image, calculates the rotation matrix and translation matrix for converting the standard three-dimensional coordinates of face key points in the world coordinate system of three-dimensional space into image two-dimensional coordinates, and based on the rotation matrix and translation matrix, converts the standard three-dimensional coordinates of face key points in the world coordinate system of three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system. Since only the rotation matrix and translation matrix between the image two-dimensional coordinates and the standard three-dimensional coordinates are used to calculate the camera three-dimensional coordinates, the image two-dimensional coordinates do not need to contain depth information, nor do they require 3D face annotation. Therefore, the vehicle face coordinate recognition method of the present invention does not require depth information or 3D face annotation, and can directly obtain the camera three-dimensional coordinates of the face in the camera coordinate system based on the face image obtained by the monocular camera, providing real camera three-dimensional coordinates for the passenger monitoring system. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a working flowchart of a vehicle face coordinate recognition method according to an embodiment of the present invention;
[0042] Figure 2 It is a working flowchart of a vehicle face coordinate recognition method according to another embodiment of the present invention;
[0043] Figure 3 It is a schematic diagram of the relationship between a world coordinate system, an image, and a camera coordinate system according to an embodiment of the present invention;
[0044] Figure 4 It is a schematic diagram of an untrustworthy face three-dimensional pose according to an example of the present invention;
[0045] Figure 5 It is a schematic diagram of image recognition according to an example of the present invention;
[0046] Figure 6 It is a flowchart of a vehicle face coordinate recognition method according to the best embodiment of the present invention;
[0047] Figure 7Schematic diagram of a vehicle face coordinate recognition device according to an embodiment of the present invention;
[0048] Figure 8 Hardware structure schematic diagram of an electronic device according to the present invention;
[0049] Figure 9 Schematic diagram of three-dimensional pose information of a human face according to the present invention. Detailed implementation manners
[0050] The following further describes the detailed implementation manners of the present invention with reference to the accompanying drawings. The same components are denoted by the same reference numerals. It should be noted that the terms "front", "rear", "left", "right", "upper" and "lower" used in the following description refer to the directions in the drawings, and the terms "inner" and "outer" respectively refer to the directions towards or away from the geometric center of a specific component.
[0051] As Figure 1 shown is a flowchart of a vehicle face coordinate recognition method according to an embodiment of the present invention, including:
[0052] Step S101, obtaining a face image captured by a camera;
[0053] Step S102, identifying the image two-dimensional coordinates of face key points in the image coordinate system from the face image, where the image coordinate system is a coordinate system established on the face image;
[0054] Step S103, calculating a rotation matrix and a translation matrix of the camera external parameters for converting the standard three-dimensional coordinates of face key points in the world coordinate system in three-dimensional space into the image two-dimensional coordinates according to the camera internal parameter matrix;
[0055] Step S104, converting the standard three-dimensional coordinates of face key points in the world coordinate system in three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system according to the rotation matrix and the translation matrix.
[0056] Specifically, the present invention can be applied to an electronic device with processing capabilities, such as an electronic control unit (ECU) of a vehicle or an extended domain control unit (XCU).
[0057] The electronic device first executes step S101 to obtain a face image. Preferably, the face image is captured by a monocular camera.
[0058] Then the electronic device executes step S102 to identify the image two-dimensional coordinates of face key points in the image coordinate system from the face image, where the image coordinate system is a coordinate system established on the face image.
[0059] Specifically, first, facial key points are identified from the facial image, and then the two-dimensional image coordinates of the facial key points in the image coordinate system are determined. The image coordinate system is a coordinate system established on the facial image. Specifically, the two-dimensional image coordinates of the facial key points can be determined according to the positions of the pixels of the facial key points in the facial image.
[0060] There are multiple facial key points, and the two-dimensional image coordinates of each facial key point are determined.
[0061] Existing general key point detection algorithms in the prior art can be used to identify facial key points and determine the two-dimensional image coordinates of the facial key points in the image coordinate system.
[0062] In some embodiments, a Convolutional Neural Networks (CNN) model is used to detect facial key points, and the pixel coordinates of 68 facial key points of the face are predicted. The pixel coordinates are the two-dimensional image coordinates of the facial key points in the image coordinate system.
[0063] Then, step S103 is executed. According to the camera internal parameter matrix, the rotation matrix and the translation matrix of the camera external parameters for converting the standard three-dimensional coordinates of the facial key points in the world coordinate system in the three-dimensional space into the two-dimensional image coordinates are calculated.
[0064] Ignoring the differences in the facial shapes of different people, each facial key point has a predetermined standard three-dimensional coordinate in the world coordinate system in the three-dimensional space. Therefore, each facial key point has a standard three-dimensional coordinate in the world coordinate system and a two-dimensional image coordinate in the image coordinate system. By calculating the rotation matrix and the translation matrix for converting the standard three-dimensional coordinates of the facial key points in the world coordinate system in the three-dimensional space into the two-dimensional image coordinates according to the camera internal parameter matrix, the relevant parameters of the camera are reflected in the rotation matrix and the translation matrix.
[0065] After that, step S104 is executed. According to the rotation matrix and the translation matrix, the standard three-dimensional coordinates of the facial key points in the world coordinate system in the three-dimensional space are converted into the three-dimensional camera coordinates in the camera coordinate system.
[0066] Since the rotation matrix and the translation matrix already contain the relevant parameters of the camera, based on the rotation matrix and the translation matrix, the standard three-dimensional coordinates of the facial key points in the world coordinate system in the three-dimensional space are converted into the three-dimensional camera coordinates in the camera coordinate system, and the three-dimensional camera coordinates of the facial key points in the camera coordinate system are obtained. Since the standard three-dimensional coordinates of the facial key points in the world coordinate system in the three-dimensional space already reflect the depth information, the depth information does not need to be obtained when acquiring the image during the whole process, nor is 3D face annotation required.
[0067] The finally obtained three-dimensional camera coordinates are the three-dimensional camera coordinates of the face key points in the camera coordinate system. For example Figure 9 As shown, through the three-dimensional camera coordinates 91 of multiple face key points, the three-dimensional face pose information can be obtained, realizing the restoration of the three-dimensional face pose based on a monocular camera.
[0068] The present invention obtains a face image through a monocular camera, identifies the two-dimensional image coordinates of the face key points in the image coordinate system from the face image, calculates the rotation matrix and translation matrix for converting the standard three-dimensional coordinates of the face key points in the world coordinate system in three-dimensional space into two-dimensional image coordinates, and based on the rotation matrix and translation matrix, converts the standard three-dimensional coordinates of the face key points in the world coordinate system in three-dimensional space into the three-dimensional camera coordinates in the camera coordinate system. Since the three-dimensional camera coordinates are calculated only based on the rotation matrix and translation matrix of the two-dimensional image coordinates and the standard three-dimensional coordinates, the two-dimensional image coordinates do not need to contain depth information, nor do they require 3D face annotation. Therefore, the vehicle face coordinate recognition method of the present invention does not require depth information or 3D face annotation, can directly obtain the three-dimensional camera coordinates of the face on the camera coordinate system based on the face image obtained by the monocular camera, and provides real three-dimensional camera coordinates for the passenger monitoring system.
[0069] For example Figure 2 As shown is the workflow diagram of a vehicle face coordinate recognition method in another embodiment of the present invention, including:
[0070] Step S201, obtain a face image captured by a camera.
[0071] Step S202, perform classification and recognition on the face image to obtain the recognition category of the face image. When the recognition category is a non-front face angle type, set the confidence level for recognizing the three-dimensional camera coordinates to an untrusted confidence level.
[0072] In one embodiment, the performing classification and recognition on the face image to obtain the recognition category of the face image includes: identifying a head detection frame and a face detection frame from the face image, performing classification and recognition on the image within the head detection frame to obtain the recognition category of the image within the head detection frame, and the range of the head detection frame is larger than the range of the face detection frame.
[0073] Step S203, identify the two-dimensional image coordinates of the face key points in the image coordinate system from the face image, and the image coordinate system is a coordinate system established on the face image.
[0074] In one embodiment, identifying the image two-dimensional coordinates of the facial key points in the image coordinate system from the facial image includes: identifying the image two-dimensional coordinates of the facial key points in the image coordinate system within the facial detection box.
[0075] Step S204, according to the camera internal parameter matrix, calculate the rotation matrix and the translation matrix of the camera external parameters for converting the standard three-dimensional coordinates of the facial key points in the world coordinate system of the three-dimensional space into the image two-dimensional coordinates.
[0076] In one embodiment, the calculating the rotation matrix and the translation matrix of the camera external parameters for converting the standard three-dimensional coordinates of the facial key points in the world coordinate system of the three-dimensional space into the image two-dimensional coordinates according to the camera internal parameter matrix includes:
[0077] Substitute the standard three-dimensional coordinates of the facial key points in the world coordinate system of the three-dimensional space and the image two-dimensional coordinates into the formula: s*p = F*[R|t]P, where s is the scale coefficient, F is the camera internal parameter matrix, p is the image two-dimensional coordinates, P is the standard three-dimensional coordinates, R is the rotation matrix of the camera external parameters, and t is the translation matrix of the camera external parameters;
[0078] Calculate the rotation matrix and the translation matrix.
[0079] Step S205, according to the rotation matrix and the translation matrix, convert the standard three-dimensional coordinates of the facial key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system.
[0080] In one embodiment, the converting the standard three-dimensional coordinates of the facial key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system according to the rotation matrix and the translation matrix includes:
[0081] Calculate the camera three-dimensional coordinates of the facial key points in the camera coordinate system as: P' = R*P + t, where P' is the camera three-dimensional coordinates, R is the rotation matrix, t is the translation matrix, and P is the standard three-dimensional coordinates.
[0082] Specifically, first execute step S201 to obtain a facial image captured by a monocular camera.
[0083] Then execute step S202 to perform classification and recognition on the facial image, and when it is determined to be a non-frontal face angle category, set the confidence level of identifying the camera three-dimensional coordinates to an untrusted confidence level.
[0084] In some embodiments, the performing classification and recognition on the facial image includes:
[0085] Input a face image into the CNN layer. Extract image features through the CNN layer. After extracting the image features, enter the classification layer and output the scores for each classification category. Select the classification category with the highest score as the recognition category of the face image. The classification categories include a frontal face angle category and one or more non-frontal face angle categories.
[0086] Among them, the CNN layer can be implemented using ResNet-18. The classification layer uses a cls head layer, usually a two-layer fully connected layer. The classification categories include a frontal face angle category, a large-angle left turn category, a large-angle right turn category, a large-angle head-up category, and a large-angle head-down category. Among them, the frontal face angle category is the normal category, and the large-angle left turn category (left), the large-angle right turn category (right), the large-angle head-up category (up), and the large-angle head-down category (down) are large-angle orientation categories, that is, non-frontal face angle categories. The model finally outputs the scores for 5 categories. The final category is i = argmax(scores).
[0087] The definition of the large-angle category is usually related to the actual scenario and project, and is also related to the face key point failure condition. In some embodiments, the definition is as follows:
[0088] left, when turning the head left until more than half of the face is not visible, it is considered a large-angle left turn;
[0089] right, when turning the head right until more than half of the face is not visible, it is considered a large-angle right turn;
[0090] up, when raising the head until the eyes are not visible, it is considered a large-angle head-up;
[0091] down, when lowering the head until the mouth is not visible, it is considered a large-angle head-down.
[0092] When the recognition category is a non-frontal face angle type, set the confidence level of the recognized camera three-dimensional coordinates to an untrusted confidence level.
[0093] The untrusted confidence level is a relatively small confidence level value, for example, it can be set to 0.1.
[0094] By setting the confidence level to the untrusted confidence level, it can indicate that the subsequent obtained camera three-dimensional coordinates are untrusted. For example, for Figure 4 the recognized camera three-dimensional coordinates of the shown face image, that is, the face three-dimensional pose, are untrusted. When the recognition category is the frontal face angle category, that is, the normal category, the confidence level can be set to a relatively high confidence level value, indicating that the obtained face three-dimensional pose is trusted.
[0095] Such as Figure 6 shown is the flowchart of a vehicle face coordinate recognition method according to the best embodiment of the present invention, including:
[0096] Step S601, obtain an image;
[0097] Step S602, perform head / face detection;
[0098] Step S603, classify the images within the head detection box by large head angles and set the confidence level;
[0099] Step S604, identify the key points of the face in the images within the face detection box;
[0100] Step S605, perform 3D face pose estimation;
[0101] Step S606, output the 3D face pose and the confidence level.
[0102] In some embodiments, when the recognition category is a non-front face angle type, set the confidence level of the recognized camera 3D coordinates to an untrusted confidence level and end.
[0103] Specifically, when it is recognized that the recognition category is a non-front face angle type, set the confidence level of the recognized camera 3D coordinates to an untrusted confidence level and do not execute the subsequent steps S203 to S205.
[0104] In one of the embodiments, the classifying and recognizing the face image to obtain the recognition category of the face image includes: recognizing a head detection box and a face detection box from the face image, classifying and recognizing the images within the head detection box to obtain the recognition category of the images within the head detection box, and the range of the head detection box is larger than the range of the face detection box.
[0105] Specifically, existing image detection algorithms, such as a CNN model, can be used to perform head / face detection, and detect a head detection box 51 (red box) and a face detection box 52 (green box) as shown in Figure 5 the figure from the face image. Then, classify and recognize the images within the head detection box to obtain the recognition category of the images within the head detection box.
[0106] In this embodiment, by recognizing the head detection box and the face detection box, and using the larger-range head detection box to identify whether the head is turned, a more accurate classification and recognition can be obtained.
[0107] After that, execute step S203 to recognize the image two-dimensional coordinates of the face key points in the image coordinate system from the face image.
[0108] In one embodiment, identifying the image two-dimensional coordinates of the face key points in the image coordinate system from the face image includes: identifying the image two-dimensional coordinates of the face key points in the image coordinate system within the face detection box.
[0109] As Figure 5 shown, within the face detection box 52, identify the face key points 53 ( Figure 5 the red dots in).
[0110] In this embodiment, a smaller face detection box is used to accurately determine the face key points to obtain a more accurate three-dimensional face pose.
[0111] In some embodiments, a CNN model is used to detect face key points and predict the pixel coordinates of 68 key points of the face.
[0112] After obtaining the image two-dimensional coordinates of the face key points in the camera coordinate system, perform step S204. According to the camera intrinsic matrix, calculate the rotation matrix and translation matrix of the camera extrinsic parameters that convert the standard three-dimensional coordinates of the face key points in the world coordinate system in three-dimensional space into the image two-dimensional coordinates.
[0113] As Figure 3 shown, use the principle of pinhole imaging and 3D perspective transformation for face three-dimensional pose estimation. A 3D coordinate point P in the world coordinate system 31 is projected onto the image coordinate system, and the image coordinate system is the 2D pixel coordinate system on the image 32. The coordinate point P is converted into the coordinate point p through the camera intrinsic matrix F and the rotation matrix R and translation matrix t of the camera extrinsic parameters. Among them, the camera parameters of the camera 34 include the camera intrinsics and the camera extrinsics. The camera intrinsics is the camera intrinsic matrix F, and the camera intrinsic matrix F describes the conversion relationship between the image coordinate system and the camera coordinate system 33. The camera extrinsics includes the rotation matrix R and the translation matrix t. The rotation matrix R and the translation matrix t represent the conversion relationship between the world coordinate system 31 and the camera coordinate system 33. Therefore, a conversion formula for the three-dimensional coordinates in the world coordinate system 31 and the two-dimensional coordinates in the image coordinate system is established through the camera intrinsic matrix, the rotation matrix, and the translation matrix. Since the standard three-dimensional coordinates of the face key points in the world coordinate system in three-dimensional space and the image two-dimensional coordinates of the face key points in the image coordinate system are known, and the camera intrinsic matrix can be obtained through camera calibration, the rotation matrix R and the translation matrix t representing the conversion relationship between the world coordinate system 31 and the camera coordinate system 33 can be solved.
[0114] In one embodiment, the calculating the rotation matrix and translation matrix of the camera extrinsic parameters that convert the standard three-dimensional coordinates of the face key points in the world coordinate system in three-dimensional space into the image two-dimensional coordinates according to the camera intrinsic matrix includes:
[0115] Substitute the standard 3D coordinates of the facial key points in the world coordinate system of the three-dimensional space and the two-dimensional image coordinates into the formula: s*p = F*[R|t]P, where s is the scale factor, F is the camera internal parameter matrix, p is the two-dimensional image coordinates, P is the standard 3D coordinates, R is the rotation matrix of the camera external parameters, and t is the translation matrix of the camera external parameters;
[0116] Calculate the rotation matrix and the translation matrix.
[0117] Specifically, calculate the rotation matrix and the translation matrix according to formula (1):
[0118] s*p = F*[R|t]P (1)
[0119] where s is the scale factor, F is the camera internal parameter matrix, p is the two-dimensional image coordinates, P is the standard 3D coordinates, R is the rotation matrix of the camera external parameters, and t is the translation matrix of the camera external parameters.
[0120] The specific formula expansion is as follows:
[0121]
[0122] where (u, v) are the two-dimensional image coordinates p, is the camera internal parameter matrix F, f x represents the pixel focal length of the camera in the x direction, f y represents the pixel focal length of the camera in the y direction, c x represents the pixel offset of the camera optical axis in the x direction in the image coordinate system, c y represents the pixel offset of the camera optical axis in the y direction in the image coordinate system, is the rotation matrix R, is the translation matrix t, (X, Y, Z) are the 3D coordinates P.
[0123] As Figure 5 shown, the world coordinate system 31 is at the center position of the human head. Therefore, the standard 3D coordinates of the facial key points can be directly used (ignoring the facial shape differences of different people). The internal parameter matrix of the camera can be obtained through camera calibration. Therefore, according to the standard 3D coordinates and the corresponding two-dimensional image coordinates of several facial key points, the rotation matrix R and the translation matrix t can be solved, representing the conversion relationship between the camera coordinate system and the world coordinate system. The solution method is not unique, and a feasible method is the pnp algorithm. The solvePnP function in the OpenCV library can be used to solve it.
[0124] In this embodiment, the rotation matrix and the translation matrix are solved based on the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space and the corresponding two-dimensional image coordinates, and the conversion relationship between the camera coordinate system and the world coordinate system is obtained.
[0125] Then, step S205 is executed. According to the rotation matrix and the translation matrix, the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space are converted into the camera three-dimensional coordinates in the camera coordinate system.
[0126] Since the rotation matrix and the translation matrix represent the relative relationship between the camera coordinate system and the world coordinate system, the standard three-dimensional coordinates of the face key points in the world coordinate system are converted into the camera three-dimensional coordinates in the camera coordinate system through the rotation matrix and the translation matrix, and the three-dimensional coordinates are the camera three-dimensional coordinates.
[0127] In one embodiment, the converting the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system according to the rotation matrix and the translation matrix includes:
[0128] Calculating the camera three-dimensional coordinates of the face key points in the camera coordinate system as: P′ = R * P + t, where P′ is the camera three-dimensional coordinates, R is the rotation matrix, t is the translation matrix, and P is the standard three-dimensional coordinates.
[0129] Specifically, for each standard three-dimensional coordinate, calculate the camera three-dimensional coordinates of the face key points in the camera coordinate system as P′ = R * P + t, where P′ is the camera three-dimensional coordinates, R is the rotation matrix, t is the translation matrix, and P is the standard three-dimensional coordinates. The rotation matrix R is: The translation matrix t is:
[0130] Calculate the camera three-dimensional coordinates of N face key points as the three-dimensional face pose information, where N is the number of face key points.
[0131] Based on the camera intrinsic matrix, the present invention calculates the rotation matrix and the translation matrix of the camera extrinsic parameters for converting the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the two-dimensional image coordinates. Since the rotation matrix and the translation matrix represent the conversion relationship between the camera coordinate system and the world coordinate system, the standard three-dimensional coordinates of the face key points in the world coordinate system can be quickly and accurately converted into the camera three-dimensional coordinates in the camera coordinate system.
[0132] In one embodiment, the method further includes:
[0133] Judging the area where the passenger is located in the vehicle according to the camera three-dimensional coordinates; or
[0134] Obtain the three-dimensional camera coordinates of the passenger's lips from the three-dimensional camera coordinates of the camera, and determine whether the passenger is speaking according to the three-dimensional camera coordinates of the passenger's lips; or
[0135] Obtain the three-dimensional camera coordinates of the passenger's lips from the three-dimensional camera coordinates of the camera, and determine the position of the voice generation area according to the three-dimensional camera coordinates of the passenger's lips.
[0136] Specifically, the vehicle face coordinate recognition method of this embodiment can be applied to solutions such as passenger occupancy, multi-modal rejection recognition, and multi-modal human voice positioning.
[0137] In passenger occupancy, the three-dimensional camera coordinate information of the face key points can be used to determine which area of the vehicle the passenger is in, so as to determine which seat the passenger is sitting on (driver's seat, co-driver's seat, left rear row, right rear row, etc.).
[0138] In multi-modal rejection recognition, the three-dimensional coordinate information of the passenger's lips can be extracted to determine whether the speaker is really speaking, thereby avoiding the mis-awakening of the auxiliary voice.
[0139] In multi-modal human voice positioning, by combining the three-dimensional lip coordinate information and the voice information, it is determined whether the three-dimensional lip coordinate information is consistent with the voice information, so as to assist in judging the sound area position and help the voice perform human voice positioning.
[0140] The face three-dimensional pose estimation method based on a monocular camera in this embodiment can be used in the OMS system, and the real three-dimensional coordinates of the face can be obtained, providing three-dimensional coordinate information for functions such as passenger occupancy and multi-modal dialogue.
[0141] Based on the same inventive concept, as Figure 7 shown in the schematic diagram of a vehicle face coordinate recognition device according to an embodiment of the present invention, including:
[0142] An image acquisition module 701, configured to acquire a face image captured by a camera;
[0143] An image two-dimensional coordinate recognition module 702, configured to recognize the image two-dimensional coordinates of face key points in the image coordinate system from the face image, where the image coordinate system is a coordinate system established on the face image;
[0144] A conversion calculation module 703, configured to calculate a rotation matrix and a translation matrix of an external camera parameter for converting the standard three-dimensional coordinates of face key points in the world coordinate system in three-dimensional space into the image two-dimensional coordinates according to the camera internal parameter matrix;
[0145] A camera three-dimensional coordinate conversion module 704, configured to convert the standard three-dimensional coordinates of face key points in the world coordinate system in three-dimensional space into camera three-dimensional coordinates in the camera coordinate system according to the rotation matrix and the translation matrix.
[0146] The present invention obtains a face image through a monocular camera, identifies the image two-dimensional coordinates of face key points in the image coordinate system from the face image, calculates the rotation matrix and translation matrix for converting the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the image two-dimensional coordinates, and based on the rotation matrix and translation matrix, converts the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system. Since the camera three-dimensional coordinates are calculated only based on the rotation matrix and translation matrix of the image two-dimensional coordinates and the standard three-dimensional coordinates, the image two-dimensional coordinates do not need to contain depth information, nor do they require 3D face annotation. Therefore, the vehicle face coordinate recognition method of the present invention does not require depth information or 3D face annotation, and can directly obtain the camera three-dimensional coordinates of the face on the camera coordinate system based on the face image obtained by the monocular camera, providing real camera three-dimensional coordinates for the passenger monitoring system.
[0147] In one embodiment, the calculating the rotation matrix and translation matrix of the camera external parameters for converting the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the image two-dimensional coordinates according to the camera internal parameter matrix includes:
[0148] Substituting the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space and the image two-dimensional coordinates into the formula: s*p = F*[R|t]P, where s is the scale factor, F is the camera internal parameter matrix, p is the image two-dimensional coordinates, P is the standard three-dimensional coordinates, R is the rotation matrix of the camera external parameters, and t is the translation matrix of the camera external parameters;
[0149] Calculating the rotation matrix and the translation matrix.
[0150] In one embodiment, the converting the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system according to the rotation matrix and the translation matrix includes:
[0151] Converting the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system according to the rotation matrix and the translation matrix.
[0152] In one embodiment, the converting the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system according to the rotation matrix and the translation matrix includes:
[0153] Calculating the camera three-dimensional coordinates of the face key points in the camera coordinate system as: P' = R*P + t, where P' is the camera three-dimensional coordinates, R is the rotation matrix, t is the translation matrix, and P is the standard three-dimensional coordinates.
[0154] In one embodiment, the device further includes a type recognition module, configured to:
[0155] Classify and recognize the face image to obtain the recognition category of the face image, and when the recognition category is a non-front face angle type, set the confidence level of the recognized camera three-dimensional coordinates to an untrusted confidence level.
[0156] In one embodiment:
[0157] The classifying and recognizing the face image to obtain the recognition category of the face image includes: recognizing a head detection frame and a face detection frame from the face image, classifying and recognizing the image within the head detection frame to obtain the recognition category of the image within the head detection frame, and the range of the head detection frame is larger than the range of the face detection frame;
[0158] The recognizing the image two-dimensional coordinates of the face key points in the image coordinate system from the face image includes: recognizing the image two-dimensional coordinates of the face key points in the image coordinate system within the face detection frame.
[0159] In one embodiment, the device further includes an application module, configured to:
[0160] Judge the area where the passenger is located in the vehicle according to the camera three-dimensional coordinates; or
[0161] Obtain the camera three-dimensional coordinates of the passenger's lips from the camera three-dimensional coordinates, and judge whether the passenger is speaking according to the camera three-dimensional coordinates of the passenger's lips; or
[0162] Obtain the camera three-dimensional coordinates of the passenger's lips from the camera three-dimensional coordinates, and judge the position of the voice generation area according to the camera three-dimensional coordinates of the passenger's lips.
[0163] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0164] As Figure 8 shown is a schematic hardware structure diagram of an electronic device according to the present invention, including:
[0165] At least one processor 801; and,
[0166] A memory 802 communicatively connected to at least one of the processors 801; wherein,
[0167] The memory 802 stores instructions executable by at least one of the processors, and the instructions are executed by at least one of the processors so that at least one of the processors can execute the vehicle face coordinate recognition method as described above.
[0168] Figure 8 Take a processor 801 as an example.
[0169] The electronic device may further include: an input device 803 and a display device 804.
[0170] The processor 801, the memory 802, the input device 803, and the display device 804 may be connected through a bus or other means. In the figure, the connection through a bus is taken as an example.
[0171] As a non-volatile computer-readable storage medium, the memory 802 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the vehicle face coordinate recognition method in the embodiments of the present application. For example, Figure 1 、 Figure 2 The method flow shown. By running the non-volatile software programs, instructions, and modules stored in the memory 802, the processor 801 executes various functional applications and data processing, that is, implements the vehicle face coordinate recognition method in the above embodiments.
[0172] The memory 802 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the vehicle face coordinate recognition method, etc. In addition, the memory 802 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 802 may optionally include a memory remotely set relative to the processor 801, and these remote memories can be connected to the device executing the vehicle face coordinate recognition method through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.
[0173] The input device 803 can receive the input user clicks and generate signal inputs related to the user settings and function controls of the vehicle face coordinate recognition method. The display device 804 may include a display screen and other display devices.
[0174] When the one or more modules are stored in the memory 802 and run by the one or more processors 801, the vehicle face coordinate recognition method in any of the above method embodiments is executed.
[0175] The present invention obtains a face image through a monocular camera, identifies the image two-dimensional coordinates of face key points in the image coordinate system from the face image, calculates the rotation matrix and translation matrix for converting the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the image two-dimensional coordinates, and based on the rotation matrix and translation matrix, converts the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system. Since the camera three-dimensional coordinates are calculated only based on the rotation matrix and translation matrix between the image two-dimensional coordinates and the standard three-dimensional coordinates, the image two-dimensional coordinates do not need to contain depth information, nor do they require 3D face annotation. Therefore, the vehicle face coordinate recognition method of the present invention does not require depth information or 3D face annotation, and can directly obtain the camera three-dimensional coordinates of the face in the camera coordinate system based on the face image obtained by the monocular camera, providing real camera three-dimensional coordinates for the passenger monitoring system.
[0176] An embodiment of the present invention provides a storage medium that stores computer instructions, which are used to execute all steps of the vehicle face coordinate recognition method as described above when the computer executes the computer instructions.
[0177] In the context of the present disclosure, the storage medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The storage medium may be a machine-readable signal medium or a machine-readable storage medium. Optionally, the storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be ROM, random access memory (Random Access Memory, RAM), compact disc read-only memory (Compact Disc ROM, CD-ROM), magnetic tape, floppy disk, and optical data storage devices, etc.
[0178] An embodiment of the present invention provides a vehicle, including the vehicle face coordinate recognition device as described above, or the electronic device as described above. It can be understood that the vehicle may also include: a processor, a memory, and a computer program. Among them, the computer program is stored in the memory and is configured to be executed by the processor to implement the vehicle face coordinate recognition method provided by the embodiments of the present disclosure. Among them, the processor and the memory have been described in the Figure 7 illustrated embodiments and will not be elaborated here.
[0179] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.
Claims
1. A vehicle face coordinate recognition method, characterized in that: include: Get the face image captured by the camera; Identifying two-dimensional image coordinates of facial key points in an image coordinate system from the facial image, wherein the image coordinate system is a coordinate system established on the facial image; According to the camera intrinsic parameter matrix, a rotation matrix and a translation matrix of the camera extrinsic parameters are calculated to convert the standard three-dimensional coordinates of the facial key points in the world coordinate system of the three-dimensional space into the two-dimensional coordinates of the image; According to the rotation matrix and the translation matrix, the standard three-dimensional coordinates of the facial key points in the world coordinate system of the three-dimensional space are converted into the camera three-dimensional coordinates in the camera coordinate system.
2. The vehicle face coordinate recognition method according to claim 1, characterized in that: The step of calculating the rotation matrix and translation matrix of the camera extrinsic parameters that convert the standard three-dimensional coordinates of the facial key points in the world coordinate system of the three-dimensional space into the two-dimensional coordinates of the image according to the camera intrinsic parameter matrix includes: Substitute the standard three-dimensional coordinates of the facial key points in the world coordinate system of the three-dimensional space and the two-dimensional coordinates of the image into the formula: s*p=F*[R|t]P, where s is the scale factor, F is the camera intrinsic parameter matrix, p is the two-dimensional coordinates of the image, P is the standard three-dimensional coordinates, R is the rotation matrix of the camera extrinsic parameters, and t is the translation matrix of the camera extrinsic parameters; The rotation matrix and the translation matrix are calculated.
3. The vehicle face coordinate recognition method according to claim 1, characterized in that: The step of converting the standard three-dimensional coordinates of the facial key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system according to the rotation matrix and the translation matrix comprises: According to the rotation matrix and the translation matrix, the standard three-dimensional coordinates of the facial key points in the world coordinate system of the three-dimensional space are converted into the camera three-dimensional coordinates in the camera coordinate system.
4. The vehicle face coordinate recognition method according to claim 3, characterized in that: The step of converting the standard three-dimensional coordinates of the facial key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system according to the rotation matrix and the translation matrix comprises: The camera three-dimensional coordinates of the facial key points in the camera coordinate system are calculated as: P′=R*P+t, where P′ is the camera three-dimensional coordinates, R is the rotation matrix, t is the translation matrix, and P is the standard three-dimensional coordinates.
5. The vehicle face coordinate recognition method according to claim 1, characterized in that: Before identifying the two-dimensional image coordinates of facial key points in the image coordinate system from the facial image, the method further includes: The face image is classified and identified to obtain an identification category of the face image. When the identification category is a non-frontal face angle type, the confidence level of the identification camera three-dimensional coordinates is set to an untrustworthy confidence level.
6. The vehicle face coordinate recognition method according to claim 5, characterized in that: The classifying and identifying the face image to obtain the recognition category of the face image includes: identifying a head detection frame and a face detection frame from the face image, classifying and identifying the image in the head detection frame to obtain the recognition category of the image in the head detection frame, wherein the range of the head detection frame is larger than the range of the face detection frame; The step of identifying the two-dimensional image coordinates of the key points of the face in the image coordinate system from the face image includes: identifying the two-dimensional image coordinates of the key points of the face in the image coordinate system within the face detection frame.
7. The vehicle face coordinate recognition method according to any one of claims 1 to 6, characterized in that: The method further comprises: Determine the area where the passenger is located in the vehicle according to the three-dimensional coordinates of the camera; or Obtaining the three-dimensional coordinates of the passenger's lip camera from the three-dimensional coordinates of the camera, and judging whether the passenger is speaking according to the three-dimensional coordinates of the passenger's lip camera; or The three-dimensional coordinates of the passenger's lip camera are obtained from the three-dimensional coordinates of the camera, and the position of the speech generation area is determined according to the three-dimensional coordinates of the passenger's lip camera.
8. A vehicle face coordinate recognition device, characterized in that: include: An image acquisition module is used to acquire a face image taken by a camera; An image two-dimensional coordinate recognition module, used to recognize the image two-dimensional coordinates of facial key points in an image coordinate system from the facial image, wherein the image coordinate system is a coordinate system established on the facial image; A conversion calculation module, used to calculate a rotation matrix and a translation matrix of camera extrinsic parameters that convert the standard three-dimensional coordinates of the key points of the face in the world coordinate system of the three-dimensional space into the two-dimensional coordinates of the image according to the camera intrinsic parameter matrix; The camera three-dimensional coordinate conversion module is used to convert the standard three-dimensional coordinates of the face key points in the world coordinate system of the three-dimensional space into the camera three-dimensional coordinates in the camera coordinate system according to the rotation matrix and the translation matrix.
9. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by at least one of the processors, and the instructions are executed by at least one of the processors so that at least one of the processors can execute the vehicle face coordinate recognition method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium stores computer instructions, and when a computer executes the computer instructions, it is used to execute all steps of the vehicle face coordinate recognition method as described in any one of claims 1 to 7.
11. A vehicle, characterized in that: It includes the vehicle face coordinate recognition device as described in claim 8, or the electronic device as described in claim 9.
Citation Information
Cited By
Hovering tracking method, hovering tracking device, hovering tracking equipment and storage medium
CN116012910A
Vehicle-mounted fraud identification method and device based on face detection, equipment and medium
CN121392923A