A method, device, equipment and storage medium for obtaining the true value of the head pose
Through multiple image acquisition devices, a coordinate system is established to obtain the true value of the head posture, the problem of wearing equipment affecting accuracy and high-precision data analysis is achieved.
Patent Information
- Application Number
- CN202111501742.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-12-09
AI Technical Summary
The wearable angle of the wearable device can easily affect the accuracy of the true value of the head posture, resulting in insufficient accuracy of the data analysis results.
Through multiple image acquisition devices, the image of the target object is collected simultaneously, the key points of the face are marked, the three-dimensional key points of the face are reconstructed, the face coordinate system is established, the true value of the head posture is obtained, and the wearable device is avoided.
It improves the accuracy of the true value of the head posture and ensures the accuracy of the data analysis results. It is suitable for driver fatigue analysis, virtual reality somatosensory games, product purchase analysis and face verification.
Smart Images

Figure CN114220149B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, device, equipment and storage medium for obtaining a true value of a head posture. Background Art
[0002] The true value of head posture includes yaw angle, pitch angle and roll angle. The yaw angle, pitch angle and roll angle are the angles of rotation relative to the y-axis, x-axis and z-axis in the Euler angle vector coordinate system. The true value of head posture is used in many fields, such as driver fatigue analysis, virtual reality somatosensory games, commodity purchase desire analysis, face verification and other fields.
[0003] Currently, the true value of head posture can be obtained through sensors in wearable devices. However, the wearing angle of the wearable device can easily affect the accuracy of the obtained true value of head posture. Combined with the many application fields mentioned above, if the accuracy of the true value of head posture is not good, it can easily affect the accuracy of the data analysis results. For example, due to the poor accuracy of the obtained true value of the driver's head posture, a driver with a high degree of fatigue is mistakenly judged as not fatigued, and then no timely and accurate voice reminder is made. It can be seen that improving the accuracy of obtaining the true value of head posture is a technical problem that needs to be solved urgently. Summary of the invention
[0004] Based on the above problems, the present application provides a method, device, equipment and storage medium for obtaining the true value of head posture.
[0005] The embodiments of the present application disclose the following technical solutions:
[0006] In a first aspect, the present application provides a method for obtaining a true value of a head posture, comprising:
[0007] Acquire images of a target object captured by multiple image acquisition devices at the same time;
[0008] Marking facial key points in the images captured by each of the multiple image capture devices to obtain facial key point information corresponding to each of the image capture devices;
[0009] Reconstructing the three-dimensional key point information of the face corresponding to each of the image acquisition devices based on the key point information of the face corresponding to each of the image acquisition devices and the parameters of each of the image acquisition devices;
[0010] Establishing a facial coordinate system based on the three-dimensional key point information of the face corresponding to the target image acquisition device, wherein the target image acquisition device is one of the multiple image acquisition devices;
[0011] According to the coordinate system of the target image acquisition device and the face coordinate system, obtain the ground truth of the head pose of the target object corresponding to the target image, where the target image is an image of the target object acquired by the target image acquisition device at the moment.
[0012] In an alternative implementation, the obtaining the ground truth of the head pose of the target object corresponding to the target image according to the coordinate system of the target image acquisition device and the face coordinate system includes:
[0013] According to the face coordinate system and the coordinate system of the target image acquisition device, obtain the rotation matrix of the face coordinate system relative to the coordinate system of the target image acquisition device;
[0014] Obtain the ground truth of the head pose of the target object corresponding to the target image according to the rotation matrix.
[0015] In an alternative implementation, the establishing the face coordinate system based on the three-dimensional face key point information corresponding to the target image acquisition device includes:
[0016] Determine the face plane based on the three-dimensional face key point information corresponding to the target image acquisition device;
[0017] Determine the normal vector of the face plane according to the face plane;
[0018] Establish the face coordinate system of the target object based on the face plane and the normal vector of the face plane.
[0019] In an alternative implementation, the reconstructing the three-dimensional face key point information corresponding to each image acquisition device based on the face key point information corresponding to each image acquisition device and the parameters of each image acquisition device includes:
[0020] According to the face key point information corresponding to the target image acquisition device, the face key point information corresponding to the reference image acquisition device, the internal parameters of the target image acquisition device, the internal parameters of the reference image acquisition device, and the external parameters between the target image acquisition device and the reference image acquisition device, reconstruct the three-dimensional face key point information corresponding to the target image acquisition device by the triangulation reconstruction method; the reference image acquisition device is any one of the multiple image acquisition devices other than the target image acquisition device;
[0021] Obtain the three-dimensional face key point information corresponding to the reference image acquisition device according to the three-dimensional face key point information corresponding to the target image acquisition device and the external parameters.
[0022] In an alternative implementation, the method for obtaining the ground truth of the head pose further includes:
[0023] Storing the mutually corresponding images and the ground truth of the head pose as an image-ground truth pair.
[0024] In an alternative implementation, obtaining the images of the target object collected by multiple image acquisition devices at the same time includes:
[0025] Providing various different lighting conditions to the space where the target object is located;
[0026] Obtaining the images of the target object collected by the multiple image acquisition devices at the same time under the same lighting condition.
[0027] In an alternative implementation, when the multiple image acquisition devices collect the images of the target object, the target object is sitting on a seat, and the seat is used to simulate the seat in the real vehicle environment; the setting parameters of the image acquisition device relative to the seat are determined based on the setting parameters of the simulated object relative to the seat in the real vehicle environment;
[0028] The simulated object includes at least one of the following:
[0029] A-pillar, B-pillar, instrument panel, front windshield or left side window glass.
[0030] In a second aspect, the present application provides a device for obtaining the ground truth of the head pose, including:
[0031] An image acquisition module, configured to obtain images of a target object collected by multiple image acquisition devices at the same time;
[0032] A key point annotation module, configured to annotate the face key points in the images collected by each of the multiple image acquisition devices to obtain the face key point information corresponding to each of the image acquisition devices;
[0033] A three-dimensional key point reconstruction module, configured to reconstruct the face three-dimensional key point information corresponding to each of the image acquisition devices based on the face key point information corresponding to each of the image acquisition devices and the parameters of each of the image acquisition devices;
[0034] A coordinate system establishment module, configured to establish a face coordinate system based on the face three-dimensional key point information corresponding to the target image acquisition device, where the target image acquisition device is one of the multiple image acquisition devices;
[0035] A true value acquisition module, configured to obtain the true value of the head pose of the target object corresponding to the target image according to the coordinate system of the target image acquisition device and the face coordinate system, where the target image is an image of the target object acquired by the target image acquisition device at the moment.
[0036] In a third aspect, the present application provides a device for acquiring the true value of the head pose, including a processor and a memory; the memory is used to store a computer program; the processor is configured to execute the method for acquiring the true value of the head pose provided in the first aspect according to the computer program.
[0037] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, where the computer program, when run by a processor, executes the method for acquiring the true value of the head pose provided in the first aspect.
[0038] Compared with the prior art, the present application has the following beneficial effects:
[0039] The present application provides a method for acquiring the true value of the head pose. Images of a target object at the same moment are acquired by multiple image acquisition devices, and the two-dimensional face key points marked in the multiple acquired images are used as the data basis for three-dimensional reconstruction of the face key points of the target object, and then the true value of the head pose of the target object corresponding to the image is acquired. Since this solution does not need to rely on a wearable device to acquire the true value of the head pose, it is not affected by the wearing angle. By using multiple image acquisition devices at the same moment and three-dimensional reconstruction of face key points, the accuracy of the acquired true value of the head pose is guaranteed. Furthermore, when the true value of the head pose is applied to fields such as driver fatigue degree analysis, virtual reality somatosensory games, commodity purchase desire analysis, face verification, etc., the analysis results of the data can be made more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 It is a flowchart of a method for acquiring the true value of the head pose provided by the embodiment of the present application;
[0042] Figure 2 It is a schematic diagram of a scenario for acquiring images of a target object by multiple image acquisition devices provided by the embodiment of the present application;
[0043] Figure 3Schematic structural diagram of an apparatus for obtaining true values of head postures provided by an embodiment of the present application. Detailed implementation manners
[0044] As described above, currently, obtaining true values of head postures usually relies on sensors of wearable devices. However, the angle at which a person wears the device may affect the accuracy of the obtained true values of head postures, resulting in an impact on the accuracy of the analysis results when using the true values of head postures for data analysis. After research, the inventor proposes a technical solution for simultaneously collecting images of a person by multiple image acquisition devices and obtaining true values of head postures from these images. This solution does not require a person to wear a wearable device and can obtain high-precision true values of head postures only by virtue of images. Furthermore, the accuracy of the data analysis results obtained when using the true values of head postures for data analysis is ensured.
[0045] To enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0046] Refer to Figure 1 , which is a flowchart of a method for obtaining true values of head postures provided by an embodiment of the present application. As Figure 1 shown, the method includes the following steps:
[0047] S101: Obtain images of a target object collected by multiple image acquisition devices at the same time.
[0048] In the embodiment of the present application, it is proposed to collect images of a target object at the same time by multiple image acquisition devices. The target object refers to the subject for which true values of head postures need to be obtained. For example, when it is necessary to obtain the true values of the head postures of Mr. A, Mr. A is used as the target object, and multiple image acquisition devices are used to collect images of Mr. A at the same time.
[0049] The image acquisition device can be any device capable of taking and forming images, such as an RGB camera. The type and model of the image acquisition device are not limited herein. The multiple image acquisition devices are specifically two or more image acquisition devices. That is, when it is required to collect images of the target object at the same time, at least two image acquisition devices are used. The setting parameters of the multiple image acquisition devices are different or not completely the same. The setting parameters include position, height, angle, etc.
[0050] S102: Label the facial key points in the images collected by each of the multiple image acquisition devices to obtain the facial key point information corresponding to each image acquisition device.
[0051] In this step, facial key points are labeled for the images collected by each image acquisition device. The labeled facial key points may include, but are not limited to: the eyebrows, the peaks of the eyebrows, the ends of the eyebrows, the inner corners of the eyes, the outer corners of the eyes, the centers of the pupils, the wings of the nose, the nostrils, the corners of the mouth, etc. As an example, 70 key points of the face in the image are labeled. The number of key points labeled here is not limited.
[0052] Identifying and marking facial key points in an image containing a face belongs to a relatively mature technology in this field, so the specific implementation method of this step is not limited here. The facial key point information may include: the pixel coordinates of the key points in the image where they are located. Since the image is a two-dimensional image, the labeled facial key point information also refers to the information of facial key points in the two-dimensional image.
[0053] S103: Based on the facial key point information corresponding to each image acquisition device and the parameters of each image acquisition device, reconstruct the three-dimensional facial key point information corresponding to each image acquisition device.
[0054] The purpose of this step is to construct the information of three-dimensional facial key points by marking two-dimensional facial key points in multiple images collected by different image acquisition devices at the same time. In specific implementation, for each image, a set of corresponding three-dimensional facial key point information needs to be obtained. Since the image and the image acquisition device have a correspondence, it can be understood that the image acquisition device and the obtained three-dimensional facial key point information have a correspondence. Taking a certain acquisition moment as an example, in this step, the three-dimensional facial key point information corresponding to each image acquisition device needs to be reconstructed. For example, there are 14 image acquisition devices. For the same acquisition moment, the three-dimensional facial key point information corresponding to each of the 14 image acquisition devices needs to be obtained respectively.
[0055] Among the 14 image acquisition devices, different image acquisition devices have unique labels, such as device No. 1, device No. 2,... device No. 14. The order of the labels is not limited here. For example, the labels can be made in the way of increasing the ordinal number sequentially along a direction.
[0056] The following introduces an exemplary implementation method of S103:
[0057] According to the facial key point information corresponding to the target image acquisition device, the facial key point information corresponding to the reference image acquisition device, the internal parameters of the target image acquisition device, the internal parameters of the reference image acquisition device, and the external parameters between the target image acquisition device and the reference image acquisition device, the three-dimensional facial key point information corresponding to the target image acquisition device is reconstructed by the triangulation reconstruction method. The reference image acquisition device is any image acquisition device other than the target image acquisition device among the multiple image acquisition devices. According to the three-dimensional facial key point information corresponding to the target image acquisition device and the external parameters, the three-dimensional facial key point information corresponding to the reference image acquisition device is obtained. The three-dimensional facial key point information may include the three-dimensional coordinates of the facial key points in the coordinate system of the image acquisition device where they are located.
[0058] Assume that device No. 4 is the target image acquisition device and device No. 5 is the reference image acquisition device. When performing this step, the facial key point information corresponding to device No. 4, the facial key point information corresponding to device No. 5, the internal parameters of device No. 4, the internal parameters of device No. 5, and the external parameters between device No. 4 and device No. 5 are used to reconstruct the information of the three-dimensional facial key points corresponding to device No. 4 by the triangulation reconstruction method. When triangulating and reconstructing the three-dimensional facial key point information, it can be realized through a triangulation reconstruction function. Triangulation reconstruction is a relatively mature technology in this field, so the specific implementation process is not elaborated here. After obtaining the three-dimensional facial key point information corresponding to device No. 4, since the external parameters between different image acquisition devices have been pre-calibrated, therefore, with the help of the external parameters between other image acquisition devices and device No. 4, the three-dimensional facial key point information corresponding to device No. 4 can be converted into the coordinate system of other image acquisition devices to obtain the three-dimensional facial key point information corresponding to other image acquisition devices. In this way, the three-dimensional facial key point information corresponding to each of the 14 image acquisition devices can be obtained.
[0059] In the following, only the three-dimensional facial key point information corresponding to one target image acquisition device is taken as an example to introduce the acquisition method of the head pose ground truth. The three-dimensional facial key point information corresponding to other image acquisition devices can also perform corresponding operations through the same implementation method. For ease of explanation, in the embodiments of the present application, the image of the target object acquired by the target image acquisition device at the above-mentioned moment is referred to as the target image.
[0060] S104: Establish a facial coordinate system based on the three-dimensional facial key point information corresponding to the target image acquisition device, where the target image acquisition device is one of the multiple image acquisition devices.
[0061] When this step is specifically implemented, the face plane can be determined based on the three-dimensional key point information of the face corresponding to the target image acquisition device. As an example, the face plane can be constructed based on the coordinates of 3 to 4 face key points in the coordinate system of the target image acquisition device. The specific face key points selected are not limited.
[0062] Next, the normal vector of the face plane is determined according to the face plane. As mentioned before, the coordinates of 3 to 4 face key points are used when determining the face plane. Of course, two non-parallel spatial vectors can also be obtained through the coordinates of these 3 to 4 face key points. When specifically implemented, the normal vector of the face plane can be obtained by taking the cross product of these two vectors. In addition, the unit vector of the normal vector can also be determined through the cross function of NumPy (full name: Numerical Python, an open-source numerical computing extension of Python).
[0063] Combined with the previous text, the face plane and the normal vector of the face plane have been obtained. Furthermore, based on this plane and the normal vector of this plane, a face coordinate system of the target object can be established. It can be understood that since the three-dimensional key point information of the face is obtained in the coordinate system of the target image acquisition device, the face coordinate system established based on the three-dimensional key point information of the face is also based on the coordinate system of the target image acquisition device.
[0064] S105: According to the coordinate system of the target image acquisition device and the face coordinate system, obtain the true value of the head pose of the target object corresponding to the target image, where the target image is the image acquired by the target image acquisition device at a certain moment of the target object.
[0065] Based on the known coordinate system of the target image acquisition device and the face coordinate system, the rotation matrix of the face coordinate system relative to the coordinate system of the target image acquisition device can be obtained according to the face coordinate system and the coordinate system of the target image acquisition device. Since there is an association relationship between this rotation matrix and the three angles (yaw angle Yaw, pitch angle Pitch, and roll angle Roll) in the true value of the head pose, therefore, according to the rotation matrix, the true value of the head pose of the target object can be obtained based on this association relationship. Since this true value of the head pose is obtained based on the coordinate system of the target image acquisition device and the face coordinate system, and the face coordinate system is also based on the coordinate system of the target image acquisition device, and the target image acquisition device has a correspondence with the target image captured at this moment, the true value of the head pose obtained in this step can be associated with the target image.
[0066] For example, the image and the obtained ground truth of the head pose can also be stored as an image-ground truth pair. In this way, a one-to-one correspondence between the image and the ground truth is established. Storing the image-ground truth pair facilitates the use of the ground truth of the head pose in subsequent applications, such as training a model. As an example, the model can be a model that determines the ground truth of the head pose from an image, or further, the model can be a model that determines the driving safety factor or analyzes the degree of driver fatigue from an image, etc. The specific functions of the trained model are not limited here.
[0067] The above is the method for obtaining the ground truth of the head pose provided by the embodiments of the present application. In this method, multiple image acquisition devices are used to acquire images of the target object at the same time, and the two-dimensional face key points marked in the multiple acquired images are used as the data basis for three-dimensionally reconstructing the face key points of the target object, and then the ground truth of the head pose of the target object corresponding to the image is obtained. Since this solution does not need to rely on a wearable device to obtain the ground truth of the head pose, it is not affected by the wearing angle. By using multiple image acquisition devices at the same time and three-dimensionally reconstructing the face key points, the accuracy of the obtained ground truth of the head pose is guaranteed. Furthermore, when the ground truth of the head pose is applied to fields such as driver fatigue analysis, virtual reality somatosensory games, analysis of the desire to purchase goods, face verification, etc., the analysis results of the data can be made more accurate.
[0068] In a specific application, in order to provide more diverse and rich data for the subsequent usage scenarios of the ground truth of the head pose, the embodiments of the present application propose that, when implementing S101, multiple different lighting conditions can be provided to the space where the target object is located. Images of the target object acquired by multiple image acquisition devices at the same time under the same lighting condition are obtained. As an example, images acquired by all image acquisition devices at the same time are obtained under a primary lighting condition, images acquired by all image acquisition devices at the same time are obtained under a secondary lighting condition, and images acquired by all image acquisition devices at the same time are obtained under a tertiary lighting condition. In practical applications, even if it is the same image acquisition device and the target object maintains a fixed head pose, the ground truth of the head pose obtained by executing this method under different lighting conditions may be different. By acquiring images under different lighting conditions and respectively obtaining the ground truth of the head pose, it can help improve the accuracy of data analysis. For example, when data analysis needs to be performed under a primary lighting condition, since images are acquired and the ground truth of the head pose is obtained under the primary lighting condition in advance, rather than only acquiring images and obtaining the ground truth of the head pose under other lighting conditions, the data analysis results for this condition are made more accurate.
[0069] The first-level, second-level, and third-level lighting conditions introduced above are used as examples of different lighting conditions, and the levels of different lighting conditions are not limited here. For example, it can include a 4-level lighting condition. The higher the level, the weaker the lighting intensity; the lower the level, the higher the lighting intensity. In another example, the lighting conditions are divided into sunlight lighting conditions, infrared lighting conditions, ultraviolet lighting conditions, etc. The above lighting conditions can be realized by the natural environment or by lighting devices. For example, different lighting conditions are achieved by adjusting or controlling the selection of light source types, the opening and closing of lights, the strength, and the irradiation angle in the lighting device.
[0070] When the true value of the head pose needs to be applied to the data analysis scenario in the driving field, the embodiments of the present application can take adaptive measures in the image acquisition stage. For example, when multiple image acquisition devices acquire images of the target object, the target object is sitting on a seat, and this seat is used to simulate the seat in the real vehicle environment. For example, the target object represents a driver, and the seat where it is located represents the driver's seat. The setting parameters of the image acquisition device relative to the seat are determined according to the setting parameters of the simulated object relative to the seat in the real vehicle environment. The setting parameters include: position, height, angle, etc. The simulated object includes at least one of the following: A-pillar, B-pillar, dashboard, front windshield, or left side window glass. As an example, Device 1 is set to simulate the A-pillar, Device 2 is set to simulate the B-pillar, Device 3 is set to simulate the dashboard, Device 4 is set to simulate the front windshield, and Device 5 is set to simulate the left side window glass.
[0071] Figure 2 It is a schematic diagram of a scenario for acquiring images of a target object by multiple image acquisition devices provided by the embodiments of the present application. Figure 2 It shows that 14 cameras acquire images of the target object on the seat, and there is a lighting device surrounding the target object in the scenario, which is used to change the lighting conditions.
[0072] In addition to the advantages mentioned above, there are also many advantages in the technical solutions of the embodiments of the present application:
[0073] In the embodiments of the present application, when multiple image acquisition devices perform acquisition, they can record the videos of the same person at the same time, and calculate the continuous true value of the head pose according to the three-dimensional key point information of the face in the video frame number. Therefore, the technical solution of the present application has continuity in obtaining the true value of the head pose.
[0074] In this method, there are image acquisition devices that capture the side face images of people. Therefore, this solution has the ability to acquire the true value of the head pose at large angles.
[0075] This method adopts a non-contact head pose acquisition method, and the acquisition method is relatively simple and convenient to implement.
[0076] Based on the method for obtaining the ground truth of head pose provided in the foregoing embodiments, correspondingly, the present application also provides a device for obtaining the ground truth of head pose. The following describes the specific implementation of this device in conjunction with the embodiments. Figure 3 It is a schematic structural diagram of a device for obtaining the ground truth of head pose. As Figure 3 shown, the device 300 includes:
[0077] An image acquisition module 301, configured to acquire images of a target object collected by a plurality of image acquisition devices at the same time;
[0078] A key point annotation module 302, configured to annotate the face key points in the images collected by each of the plurality of image acquisition devices, and obtain the face key point information corresponding to each of the image acquisition devices;
[0079] A three-dimensional key point reconstruction module 303, configured to reconstruct the face three-dimensional key point information corresponding to each of the image acquisition devices based on the face key point information corresponding to each of the image acquisition devices and the parameters of each of the image acquisition devices;
[0080] A coordinate system establishment module 304, configured to establish a face coordinate system based on the face three-dimensional key point information corresponding to the target image acquisition device, where the target image acquisition device is one of the plurality of image acquisition devices;
[0081] A ground truth acquisition module 305, configured to obtain the ground truth of the head pose of the target object corresponding to the target image according to the coordinate system of the target image acquisition device and the face coordinate system, where the target image is an image of the target object collected by the target image acquisition device at the moment;
[0082] Multiple image acquisition devices collect images of a target object at the same time, use the two-dimensional face key points marked in the multiple collected images as the data basis for three-dimensional reconstruction of the face key points of the target object, and further obtain the ground truth of the head pose of the target object corresponding to the image. Since this solution does not need to rely on a wearable device to obtain the ground truth of head pose, it is not affected by the wearing angle. By using multiple image acquisition devices at the same time and three-dimensional reconstruction of face key points, the accuracy of the obtained ground truth of head pose is guaranteed. Furthermore, when the ground truth of head pose is applied to fields such as driver fatigue degree analysis, virtual reality somatosensory games, commodity purchase desire analysis, and face verification, the analysis results of the data can be made more accurate.
[0083] Optionally, the ground truth acquisition module 305 includes:
[0084] A rotation matrix acquisition unit, configured to obtain a rotation matrix of the face coordinate system relative to the coordinate system of the target image acquisition device according to the face coordinate system and the coordinate system of the target image acquisition device;
[0085] A ground truth acquisition unit, configured to obtain the ground truth of the head pose of the target object corresponding to the target image according to the rotation matrix.
[0086] Optionally, the coordinate system establishment module 304 includes:
[0087] A face plane determination unit, configured to determine a face plane based on the three-dimensional key point information of the face corresponding to the target image acquisition device;
[0088] A normal vector determination unit, configured to determine a normal vector of the face plane according to the face plane;
[0089] A coordinate system establishment unit, configured to establish a face coordinate system of the target object based on the face plane and the normal vector of the face plane.
[0090] Optionally, the three-dimensional key point reconstruction module 303 includes:
[0091] Reconstruct the three-dimensional key point information of the face corresponding to the target image acquisition device by means of triangulation reconstruction according to the face key point information corresponding to the target image acquisition device, the face key point information corresponding to the reference image acquisition device, the internal parameters of the target image acquisition device, the internal parameters of the reference image acquisition device, and the external parameters between the target image acquisition device and the reference image acquisition device; the reference image acquisition device is any one of the multiple image acquisition devices other than the target image acquisition device;
[0092] Obtain the three-dimensional key point information of the face corresponding to the reference image acquisition device according to the three-dimensional key point information of the face corresponding to the target image acquisition device and the external parameters.
[0093] Optionally, the apparatus 300 for obtaining the ground truth of the head pose further includes:
[0094] A storage module, configured to store the mutually corresponding images and the ground truth of the head pose as an image-ground truth pair.
[0095] Optionally, the image acquisition module 301 includes:
[0096] A lighting unit, configured to provide a variety of different lighting conditions to the space where the target object is located;
[0097] An image acquisition unit, configured to acquire images of the target object collected by the multiple image acquisition devices at the same moment under the same lighting conditions.
[0098] Based on the method and device for obtaining the ground truth of head pose provided in the foregoing embodiments, correspondingly, the present application further provides a device for obtaining the ground truth of head pose, including a processor and a memory; the memory is used for storing a computer program; the processor is used for executing the method for obtaining the ground truth of head pose provided in the foregoing method embodiments according to the computer program. Additionally, the processor in this device can also be used to control the lighting device to provide variable lighting conditions.
[0099] Based on the method, device and equipment for obtaining the ground truth of head pose provided in the foregoing embodiments, correspondingly, the present application further provides a computer-readable storage medium for storing a computer program, and the computer program, when run by a processor, executes the method for obtaining the ground truth of head pose provided in the foregoing method embodiments.
[0100] As described above, this is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for obtaining the true value of the head pose, characterized in that, Including: Obtaining images of a target object collected by multiple image acquisition devices at the same moment; Labeling face key points in the images collected by each of the multiple image acquisition devices to obtain face key point information corresponding to each of the image acquisition devices; Based on the face key point information corresponding to each of the image acquisition devices and the parameters of each of the image acquisition devices, reconstructing face three-dimensional key point information corresponding to each of the image acquisition devices; Establishing a face coordinate system based on the face three-dimensional key point information corresponding to a target image acquisition device, where the target image acquisition device is one of the multiple image acquisition devices; According to the coordinate system of the target image acquisition device and the face coordinate system, obtaining the true value of the head pose of the target object corresponding to a target image, where the target image is the image of the target object collected by the target image acquisition device at the moment; The reconstructing the face three-dimensional key point information corresponding to each of the image acquisition devices based on the face key point information corresponding to each of the image acquisition devices and the parameters of each of the image acquisition devices includes: According to the face key point information corresponding to the target image acquisition device, the face key point information corresponding to a reference image acquisition device, the internal parameters of the target image acquisition device, the internal parameters of the reference image acquisition device, and the external parameters between the target image acquisition device and the reference image acquisition device, reconstructing the face three-dimensional key point information corresponding to the target image acquisition device by using a triangulation reconstruction method; the reference image acquisition device is any one of the multiple image acquisition devices other than the target image acquisition device; According to the face three-dimensional key point information corresponding to the target image acquisition device and the external parameters, obtaining the face three-dimensional key point information corresponding to the reference image acquisition device.
2. The method according to claim 1, characterized in that, The obtaining the true value of the head pose of the target object corresponding to a target image according to the coordinate system of the target image acquisition device and the face coordinate system includes: According to the face coordinate system and the coordinate system of the target image acquisition device, obtaining a rotation matrix of the face coordinate system relative to the coordinate system of the target image acquisition device; According to the rotation matrix, obtaining the true value of the head pose of the target object corresponding to the target image.
3. The method according to claim 1, wherein The establishing the face coordinate system based on the face three-dimensional key point information corresponding to the target image acquisition device includes: Determining a face plane based on the face three-dimensional key point information corresponding to the target image acquisition device; Determining a normal vector of the face plane according to the face plane; Establishing the face coordinate system of the target object based on the face plane and the normal vector of the face plane.
4. The method according to any one of claims 1 to 3, characterized in that, Also including: Storing the mutually corresponding images and the true values of the head poses as image-true value pairs.
5. The method according to any one of claims 1-3, characterized in that The obtaining the images of the target object collected by multiple image acquisition devices at the same moment includes: Providing various different lighting conditions to the space where the target object is located; Obtaining the images of the target object collected by the multiple image acquisition devices at the same moment under the same lighting conditions.
6. The method according to any one of claims 1-3, characterized in that, When the multiple image acquisition devices acquire images of the target object, the target object is sitting on a seat, and the seat is used to simulate the seat in a real vehicle environment; the setting parameters of the image acquisition devices relative to the seat are determined based on the setting parameters of the simulated object relative to the seat in the real vehicle environment; The simulated object includes at least one of the following: A-pillar, B-pillar, instrument panel, front windshield or left side window glass.
7. An apparatus for obtaining the true value of a head pose, characterized in that, Comprising: An image acquisition module, configured to acquire images of the target object acquired by multiple image acquisition devices at the same moment; A key point annotation module, configured to annotate the face key points in the images acquired by each of the multiple image acquisition devices to obtain the face key point information corresponding to each of the image acquisition devices; A three-dimensional key point reconstruction module, configured to reconstruct the face three-dimensional key point information corresponding to each of the image acquisition devices based on the face key point information corresponding to each of the image acquisition devices and the parameters of each of the image acquisition devices; A coordinate system establishment module, configured to establish a face coordinate system based on the face three-dimensional key point information corresponding to the target image acquisition device, and the target image acquisition device is one of the multiple image acquisition devices; A ground truth acquisition module, configured to obtain the ground truth of the head pose of the target object corresponding to the target image according to the coordinate system of the target image acquisition device and the face coordinate system, and the target image is the image of the target object acquired by the target image acquisition device at the moment; The three-dimensional key point reconstruction module is specifically configured to: Reconstruct the face three-dimensional key point information corresponding to the target image acquisition device by means of triangulation reconstruction according to the face key point information corresponding to the target image acquisition device, the face key point information corresponding to the reference image acquisition device, the internal parameters of the target image acquisition device, the internal parameters of the reference image acquisition device, and the external parameters between the target image acquisition device and the reference image acquisition device; The reference image acquisition device is any one of the multiple image acquisition devices other than the target image acquisition device; Obtain the face three-dimensional key point information corresponding to the reference image acquisition device according to the face three-dimensional key point information corresponding to the target image acquisition device and the external parameters.
8. An apparatus for obtaining the ground truth of a head pose, characterized in that, Comprising a processor and a memory; the memory is used to store a computer program; the processor is used to execute the method for obtaining the ground truth of the head pose according to any one of claims 1-6 according to the computer program.
9. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program, when run by a processor, executes the method for obtaining the ground truth of the head pose according to any one of claims 1-6.
Citation Information
Patent Citations
Data processing method and device, equipment and storage medium
CN113630646A