Line-of-sight direction data collection method, apparatus, device, and storage medium
By acquiring images of the target object and the user's face using a first depth camera and a second depth camera, and determining and transforming the three-dimensional coordinates of the gaze direction data, the problems of low efficiency and low accuracy in the existing technology are solved, and efficient and accurate gaze direction data acquisition is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG ZEEKR INTELLIGENT TECH CO LTD
- Filing Date
- 2022-10-24
- Publication Date
- 2026-04-21
AI Technical Summary
Existing line-of-sight data acquisition schemes are inefficient and inaccurate, failing to meet the training requirements of deep learning algorithms and being limited by the field of view of depth cameras.
The first depth camera captures an image of the target object, and the second depth camera captures an image of the user's face. The three-dimensional coordinates of the object and the three-dimensional coordinates of the eyes are determined, and these coordinates are transformed in the coordinate system of each camera to determine the gaze direction data.
It improves the efficiency and accuracy of collecting line-of-sight data, expands the line-of-sight range, and meets the training requirements of deep learning algorithms.
Smart Images

Figure CN115588052B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data acquisition technology, and in particular to a method, apparatus, device, and storage medium for acquiring line-of-sight data. Background Technology
[0002] Existing line-of-sight data acquisition schemes generally use a single acquisition method, which can only acquire one line-of-sight data at a time. This results in low data acquisition efficiency, which cannot meet the training requirements of deep learning algorithms. Moreover, due to the influence of the depth camera's field of view, there are significant limitations on the data acquisition location, leading to low accuracy of the acquired line-of-sight data.
[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this invention is to provide a method, apparatus, device, and storage medium for acquiring line-of-sight data, aiming to solve the technical problems of low efficiency and accuracy in acquiring line-of-sight data in the prior art.
[0005] To achieve the above objectives, the present invention provides a method for acquiring line-of-sight direction data, the method comprising the following steps:
[0006] The target object is captured by a first depth camera, and its three-dimensional coordinates in the first depth camera coordinate system are determined based on the object image. The target object is the object that the user is looking at.
[0007] The user's facial image is acquired by a second depth camera, and the three-dimensional coordinates of the user's eyes in the second depth camera coordinate system are determined based on the facial image.
[0008] Determine the object's three-dimensional coordinates and the eye's three-dimensional coordinates in the coordinate systems of each acquisition camera;
[0009] The user's gaze direction data is determined based on the object coordinates and eye coordinates in the coordinate systems of each acquisition camera.
[0010] Optionally, determining the object's three-dimensional coordinates and the eye's three-dimensional coordinates in each acquisition camera coordinate system includes:
[0011] When the first depth camera coordinate system and the second depth camera coordinate system are the same, the object coordinates in each acquisition camera coordinate system are determined according to the external parameter matrix from the first depth camera coordinate system to each acquisition camera coordinate system.
[0012] The eye coordinates in each acquisition camera coordinate system are determined based on the extrinsic parameter matrix from the second depth camera coordinate system to each acquisition camera coordinate system.
[0013] Optionally, determining the object's three-dimensional coordinates and the eye's three-dimensional coordinates in each acquisition camera coordinate system further includes:
[0014] When the first depth camera coordinate system and the second depth camera coordinate system are not consistent, the three-dimensional coordinates of the object in the first preset calibration plate coordinate system are determined according to the extrinsic parameter matrix from the first depth camera coordinate system to the first preset calibration plate coordinate system.
[0015] The eye's three-dimensional coordinates in the first preset calibration plate coordinate system are determined based on the extrinsic parameter matrix from the second depth camera coordinate system to the first preset calibration plate coordinate system.
[0016] The object coordinates and eye coordinates in each acquisition camera coordinate system are determined based on the extrinsic parameter matrix from the first preset calibration plate coordinate system to each acquisition camera coordinate system.
[0017] Optionally, before determining the object's three-dimensional coordinates and the eye's three-dimensional coordinates in each acquisition camera coordinate system, the method further includes:
[0018] Determine the first extrinsic parameter matrix from the first depth camera coordinate system to the second preset calibration plate coordinate system, and determine the second extrinsic parameter matrix from the second depth camera coordinate system to the second preset calibration plate coordinate system;
[0019] Determine the first three-dimensional object coordinates of the target reference object in the first depth camera coordinate system, and determine the second three-dimensional object coordinates of the target reference object in the second depth camera coordinate system;
[0020] Determine whether the first depth camera coordinate system and the second depth camera coordinate system are consistent based on the first 3D object coordinates, the second 3D object coordinates, the first extrinsic parameter matrix, and the second extrinsic parameter matrix.
[0021] Optionally, determining the first three-dimensional object coordinates of the target reference object in the first depth camera coordinate system and determining the second three-dimensional object coordinates of the target reference object in the second depth camera coordinate system includes:
[0022] A first reference object image of the target reference object is acquired by a first depth camera, and the first three-dimensional object coordinates of the target reference object in the first depth camera coordinate system are determined based on the first reference object image.
[0023] A second reference object image of the target reference object is acquired by a second depth camera, and the second three-dimensional object coordinates of the target reference object in the second depth camera coordinate system are determined based on the second reference object image.
[0024] Optionally, determining whether the first depth camera coordinate system and the second depth camera coordinate system are consistent based on the first three-dimensional object coordinates, the second three-dimensional object coordinates, the first extrinsic parameter matrix, and the second extrinsic parameter matrix includes:
[0025] Based on the first extrinsic parameter matrix, the coordinates of the first three-dimensional object are converted into the coordinates of the first object in the second preset calibration plate coordinate system;
[0026] The second three-dimensional object coordinates are converted into second object coordinates in the second preset calibration plate coordinate system based on the second extrinsic parameter matrix.
[0027] When the coordinates of the first object are consistent with the coordinates of the second object, it is determined that the first depth camera coordinate system and the second depth camera coordinate system are the same; and
[0028] When the coordinates of the first object are inconsistent with the coordinates of the second object, it is determined that the coordinate systems of the first depth camera and the second depth camera are not consistent.
[0029] Optionally, determining the user's gaze direction data based on the object coordinates and eye coordinates in each acquisition camera coordinate system includes:
[0030] Determine the line-of-sight vector in each acquisition camera coordinate system based on the object coordinates and eye coordinates in each acquisition camera coordinate system;
[0031] The user's gaze direction data is determined based on the gaze vectors in the coordinate systems of each acquisition camera.
[0032] Furthermore, to achieve the above objectives, the present invention also proposes a line-of-sight direction data acquisition device, the device comprising:
[0033] The first coordinate determination module is used to acquire an object image of the target object through the first depth camera, and determine the three-dimensional coordinates of the target object in the first depth camera coordinate system based on the object image, wherein the target object is the object being viewed by the user.
[0034] The second coordinate determination module is used to acquire the user's facial image through the second depth camera, and determine the three-dimensional coordinates of the user's eyes in the second depth camera coordinate system based on the facial image.
[0035] The third coordinate determination module is used to determine the object's three-dimensional coordinates and the eye's three-dimensional coordinates in the coordinate systems of each acquisition camera;
[0036] The direction data determination module is used to determine the user's line-of-sight direction data based on the object coordinates and eye coordinates in the coordinate system of each acquisition camera.
[0037] Furthermore, to achieve the above objectives, the present invention also proposes a line-of-sight data acquisition device, the device comprising: a memory, a processor, and a line-of-sight data acquisition program stored in the memory and executable on the processor, the line-of-sight data acquisition program being configured to implement the steps of the line-of-sight data acquisition method described above.
[0038] In addition, to achieve the above objectives, the present invention also proposes a storage medium storing a line-of-sight data acquisition program, which, when executed by a processor, implements the steps of the line-of-sight data acquisition method described above.
[0039] This invention acquires an image of a target object using a first depth camera and determines the object's three-dimensional coordinates in the first depth camera coordinate system based on the image. The target object is the object being gazed at by the user. A second depth camera acquires an image of the user's face and determines the user's eye's three-dimensional coordinates in the second depth camera coordinate system based on the facial image. The invention then determines the object's three-dimensional coordinates and the eye's three-dimensional coordinates in each acquisition camera coordinate system. Finally, it determines the user's gaze direction data based on the object's coordinates and the eye's coordinates in each acquisition camera coordinate system. This invention acquires an image of the target object and an image of the user's face using both the first and second depth cameras, determines the object's three-dimensional coordinates and the eye's three-dimensional coordinates based on these images, transforms them to the respective acquisition camera coordinate systems, and determines the gaze direction data based on these coordinates. The increased range of gaze acquisition makes the gaze direction data more accurate, and the ability to acquire multiple gaze direction data points at once improves the acquisition efficiency. Attached Figure Description
[0040] Figure 1 This is a schematic diagram of the structure of the line-of-sight data acquisition device in the hardware operating environment involved in the embodiments of the present invention;
[0041] Figure 2 This is a flowchart illustrating the first embodiment of the line-of-sight data acquisition method of the present invention;
[0042] Figure 3This is a schematic diagram illustrating the acquisition of line-of-sight direction data in one embodiment of the line-of-sight direction data acquisition method of the present invention;
[0043] Figure 4 This is a flowchart illustrating the second embodiment of the line-of-sight data acquisition method of the present invention;
[0044] Figure 5 This is a schematic diagram of a first preset calibration plate being photographed by a camera in one embodiment of the line-of-sight data acquisition method of the present invention;
[0045] Figure 6 This is a flowchart illustrating the third embodiment of the line-of-sight direction data acquisition method of the present invention;
[0046] Figure 7 This is a schematic diagram of a second preset reference plate being photographed by a depth camera in one embodiment of the line-of-sight data acquisition method of the present invention;
[0047] Figure 8 This is a schematic diagram of a target reference object being photographed by a depth camera in one embodiment of the line-of-sight data acquisition method of the present invention;
[0048] Figure 9 This is a structural block diagram of the first embodiment of the line-of-sight direction data acquisition device of the present invention.
[0049] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0050] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0051] Reference Figure 1 , Figure 1 This is a schematic diagram of the line-of-sight data acquisition device structure in the hardware operating environment involved in the embodiments of the present invention.
[0052] like Figure 1As shown, the line-of-sight data acquisition device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be high-speed random access memory (RAM) or stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0053] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the line-of-sight data acquisition device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0054] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a network communication module, a user interface module, and a line-of-sight data acquisition program.
[0055] exist Figure 1 In the line-of-sight data acquisition device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the line-of-sight data acquisition device of the present invention can be set in the line-of-sight data acquisition device, and the line-of-sight data acquisition device calls the line-of-sight data acquisition program stored in the memory 1005 through the processor 1001 and executes the line-of-sight data acquisition method provided in the embodiment of the present invention.
[0056] This invention provides a method for acquiring line-of-sight direction data, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the line-of-sight data acquisition method of the present invention.
[0057] In this embodiment, the line-of-sight direction data acquisition method includes the following steps:
[0058] Step S10: Acquire an image of the target object using the first depth camera, and determine the three-dimensional coordinates of the target object in the first depth camera coordinate system based on the image. The target object is the object being viewed by the user.
[0059] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or line-of-sight data acquisition device capable of performing the above functions. The following description uses a line-of-sight data acquisition device (hereinafter referred to as the acquisition device) as an example to illustrate this embodiment and the subsequent embodiments.
[0060] It is understood that the depth camera used to acquire the object image of the target object is collectively referred to as the first depth camera. The number of first depth cameras can be one or more. The three-dimensional coordinates of the object can be the three-dimensional coordinates of the target object being viewed by the user in the object image under the coordinate system of the first depth camera. After the first depth camera captures the object image, it can automatically obtain the three-dimensional coordinates of the target object in the image. The number of target objects can be multiple. The target objects can be objects set on the background plate, or light spots or other objects projected onto the background plate by a laser device. This embodiment does not impose any restrictions here.
[0061] In this embodiment, the objects on the background panel for the user to gaze at can also be called gaze points. The gaze points can be fixed to the background panel. A depth image of the background panel can be acquired using a first depth camera, and the 3D coordinates of the objects at each gaze point on the background panel can be predetermined based on the depth image. When collecting gaze direction data, the target gaze point of the user can be determined, and the 3D coordinates of the object corresponding to the target gaze point can be determined based on the predetermined 3D coordinates of the object. Since existing acquisition schemes generally use moving gaze points, each acquired image requires coordinate calibration, which is time-consuming. This scheme uses fixed gaze points, ensuring that the coordinates of the gaze points in each image acquired by the first depth camera are fixed. Only the coordinates of the gaze points need to be calibrated once, and subsequent data can be directly reused, achieving automated coordinate annotation and reducing the time spent on data annotation. In this embodiment, the gaze points can also be moving. When collecting gaze direction data, the target gaze point is acquired using the first depth camera, resulting in a wider range of gaze data and richer gaze direction data, thereby improving the accuracy of subsequent deep learning algorithm training. Whether the objects set on the background board are fixed or moving can be determined according to the specific scenario; this embodiment does not impose any restrictions on this.
[0062] Step S20: Acquire the user's facial image using the second depth camera, and determine the three-dimensional coordinates of the user's eyes in the second depth camera coordinate system based on the facial image.
[0063] It is understood that the depth camera used to capture the user's facial image is collectively referred to as the second depth camera. The number of second depth cameras can be one or more. The three-dimensional coordinates of the eyes can be the three-dimensional coordinates of the eyes in the user's facial image under the coordinate system of the second depth camera when the user is looking at the target object. After the second depth camera captures the user's facial image, it can automatically obtain the three-dimensional coordinates of the eyes in the image. The three-dimensional coordinates of the eyes include the three-dimensional coordinates of the left eye and the three-dimensional coordinates of the right eye, which can be selected according to the specific scene. This embodiment does not impose any restrictions on this.
[0064] Step S30: Determine the object coordinates and eye coordinates in the coordinate systems of each acquisition camera.
[0065] It is understood that the acquisition camera can be a camera used to acquire user images, including red-green-blue cameras (RGB cameras) and infrared cameras; determining the object's three-dimensional coordinates and the eye's three-dimensional coordinates in each acquisition camera coordinate system can be achieved by transforming the object's three-dimensional coordinates in the first depth camera coordinate system to the respective acquisition camera coordinate systems to obtain the object coordinates in each acquisition camera coordinate system, and transforming the eye's three-dimensional coordinates in the second depth camera coordinate system to the respective acquisition camera coordinate systems to obtain the eye coordinates in each acquisition camera coordinate system; after the above coordinate transformation, the object coordinates and eye coordinates in the acquisition camera coordinate system corresponding to the user images acquired by each acquisition camera can be obtained, wherein the user images acquired by the acquisition cameras can be user facial images.
[0066] Step S40: Determine the user's gaze direction data based on the object coordinates and eye coordinates in each acquisition camera coordinate system.
[0067] It is understandable that determining the user's gaze direction data based on the object coordinates and eye coordinates in each acquisition camera's coordinate system can be achieved by determining the gaze direction data corresponding to the user images captured by each acquisition camera based on the object coordinates and eye coordinates in each acquisition camera's coordinate system; for example, if the number of acquisition cameras is N, then gaze direction data corresponding to N user images can be obtained.
[0068] In the specific implementation, the acquisition device acquires an image of the target object being gazed at by the user using a first depth camera, and determines the 3D coordinates of the target object in the first depth camera coordinate system based on the target object image in the object image. A second depth camera acquires a facial image of the user gazing at the target object, and determines the 3D coordinates of the eyes in the second depth camera coordinate system based on the eye image in the facial image. While acquiring the 3D coordinates of the object and eyes using the depth cameras, several acquisition cameras also acquire facial images of the user, resulting in several user facial images. The 3D coordinates of the object in the first depth camera coordinate system are transformed to the coordinate systems of each acquisition camera, obtaining several object coordinates. The 3D coordinates of the eyes in the second depth camera coordinate system are also transformed to the coordinate systems of each acquisition camera, obtaining several eye coordinates. Based on the object coordinates and eye coordinates in each acquisition camera coordinate system, the gaze direction data corresponding to each user facial image is determined.
[0069] For example, refer to Figure 3 , Figure 3This diagram illustrates the acquisition of line-of-sight data. Assume there is one first depth camera and one second depth camera, denoted by A and B respectively. First depth camera A is positioned behind the user, and second depth camera B is positioned in front of the user. Several objects for the user to gaze at are placed on a background wall. The user is positioned in front of the background wall. N acquisition cameras, including RGB cameras and infrared cameras, are positioned between the user and the background wall. These cameras are located in front of the user (e.g., within 0.5m to 1m in front of the user) and distributed within a preset angle range in front of the user (e.g., within ±45 degrees around the user). First depth camera A maintains a first preset distance from the background wall (e.g., the first preset distance is set to 1.5m to 3m). The second depth camera B maintains a second preset distance from the user (e.g., set to 0.5m to 1m). The first depth camera captures a depth image of the background wall. Based on this image, the 3D coordinates of each object on the background wall in the first depth camera coordinate system can be determined. These object coordinates can then be directly reused, saving data annotation time and further improving the efficiency of gaze direction data acquisition. Assuming the user is looking at object 1 on the background wall, the 3D coordinates of object 1 in the first depth camera coordinate system are obtained. The second depth camera B captures a facial image of the user, and the 3D coordinates of the user's eyes in the second depth camera coordinate system are determined based on the eye image in the facial image. While acquiring images with the depth cameras, N user facial images are also acquired using N acquisition cameras. The 3D coordinates of the objects and eyes in the depth camera coordinate system are transformed to the coordinate systems of the N acquisition cameras, obtaining the object coordinates and eye coordinates in the N acquisition camera coordinate systems. The gaze direction data corresponding to the N user facial images is then determined based on the object coordinates and eye coordinates in each acquisition camera coordinate system.
[0070] Furthermore, in order to improve the accuracy of the acquired gaze direction data, step S30 includes: when the first depth camera coordinate system and the second depth camera coordinate system are aligned, determining the object coordinates in each acquisition camera coordinate system based on the extrinsic parameter matrix from the first depth camera coordinate system to each acquisition camera coordinate system; and determining the eye coordinates in each acquisition camera coordinate system based on the extrinsic parameter matrix from the second depth camera coordinate system to each acquisition camera coordinate system.
[0071] It is understandable that the extrinsic matrix from the first depth camera coordinate system and the second depth camera coordinate system to each acquisition camera coordinate system can be obtained through extrinsic parameter calibration; the parameter matrices from the first depth camera coordinate system and the second depth camera coordinate system to each acquisition camera coordinate system are different.
[0072] Furthermore, in order to improve the efficiency of acquiring gaze direction data, step S40 includes: determining the gaze vector in each acquisition camera coordinate system based on the object coordinates and eye coordinates in each acquisition camera coordinate system; and determining the user's gaze direction data based on the gaze vector in each acquisition camera coordinate system.
[0073] In practical implementation, for example, extrinsic parameter calibration is performed between the acquisition camera and the second depth camera to obtain the extrinsic parameter matrix (R2, T2) from the second depth camera coordinate system to each acquisition camera coordinate system. Extrinsic parameter calibration is performed between the acquisition camera and the first depth camera to obtain the extrinsic parameter matrix (R1, T1) from the first depth camera coordinate system to each acquisition camera coordinate system. A depth image of the background wall is captured by the first depth camera A, obtaining the 3D coordinates of all objects on the background wall in the first depth camera coordinate system. Assuming the user is looking at object 1 on the background wall, the 3D coordinates of object 1 in the first depth camera coordinate system can be represented as deep_cam_point1. A depth image of the user's face is captured by the second depth camera B, and the 3D coordinates of the user's left eye in the second depth camera coordinate system are obtained: deep_cam_left_eye1. According to Formula 1, the 3D coordinates in the depth camera coordinate system can be converted to coordinates in the acquisition camera coordinate system.
[0074]
[0075] In the formula, X, Y, and Z represent the coordinates in the depth camera coordinate system, X_new, Y_new, and Z_new represent the coordinates in the acquisition camera coordinate system, and R and T are the external parameter matrices from the depth camera coordinate system to the acquisition camera coordinate system obtained through calibration. After coordinate transformation, the object's 3D coordinates deep_cam_point1 and the left eye's 3D coordinates deep_cam_left_eye1 can be obtained in the object coordinates cam_point1 and the left eye coordinates cam_left_eye1 in each acquisition camera coordinate system. The left eye's line of sight direction vector in each acquisition camera coordinate system can be expressed as cam_point1-cam_left_eye1.
[0076] This embodiment acquires an image of a target object using a first depth camera, and determines the object's 3D coordinates in the first depth camera coordinate system based on the image. The target object is the object the user is looking at. A second depth camera acquires an image of the user's face, and determines the user's eye 3D coordinates in the second depth camera coordinate system based on the image. The object's 3D coordinates and the eye's 3D coordinates are then determined in the coordinate systems of each acquisition camera. Finally, the user's gaze direction data is determined based on the object and eye coordinates in each acquisition camera coordinate system. This embodiment acquires an image of the target object and an image of the user's face using both the first and second depth cameras, and determines the object's 3D coordinates and the eye's 3D coordinates based on these images. These coordinates are then transformed to the coordinate systems of each acquisition camera to obtain the object and eye coordinates. The gaze direction data is determined based on these coordinates. The increased range of gaze data acquisition makes the gaze direction data more accurate, and the ability to acquire multiple gaze direction data points at once improves the acquisition efficiency.
[0077] refer to Figure 4 , Figure 4 This is a flowchart illustrating the second embodiment of the line-of-sight data acquisition method of the present invention.
[0078] Based on the first embodiment described above, in this embodiment, step S30 includes:
[0079] Step S301: When the first depth camera coordinate system and the second depth camera coordinate system are not consistent, determine the object's three-dimensional coordinates in the first preset calibration plate coordinate system based on the extrinsic parameter matrix from the first depth camera coordinate system to the first preset calibration plate coordinate system.
[0080] It is understandable that when the first depth camera coordinate system and the second depth camera coordinate system are not consistent, directly converting the three-dimensional coordinates under the depth camera coordinate system to the coordinate system of each acquisition camera will result in inaccurate line-of-sight direction data. The first preset calibration plate coordinate system can be the coordinate system corresponding to the first calibration plate that has been set in advance.
[0081] Step S302: Determine the eye calibration three-dimensional coordinates in the first preset calibration plate coordinate system based on the extrinsic parameter matrix from the second depth camera coordinate system to the first preset calibration plate coordinate system;
[0082] Step S303: Determine the object coordinates and eye coordinates in each acquisition camera coordinate system based on the extrinsic parameter matrix from the first preset calibration plate coordinate system to each acquisition camera coordinate system.
[0083] In specific implementation, refer to Figure 5 , Figure 5 To capture a schematic diagram of the first preset calibration board using a camera, a checkerboard calibration board is pre-set. The checkerboard calibration board is captured by a depth camera and various acquisition cameras. Extrinsic parameter calibration is performed between the depth camera and the checkerboard calibration board, obtaining the extrinsic parameter matrix from the depth camera coordinate system to the checkerboard calibration board coordinate system. Extrinsic parameter calibration is then performed between each acquisition camera and the checkerboard calibration board, obtaining the extrinsic parameter matrix from each acquisition camera coordinate system to the checkerboard calibration board coordinate system. Formula 2 is used to transform the object's 3D coordinates in the first depth camera coordinate system to the checkerboard calibration board coordinate system, obtaining the object's calibrated 3D coordinates in the checkerboard calibration board coordinate system. Assuming the eye's 3D coordinates are those of the left eye, Formula 2 is used to transform the left eye's 3D coordinates in the second depth camera coordinate system to the checkerboard calibration board coordinate system, obtaining the left eye's calibrated 3D coordinates in the checkerboard calibration board coordinate system. Formula 1 is used to transform the object's calibrated 3D coordinates in the checkerboard calibration board coordinate system to the coordinates of each acquisition camera, obtaining the object's coordinates in each acquisition camera coordinate system. Formula 1 is used to transform the left-eye 3D coordinates in the checkerboard calibration board coordinate system to the coordinate systems of each acquisition camera, thus obtaining the left-eye 3D coordinates in each acquisition camera coordinate system. The left-eye gaze direction vector can be determined based on the object coordinates and the left-eye 3D coordinates in each acquisition camera coordinate system; the process for determining the right-eye gaze direction vector is similar to that for the left-eye gaze direction vector, and will not be repeated here in this embodiment.
[0084]
[0085] In the formula, X_w, Y_w, and Z_w are the three-dimensional coordinates in the coordinate system of the checkerboard calibration board; X_c, Y_c, and Z_c are the coordinates in the coordinate system of each acquisition camera.
[0086] In this embodiment, when the coordinate systems of the first and second depth cameras are not unified, the three-dimensional coordinates of the object in the first depth camera coordinate system are transformed to the first preset calibration plate coordinate system, and the three-dimensional coordinates of the eye in the second depth camera coordinate system are transformed to the first preset calibration plate coordinate system. Then, the coordinates in the first preset calibration plate coordinate system are transformed to the coordinate systems of each acquisition camera to obtain the object coordinates and eye coordinates. This allows for the unification of depth camera coordinates before determining the line-of-sight direction data when the depth camera coordinate systems are not unified, thus improving the accuracy of the acquired line-of-sight direction data.
[0087] refer to Figure 6 , Figure 6 This is a flowchart illustrating the third embodiment of the line-of-sight data acquisition method of the present invention.
[0088] Based on the above embodiments, in this embodiment, before step S30, the method further includes:
[0089] Step S01: Determine the first extrinsic parameter matrix from the first depth camera coordinate system to the second preset calibration plate coordinate system, and determine the second extrinsic parameter matrix from the second depth camera coordinate system to the second preset calibration plate coordinate system.
[0090] It is understandable that the second preset calibration plate coordinate system can be the coordinate system corresponding to the second preset calibration plate that has been set in advance; the first extrinsic parameter matrix from the first depth camera coordinate system to the second preset calibration plate coordinate system and the second extrinsic parameter matrix from the second depth camera coordinate system to the second preset calibration plate coordinate system can be obtained through extrinsic parameter calibration.
[0091] Step S02: Determine the first three-dimensional object coordinates of the target reference object in the first depth camera coordinate system, and determine the second three-dimensional object coordinates of the target reference object in the second depth camera coordinate system.
[0092] It is understandable that the target reference object can be an object used to determine whether the first depth camera coordinate system and the second depth camera coordinate system are consistent; the first three-dimensional object coordinates can be the coordinates of the target reference object in the first depth camera coordinate system; and the second three-dimensional object coordinates can be the coordinates of the target reference object in the second depth camera coordinate system.
[0093] Step S03: Determine whether the first depth camera coordinate system and the second depth camera coordinate system are consistent based on the first three-dimensional object coordinates, the second three-dimensional object coordinates, the first extrinsic parameter matrix, and the second extrinsic parameter matrix.
[0094] In practice, extrinsic parameter calibration determines the first and second extrinsic parameter matrices from the first and second depth cameras to the second preset calibration plate coordinate system. The first and second 3D object coordinates of the target reference object are determined in the first and second depth camera coordinate systems, respectively. The first 3D object coordinates are transformed to the second preset calibration plate coordinate system based on the first extrinsic parameter matrix, and the second 3D object coordinates are transformed to the second preset calibration plate coordinate system based on the second extrinsic parameter matrix. Finally, the consistency between the first and second depth camera coordinate systems is determined based on the coordinate system under the second preset calibration plate.
[0095] Furthermore, in order to determine whether the coordinate systems of the depth cameras are unified, step S02 includes: acquiring a first reference object image of the target reference object using a first depth camera, and determining the first three-dimensional object coordinates of the target reference object in the first depth camera coordinate system based on the first reference object image; acquiring a second reference object image of the target reference object using a second depth camera, and determining the second three-dimensional object coordinates of the target reference object in the second depth camera coordinate system based on the second reference object image.
[0096] Furthermore, in order to determine whether the depth camera coordinate system is unified, step S03 includes: converting the first three-dimensional object coordinates into first object coordinates in the second preset calibration plate coordinate system according to the first extrinsic parameter matrix; converting the second three-dimensional object coordinates into second object coordinates in the second preset calibration plate coordinate system according to the second extrinsic parameter matrix; determining that the first depth camera coordinate system and the second depth camera coordinate system are unified when the first object coordinates and the second object coordinates are consistent; and determining that the first depth camera coordinate system and the second depth camera coordinate system are not unified when the first object coordinates and the second object coordinates are inconsistent.
[0097] It is understandable that the coordinates of the first three-dimensional object and the coordinates of the second three-dimensional object can be converted into the coordinates of the first object and the coordinates of the second object using Formula 1.
[0098] In specific implementation, refer to Figure 7 and Figure 8 , Figure 7 This is a schematic diagram of a second preset reference board being photographed using a depth camera. Figure 8 To illustrate the use of depth cameras to capture images of a target reference object, for example, during coordinate system verification of a first depth camera and a second depth camera, images of a second preset calibration board are captured using both cameras. Extrinsic parameter calibration is used to obtain the first and second extrinsic parameter matrices from the first and second depth camera coordinate systems to the second preset calibration board coordinate system. Images of the target reference object are captured using both cameras, resulting in first and second reference object images. The coordinates of the first 3D object in the first depth camera coordinate system and the second 3D object in the second depth camera coordinate system are determined based on these images. The object coordinates in the depth camera coordinate system are transformed to the second preset calibration board coordinate system using Formula 1 and the first extrinsic parameter matrix, yielding the first and second object coordinates. If the first and second object coordinates are consistent, the coordinate system of the two depth cameras is determined to be unified; otherwise, the coordinate systems of the two depth cameras are determined to be inconsistent.
[0099] This embodiment determines the extrinsic parameter matrix from the first depth camera coordinate system and the second depth camera coordinate system to the second preset calibration plate coordinate system. Based on the extrinsic parameter matrix, the coordinates of the three-dimensional objects in the first depth camera coordinate system and the second depth camera coordinate system are transformed to the second preset calibration plate coordinate system. When the coordinates of the two objects in the second preset calibration plate coordinate system are consistent, the two depth camera coordinate systems are determined to be of the same type. This improves the accuracy of the subsequent acquisition of line-of-sight direction data while improving the determination of the depth camera coordinate system.
[0100] Furthermore, this embodiment of the invention also proposes a storage medium storing a line-of-sight data acquisition program, which, when executed by a processor, implements the steps of the line-of-sight data acquisition method described above.
[0101] Reference Figure 9 , Figure 9 This is a structural block diagram of the first embodiment of the line-of-sight direction data acquisition device of the present invention.
[0102] like Figure 9 As shown, the line-of-sight direction data acquisition device proposed in this embodiment of the invention includes:
[0103] The first coordinate determination module 10 is used to acquire an object image of a target object through a first depth camera, and determine the three-dimensional coordinates of the target object in the first depth camera coordinate system based on the object image, wherein the target object is the object being viewed by the user.
[0104] The second coordinate determination module 20 is used to acquire the user's facial image through the second depth camera, and determine the three-dimensional coordinates of the user's eyes in the second depth camera coordinate system based on the facial image.
[0105] The third coordinate determination module 30 is used to determine the object's three-dimensional coordinates and the eye's three-dimensional coordinates in the coordinate systems of each acquisition camera;
[0106] The direction data determination module 40 is used to determine the user's line-of-sight direction data based on the object coordinates and eye coordinates in the coordinate system of each acquisition camera.
[0107] This embodiment acquires an image of a target object using a first depth camera, and determines the object's 3D coordinates in the first depth camera coordinate system based on the image. The target object is the object the user is looking at. A second depth camera acquires an image of the user's face, and determines the user's eye 3D coordinates in the second depth camera coordinate system based on the image. The object's 3D coordinates and the eye's 3D coordinates are then determined in the coordinate systems of each acquisition camera. Finally, the user's gaze direction data is determined based on the object and eye coordinates in each acquisition camera coordinate system. This embodiment acquires an image of the target object and an image of the user's face using both the first and second depth cameras, and determines the object's 3D coordinates and the eye's 3D coordinates based on these images. These coordinates are then transformed to the coordinate systems of each acquisition camera to obtain the object and eye coordinates. The gaze direction data is determined based on these coordinates. The increased range of gaze data acquisition makes the gaze direction data more accurate, and the ability to acquire multiple gaze direction data points at once improves the acquisition efficiency.
[0108] Based on the first embodiment of the line-of-sight data acquisition device of the present invention, a second embodiment of the line-of-sight data acquisition device of the present invention is proposed.
[0109] In this embodiment, the third coordinate determination module 30 is further configured to, when the first depth camera coordinate system and the second depth camera coordinate system are aligned, determine the object coordinates of the three-dimensional coordinates of the object in each acquisition camera coordinate system based on the extrinsic parameter matrix from the first depth camera coordinate system to each acquisition camera coordinate system; and determine the eye coordinates of the three-dimensional coordinates of the eye in each acquisition camera coordinate system based on the extrinsic parameter matrix from the second depth camera coordinate system to each acquisition camera coordinate system.
[0110] The third coordinate determination module 30 is further configured to, when the first depth camera coordinate system and the second depth camera coordinate system are not consistent, determine the object calibration three-dimensional coordinates in the first preset calibration plate coordinate system based on the extrinsic parameter matrix from the first depth camera coordinate system to the first preset calibration plate coordinate system; determine the eye calibration three-dimensional coordinates in the first preset calibration plate coordinate system based on the extrinsic parameter matrix from the second depth camera coordinate system to the first preset calibration plate coordinate system; and determine the object coordinates and eye coordinates in each acquisition camera coordinate system based on the extrinsic parameter matrix from the first preset calibration plate coordinate system to each acquisition camera coordinate system.
[0111] The third coordinate determination module 30 is further configured to determine the first extrinsic parameter matrix from the first depth camera coordinate system to the second preset calibration plate coordinate system, and to determine the second extrinsic parameter matrix from the second depth camera coordinate system to the second preset calibration plate coordinate system; determine the first three-dimensional object coordinates of the target reference object in the first depth camera coordinate system, and to determine the second three-dimensional object coordinates of the target reference object in the second depth camera coordinate system; and determine whether the first depth camera coordinate system and the second depth camera coordinate system are consistent based on the first three-dimensional object coordinates, the second three-dimensional object coordinates, the first extrinsic parameter matrix, and the second extrinsic parameter matrix.
[0112] The third coordinate determination module 30 is further configured to acquire a first reference object image of the target reference object using a first depth camera, and determine the first three-dimensional object coordinates of the target reference object in the first depth camera coordinate system based on the first reference object image; acquire a second reference object image of the target reference object using a second depth camera, and determine the second three-dimensional object coordinates of the target reference object in the second depth camera coordinate system based on the second reference object image.
[0113] The third coordinate determination module 30 is further configured to convert the first three-dimensional object coordinates into first object coordinates in the second preset calibration plate coordinate system according to the first extrinsic parameter matrix; convert the second three-dimensional object coordinates into second object coordinates in the second preset calibration plate coordinate system according to the second extrinsic parameter matrix; determine that the first depth camera coordinate system and the second depth camera coordinate system are consistent when the first object coordinates and the second object coordinates are consistent; and determine that the first depth camera coordinate system and the second depth camera coordinate system are inconsistent when the first object coordinates and the second object coordinates are inconsistent.
[0114] The direction data determination module 40 is further configured to determine the gaze vector in each acquisition camera coordinate system based on the object coordinates and eye coordinates in each acquisition camera coordinate system; and to determine the user's gaze direction data based on the gaze vector in each acquisition camera coordinate system.
[0115] Other embodiments or specific implementations of the line-of-sight data acquisition device of the present invention can be referred to the above-described method embodiments, and will not be repeated here.
[0116] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0117] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0118] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory / random access memory, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0119] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A line-of-sight direction data collection method, characterized by, The method includes: The first depth camera captures an image of the target object fixed on the background board, and determines the three-dimensional coordinates of the target object in the coordinate system of the first depth camera based on the image. The first depth camera is set behind the user, and the target object is the object that the user is looking at. The user's facial image is acquired by a second depth camera, and the three-dimensional coordinates of the user's eyes in the second depth camera coordinate system are determined based on the facial image. Determine the object's three-dimensional coordinates and the eye's three-dimensional coordinates in the coordinate systems of each acquisition camera; The user's gaze direction data is determined based on the object coordinates and eye coordinates in the coordinate systems of each acquisition camera. The process of determining the user's gaze direction data based on the object coordinates and eye coordinates in each acquisition camera coordinate system includes: The user's facial image is captured by each of the cameras. Determine the line-of-sight vector in each acquisition camera coordinate system based on the object coordinates and eye coordinates in each acquisition camera coordinate system; The gaze direction data corresponding to each user's facial image is determined based on the gaze vector in the coordinate system of each acquisition camera.
2. The method of claim 1, wherein, Determining the object's three-dimensional coordinates and the eye's three-dimensional coordinates in the coordinate systems of each acquisition camera includes: When the first depth camera coordinate system and the second depth camera coordinate system are the same, the object coordinates in each acquisition camera coordinate system are determined according to the external parameter matrix from the first depth camera coordinate system to each acquisition camera coordinate system. The eye coordinates in each acquisition camera coordinate system are determined based on the extrinsic parameter matrix from the second depth camera coordinate system to each acquisition camera coordinate system.
3. The method of claim 1, wherein, The step of determining the object's three-dimensional coordinates and the eye's three-dimensional coordinates in the coordinate systems of each acquisition camera further includes: When the first depth camera coordinate system and the second depth camera coordinate system are not consistent, the three-dimensional coordinates of the object in the first preset calibration plate coordinate system are determined according to the extrinsic parameter matrix from the first depth camera coordinate system to the first preset calibration plate coordinate system. The eye's three-dimensional coordinates in the first preset calibration plate coordinate system are determined based on the extrinsic parameter matrix from the second depth camera coordinate system to the first preset calibration plate coordinate system. The object coordinates and eye coordinates in each acquisition camera coordinate system are determined based on the extrinsic parameter matrix from the first preset calibration plate coordinate system to each acquisition camera coordinate system.
4. The method according to any one of claims 1 to 3, characterized in that, Before determining the object's three-dimensional coordinates and the eye's three-dimensional coordinates in the coordinate systems of each acquisition camera, the method further includes: Determine the first extrinsic parameter matrix from the first depth camera coordinate system to the second preset calibration plate coordinate system, and determine the second extrinsic parameter matrix from the second depth camera coordinate system to the second preset calibration plate coordinate system; Determine the first three-dimensional object coordinates of the target reference object in the first depth camera coordinate system, and determine the second three-dimensional object coordinates of the target reference object in the second depth camera coordinate system; Determine whether the first depth camera coordinate system and the second depth camera coordinate system are consistent based on the first 3D object coordinates, the second 3D object coordinates, the first extrinsic parameter matrix, and the second extrinsic parameter matrix.
5. The method of claim 4, wherein, Determining the first three-dimensional object coordinates of the target reference object in the first depth camera coordinate system and determining the second three-dimensional object coordinates of the target reference object in the second depth camera coordinate system includes: A first reference object image of the target reference object is acquired by a first depth camera, and the first three-dimensional object coordinates of the target reference object in the first depth camera coordinate system are determined based on the first reference object image. A second reference object image of the target reference object is acquired by a second depth camera, and the second three-dimensional object coordinates of the target reference object in the second depth camera coordinate system are determined based on the second reference object image.
6. The method of claim 5, wherein, The step of determining whether the first depth camera coordinate system and the second depth camera coordinate system are consistent based on the first 3D object coordinates, the second 3D object coordinates, the first extrinsic parameter matrix, and the second extrinsic parameter matrix includes: Based on the first extrinsic parameter matrix, the coordinates of the first three-dimensional object are converted into the coordinates of the first object in the second preset calibration plate coordinate system; The second three-dimensional object coordinates are converted into second object coordinates in the second preset calibration plate coordinate system based on the second extrinsic parameter matrix. When the coordinates of the first object are consistent with the coordinates of the second object, it is determined that the first depth camera coordinate system and the second depth camera coordinate system are the same; and When the coordinates of the first object are inconsistent with the coordinates of the second object, it is determined that the coordinate systems of the first depth camera and the second depth camera are not consistent.
7. A line-of-sight data collection device, comprising: The device includes: The first coordinate determination module is used to acquire an object image of a target object fixed on a background plate through a first depth camera, and determine the three-dimensional coordinates of the target object in the first depth camera coordinate system based on the object image. The first depth camera is set behind the user, and the target object is the object that the user is looking at. The second coordinate determination module is used to acquire the user's facial image through the second depth camera, and determine the three-dimensional coordinates of the user's eyes in the second depth camera coordinate system based on the facial image. The third coordinate determination module is used to determine the object's three-dimensional coordinates and the eye's three-dimensional coordinates in the coordinate systems of each acquisition camera; The orientation data determination module is used to determine the user's line-of-sight orientation data based on the object coordinates and eye coordinates in the coordinate system of each acquisition camera. The direction data determination module is further configured to acquire user facial images through each acquisition camera; determine the gaze vector in each acquisition camera coordinate system based on the object coordinates and eye coordinates in each acquisition camera coordinate system; and determine the gaze direction data corresponding to each user facial image based on the gaze vector in each acquisition camera coordinate system.
8. A line-of-sight data collection device, comprising: The device includes: a memory, a processor, and a line-of-sight data acquisition program stored in the memory and executable on the processor, the line-of-sight data acquisition program being configured to implement the steps of the line-of-sight data acquisition method as described in any one of claims 1 to 6.
9. A storage medium, characterized by The storage medium stores a line-of-sight data acquisition program, which, when executed by a processor, implements the steps of the line-of-sight data acquisition method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Sight line calibration method and device, equipment, computer readable storage medium, system and vehicle
CN113661495A