Method, apparatus, device, and storage medium for acquiring gaze direction data
The method uses multiple deep cameras to determine three-dimensional coordinates of objects and eyes, aligning camera systems for accurate gaze direction data collection, addressing low efficiency and accuracy issues in conventional methods.
Patent Information
- Application Number
- JP2025520771
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-10-24
- Filing Date
- 2023-08-08
- Publication Date
- 2026-08-26
- Estimated Expiration
- 2043-08-08
AI Technical Summary
Conventional line-of-sight direction data collection methods suffer from low efficiency and accuracy, limited data collection locations due to viewing angle constraints, and inability to meet the training requirements of deep-learning algorithms.
A method involving multiple deep cameras to collect object and facial images, determining three-dimensional coordinates of objects and eyes, and aligning these coordinates across different camera systems to accurately calculate gaze direction data, enhancing data collection efficiency and accuracy.
Improves gaze direction data collection efficiency and accuracy by allowing multiple data points to be collected simultaneously, increasing the range of gaze data and enhancing the training of deep-learning algorithms.
Smart Images

Figure 0007911633000003 
Figure 0007911633000004 
Figure 0007911633000005
Abstract
Description
Technical Field
[0003] , , ,
[0005] , , , ,
[0006] , , ,
[0004] , , , ,
[0001] (Cross - reference to related applications) This application claims the priority of the Chinese patent application with the application number 202211305842.9 filed on October 24, 2022, and all of its content is incorporated herein by reference.<00><000006>
[0002] This application relates to the field of data collection technology, and particularly to a method, apparatus, device, and storage medium for collecting line - of - sight direction data.
Background Art
[0003] Conventional line - of - sight direction data collection solutions generally can only collect one line - of - sight direction data each time by a single - collection method, with low data collection efficiency, unable to meet the training requirements of deep - learning algorithms. Moreover, affected by the viewing angle of the deep camera, there are great limitations on the data collection location, and the accuracy of the collected line - of - sight direction data is low.
[0004] The above content is only for assisting the technical understanding of this application, and it does not admit that the above content is prior art.
Summary of the Invention
Problems to be Solved by the Invention
[0005] The main objective of this application is to provide a method, apparatus, device, and storage medium for collecting line - of - sight direction data to solve the technical problems in the prior art, such as low collection efficiency and accuracy of line - of - sight direction data.
Means for Solving the Problems
[0007] In one embodiment, the step of determining the object coordinates and eye coordinates on the respective camera coordinate systems for the three-dimensional coordinates of the object and the three-dimensional coordinates of the eye is as follows: When the first deep camera coordinate system and the second deep camera coordinate system are unified, the steps include determining the object coordinates on each acquisition camera coordinate system based on the external parameter matrix from the first deep camera coordinate system to each acquisition camera coordinate system matrix, and The process includes the step of determining the eye coordinates on each acquisition camera coordinate system based on an external parameter matrix from the second deep camera coordinate system to each acquisition camera coordinate system matrix, thereby determining the eye coordinates on each acquisition camera coordinate system of the three-dimensional eye coordinates.
[0008] In one embodiment, the step of determining the object coordinates and eye coordinates on the respective camera coordinate systems for the three-dimensional coordinates of the object and the three-dimensional coordinates of the eye is as follows: If the first deep camera coordinate system and the second deep camera coordinate system are not unified, the steps include determining the object calibration three-dimensional coordinates of the object three-dimensional coordinates on the first preset calibration plate coordinate system based on an external parameter matrix from the first deep camera coordinate system to a first preset calibration plate coordinate system, A step of determining the eye calibration three-dimensional coordinates of the eye on the first preset calibration plate coordinate system based on an external parameter matrix from the second deep camera coordinate system to the first preset calibration plate coordinate system, The method further includes the step of determining the object coordinates and eye coordinates on each acquisition camera coordinate system based on the external parameter matrix from the first preset calibration plate coordinate system to each acquisition camera coordinate system matrix, and the object coordinates and eye coordinates on each acquisition camera coordinate system, respectively, for the object calibration three-dimensional coordinates and the eye calibration three-dimensional coordinates.
[0009] In one embodiment, before the step of determining the object coordinates and eye coordinates on the respective camera coordinate systems for the three-dimensional coordinates of the object and the three-dimensional coordinates of the eye, The steps include determining a first external parameter matrix from a first deep camera coordinate system to a second pre-configured calibration plate coordinate system, and determining a second external parameter matrix from a second deep camera coordinate system to a second pre-configured calibration plate coordinate system, The steps include determining the first three-dimensional object coordinates of the target reference object in the first deep camera coordinate system, and determining the second three-dimensional object coordinates of the target reference object in the second deep camera coordinate system, The method further includes the step of determining whether the first deep camera coordinate system and the second deep camera coordinate system are unified based on the first three-dimensional object coordinates, the second three-dimensional object coordinates, the first external parameter matrix, and the second external parameter matrix.
[0010] In one embodiment, the steps of determining the first three-dimensional object coordinates of the target reference object on the first deep camera coordinate system and determining the second three-dimensional object coordinates of the target reference object on the second deep camera coordinate system are as follows: The steps include: collecting a first reference object image of the target reference object using a first deep camera, and determining the first three-dimensional object coordinates of the target reference object on the first deep camera coordinate system based on the first reference object image; The method includes the steps of: collecting a second reference object image of the target reference object using a second deep camera; and determining the second three-dimensional object coordinates of the target reference object on the second deep camera coordinate system based on the second reference object image.
[0011] In one embodiment, the step of determining whether the first deep camera coordinate system and the second deep camera coordinate system are unified based on the first three-dimensional object coordinates, the second three-dimensional object coordinates, the first external parameter matrix, and the second external parameter matrix is: A step of converting the first three-dimensional object coordinates to the first object coordinates on the second preset calibration plate coordinate system based on the first external parameter matrix, The steps include transforming the second three-dimensional object coordinates into second object coordinates on the second pre-defined calibration plate coordinate system based on the second external parameter matrix, If the coordinates of the first object and the coordinates of the second object coincide, the first deep camera coordinate system and the second deep camera coordinate system are determined to be unified. The method includes the step of determining that the first deep camera coordinate system and the second deep camera coordinate system are not unified if the first object coordinate system and the second object coordinate system do not coincide.
[0012] In one embodiment, the step of determining the user's line of sight direction data based on the object coordinates and eye coordinates on each acquired camera coordinate system is: The steps include determining the line-of-sight vector in each camera coordinate system based on the object coordinates and eye coordinates in each camera coordinate system, The process includes the step of determining the user's line of sight direction data based on the line of sight vector in each acquired camera coordinate system.
[0013] Furthermore, in order to achieve the above objective, this application provides a gaze direction data acquisition device, and the gaze direction data acquisition device is A first coordinate determination module for collecting an object image of a target object, which is an object that the user is focusing on, using a first deep camera, and determining the three-dimensional coordinates of the target object on the first deep camera coordinate system based on the object image, A second coordinate determination module for collecting the user's facial image using a second deep camera and determining the three-dimensional coordinates of the user's eye on the second deep camera coordinate system based on the facial image, A third coordinate determination module for determining the object coordinates and eye coordinates on the respective camera coordinate systems for the three-dimensional coordinates of the object and the three-dimensional coordinates of the eye, The system includes a direction data determination module for determining the user's line of sight direction data based on object coordinates and eye coordinates in each acquired camera coordinate system.
[0014] To achieve the above objective, this application further provides a gaze direction data acquisition device comprising a memory, a processor, and a gaze direction data acquisition program stored in the memory and executable on the processor, wherein the gaze direction data acquisition program is configured to implement the steps of the gaze direction data acquisition method described above.
[0015] To achieve the above objective, this application further provides a storage medium in which a gaze direction data acquisition program is stored, and when the gaze direction data acquisition program is executed by a processor, the steps of the gaze direction data acquisition method described above are realized. [Effects of the Invention]
[0016] In this application, an object image of a target object that the user is gazing at is collected by a first deep camera, and based on the object image, the object three-dimensional coordinates of the target object in the first deep camera coordinate system are determined. A facial image of the user is collected by a second deep camera, and based on the facial image, the eye three-dimensional coordinates of the user's eyes in the second deep camera coordinate system are determined. The object coordinates and eye coordinates in the collection camera coordinate systems of the object three-dimensional coordinates and the eye three-dimensional coordinates are determined, and based on the object coordinates and eye coordinates in each collection camera coordinate system, the gaze direction data of the user is determined. In this application, an object image of a target object and a facial image of a user are respectively collected by a first deep camera and a second one. Based on the object image and the facial image, the object three-dimensional coordinates and the eye three-dimensional coordinates are determined, the object three-dimensional coordinates and the eye three-dimensional coordinates are converted into each collection camera coordinate system, the object coordinates and the eye coordinates are obtained, and based on the object coordinates and eye coordinates in each collection camera coordinate system, the gaze direction data is determined. By increasing the collected gaze range, the gaze direction data is made more accurate, and moreover, multiple gaze direction data are collected each time, and the collection efficiency of the gaze direction data can be improved.
Brief Description of the Drawings
[0017] [Figure 1] It is a schematic diagram of the structure of a gaze direction data collection device in a hardware execution environment according to an embodiment aspect of this application. [Figure 2] It is a flowchart of a first embodiment of the gaze direction data collection method of this application. [Figure 3] It is a schematic diagram of collecting gaze direction data in an embodiment of the gaze direction data collection method of this application. [Figure 4] It is a flowchart of a second embodiment of the gaze direction data collection method of this application. [Figure 5] It is a schematic diagram of photographing a first preset calibration plate by a camera in an embodiment of the gaze direction data collection method of this application. [Figure 6] It is a flowchart of a third embodiment of the gaze direction data collection method of this application. [Figure 7]This is a schematic diagram showing a second preset reference plate being photographed by a deep camera in one embodiment of the line-of-sight direction data acquisition method of the present application. [Figure 8] This is a schematic diagram illustrating one embodiment of the line-of-sight direction data acquisition method of this application, in which a target reference object is photographed by a deep camera. [Figure 9] This is a structural block diagram of the first embodiment of the line-of-sight data acquisition device of this application. [Modes for carrying out the invention]
[0018] The achievement of the objectives, functional features, and advantages of this application will be further described in conjunction with the examples and with reference to the drawings. It should be understood that the specific embodiments described herein are used solely for the purpose of illustrating this application and are not intended to limit it.
[0019] Referring to Figure 1, Figure 1 is a schematic diagram of the structure of a line-of-sight data acquisition device in a hardware execution environment according to an embodiment of this application.
[0020] As shown in Figure 1, this gaze direction data acquisition device may include a processor 1001, for example, a Central Processing Unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable connection communication between these components. The user interface 1003 may include a display, an input unit, for example, a keyboard, and may further include a standard wired interface and a wireless interface. The network interface 1004 may include a standard wired interface and a wireless interface (for example, a Wireless-Fidelity (WI-FI) interface). The memory 1005 may be high-speed random access memory (RAM), or stable non-volatile memory (NVM), for example, magnetic disk memory. The memory 1005 may be a storage device independent of the processor 1001.
[0021] As those skilled in the art will understand, the structure shown in Figure 1 does not constitute a limitation of the eye-line direction data acquisition device, and may include more or fewer components than shown, or may include combinations of several components or configurations with different component arrangements.
[0022] As shown in Figure 1, the memory 1005 of the storage medium may include an operating system, a network communication module, a user interface module, and a gaze direction data acquisition program.
[0023] In the gaze direction data acquisition device shown in Figure 1, the network interface 1004 is mainly used for data communication with a network server. The user interface 1003 is mainly used for data interaction with the user. The processor 1001 and memory 1005 in the gaze direction data acquisition device of this application may be provided in the gaze direction data acquisition device, and the gaze direction data acquisition device calls the gaze direction data acquisition program stored in the memory 1005 by the processor 1001 and executes the gaze direction data acquisition method according to the embodiment of this application.
[0024] The embodiments of this application provide a method for acquiring gaze direction data, and referring to Figure 2, Figure 2 is a flowchart of the first embodiment of the gaze direction data acquisition method of this application.
[0025] In this embodiment, the method for collecting line-of-sight data includes the following steps: Step S10: An object image of the target object, which is the object the user is focusing on, is collected by the first deep camera, and the three-dimensional coordinates of the target object on the first deep camera coordinate system are determined based on the object image.
[0026] It should be noted that the implementing body of this embodiment may be a computing service device having data processing, network communication, and program execution functions such as a tablet, personal computer, or mobile phone, or an electronic device capable of realizing the above functions, such as a gaze direction data acquisition device. Hereinafter, this embodiment and each of the following embodiments will be described using a gaze direction data acquisition device (referred to as the acquisition device) as an example.
[0027] To make it easier to understand, deep cameras used to collect object images of target objects are collectively referred to as the first deep camera, the number of first deep cameras may be one or more, the three-dimensional coordinates of an object may be the three-dimensional coordinates of the target object that the user is fixated on in the object image on the coordinate system of the first deep camera, the first deep camera can automatically acquire the three-dimensional coordinates of the target object in the image after capturing the object image, the number of target objects may be multiple, the target objects may be objects placed on a background plate, spots or other objects irradiated onto the background plate by a laser device, and this embodiment is not limited thereto.
[0028] In this embodiment, an object placed on the background plate for the user to focus on may be called a point of focus. The point of focus may be fixed to the background plate, and a deep image of the background plate can be collected by the first deep camera. Based on the deep image, the three-dimensional coordinates of the object at each point of focus on the background plate can be predetermined. When collecting line-of-sight data, the target point of focus that the user is focusing on can be determined, and based on the predetermined three-dimensional coordinates of the object, the three-dimensional coordinates of the object corresponding to the target point of focus can be determined. Conventional data collection techniques generally employ moving points of focus, requiring coordinate calibration for each collected picture, which is time-consuming. In this technique, by employing fixed points of focus, the coordinates of the points of focus in each picture collected by the first deep camera are fixed, requiring only one calibration of the point of focus coordinates. Subsequent data can be directly multiplexed, enabling automated coordinate marking and reducing the time required for data marking. In this embodiment, the point of fixation may be a moving point of fixation. When collecting line-of-sight direction data, the target point of fixation is collected by the first deep camera, widening the collected line-of-sight range and enriching the collected line-of-sight direction data, thereby improving the accuracy of subsequent deep learning algorithm training. Based on a specific scene, it is possible to determine whether the object placed on the background plate is fixed or moving, and this embodiment is not limited thereto.
[0029] Step S20: A facial image of the user is collected by a second deep camera, and the three-dimensional coordinates of the user's eyes on the second deep camera coordinate system are determined based on the facial image.
[0030] To make it easier to understand, deep cameras used to collect images of the user's face are collectively referred to as second deep cameras, and there may be one or more second deep cameras. The three-dimensional coordinates of the eye may be the three-dimensional coordinates of the eye in the second deep camera coordinate system in the user's face image when the user is fixated on a target object. The second deep camera can automatically acquire the three-dimensional coordinates of the eye in the image after capturing the user's face image, and the three-dimensional coordinates of the eye include the three-dimensional coordinates of the left eye and the three-dimensional coordinates of the right eye, and can be selected based on the specific scene, and this embodiment is not limited thereto.
[0031] Step S30: Determine the object coordinates and eye coordinates on the respective camera coordinate systems for the object and the eye.
[0032] To make it clear, the acquisition camera may include a red-green-blue (RGB) camera and an infrared camera, and may be a camera for acquiring user pictures. Determining the object coordinates and eye coordinates on each acquisition camera coordinate system of the object's three-dimensional coordinates and the eye's three-dimensional coordinates may be done as follows: The object's three-dimensional coordinates on the first deep camera coordinate system can be converted to each acquisition camera coordinate system, and the object coordinates on each acquisition camera coordinate system can be obtained; the eye's three-dimensional coordinates on the second deep camera coordinate system can be converted to each acquisition camera coordinate system, and the eye coordinates on each acquisition camera coordinate system can be obtained after the above coordinate conversion, and the object coordinates and eye coordinates on the acquisition camera coordinate system corresponding to the user picture acquired by each acquisition camera can be obtained, and the user picture acquired by the acquisition camera may be a picture of the user's face.
[0033] Step S40: Based on the object coordinates and eye coordinates in each acquired camera coordinate system, the user's line of sight direction data is determined.
[0034] To make it easier to understand, determining the user's gaze direction data based on the object coordinates and eye coordinates in each camera coordinate system may also mean determining the gaze direction data corresponding to the user picture captured by each camera based on the object coordinates and eye coordinates in each camera coordinate system. For example, if the number of cameras is N, gaze direction data corresponding to N user pictures can be obtained.
[0035] In concrete implementation, the acquisition device acquires an object image of the target object that the user is fixated on using a first deep camera, and determines the three-dimensional coordinates of the target object under the coordinate system of the first deep camera based on the target object image in the object image. A face image of the user when they are fixated on the target object is acquired using a second deep camera, and determines the three-dimensional coordinates of the eyes on the second deep camera coordinate system based on the eye image in the face image. When acquiring the three-dimensional coordinates of the object and the three-dimensional coordinates of the eyes using the deep cameras, the user's face picture is further acquired using multiple acquisition cameras, thereby acquiring multiple user face pictures. The three-dimensional coordinates of the object on the first deep camera coordinate system are converted to the coordinate systems of each acquisition camera, thereby acquiring multiple object coordinates. The three-dimensional coordinates of the eyes on the second deep camera coordinate system are converted to the coordinate systems of each acquisition camera, thereby acquiring multiple eye coordinates. Based on the object coordinates and eye coordinates on each acquisition camera coordinate system, gaze direction data corresponding to each user face picture is determined.
[0036] For example, referring to Figure 3, which is a schematic diagram for collecting line-of-sight data, the number of first and second deep cameras is set to 1 and is represented as A and B respectively, the first deep camera A is located behind the user, the second deep camera B is located in front of the user, multiple objects for the user to gaze at are placed on the background wall, the user is in front of the background wall, and N collecting cameras are placed between the user and the background wall, the collecting cameras include RGB cameras and infrared cameras, are located in front of the user (for example, within a range of 0.5m to 1m in front of the user), are distributed within a preset angular range in front of the user (for example, distributed within a range of ±45 degrees around the user in front of the user), the first deep camera A and the background wall maintain a first preset distance (for example, the first preset distance is set to 1.5m to 3m), and the second deep camera B and the user maintain a second preset distance (for example, the second preset distance is set to 0.5m to 1m). A deep image of the background wall is captured by a first deep camera, and based on the captured deep image, the three-dimensional coordinates of each object on the background wall in the first deep camera coordinate system can be determined. Later, the three-dimensional coordinates corresponding to each object can be directly superimposed, saving data marking time and further improving the efficiency of line-of-sight data acquisition. If the object that the user is fixated on is object 1 on the background wall, the three-dimensional coordinates of object 1 in the first deep camera coordinate system are obtained, an image of the user's face is captured by a second deep camera B, and based on the image of the eyes in the face image, the three-dimensional coordinates of the eyes in the second deep camera coordinate system are determined when the user is fixated on object 1. When images are collected using a deep camera, N user face images are further collected using N acquisition cameras. The three-dimensional coordinates of the object and eye in the deep camera coordinate system are converted to the coordinate systems of the N acquisition cameras. The object and eye coordinates in the N acquisition camera coordinate systems are then obtained, and based on the object and eye coordinates in each acquisition camera coordinate system, gaze direction data corresponding to the N user face images is determined.
[0037] In one embodiment, in order to improve the accuracy of the collected line-of-sight data, step S30 includes, when the first deep camera coordinate system and the second deep camera coordinate system are unified, the steps of determining the object coordinates on each collected camera coordinate system based on an external parameter matrix from the first deep camera coordinate system to each collected camera coordinate system matrix, and determining the eye coordinates on each collected camera coordinate system based on an external parameter matrix from the second deep camera coordinate system to each collected camera coordinate system matrix.
[0038] To make it clear, the external parameter matrices from the first and second deep camera coordinate systems to each acquired camera coordinate system matrix may be obtained by external parameter calibration, and the parameter matrices from the first and second deep camera coordinate systems to each acquired camera coordinate system matrix are different.
[0039] In one embodiment, in order to improve the efficiency of collecting line-of-sight direction data, step S40 includes the steps of determining a line-of-sight vector in each collection camera coordinate system based on object coordinates and eye coordinates in each collection camera coordinate system, and determining the user's line-of-sight direction data based on the line-of-sight vector in each collection camera coordinate system.
[0040] In a concrete implementation, for example, external parameter calibration is performed between the acquisition camera and the second deep camera to obtain an external parameter matrix (R2, T2) from the second deep camera coordinate system to each acquisition camera coordinate system matrix. External parameter calibration is performed between the acquisition camera and the first deep camera to obtain an external parameter matrix (R1, T1) from the first deep camera coordinate system to each acquisition camera coordinate system matrix. By capturing a deep image of the background wall with the first deep camera A, the three-dimensional coordinates of all objects on the background wall in the first deep camera coordinate system are obtained. When the user gazes at object 1 on the background wall, the three-dimensional coordinates of object 1 in the first deep camera coordinate system may be represented as deep_cam_point1. A deep image of the user's face is captured by the second deep camera B, and the three-dimensional coordinates of the user's left eye in the second deep camera coordinate system: deep_cam_left_eye1 are obtained. Based on Equation 1, the three-dimensional coordinates in the deep camera coordinate system can be converted to coordinates in the acquisition camera coordinate system.
[0041]
number
[0042] In this embodiment, an object image of a target object that the user is fixated on is collected by a first deep camera, and the three-dimensional coordinates of the target object in the first deep camera coordinate system are determined based on the object image. An image of the user's face is collected by a second deep camera, and the three-dimensional coordinates of the user's eyes in the second deep camera coordinate system are determined based on the face image. The object coordinates and eye coordinates in the respective camera coordinate systems of the object and eye coordinates are determined, and the user's gaze direction data is determined based on the object coordinates and eye coordinates in the respective camera coordinate systems of the object and eye. In this embodiment, an object image of the target object and an image of the user's face are collected by a first and second deep camera, respectively. Based on the object image and face image, the three-dimensional coordinates of the object and the three-dimensional coordinates of the eye are determined. These three-dimensional coordinates are converted to the coordinate systems of each collecting camera to obtain the object coordinates and eye coordinates. Based on the object coordinates and eye coordinates in each collecting camera coordinate system, gaze direction data is determined. By increasing the range of gaze collected, the gaze direction data becomes more accurate, and multiple gaze direction data can be collected each time, improving the efficiency of gaze direction data collection.
[0043] Referring to Figure 4, Figure 4 is a flowchart of a second embodiment of the gaze direction data acquisition method of this application.
[0044] Based on the first embodiment described above, in this embodiment, step S30 includes the following steps.
[0045] Step S301: If the first deep camera coordinate system and the second deep camera coordinate system are not unified, the object calibration three-dimensional coordinates on the first preset calibration plate coordinate system are determined based on the external parameter matrix from the first deep camera coordinate system to the first preset calibration plate coordinate system.
[0046] To make it clear, if the first deep camera coordinate system and the second deep camera coordinate system are not unified, directly converting the three-dimensional coordinates on the deep camera coordinate system to the coordinate system of each acquisition camera will result in inaccurate line-of-sight data being acquired. The first pre-configured calibration plate coordinate system may be the coordinate system corresponding to the first pre-configured calibration plate.
[0047] Step S302: Based on the external parameter matrix from the second deep camera coordinate system to the first preset calibration plate coordinate system, the eye region calibration three-dimensional coordinates on the first preset calibration plate coordinate system are determined for the eye region three-dimensional coordinates.
[0048] Step S303: Based on the external parameter matrix from the first preset calibration plate coordinate system to each acquisition camera coordinate system matrix, the object coordinates and eye coordinates on each acquisition camera coordinate system are determined for the object calibration three-dimensional coordinates and the eye calibration three-dimensional coordinates.
[0049] In the specific implementation, referring to Figure 5, which is a schematic diagram of a first pre-set calibration plate being photographed by a camera, the pre-set calibration plate is a grid calibration plate, the grid calibration plate is photographed by a deep camera and each acquisition camera, external parameter calibration is performed between the deep camera and the grid calibration plate, and an external parameter matrix is obtained from the deep camera coordinate system to the grid calibration plate coordinate system. External parameter calibration is performed between each acquisition camera and the grid calibration plate, and an external parameter matrix is obtained from each acquisition camera coordinate system to the grid calibration plate coordinate system. The three-dimensional coordinates of the object on the first deep camera coordinate system are converted to the grid calibration plate coordinate system using Equation 2, and the three-dimensional calibration coordinates of the object on the grid calibration plate coordinate system are obtained. If the three-dimensional coordinates of the eye are the three-dimensional coordinates of the left eye, then the three-dimensional coordinates of the left eye on the second deep camera coordinate system are converted to the grid calibration plate coordinate system using Equation 2, and the three-dimensional calibration coordinates of the left eye on the grid calibration plate coordinate system are obtained. Equation 1 converts the object calibration 3D coordinates on the grid calibration plate coordinate system to the coordinates of each acquisition camera, and obtains the object coordinates in each acquisition camera coordinate system. Equation 1 converts the left eye 3D coordinates on the grid calibration plate coordinate system to the coordinate systems of each acquisition camera, and obtains the left eye 3D coordinates in each acquisition camera coordinate system. Based on the object coordinates and left eye 3D coordinates in each acquisition camera coordinate system, the left eye line of sight direction vector can be determined, and since the process for determining the right eye line of sight direction vector is similar to that for the left eye line of sight direction vector, this embodiment will not be explained further here.
[0050]
number
[0051] In this embodiment, if the first deep camera coordinate system and the second deep camera coordinate system are not unified, the three-dimensional coordinates of the object on the first deep camera coordinate system are converted to the first pre-set calibration plate coordinate system, the three-dimensional coordinates of the eye on the second deep camera coordinate system are converted to the first pre-set calibration plate coordinate system, and the coordinates on the first pre-set calibration plate coordinate system are converted to each acquisition camera coordinate system to obtain the object coordinates and eye coordinates. If the deep camera coordinate systems are not unified, the deep camera coordinates are unified first, and then the line-of-sight direction data is determined, thereby improving the accuracy of the acquired line-of-sight direction data.
[0052] Referring to Figure 6, Figure 6 is a flowchart of a third embodiment of the gaze direction data acquisition method of this application.
[0053] Based on the above embodiments, in this embodiment, prior to step S30, the method further includes the following steps.
[0054] Step S01: Determine the first external parameter matrix from the first deep camera coordinate system to the second pre-configured calibration plate coordinate system, and determine the second external parameter matrix from the second deep camera coordinate system to the second pre-configured calibration plate coordinate system.
[0055] To make it clearer, the second pre-configured calibration plate coordinate system may be a coordinate system corresponding to a pre-configured second calibration plate, and the first extrinsic parameter matrix from the first deep camera coordinate system to the second pre-configured calibration plate coordinate system and the second extrinsic parameter matrix from the second deep camera coordinate system to the second pre-configured calibration plate coordinate system may be obtained by extrinsic parameter calibration.
[0056] Step S02: Determine the first three-dimensional object coordinates of the target reference object in the first deep camera coordinate system, and determine the second three-dimensional object coordinates of the target reference object in the second deep camera coordinate system.
[0057] To make it easier to understand, the reference object may be an object used to determine whether the first deep camera coordinate system and the second deep camera coordinate system are unified, the coordinates of the first three-dimensional object may be the coordinates of the reference object on the first deep camera coordinate system, and the coordinates of the second three-dimensional object may be the coordinates of the reference object on the second deep camera coordinate system.
[0058] Step S03: Based on the first three-dimensional object coordinates, the second three-dimensional object coordinates, the first external parameter matrix, and the second external parameter matrix, it is determined whether the first deep camera coordinate system and the second deep camera coordinate system are unified.
[0059] In the specific implementation, the first and second external parameter matrices are determined by external parameter calibration from the first and second deep cameras to a second pre-configured calibration plate coordinate system. The first and second three-dimensional object coordinates of the target reference object under the first and second deep camera coordinate systems are determined. Based on the first external parameter matrix, the first three-dimensional object coordinates are transformed to the second pre-configured calibration plate coordinate system, and based on the second external parameter matrix, the second three-dimensional object coordinates are transformed to the second pre-configured calibration plate coordinate system. Based on the coordinate system under the second pre-configured calibration plate, it is determined whether the first and second deep camera coordinate systems are unified.
[0060] In one embodiment, in order to determine whether the coordinate systems of the deep cameras are unified, step S02 includes the steps of: collecting a first reference object image of the target reference object with a first deep camera and determining the first three-dimensional object coordinates of the target reference object on the first deep camera coordinate system based on the first reference object image; and collecting a second reference object image of the target reference object with a second deep camera and determining the second three-dimensional object coordinates of the target reference object on the second deep camera coordinate system based on the second reference object image.
[0061] In one embodiment, in order to determine whether or not the deep camera coordinate systems are unified, step S03 includes: converting the first three-dimensional object coordinates to the first object coordinates on the second preset calibration plate coordinate system based on the first external parameter matrix; converting the second three-dimensional object coordinates to the second object coordinates on the second preset calibration plate coordinate system based on the second external parameter matrix; determining that the first deep camera coordinate system and the second deep camera coordinate system are unified if the first object coordinates and the second object coordinates match; and determining that the first deep camera coordinate system and the second deep camera coordinate system are not unified if the first object coordinates and the second object coordinates do not match.
[0062] To facilitate understanding, the coordinates of the first and second three-dimensional objects may be converted to the coordinates of the first and second objects using Equation 1.
[0063] In specific implementation, referring to Figures 7 and 8, Figure 7 is a schematic diagram of a second pre-set reference plate being photographed by a deep camera, and Figure 8 is a schematic diagram of a target reference object being photographed by a deep camera. For example, when verifying the unification of coordinate systems for the first and second deep cameras, a second pre-set calibration plate is photographed by the first and second deep cameras, and the first and second external parameter matrices are obtained from the first and second deep camera coordinate systems to the second pre-set calibration plate coordinate system by external parameter calibration. The target reference object is photographed by the first and second deep cameras, and the first reference object image and the second reference object image are obtained. Based on the first and second reference images, the first three-dimensional object coordinates on the first deep camera coordinate system and the second three-dimensional object coordinates on the second deep camera coordinate system are determined. Equation 1 and the first external parameter matrix are used to convert the object coordinates under the deep camera coordinate system to a second, pre-configured calibration plate coordinate system, thereby obtaining the first and second object coordinates. If the first and second object coordinates match, it is determined that the coordinate systems of the two deep cameras are unified; otherwise, it is determined that the coordinate systems of the two deep cameras are not unified.
[0064] In this embodiment, an external parameter matrix is determined from the first deep camera coordinate system and the second deep camera coordinate system to a second pre-configured calibration plate coordinate system. Based on the external parameter matrix, the three-dimensional object coordinates on the first deep camera coordinate system and the second deep camera coordinate system are transformed to the second pre-configured calibration plate coordinate system. If the two object coordinates on the second pre-configured calibration plate coordinate system coincide, it is determined that the two deep camera coordinate systems have been unified, thereby improving the accuracy of the unification determination of the deep camera coordinate systems and improving the accuracy of the line-of-sight data collected later.
[0065] Furthermore, the embodiment of this application proposes a storage medium in which a gaze direction data acquisition program is stored, and when the gaze direction data acquisition program is executed by a processor, the steps of the gaze direction data acquisition method described above are realized.
[0066] Referring to Figure 9, Figure 9 is a structural block diagram of a first embodiment of the line-of-sight data acquisition device of the present application.
[0067] As shown in Figure 9, the line-of-sight direction data acquisition device proposed in this embodiment of the present application is A first coordinate determination module 10 collects an object image of a target object, which is an object that the user is focusing on, using a first deep camera, and determines the three-dimensional coordinates of the target object on the first deep camera coordinate system based on the object image. A second coordinate determination module 20 collects the user's facial image using a second deep camera and determines the three-dimensional coordinates of the user's eye on the second deep camera coordinate system based on the facial image. A third coordinate determination module 30 for determining the object coordinates and eye coordinates on the respective camera coordinate systems for the three-dimensional coordinates of the object and the three-dimensional coordinates of the eye, It includes a direction data determination module 40 for determining the user's line of sight direction data based on the object coordinates and eye coordinates on each acquired camera coordinate system.
[0068] In this embodiment, an object image of a target object that the user is fixated on is collected by a first deep camera, and the three-dimensional coordinates of the target object in the first deep camera coordinate system are determined based on the object image. An image of the user's face is collected by a second deep camera, and the three-dimensional coordinates of the user's eyes in the second deep camera coordinate system are determined based on the face image. The object coordinates and eye coordinates in the respective camera coordinate systems of the object and eye coordinates are determined, and the user's gaze direction data is determined based on the object coordinates and eye coordinates in the respective camera coordinate systems of the object and eye. In this embodiment, an object image of the target object and an image of the user's face are collected by a first and second deep camera, respectively. Based on the object image and face image, the three-dimensional coordinates of the object and the three-dimensional coordinates of the eye are determined. These three-dimensional coordinates are converted to the coordinate systems of each collecting camera to obtain the object coordinates and eye coordinates. Based on the object coordinates and eye coordinates in each collecting camera coordinate system, gaze direction data is determined. By increasing the range of gaze collected, the gaze direction data becomes more accurate, and multiple gaze direction data can be collected each time, improving the efficiency of gaze direction data collection.
[0069] Based on the first embodiment of the gaze direction data acquisition device described in this application, a second embodiment of the gaze direction data acquisition device described in this application is proposed.
[0070] In this embodiment, the third coordinate determination module 30 is further used to determine the object coordinates on each acquisition camera coordinate system based on an external parameter matrix from the first deep camera coordinate system to each acquisition camera coordinate system matrix when the first deep camera coordinate system and the second deep camera coordinate system are unified, and to determine the eye coordinates on each acquisition camera coordinate system based on an external parameter matrix from the second deep camera coordinate system to each acquisition camera coordinate system matrix.
[0071] The third coordinate determination module 30 is further used to determine the object calibration three-dimensional coordinates on the first preset calibration plate coordinate system of the object three-dimensional coordinates, based on an external parameter matrix from the first deep camera coordinate system to the first preset calibration plate coordinate system, if the first deep camera coordinate system and the second deep camera coordinate system are not unified; to determine the eye calibration three-dimensional coordinates on the first preset calibration plate coordinate system of the eye three-dimensional coordinates, based on an external parameter matrix from the second deep camera coordinate system to the first preset calibration plate coordinate system; and to determine the object coordinates and eye coordinates on each acquisition camera coordinate system of the object calibration three-dimensional coordinates and the eye calibration three-dimensional coordinates, based on an external parameter matrix from the first preset calibration plate coordinate system to each acquisition camera coordinate system matrix.
[0072] The third coordinate determination module 30 is further used to determine a first external parameter matrix from the first deep camera coordinate system to a second pre-set calibration plate coordinate system, to determine a second external parameter matrix from the second deep camera coordinate system to a second pre-set calibration plate coordinate system, to determine the first three-dimensional object coordinates of the target reference object in the first deep camera coordinate system, to determine the second three-dimensional object coordinates of the target reference object in the second deep camera coordinate system, and to determine whether the first deep camera coordinate system and the second deep camera coordinate system are unified based on the first three-dimensional object coordinates, the second three-dimensional object coordinates, the first external parameter matrix, and the second external parameter matrix.
[0073] The third coordinate determination module 30 is further used to collect a first reference object image of the target reference object using a first deep camera, determine the first three-dimensional object coordinates of the target reference object on the first deep camera coordinate system based on the first reference object image, collect a second reference object image of the target reference object using a second deep camera, and determine the second three-dimensional object coordinates of the target reference object on the second deep camera coordinate system based on the second reference object image.
[0074] The third coordinate determination module 30 is used to further convert the first three-dimensional object coordinates to first object coordinates on the second preset calibration plate coordinate system based on the first external parameter matrix, and to convert the second three-dimensional object coordinates to second object coordinates on the second preset calibration plate coordinate system based on the second external parameter matrix, and to determine that the first deep camera coordinate system and the second deep camera coordinate system are unified if the first object coordinates and the second object coordinates coincide, and to determine that the first deep camera coordinate system and the second deep camera coordinate system are not unified if the first object coordinates and the second object coordinates do not coincide.
[0075] The direction data determination module 40 is further used to determine the line-of-sight vector in each acquisition camera coordinate system based on the object coordinates and eye coordinates in each acquisition camera coordinate system, and to determine the user's line-of-sight direction data based on the line-of-sight vector in each acquisition camera coordinate system.
[0076] Other embodiments or specific examples of the gaze direction data acquisition device of this application can be found by referring to the embodiments of each method described above and will not be described further here.
[0077] It should be explained that, in this specification, the terms “include,” “incorporate,” or any other variation thereof are intended to cover non-exclusive inclusion such that a process, method, article, or system containing a set of elements includes not only those elements but also other elements not expressly listed, or elements specific to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase “one…includes” does not preclude the presence of other identical elements in the process, method, article, or system containing that element.
[0078] The example numbers in this application are for illustrative purposes only and do not represent any indication of superiority or inferiority among the embodiments.
[0079] As will be clearly apparent to those skilled in the art from the above description of the embodiments, the methods of the above embodiments may be implemented by adding a general-purpose hardware platform necessary for the software, or of course, by hardware, and in many cases the former is a preferred embodiment. Based on this understanding, the parts of the invention of this application that essentially contribute to or to the prior art may be expressed in the form of a software product. This computer software product is stored on a storage medium (e.g., read-only memory / random access memory, magnetic disk, optical disk) and includes a number of instructions for causing a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to perform the method of each embodiment of this application.
[0080] The foregoing are merely selectable embodiments of this application and do not thereby limit the scope of this application. Any equivalent structure or equivalent flow transformation performed using the contents of this specification and drawings, or any application directly or indirectly to other related technical fields, is included within the scope of this application.
Claims
1. A method for collecting gaze direction data, The process involves collecting an object image of a target object, which is an object that the user is focusing on, using a first deep camera, and determining the three-dimensional coordinates of the target object on the first deep camera coordinate system based on the object image. The steps include: collecting an image of the user's face using a second deep camera, and determining the three-dimensional coordinates of the user's eyes on the second deep camera coordinate system based on the face image; The steps include determining the object coordinates and eye coordinates on the respective camera coordinate systems for the three-dimensional coordinates of the object and the three-dimensional coordinates of the eye, The steps include determining the user's line of sight direction data based on the object coordinates and eye coordinates on each acquired camera coordinate system, The step of determining the user's line of sight direction data based on the object coordinates and eye coordinates on each acquired camera coordinate system is as follows: The steps include determining the line-of-sight vector in each camera coordinate system based on the object coordinates and eye coordinates in each camera coordinate system, The process includes the step of determining the user's line of sight direction data based on the line of sight vector in each acquired camera coordinate system. Method for collecting gaze direction data.
2. The step of determining the object coordinates and eye coordinates on the respective camera coordinate systems for the three-dimensional coordinates of the object and the three-dimensional coordinates of the eye is as follows: When the first deep camera coordinate system and the second deep camera coordinate system are unified, the steps include determining the object coordinates on each acquisition camera coordinate system based on the external parameter matrix from the first deep camera coordinate system to each acquisition camera coordinate system matrix, and The method according to claim 1, comprising the step of determining the eye coordinates on each acquisition camera coordinate system of the eye three-dimensional coordinates based on an external parameter matrix from the second deep camera coordinate system to each acquisition camera coordinate system matrix.
3. The step of determining the object coordinates and eye coordinates on the respective camera coordinate systems for the three-dimensional coordinates of the object and the three-dimensional coordinates of the eye is as follows: If the first deep camera coordinate system and the second deep camera coordinate system are not unified, the steps include determining the object calibration three-dimensional coordinates of the object three-dimensional coordinates on the first preset calibration plate coordinate system based on an external parameter matrix from the first deep camera coordinate system to a first preset calibration plate coordinate system, A step of determining the eye calibration three-dimensional coordinates of the eye on the first preset calibration plate coordinate system based on an external parameter matrix from the second deep camera coordinate system to the first preset calibration plate coordinate system, The method according to claim 1, further comprising the step of determining the object coordinates and eye coordinates on each acquisition camera coordinate system of the object calibration three-dimensional coordinates and the eye calibration three-dimensional coordinates based on an external parameter matrix from the first preset calibration plate coordinate system to each acquisition camera coordinate system matrix.
4. Before the step of determining the object coordinates and eye coordinates on the respective camera coordinate systems for the three-dimensional coordinates of the object and the three-dimensional coordinates of the eye, The steps include determining a first external parameter matrix from a first deep camera coordinate system to a second pre-configured calibration plate coordinate system, and determining a second external parameter matrix from a second deep camera coordinate system to a second pre-configured calibration plate coordinate system, The steps include determining the first three-dimensional object coordinates of the target reference object on the first deep camera coordinate system, and determining the second three-dimensional object coordinates of the target reference object on the second deep camera coordinate system, The method according to claim 1, further comprising the step of determining whether the first deep camera coordinate system and the second deep camera coordinate system are unified based on the first three-dimensional object coordinates, the second three-dimensional object coordinates, the first external parameter matrix, and the second external parameter matrix.
5. The steps of determining the first three-dimensional object coordinates of the target reference object on the first deep camera coordinate system and determining the second three-dimensional object coordinates of the target reference object on the second deep camera coordinate system are as follows: The steps include: collecting a first reference object image of the target reference object using a first deep camera, and determining the first three-dimensional object coordinates of the target reference object on the first deep camera coordinate system based on the first reference object image; The method according to claim 4, comprising the steps of: collecting a second reference object image of a target reference object using a second deep camera; and determining the second three-dimensional object coordinates of the target reference object on the second deep camera coordinate system based on the second reference object image.
6. The step of determining whether the first deep camera coordinate system and the second deep camera coordinate system are unified based on the first three-dimensional object coordinates, the second three-dimensional object coordinates, the first external parameter matrix, and the second external parameter matrix is: A step of converting the first three-dimensional object coordinates to the first object coordinates on the second preset calibration plate coordinate system based on the first external parameter matrix, A step of converting the second three-dimensional object coordinates to the second object coordinates on the second pre-set calibration plate coordinate system based on the second external parameter matrix, If the coordinates of the first object and the coordinates of the second object coincide, the first deep camera coordinate system and the second deep camera coordinate system are determined to be unified. The method according to claim 5, further comprising the step of determining that the first deep camera coordinate system and the second deep camera coordinate system are not unified if the first object coordinate system and the second object coordinate system do not coincide.
7. A gaze direction data acquisition device, A first coordinate determination module for collecting an object image of a target object, which is an object that the user is focusing on, using a first deep camera, and determining the three-dimensional coordinates of the target object on the first deep camera coordinate system based on the object image, A second coordinate determination module for collecting the user's facial image using a second deep camera and determining the three-dimensional coordinates of the user's eye on the second deep camera coordinate system based on the facial image, A third coordinate determination module for determining the object coordinates and eye coordinates on the respective camera coordinate systems for the three-dimensional coordinates of the object and the three-dimensional coordinates of the eye, Includes a direction data determination module for determining the user's line of sight direction data based on object coordinates and eye coordinates on each acquired camera coordinate system, Determining the user's line of sight direction data based on the object coordinates and eye coordinates in each acquired camera coordinate system is: Based on the object coordinates and eye coordinates in each camera coordinate system, the line-of-sight vector in each camera coordinate system is determined. This includes determining the user's line of sight direction data based on the line of sight vector in each acquired camera coordinate system, Eye-line direction data acquisition device.
8. Memory and Processor and A gaze direction data acquisition device comprising: a gaze direction data acquisition program stored in the memory and executable on the processor, The gaze direction data acquisition device is configured to implement the steps of the gaze direction data acquisition method described in any one of claims 1 to 6, wherein the gaze direction data acquisition program is configured to implement the steps of the gaze direction data acquisition method described in any one of claims 1 to 6.
9. A storage medium in which a gaze direction data acquisition program is stored, wherein when the gaze direction data acquisition program is executed by a processor, the steps of the gaze direction data acquisition method described in any one of claims 1 to 6 are realized.
Citation Information
Patent Citations
Gaze direction feature collection method and device, computer equipment and storage medium
CN113553920A
System and method for realizing sight line estimation and attention analysis based on recursive convolutional neural network
CN114387679A
eye tracking system
JP2018512665A