Sight calibration method and device, equipment, computer-readable storage medium, system, vehicle
By obtaining the three-dimensional position of the user's eyes and gaze point, calibrating the gaze direction in combination with camera parameters, and using small sample learning to optimize the gaze tracing model, the problem of inaccurate gaze estimation is solved and the user experience is improved.
Patent Information
- Application Number
- CN202180001805.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-06-28
AI Technical Summary
In the prior art, the line of sight tracing model causes inaccurate line of sight estimation due to individual differences and camera installation errors, which affects the user experience of applications such as line of sight interaction in the smart cockpit.
By obtaining the three-dimensional position of the user's eyes and the three-dimensional position of the gaze point, combining the external and internal references of the camera, calibrating the gaze direction, and using a small sample learning method to optimize the gaze tracing model to improve the accuracy of gaze estimation.
It improves the accuracy of line of sight data, improves the user experience of line of sight interaction and other applications in the smart cockpit, and solves the problems of low accuracy and optimization of the line of sight tracing model on individual users.
Smart Images

Figure CN113661495B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent driving, and particularly to a method and device for gaze calibration, a device, a computer-readable storage medium, a system, and a vehicle. Background Art
[0002] Gaze tracking is an important support for upper-layer applications such as distraction detection, takeover level estimation, and gaze interaction in an intelligent cockpit. Due to the differences in the external features of people's eyes and the internal structures of the eyeballs, it is usually impossible to train a gaze tracking model that is accurate for "everyone". At the same time, due to reasons such as camera installation errors, there will be a certain loss of accuracy in the gaze angle directly output by the gaze tracking model, resulting in inaccurate gaze estimation. If the error in gaze estimation can be corrected, the user experience of upper-layer applications based on gaze tracking can be effectively improved. Summary of the Invention
[0003] In view of the above problems in the related art, this application provides a method and device for gaze calibration, a device, a computer-readable storage medium, a system, and a vehicle, which can effectively improve the accuracy of gaze estimation for specific users.
[0004] To achieve the above object, a first aspect of this application provides a method for gaze calibration, including: obtaining the three-dimensional position of the user's eyes and the first gaze direction according to the first image containing the user's eyes collected by the first camera; obtaining the fixation area of the user in the second image according to the three-dimensional position of the eyes, the first gaze direction, the external parameters of the first camera, and the external and internal parameters of the second camera, where the second image is collected by the second camera and contains the external scene seen by the user; obtaining the position of the fixation point of the user in the second image according to the fixation area of the user in the second image and the second image; obtaining the three-dimensional position of the fixation point of the user according to the position of the fixation point and the internal parameters of the second camera; and obtaining the second gaze direction of the user according to the three-dimensional position of the fixation point and the three-dimensional position of the eyes, where the second gaze direction is used as the calibrated gaze direction.
[0005] Thus, the gaze direction of the user can be calibrated using the second image to obtain a second gaze direction with higher accuracy, effectively improving the accuracy of the user's gaze data, and further improving the user experience of upper-layer applications based on gaze tracking.
[0006] As a possible implementation manner of the first aspect, the first gaze direction is extracted from the first image based on a gaze tracking model.
[0007] Thus, the initial gaze direction of the user can be efficiently obtained.
[0008] As a possible implementation of the first aspect, based on the three-dimensional position of the eyes, the first line-of-sight direction, the external parameters of the first camera, and the external and internal parameters of the second camera, the fixation area of the user in the second image is obtained, including: obtaining the fixation area of the user in the second image according to the three-dimensional position of the eyes, the first line-of-sight direction, the external parameters of the first camera, the external and internal parameters of the second camera, and the accuracy of the line-of-sight tracking model.
[0009] Thus, the error caused by the accuracy limitation of the line-of-sight tracking model can be eliminated in the finally obtained second line-of-sight direction.
[0010] As a possible implementation of the first aspect, it further includes: using the second line-of-sight direction of the user and the first image as the optimization samples of the user, and optimizing the line-of-sight tracking model based on the small-sample learning method.
[0011] Thus, the line-of-sight estimation accuracy of the line-of-sight tracking model for a specific user can be continuously improved with a small number of samples and small-scale training, and then a user-level line-of-sight tracking model can be obtained.
[0012] As a possible implementation of the first aspect, it further includes: screening the fixation point or the second line-of-sight direction according to the confidence of the fixation point of the user in the second image.
[0013] Thus, the amount of computation can be reduced, and the processing efficiency and the accuracy of line-of-sight calibration can be improved.
[0014] As a possible implementation of the first aspect, the position of the fixation point of the user in the second image is obtained by using the fixation point calibration model according to the fixation area of the user in the second image and the second image.
[0015] Thus, the fixation point of the user in the second image can be obtained efficiently, accurately, and stably.
[0016] As a possible implementation of the first aspect, the fixation point calibration model simultaneously provides the probability value of the fixation point of the user in the second image, and the confidence is determined by the probability value.
[0017] Thus, the data provided by the fixation point calibration model can be fully utilized to improve the processing efficiency.
[0018] The second aspect of the present application provides a line-of-sight calibration method, including:
[0019] In response to the user's fixation operation on the reference point in the display screen, obtaining the three-dimensional position of the user's fixation point;
[0020] According to the first image collected by the first camera and including the user's eyes, obtaining the three-dimensional position of the user's eyes;
[0021] According to the three-dimensional position of the fixation point and the three-dimensional position of the eyes, obtaining the second line-of-sight direction of the user.
[0022] Thus, the accuracy of the user's line-of-sight data can be effectively improved, thereby enhancing the user experience of upper-layer applications based on line-of-sight tracking.
[0023] As a possible implementation of the second aspect, the display screen is an augmented reality head-up display.
[0024] Thus, the line-of-sight calibration of the driver can be achieved without affecting the safe driving of the driver.
[0025] As a possible implementation of the second aspect, the method further includes: using the second line-of-sight direction and the first image of the user as the optimization sample of the user, and optimizing the line-of-sight tracking model based on the small-sample learning method.
[0026] Thus, the line-of-sight estimation accuracy of the line-of-sight tracking model for a specific user can be continuously improved with a small number of samples and small-scale training, and then a user-level line-of-sight tracking model can be obtained.
[0027] The third aspect of the present application provides a line-of-sight calibration device, including:
[0028] An eye position determination unit configured to obtain the three-dimensional position of the user's eyes according to the first image including the user's eyes collected by the first camera;
[0029] A first line-of-sight determination unit configured to obtain the first line-of-sight direction of the user according to the first image including the user's eyes collected by the first camera;
[0030] A fixation area unit configured to obtain the fixation area of the user in the second image according to the three-dimensional position of the eyes, the first line-of-sight direction, the external parameters of the first camera, and the external and internal parameters of the second camera, where the second image is collected by the second camera and includes the external scene seen by the user;
[0031] A fixation point calibration unit configured to obtain the position of the fixation point of the user in the second image according to the fixation area of the user in the second image and the second image;
[0032] A fixation point conversion unit configured to obtain the three-dimensional position of the user's fixation point according to the position of the fixation point and the internal parameters of the second camera;
[0033] A second line-of-sight determination unit configured to obtain the second line-of-sight direction of the user according to the three-dimensional position of the fixation point and the three-dimensional position of the eyes, and the second line-of-sight direction is used as the calibrated line-of-sight direction.
[0034] Thus, the line-of-sight direction of the user can be calibrated using the second image to obtain a second line-of-sight direction with higher accuracy, effectively improving the accuracy of the user's line-of-sight data, and further enhancing the user experience of upper-layer applications based on line-of-sight tracking.
[0035] As a possible implementation of the third aspect, the first line-of-sight direction is extracted from the first image based on a line-of-sight tracking model.
[0036] Thus, the initial line-of-sight direction of the user can be efficiently obtained.
[0037] As a possible implementation of the third aspect, the fixation area unit is configured to obtain the fixation area of the user in the second image according to the three-dimensional position of the eyes, the first line-of-sight direction, the external parameters of the first camera, the external and internal parameters of the second camera, and the accuracy of the line-of-sight tracking model.
[0038] Thus, the error caused by the accuracy limitation of the line-of-sight tracking model can be eliminated in the finally obtained second line-of-sight direction.
[0039] As a possible implementation of the third aspect, it further includes: an optimization unit configured to use the second line-of-sight direction of the user and the first image as the optimization samples of the user, and optimize the line-of-sight tracking model based on the small-sample learning method.
[0040] Thus, the line-of-sight estimation accuracy of the line-of-sight tracking model for a specific user can be continuously improved with a small number of samples and small-scale training, and then a user-level line-of-sight tracking model can be obtained.
[0041] As a possible implementation of the third aspect, the fixation point calibration unit is further configured to screen the fixation points according to the confidence of the fixation points of the user in the second image; and / or, the optimization unit is further configured to screen the second line-of-sight direction according to the confidence of the fixation points of the user in the second image.
[0042] Thus, the amount of computation can be reduced, and the processing efficiency and the accuracy of line-of-sight calibration can be improved.
[0043] As a possible implementation of the third aspect, the position of the fixation point of the user in the second image is obtained by using a fixation point calibration model according to the fixation area of the user in the second image and the second image.
[0044] Thus, the fixation point of the user in the second image can be efficiently, accurately and stably obtained.
[0045] As a possible implementation of the third aspect, the fixation point calibration model simultaneously provides the probability value of the fixation point of the user in the second image, and the confidence is determined by the probability value.
[0046] Thus, the data provided by the fixation point calibration model can be fully utilized to improve the processing efficiency.
[0047] The fourth aspect of the present application provides a line-of-sight calibration device, including:
[0048] A fixation point position determination unit, configured to obtain the three-dimensional position of the user's fixation point in response to the user's fixation operation on a reference point in the display screen;
[0049] An eye position determination unit, configured to obtain the three-dimensional position of the user's eyes according to a first image including the user's eyes collected by a first camera;
[0050] A second line-of-sight determination unit, configured to obtain the second line-of-sight direction of the user according to the three-dimensional position of the fixation point and the three-dimensional position of the eyes.
[0051] Thus, the accuracy of the user's line-of-sight data can be effectively improved, and further the user experience of the upper-layer application based on line-of-sight tracking can be improved.
[0052] As a possible implementation manner of the fourth aspect, the display screen is the display screen of an augmented reality head-up display system.
[0053] Thus, its line-of-sight calibration can be achieved without affecting the safe driving of the driver.
[0054] As a possible implementation manner of the fourth aspect, the device further includes: an optimization unit, configured to use the second line-of-sight direction of the user and the first image as the user's optimization samples, and optimize the line-of-sight tracking model based on the small-sample learning method.
[0055] Thus, the line-of-sight estimation accuracy of the line-of-sight tracking model for a specific user can be continuously improved with a small number of samples and small-scale training, and further a user-level line-of-sight tracking model can be obtained.
[0056] The fifth aspect of the present application provides a computing device, including:
[0057] At least one processor; and,
[0058] At least one memory, which stores program instructions, and when the program instructions are executed by the at least one processor, the at least one processor is caused to execute the above-mentioned line-of-sight calibration method.
[0059] The sixth aspect of the present application provides a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a computer, the computer is caused to execute the above-mentioned line-of-sight calibration method.
[0060] The seventh aspect of the present application provides a driver monitoring system, including:
[0061] A first camera, configured to collect a first image including the user's eyes;
[0062] A second camera, configured to collect a second image including the scene outside the vehicle seen by the user;
[0063] At least one processor; and,
[0064] At least one memory storing program instructions which, when executed by at least one processor, cause the at least one processor to execute the line-of-sight calibration method of the first aspect above.
[0065] Thereby, the accuracy of line-of-sight estimation of users such as drivers in the vehicle cockpit scenario can be effectively improved, and further, the user experience of the driver monitoring system and the user experience of upper-layer applications such as distraction detection, takeover level estimation, and line-of-sight interaction in the intelligent cockpit can be improved.
[0066] As a possible implementation of the seventh aspect, the driver monitoring system further includes: a display screen configured to display a reference point to the user; and program instructions which, when executed by at least one processor, cause the at least one processor to execute the line-of-sight calibration method of the second aspect.
[0067] The eighth aspect of the present application provides a vehicle including the above-mentioned driver monitoring system.
[0068] Thereby, the accuracy of line-of-sight estimation of users such as drivers in the vehicle cockpit scenario can be effectively improved, and further, the user experience of upper-layer applications such as distraction detection, takeover level estimation, and line-of-sight interaction in the vehicle cockpit can be improved, and finally, the safety of vehicle intelligent driving can be improved.
[0069] In the embodiments of the present application, the three-dimensional position of the user's eyes is obtained from the first image including the user's eyes, and the three-dimensional position of the user's fixation point is obtained from the calibration position on the display screen or the second image including the external scene seen by the user, and then a second line-of-sight direction with higher accuracy is obtained, effectively improving the accuracy of user line-of-sight estimation and being applicable to the cockpit scenario. In addition, the second line-of-sight direction and the first image can also be used as personalized samples of the user to optimize the line-of-sight tracking model. Thereby, a line-of-sight tracking model for a specific user can be obtained, thus solving the problems of difficult optimization of the line-of-sight tracking model and low line-of-sight estimation accuracy for some users.
[0070] These and other aspects of the present invention will become more readily apparent in the following description of the (multiple) embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The following further illustrates the various features of the present invention and the relationships between the various features with reference to the drawings. The drawings are all exemplary, some features are not shown in actual proportion, and in some drawings, the conventional and non-essential features in the field related to the present application may be omitted, or non-essential features for the present application may be additionally shown. The combination of the various features shown in the drawings is not intended to limit the present application. In addition, throughout the present specification, the content referred to by the same reference numerals is also the same. The specific description of the drawings is as follows:
[0072] Figure 1It is a schematic diagram of the exemplary architecture of the system in an embodiment of the present application.
[0073] Figure 2 It is a schematic diagram of the installation position of the sensor in an embodiment of the present application.
[0074] Figure 3 It is a schematic flowchart of the line-of-sight calibration method in an embodiment of the present application.
[0075] Figure 4 It is an example diagram of the eye reference point in an embodiment of the present application.
[0076] Figure 5 It is a schematic flowchart of the three-dimensional eye position estimation in an embodiment of the present application.
[0077] Figure 6 It is an example diagram of the cockpit scene applicable to the embodiment of the present application.
[0078] Figure 7 It is Figure 6 A schematic diagram of the fixation area in the reference coordinate system of the scene.
[0079] Figure 8 It is Figure 6 A schematic diagram of the fixation area in the second image of the scene.
[0080] Figure 9 It is a schematic flowchart of determining the fixation area of the user in the second image in an embodiment of the present application.
[0081] Figure 10 It is an example diagram of the projection between the fixation area in the reference coordinate system and the fixation area in the second image.
[0082] Figure 11 It is a schematic diagram of the structure of the fixation point calibration model in an embodiment of the present application.
[0083] Figure 12 It is a schematic flowchart of obtaining the three-dimensional position of the fixation point in an embodiment of the present application.
[0084] Figure 13 It is an exemplary flowchart of optimizing the eye-tracking model in an embodiment of the present application.
[0085] Figure 14 It is a schematic diagram of the process of line-of-sight calibration and model optimization of the driver in the cockpit scene.
[0086] Figure 15 It is a schematic diagram of the structure of the line-of-sight calibration device in an embodiment of the present application.
[0087] Figure 16 It is a schematic diagram of the exemplary architecture of the system in another embodiment of the present application.
[0088] Figure 17 It is a schematic flowchart of the line-of-sight calibration method in another embodiment of the present application.
[0089] Figure 18 It is a schematic structural diagram of the line-of-sight calibration device in another embodiment of the present application.
[0090] Figure 19 It is a schematic structural diagram of the computing device in an embodiment of the present application. Detailed implementation manners
[0091] Terms such as "first, second, third, etc." or similar terms like module A, module B, module C, etc. in the description and claims are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that, where permitted, the specific order or sequence can be interchanged so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0092] In the following descriptions, the reference numerals of the steps involved, such as S110, S120... etc., do not necessarily mean that the steps will be executed in this order. Where permitted, the order of the front and rear steps can be interchanged, or they can be executed simultaneously.
[0093] The term "comprising" used in the description and claims should not be construed as being limited to the content listed thereafter; it does not exclude other elements or steps. Therefore, it should be interpreted as specifying the existence of the mentioned features, wholes, steps or components, but does not exclude the existence or addition of one or more other features, wholes, steps or components and their groups. Therefore, the expression "a device comprising device A and B" should not be limited to a device consisting only of components A and B.
[0094] "An embodiment" or "embodiments" mentioned in this specification means that the specific features, structures or characteristics described in connection with the embodiment are included in at least one embodiment of the present invention. Therefore, the phrases "in an embodiment" or "in embodiments" that appear throughout this specification do not necessarily refer to the same embodiment, but may refer to the same embodiment. In addition, in one or more embodiments, the various specific features, structures or characteristics can be combined in any suitable manner, as will be apparent to those of ordinary skill in the art from this disclosure.
[0095] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. In case of inconsistency, the meaning stated in this specification or the meaning derived from the content recorded in this specification shall prevail. In addition, the terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0096] For the purpose of accurately describing the technical content in this application and for accurately understanding the present invention, the following explanations or definitions of the terms used in this specification are given before describing the specific embodiments.
[0097] Eye tracking / gaze tracking is a technology for measuring the direction of a person's eye gaze or the fixation point.
[0098] An eye tracking / gaze tracking model is a machine learning model that can estimate the direction of a person's eye gaze or the fixation point through an image containing the human eye or face. For example, a neural network model, etc.
[0099] A Driver Monitoring System (DMS) is a system that monitors the status of the driver in the vehicle based on image processing technology, speech processing technology, etc. It includes components such as an in-vehicle camera, a processor, a fill light, etc. installed in the vehicle cockpit. The in-vehicle camera can capture images containing the driver's face, head, and part of the torso (e.g., arms) (i.e., the DMS images in this article).
[0100] An out-of-vehicle camera, also known as a front camera, is used to capture images containing the out-of-vehicle scene (especially the scene in front of the vehicle), and the out-of-vehicle scene seen by the driver is included in this image.
[0101] A Red Green Blue (RGB) camera forms a color image of an object by sensing natural light or near-infrared light reflected by the object.
[0102] A Time of Flight (TOF) camera emits light pulses towards the target object, simultaneously records the reflection movement time of the light pulses, calculates the distance between the light pulse emitter and the target object, and generates a 3D image of the target object based on this. This 3D image includes the depth information of the target object and the information of the reflected light intensity.
[0103] PnP (Perspective-n-Point) is a problem of calculating the projection relationship through N feature points in the world coordinate system and N image points in the image coordinate system, so as to obtain the pose of the camera or the object. PnP solution means: given the matching point pairs of n 3D reference points {c1, c2,..., Cn} to the 2D projection points {u1, u2,..., un} on the camera image, knowing the coordinates of the 3D reference points in the world coordinate system, the coordinates of the 2D points in the image coordinate system, and knowing the internal parameters K of the camera, to find the pose transformation {R|t} between the world coordinate system and the camera coordinate system, where R is the rotation matrix and t represents the translation variable.
[0104] The Landmark algorithm is a kind of human face feature point extraction technology.
[0105] The world coordinate system, also known as the measurement coordinate system and the objective coordinate system, is a three-dimensional rectangular coordinate system. Based on it, the three-dimensional positions of the camera and the object to be measured can be described. It is the absolute coordinate system of the objective three-dimensional world and is usually represented by Pw(Xw, Yw, Zw) for its coordinate values.
[0106] The camera coordinate system is a three-dimensional rectangular coordinate system with the optical center of the camera as the coordinate origin, the Z-axis as the camera optical axis, and the X-axis and Y-axis parallel to the X-axis and Y-axis in the image coordinate system respectively. It is usually represented by Pc(Xc, Yc, Zc) for its coordinate values.
[0107] The external parameters of the camera can determine the relative position relationship between the camera coordinate system and the world coordinate system, and are the parameters for converting from the world coordinate system to the camera coordinate system, including the rotation matrix R and the translation vector T. Taking pinhole imaging as an example, the external parameters of the camera, world coordinates, and camera coordinates satisfy the relationship formula (1): Pc = RPw + T (1); where, Pw is the world coordinate, Pc is the camera coordinate, T = (Tx, Ty, Tz) is the translation vector, and R = R(α, β, γ) is the rotation matrix, which are the rotation angles of γ around the Z-axis, β around the Y-axis, and α around the X-axis of the camera coordinate system respectively. These 6 parameters, namely α, β, γ, Tx, Ty, Tz, constitute the external parameters of the camera.
[0108] The internal parameters of the camera determine the projection relationship from three-dimensional space to two-dimensional images and are only related to the camera. Taking the pinhole imaging model as an example, without considering image distortion, the internal parameters can include the scale factors of the camera in the u and v directions of the image coordinate system, the principal point coordinates (x0, y0) relative to the imaging plane coordinate system, and the coordinate axis tilt parameter s. The scale factor of the u-axis is the ratio of the physical length of each pixel in the x direction in the image coordinate system to the camera focal length f, and the scale factor of the v-axis is the ratio of the physical length of the pixel in the y direction in the image coordinate system to the camera focal length. If image distortion is considered, the internal parameters can include the scale factors of the camera in the u and v directions of the image coordinate system, the principal point coordinates relative to the imaging plane coordinate system, the coordinate axis tilt parameter, and the distortion parameters. The distortion parameters can include three radial distortion parameters and two tangential distortion parameters of the camera.
[0109] The internal parameters and external parameters of the camera can be obtained through Zhang Zhengyou calibration. In the embodiments of the present application, the internal parameters and external parameters of the first camera and the internal parameters and external parameters of the second camera are calibrated in the same world coordinate system.
[0110] The imaging plane coordinate system, i.e., the image coordinate system, takes the center of the image plane as the coordinate origin, and the X-axis and Y-axis are respectively parallel to the two perpendicular sides of the image plane. Its coordinate values are usually represented by P(x, y). The image coordinate system represents the position of pixels in the image in physical units (e.g., millimeters).
[0111] The pixel coordinate system, i.e., the image coordinate system in pixel units, takes the upper left vertex of the image plane as the origin, and the X-axis and Y-axis are respectively parallel to the X-axis and Y-axis of the image coordinate system. Its coordinate values are usually represented by p(u, v). The pixel coordinate system represents the position of pixels in the image in pixel units.
[0112] Taking the pinhole camera model as an example, the coordinate values in the pixel coordinate system and the coordinate values in the camera coordinate system satisfy equation (2).
[0113]
[0114] Among them, (u, v) represents the coordinates of the image coordinate system in pixel units, (Xc, Yc, Zc) represents the coordinates in the camera coordinate system, and K represents the matrix of the camera internal parameters.
[0115] Few-shot learning refers to the situation where, after a neural network has pre-learned a large number of samples of certain known categories, for new categories, it can achieve fast learning with only a small number of labeled samples.
[0116] Meta-learning is an important branch in the research of few-shot learning. Its main idea is that when the number of training samples for the target task is small, a neural network is trained using a large number of small-sample tasks similar to the target small-sample task, so that the trained neural network has a good initial value for the target task, and then the trained neural network is adjusted using a small number of training samples of the target small-sample task.
[0117] The Model-agnostic meta-learning (MAML) algorithm, a specific algorithm of meta-learning, whose idea is to train the initialization parameters of a machine learning model so that the machine learning model can perform better after one or more parameter learning on a small amount of data from a new task.
[0118] Soft argmax is an algorithm or function that can obtain the key point coordinates from a heat map. It can be implemented using the layers of a neural network. The layer that implements soft argmax can be called the soft argmax layer.
[0119] Binary cross-entropy is a type of loss function.
[0120] Monocular depth estimation (Fast Depth) is a method of estimating the distance of each pixel in an image relative to the shooting source using a single RGB image or an image from a unique perspective.
[0121] A head-up display system (HUD), also known as a parallel display system, can project important driving information such as speed, engine speed, battery level, and navigation onto the windshield in front of the driver, allowing the driver to view vehicle parameters and driving information such as speed, engine speed, battery level, and navigation through the windshield display area without lowering or turning their head.
[0122] An augmented reality head-up display system (AR-HUD) precisely combines image information with the actual traffic conditions through a specially designed internal optical system, and projects information such as tire pressure, speed, and engine speed onto the windshield, enabling the vehicle owner to view vehicle-related information without lowering their head while driving.
[0123] The first possible implementation is to collect a large amount of eye gaze data to train an eye gaze tracking model, deploy the trained eye gaze tracking model on the vehicle side, and the vehicle side uses this eye gaze tracking model to process the real-time collected images to finally obtain the eye gaze direction of the user. The main defects of this implementation are as follows: There may be significant individual differences between the samples used to train the eye gaze tracking model and the current user (for example, individual differences in the internal structure of the human eye, etc.), which makes the matching degree of the eye gaze tracking model to the current user not high, resulting in inaccurate eye gaze estimation of the current user.
[0124] The second possible implementation is: Use the screen to display a specific image, and calibrate the eye gaze tracking device through the interaction between the user and the specific image on the screen to obtain parameters for this user, thereby improving the accuracy of the eye gaze tracking device for this user. The main defects of this implementation are as follows: It depends on the active cooperation of the user, the operation is cumbersome, and calibration errors may occur due to improper human operation, ultimately affecting the accuracy of the eye gaze tracking device for the user. At the same time, in a vehicle environment, it is difficult to deploy a large enough display screen directly in front of the driver in the cockpit, so this implementation is not applicable to the cockpit scenario.
[0125] The third possible implementation is that when using the screen to display a playback image, first use a basic eye gaze tracking model to predict the initial eye gaze direction, obtain the initial fixation area on the screen based on the initial eye gaze direction, and correct the predicted fixation area by combining the initial fixation area with the currently playing image, thereby improving the accuracy of eye gaze estimation. The main defects of this implementation are as follows: It is only applicable to the scenario of looking at the screen, and for scenarios where the fixation point is constantly changing, the accuracy is relatively low.
[0126] All of the above implementation manners have the problem of inaccurate line-of-sight estimation in the cockpit scenario. In view of this, the embodiments of the present application propose a line-of-sight calibration method, device, equipment, computer-readable storage medium, system, and vehicle. The three-dimensional position of the user's eyes is obtained from the first image including the user's eyes, the three-dimensional position of the fixation point is obtained from the calibration position on the display screen or the second image including the external scene seen by the user, and the second line-of-sight direction with higher accuracy is obtained from the three-dimensional position of the user's eyes and the three-dimensional position of the fixation point. Therefore, this embodiment can effectively improve the accuracy of user line-of-sight estimation and is applicable to the cockpit scenario. In addition, the embodiments of the present application also use the optimized samples including the user's second line-of-sight direction and its first image to optimize the line-of-sight tracking model through the small-sample learning method, so as to improve the line-of-sight estimation accuracy of the line-of-sight tracking model for the user, thereby obtaining a user-level line-of-sight tracking model, and solving the problems of difficult optimization of the line-of-sight tracking model and low line-of-sight estimation accuracy for some users.
[0127] The embodiments of the present application are applicable to any application scenario that requires real-time calibration or estimation of the line-of-sight direction of a person. In some examples, the embodiments of the present application are applicable to the line-of-sight calibration or estimation of drivers and / or passengers in the cockpit environment of transportation tools such as vehicles, ships, and aircraft. In other examples, the embodiments of the present application are also applicable to other scenarios, such as calibrating or estimating the line-of-sight of a person wearing wearable eyes or other devices. Of course, the embodiments of the present application can also be applied to other scenarios, which will not be listed one by one here.
[0128]
Embodiment 1
[0129] First, the system applicable to the embodiment will be exemplarily described below.
[0130] Figure 1 The architecture diagram of the exemplary system 100 in this embodiment in the cockpit environment is shown. Refer to Figure 1 As shown, the exemplary system 100 may include: a first camera 110, a second camera 120, an image processing system 130, and a model optimization system 140.
[0131] The first camera 110 is responsible for collecting the user's eye image (i.e., the first image below). Refer to Figure 1 As shown, taking the cockpit scenario as an example, the first camera 110 may be an in-vehicle camera in the DMS, and the in-vehicle camera is used to photograph the driver in the cockpit. Taking the driver as an example, refer to Figure 2 In the example of, the in-vehicle camera can be installed on the A-pillar of the car ( Figure 2at the position ① in it) or a DMS camera near the steering wheel. The DMS camera is preferably an RGB camera with a relatively high resolution. Here, the human eye image (i.e., the first image hereinafter) generally refers to various types of images containing human eyes, such as a face image, a half-body image containing a face, etc. In some embodiments, to facilitate obtaining the position of the human eye through the first image, obtaining other information of the user, and reducing the amount of image data at the same time, the human eye image (i.e., the first image hereinafter) can be a face image.
[0132] The second camera 120 is responsible for collecting a scene image (i.e., the second image hereinafter). The scene image contains the scene outside the vehicle seen by the user. That is, the field of view of the second camera 120 at least partially coincides with the field of view of the user. Refer to Figure 2 As shown, taking the cockpit scene and the driver as an example, the second camera 120 can be an external vehicle camera, and the external vehicle camera can be used to capture the scene in front of the vehicle seen by the driver. Refer to Figure 2 For example, the external vehicle camera can be a front camera installed above the front windshield of the vehicle ( Figure 2 at the position ② in it). It can capture the scene in front of the vehicle, that is, the scene outside the vehicle seen by the driver. The front camera is preferably a TOF camera, which can collect depth images to facilitate obtaining the distance between the vehicle and the target object in front of it (such as the object the user is looking at) through the image.
[0133] The image processing system 130 is an image processing system capable of processing DMS images and scene images. It can run a gaze tracking model to obtain the user's preliminary gaze data and use the preliminary gaze data (i.e., the first gaze direction hereinafter) to perform the gaze calibration method described below to obtain the user's calibrated gaze data (i.e., the second gaze direction hereinafter), thereby improving the accuracy of the user's gaze data.
[0134] The model optimization system 140 can be responsible for optimizing the gaze tracking model. It can optimize the gaze tracking model using the calibrated gaze data of the user provided by the image processing system 130 and provide the optimized gaze tracking model to the image processing system 130, thereby improving the gaze estimation accuracy of the gaze tracking model for the user.
[0135] In practical applications, the first camera 110, the second camera 120, and the image processing system 130 can all be deployed on the vehicle side, that is, in the vehicle. The model optimization system 140 can be deployed on the vehicle side and / or the cloud according to needs. The image processing system 130 and the model optimization system 140 can communicate through a network.
[0136] In some embodiments, the exemplary system 100 may further include a model training system 150, which is responsible for training a gaze tracking model that can be deployed in the cloud. In actual applications, the model optimization system 140 and the model training system 150 may be implemented by the same system.
[0137] Referring to Figure 2 as shown, the camera coordinate system of the first camera 110 may be a rectangular coordinate system Xc1 - Yc1 - Zc1, and the camera coordinate system of the second camera 120 may be a rectangular coordinate system Xc2 - Yc2 - Zc2. The image coordinate system and pixel coordinate system of the first camera 110 and the second camera 120 are not shown in Figure 2 In this embodiment, to facilitate optimizing the gaze tracking model using the calibrated second gaze direction, the camera coordinate system of the first camera 110 is used as the reference coordinate system. The gaze direction, the three - dimensional position of the fixation point, and the three - dimensional position of the eyes can all be represented by coordinates and / or angles in the camera coordinate system of the first camera 110. In specific applications, the reference coordinate system can be freely selected according to various factors such as actual requirements, specific application scenarios, and requirements for computational complexity, and is not limited to this. For example, the cockpit coordinate system of the vehicle can also be used as the reference coordinate system.
[0138] The gaze calibration method of this embodiment will be described in detail below.
[0139] Figure 3 shows an exemplary process of the gaze calibration method of this embodiment. Referring to Figure 3 as shown, an exemplary gaze calibration method of this embodiment may include the following steps:
[0140] Step S301, obtaining the three - dimensional position of the user's eyes and the first gaze direction according to the first image containing the user's eyes collected by the first camera 110;
[0141] Step S302, obtaining the fixation area of the user in the second image according to the three - dimensional position of the eyes, the first gaze direction, the external parameters of the first camera 110, and the external and internal parameters of the second camera 120. The second image is collected by the second camera 120 and contains the external scene seen by the user;
[0142] Step S303, obtaining the position of the fixation point of the user in the second image according to the fixation area of the user in the second image and the second image;
[0143] Step S304, obtaining the three - dimensional position of the user's fixation point according to the position of the fixation point of the user in the second image and the internal parameters of the second camera 120;
[0144] Step S305, obtaining the second gaze direction of the user according to the three - dimensional position of the fixation point and the three - dimensional position of the eyes, and the second gaze direction is used as the calibrated gaze direction.
[0145] In the line-of-sight calibration method of this embodiment, the line-of-sight direction of the user can be calibrated by using the second image to obtain a second line-of-sight direction with relatively high accuracy, effectively improving the accuracy of the user's line-of-sight data, and further improving the user experience of the upper-layer application based on line-of-sight tracking.
[0146] The first line-of-sight direction is extracted from the first image based on the line-of-sight tracking model. Taking the system 100 as an example, the line-of-sight tracking model can be trained by the model training system 150 deployed in the cloud and provided to the image processing system 130 deployed on the user's vehicle side. The image processing system 130 runs the line-of-sight tracking model to process the first image containing the user's eyes to obtain the first line-of-sight direction of the user.
[0147] The three-dimensional position of the eyes can be represented as the coordinates of a pre-selected eye reference point in the reference coordinate system. In at least some embodiments, the eye reference point can be selected according to the requirements of the application scenario, the usage of the line-of-sight direction, the requirements of the computational complexity, the hardware performance, and the user's own requirements. Figure 4 The example diagram of the eye reference point is shown. The eye reference point can include, but is not limited to, one or more of the midpoint O between the two eyes, the left eye center O1, and the right eye center O2. Here, the eye center can be the pupil center, the eyeball center, the corneal center, or other positions of the eye, and can be freely selected according to needs.
[0148] For the user in the cockpit scenario, the distance between the fixation point and the two eyes will be much greater than the inter-pupillary distance. At this time, the midpoint O between the two eyes can be selected as the eye reference point. In this way, the data volume can be reduced without affecting the accuracy of line-of-sight estimation, the computational complexity can be reduced, and the processing efficiency can be improved. If it is necessary to use the second line-of-sight direction to optimize the line-of-sight tracking model, and the user expects a relatively high accuracy of the line-of-sight tracking model, the left eye center O1 and the right eye center O2 can be selected as the eye reference points.
[0149] The line-of-sight direction can be represented by the viewing angle and / or the line-of-sight vector in the reference coordinate system. The viewing angle can be the angle between the line of sight and the eye axis, and the intersection position of the line of sight and the eye axis is the three-dimensional position of the user's eyes. The line-of-sight vector is a direction vector with the position of the eye in the reference coordinate system as the starting point and the position of the fixation point in the reference coordinate system as the ending point. The three-dimensional coordinates of the eye reference point in the reference coordinate system and the three-dimensional coordinates of the fixation point in the reference coordinate system can be included in this direction vector.
[0150] The fixation point refers to the point where the user's eyes are fixated. Taking the cockpit scenario as an example, the fixation point of the driver is the specific position that the driver's eyes are looking at. A fixation point can be represented by its position in space. In this embodiment, the three-dimensional position of the fixation point is represented by the three-dimensional coordinates of the fixation point in the reference coordinate system.
[0151] In step S301, the three-dimensional position of the user's eyes can be determined by various applicable methods. In some implementation manners, the three-dimensional position of the eyes can be obtained by combining a face feature point detection algorithm with a pre-constructed 3D face model. In some implementation manners, the three-dimensional position of the eyes can be obtained by using the Landmark algorithm with the two-dimensional position obtained from the first image combined with the depth information of the first image. It can be understood that any method that can obtain the three-dimensional position of a certain point in an image through image data can be applied to determine the three-dimensional position of the user's eyes in step S301, and no further enumeration will be given here.
[0152] Figure 5 An exemplary process of estimating the three-dimensional position of the eyes is shown. Refer to Figure 5 As shown, the exemplary process of estimating the three-dimensional position of the eyes may include: Step S501, using a face detection algorithm and a face feature point detection algorithm to process the first image to obtain the positions of the user's face feature points in the first image; S502, performing PnP solution by combining the positions of the user's face feature points in the first image with a pre-obtained standard 3D face model to solve the 3D coordinates of the user's face feature points in the reference coordinate system; S503, extracting the 3D coordinates of the user's eye reference points from the 3D coordinates of the user's face feature points in the reference coordinate system as the 3D coordinates of the user's eyes. It should be noted that Figure 5 This is only an example and is not used to limit the specific implementation manner of estimating the three-dimensional position of the eyes in this embodiment.
[0153] In step S302, according to the three-dimensional position of the user's eyes, the first line-of-sight direction, the external parameters of the first camera 110, the internal parameters and external parameters of the second camera 120, the camera perspective projection model can be used to determine the user's fixation region in the second image (hereinafter, the "fixation region in the second image" is abbreviated as the "second fixation region"). Here, the camera perspective projection model can be a pinhole imaging model or a non-linear perspective projection model.
[0154] To obtain a more accurate second fixation region, step S302 may include: obtaining the user's fixation region in the second image according to the three-dimensional position of the user's eyes, the first line-of-sight direction, the external parameters of the first camera 110, the internal parameters and external parameters of the second camera 120, and the accuracy of the line-of-sight tracking model. Thus, the error caused by the accuracy limitation of the line-of-sight tracking model can be eliminated in the finally obtained second line-of-sight direction.
[0155] The process of obtaining the second fixation region will be described in detail below in combination with a specific scenario.
[0156] Figure 6 A scenario where a driver (not shown in the figure) in the cockpit environment looks at a pedestrian in the crosswalk in front of the vehicle is shown.
[0157] Figure 9 An exemplary process for determining the second fixation area of a user is shown. Refer to Figure 9 As shown, the process of obtaining the second fixation area of the user may include the following steps:
[0158] Step S901, determine the fixation area S1 of the user in the reference coordinate system according to the three-dimensional position of the user's eyes and the first line of sight direction.
[0159] Specifically, according to the coordinates (Xc1, Yc1, Zc1) of the user's eye reference point in the reference coordinate system and the first line of sight direction ON (viewing angle θ) obtained from the first image, the line of sight ON of the user in the reference coordinate system is obtained. Assume that the average accuracy value of the line of sight tracking model is expressed as: ±α, where α represents the error value of the viewing angle. The lower the accuracy of the line of sight tracking model, the larger the value of α. In this step, the line of sight angle θ can be adjusted to the interval value [θ - α, θ + α], and the cone formed by the line of sight with the line of sight angle θ - α and the line of sight with the line of sight angle θ + α is used as the fixation area S1 of the user in the reference coordinate system.
[0160] Figure 7 is shown Figure 6 A visualization graph of the fixation area S1 of the driver in the reference coordinate system in the shown scenario. O represents the three-dimensional position of the eyes, the solid line with an arrow represents the first line of sight direction ON, θ represents the viewing angle of the first line of sight direction ON, α represents the average accuracy value of the line of sight tracking model, and the dashed cone represents the fixation area S1 of the user in the reference coordinate system.
[0161] Step S902, project the fixation area S1 of the user in the reference coordinate system onto the pixel coordinate system of the second camera 120 to obtain the second fixation area Q of the user.
[0162] Figure 8 is shown the second image captured by the second camera Figure 6 in the shown scenario, where only the part that the driver is fixating on is shown, and the content irrelevant to this embodiment in Figure 6 the shown scenario is omitted, and Figure 8 the second fixation area Q of the user is marked in
[0163] Taking the pinhole imaging model as an example, in combination with Figures 6 - 8For an example of this, the projection process of this step can be achieved through Equation (1) and Equation (2). Specifically, first, based on the extrinsic parameters of the first camera 110 and the extrinsic parameters of the second camera 120, the fixation area S1 is converted into the camera coordinate system of the second camera 120 according to Equation (1) to obtain the fixation area S2; then, based on the intrinsic parameters of the second camera 120, the fixation area S2 is projected into the pixel coordinate system of the second camera 120 according to the relational expression (2) to obtain the second fixation area Q of the user. Here, the extrinsic parameters of the first camera 110 and the extrinsic parameters of the second camera 120 are calibrated in the same world coordinate system.
[0164] The fixation area S1 is projected as a quadrilateral second fixation area Q on the imaging plane of the second camera 120 through the extrinsic parameters of the first camera 110 and the intrinsic and extrinsic parameters of the second camera 120. Generally, the lower the accuracy of the gaze tracking model, the larger the value of α, the larger the angle of the fixation area S1 of the user in the reference coordinate system, and the wider the width of the quadrilateral second fixation area Q.
[0165] Figure 10 Fig. shows a projection example diagram of a line of sight OX. Refer to Figure 10 As shown, the projections of points x with different depths on the line of sight OX on the imaging plane of the second camera 120 are O’X’. As Figure 10 shown, taking O on the left as the origin of the human eye line of sight in space, OX as the first line of sight direction L, mapping to the camera imaging plane of the second camera 120, the mapping point of the origin of the human eye line of sight is O’, and the first line of sight direction L is mapped to the line of sight L’.
[0166] It should be noted that Figures 7 - 10 The method shown is only an example, and the method for obtaining the second fixation area in the embodiments of the present application is not limited to this.
[0167] The second fixation area can be characterized by grayscale image data. The pixel points in the grayscale image data of the second fixation area correspond one by one to the pixel points in the second image, and the grayscale value of each pixel point can indicate whether it belongs to the fixation area. Refer to the example below Figure 11 For an example, assume that the visual representation of the second image is Fig1, and the black and white image Fig2 is the visual representation of the second fixation area. The black pixel points in the black and white image Fig2 do not belong to the second fixation area, and the white pixel points belong to the second fixation area. Taking the cockpit scenario as an example, when the second camera uses a TOF camera, the second image is a TOF image, and the grayscale value of each pixel in the second image can indicate the distance from the corresponding point of the target object to the second camera.
[0168] In step S303, the position of the user's fixation point in the second image (hereinafter referred to as the "second fixation point" for short) can be obtained based on the second fixation area and the second image through a pre-trained fixation point calibration model. The fixation point calibration model can be any machine learning model applicable to image processing. Considering the high accuracy and good stability of neural networks, in the embodiments of the present application, the fixation point calibration model is preferably a neural network model.
[0169] The exemplary implementation manner of the fixation point calibration model will be described in detail below.
[0170] Figure 11 The exemplary network structure of the fixation point calibration model is shown. Refer to Figure 11 As shown, the fixation point calibration model can be a neural network model with an encoder-decoder (encoder-decoder structure). Refer to Figure 11 As shown, the fixation point calibration model may include a channel-wise concat layer, a ResNet-18 based encoder, a Convolutional GRU Cell, a ResNet-18 based decoder, and a soft-argmax+scaling layer.
[0171] Refer to Figure 11 As shown, the processing process of the fixation point calibration model includes: at the input end of the fixation point calibration model, first, the image of the second fixation area and the second image are merged in the channel direction into a new image through the channel-wise concat layer. If both the second image and the image of the second fixation area are single-channel grayscale images, the merged new image has 2 channels. If the second image is an RGB three-channel color image and the image of the second fixation area is a single-channel grayscale image, the merged new image has 4 channels, that is, a 4-channel image. The merged new image is input into the encoding network and processed successively through the encoding network, the Convolutional GRU Cell, and the decoding network. The decoding network outputs a heatmap Fig3, and the gray value of each pixel in the heatmap Fig3 indicates the probability that the corresponding pixel point is a fixation point. After the decoding network outputs the heatmap Fig3, the heatmap Fig3 is calculated by the soft-argmax normalization layer to obtain the position of the fixation point in the second image, that is, the coordinates (x, y) of the corresponding pixel point of the fixation point in the second image. Generally, a line of sight has one fixation point, and each fixation point may include one or more pixel points in the second image.
[0172] The fixation point calibration model can be pre-trained. During training, the scene image and its corresponding grayscale image of the fixation area (the range of the fixation area in the grayscale image of the fixation area is a set value) are used as samples, and the true fixation area of the sample is known. During the training process, the ResNet part and the soft-argmax standard layer are trained simultaneously but use different loss functions. The embodiments of the present application do not limit the specific loss function to be used. For example, the loss function of the ResNet part can be binary cross-entropy (BCE loss), and the loss function of the soft-argmax standard layer can be mean square error (MSE loss).
[0173] In some examples, the decoding network in the ResNet part can use pixel-level binary cross-entropy as the loss function, and the expression is as shown in the following formula (3).
[0174]
[0175] Among them, is the label of whether pixel i is a fixation point, taking 1 when it is a fixation point and 0 when it is not a fixation point, and p(y i ) is the probability value that pixel i in the heat map Fig3 output by the decoding network is a fixation point, and N is the total number of pixels in the second image Fig1, that is, the total number of pixels in the heat map Fig3. Figure 11 In an example, the specification of the second image is 128×72, and its total number of pixels N = 128 * 72 = 9216.
[0176] In step S304, there are various specific implementation methods for obtaining the three-dimensional position of the user's fixation point according to the position of the fixation point of the user in the second image and the internal parameters of the second camera 120. The three-dimensional position of the fixation point is the three-dimensional coordinates of the fixation point in the reference coordinate system (the camera coordinate system of the first camera 110). It can be understood that any algorithm for obtaining the position of a point in space based on its position in the image can be applied to step S304.
[0177] Considering that the inverse perspective transformation is relatively mature and has a low computational complexity, in step S304, it is preferably to obtain the three-dimensional position of the fixation point through the inverse perspective transformation. Specifically, in step S304, only by obtaining the depth of the second fixation point can the Z-axis coordinate of the fixation point in the reference coordinate system be obtained. Combining the position of the second fixation point obtained in step S303, that is, the pixel coordinates (u, v), the three-dimensional coordinates of the fixation point in the reference coordinate system, that is, the three-dimensional position of the fixation point, can be obtained through a simple inverse perspective transformation.
[0178] Figure 12 Shows an exemplary specific implementation process of step S304. See Figure 12As shown, step S304 may include: step S3041, obtaining the depth of the second fixation point based on a monocular depth estimation algorithm using the second image. This depth is the distance h of the fixation point relative to the second camera 120, and the Z-axis coordinate Zc2 of the fixation point in the camera coordinate system of the second camera is estimated from the distance h; step S3042, obtaining the three-dimensional coordinates of the fixation point in the reference coordinate system based on the position of the second fixation point, i.e., the pixel coordinates (u, v), and the Z-axis coordinate of the fixation point in the camera coordinate system of the second camera, based on the internal and external parameters of the second camera 120 and the external parameters of the first camera 110.
[0179] In step S3041, the distance h of each pixel in the second image relative to the second camera 120 can be calculated using the second image through a monocular depth estimation algorithm such as FastDepth, and the distance h of the second fixation point relative to the second camera 120 can be extracted from it according to the position of the second fixation point, i.e., the pixel coordinates. Here, various applicable algorithms can be used for depth estimation. In one example, it is preferably to calculate the depth of each pixel point in the second image through a monocular depth estimation (FastDepth) algorithm. This algorithm has a low computational complexity, high processing efficiency, mature and stable algorithm, relatively low requirements for hardware performance, and is convenient to be implemented by vehicle-end devices with relatively low computing power.
[0180] In step S3042, according to the position of the second fixation point, i.e., the pixel coordinates (u, v), the Z-axis coordinate Zc of the fixation point in the reference coordinate system, and the internal parameters of the second camera 120, the coordinate values (Xc2, Yc2, Zc2) of the fixation point in the camera coordinate system of the second camera are deduced backward through Equation (2). Then, based on the external parameters of the second camera 120 and the external parameters of the first camera 110, the coordinate values (Xc1, Yc1, Zc1) of the fixation point in the camera coordinate system of the first camera 110 are deduced through Equation (1) from the coordinate values (Xc2, Yc2, Zc2) of the fixation point in the camera coordinate system of the second camera. The coordinate values (Xc1, Yc1, Zc1) are the three-dimensional position of the fixation point.
[0181] Generally, a line of sight has one fixation point, but due to accuracy limitations, multiple fixation points may be obtained corresponding to the same line of sight. At this time, the fixation points can be screened according to the confidence of the fixation points of the user in the second image. In this way, only the screened fixation points need to be executed in the subsequent steps to obtain the second line-of-sight direction, which can reduce the amount of calculation while ensuring the accuracy of the second line-of-sight direction and improve the processing efficiency. Here, the screening of the fixation points can be performed before step S304 or after step S304.
[0182] In step S303, the fixation calibration model simultaneously provides the probability value of the second fixation point, and the confidence level of the second fixation point can be determined by this probability value. In some embodiments, the heat map provided by the fixation calibration model includes the probability value of the second fixation point. This probability value represents the probability that the second fixation point is the true fixation point. The higher the probability value, the higher the likelihood that the corresponding second fixation point is the true fixation point. The probability value can be directly used as the confidence level of the second fixation point or the proportional function value of this probability value can be used as the confidence level of the second fixation point. Thus, the confidence level of the second fixation point can be obtained without separate calculation, which can improve the processing efficiency and reduce the computational complexity at the same time.
[0183] There can be various specific implementation manners for screening fixation points based on the confidence level. In some examples, only the fixation points whose confidence levels of the second fixation point exceed a preset first confidence level threshold (for example, 0.9) or the fixation points with the relatively highest confidence level can be selected. If there are still multiple fixation points whose confidence levels of the second fixation point are relatively the highest or exceed the first confidence level threshold, one or more can be randomly selected from these fixation points. Of course, if there are still multiple fixation points whose confidence levels of the second fixation point exceed the first confidence level threshold or the fixation points with the relatively highest confidence level of the second fixation point, these multiple fixation points can also be retained simultaneously. In this way, through screening, not only can the accuracy of the finally obtained second line of sight direction be ensured to be higher, but also the amount of computation and data volume in step S304, step S305, and the following step S306 can be reduced, thereby effectively improving the processing efficiency and reducing the hardware loss at the same time, which is convenient for implementation by vehicle terminal devices with relatively low computing power and limited storage capacity.
[0184] In step S305, the second line of sight direction can be represented by a vector or an angle of view determined by the three-dimensional position of the fixation point and the three-dimensional position of the eye. In some embodiments, in the camera coordinate system of the first camera, the second line of sight direction can be characterized by a vector with the three-dimensional position of the eye as the starting point and the three-dimensional position of the fixation point as the ending point. In some embodiments, in the camera coordinate system of the first camera, the second line of sight direction can be characterized by the angle (i.e., the angle of view) between the line of sight starting from the three-dimensional position of the eye and pointing to the three-dimensional position of the fixation point and the axis of the user's eye reference point.
[0185] In the embodiments of the present application, the line of sight calibration of steps S301 to S305 can be executed by the image processing system 130 in the system 100.
[0186] Generally, deep learning models can use a small number of samples for "few-shot learning" to improve the model accuracy for specific users. However, for the line of sight tracking model, the data required is the line of sight data (such as the line of sight angle) of the user in the camera coordinate system. This type of numerical data is difficult to directly obtain in a general environment, which makes it difficult to optimize the line of sight tracking model at the user level. In view of this, the result obtained in step S305 can be used to optimize the line of sight tracking model.
[0187] After step S305, the line-of-sight calibration method according to the embodiment of the present application may further include: step S306, using the user's second line-of-sight direction and the first image as the user's optimization samples, and optimizing the line-of-sight tracking model based on the few-shot learning method. Thus, the line-of-sight estimation accuracy of the line-of-sight tracking model for a specific user can be continuously improved with a small number of samples and small-scale training, and a user-level line-of-sight tracking model can be obtained.
[0188] Taking the above Figure 1 exemplary system as an example, Figure 13 an exemplary implementation process of optimizing the line-of-sight tracking model in step S306 is shown. Refer to Figure 13 As shown, this exemplary process may include: step S3061, the image processing system 130 stores the second line-of-sight direction and its corresponding first image as the user's optimization samples in the user's sample library. This sample library can be associated with user information (for example, user identity information) for easy query and is deployed in the model optimization system 140. Step S3062, the model optimization system 140 uses the newly added optimization samples in the user's sample library to optimize the user's line-of-sight tracking model obtained from the previous optimization based on the few-shot learning method. Step S3063, the model optimization system 140 sends the user's line-of-sight tracking model obtained from this optimization to the image processing system 130 on the user's vehicle side, so that the image processing system 130 can use the optimized line-of-sight tracking model to obtain its first line-of-sight direction in the user's next line-of-sight calibration. Among them, the parameter data of the user's line-of-sight tracking model obtained from the previous optimization and the user's sample library can both be associated with user information (for example, user identity information), so that the optimization samples and the parameter data of the line-of-sight tracking model obtained from the previous optimization can be directly queried through the user information during this optimization. In this way, the user's optimization samples can be collected in real time without the user's awareness and the line-of-sight tracking model can be continuously optimized. The longer the user uses the line-of-sight tracking model and the higher the frequency, the more accurate the line-of-sight estimation of the line-of-sight tracking model for the user, and the better the user experience. At the same time, the technical problem that the line-of-sight tracking model has low line-of-sight estimation accuracy and is difficult to optimize for some users is solved while improving the accuracy of the user's line-of-sight estimation in real time.
[0189] In practical applications, the optimization in step S3062 can be performed regularly or when the newly added optimization samples reach a certain number or other preset conditions are met. When the image processing system 130 and the model optimization system 140 can communicate normally, the update of the sample library in step S3061 can be performed in real time.
[0190] Optionally, in step S3061, the optimized samples of the user can be selectively uploaded to improve the quality of the optimized samples, reduce unnecessary optimization operations, and reduce the hardware loss caused by model optimization. Specifically, the second line-of-sight directions can be screened according to the second fixation confidence, and only the optimized samples formed by the screened second line-of-sight directions and their corresponding first images are uploaded. Here, the screening of the second line-of-sight directions can include, but is not limited to: 1) selecting the second line-of-sight directions with the confidence of the second fixation points greater than a preset second confidence threshold (for example, 0.95); 2) selecting the second line-of-sight directions with the relatively highest confidence of the second fixation points. Here, regarding the confidence of the second fixation points, reference can be made to the relevant descriptions above and will not be elaborated here.
[0191] The few-shot learning method can be implemented by any algorithm that can optimize the gaze tracking model with a small number of samples. For example, the MAML algorithm can be used to optimize the gaze tracking model with the optimized samples of the user to achieve the optimization of the gaze tracking model based on the few-shot learning method. Thus, a gaze tracking model that better fits the specific individual characteristics of the user can be obtained with a small number of samples, with a small amount of data and low computational complexity, which is beneficial to reducing hardware loss and lowering hardware costs.
[0192] The following takes the cockpit scenario as an example to illustrate the specific implementation manner of this embodiment.
[0193] Figure 14 An exemplary processing flow for the system 100 to perform gaze calibration and model optimization in the cockpit environment is shown. Refer to Figure 14As shown in the figure, the processing flow may include: Step S1401, the in-vehicle camera of vehicle G captures a DMS image (i.e., the first image) of driver A in the vehicle cockpit. The DMS image contains the face of driver A. The in-vehicle image processing system 130 of vehicle G runs a gaze tracking model to infer the initial gaze direction (i.e., the first gaze direction). At the same time, the three-dimensional position of driver A's eyes is estimated using the DMS image. Step S1402, the image processing system 130 combines the out-of-vehicle image (i.e., the second image) captured by the out-of-vehicle camera and the gaze area of the initial gaze direction for inference to obtain the calibrated gaze direction of driver A (i.e., the second gaze direction). The out-of-vehicle image contains the scene that driver A currently sees, and the out-of-vehicle image is captured synchronously with the above DMS image. Step S1403, when it is determined that the credibility of the calibrated gaze direction is relatively high (for example, the confidence of the second fixation point meets the relevant requirements above), the image processing system 130 uploads the DMS image of driver A and the calibrated gaze direction as the personalized data of driver A (i.e., the optimized sample) to the model optimization system 140. The model optimization system 140 optimizes the gaze tracking model of driver A using the few-shot learning method to obtain the gaze tracking model of driver A and downloads it to the in-vehicle image processing system 130 of vehicle G. It can be seen that in this embodiment, the out-of-vehicle image is used to calibrate the initial gaze data estimated by the gaze tracking model to improve the accuracy of the gaze data, and the obtained calibrated gaze data is used as the personalized gaze data of the user to optimize the gaze tracking model, improving the gaze estimation accuracy of the gaze tracking model for the corresponding user. Thus, this embodiment can not only solve the problem that the gaze estimation result of the gaze tracking model is inaccurate in actual use in the cockpit scenario, but also solve the technical problem that it is difficult to optimize the gaze tracking model due to the inability to obtain the user's gaze data in the cockpit scenario. Moreover, the system is also growth-oriented. In the in-vehicle scenario, the above processing flow can continue without the user's awareness. The more the user uses the system, the more accurate the gaze estimation of the system for the user, and the higher the accuracy of the gaze tracking model for the user.
[0194] Figure 15 The exemplary structure of the gaze calibration device 1500 provided in this embodiment is shown. Refer to Figure 15 As shown in the figure, the gaze calibration device 1500 of this embodiment may include:
[0195] An eye position determination unit 1501, configured to obtain the three-dimensional position of the user's eyes according to the first image containing the user's eyes captured by the first camera;
[0196] A first gaze determination unit 1502, configured to obtain the first gaze direction of the user according to the first image containing the user's eyes captured by the first camera;
[0197] The fixation area unit 1503 is configured to obtain the fixation area of the user in the second image according to the three-dimensional eye position, the first line-of-sight direction, the external parameters of the first camera, and the external and internal parameters of the second camera, where the second image is captured by the second camera and includes the external scene seen by the user;
[0198] The fixation point calibration unit 1504 is configured to obtain the position of the fixation point of the user in the second image according to the fixation area of the user in the second image and the second image;
[0199] The fixation point conversion unit 1505 is configured to obtain the three-dimensional fixation point position of the user according to the position of the fixation point of the user in the second image and the internal parameters of the second camera;
[0200] The second line-of-sight determination unit 1506 is configured to obtain the second line-of-sight direction of the user according to the three-dimensional fixation point position and the three-dimensional eye position, and the second line-of-sight direction is used as the calibrated line-of-sight direction.
[0201] In some embodiments, the first line-of-sight direction is extracted from the first image based on a line-of-sight tracking model.
[0202] In some embodiments, the fixation area unit 1503 is configured to obtain the fixation area of the user in the second image according to the three-dimensional eye position, the first line-of-sight direction, the external parameters of the first camera, and the external and internal parameters of the second camera, including: obtaining the fixation area of the user in the second image according to the three-dimensional eye position, the first line-of-sight direction, the external parameters of the first camera, the external and internal parameters of the second camera, and the accuracy of the line-of-sight tracking model.
[0203] In some embodiments, the line-of-sight calibration device further includes: an optimization unit 1507 configured to use the second line-of-sight direction of the user and the first image as the optimization samples of the user, and optimize the line-of-sight tracking model based on a small-sample learning method.
[0204] In some embodiments, the fixation point calibration unit 1504 may also be configured to screen the fixation points according to the confidence of the fixation points of the user in the second image; and / or, the optimization unit 1507 is also configured to screen the second line-of-sight direction according to the confidence of the fixation points of the user in the second image.
[0205] In some embodiments, the position of the fixation point of the user in the second image is obtained by using a fixation point calibration model according to the fixation area of the user in the second image and the second image.
[0206] In some embodiments, the fixation point calibration model simultaneously provides the probability value of the fixation point of the user in the second image, and the confidence is determined by the probability value.
[0207]
Embodiment 2
[0208] Figure 16 An exemplary architecture of the system 1600 applicable to this embodiment is shown. Refer to Figure 16 As shown, the exemplary system 1600 of this embodiment is basically the same as the system 100 of Embodiment 1. The difference is that in the exemplary system 1600 of this embodiment, the second camera 120 is an optional component, which includes a display screen 160. The display screen 160 can be deployed at the vehicle end and implemented through the existing display components in the vehicle-end device. Other parts of the system 1600 in this embodiment, namely the first camera 110, the image processing system 130, the model optimization system 140, and the model training system 150, have basically the same functions as the corresponding parts in the system 100 of Embodiment 1 and will not be elaborated here. In this embodiment, the display screen 160 marked with the positional relationship with the first camera 110 (i.e., the in-vehicle camera) is used. By relying on the reference point that the user gazes at the display screen 160, the calibration of the user's line of sight is achieved and its optimization sample is obtained. The small-sample learning is performed on the line-of-sight tracking model using this optimization sample to improve its accuracy.
[0209] The line-of-sight calibration method of this embodiment will be described in detail below.
[0210] Figure 17 An exemplary process of the line-of-sight calibration method in this embodiment is shown. Refer to Figure 17 As shown, the line-of-sight calibration method of this embodiment may include the following steps:
[0211] Step S1701, in response to the user's gaze operation on the reference point in the display screen 160, obtain the three-dimensional position of the user's gaze point;
[0212] Before this step, it may further include: controlling the display screen 160 to provide a line-of-sight calibration interface for the user. The line-of-sight calibration interface includes a visual prompt for reminding the user to gaze at the reference point, so that the user can perform the corresponding gaze operation according to this visual prompt. Here, the specific form of the line-of-sight calibration interface is not limited in this embodiment.
[0213] In this step, the gaze operation can be any operation related to the user's gaze at the reference point in the display screen 160. The specific implementation manner or manifestation form of the gaze operation is not limited in the embodiments of this application. For example, the gaze operation may include the user inputting confirmation information in the line-of-sight calibration interface while gazing at the reference point in the line-of-sight calibration interface.
[0214] Taking the cockpit scenario as an example, the display screen 160 can be, but is not limited to, the AR-HUD of the vehicle, the instrument panel of the vehicle, the user's portable electronic device, or others. Generally, the line-of-sight calibration in the cockpit scenario is mainly for the driver or the co-driver. Therefore, to ensure that the line-of-sight calibration does not affect safe driving, the display screen 160 is preferably the AR-HUD.
[0215] In this step, the three-dimensional coordinates of each reference point in the display screen 160 in the camera coordinate system of the first camera 110 can be pre-calibrated through the positional relationship between the display screen 160 and the first camera 110. Thus, when the user gazes at a reference point, that reference point is the user's fixation point, and the three-dimensional coordinates of that reference point in the camera coordinate system of the first camera 110 are the three-dimensional position of the user's fixation point.
[0216] Step S1702: Obtain the three-dimensional position of the user's eyes based on the first image captured by the first camera 110 and containing the user's eyes;
[0217] The specific implementation manner of this step is the same as the specific implementation manner of the three-dimensional position of the eyes in step S301 in the first embodiment, and will not be elaborated here.
[0218] Step S1703: Obtain the second line-of-sight direction of the user based on the three-dimensional position of the fixation point and the three-dimensional position of the eyes.
[0219] The specific implementation manner of this step is the same as step S305 in the first embodiment, and will not be elaborated here.
[0220] In the line-of-sight calibration method of this embodiment, the three-dimensional position of the user's fixation point can be obtained by using the reference point, and at the same time, the three-dimensional position of the user's eyes is obtained in combination with the first image, that is, the second line-of-sight direction with relatively high accuracy is obtained. It can be seen that the line-of-sight calibration method of this embodiment can not only effectively improve the accuracy of user line-of-sight estimation, but also is simple to operate, has low computational complexity, and high processing efficiency, and is applicable to the cockpit environment.
[0221] The method of this embodiment preferably uses the camera coordinate system of the first camera 110 as the reference coordinate system. The second line-of-sight direction obtained thereby can be directly used for optimizing the line-of-sight tracking model. The three-dimensional position of the fixation point and the three-dimensional position of the eyes are both represented by the three-dimensional coordinate values in the camera coordinate system of the first camera 110. The second line-of-sight direction can be represented by the viewing angle or the direction vector in the camera coordinate system of the first camera 110. For detailed details, refer to the relevant description in the first embodiment, and will not be elaborated here.
[0222] Similarly to the embodiments, the line-of-sight calibration method of this embodiment may further include: Step S1704, using the user's second line-of-sight direction and the first image as the user's optimization samples, and optimizing the line-of-sight tracking model based on the few-shot learning method. Thus, the line-of-sight estimation accuracy of the line-of-sight tracking model for a specific user can be continuously improved with a small number of samples and small-scale training, and a user-level line-of-sight tracking model can be obtained. The specific implementation of this step is the same as that of Step S306 in Embodiment 1 and will not be elaborated here. Since the three-dimensional position of the fixation point in this step is obtained through calibration and its accuracy is relatively high, there is no need to screen the second line-of-sight direction before Step S1704 of this embodiment.
[0223] Figure 18 Fig. shows an exemplary structure of the line-of-sight calibration device 1800 provided in this embodiment. Refer to Figure 18 As shown, the line-of-sight calibration device 1800 of this embodiment may include:
[0224] The fixation point position determination unit 1801 is configured to obtain the three-dimensional position of the user's fixation point in response to the user's fixation operation on the reference point in the display screen;
[0225] The eye position determination unit 1501 is configured to obtain the three-dimensional position of the user's eyes according to the first image including the user's eyes collected by the first camera;
[0226] The second line-of-sight determination unit 1506 is configured to obtain the second line-of-sight direction of the user according to the three-dimensional position of the fixation point and the three-dimensional position of the eyes.
[0227] In some embodiments, the display screen is an augmented reality head-up display.
[0228] In some embodiments, the device further includes: an optimization unit 1507 configured to use the user's second line-of-sight direction and the first image as the user's optimization samples, and optimize the line-of-sight tracking model based on the few-shot learning method.
[0229] Next, the computing device and computer-readable storage medium of the embodiments of the present application will be described.
[0230] Figure 19 Fig. is a structural schematic diagram of a computing device 1900 provided in an embodiment of the present application. The computing device 1900 includes: a processor 1910 and a memory 1920.
[0231] The computing device 1900 may further include a communication interface 1930 and a bus 1940. It should be understood that Figure 19 The communication interface 1930 in the shown computing device 1900 can be used for communication with other devices. The memory 1920 and the communication interface 1930 can be connected to the processor 1910 through the bus 1940. For the sake of convenience of representation,Figure 19 It is represented by only one line, but it does not mean that there is only one bus or one type of bus.
[0232] Among them, the processor 1910 can be connected to the memory 1920. The memory 1920 can be used to store the program code and data. Therefore, the memory 1920 can be a storage unit inside the processor 1910, or an external storage unit independent of the processor 1910, or a component including a storage unit inside the processor 1910 and an external storage unit independent of the processor 1910.
[0233] It should be understood that in the embodiments of the present application, the processor 1910 can adopt a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. Or the processor 1910 adopts one or more integrated circuits to execute relevant programs to implement the technical solutions provided by the embodiments of the present application.
[0234] The memory 1920 can include a read-only memory and a random access memory, and provide instructions and data to the processor 1910. A part of the processor 1910 can also include a non-volatile random access memory. For example, the processor 1910 can also store information about the device type.
[0235] When the computing device 1900 is running, the processor 1910 executes the computer-executable instructions in the memory 1920 to perform the operation steps of the sight calibration method in the above-mentioned embodiments.
[0236] It should be understood that the computing device 1900 according to the embodiments of the present application can correspond to the corresponding subject executing the methods according to the embodiments of the present application, and the above and other operations and / or functions of each module in the computing device 1900 respectively correspond to the corresponding processes of the methods in this embodiment. For the sake of brevity, they will not be described in detail here.
[0237] Next, the system architecture and related applications of the embodiments of the present application will be exemplarily described.
[0238] The embodiment of the present application further provides a driver monitoring system, which includes the above-mentioned first camera 110, second camera 120, and computing device 1900.
[0239] In some embodiments, the first camera 110 is configured to capture a first image including the user's eyes, the second camera 120 is configured to capture a second image including the scene seen by the user, and both the first camera 110 and the second camera 120 can communicate with the computing device 1900. In the computing device 1900, the processor 1910 uses the first image provided by the first camera 110 and the second image provided by the second camera 120 to execute the computer-executable instructions in the memory 1920 to perform the operation steps of the line-of-sight calibration method in the first embodiment above.
[0240] In some embodiments, the driver monitoring system may further include a display screen configured to display a reference point to the user. In the computing device 1900, the processor 1910 uses the first image provided by the first camera 110 and the three-dimensional position of the reference point displayed on the display screen to execute the computer-executable instructions in the memory 1920 to perform the operation steps of the line-of-sight calibration method in the second embodiment above.
[0241] In some embodiments, the driver monitoring system may further include a cloud server, which may be configured to use the second line-of-sight direction and the first image of the user provided by the computing device 1900 as an optimization sample of the user, optimize the line-of-sight tracking model based on the small-sample learning method, and provide the optimized line-of-sight tracking model to the computing device 1900, thereby improving the line-of-sight estimation accuracy of the line-of-sight tracking model for the user.
[0242] Specifically, the architecture of the driver monitoring system can refer to the system shown in the first embodiment Figure 1 and the system shown in the second embodiment Figure 16 Among them, the image processing system 130 can be deployed in the computing device 1900, and the model optimization system 140 described above can be deployed in the cloud server.
[0243] The embodiment of the present application further provides a vehicle, which may include the above-mentioned driver monitoring system. In specific applications, the vehicle is a motor vehicle, which may be, but is not limited to, a sport utility vehicle, a large bus, a large truck, a passenger vehicle of various commercial vehicles, and may also be, but is not limited to, various boats, ships, aircraft, etc. It may also be, but is not limited to, a hybrid vehicle, an electric vehicle, a plug-in hybrid electric vehicle, a hydrogen-powered vehicle, and other alternative fuel vehicles. Among them, a hybrid vehicle can be any vehicle with two or more power sources, such as a vehicle with both gasoline and electric power sources.
[0244] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0245] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0246] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0247] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0248] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0249] When the above functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of this application, in essence, or the part that makes contributions to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0250] An embodiment of this application also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it is used to execute a line-of-sight calibration method, and this method includes at least one of the solutions described in the above various embodiments.
[0251] The computer storage medium of the embodiments of this application can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0252] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0253] The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0254] The computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0255] Note that the above is only the preferred embodiment of this application and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although this application has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, all of which fall within the protection scope of the present invention.
Claims
1. A line of sight calibration method, characterized in that, Including: Obtaining the three-dimensional position of the user's eyes and the first line-of-sight direction based on a first image collected by a first camera and including the user's eyes; Obtaining the fixation area of the user in a second image according to the three-dimensional position of the eyes, the first line-of-sight direction, the extrinsic parameters of the first camera, and the extrinsic and intrinsic parameters of a second camera, where the second image is collected by the second camera and includes the external scene seen by the user; Obtaining the position of the fixation point of the user in the second image according to the fixation area of the user in the second image and the second image; Obtaining the three-dimensional position of the fixation point of the user according to the position of the fixation point and the intrinsic parameters of the second camera; Obtaining the second line-of-sight direction of the user according to the three-dimensional position of the fixation point and the three-dimensional position of the eyes, where the second line-of-sight direction is used as the calibrated line-of-sight direction.
2. The line-of-sight calibration method according to claim 1, wherein The first line-of-sight direction is extracted from the first image based on a line-of-sight tracking model.
3. The line-of-sight calibration method according to claim 2, wherein Obtaining the fixation area of the user in the second image according to the three-dimensional position of the eyes, the first line-of-sight direction, the extrinsic parameters of the first camera, and the extrinsic and intrinsic parameters of the second camera includes: obtaining the fixation area of the user in the second image according to the three-dimensional position of the eyes, the first line-of-sight direction, the extrinsic parameters of the first camera, the extrinsic and intrinsic parameters of the second camera, and the accuracy of the line-of-sight tracking model.
4. The line-of-sight calibration method according to claim 2 or 3, characterized in that Further including: Using the second line-of-sight direction of the user and the first image as the optimization samples of the user, and optimizing the line-of-sight tracking model based on a small-sample learning method.
5. The line-of-sight calibration method according to any one of claims 1 to 3, characterized in that, Further including: Screening the fixation point or the second line-of-sight direction according to the confidence level of the fixation point of the user in the second image.
6. The line-of-sight calibration method according to any one of claims 1 to 3, characterized in that The position of the fixation point of the user in the second image is obtained by using a fixation point calibration model according to the fixation area of the user in the second image and the second image.
7. The line-of-sight calibration method according to claim 6, wherein, The fixation point calibration model simultaneously provides the probability value of the fixation point of the user in the second image, and the confidence level is determined by the probability value.
8. A line of sight calibration device, characterized in that, Including: An eye position determination unit configured to obtain the three-dimensional position of the user's eyes according to a first image collected by a first camera and including the user's eyes; A first line-of-sight determination unit configured to obtain the first line-of-sight direction of the user according to a first image collected by a first camera and including the user's eyes; A fixation area unit configured to obtain the fixation area of the user in a second image according to the three-dimensional position of the eyes, the first line-of-sight direction, the extrinsic parameters of the first camera, and the extrinsic and intrinsic parameters of a second camera, where the second image is collected by the second camera and includes the external scene seen by the user; A fixation point calibration unit configured to obtain the position of the fixation point of the user in the second image according to the fixation area of the user in the second image and the second image; A fixation point conversion unit configured to obtain the three-dimensional position of the fixation point of the user according to the position of the fixation point and the intrinsic parameters of the second camera; A second line-of-sight determination unit configured to obtain the second line-of-sight direction of the user according to the three-dimensional position of the fixation point and the three-dimensional position of the eyes.
9. The line-of-sight calibration device according to claim 8, characterized in that The first line-of-sight direction is extracted from the first image based on a line-of-sight tracking model.
10. The line-of-sight calibration device according to claim 9, characterized in that, The gaze area unit is configured to obtain the user's gaze area in the second image according to the three-dimensional position of the eyes, the first line-of-sight direction, the external parameters of the first camera, the external and internal parameters of the second camera, and the accuracy of the line-of-sight tracking model.
11. The line-of-sight calibration device according to claim 9 or 10, characterized in that, It further includes: An optimization unit configured to use the user's second line-of-sight direction and the first image as the user's optimization samples to optimize the line-of-sight tracking model based on the few-shot learning method.
12. The line-of-sight calibration device according to any one of claims 8 to 10, characterized in that The gaze point calibration unit is further configured to screen the gaze points according to the confidence of the user's gaze points in the second image; and / or, the optimization unit is further configured to screen the second line-of-sight direction according to the confidence of the user's gaze points in the second image.
13. The line-of-sight calibration device according to any one of claims 8 to 10, characterized in that The position of the user's gaze point in the second image is obtained by using a gaze point calibration model based on the user's gaze area in the second image and the second image.
14. The line-of-sight calibration device according to claim 13, characterized in that, The gaze point calibration model simultaneously provides the probability value of the user's gaze point in the second image, and the confidence is determined by the probability value.
15. A computing device, characterized in that, It includes: At least one processor; And At least one memory storing program instructions that, when executed by the at least one processor, cause the at least one processor to execute the method according to any one of claims 1 to 7.
16. A computer-readable storage medium having program instructions stored thereon, characterized in that, The program instructions, when executed by a computer, cause the computer to execute the method according to any one of claims 1 to 7.
17. A driver monitoring system, characterized in that, It includes: A first camera configured to capture a first image including the user's eyes; A second camera configured to capture a second image including the external scene seen by the user; At least one processor; And At least one memory storing program instructions that, when executed by the at least one processor, cause the at least one processor to execute the method according to any one of claims 1 to 7.
18. A vehicle, characterized in that, It includes the driver monitoring system according to claim 17.
Citation Information
Patent Citations
Information providing method, device and system
CN109849788A
Eyeball tracking method and device, vehicle and storage medium
CN110341617A