A method, device, and storage medium for line-of-sight detection in assisted driving

Through infrared cameras and depth cameras, a line of sight detection model is built, which solves the problem of difficulty in sample collection and high computational complexity in line of sight detection, and improves the accuracy and simplification of line of sight detection, and improves driving safety.

CN114067422BActive Publication Date: 2025-07-29BEIJING YINWO AUTOMOBILE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111436975.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2025-07-29
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

The existing line of sight detection technology has problems such as difficulty in collecting samples, inaccurate line of sight detection results and high computational complexity, especially the lack of reliable line of sight data sets and line of sight data sets of Asian eyes, which leads to unsatisfactory detection results.

Method used

Through infrared cameras and depth cameras, a line of sight detection model is constructed, including eye model inference module and line of sight regression module, the eyeball binary map and pupil contour are used to calculate the line of sight direction, build a line of sight data set and train the model to simplify the calculation process.

Benefits of technology

It achieves improved accuracy and reduced computational complexity of line of sight detection, and can build a complete and balanced line of sight data set faster and more conveniently, improving driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067422B_ABST
    Figure CN114067422B_ABST
Patent Text Reader

Abstract

The present invention provides a line-of-sight detection method, device and storage medium for assisted driving, which can obtain reliable line-of-sight data, is convenient to calculate, improves the accuracy of line-of-sight detection, and has a low computational complexity. It collects data through a camera, enables the line of sight of the person being collected to follow the target object, obtains an image of the human eye following the target object moving, and calculates the line-of-sight direction data; constructs a line-of-sight detection model including an eyeball model inference module and a line-of-sight regression module, inputs the image into the eyeball model inference module to output an eyeball binary map, inputs the eyeball binary map into the line-of-sight regression module, and outputs the inferred line-of-sight direction; constructs a line-of-sight data set from the collected images of the human eye following the target object moving and the line-of-sight direction data, trains the line-of-sight detection model through the line-of-sight data set to obtain a trained line-of-sight detection model; inputs the collected human eye image to be detected into the trained line-of-sight detection model, and outputs the line-of-sight direction of the human eye to be detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of assisted driving and artificial intelligence, and particularly relates to a method, device and storage medium for line-of-sight detection for assisted driving. Background Art

[0002] Nowadays, with the continuous improvement of people's living standards, cars have gradually become essential items in people's lives, and safe driving has attracted more and more attention. Therefore, the research on safe driving has become more and more in-depth. One of the most influential factors in driving is the line-of-sight direction of the driver. By monitoring the line-of-sight direction of the driver, the driver's attention can be improved, and the safety of traffic driving can be enhanced.

[0003] The current line-of-sight detection technologies have the following two difficulties: difficult sample collection and inaccurate line-of-sight detection results. In the current publicly available datasets, there are almost no line-of-sight datasets based on infrared cameras, and most of them are for European and American eyes, and there are almost no line-of-sight datasets for Asian eyes. Line-of-sight data is different from the datasets for detection and classification, and it is difficult to manually annotate. It is necessary to re-collect line-of-sight data. Since the line-of-sight data in space is three-dimensional information, it is not easy to collect. The accuracy of the line-of-sight detection results is affected by the dataset and the model. In the case of a lack of reliable datasets, the line-of-sight detection results will also be unsatisfactory. At the same time, some existing methods also have the problem of high computational complexity. Summary of the Invention

[0004] In view of the above problems, the present invention provides a method, device and storage medium for line-of-sight detection for assisted driving, which can obtain reliable line-of-sight data as samples, is convenient for calculation operations, improves the accuracy of line-of-sight detection, and has a low computational complexity.

[0005] The technical solution is as follows: A method for line-of-sight detection for assisted driving, characterized by comprising the following steps: collecting data through a camera, enabling the line of sight of the person being collected to follow a target object to move, and obtaining an image of the human eye following the target object to move;

[0006] Calculating line-of-sight direction data according to the obtained image of the human eye following the target object to move;

[0007] Constructing a line-of-sight detection model, the line-of-sight detection model including an eyeball model inference module and a line-of-sight regression module, inputting the image into the eyeball model inference module, and outputting an eyeball binary map, the eyeball binary map including an eyeball contour and a pupil contour, inputting the eyeball binary map into the line-of-sight regression module, and outputting the inferred line-of-sight direction;

[0008] Construct a line-of-sight dataset from the collected images of the human eye moving with the target object and the calculated line-of-sight direction data, and train the line-of-sight detection model through the line-of-sight dataset until the model converges to obtain a trained line-of-sight detection model;

[0009] Input the collected image of the human eye to be detected into the trained line-of-sight detection model, and output the line-of-sight direction of the human eye to be detected.

[0010] Further, the data collection by the camera includes: placing an infrared camera and a depth camera in the same horizontal plane, placing a target object in front of the person to be collected, moving the target object, the person to be collected observes the target object with the human eye, the line of sight follows the target object to move, and synchronously collect images through the infrared camera and the depth camera, so that the moving trajectory of the target object covers the entire area of the image.

[0011] Further, the person to be collected observes the target object with the human eye, including: the head pose of the person to be collected maintains the face facing one direction unchanged, move the target object, the person to be collected observes the target object with the human eye, the line of sight follows the target object to move, and collect images through the infrared camera and the depth camera, so that the moving trajectory of the target object covers the entire area of the image.

[0012] Further, the person to be collected observes the target object with the human eye, including: the line of sight of the person to be collected follows the target object to move, the head of the person to be collected is in the same direction as the line of sight, and moves with the target object, and collect images through the infrared camera and the depth camera, so that the moving trajectory of the target object covers the entire area of the image.

[0013] Further, the images collected by the infrared camera and the depth camera at least include the images of the facial orientations of the upper left, upper, upper right, left, front, right, lower left, lower, and lower right.

[0014] Further, the target object uses a spherical object.

[0015] Further, the calculation of the line-of-sight direction data according to the obtained images of the human eye moving with the target object includes: obtaining the coordinates of the pupil center of the human eye in the depth camera coordinate system, obtaining the coordinates of the pupil center in the infrared camera coordinate system through the coordinate transformation relationship between the depth camera and the infrared camera, obtaining the coordinates of the target object in the depth camera coordinate system, obtaining the coordinates of the target object in the infrared camera coordinate system through the coordinate transformation relationship between the depth camera and the infrared camera, subtracting the target object coordinates from the pupil center coordinates to obtain a line-of-sight vector for representing the line-of-sight direction.

[0016] Further, obtaining the coordinates of the pupil center of the human eye in the depth camera coordinate system includes: obtaining the depth information Zc of the pupil through the shooting of the depth camera, obtaining the coordinates (u, v) of the pupil center in the image coordinate system through image annotation, and according to the conversion formula between the image coordinates and the depth camera coordinates:

[0017]

[0018] where is the internal parameter matrix of the depth camera, and the coordinates (Xc, Yc, Zc) of the pupil center in the depth camera coordinate system are obtained.

[0019] Further, the coordinate transformation relationship between the depth camera and the infrared camera is obtained through the following steps: performing binocular calibration on the infrared camera and the depth camera to obtain the translation matrix T and the rotation matrix R. The transformation relationship of the same coordinate point in the depth camera coordinate system and the infrared camera coordinate system is expressed as P1 = R * P2 + T, where P1 is the coordinate point in the infrared camera coordinate system and P2 is the coordinate point in the depth camera coordinate system.

[0020] Further, the eyeball model inference module includes a histogram equalization layer, a ResNet network layer, and a 1*1 convolutional filter. The histogram equalization layer is used to increase the contrast of the input image. The ResNet network layer includes three ResNet networks for extracting human eye features. The 1*1 convolutional filter is used to convert the extracted features into an eyeball binary map;

[0021] The gaze regression module includes a DenseNet network layer and a fully connected layer. The DenseNet network layer includes three residual modules. The gaze regression module obtains the connection line between the eyeball center and the pupil center through the input eyeball binary map and outputs it as the gaze direction.

[0022] Further, training the gaze detection model with the gaze dataset includes:

[0023] Taking the images in the gaze dataset as samples and the gaze data as labels, inputting them into the eyeball model inference module of the gaze detection model, and outputting the inferred eyeball binary map. The optimization loss function is expressed as follows:

[0024]

[0025] where is a constant, p represents the coordinates of each pixel, P represents the pixel coordinates of the entire image, represents the eyeball binary map predicted by the eyeball model, and m(p) represents the eyeball binary map of the image;

[0026] Input the binary image of the eyeball obtained from the inference into the gaze regression module of the gaze detection model to output the inferred gaze direction. Compare the inferred gaze direction with the true gaze direction to optimize the loss function, which is expressed as follows:

[0027]

[0028] where w and ε are constants, g label is the estimated gaze direction, ĝ is the gaze direction inferred by the gaze detection model, and ln() is the natural logarithm.

[0029] An apparatus for a gaze detection method for assisted driving, characterized in that it includes: a processor, a memory, and a program;

[0030] The program is stored in the memory, and the processor calls the program stored in the memory to execute the above-mentioned gaze detection method for assisted driving.

[0031] A computer-readable storage medium, characterized in that: the computer-readable storage medium is configured to store a program, and the program is configured to execute the above-mentioned gaze detection method for assisted driving.

[0032] The gaze detection method for assisted driving provided by the present invention only needs to provide a simple target object and a camera to obtain an image of the human eye following the movement of the target object, and the acquisition process involves various gaze angles, and the gaze range involved in the entire data set is more complete and balanced; through the cooperation of an infrared camera and a depth camera, the calculation from the acquired image to the obtained gaze data can be simplified, so that the data set can be constructed faster and more conveniently. For the gaze detection model used in gaze detection, in the present application, through the eyeball model inference module, a binary image of the eyeball is obtained. The binary image of the eyeball includes the eyeball contour and the pupil contour. The eyeball and the pupil are mapped onto the image as a circle and an ellipse respectively. These two shapes are combined into a binary image, and the gaze direction is the connection line between the center of the eyeball and the center of the pupil. Input the binary image of the eyeball into the gaze regression module to output the inferred gaze direction. When calculating, inferring the gaze on the binary image is simpler than calculating the gaze on the original image. Since the eyeball model inference module is added, the computational complexity of gaze detection is simplified, and it is robust to changes in head pose and image quality in the data set. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 A schematic diagram of the steps of a gaze detection method for assisted driving in an embodiment;

[0034] Figure 2 A schematic diagram of obtaining a binary image of the eyeball in an embodiment;

[0035] Figure 3 The eye ball pictures generated according to the line-of-sight directions in the line-of-sight data set during the training of the model in the embodiment;

[0036] Figure 4 The schematic diagram of the line-of-sight detection model in the embodiment;

[0037] Figure 5 The internal structure diagram of the computer device in one embodiment. Detailed implementation manners

[0038] See Figure 1 A line-of-sight detection method for assisted driving according to the present invention includes the following steps:

[0039] Step S1: Collect data through a camera, so that the line of sight of the person being collected follows the target object to move, and obtain an image of the human eye following the target object to move;

[0040] Step S2: Calculate the line-of-sight direction data according to the obtained image of the human eye following the target object to move;

[0041] Step S3: Construct a line-of-sight detection model. The line-of-sight detection model includes an eyeball model inference module and a line-of-sight regression module. Input the image into the eyeball model inference module, and output an eyeball binary map. The eyeball binary map includes the eyeball contour and the pupil contour. Input the eyeball binary map into the line-of-sight regression module, and output the inferred line-of-sight direction;

[0042] Step S4: Construct a line-of-sight data set from the collected images of the human eye following the target object and the calculated line-of-sight direction data, and train the line-of-sight detection model through the line-of-sight data set until the model converges, and obtain a trained line-of-sight detection model;

[0043] Step S5: Input the image of the human eye to be detected collected into the trained line-of-sight detection model, and output the line-of-sight direction of the human eye to be detected.

[0044] In one embodiment of the present invention, in step S1, collecting data through a camera includes: placing an infrared camera and a depth camera in the same horizontal plane, placing a target object in front of the person being collected, moving the target object, and the person being collected observes the target object through the human eye, and the line of sight follows the target object to move. Synchronously collect images through the infrared camera and the depth camera, so that the moving trajectory of the target object covers the entire area of the image.

[0045] In one embodiment of the present invention, the target object is preferably a spherical object, which is circular when viewed from any direction, convenient for annotation, and easy to determine the center point of the small ball. Specifically, it can be a small ball connected with a rope, which can be suspended in the air and can be driven to move by a human or a device.

[0046] In an embodiment of the present invention, during data acquisition, the infrared camera and the depth camera are preferably placed on the same horizontal plane for more accurate calibration. At the same time, the two cameras need to be synchronized during data acquisition to capture images at the same moment.

[0047] In an embodiment of the present invention, the captured images may include various situations:

[0048] Situation 1: The head posture of the person being captured maintains the face facing in one direction unchanged, that is, the head does not move, and the target object is moved. The person being captured observes the target object through the human eye, and the line of sight follows the movement of the target object. Images are captured by the infrared camera and the depth camera so that the trajectory of the moving target object covers the entire area of the image.

[0049] In this situation, the images captured by the infrared camera and the depth camera include at least the images of the face orientations in the upper left, upper, upper right, left, front, right, lower left, lower, and lower right directions. The images contain the human eye and the target object.

[0050] Taking the face facing the upper left as an example, slowly move the small ball serving as the target object, and the line of sight should follow the ball until the trajectory of the small ball covers the entire captured image. Here, one capture action is completed; then turn the face upward and perform one capture action; the face should perform at least 9 poses of capture actions. There is no regulation on the initial position of the small ball in the image. It is just that during one capture action, the trajectory of the small ball should cover the entire image.

[0051] Situation 2: The line of sight of the person being captured follows the movement of the target object, and the head of the person being captured is in the same direction as the line of sight and moves together with the target object. Images are captured by the infrared camera and the depth camera so that the trajectory of the moving target object covers the entire area of the image.

[0052] Generally, in this situation, the images captured by the infrared camera and the depth camera include at least the images of the face orientations in the upper left, upper, upper right, left, front, right, lower left, lower, and lower right directions. The images contain the human eye, and the capture process involves various line-of-sight angles, making the line-of-sight range involved in the entire dataset more complete and balanced.

[0053] In step S2, according to the obtained images of the human eye following the movement of the target object, calculate the line-of-sight direction data, which specifically includes: obtaining the coordinates of the pupil center of the human eye in the depth camera coordinate system, obtaining the coordinates of the pupil center in the infrared camera coordinate system through the coordinate transformation relationship between the depth camera and the infrared camera, obtaining the coordinates of the target object in the depth camera coordinate system, obtaining the coordinates of the target object in the infrared camera coordinate system through the coordinate transformation relationship between the depth camera and the infrared camera, subtracting the target object coordinates from the pupil center coordinates to obtain the line-of-sight vector, which is used to represent the line-of-sight direction.

[0054] Specifically, in one embodiment of the present invention, obtaining the coordinates of the pupil center of the human eye in the depth camera coordinate system includes: obtaining the depth information Zc of the pupil through the shooting of the depth camera, which can be directly obtained by the depth camera after collecting the picture;

[0055] By annotating the image, the coordinates (u, v) of the pupil center in the image coordinate system can be obtained. According to the conversion formula between the image coordinates and the depth camera coordinates:

[0056]

[0057] where, The internal parameter matrix of the depth camera, which has been obtained through calibration before. According to the above formula, only Xc and Yc are unknown. Since other variables are known, these two values can be directly calculated. Thus, the coordinate values of the pupil center in the depth camera coordinate system are obtained, that is, (Xc, Yc, Zc).

[0058] For the coordinate transformation relationship between the depth camera and the infrared camera, it is obtained through the following steps: performing binocular calibration on the infrared camera and the depth camera to obtain the translation matrix T and the rotation matrix R. The transformation relationship of the same coordinate point in the depth camera coordinate system and the infrared camera coordinate system is expressed as P1 = R * P2 + T, where P1 is the coordinate point in the infrared camera coordinate system and P2 is the coordinate point in the depth camera coordinate system. Through the transformation relationship between the depth camera and the infrared camera, the coordinate values of the pupil center in the infrared camera coordinate system can be obtained.

[0059] At the same time, it is necessary to annotate the small ball. Similarly, through the above method, the coordinate values of the target object in the infrared camera coordinate system can be obtained. The depth information of the small ball is directly obtained after the depth camera collects the image. The coordinate transformation formula of the small ball is expressed as:

[0060]

[0061] where, u, v are the coordinate values of the small ball in the image coordinate system, X c , Y c , Z c +R are the coordinate values of the small ball in the depth camera coordinate system, and R is the radius of the small ball.

[0062] Then, the three-dimensional coordinates of the pupil center and the three-dimensional coordinates of the small ball are subtracted to obtain a vector, which can represent the line-of-sight direction information. Because when collecting data, the eyes are always looking at the small ball, so after obtaining the coordinate values of the pupil center and the small ball center in the infrared camera coordinate system, subtracting the coordinate values of the pupil center from the coordinate values of the small ball is the line-of-sight direction of the eyes in the infrared camera coordinate system.

[0063] Step S2 is for each image. For example, if 100 images are collected at a time, the line-of-sight direction can be calculated for each image and used as a sample for subsequent training.

[0064] Specifically in step S3, the constructed line-of-sight detection model includes an eyeball model inference module and a line-of-sight regression module.

[0065] Among them, the eyeball model inference module includes a histogram equalization layer, a ResNet network layer, and a 1*1 convolutional filter. The histogram equalization layer is used to increase the contrast of the input image. The ResNet network layer includes three ResNet networks and is used to extract human eye features. The 1*1 convolutional filter is used to convert the extracted features into an eyeball binary map. The eyeball model inference module extracts deep features from the eye image through convolution pooling and other operations, simulates the positions of the eyeball and the pupil to obtain the eyeball binary map, as shown in Figure 2 shown;

[0066] The line-of-sight regression module includes a DenseNet network layer and a fully connected layer. The DenseNet network layer includes three residual modules. The line-of-sight regression module obtains the connection line between the eyeball center and the pupil center through the input eyeball binary map and outputs it as the line-of-sight direction.

[0067] In this embodiment, the line-of-sight detection model includes an eyeball model inference module and a line-of-sight regression module. The input size of the eye image is 128*64. After preprocessing by histogram equalization, the details of the eyes are more clearly contrasted. After the image is first passed through 3 eye_blocks to extract features, it passes through a 1*1 convolutional filter to obtain the eyeball binary map. The extracted features are then connected to the eyeball binary map through a 1*1 convolutional filter and input into the line-of-sight regression module. In this way, both the binary map of the eyeball model is inferred and the feature information passed from the upper layer can be retained. The two pieces of information are fused as the input of the line-of-sight regression module. In the line-of-sight regression module, after passing through 3 residual modules and a fully connected layer, the line-of-sight vector is inferred. The overall network structure is as follows Figure 4 shown. Although the entire network can be divided into two parts, the training is end-to-end and the training process is not separated.

[0068] Specifically in step S4, it includes constructing a line-of-sight data set from the images of the human eye moving with the target object collected and the calculated line-of-sight direction data, and training the line-of-sight detection model through the line-of-sight data set until the model converges to obtain a trained line-of-sight detection model.

[0069] Specifically, the images in the line-of-sight data set are input into the line-of-sight detection model, and the inferred line-of-sight direction is output. The input is the eye image and the true line-of-sight direction as the label. When preprocessing the input data, an eyeball binary map will be generated according to the line-of-sight direction, as shown in Figure 2As shown, the relationship for generating the eye binary map from the line of sight direction is as shown in the following formula. The dimension of the input image is m*n, the diameter of the eye on the image is 2r, and 2r = 1.2n is satisfied. The center of the eye is the center of the image. The coordinates of the pupil center are calculated as follows:

[0070]

[0071]

[0072] where r, = r cos(sin 0.5 -1 ), and the line of sight direction as a label is expressed as θ is the pitch angle, is the navigation angle. The conversion of the line of sight direction from vector representation to angle representation can be achieved according to existing formulas.

[0073] Both the pupil center and the eye center are obtained from the eye model inference module. To generate the eye binary map, the loss function is as follows:

[0074]

[0075] where is a constant, usually taken as 10 -5 , p represents the coordinates of each pixel, P represents the pixel coordinates of the entire image, represents the eye binary map predicted by the eye model.

[0076] The inferred eye binary map is input into the line of sight regression module of the line of sight detection model, and the inferred line of sight direction is output. The inferred line of sight direction is compared with the true line of sight direction to optimize the loss function, which is expressed as follows:

[0077]

[0078] where w and ε are constants, g label is the estimated line of sight direction, g is the line of sight direction inferred by the line of sight detection model, and ln() is the logarithm. The loss function in this embodiment draws on the loss function of key point detection, and the range of the predicted value is between [0,1]. When error is extremely large, the gradient is a constant. When error is relatively small, the gradient is larger than both L1 and MSE. Therefore, at small errors, error can be amplified to obtain better results.

[0079] In step S5, after obtaining a trained line-of-sight detection model, the image of the human eye to be detected collected can be input into the trained line-of-sight detection model, and the line-of-sight direction of the human eye to be detected can be output, which can be applied to the driver monitoring system. Through the method of this embodiment, the gazing direction of the driver's line of sight can be detected, and distraction judgment can be made based on the gazing direction of the line of sight, which can assist in improving the safety performance of the driver driving the vehicle and reducing the occurrence of traffic accidents.

[0080] For the method provided in this embodiment, only a small ball is required as the target object, which is used in cooperation with a depth camera and an infrared camera to accurately collect line-of-sight data. Two cameras are used for data collection, one is an ordinary infrared camera, and the other is a depth camera. The three-dimensional coordinates of the eyes and the small ball are obtained by the depth camera and then converted into the coordinate system of the infrared camera. The difference between the three-dimensional coordinates of the eyes and the three-dimensional coordinates of the small ball is calculated to obtain the vector information of the line of sight. Moreover, the collection process covers various line-of-sight angles, and the line-of-sight range involved in the entire data set is more complete and balanced.

[0081] For the line-of-sight detection model of the solution provided in this embodiment, an eyeball model is constructed through the eyeball model inference module. The eyeball model is a shape combination of the eyeball and the pupil mapped onto the image. The eyeball and the pupil mapped onto the image are circular and elliptical respectively, and these two shapes are combined into a binary image. The gazing direction of the line of sight is defined as the connecting line between the center of the eyeball and the center of the pupil. The change in the line-of-sight direction will cause the change in the ellipse positioning. When the line-of-sight regression module calculates, it is much simpler to infer the line of sight on the binary image than to calculate the line of sight on the original image. Due to the addition of the eyeball model inference module, the computational complexity of line-of-sight detection is simplified, and it is robust to changes in head pose and image quality.

[0082] In an embodiment of the present invention, a line-of-sight detection device for assisted driving is further provided, which includes: a processor, a memory, and a program;

[0083] The program is stored in the memory, and the processor calls the program stored in the memory to execute the above-mentioned line-of-sight detection method for assisted driving.

[0084] This computer device can be a terminal, and its internal structure diagram can be as Figure 4As shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected by a bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a line-of-sight detection method for assisted driving. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0085] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electric Erasable Programmable Read-Only Memory (EEPROM), etc. Among them, the memory is used to store programs, and the processor executes the programs after receiving execution instructions.

[0086] A processor may be an integrated circuit chip with the ability to process signals. The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc. The processor may also be other general-purpose processors, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0087] Those skilled in the art can understand that Figure 5 the structure shown in [the figure] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0088] In an embodiment of the present invention, there is also provided a computer-readable storage medium, which is configured to store a program, and the program is configured to execute the above-mentioned method for detecting a line of sight for assisted driving.

[0089] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a computer device, or a computer program product. Therefore, the embodiments of the present invention can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0090] Embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, computer devices, or computer program products according to embodiments of the present invention. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a device for implementing the functions specified in the flowchart and / or block diagram.

[0091] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in the flowchart.

[0092] The above has introduced in detail the application of the present invention in a line-of-sight detection method, a computer device, and a computer-readable storage medium for assisted driving. Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A line-of-sight detection method for assisted driving, characterized in that, It includes the following steps: collecting data through a camera, causing the line of sight of the person being collected to follow the target object to move, and obtaining an image of the human eye following the target object to move; Calculating line-of-sight direction data according to the obtained image of the human eye following the target object to move; Constructing a line-of-sight detection model, the line-of-sight detection model includes an eyeball model inference module and a line-of-sight regression module, inputting the image into the eyeball model inference module, and outputting an eyeball binary map, the eyeball binary map includes an eyeball contour and a pupil contour, inputting the eyeball binary map into the line-of-sight regression module, and outputting the inferred line-of-sight direction; Constructing a line-of-sight data set from the collected images of the human eye following the target object and the calculated line-of-sight direction data, training the line-of-sight detection model through the line-of-sight data set until the model converges, and obtaining a trained line-of-sight detection model; Inputting the collected image of the human eye to be detected into the trained line-of-sight detection model, and outputting the line-of-sight direction of the human eye to be detected; The calculating the line-of-sight direction data according to the obtained image of the human eye following the target object includes: obtaining the coordinates of the pupil center of the human eye in the depth camera coordinate system, obtaining the coordinates of the pupil center in the infrared camera coordinate system through the coordinate transformation relationship between the depth camera and the infrared camera, obtaining the coordinates of the target object in the depth camera coordinate system, obtaining the coordinates of the target object in the infrared camera coordinate system through the coordinate transformation relationship between the depth camera and the infrared camera, subtracting the target object coordinates from the pupil center coordinates, and obtaining a line-of-sight vector for representing the line-of-sight direction.

2. The method for detecting a line of sight for assisted driving according to claim 1, wherein, The collecting data through the camera includes: placing the infrared camera and the depth camera in the same horizontal plane, placing a target object in front of the person being collected, moving the target object, the person being collected observes the target object through the human eye, the line of sight follows the target object to move, and synchronously collecting images through the infrared camera and the depth camera, so that the moving trajectory of the target object covers the entire area of the image.

3. The method for detecting line of sight for assisted driving according to claim 2, wherein The person being collected observes the target object through the human eye includes: the head pose of the person being collected, maintaining the face facing one direction unchanged, moving the target object, the person being collected observes the target object through the human eye, the line of sight follows the target object to move, and collecting images through the infrared camera and the depth camera, so that the moving trajectory of the target object covers the entire area of the image.

4. A method for detecting a line of sight for assisted driving according to claim 2, characterized in that, The person being collected observes the target object through the human eye includes: the line of sight of the person being collected follows the target object to move, the head of the person being collected is in the same direction as the line of sight, and moves together with the target object, and collecting images through the infrared camera and the depth camera, so that the moving trajectory of the target object covers the entire area of the image.

5. A method for detecting a line of sight for assisted driving according to claim 3 or 4, characterized in that, Among the images collected by the infrared camera and the depth camera, at least include images of the face orientations of the upper left, upper, upper right, left, front, right, lower left, lower, and lower right.

6. The vision detection method for assisted driving according to claim 1, wherein: The target object uses a spherical object.

7. A line-of-sight detection method for assisted driving according to claim 1, characterized in that, The obtaining the coordinates of the pupil center of the human eye in the depth camera coordinate system includes: obtaining the depth information Zc of the pupil through the shooting of the depth camera, obtaining the coordinates (u, v) of the pupil center in the image coordinate system through image annotation, according to the conversion formula between the image coordinate and the depth camera coordinate: ; Among them, the internal parameter matrix of the depth camera is used to obtain the coordinates (Xc, Yc, Zc) of the pupil center in the depth camera coordinate system.

8. A line-of-sight detection method for assisted driving according to claim 1, characterized in that, The coordinate system conversion relationship between the depth camera and the infrared camera is obtained through the following steps: Perform binocular calibration on the infrared camera and the depth camera to obtain the translation matrix T and the rotation matrix R. The conversion relationship of the same coordinate point in the depth camera coordinate system and the infrared camera coordinate system is expressed as , where P1 is the coordinate point in the infrared camera coordinate system, and P2 is the coordinate point in the depth camera coordinate system.

9. A line-of-sight detection method for assisted driving according to claim 1, characterized in that: The eye model inference module includes a histogram equalization layer, a ResNet network layer, and a 1×1 convolutional filter. The histogram equalization layer is used to increase the contrast of the input image. The ResNet network layer includes three ResNet networks for extracting human eye features. The 1×1 convolutional filter is used to convert the extracted features into a binary eye image; The gaze regression module includes a DenseNet network layer and a fully connected layer. The DenseNet network layer includes three residual modules. The gaze regression module obtains the connection line between the center of the eye ball and the center of the pupil from the input binary eye image as the gaze direction output.

10. A method for line of sight detection for assisted driving according to claim 1, characterized in that: Training the gaze detection model through a gaze data set includes: Taking the images in the gaze data set as samples and the gaze data as labels, inputting them into the eye model inference module of the gaze detection model, and outputting the inferred binary eye image. The loss function is optimized as follows: ; Among them, is a constant, represents the coordinates of each pixel, represents the pixel coordinates of the entire image, represents the binary eye image predicted by the eye model, represents the binary eye image of the image; Inputting the inferred binary eye image into the gaze regression module of the gaze detection model, outputting the inferred gaze direction, comparing the inferred gaze direction with the true gaze direction, and optimizing the loss function as follows: ; where w and are constants, , is the estimated line-of-sight direction, g is the line-of-sight direction inferred by the line-of-sight detection model, and ln() is the logarithm.

11. An apparatus for a line-of-sight detection method for assisted driving, characterized in that, It includes: a processor, a memory, and a program; The program is stored in the memory. The processor calls the program stored in the memory to execute the gaze detection method for assisted driving according to any one of claims 1 to 10.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium is configured to store a program, and the program is configured to execute the gaze detection method for assisted driving according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Sight line-based driver attention detection method

    CN107818310A

  • Sight tracking method and device, computer equipment and storage medium

    CN112749655A