Method, apparatus and storage medium for generating line-of-sight regression infrared imaging sample data

Through the combination of 3D modeling and GAN model, infrared virtual pictures are generated and line-of-sight regression model is trained, which solves the problem of scarcity and generation of line-of-sight regression data in DMS scenarios, and realizes efficient line-of-sight regression parameter generation.

CN113077547BActive Publication Date: 2025-06-13KEY (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110436936.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-22
Publication Date
2025-06-13
Estimated Expiration
2041-04-22

AI Technical Summary

Technical Problem

In DMS scenarios, line-of-sight regression data is very small and generation is difficult, and the existing technology is difficult to effectively solve this problem.

Method used

The eye model is established through 3D modeling software, multiple 3D pictures are generated, and these 3D pictures are rendered using pre-trained GAN models to generate infrared virtual pictures. Then, the real DMS scene picture is input to the pre-trained line of sight regression model to obtain the line of sight regression parameters.

Benefits of technology

The problem of scarcity and generation of line-of-sight regression data in DMS scenarios was effectively overcome, the training efficiency of line-of-sight regression model was improved, and the line-of-sight regression parameters were generated that met the requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113077547B_ABST
    Figure CN113077547B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, and storage medium for generating line-of-sight regression infrared imaging sample data. The method includes: establishing an eye model using 3D modeling software and generating a plurality of 3D pictures; rendering the 3D pictures using a pre-trained GAN model to generate infrared virtual pictures; inputting real DMS scene pictures into a pre-trained line-of-sight regression model to obtain line-of-sight regression parameters in the real DMS scene; the line-of-sight regression model is trained using the infrared virtual pictures; wherein the line-of-sight regression parameters include the X coordinate and Y coordinate of the center point of the pupil, and the pitch angle and heading angle in the line-of-sight direction. This solution effectively overcomes the technical problems in the prior art solutions, where the line-of-sight regression data in the DMS scene is very small and difficult to generate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of generating sample data for gaze regression infrared imaging, and particularly relates to a method, device, and storage medium for generating sample data for gaze regression infrared imaging. Background Art

[0002] Currently, the mainstream gaze estimation methods can be divided into two categories: Geometry Based Methods and Appearance Based Methods. The basic idea of the geometry-based method is to detect some features of the eyes (such as key points like the corners of the eyes and the position of the pupils), and then calculate the gaze direction based on these features. The appearance-based method is to directly learn a model that maps appearance to the gaze direction. Generally speaking, the geometry-based method is relatively more accurate and stable in different scenarios, but this type of method has high requirements for the quality and resolution of the pictures; the appearance-based method performs better on low-resolution and high-noise images, but the training of the model requires a large amount of data and is prone to overfitting to the scenes in the training set.

[0003] In recent years, with the rise of deep learning and the public release of a large number of data sets, the appearance-based method has received more and more attention, and the effect is significantly better than the geometry-based method, gradually becoming the mainstream. The appearance-based gaze tracking method using deep learning requires a large amount of data to train the model, which brings two problems: (1) Currently, the open-source data sets are all color photos taken under natural light. However, in the DMS (Driver Fatigue Monitor System) scenario, in order to have clear imaging both during the day and at night, the acquisition device used is an infrared imaging camera. There are significant differences in the imaging of infrared cameras and color cameras, especially the characteristics of the pupil part are significantly different. The pupil is darker in the color image and brighter in the infrared image, as shown in Figure 1a and 1b Therefore, the models trained using the open-source public color data sets have relatively poor performance in directly performing gaze estimation in the infrared scenario. (2) It is not difficult to collect infrared imaging photos, but it is very difficult and costly to label these photos with corresponding data tags. Collecting data tags means annotating the direction information of the gaze in space for the photos. First, this direction is a three-dimensional vector in the world coordinate system, and it is not easy to measure this vector. Secondly, to map this vector to the two-dimensional space of the photo, the conversion relationship between the camera coordinate system and the world coordinate system is also required. The existing open data set acquisition schemes are all very complex, resulting in very little gaze regression data in the DMS scenario and being difficult to generate. Summary of the Invention

[0004] Therefore, the technical problem to be solved by the present invention is to overcome the phenomenon that the line-of-sight regression data in the DMS scenario is very scarce and difficult to generate, so as to provide a technical solution for generating line-of-sight regression infrared imaging sample data.

[0005] In a first aspect, according to the method for generating line-of-sight regression infrared imaging sample data provided by an embodiment of the present invention, it includes:

[0006] Use 3D modeling software to establish an eye model and generate multiple 3D pictures;

[0007] Use a pre-trained GAN model to render the 3D pictures to generate infrared virtual pictures;

[0008] Input real DMS scenario pictures into a pre-trained line-of-sight regression model to obtain line-of-sight regression parameters in the real DMS scenario; the line-of-sight regression model is trained by the infrared virtual pictures;

[0009] Wherein, the line-of-sight regression parameters include the X coordinate and Y coordinate of the pupil center point, and the pitch angle and heading angle in the line-of-sight direction.

[0010] Preferably, the training method of the pre-trained line-of-sight regression model is:

[0011] Use 3D modeling software to establish an eye model, generate multiple 3D pictures, and render the generated 3D pictures to obtain multiple infrared virtual pictures, and calculate the first label information of each infrared virtual picture;

[0012] Select the first part of the infrared virtual pictures as the input of the line-of-sight regression model, and calculate the second label information corresponding to the input infrared virtual pictures through the line-of-sight regression model;

[0013] Use the mean square error as the first loss function to determine the mean square error between the first label information and the second label information of the first part of the infrared virtual pictures;

[0014] When the mean square error meets the preset requirements, select the second part of the infrared virtual pictures as test samples to test the line-of-sight regression model;

[0015] When the test result meets the second requirement, determine that the line-of-sight regression model is a trained line-of-sight regression model.

[0016] Preferably, the line-of-sight regression model is formed by stacking several aggregation blocks and then going through the Fc4 process;

[0017] Wherein, between two adjacent aggregation blocks, the output of the previous aggregation block is used as the input of the next aggregation block.

[0018] Preferably, the aggregation block is:

[0019] Perform a 3×3 convolution on the input sample to obtain a first convolutional sample;

[0020] Retain a copy of the first convolutional sample, and perform a 3×3 operation on the first convolutional sample to obtain a second convolutional sample;

[0021] Perform a 3×3 convolution on the copy of the first convolutional sample and the second convolutional sample to obtain an output sample.

[0022] Preferably, the training method of the pre-trained GAN model is as follows:

[0023] Alternately optimize the generator network and the discriminator network;

[0024] Among them, the second loss function is used to optimize the generator network, and the third loss function is used to optimize the discriminator network;

[0025] When the values of the second loss function and the third loss function reach the corresponding requirements, determine that the GAN model composed of the corresponding generator network and discriminator network is a trained GAN model.

[0026] Preferably,

[0027] The second loss function is:

[0028]

[0029] The third loss function is:

[0030]

[0031] Among them, m is the number of pictures input in each batch when training the GAN model, X (i) is the i-th picture in each batch, and θ g is the generator network parameter.

[0032] Preferably, the method of establishing an eye model by using 3D modeling software and generating a plurality of 3D pictures includes:

[0033] Obtain the set camera parameters and eye parameters;

[0034] Based on the eye model established by the 3D modeling software, combine the camera parameters and the eye parameters to generate 3D pictures corresponding to the camera parameters and the eye parameters.

[0035] According to an apparatus for generating gaze regression infrared imaging sample data provided by an embodiment of the present invention, it includes:

[0036] A 3D picture generation module, configured to establish an eye model by using 3D modeling software and generate a plurality of 3D pictures;

[0037] A rendering module, configured to render the 3D picture by using a pre-trained GAN model to generate an infrared virtual picture;

[0038] A gaze regression data generation module, configured to input a real DMS scene picture into a pre-trained gaze regression model to obtain gaze regression parameters in a real DMS scene; the gaze regression model is trained by using the infrared virtual picture;

[0039] Wherein, the gaze regression parameters include the X coordinate and Y coordinate of the center point of the pupil, and the pitch angle and yaw angle in the gaze direction.

[0040] In a third aspect, according to an embodiment of the present invention, a device for generating gaze regression infrared imaging sample data is provided, including a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the method for generating gaze regression infrared imaging sample data according to any one of the above.

[0041] In a fourth aspect, according to an embodiment of the present invention, a computer-readable storage medium is provided, characterized in that the computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method for generating gaze regression infrared imaging sample data according to any one of the above.

[0042] The method, device and storage medium for generating gaze regression infrared imaging sample data provided by the embodiments of the present invention have at least the following beneficial effects:

[0043] By using the method provided by the embodiment of the present invention, an eye model is established through 3D modeling software and a plurality of 3D pictures are generated, and the 3D pictures are rendered by using a pre-trained GAN model to generate infrared virtual pictures, and then the infrared virtual pictures are input into a pre-trained gaze regression model to obtain gaze regression parameters in a real DMS scene. A large number of 3D pictures can be generated through 3D modeling software, effectively overcoming the technical problems that the gaze regression data in the DMS scene is very small and difficult to generate in the prior art solution.

[0044] In addition, based on the 3D modeling software, the label information of the 3D pictures can be directly calculated without manually calculating the label information of each 3D picture, and the label information of each 3D picture must be used in the training of the gaze regression model. Therefore, this solution lays a foundation for the subsequent training to form a gaze regression model, and effectively improves the efficiency of training to form a gaze regression model that meets the requirements. Description of the Drawings

[0045] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0046] Figure 1a It is a schematic diagram of a relatively dark pupil after converting a color picture into a black-and-white picture;

[0047] Figure 1b It is a schematic diagram of a relatively bright pupil in an infrared imaging picture;

[0048] Figure 2 It is a flowchart of a method for generating gaze regression infrared imaging sample data provided by an embodiment of the present invention;

[0049] Figure 3 It is a schematic diagram of the label information in a 3D picture in a 3D modeling software in an embodiment of the present invention;

[0050] Figure 4 It is a flowchart of the training method of a pre-trained gaze regression model in an embodiment of the present invention;

[0051] Figure 5 It is a schematic diagram of a gaze regression model formed by stacking a number of aggregation blocks in an embodiment of the present invention;

[0052] Figure 6 It is a flowchart of the operations performed by each aggregation block;

[0053] Figure 7 It is a schematic diagram of generating an infrared image using a generator network in a specific example;

[0054] Figure 8 It is a schematic diagram of training a trained GAN model through a generator network and a discriminator network;

[0055] Figure 9a It is a schematic diagram of the infrared image obtained after rendering a 3D picture using a trained GAN model;

[0056] Figure 9b It is a schematic diagram of an infrared virtual image generated after rendering a generated 3D picture using a trained GAN model;

[0057] Figure 10 It is a module diagram of a device for generating gaze regression infrared imaging sample data provided by an embodiment of the present invention;

[0058] Figure 11Bus diagram of a device for generating line-of-sight regression infrared imaging sample data provided by an embodiment of the present invention. Detailed implementation manners

[0059] The technical solutions of the present invention will be described clearly and completely below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0060] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0061] Embodiment 1

[0062] An embodiment of the present invention provides a method for generating line-of-sight regression infrared imaging sample data. As shown in Figure 2 the following, it includes:

[0063] Step S22: Establish an eye model using 3D modeling software and generate a plurality of 3D pictures;

[0064] In the embodiment of the present invention, an eye model is established using 3D modeling software. Preferably, the embodiment of the present invention takes a monocular picture as an example for illustration. The specific modeling software can be selected according to actual needs. For example, the open-source software Unity Eye of the University of Cambridge can be used to manufacture 3D pictures of the eye. Since the eye model has been established based on the software, only the camera parameters and eye parameters need to be set, and the camera parameters and eye parameters can be set according to actual needs. For example, 0, 0, 20, 40, 0, 0, 30, 30 can be set in sequence. It should be noted here that the specific camera parameters and eye parameters can be set according to actual needs, and are not specific limitations on the camera parameters and eye parameters.

[0065] In the embodiment of the present invention, a preset number of 3D eye pictures can be generated according to actual needs through 3D modeling software. Moreover, in the 3D modeling software, the position and angle of the camera are known and can be obtained, and each 3D picture has a corresponding label file when generated. Specifically, the label information included in the label file is shown in Figure 3 the following. The label information of these pictures is recorded through the label file for subsequent generation of line-of-sight regression infrared imaging sample data according to the label information, without additional manual annotation. Obtaining the label information of 3D pictures in this way can save a large amount of annotation costs.

[0066] In the embodiment of the present invention, the iris key point data can be obtained according to the label information "iris_2d"; and then the X coordinate and Y coordinate of the center point of the through hole can be obtained.

[0067] The data recorded in the label information "look_vce" is the line-of-sight direction data. It should be noted that the line-of-sight direction data recorded in the label information "look_vce" is in the world coordinate system and can be converted to the camera coordinate system through the conversion formula (1-1) in the configured virtual camera, so as to obtain the pitch angle (also known as the pitch angle) and the heading angle (also known as the yaw angle) in the required line-of-sight direction. Among them, the conversion formula (1-1) is:

[0068]

[0069] In formula (1-1), X, Y, and Z are the three coordinates of the world coordinate system, Xc, Yc, and Zc are the three coordinates in the camera coordinate system respectively, and Lw is the external camera parameter matrix. Among them, R represents the rotation matrix and t represents the offset vector.

[0070] is the transpose matrix.

[0071] Step S24: Render the 3D picture using the pre-trained GAN model to generate an infrared virtual picture;

[0072] Step S26: Input the real DMS scene picture into the pre-trained line-of-sight regression model to obtain the line-of-sight regression parameters in the real DMS scene; the line-of-sight regression model is trained through the infrared virtual picture;

[0073] Among them, the line-of-sight regression parameters include the X coordinate and Y coordinate of the center point of the pupil, and the pitch angle (pitch angle) and heading angle (yaw angle) in the line-of-sight direction.

[0074] In the embodiment of the present invention, refer to Figure 4 As shown, the training method of the pre-trained line-of-sight regression model is:

[0075] Step S41: Use 3D modeling software to establish an eye model, generate multiple 3D pictures, render the generated 3D pictures to obtain multiple infrared virtual pictures, and calculate the first label information of each infrared virtual picture;

[0076] Step S42: Select the first part of the infrared virtual pictures as the input of the line-of-sight regression model, and calculate the second label information corresponding to the input infrared virtual pictures through the line-of-sight regression model;

[0077] Step S43: Use the mean square error as the first loss function to determine the mean square error between the first label information and the second label information of the first part of the infrared virtual pictures;

[0078] Step S44: When the mean square error meets the preset requirements, select the second part of the infrared virtual pictures as test samples to test the line-of-sight regression model;

[0079] Step S45: When the test result meets the second requirement, determine that the line-of-sight regression model is a trained line-of-sight regression model.

[0080] In the embodiment of the present invention, refer to Figure 5 As shown, the line-of-sight regression model is formed by stacking several aggregation blocks and then going through the Fc4 process; among them, Fc4 refers to a fully connected layer with an output dimension of 4.

[0081] Among them, in two adjacent aggregation blocks, the output of the previous aggregation block is used as the input of the next aggregation block. Here, it is pointed out that in the figure, six aggregation blocks including block1, …… block6 are taken as an example for display, which is not a specific limitation on the number of aggregation blocks, and the number of aggregation blocks can be set according to actual needs.

[0082] Further, refer to Figure 6 As shown, in the embodiment of the present invention, the operations performed by the aggregation block are as follows:

[0083] Step S61: Perform 3*3 convolution on the input sample to obtain the first convolution sample;

[0084] Step S62: Retain a copy of the first convolution sample, and perform 3*3 on the first convolution sample to obtain the second convolution sample;

[0085] Step S63: Perform 3*3 convolution on the copy of the first convolution sample and the second convolution sample to obtain the output sample.

[0086] Here, it is pointed out that the convolution size, order, and number in each aggregation block in this application are the same. If only the convolution size, number, or order is changed through a limited number of trials without making a creative change to the technical solution of this application, it still belongs to the protection scope of the embodiment of the present invention.

[0087] In the embodiment of the present invention, the training method of the pre-trained GAN model is as follows:

[0088] 1) Alternately optimize the generator network (Generator) and the discriminator network (Discriminator); among them, the second loss function is used to optimize the generator network, and the third loss function is used to optimize the discriminator network;

[0089] 2) When the values of the second loss function and the third loss function reach the corresponding requirements, determine that the GAN model composed of the corresponding generator network and discriminator network is a trained GAN model.

[0090] In the embodiment of the present invention, the method for generating an infrared virtual picture by a generator network is to perform a series of convolutional layers on the input picture to obtain a matrix, which is the finally output image. See Figure 7 For example, a random vector z with a length of 100 is passed through a series of convolutional layers (Cov1, Cov2, Cov3, and Cov4) to obtain a 64*64 matrix, which is the finally output image.

[0091] The process of rendering a 3D picture into an infrared virtual picture by using a generator network and a discriminator network is shown in Figure 8 The Y-style picture after the style picture z is converted by the generator is shown. Then, the converted Y-style picture and the real Y-style picture are input into the discriminator network, and the discriminator network and the generator network are alternately optimized to determine the final GAN model.

[0092] Furthermore, in the implementation of the present invention, the second loss function corresponding to the generator is:

[0093]

[0094] And the third loss function corresponding to the discriminator is:

[0095]

[0096] where m is the number of pictures input in each batch when training the GAN model, X (i) is the i-th picture in each batch, and θ g is the generator network parameter.

[0097] By optimizing the generator network and the discriminator network through the first loss function and the second loss function, a generator network and a discriminator network that meet the requirements can be determined. Furthermore, a GAN model composed of the generator network and the discriminator network can be determined, which is the trained GAN model.

[0098] Specifically, the 3D pictures generated by a 3D modeling software are as Figure 9a , and the infrared virtual pictures generated by rendering the generated 3D pictures by the trained GAN model are shown in Figure 9b as follows.

[0099] In the embodiment of the present invention, the method of using a 3D modeling software to establish an eye model and generate multiple 3D pictures includes:

[0100] 1) Obtain the set camera parameters and eye parameters;

[0101] 2) Based on the eye model established by 3D modeling software, combined with the camera parameters and the eye parameters, generate 3D pictures corresponding to the camera parameters and the eye parameters.

[0102] Embodiment 2

[0103] The embodiment of the present invention also provides a device for generating line-of-sight regression infrared imaging sample data. Refer to Figure 10 as shown, including:

[0104] A 3D picture generation module 101, configured to establish an eye model using 3D modeling software and generate a plurality of 3D pictures;

[0105] A rendering module 102, configured to render the 3D pictures using a pre-trained GAN model to generate infrared virtual pictures;

[0106] A line-of-sight regression data generation module 103, configured to input a real DMS scene picture into a pre-trained line-of-sight regression model to obtain line-of-sight regression parameters in a real DMS scene; the line-of-sight regression model is trained by the infrared virtual pictures;

[0107] Wherein, the line-of-sight regression parameters include the X coordinate and Y coordinate of the pupil center point, and the pitch angle and heading angle in the line-of-sight direction.

[0108] Preferably, it further includes a line-of-sight regression model training module, configured to train and form a pre-trained line-of-sight regression model. The specific training method is:

[0109] Establish an eye model using 3D modeling software, generate a plurality of 3D pictures, and calculate the first label information of each 3D picture;

[0110] Select the first part of the 3D pictures as the input of the line-of-sight regression model, and calculate the second label information corresponding to the input 3D pictures through the line-of-sight regression model;

[0111] Use the mean square error as the loss function to determine the mean square error between the first label information and the second label information of the selected 3D pictures;

[0112] When the mean square error meets the preset requirements, select the second part of the 3D pictures as test samples to test the line-of-sight regression model;

[0113] When the test result meets the second requirement, determine that the line-of-sight regression model is a trained line-of-sight regression model.

[0114] Preferably, the line-of-sight regression model is formed by stacking several aggregation blocks and then going through the Fc4 process; where Fc4 refers to a fully connected layer with an output dimension of 4 dimensions.

[0115] Among them, in two adjacent aggregation blocks, the output of the previous aggregation block serves as the input of the subsequent aggregation block.

[0116] Furthermore, the aggregation block is as follows:

[0117] Perform a 3×3 convolution on the input sample to obtain a first convolutional sample;

[0118] Retain a copy of the first convolutional sample, and perform 3×3 on the first convolutional sample to obtain a second convolutional sample;

[0119] Perform a 3×3 convolution on the copy of the first convolutional sample and the second convolutional sample to obtain an output sample.

[0120] Preferably, it further includes a GAN model training module for training and forming a pre-trained GAN model. The specific training method is as follows:

[0121] Alternately optimize the generator network and the discriminator network;

[0122] Among them, the second loss function is used to optimize the generator network, and the third loss function is used to optimize the discriminator network;

[0123] When the values of the second loss function and the third loss function reach the corresponding requirements, determine that the GAN model composed of the corresponding generator network and discriminator network is a trained GAN model.

[0124] Specifically,

[0125] The second loss function is:

[0126]

[0127] The third loss function is:

[0128]

[0129] Among them, m is the number of pictures input in each batch when training the GAN model, X (i) is the i-th picture in each batch, and θ g is the parameter of the generator network.

[0130] Preferably, the 3D picture generation module is further used for:

[0131] Obtain the set camera parameters and eye parameters;

[0132] Based on the eye model established by 3D modeling software, combine the camera parameters and the eye parameters to generate a 3D picture corresponding to the camera parameters and the eye parameters.

[0133] Embodiment 3

[0134] This embodiment provides a device for generating line-of-sight regression infrared imaging sample data, such as Figure 11 shown. The device for generating line-of-sight regression infrared imaging sample data includes a processor 111 and a memory 112, where the processor 111 and the memory 112 can be connected through a bus or other means. Figure 11 Taking the connection through the bus as an example.

[0135] The processor 111 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), graphics processing units (GPUs), embedded neural network processors (NPUs), or other dedicated deep learning coprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or a combination of the above types of chips.

[0136] The memory 112, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for generating line-of-sight regression infrared imaging sample data in the embodiments of the present invention (such as Figure 10 shown 3D picture generation module 101, rendering module 102, line-of-sight regression data generation module 103). The processor 111 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory 112, that is, implements the method for generating line-of-sight regression infrared imaging sample data in the above method embodiments.

[0137] The memory 112 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created by the processor 111 and the like. In addition, the memory 112 may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 112 may optionally include a memory remotely disposed relative to the processor 111, and these remote memories may be connected to the processor 111 through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0138] The one or more modules are stored in the memory 112 and, when executed by the processor 111, perform as Figure 2 shown in the method for generating line-of-sight regression infrared imaging sample data.

[0139] In this embodiment, the memory 112 stores program instructions or modules of the method for generating line-of-sight regression infrared imaging sample data. When the processor 111 executes the program instructions or modules stored in the memory 112, it uses 3D modeling software to establish an eye model and generate a plurality of 3D pictures; uses a pre-trained GAN model to render the 3D pictures to generate infrared virtual pictures; inputs the infrared virtual pictures into a pre-trained line-of-sight regression model to obtain line-of-sight regression parameters in a real DMS scenario. The line-of-sight regression parameters include the X coordinate and Y coordinate of the center point of the pupil, and the pitch angle (pitch angle) and yaw angle (yaw angle) in the line-of-sight direction. Thus, it effectively overcomes the technical problems in the prior art solutions that result in very little line-of-sight regression data in the DMS scenario and is difficult to generate.

[0140] The embodiment of the present invention also provides a non-transitory computer storage medium. The computer storage medium stores computer-executable instructions, and these computer-executable instructions can execute the standard dynamic monitoring method in any of the above method embodiments. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (abbreviation: HDD), or a solid-state drive (SSD), etc.; the storage medium may also include a combination of the above types of memories.

[0141] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, an apparatus, or a computer-readable storage medium, all of which may involve or include a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.

[0142] Obviously, the above embodiments are merely examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A method for generating line-of-sight regression infrared imaging sample data, characterized in that, it includes: Using 3D modeling software to build an eye model and generate multiple 3D pictures; Using a pre-trained GAN model to render the 3D pictures to generate infrared virtual pictures; Inputting real DMS scene pictures into a pre-trained line-of-sight regression model to obtain line-of-sight regression parameters in the real DMS scene; The line-of-sight regression model is trained through the infrared virtual pictures; Wherein, the line-of-sight regression parameters include the X coordinate and Y coordinate of the pupil center point, and the pitch angle and heading angle in the line-of-sight direction; The DMS is a fatigue driving warning system; The step of using 3D modeling software to build an eye model and generate multiple 3D pictures includes: Obtaining the set camera parameters and eye parameters; Based on the eye model established by the 3D modeling software, combining the camera parameters and the eye parameters to generate 3D pictures corresponding to the camera parameters and the eye parameters; The training method of the pre-trained line-of-sight regression model is: Using 3D modeling software to build an eye model, generating multiple 3D pictures, rendering the generated 3D pictures to obtain multiple infrared virtual pictures, and calculating the first label information of each infrared virtual picture; Selecting a first part of the infrared virtual pictures as the input of the line-of-sight regression model, and calculating the second label information corresponding to the input infrared virtual pictures through the line-of-sight regression model; Using the mean square error as the first loss function to determine the mean square error between the first label information and the second label information of the first part of the infrared virtual pictures; When the mean square error meets the preset requirements, selecting a second part of the infrared virtual pictures as test samples to test the line-of-sight regression model; When the test result meets the second requirement, determining the line-of-sight regression model as the trained line-of-sight regression model.

2. The method according to claim 1, characterized in that, The line-of-sight regression model is formed by stacking several aggregation blocks and then going through the Fc4 process; Wherein, in two adjacent aggregation blocks, the output of the previous aggregation block is used as the input of the next aggregation block.

3. The method according to claim 2, characterized in that, The aggregation block is: Performing 3*3 convolution on the input sample to obtain a first convolution sample; Retaining a copy of the first convolution sample and performing 3*3 on the first convolution sample to obtain a second convolution sample; Performing 3*3 convolution on the copy of the first convolution sample and the second convolution sample to obtain an output sample.

4. The method according to claim 1, characterized in that, The training method of the pre-trained GAN model is: Alternately optimizing the generator network and the discriminator network; Wherein, the second loss function is used to optimize the generator network, and the third loss function is used to optimize the discriminator network; When the values of the second loss function and the third loss function reach the corresponding requirements, determining the GAN model composed of the corresponding generator network and discriminator network as the trained GAN model.

5. The method according to claim 4, characterized in that, The second loss function is: The third loss function is: Among them, m is the number of pictures input in each batch during the training to form the GAN model, and X (i) is the i-th picture in each batch, and θ g is the generator network parameter.

6. An apparatus for generating gaze regression infrared imaging sample data, characterized in that, comprising: a 3D image generation module for establishing an eye model using 3D modeling software and generating a plurality of 3D images; a rendering module for rendering the 3D images using a pre-trained GAN model to generate infrared virtual images; a gaze regression data generation module for inputting a real DMS scene image into a pre-trained gaze regression model to obtain gaze regression parameters in a real DMS scene; the gaze regression model is trained by the infrared virtual images; wherein, the gaze regression parameters include the X coordinate and Y coordinate of the center point of the pupil, and the pitch angle and yaw angle in the gaze direction; the DMS is a driver fatigue warning system; the establishing an eye model using 3D modeling software and generating a plurality of 3D images includes: obtaining set camera parameters and eye parameters; generating 3D images corresponding to the camera parameters and the eye parameters based on the eye model established by the 3D modeling software in combination with the camera parameters and the eye parameters; the training method of the pre-trained gaze regression model is: establishing an eye model using 3D modeling software, generating a plurality of 3D images, rendering the generated 3D images to obtain a plurality of infrared virtual images, and calculating first label information of each infrared virtual image; selecting a first part of the infrared virtual images as the input of the gaze regression model, and calculating second label information corresponding to the input infrared virtual images through the gaze regression model; using the mean square error as the first loss function to determine the mean square error between the first label information and the second label information of the first part of the infrared virtual images; when the mean square error meets a preset requirement, selecting a second part of the infrared virtual images as test samples to test the gaze regression model; when the test result meets a second requirement, determining the gaze regression model as a trained gaze regression model.

7. An apparatus for generating gaze regression infrared imaging sample data, characterized in that, comprising a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the method for generating gaze regression infrared imaging sample data according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores computer instructions for causing a computer to execute the method for generating gaze regression infrared imaging sample data according to any one of claims 1-5.

Citation Information

Patent Citations

  • Sight line estimation method based on depth regression network

    CN106599994A

  • Sight line tracking and training methods, apparatuses and systems, electronic device and storage medium

    CN108229284A