A video pupil positioning and tracking method and device

Through the method of combining adversarial network and segmentation network, pupil positioning tracking is used to use the correlation of adjacent frames, which solves the problems of insufficient accuracy and large labeling in the prior art, and achieves high-precision and real-time pupil positioning.

CN114708294BActive Publication Date: 2025-07-29BEIJING SHENRUI BOLIAN TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210147276.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-17
Publication Date
2025-07-29
Estimated Expiration
2042-02-17

AI Technical Summary

Technical Problem

The prior art has insufficient accuracy in the positioning of human pupils, especially in the case of light changes and occlusion, and insufficient real-time and anti-object occlusion ability, and requires a large amount of image annotation.

Method used

Using a combination of adversarial network and segmentation network, by obtaining the eye image of the current frame and the previous frame for prediction, and using the correlation between adjacent frames to realize automatic positioning and tracking of the eye pupil, only the first two frames of images need to be marked.

Benefits of technology

It effectively removes the influence of noise, realizes high-precision pupil positioning, reduces the number of image labels, and improves real-time and anti-occlusion capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708294B_ABST
    Figure CN114708294B_ABST
Patent Text Reader

Abstract

The present invention provides a video pupil positioning and tracking method and device. The method includes: obtaining the current frame eye image Hi, inputting the previous frame eye images Hi−1 and Hi into a confrontation network to obtain the predicted next frame eye image Hi+1; inputting Hi+1 into a first segmentation network to obtain the segmentation mask Yi+1 of the predicted next frame pupil; inputting Hi and Yi into a second segmentation network to obtain the pupil segmentation mask of the current frame; repeating the above steps until all frame eye images are processed. By setting up a confrontation network to predict the next frame eye image using the correlation between adjacent two frame eye images, the present invention can effectively remove the influence of noise and obtain the position of the segmentation mask of the predicted next frame pupil. By using the correlation between adjacent front and rear two frame images, only the first two frame images need to be labeled to achieve automatic tracking of the pupil, greatly reducing the number of labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video tracking, and particularly relates to a video eye pupil positioning and tracking method and device. Background Technique

[0002] Eye pupil positioning not only plays a very important role in safe driving detection, but also plays a crucial role in fields such as VR (Virtual Reality) and AR (Augmented Reality). Eye pupil positioning generally includes the following three steps: face detection, human eye area positioning, and human eye pupil positioning.

[0003] Existing pupil positioning methods generally adopt traditional image processing methods, such as the hough transform method, the ellipse fitting method, and the gradient vector method, etc. Although these traditional image processing methods have certain advantages in processing speed, their accuracy is not satisfactory; especially when the human eye area is affected by light changes and occlusion, it is very difficult for traditional image processing methods to locate the position of the pupil, thus greatly reducing the positioning accuracy. The invention patent with the application number 202010263340.9 proposes a method for positioning eye pupils based on deep learning. Under a mature face detection model and a face feature point detection model, this method establishes an eye pupil positioning model by combining the powerful feature learning ability of a deep neural network. Since this method only simply utilizes the feature learning ability of the deep neural network and establishes an eye pupil positioning model based on the training of a single image, without considering the correlation information between the upper and lower frames of the video, although the algorithm accuracy has been improved, the real-time performance, the ability to resist object occlusion, etc. still need to be improved, and a large amount of image annotation is required. Summary of the Invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a video eye pupil positioning and tracking method and device.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions.

[0006] In the first aspect, the present invention provides a video eye pupil positioning and tracking method, including the following steps:

[0007] Obtain the current frame eye image H i , and input the previous frame eye image H i-1 and H i into an adversarial network to obtain the predicted next frame eye image H i+1 ;

[0008] Take Hi+1 Input it into the first segmentation network to obtain the predicted segmentation mask of the pupil in the next frame Y i+1 , and Y i+1 store it in the pupil mask pool;

[0009] Input H i and the predicted segmentation mask of the pupil in the current frame Y i into the second segmentation network to obtain the segmentation mask of the pupil in the current frame; i When Y i = 1, 2, i When Y i obtain it from the pupil mask pool;

[0010] Repeat the above steps until all frame eye images are processed.

[0011] Furthermore, the method further includes: after obtaining the fine segmentation mask of the pupil in the current frame, delete Y i from the pupil mask pool.

[0012] Furthermore, the pupil masks H 1 of the first frame eye image H and the second frame eye image Y 1, Y 2 are obtained through manual annotation, or are automatically obtained by respectively inputting H 1, H 2 into a trained segmentation network model.

[0013] Even further, the adversarial network is a CycleGAN network.

[0014] Furthermore, both the first segmentation network and the second segmentation network are UNet networks.

[0015] In a second aspect, the present invention provides a video pupil positioning and tracking device, including:

[0016] An image prediction module, configured to obtain the current frame eye image H i , input the previous frame eye image H i-1 and H i into the adversarial network to obtain the predicted next frame eye image H i+1 ;

[0017] A rough segmentation module for H i+1 inputting into the first segmentation network to obtain a predicted segmentation mask of the pupil in the next frame Y i+1 and storing Y i+1 it in the pupil mask pool;

[0018] A fine segmentation module for H i inputting and the predicted segmentation mask of the pupil in the current frame Y i into the second segmentation network to obtain the pupil segmentation mask of the current frame; i When = 1, 2, Y i it is a calibrated pupil segmentation mask; i When > 2, Y i it is obtained from the pupil mask pool;

[0019] A loop execution module for repeating the above steps until all frame eye images are processed.

[0020] Furthermore, the device further includes a mask pool update module for deleting Y i from the pupil mask pool after obtaining the fine pupil segmentation mask of the current frame.

[0021] Furthermore, the pupil masks H 1 of the first frame eye image H and the pupil masks Y 1, Y 2 of the second frame eye image H 1, H 2 are obtained through manual annotation or by automatically inputting

[0022] 1,

[0023] 2 into a trained segmentation network model respectively.

[0024] Compared with the prior art, the present invention has the following beneficial effects.

[0025] The present invention realizes automatic positioning and tracking of the pupil by obtaining the current-frame eye image, inputting the previous-frame eye image and the current-frame eye image into an adversarial network to obtain the predicted next-frame eye image, inputting the predicted next-frame eye image into a first segmentation network to obtain the segmentation mask of the predicted next-frame pupil and storing it in the pupil mask pool, inputting the predicted segmentation mask of the current-frame pupil and the current-frame eye image into a second segmentation network to obtain the segmentation mask of the current-frame pupil, and repeating the above steps until all frame eye images are processed. By setting up an adversarial network, the present invention utilizes the correlation between adjacent two-frame eye images to predict the next-frame eye image, thereby obtaining the position of the segmentation mask of the predicted next-frame pupil, which can effectively remove the influence of noise. By utilizing the correlation between adjacent front and rear two-frame images, only the first two frames of images need to be labeled to realize automatic tracking of the pupil, greatly reducing the number of labels. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a flowchart of a video pupil positioning and tracking method according to an embodiment of the present invention.

[0027] Figure 2 It is a schematic diagram of the overall structure of the network model according to an embodiment of the present invention.

[0028] Figure 3 It is a schematic diagram of the structure of the generative adversarial network GAN.

[0029] Figure 4 It is a schematic diagram of the structure of the CycleGAN network.

[0030] Figure 5 It is a schematic diagram of the structure of the UNet network.

[0031] Figure 6 It is a block diagram of a video pupil positioning and tracking device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the scope of protection of the present invention.

[0033] Figure 1 It is a flowchart of a video pupil positioning and tracking method according to an embodiment of the present invention, including the following steps:

[0034] Step 101, obtain the current-frame eye image H i , the previous-frame eye image Hi-1 and H i are input into an adversarial network to obtain a predicted next-frame eye image H i+1 ;

[0035] Step 102: Input H i+1 into a first segmentation network to obtain a predicted segmentation mask of the next-frame eye pupil Y i+1 , and Y i+1 store it in the eye pupil mask pool;

[0036] Step 103: Input H i and the predicted segmentation mask of the current-frame eye pupil Y i into a second segmentation network to obtain the segmentation mask of the current-frame eye pupil; i When Y i is a calibrated segmentation mask of the eye pupil; i When Y i obtain it from the eye pupil mask pool;

[0037] Step 104: Repeat Steps 101-103 until all frame eye images are processed.

[0038] This embodiment provides a video eye pupil positioning and tracking method, and the method is implemented by executing Steps 101-104. The overall network model structure for implementing the method is as shown in Figure 2 shown.

[0039] In this embodiment, Step 101 is mainly used to predict the next-frame eye image based on the eye images of the previous two frames. Based on the correlation between adjacent two-frame video images, this embodiment predicts the next-frame eye image by inputting them into an adversarial network. The generative adversarial network GAN (Generative Adversarial Networks) is a typical adversarial network, and its structural schematic diagram is as shown in Figure 3As shown in the figure, it mainly consists of two neural networks: the generator network G (Generator) and the discriminator network D (Discriminator). G is a network that generates images. It receives a random noise z and generates an image through this noise, denoted as G(z). D is a discriminative network that determines whether an image is "real". Its input parameter is x, where x represents an image, and the output D(x) represents the probability that x is a real image. If the output is 1, it means the image is 100% real; if the output is 0, it means the image cannot be real. During the training process, the goal of the generator network G is to generate as real images as possible to deceive the discriminator network D; while the goal of D is to distinguish the images generated by G from the real images. In this way, G and D constitute a dynamic game process or an adversarial process.

[0040] In this embodiment, step 102 is mainly used to generate a segmentation mask of the pupil in the predicted next-frame eye image H i+1 Y i+1 In this embodiment, by inputting the eye image H i+1 into a segmentation network, namely the first segmentation network, a segmentation mask is obtained Y i+1 The input of the first segmentation network is a single-frame eye image, and it outputs a rough segmentation mask. The rough segmentation mask output by the first segmentation network Y i+1 is stored in the pupil mask pool and used as an input for the second segmentation network to determine the approximate position of the pupil mask, enabling the second segmentation network to output a fine pupil segmentation mask. The pupil mask pool is a memory for storing rough segmentation masks, as Figure 2 shown.

[0041] In this embodiment, step 103 is mainly used to obtain the segmentation mask of the pupil in the current frame. In this embodiment, by inputting the current-frame eye image H i and the segmentation mask of the predicted pupil in the current frame Y i into the second segmentation network, the segmentation mask of the pupil in the current frame is obtained. Different from the first segmentation network, the second segmentation network has two input terminals. One input is the current-frame eye image H i , and the other is the segmentation mask of the predicted pupil in the current frame Y i ​As described above, since the approximate position of the pupil mask can be given, the second segmentation network processes the current frame of eye image and the segmentation mask of the currently measured pupil of the current frame, and can output a high-precision, that is, a fine pupil segmentation mask. It should be noted that only the first two frames of eye images (the first frame and the second frame) need to be labeled, that is, the segmentation mask of the pupil is marked; starting from the third frame of eye image, labeling is no longer required. Instead, the adversarial network generates a predicted eye image for the next frame based on the first two frames of eye images, and then obtains its rough pupil segmentation mask, which is then input into the second segmentation network together with the input current frame of eye image to obtain the pupil segmentation mask of the current frame, realizing the positioning and tracking of the pupil. Therefore, the number of images that need to be labeled is greatly reduced (only two frames need to be labeled).

[0042] In this embodiment, step 104 is mainly used to implement pupil tracking for all frames of eye images by repeatedly executing steps 101 to 103. Each time steps 101 to 103 are executed, a new frame of eye image is input and i is updated, that is, i is replaced with i +1 until all the eye images that need to be processed are processed.

[0043] As an optional embodiment, the method further includes: after obtaining the fine pupil segmentation mask of the current frame, deleting Y i from the pupil mask pool.

[0044] This embodiment provides a technical solution for updating the pupil mask pool. As described above, the pupil mask pool is a memory for storing rough segmentation masks. In order to reduce the storage space occupied by the pupil mask pool, in this embodiment, after each execution of step 103, the predicted current pupil mask frame Y i is deleted from the pupil mask pool. After the above operation, at most two pupil mask frames, that is, Y i , Y i+1 , are stored in the pupil mask pool, which can greatly reduce the space occupied by storing the pupil mask.

[0045] As an optional embodiment, the pupil masks H 1 of the first frame of eye image H 2 and the second frame of eye image Y 1, Y 2 are obtained through manual labeling, or are automatically obtained by separately inputting H 1, H 2 into a trained segmentation network model.

[0046] This embodiment provides a method for annotating the first two frames of eye images. As described above, only the first two frames of eye images need to be annotated in this embodiment. For the annotation of the first frame and the second frame of eye images, manual annotation or automatic annotation can be used. The automatic annotation method requires constructing a pupil mask segmentation network model. By inputting the first frame and the second frame of eye images into the trained segmentation network model respectively, automatic annotation can be achieved.

[0047] As an alternative embodiment, the adversarial network is a CycleGAN network.

[0048] This embodiment provides a technical solution for the adversarial network. The adversarial network in this embodiment selects a CycleGAN network. CycleGAN is a type of GAN model for cross-domain image style transfer. The CycleGAN model is based on the idea of dual learning and can learn the mapping relationship between the source domain and the target domain without one-to-one correspondence, and can perform image style transfer tasks even without a paired training set. The CycleGAN model first maps from the source domain to the target domain and then can be converted back from the target domain. In this way, the limitation of paired training images can be eliminated. For a single GAN model, the generator and the discriminator play against each other. The generator learns the data feature distribution from the sample data, and the discriminator distinguishes between real images and generated images. The generator and the discriminator are optimized through mutual adversarial training, so that data that is completely approximated to the actual distribution can be generated finally. For this training method, there is a problem in cross-domain image style transfer tasks. The network model may map the source domain to an uncertain combination in the target domain, so it is even possible to map all source domains to the same image in the target domain. Just through the separate adversarial loss, the desired output result of mapping the source domain to the target domain cannot be achieved. To solve this problem, the CycleGAN model adopts a cyclic consistency constraint. After the data in the source domain is transformed twice, it should match the data features in the source domain distribution. The CycleGAN model uses the first mapping G to transform the data in the X domain into the Y domain, and then uses the second mapping F to transform it back. In this way, the situation where the X domain may all be mapped to the same picture in the Y domain is solved, as Figure 4 shown. The structure of the CycleGAN model can be regarded as a dual-generator adversarial mode, which is like a circular network in structure.

[0049] As an alternative embodiment, both the first segmentation network and the second segmentation network are UNet networks.

[0050] In this embodiment, both the first segmentation network and the second segmentation network adopt UNet networks. The structure of the UNet network is in a "U" shape, and its structure schematic diagram is as Figure 5As shown in the figure. Unet draws on the fully convolutional neural network FCN (Fully Convolution Network). Its network structure includes two symmetric parts: the front part of the network is the same as the ordinary convolutional network, using 3×3 convolution and pooling downsampling, which can capture the context information in the image (that is, the relationship between pixels); the back part of the network is basically symmetric with the front part, using 3×3 convolution and upsampling to achieve the purpose of outputting image segmentation. In addition, feature fusion is also used in the network to fuse the features of the downsampling network in the front part with the features of the upsampling part in the back to obtain more accurate context information and achieve better segmentation effects.

[0051] Figure 6 The following is a schematic diagram of the composition of a video eye pupil positioning and tracking device according to an embodiment of the present invention. The device includes:

[0052] An image prediction module 11, configured to obtain the current frame of eye image H i , and input the previous frame of eye image H i-1 and H i into the adversarial network to obtain the predicted next frame of eye image H i+1 ;

[0053] A rough segmentation module 12, configured to input H i+1 into the first segmentation network to obtain the segmentation mask of the predicted next frame of eye pupil Y i+1 , and store Y i+1 in the eye pupil mask pool;

[0054] A fine segmentation module 13, configured to input H i and the segmentation mask of the predicted current frame of eye pupil Y i into the second segmentation network to obtain the eye pupil segmentation mask of the current frame; i When i = 1, 2, Y i is the calibrated eye pupil segmentation mask; i When i > 2, Y i is obtained from the eye pupil mask pool;

[0055] A loop execution module 14, configured to repeat the above steps until all frames of eye images are processed.

[0056] The device of this embodiment can be used to execute Figure 1The technical solutions of the method embodiments shown have similar implementation principles and technical effects, which will not be elaborated here. The same applies to the subsequent embodiments and will not be further explained.

[0057] As an optional embodiment, the device further includes a mask pool update module, which is used to delete from the pupil mask pool after obtaining the fine pupil segmentation mask of the current frame Y i .

[0058] As an optional embodiment, the first frame of eye image H 1 and the second frame of eye image H 2 of the pupil masks Y 1, Y 2 are obtained by manual annotation, or by respectively inputting H 1, H 2 into a trained segmentation network model to be automatically obtained.

[0059] As an optional embodiment, the adversarial network is a CycleGAN network.

[0060] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A video pupil positioning and tracking method, characterized in that, Including the following steps: Step 101, obtain the current frame of eye image H i , input the previous frame of eye image H i-1 and H i into the adversarial network to obtain the predicted next frame of eye image H i+1 ; Step 102: Input H i+1 into the first segmentation network to obtain the predicted segmentation mask of the next-frame eye pupil Y i+1 , and store Y i+1 in the eye pupil mask pool; Step 103, input H i and Y i into two input ends of the second segmentation network to obtain the pupil segmentation mask of the current frame; i When Y i is the calibrated pupil segmentation mask; i When Y i is the predicted segmentation mask of the pupil of the current frame obtained from the pupil mask pool. After obtaining the pupil segmentation mask of the current frame, delete Y i from the pupil mask pool; After each execution of steps 101 to 103, a new eye image is input and updated once i , that is, use i +1 to replace i , until all eye images of all frames are processed.

2. The video pupil positioning and tracking method according to claim 1, wherein, Y 2 is obtained by manual annotation or by H automatically obtaining 2 by inputting it into a trained segmentation network model.

3. The video pupil positioning and tracking method according to claim 2, wherein The adversarial network is a CycleGAN network.

4. The video pupil positioning and tracking method according to claim 1, wherein, Both the first segmentation network and the second segmentation network are UNet networks.

5. A video pupil positioning and tracking device, characterized in that, Including: An image prediction module for obtaining the current frame of eye image H i , input the previous frame of eye image H i-1 and H i into the adversarial network to obtain the predicted next frame of eye image H i+1 ; Rough The segmentation module is used to H i+1 input into the first segmentation network to obtain the predicted segmentation mask of the next-frame eye pupil Y i+1 and Y i+1 store it in the eye pupil mask pool; The fine segmentation module inputs H i and Y i into the two input ends of the second segmentation network to obtain the pupil segmentation mask of the current frame; i When = 2, Y i is the calibrated pupil segmentation mask; i When > 2, Y i is the segmentation mask of the predicted current frame pupil obtained from the pupil mask pool. After obtaining the pupil segmentation mask of the current frame, Y i is deleted from the pupil mask pool; A loop execution module, which is used to input a new eye image and update once every time the steps corresponding to the above module are executed once i , that is, use i +1 to replace i , until all eye images of all frames are processed.

6. The video eye pupil positioning and tracking device according to claim 5, characterized in that Y 2 is obtained by manual annotation or by H automatically obtaining 2 by inputting it into a trained segmentation network model.

7. The video eye pupil positioning and tracking device according to claim 6, wherein, The adversarial network is a CycleGAN network.

8. The video pupil positioning and tracking device according to claim 5, characterized in that, Both the first segmentation network and the second segmentation network are UNet networks.

Citation Information

Patent Citations

  • Pupil positioning method based on deep learning

    CN111428680A

  • Video frame prediction method, terminal and computer storage medium

    CN108810551A

  • Real-time pupil tracking method and system based on video image

    CN112070806A

  • Real-time hair segmentation method and system

    CN112529914A

  • Liver segment segmentation method and device based on deep learning

    CN113658186A