Method and apparatus for training monocular line-of-sight redirection model, and line-of-sight redirection method and apparatus

By using a monocular gaze redirection model, the original monocular images are processed by a generator and a discriminator, which solves the problem of limited coverage of the gaze dataset, improves the reliability and training efficiency of the gaze estimation model, automatically increases the number of training samples, and simplifies the annotation process.

WO2025251559A1PCT designated stage Publication Date: 2025-12-11BEIJING MOMENTA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/135290
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-05
Filing Date
2024-11-28
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

The limited coverage of existing gaze datasets leads to low reliability of trained gaze estimation models, and the difficulty in manually annotating 3D gazes results in time-consuming and labor-intensive data collection and annotation.

Method used

By using a monocular gaze redirection model, the original monocular image is processed using a generator and a discriminator to generate redirected and reconstructed images. The model is then trained using ground truth gaze angles to automatically expand the number of training samples and the gaze range, thereby enabling automatic annotation of gazes on newly added training samples.

Benefits of technology

It improves the reliability and training efficiency of the gaze estimation model, automatically expands the number of training samples and the gaze range, and reduces the need for manual annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135290_11122025_PF_FP_ABST
    Figure CN2024135290_11122025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a method and apparatus for training a monocular line-of-sight redirection model, and a line-of-sight redirection method and apparatus. The method for training a model comprises: acquiring a training sample set; on the basis of a generator in an initial model, processing true values of line-of-sight angles of an original monocular image and a target monocular image, so as to obtain a redirected monocular image; on the basis of the generator, processing true values of line-of-sight angles of the redirected monocular image and the original monocular image, so as to obtain a reconstructed monocular image; on the basis of a discriminator in the initial model, processing the target monocular image, the redirected monocular image and the reconstructed monocular image, so as to obtain a true / false discrimination result and a predicted value of a line-of-sight angle of each of the target monocular image, the redirected monocular image and the reconstructed monocular image; on the basis of said information, calculating the current loss value of an initial monocular line-of-sight redirection model; and on the basis of the current loss value, adjusting model parameters of the initial monocular line-of-sight redirection model until a convergence condition is met, so as to obtain a target monocular line-of-sight redirection model.
Need to check novelty before this filing date? Find Prior Art

Description

Training method of monocular gaze redirection model, gaze redirection method and device TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a training method of monocular gaze redirection model, a gaze redirection method and device. BACKGROUND

[0002] Gaze redirection is a technology that changes the line of sight by changing the pupil position or opening / closing degree of the eye. Learning-based gaze estimation based on gaze dataset has made great progress, but still faces some problems: due to the acquisition device and other reasons, the acquired gaze dataset only covers a certain range of gaze angles, thereby making the reliability of the trained gaze estimation model low; three-dimensional gaze is difficult to manually annotate, and a set of complex system or process needs to be specially set to realize it. Therefore, the current collection and annotation of gaze data is time-consuming and laborious, and how to train a reliable gaze estimation model in the case of data shortage has become a problem to be solved. SUMMARY

[0003] The present application provides a training method of monocular gaze redirection model, a gaze redirection method and device, which can perform gaze redirection on the eyes in the original face image through the monocular redirection model, thereby not only automatically expanding the number of training samples and the gaze range in the training set of the gaze estimation model, but also realizing automatic annotation of the new training sample gaze, and further improving the reliability of the gaze estimation model.

[0004] The specific technical solutions are as follows:

[0005] In a first aspect, the present application provides a training method of monocular gaze redirection model, which comprises:

[0006] Obtaining a training sample set, wherein the training sample set comprises a plurality of original monocular images, a target monocular image corresponding to each original monocular image, a gaze angle true value of each original monocular image, and a gaze angle true value of each target monocular image;

[0007] Processing the gaze angle true value of each original monocular image and its corresponding target monocular image based on a generator in an initial monocular gaze redirection model, to obtain a redirection monocular image corresponding to each original monocular image;

[0008] Processing the gaze angle true value of each redirection monocular image and its corresponding original monocular image based on the generator, to obtain a reconstruction monocular image corresponding to each original monocular image;

[0009] processing each of the target monocular image, the redirected monocular image and the reconstructed monocular image based on the discriminator in the initial monocular view redirection model, to obtain a true or false discrimination result and a view angle prediction value of each of the target monocular image, the redirected monocular image and the reconstructed monocular image;

[0010] According to the true or false discrimination result of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the view angle prediction value of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the view angle true value of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the original monocular image and the reconstructed monocular image, a current loss value of the initial monocular view redirection model is calculated, wherein the view angle true value of the redirected monocular image is the view angle true value of the target monocular image, and the view angle true value of the reconstructed monocular image is the view angle true value of the original monocular image.

[0011] When the current loss value of the initial monocular view redirection model does not satisfy the convergence condition, after adjusting the model parameters of the initial monocular view redirection model, the initial monocular view redirection model after the parameter adjustment is continuously trained until the monocular view redirection model obtained by the current training satisfies the convergence condition, and the monocular view redirection model obtained by the current training is determined as a target monocular view redirection model.

[0012] Through the above scheme, it can be known that the embodiments of the present application can train the monocular view redirection capability of the generator through the view angle true value of the original monocular image and the corresponding target monocular image, train the capability of the generator to reconstruct the same eye texture and structure as the original monocular image by inputting the redirected monocular image output by the generator and the view angle true value of the corresponding original monocular image into the generator again, train the view angle prediction and true or false image discrimination capability of the discriminator through the target monocular image, the corresponding redirected monocular image and reconstructed monocular image, and continue to iterate and train the model after adjusting the model parameters through the loss value calculated based on the view angle true value and the view angle prediction value, so as to finally obtain a target monocular view redirection model capable of redirecting the view of the monocular image, so as to directly redirect the view of the eyes in the original face image based on the target monocular view redirection model in the subsequent process, so as to not only automatically expand the number of training samples and the view range in the training set of the view estimation model, but also realize the automatic labeling of the new training sample view, thereby improving the reliability of the view estimation model.

[0013] In a possible implementation, the calculation of the current loss value of the initial monocular line-of-sight redirection model according to the true-or-false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the line-of-sight angle prediction value of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the line-of-sight angle true value of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the original monocular image and the reconstructed monocular image comprises:

[0014] The calculation of the current loss value of the generator according to the true-or-false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the line-of-sight angle true value and the line-of-sight angle prediction value of the redirected monocular image, the line-of-sight angle true value and the line-of-sight angle prediction value of the reconstructed monocular image, the original monocular image and the reconstructed monocular image;

[0015] The calculation of the current loss value of the discriminator according to the true-or-false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the line-of-sight angle prediction value of the target monocular image and the line-of-sight angle true value of the target monocular image;

[0016] The determination of the current loss value of the initial monocular line-of-sight redirection model according to the current loss value of the generator and the current loss value of the discriminator.

[0017] It can be known from the above scheme that, compared with the direct calculation of the current loss value of the initial monocular line-of-sight redirection model, the current loss value of the generator and the current loss value of the discriminator are calculated first, and then the current loss value of the initial monocular line-of-sight redirection model is determined according to the current loss value of the generator and the current loss value of the discriminator, so that the accuracy of loss calculation can be improved.

[0018] In a possible implementation, the calculation of the current loss value of the generator according to the true-or-false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the line-of-sight angle true value and the line-of-sight angle prediction value of the redirected monocular image, the line-of-sight angle true value and the line-of-sight angle prediction value of the reconstructed monocular image, the original monocular image and the reconstructed monocular image comprises:

[0019] The determination of the adversarial loss of the generator according to the true-or-false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image;

[0020] The determination of the line-of-sight accuracy loss of the generator according to the line-of-sight angle true value and the line-of-sight angle prediction value of the redirected monocular image and the line-of-sight angle true value and the line-of-sight angle prediction value of the reconstructed monocular image corresponding to the redirected monocular image.

[0021] determine a reconstruction similarity loss of the generator according to the original monocular image and the corresponding reconstructed monocular image of the original monocular image;

[0022] calculate a current loss value of the generator according to the adversarial loss of the generator, the line-of-sight accuracy loss of the generator, a line-of-sight accuracy loss weight, the reconstruction similarity loss, and a reconstruction similarity loss weight.

[0023] According to the above scheme, the current loss value of the generator is determined from multiple angles such as the adversarial loss of the generator, the line-of-sight accuracy loss of the generator, the line-of-sight accuracy loss weight, the reconstruction similarity loss, and the reconstruction similarity loss weight, which can improve the accuracy of the generator loss calculation.

[0024] In a possible implementation, the calculating a current loss value of the generator according to the adversarial loss of the generator, the line-of-sight accuracy loss of the generator, a line-of-sight accuracy loss weight, the reconstruction similarity loss, and a reconstruction similarity loss weight includes:

[0025] calculating the current loss value L of the generator according to the first formula G ;

[0026] The first formula includes:

[0027] wherein, the L adv represents the adversarial loss of the generator, the γ gaze represents the line-of-sight accuracy loss weight, the represents the line-of-sight accuracy loss of the generator, the γ rec represents the reconstruction similarity loss weight, the L rec represents the reconstruction similarity loss.

[0028] In a possible implementation, the calculating a current loss value of the discriminator includes:

[0029] determining the adversarial loss of the generator according to the true or false discrimination results of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image;

[0030] determining the line-of-sight accuracy loss of the discriminator according to the line-of-sight angle predicted value of the target monocular image and the line-of-sight angle true value of the target monocular image.

[0031] The current loss value of the discriminator is calculated according to the adversarial loss of the generator, the line-of-sight accuracy loss of the discriminator, and a line-of-sight accuracy loss weight.

[0032] It can be known from the above solution that the current loss value of the discriminator is determined from multiple angles by the adversarial loss of the generator, the line-of-sight accuracy loss of the discriminator, and the line-of-sight accuracy loss weight, so that the accuracy of the loss calculation of the discriminator can be improved.

[0033] In a possible implementation, the current loss value of the discriminator is calculated according to the adversarial loss of the generator, the line-of-sight accuracy loss of the discriminator, and a line-of-sight accuracy loss weight, and includes:

[0034] The current loss value L of the discriminator is calculated according to the second formula D ;

[0035] The second formula includes:

[0036] wherein the L adv represents the adversarial loss of the generator, the γ gaze represents the line-of-sight accuracy loss weight, and the represents the line-of-sight accuracy loss of the discriminator.

[0037] In a second aspect, the embodiments of the present application provide a line-of-sight redirection method, and the method includes:

[0038] Obtaining a first original monocular image to be redirected and a target line-of-sight angle;

[0039] Processing the first original monocular image and the target line-of-sight angle based on a generator in a target monocular line-of-sight redirection model to obtain a first redirected monocular image, the target monocular line-of-sight redirection model being trained according to the method in any possible implementation of the first aspect;

[0040] Processing the first redirected monocular image based on a discriminator in the target monocular line-of-sight redirection model to obtain a line-of-sight angle of the first redirected monocular image.

[0041] It can be known from the above scheme that the embodiment of the present application can process the first original monocular image and the target gaze angle based on the generator in the pre-trained target monocular gaze redirection model to obtain a first redirected monocular image. The gaze angle of the first redirected monocular image can be obtained by processing the first redirected monocular image through the discriminator in the target monocular gaze redirection model, so that the gaze of the eye in the original face image can be redirected according to the redirected monocular image in the subsequent process. Therefore, not only the number of training samples and the gaze range in the training set of the gaze estimation model are automatically expanded, but also the automatic labeling of the new training sample gaze is realized, thereby improving the reliability of the gaze estimation model.

[0042] In a possible implementation, when the first original monocular image is a monocular image cut from an original face image to be redirected, the method further includes:

[0043] mirroring the first redirected monocular image to obtain a second redirected monocular image corresponding to a second original monocular image in the original face image and a gaze angle of the second redirected monocular image;

[0044] replacing the first original monocular image in the original face image with the first redirected monocular image, replacing the second original monocular image in the original face image with the second redirected monocular image, and marking the gaze angle of the first redirected monocular image and the gaze angle of the second redirected monocular image to obtain a redirected face image after the original face image is redirected and the gaze angle is marked.

[0045] It can be known from the above scheme that the embodiment of the present application only redirects the gaze of one eye in the original face image through the target monocular gaze redirection model, and the gaze of the other eye can be obtained by mirroring the first redirected monocular image without using the target monocular gaze redirection model for gaze redirection. Since the efficiency of mirroring is much higher than that of the target monocular gaze redirection model, the efficiency of redirecting the gaze of both eyes in the original face image is improved, thereby improving the efficiency of obtaining the redirected face image. In addition, since the redirected face image obtained by the embodiment of the present application is marked with a gaze angle, it can be directly used as a training sample of a gaze estimation model without artificial labeling or complex labeling process, thereby improving the efficiency of gaze angle labeling.

[0046] In a third aspect, the embodiment of the present application provides a device for training a monocular gaze redirection model, the device comprising:

[0047] The acquisition unit is configured to acquire a training sample set, wherein the training sample set comprises a plurality of original monocular images, a target monocular image corresponding to each of the original monocular images, a line-of-sight angle ground truth of each of the original monocular images, and a line-of-sight angle ground truth of each of the target monocular images.

[0048] The redirection unit is configured to process, based on a generator in the initial monocular line-of-sight redirection model, the line-of-sight angle ground truth of each of the original monocular images and the target monocular image corresponding thereto, to obtain a redirected monocular image corresponding to each of the original monocular images.

[0049] The reconstruction unit is configured to process, based on the generator, the line-of-sight angle ground truth of each of the redirected monocular images and the original monocular image corresponding thereto, to obtain a reconstructed monocular image corresponding to each of the original monocular images.

[0050] The discriminant prediction unit is configured to process, based on a discriminator in the initial monocular line-of-sight redirection model, each of the target monocular image, the redirected monocular image, and the reconstructed monocular image corresponding thereto, to obtain a true-or-false discrimination result and a line-of-sight angle prediction value of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image.

[0051] The loss calculation unit is configured to calculate a current loss value of the initial monocular line-of-sight redirection model according to the true-or-false discrimination result of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image, the line-of-sight angle prediction value of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image, the line-of-sight angle ground truth of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image, the original monocular image, and the reconstructed monocular image, wherein the line-of-sight angle ground truth of the redirected monocular image is the line-of-sight angle ground truth of the target monocular image, and the line-of-sight angle ground truth of the reconstructed monocular image is the line-of-sight angle ground truth of the original monocular image.

[0052] The adjustment training unit is configured to, when the current loss value of the initial monocular line-of-sight redirection model does not satisfy a convergence condition, continue to train the initial monocular line-of-sight redirection model after adjusting model parameters of the initial monocular line-of-sight redirection model, until a monocular line-of-sight redirection model obtained through current training satisfies the convergence condition, and determine the monocular line-of-sight redirection model obtained through the current training as a target monocular line-of-sight redirection model.

[0053] In a possible implementation, the loss calculation unit comprises:

[0054] a first loss calculation module, configured to calculate a current loss value of the generator according to a true or false discrimination result of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the line-of-sight angle true value and the line-of-sight angle predicted value of the redirected monocular image, the line-of-sight angle true value and the line-of-sight angle predicted value of the reconstructed monocular image, the original monocular image and the reconstructed monocular image;

[0055] a second loss calculation module, configured to calculate a current loss value of the discriminator according to the true or false discrimination result of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the line-of-sight angle predicted value of the target monocular image and the line-of-sight angle true value of the target monocular image;

[0056] a determination module, configured to determine a current loss value of the initial monocular line-of-sight redirection model according to the current loss value of the generator and the current loss value of the discriminator.

[0057] In a possible implementation, the first loss calculation module is configured to:

[0058] determine an adversarial loss of the generator according to the true or false discrimination result of each of the target monocular image, the redirected monocular image and the reconstructed monocular image;

[0059] determine a line-of-sight accuracy loss of the generator according to the line-of-sight angle true value and the line-of-sight angle predicted value of the redirected monocular image and the line-of-sight angle true value and the line-of-sight angle predicted value of the reconstructed monocular image corresponding to the redirected monocular image;

[0060] determine a reconstruction similarity loss of the generator according to the original monocular image and the reconstructed monocular image corresponding thereto;

[0061] calculate the current loss value of the generator according to the adversarial loss of the generator, the line-of-sight accuracy loss of the generator, a line-of-sight accuracy loss weight, the reconstruction similarity loss and a reconstruction similarity loss weight.

[0062] In a possible implementation, the first loss calculation module is configured to calculate the current loss value L of the generator according to a first formula G ;

[0063] The first formula comprises:

[0064] wherein the L adv represents the adversarial loss of the generator, the γ gaze represents the line-of-sight accuracy loss weight, and the represents the line-of-sight accuracy loss of the generator, the γ rec represents the reconstruction similarity loss weight, the L rec represents the reconstruction similarity loss.

[0065] In a possible implementation, the second loss calculation module is configured to:

[0066] determine, according to the true or false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, an adversarial loss of the generator;

[0067] determine, according to the line-of-sight angle prediction value of the target monocular image and the line-of-sight angle true value of the target monocular image, a line-of-sight accuracy loss of the discriminator;

[0068] calculate, according to the adversarial loss of the generator, the line-of-sight accuracy loss of the discriminator and a line-of-sight accuracy loss weight, a current loss value of the discriminator.

[0069] In a possible implementation, the second loss calculation module is configured to calculate the current loss value L D of the discriminator according to a second formula.

[0070] The second formula comprises:

[0071] wherein the L adv represents the adversarial loss of the generator, the γ gaze represents the line-of-sight accuracy loss weight, the represents the line-of-sight accuracy loss of the discriminator.

[0072] Through the above scheme, the embodiment of the present application can train the monocular line-of-sight redirection capability of the generator through the line-of-sight angle true value of the original monocular image and the corresponding target monocular image, re-input the line-of-sight angle true value of the redirected monocular image output by the generator and the corresponding original monocular image into the generator, train the capability of the generator to reconstruct the same eye texture and structure as the original monocular image, train the line-of-sight angle prediction and true or false image discrimination capability of the discriminator through the target monocular image, the corresponding redirected monocular image and the reconstructed monocular image, and continue to iterate and train the model after adjusting the model parameters through the loss value calculated by the line-of-sight angle true value, the line-of-sight angle prediction value and the like. Finally, a target monocular line-of-sight redirection model capable of redirecting the line-of-sight of a monocular image can be obtained, so as to directly redirect the line-of-sight of the eyes in the original face image based on the target monocular line-of-sight redirection model in the subsequent process. Therefore, not only the number of training samples and the line-of-sight range in the training set of the line-of-sight estimation model are automatically expanded, but also the automatic labeling of the new training sample line-of-sight is realized, thereby improving the reliability of the line-of-sight estimation model.

[0073] In a fourth aspect, the embodiment of the present application provides a line-of-sight redirection device, the device comprising:

[0074] An acquisition unit is configured to acquire a first original monocular image to be redirected and a target line-of-sight angle.

[0075] A redirection unit is configured to process the first original monocular image and the target line-of-sight angle based on a generator in a target monocular line-of-sight redirection model to obtain a first redirected monocular image, wherein the target monocular line-of-sight redirection model is trained according to the method of any possible implementation manner of the first aspect.

[0076] A line-of-sight angle prediction unit is configured to process the first redirected monocular image based on a discriminator in the target monocular line-of-sight redirection model to obtain a line-of-sight angle of the first redirected monocular image.

[0077] In a possible implementation manner, the device further comprises:

[0078] A mirror processing unit is configured to, when the first original monocular image is a monocular image cut from an original face image to be redirected, mirror process the first redirected monocular image to obtain a second redirected monocular image corresponding to a second original monocular image in the original face image and a line-of-sight angle of the second redirected monocular image.

[0079] The replacement marking unit is configured to replace the first original monocular image in the original face image with the first redirected monocular image, replace the second original monocular image in the original face image with the second redirected monocular image, and mark the line-of-sight angle of the first redirected monocular image and the line-of-sight angle of the second redirected monocular image, to obtain a redirected face image after line-of-sight redirection and line-of-sight angle marking of the original face image.

[0080] According to the above scheme, the first original monocular image and the target line-of-sight angle can be processed based on the generator in the target monocular line-of-sight redirection model, to obtain the first redirected monocular image. The line-of-sight angle of the first redirected monocular image can be obtained by processing the first redirected monocular image through the discriminator in the target monocular line-of-sight redirection model, so that the line-of-sight of the original face image can be redirected according to the redirected monocular image in the subsequent process. Therefore, not only the number of training samples and the line-of-sight range in the training set of the line-of-sight estimation model are automatically expanded, but also the automatic annotation of the line-of-sight of the new training sample is realized, thereby improving the reliability of the line-of-sight estimation model.

[0081] In a fifth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the method in any possible implementation manner of the first aspect or the second aspect.

[0082] In a sixth aspect, an electronic device is provided, and the electronic device includes:

[0083] one or more processors;

[0084] The processor is coupled with a storage device, and the storage device is configured to store one or more programs;

[0085] When the one or more programs are executed by the one or more processors, the electronic device implements the method in any possible implementation manner of the first aspect or the second aspect.

[0086] In a seventh aspect, a computer program product is provided, and the computer program product includes instructions. When the instructions are executed on a computer or a processor, the computer or the processor performs the method in any possible implementation manner of the first aspect or the second aspect. BRIEF DESCRIPTION OF DRAWINGS

[0087] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description only represent some embodiments of the present application. Those skilled in the art can also obtain other accompanying drawings without creative effort based on these accompanying drawings.

[0088] FIG. 1 is a flowchart of a training method of a monocular gaze redirection model according to an embodiment of the present application;

[0089] FIG. 2 is an architecture diagram of a monocular gaze redirection model according to an embodiment of the present application;

[0090] FIG. 3 is a flowchart of a gaze redirection method according to an embodiment of the present application;

[0091] FIG. 4 is a block diagram of a training device of a monocular gaze redirection model according to an embodiment of the present application;

[0092] FIG. 5 is a block diagram of a gaze redirection device according to an embodiment of the present application;

[0093] FIG. 6 is a structural diagram of an electronic device or computer device according to an embodiment of the present application. DETAILED DESCRIPTION

[0094] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments only represent some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort fall within the scope of protection of the present application.

[0095] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The terms “include” and “have” and any variations thereof in the embodiments of the present application and the accompanying drawings are intended to cover non-exclusive inclusion. For example, the processes, methods, systems, products or devices including a series of steps or units are not limited to the listed steps or units, but can optionally include steps or units not listed or can optionally include other steps or units inherent to these processes, methods, products or devices.

[0096] In order to improve the acquisition efficiency of the training set of the gaze estimation model and the gaze angle annotation efficiency, and expand the gaze range of the training samples, the present application provides a training method of a monocular gaze redirection model, which can be applied to an electronic device or a computer device. The method will be described in detail below with reference to the flowchart shown in FIG. 1 and the model architecture diagram shown in FIG. 2.

[0097] S110: Obtain a training sample set.

[0098] The training sample set includes multiple original monocular images, a target monocular image corresponding to each original monocular image, a line-of-sight angle true value of each original monocular image, and a line-of-sight angle true value of each target monocular image.

[0099] Each original monocular image and the corresponding target monocular image come from the same eye of the same object (such as a person, an animal, etc.) in the same head posture. In an embodiment, first, multiple images containing faces can be collected for different objects, respectively. Then, a target detection model is used to detect and cut out monocular images containing only one eye from each image, respectively. For each object, a part of the monocular images are used as original monocular images, and the other part are used as target monocular images. The line-of-sight angle true value of each monocular image is labeled. The target detection model is trained according to face images (or images containing faces) containing eye region labels and line-of-sight labels.

[0100] In the training sample set, the eyes of different objects can be different sides or the same side, for example, the monocular images of object 1 are left eyes, the monocular images of object 2 are right eyes, or the monocular images of object 1 and object 2 are right eyes.

[0101] S120: Process the line-of-sight angle true value of each original monocular image and the corresponding target monocular image based on the generator in the initial monocular line-of-sight redirection model, to obtain a redirected monocular image corresponding to each original monocular image.

[0102] The initial monocular line-of-sight redirection model includes a generator and a discriminator. The embodiment of the present application can train the generator by inputting the line-of-sight angle true value of each original monocular image and the corresponding target monocular image into the generator, and output a redirected monocular image corresponding to each original monocular image. That is, the generator is used to learn to redirect each original monocular image to an image with the line-of-sight angle true value of the target monocular image, for example, the generator learns to redirect an original monocular image with a line-of-sight angle of 30 degrees to a line-of-sight angle of 60 degrees.

[0103] As shown in FIG. 2, step (1) represents inputting the line-of-sight angle true value of the original monocular image and the corresponding target monocular image into the generator, and step (2) represents the generator outputting a redirected monocular image.

[0104] S130: Process the line-of-sight angle true value of each redirected monocular image and the corresponding original monocular image based on the generator, to obtain a reconstructed monocular image corresponding to each original monocular image.

[0105] After obtaining the corresponding redirected monocular image of each original monocular image, the line-of-sight angle true value of each redirected monocular image and the corresponding original monocular image can be re-input into the generator, so that the generator learns the following ability: according to the line-of-sight angle true value of the redirected monocular image and the original monocular image, the reconstructed monocular image which is completely similar to the eye texture and structure in the original monocular image is reconstructed.

[0106] As shown in FIG. 2, step (3) represents inputting the redirected monocular image and the line-of-sight angle true value of the corresponding original monocular image into the generator, and step (4) represents the generator outputting the reconstructed monocular image. In the figure, solid lines and dashed lines are used to distinguish the first input and output and the second input and output of the generator, without containing other special meanings.

[0107] It should be noted that the original monocular image, the target monocular image, the redirected monocular image and the reconstructed monocular image can be one-to-one corresponding.

[0108] S140: Based on the discriminator in the initial monocular line-of-sight redirection model, each target monocular image, the corresponding redirected monocular image and the reconstructed monocular image are processed to obtain the true or false discrimination result and the line-of-sight angle prediction value of each image in the target monocular image, the redirected monocular image and the reconstructed monocular image.

[0109] After obtaining the redirected monocular image and the reconstructed monocular image, each target monocular image, the corresponding redirected monocular image and the reconstructed monocular image can be input into the discriminator, so that the discriminator learns to judge the target monocular image as true and the redirected monocular image and the reconstructed monocular image as false, and learns to predict the line-of-sight angle of each monocular image. Through learning and training, the discriminator outputs the true or false discrimination result and the line-of-sight angle prediction value of each image in the target monocular image, the redirected monocular image and the reconstructed monocular image.

[0110] As shown in FIG. 2, step (5) represents inputting the target monocular image, the corresponding redirected monocular image and the reconstructed monocular image into the discriminator, and step (6) represents the discriminator outputting the true or false discrimination result and the line-of-sight angle prediction value of each monocular image.

[0111] S150: According to the true or false discrimination result of each image in the target monocular image, the redirected monocular image and the reconstructed monocular image, the line-of-sight angle prediction value of each image, the line-of-sight angle true value of each image, the original monocular image and the reconstructed monocular image, the current loss value of the initial monocular line-of-sight redirection model is calculated.

[0112] The line-of-sight angle true value of the redirected monocular image is the line-of-sight angle true value of the target monocular image, and the line-of-sight angle true value of the reconstructed monocular image is the line-of-sight angle true value of the original monocular image.

[0113] The current loss value of the initial monocular view redirection model can be determined according to the current loss value of the generator and the current loss value of the discriminator, and the implementation manner comprises steps A1-A3:

[0114] A1: calculating the current loss value of the generator according to the true or false discrimination results of each image in the target monocular image, the redirected monocular image and the reconstructed monocular image, the true value and the predicted value of the view angle of the redirected monocular image, the true value and the predicted value of the view angle of the corresponding reconstructed monocular image of the redirected monocular image, the original monocular image and the reconstructed monocular image.

[0115] Specifically, the adversarial loss of the generator is determined according to the true or false discrimination results of each image in the target monocular image, the redirected monocular image and the reconstructed monocular image; the view accuracy loss of the generator is determined according to the true value and the predicted value of the view angle of the redirected monocular image and the true value and the predicted value of the view angle of the corresponding reconstructed monocular image of the redirected monocular image; the reconstruction similarity loss of the generator is determined according to the original monocular image and the corresponding reconstructed monocular image thereof; and the current loss value of the generator is calculated according to the adversarial loss of the generator, the view accuracy loss of the generator, the view accuracy loss weight, the reconstruction similarity loss and the reconstruction similarity loss weight.

[0116] After obtaining the adversarial loss of the generator, the view accuracy loss of the generator, the view accuracy loss weight, the reconstruction similarity loss and the reconstruction similarity loss weight, the current loss value L G of the generator can be calculated according to the first formula.

[0117] The first formula comprises:

[0118] wherein, L adv adversarial loss of the generator, γ gaze view accuracy loss weight, view accuracy loss of the generator, γ rec reconstruction similarity loss weight, L rec reconstruction similarity loss.

[0119] -L adv for maximizing the output of the discriminator for the fake monocular image, and minimizing the output of the discriminator for the real monocular image.

[0120] The calculation method of the line-of-sight accuracy loss of the generator includes: calculating the difference between the line-of-sight angle true value of each redirected monocular image and the line-of-sight angle predicted value thereof as a first difference value, calculating the difference between the line-of-sight angle true value of the corresponding reconstructed monocular image of each redirected monocular image and the line-of-sight angle predicted value thereof as a second difference value, and then calculating the L2 norm of all the first difference values and all the second difference values as the line-of-sight accuracy loss of the generator. Of course, the mean of all the first difference values and all the second difference values can also be directly calculated as the line-of-sight accuracy loss of the generator.

[0121] The calculation method of the reconstruction similarity loss of the generator includes: calculating the difference between each original monocular image and the corresponding reconstructed monocular image thereof as a third difference value, and taking the L1 norm of all the third difference values as the reconstruction similarity loss of the generator.

[0122] The line-of-sight accuracy loss weight and the reconstruction similarity loss weight can be pre-set according to actual experience, for example, the line-of-sight accuracy loss weight is 5 and the reconstruction similarity loss weight is 50, or an initial value can be set first and then adjusted through model training.

[0123] A2: According to the true or false discrimination results of each image in the target monocular image, the redirected monocular image and the reconstructed monocular image, the line-of-sight angle predicted value of the target monocular image and the line-of-sight angle true value of the target monocular image, the current loss value of the discriminator is calculated.

[0124] Specifically, the adversarial loss of the generator is determined according to the true or false discrimination results of each image in the target monocular image, the redirected monocular image and the reconstructed monocular image; the line-of-sight accuracy loss of the discriminator is determined according to the line-of-sight angle predicted value of the target monocular image and the line-of-sight angle true value of the target monocular image; and the current loss value of the discriminator is calculated according to the adversarial loss of the generator, the line-of-sight accuracy loss of the discriminator and the line-of-sight accuracy loss weight.

[0125] After obtaining the adversarial loss of the generator, the line-of-sight accuracy loss of the discriminator and the line-of-sight accuracy loss weight, the current loss value L of the discriminator can be calculated according to the second formula D ;

[0126] The second formula includes:

[0127] wherein L adv represents the adversarial loss of the generator, γ gaze represents the line-of-sight accuracy loss weight, and represents the line-of-sight accuracy loss of the discriminator.

[0128] L adv is used to minimize the output of the discriminator for the fake monocular image and maximize the output of the discriminator for the real monocular image.

[0129] The method for calculating the line-of-sight accuracy loss of the discriminator includes: calculating the difference between the true value of the line-of-sight angle of each target monocular image and the predicted value of the line-of-sight angle corresponding to the true value as a fourth difference value, and taking the L2 norm of all fourth difference values as the line-of-sight accuracy loss of the generator.

[0130] The line-of-sight accuracy loss weight in the second formula can be the same as the line-of-sight accuracy loss weight in the first formula.

[0131] A3: determining the current loss value of the initial monocular line-of-sight redirection model according to the current loss value of the generator and the current loss value of the discriminator.

[0132] After obtaining the current loss value of the generator and the current loss value of the discriminator, the combination of the current loss value of the generator and the current loss value of the discriminator can be directly taken as the current loss value of the initial monocular line-of-sight redirection model, or the current loss value obtained by calculating the current loss value of the generator and the current loss value of the discriminator through a third formula can be taken as the current loss value of the initial monocular line-of-sight redirection model. The third formula can be a mean formula, or a maximum value formula, etc.

[0133] S160: When the current loss value of the initial monocular line-of-sight redirection model does not meet the convergence condition, after adjusting the model parameters of the initial monocular line-of-sight redirection model, the initial monocular line-of-sight redirection model with adjusted parameters is continuously trained until the monocular line-of-sight redirection model obtained through current training meets the convergence condition, and the monocular line-of-sight redirection model obtained through current training is determined as the target monocular line-of-sight redirection model.

[0134] When the current loss value of the initial monocular visual line redirection model is a combination of the current loss value of the generator and the current loss value of the discriminator, the convergence condition includes that the current loss value of the generator is less than or equal to a first loss threshold, and the current loss value of the discriminator is less than or equal to a second loss threshold, which can be the same or different, and is determined according to actual experience. Therefore, when the current loss value of the generator is greater than the first loss threshold, or the current loss value of the discriminator is greater than the second loss threshold, it is determined that the current loss value of the initial monocular visual line redirection model does not satisfy the convergence condition. When the current loss value of the initial monocular visual line redirection model does not satisfy the convergence condition, the model parameters of the generator can be adjusted according to the current loss value of the generator, and the model parameters of the discriminator can be adjusted according to the current loss value of the discriminator by using the back propagation method. After adjusting the model parameters, the initial monocular visual line redirection model after adjusting the model parameters is continuously trained, that is, steps S110-S150 are executed again, until the monocular visual line redirection model obtained by the current training satisfies the convergence condition, and the monocular visual line redirection model obtained by the current training is determined as the target monocular visual line redirection model.

[0135] When the current loss value of the initial monocular visual line redirection model is a current loss value obtained by calculating the current loss value of the generator and the current loss value of the discriminator by the third formula, the convergence condition includes that the current loss value obtained by the third formula is less than or equal to a third loss threshold, and the specific value of the third loss threshold can be determined according to actual experience. Therefore, when the current loss value obtained by the third formula is greater than the third loss threshold, it is determined that the current loss value of the initial monocular visual line redirection model does not satisfy the convergence condition. When the current loss value of the initial monocular visual line redirection model does not satisfy the convergence condition, the model parameters of the generator and the model parameters of the discriminator can be adjusted according to the current loss value obtained by the third formula by using the back propagation method. After adjusting the model parameters, the initial monocular visual line redirection model after adjusting the model parameters is continuously trained, that is, steps S110-S150 are executed again, until the monocular visual line redirection model obtained by the current training satisfies the convergence condition, and the monocular visual line redirection model obtained by the current training is determined as the target monocular visual line redirection model.

[0136] The training method of the monocular view redirection model provided by the embodiments of the present application can train the monocular view redirection capability of the generator through the original monocular image and the corresponding target monocular image, train the capability of the generator to reconstruct the same eye texture and structure as the original monocular image through the redirected monocular image output by the generator and the line-of-sight angle ground truth of the corresponding original monocular image, train the line-of-sight angle prediction and true or false image discrimination capability of the discriminator through the target monocular image, the corresponding redirected monocular image and reconstructed monocular image, and continue to iteratively train the model after adjusting the model parameters according to the loss value calculated based on the line-of-sight angle ground truth and the line-of-sight angle prediction value. Finally, a target monocular view redirection model capable of redirecting the line of sight of a monocular image can be obtained, so that the line of sight of the eyes in the original face image can be directly redirected based on the target monocular view redirection model in the subsequent process. As a result, not only is the number of training samples and the line-of-sight range of the training set of the line-of-sight estimation model automatically expanded, but also the automatic labeling of the line of sight of the newly added training samples is realized, thereby improving the reliability of the line-of-sight estimation model.

[0137] Based on the above method embodiments, another embodiment of the present application provides a line-of-sight redirection method, as shown in FIG. 3, which comprises the following steps:

[0138] S210: obtaining a first original monocular image to be redirected and a target line-of-sight angle.

[0139] An original image to be redirected is obtained, the original image comprising a face region. A target detection model is used to detect and intercept a certain monocular image in the original image as a first original monocular image, for example, the right eye image is intercepted as the first original monocular image. The target line-of-sight angle can be a line-of-sight angle input by a user.

[0140] S220: processing the first original monocular image and the target line-of-sight angle based on the generator in the target monocular view redirection model to obtain a first redirected monocular image.

[0141] The target monocular view redirection model is trained according to the training method of the monocular view redirection model described in any of the above embodiments.

[0142] The target monocular view redirection model comprises a generator and a discriminator. After the first original monocular image and the target line-of-sight angle are obtained, the first original monocular image and the target line-of-sight angle can be input into the generator for redirection, and the first redirected monocular image is output.

[0143] S230: processing the first redirected monocular image based on the discriminator in the target monocular view redirection model to obtain the line-of-sight angle of the first redirected monocular image.

[0144] After obtaining the first redirected monocular image of the generator output, the first redirected monocular image can be input into the discriminator for gaze angle prediction, and the gaze angle of the first redirected monocular image is output. In addition, the discriminator can also output a true or false discrimination result of the first redirected monocular image.

[0145] In a possible implementation, the monocular image of each eye can be respectively cut from the original face image to be redirected, and each eye is respectively taken as the first original monocular image. By performing steps S210-S230, the redirected monocular image of each eye and the gaze angle are obtained. Then, the redirected monocular image of each eye is respectively replaced with the corresponding original monocular image, that is, the redirected left eye image is replaced with the original left eye image, the redirected right eye image is replaced with the original right eye image, and the corresponding gaze angle is marked, so as to obtain a redirected face image with marked eye region and gaze angle as a new sample of the gaze estimation model. The target gaze angle corresponding to the first original monocular image of different sides can be different.

[0146] The gaze redirection method provided by the embodiments of the present application can process the first original monocular image and the target gaze angle based on the generator in the target monocular gaze redirection model to obtain the first redirected monocular image. By processing the first redirected monocular image through the discriminator in the target monocular gaze redirection model, the gaze angle of the first redirected monocular image can be obtained. Subsequently, the gaze of the eye in the original face image can be redirected based on the redirected monocular image, so as to not only automatically expand the number of training samples and the gaze range in the training set of the gaze estimation model, but also realize automatic labeling of the gaze of the new training sample, thereby improving the reliability of the gaze estimation model.

[0147] In a possible implementation, in order to improve the efficiency of obtaining the redirected face image, when the first original monocular image is the monocular image cut from the original face image to be redirected, the first redirected monocular image can also be mirror processed to obtain the second redirected monocular image corresponding to the second original monocular image in the original face image and the gaze angle of the second redirected monocular image. The first redirected monocular image is replaced with the first original monocular image in the original face image, the second redirected monocular image is replaced with the second original monocular image in the original face image, and the gaze angle of the first redirected monocular image and the gaze angle of the second redirected monocular image are marked, to obtain the redirected face image after the gaze redirection and the gaze angle marking of the original face image.

[0148] The embodiment of the application only performs line-of-sight redirection on one eye in the original face image through the target monocular line-of-sight redirection model, and the other eye can be obtained through mirror processing of the first redirected monocular image, without using the target monocular line-of-sight redirection model for line-of-sight redirection. Since the efficiency of mirror processing is much higher than that of the target monocular line-of-sight redirection model, the efficiency of double-eye redirection of the original face image is improved, and the efficiency of obtaining the redirected face image is further improved. In addition, since the redirected face image obtained by the embodiment of the application is marked with a line-of-sight angle, it can be directly used as a training sample for a line-of-sight estimation model, without the need for artificial annotation or complex annotation process, thereby improving the efficiency of line-of-sight angle annotation.

[0149] Based on the above method embodiment, another embodiment of the application provides a training device of a monocular line-of-sight redirection model, as shown in FIG. 4, the device comprises:

[0150] The acquisition unit 310 is configured to acquire a training sample set, wherein the training sample set comprises a plurality of original monocular images, a target monocular image corresponding to each original monocular image, a line-of-sight angle true value of each original monocular image, and a line-of-sight angle true value of each target monocular image;

[0151] The redirection unit 320 is configured to process the line-of-sight angle true value of each original monocular image and its corresponding target monocular image based on a generator in an initial monocular line-of-sight redirection model, to obtain a redirected monocular image corresponding to each original monocular image;

[0152] The reconstruction unit 330 is configured to process the line-of-sight angle true value of each redirected monocular image and its corresponding original monocular image based on the generator, to obtain a reconstructed monocular image corresponding to each original monocular image;

[0153] The discriminant prediction unit 340 is configured to process each target monocular image and its corresponding redirected monocular image and reconstructed monocular image based on a discriminator in the initial monocular line-of-sight redirection model, to obtain a true-false discrimination result and a line-of-sight angle prediction value of each image in the target monocular image, the redirected monocular image and the reconstructed monocular image;

[0154] The loss calculation unit 350 is configured to calculate a current loss value of the initial monocular view redirection model according to the target monocular image, the redirected monocular image, the reconstructed monocular image, the true or false discrimination result of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the predicted view angle value of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the true view angle value of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the original monocular image and the reconstructed monocular image, wherein the true view angle value of the redirected monocular image is the true view angle value of the target monocular image, and the true view angle value of the reconstructed monocular image is the true view angle value of the original monocular image.

[0155] The adjustment training unit 360 is configured to, when the current loss value of the initial monocular view redirection model does not satisfy the convergence condition, continue to train the initial monocular view redirection model after adjusting the model parameters of the initial monocular view redirection model, until the monocular view redirection model obtained through current training satisfies the convergence condition, and determine the monocular view redirection model obtained through current training as the target monocular view redirection model.

[0156] In a possible implementation, the loss calculation unit 350 comprises:

[0157] The first loss calculation module is configured to calculate a current loss value of the generator according to the target monocular image, the redirected monocular image, the reconstructed monocular image, the true or false discrimination result of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the true view angle value and the predicted view angle value of the redirected monocular image, the true view angle value and the predicted view angle value of the reconstructed monocular image, the original monocular image and the reconstructed monocular image.

[0158] The second loss calculation module is configured to calculate a current loss value of the discriminator according to the target monocular image, the redirected monocular image, the reconstructed monocular image, the true or false discrimination result of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the predicted view angle value of the target monocular image and the true view angle value of the target monocular image.

[0159] The determination module is configured to determine the current loss value of the initial monocular view redirection model according to the current loss value of the generator and the current loss value of the discriminator.

[0160] In a possible implementation, the first loss calculation module is configured to:

[0161] determine an adversarial loss of the generator according to the true or false discrimination result of each of the target monocular image, the redirected monocular image and the reconstructed monocular image;

[0162] determine, according to the line-of-sight angle ground truth value of the redirected monocular image and the line-of-sight angle predicted value of the redirected monocular image, the line-of-sight accuracy loss of the generator;

[0163] determine, according to the original monocular image and the corresponding reconstructed monocular image, the reconstruction similarity loss of the generator;

[0164] calculate the current loss value of the generator according to the adversarial loss of the generator, the line-of-sight accuracy loss of the generator, a line-of-sight accuracy loss weight, the reconstruction similarity loss, and a reconstruction similarity loss weight.

[0165] In a possible implementation, the first loss calculation module is configured to calculate the current loss value L G of the generator according to a first formula.

[0166] The first formula includes:

[0167] wherein the L adv adversarial loss of the generator, the γ gaze line-of-sight accuracy loss weight, the L line-of-sight accuracy loss of the generator, the γ rec reconstruction similarity loss weight, the L rec reconstruction similarity loss.

[0168] In a possible implementation, the second loss calculation module is configured to:

[0169] determine the adversarial loss of the generator according to the true or false discrimination results of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image;

[0170] determine the line-of-sight accuracy loss of the discriminator according to the line-of-sight angle predicted value of the target monocular image and the line-of-sight angle ground truth value of the target monocular image;

[0171] calculate the current loss value of the discriminator according to the adversarial loss of the generator, the line-of-sight accuracy loss of the discriminator, and a line-of-sight accuracy loss weight.

[0172] In a possible implementation, the second loss calculation module is configured to calculate the current loss value L D of the discriminator according to a second formula.

[0173] The second formula includes:

[0174] wherein, the L adv denotes the adversarial loss of the generator, and the γ gaze denotes the line-of-sight accuracy loss weight, and the denotes the line-of-sight accuracy loss of the discriminator.

[0175] The training device of the monocular line-of-sight redirection model provided by the embodiment of the present application can train the monocular line-of-sight redirection capability of the generator through the original monocular image and the line-of-sight angle true value of the corresponding target monocular image, train the capability of the generator to reconstruct the same eye texture and structure as the original monocular image by inputting the redirected monocular image output by the generator and the line-of-sight angle true value of the corresponding original monocular image into the generator again, train the line-of-sight angle prediction and true-false image discrimination capability of the discriminator through the target monocular image, the corresponding redirected monocular image and reconstructed monocular image, and continue to iterate the model after adjusting the model parameters through the loss value calculated by the line-of-sight angle true value, the line-of-sight angle prediction value and other information. Finally, a target monocular line-of-sight redirection model capable of redirecting the line-of-sight of a monocular image can be obtained, so that the line-of-sight of the eyes in the original face image can be directly redirected based on the target monocular line-of-sight redirection model in the subsequent process. As a result, not only the number of training samples and the line-of-sight range in the training set of the line-of-sight estimation model are automatically expanded, but also the automatic labeling of the new training sample line-of-sight is realized, thereby improving the reliability of the line-of-sight estimation model.

[0176] Based on the above method embodiment, another embodiment of the present application provides a line-of-sight redirection device, as shown in FIG. 5, the device comprises:

[0177] The acquisition unit 410 is configured to acquire a first original monocular image to be redirected and a target line-of-sight angle.

[0178] The redirection unit 420 is configured to process the first original monocular image and the target line-of-sight angle based on the generator in the target monocular line-of-sight redirection model to obtain a first redirected monocular image, wherein the target monocular line-of-sight redirection model is trained according to the training method of the monocular line-of-sight redirection model in any of the above embodiments.

[0179] The line-of-sight angle prediction unit 430 is configured to process the first redirected monocular image based on the discriminator in the target monocular line-of-sight redirection model to obtain the line-of-sight angle of the first redirected monocular image.

[0180] In a possible implementation manner, the device further comprises:

[0181] The mirror processing unit is configured to mirror process the first redirected monocular image to obtain a second redirected monocular image corresponding to a second original monocular image in the original face image and a line-of-sight angle of the second redirected monocular image when the first original monocular image is a monocular image cut from the original face image to be redirected.

[0182] The replacement marking unit is configured to replace the first redirected monocular image with the first original monocular image in the original face image, replace the second redirected monocular image with the second original monocular image in the original face image, and mark the line-of-sight angle of the first redirected monocular image and the line-of-sight angle of the second redirected monocular image, to obtain a redirected face image after line-of-sight redirection and line-of-sight angle marking of the original face image.

[0183] The line-of-sight redirection device provided by the embodiments of the present application can process the first original monocular image and the target line-of-sight angle based on the generator in the target monocular line-of-sight redirection model to obtain the first redirected monocular image, and the discriminator in the target monocular line-of-sight redirection model can process the first redirected monocular image to obtain the line-of-sight angle of the first redirected monocular image, so that the line-of-sight of the original face image can be redirected according to the redirected monocular image in the subsequent process, thereby not only automatically expanding the number of training samples and the line-of-sight range in the training set of the line-of-sight estimation model, but also realizing automatic annotation of the new training sample line-of-sight, and further improving the reliability of the line-of-sight estimation model.

[0184] Based on the above method embodiments, another embodiment of the present application provides a computer readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the method according to any of the above embodiments.

[0185] Based on the above method embodiments, another embodiment of the present application provides an electronic device or a computer device, as shown in FIG. 6, which includes:

[0186] one or more processors 510;

[0187] The processor 510 is coupled to a storage device 520, and the storage device 520 is configured to store one or more programs;

[0188] When the one or more programs are executed by the one or more processors 510, the electronic device or the computer device implements the method according to any of the above embodiments.

[0189] Based on the above method embodiments, another embodiment of the present application provides a vehicle, which includes the device according to any of the above embodiments, or includes the electronic device as described above.

[0190] Based on the above embodiments, another embodiment of the present application provides a computer program product containing instructions, which, when executed on a computer or processor, cause the computer or processor to perform the method according to any one of the above embodiments.

[0191] The device embodiments described above correspond to the method embodiments and have the same technical effects as the method embodiments. For specific descriptions, refer to the method embodiments. The device embodiments are based on the method embodiments, and specific descriptions can be found in the method embodiments, which will not be repeated here. Those skilled in the art can understand that the drawings are only schematic diagrams of an embodiment, and the modules or flows in the drawings are not necessarily required to implement the present application.

[0192] Those skilled in the art can understand that the modules in the device in the embodiments can be distributed in the device in the embodiments as described in the embodiments, or can be located in one or more devices different from the embodiments. The modules in the above embodiments can be combined into one module, or can be further split into multiple sub-modules.

[0193] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for training a monocular visual line redirection model, comprising: The method comprises: obtaining a training sample set, wherein the training sample set comprises a plurality of original monocular images, a target monocular image corresponding to each of the original monocular images, a line-of-sight angle true value of each of the original monocular images, and a line-of-sight angle true value of each of the target monocular images; processing, by a generator in an initial monocular line-of-sight redirection model, the line-of-sight angle true value of each of the original monocular images and the corresponding target monocular image to obtain a redirected monocular image corresponding to each of the original monocular images; processing, by the generator, the line-of-sight angle true value of each of the redirected monocular images and the corresponding original monocular image to obtain a reconstructed monocular image corresponding to each of the original monocular images; processing, by a discriminator in the initial monocular line-of-sight redirection model, each of the target monocular image, the redirected monocular image, and the reconstructed monocular image to obtain a true-false discrimination result and a line-of-sight angle prediction value of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image; calculating a current loss value of the initial monocular line-of-sight redirection model according to the true-false discrimination result of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image, the line-of-sight angle prediction value of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image, the line-of-sight angle true value of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image, the original monocular image, and the reconstructed monocular image, wherein the line-of-sight angle true value of the redirected monocular image is the line-of-sight angle true value of the target monocular image, and the line-of-sight angle true value of the reconstructed monocular image is the line-of-sight angle true value of the original monocular image; when the current loss value of the initial monocular line-of-sight redirection model does not satisfy a convergence condition, continuing to train the initial monocular line-of-sight redirection model after adjusting model parameters of the initial monocular line-of-sight redirection model until a monocular line-of-sight redirection model obtained through current training satisfies the convergence condition, and determining the monocular line-of-sight redirection model obtained through the current training as a target monocular line-of-sight redirection model.

2. The method of claim 1, wherein, The calculation of the current loss value of the initial monocular line-of-sight redirection model according to the true-false discrimination result of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image, the line-of-sight angle prediction value of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image, the line-of-sight angle true value of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image, the original monocular image, and the reconstructed monocular image comprises: calculating a current loss value of the generator according to the true-false discrimination result of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image, the line-of-sight angle true value and the line-of-sight angle prediction value of the redirected monocular image, the line-of-sight angle true value and the line-of-sight angle prediction value of the reconstructed monocular image, the original monocular image, and the reconstructed monocular image. According to the true and false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the line-of-sight angle prediction value of the target monocular image and the line-of-sight angle true value of the target monocular image, a current loss value of the discriminator is calculated; According to the current loss value of the generator and the current loss value of the discriminator, a current loss value of the initial monocular line-of-sight redirection model is determined.

3. The method of claim 2, wherein, The current loss value of the generator is calculated according to the true and false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the line-of-sight angle true value and the line-of-sight angle prediction value of the redirected monocular image, the line-of-sight angle true value and the line-of-sight angle prediction value of the reconstructed monocular image, the original monocular image and the reconstructed monocular image, and the current loss value of the generator is calculated. According to the true and false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, an adversarial loss of the generator is determined. According to the line-of-sight angle true value and the line-of-sight angle prediction value of the redirected monocular image, and the line-of-sight angle true value and the line-of-sight angle prediction value of the reconstructed monocular image corresponding to the redirected monocular image, a line-of-sight accuracy loss of the generator is determined. According to the original monocular image and the reconstructed monocular image corresponding thereto, a reconstruction similarity loss of the generator is determined. The current loss value of the generator is calculated according to the adversarial loss of the generator, the line-of-sight accuracy loss of the generator, a line-of-sight accuracy loss weight, the reconstruction similarity loss and a reconstruction similarity loss weight.

4. The method of claim 3, wherein, The current loss value of the generator is calculated according to the adversarial loss of the generator, the line-of-sight accuracy loss of the generator, a line-of-sight accuracy loss weight, the reconstruction similarity loss and a reconstruction similarity loss weight, and the current loss value of the generator is calculated. According to a first formula, a current loss value L of the generator is calculated G ; The first formula includes: wherein the L adv denotes the adversarial loss of the generator, the γ gaze denotes the line of sight accuracy loss weight, the denotes the line of sight accuracy loss of the generator, the γ rec denotes the reconstruction similarity loss weight, the L rec denotes the reconstruction similarity loss.

5. The method according to any one of claims 2-4, characterized in that, The current loss value of the discriminator is calculated according to the true and false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the line-of-sight angle prediction value of the target monocular image and the line-of-sight angle true value of the target monocular image, and the current loss value of the discriminator is calculated. According to the true and false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, an adversarial loss of the generator is determined. According to the line-of-sight angle prediction value of the target monocular image and the line-of-sight angle true value of the target monocular image, a line-of-sight accuracy loss of the discriminator is determined. The current loss value of the discriminator is calculated according to the adversarial loss of the generator, the line-of-sight accuracy loss of the discriminator and a line-of-sight accuracy loss weight.

6. The method of claim 5, wherein, The current loss value of the discriminator is calculated according to the adversarial loss of the generator, the line-of-sight accuracy loss of the discriminator and a line-of-sight accuracy loss weight, and the current loss value of the discriminator is calculated. calculating a current loss value L of the discriminator according to a second formula D ; The second formula includes: wherein the L adv denotes the adversarial loss of the generator, the γ gaze denotes the line-of-sight accuracy loss weight, the The line-of-sight accuracy loss of the discriminator is represented.

7. A line of sight redirection method, characterized by, The method comprises: obtaining a first original monocular image to be redirected and a target line-of-sight angle; processing the first original monocular image and the target line-of-sight angle based on a generator in a target monocular line-of-sight redirection model, to obtain a first redirected monocular image, the target monocular line-of-sight redirection model being trained according to the method in any one of claims 1-6; processing the first redirected monocular image based on a discriminator in the target monocular line-of-sight redirection model, to obtain a line-of-sight angle of the first redirected monocular image.

8. The method of claim 7, wherein, When the first original monocular image is a monocular image cut from an original face image to be redirected, the method further comprises: mirroring the first redirected monocular image, to obtain a second redirected monocular image corresponding to a second original monocular image in the original face image and a line-of-sight angle of the second redirected monocular image; replacing the first original monocular image in the original face image with the first redirected monocular image, replacing the second original monocular image in the original face image with the second redirected monocular image, and marking the line-of-sight angle of the first redirected monocular image and the line-of-sight angle of the second redirected monocular image, to obtain a redirected face image after line-of-sight redirection and line-of-sight angle marking of the original face image.

9. An apparatus for training a monocular view line redirection model, comprising: The apparatus comprises: an acquisition unit configured to acquire a training sample set, wherein the training sample set comprises a plurality of original monocular images, a target monocular image corresponding to each of the original monocular images, a line-of-sight angle ground truth of each of the original monocular images, and a line-of-sight angle ground truth of each of the target monocular images; a redirection unit configured to process each of the original monocular images and the line-of-sight angle ground truth of the corresponding target monocular image based on a generator in an initial monocular line-of-sight redirection model, to obtain a redirected monocular image corresponding to each of the original monocular images; a reconstruction unit configured to process each of the redirected monocular images and the line-of-sight angle ground truth of the corresponding original monocular image based on the generator, to obtain a reconstructed monocular image corresponding to each of the original monocular images; a discriminant prediction unit configured to process each of the target monocular images, the corresponding redirected monocular image, and the corresponding reconstructed monocular image based on a discriminator in the initial monocular line-of-sight redirection model, to obtain a true or false discrimination result and a line-of-sight angle prediction value of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image; a loss calculation unit configured to calculate a current loss value of the initial monocular line-of-sight redirection model according to the true or false discrimination result of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image, the line-of-sight angle prediction value of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image, the line-of-sight angle ground truth of each of the target monocular image, the redirected monocular image, and the reconstructed monocular image, the original monocular image, and the reconstructed monocular image, wherein the line-of-sight angle ground truth of the redirected monocular image is the line-of-sight angle ground truth of the target monocular image, and the line-of-sight angle ground truth of the reconstructed monocular image is the line-of-sight angle ground truth of the original monocular image. The adjusting training unit is configured to, when a current loss value of the initial monocular view redirection model does not satisfy a convergence condition, continue training the initial monocular view redirection model after adjusting model parameters of the initial monocular view redirection model, until a current training obtained monocular view redirection model satisfies the convergence condition, and determine the current training obtained monocular view redirection model as a target monocular view redirection model.

10. The apparatus of claim 9, wherein, The loss calculation unit comprises: The first loss calculation module is configured to calculate a current loss value of the generator according to true or false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the view angle true value and the view angle predicted value of the redirected monocular image, the view angle true value and the view angle predicted value of the reconstructed monocular image, the original monocular image and the reconstructed monocular image; The second loss calculation module is configured to calculate a current loss value of the discriminator according to the true or false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image, the view angle predicted value of the target monocular image and the view angle true value of the target monocular image; The determination module is configured to determine the current loss value of the initial monocular view redirection model according to the current loss value of the generator and the current loss value of the discriminator.

11. The apparatus of claim 10, wherein, The first loss calculation module is configured to: determine an adversarial loss of the generator according to the true or false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image; determine a view accuracy loss of the generator according to the view angle true value and the view angle predicted value of the redirected monocular image and the view angle true value and the view angle predicted value of the reconstructed monocular image corresponding to the redirected monocular image; determine a reconstruction similarity loss of the generator according to the original monocular image and the reconstructed monocular image corresponding thereto; calculate the current loss value of the generator according to the adversarial loss of the generator, the view accuracy loss of the generator, a view accuracy loss weight, the reconstruction similarity loss and a reconstruction similarity loss weight.

12. The apparatus of claim 11, wherein, The first loss calculation module is configured to calculate a current loss value L of the generator according to a first formula G ; The first formula includes: wherein the L adv denotes the adversarial loss of the generator, the γ gaze denotes the line of sight accuracy loss weight, the denotes the line of sight accuracy loss of the generator, the γ rec denotes the reconstruction similarity loss weight, the L rec denotes the reconstruction similarity loss.

13. The apparatus of any one of claims 10-12, wherein, The second loss calculation module is configured to: determine an adversarial loss of the generator according to the true or false discrimination results of each of the target monocular image, the redirected monocular image and the reconstructed monocular image; determine a view accuracy loss of the discriminator according to the view angle predicted value of the target monocular image and the view angle true value of the target monocular image; calculate the current loss value of the discriminator according to the adversarial loss of the generator, the view accuracy loss of the discriminator and a view accuracy loss weight.

14. The apparatus of claim 13, wherein, The second loss calculation module is configured to calculate a current loss value L of the discriminator according to a second formula D ; The second formula includes: wherein the L adv denotes the adversarial loss of the generator, the γ gaze denotes the line-of-sight accuracy loss weight, the The view accuracy loss of the discriminator is represented.

15. A line of sight redirecting device, characterized in that, The apparatus comprises: an acquisition unit configured to acquire a first original monocular image to be subjected to view redirection and a target view angle; a redirection unit, configured to process the first original monocular image and the target view angle based on a generator in a target monocular view redirection model to obtain a first redirected monocular image, the target monocular view redirection model being trained according to the method in any one of claims 1-6; a view angle prediction unit, configured to process the first redirected monocular image based on a discriminator in the target monocular view redirection model to obtain a view angle of the first redirected monocular image.

16. The apparatus of claim 15, wherein, The apparatus further includes: a mirroring processing unit, configured to perform mirroring processing on the first redirected monocular image to obtain a second redirected monocular image corresponding to a second original monocular image in the original face image and a view angle of the second redirected monocular image, when the first original monocular image is a monocular image cut from the original face image to be view redirected; a replacement marking unit, configured to replace the first original monocular image in the original face image with the first redirected monocular image, replace the second original monocular image in the original face image with the second redirected monocular image, and mark the view angle of the first redirected monocular image and the view angle of the second redirected monocular image, to obtain a redirected face image after view redirection and view angle marking on the original face image.

17. A computer readable storage medium having stored thereon a computer program, characterized in that The program, when executed by a processor, implements the method in any one of claims 1-6, or any one of claims 7-8.

18. An electronic device, comprising: The electronic device includes: one or more processors; the processor is coupled with a storage device, and the storage device is configured to store one or more programs; when the one or more programs are executed by the one or more processors, the electronic device implements the method in any one of claims 1-6, or any one of claims 7-8.

Citation Information

Patent Citations

  • Sight line prediction method, device and system and readable storage medium

    CN110008835A

  • Pseudo RGB-d for self-improving monocular slam and depth prediction

    US20210065391A1