Eye expression correction method and system based on codec and feature decoupling

Through an eye correction method based on codec and feature decoupling, the user's static attributes and eye posture features are independently encoded, and re-parameterization processing is performed using a multi-layer neural network. This solves the problems of high equipment cost and low precision in existing technologies, and achieves efficient and natural eye correction effects, which is suitable for scenarios such as video conferencing and virtual reality.

CN120853236AActive Publication Date: 2025-10-28ZHAOYI INFORMATION TECH (SHANGHAI) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510957643.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-28
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

Existing eye correction technology has problems such as high equipment cost, poor user experience, insufficient flexibility, low accuracy and low computational efficiency, making it difficult to be widely used in mass scenarios such as video conferencing.

Method used

Using a method based on codec and feature decoupling, the user's static attributes and eye posture features are independently encoded through image acquisition, feature extraction, feature transformation and image generation. A multi-layer neural network is used for re-parameterization processing to generate natural and accurate eye correction images.

Benefits of technology

It achieves high-precision and natural eye correction effects, reduces computational overhead, and is suitable for mobile devices and real-time interactive scenarios such as video conferencing and virtual reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853236A_ABST
    Figure CN120853236A_ABST
Patent Text Reader

Abstract

The invention provides an eye correction method and system based on a codec and feature decoupling. The method comprises the following steps: step a, acquiring an original face image I, and acquiring an eye image Ic and head posture information Hgt of a user from the original face image; b, feature extraction is conducted on Ic and Hgt, and a vector attribute code zi representing static attribute features of a user, current eye posture information Gpre representing a current eye posture and a rotation attribute code zr representing eye rotation attributes are obtained; c, performing three-dimensional transformation on the zr according to the Gpre and preset target eye posture information Gtar to obtain a rotation attribute code zpre corresponding to the Gtar; step d, generating a target eye image Ipre according to the zi and the zpre; and e, pasting the Ipre back to the original face image I, and outputting the face image after eye correction. According to the method, the problem of image distortion caused by confusion of attribute and attitude features in a traditional method is effectively solved, and the accuracy and naturalness of eye correction are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and artificial intelligence, and in particular to a method and system for eye movement correction based on codec and feature decoupling. Background Technology

[0002] Gaze correction, an innovative technology applied to facial feature processing, aims to align a user's gaze with a target's line of sight by adjusting the direction or position of the eyeballs. This technology has wide-ranging applications in video conferencing, virtual reality, and augmented reality.

[0003] Currently, the technical solutions for achieving eye expression correction are mainly divided into two categories: hardware methods and software methods.

[0004] Traditional eye-tracking correction methods rely heavily on specialized hardware devices, such as eye trackers or infrared cameras. These devices capture the trajectory and position of eye movements to track the user's gaze in real time and correct the direction of the gaze through optical correction or hardware adjustments. However, these methods often have high equipment costs and require users to wear or fixation devices, resulting in a poor user experience. Furthermore, their application is often limited, making them difficult to use in general settings such as video conferencing.

[0005] Software-based methods primarily utilize traditional image processing algorithms. By analyzing facial feature points (such as eyes, nose, and mouth) in an image, they perform geometric transformations or pixel replacements on the eye region to achieve eye expression correction. This method offers high flexibility and speed and requires no dedicated hardware support. However, these traditional algorithms are typically based on static rules, have limited processing power, lack adaptive optimization capabilities, and struggle to handle diverse user needs.

[0006] With the development of deep learning models such as Convolutional Neural Networks (CNNs) and Generative Adversarial Networks (GANs), eye-correction technology based on neural networks has gradually become a research hotspot. However, existing eye-correction technologies based on neural networks still suffer from problems such as insufficient accuracy, poor naturalness, and low computational efficiency. Summary of the Invention

[0007] To address the aforementioned shortcomings of existing technologies, this invention provides a method and system for eye movement correction based on codec and feature decoupling.

[0008] One aspect of the present invention proposes an eye-tracking correction method based on codec and feature decoupling, comprising the following steps:

[0009] Step a: Acquire an original face image I using an image acquisition device, and obtain the user's eye image I from the original face image. c and head pose information H gt ;

[0010] Step b: For the eye image I c and head pose information H gt Feature extraction is performed to obtain the vector attribute code z representing the user's static attribute features. i Current eye pose information G, representing the current eye pose. pre And the rotation attribute code z, which characterizes the rotation attribute of the eye. r ;

[0011] Step c: Based on the current eye pose information G pre and the pre-defined target eye pose information G tar The rotation attribute is encoded as z r Perform a three-dimensional transformation to obtain the target eye pose information G. tar The corresponding rotation attribute code z pre ;

[0012] Step d: Encode z according to the vector attribute. i and the rotation attribute encoding z pre Generate target eye image I pre ;as well as

[0013] Step e: Transfer the target eye image I pre Paste back the original face image I and output the face image after eye correction.

[0014] Another aspect of the present invention provides an eye-correction system based on codec and feature decoupling, comprising:

[0015] The data input module acquires a raw face image I using an image acquisition device, and obtains the user's eye image I from the raw face image. c and head pose information H gt

[0016] The feature extraction module extracts features from the eye image I. c and head pose information H gt Feature extraction is performed to obtain the vector attribute code z representing the user's static attribute features. i Current eye pose information G, representing the current eye pose. pre And the rotation attribute code z, which characterizes the rotation attribute of the eye. r ;

[0017] The feature transformation module, based on the current eye pose information Gpre and the pre-defined target eye pose information G tar The rotation attribute is encoded as z r Perform a three-dimensional transformation to obtain the target eye pose information G. tar The corresponding rotation attribute code z pre ;

[0018] The decoding module encodes z according to the vector attribute. i and the rotation attribute encoding z pre Generate target eye image I pre ;

[0019] The image generation module generates the target eye image I. pre Paste back the original face image I and output the face image after eye correction.

[0020] The present invention has the following beneficial effects:

[0021] 1. High correction accuracy and natural visual effect: This invention uses feature encoding decoupling technology to ensure that the generated image can accurately and naturally adjust the gaze while maintaining personalized features. At the same time, it combines generative adversarial networks to optimize the generated texture details, ensuring that the corrected image is visually natural and realistic.

[0022] 2. High real-time performance and adaptability to low-resource devices: By reparameterizing the trained multi-layer neural convolutional network, this invention significantly reduces computational overhead during eye correction testing, enabling real-time correction in mobile devices. It is suitable for scenarios with high requirements for interactivity and real-time performance, such as video conferencing and virtual reality. Attached Figure Description

[0023] Figure 1 This is a flowchart of a preferred embodiment of the eye movement correction method based on codec and feature decoupling according to the present invention.

[0024] Figure 2 This is a block diagram of an eye-correction system based on codec and feature decoupling according to a preferred embodiment of the present invention.

[0025] Figure 3 This is a schematic diagram of the multi-layer neural convolutional network of the encoding module and decoding module according to a preferred embodiment of the present invention, using reparameterization processing. Detailed Implementation

[0026] The present invention will be further illustrated below through examples, the purpose of which is only to better understand the research content of the present invention and not to limit the scope of protection of the present invention.

[0027] like Figure 1As shown, a preferred embodiment of the method of the present invention includes the following steps a to e.

[0028] First, step a: acquire an original face image I using an image acquisition device such as a camera. This original face image includes facial information, and obtain the user's eye image I from the original face image. c and head pose information H gt .

[0029] In a preferred embodiment, obtaining the user's eye image in step a further includes the following steps:

[0030] Step a1: Identify the corresponding 106 facial key points from the original face image I;

[0031] Step a2: Locate the eye region for the 106 facial key points; preferably, the eye region is maintained with the left and right edges 16 pixels from the left and right corners of the eyes and the lower edge 16 pixels from the lower eye socket;

[0032] Step a3: Crop the original face image according to the eye region to generate the eye image I. c Preferably, the image height is 96 pixels and the depth is 64 pixels.

[0033] Furthermore, in step a, the head pose information H gt This includes the pitch angle (up and down) and the yaw angle (left and right).

[0034] Next is step b: for the eye image I c and head pose information H gt Feature extraction is performed to obtain the vector attribute code z representing the user's static attribute features. i Current eye pose information G, representing the current eye pose. pre And the rotation attribute code z, which characterizes the rotation attribute of the eye. r .

[0035] Preferably, in step b, the feature extraction module uses a reparameterized convolutional neural network (CNN) for feature extraction. The CNN consists of multiple cascaded convolutional layers and activation functions. Preferably, the user's static attribute features are personalized user features, including gender, age, skin color, etc. Preferably, the vector attribute encoding z... i It is a 256-dimensional vector. The current eye pose information G before correction. pre =(P pre ,Y pre P pre Y represents the pitch angle. pre This represents the yaw angle. Preferably, the rotation attribute is encoded as z. r It has 48 dimensions.

[0036] Step c: Based on the current eye pose information G pre and the pre-defined target eye pose information G tar The rotation attribute is encoded as z r Perform a three-dimensional transformation to obtain the target eye pose information G. tar The corresponding rotation attribute code z pre .

[0037] Preferably, the target eye pose information G needs to be preset before this step. tar =(P tar ,Y tar ), where P tar Y represents the pitch angle. tar Yaw angles are expressed in radians. Preferably, in a video call scenario, the target attitude can be set to the user's facing direction, i.e., both pitch and yaw angles are 0 degrees, resulting in a target attitude of (0,0). In other scenarios, the target attitude can be manually specified by the user, for example, a target looking upwards at 15° and to the left at 10°. The target is looking down 15° and to the right 10°.

[0038] Preferably, step c further includes the following steps c1-c3. Before steps c1-c3, z is first... r The size was reorganized into 3*16, which were used as 16 three-dimensional representations.

[0039] Step c1: Obtain the current eye pose information G pre The corresponding three-dimensional rotation matrix R pre :

[0040]

[0041] Step c2: Obtain the target eye pose information G tar The corresponding three-dimensional rotation matrix R tar :

[0042]

[0043] The two matrices represent left-right rotation and up-down rotation, respectively.

[0044] Step c3: For R pre Rotation inversion operation, and through R tar The rotation matrix is ​​transformed to obtain the rotation attribute code corresponding to the target eye pose. Finally, take the 3*16 size z pre Recombined into 48 dimensions, thus relating to z rMaintaining consistency. This process completes the eye pose change from the current direction G in the latent space. pre To the target direction G tar The natural adjustment.

[0045] The next step is d: Encode z according to the vector attribute. i and the rotation attribute encoding z pre Generate target eye image I pre .

[0046] Step d further includes: a decoding module composed of multiple convolutional neural networks encoding the vector attribute z. i and the rotation attribute encoding z pre Feature extraction is performed to generate the target eye image I. pre .

[0047] Preferably, in step d, the decoding module uses a reparameterized convolutional neural network for feature extraction. This convolutional neural network consists of bilinear upsampling layers, multiple cascaded convolutional layers, and activation functions. The bilinear upsampling operation employed in the decoding module demonstrates excellent performance in preserving detail and smoothing edges in the generated image, effectively avoiding artifacts caused by deconvolution methods.

[0048] Finally, step e: Image I of the target eye. pre Paste back the original face image I, and output the face image after eye correction. The system generates the target eye image I based on key points. pre The image is then pasted back to the original face image I, achieving seamless integration and avoiding any sense of disjointedness in the overall appearance. The corrected, complete face image will be displayed on the screen in real time, allowing users to directly view the adjusted eye expression. In scenarios such as video communication, the adjusted image can be directly used for real-time transmission.

[0049] As described above, this invention uses a feature extraction module to independently encode the user's personalized attributes and eye pose features, generating attribute codes z respectively. i and rotation attribute encoding z r This achieves feature decoupling. Specifically, attribute encoding z i Representing the user's personalized characteristics, such as static information like facial gender, age, and skin color; rotation attribute encoding z. r This is used to describe the dynamic rotation characteristics of the user's eyes. During eye gaze correction, only the rotation attribute encoding z is adjusted. r It simulates the natural rotation of the eyeball through three-dimensional pose transformation. Simultaneously, attribute encoding z... i This process remains unchanged, ensuring that the generated image retains personalized features while accurately and naturally adjusting the gaze.

[0050] An embodiment of the present invention also provides an eye-tracking correction system 20 based on codec and feature decoupling, comprising:

[0051] Data input module 21 acquires an original face image I through an image acquisition device, and obtains the user's eye image I from the original face image. c and head pose information H gt ;

[0052] Feature extraction module 22, for the eye image I c and head pose information H gt Feature extraction is performed to obtain the vector attribute code z representing the user's static attribute features. i Current eye pose information G, representing the current eye pose. pre And the rotation attribute code z, which characterizes the rotation attribute of the eye. r ;

[0053] Feature transformation module 23, based on the current eye pose information G pre and the pre-defined target eye pose information G tar The rotation attribute is encoded as z r Perform a three-dimensional transformation to obtain the target eye pose information G. tar The corresponding rotation attribute code z pre ;

[0054] Decoding module 24 encodes z according to the vector attribute. i and the rotation attribute encoding z pre Generate target eye image I pre ;

[0055] Image generation module 25 generates the target eye image I pre Paste back the original face image I and output the face image after eye correction.

[0056] The data input module 21 further includes: a face recognition module 211 and a head posture prediction module 212.

[0057] The facial recognition module 211 performs the following steps:

[0058] Step a1: Identify the corresponding 106 facial key points from the original face image I;

[0059] Step a2: Locate the eye region for the 106 facial key points; preferably, the eye region is maintained with the left and right edges 16 pixels from the left and right corners of the eyes and the lower edge 16 pixels from the lower eye socket;

[0060] Step a3: Crop the original face image according to the eye region to generate the eye image I. c .

[0061] The head pose prediction module 212 is used to predict the head pose H of the corresponding face in image I. gt This includes the vertical yaw angle (pitch angle) and the horizontal yaw angle (yaw angle).

[0062] Preferably, the multi-layer neural convolutional network of the feature extraction module (encoding module) 22 and the decoding module 24 of the present invention adopts a simplified and optimized network design, thereby solving the problem that the prior art cannot achieve real-time performance under resource-constrained conditions.

[0063] In a preferred embodiment, the feature extraction module (encoding module) 22 and the decoding module 24 employ reparameterization processing on the trained multi-layer neural convolutional network, thereby reducing computational load and memory read / write operations during testing. Specifically, for a convolutional module, the module structure during training is as follows: Figure 3 As shown, there are three branches: the first branch is a 3x3 convolutional layer followed by a batch normalization (BN) layer; the second branch is a 1x1 convolutional layer followed by a batch normalization (BN) layer; and the third branch is a residual connection (assuming the input and output channels are the same; otherwise, there is no third branch). During training, all three branches are trained together, effectively simulating the residual structure and spatially independent feature transformation structure. During testing (eye-correction processing), the 1x1 convolution can be viewed as a 3x3 convolutional kernel with zeros on the outer edges, and the residual structure can be viewed as a 3x3 convolutional kernel with a 1 at the center and zeros on the outer edges. The parameters of the batch normalization layer can then be merged into the convolutional kernels. Finally, the parameters of the three kernels are added together, merging two convolutions, two batch normalization layers, and one residual connection into a single convolutional layer (3x3 convolution), significantly reducing the number of layers, computational cost, and memory usage.

[0064] In a preferred embodiment, all convolutional layers use 3x3 convolutional kernels. This is because 3x3 convolutional kernels are far more computationally efficient than other convolutional kernels, while also taking into account neighborhood information.

[0065] In a preferred embodiment, all convolutional modules are single-path architectures without any side branches during testing. This is because operations such as residuals and cross-layer connections, while computationally inexpensive, consume significant amounts of additional memory, thus reducing computational efficiency. Therefore, this invention does not employ any residual or cross-layer connection structures in practical use.

[0066] In a preferred embodiment, the eye-correction system 20 further includes a supervised optimization module 26, which optimizes the parameters of each module in the system through data-driven training. This module 26 uses five loss functions to improve the quality of the generated images and the accuracy of pose adjustment.

[0067] (1). Reconstruction loss: calculated by reconstructing the predicted image I during training. pre and target image I gt The pixel error measures the difference between the generated image and the real image at the pixel level.

[0068] (2). Perceptual loss: The input image I during training is calculated using the pre-trained VGG16 network. pre and target image I gt Differences in perceptual features of generated images.

[0069] (3) Adversarial Loss: Based on the discriminator feedback from a Generative Adversarial Network (GAN), the discriminator predicts the image I through training. pre For the false target image I gt The generator attempts to confuse the discriminator to improve the generated image I. pre The sense of realism.

[0070] (4). Pose loss: calculated by measuring the pose information G during training. pre and target attitude G tar The squared angular distance measures the angular error between the currently predicted eye pose and the target pose.

[0071] (5) Consistency Loss: Calculate the attribute encoding z of the input image. i Attribute encoding z of the target image of the same person t Cosine similarity is used to ensure that different images of the same person have similar attribute encodings.

[0072] The supervised optimization module 26 jointly optimizes five loss functions, and the system generates corrected images that are not only visually realistic, but also accurate in posture adjustment.

[0073] Obviously, those skilled in the art should recognize that the above embodiments are only used to illustrate the present invention and are not intended to limit the present invention. Any changes or modifications to the above embodiments that are within the essential spirit of the present invention will fall within the scope of the claims of the present invention.

Claims

1. A method for eye movement correction based on encoder-decoder and feature decoupling, characterized in that, The steps include: Step a: Acquire an original face image I using an image acquisition device, and obtain the user's eye image I from the original face image. c and head pose information H gt ; Step b: For the eye image I c and head pose information H gt Feature extraction is performed to obtain the vector attribute code z representing the user's static attribute features. i Current eye pose information G, representing the current eye pose. pre And the rotation attribute code z, which characterizes the rotation attribute of the eye. r ; Step c: Based on the current eye pose information G pre and the pre-defined target eye pose information G tar The rotation attribute is encoded as z r Perform a three-dimensional transformation to obtain the target eye pose information G. tar The corresponding rotation attribute code z pre ; Step d: Encode z according to the vector attribute. i and the rotation attribute encoding z pre Generate target eye image I pre ;as well as Step e: Transfer the target eye image I pre Paste back the original face image I and output the face image after eye correction.

2. The eye movement correction method based on encoder-decoder and feature decoupling according to claim 1, characterized in that, Step a, obtaining the user's eye image, further includes the following steps: Step a1: Identify multiple corresponding facial key points from the original face image I; Step a2: Locate the eye region for the multiple facial key points; Step a3: Crop the original face image according to the eye region to generate the eye image I. c .

3. The eye movement correction method based on encoder-decoder and feature decoupling according to claim 1, characterized in that, In step a, the head pose information H gt This includes the pitch angle (up and down) and the yaw angle (left and right).

4. The eye movement correction method based on encoder-decoder and feature decoupling according to claim 1, characterized in that, In step b, the feature extraction module uses a reparameterized convolutional neural network to perform feature extraction. The convolutional neural network consists of multiple cascaded convolutional layers and activation functions.

5. The eye movement correction method based on encoder-decoder and feature decoupling according to claim 1, characterized in that, Step c further includes the following steps: Step c1: Obtain the current eye pose information G pre The corresponding three-dimensional rotation matrix R pre ; Step c2: Obtain the target eye pose information G tar The corresponding three-dimensional rotation matrix R tar ; Step c3: For R pre Rotation inversion operation, and through R tar The rotation matrix is ​​transformed to obtain the rotation attribute code corresponding to the target eye pose.

6. The eye movement correction method based on encoder-decoder and feature decoupling according to claim 5, characterized in that, In step c, In the above formula, the two matrices represent left-right rotation and up-down rotation, respectively.

7. The eye movement correction method based on encoder-decoder and feature decoupling according to claim 1, characterized in that, In step d, the decoding module uses a reparameterized convolutional neural network to perform feature extraction. The convolutional neural network consists of a bilinear upsampling layer, multiple cascaded convolutional layers, and activation functions.

8. The eye-correction system based on encoder-decoder and feature decoupling according to claim 1, characterized in that, include: The data input module acquires a raw face image I using an image acquisition device, and obtains the user's eye image I from the raw face image. c and head pose information H gt The feature extraction module extracts features from the eye image I. c and head pose information H gt Feature extraction is performed to obtain the vector attribute code z representing the user's static attribute features. i Current eye pose information G, representing the current eye pose. pre And the rotation attribute code z, which characterizes the rotation attribute of the eye. r ; The feature transformation module, based on the current eye pose information G pre and the pre-defined target eye pose information G tar The rotation attribute is encoded as z r Perform a three-dimensional transformation to obtain the target eye pose information G. tar The corresponding rotation attribute code z pre ; The decoding module encodes z according to the vector attribute. i and the rotation attribute encoding z pre Generate target eye image I pre ; The image generation module generates the target eye image I. pre Paste back the original face image I and output the face image after eye correction.

9. The eye-tracking correction system based on encoder-decoder and feature decoupling according to claim 8, characterized in that, The feature processing module and the decoding module reparameterize the trained multi-layer neural convolutional network for eye correction.

10. The eye-tracking correction system based on encoder-decoder and feature decoupling according to claim 8, characterized in that, It also includes a supervised optimization module, which optimizes the parameters in the neural convolutional networks of the feature processing module and the decoding module through a data-driven training method. The supervised optimization module uses a variety of loss functions.

Citation Information

Patent Citations

  • Face image sight correction method and device, equipment and storage medium

    CN112733795A

  • System and method for eye tracking

    CN113260299A

  • Virtual character eye posture adjusting method and device, equipment and storage medium

    CN117115321A

  • Self-calibration method, device and equipment for human eye sight line deviation parameter and medium

    CN117315012A

  • Generation method and device of line-of-sight image sample, electronic equipment and storage medium

    CN118247830A