Video eye correction method and system based on image processing
By using a multimodal joint inference network and an automatic scoring model, the eye-tracking correction parameters are dynamically adjusted, which solves the problems of stiff eye movements and insufficient interactive adaptability in existing technologies. This achieves high-quality, adaptive video eye-tracking correction effects and improves the naturalness and trustworthiness of human-computer interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CLOUD ATTACK NETWORK TECH HEBEI CO LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-05-01
AI Technical Summary
Existing video eye-correction technology cannot dynamically adjust the vividness and facial expression details of the eyes, resulting in stiff video images that cannot meet the needs of high-quality human-computer interaction and personalized emotional expression. It also lacks a continuous quality self-feedback mechanism, and its interactive adaptability is severely limited.
By acquiring the input video stream frame by frame and simultaneously acquiring the audio signal, and combining facial image features and speech expression data, a multimodal joint inference network is used to dynamically calculate the eye image deformation parameters and correct eye movements. The automatic scoring model and user feedback are used for closed-loop optimization to achieve adaptive adjustment of the eye movement adjustment factor.
It achieves dynamic response between eye gaze correction results and context and emotion, improves the naturalness of eye gaze and the friendliness of interaction, increases user naturalness score, enhances the system's adaptability and robustness, breaks through the bottleneck of traditional static rules, and achieves a higher level of naturalness in human-computer interaction.
Smart Images

Figure CN121967785A_ABST
Abstract
Description
A video eye-tracking correction method and system based on image processing Technical Field
[0001] This invention relates to the field of "image processing and multimodal intelligent human-computer interaction technology", and in particular to a video eye-tracking correction method and system based on image processing. Background Technology
[0002] During video expression, users' gaze may wander due to non-technical reasons such as nervousness, looking down, or thinking. While it's necessary to adjust the gaze to look directly at the camera, blindly adjusting the gaze entirely to "look directly at the camera" can easily violate emotional logic, appearing stiff or unrealistic. When correcting based on image distortion, the failure to dynamically adjust the correction amplitude and expression according to voice / semantics and emotional drivers resulted in stiff generated effects that detached from the real context.
[0003] Most existing video facial correction technologies focus on geometric correction effects, with correction parameters typically fixed by static rules or models. This lacks the dynamic adaptability to complex contexts, genuine emotional expression, and the natural flow of speech. Specifically, when speakers interact with the device in different contexts, language styles, or emotional states, existing eye-tracking correction methods often fail to simultaneously adjust the liveliness of the eyes, the details of their expressions, and their natural coordination with facial expressions. This results in stiff eye movements, disconnected blinking from facial expressions, and misaligned contexts in the generated video, weakening the viewer's trust and immersion in the virtual interactive subject, and failing to meet the actual needs of high-quality human-computer interaction and personalized emotional expression. Furthermore, mainstream solutions generally lack continuous quality feedback mechanisms, failing to dynamically optimize parameter settings based on user subjective evaluations or the needs of the interactive scenario, severely limiting interactive adaptability. Summary of the Invention
[0004] This application provides a video eye-correction method and system based on image processing, which aims to solve one of the problems or issues of the prior art mentioned in the background.
[0005] This application provides a video eye-tracking correction method and system based on image processing, specifically including:
[0006] S1: Capture the input video stream frame by frame, and simultaneously acquire the corresponding real-time audio signal to obtain facial image features and speech expression data.
[0007] S2: Based on the acquired facial image features and audio signals, perform denoising and normalization processes respectively to eliminate environmental interference and standardize the data format required for subsequent feature extraction.
[0008] S3: For pre-processed video frames, a facial landmark localization algorithm is used to accurately extract head pose parameters and eye region features. At the same time, scene labels are archived using different shooting angles and lighting conditions to achieve multi-scene differentiated modeling of facial features.
[0009] S4: Input the preprocessed audio samples into the emotion recognition model to obtain the emotional state, tone, speech rate and semantic unit information of the speech, and establish speech context feature label mapping for different language expression styles.
[0010] S5: Based on facial key point parameters, head pose parameters, eye region features, and speech context feature labels, the above multi-source data is input into a multimodal joint inference network to encode the complete context-aware features of the current speaker.
[0011] S6: Based on the context-aware features obtained from the encoding, calculate the eye image deformation parameters, expand them into parameterized gaze accommodation amplitude and expression change coefficient control factors, and set dynamic adjustment thresholds according to different emotional scenarios to improve the adaptability of eye naturalness.
[0012] S7: For each frame of facial image, perform eye deformation, shading adjustment and image compositing processing based on differentiable rendering according to the generated eye image deformation parameters and expression change coefficients, and maintain dynamic coordination with head posture and facial expression, and output the image and video after eye correction.
[0013] S8: Perform continuous naturalness and emotional consistency quality assessment on the synthesized video clips. Use an automatic scoring model combined with front-end user feedback to dynamically optimize the eye adjustment factor setting according to different scene tags, so as to achieve an adaptive quality closed loop for the generated results.
[0014] This application also provides a video eye-correction system based on image processing, which uses the above-mentioned video eye-correction method based on image processing to correct video eye movements.
[0015] This application provides a video eye-tracking correction method and system based on image processing, which has the following beneficial effects:
[0016] (1) This invention uses multimodal joint modeling technology to deeply correlate the speaker's emotional state (such as happiness, doubt, sadness), intonation changes, speech rate, semantic content with head posture, eye features, etc., so that the eye correction results can dynamically respond to the dialogue context and emotional flow. Compared with traditional static rules or single vision-driven models, it can continuously follow the context to achieve realistic and natural adjustments to eye amplitude, dynamic expressions, blinking, pupil changes, etc., effectively avoiding the problems of "stiff expressions, stiff correction, and disconnect from the scene" in previous video outputs. The subjective user naturalness score can be improved by more than 30%.
[0017] (2) By encoding multi-source facial features and audio context parameters, the invention realizes adaptive collaborative calculation of multiple factors such as eye deformation control, emotional expression, speech rhythm, and head dynamics, ensuring that the eyes and facial movements in each frame are highly synchronized with the real scene, and that the phenomenon of "homogenization", "drift" or separation from the context no longer occurs, thereby improving the affinity and trust in the interaction.
[0018] (3) Through the automatic scoring model and user front-end feedback loop, the system can autonomously identify low-quality segments and private emotional misalignment in the correction process, and dynamically optimize eye adjustment parameters to achieve continuous self-learning and effect iteration. Compared with the traditional static parameter setting method, it realizes adaptive, multi-scene, and multi-context high-quality video output, effectively improving the system's generalization and robustness.
[0019] This invention, from a technical perspective, achieves dynamic integration and high adaptability of eye-tracking correction across all aspects of "emotion, context, and expression," breaking through the bottleneck of existing methods that primarily rely on "static correction of direction / amplitude," and achieving a higher level of naturalness in human-computer interaction. Its technical effects represent a fundamental leap forward for intelligent video interaction and virtual human technology. Attached Figure Description
[0020] Figure 1 is the main flowchart of a video eye-correction method and system based on image processing.
[0021] Figure 2 is a sub-flowchart of a video eye-correction method and system based on image processing.
[0022] Figure 3 is another sub-flowchart of a video eye-correction method and system based on image processing. Detailed Implementation
[0023] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0024] The following disclosure provides many different embodiments or examples for implementing various structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed. In addition, examples of various specific processes and materials are provided in this invention, but those skilled in the art will recognize the application of other processes and the use of other materials.
[0025] As shown in Figure 1, this application provides a video eye-tracking correction method and system based on image processing, specifically including:
[0026] S1: Capture the input video stream frame by frame, and simultaneously acquire the corresponding real-time audio signal to obtain facial image features and speech expression data.
[0027] S2: Based on the acquired facial image features and audio signals, perform denoising and normalization processes respectively to eliminate environmental interference and standardize the data format required for subsequent feature extraction.
[0028] S3: For pre-processed video frames, a facial landmark localization algorithm is used to accurately extract head pose parameters and eye region features. At the same time, scene labels are archived using different shooting angles and lighting conditions to achieve multi-scene differentiated modeling of facial features.
[0029] S4: Input the preprocessed audio samples into the emotion recognition model to obtain the emotional state, tone, speech rate and semantic unit information of the speech, and establish speech context feature label mapping for different language expression styles.
[0030] S5: Based on facial key point parameters, head pose parameters, eye region features, and speech context feature labels, the above multi-source data is input into a multimodal joint inference network to encode the complete context-aware features of the current speaker.
[0031] S6: Based on the context-aware features obtained from the encoding, calculate the eye image deformation parameters, expand them into parameterized gaze accommodation amplitude and expression change coefficient control factors, and set dynamic adjustment thresholds according to different emotional scenarios to improve the adaptability of eye naturalness.
[0032] S7: For each frame of facial image, perform eye deformation, shading adjustment and image compositing processing based on differentiable rendering according to the generated eye image deformation parameters and expression change coefficients, and maintain dynamic coordination with head posture and facial expression, and output the image and video after eye correction.
[0033] S8: Perform continuous naturalness and emotional consistency quality assessment on the synthesized video clips. Use an automatic scoring model combined with front-end user feedback to dynamically optimize the eye adjustment factor setting according to different scene tags, so as to achieve an adaptive quality closed loop for the generated results.
[0034] Step S1: The input video stream is captured frame by frame, and the corresponding real-time audio signal is simultaneously captured to obtain facial image features and speech expression data. Specifically, this includes:
[0035] S1.1: Perform frame-level decoding on the original video stream to generate a continuous and indexable sequence of video frames, ensuring that each frame has a unique temporal identifier, and realizing the temporal extraction of facial features.
[0036] The raw video data stream is input via a video capture device to obtain unprocessed continuous image frame data.
[0037] A frame-level decoding processing algorithm (parameters: input video format, decoder standard H.264 / H.265, decoding buffer size, decoding resolution, etc.) is adopted to realize the frame-by-frame parsing operation of the input video stream data packets, and to parse the continuous video stream bitstream into independent image frame buffer structures.
[0038] Furthermore, a unique and incremental frame timing identifier is assigned to each decoded video frame using a video timestamp allocation algorithm (parameters: decoder system clock, frame rate FR, decoding delay correction factor Δt), thereby achieving a fully sequential temporal mapping relationship.
[0039] A frame index and timestamp synchronization strategy is adopted (parameters: frame number N, system start time T0) to establish a frame index table in the output video frame sequence, so as to realize the reversible number query and time sequence tracing of all image frames.
[0040] Furthermore, a redundant frame discrimination and elimination algorithm (parameters: bitstream integrity identifier IDR / I frame, frame loss detection threshold θ) is used to screen and eliminate duplicate, corrupt or abnormal decoded frames, ensuring the continuity, integrity and consistency of the output frame sequence.
[0041] By employing a buffered output and asynchronous scheduling mechanism, decoded video frames are written to the frame buffer queue in real time, providing a high-throughput, low-latency input data channel for subsequent face detection and key point extraction algorithms.
[0042] By using the above-mentioned frame-level decoding and temporal processing methods, the original video data stream is effectively converted into a standardized video frame sequence with a unique temporal identifier and indexable management, which forms the technical basis for the subsequent temporal and structured extraction of facial features.
[0043] For example, in a high-definition video capture scenario, a raw video stream at 30 frames per second, 1920×1080 resolution, and H.264 encoding format is used as input, and decoding is achieved through a combination of hardware and software. The decoder buffer size is set to 128MB, the FFmpeg decoding framework is used, and the initial time T0 is the system UTC timestamp. The average decoding latency per frame is approximately 33ms (frame rate 30fps), and a unique timing identifier is assigned to each video frame using the following algorithm:
[0044]
[0045] in, The timestamp of the nth frame. This is the start time of decoding. For frame rate, This is the delay correction value for the decoding queue scheduling, where n is the frame index number.
[0046] For the decoded image frames, duplicate and missing frames are removed, and valid frames are saved as a mapping table of frame number-image data-timestamp triplet. Statistically, an average of 30 sets of non-redundant video frame sequences are stably output per second, with an inter-frame timing error of less than 0.5ms. This meets the requirements for high-precision timing consistency and high-reliability structured image input, providing a solid data foundation and verification support for subsequent face detection and audio-video synchronization processing.
[0047] S1.2: Based on the generated frame-level decoding results, real-time audio signals are acquired synchronously, and multi-channel audio acquisition technology is used to accurately align the timestamps of each frame to form a correspondence between video frames and audio segments.
[0048] Using the indexable video frame sequence generated by frame-level decoding as the input data, a multi-channel audio acquisition method (parameters: sampling rate Fs, number of channels C, quantization precision Qbit, audio buffer size B) is adopted to achieve high-precision parallel acquisition of real-time audio signals.
[0049] Furthermore, by using a hardware-level timing synchronization-based audio and video clock calibration algorithm (parameters: video timestamp tn, audio acquisition system clock ts, maximum synchronization error Δτ), frame-level timestamp indexing is established on the acquired audio signal to achieve high-precision correction with the time reference of each video frame.
[0050] Furthermore, an audio data frame-slicing algorithm (parameters: frame length Lwin, frame shift Lhop, overlap Ol) is applied, using the time interval corresponding to each video frame as the slicing window, to continuously process the real-time acquired audio into frames, dividing the audio signal into a set of audio segments that completely correspond to the video frames.
[0051] Furthermore, an index binding and mapping relationship generation algorithm (parameters: mapping table T{A→V}, mapping fault tolerance δ, mapping table T receives an input about audio A and then outputs a corresponding value about video V) is adopted to establish a one-to-one index mapping relationship between the unique time sequence identifier of each video frame and the time label of the corresponding audio segment, and generate a video frame-audio segment synchronization association table.
[0052] Furthermore, a data integrity verification algorithm (parameters: audio and video synchronization flag I_{sync}, frame loss detection σ) is applied to backtrack and verify the audio and video alignment errors in the synchronization association table, ensuring the continuity, synchronization accuracy, and fault tolerance of the output results.
[0053] By using a precise multi-channel audio acquisition and audio-video time reference synchronization mechanism, the original frame-level video data and the real-time audio main channel signal are bound at the full-sequence index level. This lays a high-precision, low-latency data foundation for the subsequent temporal calibration of frame-by-frame facial features and segmented speech data, as well as cross-modal feature fusion, achieving physical-level synchronization of audio and video multimodal inputs.
[0054] For example, in a 1080p high-definition video and audio acquisition scenario, two array microphone channels are used, with a sampling rate of 48kHz, quantization precision of 16bit, and a buffer size of 512KB. The video frame rate is 30fps, and the initial synchronization error Δτ is set to ±2ms. The audio acquisition end uses a hardware timing signal to trigger audio sampling points at the start of each frame's acquisition, performs audio slicing every 33ms time window, and assigns a corresponding timestamp to each frame. The alignment of frames and audio segments is achieved using the following synchronization formula:
[0055]
[0056] in, The audio-video synchronization deviation of the i-th frame. Let i be the temporal identifier of the i-th image frame. The audio sampling time starts within the same window. In actual testing, the synchronization error fluctuation between the audio segment and the corresponding video frame is less than 0.9ms, the frame drop rate is less than 0.01%, and the synchronization index table integrity reaches 100%. Through the chained derivation output of this step, a structured "frame number—image frame—audio segment—synchronization flag" quadruple is obtained. This provides a highly consistent temporal reference and reliable frame mapping for subsequent facial region localization and multimodal data fusion, achieving data integrity for audio and video feature extraction and the technical effect of cross-modal synchronization processing.
[0057] S1.3: Perform preliminary face region localization processing on the video frame sequence, and output candidate boxes of key facial regions through a convolutional neural network face detection algorithm to lay the data foundation for subsequent facial feature acquisition.
[0058] S1.4: Based on the audio acquisition results, extract the main channel data of the original speech signal, and standardize the signal amplitude, duration and sampling rate to ensure the consistency and usability of the audio expression data format.
[0059] S1.5: Utilizing the correspondence between video frames and audio segments, the normalized face region candidate boxes and standardized audio expression data are temporally calibrated to generate frame-by-frame facial feature data and segmented speech expression data, thus establishing a unified temporal data input foundation for the multimodal feature fusion stage.
[0060] Step S2: Based on the acquired facial image features and audio signals, denoising and normalization processes are performed respectively to eliminate environmental interference and standardize the data format required for subsequent feature extraction. Specifically, this includes:
[0061] S2.1: Spatial domain filtering is performed on the original facial feature image data for denoising. A multi-scale feature extraction algorithm using a convolutional neural network effectively removes lighting artifacts and noise textures from the facial image to obtain a high-purity facial pixel matrix. This high-purity pixel matrix provides optimized input conditions for subsequent facial keypoint localization.
[0062] For the acquired raw facial image data, a spatial domain filtering denoising method (parameters: filter kernel size kh, filter type type=Gaussian / median, filter intensity σ) is used to initially suppress high-frequency noise and local uneven illumination components in the facial region, and generate a set of low-noise, high-fidelity facial images.
[0063] Furthermore, a multi-scale feature response calculation is performed on the preliminarily denoised facial image using a convolutional neural network (CNN) multi-scale feature extraction algorithm (parameters: network structure ResNet-18 / 50, number of feature channels Nc, scale pyramid level L), thereby achieving depth discrimination of illumination artifacts, texture noise, and local blur at different spatial scales.
[0064] Furthermore, through a multi-path feature aggregation mechanism (parameters: path weight αi, fusion strategy Weighted-Add), a dynamic balance is achieved between high-frequency noise suppression and low-frequency structural information preservation, and the pixel mapping matrix after comprehensive response is output.
[0065] Furthermore, a pixel-level adaptive denoising residual compensation method (parameters: residual constraint threshold ε, compensation activation function ReLU / Swish) is adopted to perform targeted correction on the remaining high noise areas, ensuring that key facial structures (such as the corners of the eyes, bridge of the nose, and corners of the mouth) remain intact and clear after denoising.
[0066] Furthermore, the mean square error of pixel distribution (MSE) evaluation function is used to statistically analyze the pixel mean difference in each key facial region of the image before and after denoising. If the MSE is less than the set threshold θy, the denoising is considered effective.
[0067] Through the above chain derivation process, the original facial image data is transformed into a high-purity facial pixel matrix, providing optimized input conditions for subsequent facial key point localization and multimodal feature extraction, and effectively eliminating image input noise and illumination interference.
[0068] For example, in a single-frame facial image input scenario with a resolution of 1920×1080, a Gaussian filter kernel size of kh=7×7 and a standard deviation of σ=1.2 was selected. A multi-scale convolutional network with a ResNet-18 structure was used, with 64 feature channels Nc, and a three-layer scale pyramid L=3 was constructed. In the actual processing flow, the PSNR after filtering and denoising was improved to over 38dB, the MSE evaluation function was set to a threshold of θy=10, and the residual MSE in the denoised block decreased to 8.1. The final high-purity facial pixel matrix output showed a signal-to-noise ratio improvement of over 25% in areas such as facial edges and eye details, effectively ensuring the data quality of subsequent key point localization and emotion feature analysis. For different shooting lighting scenarios (bright, backlit, complex textured backgrounds), the adaptive denoising parameters were automatically adjusted, achieving a significant improvement in SSIM of 0.18–0.27. Through the standardized output of this step, the high-purity facial pixel matrix can be directly used as a high-quality input for subsequent normalization and feature extraction algorithms, significantly improving the robustness and consistency of visual input under multi-scene adaptation.
[0069] S2.2: Based on a high-purity facial pixel matrix, an image normalization algorithm (such as Z-score normalization) is used to perform a uniform distribution transformation on all pixels, making the facial image feature distribution adapt to the input specifications of the multimodal feature fusion network, thus obtaining normalized facial image feature data. This normalized feature data serves as the baseline for subsequent multi-scene modeling and visual feature extraction.
[0070] S2.3: Frequency domain filtering and denoising are performed on the original audio signal, and wavelet transform algorithm is used to separate environmental noise from the main speech signal, thereby generating high signal-to-noise ratio (SNR) speech spectral features. These high SNR speech spectral features are key inputs for subsequent emotion recognition models and semantic analysis modules.
[0071] S2.4: Based on the high signal-to-noise ratio speech spectral characteristics, audio feature normalization processing (such as standardizing MFCC parameters) is adopted to uniformly adjust the distribution range of each speech parameter, resulting in a normalized audio feature vector. The normalized audio feature vector can maintain the same source standard as image feature data in the multimodal inference network, effectively supporting the joint coding process.
[0072] S2.5: Normalized facial image feature data and normalized audio feature vectors are structurally verified and time-synchronized mapped. Relying on the multimodal input data formatting module, synchronized labeled perceptual input is generated. The synchronized labeled perceptual input serves as the data entry point for the multi-source feature joint inference network, providing clear and traceable standard input signals for the subsequent context-aware encoding stage.
[0073] Step S3: For the preprocessed video frames, a facial landmark localization algorithm is used to accurately extract head pose parameters and eye region features. Simultaneously, scene labeling is performed using different shooting angles and lighting conditions to achieve multi-scene differentiated modeling of facial features. As shown in Figure 2, this specifically includes:
[0074] S3.1: Based on the preprocessed video frame data, a facial landmark localization algorithm is applied to extract landmark coordinates for each frame to obtain facial landmark parameters, including feature points such as the corners of the eyes, eyebrows, nose tip, and mouth corners. This step uses standardized image data as input and employs a deep learning-based facial landmark detection network, extracting 100 facial landmarks and outputting high-precision facial landmark parameters, laying the foundation for subsequent head pose estimation and eye region feature extraction.
[0075] S3.2: Using the extracted facial keypoint parameters, spatial mapping and geometric calculation algorithms are applied to calculate the 3D head pose parameters, including yaw, pitch, and roll angles. This step takes the facial keypoint parameters as input, employs geometric transformation and pose estimation models, and outputs 3D head pose parameters, providing a precise data foundation for eye region localization and gaze deviation determination.
[0076] For the extracted facial key point parameters, spatial mapping and geometric calculation algorithms are used to calculate the pose of the face in three-dimensional space. The goal is to obtain high-precision three-dimensional pose parameters that can truly reflect the dynamic features of the head pose in the current video frame.
[0077] A three-dimensional spatial mapping algorithm (parameters: World Coordinate System (WCS), Image Coordinate System (ICS), and Mapping Matrix (M3D)) is used to accurately project two-dimensional facial key points (such as the left and right corners of the eyes, the tip of the nose, the corners of the mouth, and the eyebrows) from the image plane coordinates to the three-dimensional face space model, thereby realizing the spatial coordinate transformation of key points from ICS to WCS.
[0078] Furthermore, by employing the PnP (Perspective-n-Point) pose estimation algorithm (parameters: standard face 3D reference model point set P, detected 2D image point set p, camera intrinsic matrix K, distortion coefficient d), the rotation vector of the face is estimated from the correspondence between multiple sets of 2D detection points and 3D model points. With translation vector The core of the algorithm is based on the following mathematical relationship:
[0079]
[0080] in, Let K be the two-dimensional projection point in camera coordinates, and K be the camera intrinsic parameter matrix. SP is a concatenation of a rotation matrix and a translation vector, where SP is a point in three-dimensional space.
[0081] Furthermore, by adjusting the rotation vector Perform a Rodrigues transform to obtain three-dimensional Euler angle parameters, which represent the yaw angles of the head. Pitch angle Roll angle The calculation process follows the following relationship:
[0082]
[0083] Furthermore, a residual minimization optimization algorithm (such as Levenberg-Marquardt, parameter: convergence threshold) is employed. Maximum number of iterations The reprojection error in the attitude calculation process is iteratively optimized to reduce the distance between the actual key points and the projection points, thereby improving the fitting accuracy of the attitude parameters.
[0084] Furthermore, consistency checks and time-series smoothing algorithms (such as Kalman filtering, parameter: observation noise covariance) are used to further refine the results. Process noise covariance The head pose parameter sequence of consecutive video frames is dynamically smoothed to eliminate short-term jitter and abrupt changes, ensuring the smoothness and physical rationality of the parameters over time.
[0085] Through the above processing flow, the original facial key point parameters are transformed into high-precision head posture data in three-dimensional space, including yaw angle, pitch angle and roll angle, which effectively supports the robust implementation of subsequent eye area localization and gaze deviation determination in multiple scenarios and multiple angles, and achieves accurate depiction of the speaker's facial behavior and natural interaction state.
[0086] For example, in a video frame processing scenario at 1080p resolution, a 68-point Dlib facial landmark detection model is selected. Based on the 2017 standard 3D face model, six stable points are selected (left and right corners of the eyes, tip of the nose, left and right corners of the mouth, and chin). The OpenCV solvePnP algorithm is used, with the camera focal length fx=fy=1200, optical center (cx,cy)=(960,540), and RANSAC iterations N. max =100, convergence threshold .
[0087] Kalman filtering was used for temporal smoothing, with observation noise covariance Q=0.01 and process noise covariance R=0.001. In actual multi-frame video sequence processing, the accuracy of 3D head Euler angle calculation reached 99.4%, the average reprojection error was less than 1.6 pixels, and the smoothness of the pose parameter sequence was improved by more than 20%. Through the above implementation method, high-reliability 3D head pose parameters can be stably output in real-time video streams with different angles, lighting conditions, and even partial occlusion, providing a solid data foundation for accurate eye region segmentation and subsequent emotion-driven dynamic gaze correction.
[0088] S3.3: Combining facial keypoint parameters and 3D head pose parameters, the eye region is accurately extracted based on a region segmentation algorithm to obtain eye region features. This step takes facial keypoint parameters and head pose parameters as input, and uses a geometrically constrained region segmentation technique to output eye coordinate boundaries and structural feature parameters, providing structured feature input for subsequent eye image deformation and emotion-driven deformation control.
[0089] The input data consists of standardized facial key point parameters and 3D head pose parameters obtained through spatial mapping analysis.
[0090] A region segmentation algorithm based on geometric constraints is adopted (parameters: key point category (corner of the eye, upper and lower boundaries of the eye socket, and bridge of the nose reference point), and the Euler angles [θ] of the head's 3D pose are used. yaw , θ pitch ,θ roll [ ] , rotation projection matrix R), to achieve preliminary localization of the eye region in the facial image of the current video frame.
[0091] Furthermore, by using affine transformation and projection correction algorithms, the coordinates of two-dimensional key points are spatially corrected based on the head pose, generating eye candidate region coordinates that accurately match the physical space projection, thereby enhancing the segmentation accuracy at different angles.
[0092]
[0093] Where [X, Y, Z] are the reference points of the eye structure in three-dimensional space, and [x', y'] are the two-dimensional image coordinates after head pose mapping.
[0094] Furthermore, a multi-layer segmentation mask generation algorithm is adopted (parameter: multi-scale window size r). i Extract the number of channels (c) and edge enhancement threshold (e), and dynamically stitch the contours based on key points such as the corner of the eye, upper and lower eyelids, and eye socket curves to form a high-precision eye segmentation mask.
[0095] Furthermore, by using structured feature extraction methods (such as Histogram of Oriented Gradients (HOG) and deep feature coding networks, with parameters such as kernel size k and number of layers L), the segmented eye region is structurally described, and structural feature parameters such as eyelid opening and closing, iris center, and pupil boundary are extracted to form a quantitative eye region feature set.
[0096] Furthermore, a boundary consistency check and temporal smoothing strategy (parameters: boundary IoU threshold γ, temporal weighting factor λ) are adopted to perform consistency checks and dynamic optimization on the eye region segmentation results in multi-frame video sequences, suppress jitter and false detections, and ensure that the eye region is accurately synchronized with facial expressions and head movements.
[0097] Through the above processing method, facial key point parameters and head pose parameters are transformed into high-precision eye region coordinate boundaries and structured feature parameters, providing structured input for subsequent adaptive control of eye image deformation and emotion-driven processing, and realizing stable extraction of eye features under multiple scenes and dynamic expressions.
[0098] For example, in a 1080p resolution video stream where the speaker's head has dynamic yaw of ±30 degrees, pitch of ±15 degrees, and roll of ±10 degrees, based on a standard facial key point set of 68 points and combined with Euler angle parameters obtained from PnP pose estimation, the basic coordinates of the left and right eye regions are projected onto the image region through affine mapping correction (left eye rectangle: [x1=420, y1=330, x2=510, y2=390], right eye rectangle: [x1=700, y1=330, x2=790, y2=390]). Multi-layer HOG feature encoding with a kernel size of k=5 and channel number c=32 was employed. Gradient distribution features at the eyelid edges were extracted for iris region localization. By using a dynamic IoU threshold γ=0.85 combined with a temporal smoothing factor λ=0.7, after processing a 5-second video clip (150 frames), the average IoU overlap with the ground truth annotation reached 0.92 under different expressions and lighting conditions. The eyelid opening / closing detection error was <3%, and the iris center deviation was <2 pixels. Output data includes: a multi-frame continuous set of eye region coordinates, iris / pupil segmentation features, and structured eye state parameters, effectively supporting high-precision localization and smooth synthesis of gaze deviation and emotional expression dynamics during adaptive eye gaze correction.
[0099] S3.4: Based on the shooting angle and illumination distribution characteristics of video frames, a scene label archiving algorithm is applied to classify scene features for each frame and generate scene label parameters. This step takes the video frame's metadata, image brightness statistics, and shooting angle information as input, and uses a multi-feature scene classification model to output scene label parameters, which are used to guide the setting of differentiated facial feature modeling and eye correction strategies for multiple scenes.
[0100] The input data includes a single-frame standardized facial image, the corresponding video frame metadata (frame number, timestamp, camera parameters), extracted 3D head pose parameters and average brightness values of the facial region, as well as an existing scene category label library for the acquisition scene.
[0101] A multi-feature extraction algorithm is employed (parameters: image brightness distribution histogram level B, global average brightness μ). y Local contrast σ y The system utilizes a shooting angle range Δθ and a frame data storage structure M to achieve synchronous quantization and extraction of global and local illumination states and imaging angle characteristics of video frames.
[0102] Furthermore, through image brightness clustering and normalization algorithms (parameter: k-value, cluster center C), k Normalized scale λ g Multicenter clustering is performed on the brightness statistics parameters of the facial region in each frame to extract the main mode type of brightness condition, and the differences are standardized to obtain the illumination classification label L.
[0103] Furthermore, the shooting angle is used to distinguish the models (parameter: yaw angle θ). yaw Pitch angle θ pitch Classification threshold T θ The head's three-dimensional pose parameters are mapped to a preset shooting angle category space to achieve the classification and grouping of angle features.
[0104] Furthermore, a multi-dimensional feature joint encoding algorithm (parameters: brightness classification label L, angle category label A, frame metadata M) is used to jointly encode the lighting conditions, imaging angle, and other environmental metadata of each frame to generate a feature vector V. s This ensures that scene diversity is preserved during the facial modeling stage.
[0105] Employ multi-feature scene classification models (such as decision trees, random forests, or deep scene recognition networks), with parameters: feature vector input V. s Number of categories N, confidence threshold P c The multidimensional feature vectors are hierarchically archived to automatically determine the scene category to which the frame belongs and output the scene label parameter S. tag .
[0106] Through the above processing method, the shooting angle, brightness status and multi-dimensional environmental parameters corresponding to a single frame video image are transformed into standardized scene labels, which serve as the data basis for subsequent multi-scene facial feature modeling and adaptive eye correction strategy setting, thereby realizing efficient archiving and application of multi-source environmental information.
[0107] For example, in a 1920×1080 resolution video conferencing scenario, the input frame data includes the camera model (FOV=82 degrees), frame number, and acquisition time. The yaw angle θ is obtained after face pose analysis. yaw =+25°, pitch angle θ pitch =-8°. Calculate the global average brightness μ for pixels in the facial region. y At a brightness level of λ=142 and B=6, the cluster centers are C1=90, C2=145, and C3=210. A normalized scale λ is used. g =50, and the brightness was classified into category L=2 by k-means clustering. Shooting angle parameter θ yaw An angle >20° is classified as "side-front" scene, pitch angle θ pitch A value <-5° is grouped into the "slight head tilt" group. Finally, a multi-feature scene archiving network is used to output the scene label S. tag ={"Lighting Type": "Normal Indoor Light", "Shooting Angle": "Side Front - Slightly Downward", "Device Type": "HD Webcam"}. Under other conditions such as backlighting or high-contrast lighting, the normalized brightness μ... yA score dropping to 90 or rising to 200 can be categorized as either "backlighting" or "highlighting," respectively. Verification showed that, under mixed scene conditions, the multi-feature archiving accuracy reached 98.3%, achieving automatic classification of differences in lighting, imaging angle, and shooting equipment. This provides a reliable scene discrimination basis for subsequent multi-scene adaptive facial feature modeling and natural gaze correction strategy selection.
[0108] S3.5: Feature fusion processing is performed on facial key point parameters, head pose parameters, eye region features, and scene label parameters to construct a multi-scene differentiated facial feature structure. This step uses the above four types of parameters as input, employs feature fusion and normalization algorithms, and outputs a multi-scene differentiated facial feature structure, providing unified and highly adaptable underlying data support for subsequent multimodal joint modeling and context-driven eye correction algorithms.
[0109] Step S4: Input the preprocessed audio samples into the speech emotion recognition model to obtain the emotional state, intonation, speech rate, and semantic unit information of the speech, and establish speech context feature label mappings for different language expression styles. As shown in Figure 3, this specifically includes:
[0110] S4.1: Perform frame-level segmentation on the preprocessed audio signal to generate temporally sequenced audio frames, ensuring that the emotion recognition model can establish dynamic temporal analysis conditions for the emotion state for each frame input.
[0111] The input is a preprocessed and normalized audio signal, including the main audio channel waveform data after sampling rate normalization, noise reduction, and amplitude normalization.
[0112] Common speech emotion recognition models include: Hidden Markov Model, Transformer Model, CNN-LSTM Model, OpenAI's Whisper Model, Bidirectional Long Short-Term Memory Network (Bi-LSTM), Convolutional Recurrent Hybrid Network (CRNN), etc.
[0113] A fixed-time-window-based frame segmentation algorithm is adopted (parameter: window length T). frame Frame shift T shift This enables frame-level segmentation of the entire audio signal.
[0114] Furthermore, by using a windowing function (parameters: window type, such as Hann window, Hamming window, window length N), smooth weighting is applied to each frame of audio signal to suppress inter-frame edge effects and enhance the temporal localization of short-term speech features.
[0115] Furthermore, an algorithm for aligning frame numbers with input audio timestamps is employed (parameter: sampling frequency f). sThe first sample number (n0) is used to realize the absolute time positioning of each frame and its original audio stream, ensuring the temporal synchronization of subsequent emotional features.
[0116] Furthermore, through energy threshold detection and silence rejection methods (parameter: short-time energy threshold E), thresh Minimum number of consecutive frames N min It detects and filters silent segments and background noise frames, thereby enhancing the continuity and contextual relevance of the analysis of effective speech segments.
[0117] By employing the aforementioned methods of framing, windowing, alignment, and silence removal, the standardized long-duration audio signal is transformed into structured temporal audio frame data that covers the entire time domain and has absolute time stamps, providing a high-spatiotemporal-precision and consistent input for subsequent emotion recognition and multimodal fusion.
[0118] For example, in a 2-minute video-audio stream scenario with a 48kHz sampling rate and single-channel input, the frame length T is set... frame =25ms (i.e., 1200 points), frame shift T shift =10ms (i.e., 480 points), and windowing smoothing is performed using the Hamming window function. Through frame alignment calculation, each frame is assigned its absolute start and end time in the original audio (e.g., frame 120 corresponds to 1.2s~1.225s). The total energy of each frame is calculated using a short-time energy algorithm, and an energy threshold E is set. thresh =0.01, consecutive frames below this threshold are automatically marked as silent and discarded, retaining the set of valid speech segments for subsequent analysis. If the number of frames is less than N... min Five-frame (50ms) segments are further removed to ensure that all output frames are coherent speech segments. The actual number of structured output frames is approximately 9500 (100 frames per second). Each frame is dynamically associated with a timestamp and energy parameters, resulting in an audio frame misalignment rate of less than 0.1%. After silence removal, the effective speech ratio of the signal is increased to over 95%. This method guarantees the temporal accuracy of frame-level sentiment analysis and provides a highly adaptable temporal audio data foundation for multimodal temporal feature fusion.
[0119] S4.2: Based on time-series audio frames, a neural network-based emotion recognition algorithm is used to extract subjective emotional states (such as happiness, sadness, doubt, etc.) to obtain frame-level emotion feature parameters, providing a basis for subsequent tone and speed analysis.
[0120] The input is time-series audio frame data after silence removal and frame-level segmentation, which includes the main channel speech signal with frame-by-frame amplitude normalization.
[0121] An emotion recognition algorithm based on deep neural networks is adopted (parameters: network structure type, such as Bi-LSTM or CRNN, input frame length T). frame =25ms, frame shift T shift =10ms, audio feature types include MFCC, Chroma, Spectral-Contrast, etc.), to realize the function of extracting emotional state features of each frame of audio data.
[0122] Furthermore, the front-end multidimensional acoustic feature extraction module (parameters: MFCC order M=13, including first and second order differences, number of Chroma channels C=12, spectral contrast channels S=6) performs time-frequency domain feature encoding on each frame of audio signal to form a D-dimensional frame-level acoustic feature vector.
[0123] Furthermore, the standardized frame-level acoustic feature vectors are used as continuous inputs to the neural network emotion recognition model. End-to-end training and inference are performed using a labeled emotion corpus (parameters: emotion category set E={happy, sad, confused, neutral}, category labels are one-hot encoded, and training loss is the cross-entropy loss function), outputting the probability distribution of subjective emotion state corresponding to each frame.
[0124] Furthermore, through a backend sentiment state classifier (such as a fully connected layer + Softmax activation, with a sentiment classification threshold θ)... c =0.5), which maps the output probability vector to specific frame-level sentiment feature parameters, including: representative subjective sentiment state labels, sentiment intensity levels (such as three-level stratification: low-medium-high), confidence scores, etc.
[0125] Furthermore, a temporal consistency check and dynamic smoothing processing strategy (such as majority voting sliding window smoothing method and sliding window length L=5 frames) is adopted to remove noise and suppress short-term fluctuations in the continuously output frame-level sentiment tag sequence, so as to ensure the temporal stability of sentiment feature parameters.
[0126] Through the above chain-like derivation steps, the temporal audio frame data is transformed into frame-level emotional feature parameters with high spatiotemporal resolution, realizing the dynamic extraction of emotional state in multi-frame speech streams, and providing accurate dependencies and temporal benchmarks for subsequent tone, speed and semantic unit analysis.
[0127] For example, in a real-time dialogue scenario with a sampling rate of 48kHz, a frame length of 25ms, and a frame shift of 10ms, a deep Bi-LSTM network was selected as the emotion recognition model. The front end used MFCC parameters M=13 and Chroma features C=12 for feature concatenation, with an input feature dimension D=25. The publicly available RAVDESS speech emotion corpus was used as the training set, with emotion tags including six categories: happy, sad, angry, fearful, calm, and neutral. The batch size was 128, the training epochs were 40, and the early stopping tolerance was 5. In practical applications, each second of audio can be decomposed into 100 frames for emotion prediction, with an average processing latency of less than 12ms per frame and an output emotion accuracy of 90.7%. For continuous speech segments with dynamic emotion switching such as "happy-neutral-confused," after majority voting window smoothing, the frame-level mutation rate decreased to 1.4%, ensuring the dynamic consistency and recognition stability of the emotion state, and providing high-quality frame-level emotion feature parameters for subsequent tone feature layering and context label fusion.
[0128] S4.3: Utilizing frame-level sentiment feature parameters, the deep feature analysis module is applied to calculate the intonation layer features (such as pitch curves and energy changes) of the current audio frame to output refined intonation parameters, providing data support for speech rate and semantic unit detection.
[0129] The input consists of frame-level emotion feature parameters obtained through a deep neural network emotion recognition model, including the subjective emotion state label and corresponding confidence level for each frame of audio signal.
[0130] Multidimensional acoustic feature analysis method is adopted (parameters: Short Time Fourier Transform (STFT) window length T). stft =25ms, FFT points N fft =512, sliding step size T shift =10ms), to realize the frequency domain transformation of each frame of audio waveform, obtain the amplitude spectrum and phase spectrum, and realize the basic feature analysis of pitch profile and energy time-varying.
[0131] Furthermore, pitch extraction algorithms (such as autocorrelation or YIN algorithm, parameter: minimum pitch f) are used. min =50Hz, maximum pitch f max =500Hz, detection window width T w =40ms), perform fundamental frequency detection on the spectrum signal, extract the fundamental frequency F0 of each frame, and generate pitch curve features.
[0132] Furthermore, a short-time energy statistics algorithm is employed (parameter: energy calculation window T per frame). e =25ms), calculate the energy value E(n) for each frame to achieve fine quantization of the dynamic energy profile. The specific calculation formula is as follows:
[0133]
[0134] Where E(n) represents the energy of the nth frame, x n (i) represents the amplitude of the i-th sampling point within the n-th frame.
[0135] Furthermore, based on the hierarchical feature analysis method (parameter: intonation volatility S) var Average energy E mean The pitch change rate ΔF0) generates frame-level intonation layering parameters that clearly reflect fine-grained intonation content such as pitch rise and fall and energy fluctuation. Specific intonation parameters may include: average pitch, pitch variance, pitch slope, energy mean, energy standard deviation, etc.
[0136] Furthermore, a multi-parameter fusion strategy (parameters: feature splicing weights α and β) is adopted to splice the pitch curve parameters and energy parameters of each frame into a complete and refined intonation feature vector, providing an input basis for subsequent speech rate detection and semantic unit recognition.
[0137] Through the chain-like feature extraction and hierarchical modeling process described above, frame-level emotional feature parameters are mapped to multi-dimensional refined intonation parameters, effectively supporting subsequent analysis and processing of temporal speech, and realizing in-depth quantitative expression of speech expression style and emotional state.
[0138] For example, at a sampling rate of 48kHz and a frame length T stft =25ms, FFT points N fft In a speech segment with a frequency of 1024, for contexts containing three emotional states—"happy," "confused," and "neutral"—STFT transformation is performed to obtain the amplitude spectrum of each frame, and f is set based on the YIN algorithm. min =80Hz, f max =400Hz pitch curve extraction. Energy value is calculated synchronously for each frame, with average energy E mean =0.43, pitch mean F0=196Hz, when happy the F0 mean rises to 220Hz, energy standard deviation σ E The value is 0.12. Based on the statistical analysis of the pitch change rate ΔF0, the ΔF0 in the happy to doubt range increases to 28 Hz / s, and the energy fluctuation rate S... var The value was increased to 0.35 (higher than 0.18 for neutral paragraphs). After concatenating the parameters of each layer, a refined intonation parameter vector with a dimension of 8 was generated for each frame. In actual detection, the intonation rise and fall and energy features under different emotional states were distinguished, achieving high spatiotemporal accuracy in intonation feature modeling, which provides a refined feature foundation for subsequent speech rate and semantic unit discrimination.
[0139] S4.4: Input the intonation parameters and audio frame sequence into a temporal feature modeling algorithm (such as LSTM or Transformer structure) to calculate speech rate index and semantic unit boundary information to generate speech rate feature parameters and clustered semantic unit sequence.
[0140] S4.5: Based on emotion feature parameters, intonation parameters, speech rate feature parameters and semantic unit sequences, perform multi-dimensional feature label hierarchical mapping, map the above parameters into speech context feature labels, and output a structured label set for context-aware input of the subsequent multimodal joint inference network.
[0141] Step S5: Based on facial key point parameters, head pose parameters, eye region features, and speech context feature labels, the above multi-source data is input into a multimodal joint inference network to encode the complete context-aware features of the current speaker. Specifically, this includes:
[0142] S5.1: Facial key point parameters, head pose parameters, and eye region features are processed by unified feature vector encoding. A multi-dimensional vector concatenation algorithm is used to standardize each parameter according to the facial feature data structure. All face-related parameters are integrated into a unified multi-source facial feature vector as the initial input for multimodal inference.
[0143] S5.2: Use a contextual feature mapping algorithm to embed the speech contextual feature labels and audio-related parameters (such as the emotional state, tone, speech rate, and semantic unit information of the speech) into a contextual embedding vector that is consistent with the facial feature dimension, and ensure that the vector can express the diversity of emotional drive and language expression style.
[0144] The input data includes speech context feature labels and audio-related parameters obtained through preprocessing. The audio parameters cover multi-dimensional speech expression data such as emotional state, intonation, speech rate, and semantic units.
[0145] Employing a contextual feature mapping algorithm (parameter: number of label categories N) c Audio parameter dimension D a Target embedding dimension D e This enables the linked embedding of speech context feature labels and audio-related parameters.
[0146] Furthermore, through a multi-layer label coding network (parameters: layer coding depth L=2~4, activation function type such as ReLU or GELU), independent label vectors are constructed for sentiment state labels, intonation parameters, speech rate parameters, and semantic unit labels respectively. Discrete labels and continuous audio parameters are jointly projected into a unified high-dimensional space to achieve the coding alignment of discrete and continuous context factors.
[0147] Furthermore, a label-parameter interactive mapping module is employed to perform feature-level fusion of the original label encoding vector and the audio feature vector, while simultaneously introducing a weighted modulation function f. adm (·) To achieve adaptive adjustment of expression style under different context types, the specific mapping formula is as follows:
[0148]
[0149] in, To output context embedding vectors, Embed vectors for sentiment state labels. For intonation feature vectors, For speech rate feature vectors, For semantic unit label embedding vectors, This indicates a splicing operation. For trainable fusion matrix, This is an adaptive activation function.
[0150] Furthermore, style enhancement constraints are applied (parameter: diversity regularization coefficient λ). div Number of style categories K style Add a diversity loss auxiliary term to enhance the output context embedding vector. It enhances the ability to differentiate between different expressive styles, thereby improving adaptability to changes in actual context and facial expression-driven responses.
[0151] Through on-chain derivation, multi-dimensional contextual labels and audio features are transformed into high-dimensional contextual embedding vectors consistent with facial feature dimensions through multi-layer embedding and interactive mapping, ensuring that the vectors can express rich emotional drive, language expression style and semantic change characteristics.
[0152] For example, in a scenario involving the correction of real-person emotional expression videos, the emotional category set N is selected. c =6 (happy, sad, confused, angry, calm, neutral), audio parameter input dimension D a =16 (including features such as pitch, pitch variance, energy, and speech rate / rhythm), target context embedding dimension D e =64. A three-layer label encoding network is used, with ReLU activation function in each layer and an interaction projection matrix W. int The dimensions are 22×64. The style enhancement constraint sets λ. div =0.15, number of style categories K style =5. Through actual mapping processing, the embedding vectors formed by different combinations of emotion, intonation, speech rate, and semantic units produce clearly separable style clusters in spatial distribution, with an average Cosine distance improvement of 20%. For the same facial feature vector input, switching the output V... ctxThis process effectively drives the adjustment of eye deformation parameters to present natural facial expression changes, improving the subjective score of emotional fit in the corrected video from 3.5 points to 4.6 points (out of 5). This processing workflow provides a highly discriminative and adaptable context-driven representation for subsequent multimodal feature fusion and natural eye expression correction.
[0153] S5.3: Perform multimodal feature fusion processing on the above-mentioned uniformly encoded facial feature vector and context embedding vector. Use a cross-modal fusion network algorithm to hierarchically aggregate facial features and contextual expression features, and add scene label archive information during the fusion process to achieve feature depth supplementation and association enhancement for different scene distributions.
[0154] The input objects are normalized multi-source facial feature vectors and high-dimensional context embedding vectors, which respectively cover facial key point parameters, head pose parameters, eye region structural features, and speech context feature labels and speech expression dimension features.
[0155] A cross-modal feature fusion network algorithm is adopted (parameters: network type Multi-Modal FusionTransformer, number of fusion layers L). f =4, feature mapping dimension D f =128), realizing hierarchical joint aggregation of multi-source facial feature vectors and context embedding vectors, thereby completing the signal space alignment and interaction modeling of facial representation and context expression.
[0156] Furthermore, by fusing attention mechanisms (parameter: number of multi-head attention heads N) head =8, normalization method LayerNorm), to achieve dynamic weighting of feature dependencies between facial features and contextual features in each fusion layer, improving the ability to model correlations between different modalities and enhancing the ability to perceive fine-grained contextual changes. A cross-modal residual connection module is applied (parameter: residual channel dimension R). d =64), ensuring effective feature gradient transfer during deep aggregation, suppressing information loss introduced by multi-source signal fusion, and providing a stable convergence path for the hierarchical output structure.
[0157] Furthermore, scene tag archive information is used to assist in the fusion processing (parameter: scene tag category S). c =12, label embedding dimension S e =16), performing feature concatenation and dynamic gating mapping through scene label vectors and cross-modal fusion layer output (parameter: gating threshold T). g =0.35), enabling feature depth supplementation and result fine-tuning under different scenario conditions.
[0158] Furthermore, the fusion activation and normalization steps are performed (parameters: activation function Swish / LayerNorm normalization standard), outputting the fused high-dimensional multimodal representation vector, which lays a structured feature foundation for its subsequent context-aware feature encoding and correction parameter generation.
[0159] By using cross-modal fusion networks and scene label supplementary processing, the uniformly encoded facial feature vectors and context embedding vectors are transformed into fusion feature expressions adapted to multiple scene distributions, realizing deep aggregation of face-context features in rich scenes. This can support subsequent accurate modeling that distinguishes emotions and restores real-world context perception features.
[0160] For example, in a real-world video correction scenario, the facial feature vector dimension is set to D. m =72, the context embedding vector dimension is set to D. e =64, Scene label category S c =8, Tag embedding dimension S e =16. The fusion network adopts a 4-layer Transformer structure, with a single layer feature dimension of 128, and uses 8-head multi-head attention. The input data includes head pose changes such as yaw angle -15°~+15° and pitch angle -10°~+10°, as well as contextual features corresponding to three sentiment labels: "happy," "confused," and "neutral." Scene labels such as "indoor backlight," "low illumination," and "outdoor bright" are dynamically concatenated into each fusion network layer after label embedding. The output dependency weights of each fusion layer can be calculated using the following attention mechanism:
[0161]
[0162] in, For query vectors (Query). The key vector (Key) Value vector. The feature dimension is defined as follows. After Attention weighting, the interaction and fusion of facial features and contextual features are achieved. Residual connections ensure the stability of the output of each layer.
[0163] After actual processing, the root mean square amplitude of the fused feature vector increased by 12%, and the contextual discrimination improved by 18%. With the assistance of multi-scene labels, the fusion network's ability to adjust to changes in emotion and scene is significantly enhanced, achieving dynamic and accurate adaptation of eye correction parameters to context, emotion, and shooting environment, effectively avoiding the rigid output under single emotion or standard scene conditions. The final output fused feature vector provides highly compatible input for subsequent context-aware feature modeling and eye image deformation parameter generation, significantly improving the naturalness and trustworthiness of the overall eye correction.
[0164] S5.4: Based on the fused multimodal feature output, a context-aware feature encoding algorithm is executed to perform temporal modeling and dynamic contextualization on the aggregated feature vectors in order to obtain complete context-aware features that can reflect the current speaker's eye gaze, facial expression and emotional state, and to establish a standard expression space for the subsequent calculation of eye image deformation parameters.
[0165] S5.5: The encoded context-aware features are used as the final output for the eye image deformation parameter generation module to call, so as to realize the continuous input of the basic features of eye adjustment driven by facial structure, dynamic expression and emotional context, and ensure the accurate coordination between eye image deformation and dynamic emotional changes.
[0166] Step S6: Based on the encoded context-aware features, calculate the eye image deformation parameters, expand them into parameterized gaze accommodation amplitude and expression change coefficient control factors, and set dynamic adjustment thresholds according to different emotional scenarios to improve the adaptability of eye naturalness. Specifically, this includes:
[0167] S6.1: The context-aware feature vector is decoded, and a multimodal adaptive fusion algorithm is used to obtain a structured set of context-related control factors, providing basic data support for subsequent eye deformation parameterization calculations.
[0168] S6.2: Using the context-related control factor as input, an eyeball deformation generation model based on differentiable rendering is applied to calculate the standard eyeball offset parameters and form preliminary gaze accommodation amplitude features as independent variables for deformation processing.
[0169] Using a structured set of context-related control factors as input, standardized eyeball deformation control parameters are generated. An eyeball deformation generation model based on differentiable rendering is employed (parameter settings include the three-dimensional geometric constraints θ of the eyeball). e Rendering resolution R e =128×128, angle mapping range Δα=[-25°, +25°]), and parameterized deformation calculation of the eye region image is achieved based on control factors. This is achieved through a three-dimensional spherical mapping algorithm (parameter: sphere center coordinates C). e radius r e Initial line of sight This establishes a correspondence between the current physical position of the eyeball and the standard emmetropic coordinate system, and obtains the original gaze vector. .
[0170] Furthermore, by combining the input multimodal context factors, an adaptive neural deformable network (parameters: number of neural units N_h=128, activation function Swish, input layer splicing control factor and eye structure features) is used to predict the required gaze offset angle Δα in the differentiable rendering module, thus obtaining the standard eyeball offset parameters.
[0171] The mapping between spherical coordinates and offset angles is accomplished using the following formula:
[0172]
[0173] in, This is the yaw angle rotation matrix. This is the pitch angle rotation matrix. To correct the direction of the line of sight.
[0174] Furthermore, the standard eyeball deviation parameter ( , Input to the differentiable image deformation pipeline (parameter: affine transformation matrix T) e The sampling interpolation method is bilinear interpolation, which transforms the pixel coordinates of the eye region to generate a preliminary feature map of gaze accommodation amplitude. An affine transformation is performed on the pixel positions according to the following formula:
[0175]
[0176] in, These are the original pixel coordinates. Let be the affine transformation matrix. It is a translation vector.
[0177] Through the above chain processing, the generated standard eyeball deviation parameters and preliminary gaze accommodation amplitude features not only accurately reflect changes in facial structure, but also respond to emotion and context-driven data input, achieving real-time adaptive deformation.
[0178] By using a differentiable rendering-based eye deformation generation model and multi-step geometric parameter mapping, the input contextual correlation control factors are accurately transformed into standard eye deviation parameters and preliminary gaze accommodation amplitude features, significantly improving the naturalness and emotional adaptability of eye correction.
[0179] For example, in a scenario where facial expression correction is applied to real-person speech, the input contextual association control factors include an emotion embedding vector (16 dimensions), scene labels (8 dimensions), and head posture parameters (pitch angle -5°, yaw angle 10°). Based on a three-dimensional spherical mapping model, the eyeball radius r is set. e =12mm, initial line of sight After training, the differentiable neural network can output standard eye deviation parameters for contexts corresponding to "happy-positive" emotions. Affine transformation pipelines are used ( Based on the above perspective, using bilinear interpolation, deformation processing was performed on the right eye region image at a resolution of 128×128, outputting the corrected eye image. Test results show that the corrected eye expression naturally matches the emotional atmosphere, and the user's subjective naturalness score improved from 3.2 points before correction to 4.7 points (out of 5) after deformation, demonstrating high contextual adaptability and parameter responsiveness.
[0180] S6.3: Based on the preliminary gaze accommodation amplitude features and context-related control factors, the emotion-driven expression generation algorithm is used to calculate expression change coefficients (such as blinking frequency and pupil dilation) to supplement eyeball deformation parameters, thereby achieving adaptive expansion of dynamic facial expression details.
[0181] S6.4: Combining the emotional change coefficient and the standard gaze adjustment amplitude characteristics, a scene-adaptive threshold setting strategy is adopted for different emotional scene labels to dynamically allocate the limitation range of gaze adjustment amplitude and facial expression detail changes, thereby achieving personalized parameter optimization.
[0182] S6.5: The final obtained parametric gaze adjustment range, expression change coefficient, and scene adaptive threshold are constructed into a continuous and controllable eye image deformation parameter sequence for use in the next step of eye rendering and image compositing module, so as to achieve overall naturalness and emotional consistency.
[0183] Step S7: For each frame of the facial image, according to the generated eye image deformation parameters and expression change coefficients, perform eye deformation, shading adjustment, and image compositing processing based on differentiable rendering, while maintaining dynamic coordination with head posture and facial expression. Specifically, this includes:
[0184] S7.1: Using a differentiable rendering algorithm, the input facial image frame and the synchronously acquired eye image deformation parameters are processed to perform geometric deformation operations on the eye region to adjust the direction and shape of the eyeball, thereby changing the line of sight deviation in the image and obtaining the deformed eye structure feature map.
[0185] S7.2: Based on the deformed eye structure feature map and expression change coefficient, a context-adaptive color adjustment algorithm is used to refine the color processing of the iris brightness, pupil color level, blinking state, etc., in order to restore natural emotional expression and obtain an eye color feature map with enhanced realism.
[0186] S7.3: The realistic eye color feature map is integrated and fused with the original facial image at the region level. Through a dynamic coordination fusion algorithm driven by head pose parameters and facial expression, the eye region is ensured to be generated in fine-grained synchronization with other parts during the overall facial expression dynamic process, so as to obtain a dynamically coordinated facial composite frame.
[0187] Using the enhanced eye color feature map and the original facial image as input data, the regional integration and fusion module starts working to realize the spatial set of facial features in different regions.
[0188] A region fusion algorithm based on multi-scale weighted masks is employed (parameter settings: fusion mask resolution consistent with eye coloring features, boundary smoothing coefficient β=0.85). This algorithm achieves physical masking and content fusion through dynamic mask calculation, targeting the pixel space of the eye region and the original face. Mask weights are used. Adjust dynamically according to the following formula:
[0189]
[0190] in This is the Euclidean distance from the current pixel to the center pixel of the eye. and These represent the local mean and standard deviation of the mask region, respectively. It is a fusion modulator driven by facial expressions.
[0191] Furthermore, by using head pose parameters (such as pitch and yaw angles) and facial expression driving vectors (such as low-dimensional expression coefficients p)... exp The dynamic synchronization fusion algorithm [n] analyzes the overall head movement trend and changes in the main components of facial expressions in real time, and realizes the motion vector field correction between the eye area and the dynamic enhancement area of the whole face.
[0192] A time-coordinated motion compensation mapper (parameters: motion compensation window length L=5 frames, interpolation method is bicubic interpolation) dynamically corrects the spatial pose of the eye region during the fusion stage, ensuring that it is fully aligned with the head pose change, thus resolving motion artifacts and dynamic fragmentation caused by local deformation.
[0193] Furthermore, a fine-grained synchronous generation controller based on contextual conditions (parameter: contextual label dimension d=8) is adopted to dynamically adjust the fusion coefficient and blending mask area according to the emotion and scene parameters of the previous stage, thereby enhancing the synchronous rendering performance of dynamic appearance changes such as blinking and eye muscle contraction.
[0194] By leveraging a high-efficiency GPU-parallel regional pixel mixing pipeline, dynamically coordinated facial composite frames are output, laying the foundation for spatial and lighting consistency in subsequent sequential frame correlation and the output of multi-frame continuous and consistent eye-correction video clips.
[0195] Through the above chain fusion process, the realistic eye color feature map and the original facial image are transformed into a dynamic facial composite frame with a high degree of consistency in global structure and dynamic expression, achieving a seamless integration of eye correction and context-adaptive visual effects.
[0196] For example, in a video clip of a speaker looking down and smiling briefly, the eye color feature map resolution is 128×128, the fused mask resolution is 128×128, β=0.85, the head pose parameters are pitch angle -10°, and the expression principal component coefficients p exp [smile]=0.6. Using the aforementioned fusion algorithm, the mask... The fusion exhibits Gaussian attenuation in the peripheral eye region, while the central fusion weight remains between 0.97 and 1.00. The motion compensation mapper continuously corrects spatial alignment using a 5-frame sliding window, achieving a seamless transition between the eye region and the overall head dynamics during speech. The naturalness score of the synthesized frame's subjective visual evaluation significantly improved from 3.6 (out of 5) to 4.8, and the inter-frame expression synchronization error decreased from 8.3 pixels to 1.2 pixels, effectively validating the technical effectiveness and fine control capability of the region-level dynamic coordination fusion strategy in high-naturalness eye expression correction scenarios.
[0197] S7.4: Perform continuous frame temporal correlation processing on the dynamically coordinated facial composite frames and the corresponding facial key point parameters and head pose parameters. Use a dynamic temporal fusion algorithm to ensure that the inter-frame motion trajectory after eye correction is completely consistent with the facial expression changes, and output multi-frame continuous and consistent eye correction video clips.
[0198] S7.5: Based on the output multi-frame continuous and consistent eye-correction video clips, the adaptive quality assessment module automatically scores the naturalness and emotional expression of each frame, and uses the scoring results as feedback to adjust the subsequent eye image deformation parameters and expression change coefficient settings, so as to achieve collaborative optimization of image synthesis and context-driven processing.
[0199] Step S8: The synthesized video segments are subjected to a continuous naturalness and emotional consistency quality assessment. An automatic scoring model combined with front-end user feedback is used to dynamically optimize the eye-tracking adjustment factor based on different scene tags, achieving an adaptive quality closed loop for the generated results. Specifically, this includes:
[0200] S8.1: Perform frame-level feature extraction on the synthesized video clips, and output facial expression parameters and eye movement parameters based on the differentiable rendering results to obtain the data basis for subsequent quality assessment.
[0201] S8.2: Input frame-level facial expression parameters, eye movement parameters, and scene labels into the automatic scoring model, and generate a quality score vector for each frame by quantifying naturalness and emotional consistency based on computer vision and speech analysis algorithms.
[0202] S8.3: Based on the output of the automatic scoring model, perform continuous segment-level statistical analysis, including naturalness fluctuation analysis and sentiment consistency temporal assessment, to identify low-quality segments and distorted segments in video clips.
[0203] S8.4: Collect front-end user feedback, including subjective naturalness scores and contextual fit markings, and fuse user reviews with the output of the automatic scoring model to form a multi-dimensional quality perception feature set.
[0204] S8.5: The multi-dimensional quality-perceived feature set and video scene label input parameter optimization module adjust the eye adjustment factor setting based on the optimization algorithm to achieve adaptive optimization of parameters for subsequent synthesized segments.
[0205] S8.6: The optimized eye expression adjustment factor is dynamically fed back to the eye deformation and image compositing module to form an adaptive quality closed-loop mechanism, continuously improving the naturalness and emotional consistency of the generated results.
[0206] For those skilled in the art, various other corresponding changes and modifications can be made based on the technical solutions and concepts described above, and all such changes and modifications should fall within the protection scope of the claims of this invention.
[0207] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains. The terms “first,” “second,” “third,” and similar terms used in this patent application specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “comprising” or “including” and similar terms mean that the elements or objects preceding “comprising” or “including” encompass the elements or objects listed following “comprising” or “including” and their equivalents, and do not exclude other elements or objects. The “multiple” mentioned in the embodiments of this application refers to two or more. A and / or B indicate three possibilities: A; B; and A and B.
[0208] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A video eye-tracking correction method based on image processing, specifically including: S1: Acquire facial image features frame by frame from the input video stream, and simultaneously acquire the corresponding real-time audio signal; S2: Perform denoising and normalization processing on the facial image features and audio signals, respectively; S3: Extract head pose parameters and eye region features from the preprocessed video frames, and archive scene labels using different shooting angles and lighting conditions. S4: Input the preprocessed audio samples into the speech emotion recognition model to obtain the emotional state, tone, speech rate and semantic unit information of the speech, and establish speech context feature label mapping for different language expression styles; S5: Input multi-source data, including facial key point parameters, head posture parameters, eye region features, and speech context feature labels, into a multimodal joint inference network to encode the complete context perception features of the current speaker; S6: Calculate eye image deformation parameters based on the obtained context perception features, expand them into parameterized gaze accommodation amplitude and expression change coefficient control factors, and set dynamic adjustment thresholds according to different emotional scenarios; S7: For each frame of facial image, perform eye deformation, color adjustment and image compositing processing according to the generated eye image deformation parameters and expression change coefficients, and maintain dynamic coordination with head posture and facial expression, and output the image and video after eye correction.
2. The video eye-tracking correction method based on image processing according to claim 1, characterized in that: Step S7 is followed by S8: continuously naturalness and emotional consistency quality assessment of the synthesized video segments, dynamically optimizing the eye adjustment factor setting according to different scene labels to achieve an adaptive quality closed loop of the generated results.
3. The video eye-tracking correction method based on image processing according to claim 1, characterized in that: In step S2, facial image denoising is performed using median filtering and multi-scale feature extraction via a convolutional neural network.
4. The video eye-tracking correction method based on image processing according to claim 1, characterized in that: The speech emotion recognition model in step S4 adopts the Transformer model.
5. The video eye-tracking correction method based on image processing according to claim 4, characterized in that: The voice context feature labels are divided into six categories: happy, sad, confused, neutral, and angry.
6. The video eye-tracking correction method based on image processing according to claim 1, characterized in that: Step S7 includes spatial motion correction based on head posture parameters, facial key point coordinates, and facial expression principal component vectors to ensure that the eye deformation area is synchronized with other facial areas in time and space.
7. The video eye-tracking correction method based on image processing according to claim 1, characterized in that: It also includes the structured storage of the timestamp, emotion tag, and quality score of each frame of the processed eye-correction video output, in order to support subsequent segmented backtracking, content indexing, and personalized playback.
8. The video eye-tracking correction method based on image processing according to claim 1, characterized in that, Step S3 specifically includes: based on the preprocessed video frame data, applying a facial key point localization algorithm to extract key point coordinates for each frame to obtain facial key point parameters; applying spatial mapping and geometric calculation algorithms to calculate three-dimensional head pose parameters for the extracted facial key point parameters; combining facial key point parameters and three-dimensional head pose parameters to accurately extract the eye region to obtain eye region features; applying a scene label archiving algorithm to classify scene features for each frame based on the shooting angle and illumination distribution characteristics of the video frames, generating scene label parameters; and performing feature fusion processing on facial key point parameters, head pose parameters, eye region features, and scene label parameters to construct a multi-scene differentiated facial feature structure.
9. A video eye-tracking correction method based on image processing according to claim 5, characterized in that, The facial landmark detection model is a deep learning network, and 100 landmarks are extracted.
10. A video eye-tracking correction system based on image processing, wherein the video eye-tracking correction method based on image processing as described in any one of claims 1-9 is used for video eye-tracking correction.