Intelligent sight interaction system for emotional communication

Through the intelligent line of sight interaction system with multimodal data fusion, the problem of insufficient accuracy and real-time performance of line of sight interaction system is solved, high-precision line of sight tracing and natural emotional communication are realized, and the mimicry of human-computer interaction is improved.

CN120276595APending Publication Date: 2025-07-08HUNAN SANY IND VOCATIONAL & TECH COLLEGE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510352159.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing line-of-sight interaction system has shortcomings in accuracy and real-timeness, making it difficult to accurately capture user vision and emotions, resulting in poor interaction experience.

Method used

The sight capture module, head posture recognition module, emotion analysis module, digital life generation module and interaction control module are adopted to generate user emotional tags and drive digital people's sight, expressions and actions through multi-modal data fusion to achieve natural eye contact and emotional communication.

Benefits of technology

Significantly improve the accuracy of gaze tracing, enhance the comprehensiveness of emotional recognition, optimize the synergy of multimodal output, and improve the mimicry degree of human-computer interaction and the depth of emotional communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276595A_ABST
    Figure CN120276595A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent sight interaction system for emotional communication. The system comprises a sight capture module, a head posture recognition module, an emotion analysis module, a digital human generation module, an interaction control module and a multi-mode output module. According to the system, sight line tracking is carried out through collected eye images, head posture correction is carried out based on face key points and a 3D model, sight line / expression / posture cooperative processing is carried out through a multi-mode sentiment analysis model, and the technical means of dynamically generating digital human sight line and expression parameters is combined; the problems of tracking errors and single emotion feedback caused by head movement in traditional sight line interaction are solved, high-precision sight line capture and emotional digital human real-time response are achieved, man-machine interaction reaches the naturalness degree close to real sight contact, and meanwhile the adaptability to the emotional state of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of human-computer interaction and affective computing, and particularly relates to an intelligent gaze interaction system for emotional communication. Background Art

[0002] In the current field of human-computer interaction, visual interaction is an important way to achieve natural and user-friendly interaction. As an important non-verbal signal in the communication process, gaze can convey attention and emotional information. However, existing interaction systems have limited support for gaze interaction, lacking accurate capture of the user's gaze and natural generation of the digital human's gaze, making it difficult to achieve real eye contact and emotional communication.

[0003] Traditional gaze tracking technologies have deficiencies in terms of accuracy and real-time performance, being restricted by factors such as environmental light and device performance. The gaze control of digital humans often adopts preset rules, lacking real-time response to the user's gaze and emotional state, resulting in a poor interaction experience.

[0004] Therefore, there is a need for an innovative gaze interaction technology that can accurately capture the user's gaze and emotions, and achieve natural generation of the digital human's gaze and emotional interaction. Summary of the Invention

[0005] Based on this, it is necessary to provide an intelligent gaze interaction system for emotional communication in view of the above technical problems.

[0006] In a first aspect, the present application provides an intelligent gaze interaction system for emotional communication, which includes: a gaze capture module, a head pose recognition module, an emotion analysis module, a digital human generation module, an interaction control module, and a multimodal output module;

[0007] The gaze capture module is configured to perform gaze capture processing based on the user's eye image, and generate a gaze direction and fixation point coordinates;

[0008] The head pose recognition module is configured to perform head pose recognition processing based on the user's face image to obtain a head pose recognition result; and correct the gaze direction and fixation point coordinates according to the head pose recognition result to generate a corrected gaze direction and corrected fixation point coordinates;

[0009] The emotion analysis module is configured to perform multimodal emotion analysis processing on the corrected gaze direction, corrected fixation point coordinates, head pose recognition result, and the user's facial expression to generate a user emotion label;

[0010] The digital human generation module is configured to generate a digital human gaze direction, digital human expression parameters, and digital human action instructions according to the corrected gaze direction, corrected fixation point coordinates, and user emotion label;

[0011] An interaction control module, configured to generate an interaction control strategy and corresponding digital human speech text based on user emotion tags and corrected gaze point coordinates, and perform interaction control on the digital human according to the interaction control strategy; the control objects of the interaction control include: the gaze direction of the digital human, the expression parameters of the digital human, and the motion instructions of the digital human;

[0012] A multimodal output module, configured to perform multimodal output on the gaze direction of the digital human, the expression parameters of the digital human, the motion instructions of the digital human, and the digital human speech text, and generate audiovisual interaction content.

[0013] In a second aspect, the present application further provides an intelligent gaze interaction method for emotional communication, applied to the system in the first aspect, including:

[0014] S1: The gaze capture module performs gaze capture processing based on the user's eye image to generate a gaze direction and gaze point coordinates;

[0015] S2: The head pose recognition module performs head pose recognition processing based on the user's face image to obtain a head pose recognition result; according to the head pose recognition result, correct the gaze direction and gaze point coordinates to generate a corrected gaze direction and corrected gaze point coordinates;

[0016] S3: The emotion analysis module performs multimodal emotion analysis processing on the corrected gaze direction, corrected gaze point coordinates, head pose recognition result, and user facial expression to generate user emotion tags;

[0017] S4: The digital human generation module generates the gaze direction of the digital human, the expression parameters of the digital human, and the motion instructions of the digital human according to the corrected gaze direction, corrected gaze point coordinates, and user emotion tags;

[0018] S5: The interaction control module generates an interaction control strategy and corresponding digital human speech text based on user emotion tags and corrected gaze point coordinates, and performs interaction control on the digital human according to the interaction control strategy; the control objects of the interaction control include: the gaze direction of the digital human, the expression parameters of the digital human, and the motion instructions of the digital human;

[0019] S6: The multimodal output module performs multimodal output on the gaze direction of the digital human, the expression parameters of the digital human, the motion instructions of the digital human, and the digital human speech text, and generates audiovisual interaction content.

[0020] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements an intelligent gaze interaction method for emotional communication as described in the second aspect.

[0021] Fourthly, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements an intelligent gaze interaction method for emotional communication as described in the second aspect.

[0022] The above intelligent gaze interaction system for emotional communication captures gaze by processing the user's eye images to generate the gaze direction and fixation point coordinates; performs head pose recognition processing on the user's face images to obtain the head pose recognition result; corrects the gaze direction and fixation point coordinates according to the head pose recognition result to generate the corrected gaze direction and corrected fixation point coordinates; performs multi-modal emotion analysis processing on the corrected gaze direction, corrected fixation point coordinates, head pose recognition result, and the user's facial expression to generate the user emotion label; generates the digital human gaze direction, digital human expression parameters, and digital human action instructions according to the corrected gaze direction, corrected fixation point coordinates, and user emotion label; generates an interaction control strategy and corresponding digital human speech text based on the user emotion label and corrected fixation point coordinates, and controls the digital human according to the interaction control strategy; performs multi-modal output on the digital human gaze direction, digital human expression parameters, digital human action instructions, and digital human speech text to generate visual and auditory interaction content. Through the above technical means, the following technical effects are achieved: significantly improving the accuracy of gaze tracking, overcoming the problem of gaze deviation caused by head movement and environmental interference; enhancing the comprehensiveness of emotion recognition, accurately capturing the user's emotion changes through multi-modal data fusion; dynamically generating digital human interaction behaviors that match the user's gaze and emotion, achieving natural eye contact and emotional resonance; optimizing the coordination of multi-modal output to ensure a high degree of synchronization of visual and auditory feedback, and ultimately improving the fidelity and depth of emotional communication in human-computer interaction. Description of the Drawings

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is a schematic structural diagram of an intelligent gaze interaction system for emotional communication provided by the present invention;

[0025] Figure 2 It is a schematic flow diagram of an intelligent gaze interaction method for emotional communication provided by the present invention. Detailed Embodiments

[0026] In order to make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not intended to limit this application.

[0027] Reference Figure 1 , which shows a schematic structural diagram of an intelligent gaze interaction system 10 for emotional communication provided by this application. The system includes: a gaze capture module 11, a head pose recognition module 12, an emotion analysis module 12, a digital human generation module 14, an interaction control module 15, and a multimodal output module 16.

[0028] The gaze capture module 11 is used to perform gaze capture processing based on the user's eye image and generate a gaze direction and fixation point coordinates.

[0029] Specifically, in actual operation, the image data of the user's eye region can be continuously captured by a high-resolution camera at a frequency of 30 frames per second. The image data is transmitted to the image processing unit. First, it can be grayscale processed to reduce the data dimension and highlight key features. Subsequently, an eye movement tracking algorithm based on a convolutional neural network (CNN, Convolutional Neural Network) can be used. This algorithm is pre-trained on a large-scale labeled eye movement dataset and can accurately identify key feature points such as the contour of the eyeball, the center of the pupil, and the corners of the eyes. According to the relative positions and geometric relationships of the feature points, the gaze direction of the user is calculated and represented in the three-dimensional space in the form of an angle. At the same time, the two-dimensional coordinates of the fixation point in the screen or the real scene are determined.

[0030] The head pose recognition module 12 is used to perform head pose recognition processing based on the user's face image to obtain a head pose recognition result; and correct the gaze direction and fixation point coordinates according to the head pose recognition result to generate a corrected gaze direction and corrected fixation point coordinates.

[0031] Specifically, during the implementation process, an infrared depth camera can be used to capture the depth image and color image of the user's face. The depth image is used to obtain the three-dimensional structure information of the face, while the color image is used to extract appearance features such as texture. After fusing the two images, they are input into a pre-trained head pose estimation algorithm. Based on the 3D face model regression technology in deep learning, this algorithm can estimate the pitch angle, yaw angle, and roll angle of the head in real time, thereby obtaining a comprehensive head pose recognition result. After obtaining the head pose recognition result, it is fused and analyzed with the line-of-sight direction and fixation point coordinates generated by the line-of-sight capture module. For example, when the user's head tilts at a certain angle, the projection of the line-of-sight direction on the two-dimensional plane will change accordingly. By correcting this change using the head pose information and applying the principles of trigonometric functions and geometric transformations, the accurate angle value of the corrected line-of-sight direction in three-dimensional space and the precise pixel coordinates of the corrected fixation point on the screen can be recalculated.

[0032] The emotion analysis module 13 is used to perform multi-modal emotion analysis and processing on the corrected line-of-sight direction, corrected fixation point coordinates, head pose recognition result, and the user's facial expression, and generate user emotion labels.

[0033] Specifically, the changes in the line-of-sight direction and fixation point coordinates can be used as emotion clues. For example, if the user stares at a specific area for a long time, it may indicate a high degree of attention or interest in the content of that area, and frequent eye movement may reflect emotions such as hesitation or anxiety; changes in head pose can also reflect emotional tendencies, such as a lowered head may indicate depression or loss. At the same time, for facial expression images, deep learning-based facial expression recognition technology can be used. By using a CNN model pre-trained on a large amount of expression data, the facial muscle movement features can be extracted to identify basic emotion categories such as happiness, anger, sadness, and joy, as well as compound emotion states of different degrees. By fusing and analyzing the above multi-modal data, through the constructed emotion calculation model, the contribution degree of each data dimension to emotion can be comprehensively weighed, and finally accurate user emotion labels such as "joy - mild" and "anger - moderate" can be generated.

[0034] The digital human generation module 14 is used to generate the digital human's line-of-sight direction, digital human's expression parameters, and digital human's action instructions according to the corrected line-of-sight direction, corrected fixation point coordinates, and user emotion labels.

[0035] Specifically, according to the corrected line-of-sight direction, the animation parameters of the digital human model's eyes can be adjusted so that the digital human's gaze can naturally and accurately correspond to the user's line of sight, creating a realistic eye contact effect. For example, if the user's line of sight is biased to the right, the digital human's gaze will also turn to the right accordingly, with the angle coordinated with the user. At the same time, based on the user's emotion label, an expression parameter set matching this emotion can be extracted from the preset expression parameter library, and the shapes of the digital human's facial muscles, mouth corners, eyebrows and other parts can be adjusted to make the digital human present an expression corresponding to the user's emotion, enhancing emotional resonance. In addition, combined with the gaze point coordinate information, gesture, body posture and other action instructions of the digital human can be planned. For example, when the gaze point is concentrated on a certain interaction button, the digital human can make a gesture of pointing to the button and inviting a click, realizing a more natural and smooth human-computer interaction experience.

[0036] The interaction control module 15 is used to generate an interaction control strategy and the corresponding digital human speech text based on the user's emotion label and the corrected gaze point coordinates, and perform interaction control on the digital human according to the interaction control strategy; the control objects of the interaction control include: the digital human's line-of-sight direction, the digital human's expression parameters, and the digital human's action instructions.

[0037] Specifically, driven by the emotion label, if it is recognized that the user is in a negative emotion state, such as sadness or anger, the interaction control strategy can preferentially select soothing and calming interaction processes, such as guiding the digital human to express concern in a gentle tone, providing soothing music or relaxing interactive games, etc.; while when the user is in a positive emotion, more energetic and interactive strategies can be adopted, such as recommending interesting information, starting more challenging interactive tasks, etc. At the same time, according to the corrected gaze point coordinates, the interface area or interaction element that the user is currently focusing on can be determined, and the digital human's interaction behavior can be accurately controlled to focus on this area. For example, when the user is looking at a product display area, the digital human can actively introduce the detailed information, preferential activities and other content of the product, and guide the user to perform further operations, such as purchasing or viewing details, etc. The interaction control module converts the generated interaction strategy into specific control instructions to comprehensively regulate the digital human's behavior performance, including the dynamic adjustment of the line-of-sight direction, the real-time update of the expression parameters, and the orderly execution of the action instructions, etc.

[0038] The multi-modal output module 16 is used to perform multi-modal output on the digital human's line-of-sight direction, the digital human's expression parameters, the digital human's action instructions, and the digital human's speech text to generate visual and auditory interaction content.

[0039] Specifically, it receives the digital human's line-of-sight direction, expression parameters, and action instructions output by the digital human generation module, as well as the digital human speech text generated by the interaction control module. In terms of visual output, a graphics rendering engine can be used to render the dynamic eye effects of the digital human in real time according to the line-of-sight direction parameters; the texture and shape of each part of the digital human's face can be adjusted based on the expression parameters; the limbs and body of the digital human can be driven to perform corresponding action performances according to the action instructions, such as waving, nodding, making gestures, etc., and presented to the user in the form of a high-definition video stream through the screen, creating a realistic visual interaction experience. In terms of auditory output, the digital human speech text can be input into a speech synthesis system. Based on deep learning-based speech synthesis technology, this system can select appropriate intonation, speech rate, and timbre according to the user's emotion label. For example, for users with a sad emotion, a low and gentle intonation is used, and for users with an excited emotion, a cheerful and bright intonation is used to synthesize natural and fluent speech audio, which is played to the user through the speaker to achieve multi-modal emotional interaction coordinated with the visual output, comprehensively improving the quality of emotional communication and interaction effect between the user and the system.

[0040] The above intelligent line-of-sight interaction system for emotional communication captures and processes the line of sight based on the user's eye image to generate the line-of-sight direction and fixation point coordinates; performs head pose recognition processing based on the user's face image to obtain the head pose recognition result; corrects the line-of-sight direction and fixation point coordinates according to the head pose recognition result to generate the corrected line-of-sight direction and corrected fixation point coordinates; performs multi-modal emotional analysis processing on the corrected line-of-sight direction, corrected fixation point coordinates, head pose recognition result, and the user's facial expression to generate the user's emotion label; generates the digital human's line-of-sight direction, digital human expression parameters, and digital human action instructions according to the corrected line-of-sight direction, corrected fixation point coordinates, and the user's emotion label; generates an interaction control strategy and the corresponding digital human speech text based on the user's emotion label and corrected fixation point coordinates, and controls the digital human according to the interaction control strategy; performs multi-modal output on the digital human's line-of-sight direction, digital human expression parameters, digital human action instructions, and digital human speech text to generate visual and auditory interaction content. Through the above technical means, the following technical effects are achieved: significantly improving the accuracy of line-of-sight tracking, overcoming the problem of line-of-sight deviation caused by head movement and environmental interference; enhancing the comprehensiveness of emotion recognition, accurately capturing the user's emotion changes through multi-modal data fusion; dynamically generating digital human interaction behaviors that match the user's line of sight and emotion, achieving natural eye contact and emotional resonance; optimizing the coordination of multi-modal output to ensure a high degree of synchronization of visual and auditory feedback, ultimately improving the realism of human-computer interaction and the depth of emotional communication.

[0041] In an alternative embodiment, the line-of-sight capture module 11 includes:

[0042] The eye image acquisition unit 111 is used to acquire an image of the user's eye region to obtain the original eye image.

[0043] Specifically, in practical applications, a high-resolution color camera with a resolution of not less than 1280×720 pixels can be used, and the frame rate is stable at more than 30 frames per second. The camera can be fixed at the central position of the upper edge of the display directly in front of the user, and the horizontal distance from the user's eyes can be kept between 50 and 60 centimeters, and the vertical angle is slightly tilted downward by 15 to 30 degrees to capture the complete images of the user's both eyes to the greatest extent while reducing the possibility of eye occlusion. During the acquisition process, the camera converts the optical signal of the user's eye region into a digital image signal through its sensor to generate the original eye image, and the image data includes information such as the contours of the user's eyes, irises, pupils, eyelids, and facial features around the eyeballs.

[0044] The preprocessing unit 112 is used to perform grayscale conversion and smoothing filtering on the original eye image to generate a preprocessed eye image as the user's eye image.

[0045] Specifically, first, the original eye image is grayscale-converted, and the RGB three-channel color image is converted into a grayscale image. This process is based on the difference in human eye sensitivity to different colors, and the grayscale value of each pixel point can be calculated according to a specific weighted average formula (such as Y = 0.299R + 0.587G + 0.114B), where R, G, and B respectively represent the intensity values of the red, green, and blue color components, and Y is the final grayscale value. The grayscale conversion can not only reduce the data volume (from the original 3 channels to 1 channel), but also make important features such as edges and contours in the image more prominent, facilitating subsequent processing.

[0046] After the grayscale conversion is completed, the grayscale image is further subjected to a smoothing filtering operation to remove noise interference in the image. The Gaussian filtering algorithm can be used, which uses the convolution kernel generated by the Gaussian function to perform weighted average processing on the image pixels. The Gaussian function has good locality and decay, making the weight of the center of the convolution kernel larger and the surrounding weights gradually decreasing, so that while smoothing the image, the edge information is better retained. By reasonably selecting the kernel size (such as 3×3, 5×5, etc.) and standard deviation parameters of the Gaussian filter, common noises such as salt-and-pepper noise and Gaussian noise in the image can be effectively eliminated, making the image clearer and smoother, providing a better image basis for subsequent pupil center positioning and other operations. The image after grayscale conversion and smoothing filtering is the preprocessed eye image, which is the key input data for subsequent processing.

[0047] The pupil center positioning unit 113 is used to perform pupil center positioning on the preprocessed eye image based on the iris edge detection algorithm to generate the pupil center coordinates.

[0048] Specifically, the Canny edge detection operator can be used to extract the edges of the preprocessed eye image. By calculating the gradient magnitude and direction of the image, the Canny operator finds the regions in the image where the gray level changes drastically, i.e., the edge points. By setting appropriate low and high thresholds, the edge contours between the iris and the sclera and other edge features such as the eyelids can be effectively extracted.

[0049] After obtaining the edge image, a circle detection algorithm based on the Hough transform can be further used to determine the edge circle of the iris. The Hough transform can convert the circles in the image space to the cumulative peak points in the parameter space. By searching for the point with the maximum cumulative value in the parameter space, the center coordinates and radius of the iris edge circle can be obtained. Since the pupil is usually located in the central region of the iris and its color is significantly different from that of the iris, on the basis of determining the iris edge circle, the contour of the pupil is further searched inside the iris. By performing threshold segmentation on the internal region of the iris, the pupil region (usually a black or dark region) is distinguished from other parts of the iris, and then the centroid calculation method is used to determine the center coordinates of the pupil. Specifically, the pixel coordinates of the segmented pupil region are statistically analyzed, and the average values of its horizontal and vertical coordinates are calculated to obtain the precise coordinates (x_pupil, y_pupil) of the pupil center, which represent the position of the pupil in the image coordinate system.

[0050] The line-of-sight direction generation unit 114 is used to calculate the line-of-sight direction vector according to the pupil center coordinates and a preset eyeball geometric model, and generate a three-dimensional line-of-sight direction according to the said line-of-sight direction vector.

[0051] Specifically, a simplified eyeball geometric model can be established. This model approximates the eyeball as a sphere, assuming that the center of the eyeball is located at a fixed position on the user's head, denoted as point O(x_eye, y_eye, z_eye). The coordinates of this point can be preset through a three-dimensional scan of the user's head or a statistical model based on average human data. At the same time, the average radius of the eyeball is set as r_eye, and this radius value can be preset according to human anatomy data.

[0052] After obtaining the pupil center coordinates (x_pupil, y_pupil), they are converted from the two-dimensional image coordinate system to the three-dimensional eyeball coordinate system. This conversion process is based on the camera's internal parameter matrix (including focal length, principal point coordinates, distortion coefficients, etc.) and external parameter matrix (describing the position and orientation of the camera relative to the eyeball coordinate system). By inverse projection transformation, the two-dimensional image point is mapped to the light ray in the three-dimensional space, and this light ray intersects the eyeball surface at a point, which is the corresponding point P of the pupil center in the three-dimensional space. Then, calculate the vector OP between point P and the eyeball center point O, which is the line-of-sight direction vector. The direction of this vector is the user's line-of-sight direction, which can be represented by angles (such as pitch angle, yaw angle) or unit vector form in the three-dimensional space. By normalizing the line-of-sight direction vector, the unit line-of-sight direction vector (lx, ly, lz) is obtained, and this vector represents the direction of the user's line of sight in the three-dimensional space.

[0053] The fixation point coordinate generation unit 115 is used to map the line-of-sight direction to the screen coordinate system and generate fixation point coordinates.

[0054] Specifically, in practical applications, a mapping relationship between the line-of-sight direction and the screen coordinate system can be established. First, obtain the physical parameters of the screen, including the size of the screen (width W and height H), resolution (number of horizontal pixels Rw and number of vertical pixels Rh), and the position of the screen center point in the three-dimensional space (x_screen_center, y_screen_center, z_screen_center). At the same time, determine the plane equation of the screen. Usually, assuming that the relative position and orientation between the screen and the user's eyes are known, the unit vector (nx, ny, nz) of the screen normal direction and the conversion matrix between the screen coordinate system and the three-dimensional space coordinate system can be determined in advance through the calibration process.

[0055] After obtaining the above parameters, the line-of-sight direction vector (lx, ly, lz) is converted from the three-dimensional space coordinate system to the screen coordinate system. This conversion process projects the line-of-sight direction vector onto the screen plane and performs coordinate transformation according to the physical parameters and position information of the screen. Specifically, calculate the projection point coordinates (u, v) of the line-of-sight direction vector on the screen plane, where u represents the coordinate in the horizontal direction and v represents the coordinate in the vertical direction. Then, normalize the (u, v) coordinates to the pixel coordinate range according to the resolution of the screen to obtain the pixel coordinates (x_gaze, y_gaze) of the fixation point, that is, the position of the user's fixation point on the screen.

[0056] In an alternative embodiment, the head pose recognition module 12 includes:

[0057] The facial image acquisition unit 121 is used to acquire the facial area image of the user to obtain the original facial image.

[0058] Specifically, in practical applications, a high-resolution color camera with a resolution of no less than 1920×1080 pixels can be used. Its frame rate is stable at more than 30 frames per second to ensure that subtle facial expressions and movement changes of the user can be captured. The camera can be fixed at the central position of the upper edge of the display directly in front of the user, and the horizontal distance from the user's face is maintained between 50 and 60 centimeters, with the vertical angle slightly tilted downward by 15 to 30 degrees to capture the complete image of the user's face to the greatest extent while reducing the possibility of occlusion. During the acquisition process, the camera converts the optical signal of the user's face into a digital image signal through its sensor to generate an original facial image, and the image data contains information such as the contour, facial features, and expressions of the user's face.

[0059] The face key point detection unit 122 is used to process the original facial image based on the face key point detection algorithm to generate facial key point data containing multiple feature points.

[0060] Specifically, in the implementation process, a face key point detection algorithm based on deep learning can be used. This algorithm is pre-trained on a large-scale labeled face dataset and can identify 68 key feature points on the face, including the starting point, ending point, and highest point of the eyebrows, the inner and outer corners of the eyes, the pupil center, the wings and tip of the nose, the contour points of the mouth, and the key points of the facial contour, etc. Specifically, the original facial image is input into a trained convolutional neural network (CNN) model. The model extracts and transforms features through multiple layers of neural networks and outputs a key point heat map corresponding to the input image. In the heat map, the position of each key point is determined by the peak point in its corresponding heat map channel. By finding the coordinates of the peak point, the facial key point data of the user can be obtained, and this data is stored in a two-dimensional array in coordinate form.

[0061] The 3D pose calculation unit 123 is used to calculate the head rotation angle based on the facial key point data and a preset 3D face model through the PnP algorithm; generate head pose parameters according to the head rotation angle as the head pose recognition result.

[0062] Specifically, a preset 3D face model is established, which contains 3D coordinate points corresponding to the detected 2D facial key points. The 3D coordinate points define the shape and structure of a standard face. In actual calculations, the detected 2D facial key points are matched with the corresponding points in the 3D face model, and then the PnP (Perspective-n-Point) algorithm is used to solve for the rotation matrix and translation vector of the head. The PnP algorithm is based on the perspective projection model and estimates the pose of the camera relative to the 3D model by minimizing the reprojection error. Specifically, through an iterative optimization algorithm (such as the Levenberg-Marquardt algorithm), the rotation matrix R and translation vector t that minimize the error between the position where the 3D model points are projected onto the image plane after rotation and translation and the position of the detected 2D key points are solved. The rotation matrix R can be converted into a rotation vector through the Rodriguez formula, and then the rotation angles of the head are obtained, including the pitch angle, yaw angle, and roll angle. Head pose parameters are generated based on the rotation angles, and these parameters represent the pose direction of the head in three-dimensional space in vector form.

[0063] The pose correction unit 124 is used to perform spatial coordinate transformation processing on the line-of-sight direction according to the head pose parameters to generate a corrected line-of-sight direction; map the corrected line-of-sight direction to the screen coordinate system to generate corrected fixation point coordinates.

[0064] Specifically, in actual operation, the rotation matrix R in the head pose parameters can be obtained. This matrix represents the rotation of the head relative to the standard pose. Then, the original line-of-sight direction vector (lx, ly, lz) generated by the line-of-sight capture module is multiplied by the rotation matrix R to perform spatial coordinate transformation, obtaining a corrected line-of-sight direction vector (lx', ly', lz'). This process is actually a rotation correction of the original line-of-sight direction, eliminating the influence of head pose changes on the line-of-sight direction, so that the corrected line-of-sight direction is more in line with the user's actual visual attention direction. For example, when the user's head tilts to the left, the original line-of-sight direction may shift to the right. After pose correction, the corrected line-of-sight direction will accurately point to the area that the user is really concerned about. Next, the corrected line-of-sight direction vector is mapped to the screen coordinate system. This mapping process takes into account parameters such as the position, size, and resolution of the screen, and converts the three-dimensional line-of-sight direction vector into two-dimensional fixation point coordinates (x_gaze_corrected, y_gaze_corrected) on the screen through projective transformation. After correction, these coordinates can more accurately reflect the user's true fixation position on the screen, providing a reliable data basis for the subsequent emotion analysis and interaction control modules, ensuring that the system can perform precise emotional communication and interaction responses according to the user's true intentions.

[0065] In an alternative embodiment, the emotion analysis module 13 includes:

[0066] The line-of-sight behavior feature extraction unit 131 is configured to extract the line-of-sight movement frequency of the corrected line-of-sight direction and the dwell duration of the coordinates of the corrected fixation point, and generate line-of-sight behavior features based on the line-of-sight movement frequency and the dwell duration.

[0067] Specifically, in a specific implementation, the corrected line-of-sight direction is sampled at a fixed time interval (such as 10 times per second), and the line-of-sight movement frequency is obtained by calculating the change rate of the angle between the line-of-sight direction vectors of adjacent sampling points. For example, if the total change in the angle of the line-of-sight direction vectors within 1 second of continuous sampling is 180 degrees, the line-of-sight movement frequency is 180 degrees per second. At the same time, the dwell duration analysis is performed on the sequence of coordinates of the corrected fixation point, and the dwell duration is determined by detecting the stable residence time of the fixation point coordinates within a specific area of the screen. Specifically, a pixel threshold (such as 20 pixels) is set, and when the fixation point coordinates are always within the threshold range during continuous sampling, the cumulative dwell duration is calculated. Based on the line-of-sight movement frequency and the dwell duration, line-of-sight behavior features are generated and combined into a feature vector, such as [f_gaze_move, t_gaze_dwell], where f_gaze_move represents the line-of-sight movement frequency and t_gaze_dwell represents the dwell duration, providing feature inputs from the perspective of line-of-sight behavior for subsequent sentiment analysis.

[0068] The expression feature extraction unit 132 is configured to perform convolutional neural network processing on the user's facial image, extract the curvature of the mouth corners, the shape of the eyebrows, and the degree of eyelid opening and closing, and generate expression features based on the curvature of the mouth corners, the shape of the eyebrows, and the degree of eyelid opening and closing.

[0069] Specifically, in actual operation, the user's facial image is input into a pre-trained convolutional neural network (CNN) model. This model is trained based on a large-scale facial expression dataset and can automatically learn and extract high-level semantic features of facial expressions. After gradually extracting local features of the face in the convolutional layer and pooling layer of the network, the local features are integrated into a global facial expression feature representation in the fully connected layer. Further, through specific network branches or post-processing operations, key facial expression parameters such as the curvature of the mouth corners, the shape of the eyebrows, and the degree of eyelid opening and closing are accurately extracted from the global facial expression features. For example, the curvature of the mouth corners is obtained by calculating the curvature change of the pixel points in the mouth corner area, the shape of the eyebrows is classified based on the shape features of the eyebrow contour points (such as raised, lowered, relaxed, etc.), and the degree of eyelid opening and closing is determined by measuring the change in the pixel distance between the upper and lower eyelids in the eye area. The above parameters are combined into a facial expression feature vector, such as [curvature_mouth, shape_eyebrow, aperture_eyelid], where curvature_mouth is the curvature of the mouth corners, shape_eyebrow is the shape of the eyebrows, and aperture_eyelid is the degree of eyelid opening and closing. In this way, feature data from the perspective of facial expressions is provided for subsequent sentiment analysis, which helps to accurately capture the emotional changes of the user.

[0070] The head movement feature extraction unit 133 is used to calculate the head movement amplitude based on the head pose parameters and generate head movement features based on the head movement amplitude.

[0071] Specifically, during implementation, the head movement amplitude in a continuous time series can be calculated based on the head pose parameters output by the head pose recognition module. Among them, the head pose parameters include the rotation angles of the head (pitch angle, yaw angle, roll angle). Specifically, by summing or averaging the differences in the head rotation angles at adjacent time points, the head movement amplitude values in each direction are obtained. For example, within a time window (such as 1 second), the total change in the pitch angle is calculated to obtain the head up and down movement amplitude; similarly, the changes in the yaw angle and roll angle are calculated to obtain the head left and right and rotational movement amplitudes. The movement amplitude values are combined into a head movement feature vector, such as [amplitude_pitch, amplitude_yaw, amplitude_roll], where amplitude_pitch represents the head up and down movement amplitude, amplitude_yaw represents the head left and right movement amplitude, and amplitude_roll represents the head rotational movement amplitude, providing feature input from the perspective of head movement for subsequent sentiment analysis.

[0072] The feature fusion unit 134 is used to fuse the gaze behavior features, facial expression features, and head movement features through an attention mechanism to generate multi-modal fusion features.

[0073] Specifically, in specific implementation, an attention mechanism can be adopted to fuse the gaze behavior features, facial expression features, and head movement features. First, the above three types of feature vectors are respectively input into independent fully connected layers to map them into a space of the same dimension, ensuring their comparability and fusibility. Then, through a shared attention weight calculation module, which dynamically calculates the attention weight of each feature vector based on the correlation between the feature vectors. The calculation of the attention weight can be implemented by a multi-layer perceptron (MLP) or the self-attention mechanism in a Transformer. Its core is to enable the model to automatically learn which features are more discriminative in the current sentiment analysis task, and thus assign higher weights. For example, when the user's emotion is mainly reflected by facial expression features, the attention weight of the facial expression features will be relatively high; while when the user frequently expresses emotions through head movements, the weight of the head movement features will be increased. Finally, the weighted three types of feature vectors are concatenated or summed to generate a multi-modal fusion feature. This feature vector synthesizes multi-modal information such as gaze, facial expression, and head movement, and can more comprehensively and accurately reflect the user's emotional state.

[0074] The emotion classification unit 135 is used to input the multi-modal fusion feature into a Softmax classifier to generate the probability distribution of the user's emotional state as the user's emotion label.

[0075] Specifically, in practical applications, the multi-modal fusion feature is input into a Softmax classifier. This classifier is pre-trained on a large-scale labeled sentiment dataset and can map the input features to predefined emotion categories, such as basic emotions like joy, anger, sadness, surprise, fear, disgust, etc., as well as compound emotion states of different degrees. The Softmax classifier generates the probability distribution of the user's emotional state by calculating the probability values of each emotion category under the input features. Specifically, the Softmax function maps the multi-modal fusion feature vector to a probability vector, where each element represents the probability value of the corresponding emotion category, and the sum of the probability values is 1. For example, if the probability distribution is [0.7, 0.2, 0.1], it means that the probability of the user's emotion being "joy" is 70%, the probability of "anger" is 20%, and the probability of "sadness" is 10%. According to the maximum probability value in the probability distribution, the final user emotion label, such as "joy", is determined and used as the output result of the sentiment analysis module.

[0076] In an optional embodiment, the digital human generation module 14 includes:

[0077] The line-of-sight matching unit 141 is configured to drive the digital human's eyeballs to turn to the corresponding screen position according to the corrected line-of-sight direction and the corrected fixation point coordinates, and generate the digital human's line-of-sight direction.

[0078] Specifically, in specific implementation, first establish a digital human eyeball model, which includes the geometric shape of the eyeball, texture mapping, and parameters related to eyeball movement, such as the rotation center and rotation axis of the eyeball. Then, obtain the corrected line-of-sight direction vector (lx', ly', lz') and the corrected fixation point coordinates (x_gaze_corrected, y_gaze_corrected), and convert the above data into control parameters for the digital human's eyeball movement. Through inverse kinematics algorithms or direct rotation transformations, calculate the angles and directions that the digital human's eyeballs need to rotate, so that the digital human's gaze can accurately point to the screen position where the user is looking. For example, if the corrected line-of-sight direction vector of the user points to the right side of the screen, the digital human eyeball model will rotate a certain angle to the right accordingly, so that the digital human's gaze is consistent with the user's visual focus, creating a realistic eye contact effect and enhancing the naturalness and authenticity of emotional communication.

[0079] The expression generation unit 142 is configured to adjust the parameters of the digital human's facial muscle model based on the user's emotion label and generate the digital human's expression parameters.

[0080] Specifically, in actual operation, a digital human facial muscle system based on the Blendshape model can be constructed. This system defines multiple basic expression units, such as raised eyebrows, lowered eyebrows, raised corners of the mouth, lowered corners of the mouth, widened eyes, narrowed eyes, etc. Each expression unit corresponds to specific facial muscle movement parameters. When receiving the user's emotion label, according to the emotion type and intensity, retrieve the corresponding Blendshape weight combination from the preset expression parameter library. For example, for the emotion label of "joy - moderate", it may correspond to Blendshape weight settings such as slightly raised eyebrows, greatly raised corners of the mouth, and slightly narrowed eyes. By applying the weights to the digital human's facial muscle model, drive the morphological changes of each part of the face to generate natural and smooth digital human expressions, enabling the digital human to intuitively respond to the user's emotional state in the form of facial expressions, and further enhancing the interactivity and infectivity of emotional communication.

[0081] The action generation unit 143 is configured to select a preset limb action template according to the user's emotion label and the interaction scenario, and generate a digital human action instruction.

[0082] Specifically, during implementation, a rich library of body movement templates can be established, covering various common emotional expression movements and interaction movements, such as waving to say hello, clapping to show approval, putting hands on the hips to show anger, clapping hands to show excitement, etc. Each movement template defines in detail the starting posture, intermediate key-frame postures, ending posture, as well as parameters such as movement duration and speed change. When receiving the user's emotion label and interaction scenario information, the most suitable movement template for the current situation is selected from the movement template library through a matching algorithm. For example, in a video conference interaction scenario, if the user's emotion label is "excited", the movement template of "raising both hands and clapping" may be selected; if the user's emotion label is "frustrated", the movement template of "lowering the head and putting hands on the hips" may be selected. Further, according to the intensity of the user's emotion and the specific requirements of the interaction scenario, the selected movement template is fine-tuned, such as adjusting parameters such as the amplitude and speed of the movement, to generate the final digital human movement instruction, making the digital human's body movements more personalized and vivid, capable of accurately conveying emotional information, and enhancing the interest and naturalness of human-computer interaction.

[0083] In an alternative embodiment, the interaction control module 15 includes:

[0084] An emotion selection unit 151, configured to select matching voice intonation parameters from the policy library based on the user's emotion label and generate an emotional voice template.

[0085] Specifically, during specific implementation, a rich emotional voice policy library can be pre-constructed, which covers voice intonation parameters corresponding to various emotion types (such as joy, anger, sadness, etc.) and their different intensity levels (such as mild, moderate, severe). The parameters include dimensions such as speech rate, intonation, volume, and timbre. For example, the emotion of joy usually corresponds to a faster speech rate, a higher intonation, and a brighter timbre, while the emotion of sadness corresponds to a slower speech rate, a lower intonation, and a more mellow timbre. When receiving the user's emotion label output by the emotion analysis module, the most suitable combination of voice intonation parameters for the current emotion is matched in the policy library through a search algorithm to generate an emotional voice template. For example, if the user's emotion label is "joy - moderate", the corresponding parameters such as a speech rate of 180 words per minute, an intonation rise of 20%, a moderate volume, and a bright timbre are extracted from the policy library to construct an emotional voice template, providing an emotional voice basis for subsequent speech synthesis and interaction control, enabling the digital human's voice output to echo the user's emotion and enhancing the appeal of emotional communication.

[0086] An interaction content priority sorting unit 152, configured to determine the priority of the current interaction content based on the corrected fixation point coordinates and historical interaction records and generate a content sorting list.

[0087] Specifically, in actual operation, the interactive content on the screen is first divided into multiple regions or elements, each of which has a unique identifier and a corresponding initial value of priority weight. After obtaining the corrected fixation point coordinates, the region or element of the interactive content that the user is currently fixating on is determined through coordinate mapping. At the same time, referring to the historical interaction records, including information such as the interaction frequency, interaction duration, and time interval of the last interaction of the user with each piece of interactive content over a period of time in the past, the priority weights of each piece of interactive content are dynamically adjusted. For example, if the user fixates on a specific interactive element for a long time and frequently, and has less interaction recently, the priority weight of this element is appropriately increased; conversely, if the user only briefly glances at a certain region and has interacted many times recently, its priority weight is decreased. Through a weighted sorting algorithm (such as descending order based on priority weight), a content sorting list is generated, so that during the interaction process, the interactive content that the user is most likely to focus on is presented and responded to first, improving the interaction efficiency and user experience.

[0088] The dynamic intervention judgment unit 153 is used to trigger a voice guidance instruction and generate a dynamic intervention strategy when it detects that the residence duration of the corrected fixation point coordinates exceeds a preset threshold.

[0089] Specifically, during the implementation process, a fixation residence duration threshold (such as 3 seconds) is set, and the change of the user's fixation point coordinates is continuously tracked through a timer and a coordinate monitoring module. Whenever the residence time of the fixation point coordinates in a specific region reaches or exceeds this threshold, it is judged that the user may be confused, hesitant or in need of further information support for the content in this region. At this time, a voice guidance instruction is immediately triggered. This instruction will be passed to the subsequent interactive strategy generation unit, and at the same time, according to the current interaction scenario and the user's emotion label, a suitable intervention method is selected from the preset dynamic intervention strategy library, such as providing a detailed description of the content in this region, guiding the user to perform relevant operations, asking the user if they need help, etc., to generate a specific dynamic intervention strategy, so as to actively guide the user to complete the interaction process, avoid the interruption or low efficiency of the interaction caused by the user's confusion or helplessness, and improve the intelligence and user-friendliness of the system.

[0090] The interactive strategy generation unit 154 is used to integrate and process the emotional voice template, the content sorting list, and the dynamic intervention strategy to generate an interactive control strategy and the corresponding digital human voice text.

[0091] Specifically, in the specific implementation, first, based on the content sorted list, determine the presentation order and key points of the interaction content to ensure that high-priority interaction content is responded to first. Then, integrate the emotional voice templates into the interaction content, and according to different interaction scenarios and emotion tags, match appropriate voice expression methods for each interaction content segment to generate emotional voice texts. For example, for the interaction content with the highest priority, combine the emotional voice template of the joy emotion to generate enthusiastic and lively voice guidance texts; for the inquiry and help content in the dynamic intervention strategy, match gentle and soothing voice texts according to the user's current emotion (such as anxiety). Integrate the above elements into a coherent and fluent interaction script through natural language processing technology to form a complete interaction control strategy, clarify the behavior and voice output of the digital human during the interaction process, make the interaction process both in line with the user's emotion and efficient and orderly, and comprehensively improve the user's emotional experience and interaction satisfaction.

[0092] The strategy execution unit 155 is used to perform interaction control on the digital human according to the interaction control strategy.

[0093] Specifically, in the actual operation, parse the instructions in the interaction control strategy, including the adjustment of the digital human's line of sight direction, expression changes, execution of limb movements, and voice text output, etc. The above instructions can be converted into specific control signals and sent to the digital human generation module and the multimodal output module. For example, control the digital human generation module to adjust the digital human's eyesight to make it look at the area concerned by the user, drive the facial muscle model to generate an expression matching the interaction strategy, and execute the preset limb movement template; at the same time, transfer the generated digital human voice text to the speech synthesis system, convert it into a voice audio signal through speech synthesis technology, and play it to the user through the speaker. During the entire interaction process, the strategy execution unit monitors the user's feedback and changes in the interaction state in real time, such as the user's new line of sight direction, expression changes, etc., and dynamically adjusts the control signal to ensure the coherence and adaptability of the interaction process, so that the digital human can flexibly and intelligently respond to the user's interaction needs, create a natural, fluent and emotional human-computer interaction environment, and promote the smooth progress of the interaction process until the interaction task is completed.

[0094] In an optional embodiment, the multimodal output module 16 includes:

[0095] The digital human image generation unit 161 is used to drive the three-dimensional digital human model through the rendering engine according to the digital human's line of sight direction, digital human expression parameters, and digital human movement instructions to generate a dynamic digital human image.

[0096] Specifically, in specific implementations, physically based rendering (PBR) technology can be adopted to accurately simulate the material characteristics on the surface of the digital human model, such as the gloss of the skin, the texture of the clothing, etc., making the digital human image more realistic and credible. First, input the digital human's line-of-sight direction parameters into the rendering engine. By adjusting the bones and texture maps of the digital human model's eyes, drive the digital human's gaze to turn to the specified direction, so that the gaze is coordinated with the user's line-of-sight direction. For example, if the digital human's line-of-sight direction parameters indicate a gaze towards the upper right, the rendering engine will correspondingly adjust the digital human's eyeballs and eyelids to present a natural upper-right gaze state. Next, according to the digital human's expression parameters, use the Blendshape model to finely control the digital human's facial muscles, such as adjusting the upward arc of the corners of the mouth, the morphological changes of the eyebrows, etc., to generate a rich variety of expressions. For example, when the expression parameters indicate "smile", the rendering engine will drive the facial muscle model to slightly raise the corners of the mouth and slightly squint the eyes, presenting a kind and natural smiling expression. At the same time, according to the digital human's action instructions, drive the digital human's limb bones and muscle system through the animation control system to perform corresponding actions. For example, if the action instruction is "wave to say hello", the rendering engine will, according to the preset animation key frames and interpolation algorithms, make the digital human's arm perform a natural and smooth waving action. During the entire rendering process, the rendering engine calculates the effects of lighting, shadows, and reflections in real time, making the visual effect of the digital human image realistic and natural in different environments, providing an immersive visual experience for users.

[0097] The speech synthesis unit 162 is used to perform emotional speech synthesis processing on the digital human speech text to generate speech waveform data.

[0098] Specifically, in actual operation, advanced deep learning speech synthesis technologies can be adopted, such as models like Tacotron 2 or FastSpeech 2 based on the Transformer architecture. The above models can learn rich emotional expressions and natural intonation changes of speech. First, perform text preprocessing on the input digital human speech text, including operations such as word segmentation, part-of-speech tagging, and semantic analysis, to understand the emotional tendency and semantic structure of the text. For example, for text containing exclamation words or positive vocabulary, judge that it may correspond to a joyous emotion; for text containing interrogative words or uncertain vocabulary, judge that it may correspond to a confused emotion, etc. Then, combine the speech intonation parameters in the emotional speech template output by the emotion selection unit, such as speech rate, intonation, volume, etc., to adjust the parameters of the speech synthesis model. For example, if the emotional speech template indicates "sad - mild", then set the speech rate to be slower, the intonation to be low, and the volume to be moderate, etc. Next, input the processed text and parameters into the speech synthesis model. The model, through an encoder-decoder structure, converts the text features into a sequence of speech features, and then generates the corresponding speech waveform data through a vocoder. The generated speech waveform data has an emotional expression, can echo the visual performance of the digital human, and enhance the effect of emotional communication. For example, when the digital human smiles, the speech generated by the speech synthesis unit also has a pleasant intonation; when the digital human has a serious expression, the speech appears more steady and solemn, enabling the user to also feel the emotional information conveyed by the digital human auditorily.

[0099] The visual-auditory synchronization unit 163 is used to perform timestamp alignment processing on the dynamic digital human image and the speech waveform data to generate visual-auditory interaction content with visual-auditory synchronization.

[0100] Specifically, a timestamp-based synchronization mechanism can be adopted in the specific implementation. First, during the process of generating a dynamic digital human image and voice waveform data by the digital human image generation unit and the speech synthesis unit respectively, accurate timestamps are added to each key frame and speech segment. For example, during the digital human image generation process, every time a frame of image is rendered (such as one frame every 33 milliseconds, corresponding to 30 frames per second), a corresponding timestamp is added; during the speech synthesis process, every time a segment of speech waveform is generated (such as taking 10 milliseconds as a segment), a corresponding timestamp is also added. Then, the timing of the two is adjusted through a synchronization algorithm, so that during playback, the lip movements of the digital human perfectly match the speech content, and the limb movements and expression changes are also coordinated with the rhythm of the speech. For example, when the speech synthesis unit generates a speech segment of "Hello" with a start timestamp of 0 seconds and a duration of 1 second, the visual-auditory synchronization unit will ensure that the digital human image performs the corresponding lip movements, expression changes and limb movements (such as waving) during the time period from 0 seconds to 1 second. Through precise timestamp alignment processing, when the generated visual-auditory interaction content is played, users can feel that the words "spoken" by the digital human are completely synchronized with its expressions and actions, as if a real virtual character is communicating with them, greatly enhancing the realism and naturalness of the interaction, enabling users to be more immersed in the human-computer interaction process and enjoy the fun of emotional communication.

[0101] The above intelligent gaze interaction system for emotional communication generates the gaze direction by collecting eye images for iris feature localization and eyeball geometric model calculation, eliminates gaze deviation by combining facial key point detection and 3D model head pose correction, generates user emotion labels by fusing multimodal data analysis of gaze behavior, facial expressions and head movements, drives the digital human to generate matching gaze directions, expression parameters and limb movements based on the dynamic emotional state, triggers an adaptive interaction strategy by combining the fixation point coordinates and emotion labels, and finally realizes high-precision gaze tracking, real-time response of emotional digital humans and natural and realistic interaction feedback through audio-visual multimodal synchronization output technology, solves the problem of rigid interaction caused by environmental interference, head movement and single emotion analysis in traditional systems, and significantly improves the realism and emotional communication depth of human-computer interaction.

[0102] Based on the same inventive concept, the embodiments of the present application also provide a method applied to the above-mentioned intelligent gaze interaction system for emotional communication. The implementation solutions provided by this method to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the intelligent gaze interaction method for emotional communication provided below can refer to the limitations on an intelligent gaze interaction system for emotional communication in the above text, and will not be repeated here.

[0103] In an exemplary embodiment, such as Figure 2As shown, an intelligent gaze interaction method for emotional communication is provided, which is applied to the above-mentioned intelligent gaze interaction system for emotional communication. The method includes the following steps:

[0104] S1: The gaze capture module performs gaze capture processing based on the user's eye image to generate the gaze direction and the coordinates of the fixation point.

[0105] S2: The head pose recognition module performs head pose recognition processing based on the user's face image to obtain the head pose recognition result; the gaze direction and the coordinates of the fixation point are corrected according to the head pose recognition result to generate the corrected gaze direction and the corrected coordinates of the fixation point.

[0106] S3: The emotion analysis module performs multi-modal emotion analysis processing on the corrected gaze direction, the corrected coordinates of the fixation point, the head pose recognition result, and the user's facial expression to generate the user emotion label.

[0107] S4: The digital human generation module generates the digital human gaze direction, the digital human expression parameters, and the digital human action instructions according to the corrected gaze direction, the corrected coordinates of the fixation point, and the user emotion label.

[0108] S5: The interaction control module generates an interaction control strategy and the corresponding digital human speech text based on the user emotion label and the corrected coordinates of the fixation point, and performs interaction control on the digital human according to the interaction control strategy; the control objects of the interaction control include: the digital human gaze direction, the digital human expression parameters, and the digital human action instructions.

[0109] S6: The multi-modal output module performs multi-modal output on the digital human gaze direction, the digital human expression parameters, the digital human action instructions, and the digital human speech text to generate the visual and auditory interaction content.

[0110] It should be understood that although each step in the flowcharts involved in the above-mentioned embodiments is shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0111] An embodiment of the present application further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the foregoing method embodiments are implemented.

[0112] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.

[0113] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are only illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0114] The above embodiments only represent several implementation manners of the embodiments of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the application embodiments. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the embodiments of the present application.

Claims

1. An intelligent gaze interaction system for emotional communication, characterized in that, The system includes: a line-of-sight capture module, a head pose recognition module, an emotion analysis module, a digital human generation module, an interaction control module, and a multimodal output module; The line-of-sight capture module is used to perform line-of-sight capture processing based on the user's eye image, and generate a line-of-sight direction and fixation point coordinates; The head pose recognition module is used to perform head pose recognition processing based on the user's face image to obtain a head pose recognition result; and correct the line-of-sight direction and fixation point coordinates according to the head pose recognition result to generate a corrected line-of-sight direction and corrected fixation point coordinates; The emotion analysis module is used to perform multimodal emotion analysis processing on the corrected line-of-sight direction, the corrected fixation point coordinates, the head pose recognition result, and the user's facial expression to generate a user emotion label; The digital human generation module is used to generate a digital human line-of-sight direction, digital human expression parameters, and digital human action instructions according to the corrected line-of-sight direction, the corrected fixation point coordinates, and the user emotion label; The interaction control module is used to generate an interaction control strategy and a corresponding digital human speech text based on the user emotion label and the corrected fixation point coordinates, and perform interaction control on the digital human according to the interaction control strategy; the control objects of the interaction control include: the digital human line-of-sight direction, the digital human expression parameters, and the digital human action instructions; The multimodal output module is used to perform multimodal output on the digital human line-of-sight direction, the digital human expression parameters, the digital human action instructions, and the digital human speech text to generate visual and auditory interaction content.

2. The system according to claim 1, wherein The line-of-sight capture module includes: An eye image acquisition unit for acquiring an eye region image of the user to obtain an original eye image; A preprocessing unit for performing grayscale conversion and smoothing filtering on the original eye image to generate a preprocessed eye image as the user's eye image; A pupil center positioning unit for performing pupil center positioning on the preprocessed eye image based on an iris edge detection algorithm to generate pupil center coordinates; A line-of-sight direction generation unit for calculating a line-of-sight direction vector according to the pupil center coordinates and a preset eyeball geometric model, and generating the three-dimensional line-of-sight direction according to the line-of-sight direction vector; A fixation point coordinate generation unit for mapping the line-of-sight direction to a screen coordinate system to generate the fixation point coordinates.

3. The system according to claim 2, characterized in that, The head pose recognition module includes: A face image acquisition unit for acquiring a face region image of the user to obtain an original face image; A face key point detection unit for processing the original face image based on a face key point detection algorithm to generate face key point data including multiple feature points; A three-dimensional pose calculation unit for calculating a head rotation angle based on the face key point data and a preset 3D face model through a PnP algorithm; and generating head pose parameters as the head pose recognition result according to the head rotation angle; The posture correction unit is used to perform spatial coordinate transformation processing on the line-of-sight direction according to the head posture parameters to generate the corrected line-of-sight direction; map the corrected line-of-sight direction to the screen coordinate system to generate the corrected fixation point coordinates.

4. The system according to claim 3, wherein The emotion analysis module includes: The line-of-sight behavior feature extraction unit is used to extract the line-of-sight movement frequency of the corrected line-of-sight direction and the dwell time of the corrected fixation point coordinates, and generate line-of-sight behavior features based on the line-of-sight movement frequency and the dwell time; The expression feature extraction unit is used to perform convolutional neural network processing on the user's facial image, extract the mouth corner arc, eyebrow shape, and eyelid opening degree, and generate expression features based on the mouth corner arc, the eyebrow shape, and the eyelid opening degree; The head movement feature extraction unit is used to calculate the head movement amplitude based on the head posture parameters and generate head movement features based on the head movement amplitude; The feature fusion unit is used to fuse the line-of-sight behavior features, the expression features, and the head movement features through an attention mechanism to generate multi-modal fusion features; The emotion classification unit is used to input the multi-modal fusion features into a Softmax classifier to generate the probability distribution of the user's emotion state as the user's emotion label.

5. The system according to claim 1, wherein The digital human generation module includes: The line-of-sight matching unit is used to drive the digital human's eyeballs to turn to the corresponding screen position according to the corrected line-of-sight direction and the corrected fixation point coordinates to generate the digital human's line-of-sight direction; The expression generation unit is used to adjust the parameters of the digital human's facial muscle model based on the user's emotion label to generate the digital human's expression parameters; The action generation unit is used to select a preset limb action template according to the user's emotion label and the interaction scenario to generate the digital human's action instruction.

6. The system according to claim 1, wherein The interaction control module includes: The emotion selection unit is used to select matching voice intonation parameters from the policy library based on the user's emotion label to generate an emotional voice template; The interaction content priority ranking unit is used to determine the current interaction content priority according to the corrected fixation point coordinates and the historical interaction records to generate a content ranking list; The dynamic intervention judgment unit is used to trigger a voice guidance instruction and generate a dynamic intervention strategy when it detects that the dwell time of the corrected fixation point coordinates exceeds a preset threshold; The interaction strategy generation unit is used to integrate the emotional voice template, the content ranking list, and the dynamic intervention strategy to generate the interaction control strategy and the corresponding digital human voice text; The strategy execution unit is used to perform interaction control on the digital human according to the interaction control strategy.

7. The system according to any one of claims 1 to 6, characterized in that, The multi-modal output module includes: The digital human image generation unit is used to drive a three-dimensional digital human model through a rendering engine according to the digital human's line-of-sight direction, the digital human's expression parameters, and the digital human's action instruction to generate a dynamic digital human image; The voice synthesis unit is used to perform emotional voice synthesis processing on the digital human voice text to generate voice waveform data; An audiovisual synchronization unit for timestamp alignment processing of the dynamic digital human image and the speech waveform data to generate the audiovisual interaction content with audiovisual synchronization.

8. An intelligent line-of-sight interaction method for emotional communication, applied to the system according to any one of claims 1 to 7, characterized in that, The method includes: S1: A line-of-sight capture module performs line-of-sight capture processing based on the user's eye image to generate a line-of-sight direction and fixation point coordinates; S2: A head pose recognition module performs head pose recognition processing based on the user's face image to obtain a head pose recognition result; the line-of-sight direction and fixation point coordinates are corrected according to the head pose recognition result to generate a corrected line-of-sight direction and corrected fixation point coordinates; S3: An emotion analysis module performs multi-modal emotion analysis processing on the corrected line-of-sight direction, the corrected fixation point coordinates, the head pose recognition result, and the user's facial expression to generate a user emotion label; S4: A digital human generation module generates a digital human line-of-sight direction, digital human expression parameters, and digital human action instructions according to the corrected line-of-sight direction, the corrected fixation point coordinates, and the user emotion label; S5: An interaction control module generates an interaction control strategy and a corresponding digital human speech text based on the user emotion label and the corrected fixation point coordinates, and performs interaction control on the digital human according to the interaction control strategy; the control objects of the interaction control include: the digital human line-of-sight direction, the digital human expression parameters, and the digital human action instructions; S6: A multi-modal output module performs multi-modal output on the digital human line-of-sight direction, the digital human expression parameters, the digital human action instructions, and the digital human speech text to generate audiovisual interaction content.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the method described in claim 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method described in claim 8 is implemented.

Citation Information

Cited By

  • Mutual gaze detection method based on space-time diagram neural network and attention mechanism

    CN122067165A

  • Fused image reprojection and optical distortion correction

    US12718485B1