Screen interaction method and system based on face recognition

By extracting the dynamic behavior characteristics of the user's face and the coordinates of the screen gaze area, combining the preset interaction mode library and the pre-trained facial feature fusion model, the presentation parameters of the dynamic interactive content are adjusted in real time, and the problems of lag in emotional judgment and misalignment with user needs in the existing technology are solved, achieving more efficient and personalized screen interaction.

CN120045076AActive Publication Date: 2025-05-27CHINA GUANGSHEN OPTOELECTRONICS (SHENZHEN) CO LTD

Patent Information

Application Number
CN202510512628.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-27
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

Existing screen interaction technology based on face recognition cannot effectively capture dynamic changes in facial muscles, resulting in lag in emotional judgment and being unable to accurately correlate user identity preferences and real-time psychological state, resulting in misalignment of interaction response and user needs.

Method used

By obtaining the user's face image sequence in the screen interaction scene, extracting facial dynamic behavior characteristics and screen gaze area coordinates, combining the preset interactive mode library to match the target interaction mode, generating dynamic interactive content corresponding to the user's gaze focus, and calling the pretrained facial feature fusion model to extract user authentication vectors and emotional state vectors, and adjusting the presentation parameters of dynamic content in real time.

Benefits of technology

It significantly improves the matching accuracy of screen interactive content and user needs, enhances the flexibility and adaptability of content presentation, accurately recognizes the user's true intentions, and improves interaction efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045076A_ABST
    Figure CN120045076A_ABST
Patent Text Reader

Abstract

The invention provides a screen interaction method and system based on face recognition, and the method comprises the steps: obtaining a face image sequence triggered by a target user in a screen interaction scene, carrying out the matching of a target interaction mode from a preset interaction mode library according to the face dynamic behavior characteristics, and carrying out the matching of a target interaction mode based on the coordinates of a screen fixation region and the target interaction mode. Generating dynamic interaction content corresponding to the sight focus of the user, calling a pre-trained facial feature fusion model, performing multi-level feature extraction on a facial region in the facial image sequence, generating a user identity verification vector and an emotional state vector, and performing real-time adjustment on presentation parameters of the dynamic interaction content based on the multi-level feature extraction. And outputting the adjusted screen interaction signal. According to the invention, the user experience can be optimized while the interaction efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular, to a screen interaction method and system based on face recognition. Background Art

[0002] The screen interaction technology based on face recognition aims to achieve intelligent response and personalized adaptation of screen content by analyzing user biometric features. In the prior art, static face recognition is usually used for user identity verification or preset interaction commands are triggered through an expression classification algorithm. For example, the identity library is matched based on the geometric features of the face contour and then a fixed interface configuration is loaded, or the screen display mode is switched according to basic expression tags. However, such methods have significant defects: static recognition cannot capture the dynamic changes of facial muscles, resulting in a lag in emotion judgment. Simple data processing is difficult to associate user identity preferences with real-time psychological states. The independent expression classification rules lack collaborative analysis with the line-of-sight focus and behavioral intentions, resulting in a mismatch between the interaction response and the real needs of the user. At the same time, the fixed content adjustment mechanism cannot adapt to the dynamic interaction requirements in different scenarios, resulting in insufficient personalization and low interaction efficiency. Summary of the Invention

[0003] The present invention provides a screen interaction method and system based on face recognition.

[0004] In a first aspect, an embodiment of the present invention provides a screen interaction method based on face recognition, the method including: obtaining a sequence of face images triggered by a target user in a screen interaction scenario, the sequence of face images including the facial dynamic behavior features of the user and the coordinates of the screen gaze area; matching a target interaction mode from a preset interaction mode library according to the facial dynamic behavior features, the target interaction mode including an interaction response rule associated with the user's expression intensity; generating dynamic interaction content corresponding to the user's line-of-sight focus based on the coordinates of the screen gaze area and the target interaction mode; calling a pre-trained facial feature fusion model to perform multi-level feature extraction on the facial area in the sequence of face images to generate a user identity verification vector and an emotion state vector; and adjusting the presentation parameters of the dynamic interaction content in real time according to the user identity verification vector and the emotion state vector, and outputting an adjusted screen interaction signal.

[0005] In a second aspect, an embodiment of the present invention provides a screen interaction system, including: a memory in which a computer program is stored; and a processor configured to load the computer program to implement the screen interaction method based on face recognition as described above.

[0006] The screen interaction method based on face recognition provided by the present invention obtains the facial dynamic behavior features and screen fixation area coordinates in the face image sequence triggered by the target user in the screen interaction scenario, combines with a preset interaction mode library to match the target interaction mode associated with the user's expression intensity, generates dynamic interaction content corresponding to the line of sight focus, and calls a pre-trained facial feature fusion model to extract the user authentication vector and emotional state vector. Based on the identity features and emotional features, the presentation parameters of the dynamic content are adjusted in real time, and finally a screen interaction signal adapted to the user state is output. In this way, the facial dynamic behavior features can capture the user's real-time expression changes and muscle movement trends, the screen fixation area coordinates can accurately locate the user's visual focus of attention, the authentication vector provides a benchmark for the user's historical preferences for personalized interaction, and the emotional state vector dynamically reflects the user's current mental state. Through the collaborative analysis and fusion of multi-dimensional data, the matching accuracy of the screen interaction content and the user's needs can be significantly improved; at the same time, dynamically adjusting the parameters of the dynamic interaction content based on real-time emotions and identity features can enhance the flexibility and adaptability of content presentation, effectively solving the problems of response latency and insufficient personalization in traditional interaction technologies; in addition, through the joint analysis of facial dynamic behavior and screen fixation area, the user's true intention can be accurately identified, avoiding ineffective interactions caused by accidental touches or distractions, thereby optimizing the user experience while improving the interaction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0008] Figure 1 is a flowchart of a screen interaction method based on face recognition provided by an embodiment of the present invention.

[0009] Figure 2 is a schematic diagram of the composition of a screen interaction system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0010] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0011] Please refer to Figure 1 ,Figure 1 The following is a flowchart of a screen interaction method based on face recognition provided by an embodiment of the present invention. This screen interaction method based on face recognition can be executed by a screen interaction system, and the method may include the following steps: Step S100: Obtain a sequence of face images triggered by a target user in a screen interaction scenario. The sequence of face images includes the user's facial dynamic behavior characteristics and the screen fixation area coordinates.

[0012] In this embodiment, the sequence of face images refers to a set of images containing the user's facial information arranged in chronological order. The facial dynamic behavior characteristics refer to the action changes of the user's face over a period of time, such as expression changes, muscle movements, etc. The screen fixation area coordinates refer to the position information of the screen area where the user's line of sight focuses during the screen interaction process. Obtaining the sequence of face images triggered by a target user in a screen interaction scenario can be specifically implemented in various ways. For example, an embedded camera can be used to capture the user's facial images.

[0013] Specifically, this step may include the following detailed steps: Step S110: Capture the user's raw facial image stream through an embedded camera and perform frame rate synchronization processing on the raw image stream.

[0014] An embedded camera is a camera installed on the screen device for real-time capturing of the user's facial images. The raw image stream is a series of image data continuously captured by the camera. Frame rate synchronization processing is to ensure that the time intervals of each image frame are uniform and meet the requirements of subsequent processing. For example, if the shooting frame rate of the camera is unstable, it may cause errors when analyzing the facial dynamic behavior characteristics later. Through frame rate synchronization processing, the frame rate of the raw image stream can be adjusted to a stable value. For example, the frame rate can be uniformly adjusted to 30 frames per second, so that when analyzing the user's facial movements later, the change process of the movements can be captured more accurately.

[0015] Step S120: Detect the face bounding box coordinates in each frame of the image and perform image cropping and size normalization processing based on the bounding box coordinates.

[0016] The coordinates of the face bounding box refer to the coordinate information of the rectangular box that can frame the face area. By determining these coordinates, the position of the face in the image can be accurately located. Image cropping is to intercept the face area from the original image according to the detected coordinates of the face bounding box, removing the irrelevant background information. The size normalization process is to adjust the cropped face image to a unified size for subsequent feature extraction and analysis. For example, the cropped face images are uniformly adjusted to a size of 224×224 pixels, which can ensure that different images have the same input size during feature extraction, improving the accuracy and consistency of feature extraction.

[0017] Specifically, detecting the coordinates of the face bounding box in each frame of the image may include the following steps: Step S121: Call the pre-trained face detection model to perform an initial prediction of the face area in the current frame image, generating a set of candidate bounding boxes.

[0018] The pre-trained face detection model is a model trained with a large amount of face image data and can identify the face area in the image. For example, a commonly used face detection model is MTCNN (Multi-task Cascaded Convolutional Networks), which can quickly and accurately detect the approximate position of the face in the image and generate multiple candidate bounding boxes. These candidate bounding boxes may contain face areas of different sizes and positions, providing a basis for subsequent screening.

[0019] Step S122: Calculate the confidence scores of each candidate bounding box, and screen and merge the candidate boxes with scores exceeding the second threshold.

[0020] The confidence score refers to the evaluation score of the model's possibility that each candidate bounding box contains a face. The second threshold is a preset score standard for screening candidate bounding boxes with higher confidence. For example, the second threshold is set to 0.8, and only the candidate bounding boxes with confidence scores exceeding 0.8 will be retained. The merging process is to merge the screened candidate bounding boxes, removing the overlapping parts to obtain a more accurate face bounding box. This can avoid the situation where multiple overlapping bounding boxes are all considered as faces, improving the accuracy of face positioning.

[0021] Step S123: Determine the center point coordinates and width-to-height ratio of the final face bounding box according to the coordinates of the merged candidate boxes.

[0022] The coordinates of the merged candidate bounding boxes contain the approximate location information of the face region. By calculating these coordinates, the center point coordinates and the width-to-height ratio of the final face bounding box can be determined. The center point coordinates can accurately represent the position of the face in the image, while the width-to-height ratio can reflect the shape characteristics of the face. For example, by calculating the average value of the upper left and lower right coordinates of the merged candidate bounding box, the center point coordinates can be obtained; by calculating the ratio of the width and height of the candidate bounding box, the width-to-height ratio can be obtained. This information is very important for subsequent image cropping and size normalization processing.

[0023] Step S124: Predict the position of the bounding box in the current frame based on the face movement trajectory of the historical frame. When the deviation between the predicted position and the actual detection position exceeds the third threshold, trigger the face tracking model to perform position correction.

[0024] The face movement trajectory of the historical frame refers to the position change information of the face in the previous image frames. By analyzing these trajectories, the possible position of the face in the current frame can be predicted. The third threshold is a pre-set deviation standard used to determine whether the difference between the predicted position and the actual detection position is too large. For example, if the third threshold is set to 10 pixels, and the deviation between the predicted position and the actual detection position exceeds 10 pixels, it means the prediction is inaccurate and the face tracking model needs to be triggered for position correction. The face tracking model can perform more accurate tracking and correction of the face position based on the information of the current frame and the historical frame to ensure the accurate position of the face bounding box.

[0025] Step S125: Perform a ratio conversion on the corrected bounding box coordinates and the screen resolution to generate standardized coordinate data adapted to the current screen size.

[0026] The screen resolution refers to the pixel size of the screen, and different screen devices may have different resolutions. Performing a ratio conversion on the corrected bounding box coordinates and the screen resolution can convert the bounding box coordinates into standardized coordinate data suitable for the current screen size. For example, if the screen resolution is 1920×1080 pixels and the corrected bounding box coordinates are (x1, y1, x2, y2), by dividing these coordinates by the width and height of the screen, standardized coordinate data can be obtained, so that the face region can be accurately located on screens with different resolutions.

[0027] Step S130: Perform illumination equalization processing on the normalized image to eliminate the influence of environmental light fluctuations on the image quality and obtain an equalized image.

[0028] Lighting equalization processing is an image processing technique used to adjust the brightness and contrast of an image, enabling the image to maintain good quality under different lighting conditions. Fluctuations in ambient light may cause over-bright or over-dark areas in the image, affecting the extraction and analysis of facial features. Through lighting equalization processing, these effects can be eliminated, making the brightness of the image more uniform. For example, the method of histogram equalization can be adopted to process the normalized image, adjusting the gray value distribution of the image to a more uniform range, thereby obtaining an equalized image. In this way, during the subsequent facial feature extraction process, the detailed features of the face can be more accurately recognized.

[0029] Step S140: Use a background segmentation algorithm to extract the foreground face region in the equalized image and remove invalid frames containing occlusions or blurred areas.

[0030] The background segmentation algorithm is an algorithm used to separate foreground objects (such as faces) in an image from the background. Through this algorithm, the foreground face region in the equalized image can be extracted, removing irrelevant background information. Occlusions or blurred areas may affect the accurate extraction of facial features, so it is necessary to remove invalid frames containing these areas. For example, a background segmentation algorithm based on deep learning, such as the U-Net network, can be used to process the equalized image, separating the face region from the background. Then, by analyzing the clarity and occlusion situation of the image, it is judged whether there are occlusions or blurred areas. If so, the frame image is marked as an invalid frame and removed. This can ensure that the face images used for subsequent analysis are clear and unobstructed, improving the accuracy of facial feature extraction.

[0031] Step S150: Arrange the processed valid image frames in chronological order to generate a sequence of face images, and add a timestamp and a screen touch event mark to each frame of the image.

[0032] The processed valid image frames refer to the face image frames that meet the requirements after processes such as cropping, normalization, lighting equalization, and background segmentation. Arranging these image frames in chronological order can generate a sequence of face images containing the dynamic information of the user's face. The timestamp is used to record the shooting time of each frame of the image. Through the timestamp, the chronological order and change process of the user's facial movements can be understood. The screen touch event mark is used to record whether the user performed a screen touch operation when the frame of the image was taken, as well as information such as the type and position of the operation. For example, when the user clicks a certain button on the screen, the corresponding screen touch event mark is added to the corresponding image frame. These mark information can provide important reference basis for subsequent interaction analysis.

[0033] Step S200: According to the facial dynamic behavior characteristics, match the target interaction mode from the preset interaction mode library, where the target interaction mode includes interaction response rules associated with the user's expression intensity.

[0034] In this embodiment, the facial dynamic behavior characteristics refer to the action changes of the user's face over a period of time, such as expression changes, muscle movements, etc. The preset interaction mode library is a database that stores multiple interaction modes in advance, and each interaction mode includes facial action trigger conditions and corresponding screen response strategies. The target interaction mode refers to the most suitable interaction mode matched from the interaction mode library according to the user's facial dynamic behavior characteristics. The interaction response rule refers to the response method that the screen should make when the user's expression intensity meets the preset conditions. Specifically, matching the target interaction mode from the preset interaction mode library according to the facial dynamic behavior characteristics may include the following steps: Step S210: Input the facial dynamic behavior characteristics into a pre-trained spatio-temporal convolutional network to extract facial muscle movement trajectory characteristics and eye opening and closing frequency characteristics.

[0035] The pre-trained spatio-temporal convolutional network is a neural network model trained with a large amount of data and can process sequence data containing time and space information. The facial muscle movement trajectory characteristics refer to the movement path and changes of facial muscles over a period of time, which can reflect the user's expression changes. The eye opening and closing frequency characteristics refer to the number of times the user's eyes open and close per unit time, which can reflect the user's attention to the screen content. For example, a 3D convolutional neural network (3D CNN) can be used as the pre-trained spatio-temporal convolutional network, and the facial dynamic behavior characteristics are input into this network for processing.

[0036] Specifically, inputting the facial dynamic behavior characteristics into the pre-trained spatio-temporal convolutional network to extract facial muscle movement trajectory characteristics and eye opening and closing frequency characteristics may include the following steps: Step S211: Perform facial key point detection on each frame image in the face image sequence to generate a key point distribution map including the coordinate information of the eyebrow area, cheek area, and mouth corner area.

[0037] Facial key point detection refers to accurately locating key points in a face image, such as the coordinates of parts like eyebrows, cheeks, and mouth corners. By detecting the positions of these key points, a key point distribution map containing this coordinate information can be generated. For example, a deep learning-based facial key point detection algorithm, such as the 68-point face key point detection model in the Dlib library, can be used to process each frame image in the face image sequence to obtain the coordinate information of the eyebrow area, cheek area, and mouth corner area in each image frame, and integrate this coordinate information into the key point distribution map. This can intuitively display the position changes of each part of the face.

[0038] Step S212: Based on the key-point displacement vectors between adjacent frames, determine the motion direction consistency parameters of each facial region within a preset time window. The motion direction consistency parameters are used to characterize the stability of the coordinated movement of muscle groups.

[0039] The key-point displacement vectors between adjacent frames refer to the position change vectors of the same key point in two adjacent image frames. By analyzing these displacement vectors, the motion direction consistency parameters of each facial region within a preset time window can be determined. The preset time window is a preset time period used to analyze the motion of facial regions. For example, if the preset time window is set as the time period of 10 image frames, calculate the displacement vectors of the key points of each facial region within this time period, and then count whether the directions of these displacement vectors are consistent. If the directions of most displacement vectors are the same, it indicates that the coordinated movement of the muscle group in this facial region is relatively stable and the motion direction consistency parameter is high; otherwise, it indicates that the coordinated movement of the muscle group is unstable and the motion direction consistency parameter is low. This parameter can help determine whether the user's facial expression changes are natural and stable.

[0040] Step S213: Perform optical flow field analysis on the key-point distribution map to generate a dynamic change curve reflecting the muscle contraction intensity, and extract the periodic characteristics of facial muscle movement according to the occurrence frequencies of the peaks and valleys in the dynamic change curve.

[0041] Optical flow field analysis is a technique for analyzing the movement of objects in an image. By performing optical flow field analysis on the key-point distribution map, the motion speed and direction information of facial muscles can be obtained. Based on this information, a dynamic change curve reflecting the muscle contraction intensity can be generated. The peaks and valleys in the dynamic change curve represent the maximum and minimum values of the muscle contraction intensity. By counting the occurrence frequencies of the peaks and valleys, the periodic characteristics of facial muscle movement can be extracted. For example, use the Lucas-Kanade optical flow algorithm to analyze the key-point distribution map, calculate the optical flow vectors of each key point, and then generate a dynamic change curve according to the magnitude and direction of the optical flow vectors. If the occurrence frequencies of the peaks and valleys in the dynamic change curve are high, it indicates that the facial muscle movement has strong periodicity; otherwise, it indicates weak periodicity. These periodic characteristics can reflect the rules of the user's facial expression changes.

[0042] Step S214: Perform block gray value comparison on the image blocks in the eye region of the face image sequence, count the pixel area change rate of the eyelid covering the pupil region per unit time, and generate an eye opening and closing state waveform diagram with time as the horizontal axis.

[0043] An image patch of the eye region refers to an image area containing eye information extracted from a sequence of face images. Block gray value comparison means dividing the image patch of the eye region into multiple small blocks and then comparing the gray value changes of each small block. By statistically calculating the pixel area change rate of the pupil region covered by the eyelid within a unit time, the opening and closing state of the eye can be understood. An opening and closing state waveform diagram of the eye with time as the horizontal axis can intuitively show the change of the opening and closing state of the eye over time. For example, the image patch of the eye region is divided into 8×8 small blocks, and the gray value changes of each small block between adjacent frames are compared. When the gray value change exceeds a preset threshold, it is considered that the state of the small block has changed. Then, the pixel area change rate of the pupil region covered by the eyelid is statistically calculated, and these change rates are plotted on the waveform diagram in chronological order to obtain the opening and closing state waveform diagram of the eye. This waveform diagram can help analyze the eye movement and attention concentration of the user.

[0044] Step S215: Calculate the coordination index between the muscle movement and the eye movement according to the phase difference between the periodic feature and the opening and closing state waveform diagram of the eye, and perform weighted fusion of the coordination index and the movement direction consistency parameter to generate the facial muscle movement trajectory feature.

[0045] The periodic feature refers to the periodic law of facial muscle movement, and the opening and closing state waveform diagram of the eye reflects the change of the opening and closing state of the eye over time. The phase difference refers to the time offset between the periodic feature and the opening and closing state waveform diagram of the eye. By calculating the phase difference, the coordination degree between the muscle movement and the eye movement can be understood. The coordination index is a numerical value used to measure the coordination degree between the muscle movement and the eye movement. Performing weighted fusion of the coordination index and the movement direction consistency parameter can comprehensively consider the stability of the coordinated movement of the muscle group and the coordination degree between the muscle movement and the eye movement to generate the facial muscle movement trajectory feature. For example, calculate the coordination index according to the phase difference between the periodic feature and the opening and closing state waveform diagram of the eye, assign different weights to the coordination index and the movement direction consistency parameter respectively, and then perform weighted summation on them to obtain the facial muscle movement trajectory feature. This feature can more comprehensively reflect the facial dynamic behavior of the user.

[0046] Step S216: Perform frequency domain transformation on the opening and closing state waveform diagram of the eye, extract the effective waveform segments with amplitudes exceeding the preset noise threshold, and generate the eye opening and closing frequency feature according to the interval duration and duration ratio of the effective waveform segments.

[0047] Frequency domain transformation is a technique that converts a time-domain signal into a frequency-domain signal. By performing frequency domain transformation on the waveform diagram of the eye opening and closing state, the waveform diagram can be transformed from the time domain to the frequency domain, thereby analyzing its frequency components. The preset noise threshold is a pre-set amplitude standard used to screen out valid waveform segments. A valid waveform segment refers to the part of the waveform whose amplitude exceeds the preset noise threshold. The interval duration is the time interval between adjacent valid waveform segments, and the duration ratio is the ratio of the duration of the valid waveform segment to the interval duration. By analyzing the interval duration and duration ratio of the valid waveform segments, the eye opening and closing frequency characteristics can be generated. For example, the fast Fourier transform (FFT) is used to perform frequency domain transformation on the waveform diagram of the eye opening and closing state, converting the waveform diagram into a frequency-domain signal. Then, a preset noise threshold is set to screen out the valid waveform segments whose amplitude exceeds this threshold. The interval duration and duration ratio of the valid waveform segments are statistically analyzed, and this information is integrated into the eye opening and closing frequency characteristics. This characteristic can accurately reflect the eye opening and closing frequency of the user.

[0048] Step S220: According to the eye opening and closing frequency characteristics, determine the user's focus level on the screen content, and combine the muscle movement trajectory characteristics to generate a user intention prediction vector.

[0049] In this embodiment, the eye opening and closing frequency characteristics can reflect the user's eye activities. Generally, a lower eye opening and closing frequency indicates that the user is more focused on the screen content, while a higher frequency may indicate that the user's attention is dispersed. The focus level is a quantitative division of the user's focus degree based on the eye opening and closing frequency characteristics. For example, it can be divided into three levels: high, medium, and low. The muscle movement trajectory characteristics can reflect the user's facial expression changes and potential intentions. The user intention prediction vector is a vector containing multiple dimensions of information used to predict the user's intention to interact with the screen. Specifically, according to the eye opening and closing frequency characteristics, determining the user's focus level on the screen content and combining the muscle movement trajectory characteristics to generate a user intention prediction vector may include the following steps: Step S221: According to the interval duration of the valid waveform segments in the eye opening and closing frequency characteristics, determine the cumulative duration of the user's line of sight leaving the screen per unit time, and combine the duration ratio of the valid waveform segments to generate an initial focus score.

[0050] The duration of the effective waveform segment interval refers to the time interval between adjacent effective waveform segments, which can reflect the time when the user's line of sight leaves the screen. By counting the total duration of the effective waveform segment intervals within a unit time, the cumulative duration of the user's line of sight leaving the screen within the unit time can be determined. The duration ratio of the effective waveform segment refers to the ratio of the duration of the effective waveform segment to the interval duration, which can reflect the stability of the user's eye opening and closing. Combining these two factors, an initial concentration score can be generated. For example, if the cumulative duration of the line of sight leaving the screen within a unit time is short and the duration ratio of the effective waveform segment is large, it indicates that the user's attention is relatively concentrated and the initial concentration score is high; otherwise, the score is low. In this way, the user's concentration on the screen content can be preliminarily evaluated.

[0051] Step S222: Detect the displacement direction and amplitude of the coordinates in the mouth corner area of the muscle movement trajectory characteristics. When a continuous upward movement appears in the mouth corner area, trigger the concentration compensation mechanism to make a positive correction to the initial concentration score.

[0052] The displacement direction and amplitude of the coordinates in the mouth corner area can reflect the user's facial expression changes. When a continuous upward movement appears in the mouth corner area, it usually indicates that the user is in a pleasant or concentrated state. The concentration compensation mechanism is a strategy for adjusting the initial concentration score. When a continuous upward movement is detected in the mouth corner area, it will make a positive correction to the initial concentration score to improve the accuracy of the score. For example, if it is detected that the mouth corner area has an upward movement in three consecutive frames of images, the concentration compensation mechanism is triggered to increase the initial concentration score by one level. In this way, the actual concentration level of the user can be more accurately reflected.

[0053] Step S223: Map the corrected concentration score to a preset grade division interval to determine the concentration level.

[0054] The preset grade division interval is a range preset for dividing the concentration level. For example, the concentration level can be divided into three levels: high, medium, and low, and the corresponding score intervals are [80, 100], [50, 79], and [0, 49] respectively. Mapping the corrected concentration score into these grade division intervals can determine the user's concentration level. For example, if the corrected concentration score is 85 points, it is mapped to the high-grade interval to determine that the user's concentration level is high. In this way, the user's concentration on the screen content can be intuitively understood.

[0055] Step S224: Perform a trend prediction on the movement direction consistency parameter in the eyebrow area of the muscle movement trajectory characteristics to generate a muscle activity intensity vector reflecting the user's potential attention.

[0056] The consistency parameter of the movement direction in the eyebrow area can reflect the stability of the coordinated movement of the muscle groups in the eyebrow area. By predicting the trend of this parameter, the changing trend of the muscle activities in the user's eyebrow area can be understood. The muscle activity intensity vector is a vector containing multi-dimensional information, which is used to reflect the user's potential attention. For example, the time series analysis method is used to predict the trend of the consistency parameter of the movement direction in the eyebrow area, and the muscle activity intensity vector is generated according to the prediction result. If the prediction result shows that the parameter is on the rise, it indicates that the user's potential attention is increasing, and the corresponding dimension value in the muscle activity intensity vector will also increase; otherwise, it indicates that the potential attention is decreasing. This vector can provide important reference information for subsequent user intention prediction.

[0057] Step S225: Input the concentration level and the muscle activity intensity vector into a pre-trained long short-term memory network for time series feature matching, and extract the attention transfer patterns that are strongly related to the screen interaction intention.

[0058] The pre-trained long short-term memory network (LSTM) is a neural network model that can process sequential data and capture the long-term dependencies in the sequential data. Inputting the concentration level and the muscle activity intensity vector into this network for time series feature matching can find the attention transfer patterns that are strongly related to the screen interaction intention. The attention transfer pattern refers to the changing rule of the user's attention during the interaction with the screen. For example, when the user's concentration level changes and the muscle activity intensity vector also changes accordingly, the LSTM network can learn this changing pattern and use it as the attention transfer pattern that is strongly related to the screen interaction intention. These patterns can help predict the user's next interaction intention.

[0059] Step S226: According to the movement priorities of different facial areas in the attention transfer pattern, weight the intention correlation of the muscle movement trajectory features to generate a user intention prediction vector containing multi-dimensional weight coefficients.

[0060] The movement priorities of different facial areas in the attention transfer pattern refer to the influence degree of the movements of different facial areas on the interaction intention during the attention transfer. According to these priorities, weighting the intention correlation of the muscle movement trajectory features can highlight the feature information of the facial areas related to the interaction intention. The multi-dimensional weight coefficients refer to the different weight values assigned to each dimension of the muscle movement trajectory features. Through the weighting process, a user intention prediction vector containing multi-dimensional weight coefficients is generated. For example, if the attention transfer pattern shows that the movement of the eyebrow area has a greater influence on the interaction intention, a higher weight value will be assigned to the muscle movement trajectory features of the eyebrow area during the weighting process. The user intention prediction vector generated in this way can more accurately reflect the user's interaction intention.

[0061] Step S230: Perform a similarity match between the user intention prediction vector and the pattern features in the interaction pattern library, and filter out a set of candidate interaction patterns whose similarity exceeds the first threshold.

[0062] In this embodiment, the user intention prediction vector is a vector containing multi-dimensional weight coefficients, which is used to predict the user's intention to interact with the screen. The pattern features in the interaction pattern library refer to the feature vectors corresponding to each interaction pattern, which contain information such as the facial action trigger conditions and screen response strategies of the interaction pattern. The similarity match refers to calculating the similarity between the user intention prediction vector and the pattern features. Common similarity calculation methods include cosine similarity, Euclidean distance, etc. The first threshold is a pre-set similarity standard used to filter out interaction patterns with a relatively high similarity to the user intention prediction vector. For example, use cosine similarity to calculate the similarity between the user intention prediction vector and each pattern feature in the interaction pattern library, and filter out the interaction patterns with a similarity exceeding 0.8 to form a set of candidate interaction patterns. This can find interaction patterns that are more in line with the user's current intention.

[0063] Step S240: Based on the focus level, perform weight correction on each candidate interaction pattern in the set of candidate interaction patterns, and select the candidate interaction pattern with the highest weight as the target interaction pattern; where each interaction pattern in the interaction pattern library is associated with at least one facial action trigger condition and the corresponding screen response strategy.

[0064] The focus level can reflect the user's attention level to the screen content. Different focus levels may require different interaction patterns to meet the user's needs. Performing weight correction on each candidate interaction pattern in the set of candidate interaction patterns based on the focus level means assigning different weight values to each candidate interaction pattern according to the focus level. For example, when the user's focus level is high, they may be more inclined to choose complex and interesting interaction patterns, so a higher weight is assigned to such interaction patterns; while when the focus level is low, simple and direct interaction patterns may be more suitable, and a higher weight is assigned to them. Selecting the candidate interaction pattern with the highest weight as the target interaction pattern can ensure that the selected interaction pattern best meets the user's current state and needs. Each interaction pattern in the interaction pattern library is associated with at least one facial action trigger condition and the corresponding screen response strategy. For example, when the user makes a smiling expression (facial action trigger condition), a cute animation can pop up on the screen (screen response strategy).

[0065] Step S300: Based on the screen gaze area coordinates and the target interaction pattern, generate dynamic interaction content corresponding to the user's line of sight focus.

[0066] In this embodiment, the screen gaze area coordinates refer to the position information of the screen area where the user's line of sight focuses during the screen interaction process. The target interaction mode is the most suitable interaction mode matched from a preset interaction mode library according to the user's facial dynamic behavior characteristics. The dynamic interaction content refers to the interaction content with dynamic effects generated based on the user's line of sight focus and the target interaction mode.

[0067] Specifically, based on the screen gaze area coordinates and the target interaction mode, generating dynamic interaction content corresponding to the user's line of sight focus may include the following steps: Step S310: Determine the current focused control identifier of the user in the screen interface according to the screen gaze area coordinates.

[0068] The controls in the screen interface refer to various interactive elements on the screen, such as buttons, text boxes, icons, etc. Through the screen gaze area coordinates, the position of the control where the user's current line of sight focuses can be determined, and then the identifier of the control can be determined. For example, if the screen gaze area coordinates fall within the area of a certain button, it can be determined that the identifier of the control currently focused by the user is the identifier of that button. This can clarify the interactive element that the user is currently concerned about.

[0069] Step S320: Obtain the historical interaction data associated with the currently focused control. The historical interaction data includes the number of times the user triggers the control, the stay duration, and the associated operation records.

[0070] The historical interaction data refers to the data generated when the user interacted with the currently focused control in the past. The number of triggers refers to the number of times the user clicks or operates the control. The stay duration refers to the time the user's line of sight stays on the control. The associated operation records refer to the records of other relevant operations performed by the user when operating the control. For example, if the currently focused control is a shopping cart button, the historical interaction data may include the number of times the user clicks the button, the time spent on the shopping cart page after each click, and the operation records such as adding products and deleting products in the shopping cart. These historical interaction data can reflect the user's usage habits and preferences for the control.

[0071] Step S330: Generate a dynamic content update instruction corresponding to the currently focused control based on the response rule in the target interaction mode.

[0072] The response rules in the target interaction mode refer to the response methods that the screen should adopt when the user's expression intensity meets the preset conditions. Based on these response rules and combined with the information of the currently focused control, dynamic content update instructions corresponding to the control can be generated. For example, if the target interaction mode stipulates that when the user smiles, the color of the currently focused button changes to green and a prompt message is displayed, the dynamic content update instructions generated according to this rule may include operations such as setting the button color to green and displaying the prompt message. In this way, the content of the currently focused control can be dynamically updated according to the user's state and the target interaction mode.

[0073] Step S340: Adjust the element attributes in the dynamic content update instructions according to the emotion state vector and historical interaction data. The element attributes include color fade rate, content switching frequency, and interaction feedback intensity.

[0074] The emotion state vector is obtained by performing multi-level feature extraction on a sequence of face images, and it can reflect the user's current emotion state. The historical interaction data includes the user's usage habits and preferences for the currently focused control. By adjusting the element attributes in the dynamic content update instructions based on these two factors, the dynamic interaction content can be made more in line with the user's needs and emotion state. The color fade rate refers to the speed of color change, the content switching frequency refers to the time interval of content update, and the interaction feedback intensity refers to the intensity of the feedback effect obtained when the user operates the control. For example, if the emotion state vector shows that the user is in an excited state and the historical interaction data shows that the user likes fast content switching, the content switching frequency in the dynamic content update instructions can be increased; if the user's previous operation habit is to like strong interaction feedback, the interaction feedback intensity can be increased.

[0075] Specifically, adjusting the element attributes in the dynamic content update instructions according to the emotion state vector and historical interaction data may include the following steps: Step S341: Dynamically allocate weights to the emotion dimension parameters in the emotion state vector and the user operation preference parameters in the historical interaction data to generate an emotion-dominated adjustment parameter and a history-dominated adjustment parameter.

[0076] The emotional dimension parameters in the emotional state vector include excitability, anxiety level, pleasure level, etc. These parameters can reflect the user's current emotional state. The user operation preference parameters in the historical interaction data include color sensitivity, browsing speed, touch pressure, etc. These parameters can reflect the user's usage habits and preferences. Dynamic weight assignment means assigning different weight values to the emotional dimension parameters and user operation preference parameters according to different situations. For example, when the user's emotions change significantly, a higher weight is assigned to the emotional dimension parameters; when the user's operation habits are relatively stable, a higher weight is assigned to the user operation preference parameters. Through dynamic weight assignment, an emotion-dominated adjustment parameter and a history-dominated adjustment parameter are generated. The emotion-dominated adjustment parameter mainly considers the user's emotional state, and the history-dominated adjustment parameter mainly considers the user's historical operation habits.

[0077] Step S342: According to the element attribute types to be adjusted in the dynamic content update instruction, perform attribute association mapping on the emotion-dominated adjustment parameter and the history-dominated adjustment parameter, where the color fade rate is associated with the excitability parameter in the emotional dimension and the color sensitivity parameter in the user operation preference, the content switching frequency is associated with the anxiety level parameter in the emotional dimension and the browsing speed parameter in the user operation preference, and the interaction feedback intensity is associated with the pleasure level parameter in the emotional dimension and the touch pressure parameter in the user operation preference.

[0078] Attribute association mapping means corresponding and associating the emotion-dominated adjustment parameter and the history-dominated adjustment parameter with the element attribute types to be adjusted in the dynamic content update instruction. For example, the color fade rate is associated with the excitability parameter in the emotional dimension and the color sensitivity parameter in the user operation preference. When the excitability is high and the color sensitivity is strong, the color fade rate can be increased; the content switching frequency is associated with the anxiety level parameter in the emotional dimension and the browsing speed parameter in the user operation preference. When the anxiety level is high and the browsing speed is fast, the content switching frequency can be increased; the interaction feedback intensity is associated with the pleasure level parameter in the emotional dimension and the touch pressure parameter in the user operation preference. When the pleasure level is high and the touch pressure is large, the interaction feedback intensity can be enhanced. Through this association mapping, the element attributes can be adjusted more accurately according to the user's emotional state and operation preferences.

[0079] Step S343: Call the preset dynamic balance rule to sort the parameters after attribute association mapping. When the change range of the emotional dimension parameters exceeds the stability threshold of the historical operation parameters, the emotion-dominated adjustment parameter is preferentially used to instantaneously adjust the element attributes; otherwise, the element attributes are gradually adjusted based on the history-dominated adjustment parameter.

[0080] The preset dynamic balance rule is a rule preset for determining parameter priorities. The parameters after attribute association mapping include emotion-dominated adjustment parameters and history-dominated adjustment parameters. Priority sorting refers to sorting these parameters according to the dynamic balance rule to determine which parameter has a higher priority when adjusting element attributes. The stability threshold of historical operation parameters is a preset standard for judging the stability of historical operation parameters. When the change range of emotion dimension parameters exceeds the stability threshold of historical operation parameters, it indicates that the user's emotion has changed greatly. At this time, the emotion-dominated adjustment parameters are preferentially used to instantaneously adjust the element attributes to quickly respond to the user's emotion change; otherwise, the element attributes are gradually adjusted based on the history-dominated adjustment parameters to maintain consistency with the user's historical operation habits. For example, if the excitement parameter in the emotion dimension suddenly increases and exceeds the stability threshold of historical operation parameters, the emotion-dominated adjustment parameter is immediately used to increase the color fade rate; if the emotion change is small, the color fade rate is gradually adjusted according to the history-dominated adjustment parameter.

[0081] Step S344: According to the priority sorting result, generate a set of element attribute adjustment instructions including the adjustment amplitude and the effective timing, and verify whether each parameter in the adjustment instruction set exceeds the safe execution range of the screen rendering engine.

[0082] The priority sorting result determines the parameters that should be preferentially used when adjusting element attributes. Based on this result, a set of element attribute adjustment instructions including the adjustment amplitude and the effective timing is generated. The adjustment amplitude refers to the degree to which the element attributes need to be adjusted, and the effective timing refers to the time sequence when the adjustment instruction takes effect. For example, if the priority sorting result indicates that the emotion-dominated adjustment parameter should be preferentially used to adjust the color fade rate, the adjustment amplitude is to increase by 50%, and the effective timing is to take effect immediately, then the generated set of element attribute adjustment instructions should include this information. Verifying whether each parameter in the adjustment instruction set exceeds the safe execution range of the screen rendering engine is to ensure that the adjustment operation will not have an adverse impact on the screen display. For example, if the adjusted color fade rate is too fast, it may cause the screen to flicker. At this time, the adjustment amplitude needs to be appropriately adjusted to keep it within the safe execution range.

[0083] Step S345: Perform instruction-level fusion on the verified adjustment instruction set and the dynamic content update instruction, perform an overwriting update on the effective logic of the element attributes, generate a content rendering strategy adapted to the current user state, and trigger the screen interaction module to load the updated interactive content stream.

[0084] Instruction layer fusion refers to combining the verified set of adjustment instructions with the dynamic content update instructions to form a unified instruction set. Overwriting the effective logic of element attributes means replacing the original effective logic with the adjusted effective logic to ensure that the element attributes are updated according to the new adjustment requirements. Generating a content rendering strategy adapted to the current user state is to formulate the most suitable content rendering plan for the current user based on the user's emotional state, operation preferences, and the adjusted instruction set. Triggering the screen interaction module to load the updated interactive content stream means sending the generated content rendering strategy to the screen interaction module so that it loads and displays the updated interactive content. For example, integrating information such as the adjusted color fade rate, content switching frequency, and interaction feedback intensity into the dynamic content update instructions, updating the effective logic of element attributes, then generating a new content rendering strategy, and finally having the screen interaction module load and display the interactive content generated according to this strategy.

[0085] Step S350: Send the adjusted dynamic content update instructions to the screen rendering engine to generate an interactive content stream with three-dimensional visual effects.

[0086] The screen rendering engine is a software module used to convert dynamic content update instructions into actual screen display content. Sending the adjusted dynamic content update instructions to the screen rendering engine, it will generate an interactive content stream with three-dimensional visual effects according to these instructions. The three-dimensional visual effects can enhance the three-dimensional sense and realism of the interactive content and improve the user's interaction experience. For example, the screen rendering engine can generate a three-dimensional interactive scene with gradient colors, dynamically switched content, and strong interaction feedback according to parameters such as the adjusted color fade rate, content switching frequency, and interaction feedback intensity, and then output it in the form of an interactive content stream for the screen display module to display.

[0087] Step S400: Call the pre-trained facial feature fusion model to perform multi-level feature extraction on the facial regions in the sequence of face images to generate a user authentication vector and an emotional state vector.

[0088] In this embodiment, the pre-trained facial feature fusion model is a model trained with a large amount of face image data and can extract multi-level feature information from face images. The sequence of face images is a set of images containing user facial information arranged in chronological order. Multi-level feature extraction means extracting features from the facial regions in the sequence of face images from different levels and angles to obtain more comprehensive and accurate feature information. The user authentication vector is a feature vector used to verify the user's identity, and the emotional state vector is a feature vector used to reflect the user's current emotional state. Specifically, calling the pre-trained facial feature fusion model to perform multi-level feature extraction on the facial regions in the sequence of face images may include the following steps: Step S410: Divide each frame image in the face image sequence into a first facial region, a second facial region, and a third facial region. The first facial region contains eye contour coordinates, the second facial region contains mouth contour coordinates, and the third facial region contains overall face contour coordinates.

[0089] Dividing each frame image in the face image sequence into regions can extract features from different facial regions more specifically. The first facial region contains eye contour coordinates. The eyes are an important part of the face for expressing emotions and attention. By extracting the features of this region, the user's eye movement and attention can be understood. The second facial region contains mouth contour coordinates. The movements and expressions of the mouth can reflect the user's emotions and language expression intentions. The third facial region contains overall face contour coordinates, which can provide information about the overall shape and structure of the face. For example, using a deep learning-based facial key point detection algorithm, detect the contour coordinates of the eyes, mouth, and overall face in the face image, and then divide the image into three regions according to these coordinates. In this way, the features of different regions can be extracted and analyzed separately.

[0090] Step S420: Through the first feature extraction branch of the facial feature fusion model, perform local texture analysis on the first facial region to generate a first region feature vector containing the pupil movement trajectory.

[0091] The first feature extraction branch of the facial feature fusion model is a sub-model dedicated to processing the first facial region. Local texture analysis refers to the extraction and analysis of local texture features of the first facial region, such as the direction and density of the texture. The pupil movement trajectory refers to the position change of the pupil over a period of time, which can reflect the user's line of sight direction and attention focus. By performing local texture analysis on the first facial region, a first region feature vector containing the pupil movement trajectory is generated. For example, using a convolutional neural network (CNN) as the first feature extraction branch, perform convolutional operations on the first facial region to extract local texture features. At the same time, by tracking the position change of the pupil, the pupil movement trajectory information is incorporated into the first region feature vector. This feature vector can provide important eye-related information for subsequent user authentication and emotional state analysis.

[0092] Step S430: Through the second feature extraction branch of the facial feature fusion model, perform dynamic deformation monitoring on the second facial region to generate a second region feature vector containing the lip opening and closing amplitude.

[0093] The second feature extraction branch of the facial feature fusion model is a sub-model for processing the second facial region. Dynamic deformation monitoring refers to monitoring the dynamic changes in the second facial region, such as the opening and closing of the mouth, the upward or downward movement of the corners of the mouth, etc. The lip opening and closing amplitude refers to the maximum distance during the opening and closing process of the lips, which can reflect the user's expression and language expression. By performing dynamic deformation monitoring on the second facial region, a second region feature vector containing the lip opening and closing amplitude is generated. For example, using a dynamic deformation monitoring algorithm based on the optical flow method to process the second facial region and calculate the lip opening and closing amplitude. Then, the lip opening and closing amplitude and other relevant dynamic deformation information are integrated into the second region feature vector. This feature vector can provide important mouth-related information for emotion state analysis and user authentication.

[0094] Step S440: Through the third feature extraction branch of the facial feature fusion model, perform global illumination compensation processing on the third facial region to generate a third region feature vector containing the skin color change trend.

[0095] The third feature extraction branch of the facial feature fusion model is a sub-model for processing the third facial region. Global illumination compensation processing refers to adjusting the illumination of the third facial region to eliminate the impact of uneven illumination on the image quality. The skin color change trend refers to the change of skin color over a period of time, which can reflect the user's health status and emotion state. By performing global illumination compensation processing on the third facial region, a third region feature vector containing the skin color change trend is generated. For example, using histogram equalization and adaptive illumination compensation algorithms to adjust the illumination of the third facial region to make the image brightness more uniform. Then, by analyzing the color features of the image, the skin color change trend information is extracted and integrated into the third region feature vector. This feature vector can provide important overall facial-related information for user authentication and emotion state analysis.

[0096] Step S450: Cross-channel fuse the first region feature vector, the second region feature vector, and the third region feature vector to generate a user authentication vector and an emotion state vector.

[0097] Cross-channel fusion refers to the merging and integration of feature vectors from different regions to obtain more comprehensive and integrated feature information. By performing cross-channel fusion on the first-region feature vector, the second-region feature vector, and the third-region feature vector, the feature information of different facial regions can be fully utilized to generate more accurate user authentication vectors and emotion state vectors. For example, a fully connected layer is used to concatenate the feature vectors of the three regions, and then through a non-linear transformation, they are mapped to a new feature space to generate user authentication vectors and emotion state vectors. The user authentication vector can be used to verify the user's identity, ensuring that only authorized users can perform interactive operations; the emotion state vector can reflect the user's current emotion, providing a basis for subsequent adjustment of interactive content.

[0098] In an embodiment of the present invention, the training process of the facial feature fusion model may include the following steps: Step S401: Obtain a training image set labeled with user identity labels and emotion labels, where each image in the training image set is labeled with eye region coordinates, mouth region coordinates, and overall facial region coordinates.

[0099] The training image set is a set of image data used to train the facial feature fusion model. The training image set labeled with user identity labels and emotion labels means that corresponding user identity information and emotion state information are labeled for each image. The eye region coordinates, mouth region coordinates, and overall facial region coordinates are information used to determine the positions of different facial regions in the image. For example, a large number of face images are collected, and each image is labeled with the user's identity identifier (such as name, number, etc.) and emotion label (such as happy, sad, angry, etc.), and at the same time, the region coordinates of the eyes, mouth, and overall face are labeled using a facial key point detection algorithm. These labeled information can provide supervision signals for the training of the model, enabling the model to learn the relationship between the features of different facial regions and the user's identity and emotion state.

[0100] Step S402: Construct an initial fusion model containing three parallel feature extraction branches, where the three branches correspond to the feature extraction of the first facial region, the second facial region, and the third facial region respectively.

[0101] The initial fusion model is the initial version of the facial feature fusion model, which contains three parallel feature extraction branches. Each branch is used to process the feature extraction of the first facial region, the second facial region, and the third facial region respectively. For example, a convolutional neural network (CNN) is used to construct each feature extraction branch, and each branch has a different convolutional layer and pooling layer structure to adapt to the feature characteristics of different facial regions. The three branches work in parallel, and can simultaneously extract the features of different facial regions, improving the processing efficiency of the model.

[0102] Step S403: Perform region division on each image in the training image set, extract the training region images corresponding to each branch respectively, and generate the region feature vectors of each branch.

[0103] Performing region division on each image in the training image set is based on the previously marked eye region coordinates, mouth region coordinates, and overall face region coordinates, and dividing the image into three regions. Extracting the training region images corresponding to each branch respectively is to input the divided region images into the corresponding feature extraction branches for processing. Generating the region feature vectors of each branch is to perform feature extraction on the training region images through the feature extraction branches to obtain the feature vectors of each branch. For example, divide each image in the training image set into the first face region, the second face region, and the third face region according to the marked coordinates, then input the images of each region into the corresponding feature extraction branches respectively, and extract features through convolutional operations and pooling operations to generate the region feature vectors of each branch.

[0104] Step S404: Input the region feature vectors of each branch into the fully connected layer for feature concatenation to generate the fused feature vector.

[0105] The fully connected layer is a neural network layer that connects all the input neurons to all the output neurons. Inputting the region feature vectors of each branch into the fully connected layer for feature concatenation is to combine the feature vectors of the three branches into one vector to generate the fused feature vector. For example, arrange the first region feature vector, the second region feature vector, and the third region feature vector in sequence, then input them into the fully connected layer, and concatenate them into a longer vector through linear transformation to obtain the fused feature vector. This fused feature vector contains the comprehensive feature information of different face regions.

[0106] Step S405: Obtain the first cross-entropy loss between the fused feature vector and the user identity label, and the second cross-entropy loss between the fused feature vector and the emotion label.

[0107] The cross-entropy loss is a loss function used to measure the difference between the model's prediction results and the true labels. Obtaining the first cross-entropy loss between the fused feature vector and the user identity label is to calculate the degree of difference between the user identity predicted by the fused feature vector and the true user identity label. Obtaining the second cross-entropy loss between the fused feature vector and the emotion label is to calculate the degree of difference between the emotion state predicted by the fused feature vector and the true emotion label. For example, use the softmax function to convert the fused feature vector into a probability distribution, and then calculate the cross-entropy loss with the true user identity label and emotion label. These two loss values can reflect the accuracy of the model in user identity recognition and emotion state prediction.

[0108] Step S406: Optimize the parameters of the initial fusion model through backpropagation based on the weighted sum of the first cross-entropy loss and the second cross-entropy loss until the loss converges.

[0109] Backpropagation optimization is an optimization method that calculates the gradient of the loss function with respect to the model parameters and updates the model parameters according to the gradient. Optimizing the parameters of the initial fusion model through backpropagation based on the weighted sum of the first cross-entropy loss and the second cross-entropy loss means adding the two loss values according to preset weights to obtain a total loss value, and then using the backpropagation algorithm to update the model parameters to continuously reduce the total loss value. For example, different weights, such as 0.6 and 0.4, are assigned to the first cross-entropy loss and the second cross-entropy loss respectively, and they are added together to obtain the total loss value. Then, the Stochastic Gradient Descent (SGD) algorithm is used to update the model parameters according to the total loss value, and the iteration is continued until the loss value converges to a smaller value. This can enable the model to have good performance in both user identity recognition and emotion state prediction.

[0110] Step S407: Solidify the optimized model parameters to generate a pre-trained facial feature fusion model.

[0111] Solidifying the optimized model parameters means saving the model parameters after backpropagation optimization so that they no longer change. Generating a pre-trained facial feature fusion model is applying the solidified model parameters to the initial fusion model to obtain a pre-trained model that can be directly used for feature extraction. For example, the optimized model parameters are saved as a file, and then when using the facial feature fusion model, this file is loaded and the parameters are assigned to the initial fusion model to make it a pre-trained model. This pre-trained model can quickly and accurately extract the feature information of face images in practical applications.

[0112] Step S500: According to the user identity authentication vector and the emotion state vector, adjust the presentation parameters of the dynamic interaction content in real time and output the adjusted screen interaction signal.

[0113] In this embodiment, the user identity authentication vector is used to verify the user's identity, and different users may have different preferences and needs. The emotion state vector reflects the user's current emotion state, and the change of emotion may affect the user's feelings and needs for the interaction content. The presentation parameters of the dynamic interaction content include interface brightness, content layout density, font size, color contrast, etc., and these parameters will affect the display effect and user experience of the interaction content. Real-time adjustment means adjusting the presentation parameters in a timely manner according to the changes of the user identity authentication vector and the emotion state vector. The adjusted screen interaction signal refers to the interaction content signal after adjusting the presentation parameters, which can make the screen display interaction content more in line with the user's current state.

[0114] Specifically, according to the user authentication vector and the emotional state vector, the real-time adjustment of the presentation parameters of the dynamic interaction content may include the following steps: Step S510: Monitor the user's facial deflection angle and the distance parameter from the screen, and generate a screen perspective correction parameter.

[0115] The facial deflection angle refers to the tilt angle of the user's face relative to the screen, which affects the user's visual perception of the screen content. The distance parameter from the screen refers to the actual distance between the user and the screen. Different distances may require different display parameters. By monitoring the user's facial deflection angle and the distance parameter from the screen, a screen perspective correction parameter can be generated. For example, a depth camera or an infrared sensor is used to monitor the user's facial deflection angle and the distance from the screen, and the screen perspective correction parameter is calculated based on this data. If the user's facial deflection angle is large, the display direction or perspective of the screen content may need to be adjusted; if the user is far from the screen, the font size and icon size may need to be increased. These correction parameters can ensure that the user can clearly see the screen content at different perspectives and distances.

[0116] Step S520: According to the user authentication vector, retrieve the preference setting information from the user portrait database. The preference setting information includes the font size preference, the color contrast threshold, and the animation playback speed range.

[0117] The user portrait database is a database that stores various preference information of users. According to the user authentication vector, the preference setting information corresponding to the current user can be retrieved from this database. The font size preference refers to the font size that the user likes, the color contrast threshold refers to the range of color contrast that the user can accept, and the animation playback speed range refers to the interval of animation playback speed that the user likes. For example, after the user authentication vector is verified, according to the user identity information in the vector, the corresponding preference setting information is searched from the user portrait database. If the user has previously set a preference for a larger font size, a higher color contrast, and a faster animation playback speed, then this information can be retrieved to provide a basis for subsequent adjustment of the presentation parameters.

[0118] Step S530: Input the emotional state vector into the pre-trained parameter mapping network to generate an interface brightness adjustment value and a content layout density value that match the current emotion.

[0119] The pre-trained parameter mapping network is a neural network model trained with a large amount of data. It can map the emotional state vector to the corresponding interface brightness adjustment value and content layout density value. The interface brightness adjustment value refers to the value for adjusting the brightness of the screen interface, and the content layout density value refers to the distribution density of the content on the screen. For example, using a long short-term memory network (LSTM) as the pre-trained parameter mapping network, the emotional state vector is input into the network for processing. If the emotional state vector indicates that the user is in an excited state, the network may generate a higher interface brightness adjustment value and a lower content layout density value to create a lively and open interaction atmosphere; if the user is in a calm state, it may generate a moderate interface brightness adjustment value and a moderate content layout density value. These adjustment values can dynamically adjust the display effect of the interactive content according to the user's emotional state.

[0120] Step S540: Based on the screen perspective correction parameter, preference setting information, and interface brightness adjustment value, determine the final display parameter combination of the dynamic interactive content.

[0121] The final display parameter combination refers to the optimal set of display parameters for the dynamic interactive content determined by comprehensively considering the screen perspective correction parameter, preference setting information, and interface brightness adjustment value. For example, according to the screen perspective correction parameter, adjust the display direction and perspective of the screen content; according to the preference setting information, determine parameters such as font size, color contrast, and animation playback speed; according to the interface brightness adjustment value, adjust the brightness of the screen interface. Then combine these parameters to form the final display parameter combination. This combination can ensure that the dynamic interactive content is displayed with the best effect under different perspectives, user preferences, and emotional states.

[0122] Step S550: According to the content layout density value, reorganize the information elements in the interactive content stream and trigger the screen display module to perform a parameter update operation.

[0123] The content layout density value determines the distribution density of the information elements on the screen. Reorganizing the information elements in the interactive content stream according to this value means adjusting the arrangement and spacing of the information elements to achieve an appropriate layout density. For example, if the content layout density value is low, it may be necessary to increase the spacing between the information elements to make the content more open; if the content layout density value is high, it may be necessary to reduce the spacing and increase the number of information elements. Triggering the screen display module to perform a parameter update operation means sending the final display parameter combination and the reorganized interactive content stream to the screen display module so that it can display according to the new parameters and content. This can achieve real-time adjustment and update of the dynamic interactive content and improve the user's interaction experience.

[0124] In the embodiment of the present invention, the training process of the above parameter mapping network may include the following steps: Step S501: Collect feedback data of multiple groups of users on screen interaction content in different emotional states. The feedback data includes records of users' active adjustment of interface parameters and physiological signal monitoring data.

[0125] Collecting feedback data of multiple groups of users on screen interaction content in different emotional states is to obtain users' requirements and reactions to interface parameters in different emotions. The record of users' active adjustment of interface parameters refers to the operation records of users manually adjusting parameters such as interface brightness, font size, and color contrast during the interaction with the screen. Physiological signal monitoring data refers to the physiological signals of users monitored by sensors, such as heart rate, blood pressure, and skin conductance response. These signals can reflect the emotional state of users. For example, use the methods of questionnaire survey and sensor monitoring to collect feedback data of multiple groups of users on screen interaction content in different emotional states such as happiness, sadness, and anger. These data can provide rich samples for the training of the parameter mapping network.

[0126] Step S502: Construct an initial mapping network including an emotional feature input layer and a parameter prediction output layer. The dimension of the emotional feature input layer is consistent with the dimension of the emotional state vector.

[0127] The initial mapping network is the initial version of the parameter mapping network. It includes an emotional feature input layer and a parameter prediction output layer. The emotional feature input layer is used to receive the emotional state vector, and its dimension is consistent with the dimension of the emotional state vector to ensure accurate input of emotional information. The parameter prediction output layer is used to output the interface parameter adjustment values corresponding to the emotional state, such as the interface brightness adjustment value and the content layout density value. For example, use a multi-layer perceptron (MLP) to construct the initial mapping network, set the number of neurons in the emotional feature input layer to be the same as the dimension of the emotional state vector, and set the number of neurons in the parameter prediction output layer to be the number of interface parameters to be predicted. In this way, a network model that can map the emotional state to the interface parameters can be constructed.

[0128] Step S503: Align the emotional state vector and the user physiological signal data in time series to generate a training sample set with time series labels.

[0129] Time series alignment refers to matching the emotional state vector and the user's physiological signal data in terms of time so that they correspond to information in the same time period. Generating a training sample set with time sequence tags is to combine the aligned emotional state vector and the user's physiological signal data with the corresponding interface parameter adjustment values to form a training sample set containing time sequence information. For example, use timestamp information to align the emotional state vector and the user's physiological signal data, and then combine them with the records of the user's active adjustment of interface parameters or the interface parameter adjustment values predicted based on physiological signals to form training samples, and add time tags to each sample. This can ensure that the samples in the training sample set have time sequence information, facilitating the model to learn the dynamic relationship between the emotional state and the interface parameters.

[0130] Step S504: Segment the training sample set through a sliding window mechanism, and extract the average value of emotional features and the parameter adjustment trend vector for each time period.

[0131] The sliding window mechanism is a method for processing time series data. It divides the data into multiple small segments by sliding a window of a fixed size on the time series. Extracting the average value of emotional features and the parameter adjustment trend vector for each time period means performing statistical analysis on the emotional features and parameter adjustment values within each window, calculating the average value of the emotional features and the change trend of the parameter adjustment values. For example, set a sliding window with a size of 10 time steps, slide this window on the training sample set, calculate the average value of the emotional state vector within each window to obtain the average value of emotional features for this time period. At the same time, analyze the change situation of the interface parameter adjustment values within the window to generate the parameter adjustment trend vector. These average values and trend vectors can reflect the change characteristics of the emotional state and the interface parameters in different time periods.

[0132] Step S505: Use a long short-term memory network to perform sequence modeling on the parameter adjustment trend vector to generate parameter prediction values.

[0133] The long short-term memory network (LSTM) is a neural network model that can process sequence data and can capture long-term dependencies in sequence data. Using LSTM to perform sequence modeling on the parameter adjustment trend vector is to input the parameter adjustment trend vector into the LSTM network, allowing the network to learn the sequence pattern of parameter adjustment, thereby generating parameter prediction values. For example, take the parameter adjustment trend vector as the input of the LSTM network, and after the processing and learning of the network, output the predicted interface parameter adjustment values. These prediction values can be used to guide subsequent interface parameter adjustments.

[0134] Step S506: Calculate the mean square error loss between the parameter prediction values and the actual adjusted parameters, and use the gradient descent algorithm to optimize the weight parameters of the initial mapping network.

[0135] The mean squared error loss is a loss function used to measure the difference between the predicted values of a model and the actual values. Calculating the mean squared error loss between the predicted parameter values and the actual adjusted parameters involves comparing the parameter predicted values generated by the LSTM network with the actual interface parameter adjustment values and computing the mean squared error between them. Optimizing the weight parameters of the initial mapping network using the gradient descent algorithm is based on the gradient of the mean squared error loss to update the weight parameters of the initial mapping network, thereby continuously reducing the loss value. For example, the mean squared error loss function is used to calculate the loss value between the predicted parameter values and the actual adjusted parameters, and then the stochastic gradient descent (SGD) algorithm is used to update the weight parameters of the initial mapping network according to the gradient of the loss value. Through continuous iteration, the prediction performance of the model is continuously improved.

[0136] Step S507: When the mean squared error loss is lower than the preset threshold, stop the training and save the network parameters to generate a pre-trained parameter mapping network.

[0137] The preset threshold is a pre-set loss value standard. When the mean squared error loss is lower than this threshold, it indicates that the prediction performance of the model has reached the preset requirements. Stopping the training and saving the network parameters to generate a pre-trained parameter mapping network means saving the current network parameters to form a pre-trained model that can be directly used for parameter mapping. For example, when the mean squared error loss is lower than 0.01, stop the training process and save the weight parameters of the network as a file. In practical applications, load this file and assign the parameters to the initial mapping network to make it a pre-trained parameter mapping network. This pre-trained model can accurately predict the corresponding interface parameter adjustment values according to the user's emotional state.

[0138] As an implementation, after outputting the adjusted screen interaction signal in step S500, the method provided by the embodiments of the present invention may further include the following steps: Step S600: Continuously monitor the response behavior of the user to the adjusted interaction content, where the response behavior includes the line-of-sight movement trajectory, facial micro-expression changes, and touch operation frequency.

[0139] Continuously monitoring the user's response behavior to the adjusted interactive content is to understand the user's satisfaction and usage of the adjusted interactive content. The eye movement trajectory refers to the movement path of the user's eyes on the screen, which can reflect the focus of the user's attention on the screen content and the browsing order. Facial micro-expression changes refer to the subtle changes in the user's facial expressions, such as smiling, frowning, etc., which can reflect the user's emotional reactions. The touch operation frequency refers to the number of touch operations performed by the user on the screen, which can reflect the degree of interaction between the user and the screen. For example, an eye tracker is used to monitor the user's eye movement trajectory, a camera is used to capture the user's facial micro-expression changes, and a sensor is used to record the user's touch operation frequency. By continuously monitoring these response behaviors, the feedback and needs of the user for the interactive content can be discovered in a timely manner.

[0140] Step S700: Perform a matching degree analysis on the response behavior and the preset expected feedback pattern to generate an interactive effect evaluation index.

[0141] The preset expected feedback pattern is the ideal response pattern of the user to the interactive content set in advance, which includes the ideal states in aspects such as eye movement trajectory, facial micro-expression changes, and touch operation frequency. Performing a matching degree analysis on the response behavior and the preset expected feedback pattern is to evaluate the effect of the interactive content by calculating the similarity between the actual response behavior and the expected feedback pattern. The interactive effect evaluation index is a quantitative value used to measure the quality of the interactive content. For example, for the eye movement trajectory, the coincidence degree between the actual trajectory and the expected trajectory can be calculated; for the facial micro-expression changes, the similarity between the actual expression and the expected expression can be judged; for the touch operation frequency, the difference between the actual frequency and the expected frequency can be compared. Then, by synthesizing the matching degrees in these aspects, an interactive effect evaluation index is generated. If the matching degree is high, the interactive effect evaluation index is high, indicating that the interactive content meets the user's needs relatively well; otherwise, it means that the interactive content may need further optimization.

[0142] Step S800: When the evaluation index is lower than the fourth threshold, trigger the interactive mode optimization engine to iteratively update the response rules of the target interactive mode.

[0143] The fourth threshold is a pre-set evaluation index standard used to determine whether the interaction effect has reached an acceptable level. When the interaction effect evaluation index is lower than the fourth threshold, it indicates that the current interaction content has a poor effect, and the response rules of the target interaction mode need to be adjusted. The interaction mode optimization engine is a module dedicated to optimizing the interaction mode. It can iteratively update the response rules of the target interaction mode according to the user's response behavior and evaluation index. For example, if the evaluation index shows that the user's attention to the current interaction content is low, the interaction mode optimization engine can adjust the response rules to increase the interest or attraction of the interaction content. Through continuous iterative updates, the target interaction mode can better meet the user's needs and improve the interaction effect.

[0144] Step S900: Record the characteristic data generated during the current interaction and synchronize the updated interaction mode characteristics to the local mode libraries of all associated devices.

[0145] Recording the characteristic data generated during the current interaction is for subsequent analysis and optimization. These characteristic data include the user's facial dynamic behavior characteristics, screen gaze area coordinates, interaction effect evaluation index, etc. By recording these data, the user's behavior patterns and preferences during the interaction can be understood, providing a basis for further optimizing the interaction mode. Synchronizing the updated interaction mode characteristics to the local mode libraries of all associated devices is to ensure that all associated devices can use the latest interaction mode. For example, if the user interacts with the screen on one device and the interaction mode optimization engine updates the interaction mode, then synchronize the updated interaction mode characteristics to the local mode libraries of the user's other associated devices (such as mobile phones, tablets, etc.), so that the user can also enjoy the optimized interaction experience on other devices.

[0146] Step S1000: Generate a personalized interaction log based on the user authentication vector and encrypt and store the log data in the distributed database node.

[0147] The personalized interaction log is a log file generated based on the user's identity information and data during the interaction process. It contains detailed information about the user during the interaction, such as interaction time, interaction content, response behavior, etc. Generating the personalized interaction log according to the user identity authentication vector can ensure the association between the log file and the user's identity, facilitating subsequent query and analysis. Encrypting and storing the log data in the distributed database nodes is to ensure the security and reliability of the log data. The distributed database nodes refer to multiple database servers distributed in different geographical locations. Through encrypted storage, it can prevent the log data from being illegally obtained and tampered with. For example, using a symmetric encryption algorithm to encrypt the personalized interaction log data and then storing the encrypted data in the distributed database nodes. This can provide secure and reliable data support for subsequent data analysis and user behavior research.

[0148] As an implementation manner, after outputting the adjusted screen interaction signal in step S500, the method provided by the embodiments of the present invention may further include: Step S1100: Real-time collect the feedback data of the user on the adjusted screen interaction signal. The feedback data includes the facial expression change rate, the pupil focus area offset, and the touch operation response delay duration.

[0149] Real-time collecting the feedback data of the user on the adjusted screen interaction signal is to timely understand the user's reaction to the adjusted interaction content. The facial expression change rate refers to the degree of change of the user's facial expression per unit time, which can reflect the fluctuation of the user's emotion. The pupil focus area offset refers to the offset distance of the area where the user's pupil focuses relative to the normal position, which can reflect the degree of the user's attention concentration. The touch operation response delay duration refers to the time interval between the user's touch operation and the screen's response, which can reflect the response speed of the interaction system. For example, using a high-speed camera to capture the change of the user's facial expression in real time and calculate the facial expression change rate; using an eye tracker to monitor the pupil focus area of the user and measure the offset; using a sensor to record the time of the touch operation and the time of the screen response and calculate the touch operation response delay duration. By real-time collecting these feedback data, problems existing in the interaction content can be discovered in a timely manner.

[0150] Step S1200: Align the feedback data and the emotion state vector in time series, extract the correlation features between the emotion fluctuation and the screen interaction content, and generate an interaction effect fluctuation curve.

[0151] Temporally aligning the feedback data with the emotional state vector means matching the feedback data and the emotional state vector in terms of time so that they correspond to information in the same time period. Extracting the correlation features between emotional fluctuations and screen interaction content is to find the correlation patterns between emotional fluctuations and screen interaction content by analyzing the relationship between the feedback data and the emotional state vector. The interaction effect fluctuation curve is a curve with time as the horizontal axis and the interaction effect evaluation value as the vertical axis, which can intuitively show the change of the interaction effect over time. For example, arrange feedback data such as the facial expression change rate, the pupil focus area offset, and the touch operation response delay duration in chronological order with the emotional state vector, and then use the correlation analysis method to find the correlation between emotional fluctuations and these feedback data. Based on these correlations, calculate the interaction effect evaluation value at each time point and draw the interaction effect fluctuation curve. By analyzing this curve, we can understand the change of the interaction effect in different time periods and the impact of emotional fluctuations on the interaction effect.

[0152] Step S1300: Based on the user authentication vector, extract the benchmark interaction pattern that matches the current user from the historical interaction logs, and compare the similarity between the interaction effect fluctuation curve and the expected effect curve in the benchmark interaction pattern to generate an interaction deviation index.

[0153] Extracting the benchmark interaction pattern that matches the current user from the historical interaction logs based on the user authentication vector is to find, according to the user's identity information, the interaction pattern that the user has used before and has good effects from the historical interaction logs. The expected effect curve in the benchmark interaction pattern refers to the curve of the change of the interaction effect over time under ideal conditions for this interaction pattern. Comparing the similarity between the interaction effect fluctuation curve and the expected effect curve in the benchmark interaction pattern is to evaluate the degree of difference between the current interaction effect and the expected effect by calculating the similarity between the two curves. The interaction deviation index is a quantitative value used to represent this degree of difference. For example, use the dynamic time warping (DTW) algorithm to calculate the similarity between the interaction effect fluctuation curve and the expected effect curve, and convert the similarity into an interaction deviation index. If the interaction deviation index is small, it means that the current interaction effect is relatively close to the expected effect; otherwise, it means that there is a large deviation and the interaction content needs to be adjusted.

[0154] Step S1400: Determine the type of interaction element that needs to be optimized first according to the deviation degree of different feedback dimensions in the interaction deviation index, and generate a dynamic optimization strategy including the element type weight and the optimization direction.

[0155] The degree of deviation in different feedback dimensions of the interaction deviation indicator refers to the degree of difference between the current interaction effect and the expected effect in different feedback dimensions such as the rate of change of facial expressions, the offset of the pupil focus area, and the response delay duration of touch operations. Based on these degrees of deviation, determining the types of interaction elements that need to be optimized first is to identify the interaction elements that have a greater impact on the interaction effect, such as interface color, content layout, interaction feedback, etc. Generating a dynamic optimization strategy that includes element type weights and optimization directions is to assign different weights to each type of interaction element that needs to be optimized and determine the optimization direction, such as increasing or decreasing the parameter value of a certain element. For example, if the interaction deviation indicator shows a large degree of deviation in the offset of the pupil focus area, indicating that the user's attention is easily distracted, then it may be necessary to optimize the content layout of the interface first to make important information more prominent. Assign a higher weight to the content layout element and determine the optimization direction as adjusting the arrangement and spacing of information. This can optimize the interaction content in a targeted manner and improve the interaction effect.

[0156] Step S1500: Perform rule integration on the dynamic optimization strategy and the response rules of the target interaction mode, and update the trigger conditions and content generation logic of the corresponding mode in the interaction mode library.

[0157] Performing rule integration on the dynamic optimization strategy and the response rules of the target interaction mode is to incorporate the element type weights and optimization directions in the dynamic optimization strategy into the response rules of the target interaction mode, enabling the response rules to be adjusted according to the optimization strategy. Updating the trigger conditions and content generation logic of the corresponding mode in the interaction mode library is to modify the trigger conditions and content generation logic of the corresponding mode in the interaction mode library according to the integrated rules. For example, if the dynamic optimization strategy requires increasing the contrast of the interface color, then add the corresponding adjustment logic to the response rules of the target interaction mode, and automatically adjust the contrast of the interface color when the preset conditions are met. At the same time, update the trigger conditions and content generation logic of the corresponding mode in the interaction mode library to ensure that during subsequent interaction processes, interaction content can be generated according to the optimized rules.

[0158] Step S1600: Re-match the user's current facial dynamic behavior characteristics according to the updated interaction mode library, generate a secondarily adjusted screen interaction signal, and overwrite the signal execution result output previously.

[0159] Rematching the user's current facial dynamic behavior characteristics according to the updated interaction pattern library is to perform a similarity match between the user's current facial dynamic behavior characteristics and the patterns in the updated interaction pattern library to find the most suitable interaction pattern. Generating the screen interaction signal after secondary adjustment is to re-adjust the presentation parameters of the dynamic interaction content according to the matched interaction pattern to generate a new screen interaction signal. Overwriting the signal execution result of the previous output is to replace the signal of the previous output with the screen interaction signal after secondary adjustment, so that the screen displays the interaction content further optimized. For example, input the dynamic behavior characteristics such as the user's current facial expression and muscle movement into the updated interaction pattern library for matching. If a new interaction pattern is matched, adjust the presentation parameters such as the interface brightness and content layout according to the response rules of the pattern to generate the screen interaction signal after secondary adjustment. Then send this signal to the screen display module to overwrite the signal of the previous output, so that the screen displays the interaction content more in line with the user's needs. Through this continuous optimization and adjustment process, the effect of screen interaction and the user experience can be continuously improved.

[0160] In summary, the screen interaction method based on face recognition provided by the embodiments of the present invention, by obtaining the user's face image sequence, extracting facial dynamic behavior characteristics and emotional state information, matching the target interaction pattern from the preset interaction pattern library, generating dynamic interaction content corresponding to the user's line of sight focus, and adjusting the presentation parameters in real time according to the user's identity and emotional state. At the same time, by continuously monitoring the user's response behavior, optimizing and updating the interaction pattern, and continuously improving the interaction effect and user experience. This method makes full use of face recognition technology and the user's emotional information to achieve more personalized and intelligent screen interaction, and has broad application prospects and practical value. In practical applications, this method can be applied to fields such as smart TVs, advertising screens, and interactive games to provide users with richer and more interesting interaction experiences. For example, in a smart TV, automatically adjust the recommendation and display effect of program content according to the user's expression and emotional state; in an advertising screen, adjust the advertising content and display method in real time according to the audience's reaction to improve the attractiveness and effect of the advertisement; in an interactive game, dynamically adjust the game difficulty and scene according to the player's emotional changes to enhance the fun and immersion of the game. With the continuous development and improvement of technology, this method can also be combined with other technologies (such as virtual reality, augmented reality, etc.) to create more novel and unique interaction experiences.

[0161] Please refer to Figure 2 , Figure 2Schematic structural diagram of a screen interaction system provided by an embodiment of the present invention. The screen interaction system is, for example, a system embedded in an interactive screen or the interactive screen itself, and at least includes a processor 101, a communication interface 102, and a memory 103. Among them, the processor 101, the communication interface 102, and the memory 103 can be connected through a bus or other means. Among them, the processor 101 (or Central Processing Unit (CPU)) is the computing core and control core of the screen interaction system, which can parse various instructions in the screen interaction system and process various data of the screen interaction system. The communication interface 102 can optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 101; the communication interface 102 can also be used for the transmission and interaction of internal data of the screen interaction system. The memory 103 (Memory) is a memory device in the screen interaction system, used to store programs and data. It can be understood that the memory 103 here can include both the built-in memory of the screen interaction system and, of course, the extended memory supported by the screen interaction system. The memory 103 provides a storage space, and the operating system of the screen interaction system is stored in this storage space, which can include but is not limited to: Android system, iOS system, Windows Phone system, etc., and the present invention does not make any limitations in this regard. In one embodiment, the processor 101 executes the screen interaction method based on face recognition provided above in the embodiments of the present invention by running the computer program in the memory 103.

Claims

1. A screen interaction method based on face recognition, characterized in that: The method comprises: Acquire a facial image sequence triggered by a target user in a screen interaction scenario, wherein the facial image sequence includes the user's facial dynamic behavior features and screen gaze area coordinates; According to the facial dynamic behavior characteristics, matching a target interaction mode from a preset interaction mode library, wherein the target interaction mode includes an interaction response rule associated with the intensity of the user's expression; Based on the screen gaze area coordinates and the target interaction mode, generating dynamic interactive content corresponding to the user's visual focus; Calling a pre-trained facial feature fusion model to perform multi-level feature extraction on the facial region in the facial image sequence to generate a user identity verification vector and an emotional state vector; According to the user identity verification vector and the emotional state vector, the presentation parameters of the dynamic interactive content are adjusted in real time, and an adjusted screen interaction signal is output.

2. The method according to claim 1, characterized in that The matching of the target interaction mode from a preset interaction mode library according to the facial dynamic behavior characteristics includes: Inputting the facial dynamic behavior features into a pre-trained spatiotemporal convolutional network to extract facial muscle movement trajectory features and eye opening and closing frequency features; Determining the user's concentration level on the screen content according to the eye opening and closing frequency characteristics, and generating a user intention prediction vector in combination with the muscle movement trajectory characteristics; Performing similarity matching between the user intention prediction vector and the pattern features in the interaction pattern library, and screening out a set of candidate interaction patterns whose similarity exceeds a first threshold; Based on the concentration level, weight correction is performed on each candidate interaction mode in the candidate interaction mode set, and the candidate interaction mode with the highest weight is selected as the target interaction mode; Each interaction mode in the interaction mode library is associated with at least one facial action triggering condition and a corresponding screen response strategy.

3. The method according to claim 2, characterized in that The step of inputting the facial dynamic behavior features into a pre-trained spatiotemporal convolutional network to extract facial muscle movement trajectory features and eye opening and closing frequency features includes: Performing facial key point detection on each frame image in the facial image sequence to generate a key point distribution map including eyebrow region coordinates, cheek region coordinates and mouth corner region coordinates; Based on the key point displacement vectors between adjacent frames, determining the motion direction consistency parameter of each facial region within a preset time window, wherein the motion direction consistency parameter is used to characterize the stability of the coordinated motion of the muscle groups; Performing optical flow field analysis on the key point distribution map to generate a dynamic change curve reflecting the muscle contraction intensity, and extracting the periodic characteristics of facial muscle movement according to the frequency of occurrence of peaks and troughs in the dynamic change curve; Comparing the grayscale values ​​of the image blocks of the eye area in the face image sequence, calculating the pixel area change rate of the eyelid covering the pupil area per unit time, and generating an eye opening and closing state waveform diagram with time as the horizontal axis; Calculating the synergy index of muscle movement and eye movement according to the phase difference between the periodic feature and the eye opening and closing state waveform, and weightedly fusing the synergy index with the movement direction consistency parameter to generate the facial muscle movement trajectory feature; The eye opening and closing state waveform is transformed in the frequency domain to extract the effective waveform segments whose amplitude exceeds the preset noise threshold, and the eye opening and closing frequency characteristics are generated according to the interval length and duration ratio of the effective waveform segments.

4. The method according to claim 2, characterized in that: The calling of the pre-trained facial feature fusion model to perform multi-level feature extraction on the facial region in the facial image sequence includes: Dividing each frame image in the facial image sequence into a first facial region, a second facial region and a third facial region, wherein the first facial region includes eye contour coordinates, the second facial region includes mouth contour coordinates, and the third facial region includes overall facial contour coordinates; Performing local texture analysis on the first facial region through the first feature extraction branch of the facial feature fusion model to generate a first region feature vector including a pupil movement trajectory; Performing dynamic deformation monitoring on the second facial region through the second feature extraction branch of the facial feature fusion model to generate a second region feature vector including the lip opening and closing amplitude; Performing global illumination compensation processing on the third facial region through the third feature extraction branch of the facial feature fusion model to generate a third region feature vector including a skin color change trend; Cross-channel fusion of the first region feature vector, the second region feature vector, and the third region feature vector to generate the user identity verification vector and the emotional state vector; The training process of the facial feature fusion model includes: Obtaining a training image set annotated with a user identity label and an emotion label, wherein each image in the training image set is annotated with eye region coordinates, mouth region coordinates, and overall facial region coordinates; constructing an initial fusion model including three parallel feature extraction branches, wherein the three branches correspond to feature extraction of the first facial region, the second facial region, and the third facial region, respectively; Performing region division on each image in the training image set, extracting training region images corresponding to each branch respectively, and generating region feature vectors of each branch; The regional feature vectors of each branch are input into the fully connected layer for feature concatenation to generate a fused feature vector; Obtaining a first cross entropy loss between the fused feature vector and the user identity label, and a second cross entropy loss between the fused feature vector and the emotion label; Performing back propagation optimization on the parameters of the initial fusion model according to a weighted sum of the first cross entropy loss and the second cross entropy loss until the loss converges; The optimized model parameters are solidified to generate the pre-trained facial feature fusion model.

5. The method according to claim 1, characterized in that The generating of dynamic interactive content corresponding to the user's sight focus based on the screen gaze area coordinates and the target interaction mode includes: Determining the current focus control identifier of the user in the screen interface according to the screen gaze area coordinates; Obtaining historical interaction data associated with the currently focused control, wherein the historical interaction data includes the number of times the user triggers the control, the length of time the user stays there, and associated operation records; Based on the response rule in the target interaction mode, generating a dynamic content update instruction corresponding to the currently focused control; According to the emotional state vector and the historical interaction data, adjusting the element attributes in the dynamic content update instruction, the element attributes including color gradient rate, content switching frequency and interaction feedback intensity; The adjusted dynamic content update instructions are sent to the screen rendering engine to generate an interactive content stream containing three-dimensional visual effects.

6. The method according to claim 5, characterized in that The real-time adjustment of presentation parameters of dynamic interactive content according to the user identity verification vector and the emotional state vector includes: Monitor the user's facial deflection angle and the distance between the user and the screen to generate screen viewing angle correction parameters; Retrieving preference information from a user portrait database according to the user identity verification vector, the preference information including font size preference, color contrast threshold, and animation playback speed range; Inputting the emotional state vector into a pre-trained parameter mapping network to generate an interface brightness adjustment value and a content layout density value that match the current emotion; Determining a final display parameter combination of the dynamic interactive content based on the screen viewing angle correction parameter, the preference setting information and the interface brightness adjustment value; Reorganize the information elements in the interactive content flow according to the content layout density value, and trigger the screen display module to perform a parameter update operation; The training process of the parameter mapping network includes: Collecting feedback data of multiple groups of users on screen interactive content in different emotional states, wherein the feedback data includes records of users actively adjusting interface parameters and physiological signal monitoring data; Constructing an initial mapping network including an emotion feature input layer and a parameter prediction output layer, wherein the dimension of the emotion feature input layer is consistent with the dimension of the emotion state vector; Performing time series alignment on the emotional state vector and the user's physiological signal data to generate a training sample set with time series labels; The training sample set is segmented by a sliding window mechanism to extract the mean value of the emotion feature and the parameter adjustment trend vector of each time period; Using a long short-term memory network to perform sequence modeling on the parameter adjustment trend vector to generate parameter prediction values; Calculating the mean square error loss between the parameter prediction value and the actual adjustment parameter, and optimizing the weight parameters of the initial mapping network using a gradient descent algorithm; When the mean square error loss is lower than a preset threshold, the training is stopped and the network parameters are saved to generate the pre-trained parameter mapping network.

7. The method according to claim 1, characterized in that The step of obtaining a facial image sequence triggered by a target user in a screen interaction scenario includes: Capturing the user's facial original image stream through an embedded camera, and performing frame rate synchronization processing on the original image stream; Detecting the coordinates of the face bounding box in each frame image, and performing image cropping and size normalization processing based on the bounding box coordinates; Perform illumination equalization processing on the normalized image to eliminate the influence of ambient light fluctuation on image quality and obtain a balanced image; A background segmentation algorithm is used to extract the face foreground area in the equalized image and remove invalid frames containing occlusions or blurred areas; The processed valid image frames are arranged in time sequence to generate the face image sequence, and a timestamp and a screen touch event mark are added to each image frame.

8. The method according to claim 7, characterized in that The detecting of the coordinates of the face boundary box in each frame of image includes: Call the pre-trained face detection model to perform initial face area prediction on the current frame image and generate a set of candidate bounding boxes; Calculate the confidence score of each candidate bounding box, and select the candidate boxes whose scores exceed the second threshold for merging; According to the merged candidate frame coordinates, determine the center point coordinates and width-to-height ratio of the final face bounding box; Predict the position of the bounding box of the current frame based on the face motion trajectory of the historical frame, and when the deviation between the predicted position and the actual detected position exceeds a third threshold, trigger the face tracking model to correct the position; The corrected bounding box coordinates are proportionally converted to the screen resolution to generate standardized coordinate data that fits the current screen size.

9. The method according to claim 1, characterized in that: After outputting the adjusted screen interaction signal, the method further includes: Continuously monitoring the user's response behavior to the adjusted interactive content, wherein the response behavior includes eye movement trajectory, facial micro-expression changes, and touch operation frequency; Performing a matching analysis between the response behavior and the preset expected feedback mode to generate an interaction effect evaluation index; When the evaluation index is lower than a fourth threshold, triggering the interaction mode optimization engine to iteratively update the response rule of the target interaction mode; Record the feature data generated during the current interaction and synchronize the updated interaction mode features to the local mode library of all associated devices; Generate personalized interaction logs based on user authentication vectors, and encrypt and store log data in distributed database nodes.

10. A screen interactive system, characterized in that: include: a memory, wherein a computer program is stored in the memory; A processor, used to load the computer program to implement the screen interaction method based on face recognition as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Eye movement interaction method, system and device based on eye movement tracking technology

    CN111949131A

  • Multi-modal face emotion recognition method and device

    CN114399818A

  • System switching method and system of family education learning machine

    CN117666786A

  • Student participation degree analysis method in talent social practice teaching based on VR

    CN118674168A

  • Multi-modal emotion recognition method and psychological intervention system

    CN119763175A

Cited By

  • Holographic projection image intelligent generation method and system applied to stage virtual interaction

    CN120909439A

  • Television terminal education application multi-mode user authentication method and system

    CN120956974A

  • Multi-modal user authentication method and system for television-based education applications

    CN120956974B

  • Computer screen automatic rotation adjusting method and device based on face recognition and angle measurement

    CN121501085A