A screen interaction method and system based on face recognition

By obtaining the facial dynamic behavior features and screen gaze area coordinates in the face image sequence, and combining the pre-trained model to extract the user identity and emotional state vector, real-time dynamic adjustment of the screen interactive content is achieved, solving the problems of delayed emotional judgment and insufficient personalization in the existing technology, and improving interaction efficiency and user experience.

CN120045076BActive Publication Date: 2025-09-19CHINA GUANGSHEN OPTOELECTRONICS (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510512628.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-19
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

Existing screen interaction technology based on face recognition cannot effectively capture the dynamic changes of facial muscles, resulting in delayed emotional judgment. Simple data processing makes it difficult to associate user identity preferences with real-time psychological state. Independent expression classification rules lack coordinated analysis with visual focus and behavioral intentions, resulting in a misalignment between interactive responses and users' real needs, insufficient personalization and low interaction efficiency.

Method used

By obtaining the facial dynamic behavior features and screen gaze area coordinates in the face image sequence, combining the preset interaction mode library to match the target interaction mode, calling the pre-trained facial feature fusion model to extract the user authentication vector and emotional state vector, and based on these vectors, adjusting the presentation parameters of the dynamic interactive content in real time to generate a screen interaction signal that matches the user status.

Benefits of technology

It significantly improves the matching accuracy between screen interactive content and user needs, enhances the flexibility and adaptability of content presentation, accurately identifies the user's true intentions, avoids invalid interactions caused by accidental touches or distractions, and optimizes the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045076B_ABST
    Figure CN120045076B_ABST
Patent Text Reader

Abstract

The present invention provides a screen interaction method and system based on face recognition. The method obtains a sequence of facial images triggered by a target user in a screen interaction scenario, matches a target interaction pattern from a preset interaction pattern library based on facial dynamic behavior characteristics, generates dynamic interaction content corresponding to the user's gaze focus based on the screen gaze area coordinates and the target interaction pattern, invokes a pre-trained facial feature fusion model, performs multi-level feature extraction on the facial area in the facial image sequence, generates a user identity verification vector and an emotional state vector, and adjusts the presentation parameters of the dynamic interaction content in real time based on these vectors, and outputs the adjusted screen interaction signal. The present invention can improve interaction efficiency while optimizing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a screen interaction method and system based on face recognition. Background Art

[0002] Screen interaction technology based on facial recognition aims to achieve intelligent response and personalized adaptation of screen content by analyzing user biometrics. Existing technologies typically use static facial recognition for user authentication or trigger preset interaction commands through expression classification algorithms. For example, these methods rely on matching facial contour geometry to an identity library to load a fixed interface configuration, or switch screen display modes based on basic expression tags. However, these methods have significant flaws: static recognition cannot capture the dynamic changes in facial muscles, resulting in delayed emotional judgment; simple data processing makes it difficult to correlate user identity preferences with real-time psychological states; independent expression classification rules lack coordinated analysis of gaze focus and behavioral intentions, resulting in a misalignment between interactive responses and actual user needs; and fixed content adjustment mechanisms cannot adapt to the dynamic interaction needs in different scenarios, resulting in insufficient personalization and low interaction efficiency. Summary of the Invention

[0003] The present invention provides a screen interaction method and system based on face recognition.

[0004] In the first aspect, an embodiment of the present invention provides a screen interaction method based on face recognition, the method comprising: obtaining a face image sequence triggered by a target user in a screen interaction scene, the face image sequence comprising the user's facial dynamic behavior characteristics and screen gaze area coordinates; matching a target interaction mode from a preset interaction mode library based on the facial dynamic behavior characteristics, the target interaction mode comprising interaction response rules associated with the intensity of the user's expression; generating dynamic interaction content corresponding to the user's visual focus based on the screen gaze area coordinates and the target interaction mode; calling a pre-trained facial feature fusion model to perform multi-level feature extraction on the facial area in the face image sequence to generate a user identity authentication vector and an emotional state vector; adjusting the presentation parameters of the dynamic interaction content in real time based on the user identity authentication vector and the emotional state vector, and outputting the adjusted screen interaction signal.

[0005] In a second aspect, an embodiment of the present invention provides a screen interaction system, comprising: a memory storing a computer program; and a processor for loading the computer program to implement the screen interaction method based on face recognition as described above.

[0006] The screen interaction method based on face recognition provided by the present invention obtains the facial dynamic behavior characteristics and screen gaze area coordinates in the face image sequence triggered by the target user in the screen interaction scene, combines the preset interaction pattern library to match the target interaction pattern associated with the user's expression intensity, generates dynamic interaction content corresponding to the focus of sight, and calls the pre-trained facial feature fusion model to extract the user identity authentication vector and emotional state vector, adjusts the presentation parameters of the dynamic content in real time based on the identity characteristics and emotional characteristics, and finally outputs a screen interaction signal adapted to the user status. In this way, facial dynamic behavior features can capture the user's real-time expression changes and muscle movement trends, the screen gaze area coordinates can accurately locate the user's visual focus, the identity verification vector provides the user's historical preference benchmark for personalized interaction, and the emotional state vector dynamically reflects the user's current psychological state. Through the collaborative analysis and fusion of multi-dimensional data, the matching accuracy of screen interactive content and user needs can be significantly improved; at the same time, dynamic adjustment of the parameters of dynamic interactive content based on real-time emotional and identity characteristics can enhance the flexibility and adaptability of content presentation, and effectively solve the problems of delayed response and insufficient personalization in traditional interactive technologies; in addition, through the joint analysis of facial dynamic behavior and screen gaze area, the user's true intention can be accurately identified, and invalid interaction caused by accidental touch or distraction can be avoided, thereby improving interaction efficiency while optimizing user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0008] Figure 1 This is a flow chart of a screen interaction method based on face recognition provided by an embodiment of the present invention.

[0009] Figure 2 It is a schematic diagram of the composition of a screen interaction system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0010] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0011] See also Figure 1 , Figure 1 A flowchart of a screen interaction method based on face recognition provided by an embodiment of the present invention. The screen interaction method based on face recognition can be executed by a screen interaction system. The screen interaction method based on face recognition may include the following steps:

[0012] Step S100: Acquire a facial image sequence triggered by a target user in a screen interaction scene, where the facial image sequence includes the user's facial dynamic behavior features and screen gaze area coordinates.

[0013] In this embodiment, a facial image sequence refers to a chronological collection of images containing user facial information. Facial dynamic behavioral features refer to changes in a user's facial movements over a period of time, such as changes in expression and muscle movement. Screen gaze area coordinates refer to the location of the screen area where the user's gaze is focused during screen interaction. Acquiring a facial image sequence triggered by a target user in a screen interaction scenario can be accomplished in a variety of ways, such as using an embedded camera to capture the user's facial image.

[0014] Specifically, this step may include the following detailed steps:

[0015] Step S110: Capture the user's facial original image stream through the embedded camera, and perform frame rate synchronization processing on the original image stream.

[0016] An embedded camera is a camera installed on a display device, used to capture the user's facial image in real time. The raw image stream is a series of image data continuously captured by the camera. Frame rate synchronization ensures that the time intervals between image frames are uniform and meet the requirements of subsequent processing. For example, an unstable camera frame rate can cause errors in the subsequent analysis of dynamic facial behavior characteristics. Frame rate synchronization can adjust the frame rate of the raw image stream to a stable value, such as a uniform 30 frames per second. This allows for more accurate capture of the user's facial movements during subsequent analysis.

[0017] Step S120: Detecting the face bounding box coordinates in each frame image, and performing image cropping and size normalization processing based on the bounding box coordinates.

[0018] Face bounding box coordinates refer to the coordinate information of the rectangular frame that encloses the face area. By determining these coordinates, the face's position in the image can be accurately located. Image cropping is performed by extracting the face area from the original image based on the detected face bounding box coordinates, removing irrelevant background information. Size normalization adjusts the cropped face image to a uniform size for subsequent feature extraction and analysis. For example, resizing the cropped face image to a uniform size of 224×224 pixels ensures that different images have the same input size when performing feature extraction, improving the accuracy and consistency of feature extraction.

[0019] Specifically, detecting the coordinates of the face bounding box in each frame image may include the following steps:

[0020] Step S121: Calling a pre-trained face detection model to perform initial face region prediction on the current frame image and generate a set of candidate bounding boxes.

[0021] Pretrained face detection models are trained on a large amount of facial image data and are capable of identifying facial regions within images. For example, a commonly used face detection model is MTCNN (Multi-task Cascaded Convolutional Networks), which can quickly and accurately detect the approximate location of a face within an image and generate multiple candidate bounding boxes. These candidate bounding boxes may contain facial regions of varying sizes and positions, providing a foundation for subsequent screening.

[0022] Step S122: Calculate the confidence score of each candidate bounding box, and select candidate boxes with scores exceeding a second threshold for merging.

[0023] The confidence score is the model's assessment of the likelihood that each candidate bounding box contains a face. The second threshold is a pre-set score used to select candidate bounding boxes with higher confidence scores. For example, if the second threshold is set to 0.8, only candidate bounding boxes with a confidence score exceeding 0.8 will be retained. Merging combines the selected candidate bounding boxes, removing any overlap to obtain a more accurate face bounding box. This prevents multiple overlapping bounding boxes from being identified as faces, improving face localization accuracy.

[0024] Step S123: Determine the center point coordinates and aspect ratio of the final face bounding box based on the merged candidate box coordinates.

[0025] The coordinates of the merged candidate box contain the approximate location of the face region. By calculating these coordinates, we can determine the center coordinates and aspect ratio of the final face bounding box. The center coordinates accurately represent the face's position in the image, while the aspect ratio reflects the facial shape. For example, the center coordinates can be obtained by averaging the coordinates of the top-left and bottom-right corners of the merged candidate box; the aspect ratio can be obtained by calculating the ratio of the candidate box's width to height. This information is crucial for subsequent image cropping and resizing.

[0026] Step S124: predicting the bounding box position of the current frame based on the face motion trajectory of the historical frame, and triggering the face tracking model to perform position correction when the deviation between the predicted position and the actual detected position exceeds a third threshold.

[0027] The facial motion trajectory of a historical frame refers to the positional changes of the face in the previous image frame. By analyzing these trajectories, the possible position of the face in the current frame can be predicted. The third threshold is a pre-set deviation standard used to determine whether the difference between the predicted position and the actual detected position is too large. For example, if the third threshold is set to 10 pixels and the deviation between the predicted position and the actual detected position exceeds 10 pixels, the prediction is inaccurate and the face tracking model needs to be triggered to correct the position. The face tracking model can more accurately track and correct the position of the face based on the information of the current frame and historical frames, ensuring the accurate position of the face bounding box.

[0028] Step S125: Scale the corrected bounding box coordinates with the screen resolution to generate standardized coordinate data that fits the current screen size.

[0029] Screen resolution refers to the pixel dimensions of the screen, and different screen devices may have different resolutions. By scaling the corrected bounding box coordinates with the screen resolution, you can convert them into standardized coordinate data appropriate for the current screen size. For example, if the screen resolution is 1920×1080 pixels, the corrected bounding box coordinates are (x1, y1, x2, y2). By dividing these coordinates by the width and height of the screen, you can obtain standardized coordinate data, allowing you to accurately locate the face area on screens of different resolutions.

[0030] Step S130: performing illumination equalization processing on the normalized image to eliminate the influence of ambient light fluctuation on the image quality and obtain a balanced image.

[0031] Lighting equalization is an image processing technique used to adjust image brightness and contrast, maintaining good image quality under varying lighting conditions. Fluctuations in ambient light can cause overly bright or dark areas in an image, hindering facial feature extraction and analysis. Lighting equalization can eliminate these effects and achieve a more uniform image brightness. For example, histogram equalization can be applied to a normalized image to adjust the grayscale distribution to a more uniform range, resulting in a balanced image. This allows for more accurate identification of facial details during subsequent facial feature extraction.

[0032] Step S140: extracting the face foreground area in the equalized image using a background segmentation algorithm, and removing invalid frames containing occlusions or blurred areas.

[0033] Background segmentation algorithms are used to separate foreground objects (such as faces) from the background in an image. This algorithm can extract the foreground region of the face in the equalized image and remove irrelevant background information. Occlusions or blurred areas may affect the accurate extraction of facial features, so invalid frames containing these areas need to be removed. For example, a deep learning-based background segmentation algorithm, such as a U-Net network, can be used to process the equalized image to separate the face from the background. The image's clarity and occlusion levels are then analyzed to determine if any occlusions or blurred areas exist. If so, the frame is marked as invalid and removed. This ensures that the facial images used for subsequent analysis are clear and unobstructed, improving the accuracy of facial feature extraction.

[0034] Step S150: Arrange the processed valid image frames in chronological order to generate a facial image sequence, and add a timestamp and a screen touch event mark to each frame of the image.

[0035] A processed valid image frame refers to a facial image frame that meets the required requirements after undergoing processes such as cropping, normalization, illumination equalization, and background segmentation. Arranging these image frames in chronological order can generate a facial image sequence containing dynamic facial information about the user. The timestamp records the capture time of each frame, which can be used to understand the temporal sequence and evolution of the user's facial movements. The screen touch event marker records whether the user performed a touch operation on the screen when the frame was captured, as well as information such as the type and location of the operation. For example, when a user clicks a button on the screen, a corresponding screen touch event marker is added to the corresponding image frame. This marker information provides important reference for subsequent interaction analysis.

[0036] Step S200: Matching a target interaction pattern from a preset interaction pattern library according to facial dynamic behavior characteristics, where the target interaction pattern includes interaction response rules associated with the intensity of the user's expression.

[0037] In this embodiment, facial dynamic behavior characteristics refer to the changes in the user's facial movements over a period of time, such as changes in expression, muscle movements, etc. The preset interaction mode library is a database that pre-stores a variety of interaction modes, and each interaction mode includes facial movement trigger conditions and corresponding screen response strategies. The target interaction mode refers to the most suitable interaction mode matched from the interaction mode library based on the user's facial dynamic behavior characteristics. The interaction response rule refers to the way the screen should respond when the user's expression intensity meets the preset conditions. Specifically, matching the target interaction mode from the preset interaction mode library based on the facial dynamic behavior characteristics may include the following steps:

[0038] Step S210: Input the facial dynamic behavior features into the pre-trained spatiotemporal convolutional network to extract facial muscle movement trajectory features and eye opening and closing frequency features.

[0039] A pretrained spatiotemporal convolutional network (STN) is a neural network model trained on a large amount of data and capable of processing sequential data containing temporal and spatial information. Facial muscle movement trajectory features refer to the movement paths and changes of facial muscles over a period of time, reflecting changes in a user's facial expressions. Eye opening and closing frequency features refer to the number of times a user's eyes open and close per unit time, reflecting the user's attention to the screen content. For example, a 3D convolutional neural network (3D CNN) can be used as a pretrained STN, feeding dynamic facial behavior features into the network for processing.

[0040] Specifically, inputting facial dynamic behavior features into a pre-trained spatiotemporal convolutional network to extract facial muscle movement trajectory features and eye opening and closing frequency features may include the following steps:

[0041] Step S211: performing facial key point detection on each frame image in the facial image sequence, and generating a key point distribution map including eyebrow region coordinates, cheek region coordinates, and mouth corner region coordinates.

[0042] Facial landmark detection involves accurately locating key points in a facial image, such as the coordinates of eyebrows, cheeks, and corners of the mouth. By detecting the locations of these key points, a key point distribution map containing these coordinates can be generated. For example, a deep learning-based facial landmark detection algorithm, such as the 68-point facial landmark detection model in the Dlib library, can be used to process each frame in a facial image sequence, obtaining coordinate information for the eyebrows, cheeks, and corners of the mouth in each frame. This coordinate information is then integrated into a key point distribution map. This allows for a visual display of the positional changes of various facial features.

[0043] Step S212: Based on the key point displacement vectors between adjacent frames, the motion direction consistency parameter of each facial region within a preset time window is determined. The motion direction consistency parameter is used to characterize the stability of the coordinated movement of muscle groups.

[0044] The key point displacement vector between adjacent frames refers to the position change vector of the same key point in two adjacent image frames. By analyzing these displacement vectors, the motion direction consistency parameter of each facial region within a preset time window can be determined. The preset time window is a pre-set time period used to analyze the movement of the facial region. For example, the preset time window is set to a time period of 10 image frames. The displacement vectors of each key point in the facial region within this time period are calculated, and then the direction of these displacement vectors is statistically analyzed to determine whether they are consistent. If the majority of the displacement vectors have the same direction, it indicates that the coordinated movement of the muscle groups in that facial region is relatively stable, and the motion direction consistency parameter is high. Conversely, it indicates that the coordinated movement of the muscle groups is unstable, and the motion direction consistency parameter is low. This parameter can help determine whether the user's expression changes are natural and stable.

[0045] Step S213: performing optical flow analysis on the key point distribution map to generate a dynamic change curve reflecting the muscle contraction intensity, and extracting the periodic characteristics of facial muscle movement based on the frequency of occurrence of peaks and troughs in the dynamic change curve.

[0046] Optical flow analysis is a technique used to analyze the motion of objects in images. By performing optical flow analysis on a keypoint distribution map, we can obtain information about the speed and direction of facial muscle movement. Based on this information, we can generate a dynamic change curve reflecting the strength of muscle contraction. The peaks and troughs in the dynamic change curve represent the maximum and minimum values ​​of muscle contraction intensity. By counting the frequency of these peaks and troughs, we can extract the periodic characteristics of facial muscle movement. For example, we use the Lucas-Kanade optical flow algorithm to analyze the keypoint distribution map, calculate the optical flow vector for each keypoint, and then generate a dynamic change curve based on the magnitude and direction of these optical flow vectors. A high frequency of peaks and troughs in the dynamic change curve indicates strong periodicity in facial muscle movement; a low frequency indicates weak periodicity. These periodic characteristics can reflect the changing patterns of a user's facial expressions.

[0047] Step S214: performing block-by-block grayscale value comparison on the image blocks of the eye area in the facial image sequence, calculating the pixel area change rate of the eyelid covering the pupil area per unit time, and generating an eye opening and closing state waveform with time as the horizontal axis.

[0048] The eye region image block refers to the image region containing eye information extracted from a facial image sequence. Block-wise grayscale value comparison involves dividing the eye region image block into multiple smaller blocks and then comparing the grayscale value changes within each block. By calculating the rate of change in the pixel area of ​​the eyelid covering the pupil per unit time, the eye opening and closing state can be understood. A time-based eye opening and closing waveform graph can visually demonstrate how the eye opening and closing state changes over time. For example, the eye region image block can be divided into 8×8 blocks, and the grayscale value changes of each block between adjacent frames are compared. When the grayscale value change exceeds a preset threshold, the block's state is considered to have changed. The rate of change in the pixel area of ​​the eyelid covering the pupil is then calculated and plotted chronologically on a waveform graph to produce an eye opening and closing waveform graph. This waveform graph can help analyze a user's eye movements and attention levels.

[0049] Step S215: Calculate the synergy index between muscle movement and eye movement based on the periodic feature and the phase difference between the eye opening and closing state waveform, and perform weighted fusion of the synergy index and the movement direction consistency parameter to generate facial muscle movement trajectory features.

[0050] Periodicity refers to the periodic pattern of facial muscle movement, while the eye opening and closing waveform reflects how the eye opening and closing state changes over time. Phase difference refers to the time offset between the periodicity and the eye opening and closing waveform. By calculating the phase difference, we can understand the degree of coordination between muscle and eye movements. The synergy index is a numerical value used to measure the degree of coordination between muscle and eye movements. Weighted fusion of the synergy index with the motion direction consistency parameter comprehensively considers the stability of muscle group coordinated movement and the degree of coordination between muscle movement and eye movement, generating a facial muscle movement trajectory feature. For example, the synergy index is calculated based on the periodicity and the phase difference of the eye opening and closing waveform. The synergy index and the motion direction consistency parameter are assigned different weights, and then the weighted sum of these two factors is used to generate the facial muscle movement trajectory feature. This feature can more comprehensively reflect the user's dynamic facial behavior.

[0051] Step S216: Perform frequency domain transformation on the eye opening and closing state waveform, extract valid waveform segments whose amplitude exceeds a preset noise threshold, and generate eye opening and closing frequency features based on the interval length and duration ratio of the valid waveform segments.

[0052] Frequency domain transformation is a technique that converts time-domain signals into frequency-domain signals. By performing frequency domain transformation on the eye opening / closing waveform, the waveform can be converted from the time domain to the frequency domain, allowing its frequency components to be analyzed. The preset noise threshold is a pre-set amplitude standard used to filter out valid waveform segments. A valid waveform segment is a waveform segment whose amplitude exceeds the preset noise threshold. Interval duration refers to the time interval between adjacent valid waveform segments, and duration ratio refers to the ratio of the duration of a valid waveform segment to the interval duration. By analyzing the interval duration and duration ratio of valid waveform segments, an eye opening / closing frequency signature can be generated. For example, a fast Fourier transform (FFT) is used to perform frequency domain transformation on the eye opening / closing waveform, converting the waveform into a frequency domain signal. A preset noise threshold is then set to filter out valid waveform segments whose amplitude exceeds the threshold. The interval duration and duration ratio of valid waveform segments are calculated and integrated into the eye opening / closing frequency signature. This signature accurately reflects the user's eye opening / closing frequency.

[0053] Step S220: Determine the user's concentration level on the screen content based on the eye opening and closing frequency characteristics, and generate a user intention prediction vector in combination with the muscle movement trajectory characteristics.

[0054] In this embodiment, the eye opening and closing frequency feature can reflect the activity of the user's eyes. Generally, a low eye opening and closing frequency indicates that the user is more focused on the screen content, while a high frequency may indicate that the user is distracted. The concentration level is a quantitative classification of the user's concentration level based on the eye opening and closing frequency feature, for example, it can be divided into three levels: high, medium, and low. The muscle movement trajectory feature can reflect the user's expression changes and potential intentions. The user intention prediction vector is a vector containing information in multiple dimensions, which is used to predict the user's intention to interact with the screen. Specifically, based on the eye opening and closing frequency feature, the user's concentration level on the screen content is determined, and combined with the muscle movement trajectory feature, generating the user intention prediction vector can include the following steps:

[0055] Step S221: Determine the cumulative time the user's gaze is away from the screen per unit time based on the effective waveform segment interval length in the eye opening and closing frequency characteristics, and generate an initial concentration score based on the duration ratio of the effective waveform segments.

[0056] The effective waveform segment interval duration refers to the time interval between adjacent effective waveform segments and can reflect the amount of time the user's gaze is away from the screen. By summing the effective waveform segment interval durations per unit time, the cumulative amount of time the user's gaze is away from the screen per unit time can be determined. The effective waveform segment duration ratio refers to the ratio of the effective waveform segment duration to the interval duration and can reflect the stability of the user's eye opening and closing. Combining these two factors, an initial concentration score can be generated. For example, if the cumulative amount of time the user's gaze is away from the screen per unit time is short and the effective waveform segment duration is long, it indicates that the user's attention is relatively focused, resulting in a higher initial concentration score; conversely, the score is lower. This provides a preliminary assessment of the user's level of focus on the screen content.

[0057] Step S222: Detect the displacement direction and amplitude of the coordinates of the corner of the mouth area in the muscle movement trajectory characteristics. When a continuous upward movement of the corner of the mouth is detected, trigger the concentration compensation mechanism to make a positive correction to the initial concentration score.

[0058] The direction and magnitude of the displacement of the mouth corner coordinates can reflect changes in the user's facial expressions. Continuous upward movement of the mouth corners typically indicates a state of happiness or concentration. The concentration compensation mechanism is a strategy used to adjust the initial concentration score. When continuous upward movement of the mouth corners is detected, a positive correction is made to the initial concentration score to improve the accuracy of the score. For example, if upward movement of the mouth corners is detected in three consecutive image frames, the concentration compensation mechanism is triggered, increasing the initial concentration score by one level. This more accurately reflects the user's actual level of concentration.

[0059] Step S223: Mapping the corrected concentration score to a preset level division interval to determine the concentration level.

[0060] The preset grading intervals are pre-set ranges for categorizing concentration levels. For example, the concentration level can be divided into three levels: high, medium, and low, with corresponding scoring intervals of [80, 100], [50, 79], and [0, 49], respectively. By mapping the revised concentration score to these grading intervals, the user's concentration level can be determined. For example, if the revised concentration score is 85, it is mapped to the high-level interval, and the user's concentration level is determined to be high. This allows for an intuitive understanding of the user's level of focus on the screen content.

[0061] Step S224: performing trend prediction on the eyebrow region movement direction consistency parameter in the muscle movement trajectory feature to generate a muscle activity intensity vector reflecting the user's potential attention.

[0062] The eyebrow region's motion direction consistency parameter reflects the stability of the coordinated movement of the eyebrow muscle groups. Trend prediction of this parameter can reveal the changing trends of muscle activity in the user's eyebrow region. The muscle activity intensity vector, a vector containing information from multiple dimensions, reflects the user's potential attention. For example, a time series analysis method is used to perform trend prediction on the eyebrow region's motion direction consistency parameter, and a muscle activity intensity vector is generated based on the prediction results. If the prediction results show an upward trend in the parameter, it indicates that the user's potential attention is increasing, and the corresponding dimension value in the muscle activity intensity vector will also increase; conversely, it indicates that potential attention is decreasing. This vector provides important reference information for subsequent user intent prediction.

[0063] Step S225: Input the concentration level and muscle activity intensity vector into the pre-trained long short-term memory network to perform time series feature matching and extract the attention transfer pattern that is strongly related to the screen interaction intention.

[0064] A pre-trained long short-term memory (LSTM) network is a neural network model capable of processing sequential data and capturing long-term dependencies within it. By inputting concentration levels and muscle activity intensity vectors into the network and performing time series feature matching, attention shift patterns strongly correlated with screen interaction intent can be identified. Attention shift patterns refer to the patterns of changes in a user's attention during screen interaction. For example, when a user's concentration level changes and their muscle activity intensity vector also changes accordingly, the LSTM network can learn this pattern and identify it as an attention shift pattern strongly associated with screen interaction intent. These patterns can help predict the user's next interaction intent.

[0065] Step S226: According to the movement priorities of different facial areas in the attention transfer pattern, the muscle movement trajectory features are weighted with intention relevance to generate a user intention prediction vector containing multi-dimensional weight coefficients.

[0066] The motion priorities of different facial regions in the attention shift pattern refer to the degree of influence of the motion of different facial regions on the interaction intent during the attention shift process. Based on these priorities, the muscle motion trajectory features are weighted by the intention relevance, which can highlight the characteristic information of the facial regions related to the interaction intent. The multidimensional weight coefficient refers to the different weight values ​​assigned to each dimension of the muscle motion trajectory features. Through weighted processing, a user intention prediction vector containing multidimensional weight coefficients is generated. For example, if the attention shift pattern shows that the movement of the eyebrow region has a greater impact on the interaction intent, the muscle motion trajectory features of the eyebrow region are assigned a higher weight value during the weighting process. The user intention prediction vector generated in this way can more accurately reflect the user's interaction intention.

[0067] Step S230: performing similarity matching between the user intention prediction vector and the pattern features in the interaction pattern library, and screening out a set of candidate interaction patterns whose similarity exceeds a first threshold.

[0068] In this embodiment, the user intention prediction vector is a vector containing multi-dimensional weight coefficients, which is used to predict the user's intention to interact with the screen. The pattern feature in the interactive pattern library refers to the feature vector corresponding to each interactive pattern, which contains information such as the facial action triggering conditions and screen response strategy of the interactive pattern. Similarity matching refers to calculating the similarity between the user intention prediction vector and the pattern feature. Commonly used similarity calculation methods include cosine similarity, Euclidean distance, etc. The first threshold is a pre-set similarity standard used to screen out interactive patterns with a high similarity to the user intention prediction vector. For example, cosine similarity is used to calculate the similarity between the user intention prediction vector and each pattern feature in the interactive pattern library, and the interactive patterns with a similarity of more than 0.8 are screened out to form a set of candidate interactive patterns. In this way, an interactive pattern that is more closely matched with the user's current intention can be found.

[0069] Step S240: Based on the concentration level, the weight of each candidate interaction mode in the candidate interaction mode set is modified, and the candidate interaction mode with the highest weight is selected as the target interaction mode; wherein, each interaction mode in the interaction mode library is associated with at least one facial action trigger condition and a corresponding screen response strategy.

[0070] The concentration level reflects the user's level of attention to the screen content. Different concentration levels may require different interaction modes to meet the user's needs. Weighting each candidate interaction mode in the candidate interaction mode set based on the concentration level means assigning different weights to each candidate interaction mode based on the concentration level. For example, when a user's concentration level is high, they may be more inclined to choose complex and interesting interaction modes, so these interaction modes are assigned a higher weight; when the concentration level is low, simple and direct interaction modes may be more suitable, so they are assigned a higher weight. The candidate interaction mode with the highest weight is selected as the target interaction mode. This ensures that the selected interaction mode best meets the user's current state and needs. Each interaction mode in the interaction mode library is associated with at least one facial action trigger condition and a corresponding screen response strategy. For example, when the user makes a smiling expression (facial action trigger condition), a cute animation can pop up on the screen (screen response strategy).

[0071] Step S300: generating dynamic interactive content corresponding to the user's visual focus based on the screen gaze area coordinates and the target interaction mode.

[0072] In this embodiment, the screen gaze area coordinates refer to the location of the screen area where the user's gaze is focused during screen interaction. The target interaction mode is the most suitable interaction mode selected from a preset interaction mode library based on the user's facial dynamic behavior characteristics. Dynamic interaction content refers to interactive content with dynamic effects generated based on the user's gaze focus and the target interaction mode.

[0073] Specifically, based on the screen gaze area coordinates and the target interaction mode, generating dynamic interactive content corresponding to the user's gaze focus may include the following steps:

[0074] Step S310: Determine the identifier of the control currently focused on by the user in the screen interface according to the coordinates of the screen gaze area.

[0075] Controls in a screen interface refer to various interactive elements on the screen, such as buttons, text boxes, and icons. Using the coordinates of the screen gaze area, we can determine the location of the control on which the user's gaze is currently focused, and thus the identity of that control. For example, if the coordinates of the screen gaze area fall within the area of ​​a button, the identity of the control currently focused on by the user can be determined to be that button's identity. This allows us to clearly identify the interactive element the user is currently focusing on.

[0076] Step S320: Obtain historical interaction data associated with the currently focused control, where the historical interaction data includes the number of times the user triggers the control, the duration of the stay, and associated operation records.

[0077] Historical interaction data refers to data generated by past user interactions with the currently focused control. Trigger count refers to the number of times a user clicks or operates the control, dwell time refers to the length of time a user's gaze remains on the control, and associated operation records refer to records of other related operations performed by the user while operating the control. For example, if the currently focused control is a shopping cart button, historical interaction data may include the number of times the user clicked the button, the duration of time spent on the shopping cart page after each click, and records of operations such as adding and removing items from the shopping cart. This historical interaction data can reflect the user's usage habits and preferences for the control.

[0078] Step S330: Based on the response rules in the target interaction mode, a dynamic content update instruction corresponding to the currently focused control is generated.

[0079] The response rules in the target interaction mode refer to how the screen should respond when the intensity of the user's expression meets the preset conditions. Based on these response rules, combined with the information of the currently focused control, dynamic content update instructions corresponding to the control can be generated. For example, if the target interaction mode stipulates that when the user smiles, the color of the currently focused button changes to green and a prompt message is displayed, the dynamic content update instructions generated according to the rule may include operations such as setting the button color to green and displaying a prompt message. In this way, the content of the currently focused control can be dynamically updated according to the user's state and the target interaction mode.

[0080] Step S340: adjusting the element attributes in the dynamic content update instruction according to the emotional state vector and the historical interaction data, the element attributes including the color gradient rate, the content switching frequency and the interaction feedback intensity.

[0081] The emotional state vector is obtained by performing multi-level feature extraction on a sequence of facial images, and it can reflect the user's current emotional state. Historical interaction data contains the user's usage habits and preferences for the currently focused control. Based on these two factors, the element attributes in the dynamic content update instruction are adjusted to make the dynamic interactive content more in line with the user's needs and emotional state. The color gradient rate refers to the speed of color change, the content switching frequency refers to the time interval between content updates, and the interactive feedback intensity refers to the intensity of the feedback effect obtained when the user operates the control. For example, if the emotional state vector shows that the user is in an excited state and the historical interaction data shows that the user likes fast content switching, the content switching frequency in the dynamic content update instruction can be increased; if the user's previous operating habit is to like stronger interactive feedback, the interactive feedback intensity can be increased.

[0082] Specifically, adjusting the element attributes in the dynamic content update instruction according to the emotional state vector and the historical interaction data may include the following steps:

[0083] Step S341: Dynamically weight the emotion dimension parameters in the emotion state vector and the user operation preference parameters in the historical interaction data to generate emotion-dominant adjustment parameters and history-dominant adjustment parameters.

[0084] The emotional dimension parameters in the emotional state vector include excitement, anxiety, and pleasure, which can reflect the user's current emotional state. User operation preference parameters in historical interaction data include color sensitivity, browsing speed, touch strength, etc. These parameters can reflect the user's usage habits and preferences. Dynamic weight allocation refers to assigning different weight values ​​to emotional dimension parameters and user operation preference parameters according to different situations. For example, when the user's emotions fluctuate significantly, the emotional dimension parameters are assigned a higher weight; when the user's operation habits are relatively stable, the user operation preference parameters are assigned a higher weight. Through dynamic weight allocation, emotion-driven adjustment parameters and history-driven adjustment parameters are generated. Emotion-driven adjustment parameters mainly consider the user's emotional state, while history-driven adjustment parameters mainly consider the user's historical operation habits.

[0085] Step S342: According to the attribute type of the element to be adjusted in the dynamic content update instruction, attribute association mapping is performed on the emotion-dominated adjustment parameters and the history-dominated adjustment parameters, wherein the color gradient rate is associated with the excitement parameter in the emotion dimension and the color sensitivity parameter in the user operation preference, the content switching frequency is associated with the anxiety parameter in the emotion dimension and the browsing speed parameter in the user operation preference, and the interactive feedback intensity is associated with the pleasure parameter in the emotion dimension and the touch force parameter in the user operation preference.

[0086] Attribute association mapping refers to the corresponding association of emotion-driven adjustment parameters and history-driven adjustment parameters with the element attribute types to be adjusted in the dynamic content update instruction. For example, the color gradient rate is associated with the excitement parameter in the emotion dimension and the color sensitivity parameter in the user operation preference. When the excitement is high and the color sensitivity is strong, the color gradient rate can be increased; the content switching frequency is associated with the anxiety parameter in the emotion dimension and the browsing speed parameter in the user operation preference. When the anxiety is high and the browsing speed is fast, the content switching frequency can be increased; the interactive feedback intensity is associated with the pleasure parameter in the emotion dimension and the touch force parameter in the user operation preference. When the pleasure is high and the touch force is large, the interactive feedback intensity can be enhanced. Through this association mapping, element attributes can be adjusted more accurately according to the user's emotional state and operation preferences.

[0087] Step S343: Call the preset dynamic balance rules to prioritize the parameters after attribute association mapping. When the change amplitude of the emotion dimension parameters exceeds the stability threshold of the historical operation parameters, the emotion-dominant adjustment parameters are preferentially used to make instantaneous adjustments to the element attributes. Otherwise, the element attributes are gradually adjusted based on the history-dominant adjustment parameters.

[0088] The preset dynamic balancing rules are pre-defined rules for determining parameter priority. The parameters after attribute association mapping include emotion-driven adjustment parameters and history-driven adjustment parameters. Prioritization involves sorting these parameters according to the dynamic balancing rules to determine which parameter has higher priority when adjusting element attributes. The stability threshold of the historical operation parameters is a pre-defined criterion used to determine the stability of historical operation parameters. When the change in the emotion dimension parameters exceeds the stability threshold of the historical operation parameters, indicating a significant change in user emotion, the emotion-driven adjustment parameters are prioritized for instantaneous adjustment of element attributes to quickly respond to the user's emotional changes. Otherwise, the history-driven adjustment parameters are used for gradual adjustments to maintain consistency with the user's historical operating habits. For example, if the excitement parameter in the emotion dimension suddenly increases and exceeds the stability threshold of the historical operation parameters, the emotion-driven adjustment parameters are immediately used to increase the color gradient rate. If the emotion change is minor, the color gradient rate is gradually adjusted based on the history-driven adjustment parameters.

[0089] Step S344: Generate an element attribute adjustment instruction set including an adjustment range and an effective timing according to the priority sorting result, and verify whether each parameter in the adjustment instruction set exceeds the safe execution range of the screen rendering engine.

[0090] The priority sorting results determine the parameters that should be prioritized when adjusting element properties. Based on this result, a set of element property adjustment instructions is generated, including the adjustment range and effective sequence. The adjustment range refers to the degree to which the element property needs to be adjusted, and the effective sequence refers to the time sequence in which the adjustment instructions take effect. For example, if the priority sorting results indicate that the color gradient rate should be adjusted using emotion-driven adjustment parameters, with a 50% increase and an immediate effective sequence, the generated element property adjustment instruction set should include this information. Verifying that the parameters in the adjustment instruction set exceed the safe execution range of the screen rendering engine ensures that the adjustment operation does not adversely affect the screen display. For example, if the adjusted color gradient rate is too fast, it may cause screen flickering. In this case, the adjustment range needs to be appropriately adjusted to keep it within the safe execution range.

[0091] Step S345: The verified adjustment instruction set is integrated with the dynamic content update instruction at the instruction layer, the effectiveness logic of the element attributes is updated in an overlay manner, a content rendering strategy adapted to the current user status is generated, and the screen interaction module is triggered to load the updated interactive content stream.

[0092] Instruction-layer fusion refers to merging the verified adjustment instruction set with the dynamic content update instructions into a unified instruction set. Overwriting and updating the element attribute validation logic refers to replacing the original validation logic with the adjusted validation logic to ensure that the element attributes are updated according to the new adjustment requirements. Generating a content rendering strategy that adapts to the current user state involves developing a content rendering solution that best suits the current user based on the user's emotional state, operational preferences, and the adjusted instruction set. Triggering the screen interaction module to load the updated interactive content stream refers to sending the generated content rendering strategy to the screen interaction module, causing it to load and display the updated interactive content. For example, information such as the adjusted color gradient rate, content switching frequency, and interactive feedback intensity can be incorporated into the dynamic content update instructions, updating the element attribute validation logic. A new content rendering strategy is then generated, and finally, the screen interaction module is instructed to load and display the interactive content generated based on this strategy.

[0093] Step S350: Send the adjusted dynamic content update instruction to the screen rendering engine to generate an interactive content stream including a three-dimensional visual effect.

[0094] The screen rendering engine is a software module that converts dynamic content update instructions into actual on-screen display content. Adjusted dynamic content update instructions are sent to the screen rendering engine, which generates an interactive content stream with 3D visual effects based on these instructions. 3D visual effects enhance the three-dimensionality and realism of interactive content, improving the user's interactive experience. For example, based on adjusted parameters such as the color gradient rate, content switching frequency, and interactive feedback intensity, the screen rendering engine can generate a 3D interactive scene with gradient colors, dynamically switching content, and strong interactive feedback. This scene is then output as an interactive content stream for display by the screen display module.

[0095] Step S400: calling a pre-trained facial feature fusion model to perform multi-level feature extraction on the facial area in the face image sequence to generate a user identity verification vector and an emotional state vector.

[0096] In this embodiment, the pre-trained facial feature fusion model is a model obtained by training a large amount of facial image data, and can extract multi-level feature information from facial images. A facial image sequence is a collection of images containing user facial information arranged in chronological order. Multi-level feature extraction refers to extracting features from facial areas in a facial image sequence from different levels and angles to obtain more comprehensive and accurate feature information. The user identity verification vector is a feature vector used to verify the identity of the user, and the emotional state vector is a feature vector used to reflect the user's current emotional state. Specifically, calling the pre-trained facial feature fusion model and performing multi-level feature extraction on facial areas in a facial image sequence may include the following steps:

[0097] Step S410: Divide each frame image in the facial image sequence into a first facial region, a second facial region, and a third facial region, wherein the first facial region includes eye contour coordinates, the second facial region includes mouth contour coordinates, and the third facial region includes overall facial contour coordinates.

[0098] By dividing each frame in a facial image sequence into regions, we can more specifically extract features from different facial regions. The first facial region contains the coordinates of the eye contour. The eyes are a key part of the face that expresses emotion and attention. By extracting features from this region, we can understand the user's eye expression and attention. The second facial region contains the coordinates of the mouth contour. Mouth movements and expressions can reflect the user's emotions and verbal expression intent. The third facial region contains the coordinates of the overall facial contour, which provides information on the overall shape and structure of the face. For example, using a deep learning-based facial landmark detection algorithm, we can detect the coordinates of the eye, mouth, and overall facial contours in a facial image. These coordinates are then used to divide the image into three regions. This allows for the extraction and analysis of features in different regions.

[0099] Step S420: performing local texture analysis on the first facial region through the first feature extraction branch of the facial feature fusion model to generate a first region feature vector containing the pupil movement trajectory.

[0100] The first feature extraction branch of the facial feature fusion model is a sub-model specifically designed to process the first facial region. Local texture analysis involves extracting and analyzing local texture features of the first facial region, such as texture direction and density. Pupil movement trajectory refers to the changes in pupil position over time, which can reflect the user's gaze direction and focus of attention. By performing local texture analysis on the first facial region, a first-region feature vector containing the pupil movement trajectory is generated. For example, a convolutional neural network (CNN) is used as the first feature extraction branch to perform a convolution operation on the first facial region to extract local texture features. Simultaneously, by tracking pupil position changes, pupil movement trajectory information is incorporated into the first-region feature vector. This feature vector provides important eye-related information for subsequent user authentication and emotional state analysis.

[0101] Step S430: Dynamic deformation monitoring is performed on the second facial region through the second feature extraction branch of the facial feature fusion model to generate a second region feature vector including the lip opening and closing amplitude.

[0102] The second feature extraction branch of the facial feature fusion model is a sub-model used to process the second facial region. Dynamic deformation monitoring refers to monitoring dynamic changes in the second facial region, such as the opening and closing of the mouth, and the upward or downward movement of the mouth corners. The lip opening and closing amplitude refers to the maximum distance the lips move during the opening and closing process, which can reflect the user's facial expressions and language expression. By dynamically monitoring the second facial region, a second region feature vector containing the lip opening and closing amplitude is generated. For example, a dynamic deformation monitoring algorithm based on optical flow is used to process the second facial region and calculate the lip opening and closing amplitude. The lip opening and closing amplitude and other relevant dynamic deformation information are then integrated into the second region feature vector. This feature vector can provide important mouth-related information for emotional state analysis and user authentication.

[0103] Step S440: performing global illumination compensation processing on the third facial region through the third feature extraction branch of the facial feature fusion model to generate a third region feature vector containing a skin color change trend.

[0104] The third feature extraction branch of the facial feature fusion model is a sub-model used to process the third facial region. Global illumination compensation (GIC) adjusts the illumination of the third facial region to eliminate the impact of uneven illumination on image quality. Skin color trend refers to changes in skin color over time, which can reflect the user's health and emotional state. By performing GIC on the third facial region, a feature vector for the third region is generated that contains the skin color trend. For example, histogram equalization and adaptive GI compensation algorithms are used to adjust the illumination of the third facial region to achieve more uniform image brightness. Then, by analyzing the image's color features, skin color trend information is extracted and incorporated into the third region feature vector. This feature vector provides important overall facial information for user authentication and emotional state analysis.

[0105] Step S450: cross-channel fusion of the first region feature vector, the second region feature vector, and the third region feature vector to generate a user identity verification vector and an emotional state vector.

[0106] Cross-channel fusion refers to the merging and integration of feature vectors from different regions to obtain more comprehensive and integrated feature information. Cross-channel fusion of the feature vectors from the first, second, and third regions can fully utilize the feature information from different facial regions to generate more accurate user authentication vectors and emotional state vectors. For example, a fully connected layer is used to concatenate the feature vectors from the three regions, which are then mapped to a new feature space through a nonlinear transformation to generate a user authentication vector and emotional state vector. The user authentication vector can be used to verify the user's identity, ensuring that only authorized users can perform interactive operations; the emotional state vector can reflect the user's current mood, providing a basis for subsequent adjustments to interactive content.

[0107] In an embodiment of the present invention, the training process of the facial feature fusion model may include the following steps:

[0108] Step S401: obtaining a training image set annotated with a user identity label and an emotion label, wherein each image in the training image set is annotated with eye region coordinates, mouth region coordinates, and overall facial region coordinates.

[0109] The training image set is a collection of image data used to train the facial feature fusion model. A training image set labeled with user identity and emotion labels means that each image is annotated with the corresponding user identity information and emotional state information. Eye region coordinates, mouth region coordinates, and overall facial region coordinates are used to determine the locations of different facial regions within an image. For example, a large number of facial images are collected and each image is annotated with the user's identity identifier (such as name, number, etc.) and emotion label (such as happy, sad, angry, etc.). At the same time, a facial landmark detection algorithm is used to annotate the coordinates of the eye, mouth, and overall facial region. This annotation information provides supervision signals for model training, enabling the model to learn the relationship between different facial region features and the user's identity and emotional state.

[0110] Step S402: constructing an initial fusion model including three parallel feature extraction branches, wherein the three branches correspond to feature extraction of the first facial region, the second facial region, and the third facial region, respectively.

[0111] The initial fusion model, the initial version of the facial feature fusion model, consists of three parallel feature extraction branches. Each branch is responsible for extracting features from the first, second, and third facial regions. For example, a convolutional neural network (CNN) is used to construct each feature extraction branch, each with a different convolutional and pooling layer structure to adapt to the characteristics of different facial regions. The three branches operate in parallel, allowing for simultaneous feature extraction from different facial regions, improving model processing efficiency.

[0112] Step S403: performing region division on each image in the training image set, extracting the training region image corresponding to each branch, and generating a region feature vector for each branch.

[0113] Each image in the training image set is divided into three regions based on the previously labeled eye region coordinates, mouth region coordinates, and overall facial region coordinates. The training region images corresponding to each branch are extracted separately by inputting the divided region images into the corresponding feature extraction branches for processing. The region feature vectors of each branch are generated by extracting features from the training region images through the feature extraction branches to obtain the feature vectors of each branch. For example, each image in the training image set is divided into a first facial region, a second facial region, and a third facial region according to the labeled coordinates. The images of each region are then input into the corresponding feature extraction branches, and features are extracted through convolution and pooling operations to generate the region feature vectors of each branch.

[0114] Step S404: Input the regional feature vectors of each branch into the fully connected layer for feature splicing to generate a fused feature vector.

[0115] A fully connected layer is a neural network layer that connects all input neurons to all output neurons. The regional feature vectors from each branch are fed into the fully connected layer for feature concatenation, which combines the feature vectors from the three branches into a single vector to generate a fused feature vector. For example, the feature vectors for the first, second, and third regions are arranged in order and fed into the fully connected layer. These are then concatenated into a longer vector through a linear transformation to generate a fused feature vector. This fused feature vector contains comprehensive feature information from different facial regions.

[0116] Step S405: Obtain a first cross entropy loss between the fused feature vector and the user identity label, and a second cross entropy loss between the fused feature vector and the emotion label.

[0117] Cross-entropy loss is a loss function used to measure the difference between the model's predictions and the true labels. The first cross-entropy loss, obtained between the fused feature vector and the user's identity label, calculates the degree of difference between the user's identity predicted by the fused feature vector and the true user's identity label. The second cross-entropy loss, obtained between the fused feature vector and the emotion label, calculates the degree of difference between the emotional state predicted by the fused feature vector and the true emotion label. For example, the fused feature vector can be converted into a probability distribution using the softmax function, and then the cross-entropy loss is calculated with the true user identity label and emotion label. These two loss values ​​can reflect the model's accuracy in user identity recognition and emotional state prediction.

[0118] Step S406: Perform back propagation optimization on the parameters of the initial fusion model according to the weighted sum of the first cross entropy loss and the second cross entropy loss until the loss converges.

[0119] Backpropagation optimization is an optimization method that calculates the gradient of a loss function with respect to model parameters and updates the model parameters based on the gradient. Backpropagation optimization is performed on the parameters of the initial fusion model based on the weighted sum of the first and second cross-entropy losses. These two losses are summed according to preset weights to obtain a total loss. The model parameters are then updated using the backpropagation algorithm to continuously reduce the total loss. For example, different weights can be assigned to the first and second cross-entropy losses, such as 0.6 and 0.4, and the total loss is summed. The model parameters are then updated using the stochastic gradient descent (SGD) algorithm based on the total loss, iterating continuously until the loss converges to a smaller value. This results in a model with good performance in both user identity recognition and emotional state prediction.

[0120] Step S407: solidify the optimized model parameters to generate a pre-trained facial feature fusion model.

[0121] Solidifying the optimized model parameters means preserving the model parameters optimized through backpropagation so that they no longer change. Generating a pretrained facial feature fusion model involves applying the solidified model parameters to the initial fusion model, resulting in a pretrained model that can be directly used for feature extraction. For example, the optimized model parameters can be saved as a file. When using the facial feature fusion model, this file can be loaded and the parameters assigned to the initial fusion model, creating a pretrained model. This pretrained model can quickly and accurately extract feature information from facial images in real-world applications.

[0122] Step S500: adjusting the presentation parameters of the dynamic interactive content in real time according to the user identity authentication vector and the emotional state vector, and outputting the adjusted screen interaction signal.

[0123] In this embodiment, the user authentication vector is used to verify the user's identity. Different users may have different preferences and needs. The emotional state vector reflects the user's current emotional state. Emotional changes may affect the user's perception of and needs for interactive content. The presentation parameters of dynamic interactive content include interface brightness, content layout density, font size, color contrast, etc. These parameters affect the display effect of interactive content and user experience. Real-time adjustment refers to the timely adjustment of presentation parameters based on changes in the user authentication vector and emotional state vector. The adjusted screen interaction signal refers to the interactive content signal after the presentation parameters are adjusted, which can make the screen display interactive content more consistent with the user's current state.

[0124] Specifically, adjusting the presentation parameters of dynamic interactive content in real time based on the user identity verification vector and the emotional state vector may include the following steps:

[0125] Step S510: monitoring the user's facial deflection angle and the distance parameters between the user and the screen, and generating screen viewing angle correction parameters.

[0126] Facial deflection angle refers to the tilt of the user's face relative to the screen, which affects the user's visual experience of on-screen content. The distance parameter to the screen refers to the actual distance between the user and the screen. Different distances may require different display parameters. By monitoring the user's facial deflection angle and the distance parameter to the screen, screen viewing angle correction parameters can be generated. For example, a depth camera or infrared sensor can be used to monitor the user's facial deflection angle and the distance to the screen, and the screen viewing angle correction parameters are calculated based on this data. If the user's facial deflection angle is large, the display direction or viewing angle of the screen content may need to be adjusted; if the user is far away from the screen, the font size and icon size may need to be increased. These correction parameters ensure that the user can clearly see the screen content at different viewing angles and distances.

[0127] Step S520: According to the user identity verification vector, preference setting information is retrieved from the user portrait database, where the preference setting information includes font size preference, color contrast threshold, and animation playback speed range.

[0128] The user portrait database is a database that stores various user preference information. Based on the user identity verification vector, the preference setting information corresponding to the current user can be retrieved from the database. Font size preference refers to the font size that the user likes, the color contrast threshold refers to the color contrast range that the user can accept, and the animation playback speed range refers to the animation playback speed range that the user likes. For example, when the user identity verification vector is verified, the corresponding preference setting information is searched from the user portrait database based on the user identity information in the vector. If the user has previously set a preference for larger font size, higher color contrast, and faster animation playback speed, then this information can be retrieved to provide a basis for subsequent presentation parameter adjustments.

[0129] Step S530: Input the emotional state vector into the pre-trained parameter mapping network to generate an interface brightness adjustment value and a content layout density value that match the current emotion.

[0130] The pre-trained parameter mapping network is a neural network model trained on a large amount of data. It maps the emotional state vector to corresponding interface brightness adjustment values ​​and content layout density values. The interface brightness adjustment value refers to the value by which the screen brightness needs to be adjusted, and the content layout density value refers to the distribution density of the content on the screen. For example, a long short-term memory (LSTM) network is used as the pre-trained parameter mapping network, and the emotional state vector is input into the network for processing. If the emotional state vector indicates that the user is excited, the network may generate a higher interface brightness adjustment value and a lower content layout density value to create a lively and open interactive atmosphere. If the user is calm, the network may generate a moderate interface brightness adjustment value and a moderate content layout density value. These adjustment values ​​can dynamically adjust the display of interactive content based on the user's emotional state.

[0131] Step S540: Determine a final display parameter combination of the dynamic interactive content based on the screen viewing angle correction parameter, the preference setting information, and the interface brightness adjustment value.

[0132] The final display parameter combination refers to the optimal display parameter set for dynamic interactive content, determined by comprehensively considering the screen viewing angle correction parameters, preference settings, and interface brightness adjustment values. For example, the display direction and viewing angle of the screen content are adjusted based on the screen viewing angle correction parameters; parameters such as font size, color contrast, and animation playback speed are determined based on the preference settings; and the screen interface brightness is adjusted based on the interface brightness adjustment value. These parameters are then combined to form the final display parameter combination. This combination ensures that dynamic interactive content is displayed optimally across different viewing angles, user preferences, and emotional states.

[0133] Step S550: reorganize the information elements in the interactive content flow according to the content layout density value, and trigger the screen display module to perform a parameter update operation.

[0134] The content layout density value determines the distribution density of information elements on the screen. Based on this value, the information elements in the interactive content stream are reorganized, meaning their arrangement and spacing are adjusted to achieve an appropriate layout density. For example, if the content layout density value is low, the spacing between information elements may need to be increased to make the content more spacious; if the content layout density value is high, the spacing may need to be reduced, increasing the number of information elements. Triggering the screen display module to perform a parameter update operation means sending the final display parameter combination and the reorganized interactive content stream to the screen display module, causing it to display according to the new parameters and content. This enables real-time adjustment and updating of dynamic interactive content, improving the user's interactive experience.

[0135] In an embodiment of the present invention, the training process of the parameter mapping network may include the following steps:

[0136] Step S501: Collect feedback data on screen interactive content from multiple groups of users in different emotional states, where the feedback data includes records of users actively adjusting interface parameters and physiological signal monitoring data.

[0137] The purpose of collecting feedback data from multiple groups of users on screen interactive content in different emotional states is to understand users' needs and reactions to interface parameters under different emotions. Records of users actively adjusting interface parameters refer to records of users manually adjusting interface parameters such as interface brightness, font size, and color contrast during the process of interacting with the screen. Physiological signal monitoring data refers to the user's physiological signals monitored by sensors, such as heart rate, blood pressure, and skin conductivity, which can reflect the user's emotional state. For example, using questionnaires and sensor monitoring, feedback data on screen interactive content from multiple groups of users in different emotional states such as happiness, sadness, and anger can be collected. This data can provide rich samples for training the parameter mapping network.

[0138] Step S502: constructing an initial mapping network including an emotion feature input layer and a parameter prediction output layer, wherein the dimension of the emotion feature input layer is consistent with the dimension of the emotion state vector.

[0139] The initial mapping network is the initial version of the parameter mapping network. It consists of an emotion feature input layer and a parameter prediction output layer. The emotion feature input layer receives the emotion state vector, whose dimensions match the emotion state vector to ensure accurate input of emotion information. The parameter prediction output layer outputs interface parameter adjustment values ​​corresponding to the emotion state, such as interface brightness adjustment and content layout density. For example, a multi-layer perceptron (MLP) is used to construct the initial mapping network. The number of neurons in the emotion feature input layer is set to the same as the dimension of the emotion state vector, and the number of neurons in the parameter prediction output layer is set to the number of interface parameters to be predicted. This constructs a network model capable of mapping emotion states to interface parameters.

[0140] Step S503: Time series alignment of the emotional state vector and the user's physiological signal data is performed to generate a training sample set with time series labels.

[0141] Time series alignment involves temporally matching the emotional state vector and the user's physiological signal data so that they correspond to information from the same time period. Generating a training sample set with time series labels involves combining the aligned emotional state vector and user physiological signal data with the corresponding interface parameter adjustment values ​​to form a training sample set containing time sequence information. For example, timestamp information is used to align the emotional state vector and user physiological signal data. These are then combined with records of users actively adjusting interface parameters or predicted interface parameter adjustment values ​​based on physiological signals to form training samples, with each sample being time-labeled. This ensures that the samples in the training sample set have time sequence information, making it easier for the model to learn the dynamic relationship between emotional state and interface parameters.

[0142] Step S504: segmenting the training sample set using a sliding window mechanism to extract the emotional feature mean and parameter adjustment trend vector for each time period.

[0143] The sliding window mechanism is a method for processing time series data. It divides the data into multiple segments by sliding a fixed-size window across the time series. Extracting the mean emotional feature and parameter adjustment trend vector for each time period involves statistically analyzing the emotional features and parameter adjustment values ​​within each window, calculating the mean emotional feature and the trend of the parameter adjustment values. For example, a sliding window of 10 time steps is set and slid across the training sample set. The mean emotional state vector within each window is calculated to obtain the mean emotional feature for that time period. Simultaneously, the changes in the interface parameter adjustment values ​​within the window are analyzed to generate the parameter adjustment trend vector. These mean and trend vectors reflect the changing characteristics of emotional states and interface parameters over different time periods.

[0144] Step S505: Use a long short-term memory network to perform sequence modeling on the parameter adjustment trend vector to generate parameter prediction values.

[0145] A long short-term memory (LSTM) network is a neural network model capable of processing sequential data, capturing long-term dependencies within it. Using LSTM to model the sequence of parameter adjustment trend vectors involves inputting the parameter adjustment trend vector into the LSTM network, allowing the network to learn the sequential pattern of parameter adjustments and generate parameter predictions. For example, the parameter adjustment trend vector is fed into the LSTM network, which processes and learns the network to output predicted values ​​for interface parameter adjustments. These predicted values ​​can be used to guide subsequent interface parameter adjustments.

[0146] Step S506: Calculate the mean square error loss between the parameter prediction value and the actual adjustment parameter, and use the gradient descent algorithm to optimize the weight parameters of the initial mapping network.

[0147] Mean squared error loss is a loss function used to measure the difference between model predictions and actual values. To calculate the mean squared error loss between parameter predictions and actual adjusted parameters, the LSTM network's predicted values ​​are compared with the actual interface parameter adjustments and the mean squared error between them is calculated. Using the gradient descent algorithm to optimize the weight parameters of the initial mapping network involves updating the weight parameters of the initial mapping network based on the gradient of the mean squared error loss, thereby continuously reducing the loss. For example, the mean squared error loss function is used to calculate the loss between the parameter predictions and the actual adjusted parameters. Then, the stochastic gradient descent (SGD) algorithm is used to update the weight parameters of the initial mapping network based on the gradient of the loss. Through continuous iteration, the model's predictive performance is continuously improved.

[0148] Step S507: When the mean square error loss is lower than a preset threshold, stop training and save the network parameters to generate a pre-trained parameter mapping network.

[0149] The preset threshold is a pre-set loss value standard. When the mean squared error loss falls below this threshold, the model's predictive performance has met the preset requirements. Stopping training and saving the network parameters to generate a pretrained parameter mapping network involves saving the current network parameters to form a pretrained model that can be directly used for parameter mapping. For example, when the mean squared error loss falls below 0.01, training is stopped and the network weight parameters are saved to a file. In actual applications, this file is loaded and the parameters are assigned to the initial mapping network, creating a pretrained parameter mapping network. This pretrained model can accurately predict the corresponding interface parameter adjustment values ​​based on the user's emotional state.

[0150] As an implementation manner, after outputting the adjusted screen interaction signal in step S500, the method provided by the embodiment of the present invention may further include the following steps:

[0151] Step S600: Continuously monitor the user's response behavior to the adjusted interactive content, where the response behavior includes eye movement trajectory, facial micro-expression changes, and touch operation frequency.

[0152] Continuously monitoring user responses to adjusted interactive content is designed to understand user satisfaction and usage of the adjusted interactive content. Gaze tracking refers to the path of a user's gaze on the screen, which can reflect the user's focus and browsing order of on-screen content. Facial micro-expressions refer to subtle changes in a user's facial expression, such as a smile or a frown, which can reflect the user's emotional response. Touch operation frequency refers to the number of times a user performs touch operations on the screen, which can reflect the user's level of interaction with the screen. For example, an eye tracker can be used to monitor the user's gaze tracking, a camera can be used to capture changes in the user's facial micro-expressions, and sensors can be used to record the user's touch operation frequency. By continuously monitoring these responses, user feedback and needs for interactive content can be promptly identified.

[0153] Step S700: Analyze the matching degree between the response behavior and the preset expected feedback mode to generate an interaction effect evaluation index.

[0154] The preset expected feedback pattern is the user's pre-defined ideal response pattern for interactive content. It encompasses ideal conditions in terms of gaze trajectory, facial micro-expression changes, and touch operation frequency. Analyzing the match between response behavior and the preset expected feedback pattern evaluates the effectiveness of the interactive content by calculating the similarity between the actual response behavior and the expected feedback pattern. The interactive effectiveness evaluation index is a quantitative value used to measure the quality of the interactive content. For example, for gaze trajectory, the degree of overlap between the actual and expected trajectory can be calculated; for facial micro-expression changes, the similarity between the actual and expected expressions can be determined; and for touch operation frequency, the difference between the actual and expected frequencies can be compared. The degree of match across these aspects is then combined to generate an interactive effectiveness evaluation index. A high match results in a high interactive effectiveness evaluation index, indicating that the interactive content meets user needs; a low match indicates that the interactive content may require further optimization.

[0155] Step S800: When the evaluation index is lower than the fourth threshold, the interaction mode optimization engine is triggered to iteratively update the response rules of the target interaction mode.

[0156] The fourth threshold is a pre-set evaluation index standard used to determine whether the interactive effect has reached an acceptable level. When the interactive effect evaluation index is lower than the fourth threshold, it means that the current interactive content is not effective, and the response rules of the target interactive mode need to be adjusted. The interactive mode optimization engine is a module specifically used to optimize the interactive mode. It can iteratively update the response rules of the target interactive mode based on the user's response behavior and evaluation indicators. For example, if the evaluation indicators show that the user's attention to the current interactive content is low, the interactive mode optimization engine can adjust the response rules to increase the fun or attractiveness of the interactive content. Through continuous iterative updates, the target interactive mode can better meet the needs of users and improve the interactive effect.

[0157] Step S900: Record the feature data generated during the current interaction process, and synchronize the updated interaction mode features to the local mode libraries of all associated devices.

[0158] The feature data generated during the current interaction is recorded for subsequent analysis and optimization. These feature data include the user's facial dynamic behavior characteristics, screen gaze area coordinates, interaction effect evaluation indicators, etc. By recording these data, we can understand the user's behavior patterns and preferences during the interaction process, and provide a basis for further optimizing the interaction mode. The updated interaction mode features are synchronized to the local mode library of all associated devices to ensure that all associated devices can use the latest interaction mode. For example, if a user interacts with the screen on one device and the interaction mode optimization engine updates the interaction mode, the updated interaction mode features will be synchronized to the local mode library of the user's other associated devices (such as mobile phones, tablets, etc.), so that users can also enjoy the optimized interaction experience on other devices.

[0159] Step S1000: Generate a personalized interaction log based on the user identity verification vector, and encrypt and store the log data in a distributed database node.

[0160] Personalized interaction logs are log files generated based on the user's identity information and data from the interaction process. They contain detailed information about the user's interaction process, such as interaction time, interaction content, response behavior, etc. Generating personalized interaction logs based on the user's identity authentication vector ensures that the log file is associated with the user's identity, facilitating subsequent queries and analysis. Encrypting and storing log data in distributed database nodes is to ensure the security and reliability of log data. Distributed database nodes refer to multiple database servers distributed in different geographical locations. Encrypted storage can prevent log data from being illegally obtained and tampered with. For example, personalized interaction log data is encrypted using a symmetric encryption algorithm, and then the encrypted data is stored in distributed database nodes. This can provide secure and reliable data support for subsequent data analysis and user behavior research.

[0161] As an implementation manner, after outputting the adjusted screen interaction signal in step S500, the method provided by the embodiment of the present invention may further include:

[0162] Step S1100: Collecting user feedback data on the adjusted screen interaction signal in real time, where the feedback data includes the facial expression change rate, pupil focus area offset, and touch operation response delay time.

[0163] Real-time collection of user feedback on adjusted screen interaction signals is designed to promptly understand user reactions to the adjusted interactive content. The facial expression change rate refers to the degree of change in a user's facial expression per unit time, reflecting the user's emotional fluctuations. The pupil focus area offset refers to the distance the user's pupil focus area is offset from its normal position, reflecting the user's level of concentration. The touch operation response delay refers to the time interval between a user's touch operation and the screen's response, reflecting the responsiveness of the interactive system. For example, a high-speed camera is used to capture the user's facial expression changes in real time to calculate the facial expression change rate; an eye tracker is used to monitor the user's pupil focus area and measure the offset; and sensors are used to record the duration of touch operations and screen responses to calculate the touch operation response delay. By collecting this real-time feedback data, problems with interactive content can be promptly identified.

[0164] Step S1200: performing time-series alignment on the feedback data and the emotional state vector, extracting the correlation features between the emotional fluctuations and the screen interactive content, and generating an interactive effect fluctuation curve.

[0165] Temporal alignment of feedback data with the emotional state vector refers to matching the feedback data and the emotional state vector in time so that they correspond to information from the same time period. Extracting correlation features between emotional fluctuations and screen interactive content involves analyzing the relationship between feedback data and the emotional state vector to identify correlation patterns between emotional fluctuations and screen interactive content. The interactive effect fluctuation curve, with time as the horizontal axis and the interactive effect evaluation value as the vertical axis, visually demonstrates how the interactive effect changes over time. For example, feedback data such as the rate of change of facial expression, pupil focus area offset, and touch operation response delay are arranged in chronological order with the emotional state vector. Correlation analysis is then used to identify correlations between emotional fluctuations and these feedback data. Based on these correlations, the interactive effect evaluation value at each time point is calculated, and the interactive effect fluctuation curve is plotted. By analyzing this curve, we can understand how the interactive effect changes over different time periods and the impact of emotional fluctuations on the interactive effect.

[0166] Step S1300: Based on the user identity authentication vector, extract the baseline interaction pattern that matches the current user from the historical interaction log, and compare the interaction effect fluctuation curve with the expected effect curve in the baseline interaction pattern for similarity to generate an interaction deviation index.

[0167] Based on the user authentication vector, extracting a baseline interaction pattern matching the current user from historical interaction logs involves identifying previously effective interaction patterns used by the user based on their identity information. The expected effect curve in the baseline interaction pattern represents the time-varying curve of the interaction effect under ideal conditions. Comparing the interaction effect fluctuation curve with the expected effect curve in the baseline interaction pattern evaluates the degree of difference between the current interaction effect and the expected effect by calculating the similarity between the two curves. The interaction deviation index is a quantitative value used to represent this difference. For example, the dynamic time warping (DTW) algorithm is used to calculate the similarity between the interaction effect fluctuation curve and the expected effect curve, and the similarity is converted into an interaction deviation index. A small interaction deviation index indicates that the current interaction effect is close to the expected effect; otherwise, a large deviation indicates that the interaction content needs to be adjusted.

[0168] Step S1400: Determine the interactive element types that need to be optimized first according to the deviation degree of different feedback dimensions in the interactive deviation index, and generate a dynamic optimization strategy including element type weights and optimization directions.

[0169] The degree of deviation across different feedback dimensions in the interaction deviation metric refers to the difference between the current interaction effect and the expected effect across various feedback dimensions, such as the rate of change of facial expression, pupil focus area offset, and touch operation response delay. Based on these deviations, priority optimization is determined by identifying interactive elements that have a significant impact on interaction effectiveness, such as interface color, content layout, and interactive feedback. Generating a dynamic optimization strategy that includes element type weights and optimization directions assigns different weights to each interactive element type and determines optimization directions, such as increasing or decreasing the parameter value of a specific element. For example, if the interaction deviation metric shows a significant deviation in pupil focus area offset, indicating that users are easily distracted, it may be necessary to prioritize optimizing the interface's content layout to make important information more prominent. Assigning a higher weight to the content layout element and determining optimization directions to adjust the arrangement and spacing of information can help optimize interactive content in a targeted manner, improving interaction effectiveness.

[0170] Step S1500: Fusing the dynamic optimization strategy with the response rules of the target interactive mode, and updating the triggering conditions and content generation logic of the corresponding mode in the interactive mode library.

[0171] Fusing the dynamic optimization strategy with the response rules of the target interaction mode means integrating the element type weights and optimization directions in the dynamic optimization strategy into the response rules of the target interaction mode, so that the response rules can be adjusted according to the optimization strategy. Updating the trigger conditions and content generation logic of the corresponding mode in the interactive mode library means modifying the trigger conditions and content generation logic of the corresponding mode in the interactive mode library based on the merged rules. For example, if the dynamic optimization strategy requires increasing the contrast of the interface color, then add the corresponding adjustment logic to the response rules of the target interaction mode, and automatically adjust the contrast of the interface color when the preset conditions are met. At the same time, update the trigger conditions and content generation logic of the corresponding mode in the interactive mode library to ensure that in the subsequent interactive process, interactive content can be generated according to the optimized rules.

[0172] Step S1600: Re-match the user's current facial dynamic behavior characteristics according to the updated interaction pattern library, generate a second-adjusted screen interaction signal, and overwrite the signal execution result outputted previously.

[0173] Re-matching the user's current facial dynamic behavior features against the updated interaction pattern library involves performing a similarity match between the user's current facial dynamic behavior features and the patterns in the updated interaction pattern library to identify the most suitable interaction pattern. Generating a secondary adjusted screen interaction signal involves further adjusting the presentation parameters of the dynamic interactive content based on the matched interaction pattern, generating a new screen interaction signal. Overwriting the previously output signal results in replacing the previously output signal with the secondary adjusted screen interaction signal, resulting in the screen displaying further optimized interactive content. For example, the user's current dynamic behavior features, such as facial expressions and muscle movements, are input into the updated interaction pattern library for matching. If a new interaction pattern is matched, presentation parameters such as interface brightness and content layout are adjusted based on the response rules of that pattern to generate a secondary adjusted screen interaction signal. This signal is then sent to the screen display module, overwriting the previously output signal, resulting in the screen displaying interactive content that better meets the user's needs. This continuous optimization and adjustment process can continuously improve the effectiveness of screen interaction and user experience.

[0174] In summary, the face recognition-based screen interaction method provided by the embodiments of the present invention obtains a user's facial image sequence, extracts facial dynamic behavior characteristics and emotional state information, matches the target interaction mode from a preset interaction mode library, generates dynamic interactive content corresponding to the user's visual focus, and adjusts presentation parameters in real time based on the user's identity and emotional state. Furthermore, by continuously monitoring the user's response behavior, the interaction mode is optimized and updated, continuously improving the interactive effect and user experience. This method fully utilizes face recognition technology and user emotional information to achieve more personalized and intelligent screen interaction, with broad application prospects and practical value. In practical applications, this method can be applied to smart TVs, advertising screens, interactive games, and other fields, providing users with a richer and more engaging interactive experience. For example, in smart TVs, program recommendations and display effects can be automatically adjusted based on the user's expression and emotional state; in advertising screens, advertising content and presentation methods can be adjusted in real time based on audience reactions, improving the appeal and effectiveness of the ads; and in interactive games, game difficulty and scenes can be dynamically adjusted based on the player's emotional changes, enhancing the fun and immersion of the game. As technology continues to develop and improve, this method can also be combined with other technologies (such as virtual reality, augmented reality, etc.) to create more novel and unique interactive experiences.

[0175] See also Figure 2 , Figure 2This is a schematic diagram of the structure of a screen interactive system provided in an embodiment of the present invention. The screen interactive system, for example, is a system embedded in an interactive screen, or the interactive screen itself, and includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, communication interface 102, and memory 103 may be connected via a bus or other means. The processor 101 (also known as the Central Processing Unit (CPU)) is the computing and control core of the screen interactive system, capable of parsing various instructions within the screen interactive system and processing various data within the screen interactive system. The communication interface 102 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi or a mobile communication interface), and may be used to send and receive data under the control of the processor 101. The communication interface 102 may also be used for data transmission and interaction within the screen interactive system. The memory 103 is a storage device within the screen interactive system, used to store programs and data. It is understood that the memory 103 herein may include both the built-in memory of the screen interactive system and the extended memory supported by the screen interactive system. Memory 103 provides storage space for storing the operating system of the screen interaction system, which may include, but is not limited to, Android, iOS, Windows Phone, and the like, though the present invention is not limited thereto. In one embodiment, processor 101 executes the face recognition-based screen interaction method provided above in the embodiments of the present invention by running the computer program stored in memory 103.

Claims

1. A screen interaction method based on face recognition, characterized in that: The method comprises: Acquire a facial image sequence triggered by a target user in a screen interaction scenario, wherein the facial image sequence includes the user's facial dynamic behavior features and screen gaze area coordinates; According to the facial dynamic behavior characteristics, matching a target interaction mode from a preset interaction mode library, wherein the target interaction mode includes interaction response rules associated with the intensity of the user's expression; Based on the coordinates of the screen gaze area and the target interaction mode, generating dynamic interactive content corresponding to the user's visual focus; Calling a pre-trained facial feature fusion model to perform multi-level feature extraction on the facial region in the face image sequence to generate a user identity verification vector and an emotional state vector; Adjusting the presentation parameters of the dynamic interactive content in real time according to the user identity verification vector and the emotional state vector, and outputting an adjusted screen interaction signal; The step of adjusting the presentation parameters of the dynamic interactive content in real time according to the user identity verification vector and the emotional state vector includes: Monitor the user's facial deflection angle and the distance between the user and the screen to generate screen viewing angle correction parameters; Retrieving preference information from a user portrait database based on the user identity verification vector, the preference information including font size preference, color contrast threshold, and animation playback speed range; Inputting the emotional state vector into a pre-trained parameter mapping network to generate an interface brightness adjustment value and a content layout density value that match the current emotion; Determining a final display parameter combination of the dynamic interactive content based on the screen viewing angle correction parameter, the preference setting information, and the interface brightness adjustment value; Reorganize the information elements in the interactive content flow according to the content layout density value, and trigger the screen display module to perform a parameter update operation; The training process of the parameter mapping network includes: Collect feedback data from multiple groups of users on screen interactive content in different emotional states, including records of users actively adjusting interface parameters and physiological signal monitoring data; Constructing an initial mapping network comprising an emotion feature input layer and a parameter prediction output layer, wherein the dimension of the emotion feature input layer is consistent with the dimension of the emotion state vector; Performing time series alignment on the emotional state vector and the user's physiological signal data to generate a training sample set with time series labels; Segmenting the training sample set using a sliding window mechanism to extract the emotional feature mean and parameter adjustment trend vector for each time period; Performing sequence modeling on the parameter adjustment trend vector using a long short-term memory network to generate parameter prediction values; Calculating the mean square error loss between the parameter prediction value and the actual adjustment parameter, and optimizing the weight parameters of the initial mapping network using a gradient descent algorithm; When the mean square error loss is lower than a preset threshold, the training is stopped and the network parameters are saved to generate the pre-trained parameter mapping network.

2. The method according to claim 1, characterized in that The matching of the target interaction mode from a preset interaction mode library according to the facial dynamic behavior characteristics includes: Inputting the facial dynamic behavior features into a pre-trained spatiotemporal convolutional network to extract facial muscle movement trajectory features and eye opening and closing frequency features; Determining the user's concentration level on screen content based on the eye opening and closing frequency characteristics, and generating a user intention prediction vector in combination with the muscle movement trajectory characteristics; Performing similarity matching between the user intention prediction vector and the pattern features in the interaction pattern library, and screening out a set of candidate interaction patterns whose similarity exceeds a first threshold; Performing weight correction on each candidate interaction mode in the candidate interaction mode set based on the concentration level, and selecting the candidate interaction mode with the highest weight as the target interaction mode; Each interaction mode in the interaction mode library is associated with at least one facial action triggering condition and a corresponding screen response strategy.

3. The method according to claim 2, characterized in that Inputting the facial dynamic behavior features into a pre-trained spatiotemporal convolutional network to extract facial muscle movement trajectory features and eye opening and closing frequency features includes: Performing facial key point detection on each frame image in the facial image sequence to generate a key point distribution map including coordinates of the eyebrow region, the cheek region, and the mouth corner region; Determining the motion direction consistency parameter of each facial region within a preset time window based on the key point displacement vectors between adjacent frames, wherein the motion direction consistency parameter is used to characterize the stability of the coordinated movement of the muscle groups; Performing optical flow analysis on the key point distribution map to generate a dynamic change curve reflecting the strength of muscle contraction, and extracting periodic features of facial muscle movement based on the frequency of peaks and troughs in the dynamic change curve; performing block-by-block grayscale value comparison on image blocks of the eye region in the facial image sequence, calculating a rate of change in pixel area of ​​the pupil region covered by the eyelid per unit time, and generating an eye opening and closing state waveform graph with time as the horizontal axis; Calculating a synergy index between muscle movement and eye movement based on the phase difference between the periodic feature and the eye opening and closing state waveform, and performing weighted fusion on the synergy index and the movement direction consistency parameter to generate the facial muscle movement trajectory feature; The eye opening and closing state waveform is transformed in the frequency domain to extract the effective waveform segments whose amplitude exceeds the preset noise threshold, and the eye opening and closing frequency characteristics are generated based on the interval length and duration ratio of the effective waveform segments.

4. The method according to claim 2, characterized in that The calling of the pre-trained facial feature fusion model to perform multi-level feature extraction on the facial region in the facial image sequence includes: Dividing each frame image in the facial image sequence into a first facial region, a second facial region, and a third facial region, wherein the first facial region includes eye contour coordinates, the second facial region includes mouth contour coordinates, and the third facial region includes overall facial contour coordinates; Performing local texture analysis on the first facial region using a first feature extraction branch of the facial feature fusion model to generate a first region feature vector containing a pupil movement trajectory; Performing dynamic deformation monitoring on the second facial region through the second feature extraction branch of the facial feature fusion model to generate a second region feature vector including the lip opening and closing amplitude; Performing global illumination compensation processing on the third facial region through the third feature extraction branch of the facial feature fusion model to generate a third region feature vector including a skin color change trend; Cross-channel fusing the first region feature vector, the second region feature vector, and the third region feature vector to generate the user identity verification vector and the emotional state vector; The training process of the facial feature fusion model includes: Obtaining a training image set labeled with a user identity label and an emotion label, wherein each image in the training image set is labeled with eye region coordinates, mouth region coordinates, and overall facial region coordinates; constructing an initial fusion model comprising three parallel feature extraction branches, wherein the three branches correspond to feature extraction of the first facial region, the second facial region, and the third facial region, respectively; Performing regional division on each image in the training image set, extracting training region images corresponding to each branch, and generating regional feature vectors for each branch; The regional feature vectors of each branch are input into the fully connected layer for feature splicing to generate a fused feature vector; Obtaining a first cross entropy loss between the fused feature vector and the user identity label, and a second cross entropy loss between the fused feature vector and the emotion label; Performing backpropagation optimization on the parameters of the initial fusion model according to a weighted sum of the first cross entropy loss and the second cross entropy loss until the loss converges; The optimized model parameters are solidified to generate the pre-trained facial feature fusion model.

5. The method according to claim 1, wherein The generating of dynamic interactive content corresponding to the user's sight focus based on the screen gaze area coordinates and the target interaction mode includes: Determining an identifier of a currently focused control on the screen interface according to the coordinates of the screen gaze area; Obtaining historical interaction data associated with the currently focused control, where the historical interaction data includes the number of times the user triggers the control, the duration of the stay, and associated operation records; generating a dynamic content update instruction corresponding to the currently focused control based on a response rule in the target interaction mode; Adjusting element attributes in the dynamic content update instruction according to the emotional state vector and the historical interaction data, the element attributes including color gradient rate, content switching frequency, and interactive feedback intensity; The adjusted dynamic content update instructions are sent to the screen rendering engine to generate an interactive content stream containing three-dimensional visual effects.

6. The method according to claim 1, characterized in that The step of obtaining a facial image sequence triggered by a target user in a screen interaction scenario includes: Capturing a user's facial original image stream through an embedded camera, and performing frame rate synchronization processing on the original image stream; Detecting the face bounding box coordinates in each frame image, and performing image cropping and size normalization based on the bounding box coordinates; Perform illumination equalization on the normalized image to eliminate the influence of ambient light fluctuation on image quality and obtain a balanced image; Use background segmentation algorithm to extract the face foreground area in the equalized image and remove invalid frames containing occlusions or blurred areas; The processed valid image frames are arranged in chronological order to generate the face image sequence, and a timestamp and a screen touch event mark are added to each frame of the image.

7. The method according to claim 6, characterized in that The detecting of the face bounding box coordinates in each frame of image includes: Call the pre-trained face detection model to perform initial face area prediction on the current frame image and generate a set of candidate bounding boxes; Calculate the confidence score of each candidate bounding box and select candidate boxes with scores exceeding the second threshold for merging; According to the merged candidate frame coordinates, determine the center point coordinates and width-to-height ratio of the final face bounding box; The bounding box position of the current frame is predicted based on the facial motion trajectory of the historical frames. When the deviation between the predicted position and the actual detected position exceeds a third threshold, the face tracking model is triggered to perform position correction. The corrected bounding box coordinates are proportionally converted to the screen resolution to generate standardized coordinate data that adapts to the current screen size.

8. The method according to claim 1, characterized in that After outputting the adjusted screen interaction signal, the method further includes: Continuously monitoring the user's response to the adjusted interactive content, including eye movement trajectory, facial micro-expression changes, and touch operation frequency; Analyze the matching degree between the response behavior and the preset expected feedback mode to generate an interaction effect evaluation index; When the evaluation index is lower than a fourth threshold, triggering the interaction mode optimization engine to iteratively update the response rule of the target interaction mode; Record the feature data generated during the current interaction and synchronize the updated interaction mode features to the local mode library of all associated devices; Generate personalized interaction logs based on user authentication vectors and encrypt and store log data in distributed database nodes.

9. A screen interactive system, characterized in that: include: a memory storing a computer program; A processor, configured to load the computer program to implement the face recognition-based screen interaction method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • System switching method and system of family education learning machine

    CN117666786A