Digital human live broadcast attention assessment system based on multi-modal physiological feature fusion
By building a multimodal physiological feature fusion model based on Transformer architecture, combining EEG and eye movement data, the attention of digital live viewers in real time and dynamically adjusting the perspective of virtual cameras, the problem of lack of intelligent attention evaluation and optimization in the existing technology is solved, and the user experience and live broadcast effect are significantly improved.
Patent Information
- Application Number
- CN202510247382.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-13
AI Technical Summary
The existing technology lacks a systematic and intelligent way to combine multimodal physiological characteristics to evaluate the audience's attention during live broadcasts in real time, thereby providing digital people with dynamic adaptive optimization strategies.
The deep learning model based on the Transformer architecture is adopted, and the multimodal physiological feature fusion model is constructed by combining historical digital live videos, user EEG data and eye movement data to output the audience's attention distribution changes in real time, and adjust the virtual camera perspective according to the attention distribution.
Accurately evaluate the attention of digital people's live broadcast audiences, dynamically optimize live broadcast content, and significantly improve user experience and live broadcast effect.
Smart Images

Figure CN120145307A_ABST
Abstract
Description
Technical Field
[0001] The present invention specifically relates to a digital human live broadcast attention evaluation system based on multimodal physiological feature fusion. Background Art
[0002] In the context of the rapid development of current digital and artificial intelligence technologies, digital humans, as a new type of interaction technology, have been widely applied in multiple fields such as customer service, education, and healthcare, greatly improving the user experience and work efficiency. During a live broadcast, if a digital human fails to effectively guide the audience's attention, it may lead to a significant decline in the user experience. The audience may feel confused or disappointed, thus reducing their attention to the live broadcast products and affecting the live broadcast effect of the digital human. However, in the evaluation of the attention of the audience during a digital human live broadcast, the existing technologies mainly rely on traditional user behavior analysis and ignore the potential of multimodal physiological signals in optimizing the attention evaluation method.
[0003] On the other hand, the progress of image processing technologies, such as Scale-Invariant Feature Transform (SIFT), provides a powerful tool for extracting and describing image features. SIFT can identify key points in an image at different scales and rotations and extract their local features, enabling machines to better understand visual inputs. The progress of video processing technologies under architectures such as Swin Transformer has enabled deep learning models to be widely applied in various video processing scenarios. It can handle dynamic changes in videos by introducing temporal dimension modeling, improving the performance of tasks such as video classification and action recognition. In addition, with the progress of electroencephalogram (EEG) and eye-tracking technologies, researchers can effectively capture the changes in the user's attention through these physiological data, making it possible to analyze the dynamic changes in the attention of the audience when watching a digital human live broadcast. Nevertheless, how to effectively combine these physiological signals into the interactive optimization of digital humans remains an unsolved challenge. The existing technologies lack a systematic and intelligent way to combine multimodal physiological features to evaluate the attention of the audience during a digital human live broadcast in real time, so as to provide dynamic adaptive optimization strategies for digital humans. Summary of the Invention
[0004] The present invention provides a digital human live broadcast attention evaluation system based on multimodal physiological feature fusion to solve the above-mentioned technical problems, and specifically adopts the following technical solutions:
[0005] A digital human live broadcast attention evaluation system based on multimodal physiological feature fusion, characterized by comprising:
[0006] A multi-modal physiological feature fusion model construction module, which is used to construct a deep learning model based on the Transformer architecture based on historical digital human live video, electroencephalogram data, and eye movement data of users when watching videos;
[0007] A digital human live audience attention evaluation module, which is used to input the digital human live video into the multi-modal physiological feature fusion model, output the change of the attention distribution of the audience when watching the video in real time, evaluate the attention overlap and focus degree of the audience according to the change of the attention distribution, and then adjust the virtual camera angle to optimize the digital human live broadcast.
[0008] Furthermore, the multi-modal physiological feature fusion model construction module includes:
[0009] A multi-modal data collection unit, which is used to collect electroencephalogram data, eye movement data, and corresponding mobile videos of users when watching digital human live broadcasts;
[0010] A synchronization processing unit, which is used to perform time-axis synchronization processing on the eye movement data and the mobile video to ensure the timing consistency of the data;
[0011] A feature matching unit, which is used to extract the key points and their descriptors of the synchronized video frames, perform feature matching, calculate the similarity between the key point descriptors, filter out outliers, and construct a homography transformation matrix.
[0012] Furthermore, the feature matching unit calculates the similarity between the key point descriptors through the Euclidean distance and uses the random sample consensus algorithm to filter out outliers.
[0013] Furthermore, the multi-modal physiological feature fusion model construction module further includes a fixation point mapping unit, which is used to map the eye movement data in the first-person perspective to the mobile video, obtain the mapped fixation point coordinates, and perform coordinate transformation through the homography transformation matrix.
[0014] Furthermore, the multi-modal physiological feature fusion model construction module further includes a fixation point aggregation unit, which is used to clean and aggregate the mapped fixation points, generate a continuous saliency map for each frame, and perform convolution processing on the fixation data through a Gaussian mask.
[0015] Furthermore, the multi-modal physiological feature fusion model construction module further includes a visual feature map extraction unit, which captures cross-frame attention transfer information through Video Swin Transformer and extracts four-scale visual features F1, F2, F3, and F4.
[0016] Furthermore, the multi-modal physiological feature fusion model construction module further includes an electroencephalogram feature map extraction unit, which extracts the alpha wave features of the electroencephalogram signal through wavelet transform and fuses them into the visual feature F1 to form the fused feature.
[0017] Furthermore, the multi-modal physiological feature fusion model construction module further includes a global feature enhancement module for enhancing the spatial and semantic features of the fused feature, including spatio-temporal feature enhancement and channel feature enhancement.
[0018] Furthermore, the multi-modal physiological feature fusion model construction module further includes a selective feature fusion module for fusing features of different scales and refining the key gaze transfer information.
[0019] Furthermore, the multi-modal physiological feature fusion model construction module further includes a progressive saliency prediction module for using the gaze shift information and high-level semantic information to guide the low-level appearance information, refining the saliency region, and optimizing the model performance through a loss function to generate the saliency map S.
[0020] Furthermore, the digital human live broadcast audience attention evaluation module includes:
[0021] An output module for predicting the user's attention heat map through the trained multi-modal physiological feature fusion model;
[0022] An evaluation module for evaluating the rationality of attention distribution by calculating the overlap degree and entropy value between the attention heat map and the target product area;
[0023] A visual focus adjustment unit for adjusting the virtual camera view according to the evaluation result, optimizing the live broadcast content, and ensuring that the user's attention is concentrated on the key area.
[0024] The digital human live broadcast attention evaluation system based on multi-modal physiological feature fusion provided by the present invention accurately evaluates the attention of the digital human live broadcast audience, dynamically optimizes the live broadcast content, significantly improves the user experience and the live broadcast effect, and promotes the wide application of digital human technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0026] Figure 1 It is a schematic diagram of a digital human live broadcast attention evaluation system based on multi-modal physiological feature fusion of the present invention.
[0027] Figure 2 It is a schematic diagram of the spatio-temporal feature enhancement module and the channel feature enhancement module of the digital human live broadcast attention evaluation system based on multi-modal physiological feature fusion of the present invention;
[0028] Figure 3 It is a schematic diagram of the selective feature fusion module of the digital human live broadcast attention evaluation system based on multi-modal physiological feature fusion of the present invention; Detailed implementation manners
[0029] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.
[0030] As Figure 1 shown, the system of the present invention includes two main modules: a multi-modal physiological feature fusion model construction module and a digital human live broadcast audience attention evaluation module. The multi-modal physiological feature fusion model construction module is used to construct a deep learning model based on the Transformer architecture based on historical digital human live broadcast videos, electroencephalogram data and eye movement data of users when watching videos. The digital human live broadcast audience attention evaluation module is used to input the digital human live broadcast video into the multi-modal physiological feature fusion model, output the change of the attention distribution of the audience when watching the video in real time, evaluate the attention overlap degree and focus degree of the audience according to the change of the attention distribution, and then adjust the virtual camera angle to optimize the digital human live broadcast.
[0031] The specific implementation manners of each module are introduced in detail below.
[0032] 1. Multi-modal physiological feature fusion model construction module
[0033] (1) Multi-modal data collection unit
[0034] The multi-modal data collection unit is used to collect electroencephalogram data, eye movement data and corresponding mobile videos of users when watching digital human live broadcasts. Specifically, the user wears an electroencephalogram-glasses type eye movement device to watch the virtual digital human live broadcast with goods. The electroencephalogram device records the electroencephalogram signals of the subject, as well as the actions, behaviors and eye movement data (fixation points and their timestamps and coordinates) of the user from the first-person perspective, and at the same time records the mobile video content watched by the user.
[0035] (2) Synchronization processing unit
[0036] The synchronization processing unit is used to perform time-axis synchronization processing on the eye movement data and the mobile video to ensure the temporal consistency of the data. Since the eye movement data is recorded by an eye movement device worn on the user's head (the ego-centric video from the first-person perspective), its time axis may not be completely synchronized with that of the mobile video. Therefore, it is necessary to synchronize the ego-centric video frames and the mobile video frames to ensure that their time sequences match. Based on the synchronization information, pairs of ego-centric video frames and mobile video frames are obtained, and at the same time, the synchronization of the electroencephalogram data, eye movement data, and mobile video is ensured according to the time stamp information of the electroencephalogram.
[0037] (3) Feature matching unit
[0038] The feature matching unit is used to extract the key points and their descriptors of the synchronized video frames, perform feature matching, calculate the similarity between the key point descriptors, filter out outliers, and construct a homography transformation matrix. For each pair of synchronized frames of the ego-centric video frame and the mobile video, the scale-invariant feature transform (SIFT) image processing algorithm is applied to extract the key points and their descriptors. Key points are points with significant features in the image, and these points are usually stable in different images or perspectives. Even under lighting changes, perspective changes, or other interfering factors, they can still maintain recognizability. Descriptors are feature vectors extracted around each key point to describe the local image information of that point.
[0039] The similarity between each key point descriptor is calculated by computing the Euclidean distance, and the formula is as follows:
[0040]
[0041] This is the descriptor of the i-th key point in the ego-centric video frame (first-person perspective). is the descriptor of the j-th key point in the mobile video frame. is to calculate and is a measure of the distance or similarity between. When the distance between the descriptors of two key points is less than the set threshold, they are considered to be matched; otherwise, they will be regarded as false matches. By setting the threshold, the possibility of false matches can be effectively reduced, thereby improving the accuracy of subsequent analysis.
[0042] The Random Sample Consensus (RANSAC) algorithm is used to filter out outliers in the matches. This method estimates the mapping relationship by randomly selecting a subset of feature matching points, and identifies and removes outliers (noisy data). After feature matching, a system of equations is constructed based on the matched key points to calculate the homography transformation matrix H. This matrix H describes the geometric relationship between the egocentric video coordinate system and the moving video coordinate system.
[0043] (4) Fixation point mapping unit
[0044] The fixation point mapping unit is used to map the eye movement data in the first-person view to the moving video, obtain the mapped fixation point coordinates, and achieve coordinate transformation through the homography transformation matrix. Since the eye movement device worn by the user records the eye movement data from the first perspective, it is necessary to effectively map the data in the first-person video to the moving video watched by the subject. Through the homography transformation matrix H, the original fixation point coordinates in the first-person video frame are mapped to the coordinates in the moving video frame. The formula is as follows:
[0045]
[0046] The original fixation point coordinates are the original fixation point coordinates in the first-person video frame, and are the coordinates in the mapped moving video frame.
[0047] (5) Fixation point aggregation unit
[0048] The fixation point aggregation unit is used to clean and aggregate the mapped fixation points, generate a continuous saliency map for each frame, and perform convolution processing on the fixation data through a Gaussian mask. After mapping, the fixation points in the moving video are further cleaned to remove invalid fixations. The valid eye movement fixation data are aggregated in each frame of the moving video. To generate a continuous saliency map for each frame, a Gaussian mask is used to convolve the fixation data of all participants.
[0049] (6) Visual feature map extraction unit
[0050] The visual feature map extraction unit captures cross-frame attention transfer information through Video Swin Transformer and extracts visual features F1, F2, F3, and F4 at four scales. Traditional convolutional neural network (CNN)-based methods cannot effectively capture long-distance cross-frame fixation transfer features due to their fixed-size spatio-temporal feature extraction. Based on the moving video after fixation point aggregation, which contains cross-frame attention transfer information, the present invention captures this information through Video Swin Transformer, thereby improving the accuracy of saliency prediction. The feature extraction process aims to extract features F1, F2, F3, and F4 at four scales. The generation of features at each stage can be expressed by the following formula:
[0051]
[0052] where I represents the currently processed input frame or image. This frame is one frame extracted from the video sequence and serves as the starting input for feature extraction. Subsequently, the features at each stage are used to generate the features of the next stage.
[0053]
[0054] The shapes of the features output at each stage are:
[0055]
[0056] where C is the number of channels, T is the length of the input video clip, H is the video height, and W is the video width.
[0057] (7) EEG feature map extraction unit
[0058] The EEG feature map extraction unit extracts the alpha wave features of the EEG signal through wavelet transform and fuses them into the visual feature F1 to form the fused feature. The alpha wave (8 - 12 Hz) is highly correlated with individual attention, so it is selected as the EEG index. The alpha wave features E of the EEG signal are extracted through wavelet transform alpha .
[0059] (8) Feature splicing unit
[0060] The extracted EEG feature E alphaFused into the visual features. Since feature F1 usually contains relatively basic and key information, which can directly reflect the main changes or features of the task. Fusing with features such as F2, F3, F4, etc. will gradually increase the complexity of the features and may lead to a significant increase in calculation and training time. Excessive fusion may result in feature redundancy, which instead weakens the contribution of EEG features to the model. Therefore, to balance computational efficiency and model performance, only fuse EEG features in the basic features. Merge the visual features and EEG features in the channel dimension to form a richer feature representation:
[0061]
[0062] Subsequently, input the fused features into the global feature enhancement module.
[0063] (9) Global Feature Enhancement Module
[0064] The global feature enhancement module is used to enhance the spatial and semantic features of the fused features, including a spatio-temporal feature enhancement module and a channel feature enhancement module.
[0065] To further improve the expressive ability of the fused features, the global feature enhancement module aims to enhance the global characteristics of the spatial features and the comprehensiveness of the semantic features. This module specifically includes two major units: spatio-temporal feature enhancement and channel feature enhancement.
[0066] The spatio-temporal feature enhancement module, as shown in Figure 2 :
[0067] Input the features obtained from the step of feature concatenation Obtain feature maps through average pooling and max pooling in the channel dimension and .
[0068]
[0069] Concatenate these two feature maps in the channel dimension, and calculate the spatio-temporal attention weight map using a 3×3 convolutional block and the Sigmoid function :
[0070]
[0071] Through residual connection, multiply the original feature map and the attention weight map element by element to obtain the final enhanced spatio-temporal features.
[0072]
[0073] Input feature F2, repeat the above steps to obtain the spatio-temporally enhanced features .
[0074] Channel feature enhancement module, such as Figure 2 shown in
[0075] The goal of channel feature enhancement is to improve the importance of feature channels, enabling the model to pay more attention to the feature channels relevant to the task.
[0076] Two feature maps are obtained by performing global average pooling (GAP) and global max pooling (GMP) on the high-level feature F3 respectively and .
[0077]
[0078] Embed two 1×1 convolutional blocks to compress and expand the channel dimension, then concatenate these feature maps, and through and perform a 3×3 convolutional operation and sigmoid function to obtain the channel attention weight map .
[0079]
[0080] Finally, apply a residual connection to F3 to generate an enhanced channel representation . Repeat the above steps for feature F4 to obtain the channel-enhanced feature .
[0081] (10) Figure 3 shows the selective feature fusion module
[0082] The selective feature fusion module is used to fuse features of different scales and refine the key gaze transfer information. Set to fuse features of different scales to further refine the key gaze transfer information. In the upper branch, the feature is upsampled through a 1×1 convolutional layer and aligned with the feature to obtain the upsampled feature respectively. After passing through a global average pooling layer, two 1×1 convolutional layers and a sigmoid function, an attention map is generated.
[0083] In the lower branch, the spatial dimension of the feature is adjusted through an upsampling layer to obtain the adjusted feature . Subsequently,[[]] through the same steps, an attention map is generated.
[0084] To promote information interaction, the upsampled features are multiplied element-wise with the attention map to generate preliminary fused features through residual connections .
[0085]
[0086] Similarly, features are obtained in a similar manner .
[0087]
[0088] Finally, the preliminary fused features and are concatenated in the channel dimension and embedded in a 3x3 convolutional block to compress the channel dimension. Finally, the fused features can be obtained
[0089]
[0090] are generated through the processing method of the selective feature fusion module and .
[0091] Upper branch processing: The shallow features are aligned in channels and upsampled to the same resolution as the deep features to generate features . After passing through the global average pooling layer, a 1x1 convolutional layer, and the sigmoid function, the attention map is generated
[0092] In the lower branch processing, the spatial dimensions of the features and are adjusted through the upsampling layer to obtain the adjusted features and . Subsequently, and go through the same steps such as the global average pooling layer to generate the attention maps and .
[0093] To promote information interaction, the upsampled features are multiplied element-wise with the attention maps and to generate preliminary fused features , and through residual connections
[0094] The preliminary fused features and , and are concatenated in the channel dimension, and a 3x3 convolutional block is embedded to compress the channel dimension. Finally, the fused feature and can be obtained.
[0095] (11) Progressive saliency prediction module
[0096] The progressive saliency prediction module is used to utilize the fixation shift information and high-level semantic information to guide the low-level appearance information, refine the saliency region, and optimize the model performance through a loss function to generate a saliency map S. In the progressive saliency prediction module, the low-level appearance information is guided by using the fixation shift information and high-level semantic information to refine the most salient region. The specific steps are as follows:
[0097] Product information encoding: The product information (such as product category, characteristics, etc.) is converted into a vector form through an embedding layer to obtain the product information encoding .
[0098] Feature adjustment and concatenation: In the first progressive saliency prediction block, the number of channels and the spatial resolution of the product information encoding are adjusted through a 3x3 convolutional block and an upsampling layer to make it consistent with the feature . Then, the adjusted feature and the fused feature are concatenated along the time dimension to generate a hybrid feature .
[0099] Saliency information aggregation: In the subsequent progressive saliency prediction blocks, new features are gradually introduced, and the number of channels and the time dimension are gradually reduced through a 3x3 convolutional block with a time stride of 2. Through these progressive saliency prediction blocks, sufficient saliency information can be aggregated to obtain a refined saliency representation .
[0100]
[0101] Saliency map generation: Finally, multiple 3x3 convolutions and upsampling layers are used to output the saliency map S. The saliency map S is a heat map used to identify the salient regions in a visual scene, and the salient regions are usually the regions that the human visual system is most likely to focus on:
[0102]
[0103] Loss function optimization: The optimization objective of the loss function is to make the saliency map generated by the progressive saliency prediction module closer to the true attention distribution. The Kulback-Leibler (KL) divergence, correlation coefficient (CC), normalized scanpath saliency (NSS), and similarity (SIM) are used to optimize the model performance. Thereby improving the accuracy and robustness of saliency prediction. The parameters λ1, λ2, and λ3 are balance parameters, which are set to -0.1, -0.1, and -0.1 respectively.
[0104]
[0105] 2. Digital human live broadcast audience attention evaluation module
[0106] (1)Output module
[0107] The multi-modal physiological feature fusion model trained by minimizing the loss function can predict the user's attention heat map (saliency map S) after inputting the digital human live e-commerce video. According to the shape of the heat map, the interaction mode of the digital human can be optimized to ensure that the user's attention always focuses on the key interaction content and information. This optimization not only improves the user experience, but also helps content creators identify which elements can most attract the audience's attention, so as to more effectively utilize this information in future content design.
[0108] (2)Evaluation module
[0109] The method of this application is different from traditional prediction models. Its output module generates the user's attention heat map. This design makes the prediction results and attention data output by the model more intuitive, but it also needs to be further evaluated by the evaluation module to obtain a reasonable judgment on the user's attention distribution. The rationality of the user's attention distribution is judged by the attention overlap degree and focus degree of the user.
[0110] It is quantified by calculating the overlap degree between the heat map and a live e-commerce product area. The product area is defined by the method of manual annotation:
[0111]
[0112] Among them, represents the user's actual attention heat map, P represents the target product area, represents the intersection part of the user's actual attention heat map and the target product area, represents the union part of the user's actual attention heat map and the target product area, that is, all positions covering the user's attention and the target area.
[0113] The entropy of the heat map is used to measure whether the attention is concentrated in a few regions rather than evenly distributed. The more concentrated the user's attention is, the lower the entropy; the more dispersed it is, the higher the entropy:
[0114]
[0115]
[0116] Among them, p(x) represents the probability density function of the attention distribution, and H(x) is the attention value at region x, which is the sum of the attention values at all positions.
[0117] (3) Visual focus adjustment unit
[0118] When the Overlap value is lower than 0.5 or the entropy value exceeds 3.0, it indicates that the attention is very dispersed, and there is a significant deviation between the user's attention and the product area that the manager expects the user to focus on. At this time, it is necessary to adjust the digital human's live content in a timely manner.
[0119] Different from the live video of real people, the digital human live background is usually generated by a computer. When it is recognized that the user may have a distracted attention, the manager can change the background, scene, digital human actions, etc. in a timely manner to provide visual freshness and maintain the audience's attention. Specifically, in the digital human scene live broadcast, a virtual camera is usually used to create and control the scene. The virtual camera allows precise control of the viewing angle, including the angle of the lens, focal length, movement trajectory, etc. This enables some complex and dynamic camera movements and effects to be achieved in the virtual digital human live broadcast, which may be difficult to achieve in the live broadcast of real people. The manager can adjust the focus and angle of the virtual camera to highlight the content that the user is interested in, such as narrowing the camera view angle and focusing on the product in the digital human's hand or the feature display of the product; adjusting the angle to reduce the visual interference of the background area and ensure that the product details are more prominent. After the adjusted Overlap value is increased to more than 0.6 and the entropy value is decreased to 1.8, it indicates that the audience's attention is more consistent with the expected area, and this indicator means that the user's attention is more concentrated. At this time, the manager should maintain the focus and angle of the virtual camera.
[0120] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the above embodiments do not limit the present invention in any form. Any technical solutions obtained by using equivalent replacements or equivalent transformations fall within the protection scope of the present invention.
Claims
1. A digital human live broadcast attention assessment system based on multi-modal physiological feature fusion, characterized in that: include: The multimodal physiological feature fusion model building module is used to build a deep learning model based on the Transformer architecture based on historical digital human live broadcast videos and the EEG data and eye movement data of users watching videos; The digital human live broadcast audience attention assessment module is used to input the digital human live broadcast video into the multimodal physiological feature fusion model, output the changes in the audience's attention distribution when watching the video in real time, and evaluate the audience's attention overlap and focus based on the changes in attention distribution, thereby adjusting the virtual camera angle to optimize the digital human live broadcast.
2. The digital human live broadcast attention assessment system based on multi-modal physiological feature fusion according to claim 1 is characterized in that: The multimodal physiological feature fusion model building module includes: A multimodal data collection unit, used to collect EEG data, eye movement data and corresponding mobile video when users watch the live broadcast of digital humans; A synchronization processing unit is used to synchronize the eye movement data with the mobile video to ensure the timing consistency of the data; The feature matching unit is used to extract the key points and their descriptors of the synchronized video frames, perform feature matching, calculate the similarity between the key point descriptors, filter outliers and construct a homography transformation matrix.
3. The digital human live broadcast attention assessment system based on multi-modal physiological feature fusion according to claim 2 is characterized in that: The feature matching unit calculates the similarity between key point descriptors through Euclidean distance and uses a random sampling consistency algorithm to filter outliers.
4. The digital human live broadcast attention assessment system based on multi-modal physiological feature fusion according to claim 2 is characterized in that: The multimodal physiological feature fusion model building module also includes a gaze point mapping unit, which is used to map the eye movement data of the first-person perspective into the mobile video, obtain the mapped gaze point coordinates, and realize the coordinate conversion through the homography transformation matrix.
5. The digital human live broadcast attention assessment system based on multi-modal physiological feature fusion according to claim 4 is characterized in that: The multimodal physiological feature fusion model construction module also includes a gaze point aggregation unit, which is used to clean and aggregate the mapped gaze points, generate a continuous saliency map for each frame, and perform convolution processing on the gaze data through a Gaussian mask.
6. The digital human live broadcast attention assessment system based on multi-modal physiological feature fusion according to claim 2 is characterized in that: The multimodal physiological feature fusion model construction module also includes a visual feature map extraction unit, which captures cross-frame attention migration information through Video SwinTransformer and extracts four-scale visual features F1, F2, F3, and F4.
7. The digital human live broadcast attention assessment system based on multi-modal physiological feature fusion according to claim 2 is characterized in that: The multimodal physiological feature fusion model construction module also includes an EEG feature map extraction unit, which extracts the alpha wave feature of the EEG signal through wavelet transform and fuses it into the visual feature F1 to form a fused feature.
8. The digital human live broadcast attention assessment system based on multi-modal physiological feature fusion according to claim 2 is characterized in that: The multimodal physiological feature fusion model construction module also includes a global feature enhancement module, which is used to enhance the spatial features and semantic features of the fusion features, including spatiotemporal feature enhancement and channel feature enhancement; The multimodal physiological feature fusion model building module also includes a selective feature fusion module for fusing features of different scales and extracting key gaze transfer information.
9. The digital human live broadcast attention assessment system based on multi-modal physiological feature fusion according to claim 2 is characterized in that: The multimodal physiological feature fusion model construction module also includes a progressive saliency prediction module, which is used to use gaze offset information and high-level semantic information to guide low-level appearance information, refine salient areas, and optimize model performance through a loss function to generate a saliency map S.
10. The digital human live broadcast attention assessment system based on multi-modal physiological feature fusion according to claim 1 is characterized in that: The digital human live broadcast audience attention assessment module includes: The output module is used to predict the user's attention heat map through the trained multimodal physiological feature fusion model; An evaluation module is used to evaluate the rationality of attention allocation by calculating the overlap and entropy value between the attention heat map and the target product area; The visual focus adjustment unit is used to adjust the virtual camera viewing angle according to the evaluation results, optimize the live broadcast content, and ensure that the user's attention is focused on key areas.