EEG signal-assisted video experience quality evaluation system

The video experience quality evaluation system assisted by EEG signals uses an audio-free video quality measurement model established through three-stage training to solve the problem that traditional methods cannot capture users' subjective perceptions, realize efficient and automated video experience quality prediction and evaluation, and guide communication systems to optimize resource allocation.

CN120529067BActive Publication Date: 2025-09-23TSINGHUA UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511007571.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-09-23
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

Traditional video quality evaluation methods cannot effectively capture users' subjective perceptions, resulting in high costs, long time consumption, poor scalability, and lack of real-time performance.

Method used

A video experience quality evaluation system assisted by EEG signals is adopted. The user's real subjective quality score and EEG signals are obtained through the training sample generation module. The EEG feature learning sub-model and the video feature learning sub-model are used for three-stage training to establish an audioless video quality measurement model and extract video quality features that are consistent with human subjective perception.

Benefits of technology

It achieves efficient and automated video experience quality prediction and evaluation, which can reflect the user's real feelings, guide the communication system to perform coding optimization and resource allocation, reduce bit rate and save bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120529067B_ABST
    Figure CN120529067B_ABST
Patent Text Reader

Abstract

The present application provides a video experience quality evaluation system assisted by EEG signals, which relates to the field of video transmission technology. The system includes: a model training module, which performs three-stage training to obtain a trained audioless video quality measurement model; the first stage: based on the EEG signals of multiple annotated audioless video materials, the EEG feature learning sub-model in the audioless video quality measurement model is trained; the second stage: using the EEG features output by the EEG feature learning sub-model as the true value, the video feature encoder in the video feature learning sub-model in the audioless video quality measurement model is preliminarily trained; the third stage: based on multiple annotated audioless video materials, on the basis of the preliminarily trained video feature encoder, the video feature learning sub-model in the audioless video quality measurement model is trained. Through this application, efficient prediction and evaluation of audioless video experience quality that can reflect the user's true feelings is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of video transmission technology, and in particular to an EEG signal-assisted video experience quality evaluation system. Background Art

[0002] With the widespread adoption of various multimedia communication services, users are increasingly demanding higher quality and user experience for high-definition video content. Evaluating video data quality can effectively guide communication systems in coding optimization and resource allocation, reducing bitrates and conserving bandwidth while ensuring user experience.

[0003] However, traditional objective evaluation methods rely primarily on the calculation of objective metrics, such as the structural similarity index and peak signal-to-noise ratio, which fail to capture users' subjective perception of video quality. Subjective evaluation methods also require significant manpower and time to complete the scoring process, resulting in high costs, time-consuming, poor scalability, and insufficient real-time performance. Therefore, an efficient and automated video data quality of experience evaluation solution is urgently needed. Summary of the Invention

[0004] The embodiments of the present application provide an EEG signal-assisted video experience quality evaluation system, aiming to overcome the above-mentioned problems or at least partially solve the above-mentioned problems.

[0005] A first aspect of an embodiment of the present application provides an EEG signal-assisted video quality of experience evaluation system, comprising:

[0006] a training sample generation module, configured to play audioless video materials of various video types and at multiple video distortion levels, obtain users' real subjective quality ratings and users' electroencephalogram (EEG) signals, and thereby obtain a plurality of annotated audioless video materials and the EEG signals of each of the plurality of annotated audioless video materials;

[0007] The model training module is configured to perform three stages of training to obtain a trained audioless video quality measurement model; wherein the first stage is to train an EEG feature learning sub-model in the audioless video quality measurement model based on the EEG signals of each of the plurality of labeled audioless video materials;

[0008] The second stage is: taking the EEG features output by the EEG feature learning sub-model as the true value, the video feature encoder in the video feature learning sub-model in the audioless video quality measurement model is preliminarily trained; the third stage is: based on the multiple labeled audioless video materials, on the basis of the preliminarily trained video feature encoder, the video feature learning sub-model in the audioless video quality measurement model is trained.

[0009] In an optional embodiment, the training sample generation module at least includes:

[0010] The subjective quality score annotation submodule is used to play the audio-free video material of the jth video distortion level of the i-th video type for M rounds;

[0011] Obtain the subjective quality score of the audio-free video material of the kth user for the jth video distortion level of the ith video type in the mth round;

[0012] Averaging across rounds yields the average subjective quality score of the k-th user for the audio-free video material at the j-th video distortion level for the i-th video type.

[0013] Averaging from the user dimension, obtaining the true subjective quality score of the audioless video material of the jth video distortion level of the i-th video type, to obtain the multiple labeled audioless video materials;

[0014] The value of m ranges from 1 to M, the value of i ranges from 1 to I, the value of j ranges from 1 to J, and the value of k ranges from 1 to K.

[0015] In an optional embodiment, the training sample generation module at least includes:

[0016] The EEG signal processing submodule is used to obtain the EEG signal of the kth user when watching the audio-free video material of the i-th video type and the j-th video distortion level in the m-th round; average the EEG signal of the kth user watching the audio-free video material of the i-th video type and the j-th video distortion level from the round dimension; and average the EEG signal of the audio-free video material of the i-th video type and the j-th video distortion level from the user dimension to obtain the EEG signal of each of the multiple labeled audio-free video materials.

[0017] In an optional embodiment, the process of the model training module executing the first stage includes:

[0018] According to the video distortion level, labeling the plurality of labeled audioless video materials with video quality labels, wherein the video quality labels are high-quality labels or low-quality labels;

[0019] Inputting the EEG signals of the audioless video materials labeled with multiple video quality categories into the EEG feature learning sub-model to be trained;

[0020] Extracting the EEG features of the audioless video materials labeled with multiple video quality categories through the EEG feature encoder in the EEG feature learning sub-model to be trained;

[0021] Processing the EEG features of the audioless video materials labeled with the plurality of video quality categories through the fully connected layer in the EEG feature learning sub-model to be trained to obtain the video quality category of the audioless video materials labeled with the plurality of video quality categories;

[0022] According to the video quality categories of the audioless video materials after the multiple video quality categories are marked and the video quality labels they carry, the EEG feature learning sub-model to be trained is trained to obtain the EEG feature learning sub-model in the audioless video quality measurement model.

[0023] In an optional embodiment, the process of the model training module performing the second stage includes:

[0024] Performing EEG feature extraction on the EEG signals of the plurality of annotated audioless video materials using the EEG feature encoder to obtain EEG features corresponding to the plurality of annotated audioless video materials and using the EEG features as true values;

[0025] Performing video feature extraction on the plurality of annotated audioless video materials using a video feature encoder in the video feature learning sub-model to be trained to obtain first video features corresponding to each of the plurality of annotated audioless video materials;

[0026] The EEG features output by the EEG feature learning sub-model are used as true values ​​to perform preliminary training on the video feature encoder in the video feature learning sub-model in the audioless video quality measurement model, including:

[0027] Taking the EEG features corresponding to each of the multiple labeled audioless video materials as true values, the video feature encoder in the video feature learning sub-model in the audioless video quality measurement model is preliminarily trained, so that the first video feature extracted by the preliminarily trained video feature encoder is aligned with the EEG feature extracted by the EEG feature encoder.

[0028] In an optional embodiment, the EEG features corresponding to the plurality of labeled audioless video materials are used as true values ​​to perform preliminary training on a video feature encoder in a video feature learning sub-model in the audioless video quality measurement model, including:

[0029] determining an alignment loss based on the EEG features corresponding to the plurality of annotated audioless video materials and the first video features corresponding to the plurality of annotated audioless video materials;

[0030] The alignment loss is used to perform preliminary parameter update on the video feature encoder in the video feature learning sub-model to be trained, so that the first video feature extracted by the preliminarily trained video feature encoder is aligned with the EEG feature extracted by the EEG feature encoder.

[0031] In an optional embodiment, the process of the model training module executing the third stage includes:

[0032] Inputting the plurality of annotated audioless video materials into a preliminarily trained video feature encoder to obtain second video features corresponding to each of the plurality of annotated audioless video materials;

[0033] Processing the second video features corresponding to the plurality of annotated audioless video materials through a fully connected layer in the video feature learning sub-model to be trained to obtain a video quality score for each of the plurality of annotated audioless video materials;

[0034] The video feature learning sub-model to be trained is trained according to the video quality scores and the carried annotations of the multiple annotated audioless video materials to obtain the video feature learning sub-model in the audioless video quality measurement model.

[0035] In an optional embodiment, the system further includes an electroencephalogram (EEG) analysis module, which is configured to:

[0036] Traverse each candidate duration in the candidate duration range, and continuously play the audio-free video with the high-quality label and the audio-free video with the low-quality label for the traversed candidate duration;

[0037] At each candidate duration, analyze the EEG of users watching audio-free videos with high-quality labels, and analyze the EEG of users watching audio-free videos with low-quality labels to determine the target duration;

[0038] The EEG of a user watching an audio-free video with a high-quality label and the EEG of a user watching an audio-free video with a low-quality label are compared at the target duration to determine the target EEG signal channel for training the EEG feature learning sub-model.

[0039] In an optional embodiment, the EEG feature encoder in the EEG feature learning sub-model uses an improved EEGNet network, which includes: a convolutional layer, a depthwise separable convolutional layer, a separable convolutional layer, and a fully connected layer. The improved EEGNet network models the spatiotemporal relationship of EEG signals through three layers of convolution to extract high-level features associated with human subjective cognition;

[0040] The video feature encoder in the video feature learning sub-model includes a convolutional neural network backbone and a Transformer module. The video feature encoder is used to extract low-level features including clarity, brightness, and color.

[0041] In an optional embodiment, the system further includes:

[0042] The audioless video quality measurement module is used to input the target audioless video into the video feature learning sub-model in the audioless video quality measurement model to obtain the quality score of the target audioless video.

[0043] In the EEG signal-assisted video experience quality evaluation system provided in the present application, the training sample generation module plays audio-free video materials of multiple video types and multiple video distortion levels to obtain the user's real subjective quality score and the user's EEG signal to obtain multiple labeled audio-free video materials and their respective EEG signals; the model training module performs three stages of training to obtain a trained audio-free video quality measurement model; the first stage is: based on the EEG signals of each of the multiple labeled audio-free video materials, the EEG feature learning sub-model in the audio-free video quality measurement model is trained; the second stage is: using the EEG features output by the EEG feature learning sub-model as the true value, the video feature encoder in the video feature learning sub-model in the audio-free video quality measurement model is preliminarily trained; the third stage is: based on the preliminarily trained video feature encoder, based on the multiple labeled audio-free video materials, the video feature learning sub-model in the audio-free video quality measurement model is trained.

[0044] In this way, this application introduces advanced cognitive information of EEG characteristics of EEG signals into the audio-free video quality measurement model for assistance, so that the neural network can fit the characteristics of human cognition, and realize efficient prediction and evaluation of audio-free video experience quality that can reflect the user's real feelings, which helps to guide the communication system to perform coding optimization and resource allocation, reduce the bit rate and save bandwidth while ensuring user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0046] Figure 1 This is a structural diagram of a video quality of experience evaluation system assisted by EEG signals proposed in one embodiment of the present application;

[0047] Figure 2 This is a schematic diagram of the process of collecting users' real subjective quality scores and EEG signals in the EEG signal-assisted video experience quality evaluation system proposed in one embodiment of the present application;

[0048] Figure 3 This is a schematic diagram showing the distribution of average subjective scores of multiple types of audioless video materials at different video distortion levels in the EEG signal-assisted video quality of experience evaluation system proposed in one embodiment of the present application;

[0049] Figure 4 This is a schematic diagram of the prediction results of the audioless video quality measurement model of the EEG signal-assisted video experience quality evaluation system proposed in one embodiment of the present application;

[0050] Figure 5 This is a boxplot of EEG feature extraction of the EEG feature learning sub-model of the audio-free video quality measurement model of the EEG signal-assisted video experience quality evaluation system proposed in one embodiment of the present application;

[0051] Figure 6 This is a schematic diagram of the network structure of the audio-free video quality measurement model of the EEG signal-assisted video experience quality evaluation system proposed in one embodiment of the present application;

[0052] Figure 7 The EEG of a user of the EEG signal-assisted video experience quality evaluation system proposed in one embodiment of the present application watching an audio-free video with a high-quality label;

[0053] Figure 8 This is the EEG of a user of the EEG signal-assisted video experience quality evaluation system proposed in one embodiment of the present application watching an audio-free video with a low-quality label. DETAILED DESCRIPTION

[0054] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0055] In the drawings, the sizes of components, layer thicknesses, or regions may be exaggerated for clarity. Therefore, any implementation of the present disclosure is not necessarily limited to the dimensions shown in the drawings, and the shapes and sizes of components in the drawings do not reflect true proportions. Furthermore, the drawings schematically illustrate idealized examples, and any implementation of the present disclosure is not limited to the shapes or values ​​shown in the drawings.

[0056] Video quality assessment technology is rapidly developing. In many scenarios, such as surveillance video, audio-free video quality assessment is required to address numerous downstream tasks. However, current video quality of experience assessments often rely on information from a single modality, video. The resulting neural networks may extract low-level visual features like color and brightness, while ignoring the more important user's subjective perception. This application considers electroencephalogram (EEG) signals, or EEG signals, whose characteristics are closely related to the viewer's subjective perception.

[0057] Therefore, this application proposes an EEG signal-assisted video experience quality evaluation system, which uses EEG signals as an aid to automatically predict and evaluate the quality of audio-free video experience, improves the prediction efficiency of experience quality, and helps guide the communication system to perform coding optimization and resource allocation, thereby reducing the bit rate and saving bandwidth while ensuring user experience.

[0058] Reference Figure 1 , Figure 1 This is a schematic diagram of the structure of the video experience quality evaluation system assisted by EEG signals proposed in one embodiment of the present application. Figure 1 As shown, the system includes: a training sample generation module 101 and a model training module 102.

[0059] The training sample generation module 101 is used to play audio-free video materials of multiple video types and multiple video distortion levels, obtain the user's real subjective quality score and the user's electroencephalogram signal, so as to obtain multiple labeled audio-free video materials and the electroencephalogram signals of each of the multiple labeled audio-free video materials.

[0060] In this embodiment, audioless video materials of multiple video types and multiple video distortion levels are first obtained. Specifically, multiple segments of original audioless video stimulation materials of a certain length can be selected or cut out from public websites, documentaries, etc. The multiple segments of original audioless video stimulation materials have diverse content and rich scenes, covering multiple types such as daytime and nighttime scenes, animals, human activity scenes, and environments without specific target objects, so as to help improve the accuracy and robustness of video quality prediction. For the selected audioless video materials of multiple video types, the distortion level of the video is adjusted by changing the constant quality factor (CRF). Video encoding tools (such as FFmpeg) can be used to process the audioless video materials of multiple video types at multiple video distortion levels according to different pre-selected CRF values ​​to obtain audioless video materials of multiple video types and multiple video distortion levels.

[0061] For example, for the 20 categories of audioless video materials with different contents selected, the duration of each video stimulus is 3 seconds; six CRF values ​​of 1, 20, 30, 42, 47 and 51 are selected, among which CRF=1 indicates lossless compression and the best video quality, and CRF=51 indicates the highest compression ratio and the worst video quality; each category of audioless video material is processed with 6 distortion levels, and finally 120 segments (20 categories × 6 CRF levels) of audioless video stimulus materials can be obtained.

[0062] In this embodiment, Figure 2 As shown, the training sample generation module plays audioless video clips of each video type at multiple levels of video distortion in each trial. Each trial consists of three phases: a gray screen phase (G1), a video clip playback phase (V), and another gray screen feedback phase (G2). The gray screen phase G1 lasts for a preset duration (e.g., 500 milliseconds) to stabilize the user's state. Subsequently, phase V plays video clips for a certain duration to stimulate the user's senses. Finally, the gray screen feedback phase G2 displays subtitles prompting the user to rate the video clip's quality according to a rating scale (e.g., 1 to 5, with 1 representing the worst and 5 representing the best). The feedback phase G2 continues until the user clicks a button to rate, obtaining the user's true subjective quality rating. For each trial, while the user watches the clip, an EEG acquisition device with a sampling rate of 1000 Hz and a total of 64 channels can be used to record the user's EEG activity, generating the user's EEG signals. The user's actual subjective quality ratings for each audioless video clip of each video type and each distortion level are associated with the corresponding video clips and annotated to obtain multiple annotated audioless video clips. Simultaneously, the user's collected EEG signals are preprocessed to obtain the EEG signals for each of the multiple annotated audioless video clips.

[0063] In this embodiment, to improve the quality of EEG signal samples in this experiment, the collected raw EEG signals undergo a series of preprocessing steps. Specifically, data import: import EEG recording data and electrode location information. Re-referencing: re-reference the EEG signals using a global average reference to reduce noise interference. Invalid channel removal: remove invalid channels from the EEG signals to improve overall signal quality. Filtering: bandpass filter the EEG signals with a 1-40 Hz band using a finite impulse response (FIR) filter. Time windowing: starting from the moment the target video stimulus is presented, the EEG signals are clipped to 3000 milliseconds after the stimulus as valid samples. ICA (Independent Component Analysis) processing: perform independent component analysis on the EEG signals to separate artifacts. ADJUST (Artifact Detection and Joint Subtraction Technique) processing: use the ADJUST plug-in to automatically identify and remove eye movement and muscle movement interference components.

[0064] In this embodiment, in order to verify the effectiveness of collecting real subjective quality scores, the distribution of the collected subjective quality scores is also analyzed. Figure 3 shown. Figure 3 The distribution of average subjective scores of various audio-free video materials at different video distortion levels is shown, clearly reflecting the significant differences between different video distortion levels. Figure 3 It can be seen that the average subjective score gradually decreases with the increase of video distortion level, indicating that the obtained true subjective quality score can effectively capture the experience quality degradation caused by the increase of distortion, and that the true subjective quality scoring process has high reliability.

[0065] The model training module 102 is used to perform three stages of training to obtain a trained audio-free video quality measurement model; wherein, the first stage is: based on the EEG signals of each of the multiple labeled audio-free video materials, training is performed to obtain an EEG feature learning sub-model in the audio-free video quality measurement model; the second stage is: using the EEG features output by the EEG feature learning sub-model as the true value, preliminary training is performed on the video feature encoder in the video feature learning sub-model in the audio-free video quality measurement model; the third stage is: based on the multiple labeled audio-free video materials, on the basis of the preliminary trained video feature encoder, training is performed to obtain the video feature learning sub-model in the audio-free video quality measurement model.

[0066] In this embodiment, the audioless video quality measurement model includes: an EEG feature learning sub-model and a video feature learning sub-model, and the video feature learning sub-model includes a video feature encoder. The model training module obtains a trained audioless video quality measurement model through three stages of training. In the first stage, based on the EEG signals of each of a plurality of labeled audioless video materials, the EEG feature learning sub-model in the audioless video quality measurement model is trained. The trained EEG feature learning sub-model can extract effective features related to video quality from the EEG signals, and provide EEG features aligned with subjective perception for subsequent stages as supervision signals for video feature alignment. In the second stage, the EEG features extracted by the EEG feature learning sub-model after training in the first stage are used as true values ​​to perform preliminary training on the video feature encoder in the video feature learning sub-model in the audioless video quality measurement model, so that the video features extracted by it are aligned with the EEG features, which helps the audioless video quality measurement model to extract video quality features consistent with human subjective perception. In the third stage, based on multiple annotated audioless video materials and the initially trained video feature encoder, the video feature learning sub-model in the audioless video quality measurement model is trained. At this point, the audioless video quality measurement model training is completed. The audioless video quality measurement model is used to output a video quality score that can reflect the user's true perception based on the input audioless video. Figure 4 , Figure 4 The results of predicting the quality scores of different audioless videos using the audioless video quality measurement model are shown.

[0067] In the EEG signal-assisted video experience quality evaluation system provided in the present application, the model training module obtains a trained audioless video quality measurement model through three stages of training; the first stage is: based on the EEG signals of multiple labeled audioless video materials, the EEG feature learning sub-model in the audioless video quality measurement model is trained; the second stage is: using the EEG features output by the EEG feature learning sub-model as the true value, the video feature encoder in the video feature learning sub-model in the audioless video quality measurement model is preliminarily trained to align the extracted video features with the EEG features, which helps the audioless video quality measurement model to extract video quality features consistent with human subjective perception; the third stage is: based on multiple labeled audioless video materials, on the basis of the preliminarily trained video feature encoder, the video feature learning sub-model in the audioless video quality measurement model is trained. In this way, this application introduces advanced cognitive information of EEG characteristics of EEG signals into the audio-free video quality measurement model for assistance, so that the neural network can fit the characteristics of human cognition, and realize efficient and automated prediction and evaluation of audio-free video experience quality that can reflect the user's real feelings, which helps to guide the communication system to perform coding optimization and resource allocation in real time, reduce the bit rate and save bandwidth while ensuring user experience.

[0068] In one embodiment, the present application further provides an EEG signal-assisted video quality of experience evaluation system, in which the model training module performs the first phase of the process including:

[0069] According to the video distortion level, labeling the plurality of labeled audioless video materials with video quality labels, wherein the video quality labels are high-quality labels or low-quality labels;

[0070] Inputting the EEG signals of the audioless video materials labeled with multiple video quality categories into the EEG feature learning sub-model to be trained;

[0071] Extracting the EEG features of the audioless video materials labeled with multiple video quality categories through the EEG feature encoder in the EEG feature learning sub-model to be trained;

[0072] Processing the EEG features of the audioless video materials labeled with the plurality of video quality categories through the fully connected layer in the EEG feature learning sub-model to be trained to obtain the video quality category of the audioless video materials labeled with the plurality of video quality categories;

[0073] According to the video quality categories of the audioless video materials after the multiple video quality categories are marked and the video quality labels they carry, the EEG feature learning sub-model to be trained is trained to obtain the EEG feature learning sub-model in the audioless video quality measurement model.

[0074] In this embodiment, according to the video distortion level, video quality labels are respectively labeled for multiple audioless video materials with video quality labels, and multiple audioless video materials labeled with video quality categories are obtained. The video quality label can be a high-quality label or a low-quality label. For example, the EEG signals corresponding to video materials with distortion levels of 1, 2, and 3 are labeled as "high quality", and the EEG signals corresponding to video materials with distortion levels of 4, 5, and 6 are labeled as "low quality" for subsequent classification training of the EEG feature learning sub-model. The EEG feature learning sub-model to be trained is composed of an EEG feature encoder and a fully connected layer. The EEG signals of each of the audioless video materials labeled with multiple video quality categories are input into the EEG feature learning sub-model to be trained, and the EEG features of each of the audioless video materials labeled with multiple video quality categories are extracted through the EEG feature encoder in the EEG feature learning sub-model to be trained. The extracted EEG features can be represented as follows:

[0075]

[0076] Among them, F E Indicates EEG characteristics; represents the EEG feature encoder; represents the input EEG signal; c represents the channel sequence; t represents the time series.

[0077] In this embodiment, the EEG features of the EEG signal after multiple video quality category labeling of the audio-free video material are processed through the fully connected layer in the EEG feature learning sub-model to be trained. After passing through the fully connected layer, the EEG features need to be passed through the sigmoid function to predict the video quality category of the audio-free video material after multiple video quality category labeling, and then 0.5 is used as the boundary to judge the predicted category. That is, the output result R of the EEG feature learning sub-model is E for:

[0078] ;

[0079] in, is the fully connected layer, F E is the EEG characteristic. E >0.5, then predict the label is "high quality"; otherwise, the predicted label is “low quality”.

[0080] In this embodiment, based on the video quality categories of the audioless video material after multiple video quality categories of the EEG signal are labeled and the EEG signal video quality labels it carries, the binary cross-entropy loss (BCE Loss) of the model classification training is calculated. The parameters of the EEG feature learning sub-model to be trained are updated according to the calculated loss to obtain a trained EEG feature learning sub-model. The trained EEG feature learning sub-model can extract effective features related to video quality from the EEG signal, and can predict and output the video quality category of the audioless video material based on the EEG signal corresponding to the audioless video material.

[0081] The binary cross entropy loss is calculated using the following formula:

[0082]

[0083] Among them, L BCE is the binary cross entropy loss; is the predicted video quality category (predicted label); y is the actual video quality label (actual label).

[0084] In this embodiment, in order to verify the effectiveness of EEG feature extraction, 100 EEG samples corresponding to undistorted videos and 100 EEG samples corresponding to high-distortion videos are input into the EEG feature learning sub-model, and the output distribution is visualized through a box plot, as shown in Figure 2. Figure 5 The results show that there is a clear separation between the distributions of the two types of samples, and the decision boundary is located near 0.5, indicating that the EEG feature learning sub-model is effective in extracting EEG features related to high / low quality visual stimuli.

[0085] In one embodiment, the present application further provides an EEG signal-assisted video quality of experience evaluation system, in which the model training module performs the second phase of the process including:

[0086] Performing EEG feature extraction on the EEG signals of the plurality of annotated audioless video materials using the EEG feature encoder to obtain EEG features corresponding to the plurality of annotated audioless video materials and using the EEG features as true values;

[0087] Performing video feature extraction on the plurality of annotated audioless video materials using a video feature encoder in the video feature learning sub-model to be trained to obtain first video features corresponding to each of the plurality of annotated audioless video materials;

[0088] The EEG features output by the EEG feature learning sub-model are used as true values ​​to perform preliminary training on the video feature encoder in the video feature learning sub-model in the audioless video quality measurement model, including:

[0089] Taking the EEG features corresponding to each of the multiple labeled audioless video materials as true values, the video feature encoder in the video feature learning sub-model in the audioless video quality measurement model is preliminarily trained, so that the first video feature extracted by the preliminarily trained video feature encoder is aligned with the EEG feature extracted by the EEG feature encoder.

[0090] In this embodiment, after completing the training of the EEG feature learning sub-model in the first stage, the relevant network parameters are frozen, and then the EEG signals of the multiple labeled audio-free video materials are input into the EEG feature learning sub-model, and the EEG features corresponding to the multiple labeled audio-free video materials are obtained through the EEG feature encoder. The video feature learning sub-model to be trained is composed of a video feature encoder and a fully connected layer. The video feature encoder in the video feature learning sub-model to be trained is used to extract video features of the multiple labeled audio-free video materials to obtain the first video features corresponding to the multiple labeled audio-free video materials. The EEG features corresponding to the multiple labeled audio-free video materials are used as the true value, and the video feature encoder in the video feature learning sub-model to be trained is preliminarily trained so that the first video features extracted by the preliminarily trained video feature encoder are aligned with the EEG features extracted by the EEG feature encoder. The EEG features and video features are aligned in the second stage, so that the video features imply human subjective perception, which helps the audio-free video quality measurement model to extract video quality features consistent with human subjective perception.

[0091] In one embodiment, the present application further provides an EEG signal-assisted video experience quality evaluation system. In the system, the EEG features corresponding to the plurality of annotated audioless video materials are used as true values ​​to perform preliminary training on the video feature encoder in the video feature learning sub-model in the audioless video quality measurement model, including:

[0092] determining an alignment loss based on the EEG features corresponding to the plurality of annotated audioless video materials and the first video features corresponding to the plurality of annotated audioless video materials;

[0093] The alignment loss is used to perform preliminary parameter update on the video feature encoder in the video feature learning sub-model to be trained, so that the first video feature extracted by the preliminarily trained video feature encoder is aligned with the EEG feature extracted by the EEG feature encoder.

[0094] In this embodiment, an alignment loss is determined based on the EEG features corresponding to each of the multiple annotated audioless video materials and the first video features corresponding to each of the multiple annotated audioless video materials. The alignment loss is used to perform a preliminary parameter update on the video feature encoder in the video feature learning sub-model to be trained, so that the first video features extracted by the preliminarily trained video feature encoder are aligned with the EEG features extracted by the EEG feature encoder. The alignment loss is calculated using the following loss function:

[0095]

[0096] Among them, L C is the alignment loss of the second stage; F V is the first video feature; F E For EEG characteristics.

[0097] In one embodiment, the present application further provides an EEG signal-assisted video quality of experience evaluation system, in which the model training module performs the third stage of the process including:

[0098] Inputting the plurality of annotated audioless video materials into a preliminarily trained video feature encoder to obtain second video features corresponding to each of the plurality of annotated audioless video materials;

[0099] Processing the second video features corresponding to the plurality of annotated audioless video materials through a fully connected layer in the video feature learning sub-model to be trained to obtain a video quality score for each of the plurality of annotated audioless video materials;

[0100] The video feature learning sub-model to be trained is trained according to the video quality scores and the carried annotations of the multiple annotated audioless video materials to obtain the video feature learning sub-model in the audioless video quality measurement model.

[0101] In this embodiment, the third stage is the final training stage of the audioless video quality measurement model. The parameters trained in the second stage are used as the initial parameters, and the real subjective quality scores are used as labels to retrain the audioless video quality measurement model. Specifically, multiple annotated audioless video materials are input into the video feature encoder that has been preliminarily trained to obtain the second video features corresponding to the multiple annotated audioless video materials; the second video features corresponding to the multiple annotated audioless video materials are processed through the fully connected layer in the video feature learning sub-model to be trained, and the video quality scores of the multiple annotated audioless video materials are predicted. According to the video quality scores of the multiple annotated audioless video materials and the carried annotations (real subjective quality scores), the loss of the video feature learning sub-model to be trained is calculated, and the parameters of the video feature learning sub-model to be trained are updated according to the calculated loss to obtain the trained video feature learning sub-model. At this point, the training of the audioless video quality measurement model is completed. The loss is calculated by the following formula:

[0102]

[0103] Among them, L MSE is the loss in the third stage; X p is the predicted quality score; X m Rating for true subjective quality.

[0104] In one embodiment, the present application also provides an EEG signal-assisted video quality of experience assessment system, in which the EEG feature encoder in the EEG feature learning sub-model uses an improved EEGNet network, which includes: a convolutional layer, a depthwise separable convolutional layer, a separable convolutional layer, and a fully connected layer. The improved EEGNet network uses three layers of convolution to model the spatiotemporal relationship of EEG signals to extract high-level features associated with human subjective cognition;

[0105] The video feature encoder in the video feature learning sub-model includes a convolutional neural network backbone and a Transformer module. The video feature encoder is used to extract low-level features including clarity, brightness, and color.

[0106] In this embodiment, the EEG feature learning sub-model consists of an EEG feature encoder and a fully connected layer. The EEG feature encoder uses an improved EEGNet (electroencephalogram network) network, including convolutional layers, depthwise separable convolutional layers, separable convolutional layers, and fully connected layers, forming a deeper spatiotemporal feature extraction layer. By modeling the spatiotemporal relationship of EEG signals through three layers of convolution, high-level features related to human subjective cognition can be more effectively extracted. The specific network structure of the EEG feature learning sub-model is shown in Table 1:

[0107] Table 1. Network structure of the EEG feature learning sub-model

[0108]

[0109] In this embodiment, the video feature learning sub-model consists of a video encoder and a fully connected layer. The video encoder consists of a Convolutional Neural Network (CNN) backbone and a Transformer (self-attention) module. The CNN backbone is used to extract the spatial features of the video, including features such as clarity, brightness, and color. It consists of five convolutional layers with ReLU activation functions. Each convolutional layer uses a convolution kernel of size 3×3, a step size of 2, and uses maximum pooling and batch normalization operations after convolution. The Transformer module is used to extract the temporal dependency characteristics of the video. It consists of two Transformer encoder layers, each encoder layer contains 8 attention heads, and the hidden dimension is 512. The position encoding in the Transformer layer can ensure that the model recognizes the dependencies between sequences. In addition, the fully connected layer is used to transform the shape of the tensor result to obtain the prediction score (the quality score of the video). The extracted video features can be represented as follows:

[0110]

[0111] in, represents a video feature, such as the first video feature; V (x, y, t) represents an input video clip; x, y represent two-dimensional coordinates; t represents a time series.

[0112] In one embodiment, the present application also provides an EEG signal-assisted video experience quality evaluation system, in which the system at least includes: an audio-free video quality measurement module.

[0113] The audioless video quality measurement module is used to input the target audioless video into the video feature learning sub-model in the audioless video quality measurement model to obtain the quality score of the target audioless video.

[0114] In this embodiment, when it is necessary to perform a quality evaluation on the target audioless video, the target audioless video is input into the audioless video quality measurement model through the audioless video quality measurement module, and the video feature learning sub-model in the audioless video quality measurement model can extract video features that are consistent with human subjective perception, and then the quality score of the target audioless video that reflects the user's true feelings can be output. When it is necessary to perform a quality evaluation on the target audioless video for the target user, the target audioless video and the EEG signals of the user during the process of watching the target audioless video can be respectively input into the video feature learning sub-model and the EEG feature learning sub-model in the audioless video quality measurement model to obtain the quality score and video quality category of the target audioless video for the target user. The quality evaluation of video data performed by this application can effectively guide the communication system to perform coding optimization and resource allocation, reduce the bit rate and save bandwidth while ensuring the user experience.

[0115] In one embodiment, if Figure 6 As shown, Figure 6 This is a schematic diagram of the network structure of an audioless video quality measurement model according to an embodiment of the present application. The audioless video quality measurement model (audioless video quality measurement network) includes an EEG feature learning sub-model and a video feature learning sub-model. The EEG feature learning sub-model consists of an EEG feature encoder and fully connected layers, where the EEG feature encoder uses a modified EEGNet network. The video feature learning sub-model consists of a video encoder and fully connected layers, where the video encoder consists of a convolutional neural network backbone and a Transformer module. The trained audioless video quality measurement model is obtained through three stages of training. Stage 1: Training of the EEG feature learning sub-model (EEG feature learning network). In this stage, EEG signals are classified into "high quality" and "low quality" categories based on the video distortion level. The EEG signals are input and preprocessed, and the EEG feature learning network is trained for binary classification based on the corresponding EEG signal labels. After training, the network parameters are frozen, and the EEG signals corresponding to each video source are input into the network, where the corresponding EEG features are obtained through the EEG feature encoder. Phase 2: EEG and video feature alignment. In this phase, the video encoder in the video feature learning sub-model is used to extract video features. The EEG features are then used as the true values. The alignment loss is calculated based on the video and EEG features. The audioless video quality measurement network is trained to align the extracted video features with the EEG features. After training, the parameters of the audioless video quality measurement network are frozen. Phase 3: Final training of the audioless video quality measurement network. In this phase, the parameters trained in Phase 2 are used as the initial parameters, and the true values ​​of the subjective quality scores are used as labels. The audioless video quality measurement network is retrained to obtain the trained audioless video quality measurement network.

[0116] In one embodiment, the present application also provides an EEG signal-assisted video experience quality evaluation system, in which the system at least includes: an EEG analysis module.

[0117] The EEG analysis module is used to:

[0118] Traverse each candidate duration in the candidate duration range, and continuously play the audio-free video with the high-quality label and the audio-free video with the low-quality label for the traversed candidate duration;

[0119] At each candidate duration, analyze the EEG of users watching audio-free videos with high-quality labels, and analyze the EEG of users watching audio-free videos with low-quality labels to determine the target duration;

[0120] The EEG of a user watching an audio-free video with a high-quality label and the EEG of a user watching an audio-free video with a low-quality label are compared at the target duration to determine the target EEG signal channel for training the EEG feature learning sub-model.

[0121] In this embodiment, the EEG response is analyzed by the EEG analysis module to determine the target duration of the video playback and select the target EEG signal channel corresponding to the main area where the EEG response occurs. Specifically, each candidate duration in the candidate duration range (such as the candidate durations in the range of 200ms-400ms: 200ms, 300ms, 400ms) is traversed, and the audioless video with high-quality labels and the audioless video with low-quality labels are continuously played for the traversed candidate durations, and the EEG of the user watching the audioless video with high-quality labels at each candidate duration is obtained, such as Figure 7 As shown in , and the EEG of the user watching the audio-free video with low-quality labels is obtained, as Figure 8 shown.

[0122] In this embodiment, a variety of analysis methods can be used to determine the target duration and the target EEG signal channel under the target duration. For example, for each candidate EEG data at the target duration, the time domain features and frequency domain features are extracted, the difference in EEG features when watching high-quality labeled videos and low-quality labeled videos is calculated, and the candidate duration with the largest difference is selected as the target duration (e.g., 300ms); under the target duration, the EEG feature differences between high-quality labeled videos and low-quality labeled videos are calculated for all EEG signal channels, and the channels with the most significant differences are selected as the target EEG signal channels for training the EEG feature learning sub-model (e.g., selecting the five channels Oz, O1, O2, POz, and Pz located at the back of the brain in the EEG signal). Therefore, multiple segments of original audioless video stimulation materials of the target duration can be selected or cropped for playback to obtain more obvious and stable EEG signals, which helps improve the training effect and generalization ability of the subsequent model; and selecting the target EEG signal channel reduces the data dimension and focuses on the brain areas related to video quality perception, which helps improve the efficiency and accuracy of the model.

[0123] In one embodiment, the present application also provides an EEG signal-assisted video experience quality evaluation system, in which the training sample generation module at least includes: a subjective quality score annotation submodule.

[0124] The subjective quality score annotation submodule is used to play the audio-free video material of the jth video distortion level of the i-th video type for M rounds;

[0125] Obtain the subjective quality score of the audio-free video material of the kth user for the jth video distortion level of the ith video type in the mth round;

[0126] Averaging across rounds yields the average subjective quality score of the k-th user for the audio-free video material at the j-th video distortion level for the i-th video type.

[0127] Averaging from the user dimension, obtaining the true subjective quality score of the audioless video material of the jth video distortion level of the i-th video type, to obtain the multiple labeled audioless video materials;

[0128] The value of m ranges from 1 to M, the value of i ranges from 1 to I, the value of j ranges from 1 to J, and the value of k ranges from 1 to K.

[0129] In this embodiment, in order to improve the reliability of the obtained user's subjective quality rating, for each video type and each video distortion level of the audio-free video material, the subjective quality rating annotation submodule collects the subjective quality ratings of multiple users in multiple rounds in each trial and averages them to obtain the user's true subjective quality rating. The subjective quality rating corresponding to the video material is recorded as , where i, j, k, and m represent the sequence number of the video material, the distortion level of the video, the sequence number of the user, and the test round, respectively. Therefore, averaging across rounds, the average subjective quality score of the k-th user for the audio-free video material of the j-th video distortion level of the i-th video type is:

[0130]

[0131] in, is the average subjective quality score of the k-th user for the audio-free video material with the j-th video distortion level for the i-th video type; To take the average from the round dimension; is the subjective quality score of the audio-free video material of the kth user for the jth video distortion level of the i-th video type in the m-th round.

[0132] Averaged from the user dimension, the real subjective quality score of the audio-free video material of the j-th video distortion level of the i-th video type is:

[0133]

[0134] in, The real subjective quality score of the audio-free video material of the jth video distortion level for the i-th video type; To take the average from the user dimension; is the average subjective quality score of the k-th user for the audio-free video material with the j-th video distortion level for the i-th video type.

[0135] In one embodiment, the present application also provides an EEG signal-assisted video experience quality evaluation system, in which the training sample generation module at least includes: an EEG signal processing submodule.

[0136] The EEG signal processing submodule is used to obtain the EEG signal of the kth user when watching the audio-free video material of the i-th video type and the j-th video distortion level in the m-th round; average the EEG signal of the kth user watching the audio-free video material of the i-th video type and the j-th video distortion level from the round dimension; and average the EEG signal of the audio-free video material of the i-th video type and the j-th video distortion level from the user dimension to obtain the EEG signal of each of the multiple labeled audio-free video materials.

[0137] In this embodiment, in order to improve the reliability of the user's EEG signal, for each video type and each video distortion level of the audio-free video material, the EEG signal processing submodule records the EEG signals of multiple users in multiple rounds in each trial, and performs average calculation to finally obtain the user's EEG signal. Similar to the calculation of the subjective quality score annotation submodule, the EEG signal processing submodule records the EEG signal corresponding to the video material as , where i, j, k, and m represent the sequence number of the video material, the distortion level of the video, the sequence number of the user, and the round of testing, respectively. c and t represent the channel sequence number and time series of the EEG signal, respectively. Therefore, after averaging from the round dimension and the user dimension, the average EEG signal of the audio-free video material of the jth video distortion level of the i-th video type is:

[0138]

[0139] in, is the EEG signal of the audio-free video material of the i-th video type and the j-th video distortion level; To take the average from the user dimension; To take the average from the round dimension; is the EEG signal of the kth user in the mth round for the audio-free video material of the jth video distortion level of the i-th video type.

[0140] In one embodiment, the present application compares the performance of an audioless video quality measurement model with other video feature extraction methods on a test set. The comparative experimental results are shown in Table 2. Among the comparison methods, this experiment used four different video feature extraction methods for performance comparison: Msm-Net, C3D, ConvLSTM, and I3D. The comparison results in Table 2 show that the video feature extraction method for the audioless video quality measurement model provided by this application has superior performance compared to other methods, and its prediction results are more strongly correlated with the true score.

[0141] Table 2. Comparative experimental results of audioless video quality measurement models

[0142]

[0143] In one embodiment, the present application conducted an ablation experiment, and the experimental results are shown in Table 3. Referring to Table 3, it can be seen that the performance changes after removing the key components in the model, where "w / o" means that the component is removed in the method. The results show that removing the CNN backbone module will lead to extremely poor fitting results, and the PLCC and SROCC are negative values, indicating that the prediction results are opposite to the true labels, which illustrates the core role of the CNN module in extracting video spatial features. Removing the Transformer module will cause the performance of this network to decline, which shows that the Transformer module is better in extracting video temporal features and can play a certain auxiliary role. In addition, removing the EEG feature learning module (EEG feature learning sub-model) will also cause the performance of this network to decline, indicating that EEG signal assistance can indeed help improve the accuracy of subjective experience quality prediction.

[0144] Table 3. Ablation experiment results of audio-free video quality measurement model

[0145]

[0146] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0147] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the embodiments of the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0148] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0149] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0150] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0151] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0152] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0153] The above is a detailed introduction to the EEG signal-assisted video experience quality evaluation system provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this application.

Claims

1. An EEG signal-assisted video experience quality evaluation system, characterized in that: The system comprises: a training sample generation module, configured to play audioless video materials of various video types and at multiple video distortion levels, obtain users' real subjective quality ratings and users' electroencephalogram (EEG) signals, and thereby obtain a plurality of annotated audioless video materials and the EEG signals of each of the plurality of annotated audioless video materials; The model training module is configured to perform three stages of training to obtain a trained audioless video quality measurement model; wherein the first stage is: based on the EEG signals of each of the plurality of annotated audioless video materials, training an EEG feature learning sub-model in the audioless video quality measurement model; the second stage is: using the EEG features output by the EEG feature learning sub-model as the true value, preliminarily training the video feature encoder in the video feature learning sub-model in the audioless video quality measurement model; the third stage is: based on the plurality of annotated audioless video materials and on the basis of the preliminarily trained video feature encoder, training the video feature learning sub-model in the audioless video quality measurement model; The process of executing the second phase of the model training module includes: Performing EEG feature extraction on the EEG signals of each of the plurality of annotated audioless video materials using the EEG feature encoder in the EEG feature learning sub-model, obtaining EEG features corresponding to each of the plurality of annotated audioless video materials and using them as true values; Performing video feature extraction on the plurality of annotated audioless video materials using a video feature encoder in the video feature learning sub-model to be trained to obtain first video features corresponding to each of the plurality of annotated audioless video materials; The EEG features output by the EEG feature learning sub-model are used as true values ​​to perform preliminary training on the video feature encoder in the video feature learning sub-model in the audioless video quality measurement model, including: Using the EEG features corresponding to the plurality of labeled audioless video materials as true values, preliminarily training a video feature encoder in a video feature learning sub-model in the audioless video quality measurement model so that a first video feature extracted by the preliminarily trained video feature encoder is aligned with an EEG feature extracted by the EEG feature encoder; Using the EEG features corresponding to the plurality of labeled audioless video materials as true values, preliminary training is performed on the video feature encoder in the video feature learning sub-model in the audioless video quality measurement model, including: determining an alignment loss based on the EEG features corresponding to the plurality of annotated audioless video materials and the first video features corresponding to the plurality of annotated audioless video materials; The alignment loss is used to perform preliminary parameter update on the video feature encoder in the video feature learning sub-model to be trained, so that the first video feature extracted by the preliminarily trained video feature encoder is aligned with the EEG feature extracted by the EEG feature encoder.

2. The EEG signal-assisted video experience quality evaluation system according to claim 1, characterized in that: The training sample generation module at least includes: The subjective quality score annotation submodule is used to play the audio-free video material of the jth video distortion level of the i-th video type for M rounds; Obtain the subjective quality score of the audio-free video material of the kth user for the jth video distortion level of the ith video type in the mth round; Averaging across rounds yields the average subjective quality score of the k-th user for the audio-free video material at the j-th video distortion level for the i-th video type. Averaging from the user dimension, obtaining the true subjective quality score of the audioless video material of the jth video distortion level of the i-th video type, to obtain the multiple labeled audioless video materials; The value of m ranges from 1 to M, the value of i ranges from 1 to I, the value of j ranges from 1 to J, and the value of k ranges from 1 to K.

3. The EEG signal-assisted video experience quality evaluation system according to claim 2, characterized in that: The training sample generation module at least includes: The EEG signal processing submodule is used to obtain the EEG signal of the kth user when watching the audio-free video material of the i-th video type and the j-th video distortion level in the m-th round; average the EEG signal of the kth user watching the audio-free video material of the i-th video type and the j-th video distortion level from the round dimension; and average the EEG signal of the audio-free video material of the i-th video type and the j-th video distortion level from the user dimension to obtain the EEG signal of each of the multiple labeled audio-free video materials.

4. The EEG signal-assisted video experience quality evaluation system according to claim 1, characterized in that: The process of executing the first phase of the model training module includes: According to the video distortion level, labeling the plurality of labeled audioless video materials with video quality labels, wherein the video quality labels are high-quality labels or low-quality labels; Inputting the EEG signals of the audioless video materials labeled with multiple video quality categories into the EEG feature learning sub-model to be trained; Extracting the EEG features of the audioless video materials labeled with multiple video quality categories through the EEG feature encoder in the EEG feature learning sub-model to be trained; Processing the EEG features of the audioless video materials labeled with the plurality of video quality categories through the fully connected layer in the EEG feature learning sub-model to be trained to obtain the video quality category of the audioless video materials labeled with the plurality of video quality categories; According to the video quality categories of the audioless video materials after the multiple video quality categories are marked and the video quality labels they carry, the EEG feature learning sub-model to be trained is trained to obtain the EEG feature learning sub-model in the audioless video quality measurement model.

5. The EEG signal-assisted video experience quality evaluation system according to claim 1, characterized in that: The process of executing the third stage of the model training module includes: Inputting the plurality of annotated audioless video materials into a preliminarily trained video feature encoder to obtain second video features corresponding to each of the plurality of annotated audioless video materials; Processing the second video features corresponding to the plurality of annotated audioless video materials through a fully connected layer in the video feature learning sub-model to be trained to obtain a video quality score for each of the plurality of annotated audioless video materials; The video feature learning sub-model to be trained is trained according to the video quality scores and the carried annotations of the multiple annotated audioless video materials to obtain the video feature learning sub-model in the audioless video quality measurement model.

6. The EEG signal-assisted video experience quality evaluation system according to claim 4, characterized in that: The system further comprises an electroencephalogram (EEG) analysis module, which is configured to: Traverse each candidate duration in the candidate duration range, and continuously play the audio-free video with the high-quality label and the audio-free video with the low-quality label for the traversed candidate duration; At each candidate duration, analyze the EEG of users watching audio-free videos with high-quality labels, and analyze the EEG of users watching audio-free videos with low-quality labels to determine the target duration; The EEG of a user watching an audio-free video with a high-quality label and the EEG of a user watching an audio-free video with a low-quality label are compared at the target duration to determine the target EEG signal channel for training the EEG feature learning sub-model.

7. The EEG signal-assisted video experience quality evaluation system according to claim 4, characterized in that: The EEG feature encoder in the EEG feature learning sub-model uses an improved EEGNet network, which includes: a convolutional layer, a depthwise separable convolutional layer, a separable convolutional layer, and a fully connected layer. The improved EEGNet network models the spatiotemporal relationship of EEG signals through three layers of convolution to extract high-level features related to human subjective cognition; The video feature encoder in the video feature learning sub-model includes a convolutional neural network backbone and a Transformer module, and the video feature encoder is used to extract low-level features including clarity, brightness, and color.

8. The EEG signal-assisted video experience quality evaluation system according to any one of claims 1 to 7, characterized in that: The system further comprises: The audioless video quality measurement module is used to input the target audioless video into the video feature learning sub-model in the audioless video quality measurement model to obtain the quality score of the target audioless video.

Citation Information

Patent Citations

  • Visual image reconstruction system based on brain-computer interface

    CN111539331A

  • Video quality evaluation method and apparatus, electronic device and storage medium

    WO2024216777A1