Image quality assessment method based on EEG signals and structural uncertainty
By selecting images with different structural uncertainties for appropriately perceived distortion acquisition experiments, collecting subjects' EEG signals, and building an EEG signal quality evaluation network, solving the problem of failure to consider the impact of image structure uncertainty in the prior art, and achieving more accurate image quality evaluation.
Patent Information
- Application Number
- CN202310421993.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-19
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2043-04-19
AI Technical Summary
The existing image quality evaluation method based on EEG signals fails to effectively consider the impact of image structure uncertainty on human visual psychology, resulting in low accuracy of prediction results.
By selecting images with different structural uncertainties, conducting an appropriately perceived distortion acquisition experiment, subjects can view EEG signals under different quality levels, construct and train an EEG signal quality evaluation network, and obtain the quality prediction score of the image.
It improves the accuracy of image quality evaluation, can better reflect the human visual perception mechanism, and overcomes the problem of high cost of subjective evaluation.
Smart Images

Figure CN116468692B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image quality evaluation, and in particular relates to an image quality evaluation method based on electroencephalogram (EEG) signals and structural uncertainty. Background Art
[0002] Images are important visual carriers in our daily lives. Quality loss is inevitable during the acquisition, transmission, and processing processes, affecting people's perception of images. Therefore, designing effective quality evaluation methods is of great significance for the transmission of visual information in images.
[0003] Image quality assessment methods are primarily categorized as subjective and objective. Subjective image quality assessment methods rely on the statistical analysis of large amounts of data. Therefore, to ensure statistical significance, these methods require a large number of observers to participate in quality assessment experiments, resulting in high experimental costs. Furthermore, subjective image quality assessment is susceptible to influences from observer preferences, experience, and expectations during the experiment, resulting in quality assessment results that do not accurately reflect human perception of image quality. Objective image quality assessment methods primarily analyze images through the construction of algorithmic models, generating low-cost image quality scores. While these algorithmic models can account for the influence of various image factors on quality assessment, due to the complexity of the human visual system, objective image quality assessment methods cannot fully simulate the human eye's perception mechanisms, resulting in discrepancies between quality assessment scores and subjective evaluations.
[0004] With the continuous development and progress of fields such as computer science, biomedical engineering, and data science, more and more physiological measurement technologies have been developed and applied in various fields, such as electroencephalography (EEG) measurement, eye tracking, and galvanic skin response measurement. Using physiological measurement technologies, researchers can measure and analyze physiological indicators of the human body when exposed to various stimuli, thereby better understanding the physiological mechanisms of the human body. Among them, electroencephalography (EEG) measurement has shown great application prospects in the field of image quality assessment. In image quality assessment, the EEG signal response of human observers when viewing images of different quality levels is recorded non-invasively, thereby analyzing the impact of image quality on human perception and cognition.
[0005] Existing EEG research has demonstrated that image stimuli with varying structural determinism produce different responses in the brain's attention and cognition. This phenomenon, revealed through EEG event-related potentials, is conducive to further research into EEG-based image quality assessment methods. Compared to objective quality assessment, EEG-based image quality assessment can better reflect the human eye's quality perception mechanism and overcome the high cost of subjective quality assessment experiments, making it of great research value.
[0006] Tsinghua University disclosed a method for image quality assessment based on EEG features in its patent application, "Image quality assessment method and device based on EEG features" (patent application number: CN202110700519.0, application publication number: CN 113554597 A). The implementation steps of this method are as follows: first, obtain a color image to be evaluated, input the color image to be evaluated into a score prediction network, and obtain a color image quality score output by the score prediction network. The disadvantage of this method is that the structural uncertainty of the image will affect the expectations and strategies of human visual psychology when the human eye perceives distortion, and a large number of visual psychology experiments have also verified that image structural uncertainty has a significant impact on distortion perception. Therefore, when performing image quality assessment, this method does not take into account the influence of image structural uncertainty on the human visual system when perceiving image distortion, resulting in low accuracy of the model prediction results. Summary of the Invention
[0007] In order to overcome the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide an image quality evaluation method based on EEG signals and structural uncertainty. According to the structural uncertainty value calculation formula and the image stimulus selection principle, images with different structural uncertainties are selected; a just noticeable distortion acquisition experiment is performed to obtain the average just noticeable distortion of images with different structural uncertainties; an EEG signal acquisition experiment is performed to collect EEG signals generated by subjects watching image stimuli with different structural uncertainties at different quality levels; and an EEG signal quality evaluation network is constructed and trained. The evaluation method has high prediction accuracy.
[0008] In order to achieve the above object, the technical solution adopted by the present invention is:
[0009] The image quality evaluation method based on EEG signals and structural uncertainty includes the following steps:
[0010] Step 1: According to the structural uncertainty value calculation formula and the image stimulus selection principle, select images with different structural uncertainties, and select multiple images to form an image set P;
[0011] Step 2: Perform different levels of distortion processing on the images in the image set P, conduct experiments on obtaining just-perceptible distortion, and obtain the average just-perceptible distortion of images with different structural uncertainties;
[0012] Step 3: Conduct an EEG signal acquisition experiment, select the quality level corresponding to the image distortion that is just noticeable, perform distortion processing on the image stimulus, and collect the EEG signals generated by the subjects viewing the image stimuli with different structural uncertainties at different quality levels;
[0013] Step 4: Preprocess the EEG signal;
[0014] Step 5: Build and train the EEG signal quality evaluation network to obtain the quality prediction score of the image corresponding to the EEG signal.
[0015] In step 1, the multiple images are selected according to the following principles:
[0016] (1) No salient objects should appear in the selected images;
[0017] (2) The selected images should contain natural textures and no semantically meaningful content; the natural texture refers to the texture in the pictures of natural imaging (taken by a camera rather than generated by a computer);
[0018] Texture is a visual feature that reflects homogeneous phenomena in an image. It embodies the common intrinsic properties of the surface of an object and contains important information about the structural organization and arrangement of the surface of the object and its relationship with the surrounding environment.
[0019] (3) The texture distribution of the selected image is required to be basically uniform in space; the term "basic uniformity" means that the statistical features of the local area within the image appear periodically in the entire image space;
[0020] (4) The selected images are classified into three categories according to their structural uncertainty, namely, high structural uncertainty images, medium structural uncertainty images, and low structural uncertainty images;
[0021] Define H U is the structural uncertainty value of an image. According to this value, the definitions of three types of structural uncertainty images are as follows:
[0022] Low structural uncertainty image, H U <1.7, medium structural uncertainty image, H U 1.7~2.2, high structural uncertainty image, H U >2.2.
[0023] The structural uncertainty H of the image U is calculated as follows:
[0024] The predicted value of each pixel in the image is calculated pixel by pixel. The calculation formula is as follows:
[0025]
[0026] Among them, g' represents the predicted value of the center pixel, χ represents the 21x21 pixel area centered on the center pixel, and g i represents the gray value of the i-th pixel in the region χ, ε represents white noise, and the autoregressive coefficient C i The calculation formula is as follows:
[0027]
[0028] Among them, x represents the center pixel, x i represents the i-th pixel in the region χ, x k represents the kth pixel in region χ, and I represents the mutual information between pixels;
[0029] Let the original image be M, the predicted image obtained by pixel-by-pixel prediction using the above formula be M', and the difference between the original image and the predicted image be considered as the prediction error image U, which is calculated as follows:
[0030] U=MM'
[0031] The Shannon entropy formula is used to calculate the information entropy of each pixel in U and the average value is taken as the structural uncertainty value of the image:
[0032]
[0033] where x i Represents the i-th pixel in the prediction error image U, p b (x i ) represents the pixel x in image U i The probability of the gray value appearing, n represents the total number of pixels. The step 2 is specifically as follows:
[0034] 2.1 Perform different levels of distortion processing on the image;
[0035] The quality parameter (QP) of the image is adjusted using the VideoWriter tool of MATLAB software, and the images in the image set P are distorted to obtain an image sequence set;
[0036] The range of QP is 0 to 100. When QP is 100, the image is not distorted. When QP is 0, the image is most distorted. The larger the QP parameter value, the lower the degree of distortion, and the more difficult it is for the subject to perceive the image distortion.
[0037] For each image in the image set P selected in step 1, generate distorted images with quality parameters QP ranging from 11 to 80, and sort them from high to low in terms of quality to form an image sequence consisting of distorted images;
[0038] The image sequences corresponding to each image in the image set P are collected to obtain the image sequence set P'; the low structure uncertainty image set P l The image sequences corresponding to each image in are collected together to obtain the low structural uncertainty image sequence set P l '; The structural uncertainty image set P m The image sequences corresponding to each image in are collected together to obtain the image sequence set P of structural uncertainty in m '; Set the high structural uncertainty image set P h The image sequences corresponding to each image in are collected together to obtain the high structural uncertainty image sequence set P h ';
[0039] 2.2 The subjects viewed the image sequences and obtained the average just-perceptible distortion of low structural uncertainty images, medium structural uncertainty images, and high structural uncertainty images respectively.
[0040] Furthermore, each subject will watch a low structural uncertainty image sequence set P l ', the image sequence set P of medium structure uncertainty m ', high structural uncertainty image sequence set P h ' once for each image sequence in the image file. The specific process is as follows:
[0041] The subjects sequentially watch a low-structure uncertainty image sequence set P containing the image sequence l ', the image sequence set P containing the image sequence with medium structural uncertainty m ', a set of image sequences with high structural uncertainty P containing image sequences h ';
[0042] The process of viewing a single image sequence includes a fixation point presentation phase and an experimental stimulus presentation phase. During the fixation point presentation phase, a red "+" is displayed to the subject for 1 second to encourage the subject to focus on the image stimulus, signaling the start of the image sequence. Subsequently, during the experimental stimulus presentation phase, the image sequence described in step 2.1 is played, with the images arranged from high to low quality, and each image is displayed for 0.5 seconds.
[0043] During the playback of the image sequence, the subject judges the moment when the image becomes distorted. At this time, the system records the quality parameter QP of the image played when the subject makes the judgment, and records this QP as the just noticeable distortion of the image. When the subject makes the judgment, the playback process of the current image sequence ends immediately, and after waiting for 5 seconds, the playback process of the next image sequence begins.
[0044] The step 3 is specifically as follows:
[0045] 3.1 Obtain distorted images with different distortion levels;
[0046] Three quality parameters (QPs) were selected to obtain distorted images with different degrees of distortion. The three quality parameters (QPs) were obtained from the just-perceptible distortion acquisition experiment in step 2. They are QP1, QP2, and QP3, which correspond to the average just-perceptible distortion of images with low structural uncertainty, medium structural uncertainty, and high structural uncertainty, respectively. The VideoWriter tool in MATLAB software was used to perform distortion processing with quality parameters of QP1, QP2, and QP3, respectively.
[0047] 3.2 Subjects viewed distorted images;
[0048] The distorted images were used as experimental stimuli, and each subject was randomly presented with the distorted images described in 3.1. The specific experimental process was as follows:
[0049] Furthermore, the process of viewing a single trial includes the fixation point presentation stage, the experimental stimulus presentation stage and the distortion judgment stage; in the fixation point presentation stage, a red "+" is shown to the subject for 1 second to enable the subject to concentrate on watching the image stimulus; then, in the experimental stimulus presentation stage, the undistorted image is first played for 2 seconds, and then the distorted image is played for 2 seconds; in the distortion judgment stage, the subject judges whether the distortion is perceived.
[0050] 3.3 EEG signal acquisition;
[0051] The EEG signal was collected in each single trial of the subject, and the EEG signal acquisition system used was NeuroScan.
[0052] Furthermore, during the EEG data collection process, the following collection conditions need to be met: (1) the subjects are required to remain energetic to prevent inattention, frequent blinking, and drowsiness during the collection process; (2) the resistance of all electrodes of the conductive cap must be less than 20 kilo-ohms, and adhesion between the electrodes must be avoided; (3) the experimental site must be well-lit, with a suitable temperature and no noise; and (4) the distance from the subject's eyes to the image stimulus must be approximately 4 times the height of the image stimulus display.
[0053] The step 4 is specifically as follows:
[0054] Preprocess the EEG data, including reference transfer, baseline correction, filtering, and artifact removal;
[0055] 4.1 For reference;
[0056] The M1 and M2 electrodes located at the bilateral mastoid processes behind the ears are used as reference electrodes. The average value of the signals collected by the M1 and M2 electrodes is used as the reference value of the EEG signal. The EEG signals of all electrodes are recalculated based on the reference value.
[0057] 4.2 Baseline correction;
[0058] The mean of the EEG signal segment from 200 milliseconds before the image stimulus is presented to the beginning of the image stimulus is calculated as the baseline to correct the entire EEG signal;
[0059] 4.3 Bandpass filtering;
[0060] A band-pass filter was selected to intercept noise signals below 0.1 Hz and above 20 Hz, retaining the EEG signals between 0.1 and 20 Hz;
[0061] 4.4 Remove artifacts;
[0062] During the acquisition of EEG signals, artifacts similar to EEG signals may be left in the EEG signals due to interference sources such as blinking, muscle movement and heartbeat. Processing software is used to remove the artifacts.
[0063] The step 5 is specifically as follows:
[0064] 5.1 Divide the training set and test set;
[0065] Select EEG samples, one part of which is used as the training set and the other part as the test set. The dimension of each EEG signal data sample is C×T;
[0066] 5.2 Constructing an EEG signal quality evaluation network;
[0067] Construct an EEG signal quality evaluation network S, which includes modules 1, 2, and 3. Module 1 is a spatiotemporal convolution-average pooling module, module 2 is a separable convolution module, and module 3 is a classification output module. The input dimension of the EEG signal is C×T, where C represents the number of EEG signal channels and T represents the number of sampling points.
[0068] In module 1, the first layer is a temporal convolution layer, which uses eight 1×64 convolution kernels to extract the temporal features of the EEG signal. The second layer is a spatial convolution layer, which uses four C×1 depthwise convolution kernels to learn spatial filters. Batch normalization is then applied along the feature map dimensions, using an exponential linear unit as the activation function. Finally, a 1x4 average pooling layer is used for downsampling. In addition, the dropout technique is used to prevent overfitting, and the weights of each spatial filter are regularized by applying a maximum norm constraint of 1.
[0069] In module 2, four 1x16 kernels are used for separable convolution, followed by eight 1x1 kernels for pointwise convolution. An exponential linear unit is used as the activation function, and finally a 1x8 average pooling layer is used for dimensionality reduction. Dropout is used to prevent overfitting.
[0070] In module 3, a fully connected layer is used to output the output of the fully connected layer using a 4-category softmax activation function, and the output is the quality prediction score of the image;
[0071] 5.3 Training the EEG signal quality evaluation network;
[0072] During the training phase, all training processes use the Adam optimization algorithm. The formula for calculating the mean square error between the quality prediction score corresponding to each training sample and the quality score label corresponding to the training sample is:
[0073]
[0074] b represents the number of training samples randomly selected from the training sample set B without replacement during iterative training of the EEG signal quality evaluation network S, q g represents the quality score label corresponding to the g-th training sample, Represents the quality prediction score corresponding to the g-th training sample.
[0075] Beneficial effects of the present invention:
[0076] The present invention utilizes the different responses of the brain's attention and cognitive systems to viewing images with different structural uncertainties in the human visual perception mechanism. By collecting images of subjects viewing images with different structural uncertainties and different distortion levels, an image quality evaluation method based on EEG signals and structural uncertainty is designed. The evaluation method has a high prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 It is a schematic diagram of the process of the present invention.
[0078] Figure 2is an uncertainty image, where Figure 2 a is a low structural uncertainty image, Figure 2 b is the medium structure uncertainty image, Figure 2 c is an image with high structural uncertainty.
[0079] Figure 3 Schematic diagram of the experimental process for subjects to watch image sequences.
[0080] Figure 4 Schematic diagram of the experimental process for subjects to view distorted images.
[0081] Figure 5 Schematic diagram of the EEG signal quality evaluation network. DETAILED DESCRIPTION
[0082] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0083] like Figure 1 As shown:
[0084] Step 1: Select multiple images to form a picture set:
[0085] Select multiple images from the texturelib website. The images on the texturelib website are specifically designed to provide materials for experiments related to image textures. The selection principles are as follows:
[0086] (1) The selected images should not contain any salient objects, to prevent the appearance of salient objects in the images from interfering with the subjects’ visual attention and affecting their judgment of image quality;
[0087] (2) The selected images should contain natural textures and no semantically meaningful content, so as to minimize the impact of image meaning on the subjects;
[0088] (3) The texture distribution of the selected image is required to be basically uniform in space, so as to avoid large differences in structural uncertainty in different regions of the image, which makes it difficult to explore the impact of the overall structural uncertainty of the image on the perception of image quality;
[0089] (4) The selected images are classified into three categories according to their structural uncertainty, namely, high structural uncertainty images, medium structural uncertainty images, and low structural uncertainty images. Define H U is the structural uncertainty value of an image. According to this value, the definitions of three types of structural uncertainty images are shown in the following table:
[0090]
[0091] Specifically, the structural uncertainty H of an image U is calculated as follows:
[0092] The predicted value of each pixel in the image is calculated pixel by pixel. The calculation formula is as follows:
[0093]
[0094] Among them, g' represents the predicted value of the center pixel, χ represents the 21x21 pixel area centered on the center pixel, and g i Represents the grayscale value of the i-th pixel in the region χ, and ε represents white noise. Autoregressive coefficient C i The calculation formula is as follows:
[0095]
[0096] Among them, x represents the center pixel, x i represents the i-th pixel in the region χ, x k represents the kth pixel in region χ, and I represents the mutual information between pixels.
[0097] Let the original image be M, the predicted image obtained by pixel-by-pixel prediction using the above formula be M', and the difference between the original image and the predicted image be considered as the prediction error U, which is calculated as follows:
[0098] U=MM'
[0099] The Shannon entropy formula is used to calculate the information entropy of each pixel in U and the average value is taken as the structural uncertainty value of the image:
[0100]
[0101] where p b (x) represents the probability of x in U, and n represents the total number of pixels.
[0102] In this embodiment, a total of 30 images are selected according to the above-mentioned image selection principle to form an image set P, including 10 low structural uncertainty images, 10 medium structural uncertainty images, and 10 high structural uncertainty images. It is expressed as P = {P l ,P m ,P h}, where the low structural uncertainty image set P l ={P l1 ,P l2 ...P l10}, the structural uncertainty image set P m ={P m1 ,P m2 ...P m10}, high structural uncertainty image set P h ={P h1 ,P h2...P h10 The image resolution is uniformly 1920x1080. To control variables and prevent the image content and color from affecting the subjects' judgment, the images are all walls with basically the same color. Figure 2 As shown:
[0103] Using the above structural uncertainty calculation formula, we calculated the values for images with different structural uncertainties. The average structural uncertainty value for images with low structural uncertainty was 1.27, for images with medium structural uncertainty it was 2.05, and for images with high structural uncertainty it was 2.54. This shows that the structural uncertainty values for the three image groups meet the image selection principle.
[0104] Step 2: Chuck perceived distortion acquisition experiment
[0105] 2.1 Perform different levels of image distortion processing
[0106] The quality parameter (QP) of the image is adjusted using the VideoWriter tool of MATLAB software, and the 30 images in the image set P are distorted to obtain the image sequence set
[0107] The QP range is 0 to 100. A QP of 100 indicates an undistorted image, while a QP of 0 indicates the most severe image distortion. The larger the QP parameter value, the less distortion there is, and the less likely the subject will notice the image distortion. Experience has shown that image distortion is very noticeable when the QP is between 0 and 10, while it is imperceptible when the QP is between 80 and 100. Therefore, for each image in the image set P selected in step 1, 70 distorted images with QP parameters ranging from 11 to 80 are generated. These images are sorted from high to low quality, forming an image sequence consisting of 70 distorted images.
[0108] The image sequences corresponding to each image in the image set P are collected to obtain the image sequence set P'; the low structure uncertainty image set P l The image sequences corresponding to each image in are collected together to obtain the low structural uncertainty image sequence set P l '; The structural uncertainty image set P m The image sequences corresponding to each image in are collected together to obtain the image sequence set P of structural uncertainty in m '; Set the high structural uncertainty image set P h The image sequences corresponding to each image in are collected together to obtain the high structural uncertainty image sequence set P h ';
[0109] 2.2 The subjects viewed the image sequences and obtained the average just-perceptible distortion of low structural uncertainty images, medium structural uncertainty images, and high structural uncertainty images respectively.
[0110] Just-Noticeable Distortion (JND) refers to the maximum distortion that the human visual system cannot perceive. In the field of image processing, just-noticeable distortion can be used to measure the human eye's sensitivity to image distortion.
[0111] In this experiment, 6 subjects (3 males and 3 females) were selected. Each subject watched a low-structure uncertainty image sequence set P l ', the image sequence set P of medium structure uncertainty m ', high structural uncertainty image sequence set P h ', each of the image sequences is once. The specific experimental process is as follows Figure 3 As shown:
[0112] The subjects watched a low structural uncertainty image sequence set P containing 10 image sequences in sequence. l ', a set of medium structure uncertainty image sequences P containing 10 image sequences m ', a high structural uncertainty image sequence set P containing 10 image sequences h '.
[0113] Specifically, the process for viewing a single image sequence includes a fixation point presentation phase and an experimental stimulus presentation phase. During the fixation point presentation phase, a red "+" icon is displayed to the subject for one second to encourage them to focus on the stimulus image and signal the start of the image sequence. Next, during the experimental stimulus presentation phase, the image sequence described in 2.1, consisting of 70 images, is played. The images are arranged in descending order of quality, with each image displayed for 0.5 seconds.
[0114] During the playback of an image sequence, the subject judged the moment when image distortion occurred. The system recorded the quality parameter (QP) of the image being played at the time the subject made the judgment, and recorded this QP as the just-noticeable distortion of the image. When the subject made the judgment, playback of the current image sequence ended immediately, and after a 5-second wait, playback of the next image sequence began.
[0115] In this embodiment, 6 subjects watched P l '、P m '、P h 'The record table is as follows:
[0116] Just noticeable distortion to obtain experimental record sheet
[0117] Subjects <![CDATA[P l ']]> <![CDATA[P m ']]> <![CDATA[P h ']]> S1 45 38 32 S2 42 37 31 S3 43 36 29 S4 41 33 29 S5 44 35 32 S6 43 37 33 average value 43 36 31
[0118] The values in the table represent the image quality parameter (QP) at which subjects perceive just-noticeable distortion. The table shows that the average just-noticeable distortion for low structural uncertainty images occurs at a QP of 43, the average just-noticeable distortion for medium structural uncertainty images occurs at a QP of 36, and the average just-noticeable distortion for high structural uncertainty images occurs at a QP of 31.
[0119] Step 3: EEG signal acquisition experiment
[0120] 3.1 Obtaining distorted images with different distortion levels
[0121] Three quality parameters (QP) were selected to obtain distorted images with varying degrees of distortion. These three QP parameters were obtained from the just-perceptible distortion acquisition experiment in step 2: QP1, the average just-perceptible distortion for low structural uncertainty images; QP2, the average just-perceptible distortion for medium structural uncertainty images; and QP3, corresponding to the average just-perceptible distortion for high structural uncertainty images. In this example, QP1 = 43, QP2 = 36, and QP3 = 31. Using the VideoWriter tool in MATLAB software, the 30 images in image set P were distorted using quality parameters QP1, QP2, and QP3, respectively, to obtain 90 distorted images.
[0122] 3.2 Subjects viewing distorted images
[0123] This experiment selected 6 subjects (3 males and 3 females). The distorted images were used as experimental stimuli. Each subject was randomly presented with 90 distorted images as described in 3.1. Each image was repeated 10 times, and each subject was required to complete a total of 900 trials. The specific experimental process is as follows: Figure 4 As shown:
[0124] Specifically, the process for viewing a single trial includes a fixation point presentation phase, an experimental stimulus presentation phase, and a distortion judgment phase. During the fixation point presentation phase, a red "+" icon is displayed for one second to encourage participants to focus on the stimulus image. Next, during the experimental stimulus presentation phase, the undistorted image is played for two seconds, followed by the distorted image for two seconds. During the distortion judgment phase, participants determine whether they perceive distortion.
[0125] 3.3 EEG signal acquisition
[0126] EEG signals were collected during each trial of the subjects using the NeuroScan EEG signal acquisition system, which includes a Quik-Cap 64-conductor electrode cap, a SynAmps2 64-conductor EEG signal amplifier, and the professional brain signal acquisition and processing software Curry7.
[0127] During the EEG data collection process, the following collection conditions need to be met: (1) The subjects are required to remain energetic to prevent inattention, frequent blinking, and drowsiness during the collection process; (2) The resistance of all electrodes of the conductive cap must be less than 20 kilo-ohms, and adhesion between the electrodes must be avoided; (3) The experimental site must be well lit, with a suitable temperature and no noise; (4) The distance from the subject's eyes to the image stimulus must be about 4 times the height of the image stimulus display.
[0128] In this embodiment, a total of 5400 EEG signal samples were collected from 6 subjects. The dimension of each EEG signal data sample is C×T, where C represents the number of EEG signal channels and T represents the number of sampling points.
[0129] In this embodiment, C is 64 and T is 1000.
[0130] Step 4: EEG signal preprocessing
[0131] To more accurately analyze the characteristics and changes of EEG signals, raw EEG data needs to be preprocessed to achieve data cleaning and data transformation. Preprocessing steps include reference transfer, baseline correction, filtering, and artifact removal.
[0132] 4.1 Reference
[0133] The NeuroScan acquisition system used in this experiment defaults to using the Ref electrode in the Quik-Cap as the reference electrode. Because the Ref electrode often varies in location among subjects due to varying head sizes, the M1 and M2 electrodes, located bilaterally behind the ears on the mastoid process, were selected as reference electrodes. The average value of the signals collected by the M1 and M2 electrodes was used as the reference value for the EEG signals, and the EEG signals of all electrodes were recalculated based on this reference value.
[0134] 4.2 Baseline Correction
[0135] Baseline correction can avoid EEG drift caused by noise interference and imbalance between different electrodes. Therefore, we selected the EEG signal segment from 200 milliseconds before the image stimulus was presented to the beginning of the image stimulus, and calculated the mean of this EEG signal segment as the baseline for correction of the entire EEG signal.
[0136] 4.3 Bandpass Filtering
[0137] During the EEG acquisition process, signals contain a significant amount of both physiological and non-physiological noise. Physiological noise, such as myoelectric and oculoscopic noise, is high-frequency, while non-physiological noise, such as DC offset, is low-frequency. Therefore, in this experiment, a bandpass filter was used to intercept noise signals below 0.1 Hz and above 20 Hz, retaining the EEG signal between 0.1 and 20 Hz.
[0138] 4.4 Artifact Removal
[0139] During EEG signal acquisition, interference from sources such as blinking, muscle movement, and heartbeats can leave artifacts similar to those in the EEG signal. These artifacts can interfere with EEG signal interpretation and analysis, and cannot be completely removed using bandpass filtering. Therefore, in this experiment, the independent component analysis function in the processing software Curry7 was used to remove these artifacts.
[0140] Step 5: Build and train the EEG signal quality evaluation network
[0141] 5.1 Dividing the training set and test set
[0142] There are 5400 EEG samples in total, 80% of which are selected as the training set (4320) and 20% as the test set (1080). The dimension of each EEG signal data sample is C×T. In this embodiment, C is 64 and T is 1000.
[0143] 5.2 Construction of EEG signal quality evaluation network
[0144] Construct an EEG signal quality evaluation network S. The network structure is as follows Figure 5 As shown in the figure, module 1 is the spatiotemporal convolution-average pooling module, module 2 is the separable convolution module, and module 3 is the classification output module. The input dimension of the EEG signal is C×T, where C represents the number of EEG signal channels and T represents the number of sampling points.
[0145] In module 1, the first layer is a temporal convolution layer, using eight 1×64 convolution kernels to extract temporal features of the EEG signal. The second layer is a spatial convolution layer, using four C×1 depthwise convolution kernels to learn spatial filters. Batch normalization is then applied along the feature map dimensions, using an exponential linear unit as the activation function. Finally, a 1×4 average pooling layer is used for downsampling. Dropout is also used to prevent overfitting, regularizing each spatial filter weight by applying a maximum norm constraint of 1.
[0146] In module 2, four 1x16 kernels are used for separable convolution, followed by eight 1x1 kernels for pointwise convolution. An exponential linear unit is used as the activation function, and finally a 1x8 average pooling layer is used for dimensionality reduction. Dropout is used to prevent overfitting.
[0147] In module 3, a fully connected layer is used to output the output of the fully connected layer using a 4-category softmax activation function, and the output is the quality prediction score of the image.
[0148] 5.3 Training EEG signal quality evaluation network
[0149] During the training phase, the batch size is set to 64 and the initial learning rate Ir = 5×10 -4 After every 50 training iterations, the learning rate is reduced to 1 / 10 of the previous stage, and a total of 200 training iterations are run. All training processes use the Adam optimization algorithm. The formula for calculating the mean square error between the quality prediction score corresponding to each training sample and the quality score label corresponding to the training sample is:
[0150]
[0151] b represents the number of training samples randomly selected from the training sample set B without replacement during iterative training of the EEG signal quality evaluation network S, q g represents the quality score label corresponding to the g-th training sample, Represents the quality prediction score corresponding to the g-th training sample.
Claims
1. An image quality assessment method based on EEG signals and structural uncertainty, characterized in that: The following steps are included: Step 1: According to the structural uncertainty value calculation formula and the image stimulus selection principle, select images with different structural uncertainties, and select multiple images to form an image set P; Step 2: Perform different levels of distortion processing on the images in the image set P, conduct experiments on obtaining just-perceptible distortion, and obtain the average just-perceptible distortion of images with different structural uncertainties; Step 3: Conduct an EEG signal acquisition experiment, select the quality level corresponding to the image distortion that is just noticeable, perform distortion processing on the image stimulus, and collect the EEG signals generated by the subjects viewing the image stimuli with different structural uncertainties at different quality levels; Step 4: Preprocess the EEG signal; Step 5: Build and train the EEG signal quality evaluation network to obtain the quality prediction score of the image corresponding to the EEG signal; In step 1, the multiple images are selected according to the following principles: (1) No salient objects should appear in the selected images; (2) The selected images should contain natural textures and have no semantically meaningful content; The natural texture refers to the texture in a naturally imaged picture; (3) The texture distribution of the selected image is required to be basically uniform in space; the term "basic uniformity" means that the statistical features of the local area within the image appear periodically in the entire image space; (4) The selected images are classified into three categories according to their structural uncertainty, namely, high structural uncertainty images, medium structural uncertainty images, and low structural uncertainty images; Define H U is the structural uncertainty value of an image. According to this value, the definitions of three types of structural uncertainty images are as follows: Low structural uncertainty image, H U <1.7, medium structural uncertainty image, H U 1.7~2.2, high structural uncertainty image, H U >2.2; The structural uncertainty H of the image U is calculated as follows: The predicted value of each pixel in the image is calculated pixel by pixel. The calculation formula is as follows: Among them, g' represents the predicted value of the center pixel, χ represents the 21x21 pixel area centered on the center pixel, and g i represents the gray value of the i-th pixel in the region χ, ε represents white noise, and the autoregressive coefficient C i The calculation formula is as follows: Among them, x represents the center pixel, x i represents the i-th pixel in the region χ, x k represents the kth pixel in region χ, and I represents the mutual information between pixels; Let the original image be M, the predicted image obtained by pixel-by-pixel prediction using the above formula be M', and the difference between the original image and the predicted image be considered as the prediction error U, which is calculated as follows: U=MM' The Shannon entropy formula is used to calculate the information entropy of each pixel in U and the average value is taken as the structural uncertainty value of the image: where x i Represents the i-th pixel in the prediction error image U, p b (x i ) represents the pixel x in image U i The probability of the gray value appearing, n represents the total number of pixels.
2. The image quality assessment method based on EEG signals and structural uncertainty according to claim 1, characterized in that: The step 2 is specifically as follows: 2.1 Perform different levels of distortion processing on the image; The VideoWriter tool of MATLAB software is used to adjust the quality parameters of the image, and the images in the image set P are distorted to obtain an image sequence set; The range of QP is 0 to 100. When QP is 100, the image is not distorted. When QP is 0, the image is most distorted. The larger the QP parameter value, the lower the degree of distortion, and the more difficult it is for the subject to perceive the image distortion. For each image in the image set P selected in step 1, generate distorted images with quality parameters QP ranging from 11 to 80, and sort them from high to low in terms of quality to form an image sequence consisting of distorted images; The image sequences corresponding to each image in the image set P are collected to obtain the image sequence set P'; the low structure uncertainty image set P l The image sequences corresponding to each image in are collected together to obtain the low structural uncertainty image sequence set P' l ; The structural uncertainty image set P m The image sequences corresponding to each image in are collected together to obtain the image sequence set P' of structural uncertainty in m ; Set the high structural uncertainty image set P h The image sequences corresponding to each image in are collected together to obtain the high structural uncertainty image sequence set P' h ; 2.2 The subjects viewed the image sequences and obtained the average just-perceptible distortion of low structural uncertainty images, medium structural uncertainty images, and high structural uncertainty images respectively.
3. The image quality assessment method based on EEG signals and structural uncertainty according to claim 2, characterized in that: Each subject will watch a low structural uncertainty image sequence set P' l , the image sequence set P' of medium structural uncertainty m , high structural uncertainty image sequence set P' h All image sequences in are processed once. The specific process is as follows: The subjects sequentially watch a low structural uncertainty image sequence set P' containing the image sequence l , the image sequence set P' containing the image sequence with medium structural uncertainty m , a set of image sequences with high structural uncertainty P' h ; The process of viewing a single image sequence includes a fixation point presentation phase and an experimental stimulus presentation phase. During the fixation point presentation phase, a red "+" is displayed to the subject for 1 second to encourage them to focus on the image stimulus and signal the start of the image sequence. Next, during the experimental stimulus presentation phase, the image sequence described in step 2.1 is played, with the images arranged from high to low quality, and each image is displayed for 0.5 seconds. During the playback of an image sequence, if the subject perceives image distortion and makes a judgment, the system records the quality parameter QP of the image played when the subject makes the judgment, and records this QP as the just noticeable distortion of the image. When the subject makes the judgment, the playback process of the current image sequence ends immediately, and after waiting for 5 seconds, the playback process of the next image sequence begins.
4. The image quality assessment method based on EEG signals and structural uncertainty according to claim 1, characterized in that: The step 3 is specifically as follows: 3.1 Obtain distorted images with different distortion levels; Three quality parameters (QPs) were selected to obtain distorted images with different degrees of distortion. The three quality parameters (QPs) were obtained from the just-perceptible distortion acquisition experiment in step 2. They are QP1, QP2, and QP3, which correspond to the average just-perceptible distortion of images with low structural uncertainty, medium structural uncertainty, and high structural uncertainty, respectively. The VideoWriter tool in MATLAB software was used to perform distortion processing with quality parameters of QP1, QP2, and QP3, respectively. 3.2 Subjects viewed distorted images; The distorted images were used as experimental stimuli, and each subject was randomly presented with the distorted images described in 3.
1. The specific experimental process was as follows: The process of viewing a single trial included a fixation point presentation phase, an experimental stimulus presentation phase, and a distortion judgment phase. During the fixation point presentation phase, a red "+" was displayed to the subject for one second to encourage them to focus on the image stimulus. Next, during the experimental stimulus presentation phase, the undistorted image was played for two seconds, followed by the distorted image for two seconds. During the distortion judgment phase, the subject judged whether they perceived the distortion. 3.3 EEG signal acquisition; The EEG signal was collected in each single trial of the subject, and the EEG signal acquisition system used was NeuroScan.
5. The image quality assessment method based on EEG signals and structural uncertainty according to claim 1, characterized in that: The step 4 is specifically as follows: Preprocess the EEG data, including reference transfer, baseline correction, filtering, and artifact removal; 4.1 For reference; The M1 and M2 electrodes located at the bilateral mastoid processes behind the ears are used as reference electrodes. The average value of the signals collected by the M1 and M2 electrodes is used as the reference value of the EEG signal. The EEG signals of all electrodes are recalculated based on the reference value. 4.2 Baseline correction; The mean of the EEG signal segment from 200 milliseconds before the image stimulus is presented to the beginning of the image stimulus is calculated as the baseline to correct the entire EEG signal; 4.3 Bandpass filtering; A band-pass filter was selected to intercept noise signals below 0.1 Hz and above 20 Hz, retaining the EEG signals between 0.1 and 20 Hz; 4.4 Remove artifacts; During the acquisition of EEG signals, artifacts similar to EEG signals may be left in the EEG signals due to interference sources such as blinking, muscle movement and heartbeat. Processing software is used to remove the artifacts.
6. The image quality assessment method based on EEG signals and structural uncertainty according to claim 1, characterized in that: The step 5 is specifically as follows: 5.1 Divide the training set and test set; Select EEG samples, one part of which is used as the training set and the other part as the test set. The dimension of each EEG signal data sample is C×T; 5.2 Constructing an EEG signal quality evaluation network; Construct an EEG signal quality evaluation network S, which includes modules 1, 2, and 3. Module 1 is a spatiotemporal convolution-average pooling module, module 2 is a separable convolution module, and module 3 is a classification output module. The input dimension of the EEG signal is C×T, where C represents the number of EEG signal channels and T represents the number of sampling points. In module 1, the first layer is a temporal convolution layer, which uses eight 1×64 convolution kernels to extract the temporal features of the EEG signal. The second layer is a spatial convolution layer, which uses four C×1 depthwise convolution kernels to learn spatial filters. Batch normalization is then applied along the feature map dimensions, using an exponential linear unit as the activation function. Finally, a 1x4 average pooling layer is used for downsampling. In addition, the dropout technique is used to prevent overfitting, and the weights of each spatial filter are regularized by applying a maximum norm constraint of 1. In module 2, four 1x16 kernels are used for separable convolution, followed by eight 1x1 kernels for pointwise convolution. An exponential linear unit is used as the activation function, and finally a 1x8 average pooling layer is used for dimensionality reduction. Dropout is used to prevent overfitting. In module 3, a fully connected layer is used to output the output of the fully connected layer using a 4-category softmax activation function, and the output is the quality prediction score of the image; 5.3 Training the EEG signal quality evaluation network; During the training phase, all training processes use the Adam optimization algorithm. The formula for calculating the mean square error between the quality prediction score corresponding to each training sample and the quality score label corresponding to the training sample is: b represents the number of training samples randomly selected from the training sample set B without replacement during iterative training of the EEG signal quality evaluation network S, q g represents the quality score label corresponding to the g-th training sample, Represents the quality prediction score corresponding to the g-th training sample.
Citation Information
Patent Citations
Image quality evaluation method and device based on electroencephalogram characteristics
CN113554597A
Image quality evaluation method and device based on electroencephalogram characteristics
CN113554597B
Full-reference image quality evaluation method based on visual attention characteristics
CN109859157A
Information processing device, information processing method, and program
CN109891519A