Panel display visual comfort prediction method and system based on multi-modal fusion model
By generating an adversarial network to generate a method of fusion of diverse samples and multimodal features, combining physical and physiological characteristics, a multimodal fusion model is constructed, which solves the data scarcity and scene adaptability problems in the visual comfort prediction of flat panel display devices, and realizes efficient and accurate visual comfort prediction and physiologically interpretable display design guidance.
Patent Information
- Application Number
- CN202510448646.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The visual comfort prediction method of flat panel display devices in the prior art has the problems of scarce high-quality labeling samples, missing coupling relationships of single mode analysis, and poor adaptability of data and models, resulting in insufficient robustness of the model and low prediction accuracy.
A multimodal fusion model based on the generative adversarial network generation enhancement samples and original images is adopted, and a multimodal fusion model is constructed through a stacking integration framework. A multimodal fusion model is generated using the generative adversarial network to generate diversified samples that retain core features, extract heterogeneous features that conform to the working principles of human vision systems, and integrate the basic classifier output through a stacking classifier.
It improves the robustness and prediction accuracy of the model, can be efficiently generalized in different scenarios, provides physiologically interpretable visual comfort prediction results, and guides display design.
Smart Images

Figure CN120299099A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the intersection of display technology and human-computer interaction, and in particular to a method and system for predicting visual comfort of flat panel displays based on a multimodal fusion model. Background Art
[0002] With the widespread use of new flat-panel display technologies, the use time of flat-panel display devices has increased significantly in recent decades, causing more and more serious visual discomfort problems, which has become an important issue in the field of public health. Visual discomfort refers to the subjective discomfort and corresponding physiological changes caused by the human visual system during the process of viewing objects due to long-term viewing and high perceived pressure, including symptoms such as fatigue, sore eyes, and blurred vision.
[0003] The display industry traditionally evaluates visual comfort through subjective questionnaires, that is, asking test subjects to fill out questionnaires after watching the displayed content in order to adjust display parameters or determine the direction of research and development. However, this method requires a large number of repeated experiments, is time-consuming and labor-intensive, and has low efficiency, making it difficult to meet the needs of rapid iteration of research and development. With the development of artificial intelligence technology, a task model based on visual comfort indicators has been proposed, but the existing technology has the following core defects:
[0004] First, due to factors such as insufficient professional personnel and long experimental cycles, the number of high-quality annotated samples is limited. In order to train the model, we have to rely on data enhancement technology to expand the limited original samples. Currently, commonly used enhancement methods such as adjusting color transformations such as brightness / contrast, synthesizing minority oversampling, injecting random noise, etc., are essentially to generate new samples by modifying the original image features and use the same labels. However, such operations have limitations: the enhancement method is not based on the probability distribution of data, which may cause distortion of sample distribution; data transformation may distort the perception characteristics of the human visual system of real scenes, resulting in the actual visual comfort of the new samples not matching the original labels; enhanced samples are only generated within the known data neighborhood and cannot cover a wider range of visual perception scenarios, ultimately resulting in insufficient model robustness.
[0005] Second, existing research focuses on single-mode analysis of display characteristics or physiological signals. Models that use display characteristics as input rely on users' subjective evaluations and are easily affected by individual differences, environmental factors, and other factors, resulting in low data quality. Models that use physiological signals as input, although with high data quality, require high-cost psychophysical experiments. In addition, neither model captures the coupling relationship between light signals and the response of the human visual system, resulting in insufficient prediction accuracy and stability.
[0006] Thirdly, due to the over-reliance on single-modal data in existing models, the data acquisition paradigm is strongly bound to the model design, manifested as only supporting input data in a preset format and being unable to be compatible with heterogeneous data in industrial scenarios, etc. This rigid constraint between the data format and the model severely restricts its flexible deployment in the actual R & D of the display industry.
[0007] Fourthly, most existing models are proposed for stereoscopic 3D displays, but the data processing methods of existing 3D display models cannot be migrated to the flat panel display scenario. The main reason for the visual discomfort caused by 3D displays is that due to the depth information difference, the visual system needs to continuously adjust the focal length and convergence angle of both eyes, resulting in convergence adjustment conflicts. Over time, it becomes more difficult for viewers to fuse binocular images, leading to visual discomfort, mainly manifested as a sense of dizziness, which is generally related to electroencephalogram and semicircular canal activities. For flat panel display users, the visual system does not need to maintain a dynamic adjustment process to adapt to the depth-changing images, and the visual discomfort has less to do with convergence adjustment conflicts, mainly manifested as the decline of specific visual functions such as dry eyes, soreness, and blurred vision, which are generally related to eye movement and electrocardiogram activities. Obviously, the influencing factors, generation mechanisms, and physiological manifestations of visual discomfort in 3D displays and flat panel displays are quite different, and there are essential differences in feature engineering and algorithm logic between the two scenarios, and the research methods cannot be simply replicated. Summary of the Invention
[0008] The technical problem to be solved by the present invention is to provide a method and system for predicting the visual comfort of flat panel displays based on a multi-modal fusion model in view of the above-mentioned defects of the existing technology, aiming to solve the problems of scarce high-quality labeled samples, lack of coupling relationship in single-modal analysis, and poor scene adaptability of data and models in the existing technology.
[0009] To solve the above technical problems,
[0010] In the first aspect, a method for predicting the visual comfort of flat panel displays based on a multi-modal fusion model is provided, including the following steps: Step 1: Expand the image samples
[0011] Collect the original image set, where the original image set includes several original images. For each original image, use the generative adversarial network model to convert the input random sequence into enhanced materials, and then fuse the enhanced materials and the original images according to a preset ratio to obtain several new images; the new images corresponding to the original image set form a new image set.
[0012] Step 2: Obtain training samples
[0013] Play the original image set to viewers, and collect the physiological characteristics and visual comfort evaluations of the viewers; measure the physical parameters of the flat panel display screen, and input the original image set and the new image set to obtain physical characteristics; preprocess the physical characteristics and physiological characteristics to obtain physical-comfort training samples and physiological-comfort training samples.
[0014] Step 3: Train the multi-modal fusion model
[0015] Train the physical-based classifier and the physiological-based classifier based on the physical-comfort training samples and the physiological-comfort training samples respectively; train the stacked classifier based on the stacked integration framework and the trained physical-based classifier and physiological-based classifier; the basic classifier and the stacked classifier together constitute the multi-modal fusion model;
[0016] Step 4: Predict visual comfort
[0017] Input the test data into the multi-modal fusion model to obtain the visual comfort prediction result.
[0018] In one implementation, in Step 1, the generative adversarial network model includes a generator and a discriminator:
[0019] The generator is composed of a cascade of several deep 2D convolutional blocks, and each convolutional layer uses the ReLU activation function;
[0020] The discriminator is composed of a cascade of several deep 2D convolutional blocks, and each convolutional layer uses the leaky ReLU activation function;
[0021] Through unsupervised deep learning, train the generator and the discriminator simultaneously, set the training target of the generator to generate enhanced materials that are substantially equivalent to the original image in topological structure; set the training target of the discriminator to accurately distinguish the enhanced materials from the original image.
[0022] In one implementation, the original image is rotated, translated, and reflected before being input into the generative adversarial network model.
[0023] In one implementation, in Step 2, the preprocessing of the physical features is to extract imaging features and non-imaging features from the physical features; the physical features at least include the brightness, color coordinates, and spectrum of all input signal values of all pixel points of the flat display screen; the imaging features include the brightness, chromaticity, and hue matrix in the LCH color space, and the phase consistency matrix in the frequency domain; the non-imaging features include the retinal irradiance map corresponding to the retinal photoreceptors;
[0024] Thus, the physical-comfort training samples are divided into imaging-comfort training samples and non-imaging-comfort training samples; correspondingly, in Step 3, the physical-based classifier is divided into a first classifier and a second classifier, and the first classifier is trained based on the imaging-comfort training samples, and the second classifier is trained based on the non-imaging-comfort training samples.
[0025] In one embodiment, in step two, preprocessing the physiological features is to extract heart rate variability features and eye movement tracking features from the physiological features; the physiological features at least include electrocardiogram signals and eye movement tracking videos; the heart rate variability features include the average RR interval; the standard deviation of NN intervals SDNN; the root mean square of the differences between adjacent NN intervals RMSSD; among all NN intervals, the number of heartbeats NN50 where the difference between adjacent NN intervals is greater than 50 ms, the proportion pNN50 of NN50 among all NN intervals; the HRV triangular index HRVI; the width TINN of the base of the histogram of all NN intervals approximated as a triangle; the power of the very low frequency VLF, low frequency LF, and high frequency HF three frequency bands; the power ratios of LF / HF and VLF / HF; the approximate entropy ApEn, sample entropy SampEn, and Shannon entropy ShanEn; the SD1 and SD2 values of the Poincaré section; the eye movement tracking features include the blink frequency and duration calculated from the binocular images of consecutive video frames; the saccade frequency, duration, amplitude, delay period, and speed; the fixation frequency, duration, and dispersion; the pupil diameter.
[0026] Thereby, the physiological-comfort training samples are divided into heart rate-comfort training samples and eye movement tracking-comfort training samples; correspondingly, in step three, the physiological-base classifier is divided into a third classifier and a fourth classifier, and the third classifier is trained based on the heart rate-comfort training samples, and the fourth classifier is trained based on the eye movement tracking-comfort training samples.
[0027] In one embodiment, in step four, the test data includes test images and the corresponding viewer's physiological features. Before inputting the test data into the multimodal fusion model, first measure the physical parameters of the flat display screen, calculate the physical features in combination with the input test images, and then process the physical features and physiological features corresponding to the test images according to the preprocessing process in step two, including extracting imaging features and non-imaging features from the physical features, extracting heart rate variability features and eye movement tracking features from the physiological features, and then inputting them into the multimodal fusion model to obtain the visual comfort prediction result.
[0028] In one embodiment, in step four, it also includes the case of only inputting the test images or only inputting any one of the viewer's physiological features corresponding to the test images. At this time, according to the type of input data, extract imaging features and non-imaging features from the physical features, or extract heart rate variability features or eye movement tracking features from the physiological features, and then input them into the multimodal fusion model to obtain the visual comfort prediction result.
[0029] In one embodiment, in step four, when only the test image or any physiological feature of the viewer corresponding to the test image is input, according to the extracted feature type, the corresponding basic classifier is automatically matched; by disabling the output of the unmatched basic classifier, the stacked classifier adaptively fuses the matched basic classifiers and outputs the visual comfort prediction result.
[0030] In a second aspect, a flat panel display visual comfort prediction system based on a multi-modal fusion model is provided, including:
[0031] An augmented image sample module for collecting an original image set, where the original image set includes a number of original images. For each original image, a generative adversarial network model is used to convert the input random sequence into enhancement materials, and then the enhancement materials and the original images are fused according to a preset ratio to obtain a number of new images; the new images corresponding to the original image set form a new image set.
[0032] A training sample acquisition module for playing the original image set to the viewer, collecting the physiological features and visual comfort evaluations of the viewer; measuring the physical parameters of the flat panel display screen, and inputting the original image set and the new image set to obtain physical features; preprocessing the physical features and the physiological features to obtain physical-comfort training samples and physiological-comfort training samples.
[0033] A multi-modal fusion model training module for training a physical basic classifier and a physiological basic classifier respectively based on the physical-comfort training samples and the physiological-comfort training samples; training the stacked classifier based on a stacked integration framework and the trained physical basic classifier and physiological basic classifier; the basic classifier and the stacked classifier together constitute the multi-modal fusion model.
[0034] A visual comfort prediction module for inputting test data into the multi-modal fusion model to obtain a visual comfort prediction result.
[0035] In one embodiment, the visual comfort prediction module includes an adaptive unit for automatically matching the corresponding basic classifier according to the extracted feature type when only the test image or any physiological feature of the viewer corresponding to the test image is input; by disabling the output of the unmatched basic classifier, the stacked classifier adaptively fuses the matched basic classifiers and outputs the visual comfort prediction result.
[0036] The beneficial effects brought by the present invention are:
[0037] 1. The Generative Adversarial Network (GAN) model generates samples that retain its core features but have diverse details by learning the probability distribution of the original images. The present invention further proposes a proportional fusion strategy to mix the generated enhanced samples with the original images, which can not only retain the visual benchmark of the original images but also utilize the diversity of the fused samples to avoid the over-concentration of the generated content on local details, breaking through the potential mode collapse limitation of the generator. The fused samples maintain semantic consistency at the visual level, ensuring an approximate visual perception with the original images and allowing the same visual comfort labels to be assigned, thus alleviating the problem of scarce training data. At the same time, new information is introduced through the enhanced samples at the matrix value level, thereby enhancing the robustness and generalization ability of the model to input perturbations and noises, and achieving a balance between the retention of original features and the innovation of generated samples.
[0038] 2. The present invention is based on two types of data sources, physical features and physiological features, from which heterogeneous features that conform to the working principle of the human visual system are automatically extracted to construct a basic classifier. By utilizing the complementarity of physical features and physiological features, the influencing factors of visual comfort are comprehensively characterized. Subsequently, the outputs of the basic classifiers are integrated through a stacking structure, which not only improves the prediction accuracy by reducing the uncertainty between different source data but also enhances the robustness of the model by providing more comprehensive and accurate information and reducing the interference of outliers in a single data source. Moreover, since the features are extracted based on the working principle of the human visual system, the prediction results of the model trained based on these features are physiologically interpretable. Users can analyze the specific factors and manifestations of visual discomfort according to the contribution degree of different input features to the results, which has guiding significance for display design.
[0039] 3. The present invention proposes pre-training each basic classifier and then constructing a stacked classifier. During testing, the corresponding basic classifier is activated according to the input data type, and the basic classifiers of unmatched modalities are frozen. This design not only retains the domain knowledge of each modality's basic classifier during training but also avoids interference caused by modality loss or change through freezing. At the same time, the stacked classifier saves computational time and achieves efficient generalization of the input scenario, solving the limitation of traditional multi-modal models that rely on full-scale data input and is applicable to practical application scenarios with variable data collection conditions.
[0040] 4. Through the synergistic effect of generating diverse samples by the generative adversarial network, multi-modal feature extraction and stacking structure, and dynamic modal activation mechanism, the present invention realizes a positive cycle of data quality improvement - model performance enhancement - scenario generalization expansion. In particular, the complementary effect of diverse samples and the stacking structure: the diverse samples alleviate the problem of data scarcity, and the stacking structure maps the diverse features generated by the generative adversarial network to a shared feature space by integrating the output of the basic classifiers of physical and physiological modalities, avoiding both the mode collapse of the generated samples and enhancing the robustness of the model to input perturbations and noises through multi-modal information redundancy. The synergy between the two significantly improves the accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The present invention will be further described below with reference to the accompanying drawings.
[0042] Figure 1 is the overall design framework of the flat panel display visual comfort prediction method based on the multi-modal fusion model according to the embodiment of the present invention.
[0043] Figure 2 is the flow chart of sample diversification based on the generative adversarial network model according to the embodiment of the present invention.
[0044] Figure 3 is the flow chart of extracting imaging features and non-imaging features based on the physical features of the flat panel display according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. The description of at least one exemplary embodiment is actually only illustrative and in no way restrictive of the present invention and its application or use. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0046] The embodiment of the present invention provides a flat panel display visual comfort prediction method based on a multi-modal fusion model, as Figure 1 shown, including the following steps:
[0047] Step 1. Expand the image samples
[0048] Collect the original image set, which includes a number of original images. For each original image, use the generative adversarial network model to convert the input random sequence into enhancement materials, and then fuse the enhancement materials and the original images according to a preset ratio to obtain a number of new images; the new images corresponding to the original image set form a new image set;
[0049] Step 2. Obtain training samples
[0050] Play the original image set for the viewer, collect the physiological characteristics and visual comfort evaluation of the viewer; measure the physical parameters of the flat display screen, and input the original image set and the new image set to obtain physical characteristics; preprocess the physical characteristics and physiological characteristics to obtain physical-comfort training samples and physiological-comfort training samples;
[0051] Step 3: Train the multi-modal fusion model
[0052] Based on the physical-comfort training samples and physiological-comfort training samples, train the physical-based classifier and physiological-based classifier respectively; based on the stacking integration framework and the trained physical-based classifier and physiological-based classifier, train the stacking classifier; the basic classifier and the stacking classifier together constitute the multi-modal fusion model;
[0053] Step 4: Predict visual comfort
[0054] Input the test data into the multi-modal fusion model to obtain the visual comfort prediction result.
[0055] Specifically, in step 1, the generative adversarial network (GAN) model, as a type of generative artificial intelligence, its internal logic is to learn the joint probability distribution P(X,Y) of the sample (X) and the label (Y) by observing data and learn sufficient feature representations. Currently, a large number of comparative studies have confirmed that the GAN model has many advantages compared with other generative models: compared with the Variational Autoencoder (VAE), the bias is smaller; compared with the Deep Boltzmann Machine and the Generative Stochastic Network (GSN), samples can be generated at one time, and compared with the Non-linear Independent Components Estimation (NICE) and the Neural Variational Posterior (Real NVP), there is no limit on the size of the latent code.
[0056] The embodiment of the present invention does not adopt the method of directly generating new samples after inputting the original image into the generative adversarial network model because the directly generated samples are prone to mode collapse, that is, over-concentrating on local data patterns. The embodiment of the present invention adopts a proportional fusion strategy, which is specifically reflected in the mathematical formula as:
[0057] I S = ratio × I R +(1 - ratio) × I G
[0058] Where ratio is the fusion ratio, which is a random value within the range of 0.1 - 0.3 in the embodiment, and I S is the new image matrix value, and I R is the original image matrix value, and I G is the enhanced material matrix value generated by the GAN model. The purpose of this step is to select new samples that neither overly change the human visual perception nor have an obvious difference in matrix values from the original image, so as to avoid the local optimum problem that the GAN model may cause. The added enhanced material can be regarded as a perturbation to the original image, which helps the training model to be more robust in dealing with noise and transformation, and thus perform better in practical applications. The fused image maintains semantic consistency at the visual level and ensures an approximate visual perception with the original image. Therefore, for the new image, the same visual comfort label as the original image can be labeled.
[0059] Specifically, in step two, playing the original image set to the viewer and collecting the physiological characteristics and visual comfort evaluation of the viewer means arranging the viewer to watch the original image set on a flat panel display in a dark room. For each original image watched, the electrocardiogram signal (ECG) and eye movement tracking video (EM) should be recorded throughout the process, and after watching, the viewer is required to fill out a questionnaire to evaluate the visual comfort level, so as to obtain the physiological characteristics and corresponding visual comfort of the viewer when watching the original image, that is, the physiological-comfort training samples described above.
[0060] Measuring the flat panel display screen means using a photometer, colorimeter, and spectrometer to measure the red, green, and blue primary colors of the flat panel display, as well as the brightness, color coordinates, and spectra of all levels of input signals for the full-screen white field, so as to obtain the true gamma curve, color gamut, and spectral distribution of the display. On this basis, the signal values of the original image set and the new image set can be calculated to obtain the physical characteristics corresponding to the original image set and the new image set.
[0061] Since the original image has been played to the viewer, the physical characteristics corresponding to the original image set can directly label the visual comfort labels filled in by the viewer, that is, the physical-comfort training samples corresponding to the original image set; for the new image, as described above, it is generated by fusing the original image and maintains semantic consistency at the visual level with its corresponding original image, having an approximate visual perception. Therefore, the same visual comfort label as its corresponding original image can be labeled, and thus the physical-comfort training samples corresponding to the new image set are obtained; the physical-comfort training samples corresponding to the original image set and the physical-comfort training samples corresponding to the new image set together constitute the physical-comfort training samples, which are used to train the subsequent physical basis classifier.
[0062] In this embodiment, the visual comfort level labels are divided into three categories: discomfort, neutral, and comfort. In practical applications, the classification method of subjective evaluation of visual comfort can be adjusted according to needs, such as being divided into very uncomfortable, uncomfortable, average, comfortable, very comfortable, etc. Since this embodiment adopts the method of expanding image samples in Step 1, only a small amount of physiological characteristics and visual comfort evaluations of viewers need to be collected, and then the subsequent multimodal fusion model can be built, and the prediction result accuracy is good.
[0063] In one implementation, as Figure 2 shown, in Step 1, the generative adversarial network model includes a generator and a discriminator:
[0064] The generator is composed of several cascaded deep 2D convolutional blocks, and each convolutional layer uses the ReLU activation function;
[0065] The discriminator is composed of several cascaded deep 2D convolutional blocks, and each convolutional layer uses the leaky ReLU activation function;
[0066] Through unsupervised deep learning, the generator and the discriminator are trained simultaneously. The training objective of the generator is set to generate enhanced materials that are substantially equivalent to the original image in topological structure, that is, enhanced materials as similar as possible to the original image; the training objective of the discriminator is set to accurately distinguish the enhanced materials from the original image.
[0067] Specifically, the specific designs of the generator and the discriminator in this embodiment are shown in the following table:
[0068]
[0069] Specifically, the training objective of the generator is set to generate enhanced materials as similar as possible to the original image, that is, to make the loss function Loss G of the sample as small as possible, so that the generated enhanced materials can pass the judgment of the discriminator. Specifically, in the mathematical formula, it is: Loss G =-average(log(P G-R )). The training objective of the discriminator is set to distinguish the enhanced materials from the original image as accurately as possible, maximize the recognition probability of the original image and minimize the misjudgment probability of the enhanced sample, that is, to make the loss function Loss D of the sample as small as possible. Specifically, in the mathematical formula, it is: Loss D =-average(log(P F-R ))-average(log(1 - P G-R )). In the formula, P G-R is the probability that the discriminator recognizes the generated enhanced material as true, and P R-R is the probability that the discriminator recognizes the original image as true.
[0070] In one embodiment, the original image is rotated, translated, and reflected before being input into the generative adversarial network model.
[0071] The methods adopted in this embodiment are rotation (positive / negative 10°), translation (10 pixels in the left / right and up / down directions), and reflection (50% probability) to successively improve the robustness of the generative adversarial network model.
[0072] In one embodiment, in step two, preprocessing the physical features is to extract imaging features and non-imaging features from the physical features; the physical features at least include the brightness, color coordinates, and spectrum of all input signal values of all pixel points of the flat display screen; the imaging features include the brightness, chroma, hue matrix in the LCH color space, and the phase consistency matrix in the frequency domain; the non-imaging features include the retinal irradiance map corresponding to the retinal photoreceptors;
[0073] Thus, the physical-comfort training samples are divided into imaging-comfort training samples and non-imaging-comfort training samples; correspondingly, in step three, the physical-base classifier is divided into a first classifier and a second classifier, and the first classifier is trained based on the imaging-comfort training samples, and the second classifier is trained based on the non-imaging-comfort training samples.
[0074] Specifically, as Figure 3 shown, the features extracted from the physical features are not limited to the following categories, and other features can be extracted according to actual needs. For the imaging features, they are divided into two parts:
[0075] (1) Based on the measured RGB gamma curve and color coordinates, according to the order from RGB to XYZ, from XYZ to CIE La * b * , perform color space conversion in the order from CIELa * b * to LCH. In the LCH color space, extract the feature matrices of brightness, chroma, and hue;
[0076] (2) Convert the image to the frequency domain through the Fast Fourier Transform (FFT), and calculate the feature matrix of phase congruency, and the formula is as follows:
[0077]
[0078] where PC is the phase congruency, is the local phase, and A n is the local amplitude or energy. When all phases are aligned, the PC value is equal to 1.
[0079] For non-imaging features, first, based on the measured RGB spectrum, the actual spectrum of each pixel of the flat panel display is calculated by the method of adding the base color spectra. Then, according to the retinal photoreceptor sensitivity curve provided by the CIE S 026:2018 international standard of the International Commission on Illumination, the characteristic matrix of the retinal irradiance of five photoreceptor cells, namely S cone cells, M cone cells, L cone cells, rod cells, and ipRGC cells on the retina, is calculated. The formula is as follows:
[0080] E α =∫E e,λ (λ)s α (λ)dλ
[0081] where α corresponds to the five photoreceptor cells, E α is the retinal irradiance of this photoreceptor cell, E e,λ (λ) is the spectral irradiance at wavelength λ, and s α (λ) is the spectral sensitivity curve of the photoreceptor cell at wavelength λ.
[0082] Specifically, the physical-based classifier is divided into a first classifier and a second classifier, which means in this embodiment:
[0083] Input the imaging features, and classify the subjective evaluation of visual comfort by constructing a convolutional neural network (CNN network) to construct the first classifier;
[0084] Input the non-imaging features, and classify the subjective evaluation of visual comfort by constructing a convolutional neural network (CNN network) to construct the second classifier;
[0085] The structures of the first and second classifier models are the same, as shown in the following table:
[0086] Input layer Projection and reshaping layer 2D convolutional layer (3 layers in total) Fully connected layer Softmax layer with weights (the weights of classes are inversely proportional to the probabilities of samples in these classes)
[0087] Most of the calculations are performed in three 2D convolutional layers. In each layer, a filter of size 5×5 scans the entire input matrix, calculates the dot product with the pixel points, and serves as the output array. The first layer focuses on simple features. As the layers progress, the complexity gradually increases, and more scale information of the input matrix can be extracted. In the fully connected layer, each node in the output layer is directly connected to a node in the previous layer. In the last Softmax layer, the final class probability value is output:
[0088]
[0089] Among them, P is the output probability, j is the j-th classification, x is the sample vector, w is the weighted vector, and K is related to the number of nodes in the previous layer. The category with the highest probability is the prediction result.
[0090] In this embodiment, the training strategy of the CNN network adopts 10-fold cross-validation. Each time, 10% of the samples in the dataset are randomly selected as the test set, and the rest are used as the training set, repeating 10 times. In the embodiment, the first and second classifier models are trained using Stochastic Gradient Descent with Momentum (SGDM). The advantage of this method is that it adjusts the parameters according to the update trend and will not get stuck at points with relatively small current gradients, and converges relatively stably. In practical applications, a suitable model can also be selected for training according to the actual situation.
[0091] In one implementation, in step two, preprocessing the physiological features is to extract heart rate variability features and eye movement tracking features from the physiological features; the physiological features at least include electrocardiogram signals and eye movement tracking videos; the heart rate variability features include the average RR interval; the standard deviation of NN intervals SDNN; the root mean square of the differences between adjacent NN intervals RMSSD; the number of heart beats NN50 where the difference between adjacent NN intervals is greater than 50 ms among all NN intervals, the proportion pNN50 of NN50 in all NN intervals; the HRV triangular index HRVI; the width TINN of the base of the histogram approximated as a triangle of all NN intervals; the power of the very low frequency VLF, low frequency LF, and high frequency HF frequency bands; the power ratios of LF / HF and VLF / HF; the approximate entropy ApEn, sample entropy SampEn, and Shannon entropy ShanEn; the SD1 and SD2 values of the Poincaré section; the eye movement tracking features include the blink frequency and duration calculated from the binocular images of consecutive video frames; the saccade frequency, duration, amplitude, latency, and velocity; the fixation frequency, duration, and dispersion; the pupil diameter.
[0092] Thus, the physiological-comfort training samples are divided into heart rate-comfort training samples and eye movement tracking-comfort training samples; correspondingly, in step three, the physiological-base classifier is divided into a third classifier and a fourth classifier, and the third classifier is trained based on the heart rate-comfort training samples, and the fourth classifier is trained based on the eye movement tracking-comfort training samples.
[0093] Specifically, the features extracted from the physiological features are not limited to the following categories, and other features can be extracted according to actual needs. The following features are only examples:
[0094] Using software such as Matlab (accompanied by dedicated toolboxes such as ECG analysis and Eyelink), physiological features of human electrocardiogram signals and eye movement tracking are extracted. The biological features extracted from the electrocardiogram signals include heart rate (HR), mean RR interval, standard deviation of NN intervals (SDNN), root mean square of the differences between adjacent NN intervals (RMSSD), number of heartbeats with a difference between adjacent NN intervals greater than 50 ms among all NN intervals (NN50), proportion of NN50 in all NN intervals (pNN50), HRV triangular index (HRVI), width of the base of the histogram approximation of all NN intervals that is triangular (TINN), very low frequency (VLF), low frequency (LF), high frequency (HF) power in three frequency bands, power ratios of LF / HF and VLF / HF, approximate entropy (ApEn), sample entropy (SampEn), Shannon entropy (ShanEn), SD1 and SD2 values of the Poincaré section; the biological features extracted from the eye movement tracking video include, from consecutive frames of the binocular images, calculating the frequency (Hz) and duration (ms) of blinks, saccades and fixations, amplitude (°) of saccades, latency period (ms) and speed (° / s), dispersion of fixations (px), pupil diameter (mm).
[0095] Specifically, the physiology-based classifier is divided into a third classifier and a fourth classifier, which means in this embodiment:
[0096] Input the biological features extracted from the electrocardiogram signal, construct a decision tree model, classify the subjective evaluation of visual comfort, and construct a third classifier;
[0097] Input the biological features extracted from the eye movement tracking image, construct a decision tree model, classify the subjective evaluation of visual comfort, and construct a fourth classifier.
[0098] As a tree structure, the decision tree has good robustness to data sources with strong noise such as physiological signals. Each non-leaf node of the decision tree represents a test on a feature attribute, each leaf node stores a category, and each branch represents the output of this feature attribute in a certain value range. The training of the decision tree also uses 10-fold cross-validation. During training, the training set is split into multiple subsets. For each subset, starting from the root node, test the feature attributes and select the output branch according to its value until reaching the leaf node, and use the category stored in the leaf node as the classification result. When the classification labels of a subset are the same, the repeated recursion stops. In this embodiment, the classification with the minimum entropy (maximum probability) is selected as the prediction result:
[0099]
[0100] where E is entropy, i is the i-th classification, p iThe probability of the i-th class.
[0101] In practical applications, an appropriate model can also be selected for training according to the actual situation.
[0102] In one implementation, in step four, the test data includes test images and corresponding viewer physiological characteristics. Before inputting the test data into the multi-modal fusion model, first measure the physical parameters of the flat display screen, calculate the physical characteristics by combining the signal values of the input test images, and then process the physical characteristics and physiological characteristics corresponding to the test images according to the preprocessing process in step two, including extracting imaging features and non-imaging features from the physical characteristics, and extracting heart rate variability features and eye movement tracking features from the physiological characteristics, and then input them into the multi-modal fusion model to obtain the visual comfort prediction result.
[0103] Specifically, most of the existing research inputs the RGB value matrix of the display image and the original physiological data into the model without processing for training, which will lead to poor classification performance of the model. The beneficial effect of first extracting meaningful features from the measurement results and physiological data based on the existing experimental research results of visual comfort and then constructing the basic model is that it can measure the influencing factors and physiological manifestations of visual comfort from more dimensions, extract multi-source heterogeneous information from limited data, and thus improve the prediction accuracy of the stacked classifier; and because the features are extracted based on the working principle of the human visual system, the prediction results of the model trained based on the features have physiological interpretability, and users can analyze the specific factors and manifestations of visual discomfort according to the contribution degree of different inputs to the results, which has guiding significance for display design.
[0104] In one implementation, in step four, it also includes the case of only inputting the test image or only inputting any one physiological characteristic of the viewer corresponding to the test image. At this time, according to the type of input data, extract imaging features and non-imaging features from the physical characteristics, or extract heart rate variability features or extract eye movement tracking features from the physiological characteristics, and then input them into the multi-modal fusion model to obtain the visual comfort prediction result.
[0105] For example, when only the user's ECG signal is input, or only the display image matrix is input, the stacked classifier can use the corresponding classifier as the base model to output the predicted visual comfort level. The more data is input and the more base models there are, the higher the prediction accuracy will be. This application realizes information complementarity based on multi-modal feature fusion. Compared with other models that can only predict visual comfort level by inputting a certain specific type of data, this application only needs to input one or more features corresponding to the screen image according to the existing experimental data, and the algorithm automatically extracts optical and biological information, then the stacked classifier can output the predicted visual comfort level, significantly broadening the scope of application. For example, when other laboratories only collect one or several features, the method of this application can be used to predict visual comfort level based on limited experiments. Moreover, this application also significantly improves the accuracy and stability of the prediction of the stacked classifier through ensemble learning.
[0106] In one implementation, in step four, when only the test image is input or only any physiological feature of the viewer corresponding to the test image is input, according to the extracted feature type, the corresponding base classifier is automatically matched; by disabling the output of the unmatched base classifiers, the stacked classifier adaptively fuses the matched base classifiers and outputs the visual comfort level prediction result, achieving the purpose of flexibly adapting to different scenarios.
[0107] Specifically, the advantage of the stacked classifier is that the predictions (or their errors) made are uncorrelated or low-correlated. By providing more comprehensive and accurate information, it can prevent the wrong guidance of the classification result by a single modality. In this embodiment, the commonly used random forest model in the relevant field is adopted as the meta-model, and the first to fourth classifiers are used as the base models. The input is the prediction scores of the base models for the training set samples, that is, the class probability values of the three labels. The training also uses 10-fold cross-validation. When the input test data decreases, the stacked classifier can also disable the output of the unmatched base classifiers, and according to the performance and characteristics of the matched base classifiers, assign different weights to the outputs of each base classifier, so as to more flexibly fuse their results to produce the final prediction. This adaptive ability enables the stacked classifier to have better robustness when facing changes in the base classifiers or data distribution changes.
[0108] In addition, as a new artificial intelligence technology, ensemble learning has been applied in fields such as medical data analysis and security detection. Among them, Bagging considers homogeneous learners and conducts parallel learning; Boosting considers homogeneous learners and conducts sequential learning. In this embodiment, first, for datasets of different modalities, heterogeneous learners are respectively designed as the basic models and trained in parallel to save computing time. Then, based on the prediction scores of the trained basic models, a stacking classifier is trained to enable it to learn the complementary patterns between modalities. The above advantages cannot be achieved by the majority voting or averaging of bagging and boosting, and are more practically significant.
[0109] An embodiment of the present invention further provides a flat panel display visual comfort prediction system based on a multi-modal fusion model, including:
[0110] An augmented image sample module, configured to collect an original image set, the original image set including a plurality of original images. For each original image, a generative adversarial network model is used to convert an input random sequence into enhanced material, and then the enhanced material and the original image are fused according to a preset ratio to obtain a plurality of new images; the new images corresponding to the original image set form a new image set;
[0111] A training sample acquisition module, configured to play the original image set to a viewer, collect the physiological characteristics and visual comfort evaluation of the viewer; measure the physical parameters of the flat panel display screen, and input the original image set and the new image set to obtain physical characteristics; preprocess the physical characteristics and the physiological characteristics to obtain physical-comfort training samples and physiological-comfort training samples;
[0112] A multi-modal fusion model training module, configured to train a physical-based classifier and a physiological-based classifier respectively based on the physical-comfort training samples and the physiological-comfort training samples; train a stacking classifier based on a stacking integration framework and the completed physical-based classifier and physiological-based classifier; the basic classifier and the stacking classifier together constitute a multi-modal fusion model;
[0113] A visual comfort prediction module, configured to input test data into the multi-modal fusion model to obtain a visual comfort prediction result.
[0114] In an implementation manner, the visual comfort prediction module includes an adaptive unit, configured to automatically match a corresponding basic classifier according to the extracted feature type when only a test image or any one physiological characteristic of the viewer corresponding to the test image is input; by disabling the outputs of the unmatched basic classifiers, the stacking classifier adaptively fuses the matched basic classifiers and outputs a visual comfort prediction result.
[0115] The advantages of the present invention are:
[0116] 1. The generative adversarial network (GAN) model generates samples that retain its core features but have diverse details by learning the statistical laws of the original images. The present invention further proposes a proportional fusion strategy to mix the generated enhanced samples with the original images, which can not only retain the visual benchmark of the original images but also utilize the diversity of the fused samples to avoid the over - concentration of the generated content on local details, breaking through the potential mode collapse limitation of the generator. The fused samples maintain semantic consistency at the visual level, ensuring an approximate visual perception with the original images and can be labeled with the same visual comfort label to alleviate the problem of scarce training data. At the same time, new information is introduced through the enhanced samples at the matrix value level, thereby improving the robustness and generalization ability of the model to input perturbations and noises, and achieving a balance between retaining the original features and innovating the generated samples.
[0117] 2. The present invention is based on two types of data sources, physical features and physiological features, and automatically extracts heterogeneous features that conform to the working principle of the human visual system from them to construct a basic classifier. By utilizing the complementarity of physical features and physiological features, the influencing factors of visual comfort are comprehensively characterized. Then, the outputs of the basic classifiers are integrated through a stacking structure, which not only improves the prediction accuracy by reducing the uncertainty between different source data but also enhances the robustness of the model by providing more comprehensive and accurate information and reducing the interference of outliers in a single data source. And because the features are extracted based on the working principle of the human visual system, the prediction results of the model trained based on the features have physiological interpretability. Users can analyze the specific factors and manifestations of visual discomfort that cause visual discomfort according to the contribution degree of different input features to the results, which has guiding significance for display design.
[0118] 3. The present invention proposes to pre - train each basic classifier and then construct a stacked classifier. During testing, the corresponding basic classifier is activated according to the input data type, and the basic classifiers of unmatched modalities are frozen. This design not only retains the domain knowledge of each modal basic classifier in pre - training but also avoids cross - modal interference through freezing. At the same time, it saves computing time through the stacked classifier and achieves efficient generalization of the input scenario, solving the limitation of traditional multi - modal models relying on full - volume data input and being applicable to practical application scenarios with variable data acquisition conditions.
[0119] 4. Through the synergistic effect of generating diverse samples by the generative adversarial network, multi-modal feature extraction and stacking structure, and dynamic modal activation mechanism, the present invention realizes a positive cycle of data quality improvement - model performance enhancement - scenario generalization expansion. In particular, the complementary effect of diverse samples and stacking structure: diverse samples alleviate the problem of data scarcity, and the stacking structure maps the diverse features generated by the generative adversarial network to a shared feature space by integrating the output of basic classifiers of physical and physiological modalities, avoiding both the mode collapse of generated samples and enhancing the robustness of the model to input perturbations through multi-modal information redundancy. The synergy of the two significantly improves the accuracy of the model.
[0120] Verify the effects of the embodiments of the present invention according to the following method: use common evaluation metrics such as accuracy (ACC), precision (Pre), recall (Rec), and weighted F1 score (F1 W ) to measure the advantages and disadvantages of this application. The closer the above values are to 1, the higher the prediction accuracy of the multi-modal fusion model:
[0121]
[0122] (i is each classification)
[0123] Among them:
[0124] Pre w = W1×Pre1 + W2×Pre2 + W3×Pre3
[0125] Rec W = W1×Rec1 + W2×Rec2 + W3×Rec3
[0126] Among them, i refers to each classification. For category i, TP i (True Positive) refers to the weighted samples correctly predicted by the model, TN i (True Negative) refers to the samples of other classes correctly predicted by the model, FP i (False Positive) refers to the samples of other classes predicted as this class by the model, FN i (False Negative) refers to the samples of this class predicted as other classes by the model, W i refers to the weight of this class, which is proportional to the occurrence frequency of samples of classification i and the sum is 1.
[0127] The saturation-visual comfort dataset containing 1120 samples (dataset source: Yunyang Shi, Yan Tu, Lili Wang, Xin Gao. 2019. P-33: Effects of luminance, contrast and saturation of HDR QLED display on visual system based on eye movement. SIDSymposium Digest of Technical Papers, 50) was used for training and testing. The experiment corresponding to the dataset involved 14 subjects. Each subject watched pictures under 5 saturation settings of a TV monitor in a dark room in random order. The electrocardiogram signals and eye movement tracking videos were recorded throughout the process, and a psychological evaluation score of the visual comfort level under different settings was given after watching. The duration of each experiment was about 50 minutes. The dataset samples obtained included the displayed content, electrocardiogram signals, eye movement tracking videos, and the subjective evaluation of the visual comfort level as labels. The label distribution was: 21.52% discomfort, 49.37% no feeling, and 29.11% comfort.
[0128] The prediction performance of the trained model on the test set is shown in the following table:
[0129]
[0130] In the table, "no sample diversification" means directly using I R (the original samples not mixed with the GAN output) instead of I G (the diversified samples mixed with the GAN output) as the input of the "display feature extraction" part. "Does not contain human information" means only inputting physical features (imaging features and non-imaging features); "does not contain display information" means only inputting physiological features. "Does not adopt ensemble learning" means only using a single model classifier with four inputs for prediction, rather than the stacked fusion of multiple models.
[0131] To verify the performance of the model under different display usage scenarios, the model corresponding to the above table was used as a pre-trained model for transfer learning on the brightness-display quality dataset (dataset source: Xin Gao, Yan Tu, Lili Wang, Yunyang Shi, Wei Zhang. 2019. The effect of luminance on visual perception based on eyemovement and ECG. 2019 3rd International Conference on Circuits, System andSimulation (ICCSS), 221-224). In the experiment corresponding to the dataset, another 26 subjects watched pictures under 5 brightness settings of a TV monitor and evaluated the degree of visual fatigue (other settings of the experiment were the same as those of the saturation-visual comfort dataset). The obtained dataset samples included display information, ECG, and EM signals, and the label distribution was as follows: 8.15% moderate fatigue, 43.70% mild fatigue, and 48.15% no fatigue. The prediction performance of the model was: ACC = 0.78 ± 0.17, F1 W = 0.72 ± 0.24. It shows that the model of the embodiment of the present invention has the advantages of being flexible and stable and being able to adapt to different environments and tasks.
Claims
1. A method for predicting the visual comfort of flat panel displays based on a multi-modal fusion model, comprising the following steps: Step 1: Expand image samples Collect the original image set, which includes a number of original images. For each original image, use the generative adversarial network model to convert the input random sequence into enhancement materials, and then fuse the enhancement materials and the original images according to a preset ratio to obtain a number of new images; the new images corresponding to the original image set form the new image set. Step 2: Obtain training samples Play the original image set to viewers, and collect the physiological characteristics and visual comfort evaluations of the viewers; measure the physical parameters of the flat panel display screen, and input the original image set and the new image set to obtain physical characteristics. Preprocess the physical characteristics and physiological characteristics to obtain physical-comfort training samples and physiological-comfort training samples. Step 3: Train the multi-modal fusion model Train the physical-based classifier and the physiological-based classifier respectively based on the physical-comfort training samples and the physiological-comfort training samples; based on the stacked integration framework and the trained physical-based classifier and physiological-based classifier, train the stacked classifier; the basic classifier and the stacked classifier together constitute the multi-modal fusion model. Step 4: Predict visual comfort Input the test data into the multi-modal fusion model to obtain the visual comfort prediction result.
2. The flat panel display visual comfort prediction method based on a multi-modal fusion model according to claim 1, wherein: In Step 1, the generative adversarial network model includes a generator and a discriminator: The generator is composed of several cascaded deep 2D convolutional blocks, and each convolutional layer uses the ReLU activation function. The discriminator is composed of several cascaded deep 2D convolutional blocks, and each convolutional layer uses the leaky ReLU activation function. Through unsupervised deep learning, train the generator and the discriminator at the same time, set the training target of the generator to generate enhancement materials that are substantially equivalent to the original images in topological structure; set the training target of the discriminator to accurately distinguish the enhancement materials from the original images.
3. Any one of the flat panel display visual comfort prediction methods based on a multimodal fusion model according to claim 2, characterized in that: Rotate, translate, and reflect the original images and then input them into the generative adversarial network model.
4. A method for predicting the visual comfort of a flat panel display based on a multimodal fusion model according to claim 1, characterized in that: In Step 2, preprocessing the physical characteristics is to extract imaging characteristics and non-imaging characteristics from the physical characteristics; the physical characteristics at least include the brightness, color coordinates, and spectrum of all input signal values of all pixel points of the flat panel display screen. The imaging characteristics include the brightness, chromaticity, and hue matrix in the LCH color space, and the phase consistency matrix in the frequency domain. The non-imaging characteristics include the retinal irradiance map corresponding to the retinal receptors. Thus, the physical-comfort training samples are divided into imaging-comfort training samples and non-imaging-comfort training samples; correspondingly, in Step 3, the physical-based classifier is divided into a first classifier and a second classifier, and the first classifier is trained based on the imaging-comfort training samples, and the second classifier is trained based on the non-imaging-comfort training samples.
5. A flat panel display visual comfort prediction method based on a multi-modal fusion model according to claim 1, characterized in that: In Step 2, preprocessing the physiological characteristics is to extract heart rate variability characteristics and eye movement tracking characteristics from the physiological characteristics; the physiological characteristics at least include electrocardiogram signals and eye movement tracking videos. Heart rate variability features include the average RR interval; the standard deviation of NN intervals SDNN; the root mean square of the differences between adjacent NN intervals RMSSD; the number of heartbeats NN50 where the difference between adjacent NN intervals is greater than 50 ms among all NN intervals, and the proportion pNN50 of NN50 in all NN intervals; the HRV triangular index HRVI; the width TINN of the base of the histogram approximated as a triangle for all NN intervals; the powers of the very low frequency VLF, low frequency LF, and high frequency HF frequency bands; the power ratios of LF / HF and VLF / HF; the approximate entropy ApEn, sample entropy SampEn, and Shannon entropy ShanEn; the SD1 and SD2 values of the Poincaré section; Eye movement tracking features include the blink frequency and duration calculated from the binocular images of consecutive video frames; the saccade frequency, duration, amplitude, latency period, and velocity; the fixation frequency, duration, and dispersion; Pupil diameter; Thereby, the physiological-comfort training samples are divided into heart rate-comfort training samples and eye movement tracking-comfort training samples; correspondingly, in step three, the physiological-base classifier is divided into a third classifier and a fourth classifier, and the third classifier is trained based on the heart rate-comfort training samples, and the fourth classifier is trained based on the eye movement tracking-comfort training samples.
6. Any of the flat panel display visual comfort prediction methods based on a multimodal fusion model according to claim 4 or 5, characterized in that: In step four, the test data includes test images and the corresponding viewer physiological characteristics. Before inputting the test data into the multi-modal fusion model, first measure the physical parameters of the flat display screen, calculate and obtain the physical characteristics in combination with the test images, and then process the physical characteristics and physiological characteristics corresponding to the test images according to the preprocessing process in step two, including extracting imaging features and non-imaging features from the physical characteristics, and extracting heart rate variability features and eye movement tracking features from the physiological characteristics, and then input them into the multi-modal fusion model to obtain the visual comfort prediction result.
7. The method for predicting visual comfort of a flat panel display based on a multimodal fusion model according to claim 6, characterized in that: In step four, it also includes the case of only inputting the test image or only inputting any one of the viewer physiological characteristics corresponding to the test image. At this time, according to the type of input data, extract imaging features and non-imaging features from the physical characteristics, or extract heart rate variability features or eye movement tracking features from the physiological characteristics, and then input them into the multi-modal fusion model to obtain the visual comfort prediction result.
8. The method for predicting the visual comfort of a flat panel display based on a multimodal fusion model according to claim 7, characterized in that: In step four, when only the test image or any one of the viewer physiological characteristics corresponding to the test image is input, automatically match the corresponding base classifier according to the extracted feature type; by disabling the output of the unmatched base classifier, stack the classifiers to adaptively fuse the matched base classifiers and output the visual comfort prediction result.
9. A flat display visual comfort prediction system based on a multi-modal fusion model, comprising: An augmented image sample module for collecting an original image set. The original image set includes a number of original images. For each original image, use a generative adversarial network model to convert the input random sequence into enhanced materials, and then fuse the enhanced materials and the original images according to a preset ratio to obtain a number of new images; the new images corresponding to the original image set form a new image set; A training sample acquisition module, which is used to play an original image set to viewers, collect the physiological characteristics and visual comfort evaluations of the viewers; measure the physical parameters of a flat display screen, and input the original image set and a new image set to obtain physical characteristics; Preprocess the physical characteristics and physiological characteristics to obtain physical-comfort training samples and physiological-comfort training samples; A multi-modal fusion model training module, which is used to train a physical-based classifier and a physiological-based classifier respectively based on the physical-comfort training samples and the physiological-comfort training samples; train a stacked classifier based on a stacked integration framework and the trained physical-based classifier and physiological-based classifier; the basic classifier and the stacked classifier together constitute a multi-modal fusion model; A visual comfort prediction module, which is used to input test data into the multi-modal fusion model to obtain a visual comfort prediction result.
10. The flat panel display visual comfort prediction system based on a multi-modal fusion model according to claim 9, characterized in that: The visual comfort prediction module includes an adaptive unit, which is used to automatically match the corresponding basic classifier according to the extracted feature type when only the test image or any one of the physiological characteristics of the viewer corresponding to the test image is input; by disabling the outputs of the unmatched basic classifiers, the stacked classifier adaptively fuses the matched basic classifiers and outputs a visual comfort prediction result.
Citation Information
Patent Citations
Human behavior recognition method and system based on multi-mode deep Boltzmann machine
CN107886061A
Passenger comfort evaluation system and method based on multi-modal physiological data
CN116965830A
Multi-modal online car-hailing comfort evaluation network model and construction method thereof
CN119227510A
Systems, devices, and methods for generating and manipulating objects in a virtual reality or multi-sensory environment to maintain a positive state of a user
US20230012960A1