Quality intelligent evaluation method for light field 3D display device
By introducing the CLIP model and semantic text prompts, a cross-modal aligned quality assessment method is constructed, which solves the problem of insufficient semantic information in the quality assessment of existing light field 3D display devices, achieves more accurate quality prediction and optimization guidance, and improves user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIANJIN UNIV
- Filing Date
- 2026-04-24
- Publication Date
- 2026-05-29
Smart Images

Figure CN122115429A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimedia information processing, and in particular to a method for intelligent quality evaluation of light field 3D display devices. Background Technology
[0002] As an emerging technology in the field of 3D display, light field 3D display technology, by simultaneously reconstructing the intensity and direction information of light, can provide viewers with a continuous viewpoint stereoscopic visual experience without auxiliary equipment, and has become a hot topic in current research and application. However, in practical applications, the display quality of different light field 3D display devices varies, and some light field 3D display devices with poor display quality can easily lead to visual fatigue and discomfort, thus affecting the visual experience. Therefore, studying the potential correlation between the quality of light field 3D display devices and the viewer's visual experience, and realizing efficient quality evaluation of light field 3D display devices, is of great significance for optimizing the design of light field 3D display devices, improving the user viewing experience, and promoting the industrialization of light field 3D display technology.
[0003] Traditional methods for evaluating the quality of 3D display devices typically require analyzing viewer visual experience data to assess 3D display performance, resulting in a complex evaluation process with limited intelligence. To improve the intelligence of 3D display device quality evaluation, Pan et al., based on a stereo image quality evaluation network backbone, guided crosstalk feature learning through color difference maps and combined it with a spatial-channel attention mechanism to simulate binocular interaction, effectively modeling the relationship between 3D display device quality and viewer visual experience, thus achieving intelligent quality evaluation of 3D display devices. However, existing 3D display device quality evaluation methods mainly rely on data-driven visual feature learning, rarely mining semantic-level perceptual quality information and lacking an understanding of subjective perception of 3D quality. In recent years, visual language foundational models have developed rapidly, demonstrating great potential in visual understanding tasks such as object recognition, scene classification, and semantic segmentation. As one of the representative visual language foundational models, the CLIP (Contrastive Language-Image Pre-training) model, through pre-training on large-scale image-text pairs, can learn visual-semantic aligned representations and possesses strong cross-task transfer capabilities. Introducing the CLIP model into the quality evaluation task of light field 3D display devices, and making full use of its rich semantic knowledge and cross-modal alignment capabilities, is expected to improve the quality evaluation performance of light field 3D display devices.
[0004] Existing quality evaluation methods for light field 3D display devices mainly rely on data-driven visual feature learning, with less emphasis on mining semantic-level perceptual quality and stereoscopic information. This results in models lacking an understanding of the subjective perception of the quality and stereoscopic effect of light field 3D display devices, thus limiting the intelligent development of quality evaluation methods for light field 3D display devices. Summary of the Invention
[0005] This invention provides a method for intelligent quality evaluation of light field 3D display devices. This invention fully utilizes the cross-modal understanding and reasoning capabilities of the visual language basic model, and introduces overall display quality descriptions and stereoscopic quality descriptions as semantic priors to achieve intelligent quality evaluation of light field 3D display devices. See the description below for details:
[0006] A method for intelligent quality evaluation of a light field 3D display device, the method comprising:
[0007] Construct overall display quality description text prompts and stereoscopic quality description text prompts, which are used to provide explicit semantic guidance for the quality perception process of left and right viewpoint images and their corresponding parallax information, respectively;
[0008] The constructed overall display quality description text prompts and stereoscopic quality description text prompts are respectively input into the text encoder in the pre-trained CLIP model to obtain semantic features containing the overall display quality perception prior and the stereoscopic quality perception prior, respectively.
[0009] The content of the light field 3D display device is captured by a binocular camera, and the binocular stereo images containing left and right viewpoint images are used as input to the model to construct an overall display quality prediction module based on the left and right viewpoint images.
[0010] A stereoscopic quality prediction module based on parallax information is constructed to predict stereoscopic quality by introducing parallax information, thereby helping to improve the quality evaluation performance of light field 3D display devices.
[0011] An end-to-end approach is used to train a quality intelligent evaluation network for light field 3D display devices, and the trained network guides the optimization of light field 3D display devices.
[0012] The overall display quality description text prompt and the stereoscopic quality description text prompt are respectively:
[0013] The template for the overall display quality description text is "a stereoscopic image pair with {c} quality."; the template for the stereoscopic perception quality description text is "a disparity map with {s} stereoscopic perception.";
[0014] Where {c} represents the overall display quality level of the displayed content, {s} represents the stereoscopic quality level of the displayed content, and the candidate text set corresponding to the overall display quality description text template is denoted as . The candidate text set corresponding to the text template describing the three-dimensional quality is denoted as . .
[0015] The semantic features for obtaining the prior knowledge of overall display quality perception and prior knowledge of stereoscopic quality perception are as follows:
[0016] Candidate text sets describing different quality levels will be used. and The text encoder in the pre-trained CLIP model is input separately to obtain semantic features that imply an overall display quality-aware prior. and semantic features that imply prior knowledge of stereoscopic quality .
[0017] The overall display quality prediction module based on left and right viewpoint images is as follows:
[0018] Left and right viewpoint images and The image encoders in the pre-trained CLIP model are input separately to obtain the visual features of the left viewpoint image corresponding to the left and right viewpoint images. Visual features of the right viewpoint image ;
[0019] Based on the visual features of the obtained left viewpoint image Visual features of the right viewpoint image and overall display quality semantic features Calculate separately and and and Cosine similarity between the two views is used to obtain the quality mapping value of the left viewpoint image after cross-modal alignment. Image quality mapping value for right viewpoint ;
[0020] The Softmax function is used to map the quality values of the obtained left viewpoint image. Image quality mapping value for right viewpoint Normalization is performed to obtain the probability distribution of the predicted image quality for the left viewpoint. And the probability distribution of prediction quality for right viewpoint images ;
[0021] Predicting quality probability distribution of the obtained left viewpoint image And the probability distribution of prediction quality for right viewpoint images Perform weighted fusion to obtain the overall display quality probability distribution of the fused data. :
[0022] The quality probability distribution after fusion Weighted summation is performed to achieve score continuity and obtain the binocular quality prediction score. This is used as a predicted score for the overall display quality.
[0023] The stereoscopic quality prediction module based on disparity information is as follows:
[0024] disparity map The image encoder is fed into the pre-trained CLIP model to obtain the disparity visual features corresponding to the disparity information. ;
[0025] Based on the obtained parallax visual features and semantic features of stereoscopic quality ,calculate and Cosine similarity between the two modes is used to obtain the stereo quality mapping value of the disparity map after cross-modal alignment. ;
[0026] The Softmax function is used to map the stereoscopic quality values of the obtained disparity map. Normalization is performed to obtain the probability distribution of stereoscopic quality. ;
[0027] Probability distribution of stereoscopic quality Weighted summation is performed to achieve score continuity, thereby obtaining the predicted score for stereoscopic quality. .
[0028] The beneficial effects of the technical solution provided by this invention are:
[0029] 1. This invention designs a descriptive text prompt for the evaluation of stereoscopic display quality, and uses the semantic information of text related to the overall display quality and stereoscopic quality as explicit guidance to introduce the quality evaluation process of light field 3D display devices, which effectively improves the intelligence level of the quality evaluation method of light field 3D display devices;
[0030] 2. This invention constructs an overall display quality prediction module based on left and right viewpoint images. By modeling the matching relationship between the visual features of the left and right viewpoint images of the light field 3D display device and the semantic features of the overall display quality, it realizes cross-modal alignment of the visual and semantic information of the overall display quality, effectively enhancing the model's ability to perceive the overall display quality of the light field 3D display device.
[0031] 3. This invention designs a stereoscopic quality prediction module based on parallax information. By introducing parallax information containing stereoscopic quality and using stereoscopic quality semantic features as display semantic guidance, the quality evaluation performance of light field 3D display devices is further improved by mining the stereoscopic quality in light field 3D display.
[0032] 4. This invention can obtain more accurate display device quality prediction results, and therefore can be used to guide the imaging quality optimization process of light field 3D display devices. Specifically, based on the display device quality prediction results, key performance parameters affecting display device quality can be adjusted in a targeted manner, including parallax range, crosstalk-related parameters, etc., thereby improving stereoscopic performance, reducing crosstalk distortion, and enhancing the overall imaging quality of light field 3D display devices. Attached Figure Description
[0033] Figure 1 This is a flowchart of a method for intelligent quality evaluation of a light field 3D display device. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below.
[0035] To address the problems in the background art, this invention proposes an intelligent quality evaluation method for light field 3D display devices. (See [link to relevant documentation]). Figure 1 This method extracts visual features from the left and right viewpoint images of a light field 3D display device as input and aligns them cross-modally with the overall display quality description text. Simultaneously, it extracts visual features based on the disparity information corresponding to the left and right viewpoint images and aligns them cross-modally with the stereoscopic quality description text. Furthermore, it utilizes the shared visual feature space of the CLIP model to jointly predict the overall display quality and stereoscopic quality, thereby achieving intelligent quality evaluation of the light field 3D display device.
[0036] The CLIP model is trained on large-scale image-text pairs. By simultaneously bringing together matching image and text features and separating mismatched image and text features, it maps them into a unified cross-modal embedding space. Based on the similarity between image features and text features in this embedding space, the semantic consistency between images and text can be measured.
[0037] I. Design of 3D display quality description text prompts
[0038] Considering that stereoscopic quality is one of the important parameters affecting the quality of light field 3D display devices, this invention designs overall display quality description text prompts and stereoscopic quality description text prompts, which are used to provide explicit semantic guidance for the quality perception process of subsequent left and right viewpoint images and their corresponding disparity information. Specifically, the constructed overall display quality description text prompt template is "a stereoscopic image pair with {c} quality.", and the stereoscopic quality description text prompt template is "a disparity map with {s} stereoscopic perception.", where {c} represents the overall display quality level of the displayed content, {s} represents the stereoscopic quality level of the displayed content, and the candidate text set corresponding to the overall display quality description text template is denoted as . The candidate text set corresponding to the text template describing the three-dimensional quality is denoted as . The overall display quality description text prompt template is a pair of stereo images with a quality level of {c}; the stereoscopic quality description text prompt template is a disparity map reflecting a stereoscopic quality level of {s}, where {c} and {s} are both variables.
[0039] and The candidate text sets all follow the standard five-level evaluation criteria, with the level words set as "bad", "poor", "fair", "good", and "perfect". These mean that the overall display quality and stereoscopic quality level of the displayed content are "poor", "fair", "good", and "perfect", respectively. That is, their actual quality score labels and stereoscopic quality score labels are "1", "2", "3", "4", and "5", respectively, corresponding to five levels of overall display quality and stereoscopic quality from low to high.
[0040] Subsequently, candidate text sets describing different quality levels will be... and The text encoder in the pre-trained CLIP model is input separately to obtain semantic features that imply an overall display quality-aware prior. and semantic features that imply prior knowledge of stereoscopic quality Overall, it displays semantic features of quality. and semantic features of stereoscopic quality The learning process can be expressed by the following formula:
[0041]
[0042]
[0043] in, The text encoder representing the CLIP model of the overall display quality description text prompt template. A text encoder representing a CLIP model for text prompt templates that describe stereoscopic quality. , , Indicates the number of candidate texts. In this embodiment of the invention, the dimension representing the semantic features of candidate text is... Set to 5. Set to 512. Extracted overall display quality semantic features. The semantic guidance used for overall display quality judgment will be employed for subsequent cross-modal alignment with stereo image features, guiding the subsequent overall display quality prediction process based on left and right viewpoint images. Similarly, the extracted stereo quality semantic features... The semantic guidance used for stereoscopic quality judgment will be used for cross-modal alignment of disparity information features corresponding to the left and right viewpoint images, guiding the subsequent stereoscopic quality prediction process based on disparity information.
[0044] II. Constructing an overall display quality prediction module based on left and right viewpoint images
[0045] Considering that the final evaluation of light field 3D display devices is based on the subjective perception of the human eye, and the human eye's perception of light field information is essentially based on a binocular parallax fusion mechanism under a specific spatial pose, in order to simulate the real visual experience of the human eye when viewing a light field 3D display device, this embodiment of the invention uses a binocular camera to capture the content of the light field 3D display device, and obtains a binocular stereo image containing left and right viewpoint images as the input of the model, so as to fully simulate the binocular vision observation process.
[0046] Specifically, the left viewpoint image in a binocular stereo image is represented as follows: The right viewpoint image is represented as ,Will and The image encoders in the pre-trained CLIP model are input separately to obtain the visual features of the left viewpoint image corresponding to the left and right viewpoint images. Visual features of the right viewpoint image Based on the obtained visual features of the left viewpoint image Visual features of the right viewpoint image Overall display quality semantic features The overall calculation reveals semantic features of quality. Visual features of the left viewpoint image Cosine similarity between them, and overall display quality semantic features Visual features of the right viewpoint image Cosine similarity between them.
[0047] This step involves modeling the matching relationship between the visual content of the left and right viewpoint images and the overall display quality description text cues to obtain the cross-modal aligned left viewpoint image quality mapping value. Image quality mapping value for right viewpoint The specific calculation formula is as follows:
[0048]
[0049]
[0050] in, This represents the vector dot product operation. This represents the L2 norm.
[0051] Subsequently, the Softmax function is used to map the quality values of the obtained left viewpoint image. Image quality mapping value for right viewpoint The normalization process is performed separately. This step aims to map the quality prediction distribution of the binocular stereo images to a probability space, thereby transforming it into a quality probability distribution that reflects the quality level of the left viewpoint image. Quality probability distribution of image quality level from the right viewpoint The specific calculation formula is as follows:
[0052]
[0053]
[0054] in, Indicates exponentiation. This indicates the calculation of the L1 norm.
[0055] To obtain the overall display quality prediction result for the binocular stereo image, the prediction quality probability distribution of the obtained left viewpoint image is calculated. And the probability distribution of prediction quality for right viewpoint images Perform weighted fusion to obtain the overall display quality probability distribution of the fused data. :
[0056] in, This represents the weight of the probability distribution of the predicted quality of the left viewpoint image during the weighted fusion process. In this embodiment of the invention, Set it to 0.5.
[0057] Furthermore, based on the discretized real overall display quality labels corresponding to different overall display quality levels, the probability distribution of fusion quality is analyzed. Weighted summation is performed to achieve score continuity, thereby obtaining the binocular quality prediction score. This score, used as a predicted score for overall display quality, is calculated using the following formula:
[0058] in, In this embodiment of the invention, the total number of quality grades is represented. Set to 5. In the probability distribution of fusion quality, the th The predicted probability corresponding to each quality level.
[0059] III. Constructing a stereoscopic quality prediction module based on parallax information
[0060] Considering that the disparity maps generated from the left and right viewpoint images of a light field 3D display device contain rich stereoscopic information, this invention constructs a stereoscopic quality prediction module based on disparity information. By introducing disparity information to predict stereoscopic quality, it helps improve the quality evaluation performance of light field 3D display devices.
[0061] Specifically, first, the left viewpoint image and right viewpoint image The input disparity map generation network obtains the corresponding disparity map. The process can be represented as follows:
[0062] in, This represents a disparity map generation network based on PSMNet (known to those skilled in the art).
[0063] disparity map The image encoder is fed into the pre-trained CLIP model to obtain the disparity visual features corresponding to the disparity information. .
[0064] In this step, the image encoder of the CLIP model and the image encoder of the CLIP model in the stereoscopic quality prediction module based on disparity information adopt a weight sharing mechanism (known to those skilled in the art) to achieve joint prediction of overall display quality and stereoscopic quality.
[0065] Based on the obtained parallax visual features and semantic features of stereoscopic quality ,calculate and Cosine similarity between them. This step obtains the stereo quality mapping value of the disparity map after cross-modal alignment by modeling the matching relationship between the visual content of the disparity image and the stereo quality description text cues. The specific calculation formula is as follows:
[0066]
[0067] The Softmax function is used to map the stereoscopic quality values of the obtained disparity map. The normalization process aims to map the disparity map-guided stereoscopic quality prediction distribution to a probability space, thereby obtaining the stereoscopic quality probability distribution. .
[0068] Subsequently, based on the discretized true stereo quality labels corresponding to different stereo quality levels, the probability distribution of stereo quality was analyzed. Weighted summation is performed to achieve score continuity, thereby obtaining the predicted score for stereoscopic quality. This is used as a predicted score for the stereoscopic quality, and the specific calculation formula is as follows:
[0069] in, In this embodiment of the invention, the total number of quality grades is represented. Set to 5. The probability distribution of stereoscopic quality represents the th The predicted probability corresponding to each quality level.
[0070] IV. Intelligent Evaluation Network for the Quality of Training Light Field 3D Display Devices
[0071] The intelligent quality evaluation network for light field 3D display devices proposed in this invention is trained in an end-to-end manner. To improve the accuracy of the model's evaluation of overall display quality and stereoscopic quality, a fidelity loss function based on ranking and an L1 loss function based on numerical regression are introduced to jointly constrain the network during training. During training, a synchronous training strategy is adopted for the overall display quality prediction module based on left and right viewpoint images and the stereoscopic quality prediction module based on disparity information.
[0072] Specifically, a fidelity loss function is used to constrain the overall display quality prediction module. First, the Thurstone model (known to those skilled in the art) is used to calculate the predicted score for the overall display quality. Mapped to pairwise preference probabilities of overall display quality Paired preference probability This indicates that in the predicted quality output of the overall display quality prediction module based on left and right viewpoint images, the image obtained from the first shot... A stereoscopic image sample corresponding to a light field 3D display device The overall prediction quality is better than that of the first image obtained. A stereoscopic image sample corresponding to a light field 3D display device The predicted overall display quality probability is used. Pairwise preference probabilities are included in the calculation of the fidelity loss function. The fidelity loss is used to measure the distance between the predicted preference probability and the true preference probability of the overall display quality, constraining the model to learn the relative overall display quality relationship between samples.
[0073] The constructed fidelity loss function is expressed as follows:
[0074]
[0075] in, Indicates the number of photos taken. A sample of binocular stereoscopic images corresponding to a light field 3D display device. Indicates the number of photos taken. A sample of binocular stereoscopic images corresponding to a light field 3D display device. and All of these are samples to be compared. Indicates the set of stereo image sample indices The index variable pairs retrieved from the data. This indicates the introduced fidelity loss. This represents the binocular stereo image sample corresponding to the light field 3D display device obtained from the capture. Superior to the stereoscopic image samples corresponding to the light field 3D display device obtained by photography The overall probability of the true preference for display quality. This represents the binocular stereo image sample corresponding to the light field 3D display device obtained from the capture. Superior to the stereoscopic image samples corresponding to the light field 3D display device obtained by photography The overall display quality prediction preference probability.
[0076] To further improve the model's predictive performance for the stereoscopic quality of light field 3D display devices, in addition to the fidelity loss function, an L1 loss function based on numerical regression was added to constrain the stereoscopic quality prediction module based on disparity information. The expression for the L1 loss function is:
[0077] in, Indicates the number of photos taken. The quality prediction score corresponding to the binocular stereo image sample of each light field 3D display device. Indicates the number of photos taken. The true quality score corresponding to the binocular stereo image sample of each light field 3D display device. This represents the introduced L1 loss. Indicates the number of photos taken. The total number of binocular stereo image samples corresponding to each light field 3D display device.
[0078] In the stereoscopic quality prediction module based on disparity information, the fidelity loss function and the L1 loss function are weighted and fused to construct a joint loss function for stereoscopic quality prediction. To optimize the predicted values of stereoscopic quality:
[0079]
[0080] in, This indicates the weight of fidelity loss in the joint loss function for predicting stereoscopic quality. This represents the weight of the L1 loss in the joint loss function for stereoscopic quality prediction. In this embodiment of the invention, Set to 1, Set it to 0.2.
[0081] V. Application of the intelligent evaluation network for the quality of trained light field 3D display devices
[0082] Based on the network model trained using the fidelity loss function and the L1 loss function, the left and right viewpoint images corresponding to the light field 3D display device under test are captured and input into the network model. The model outputs the corresponding quality evaluation prediction score, which is used as the result of the quality evaluation of the light field 3D display device.
[0083] Furthermore, the obtained quality assessment prediction scores are used to guide the quality optimization of the light field 3D display device, thereby improving its quality. For example, performance parameters of the light field 3D display device (such as parallax range, ray reconstruction algorithm parameters, etc.) can be adjusted, and images of the light field 3D display device before and after optimization can be captured to obtain corresponding left and right viewpoint images. These images are then input into the network model, and the quality assessment prediction scores under different parameter settings are compared. The performance parameter configuration that results in a higher quality assessment prediction score is selected, thus guiding the optimization of the light field 3D display device.
[0084] In this embodiment of the invention, the 3D display device can be a light field 3D display device or an automatic stereoscopic display device, specifically a light field 3D display screen, a light field 3D display terminal (such as a mobile terminal), and an automatic stereoscopic display screen, but is not limited to the above types.
[0085] Unless otherwise specified, the model numbers of the various devices in this embodiment of the invention are not limited, and any device that can perform the above functions is acceptable.
[0086] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0087] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for intelligently evaluating the quality of a light field 3D display device, characterized in that, The method includes: Construct overall display quality description text prompts and stereoscopic quality description text prompts, which are used to provide explicit semantic guidance for the quality perception process of left and right viewpoint images and their corresponding parallax information, respectively; The constructed overall display quality description text prompts and stereoscopic quality description text prompts are respectively input into the text encoder in the pre-trained CLIP model to obtain semantic features containing the overall display quality perception prior and the stereoscopic quality perception prior, respectively. The content of the light field 3D display device is captured by a binocular camera, and the binocular stereo images containing left and right viewpoint images are used as input to the model to construct an overall display quality prediction module based on the left and right viewpoint images. A stereoscopic quality prediction module based on parallax information is constructed to predict stereoscopic quality by introducing parallax information, thereby helping to improve the quality evaluation performance of light field 3D display devices. An end-to-end approach is used to train a quality intelligent evaluation network for light field 3D display devices, and the trained network guides the optimization of light field 3D display devices.
2. The intelligent quality evaluation method for a light field 3D display device according to claim 1, characterized in that, The overall display quality description text prompts and the stereoscopic quality description text prompts are respectively: The template for the overall display quality description text is "a stereoscopic image pair with {c}quality."; The template for the stereoscopic quality description text prompt is "a disparity map with {s} stereoscopic perception."; Where {c} represents the overall display quality level of the displayed content, {s} represents the stereoscopic quality level of the displayed content, and the candidate text set corresponding to the overall display quality description text template is denoted as . The candidate text set corresponding to the text template describing the three-dimensional quality is denoted as . .
3. The intelligent quality evaluation method for a light field 3D display device according to claim 1, characterized in that, The semantic features for obtaining the prior knowledge of overall display quality perception and prior knowledge of stereoscopic quality are as follows: Candidate text sets describing different quality levels will be used. and The text encoder in the pre-trained CLIP model is input separately to obtain semantic features that imply an overall display quality-aware prior. and semantic features that imply prior knowledge of stereoscopic quality .
4. The intelligent quality evaluation method for a light field 3D display device according to claim 1, characterized in that, The overall display quality prediction module based on left and right viewpoint images is as follows: Left and right viewpoint images and The image encoders in the pre-trained CLIP model are input separately to obtain the visual features of the left viewpoint image corresponding to the left and right viewpoint images. Visual features of the right viewpoint image At the same time, it will describe candidate text sets of different quality levels. Input the text encoder from the pre-trained CLIP model to obtain overall display-quality semantic features. ; Based on the visual features of the obtained left viewpoint image Visual features of the right viewpoint image and overall display quality semantic features Calculate separately and and and Cosine similarity between the two views is used to obtain the quality mapping value of the left viewpoint image after cross-modal alignment. Image quality mapping value for right viewpoint ; The Softmax function is used to map the quality values of the obtained left viewpoint image. Image quality mapping value for right viewpoint Normalization is performed to obtain the probability distribution of the predicted image quality for the left viewpoint. And the probability distribution of prediction quality for right viewpoint images ; Predicting quality probability distribution of the obtained left viewpoint image And the probability distribution of prediction quality for right viewpoint images Perform weighted fusion to obtain the overall display quality probability distribution of the fused data. : The quality probability distribution after fusion Weighted summation is performed to achieve score continuity and obtain the binocular quality prediction score. This is used as a predicted score for the overall display quality.
5. The intelligent quality evaluation method for a light field 3D display device according to claim 1, characterized in that, The stereoscopic quality prediction module based on disparity information is as follows: disparity map The image encoder is fed into the pre-trained CLIP model to obtain the disparity visual features corresponding to the disparity information. ; Based on the obtained parallax visual features and semantic features of stereoscopic quality ,calculate and Cosine similarity between the two modes is used to obtain the stereo quality mapping value of the disparity map after cross-modal alignment. ; The Softmax function is used to map the stereoscopic quality values of the obtained disparity map. Normalization is performed to obtain the probability distribution of stereoscopic quality. ; Probability distribution of stereoscopic quality Weighted summation is performed to achieve score continuity, thereby obtaining the predicted score for stereoscopic quality. .