Multi-modal three-dimensional image quality evaluation method based on consistency-complementarity characteristics

This multimodal stereo image quality assessment method, based on consistent-complementary features, solves the problem of inconsistent subjective perception in stereo image evaluation, improves stability and generalization ability, and is applicable to quality monitoring and optimization in scenarios such as 3D video coding, virtual reality, and augmented reality.

CN122048847APending Publication Date: 2026-05-15TIANJIN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2026-01-28
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing methods for assessing the quality of stereo images fail to adequately reflect human subjective perception, and the cross-modal feature distributions are difficult to align, affecting the stability and generalization performance of the assessment.

Method used

A multimodal stereo image quality assessment method based on consistency-complementarity features is adopted. Features are extracted by a parallax-perceptual image encoder and a multi-scale attention EEG encoder. Combined with a complementarity learning module and a consistency learning module, the complementarity and consistency of visual and EEG modalities are explicitly characterized. EEG signals are used to guide feature learning during the training phase, and quality prediction is performed solely based on stereo images during the inference phase.

Benefits of technology

It improves the consistency between objective evaluation results and subjective perception, enhances the stability and generalization ability of the evaluation, reduces additional acquisition costs, and is suitable for various stereoscopic image application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122048847A_ABST
    Figure CN122048847A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal three-dimensional image quality evaluation method based on consistency-complementarity characteristics, and belongs to the technical field of three-dimensional image quality evaluation and brain-computer intelligence. The method comprises the following steps: S1, acquiring a plurality of groups of distorted stereo images and electroencephalogram signals corresponding to the distorted stereo images, preprocessing the distorted stereo images, and endowing a quality grade label to a sample according to a subjective experiment result; s2, inputting the preprocessed three-dimensional image and the electroencephalogram signal into a parallax perception image encoder and a multi-scale attention electroencephalogram encoder to obtain binocular image features and electroencephalogram spatial-temporal features; s3, inputting the binocular image features and the electroencephalogram spatial-temporal features into a complementary learning module, and extracting image complementary features, electroencephalogram complementary features and cross-modal consistency features; and S4, inputting the consistency characteristics into a consistency learning module for cross-modal alignment to obtain an optimized consistency characteristic representation, fusing the optimized consistency characteristic representation with the complementary characteristics, inputting the fused consistency characteristic representation into a classification module, and outputting a stereo image quality grade.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of image processing and brain-computer intelligence, specifically to a multimodal stereo image quality evaluation method based on consistent-complementary features. Background Technology

[0002] With the development of technologies such as 3D display, virtual reality, and augmented reality, stereoscopic images and videos are widely used in film and television entertainment, telecommunications, medical imaging, and industrial inspection. The subjective visual quality of stereoscopic images directly affects the user's viewing comfort and immersion; therefore, how to objectively and accurately evaluate the quality of stereoscopic images has become an important issue.

[0003] Existing stereo image quality assessment methods can be broadly categorized into three types: full-reference, half-reference, and no-reference. Full-reference and half-reference methods typically require a clear reference image or partial reference information as prior knowledge, making them difficult to apply in real-world transmission links where obtaining a reference image is challenging. No-reference stereo image quality assessment methods, relying solely on the distorted image itself, better meet practical application needs and have become a research hotspot.

[0004] In the assessment of stereo image quality without reference, early methods were mostly based on manually designed features. They relied on statistical analysis of distortion-sensitive indicators such as brightness, contrast, texture, disparity, and binocular adversarial behavior, and then used traditional regression models to predict quality scores. These methods are highly sensitive to distortion types and scene variations, and their generalization ability is limited. In recent years, deep learning methods such as convolutional neural networks and attention mechanisms have been introduced into stereo image quality assessment. By modeling the left and right views separately or through simple fusion, they can learn complex distortion features to some extent. However, most of these methods still rely solely on the visual information of the image itself and fail to fully reflect the subjective perception process of humans.

[0005] On the other hand, electroencephalogram (EEG) signals can reflect the neural response of the human visual cortex to external stimuli at millisecond-level temporal resolution. Existing research has attempted to utilize EEG signals to assist in the quality assessment of two-dimensional images or videos, aiming to improve the consistency between subjective and objective assessments. However, current work typically treats EEG features as supplementary inputs or simple statistical descriptions, lacking deep semantic alignment with image features. It fails to systematically model the relationship between visual and neural modalities, and even more so, it neglects the essential characteristics of their consistency and complementarity in feature space.

[0006] Specifically, existing technologies suffer from at least the following shortcomings: most no-reference stereo image quality assessment methods focus only on the visual features of stereo images, failing to incorporate EEG information reflecting human subjective perception, making it difficult to guarantee the consistency between objective evaluation results and subjective perception; some quality assessment methods incorporating EEG primarily target 2D images or videos, directly stitching or simply fusing image features with EEG features, without explicitly distinguishing between modality-specific complementary information and cross-modal shared consistency information, leading to difficulties in aligning cross-modal feature distributions; existing cross-modal learning methods generally lack fine-grained constraints and generative modeling of the consistency relationship between visual and neural modalities, failing to effectively reduce the distribution differences between the two modalities in the feature space, thus affecting the stability and generalization performance of stereo image quality assessment. Therefore, this invention proposes a multimodal no-reference stereo image quality assessment method that can simultaneously mine the consistency and complementarity between stereo image visual features and EEG signals, using EEG to guide network learning during the training phase and relying solely on stereo image input to complete quality prediction during the inference phase, thereby improving the consistency and robustness between objective evaluation results and human subjective perception. Summary of the Invention

[0007] The purpose of this invention is to propose a multimodal stereo image quality assessment method based on consistent-complementary features to solve the problem that existing technologies rely solely on image visual information and cannot fully reflect human subjective perception when assessing stereo image quality.

[0008] To achieve the above objectives, the present invention adopts the following technical solution: A multimodal stereo image quality assessment method based on consistent-complementary features includes the following steps: S1. Acquire multiple sets of distorted stereo images and their corresponding EEG signals, preprocess the obtained stereo images and EEG signals, and assign quality grade labels to the samples based on the subjective experimental results. S2. Input the preprocessed stereo image into the parallax perception image encoder to obtain binocular image features; input the preprocessed EEG signal into the multi-scale attention EEG encoder to obtain EEG spatiotemporal features. S3. Input the binocular image features and EEG spatiotemporal features into the complementary learning module to extract image complementary features, EEG complementary features, and cross-modal consistency features. S4. Input the consistency features into the consistency learning module for cross-modal alignment to obtain the optimized consistency feature representation. Then, fuse it with the complementary features and input it into the classification module to output the stereo image quality level.

[0009] Preferably, the preprocessing in S1 specifically includes: Size normalization and parallax correction are performed on the stereoscopic image; Bandpass filtering was performed on multichannel EEG signals, and independent component analysis was used to remove eye movement and electromyography artifacts. The artifact-free EEG signals were segmented according to fixed time windows, and each segment of EEG signal was aligned with the corresponding stereoscopic image and quality grade label.

[0010] Preferably, the parallax-aware image encoder in S2 performs the following operations: A dual-path convolutional feature extraction structure is used to perform multi-scale convolution on the left and right views to obtain initial monocular features; The disparity compensation module is used to perform disparity correlation calculations on the features of the left and right views to obtain a disparity attention map, and the monocular features are weighted to form disparity compensation features. The attention fusion module stitches together the disparity compensation features of the left and right views with the initial monocular features and performs attention weighting to output binocular image features.

[0011] Preferably, the multi-scale attention EEG encoder described in S2 performs the following operations: Multi-scale EEG features are extracted in both the temporal and channel dimensions using multi-scale convolutional branches. Weights are assigned to time-dimensional features through the temporal attention module, and weights are assigned to channel-dimensional features through the spatial attention module. By fusing features weighted by temporal attention and spatial attention, spatiotemporal features of EEG are obtained.

[0012] Preferably, the complementary learning module in S3 performs the following operations: Private channels are set up in the image modality and EEG modality respectively to extract image complementary features and EEG complementary features; and a shared channel is set up to extract cross-modal consistency features. By using a loss function that includes complementary constraints and consistency comparison terms, the correlation between complementary features and consistent features is constrained, thereby achieving the decoupling and alignment of modality-specific complementary information and cross-modal consistent information.

[0013] Preferably, the classification module in S4 performs the following operations: After fusing complementary features and consistent features, the data is input into the class-level contrastive learning branch and the task classification branch in the classification module. The class-level contrastive learning branch constructs positive and negative sample pairs based on quality level labels to increase the margin between features of different quality levels. The task classification branch uses cross-entropy loss to supervise the prediction results and the true quality level labels. Output the probability distribution of each quality level, and use the quality level corresponding to the highest probability as the objective quality evaluation result of the stereo image.

[0014] The present invention further protects a computer device, characterized in that the computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the instruction, program, code set or instruction set is loaded and executed by the processor to implement the method as described in any one of claims 1-6.

[0015] The present invention further protects a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the instruction, program, code set, or instruction set is loaded and executed by a processor to implement the above-described multimodal stereo image quality evaluation method based on consistent-complementary features.

[0016] Compared with existing technologies, this invention provides a multimodal stereo image quality assessment method based on consistent-complementary features, which has the following beneficial effects: (1) This invention introduces EEG signals into the multimodal no-reference stereo image quality evaluation framework and uses a consistency-complementary joint learning mechanism to jointly model the visual features of stereo images and the spatiotemporal features of EEG. It not only utilizes the distortion information of the stereo images themselves, but also explicitly depicts the neural response process of the human visual cortex to stereo stimuli, making the objective evaluation results more consistent with subjective perception. (2) In this invention, private channels and shared channels are set up in the image modality and the EEG modality respectively. Combined with complementary constraint loss, information theory contrastive learning and bidirectional cross-modal generative adversarial network, the modality-specific complementary information and cross-modal shared consistency information are explicitly decoupled and distributed and aligned. This effectively suppresses redundant correlation and feature collapse, reduces the feature distribution difference between the visual modality and the neural modality, and thus significantly improves the stability and generalization ability of stereo image quality evaluation. (3) In the training stage, the present invention uses EEG signals to guide and constrain the feature space. In the reasoning stage, it can complete the reference-free quality prediction by relying only on the stereo image to be evaluated. It does not require real-time collection of EEG signals, and takes into account both high subjective correlation and cost and deployment convenience in practical applications. It has good engineering application value. Attached Figure Description

[0017] Figure 1 is a flowchart of the multimodal stereo image quality assessment method based on consistent-complementary features proposed in this invention; Figure 2 is a structural diagram of the image encoder mentioned in Embodiment 1 of the present invention; Figure 3 is a diagram of the bidirectional cross-modal generative adversarial network structure mentioned in Embodiment 1 of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below.

[0019] This invention proposes a multimodal stereo image quality assessment method based on consistency-complementarity features, comprising: acquiring distorted stereo images and corresponding EEG signals; performing size normalization and disparity correction on the stereo images; filtering, artifact removal, and segmentation on the EEG signals; and assigning quality level labels to samples based on subjective experimental results; extracting left and right view features using a dual-channel disparity-aware image encoder, and obtaining binocular fused image features through disparity compensation and attention fusion; simultaneously extracting spatiotemporal features of EEG using a multi-scale attention EEG encoder; inputting the image features and EEG features into a complementarity learning module, and obtaining complementary features that retain modality-specific information and initial cross-modal consistency features through decoupling of a dual-branch autoencoder and private / shared channels; inputting the initial consistency features into a consistency learning module, and aligning the visual and neural domain feature distributions through a cross-modal consistency contrast loss based on mutual information and entropy constraints and a bidirectional cross-modal generative adversarial network to generate semantically consistent cross-modal representations; and inputting the fused complementary features and consistency features into a classification module to output the stereo image quality level. The inference stage relies only on the stereo image input and does not require real-time acquisition of EEG signals.

[0020] The multimodal stereo image quality evaluation method based on consistent-complementary features proposed in this invention will be described below with reference to the relevant accompanying drawings and specific examples.

[0021] Example 1: Please refer to Figure 1. This invention provides a referenceless stereo image quality assessment method based on EEG-guided consistency-complementary learning, which includes the following steps: S101: Data Acquisition and Preprocessing Multiple sets of distorted stereoscopic images and their corresponding EEG signals were acquired. In a controlled experimental environment, the stereoscopic images were presented while multi-channel EEG data were simultaneously recorded during subject viewing. Subjective quality scores were also collected and discretized into several quality level labels. Size normalization and parallax correction were performed on the left and right views, with further normalization processing as needed. Bandpass filtering was applied to the EEG signals, and independent component analysis was used to remove artifacts such as eye movement and electromyography. The signals were then segmented according to fixed time windows, and each EEG segment was aligned with its corresponding stereoscopic image sample and quality level label to construct training samples.

[0022] S102: Feature Extraction As shown in Figure 2, in the image modality, the preprocessed left and right views are input into the disparity-aware image encoder, respectively. A dual-path convolutional structure is used to extract multi-scale initial features of the left and right views. After disparity correlation calculation, a disparity attention map is obtained. After disparity compensation of the monocular features, the features of the left and right views are weighted and fused through the attention fusion module to obtain the binocular image feature representation. In the EEG modality, multi-channel EEG segments aligned with the stereo image are input into the multi-scale attention EEG encoder. Local features are extracted through multi-scale convolutional branches in the time and channel dimensions. The temporal attention module highlights key time slices, and the spatial attention module highlights key brain regions. The weighted results are then fused to obtain the spatiotemporal feature representation of the EEG.

[0023] S103: Progressive Stereo Interactive Restoration Result Binocular image features and EEG spatiotemporal features are jointly input into the consistency-complementarity learning module. Private and shared channels are set in the image and EEG modalities respectively. Modality-specific image and EEG complementary features are extracted through the private image and EEG channels, while cross-modal consistency features are extracted through the shared channel. During training, a loss function containing complementarity constraints and consistency comparison terms is constructed to constrain the correlations between complementary features and between private and consistency features, thereby decoupling modal complementarity information from cross-modal consistency information. Simultaneously, as shown in Figure 3, a bidirectional cross-modal generative adversarial network is introduced to learn the mappings from image consistency features to EEG consistency features and from EEG consistency features to image consistency features, respectively. The feature distributions of the two modalities in the consistency subspace are aligned using adversarial loss and reconstruction loss to obtain semantically aligned consistency feature representations.

[0024] S104: Quality Prediction and Inference Image complementary features, EEG complementary features, and consistency features are fused and input into the classification module. The classification module includes a class-level contrastive learning branch and a task classification branch. Class-level contrastive loss brings together features of samples with the same quality level and separates features of samples with different quality levels. Cross-entropy loss constrains the predicted quality level and the true quality level label, thus obtaining the trained quality prediction network. During the training phase, the weighted sum of complementary constraint loss, consistency contrastive and adversarial loss, and classification loss is used as the overall loss for end-to-end joint optimization of each encoder, consistency-complementary learning module, and classification module. During the inference phase, only the stereo image to be evaluated needs to be input. After passing through the disparity-aware image encoder, consistency-complementary learning module, and classification module, the corresponding quality level can be output, achieving referenceless stereo image quality evaluation.

[0025] S105: Technical Applications This invention's method can be deployed on server-side or terminal devices for quality monitoring and optimization in various stereoscopic image and video service scenarios. In 3D video encoding and transmission scenarios, it can interface with encoders or transmission links to perform real-time or offline quality evaluation of received stereoscopic image / video frames, providing a basis for bitrate allocation, adaptive adjustment of encoding parameters, and error retransmission strategies. In immersive display scenarios such as virtual reality and augmented reality, it can be combined with head-mounted display devices to evaluate the quality of rendered stereoscopic content, guiding adaptive control of rendering quality, content selection, and user experience assurance. In professional applications sensitive to stereoscopic details, such as medical stereoscopic imaging and industrial inspection, it can objectively assess the quality of acquired or reconstructed stereoscopic images, providing a reliable reference for subsequent diagnosis, analysis, or judgment. Because this invention introduces EEG for guidance during the training phase, and relies solely on stereoscopic images for quality prediction during the actual deployment phase, it can improve the subjective perception consistency and service quality of various stereoscopic image application systems without increasing additional acquisition costs.

[0026] Example 2: Please see Figures 1-3 Based on Example 1, but with some differences, the scheme in Example 1 will be further described below with specific calculation formulas and example data, as detailed in the following description: Step 1: Dataset and Label Settings This embodiment uses the BVM-SIQA2025 stereoscopic image quality assessment dataset as an example to illustrate the method of the present invention. Each subject in this dataset contains 580 distorted stereoscopic images (including left and right views) and synchronously acquired multi-channel EEG data. The EEG data has a shape of 580×62×3000, corresponding to "trial number × number of channels × number of sampling points". The stereoscopic image data has a shape of 580×3×650×555, corresponding to "trial number × number of channels × width × height". Each sample also has 5 quality labels (excellent, good, fair, poor, bad).

[0027] During training, the dataset was divided into training and testing sets in an 8:2 ratio, and a five-fold repeated experiment was used to improve the stability of the results. Each stereo image was assigned a five-level quality label, and this label was synchronously assigned to an EEG segment aligned with its time, for supervising the joint learning of image modalities and EEG modalities.

[0028] Step 2: Feature Extraction (1) Feature extraction of stereo images To extract quality-related structural information from stereo images and mitigate parallax mismatch, such as Figure 2As shown, this embodiment employs a dual-branch disparity-aware image encoder (PaIE). The left and right views are processed through two parallel hybrid convolutional blocks to obtain initial monocular features containing both global and local structures. Subsequently, a disparity compensation module is introduced: first using... Convolutional and multi-scale dilated convolutions are used to obtain multi-receptive field features. Then, cross-correlation of the left and right views is performed, followed by softmax to obtain a bidirectional disparity attention map. This attention map is multiplied by the corresponding monocular features to generate disparity compensation features. Finally, the disparity compensation features are concatenated with the original monocular features, and then... Convolutional and Block Convolutional Attention (CBAM) modules adaptively reweight left and right features and fuse them in either the channel dimension or the feature dimension to obtain the final stereo image feature representation. F img .

[0029] (2) Extraction of EEG features To characterize the spatiotemporal dynamics of EEG related to visual perception, this embodiment employs a multi-scale attention EEG encoder. Multi-channel EEG segments aligned with stereoscopic images are first processed through a multi-scale convolutional structure to extract spatial and temporal features in the channel and temporal dimensions, respectively. Subsequently, these two types of features are fed into a spatiotemporal reconstruction module: the temporal branch generates temporal attention based on learnable weights, weighting different time slices; the spatial branch generates spatial attention based on channel weights, reconstructing different brain regions. Finally, the features weighted by temporal and spatial attention are fused and processed... Convolution is used for compression to obtain spatiotemporal feature representations of EEG. F eeg .

[0030] Step 3: Design of Consistency-Complementary Joint Loss Function In obtaining binocular image features F img Spatial and temporal characteristics of EEG F eeg Subsequently, this embodiment employs a dual-channel decoupled learning strategy to jointly model cross-modal consistency and complementary information. Specifically, for each modality, a private channel and a shared channel are set up: the image private channel and the EEG private channel respectively receive the features of their respective modalities and extract complementary features that exist only in that modality. To preserve modality-specific information, the image sharing channel and the EEG sharing channel share parameters and perform unified mapping on the features of the two modalities to obtain image consistency features and EEG consistency features. It is used to characterize cross-modal consistency between vision and the nervous system.

[0031] To avoid interference between complementary and consistent features, this embodiment introduces a complementarity constraint loss to constrain the correlation between complementary features and between complementary and consistent features, making them approximately orthogonal in the feature space. The complementarity constraint loss can be simply written as the sum of squares of multinomial inner products: (1) By minimizing This allows image complementary features, EEG complementary features, and consistency features to be decoupled from each other in the subspace, thereby enabling clearer extraction of modality-specific complementary information and cross-modal shared consistency information.

[0032] Building upon this, to further enhance cross-modal consistency, this embodiment designs a consistency learning module. First, the initial consistency features are... As two sources of information, an information-theoretic consistency contrastive learning loss is constructed. By maximizing the mutual information between the two and introducing an entropy regularization term, semantic alignment is strengthened while preventing feature collapse and maintaining feature diversity. Its objective can be summarized as follows: (2) in, Indicates mutual information, Represents entropy, To balance the parameters. This objective allows for the acquisition of semantically consistent and informative consistent feature representations within a shared latent space.

[0033] Furthermore, to reduce the distribution differences of different modalities in the consistency subspace, this embodiment introduces a bidirectional cross-modal generative adversarial network (BC-GAN), as shown in Figure 3. BC-GAN employs a symmetrical dual-path structure. One path generates corresponding EEG consistency features based on image consistency features, while the other path generates corresponding image consistency features based on EEG consistency features. The discriminator judges the authenticity of the real features and the generated features, as well as their consistency with the conditional features. Through joint optimization of adversarial loss and reconstruction loss, bidirectional translation and reconstruction from image to EEG and from EEG to image are achieved, effectively aligning the distribution of the two modalities in the consistency subspace, improving the robustness and separability of consistency features, and providing high-quality cross-modal representations for subsequent quality classification.

[0034] Step 4: Loss Function Module During the training phase, this embodiment performs end-to-end joint optimization of the image encoder, EEG encoder, consistency-complementary learning module, and classification module. The overall loss function consists of three parts: complementary learning loss, consistency learning loss, and classification loss, defined as follows: (3) in, and This is an adjustable hyperparameter used to balance the contributions of the three parts of the loss.

[0035] (1) Complementary learning loss

[0036] The complementary learning loss is used to extract modality-specific complementary features and suppress redundant cross-modal information. It consists of the reconstruction loss of each modal autoencoder and the complementary constraint loss. (4) in, and These represent the reconstruction losses for the image modality and the EEG modality, respectively. Let the... Each sample in modality The input below is Features are obtained through the encoder via the corresponding decoder After reconstruction The reconstruction loss is: (5) Complementary constraint loss L comp The complementary features and the consistency features can be distinguished by formula (1). By minimizing formula (1), the complementary features of the image, the complementary features of the EEG and the consistency features are made as orthogonal as possible in the subspace, thereby preserving the modality-specific complementary information and providing a clean basic representation for subsequent consistency modeling.

[0037] (2) Consistency learning loss L ConsLM The consistency learning loss consists of information-theoretic consistency contrastive loss and bidirectional cross-modal generative adversarial loss: (6) Information-theoretic consistency contrast loss Initial consistency characteristics and Treating them as random variables with two modalities, semantic alignment and diversity preservation are achieved by maximizing mutual information and introducing entropy regularization. The consistency comparison objective can be written as: (7) (8) in, and They represent the marginal probability distributions, For balance coefficient, For feature dimensions; Represents initial consistency image features Consistent EEG characteristics with initial The joint probability distribution of ; Where is the number of samples, softmax(·) represents the probability normalization operation along the feature dimension. Bidirectional cross-modal adversarial loss To further reduce the distributional differences of consistent features across different modalities, a bidirectional cross-modal generative adversarial network (BC-GAN) is introduced, with its loss defined as: (9) Taking the image-to-EEG direction as an example, the generator loss consists of adversarial loss and reconstruction loss: (10) Among them, the EEG consistency features generated by the adversarial loss constraint In the discriminator The features should be as close to the real features as possible: (11) Reconstruction loss metric generates consistent features with real EEG data Differences between them: (12) The discriminator loss consists of two parts: real samples and generated samples. (13) Another direction The format is similar, simply by swapping the roles of image and EEG. Through bidirectional adversarial and reconstruction constraints, consistent feature translation and reconstruction from image to EEG and from EEG to image can be achieved, further improving the accuracy of cross-modal consistency modeling.

[0038] (3) Classification loss

[0039] Classification loss consists of class-level contrast loss and task classification loss composition: (14) Class-level contrast loss Record No. The final feature after fusion of samples is Its real label is The predicted label is The class-level contrastive loss constructs positive and negative sample pairs using the true and predicted labels, encouraging misclassified samples to move closer to the true class features and further away from the incorrect class features. Its form can be written as: (15) in, Indicates by and Features obtained by splicing Indicates different samples; The dot product similarity function; and These represent real labels. and prediction labels The corresponding feature representation set.

[0040] here Boundary values The contribution of this item to the loss is controlled by setting it to 0 when the prediction is correct and 1 otherwise.

[0041] Task classification loss To enhance the correlation between features and image quality and suppress task-irrelevant information, this embodiment sets up a classifier for both the visual and cognitive (EEG) modalities, employing cross-entropy loss in both. Taking the visual classifier as an example, let the first classifier be... The true quality label for each sample is The predicted probability is ,but: (16) Loss of EEG modalities With the same format, the total loss for task classification is: (17) Example 3: Based on Examples 1 and 2, but with some differences, the feasibility of the schemes in Examples 1 and 2 is verified below with specific experiments, as detailed in the following description: This experiment mainly uses the BVM-SIQA2025 dataset for five-level quality classification, and the evaluation metrics are classification accuracy (ACC) and Kappa coefficient.

[0042] On the BVM-SIQA2025 dataset, the model was independently trained on 8 subjects (S1–S8), and the performance results are shown in Table 1. The ACC of each subject ranged from 70.89% to 73.88%, with an average ACC of 72.74% and an average Kappa of 0.62, indicating that the method of this invention has good stability among different subjects.

[0043] To verify the performance of this invention, we selected several typical no-reference quality assessment methods based solely on stereo images as baselines and compared them under the same dataset and five-class classification settings. The results are shown in Table 2. The proposed method significantly outperforms all single-modal methods on both ACC and Kappa. Furthermore, we selected several existing multimodal learning methods (including cross-modal GAN ​​generation methods and collaborative learning methods) for comparison. The results under the same settings are shown in Table 3. The proposed method achieves the highest values ​​on both ACC and Kappa, indicating that the consistency-complementary learning framework can more fully utilize visual and cognitive information.

[0044] Table 1. Performance comparison among different participants

[0045] Table 2 Comparison with the single-mode method

[0046] Table 3 Comparison with multimodal methods

[0047] In summary, the experimental results of this embodiment show that, compared with various existing single-modal and multi-modal baseline methods, the multi-modal stereo image quality assessment method based on consistent-complementary features proposed in this invention has significant performance advantages on the same dataset, thus verifying the effectiveness and feasibility of the technical solutions described in Embodiments 1 and 2.

[0048] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A multimodal stereo image quality assessment method based on consistent-complementary features, characterized in that, Includes the following steps: S1. Acquire multiple sets of distorted stereo images and their corresponding EEG signals, preprocess the obtained stereo images and EEG signals, and assign quality grade labels to the samples based on the subjective experimental results. S2. Input the preprocessed stereo image into the disparity-sensing image encoder to obtain binocular image features; The preprocessed EEG signal is input into a multi-scale attention EEG encoder to obtain the spatiotemporal features of EEG. S3. Input the binocular image features and EEG spatiotemporal features into the complementary learning module to extract image complementary features, EEG complementary features, and cross-modal consistency features. S4. Input the consistency features into the consistency learning module for cross-modal alignment to obtain the optimized consistency feature representation. Then, fuse it with the complementary features and input it into the classification module to output the stereo image quality level.

2. The method according to claim 1, characterized in that, The preprocessing described in S1 specifically includes: Size normalization and parallax correction are performed on the stereoscopic image; Bandpass filtering was performed on multichannel EEG signals, and independent component analysis was used to remove eye movement and electromyography artifacts. The artifact-free EEG signals were segmented according to fixed time windows, and each segment of EEG signal was aligned with the corresponding stereoscopic image and quality grade label.

3. The method according to claim 1, characterized in that, The parallax-aware image encoder described in S2 performs the following operations: A dual-path convolutional feature extraction structure is used to perform multi-scale convolution on the left and right views to obtain initial monocular features; The disparity compensation module is used to perform disparity correlation calculations on the features of the left and right views to obtain a disparity attention map, and the monocular features are weighted to form disparity compensation features. The attention fusion module stitches together the disparity compensation features of the left and right views with the initial monocular features and performs attention weighting to output binocular image features.

4. The method according to claim 1, characterized in that, The multi-scale attention EEG encoder described in S2 performs the following operations: Multi-scale EEG features are extracted in both the temporal and channel dimensions using multi-scale convolutional branches. Weights are assigned to time-dimensional features through the temporal attention module, and weights are assigned to channel-dimensional features through the spatial attention module. By fusing features weighted by temporal attention and spatial attention, spatiotemporal features of EEG are obtained.

5. The method according to claim 1, characterized in that, The complementary learning module described in S3 performs the following operations: Private channels are set up in the image modality and the EEG modality respectively to extract image complementary features and EEG complementary features; A shared channel is set up to extract cross-modal consistency features; By using a loss function that includes complementary constraints and consistency comparison terms, the correlation between complementary features and consistent features is constrained, thereby achieving the decoupling and alignment of modality-specific complementary information and cross-modal consistent information.

6. The method according to claim 1, characterized in that, The classification module described in S4 performs the following operations: After fusing complementary features and consistent features, the data is input into the class-level contrastive learning branch and the task classification branch in the classification module. The class-level contrastive learning branch constructs positive and negative sample pairs based on quality level labels to increase the margin between features of different quality levels. The task classification branch uses cross-entropy loss to supervise the prediction results and the true quality level labels. Output the probability distribution of each quality level, and use the quality level corresponding to the highest probability as the objective quality evaluation result of the stereo image.

7. A computer device, characterized in that, The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the instruction, program, code set, or instruction set is loaded and executed by the processor to implement the multimodal stereo image quality evaluation method based on consistent-complementary features as described in any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, at least one program, code set, or instruction set, which is loaded and executed by a processor to implement the multimodal stereo image quality assessment method based on consistent-complementary features as described in any one of claims 1-6.