This invention discloses a method and
system for evaluating the quality of immersive videos based on multimodal
perception, relating to the field of video evaluation. The method includes: acquiring a multi-viewpoint texture and depth format video and extracting keyframes; extracting texture and depth features using ConvNeXt,
processing texture features through a frequency-space
texture enhancement module, and combining depth features with a cross-
modal collaborative representation module to obtain structural texture
coupling features; extracting semantic
distortion features using a semantic
distortion perception module; concatenating the
coupling features and semantic
distortion features and inputting them into a temporal modeling module to capture temporal features, then generating global
perception features through a viewpoint fusion module; and finally outputting the final
video quality score through a quality regression module. This invention achieves a more accurate and robust evaluation of immersive
video quality by fusing multi-viewpoint texture and depth information and sequentially performing frequency-space
texture enhancement, cross-
modal collaborative representation, semantic distortion perception, temporal modeling, and viewpoint fusion.