Panoramic video quality evaluation method based on human perceptual perception
By integrating texture and semantic features in the 360° panoramic video quality evaluation method and introducing human perceptual characteristics, the shortcomings of ignoring the video feature relationship and human perception in the existing methods are solved, and a more efficient and accurate video quality evaluation is achieved.
Patent Information
- Application Number
- CN202510128106.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-13
AI Technical Summary
The existing 360° panoramic video quality evaluation method ignores the relationship between video texture features and semantic features, and does not fully consider the potential characteristics of human perception, resulting in low evaluation efficiency and insufficient accuracy.
A panoramic video quality evaluation method based on human perceptual perception is proposed. By integrating video texture features and semantic features, and introducing human perceptual characteristics, using space-time consistency sampling and central crop sampling to obtain video quality features, combining perceptual perception metric optimization model.
On the premise of not losing accuracy, the efficiency of 360° panoramic video quality evaluation is significantly improved, which can better reflect human perception of video quality and improve the accuracy and reliability of evaluation.
Smart Images

Figure CN119996654A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of computer vision, video quality evaluation, etc., and specifically relates to a panoramic video quality evaluation method based on human perceptual perception. Background Art
[0002] With the continuous growth of Internet video content, more and more videos of various types are constantly being generated and collected and uploaded to the network. In particular, various new video formats, such as panoramic videos and free viewpoint videos, have penetrated into people's daily lives. Different from traditional 2D videos, 360° panoramic videos can provide users with a brand new visual experience. In order to meet the user's immersive experience quality (QoE) in 360° panoramic videos, it is very necessary to perform video quality assessment (VQA) on 360° videos. The resolution of 360° panoramic videos is much higher than that of traditional videos, which not only puts a huge burden on the storage and transmission of 360° panoramic videos, but also puts higher requirements on the quality assessment methods of 360° panoramic videos.
[0003] In recent years, with the rapid development of 360° panoramic video technology and the popularization of high-definition and ultra-high-definition videos, the field of 360° panoramic video quality evaluation has also made significant progress. Unfortunately, in existing methods, when evaluating video quality, the relationship between video texture features and semantic features is often ignored. For example, the blur of the sky has a significantly smaller impact on the perception than the blur of the tree texture. At the same time, only the quality semantics of the video itself are focused on, but the potential characteristics of human perception, such as selective perception, are not taken into account. Summary of the invention
[0004] Taking the above factors into consideration, the present invention proposes a new 360° video BVQA method, which not only integrates video texture features and semantic features, but also takes into account human perception characteristics, and achieves the best effect in the field of 360° panoramic video quality evaluation.
[0005] In view of this, the purpose of the present invention is to provide a panoramic video quality evaluation method based on human perceptual perception, which can improve the efficiency of 360° panoramic video quality evaluation without losing accuracy as much as possible.
[0006] The technical solution specifically adopted by the present invention to solve the technical problem is:
[0007] A method for evaluating panoramic video quality based on human perception: 360° panoramic videos are collected to construct a data set and preprocessed, the data set input is mapped into an equidistant rectangular projection format, and two different formats are obtained by spatiotemporal consistency sampling and center crop sampling respectively, which are input into a panoramic video quality evaluation model; the spatiotemporal consistency sampling and center crop sampling are input into the video quality feature branch of the panoramic video quality evaluation model, and the center crop sampling is input into the human perception branch of the panoramic video quality evaluation model; in the video quality feature branch, the texture feature subnet encodes the input using consistent sampling to extract video texture features; the video semantic feature subnet encodes the center crop sampling The sampled input is encoded to extract video semantic features; texture features and semantic features are fused through a feature enhancement fusion module to obtain a video quality score that fuses the two information; in the human perceptual perception branch, human perceptual bias is introduced by calculating perceptual perception metrics; a model score is obtained through linear mapping; the predicted video quality score predicted by the panoramic video quality evaluation model is compared with the real video quality score, and the monotonicity loss, linear fusion loss and mean square error are calculated to further optimize the model; the 360° panoramic video to be detected is input into the trained panoramic video quality evaluation model, and the quality of the 360° panoramic video to be detected is judged by the score predicted by the model.
[0008] Furthermore, the process of constructing the dataset is as follows: shoot a 360° panoramic original video, and have the subjects score the shot 360° panoramic original video, and take the average value after excluding outliers as the true quality score y of the current video gt ; All 360° panoramic videos in the dataset are preprocessed, including image resizing, image standardization, and 360° panoramic original video true score standardization.
[0009] Furthermore, the data set input is mapped into an equidistant rectangular projection format, and two different formats are obtained by spatiotemporal consistency sampling and center crop sampling respectively:
[0010] The pre-processed 360° panoramic videos with different projection modes are uniformly mapped into an equidistant rectangular projection format to obtain the video Video ERP ; Perform temporal and spatial uniform sampling and center cropping sampling on the preprocessed 360° panoramic video;
[0011] The calculation method of the spatiotemporal uniform sampling is as follows:
[0012]
[0013] Among them, N h and N w Respectively represent the number of samples in the video height and width, H and W represent VideoERP The height and width of rand_crop means to perform a random crop in a given area. h Indicates the height of vi h Positions, vj w Indicates the width of the vj w Locations, Video sample Indicates Video ERP The results obtained after uniform sampling in time and space;
[0014] The calculation method of the center crop sampling is as follows:
[0015] Video center =center_crop(Video ERP )
[0016] Among them, the center cropping operation in the center_crop table, Video center Indicates Video ERP The result obtained after center crop sampling.
[0017] Furthermore, in the video quality feature branch, Video sample The input texture feature subnet is encoded by convolution, Feature t_input =Conv(Video sample ), where Conv represents convolution calculation, Feature t_input Indicates the result after encoding; Feature t_input The self-attention module performs self-attention operation and extracts the first, second, and third stage feature maps respectively as the video texture quality feature map set Feature texture ;
[0018] Video center The input semantic feature subnet is encoded by convolution. c_input =Conv(Video center ), Feature c_input Indicates the result after encoding; Feature c_input The features are calculated through the residual module, and the first, second, and third stage feature maps are extracted respectively as the video semantic feature map set Feature content ;
[0019] Feature texture and Feature contentThe input feature enhancement fusion module first normalizes them to the same feature space, and then performs a self-attention operation to enhance their important features to obtain Feature texture′ and Feature content′ ; Feature texture′ and Feature content′ The fusion is performed through cross attention, and finally the addition operation is performed, and a linear mapping is used to map it to the 360° panoramic video quality branch prediction score y f_pre .
[0020] Furthermore, in the human perception branch, Video center Encoding through convolution, Feature p_input =Conv(Video center ), where Conv represents convolution calculation, Feature p_input Indicates the result after encoding; Feature p_input The CLIP model is used to calculate the similarity with the four text description prompts ["a high quality photo", "a low quality photo", "a photo contains attractive content", "a photo contains boring content"], and the similarities p1, p2, p3, and p4 are obtained respectively;
[0021] The PDM will be obtained through perceptual measurement:
[0022] PDM=(p1-p2)+(p3-p4)
[0023] Map PDM using the sigmod function:
[0024]
[0025] Get the 360° panoramic video human perception branch prediction score y p_pre .
[0026] Furthermore, the predicted video quality score predicted by the panoramic video quality evaluation model is compared with the real video quality score, and the monotonicity loss, linear fusion loss and mean square error are calculated to further optimize the model. Specifically, y f_pre and p_pre Obtain the 360° panoramic video prediction score y through a linear mapping pre ; y pre and gt Input loss function, where ypre is the final 360° panoramic video quality prediction score, y gt To obtain the true quality score of the 360° panoramic video, the monotonicity loss, linear fusion loss, and mean square error are calculated. After adding the losses, back propagation is performed to conduct supervised training of the 360° panoramic video quality evaluation model.
[0027] Furthermore, the loss function is specifically as follows: the monotonicity loss and the linear fusion loss are calculated for all the true 360° panoramic video quality scores in a training batch and the 360° panoramic video quality scores predicted by the model, and the mean square error between each 360° panoramic video quality score predicted by the model and the true 360° panoramic video quality score is calculated, and finally the losses are summed up to obtain the total loss value:
[0028]
[0029] L mae =(y gt -y pred ) 2
[0030] L total =L mo +αL li +βL mae
[0031] Where sign(·) represents the sign function, <> represents the inner product of two vectors, and s pred and gt A vector representing all the predictions and true labels in a training batch, Respectively represent s pred and gt The average value of Respectively represent the bith batch and bj batch The model predicted value, Respectively represent the bith batch and bj batch The true quality score of the 360° panoramic video, ‖·‖2 represents the square root of the value, L mo represents the monotonicity loss, L li represents the linear fusion loss, L mae represents the mean square error, L total Represents the overall loss of the model.
[0032] And, a panoramic video quality evaluation system based on human perceptual perception, comprising: a panoramic video quality evaluation model, the panoramic video quality evaluation model comprising a video quality feature branch and a human perceptual perception branch: constructing a data set by collecting 360° panoramic videos and performing preprocessing, mapping the data set input into an equidistant rectangular projection format, respectively obtaining two different formats through spatiotemporal consistency sampling and center crop sampling to input into the panoramic video quality evaluation model; inputting the spatiotemporal consistency sampling and center crop sampling into the video quality feature branch, and inputting the center crop sampling into the human perceptual perception branch; in the video quality feature branch, the texture feature subnet encodes the input using consistent sampling to extract video texture features; the video semantic feature subnet encodes the input using center crop sampling to extract video semantic features; the texture features and semantic features are fused through a feature enhancement fusion module to obtain a video quality score that fuses the information of the two; in the human perceptual perception branch, the human perceptual bias is introduced by calculating the perceptual perception metric; and the model score is obtained through linear mapping;
[0033] The training method of the panoramic video quality evaluation model is: comparing the predicted video quality score predicted by the panoramic video quality evaluation model with the real video quality score, and calculating the monotonicity loss, linear fusion loss and mean square error to further optimize the model;
[0034] The 360° panoramic video to be tested is input into the trained panoramic video quality evaluation model, and the quality of the 360° panoramic video to be tested is judged by the score predicted by the model.
[0035] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for evaluating panoramic video quality based on human perceptual perception as described above are implemented.
[0036] A non-transitory computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for evaluating panoramic video quality based on human perceptual perception as described above.
[0037] 1. Compared with the existing 360° panoramic video quality evaluation methods, the effective feature fusion module can be used to model the relationship between video texture features and video semantic features, which greatly improves the model effect;
[0038] 2. Innovatively introduce human perception characteristics into 360° video quality evaluation, and propose perceptual distance measurement to alleviate the difficulty of accurate classification of video quality;
[0039] 3. A dual-branch network is proposed, which not only integrates video texture features and semantic features, but also takes into account human perception characteristics, achieving the best results in the field of 360° video quality evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0041] Figure 1 It is a flow chart of a method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0042] In order to make the features and advantages of this patent more obvious and easy to understand, the following embodiments are specifically described in detail as follows:
[0043] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0044] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0045] Please refer to Figure 1 The present invention provides a method for evaluating panoramic video quality based on human perception, comprising the following steps:
[0046] Step S1: Capture a 360° panoramic video in a correct video format through a 360° panoramic video acquisition camera, and organize the video into a data set for preprocessing;
[0047] Step S2: Map the dataset input into an equidistant rectangular projection format, and obtain two different format inputs of the model through spatiotemporal consistency sampling and center crop sampling respectively;
[0048] Step S3: Input the spatiotemporal consistency sampling and the center crop sampling into the video quality feature branch, and input the center crop sampling into the human perceptual perception branch. In the video quality feature branch, the texture feature subnet encodes the input using consistency sampling to extract video texture features; the video semantic feature subnet encodes the input of the center crop sampling to extract video semantic features. The texture features and semantic features are fused through a fusion module to obtain a video quality score that integrates the information of the two. In the human perceptual perception branch, human perceptual bias is introduced by calculating perceptual perception metrics. Finally, the final score of the model is obtained through linear mapping. The predicted video quality score obtained by the model is compared with the actual video quality score, and their monotonicity loss, linear fusion loss and mean square error are calculated to further optimize the model;
[0049] Step S4: input the 360° panoramic video to be detected into the panoramic video quality evaluation model, and judge the quality of the 360° panoramic video to be detected by the score predicted by the model.
[0050] In this embodiment, step S1 specifically includes the following steps:
[0051] Step S11: Install a 360° panoramic video acquisition camera at the shooting site;
[0052] Step S12: Shoot a 360° panoramic original video and invite subjects to rate the shot 360° panoramic original video. After excluding outliers, take the average value as the true quality score of the current video, called y gt ;
[0053] Step S13: All 360° panoramic videos in the dataset are preprocessed, including image resizing, image standardization, and 360° panoramic original video true score standardization.
[0054] In this embodiment, step S2 specifically includes the following steps:
[0055] Step S21: The pre-processed 360° panoramic videos of different projection modes are uniformly mapped into an equidistant rectangular projection format, and the obtained video is called Video ERp ;
[0056] Step S22: Perform temporal and spatial uniform sampling on the pre-processed 360° panoramic video. The calculation method is as follows:
[0057]
[0058] Among them, N h and N w Respectively represent the number of samples in the video height and width, H and W represent Video ERPThe height and width of the , rand_crop means to perform a random crop in a given area, i h Indicates the height of vi h Positions, vj w Indicates the width of the vj w Locations, Video sample Indicates Video ERP The results obtained after uniform sampling in time and space;
[0059] Step S23: Video sample Input the 360° panoramic video quality assessment model for further processing;
[0060] In this embodiment, step S3 specifically includes the following steps:
[0061] Step S311: Video sample The input texture feature subnet is simply encoded by convolution. t_input =Conv(Video sample ), where Conv represents convolution calculation, Feature t_input Indicates the result obtained after encoding;
[0062] Step S312: Feature t_input The self-attention module performs self-attention operation and extracts the first, second, and third stage feature maps respectively as a set of video texture quality feature maps, called Feature texture ;
[0063] Step S321: Video center The input semantic feature subnet is simply encoded through convolution. c_input =Conv(Video center ), where Conv represents convolution calculation, Feature c_input Indicates the result obtained after encoding;
[0064] Step S322: Feature c_input The features are calculated through the residual module, and the first, second, and third stage feature maps are extracted respectively as a set of video semantic feature maps, called Feature content ;
[0065] Step S331: Feature texture and Feature content The input feature enhancement fusion module first normalizes them to the same feature space, and then performs a self-attention operation to enhance their important features to obtain Feature texture′and Feature content′ .
[0066] Step S332: Feature texture′ and Feature content′ The two are fused by cross attention, and finally added, and linearly mapped to the 360° panoramic video quality branch prediction score, called y f_pre ;
[0067] Step S341: Video center The input human perception branch is simply encoded through convolution, Feature p_input =Conv(Video center ), where Conv represents convolution calculation, Feature p_input Indicates the result obtained after encoding;
[0068] Step S342: Calculate the similarity between Featurep_input and the four prompts [“a high quality photo”, “alow quality photo”, “a photo contains attractive content”, “a photo contains boring content”] through the CLIP model, and obtain similarities p1, p2, p3, and p4 respectively;
[0069] Step S343: Obtain PDM through perceptual measurement:
[0070] PDM=(p1-p2)+(p3-p4)
[0071] Then use the sigmod function to map PDM
[0072]
[0073] Get the 360° panoramic video human perception branch prediction score, called y p_pre ;
[0074] Step S35: Set y f_pre and p_pre The 360° panoramic video prediction score is obtained through a simple linear mapping, called y pre ;
[0075] Step S36: y pre and gt Input loss function, where y pre is the final 360° panoramic video quality prediction score, ygt To obtain the true quality score of the 360° panoramic video, the monotonicity loss, linear fusion loss, and mean square error are calculated. After adding the losses, back propagation is performed to conduct supervised training of the 360° panoramic video quality evaluation model.
[0076] In this embodiment, step S4 specifically includes the following steps:
[0077] Step S41: The 360° panoramic video to be evaluated for quality is also processed through steps S21 and S22 to obtain a pre-processed 360° panoramic video Video process ;
[0078] Step S42: Video process The corresponding 360° panoramic video quality score is calculated by step S3, and the higher the score is, the better the quality of the corresponding video is.
[0079] This embodiment further provides a systematic implementation of the above method:
[0080] A panoramic video quality evaluation model deployed on a computer system includes a video quality feature branch and a human perception branch: a data set is constructed by collecting 360° panoramic videos and preprocessing is performed, the data set input is mapped into an equidistant rectangular projection format, and two different formats are obtained by spatiotemporal consistency sampling and center crop sampling respectively to input into the panoramic video quality evaluation model; the spatiotemporal consistency sampling and center crop sampling are input into the video quality feature branch, and the center crop sampling is input into the human perception branch; in the video quality feature branch, the texture feature subnet encodes the input using consistent sampling to extract video texture features; the video semantic feature subnet encodes the input of the center crop sampling to extract video semantic features; the texture features and the semantic features are fused through a feature enhancement fusion module to obtain a video quality score that fuses the information of the two; in the human perception branch, the human perception bias is introduced by calculating the perception metric; and the model score is obtained through linear mapping;
[0081] The training method of the panoramic video quality evaluation model is: comparing the predicted video quality score predicted by the panoramic video quality evaluation model with the real video quality score, and calculating the monotonicity loss, linear fusion loss and mean square error to further optimize the model;
[0082] The 360° panoramic video to be tested is input into the trained panoramic video quality evaluation model, and the quality of the 360° panoramic video to be tested is judged by the score predicted by the model.
[0083] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.
[0084] It needs to be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and the computer program is executed by a processor to execute the above method. The storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.
[0085] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should be understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0086] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.
[0087] This patent is not limited to the above-mentioned optimal implementation mode. Anyone can derive other forms of panoramic video quality evaluation methods based on human perception under the inspiration of this patent. All equal changes and modifications made according to the scope of the patent application of this invention should be covered by this patent.
Claims
1. A panoramic video quality assessment method based on human perception, characterized in that: Collect 360° panoramic videos to construct a dataset and perform preprocessing. Map the dataset input into an equidistant rectangular projection format, and obtain two different formats through spatiotemporal consistency sampling and center crop sampling, respectively, to input into the panoramic video quality assessment model; input the spatiotemporal consistency sampling and center crop sampling into the video quality feature branch of the panoramic video quality assessment model, and input the center crop sampling into the human perception branch of the panoramic video quality assessment model; In the video quality feature branch, the texture feature subnet encodes the input using consistent sampling to extract video texture features; The video semantic feature subnet encodes the input of the center crop sample and extracts the video semantic features; Texture features and semantic features are fused through a feature enhancement fusion module to obtain a video quality score that integrates the information of the two. In the human perceptual perception branch, human perceptual bias is introduced by calculating perceptual perception metrics. A model score is obtained through linear mapping. The predicted video quality score predicted by the panoramic video quality evaluation model is compared with the true video quality score, and the monotonicity loss, linear fusion loss and mean square error are calculated to further optimize the model. The 360° panoramic video to be detected is input into the trained panoramic video quality evaluation model, and the quality of the 360° panoramic video to be detected is judged by the score predicted by the model.
2. The method for evaluating panoramic video quality based on human perception according to claim 1, characterized in that: The specific process of constructing the dataset is as follows: shoot a 360° panoramic original video, and have the subjects score the shot 360° panoramic original video, and take the average value after excluding outliers as the true quality score y of the current video gt ; All 360° panoramic videos in the dataset are preprocessed, including image resizing, image standardization, and 360° panoramic original video true score standardization.
3. The method for evaluating panoramic video quality based on human perception according to claim 1, characterized in that: The data set input is mapped into an equidistant rectangular projection format, and two different formats are obtained by spatiotemporal consistency sampling and center crop sampling respectively: The pre-processed 360° panoramic videos with different projection modes are uniformly mapped into an equidistant rectangular projection format to obtain the video Video ERP ; Perform temporal and spatial uniform sampling and center cropping sampling on the preprocessed 360° panoramic video; The calculation method of the spatiotemporal uniform sampling is as follows: Among them, N h and N w Respectively represent the number of samples in the video height and width, H and W represent Video ERP The height and width of rand_crop means to perform a random crop in a given area. h Indicates the height of vi h Positions, vj w Indicates the width of the vj w Locations, Video sample Indicates Video ERP The results obtained after uniform sampling in time and space; The calculation method of the center crop sampling is as follows: Video center =center_crop(Video ERP ) Among them, the center cropping operation in the center_crop table, Video center Indicates Video ERP The result obtained after center crop sampling.
4. The method for evaluating panoramic video quality based on human perception according to claim 3, characterized in that: In the video quality feature branch, Video sample The input texture feature subnet is encoded by convolution, Feature t_input =Conv(Video sample ), where Conv represents convolution calculation, Feature t_input Indicates the result after encoding; Feature t_input The self-attention module performs self-attention operation and extracts the first, second, and third stage feature maps respectively as the video texture quality feature map set Feature texture ; Video center The input semantic feature subnet is encoded by convolution. c_input =Conv(Video center ), Feature c_input Indicates the result after encoding; Feature c_input The features are calculated through the residual module, and the first, second, and third stage feature maps are extracted respectively as the video semantic feature map set Feature content ; Feature texture and Feature content The input feature enhancement fusion module is first normalized to the same feature space, and then a self-attention operation is performed to enhance its own important features to obtain Feature texture' and Feature content' ; Feature texture' and Feature content' The fusion is performed through cross attention, and finally the addition operation is performed, and a linear mapping is used to map it to the 360° panoramic video quality branch prediction score y f_pre .
5. The method for evaluating panoramic video quality based on human perception according to claim 4, characterized in that: In the human perception branch, Video center Encoding through convolution, Feature p_input = Conv(Video center ), where Conv represents convolution calculation, Feature p_input Indicates the result after encoding; Feature p_input The CLIP model is used to calculate the similarity with the four text description prompts ["ahigh quality photo", "alow quality photo", "aphoto contains attractive content", "aphoto contains boring content"], and the similarities p1, p2, p3, and p4 are obtained respectively; The PDM will be obtained through perceptual measurement: PDM=(p1-p2)+(p3-p4) Map PDM using the sigmod function: Get the 360° panoramic video human perception branch prediction score y p_pre .
6. The method for evaluating panoramic video quality based on human perception according to claim 5, characterized in that: The predicted video quality score predicted by the panoramic video quality evaluation model is compared with the actual video quality score, and the monotonicity loss, linear fusion loss and mean square error are calculated to further optimize the model. Specifically, y f_pre and p_pre Obtain the 360° panoramic video prediction score y through a linear mapping pre ; y pre and gt Input loss function, where y pre is the final 360° panoramic video quality prediction score, y gt To obtain the true quality score of the 360° panoramic video, the monotonicity loss, linear fusion loss, and mean square error are calculated. After adding the losses, back propagation is performed to conduct supervised training of the 360° panoramic video quality evaluation model.
7. The method for evaluating panoramic video quality based on human perception according to claim 6, characterized in that: The loss function is specifically as follows: the monotonicity loss and linear fusion loss are calculated for all the true 360° panoramic video quality scores in a training batch and the 360° panoramic video quality scores predicted by the model, and the mean square error between each 360° panoramic video quality score predicted by the model and the true 360° panoramic video quality score is calculated, and finally the losses are summed up to obtain the total loss value: L mae (and gt -and pred ) 2 L total =L mo +αL li +βL mae Where sign(·) represents the sign function, <> represents the inner product of two vectors, and s pred and gt A vector representing all the predictions and true labels in a training batch, Respectively represent s pred and gt The average value, Respectively represent the bith batch and bj batch The model predicted value, Respectively represent the bith batch and bj batch The true quality score of the 360° panoramic video, ||·||2 represents the square root of the value, L mo represents the monotonicity loss, L li represents the linear fusion loss, L mae represents the mean square error, L total Represents the overall loss of the model.
8. A panoramic video quality assessment system based on human perception, characterized in that: include: A panoramic video quality evaluation model, the panoramic video quality evaluation model includes a video quality feature branch and a human perceptual perception branch: a data set is constructed by collecting 360° panoramic videos and preprocessing is performed, the data set input is mapped into an equidistant rectangular projection format, and two different formats are obtained by spatiotemporal consistency sampling and center crop sampling respectively to input into the panoramic video quality evaluation model; the spatiotemporal consistency sampling and center crop sampling are input into the video quality feature branch, and the center crop sampling is input into the human perceptual perception branch; in the video quality feature branch, the texture feature subnet encodes the input using consistent sampling to extract video texture features; the video semantic feature subnet encodes the input of the center crop sampling to extract video semantic features; Texture features and semantic features are fused through a feature enhancement fusion module to obtain a video quality score by integrating the information of the two; in the human perceptual perception branch, human perceptual bias is introduced by calculating perceptual perceptual metrics; and a model score is obtained through linear mapping; The training method of the panoramic video quality evaluation model is: comparing the predicted video quality score predicted by the panoramic video quality evaluation model with the real video quality score, and calculating the monotonicity loss, linear fusion loss and mean square error to further optimize the model; The 360° panoramic video to be tested is input into the trained panoramic video quality evaluation model, and the quality of the 360° panoramic video to be tested is judged by the score predicted by the model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the panoramic video quality assessment method based on human perceptual perception are implemented as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the method for evaluating panoramic video quality based on human perceptual perception as claimed in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Two-stage fine tuning and decoupling reasoning method and device for visual language model
CN122154841A