HDR video quality evaluation method, device and equipment and storage medium

CN122741684APending Publication Date: 2026-09-11MALANSHAN AUDIO & VIDEO LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611126717.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

然而,现有的带主观质量标注的视频数据大多面向SDR内容构建,HDR主观质量数据集在数据规模、内容多样性以及失真类型覆盖等方面明显不足

Benefits of technology

[0014]In this application, when evaluating the video quality of HDR videos, SDR frame data and HDR frame data based on content pairing are acquired. A first training set is constructed based on the SDR frame data and the HDR frame data. The SDR frame data in the first training set is used as input data, and the HDR frame data corresponding to the SDR frame data is used as the supervision target to train a pre-constructed residual domain adapter to obtain a target residual domain adapter. The residual domain adapter is a domain adapter based on log-odds transformation. Target video data is acquired, and the dynamic range domain identifier corresponding to the target video data is determined based on the data source. A second training set is constructed based on the target video data. An initial quality evaluation model is constructed using conditional routing, identity bypass, a quality evaluation backbone, and the target residual domain adapter. The target video data includes SDR video data and HDR video data, and the target video data carries a corresponding subjective quality score. The dynamic range domain identifier includes an SDR domain identifier and an HDR domain identifier. The conditional routing is used to determine the forward processing path from the target video data input to the quality evaluation backbone based on the dynamic range domain identifier. The path is defined as follows: When the identity bypass is used as the forward processing path, the output data of the forward processing path is the same as the input data; the quality evaluation backbone is a no-reference quality evaluation network used to obtain the video quality score corresponding to the video data; using the conditional routing, when the dynamic range domain identifier corresponding to the target video data in the second training set is the SDR domain identifier, the target video data is adapted through the target residual domain adapter to obtain adapted samples, and the adapted samples are input into the quality evaluation backbone to obtain a first predicted quality score; when the dynamic range domain identifier corresponding to the target video data in the second training set is the HDR domain identifier, the target video data is input into the quality evaluation backbone through the identity bypass to obtain a second predicted quality score; based on the predicted quality score and subjective quality score corresponding to the target video data, the quality evaluation backbone and the target residual domain adapter in the initial quality evaluation model are trained to obtain a target quality evaluation model, and the target quality evaluation model is used to obtain the target video quality score corresponding to the HDR video to be evaluated; the predicted quality score includes a first predicted quality score and a second predicted quality score. As can be seen, this application expands the range of available training data without changing the no-reference attribute in the inference stage by combining three stages: paired domain transfer pre-training, dynamic range conditional routing, and cross-domain quality joint optimization. It also reduces the feature distribution shift caused by cross-dynamic range mixed training, thereby improving the accuracy and generalization ability of HDR compressed video quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122741684A_ABST
    Figure CN122741684A_ABST
Patent Text Reader

Abstract

This application discloses an HDR video quality assessment method, apparatus, device, and storage medium, relating to the field of video processing. The method includes: constructing a first training set based on content-paired SDR and HDR frame data to train a target residual domain adapter; utilizing conditional routing, when the dynamic range domain identifier corresponding to the target video data in the second training set is an SDR domain identifier, adapting the target video data through the target residual domain adapter and inputting it into a quality assessment backbone to obtain a first predicted quality score; when the dynamic range domain identifier is an HDR domain identifier, inputting the target video data into the quality assessment backbone through an identity bypass to obtain a second predicted quality score; obtaining a target quality assessment model based on the predicted quality score and subjective quality score, and using the target quality assessment model to obtain the target video quality score of the HDR video to be evaluated. This application improves the accuracy of referenceless quality assessment of HDR compressed video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing, and in particular to a method, apparatus, device, and storage medium for evaluating HDR video quality. Background Technology

[0002] With the continuous development of video capture and display technologies, High Dynamic Range (HDR) video, with its higher peak brightness, wider color gamut, and richer tonal gradations, has gradually become an important video format in film and television production, streaming media distribution, live sports broadcasts, and terminal playback. Compared with Standard Dynamic Range (SDR) video, HDR video typically uses a coding depth of ten bits or more and employs photoelectric transfer characteristics such as Perceptual Quantization (PQ) or Hybrid Log-Gamma (HLG). The mapping relationship between its coded values ​​and display brightness differs significantly from that of SDR video; the same coded value corresponds to different display brightness, local contrast, and perceptible detail levels in the two types of signals. HDR video inevitably undergoes lossy compression processing at each stage of production, encoding, transmission, and distribution. Most no-reference video quality assessment methods in related technologies are based on deep neural networks, learning the mapping relationship from video frames to quality scores through end-to-end training on data with subjective quality labels. The predictive performance of such methods is highly dependent on the scale and diversity of the training data. However, most existing video data with subjective quality annotations are designed for SDR content, and HDR subjective quality datasets are significantly lacking in terms of data scale, content diversity, and coverage of distortion types.

[0003] Therefore, improving the accuracy of referenceless quality assessment of HDR compressed video under the condition of limited HDR subjective quality data scale is an urgent problem to be solved. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a method, apparatus, device, and storage medium for evaluating HDR video quality, which can improve the accuracy of no-reference quality evaluation of HDR compressed video under the condition of limited HDR subjective quality data scale. The specific solution is as follows: Firstly, this application discloses a method for evaluating HDR video quality, including: Acquire content-paired SDR frame data and HDR frame data, construct a first training set based on the SDR frame data and the HDR frame data, use the SDR frame data in the first training set as input data, and use the HDR frame data corresponding to the SDR frame data as the supervision target to train a pre-constructed residual domain adapter to obtain a target residual domain adapter; wherein, the residual domain adapter is a domain adapter based on log-odds transformation. The process involves acquiring target video data, determining the dynamic range domain identifier corresponding to the target video data based on its data source, constructing a second training set based on the target video data, and building an initial quality assessment model using conditional routing, identity bypass, a quality assessment backbone, and the target residual domain adapter. The target video data includes SDR and HDR video data, and carries corresponding subjective quality scores. The dynamic range domain identifier includes SDR and HDR domain identifiers. Conditional routing is used to determine the forward processing path from the target video data input to the quality assessment backbone based on the dynamic range domain identifier. When the identity bypass is used as the forward processing path, the output data of the forward processing path is the same as the input data. The quality assessment backbone is a no-reference quality assessment network used to obtain the video quality score corresponding to the video data. Using the conditional routing, when the dynamic range domain identifier corresponding to the target video data in the second training set is an SDR domain identifier, the target video data is adapted through the target residual domain adapter to obtain adapted samples. The adapted samples are then input into the quality evaluation backbone to obtain a first predicted quality score. When the dynamic range domain identifier corresponding to the target video data in the second training set is an HDR domain identifier, the target video data is input into the quality evaluation backbone through the identity bypass to obtain a second predicted quality score. The quality assessment backbone and the target residual domain adapter in the initial quality assessment model are trained based on the predicted quality score and subjective quality score corresponding to the target video data to obtain the target quality assessment model, and the target video quality score corresponding to the HDR video to be evaluated is obtained using the target quality assessment model; the predicted quality score includes a first predicted quality score and a second predicted quality score.

[0005] Optionally, obtaining the target video quality score corresponding to the HDR video to be evaluated using the target quality evaluation model includes: Obtain the HDR video to be evaluated, and preprocess the HDR video to be evaluated to obtain the corresponding HDR frame sequence; The HDR frame sequence is input into the quality assessment backbone of the target quality assessment model through the identity bypass, so as to obtain the quality-perceived features corresponding to each sampled frame in the HDR frame sequence using the quality assessment backbone, and determine the frame-level quality response corresponding to each sampled frame based on the quality-perceived features. The frame-level quality responses of the HDR frame sequence are time-aggregated to obtain the target video quality score corresponding to the HDR video to be evaluated.

[0006] Optionally, the step of constructing a first training set based on the SDR frame data and the HDR frame data, using the SDR frame data in the first training set as input data, and using the HDR frame data corresponding to the SDR frame data as the supervision target, to train a pre-constructed residual domain adapter to obtain a target residual domain adapter includes: The content correspondence between the SDR frame data and the HDR frame data is established based on the video identifier and frame sequence number. The encoded value of the SDR frame data is normalized to a preset bounded interval to obtain the normalized SDR encoded value. The encoded value of the HDR frame data is normalized to a preset bounded interval to obtain the normalized HDR encoded value. Spatial size alignment is performed on the normalized SDR encoded value and the normalized HDR encoded value to obtain the first training set. The SDR frames in the first training set are input into the pre-constructed residual domain adapter to obtain the adaptation results corresponding to the SDR frames. Based on the content correspondence, the corresponding HDR frame of the SDR frame is determined, and the adaptation parameters of the residual domain adapter are updated based on the pixel-level consistency loss between the adaptation result and the HDR frame to obtain the updated residual domain adapter. Determine whether the updated residual domain adapter meets the first training termination condition. If the updated residual domain adapter meets the first training termination condition, use the updated residual domain adapter to obtain the pixel-level consistency index corresponding to the first preset validation set. Determine the adaptation parameter corresponding to the residual domain adapter when the pixel-level consistency index is optimal as the target adaptation parameter, and determine the residual domain adapter with the target adaptation parameter as the target residual domain adapter.

[0007] Optionally, the step of adapting the target video data using the target residual domain adapter to obtain adapted samples includes: Boundary protection processing is performed on the normalized encoded values ​​of the target video data to obtain the encoded values ​​to be transformed located within a preset open interval; A log-odds transform is performed on the encoded value to be transformed, mapping the encoded value to the target transform domain to obtain the target mapping result; the target transform domain is the log-odds transform domain. In the target transform domain, the target residual domain adapter is used to apply a residual correction amount to the target mapping result to obtain the target correction result; Perform the inverse log-probability transformation on the target correction result to obtain the adapted sample located within the preset open interval.

[0008] Optionally, training the quality assessment backbone and the target residual domain adapter in the initial quality assessment model based on the predicted quality score and subjective quality score corresponding to the target video data to obtain the target quality assessment model includes: Based on the predicted quality score and subjective quality score corresponding to each target video data in the second training set, determine the target relevance constraint and target ranking constraint corresponding to the second training set; The joint training loss corresponding to the second training set is determined based on the target relevance constraint and the target ranking constraint. Backpropagation is performed based on the joint training loss to simultaneously update the target residual domain adapter and the quality assessment backbone to obtain the target quality assessment model.

[0009] Optionally, the step of performing backpropagation based on the joint training loss to simultaneously update the target residual domain adapter and the quality assessment backbone to obtain the target quality assessment model includes: Backpropagation is performed based on the joint training loss to simultaneously update the target residual domain adapter and the quality assessment backbone to obtain the updated initial quality assessment model. Determine whether the updated initial quality evaluation model meets the second training termination condition. If the updated initial quality evaluation model meets the second training termination condition, perform deterministic sampling on the second preset validation set to obtain validation sampling frames. Use the updated initial quality evaluation model to obtain the quality prediction performance index corresponding to the validation sampling frames. Determine the model parameters of the initial quality evaluation model corresponding to the optimal quality prediction performance index as the target model parameters. Determine the initial quality evaluation model with the target model parameters as the target quality evaluation model.

[0010] Optionally, the HDR video quality evaluation method further includes: If the video resolution of the target video data in the second training set is greater than a preset resolution threshold, then the target residual domain adapter and the quality evaluation backbone are updated simultaneously based on a preset resource optimization strategy. The preset resource optimization strategy includes a first optimization strategy based on micro-batch forward pass, a second optimization strategy based on gradient checkpoints, and a third optimization strategy based on a combination of micro-batch forward pass and gradient checkpoints.

[0011] Secondly, this application discloses an HDR video quality evaluation device, comprising: An adapter training module is used to acquire content-paired SDR frame data and HDR frame data, construct a first training set based on the SDR frame data and the HDR frame data, use the SDR frame data in the first training set as input data, and use the HDR frame data corresponding to the SDR frame data as the supervision target to train a pre-constructed residual domain adapter to obtain a target residual domain adapter; wherein, the residual domain adapter is a domain adapter based on log-odds transformation; The evaluation model construction module is used to acquire target video data, determine the dynamic range domain identifier corresponding to the target video data based on the data source corresponding to the target video data, and construct a second training set based on the target video data. An initial quality evaluation model is constructed using conditional routing, identity bypass, a quality evaluation backbone, and the target residual domain adapter. The target video data includes SDR video data and HDR video data, and the target video data carries a corresponding subjective quality score. The dynamic range domain identifier includes an SDR domain identifier and an HDR domain identifier. The conditional routing is used to determine the forward processing path from the target video data input to the quality evaluation backbone based on the dynamic range domain identifier. When the identity bypass is used as the forward processing path, the output data of the forward processing path is the same as the input data. The quality evaluation backbone is a no-reference quality evaluation network used to acquire the video quality score corresponding to the video data. The quality score prediction module is used to utilize the conditional routing to adapt the target video data through the target residual domain adapter to obtain adapted samples when the dynamic range domain identifier corresponding to the target video data in the second training set is an SDR domain identifier, and input the adapted samples into the quality evaluation backbone to obtain a first predicted quality score; and when the dynamic range domain identifier corresponding to the target video data in the second training set is an HDR domain identifier, the module uses the identity bypass to input the target video data into the quality evaluation backbone to obtain a second predicted quality score. The video evaluation module is used to train the quality evaluation backbone and the target residual domain adapter in the initial quality evaluation model based on the predicted quality score and subjective quality score corresponding to the target video data to obtain a target quality evaluation model, and to use the target quality evaluation model to obtain the target video quality score corresponding to the HDR video to be evaluated; the predicted quality score includes a first predicted quality score and a second predicted quality score.

[0012] Thirdly, this application discloses an electronic device, including: Memory, used to store computer programs; A processor is used to execute the computer program to implement the aforementioned HDR video quality evaluation method.

[0013] Fourthly, this application discloses a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned HDR video quality evaluation method.

[0014] In this application, when evaluating the video quality of HDR videos, SDR frame data and HDR frame data based on content pairing are acquired. A first training set is constructed based on the SDR frame data and the HDR frame data. The SDR frame data in the first training set is used as input data, and the HDR frame data corresponding to the SDR frame data is used as the supervision target to train a pre-constructed residual domain adapter to obtain a target residual domain adapter. The residual domain adapter is a domain adapter based on log-odds transformation. Target video data is acquired, and the dynamic range domain identifier corresponding to the target video data is determined based on the data source. A second training set is constructed based on the target video data. An initial quality evaluation model is constructed using conditional routing, identity bypass, a quality evaluation backbone, and the target residual domain adapter. The target video data includes SDR video data and HDR video data, and the target video data carries a corresponding subjective quality score. The dynamic range domain identifier includes an SDR domain identifier and an HDR domain identifier. The conditional routing is used to determine the forward processing path from the target video data input to the quality evaluation backbone based on the dynamic range domain identifier. The path is defined as follows: When the identity bypass is used as the forward processing path, the output data of the forward processing path is the same as the input data; the quality evaluation backbone is a no-reference quality evaluation network used to obtain the video quality score corresponding to the video data; using the conditional routing, when the dynamic range domain identifier corresponding to the target video data in the second training set is the SDR domain identifier, the target video data is adapted through the target residual domain adapter to obtain adapted samples, and the adapted samples are input into the quality evaluation backbone to obtain a first predicted quality score; when the dynamic range domain identifier corresponding to the target video data in the second training set is the HDR domain identifier, the target video data is input into the quality evaluation backbone through the identity bypass to obtain a second predicted quality score; based on the predicted quality score and subjective quality score corresponding to the target video data, the quality evaluation backbone and the target residual domain adapter in the initial quality evaluation model are trained to obtain a target quality evaluation model, and the target quality evaluation model is used to obtain the target video quality score corresponding to the HDR video to be evaluated; the predicted quality score includes a first predicted quality score and a second predicted quality score. As can be seen, this application expands the range of available training data without changing the no-reference attribute in the inference stage by combining three stages: paired domain transfer pre-training, dynamic range conditional routing, and cross-domain quality joint optimization. It also reduces the feature distribution shift caused by cross-dynamic range mixed training, thereby improving the accuracy and generalization ability of HDR compressed video quality evaluation. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0016] Figure 1 This is a flowchart of an HDR video quality evaluation method disclosed in this application; Figure 2 This is a schematic diagram illustrating the training process of a specific HDR no-reference quality assessment model disclosed in this application; Figure 3 This is a schematic diagram of a specific HDR video quality evaluation process disclosed in this application; Figure 4 This is a schematic diagram of the structure of an HDR video quality evaluation device disclosed in this application; Figure 5 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] With the continuous development of video capture and display technologies, High Dynamic Range (HDR) video, with its higher peak brightness, wider color gamut, and richer tonal gradations, has gradually become an important video format in film and television production, streaming media distribution, live sports broadcasts, and terminal playback. Compared with Standard Dynamic Range (SDR) video, HDR video typically uses a coding depth of ten bits or more and employs photoelectric transfer characteristics such as Perceptual Quantization (PQ) or Hybrid Log-Gamma (HLG). The mapping relationship between its coded values ​​and display brightness differs significantly from that of SDR video; the same coded value corresponds to different display brightness, local contrast, and perceptible detail levels in the two types of signals. HDR video inevitably undergoes lossy compression processing at each stage of production, encoding, transmission, and distribution. Most no-reference video quality assessment methods in related technologies are based on deep neural networks, learning the mapping relationship from video frames to quality scores through end-to-end training on data with subjective quality labels. The predictive performance of such methods is highly dependent on the scale and diversity of the training data. However, most existing video data with subjective quality annotations are designed for SDR content, and HDR subjective quality datasets are significantly insufficient in terms of data scale, content diversity, and distortion type coverage. To address these technical issues, this application discloses an HDR video quality assessment method that can improve the accuracy of referenceless quality assessment for HDR compressed videos under the condition of limited HDR subjective quality data scale.

[0019] See Figure 1 As shown, this embodiment of the invention discloses a method for evaluating HDR video quality, including: Step S11: Obtain content-paired SDR frame data and HDR frame data, construct a first training set based on the SDR frame data and the HDR frame data, use the SDR frame data in the first training set as input data, and use the HDR frame data corresponding to the SDR frame data as the supervision target to train the pre-constructed residual domain adapter to obtain the target residual domain adapter; wherein, the residual domain adapter is a domain adapter based on log-odds transformation.

[0020] This embodiment can include two parts: a model training phase and a quality evaluation phase. The model training phase uses paired SDR / HDR data and cross-dynamic range data with subjective quality labels as processing objects, and obtains an HDR no-reference video quality evaluation model through two-stage optimization. The quality evaluation phase uses the HDR compressed video to be evaluated as the processing object, and calls the trained model to output a video-level quality score. For example... Figure 2 As shown, the pre-training phase of the domain adapter, also known as the first training phase, uses content-paired SDR and HDR frame data. These refer to two sets of frame data whose content corresponds to each other; that is, two versions of the same video content, scene, and moment, existing in SDR and HDR formats respectively. These can originate from different dynamic range versions of the same master, or from versions obtained by converting between the two dynamic ranges through a predetermined production chain. This embodiment of the invention does not limit the specific acquisition method of the paired frame data, as long as the content correspondence between frames can be established. It is understood that the first training set used in this step does not require subjective quality labels; its supervision signal comes entirely from the pairing relationship itself, i.e., using SDR frames as source domain input and the corresponding HDR frames as target domain supervision targets. Therefore, the pre-training of the domain adapter does not consume limited HDR subjective quality labeling resources, and the scale of available paired data is not limited by the cost of subjective experiments.

[0021] In this embodiment, the residual domain adapter is a domain adapter based on log-probability transform. That is, this adapter does not directly correct the encoded value in the pixel domain, but first maps the bounded normalized encoded value to the unbounded log-probability transform domain, applies an additive residual correction within this transform domain, and then maps the correction result back to the original bounded interval through an inverse transform. This "bounded-unbounded-bounded" transform framework can achieve stronger nonlinear brightness distribution correction capability while ensuring the validity of the output encoded value.

[0022] In one specific implementation, a first training set is constructed based on SDR frame data and HDR frame data. The SDR frame data in the first training set is used as input data, and the corresponding HDR frame data is used as the supervision target. A pre-constructed residual domain adapter is trained to obtain a target residual domain adapter. This includes: establishing a content correspondence between SDR frame data and HDR frame data based on video identifiers and frame numbers; normalizing the encoded values ​​of the SDR frame data to a preset bounded interval to obtain normalized SDR encoded values; normalizing the encoded values ​​of the HDR frame data to a preset bounded interval to obtain normalized HDR encoded values; and performing spatial size alignment on the normalized SDR encoded values ​​and normalized HDR encoded values ​​to obtain the first training set. The SDR frames in the first training set are input into a pre-constructed residual domain adapter to obtain the adaptation results corresponding to the SDR frames. Based on the content correspondence, the corresponding HDR frames of the SDR frames are determined, and the adaptation parameters of the residual domain adapter are updated based on the pixel-level consistency loss between the adaptation results and the HDR frames to obtain the updated residual domain adapter. It is determined whether the updated residual domain adapter meets the first training termination condition. When the updated residual domain adapter meets the first training termination condition, the pixel-level consistency index corresponding to the first preset validation set is obtained using the updated residual domain adapter. The adaptation parameters corresponding to the residual domain adapter with the optimal pixel-level consistency index are determined as the target adaptation parameters, and the residual domain adapter with the target adaptation parameters is determined as the target residual domain adapter.

[0023] Specifically, this embodiment can establish a pairing index based on video identifiers and frame numbers. That is, a one-to-one correspondence is established between SDR frames and HDR frames in the same scene according to video names and frame numbers, so that any SDR frame can uniquely identify the corresponding HDR frame. The video identifier can be a video file name, video sequence number, or other information that can uniquely identify a video content. This embodiment does not limit this. Furthermore, considering that the encoding bit depth of SDR frame data and HDR frame data is usually different, for example, SDR frame data can be eight-bit data, and HDR frame data can be ten-bit or twelve-bit data, it is necessary to normalize the encoding values ​​of the two types of frame data to a preset bounded interval to eliminate the numerical scale difference caused by the difference in bit depth. As an optional implementation, the encoding value can be divided by the maximum value corresponding to its bit depth, thereby normalizing the encoding value to the [0,1] interval; or the encoding value can be normalized to other preset bounded intervals first, and then the interval can be further mapped to a unit interval. The normalized SDR frame can be denoted as The HDR frames corresponding to its content can be denoted as .

[0024] Furthermore, to ensure that paired frames strictly correspond in spatial location and that the network input has a uniform tensor shape, this embodiment of the invention also performs spatial size alignment on the normalized two types of encoded values. Spatial size alignment can be achieved through scaling, cropping, or a combination of scaling and cropping. For example, the frame can be scaled proportionally according to a preset short side length, and then center-cropped or randomly cropped to obtain a frame tensor of a preset input size. It should be noted that the same spatial transformation parameters should be applied to the paired SDR frames and HDR frames to ensure that the pixel positions of the two still correspond one-to-one after alignment. Specifically, the normalized and aligned SDR frames... Input the pre-built residual domain adapter A(·) to obtain the adaptation result A( The adaptation result matches the target HDR frame. They have the same tensor shape and the same value range. Then, based on the content correspondence established in the previous steps, the HDR frame corresponding to the current SDR frame is determined. And constrained the adaptation result A with pixel-level consistency loss. ) and target HDR frame Consistency between pixels. As a preferred implementation, pixel-level consistency loss can be achieved using mean squared error, which can be expressed as: ; Where N is the batch size, C, H, and W are the number of channels, height, and width of the frame tensor, respectively, and sg(·) indicates stopping the gradient propagation operation. This represents the summation over all elements in the batch dimension, channel dimension, and spatial dimension. Besides mean squared error, pixel-level consistency loss can also be achieved using mean absolute error, Charbonnier loss, or a weighted combination of the above losses; this embodiment does not limit this approach.

[0025] It should be noted that, in this embodiment, the target HDR tensor... Gradient propagation is stopped during this phase, meaning the target HDR frame is used only as a supervision target and does not participate in gradient calculation. Simultaneously, only pixel-level supervision is used in this stage; the quality evaluation backbone is not trained, and only the adaptation parameters of the residual domain adapter are involved in parameter updates. This ensures that the optimization objective in the first training phase is singular and the process is stable, resulting in an adapter with a clear adaptation direction from SDR to HDR.

[0026] During training, this embodiment iteratively updates the adaptation parameters of the residual domain adapter according to the preset input size, frame sampling number, training batch, optimization algorithm, and learning rate until the first training termination condition is met. The first training termination condition can be an early stopping condition, such as reaching a preset number of training rounds, the loss decrease being less than a preset convergence threshold, or the validation metric no longer improving within a preset number of rounds, or a combination of the above conditions. This embodiment does not limit this. Further, during training, a validation can be performed once after each training round, that is, the pixel-level consistency index is calculated on the first preset validation set using the current residual domain adapter, and the adaptation parameters when the pixel-level consistency index is optimal are saved. The pixel-level consistency index can be the pixel mean square error on the validation set, where the optimal index is when the error value is minimized; it can also be an equivalent index such as peak signal-to-noise ratio, where the optimal index is when the index value is maximized. Finally, the adaptation parameters corresponding to the optimal pixel-level consistency index are determined as the target adaptation parameters, and the residual domain adapter with the target adaptation parameters is determined as the target residual domain adapter for initialization in the second training phase.

[0027] Step S12: Obtain target video data; determine the dynamic range domain identifier corresponding to the target video data based on the data source corresponding to the target video data; construct a second training set based on the target video data; and construct an initial quality assessment model using conditional routing, identity bypass, quality assessment backbone, and the target residual domain adapter. The target video data includes SDR video data and HDR video data, and the target video data carries a corresponding subjective quality score. The dynamic range domain identifier includes an SDR domain identifier and an HDR domain identifier. The conditional routing is used to determine the forward processing path from the target video data input to the quality assessment backbone based on the dynamic range domain identifier. When the identity bypass is used as the forward processing path, the output data of the forward processing path is the same as the input data. The quality assessment backbone is a no-reference quality assessment network used to obtain the video quality score corresponding to the video data.

[0028] In this embodiment, the target video data is cross-dynamic range quality video data with subjective quality scores, which is composed of SDR source domain quality video data and HDR target domain quality video data. Since video data from the same data source has the same dynamic range attribute, the dynamic range domain identifier can be directly determined based on the data source. For example, the domain identifier of video data from the SDR quality dataset can be set to the SDR domain identifier, and the domain identifier of video data from the HDR quality dataset can be set to the HDR domain identifier. The method of determining the domain identifier based on the data source does not require additional dynamic range discrimination of the video content, which is simple to implement and the result is certain, without introducing discrimination errors. The target video data includes SDR video data and HDR video data, and the target video data carries the corresponding subjective quality scores; the dynamic range domain identifier includes the SDR domain identifier and the HDR domain identifier; conditional routing is used to determine the forward processing path from the target video data input to the quality evaluation backbone based on the dynamic range domain identifier; when the identity bypass is used as the forward processing path, the output data of the forward processing path is the same as the input data; the quality evaluation backbone is a no-reference quality evaluation network used to obtain the video quality score corresponding to the video data.

[0029] It should be noted that the dynamic range domain identifier in this embodiment is only used to select the forward processing path of the sample, and is not used as a quality regression feature input to the quality assessment backbone, nor does it participate in quality-perceived feature extraction and quality score regression. This setting can prevent the model from using the dynamic range category itself as a shortcut clue to predict the quality score, thereby ensuring that the predicted score is determined only by the observable degree of distortion in the video. The quality assessment backbone is a shared backbone, that is, SDR samples and HDR samples use the same set of quality assessment backbone parameters. This is a prerequisite for SDR supervision information to be transferred to the HDR quality assessment task. If independent backbones are set for the two types of samples, there will be no parameter sharing between the two types of data, and the content prior and general compression distortion prior in the SDR data will not be able to be applied to HDR quality prediction.

[0030] In one specific implementation, the process of acquiring target video data, determining the dynamic range domain identifier corresponding to the target video data based on the data source corresponding to the target video data, and constructing a second training set based on the target video data can be implemented as follows: First, SDR and HDR video data with subjective quality scores are acquired. SDR video data, as the source domain quality data, is typically large in scale and covers a wide range of content types, compression distortion types, and quality levels. HDR video data, as the target domain quality data, is typically smaller in scale but directly corresponds to the final evaluation task.

[0031] Then, a dynamic range domain identifier d is generated for each target video data based on the data source. The domain identifier includes at least an SDR source domain identifier and an HDR target domain identifier; for example, the SDR source domain identifier can be encoded as 0, and the HDR target domain identifier can be encoded as 1. As mentioned earlier, the domain identifier is only used to select the forward processing path and is not used as input for quality regression features.

[0032] Next, non-zero sampling weights are set for the SDR source domain quality data and the HDR target domain quality data, ensuring that both types of data can be sampled during training. The sampling weights can be configured based on the data size, training stage, or target domain performance. For example, a smaller sampling weight can be set for the SDR source domain data, and a larger sampling weight for the HDR target domain data, to ensure that the optimization process is primarily driven by target domain performance. Alternatively, a larger SDR sampling weight can be set in the early stages of training to fully utilize the diversity of content and distortion in the source domain, and the HDR sampling weight can be gradually increased as training progresses. This embodiment of the invention does not limit this approach.

[0033] Finally, a second training set is constructed, where each sample includes a normalized frame tensor X, a subjective quality score y, and a dynamic range domain identifier d. The construction process of the normalized frame tensor may include operations such as decoding, encoding value normalization, spatial size normalization, and temporal sampling. During the training phase, random frame sampling can be used to increase the content coverage, while deterministic window sampling is used during the validation phase to ensure the repeatability of the evaluation results.

[0034] In one optional implementation, cross-dynamic range mixed training can employ a homogeneous batch strategy, ensuring that all samples in a training batch originate from the same data source. It's important to note that the reason for using a homogeneous batch strategy is that subjective quality scores from different datasets often use different units and value ranges, and the input statistical characteristics of data from different dynamic ranges also differ. If samples from different sources are mixed within the same batch, the intra-batch statistics will be affected by both the differences in component units and the differences in input statistics, thus interfering with the correlation constraint terms calculated based on intra-batch statistics. With a homogeneous batch strategy, the component units within each batch remain consistent with the input statistics, reducing the aforementioned optimization interference.

[0035] Step S13: Using the conditional routing, when the dynamic range domain identifier corresponding to the target video data in the second training set is an SDR domain identifier, the target video data is adapted through the target residual domain adapter to obtain adapted samples. The adapted samples are then input into the quality evaluation backbone to obtain a first predicted quality score. When the dynamic range domain identifier corresponding to the target video data in the second training set is an HDR domain identifier, the target video data is input into the quality evaluation backbone through the identity bypass to obtain a second predicted quality score.

[0036] In this embodiment, the process of constructing the initial quality assessment model using conditional routing, identity bypass, the quality assessment backbone, and the target residual domain adapter can be implemented as follows: The initial quality assessment model includes conditional routing, a target residual domain adapter, an identity bypass, and a quality assessment backbone. The conditional routing takes the dynamic range domain identifier of the sample as input and outputs the forward processing path used by that sample. The target residual domain adapter and the identity bypass are connected in parallel before the quality assessment backbone, forming the SDR branch and HDR branch, respectively. The quality assessment backbone is a no-reference quality assessment network shared by the two branches.

[0037] For the i-th training sample, let its dynamic range domain be denoted as d_i, where d_i = d_SDR represents an SDR sample and d_i = d_HDR represents an HDR sample. Then the input x'_i of the quality assessment backbone can be expressed as: x'_i = A(x_i), when d_i = d_SDR; x'_i = x_i, when d_i = d_HDR.

[0038] Where A(·) represents the target residual domain adapter. As can be seen from the above formula, SDR source domain samples enter the shared backbone after being processed by the domain adapter, while HDR target domain samples retain their original encoded representation and enter the shared backbone through the identity bypass. The two form an asymmetric processing path, but share the same set of backbone parameters.

[0039] In this embodiment, the quality assessment backbone may include a quality-aware feature extraction network, a feature aggregation module, and a quality mapping head. The quality-aware feature extraction network extracts hierarchical features related to compression distortion from the input frame. It can be implemented using a convolutional neural network or a network based on a self-attention mechanism, such as the EfficientNetV2 series, ResNet series, or Swin Transformer series. This embodiment does not limit the specific type of feature extraction network. The feature aggregation module converts the hierarchical features into frame-level quality-aware vectors of a preset dimension. This can include operations such as global pooling, multi-level feature concatenation, and fully connected mapping. The quality mapping head maps the frame-level quality-aware vectors to scalar frame-level quality responses, which can be implemented by one or more fully connected layers. Temporal aggregation of the frame-level quality responses of the same video yields the video-level prediction score. In an optional implementation, the quality assessment backbone may also employ a multi-view approach, constructing multiple views with different scale factors for the same input frame, extracting quality-aware features from each view, and then fusing them. Because compression distortion is perceptible at different scales—for example, blockiness and color banding are more noticeable in larger-scale views, while blurring and loss of detail are more perceptible in original-scale views—multi-view implementations help to capture distortion cues at different granularities simultaneously.

[0040] It is understandable that this embodiment employs an asymmetric forward processing procedure: for SDR source domain samples, their encoded representations have a distributional offset from the HDR target domain, thus requiring adaptation processing via a target residual domain adapter to map them to an encoded representation approximately consistent with the HDR target domain before inputting them into the shared quality assessment backbone; for HDR target domain samples, their encoded representations are themselves the target domain representation for the quality assessment task, requiring no numerical transformation, and are therefore directly input into the shared quality assessment backbone via an identity bypass. The reason for using an identity bypass for HDR samples instead of applying the same transformation is that HDR samples are the target domain data for the final evaluation task, and the highlight details, shadow levels, and compression distortion cues in their original encoded representations are the direct basis for quality prediction; applying additional transformations may introduce unnecessary numerical errors and information loss, thus weakening the quality cues. Simultaneously, the existence of the identity bypass ensures that the processing flow in the quality assessment stage is completely consistent with the HDR branch in the training stage, guaranteeing consistency between training and inference.

[0041] In this embodiment, the target video data is adapted using a target residual domain adapter to obtain adapted samples. This includes: performing boundary protection processing on the normalized encoded values ​​of the target video data to obtain the encoded values ​​to be transformed within a preset open interval; performing a log-probability transformation on the encoded values ​​to be transformed to map them to the target transform domain to obtain the target mapping result; the target transform domain is a log-probability transformation domain; in the target transform domain, a residual correction amount is applied to the target mapping result using the target residual domain adapter to obtain the target correction result; and performing an inverse log-probability transformation on the target correction result to obtain adapted samples within the preset open interval.

[0042] Specifically, let the normalized SDR encoded frame be x, whose value lies within the closed interval [0,1]. Since the subsequent log-probability transformation diverges when the independent variable takes the value 0 or 1, boundary protection processing needs to be performed first to restrict the encoded value to the open interval (0,1), thus obtaining the encoded frame to be transformed. As an optional implementation, boundary protection can be achieved using truncation operations, i.e.: ; in, The preset boundary protection parameter is greater than zero and much less than 1, for example, it can be taken as... Besides truncation, boundary protection can also be implemented using linear compression, that is, linearly mapping the [0,1] interval to... The range is not limited in this embodiment.

[0043] Perform a log-odds transform on the frame x_b to be transformed, and obtain the target mapping result z located in the log-odds transform domain: ; It should be noted that the log-odds transform is a strictly monotonically increasing function defined on the open interval (0,1), with a range covering the entire real number domain. After this transform, the encoded values, which were originally limited to a bounded interval, are mapped to an unbounded transform domain. This means that subsequent additive correction is no longer restricted by the range of values, and there is no need to impose artificial constraints on the magnitude of the correction to ensure the validity of the output.

[0044] In this embodiment, the residual domain adapter includes an input mapping layer, at least one residual feature transformation unit, and an output mapping layer. The input mapping layer maps the input encoded frames to a latent feature space of a preset dimension. The residual feature transformation unit includes convolution operations, nonlinear activation, and residual connections, used to extract local and contextual features related to brightness distribution correction in the latent feature space. The output mapping layer outputs a residual correction amount consistent with the number of input channels. ,Right now: ; Where G(·) represents the mapping function consisting of the input mapping layer, the residual feature transformation unit, and the output mapping layer. The target correction result is... .

[0045] It is understandable that the output mapping layer can be zero-initialized, that is, the weights and biases of the output mapping layer are initialized to zero, so that the following conditions are met at the beginning of training. =0, thus ensuring that the adaptation result satisfies A(x)≈x. Therefore, the residual domain adapter is optimized starting from an approximate identity mapping, which can avoid drastic changes in the output in the early stage of training and improve the stability of the optimization process.

[0046] In one specific implementation, the inverse transformation of the logarithmic odds transformation is the Sigmoid function. (·), therefore the adapted sample can be represented as: ; Since the range of the Sigmoid function is an open interval (0,1), the adapted sample will necessarily be within the valid encoding range, and no additional truncation processing is required to ensure that the output encoded value is valid.

[0047] Furthermore, by performing an equivalent transformation on the above equation, we can obtain: ; Therefore, additive correction in the log-probability transform domain is equivalent to performing a multiplicative adjustment on the probability ratio of the encoded values. This property brings two technical benefits: First, the same correction amount Δ applied to different input encoded values ​​results in different absolute changes in the encoded values, meaning the mapping itself is a non-linear mapping that varies with the input encoded values, enabling adjustments of different magnitudes to be applied to dark and bright areas, thus adapting to the statistical differences in brightness and contrast between SDR and HDR; Second, regardless of the value of the correction amount Δ, the output is always constrained within the effective encoding range, thereby avoiding out-of-bounds values ​​caused by performing unconstrained addition directly in the pixel domain.

[0048] It should be noted that the above adaptation processing can be performed separately for each channel of the frame tensor, or some parameters can be shared between the channels; it can be performed on the RGB three channels, or it can be performed only on the luminance component while keeping the chrominance component constant. This embodiment does not limit this.

[0049] Step S14: Based on the predicted quality score and subjective quality score corresponding to the target video data, train the quality evaluation backbone and the target residual domain adapter in the initial quality evaluation model to obtain the target quality evaluation model, and use the target quality evaluation model to obtain the target video quality score corresponding to the HDR video to be evaluated; the predicted quality score includes a first predicted quality score and a second predicted quality score.

[0050] In this embodiment, the cross-domain quality joint training phase can also be referred to as the second training phase. In this phase, the target residual domain adapter remains trainable, its parameters are initialized from the pre-trained parameters obtained in the first training phase, and continue to be updated along with the quality regression target. Thus, the mapping implemented by the domain adapter is gradually adjusted from the pixel-level mapping of the first training phase to a task-related mapping that is beneficial to subjective quality prediction, i.e., achieving domain adaptation oriented towards quality tasks. It can be understood that after obtaining the target quality assessment model, this model can be used to perform referenceless quality assessment on the HDR video to be evaluated. At this time, the input sample already belongs to the HDR target domain, and its domain identifier is set to the HDR domain identifier. Therefore, the domain adapter does not participate in numerical transformation, and the HDR video to be evaluated directly enters the quality assessment backbone via the identity bypass, without requiring an undistorted reference video, a paired SDR video, or additional domain transformation processing.

[0051] In one specific implementation, the quality assessment backbone and target residual domain adapter in the initial quality assessment model are trained based on the predicted quality score and subjective quality score corresponding to the target video data to obtain the target quality assessment model. Specifically, this may include: determining the target relevance constraint and target ranking constraint corresponding to the second training set based on the predicted quality score and subjective quality score corresponding to each target video data in the second training set; determining the joint training loss corresponding to the second training set based on the target relevance constraint and target ranking constraint; and performing backpropagation based on the joint training loss to simultaneously update the target residual domain adapter and the quality assessment backbone to obtain the target quality assessment model. The process of performing backpropagation based on joint training loss to simultaneously update the target residual domain adapter and the quality evaluation backbone to obtain the target quality evaluation model includes: performing backpropagation based on joint training loss to simultaneously update the target residual domain adapter and the quality evaluation backbone to obtain the updated initial quality evaluation model; determining whether the updated initial quality evaluation model meets the second training termination condition; when the updated initial quality evaluation model meets the second training termination condition, performing deterministic sampling on the second preset validation set to obtain validation sampling frames; using the updated initial quality evaluation model to obtain the quality prediction performance index corresponding to the validation sampling frames; determining the model parameters of the initial quality evaluation model corresponding to the optimal quality prediction performance index as the target model parameters; and determining the initial quality evaluation model with the target model parameters as the target quality evaluation model.

[0052] Specifically, let the predicted quality score of the i-th sample in a certain training batch be . Subjective quality score is Batch size is B. Target relevance constraints. This constraint is used to constrain the correlation consistency between predicted scores and subjective scores. As a preferred implementation, the correlation constraint term can employ Pearson linear correlation coefficient loss, the expression of which is: ; in, and These represent the mean of the predicted score and the mean of the subjective score within the current batch, respectively. This is a very small positive number used to ensure numerical stability. It should be noted that the Pearson linear correlation coefficient loss is insensitive to linear shifts and scaling of predicted scores. Therefore, it can still provide effective supervision signals when the dimensions of subjective scores differ across datasets, making it particularly suitable for mixed training scenarios across datasets and dynamic ranges.

[0053] Furthermore, the target ranking constraint term This constraint is used to constrain the relative quality order among samples, ensuring that the model correctly distinguishes between higher-quality and lower-quality samples in the overall distribution. As a preferred implementation, the ranking constraint... Pairwise interval sorting loss can be used, and its expression can be: ; Where sign(·) represents the sign function. As shown in the above equation, when the model's prediction order for a sample pair is consistent with the subjective order and the difference is sufficient, the sample pair does not incur loss; otherwise, it incurs positive loss and drives model adjustment. Besides pairwise interval ranking loss, the ranking constraint term can also adopt fidelity loss or other loss forms that can characterize relative order; this embodiment does not limit this.

[0054] Joint training losses It can be a weighted sum of the relevance constraint and the ranking constraint, and its expression can be: ; in, and The weights, which are greater than zero, are used to adjust the relative contributions of the two types of constraints during the optimization process. It should be noted that the correlation constraint primarily affects the consistency of the predicted score and the subjective score in the overall trend, while the ranking constraint primarily affects the relative order between sample pairs. The two complement each other: using only the correlation constraint may result in insufficient ability to distinguish between locally adjacent quality levels; using only the ranking constraint may cause the model's predicted distribution to deviate from the overall shape of the subjective distribution. Therefore, using a weighted combination of the two can achieve more stable quality prediction performance.

[0055] During model iteration, this embodiment first loads the target adaptation parameters obtained in the first training phase to initialize the residual domain adapter and keeps the residual domain adapter in a trainable state; then, it performs forward computation according to conditional routing, inputting the adaptation results of SDR samples and the original encoded representation of HDR samples into the shared quality assessment backbone to obtain the predicted quality score q; finally, it performs forward computation based on the joint training loss. Backpropagation is performed, simultaneously updating the parameters of the residual domain adapter and the quality assessment backbone. It's important to note that when processing SDR samples, the gradient is backpropagated to the residual domain adapter via the shared quality assessment backbone, further adjusting the adapter from pixel-level mapping initialization to a task-relevant adaptation serving HDR quality prediction. When processing HDR samples, since this branch is an identity bypass, the gradient only applies to the shared quality assessment backbone. Thus, both types of samples participate in the backbone optimization, while the adapter's optimization direction is driven by the quality prediction objective.

[0056] In this embodiment, iterative updates can be performed according to a preset training batch, optimization algorithm, learning rate, and learning rate adjustment strategy. The optimization algorithm can be the Adam optimizer, AdamW optimizer, or a stochastic gradient descent optimizer with momentum; this embodiment is not limited to this. The second training termination condition can be reaching a preset number of training rounds, the decrease in joint training loss being less than a preset threshold, or an early stopping condition where the validation metric no longer improves within a preset number of rounds, or a combination of the above conditions. Further, in this embodiment, deterministic sampling is performed on the second preset validation set during the validation phase, that is, validation sampling frames are extracted from the video to be validated according to fixed sampling rules, such as sampling frames at fixed intervals or fixed window positions. The reason for using deterministic sampling is that if random frame sampling is still used in the validation phase, the results of the same model in different validation rounds will fluctuate due to sampling differences, which cannot accurately reflect the quality of the model parameters themselves; after using deterministic sampling, the validation rounds are comparable, and the parameter selection process is more reliable.

[0057] Furthermore, the quality prediction performance index can be one or more combinations of Spearman's rank correlation coefficient, Pearson's linear correlation coefficient, Kendall's rank correlation coefficient, and root mean square error. Spearman's rank correlation coefficient measures the monotonic consistency between the predicted score and the subjective score, while Pearson's linear correlation coefficient measures the linear consistency between the two. In this embodiment of the invention, a weighted sum of the above indices or the primary indices can be used as the selection criterion, and the model parameters at which the quality prediction performance index is optimal can be saved as the target model parameters. It should be noted that the second preset validation set preferably uses independent HDR compressed video data to ensure that the selected model parameters have optimal performance in the HDR target domain.

[0058] It should be noted that when processing high-resolution video data, such as video resolutions reaching or exceeding ultra-high definition resolution, the residual domain adapter needs to perform convolution operations at the same spatial resolution as the input frames. The intermediate activations consume significant GPU memory, potentially leading to insufficient GPU memory during training. To address this, this embodiment of the invention introduces a preset resource optimization strategy. Specifically, if the video resolution of the target video data in the second training set is greater than a preset resolution threshold, the target residual domain adapter and the quality evaluation backbone are updated simultaneously based on the preset resource optimization strategy. The preset resource optimization strategy includes a first optimization strategy based on micro-batch forward pass, a second optimization strategy based on gradient checkpoints, and a third optimization strategy based on a combination of micro-batch forward pass and gradient checkpoints.

[0059] The first optimization strategy based on micro-batch forward propagation involves further dividing a training batch into several micro-batches, performing forward and backward computations on each micro-batch, accumulating the gradients generated by each micro-batch, and then updating the parameters uniformly after all micro-batches have been processed. Since only the intermediate activations of a single micro-batch need to be saved at any given time, peak memory usage is significantly reduced. The second optimization strategy based on gradient checkpointing involves not saving all intermediate activations during forward computation, but only saving activations at a few checkpoints. During backward propagation, the required intermediate activations are recalculated from the checkpoints. This strategy trades a small amount of repetitive computation for a reduction in memory usage. The third optimization strategy, a combination of micro-batch forward propagation and gradient checkpointing, involves simultaneously employing both strategies to further reduce memory usage. It should also be noted that the above resource optimization strategies only change the organization of computation and the method of saving intermediate results; they do not change the computational relationship between domain adaptation and quality evaluation, nor do they change the mathematical equivalence of the model or the final training results. In this embodiment of the invention, the three optimization strategies can be selected based on the video resolution and hardware conditions. For example, the third optimization strategy can be enabled when the video resolution is greater than a preset resolution threshold, and the resource optimization strategy can be disabled when the video resolution is not greater than the preset resolution threshold.

[0060] In this embodiment, the target video quality score corresponding to the HDR video to be evaluated is obtained using a target quality assessment model, including: obtaining the HDR video to be evaluated; preprocessing the HDR video to be evaluated to obtain the corresponding HDR frame sequence; inputting the HDR frame sequence into the quality assessment backbone in the target quality assessment model through an identity bypass, so as to obtain the quality-perceived features corresponding to each sampled frame in the HDR frame sequence using the quality assessment backbone, and determining the frame-level quality response corresponding to each sampled frame based on the quality-perceived features; and performing temporal aggregation on each frame-level quality response of the HDR frame sequence to obtain the target video quality score corresponding to the HDR video to be evaluated.

[0061] In one specific implementation, such as Figure 3As shown, the quality evaluation stage only processes the HDR compressed video to be evaluated, without requiring paired SDR videos, undistorted reference videos, or subjective quality labels. The preprocessing process may include: decoding the input HDR video to be evaluated to obtain decoded frames; normalizing the encoded values ​​of the decoded frames to a preset bounded interval; performing spatial regularization on the decoded frames according to a preset spatial size; and extracting sampled frames from the video according to a preset temporal sampling rule to form an HDR frame sequence. The temporal sampling in the quality evaluation stage preferably adopts a deterministic sampling rule, such as uniformly extracting F frames from the video at fixed intervals, or extracting sampled frames according to a fixed window position. Using deterministic sampling ensures that multiple evaluation results for the same video are completely consistent, thus making the evaluation results repeatable and meeting the stability requirements of application scenarios such as encoding parameter optimization and quality monitoring. Since the input samples already belong to the HDR target domain, their dynamic range domain identifier is set to the HDR domain identifier. Therefore, the conditional routing will select the identity bypass as the forward processing path, the residual domain adapter does not participate in numerical transformation, and the HDR frame sequence retains its original encoded representation and is directly input into the trained shared quality evaluation backbone.

[0062] Furthermore, the shared quality assessment backbone extracts quality-aware features from each sampled frame and obtains the frame-level quality response of that frame via a quality mapping head. When using a multi-view implementation, the backbone further fuses the quality-perceived features corresponding to each view of the same sampled frame, and then outputs the frame-level quality response from the quality mapping head. For a video to be evaluated containing F sampled frames, a preset temporal aggregation is performed on its frame-level quality response to obtain a video-level score Q. As a preferred implementation, the temporal aggregation can use arithmetic mean aggregation, and its expression can be: ; Besides arithmetic mean aggregation, temporal aggregation can also employ weighted average aggregation, percentile aggregation, worst-frame pooling aggregation, or weighted aggregation considering temporal memory effects, etc., and this embodiment does not limit these methods. It should be noted that the video-level score obtained in this embodiment is the unreferenced objective quality evaluation result of the HDR compressed video to be evaluated. It can be used for quality comparison between different encoding methods, different quantization parameters, or different bitrate configurations, thereby providing an objective basis for encoding parameter optimization, bitrate-quality control, transmission link quality monitoring, and codec scheme evaluation.

[0063] In one specific embodiment of domain adapter pre-training, paired data consists of content-corresponding SDR and HDR frames, which are normalized and input into a residual domain adapter based on log-odds transform. The input uses... Perform boundary protection, where boundary protection parameters Pick The residual domain adapter employs one 3×3 ingress convolutional layer, two residual convolutional blocks, and one 3×3 egress convolutional layer. The hidden layer has 64 channels, and the weights and biases of the egress convolutional layer are initialized to zero, allowing the adapter to begin optimization from an approximate identity mapping. The paired frames are regularized to 512×512 pixels, with 20 frames read per video and a training batch size of 2. The Adam optimizer is used, with a learning rate of... The system was trained for 150 rounds using the pixel mean square error as the loss function. Validation was performed once in each round, and the domain adapter parameters at which the validation pixel loss was minimized were saved as the target adaptation parameters.

[0064] In a specific embodiment of cross-domain quality joint training, the mixed quality training set includes SDR source domain data and HDR target domain data, with sampling weights of 0.2 and 0.8, respectively; the SDR source domain identifier is encoded as 0, and the HDR target domain identifier is encoded as 1. Training employs a same-source batch strategy, generating a complete batch from one data source each time. The shared quality evaluation backbone uses EfficientNetV2-S as the quality-aware feature extraction network, with the dimension of the frame-level quality-aware vector set to 128, and outputting a scalar frame-level quality response through a quality mapping head. The joint loss is the sum of Pearson linear correlation coefficient loss and ranking loss, where the correlation loss weight is... The sorting loss weight is 1. The training batch size is 5. The Adam optimizer is used, and the learning rate is 5. The training process consisted of 45 epochs. For the domain adapter, a micro-batch forward pass of size 4 was used with gradient checkpointing enabled to reduce memory overhead under high-resolution input. During the validation phase, independent HDR compressed bitstream data was used, and the final model parameters were selected based on the Spearman rank correlation coefficient and the Pearson linear correlation coefficient.

[0065] In a specific embodiment of HDR quality evaluation, the encoded values ​​of the HDR compressed video to be evaluated are normalized to the [0,1] interval, a frame sequence is formed according to deterministic sampling rules, and the dynamic range domain identifier is set to 1. The HDR frames enter the shared multi-scale quality evaluation backbone through an identity bypass, and views with scale factors of 1 and 4 are constructed respectively to obtain the quality response of each frame; the arithmetic mean of the quality responses of all sampled frames is taken as the video-level no-reference quality score of the video.

[0066] Based on the above, the HDR video quality evaluation method provided in this embodiment can support at least the following scenarios in practical applications: In the encoding parameter optimization scenario, quality scores can be output for reconstructed videos of the same content under different quantization parameters, thereby selecting the encoding configuration with the best quality under a given bitrate constraint; In the bitrate-quality control scenario, the quality score can be used as a feedback quantity to participate in the bitrate allocation decision, realizing adaptive encoding oriented towards perceived quality; In the transmission link quality monitoring scenario, since the evaluation process does not require a reference video, the quality of the received HDR bitstream can be directly monitored at the distribution node or terminal side; In the encoding and decoding scheme evaluation scenario, the quality of HDR compressed videos generated by different encoders and different combinations of encoding tools can be compared on a unified scale.

[0067] As can be seen, the HDR video quality assessment method provided in this application achieves SDR to HDR coding domain adaptation by learning additive residual correction in the log-odds transform domain, establishes an asymmetric processing path of "SDR adaptation, HDR bypass, and backbone sharing" through dynamic range conditional routing, and simultaneously updates the domain adapter and quality assessment backbone through joint optimization of correlation constraints and ranking constraints. Thus, under the condition of limited HDR subjective quality data scale, it effectively utilizes the supervision information of SDR quality data, alleviates the distribution offset between the source domain and the target domain, improves the accuracy and generalization ability of HDR compressed video no-reference quality assessment, and still only requires the input HDR video to be evaluated in the quality assessment stage, maintaining no-reference attributes and a unified deployment form.

[0068] See Figure 4 As shown, this application discloses an HDR video quality evaluation device, comprising: The adapter training module 11 is used to acquire content-paired SDR frame data and HDR frame data, construct a first training set based on the SDR frame data and the HDR frame data, use the SDR frame data in the first training set as input data, and use the HDR frame data corresponding to the SDR frame data as the supervision target to train a pre-constructed residual domain adapter to obtain a target residual domain adapter; wherein, the residual domain adapter is a domain adapter based on log-odds transformation. The evaluation model construction module 12 is used to acquire target video data, determine the dynamic range domain identifier corresponding to the target video data based on the data source corresponding to the target video data, construct a second training set based on the target video data, and construct an initial quality evaluation model using conditional routing, identity bypass, quality evaluation backbone, and the target residual domain adapter. The target video data includes SDR video data and HDR video data, and the target video data carries a corresponding subjective quality score. The dynamic range domain identifier includes an SDR domain identifier and an HDR domain identifier. The conditional routing is used to determine the forward processing path from the target video data input to the quality evaluation backbone based on the dynamic range domain identifier. When the identity bypass is used as the forward processing path, the output data of the forward processing path is the same as the input data. The quality evaluation backbone is a no-reference quality evaluation network used to acquire the video quality score corresponding to the video data. The quality score prediction module 13 is used to utilize the conditional routing to adapt the target video data through the target residual domain adapter to obtain adapted samples when the dynamic range domain identifier corresponding to the target video data in the second training set is an SDR domain identifier, and input the adapted samples into the quality evaluation backbone to obtain a first predicted quality score; and when the dynamic range domain identifier corresponding to the target video data in the second training set is an HDR domain identifier, the target video data is input into the quality evaluation backbone through the identity bypass to obtain a second predicted quality score. The video evaluation module 14 is used to train the quality evaluation backbone and the target residual domain adapter in the initial quality evaluation model based on the predicted quality score and subjective quality score corresponding to the target video data to obtain a target quality evaluation model, and to use the target quality evaluation model to obtain the target video quality score corresponding to the HDR video to be evaluated; the predicted quality score includes a first predicted quality score and a second predicted quality score.

[0069] As can be seen, this application expands the range of available training data without changing the no-reference attribute in the inference stage by combining three stages: paired domain transfer pre-training, dynamic range conditional routing, and cross-domain quality joint optimization. It also reduces the feature distribution shift caused by cross-dynamic range mixed training, thereby improving the accuracy and generalization ability of HDR compressed video quality evaluation.

[0070] In one specific implementation, the video evaluation module 14 may include: The video preprocessing submodule is used to acquire the HDR video to be evaluated and preprocess the HDR video to be evaluated to obtain the corresponding HDR frame sequence. The response acquisition submodule is used to input the HDR frame sequence into the quality assessment backbone in the target quality assessment model through the identity bypass, so as to use the quality assessment backbone to obtain the quality perception features corresponding to each sampled frame in the HDR frame sequence, and determine the frame-level quality response corresponding to each sampled frame based on the quality perception features. The quality score acquisition submodule is used to perform temporal aggregation on the frame-level quality responses of the HDR frame sequence to obtain the target video quality score corresponding to the HDR video to be evaluated.

[0071] In one specific implementation, the adapter training module 11 may include: The first training set acquisition submodule is used to establish the content correspondence between the SDR frame data and the HDR frame data based on the video identifier and frame sequence number, normalize the encoding value of the SDR frame data to a preset bounded interval to obtain the normalized SDR encoding value, normalize the encoding value of the HDR frame data to a preset bounded interval to obtain the normalized HDR encoding value, and perform spatial size alignment on the normalized SDR encoding value and the normalized HDR encoding value to obtain the first training set. The adaptation result acquisition submodule is used to input the SDR frames in the first training set into the pre-constructed residual domain adapter to obtain the adaptation results corresponding to the SDR frames. The adapter update submodule is used to determine the HDR frame corresponding to the SDR frame based on the content correspondence, and update the adaptation parameters of the residual domain adapter based on the pixel-level consistency loss between the adaptation result and the HDR frame to obtain the updated residual domain adapter. The adapter determination submodule is used to determine whether the updated residual domain adapter meets the first training termination condition. When the updated residual domain adapter meets the first training termination condition, the pixel-level consistency index corresponding to the first preset validation set is obtained using the updated residual domain adapter. The adaptation parameter corresponding to the residual domain adapter when the pixel-level consistency index is optimal is determined as the target adaptation parameter, and the residual domain adapter with the target adaptation parameter is determined as the target residual domain adapter.

[0072] In one specific implementation, the quality score prediction module 13 may include: The submodule for obtaining the coded value to be transformed is used to perform boundary protection processing on the coded value after normalization of the target video data to obtain the coded value to be transformed located within a preset open interval. The log-odds transformation submodule is used to perform a log-odds transformation on the encoded value to be transformed, mapping the encoded value to be transformed to a target transformation domain to obtain a target mapping result; the target transformation domain is the log-odds transformation domain; The residual correction submodule is used to apply a residual correction amount to the target mapping result in the target transform domain using the target residual domain adapter to obtain the target correction result. The sample adaptation submodule is used to perform the inverse transformation of the log-probability transformation on the target correction result to obtain the adapted sample located within the preset open interval.

[0073] In one specific implementation, the video evaluation module 14 may include: The constraint determination submodule is used to determine the target relevance constraint and target ranking constraint corresponding to the second training set based on the predicted quality score and subjective quality score corresponding to each target video data in the second training set; The training loss determination submodule is used to determine the joint training loss corresponding to the second training set based on the target relevance constraint and the target ranking constraint. The model update submodule is used to perform backpropagation based on the joint training loss to simultaneously update the target residual domain adapter and the quality assessment backbone to obtain the target quality assessment model.

[0074] In one specific implementation, the model update submodule may include: The model update unit is used to perform backpropagation based on the joint training loss to simultaneously update the target residual domain adapter and the quality assessment backbone to obtain the updated initial quality assessment model. The model determination unit is used to determine whether the updated initial quality evaluation model meets the second training termination condition. When the updated initial quality evaluation model meets the second training termination condition, it performs deterministic sampling on the second preset validation set to obtain validation sampling frames, uses the updated initial quality evaluation model to obtain the quality prediction performance index corresponding to the validation sampling frames, determines the model parameters of the initial quality evaluation model corresponding to the optimal quality prediction performance index as the target model parameters, and determines the initial quality evaluation model with the target model parameters as the target quality evaluation model.

[0075] In one specific embodiment, the device may further include: The optimization and update module is used to simultaneously update the target residual domain adapter and the quality evaluation backbone based on a preset resource optimization strategy if the video resolution of the target video data in the second training set is greater than a preset resolution threshold. The preset resource optimization strategy includes a first optimization strategy based on micro-batch forward pass, a second optimization strategy based on gradient checkpoints, and a third optimization strategy based on a combination of micro-batch forward pass and gradient checkpoints.

[0076] Furthermore, embodiments of this application also disclose an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0077] Figure 5 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the HDR video quality evaluation method disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0078] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0079] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk, or optical disk, etc. The resources stored thereon can include an operating system 221, computer programs 222, etc., and the storage method can be temporary storage or permanent storage.

[0080] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the HDR video quality evaluation method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0081] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed HDR video quality evaluation method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0082] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0083] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0084] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0085] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0086] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for evaluating the quality of HDR video, characterized in that, include: Acquire content-paired SDR frame data and HDR frame data, construct a first training set based on the SDR frame data and the HDR frame data, use the SDR frame data in the first training set as input data, and use the HDR frame data corresponding to the SDR frame data as the supervision target to train a pre-constructed residual domain adapter to obtain a target residual domain adapter; wherein, the residual domain adapter is a domain adapter based on log-odds transformation. The process involves acquiring target video data, determining the dynamic range domain identifier corresponding to the target video data based on its data source, constructing a second training set based on the target video data, and building an initial quality assessment model using conditional routing, identity bypass, a quality assessment backbone, and the target residual domain adapter. The target video data includes SDR and HDR video data, and carries corresponding subjective quality scores. The dynamic range domain identifier includes SDR and HDR domain identifiers. Conditional routing is used to determine the forward processing path from the target video data input to the quality assessment backbone based on the dynamic range domain identifier. When the identity bypass is used as the forward processing path, the output data of the forward processing path is the same as the input data. The quality assessment backbone is a no-reference quality assessment network used to obtain the video quality score corresponding to the video data. Using the conditional routing, when the dynamic range domain identifier corresponding to the target video data in the second training set is an SDR domain identifier, the target video data is adapted through the target residual domain adapter to obtain adapted samples. The adapted samples are then input into the quality evaluation backbone to obtain a first predicted quality score. When the dynamic range domain identifier corresponding to the target video data in the second training set is an HDR domain identifier, the target video data is input into the quality evaluation backbone through the identity bypass to obtain a second predicted quality score. The quality assessment backbone and the target residual domain adapter in the initial quality assessment model are trained based on the predicted quality score and subjective quality score corresponding to the target video data to obtain the target quality assessment model, and the target video quality score corresponding to the HDR video to be evaluated is obtained using the target quality assessment model; the predicted quality score includes a first predicted quality score and a second predicted quality score.

2. The HDR video quality evaluation method according to claim 1, characterized in that, The step of obtaining the target video quality score corresponding to the HDR video to be evaluated using the target quality evaluation model includes: Obtain the HDR video to be evaluated, and preprocess the HDR video to be evaluated to obtain the corresponding HDR frame sequence; The HDR frame sequence is input into the quality assessment backbone of the target quality assessment model through the identity bypass, so as to obtain the quality-perceived features corresponding to each sampled frame in the HDR frame sequence using the quality assessment backbone, and determine the frame-level quality response corresponding to each sampled frame based on the quality-perceived features. The frame-level quality responses of the HDR frame sequence are time-aggregated to obtain the target video quality score corresponding to the HDR video to be evaluated.

3. The HDR video quality evaluation method according to claim 1, characterized in that, The step of constructing a first training set based on the SDR frame data and the HDR frame data, using the SDR frame data in the first training set as input data, and using the HDR frame data corresponding to the SDR frame data as the supervision target, to train a pre-constructed residual domain adapter to obtain a target residual domain adapter includes: The content correspondence between the SDR frame data and the HDR frame data is established based on the video identifier and frame sequence number. The encoded value of the SDR frame data is normalized to a preset bounded interval to obtain the normalized SDR encoded value. The encoded value of the HDR frame data is normalized to a preset bounded interval to obtain the normalized HDR encoded value. Spatial size alignment is performed on the normalized SDR encoded value and the normalized HDR encoded value to obtain the first training set. The SDR frames in the first training set are input into the pre-constructed residual domain adapter to obtain the adaptation results corresponding to the SDR frames. Based on the content correspondence, the corresponding HDR frame of the SDR frame is determined, and the adaptation parameters of the residual domain adapter are updated based on the pixel-level consistency loss between the adaptation result and the HDR frame to obtain the updated residual domain adapter. Determine whether the updated residual domain adapter meets the first training termination condition. If the updated residual domain adapter meets the first training termination condition, use the updated residual domain adapter to obtain the pixel-level consistency index corresponding to the first preset validation set. Determine the adaptation parameter corresponding to the residual domain adapter when the pixel-level consistency index is optimal as the target adaptation parameter, and determine the residual domain adapter with the target adaptation parameter as the target residual domain adapter.

4. The HDR video quality evaluation method according to claim 1, characterized in that, The step of adapting the target video data using the target residual domain adapter to obtain adapted samples includes: Boundary protection processing is performed on the normalized encoded values ​​of the target video data to obtain the encoded values ​​to be transformed located within a preset open interval; A log-odds transform is performed on the encoded value to be transformed, mapping the encoded value to the target transform domain to obtain the target mapping result; the target transform domain is the log-odds transform domain. In the target transform domain, the target residual domain adapter is used to apply a residual correction amount to the target mapping result to obtain the target correction result; Perform the inverse log-probability transformation on the target correction result to obtain the adapted sample located within the preset open interval.

5. The HDR video quality evaluation method according to claim 1, characterized in that, The step of training the quality assessment backbone and the target residual domain adapter in the initial quality assessment model based on the predicted quality score and subjective quality score corresponding to the target video data to obtain the target quality assessment model includes: Based on the predicted quality score and subjective quality score corresponding to each target video data in the second training set, determine the target relevance constraint and target ranking constraint corresponding to the second training set; The joint training loss corresponding to the second training set is determined based on the target relevance constraint and the target ranking constraint. Backpropagation is performed based on the joint training loss to simultaneously update the target residual domain adapter and the quality assessment backbone to obtain the target quality assessment model.

6. The HDR video quality evaluation method according to claim 5, characterized in that, The step of performing backpropagation based on the joint training loss to simultaneously update the target residual domain adapter and the quality assessment backbone to obtain the target quality assessment model includes: Backpropagation is performed based on the joint training loss to simultaneously update the target residual domain adapter and the quality assessment backbone to obtain the updated initial quality assessment model. Determine whether the updated initial quality evaluation model meets the second training termination condition. If the updated initial quality evaluation model meets the second training termination condition, perform deterministic sampling on the second preset validation set to obtain validation sampling frames. Use the updated initial quality evaluation model to obtain the quality prediction performance index corresponding to the validation sampling frames. Determine the model parameters of the initial quality evaluation model corresponding to the optimal quality prediction performance index as the target model parameters. Determine the initial quality evaluation model with the target model parameters as the target quality evaluation model.

7. The HDR video quality evaluation method according to claim 5 or 6, characterized in that, Also includes: If the video resolution of the target video data in the second training set is greater than a preset resolution threshold, then the target residual domain adapter and the quality evaluation backbone are updated simultaneously based on a preset resource optimization strategy. The preset resource optimization strategies include a first optimization strategy based on micro-batch forward pass, a second optimization strategy based on gradient checkpoints, and a third optimization strategy based on a combination of micro-batch forward pass and gradient checkpoints.

8. An HDR video quality evaluation device, characterized in that, include: An adapter training module is used to acquire content-paired SDR frame data and HDR frame data, construct a first training set based on the SDR frame data and the HDR frame data, use the SDR frame data in the first training set as input data, and use the HDR frame data corresponding to the SDR frame data as the supervision target to train a pre-constructed residual domain adapter to obtain a target residual domain adapter; wherein, the residual domain adapter is a domain adapter based on log-odds transformation; The evaluation model construction module is used to acquire target video data, determine the dynamic range domain identifier corresponding to the target video data based on the data source corresponding to the target video data, and construct a second training set based on the target video data. An initial quality evaluation model is constructed using conditional routing, identity bypass, a quality evaluation backbone, and the target residual domain adapter. The target video data includes SDR video data and HDR video data, and the target video data carries a corresponding subjective quality score. The dynamic range domain identifier includes an SDR domain identifier and an HDR domain identifier. The conditional routing is used to determine the forward processing path from the target video data input to the quality evaluation backbone based on the dynamic range domain identifier. When the identity bypass is used as the forward processing path, the output data of the forward processing path is the same as the input data. The quality evaluation backbone is a no-reference quality evaluation network used to acquire the video quality score corresponding to the video data. The quality score prediction module is used to utilize the conditional routing to adapt the target video data through the target residual domain adapter to obtain adapted samples when the dynamic range domain identifier corresponding to the target video data in the second training set is an SDR domain identifier, and input the adapted samples into the quality evaluation backbone to obtain a first predicted quality score; and when the dynamic range domain identifier corresponding to the target video data in the second training set is an HDR domain identifier, the module uses the identity bypass to input the target video data into the quality evaluation backbone to obtain a second predicted quality score. The video evaluation module is used to train the quality evaluation backbone and the target residual domain adapter in the initial quality evaluation model based on the predicted quality score and subjective quality score corresponding to the target video data to obtain a target quality evaluation model, and to use the target quality evaluation model to obtain the target video quality score corresponding to the HDR video to be evaluated; the predicted quality score includes a first predicted quality score and a second predicted quality score.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the HDR video quality evaluation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program, wherein the computer program, when executed by a processor, implements the HDR video quality evaluation method as described in any one of claims 1 to 7.