Method and apparatus for machine video encoding

Through computer processing circuit evaluation and comparison of video coding schemes for machine vision and human vision, and using methods such as BD measurement and cost measurement, the coding efficiency and quality optimization problems in the existing technology are solved, and the efficient optimization and quality improvement of the coding scheme is achieved, which is suitable for a variety of application scenarios.

CN114641998BActive Publication Date: 2025-07-25TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180006264.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-06-28
Filing Date
2021-07-01
Publication Date
2025-07-25
Estimated Expiration
2041-07-01

Smart Images

  • Figure CN114641998B_ABST
    Figure CN114641998B_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide methods and apparatuses for use in machine video coding. In some examples, an apparatus for machine video coding includes processing circuitry. The processing circuitry determines a first picture quality versus coding efficiency characteristic of a first coding scheme for machine video coding (VCM), and determines a second picture quality versus coding efficiency characteristic of a second coding scheme for machine video coding. Then, based on the first picture quality versus coding efficiency characteristic and the second picture quality versus coding efficiency characteristic, the processing circuitry determines a delta (BD) metric for comparing the first coding scheme and the second coding scheme.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Incorporation by reference

[0002] This application claims the benefit of priority of U.S. Patent Application No. 17 / 360,838, filed on Jun. 28, 2021, entitled "METHOD AND APPARATUS IN VIDEO CODING FOR MACHINES" (which claims the benefit of priority of U.S. Provisional Application No. 63 / 090,555, filed on Oct. 12, 2020, entitled "WEIGHTED PERFORMANCE METRIC FOR VIDEO CODING FOR MACHINE", and U.S. Provisional Application No. 63 / 089,858, filed on Oct. 9, 2020, entitled "PERFORMANCE METRIC FOR VIDEO CODING FOR MACHINE"), the entire disclosure of which is incorporated herein by reference. Technical field

[0003] The present disclosure describes embodiments generally related to machine video coding. Background art

[0004] The background description provided herein is for the purpose of generally presenting the content of the present disclosure. To the extent that the work of the presently named inventors, which is described in this background art section and in various aspects of this specification, was carried out, it does not imply that it was available as prior art at the time of filing of this application, and it has never been expressly or implicitly admitted as prior art to the present disclosure.

[0005] Traditionally, videos or images have been used by people for various purposes, such as entertainment, education, etc. Therefore, video coding or image coding typically utilizes the characteristics of the human visual system to improve compression efficiency while maintaining good subjective quality.

[0006] In recent years, with the rise of machine learning applications and the abundance of sensors, many platforms use videos for machine vision tasks, such as object detection, segmentation, or tracking, etc. Video or image coding for consumption by machine tasks has become an area of concern and challenge. Summary of the invention

[0007] Aspects of the present disclosure provide methods and apparatuses for use in machine video coding. In some examples, a video coding apparatus for a machine includes processing circuitry. The processing circuitry determines a first picture quality versus coding efficiency characteristic of a first coding scheme for machine video coding (VCM), and determines a second picture quality versus coding efficiency characteristic of a second coding scheme for machine video coding. Then, based on the first picture quality versus coding efficiency characteristic and the second picture quality versus coding efficiency characteristic, the processing circuitry determines a delta (BD) metric for comparing the first coding scheme and the second coding scheme.

[0008] In some embodiments, the processing circuitry calculates the BD metric using at least one of: mean average precision (mAP), bits per pixel (BPP), multi-object tracking accuracy (MOTA), average precision at an intersection over union (IoU) threshold of 50% (AP50), average precision at an IoU threshold of 75% (AP75), and average accuracy.

[0009] In some examples, the first picture quality versus coding efficiency characteristic includes a first curve in a two-dimensional plane for picture quality and coding efficiency, and the second picture quality versus coding efficiency characteristic includes a second curve in a two-dimensional plane for picture quality and coding efficiency. In one example, the processing circuitry calculates the BD metric as an average gap between the first curve and the second curve.

[0010] In some examples, the first picture quality versus coding efficiency characteristic includes a first plurality of picture quality versus coding efficiency curves for a first coding scheme in a two-dimensional plane for picture quality on a first axis and coding efficiency on a second axis, and the second picture quality versus coding efficiency characteristic includes a second plurality of picture quality versus coding efficiency curves for a second coding scheme in a two-dimensional plane for picture quality and coding efficiency. The processing circuitry calculates a first Pareto front curve for the first plurality of picture quality versus coding efficiency curves for the first coding scheme, and calculates a second Pareto front curve for the second plurality of picture quality versus coding efficiency curves for the second coding scheme. Then, the processing circuitry calculates the BD metric based on the first Pareto front curve and the second Pareto front curve.

[0011] In one example, the second plurality of picture quality versus coding efficiency characteristic curves respectively correspond to the first plurality of picture quality versus coding efficiency characteristic curves. Then, the processing circuitry calculates BD metric values respectively based on the first plurality of picture quality versus coding efficiency characteristic curves and the corresponding second plurality of picture quality versus coding efficiency characteristic curves. Then, the processing circuitry calculates a weighted sum of the plurality of BD metric values as an overall BD metric for comparing the first coding scheme and the second coding scheme.

[0012] In some examples, the first coding scheme and the second coding scheme are for video coding for machine vision and human vision. The processing circuit calculates a first BD-rate for bits per pixel (BPP) for machine vision and a second BD-rate for human vision, and calculates a weighted sum of the first BD-rate and the second BD-rate as an overall BD metric for comparing the first coding scheme and the second coding scheme.

[0013] In some examples, the first coding scheme and the second coding scheme are for video coding for machine vision and human vision. The processing circuit calculates a first overall distortion by the first coding scheme based on a weighted sum of distortions for machine vision and human vision, and calculates a first cost metric value based on the first overall distortion and first rate information of the first coding scheme. Then, the processing circuit calculates a second overall distortion by the second coding scheme based on a weighted sum of distortions for machine vision and human vision, and calculates a second cost metric value based on the second overall distortion and second rate information of the second coding scheme. The processing circuit compares the first coding scheme and the second coding scheme based on the first cost metric value and the second cost metric value.

[0014] In some examples, the first coding scheme and the second coding scheme are for video coding for machine vision and human vision. The processing circuit determines a first picture quality by the first coding scheme based on a weighted sum of distortions for machine vision and human vision, and determines a second picture quality by the second coding scheme based on a weighted sum of distortions for machine vision and human vision.

[0015] In some examples, the first coding scheme and the second coding scheme are for video coding for multiple visual tasks. The processing circuit determines a first picture quality by the first coding scheme based on a weighted sum of distortions for multiple visual tasks, and determines a second picture quality by the second coding scheme based on a weighted sum of distortions for multiple visual tasks.

[0016] Aspects of the present disclosure also provide a non-transitory computer-readable medium for storing instructions that, when executed by a computer to perform machine video coding, cause the computer to perform a method of video coding. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0018] Figure 1 A block diagram of a VCM system according to an embodiment of the present disclosure is shown.

[0019] Figure 2 A diagram showing the calculation of a BD metric for machine video coding according to some embodiments of the present disclosure is shown.

[0020] Figure 3 Another figure shows the calculation of the BD metric for machine video coding according to some embodiments of the present disclosure.

[0021] Figures 4 to 7 An example of pseudocode is shown according to some embodiments of the present disclosure.

[0022] Figure 8 A flowchart outlining a processing example is shown according to an embodiment of the present disclosure.

[0023] Figure 9 is a schematic diagram of a computer system according to an embodiment. Detailed Description

[0024] Aspects of the present disclosure provide performance metric techniques for machine video coding (VCM). The performance metric techniques can be used to evaluate the performance of a coding tool according to another coding tool for meaningful comparison.

[0025] Figure 1 A block diagram of a VCM system (100) is shown according to an embodiment of the present disclosure. The VCM system (100) can be used for various usage applications, such as, for example, augmented reality (AR) applications, autonomous driving applications, video game goggle applications, sports game animation applications, surveillance applications, and the like.

[0026] The VCM system (100) includes a VCM encoding subsystem (101) and a VCM decoding subsystem (102) connected via a network (105). In one example, the VCM encoding subsystem (101) can include one or more devices for video coding for machine functions. In one example, the VCM encoding subsystem (101) includes a single computing device, such as, for example, a desktop computer, a laptop computer, a server computer, a tablet computer, and the like. In another example, the VCM encoding subsystem (101) includes one or more data centers, one or more server farms, and the like. The VCM encoding subsystem (101) can receive video content, such as, for example, a sequence of video frames output from a sensor device, and compress the video content into an encoded bitstream according to video coding for machine vision and / or video coding for human vision. The encoded bitstream can be transmitted via the network (105) to the VCM decoding subsystem (102).

[0027] The VCM decoding subsystem (102) includes one or more devices for video encoding for machine functions. In one example, the VCM decoding subsystem (102) includes computing devices such as a desktop computer, a laptop computer, a server computer, a tablet computer, a wearable computing device, a head-mounted display (HMD), etc. The VCM decoding subsystem (102) can decode an encoded bitstream according to video encoding for machine vision and / or video encoding for human vision. The decoded video content can be used for machine vision and / or human vision.

[0028] Any suitable technology can be used to implement the VCM encoding subsystem (101). In Figure 1 an example, the VCM encoding subsystem (101) includes a processing circuit (120), an interface circuit (111), and a multiplexer (112) coupled together.

[0029] The processing circuit (120) can include any suitable processing circuit, such as one or more central processing units (CPUs), one or more graphics processing units (GPUs), application-specific integrated circuits, etc. In Figure 1 an example, the processing circuit (120) can be configured to include two encoders, such as a video encoder (130) for human vision and a feature encoder (140) for machine vision. In one example, one or more CPUs and / or GPUs can execute software to act as the video encoder (130), and one or more CPUs and / or GPUs can execute software to act as the feature encoder (140). In another example, application-specific integrated circuits can be used to implement the video encoder (130) and / or the feature encoder (140).

[0030] In some examples, the video encoder (130) can perform video encoding for human vision on a sequence of video frames and generate a first bitstream, and the feature encoder (140) can perform feature encoding for machine vision on the sequence of video frames and generate a second bitstream. The multiplexer (112) can combine the first bitstream with the second bitstream to generate an encoded bitstream.

[0031] In some examples, the feature encoder (140) includes a feature extraction module (141), a feature transformation module (142), and a feature encoding module (143). The feature extraction module (141) can detect and extract features from the sequence of video frames. The feature transformation module (142) can transform the extracted features into a suitable feature representation, such as a feature map, a feature vector, etc. The feature encoding module (143) can encode the feature representation into the second bitstream. In some embodiments, feature extraction can be performed by an artificial neural network.

[0032] The interface circuit (111) can connect the VCM encoding subsystem (101) to the network (105). The interface circuit (111) can include a receiving part for receiving signals from the network (105) and a transmitting part for transmitting signals to the network (105). For example, the interface circuit (111) can transmit a signal carrying the encoded bitstream to other devices, such as the VCM decoding subsystem (102), via the network (105).

[0033] The network (105) is suitably coupled to the VCM encoding subsystem (101) and the VCM decoding subsystem (102) through wired and / or wireless connections (such as Ethernet connection, fiber optic connection, WI-FI connection, cellular network connection, etc.). The network (105) can include network server devices, storage devices, network devices, etc. The components of the network (105) are suitably coupled together through wired and / or wireless connections.

[0034] The VCM decoding subsystem (102) is configured to decode the encoded bitstream for machine vision and / or human vision. In one example, the VCM decoding subsystem (102) can perform video decoding to reconstruct a sequence of video frames that can be displayed to human vision. In another example, the VCM decoding subsystem (102) can perform feature decoding to reconstruct a feature representation that can be used for machine vision.

[0035] Any suitable technology can be used to implement the VCM decoding subsystem (102). In Figure 1 the example, the VCM decoding subsystem (102) includes an interface circuit (161), a demultiplexer (162), and a processing circuit 170 coupled together as Figure 1 shown.

[0036] The interface circuit (161) can connect the VCM decoding subsystem (102) to the network (105). The interface circuit (161) can include a receiving part for receiving signals from the network (105) and a transmitting part for transmitting signals to the network (105). For example, the interface circuit (161) can receive a signal carrying data from the network (105), such as a signal carrying the encoded bitstream.

[0037] The demultiplexer (162) can separate the received encoded bitstream into a first encoded video bitstream and a second encoded feature bitstream.

[0038] The processing circuit (170) can include appropriate processing circuits, such as a CPU, a GPU, an application-specific integrated circuit, etc. The processing circuit (170) can be configured to include various decoders, such as a video decoder, a feature decoder, etc. For example, the processing circuit (170) is configured to include a video decoder (180) and a feature decoder (190).

[0039] In one example, the GPU is configured as a video decoder (180). In another example, the CPU can execute software instructions to act as a video decoder (180). In another example, the GPU is configured as a feature decoder (190). In another example, the CPU can execute software instructions to act as a feature decoder (190).

[0040] The video decoder (180) can decode the information in the first encoded video bitstream and reconstruct the decoded video (e.g., a sequence of picture frames). The decoded video can be displayed for human vision. In some examples, the decoded video can be provided for machine vision.

[0041] The feature decoder (190) can decode the information in the second encoded feature bitstream and reconstruct the features in a suitable representation. The decoded features can be provided for machine vision. In some examples, the decoded features can be provided for human vision.

[0042] Machine video encoding can be performed in the VCM system (100) by various encoding tools or according to various configurations of the encoding tools. Aspects of the present disclosure provide performance metric techniques to evaluate the performance of various encoding tools and / or various configurations of the encoding tools, and then the encoding tools and suitable configurations can be selected based on the performance metric techniques.

[0043] In Figure 1 an example, the VCM encoding subsystem (101) includes a controller (150) coupled to a video encoder (130) and a feature encoder (140). The controller (150) can perform a performance evaluation using the performance metric techniques and can select the encoding tools and / or configurations to be used in the video encoder (130) and the feature encoder (140) based on the performance evaluation. The controller (150) can be implemented by various techniques. In one example, the controller (150) is implemented as a processor that executes software instructions for performance evaluation and selection of encoding tools and configurations. In another example, the controller (150) is implemented using an application specific integrated circuit.

[0044] According to some aspects of the present invention, different picture quality metrics and coding efficiency metrics are used to evaluate the video / image coding quality for human vision and machine vision.

[0045] In some examples for evaluating video / image coding quality for human vision, performance metrics such as mean squared error (MSE) / peak signal-to-noise ratio (PSNR), structure similarity index measure (SSIM) / multi-scale SSIM (MS-SSIM), video multimethod assessment fusion (VMAF), etc. can be used. In one example, MSE can be used to calculate the mean squared error between the original image and the reconstructed image of the original image, and the reconstructed image is the result of operations under or with the configuration of coding tools. PSNR can be calculated as the ratio between the maximum possible power of the signal and the power of the corrupting noise that affects its rendering fidelity. PSNR can be defined based on MSE. MSE or PSNR is calculated based on absolute error.

[0046] In another example, SSIM can be used to measure the similarity between the original image and the reconstructed image of the original image. SSIM uses structural information, that is, there is a strong correlation between pixels, especially when they are spatially close. The correlation carries important information about the structure of the objects in the visual scene. MS-SSIM can be performed at multiple scales through a multi-stage subsampling process.

[0047] In some examples for evaluating video / image coding quality for machine vision, performance metrics (such as mAP, MOTA, etc.) can be used to measure the performance of machine vision tasks, such as object detection, segmentation, or tracking, etc. In addition, BPP can be used to measure the cost of storing or transmitting the generated bitstream for VCM.

[0048] Specifically, in some examples, mAP is calculated as the area under the precision-recall curve (PR curve), where the x-axis is the recall rate and the y-axis is the precision rate. In some examples, BPP is calculated according to the image resolution (such as the original image resolution).

[0049] In some examples, a picture quality versus encoding efficiency characteristic can be determined for an encoding tool for VCM, and a performance evaluation of the encoding tool can be determined based on the picture quality versus encoding efficiency characteristic. By using the encoding tool for encoding video, the picture quality versus encoding efficiency characteristic for the encoding tool represents the relationship between picture quality and encoding efficiency. In some examples, the picture quality versus encoding efficiency characteristic of the encoding tool can be represented as a curve in a two-dimensional plane, where picture quality is on the first axis and encoding efficiency is on the second axis. In some examples, the picture quality versus encoding efficiency characteristic of the encoding tool can be represented by an equation that calculates encoding efficiency based on picture quality. In some examples, a look-up table that correlates picture quality with encoding efficiency can be used to represent the picture quality and encoding efficiency characteristic of the encoding tool.

[0050] In an example, picture quality is measured based on mAP (or MOTA), and encoding efficiency is measured based on BPP. The relationship between mAP (or MOTA) and BPP can be plotted as a curve indicating the picture quality versus encoding efficiency characteristic and can be used to represent the performance of an encoding scheme (also referred to as an encoding tool) for machine vision. Additionally, before video / image encoding and cropping, the video / image can be preprocessed by padding or scaling to achieve different resolutions of the original content, e.g., 100%, 75%, 50%, and 25% of the original resolution. After decoding for a machine vision task, the decoded video / image can be scaled back to the original resolution. In some embodiments, multiple mAP (or MOTA) versus BPP curves can be plotted for the encoding scheme for machine vision.

[0051] Some aspects of the present disclosure provide techniques for calculating a single performance value (performance metric) from, for example, one or more relationship curves between mAP (or MOTA) and BPP, such that comparison of multiple encoding schemes for machine vision can be based on the performance values of the multiple encoding schemes. For example, a controller (150) can calculate performance metric values for multiple encoding schemes respectively and can select an encoding scheme from the multiple encoding schemes based on the multiple performance metric values.

[0052] It should be noted that in the following description, the mAP versus BPP relationship curve is used to describe the performance metric technique according to some aspects of the present disclosure. The performance metric technique can be used for other relationship curves. For example, mAP can be changed to other suitable performance metrics, e.g., MOTA, average precision with an intersection over union threshold of 50% (AP50), average precision with an intersection over union threshold of 75%, average accuracy, etc. In another example, when the input is video, BPP can be changed to bit rate.

[0053] According to some aspects of the present disclosure, BD metrics, such as BD mean average precision (BD-mAP), BD bits per pixel (BD-rate), can be used for performance evaluation of machine video coding.

[0054] Figure 2 A diagram (200) for calculating BD metrics for machine video coding according to some embodiments of the present disclosure is shown. Diagram (200) includes a first mAP vs. BPP curve (210) for a first VCM scheme, and a second mAP vs. BPP curve (220) for a second VCM scheme. In some examples, the mAP and BPP values can be determined according to different quantization parameter (QP) values, such as the QP for all I slices (QPISlice).

[0055] For example, to encode a video using the first VCM scheme, the controller (150) sets the QPISlice value, and the video encoder (130) and the feature encoder (140) can encode the video using the first VCM scheme based on the QPISlice value and generate a first encoded bitstream. Based on the first encoded bitstream, the mAP and BPP values associated with the QPISlice value can be determined, for example, by the controller (150). The controller (150) can set different QPISlice values for using the first VCM scheme and determine the mAP and BPP values associated with each QPISlice value. Then, the first curve (210) is formed using the mAP and BPP values. For example, for the QPISlice value, the mAP value and the BPP value associated with the QPISlice value are used to form a point on the first curve (210).

[0056] Similarly, to encode a video using the second VCM scheme, the controller (150) sets the QPISlice value, and the video encoder (130) and the feature encoder (140) can encode the video using the second VCM scheme based on the QPISlice value and generate a second encoded bitstream. Based on the second encoded bitstream, the mAP and BPP values associated with the QPISlice value can be determined, for example, by the controller (150). The controller (150) can set different QPISlice values for using the second VCM scheme and determine the mAP and BPP values associated with each QPISlice value. Then, the second curve (220) is formed using the mAP and BPP values. For example, for the QPISlice value, the mAP value and the BPP value associated with the QPISlice value are used to form a point on the second curve (220).

[0057] According to one aspect of the present invention, the BD-mAP can be determined as the average gap between a first curve (210) and a second curve (220), and the average gap can be calculated based on the area (230) (shown by the gray shading) between the first curve (210) and the second curve (220). In some examples, the controller (150) can execute software instructions corresponding to an algorithm for calculating the area of the gap between two curves (e.g., the first curve (210) and the second curve (220)), and determine the BD-mAP value based on the area, for example.

[0058] In some examples, the first VCM scheme can be a reference scheme (also referred to as an anchor), and the second VCM scheme is the scheme being evaluated (or being tested). In one example, the reference scheme is applied to an unscaled (or scaled to 100%) video, and the scheme being evaluated can be applied to a 75% scaled video of the video. The BD metric of the second VCM scheme with respect to the reference scheme can be used to compare the performance of the second VCM scheme with that of other VCM schemes (other than the first VCM scheme and the second VCM scheme).

[0059] In some embodiments, the BD rate can be determined similarly.

[0060] Figure 3 A diagram (300) for calculating the BD metric for machine video coding according to some embodiments of the present disclosure is shown. The diagram (300) includes a first BPP-versus-mAP curve (310) for a first VCM scheme, and a second BPP-versus-mAP curve (320) for a second VCM scheme. In some examples, the mAP and BPP values are determined according to different values of QPISlice in the same manner as described in the reference Figure 2 stated.

[0061] According to one aspect of the present invention, the BD rate can be determined as the average gap between a first curve (310) and a second curve (320), and the average gap can be calculated based on the area (330) between the first curve (310) and the second curve (320). In some examples, the controller (150) can execute software instructions corresponding to an algorithm for calculating the area of the gap between two curves (e.g., the first curve (310) and the second curve (320)), and determine the BD rate value based on the area, for example.

[0062] In some examples, the first VCM scheme can be a reference scheme, and the second VCM scheme is the scheme being evaluated. In one example, the reference scheme is applied to an unscaled (or scaled to 100%) video, and the scheme being evaluated can be applied to a video scaled by 75%. In one example, the average gap between the first curve (310) and the second curve (320) indicates that 14.75% fewer bits are sent or stored when using the second VCM scheme to achieve equivalent quality. The BD metric for the second VCM scheme with respect to the reference scheme can be used to compare the performance of the second VCM scheme with the performance of other VCM schemes (other than the first VCM scheme and the second VCM scheme). For example, when the BD rate of the third VCM scheme indicates that 10% fewer bits are to be sent or stored, it is determined that the second VCM scheme has better VCM performance than the third VCM scheme.

[0063] According to one aspect of the present invention, when comparing two VCM schemes, multiple BPP-versus-mAP (or mAP-versus-BPP) curves can be generated for each scheme, and some techniques can be used to determine a single performance metric that indicates an overall summary of the performance difference for performance comparison, e.g., the BD metric.

[0064] In one embodiment, a first Pareto front curve can be formed based on multiple BPP-versus-mAP (or mAP-versus-BPP) curves for the first VCM scheme, and a second Pareto front curve can be formed based on multiple mAP-versus-BPP curves for the second VCM scheme. For example, for the first VCM scheme, when a particular BPP-versus-mAP curve is always better than other BPP-versus-mAP curves, that particular BPP-versus-mAP curve can be used as the first Pareto front curve. However, when multiple BPP-versus-mAP curves may intersect, an optimal section of the multiple BPP-versus-mAP curves can be selected to form the first Pareto front curve. For the second VCM scheme, the second Pareto front curve can be formed similarly. Then, the BD metric can be calculated using the first Pareto front curve of the first scheme and the second Pareto front curve of the second scheme.

[0065] In another embodiment, the BD metric can be calculated separately for multiple BPP-versus-mAP curves of the VCM scheme, and the average (e.g., weighted average) of the multiple BD metric values can be used for performance comparison.

[0066] In some examples, a video can be pre - processed to obtain different resolutions of the original content. Then, the first VCM scheme and the second VCM scheme can be used to encode the pre - processed video respectively, and the BD - rate can be calculated separately for different resolutions. In one example, the video is pre - processed to obtain four videos with different resolutions, for example, a first video with 100% resolution, a second video with 75% resolution, a third video with 50% resolution, and a fourth video with 25% resolution. In one example, the first VCM scheme and the second VCM scheme can be applied to the first video to calculate the first BD - rate; the first VCM scheme and the second VCM scheme can be applied to the second video to calculate the second BD - rate; the first VCM scheme and the second VCM scheme can be applied to the third video to calculate the third BD - rate; the first VCM scheme and the second VCM scheme can be applied to the fourth video to calculate the fourth BD - rate. Then, the average value of the first BD - rate, the second BD - rate, the third BD - rate, and the fourth BD - rate can be calculated as the overall BD - rate for the performance comparison of the first VCM scheme and the second VCM scheme.

[0067] In some examples, the BD - rates can be weighted equally or differently to calculate the overall BD - rate. In one example, a particular scaling has more importance than other scalings, and a higher weight can be assigned to the particular scaling. In some examples, the sum of all weights is equal to 1. It should be noted that, taking the mAP curve of BPP for various resolution scalings as an example, the comparison of two VCM schemes is shown, and each scheme has multiple curves. The technique for handling multiple curves is not limited to the technique with various resolution scalings.

[0068] According to some aspects of the present invention, in certain applications, the decoded video can be consumed by machine vision and human vision, and performance metrics can be used to compare two encoding schemes while considering both usage scenarios (consumed by both machine vision and human vision) simultaneously. In one embodiment, the first BD - rate for machine vision and the second BD - rate for human vision can be calculated separately, and then the first BD - rate for machine vision and the second BD - rate for human vision can be appropriately combined to form a performance metric for the performance comparison of the two encoding schemes.

[0069] In some examples, when comparing two encoding schemes, a machine vision metric (e.g., the BPP - versus - mAP curve) can be used to calculate the first BD - rate for machine vision (denoted as BDm). Then, a human vision metric (e.g., the bit - rate curve versus PSNR, the bit - rate curve versus MS - SSIM, the bit - rate curve versus SSIM, etc.) can be used to calculate the second BD - rate for human vision (denoted as B Dh). It should be noted that, in addition to PSNR, MS - SSIM, or SSIM, similar human - vision - related performance metrics can be used to measure the encoding performance for human vision.

[0070] In some examples, the first BD rate for machine vision and the second BD rate for human vision can be combined to calculate the final performance comparison result represented by BDoverall. For example, Equation (1) can be used to combine the first BD rate BDm for machine vision and the second BD rate B Dh for human vision:

[0071] BD overall =(1 - w)×BD m + w×BD h Equation (1)

[0072] where w represents a weight within the range of [0, 1] and indicates the relative importance of human vision to the overall coding performance.

[0073] In some embodiments, a cost metric can be calculated as a performance metric for comparing coding schemes. The cost metric (denoted as C) can include a first part for distortion and a second part for rate information (denoted as RT). The distortion can be an overall distortion (denoted as D) generated as a combination of the distortion for human vision (denoted as Dh) and the distortion for machine vision (denoted as Dm). In one example, the rate information (denoted as RT) can be the bitstream length used to represent the encoded video. In one example, Equations (2) and (3) can be used to calculate the overall distortion D and the cost metric C.

[0074] D=(1 - w1)×D m + w1×D h Equation (2)

[0075] C = D+λ×RT Equation (3)

[0076] where w1 is a weight within the range of [0, 1] and is used to indicate the relative importance of human vision in the overall application of the combination of human vision and machine vision. The parameter λ is a non - negative scalar used to represent the relative importance of distortion and rate.

[0077] In certain examples, the distortion Dh of human vision can be calculated as the normalized mean error (NME), for example, calculated using Equation (4):

[0078]

[0079] where N represents the total number of pixels in the image or video; P is typically 1 or 2; ‖P‖ represents the corresponding P - norm operation; ‖P‖ maxrepresents the maximum P-norm of the pixels in the original image or video; P(i) represents the i-th pixel in the original image or video; and P'(i) represents the i-th pixel in the decoded image or video. It can be noted that NME is in the range of [0, 1].

[0080] In some examples, for instance, when the image or video is monochromatic with only a single color channel, the P-norm operation ‖P‖ can be converted to an absolute value operation. In one example, when using a color image or video, a pixel can be represented by a three-tuple, such as (R, G, B), where R, G, and B represent the three color channel values in the RGB color space. In another example, when using a color image or video, a pixel can be represented by a three-tuple, such as (Y, Cb, Cr), where Y, Cb, and Cr represent the three channel values in the YCbCr color space. Using the example of the RGB color space, for example, the P-norm operation can be calculated using formula (5):

[0081]

[0082] where w R ,w G ,w B represent the weights of each color channel, and these weights may be the same or different. In some examples, if the three channels have different resolutions, then smaller weights can be applied to the channel with lower resolution to reflect this difference.

[0083] In some examples, the normalized mean error (NME) of each channel can be calculated first, for example, using formulas (6) to (8):

[0084]

[0085]

[0086]

[0087] where N R 、N g and N b represent the number of pixels in the R, G, and B channels respectively; R max ,G max and B max represent the maximum values in the R, G, and B channels, and in some examples are usually set to the same value. (R(i), G(i), B(i)) represents the i-th pixel in the original image or video, and (R′(i), G′(i), B′(i)) represents the i-th pixel in the decoded image or video. Additionally, in one example, P is usually set to 1 or 2. The overall normalized mean error NME can be used as NMER , NME G , NME B is calculated as the weighted average of, for example, using formula (9):

[0088] NME = w R × NME R + w G × NME G + w B × NME B Formula (9)

[0089] wherein, w R , w G , w B are non - negative weights representing the relative importance of the three color channels, and w R + w G + w B = 1.

[0090] It can be noted that the above channel - based NME calculation in (6) - (8) can be appropriately converted to other color formats, for example, Y'CbCr (YUV) with 4:2:0 subsampling or 4:4:4 subsampling. In one example, the weight 1 can be assigned only to the Y component, while the weight 0 is assigned to the other two components Cb and Cr. In another example, the weights are determined based on the sample resolution of each channel. For example, in 4:2:0, since the resolution of Y is higher than that of UV, the weights of the UV components should be less than the weight of the Y component.

[0091] According to one aspect of the present invention, depending on the machine task, the distortion of machine vision can be expressed as (1 - mAP) or (1 - MOTA). It can be noted that mAP and MOTA are in the range of [0, 1]. It can also be noted that in the calculation of the distortion of machine vision, other similar machine vision performance metrics (e.g., average accuracy) can also be used to replace mAP or MOTA.

[0092] In some examples, when the distortion for human vision is greater than the threshold for human vision (denoted as Threshh), the decoded image may not be useful for the combination of human vision and machine vision. Similarly, when the distortion for machine vision is greater than the threshold for machine vision (denoted as Threshm), the decoded image becomes useless for the combination of human vision and machine vision. In either case, the overall distortion D is set to a predetermined value, for example, its maximum value (e.g., 1).

[0093] Figure 4 Shows an example of the pseudocode (400) for calculating the overall distortion D and the overall cost metric C in an example when considering both machine vision and human vision. InFigure 4 In the example, w1 is a weight within the range of [0, 1] and is used to indicate the relative importance of human vision in the overall application of the combination of human vision and machine vision. Parameter Dh is the distortion for human vision and parameter Dm is the distortion for machine vision. Parameter λ is a non - negative scalar and is used to represent the relative importance of distortion and rate. Parameter Threshh is the distortion threshold for human vision and parameter Threshm is the distortion threshold for machine vision.

[0094] In some examples, the distortion of the machine vision task can be considered more important than the human vision task, while the quality of the human vision perspective should be maintained at a minimum acceptable level.

[0095] Figure 5 An example of the pseudocode (500) for calculating the overall distortion D and the overall cost metric C is shown in an example where the distortion of the machine vision task is considered more important than the human vision task. The quality of the human vision perspective can be maintained at the lowest acceptable level. In Figure 5 In the example, w1 is a weight within the range of [0, 1] and is used to indicate the relative importance of human vision in the overall application of the combination of human vision and machine vision. Parameter Dm is the distortion for machine vision. Parameter λ is a non - negative scalar and is used to represent the relative importance of distortion and rate. Parameter Threshh is the distortion threshold for human vision.

[0096] In some examples, when the distortion of human vision is below the lowest acceptable level, the distortion of human vision can be considered.

[0097] Figure 6 An example of the pseudocode (600) for calculating the overall distortion D and the overall cost metric C is shown in an example where the distortion of the machine vision task is considered more important than the human vision task. When the distortion of human vision is below the lowest acceptable level, the distortion of human vision can be considered. In Figure 6 In the example, w1 is a weight within the range of [0, 1] and is used to indicate the relative importance of human vision in the overall application of the combination of human vision and machine vision. Parameter Dh is the distortion for human vision and parameter Dm is the distortion for machine vision. Parameter λ is a non - negative scalar and is used to represent the relative importance of distortion and rate. Parameter Threshh is the distortion threshold for human vision.

[0098] In some examples, in addition to combining the encoding performance for human vision and machine vision, the encoding performance for multiple tasks (e.g., more than two tasks) can also be combined. The multiple tasks can include object detection tasks, object segmentation tasks, object tracking tasks, etc. In an example of M tasks (where M is an integer greater than 2), the overall BD rate can be a weighted combination of the BD rates for the multiple individual tasks, e.g., using Equation (10):

[0099] BD overall = w(0) × BD0 + w(1) × BD1 + … + w(M−1) × BD M-1 Equation (10)

[0100] where w(i), i = 0, 1, …, M−1 are non-negative weight factors, and w(0) + w(1) + … + w(M−1) = 1. BD i , i = 0, …, M−1 are the corresponding BD rates for the M tasks.

[0101] In an example of M tasks (where M is an integer greater than 2), the overall distortion can be a weighted combination of the distortions for the multiple individual tasks, e.g., using Equation (11), and then the cost metric C can be calculated, e.g., using Equation (12):

[0102] D = w(0) × D0 + w(1) × D1 + … + w(M−1) × D M-1 Equation (11)

[0103] C = D + λ × RT Equation (12)

[0104] where w(i), i = 0, 1, …, M−1 are non-negative weight factors, and w(0) + w(1) + … + w(M−1) = 1. D i , i = 0, …, M−1 are the corresponding distortions for the M tasks.

[0105] In some examples, when the distortion of a task is greater than the threshold of the task, for the combination of multiple tasks, the decoded image may not be useful.

[0106] Figure 7 An example of the pseudocode (700) for calculating the overall distortion D and the overall cost metric C is shown in an example when a threshold is used separately for each task. When the distortion of a task is greater than the threshold of the task, for the combination of multiple tasks, the decoded image may not be useful. In Figure 7 the example, w(i), i = 0, 1, …, M−1 are non-negative weight factors, and w(0) + w(1) + … + w(M−1) = 1. Di, i = 0, …, M−1 are the corresponding distortions for the M tasks. Threshi, i = 0, …, M−1 are the corresponding distortions for the M tasks.

[0107] In some examples, based on the calculated overall distortion, the overall weighted precision expressed as wmAP can be calculated using, for example, formula (13):

[0108] wmAP = 1 - D Formula (13)

[0109] In addition, a wmAP vs. BPP curve can be formed. Then, the corresponding BD rate for comparing two wmAP vs. BPP curves can be calculated.

[0110] In one embodiment, to compare two coding schemes, e.g., an anchor coding scheme and a test coding scheme, the overall BD rate can be calculated. If the overall BD rate is negative, the test coding scheme has better performance than the anchor coding scheme.

[0111] In one embodiment, to compare two coding schemes, e.g., an anchor coding scheme and a test coding scheme, the corresponding cost metric can be calculated. The coding scheme with a smaller overall cost metric is considered to have better performance.

[0112] It can be noted that the above techniques can be modified appropriately. In some examples, a transformation function can be used. In one example, a machine vision quality metric, e.g., mAP or MOTA, can be used as the input to the transformation function. Then, the output of the transformation function can be used for BD rate calculation. The transformation function can include linear scaling, square root operation, log-domain transformation, etc.

[0113] Figure 8 A flowchart outlining a process (800) according to an embodiment of the present disclosure is shown. The process (800) can be used for comparing coding schemes, e.g., for a VCM system (100) etc. In embodiments, the process (800) is executed by a processing circuit such as, for example, the processing circuit (120). In some embodiments, the process (800) is implemented as software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the process (800). The process starts at (S801) and proceeds to (S810).

[0114] At (S810), determine the first picture quality vs. coding efficiency characteristic of a first coding scheme for machine video coding.

[0115] At (S820), determine the second picture quality vs. coding efficiency characteristic of a second coding scheme for machine video coding.

[0116] In some examples, the picture quality can be measured by any suitable machine vision quality metric, for example, mAP, MOTA, average precision at an intersection over union (IoU) threshold of 50% (AP50), average precision at an IoU threshold of 75% (AP75), average accuracy, etc. The coding efficiency can be measured by any suitable metric for VCM, for example, BPP, etc.

[0117] At (S830), based on the first picture quality versus coding efficiency characteristic and the second picture quality versus coding efficiency characteristic, a BD metric for comparing the first coding scheme and the second coding scheme is determined.

[0118] In some examples, the first picture quality versus coding efficiency characteristic includes a first curve in a two-dimensional plane for picture quality and coding efficiency, and the second picture quality versus coding efficiency characteristic includes a second curve in a two-dimensional plane for picture quality and coding efficiency. Then, the BD metric can be calculated as the average gap between the first curve and the second curve. The average gap can be calculated based on the area between the first curve and the second curve.

[0119] In one example, the BD metric can be calculated based on mAP, for example, denoted as BD-mAP. In another example, the BD metric is calculated based on BPP, for example, denoted as BD-BPP or BD rate. In one example, the BD metric is calculated based on MOTA, for example, denoted as BD-MOTA. In one example, the BD metric is calculated based on the average precision at an IoU threshold of 50% (AP50), for example, denoted as BD-AP50. In one example, the BD metric is calculated based on the average precision at an IoU threshold of 75% (AP75), for example, denoted as BD-AP75. In one example, the BD metric is calculated based on the average precision.

[0120] In some examples, the first picture quality versus coding efficiency characteristic includes a first plurality of picture quality versus coding efficiency curves for a first coding scheme in a two-dimensional plane for picture quality and coding efficiency, and the second picture quality versus coding efficiency characteristic includes a second plurality of picture quality versus coding efficiency curves for a second coding scheme in a two-dimensional plane for picture quality and coding efficiency. In one example, based on the first plurality of picture quality versus coding efficiency curves for the first coding scheme, a first Pareto front curve is calculated, and based on the second plurality of picture quality versus coding efficiency curves for the second coding scheme, a second Pareto front curve is calculated. Then, the BD metric is calculated based on the first Pareto front curve and the second Pareto front curve.

[0121] In another example, the second plurality of picture quality vs. coding efficiency characteristic curves respectively correspond to the first plurality of picture quality vs. coding efficiency characteristic curves. Then, based on the first plurality of picture quality vs. coding efficiency characteristic curves and the corresponding second plurality of picture quality vs. coding efficiency characteristic curves, BD metric values are respectively calculated. Then, a weighted sum of the plurality of BD metric values is calculated as the overall BD metric for comparing the first coding scheme and the second coding scheme.

[0122] In some examples, the first coding scheme and the second coding scheme are used for video coding for machine vision and human vision. In one example, a first BD rate for bits per pixel for machine vision and a second BD rate for human vision are calculated. Then, a weighted sum of the first BD rate and the second BD rate is calculated as the overall BD metric for comparing the first coding scheme and the second coding scheme.

[0123] In another embodiment, through the first coding scheme, a first overall distortion is calculated based on a weighted sum of distortions for machine vision and human vision. Then, based on the first overall distortion and the first rate information of the first coding scheme, a first cost metric value is calculated. Further, through the second coding scheme, a second overall distortion is calculated based on a weighted sum of distortions for machine vision and human vision, and based on the second overall distortion and the second rate information of the second coding scheme, a second cost metric value is calculated. The first coding scheme and the second coding scheme can be compared based on the first cost metric value and the second cost metric value.

[0124] In another embodiment, through the first coding scheme, a first picture quality is determined based on a weighted sum of distortions for machine vision and human vision. Through the second coding scheme, a second picture quality is determined based on a weighted sum of distortions for machine vision and human vision. Then, a BD metric can be calculated and used to compare the first coding scheme and the second coding scheme.

[0125] In some examples, the first coding scheme and the second coding scheme are for video coding for multiple visual tasks. Then, through the first coding scheme, a first picture quality is determined based on a weighted sum of distortions for multiple visual tasks. Through the second coding scheme, a second picture quality is determined based on a weighted sum of distortions for multiple visual tasks. Then, a BD metric can be calculated and used to compare the first coding scheme and the second coding scheme.

[0126] Then, the process proceeds to (S899) and ends.

[0127] The above techniques can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 9 A computer system (900) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0128] A computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can be subject to mechanisms such as assembly, compilation, linking, or the like to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode, etc.

[0129] The instructions can be executed on various types of computers or their components, such as including personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0130] Figure 9 The components of the computer system (900) shown are exemplary in nature and are not intended to impose any limitations on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should also not be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system (900).

[0131] The computer system (900) may include certain human - machine interface input devices. Such human - machine interface input devices can respond to one or more human users through inputs such as the following: tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, clapping), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human - machine interface devices can also be used to capture certain media not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still - image camera), video (e.g., two - dimensional video, three - dimensional video including stereoscopic video), etc.

[0132] The input human - machine interface devices may include one or more of the following (only one of each is shown): keyboard (901), mouse (902), touchpad (903), touch screen (910), data glove (not shown), joystick (905), microphone (906), scanner (907), camera (908).

[0133] A computer system (900) may include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users, for example, through haptic output, sound, light, and smell / taste. Such human-machine interface output devices may include haptic output devices (such as haptic feedback of a touch screen (910), a data glove (not shown), or a joystick (905), but may also be haptic feedback devices that are not input devices), audio output devices (such as: speakers (909), headphones (not depicted)), visual output devices (such as a screen (910) including a CRT screen, an LCD screen, a plasma screen, an OLED screen, each screen having or not having touch screen input functionality, each screen having or not having haptic feedback functionality - some of which screens are capable of outputting two-dimensional visual output or output beyond three dimensions through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted)).

[0134] The computer system (900) may also include human-accessible storage devices and their associated media, for example, including optical media such as CD / DVDROM / RW (920) with media such as CD / DVD (921), thumb drives (922), removable hard disk drives or solid-state drives (923), traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.

[0135] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0136] The computer system (900) may also include an interface (954) to one or more communication networks (955). The network can be, for example, a wireless network, a wired network, an optical network. The network can also be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Certain networks typically require an external network interface adapter (e.g., a USB port of the computer system (900)) to be connected to certain general data ports or peripheral buses (949); as described below, other network interfaces are typically integrated into the core of the computer system (900) by connecting to the system bus (e.g., an Ethernet interface in a PC computer system or a cellular network interface in a smartphone computer system). The computer system (900) can communicate with other entities using any of these networks. Such communication can be one-way reception only (e.g., broadcast television), one-way transmission only (e.g., CANbus connected to certain CANbus devices), or two-way, for example, using a local area network or a wide area digital network to connect to other computer systems. As described above, certain protocols and protocol stacks can be used on each of those networks and network interfaces.

[0137] The above-mentioned human-machine interface device, human-machine accessible storage device, and network interface can be attached to the core (940) of the computer system (900).

[0138] The core (940) can include one or more central processing units (CPUs) (941), a graphics processing unit (GPU) (942), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (943), a hardware accelerator (940) for certain tasks, a graphics adapter (950), etc. These devices, as well as a read-only memory (ROM) (945), a random access memory (946), an internal mass storage device such as an internal non-user accessible hard disk drive, SSD, etc. (947) can be connected through a system bus (948). In some computer systems, the system bus (948) can be accessed in the form of one or more physical plugs to enable expansion through additional CPUs, GPUs, etc. Peripheral devices can be directly connected to the system bus (948) of the core or connected to the system bus (1848) of the core through a peripheral bus (949). In one example, a screen (910) can be connected to the graphics adapter (950). The architecture of the peripheral bus includes PCI, USB, etc.

[0139] The CPU (941), GPU (942), FPGA (943), and accelerator (944) can execute certain instructions, which can be combined to form the above computer code. The computer code can be stored in the ROM (945) or RAM (946). Transitional data can also be stored in the RAM (946), while permanent data can be stored, for example, in the internal mass storage (947). Fast storage and retrieval to any storage device can be performed by using a cache, which can be closely associated with one or more of the following: one or more CPUs (941), GPUs (942), mass storage (947), ROM (945), RAM (946), etc.

[0140] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be media and computer code specially designed and constructed for the purposes of this disclosure, or the medium and the computer code can be of the type well-known and available to those skilled in the field of computer software.

[0141] As a non-limiting example, a computer system having an architecture (900), particularly a core (940), can provide functionality due to one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage as described above, as well as certain non-transitory memories of the core (940), such as the core internal mass storage (947) or ROM (945). The software implementing the embodiments of this disclosure can be stored in such devices and executed by the core (940). Depending on specific needs, the computer-readable media can include one or more memory devices or chips. The software can cause the core (940), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute the specific processes or specific parts of the specific processes described herein, including defining data structures stored in the RAM (946) and modifying such data structures according to the processes defined by the software. Additionally or alternatively, a computer system can provide functionality due to logic hard-wired or otherwise embodied in a circuit (e.g., accelerator (944)), which can replace the software or operate together with the software to execute the specific processes or specific parts of the specific processes described herein. In appropriate cases, portions referring to software can include logic, and vice versa. In appropriate cases, portions referring to computer-readable media can include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or including both. This disclosure encompasses any suitable combination of hardware and software.

[0142] Although the present disclosure has described a number of exemplary embodiments, there are modifications, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art will be able to devise many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure.

Claims

1. A method for video coding and decoding, comprising: Determining a first picture quality versus coding efficiency characteristic of a first coding scheme for machine video coding (VCM), the first picture quality versus coding efficiency characteristic comprising: a first plurality of picture quality versus coding efficiency curves for the first coding scheme in a two-dimensional plane for picture quality and coding efficiency; Determining a second picture quality versus coding efficiency characteristic of a second coding scheme for VCM, the second picture quality versus coding efficiency characteristic comprising: a second plurality of picture quality versus coding efficiency curves for the second coding scheme in the two-dimensional plane for picture quality and coding efficiency; The second plurality of picture quality versus coding efficiency characteristic curves respectively correspond to the first plurality of picture quality versus coding efficiency characteristic curves. Based on the first plurality of picture quality versus coding efficiency characteristic curves and the corresponding second plurality of picture quality versus coding efficiency characteristic curves, calculating BD metric values respectively; and Calculating a weighted sum of the plurality of BD metric values as an overall Bjøntegaard delta (BD) metric for comparing the first coding scheme and the second coding scheme.

2. The method according to claim 1, further comprising: Calculating the BD metric by at least one of the following: mean average precision (mAP), bits per pixel (BPP), multi-object tracking accuracy (MOTA), average precision at an intersection over union (IoU) threshold of 50% (AP50), average precision at an IoU threshold of 75% (AP75), and average accuracy.

3. The method according to claim 1, further comprising: Calculating the BD metric as an average gap between the first plurality of picture quality versus coding efficiency curves of the first coding scheme and the second plurality of picture quality versus coding efficiency curves of the second coding scheme.

4. The method according to claim 1, wherein The method further comprises: Calculating a first Pareto front curve for the first plurality of picture quality versus coding efficiency curves of the first coding scheme; and Calculating a second Pareto front curve for the second plurality of picture quality versus coding efficiency curves of the second coding scheme; and Calculating the BD metric based on the first Pareto front curve and the second Pareto front curve.

5. The method according to any one of claims 1 to 4, wherein The first coding scheme and the second coding scheme are used for video coding for machine vision and human vision. The method comprises: Calculating a first BD rate for bits per pixel for machine vision; Calculating a second BD rate for human vision; and Calculating a weighted sum of the first BD rate and the second BD rate as an overall BD metric for comparing the first coding scheme and the second coding scheme.

6. The method according to any one of claims 1 to 4, wherein The first coding scheme and the second coding scheme are used for video coding for machine vision and human vision. The method comprises: Calculating a first overall distortion based on a weighted sum of distortions for the machine vision and the human vision by the first coding scheme; Calculating a first cost metric value based on the first overall distortion and first rate information of the first coding scheme; With the second coding scheme, calculate a second overall distortion based on a weighted sum of distortions for the machine vision and the human vision; Calculate a second cost metric value based on the second overall distortion and second rate information of the second coding scheme; and Compare the first coding scheme and the second coding scheme based on the first cost metric value and the second cost metric value.

7. The method according to any one of claims 1 to 4, wherein The first coding scheme and the second coding scheme are used for video coding for machine vision and human vision, and the method includes: Determine the first picture quality based on a weighted sum of distortions for the machine vision and the human vision through the first coding scheme; and Determine the second picture quality based on a weighted sum of distortions for the machine vision and the human vision through the second coding scheme.

8. The method according to any one of claims 1 to 4, wherein, The first coding scheme and the second coding scheme are used for video coding for multiple visual tasks, and the method includes: Determine the first picture quality based on a weighted sum of distortions for the multiple visual tasks through the first coding scheme; and Determine the second picture quality based on a weighted sum of distortions for the multiple visual tasks through the second coding scheme.

9. An apparatus for video encoding and decoding, the apparatus includes a processing circuit configured to: Determine a first picture quality versus encoding efficiency characteristic of a first encoding scheme for machine video coding (VCM), the first picture quality versus encoding efficiency characteristic including: A first plurality of picture quality versus coding efficiency curves of the first coding scheme in a two-dimensional plane for picture quality and coding efficiency; Determine a second picture quality versus coding efficiency characteristic of a second coding scheme for VCM, the second picture quality versus coding efficiency characteristic including: a second plurality of picture quality versus coding efficiency curves of the second coding scheme in the two-dimensional plane for the picture quality and coding efficiency; The second plurality of picture quality versus coding efficiency characteristic curves respectively correspond to the first plurality of picture quality versus coding efficiency curves, and calculate BD metric values respectively based on the first plurality of picture quality versus coding efficiency curves and the corresponding second plurality of picture quality versus coding efficiency curves; and Calculate a weighted sum of the plurality of BD metric values as an overall Bjøntegaard delta BD metric for comparing the first coding scheme and the second coding scheme.

10. The device according to claim 9, wherein, The processing circuit is configured to: Calculate the BD metric in at least one of the following: mean average precision mAP, bits per pixel BPP, multi-object tracking accuracy MOTA, average precision AP50 with an intersection over union threshold of 50%, average precision AP75 with an intersection over union threshold of 75%, and average accuracy.

11. The device according to claim 9, wherein, The processing circuit is configured to: Calculate the BD metric as an average gap between the first plurality of picture quality versus coding efficiency curves of the first coding scheme and the second plurality of picture quality versus coding efficiency curves of the second coding scheme.

12. The apparatus according to claim 11, wherein, And the processing circuit is configured to: Calculate a first Pareto front curve of the first plurality of picture quality versus coding efficiency curves for the first coding scheme; and Calculate a second Pareto front curve of the second plurality of picture quality and coding efficiency curves for the second coding scheme; And Calculate the BD metric based on the first Pareto front curve and the second Pareto front curve.

13. The device according to any one of claims 9 to 12, wherein, The first coding scheme and the second coding scheme are used for video coding for machine vision and human vision, and the processing circuit is configured to: Calculate a first BD rate of bits per pixel for machine vision; Calculate a second BD rate for human vision; And Calculate a weighted sum of the first BD rate and the second BD rate as an overall BD metric for comparing the first coding scheme and the second coding scheme.

14. The apparatus according to any one of claims 9 to 12, wherein, The first coding scheme and the second coding scheme are used for video coding for machine vision and human vision, and the processing circuit is configured to: Calculate a first overall distortion based on a weighted sum of distortions for the machine vision and the human vision through the first coding scheme; Calculate a first cost metric value based on the first overall distortion and first rate information of the first coding scheme; Calculate a second overall distortion based on a weighted sum of distortions for the machine vision and the human vision through the second coding scheme; Calculate a second cost metric value based on the second overall distortion and second rate information of the second coding scheme; And Compare the first coding scheme and the second coding scheme based on the first cost metric value and the second cost metric value.

15. The device according to any one of claims 9 to 12, wherein The first coding scheme and the second coding scheme are used for video coding for machine vision and human vision, and the processing circuit is configured to: Determine the first picture quality based on a weighted sum of distortions for the machine vision and the human vision through the first coding scheme; And Determine the second picture quality based on a weighted sum of distortions for the machine vision and the human vision through the second coding scheme.

16. The device according to any one of claims 9 to 12, wherein, The first coding scheme and the second coding scheme are used for video coding for multiple vision tasks, and the processing circuit is configured to: Determine the first picture quality based on a weighted sum of distortions for the multiple vision tasks through the first coding scheme; And Determine the second picture quality based on a weighted sum of distortions for the multiple vision tasks through the second coding scheme.

17. A method for storing a bitstream, characterized in that, The bitstream is encoded and decoded based on the video codec method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for detecting image compression coding efficiency, and storage medium

    CN108769685A