Computing system and transcoding method for transcoding
By predicting the transcoder bitrate based on the regression function of the video quality estimator, the problem of difficulty in determining the target bitrate during the transcoding process is solved, the transcoding efficiency is improved and the computational complexity is reduced.
Patent Information
- Application Number
- CN202110339647.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-30
- Filing Date
- 2021-03-30
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2041-03-30
AI Technical Summary
The existing technology has difficulty in effectively determining the target bit rate during the transcoding process, resulting in low compression efficiency and high computational complexity of the transcoder.
The transcoder bit rate is estimated by receiving the first encoded content and based on the regression on the video quality estimator, the target bit rate is predicted by using a linear or polynomial regression function, and an appropriate video quality estimator is selected in combination with encoder information matching to optimize the transcoding parameters.
The process of determining the bit rate of the transcoder is simplified, the compression efficiency of the transcoder is improved and the calculation complexity is reduced.
Smart Images

Figure CN113473146B_ABST
Abstract
Description
Background Art
[0001] Data compression is used to reduce the number of bits used to transmit and / or store video content, etc. Referring to FIG. 1 , a data compression system according to conventional techniques is shown. The data compression system may include an encoder 110 for compressing input video content 120 into encoded data 130 for streaming to one or more users. Encoded content 130 advantageously reduces bandwidth utilization for streaming the encoded content. The encoded content 130, received by the user device, may then be decoded 140 to generate decoded content 150. Encoding with a factor of more than ten is typically lossy, affecting the visual quality of the content. Coding parameters 160 can be used to adjust the data compression performed by encoder 120 to achieve a given bit rate, visual quality, encoding latency, etc. Typically, encoding parameter values that provide a given bit rate, visual quality, encoding latency, etc. are determined through a brute-force search. In a brute-force search, the input content is encoded using a given set of encoding parameter values, then decoded, and the decoded content is compared with the original input content to determine objective quality.
[0002] In some cases, the encoded data in the first format can be converted into the second encoding format. Now referring to Figure 2, a transcoder according to conventional technology is shown. A transcoder 210 is utilized to convert from a first encoded content 220 to a second encoded content 230. Transcoding may also be a lossy process and accumulate as encoding losses. Transcoding parameters 240 can be used to adjust the transcoding process to achieve a given bit rate, visual quality, transcoding latency, etc. Typically, the original input content is not available for a brute force search of transcoding parameters. Brute force determination of the optimal target bit rate for a transcoder typically includes multiple encodings and objective quality measurements by a predictive model, which is computationally complex. Alternatively, a default bit rate higher than required can be utilized, which reduces the compression efficiency of the transcoder. Therefore, there is a continuing need for improved target bit rate determination techniques for transcoding. Summary of the Invention
[0003] The present technology may be best understood by referring to the following description and accompanying drawings, which illustrate embodiments of the present technology for transcoder bitrate prediction.
[0004] In one embodiment, a computing system may include one or more processors, a memory, and one or more transcoders. Instructions stored in the memory may cause the processors to perform a video encoding method, the method comprising: receiving first encoded content; and estimating a transcoder bitrate based on regression of a video quality estimator on the first encoded content and second encoded content. The video encoder may be configured to convert the first encoded content into second encoded content based on one or more transcoding parameter values, including the estimated transcoder bitrate.
[0005] In another embodiment, a transcoder method may include receiving first encoder content and an encoder information set for the first encoded content. A determination may be made as to whether the received encoder information set matches one of a plurality of existing encoder information sets. When the received encoder information does not match one of the plurality of existing encoder information sets, a predetermined target bit rate may be selected as a transcoder bit rate. When the received encoder information matches one of the plurality of existing encoder information sets, a video quality estimator of a corresponding one of the plurality of existing encoder information sets that matches the received encoder information set may be selected. When the received encoder information matches one of the plurality of existing encoder information sets, a transcoder bit rate may be estimated based on a regression of the matched video quality estimator on the first encoded content and the second encoded content. The first encoded content may be transcoded into the second encoded content based on one or more transcoding parameter values including the transcoder bit rate.
[0006] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] There are illustrated by way of example and not by way of limitation embodiments of the present technology in the figures of the accompanying drawings and in which like reference numerals refer to similar elements and in which:
[0008] FIG1 shows a data compression system according to conventional technology.
[0009] FIG2 shows a transcoding system according to conventional technology.
[0010] Figure 3 A transcoder method in accordance with aspects of the present technique is shown.
[0011] Figure 4 A transcoder method in accordance with aspects of the present technique is shown.
[0012] Figure 5 A computing system for transcoding content in accordance with aspects of the present technology is shown.
[0013] Figure 6 An exemplary processing core in accordance with aspects of the present technique is shown. DETAILED DESCRIPTION
[0014] Reference will now be made in detail to embodiments of the present technology, examples of which are illustrated in the accompanying drawings. Although the present technology will be described in conjunction with these embodiments, it will be understood that they are not intended to limit the present technology to these embodiments. On the contrary, the present invention is intended to encompass alternatives, modifications, and equivalents that may be included within the scope of the present invention as defined by the appended claims. In addition, in the following detailed description of the present technology, many specific details are set forth in order to provide a thorough understanding of the present technology. However, it will be understood that the present technology can be practiced without these specific details. In other cases, well-known methods, processes, components, and circuits have not yet been described in detail to avoid unnecessarily confusing various aspects of the present technology.
[0015] The following embodiments of the present technology are presented in terms of routines, modules, logic blocks, and other symbolic representations of the operations performed on data in one or more electronic devices. These descriptions and representations are the means by which the essence of their work is most effectively conveyed to other technical personnel in the art by those skilled in the art. Routine, module, logic block, and / or the like are generally conceived herein as a self-consistent sequence of processes or instructions that result in a desired result. A process is one that includes the physical manipulation of a physical quantity. Typically, although not necessarily, these physical manipulations take the form of electrical or magnetic signals that can be stored, transferred, compared, and otherwise manipulated in an electronic device. For convenience, and with reference to common usage, these signals are referred to as data, bits, values, elements, symbols, characters, terms, digits, character strings, and / or the like with reference to the embodiments of the present technology.
[0016] However, it should be kept in mind that these terms are to be interpreted as referring to physical operations and quantities and are merely convenient labels and are to be further interpreted in light of terminology commonly used in the art. Unless specifically stated otherwise as apparent from the following discussion, it should be understood that throughout the discussion of the present technology, discussions utilizing terms such as "receiving" and / or the like refer to actions and processes of electronic devices, such as electronic computing devices, that manipulate and transform data. Data is represented as physical (e.g., electronic) quantities within the logic circuits, registers, memories, and / or the like of the electronic device and is transformed into other data similarly represented as physical quantities within the electronic device.
[0017] In this application, the use of disjunctives is intended to include conjunctions. The use of definite or indefinite articles is not intended to indicate cardinality. In particular, reference to an object "the" or "an" object is intended to also represent one of a plurality of possible such objects. The use of the terms "comprise," "comprises," "includes," "contains," etc. specifies the presence of a stated element, but does not exclude the presence or addition of one or more other elements and or groups thereof. It should also be understood that although the terms first, second, etc. may be used herein to describe various elements, such elements should not be limited by these terms. These terms are used herein to distinguish one element from another. For example, without departing from the scope of the embodiment, the first element can be referred to as the second element, and similarly, the second element can be referred to as the first element. It should also be understood that when an element is referred to as "coupled" to another element, it can be directly or indirectly connected to the other element, or there can be an intermediate element. In contrast, when an element is referred to as "directly connected" to another element, there is no intermediate element. It should also be understood that the term "and or" includes any and all combinations of one or more of the associated elements. Also, it is to be understood that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting.
[0018] Now refer to Figure 3 , shows a transcoder method according to aspects of the present technology. The method may include receiving first encoder content at 310. In one implementation, the first encoder content may be video content encoded in a first encoder format.
[0019] At 320, a transcoder bitrate may be estimated based on a regression of the first encoded content and the second encoded content on a video quality estimator. The second encoded content may be a transcode of the first video content. The regression may be a linear regression or a polynomial regression of the first encoded content and the second encoded content on the video quality estimator. In one implementation, the transcoding bitrate may be estimated based on a linear regression on a video multi-method assessment fusion (VMAF) estimator. For example, a predetermined VMAF (Pred_VMAF) may be a function of the bitrate of the input encoded content (e.g., the first encoded content) and the bitrate of the output encoded content (e.g., the second encoded content) for the transcoder as expressed by Equation 1:
[0020] Pred_VMAF=f(bitrate_in_video, bitrate_out_video)
[0021] =a*ln(bitrate_in_video)+b*ln(bitrate_out_video) (1)
[0022] Where bitrate_in_video is the bitrate of the input encoded content and bitrate_out_video is the bitrate of the output encoded content. The target VMAF (Target_VMAF) may be a function of the bitrate of the input encoded content (e.g., the first encoded content) and the bitrate of the output encoded content (e.g., the second encoded content) for the transcoder as expressed by Equation 2:
[0023] Target_VMAF=f(bitrate_in_video, bitrate_out_video)
[0024] =a*ln(bitrate_in_video)+b*ln(bitrate_out_video)+c (2) The target bit rate (Bitrate_out_video) can be expressed according to Equation 3:
[0025]
[0026] In another example, the transcoding bitrate may be estimated based on a polynomial regression on a VMAF estimator. For example, the predetermined VMAF (Pred_VMAF) may be a function of the bitrate of the input encoded content (e.g., the first encoded content) and the bitrate of the output encoded content (e.g., the second encoded content) for the transcoder as expressed by Equation 4:
[0027] Pred_VMAF=f(bitrate_in_video, bitrate_out_video)
[0028] =a*x+b*x 2 +c*y+d*y 2 +f (4)
[0029] Where bitrate_in_video is the bitrate of the input encoded content, bitrate_out_video is the bitrate of the output encoded content, x=ln(bitrate_in_video), and y=ln(bitrate_out_video). The target VMAF (Target_VMAF) may be a function of the bitrate of the input encoded content (e.g., the first encoded content) and the bitrate of the output encoded content (e.g., the second encoded content) for the transcoder as expressed by Equation 5:
[0030] Target_VMAF=f(bitrate_in_video, bitrate_out_video)
[0031] =a*x+b*x 2+c*y+d*y 2 +e (5)
[0032] The target bit rate (Bitrate_out_video) can be expressed according to Equation 6:
[0033]
[0034] where f = e + a*x 2 +b*x-Target_VMAF
[0035] At 330 , the first encoded content may be transcoded into second encoded content based on one or more transcoding parameter values including the estimated transcoder bit rate.
[0036] Now refer to Figure 4 , shows another transcoder method according to aspects of the present technology. The method may include receiving an encoder information set of a first encoded content at 410. In one implementation, the encoder information may include an encoder type, encoding parameters, etc. of the encoder.
[0037] At 420, a determination may be made as to whether the received encoder information set matches one of a plurality of existing encoder information sets. In one implementation, the plurality of existing encoder information sets and corresponding video quality estimators may be determined by training a plurality of different video quality estimators on different encoding information sets. For example, a plurality of VMAF estimators may be trained based on a plurality of different encoding information sets. At 430, when the received encoder information does not match one of the plurality of existing encoder information sets, a predetermined target bitrate may be selected as the transcoder bitrate. In one implementation, the predefined target bitrate may be a default target bitrate with high redundancy.
[0038] At 440, when the received encoder information matches a corresponding one of the plurality of existing encoder information sets, a video quality estimator of the corresponding one of the plurality of existing encoder information sets may be selected, for example, a VMAF estimator of the corresponding one of the plurality of VMAF estimators that matches the encoding information of the encoder.
[0039] At 450, a first encoded content for the received encoder information may be received by the encoder. At 460, a transcoder bit rate may be estimated based on a regression of the first encoded content and the second encoded content on the matched video quality estimator. The second encoded content may be a transcode of the first video content. The regression may be a linear regression or a polynomial regression of the first encoded content and the second encoded content on the video quality estimator. In one implementation, the transcoding bit rate may be estimated based on a linear regression on a VMAF estimator. For example, the predetermined VMAF (Pred_VMAF) may be a function of the bit rate of the input encoded content (e.g., the first encoded content) and the bit rate of the output encoded content (e.g., the second encoded content) for the transcoder as expressed by Equation 1:
[0040] Pred_VMAF=f(bitrate_in_video, bitrate_out_video)
[0041] =a*ln(bitrate_in_video)+b*ln(bitrate_out_video) (1)
[0042] Where bitrate_in_video is the bitrate of the input encoded content and bitrate_out_video is the bitrate of the output encoded content. The target VMAF (Target_VMAF) may be a function of the bitrate of the input encoded content (e.g., the first encoded content) and the bitrate of the output encoded content (e.g., the second encoded content) for the transcoder as expressed by Equation 2:
[0043] Target_VMAF=f(bitrate_in_video, bitrate_out_video)
[0044] =a*ln(bitrate_in_video)+b*ln(bitrate_out_video)+c (2)
[0045] The constants a, b, and c can be obtained by training multiple different video quality estimators on different sets of coding information. The target bit rate (Bitrate_out_video) can be expressed according to Equation 3:
[0046]
[0047] In another example, the transcoding bitrate may be estimated based on a polynomial regression on a VMAF estimator. For example, the predetermined VMAF (Pred_VMAF) may be a function of the bitrate of the input encoded content (e.g., the first encoded content) and the bitrate of the output encoded content (e.g., the second encoded content) for the transcoder as expressed by Equation 4:
[0048] Pred_VMAF=f(bitrate_in_video, bitrate_out_video)
[0049] =a*x+b*x 2 +c*y+d*y 2 +f (4)
[0050] Where bitrate_in_video is the bitrate of the input encoded content, bitrate_out_video is the bitrate of the output encoded content, x=ln(bitrate_in_video), and y=ln(bitrate_out_video). The target VMAF (Target_VMAF) may be a function of the bitrate of the input encoded content (e.g., the first encoded content) and the bitrate of the output encoded content (e.g., the second encoded content) for the transcoder as expressed by Equation 5:
[0051] Target_VMAF=f(bitrate_in_video, bitrate_out_video)
[0052] =a*x+b*x 2 +c*y+d*y 2 +e (5)
[0053] The constants a, b, c, d, and e can be obtained by training multiple different video quality estimators on different sets of coding information. The target bit rate (Bitrate_out_video) can be expressed according to Equation 6:
[0054]
[0055] where f = e + a*x 2 +b*x-Target_VMAF
[0056] At 470, the first encoded content may be transcoded into the second encoded content based on one or more transcoding parameter values including a transcoder bitrate. For example, when the received encoder information does not match one of a plurality of existing encoder information sets, the first encoded content may be transcoded into the second encoded content at a predetermined bitrate. When the received encoder information matches a corresponding one of the plurality of existing encoder information sets, the first encoded content may be transcoded into the second encoded content at an estimated transcoder bitrate based on regressing the first and second encoded content on the matched video quality estimator.
[0057] Now refer to Figure 5 , illustrates a computing system for transcoding content in accordance with aspects of the present technology. Computing system 500 may include one or more processors 502, memory 504, and one or more transcoders 506. One or more transcoders 506 may be implemented in separate hardware or in software executed on one or more processors 502. In one implementation, computing system 500 may be a server computer, a data center, a cloud computing system, a streaming service system, an internet service provider system, a cellular service provider system, or the like.
[0058] The one or more processors 502 may be a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a vector processor, a memory processing unit, etc., or a combination thereof. In one implementation, the processor 502 may include peripheral components such as a PCIe 4 interface (PCIe4) 508 and an integrated circuit (IC) interface. 2 C) One or more communication interfaces of interface 510, an on-chip circuit tester such as a Joint Test Action Group (JTAG) engine 512, a direct memory access engine 514, a command processor (CP) 516, and one or more cores 518-524. One or more cores 518-524 may be coupled in a directional ring bus configuration.
[0059] Now refer to Figure 6, shows an exemplary processing core according to various aspects of the present technology. Processing core 600 may include a tensor engine (TE) 610, a pooling engine (PE) 615, a memory copy engine (ME) 620, a sequencer (SEQ) 625, an instruction buffer (IB) 630, a local memory (LM) 635, and a constant buffer (CB) 640. The local memory 635 may be pre-installed with model weights and may store activations in use in a timely manner. The constant buffer 640 may store constants for batch normalization, quantization, etc. The tensor engine 610 may be used to accelerate fused convolutions and / or matrix multiplications. The pooling engine 615 may support operations such as pooling, interpolation, and regions of interest. The memory copy engine 620 may be configured for inter-core and / or intra-core data copying, matrix transposition, etc. The tensor engine 610, the pooling engine 615, and the memory copy engine 620 may run in parallel. The sequencer 625 can orchestrate the operations of the tensor engine 610, the pooling engine 615, the memory copy engine 620, the local memory 635, and the constant buffer 640 according to the instructions from the instruction buffer 630. The processing core 600 can provide efficient computation for video compilation under the control of the operation fusion coarse-grained instructions. A detailed description of the exemplary processing unit core 600 is not necessary to understand various aspects of the present technology and will not be described further herein.
[0060] Reference again Figure 5 , one or more cores 518-524 may execute one or more computing device-executable instruction sets to perform one or more functions, including, but not limited to, a transcoder bitrate predictor 526. The one or more functions may be executed on an individual core 518-524, may be distributed across multiple cores 518-524, may be executed in conjunction with one or more other functions on one or more cores, and or the like.
[0061] The transcoder bitrate predictor 526 may be configured as described above with reference to Figure 3 and Figure 4 The one or more processors 502 may output transcoder parameters 528 including the transcoder bitrate to the one or more transcoders 506. The one or more transcoders 506 may use the transcoder parameters 528 including the transcoder bitrate to convert the first encoded content into the second content 532.
[0062] Aspects of the present technology can advantageously estimate a minimum target bitrate to be used as a transcoder bitrate that meets a predetermined quality. The estimated transcoder bitrate can advantageously improve the compression efficiency of the transcoder. Aspects of the present technology can also advantageously eliminate the need for brute-force determination of the optimal transcoder bitrate, thereby reducing computational complexity.
[0063] The foregoing descriptions of specific embodiments of the present technology have been presented for purposes of illustration and description. They are not intended to be exhaustive or to limit the technology to the precise form disclosed, and obviously, many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best illustrate the principles of the technology and its practical application, thereby enabling others skilled in the art to best utilize the technology and various embodiments with various modifications as are suitable for the particular use contemplated. It is intended that the scope of the invention be defined by the claims appended hereto and their equivalents.
Claims
1. A computing system, comprising: one or more processors; One or more non-transitory computing device-readable storage media storing computing executable instructions that, when executed by the one or more processors, perform a method comprising: Receiving input encoded content; and estimating a transcoder bitrate based on a Video Multi-Method Assessment Fusion (VMAF) estimator: Where bitrate_in_video is the bitrate of the input encoded content, bitrate_out_video is the bitrate of the output encoded content, Target_VMAF is the function of the input encoded content bitrate and the output encoded content bitrate used by the transcoder, a, b, and c are constants obtained by training multiple different video quality estimators on different encoding information sets; A transcoder is configured to convert the input encoded content into output encoded content based on one or more transcoding parameter values including the estimated transcoder bit rate.
2. The computing system of claim 1 , wherein: The input encoded content includes first encoded video content; and The output encoded content includes second encoded video content.
3. The computing system of claim 1 , wherein the method further comprises: Training multiple video multi-method assessment fusion (VMAF) estimators based on a variety of different encoding information; as well as A video multi-method assessment fusion (VMAF) estimator is selected that matches the encoding information of the input encoded content from among the plurality of video multi-method assessment fusion (VMAF) estimators.
4. The computing system of claim 3, wherein the method further comprises: Receiving encoder information of the input encoded content; determining whether the received encoder information matches one of a plurality of existing encoder information sets; selecting a predetermined target bit rate as a transcoder bit rate when the received encoder information does not match one of the plurality of existing encoder information sets; A video quality estimator is selected from the plurality of existing encoder information sets, the video quality estimator matching the received encoder information.
5. A transcoding method, comprising: receiving input encoded content and an encoder information set of a first input encoded content; determining whether the received encoder information set matches one of a plurality of existing encoder information sets; selecting a predetermined target bit rate as a transcoder bit rate when the received encoder information set does not match one of the plurality of existing encoder information sets; selecting a video quality estimator corresponding to one of the plurality of existing encoder information sets that matches the received encoder information set; Estimate the transcoder bitrate based on the Video Multi-method Assessment Fusion (VMAF) estimator: Where bitrate_in_video is the bitrate of the input encoded content, bitrate_out_video is the bitrate of the output encoded content, Target_VMAF is the function of the input encoded content bitrate and the output encoded content bitrate used by the transcoder, a, b, and c are constants obtained by training multiple different video quality estimators on different encoding information sets; as well as The input encoded content is transcoded into output encoded content based on one or more transcoding parameter values including an estimated transcoder bit rate.
6. The method according to claim 5, wherein: The input encoded content includes video content encoded in a first encoding format; and The output encoded content includes video content encoded in a second encoding format.
7. The method according to claim 5, further comprising: Training multiple video quality estimators based on multiple different encoding information; as well as A corresponding one of the plurality of video quality estimators that matches the encoding information of the input encoded content is selected.
8. A transcoding method, comprising: Receive input encoded content; Estimate the transcoder bitrate based on the Video Multi-method Assessment Fusion (VMAF) estimator: Where bitrate_in_video is the bitrate of the input encoded content, bitrate_out_video is the bitrate of the output encoded content, Target_VMAF is the function of the input encoded content bitrate and the output encoded content bitrate used by the transcoder, a, b, and c are constants obtained by training multiple different video quality estimators on different encoding information sets; A transcoder is configured to convert the input encoded content into output encoded content based on one or more transcoding parameter values including the estimated transcoder bit rate.
9. The method according to claim 8, further comprising: Training multiple video quality estimators based on multiple different encoding information; as well as A corresponding one of the plurality of video quality estimators that matches the encoding information of the input encoded content is selected.
Citation Information
Patent Citations
Transcoding of data
US20040223650A1
Method and apparatus of content-based self-adaptive video transcoding
US20150350726A1