Video data processing method and device, electronic equipment and storage medium

By introducing the collaborative work of the neural processing unit and the first processing unit in the video data processing, the high power consumption and high delay problems in the decoding process of the advanced video compression scheme are solved, and video encoding and decoding processing with low power consumption and low latency is realized.

CN120455678APending Publication Date: 2025-08-08BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510726493.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The decoding process of advanced video compression schemes has high computational complexity and high power consumption, especially on mobile devices, there are problems such as rendering resource competition and performance jitter.

Method used

The neural processing unit and the first processing unit are used to jointly construct the reconstruction image, and the data processing pressure of the first processing unit is shared by the neural processing unit to reduce power consumption and delay.

Benefits of technology

It significantly reduces the power consumption and delay of video data processing, and improves the video encoding and decoding processing efficiency of mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455678A_ABST
    Figure CN120455678A_ABST
Patent Text Reader

Abstract

The invention discloses a video data processing method and device, electronic equipment and a storage medium. The video data processing method comprises the following steps: in response to conversion between a video frame of a video and a bit stream of the video, in a process of constructing a reconstructed image corresponding to the video frame, acquiring intermediate data obtained by performing a first processing operation by a first processing unit; and performing a second processing operation on the intermediate data by a neural processing unit to obtain a reconstructed image. The video data processing method can reduce power consumption and delay of video data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a video data processing method and device, an electronic device, and a storage medium. Background Art

[0002] Advanced video compression schemes, with their extremely high compression ratios, can significantly reduce the bandwidth required for video transmission, effectively alleviating network load. However, the decoding process faces technical challenges such as high computational complexity and high power consumption. Decoding is an essential step in encoding video into a stream or presenting it to end users. For example, software decoding is often used throughout the entire process of presenting video content to end users. After software decoding, post-processing operations such as video quality enhancement are often required to enhance the visual experience. Summary of the Invention

[0003] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] At least one embodiment of the present disclosure provides a video data processing method, comprising: in response to conversion between a video frame of a video and a bit stream of the video, in a process of constructing a reconstructed image corresponding to the video frame, obtaining intermediate data obtained by a first processing unit performing a first processing operation; and performing a second processing operation on the intermediate data by a neural processing unit to obtain the reconstructed image.

[0005] At least another embodiment of the present disclosure provides a video data processing device, including: a first processing unit, configured to, in response to conversion between a video frame of a video and a bit stream of the video, obtain intermediate data obtained by performing a first processing operation in a process of constructing a reconstructed image corresponding to the video frame; and a neural processing unit, configured to perform a second processing operation on the intermediate data to obtain a reconstructed image.

[0006] At least another embodiment of the present disclosure provides an electronic device, comprising: a processor; and a memory, comprising one or more computer program instructions; the processor comprises a first processing unit and a neural processing unit, wherein the one or more computer program instructions are executed by the processing device in response to conversion between a video frame of a video and a bit stream of the video when the processing device is running, and in the process of constructing a reconstructed image corresponding to the video frame, intermediate data obtained by a first processing operation performed by the first processing unit is obtained; and the neural processing unit performs a second processing operation on the intermediate data to obtain the reconstructed image.

[0007] At least another embodiment of the present disclosure provides a computer-readable storage medium that non-temporarily stores computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, conversion between a video frame in response to a video and a bit stream of the video is implemented, and in the process of constructing a reconstructed image corresponding to the video frame, intermediate data obtained by a first processing operation performed by a first processing unit is obtained; and a second processing operation is performed on the intermediate data by a neural processing unit to obtain the reconstructed image. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0009] Figure 1 A frame diagram of an encoder is shown;

[0010] Figure 2 A frame diagram of a decoder is shown;

[0011] Figure 3 A flow chart illustrating a decoding method for presenting a video to an end user is shown;

[0012] Figure 4 A flowchart of a video data processing method provided by at least one embodiment of the present disclosure is shown;

[0013] Figure 5 A schematic diagram of a process for decoding a bit stream into a video image according to at least one embodiment of the present disclosure is shown;

[0014] Figure 6 A schematic diagram showing data flow between a CPU and an NPU provided by at least one embodiment of the present disclosure is shown;

[0015] Figure 7 An architectural diagram for displaying video on a mobile terminal provided by at least one embodiment of the present disclosure is shown;

[0016] Figure 8 A block diagram of a video data processing device provided by at least one embodiment of the present disclosure is shown;

[0017] Figure 9 A schematic structural diagram of an electronic device provided by at least one embodiment of the present disclosure is shown; and

[0018] Figure 10 A schematic structural diagram of another electronic device provided by at least one embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0019] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0020] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0021] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0022] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0023] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0024] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0025] Video coding, also known as video compression, is a method of converting the original video signal into a video signal that uses fewer bits than the uncompressed video signal. For example, it removes various redundant information in the video signal. The compressed video signal data can usually be expressed as a "code stream" or "bit stream" for transmission, storage, etc. Video decoding is the process of restoring the encoded and compressed video signal data (such as "code stream" or "bit stream") into video frames that can be displayed on a display device for users to watch. Through decoding, the video signal data compressed during encoding is restored to uncompressed video signal data or video signal data with a lower compression ratio, so that the video signal can be played normally on terminal devices (such as TVs, computers, and mobile phones). The video coding and decoding scheme enables the code stream files that conform to the scheme, which are encoded using different video encoders, to be decoded normally by decoders that conform to the scheme.

[0026] For example, a video codec solution supports multiple prediction modes to reduce spatial redundancy (intra-frame prediction) and temporal redundancy (inter-frame prediction). For example, a 16x16 pixel macroblock is used as the basic processing unit; each macroblock can be further divided into smaller sub-blocks (such as 8x8, 8x16, 16x8, etc.) to meet different motion estimation requirements. This video codec solution may include the following five parts: inter-frame and intra-frame prediction (Estimation), transform and inverse transform, quantization and inverse quantization, loop filter (Loop Filter) / deblocking filter (Deblocking Filter), entropy coding, etc.

[0027] For the aforementioned video codec solutions, for example, more flexible concepts such as Coding Unit (CU), Prediction Unit (PU), and Transform Unit (TU) can be introduced. For loop filtering, technologies such as Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) can be further adopted. This approach aims to reduce the bitrate by approximately 50% while reconstructing the same video quality, or reduce coding distortion at the same bitrate, significantly improving video quality. This approach is suitable for HD and UHD video scenarios. For example, when streaming 4K content, it can provide clearer image quality with less bandwidth usage.

[0028] Furthermore, for the above-mentioned video coding and decoding solutions, for example, new coding tools and technologies can be introduced to improve performance, such as more efficient inter-frame prediction, intra-frame prediction, transformation technology, loop filters, etc., to further improve video compression efficiency while supporting a wider range of video application requirements, such as 4K, 8K ultra-high-definition video, 360-degree panoramic video, etc.

[0029] Figure 1 A framework diagram of an encoder is shown.

[0030] like Figure 1 As shown in the figure, the input video signal is first divided into one or more coding tree units (CTUs). The CTU serves as the basic processing unit of the encoder's video codec scheme. Subsequent operations are performed on the CTU and its subunits. For example, the encoder can split the CTU to form one or more coding units (CUs). Such CUs are also commonly called "blocks" or "video blocks." Subsequent operations can be performed on the CUs.

[0031] The input video signal is divided into two paths and used for residual calculation (subtracted from the prediction signal to generate a residual signal). The residual signal is the actual input to the subsequent transform / quantization module. By transforming and quantizing the residual signal, redundant information is removed, achieving video data compression. For example, the residual signal is transformed using a discrete cosine transform (DCT) or discrete sine transform (DST) to convert it from the spatial domain to the transform domain (frequency domain). This transform converts the correlation between adjacent pixels into the correlation between coefficients in the frequency domain, creating conditions for subsequent operations such as quantization and effectively removing spatial redundancy. After the residual signal is transformed, the transformed coefficients can be quantized to adjust the amplitude of the frequency domain coefficients, effectively reducing the signal value space and achieving better compression. For example, the transformed coefficients can be mapped to a finite number of quantization values according to a certain quantization step size to achieve data compression, remove some redundant information, and output the quantized coefficients.

[0032] Next, the quantized coefficients undergo inverse quantization and inverse transformation in the inverse quantization / inverse transform module to obtain a reconstructed residual signal. This reconstructed residual signal is then added to the predicted signal to produce a reconstructed frame. This provides a reference image for subsequent filter control analysis and other encoding steps (such as intra-frame prediction and motion compensation). Inverse quantization and inverse transformation are necessary steps to restore the original signal. The quantized coefficients undergo inverse quantization to restore their approximate values in the transform domain, and then undergo an inverse transform to convert them back to the spatial domain.

[0033] The reconstructed frame enters the filter control and analysis module, which analyzes the reconstructed frame and determines the parameters required for subsequent filtering. Examples of subsequent filtering include deblocking and sample adaptive offset (SAO) filtering. Deblocking and SAO filtering can reduce blocking artifacts and enhance image quality. The processed image is stored in the decoded image buffer, providing a reference for subsequent inter-frame prediction.

[0034] The input video signal participates in inter-frame prediction in one channel. Using historical reference frames in the decoded image buffer, motion estimation and motion compensation are used to generate a prediction signal. The intra / inter mode selection determines which prediction method is used to optimize coding efficiency.

[0035] Motion estimation, for example, is to find the block that is most similar to the current block, i.e., the matching block, based on the historical reference frames (decoded and reconstructed image frames) stored in the decoded image cache; calculate the motion vector based on the matching block and the current block, and the relative displacement between the matching block and the current block is the motion vector (MV).

[0036] Motion compensation is, for example, extracting a block at a corresponding position from a historical reference frame in a decoded image buffer according to a motion vector obtained by motion estimation, and using the extracted reference block as a prediction signal for the current block.

[0037] Intra-frame prediction is used to generate a prediction signal based on the best mode selected by intra-frame estimation, using decoded pixels within the current frame. This process relies entirely on information within the current frame, without reference to other frames, effectively removing spatial redundancy from the image and improving coding efficiency. Furthermore, intra-frame estimation can be used, for example, to select the optimal intra-frame prediction mode for intra-frame prediction. Intra-frame estimation evaluates the prediction performance of different intra-frame prediction modes for the current block, identifying the optimal intra-frame prediction mode by calculating metrics such as prediction error under different intra-frame prediction modes.

[0038] Various types of coding data (quantized coefficients, motion data, intra-frame prediction data, filter control data, coding control data, etc.) are compressed and encoded by the entropy coding module, and finally the encoded bit stream is generated to complete the video coding process.

[0039] Figure 2 A decoder framework diagram is shown.

[0040] The code stream file received by the decoder (for example, Figure 1The bitstream encoded by the encoder is entropy decoded by the entropy decoder, then inversely transformed and dequantized to produce a reconstructed residual signal. Based on the entropy-decoded data, inter-frame prediction or intra-frame prediction information is performed to obtain a prediction signal. The reconstructed residual signal and the prediction signal are added together to obtain a reconstructed frame. Deblocking filtering and sample adaptive offset are then performed to obtain the optimized decoding result.

[0041] For example, an entropy decoder first performs entropy decoding on the encoded bitstream using an algorithm such as CABAC (Context-Adaptive Binary Arithmetic Coding). This process restores the encoded bitstream to its original encoded data form, including quantized coefficients. The quantized coefficients obtained from entropy decoding are then dequantized based on the quantization parameters used by the encoder to restore their original amplitudes. An inverse transform (such as an inverse discrete cosine transform) is then performed on the dequantized coefficients, converting the frequency domain coefficients back to the spatial domain to produce a reconstructed residual signal.

[0042] The prediction method used is determined by the intra / inter mode selection. If intra prediction is used, the current block is predicted using the decoded pixel information within the current frame. Based on the intra prediction mode information transmitted by the encoder, the appropriate mode is selected from multiple prediction directions (such as horizontal and vertical) to generate a prediction signal. If inter prediction is used, the corresponding position is found in the historical reference frames in the decoded image buffer based on the motion information such as the motion vector provided by the encoder, and the prediction signal is generated through motion compensation.

[0043] The prediction signal is added to the reconstructed residual signal to obtain a reconstructed frame. The reconstructed frame is then subjected to deblocking filtering and sample adaptive offset to obtain a reconstructed image. Continuous reconstructed images are used to obtain a video signal for display.

[0044] For example, Figure 1 and Figure 2 Tasks such as inverse transformation, inverse quantization, and filtering are typically performed by the central processing unit (CPU) via software, resulting in high power consumption and high CPU software computational complexity. Currently, mainstream playback platforms utilize a graphics processing unit (GPU) for post-processing after CPU software decoding. GPUs are processing devices specifically designed to accelerate the creation and rendering of images, videos, and animations (such as games).

[0045] Figure 3 A flow chart of a decoding method for presenting a video to an end user is shown.

[0046] like Figure 3As shown, in response to a terminal device (such as a mobile phone or tablet) acquiring a video bitstream, the CPU performs software decoding on the bitstream to obtain decoded data, and then the GPU performs post-processing on the decoded data to obtain a reconstructed image. Next, the GPU performs operations such as format conversion on the reconstructed image to display the reconstructed image on a display screen.

[0047] exist Figure 3 In the decoding method shown, the CPU's software decoding is computationally complex and consumes significant power. Furthermore, the GPU's post-processing experience significant latency, leading to performance jitter caused by competition for rendering resources and high GPU power consumption. Furthermore, GPUs are only suitable for processing algorithms with high computational density and simple logic, making them unsuitable for mobile algorithm optimization.

[0048] To this end, at least one embodiment of the present disclosure provides a video data processing method, comprising: in response to conversion between a video frame and a video bitstream, in the process of constructing a reconstructed image corresponding to the video frame, obtaining intermediate data obtained by a first processing unit performing a first processing operation; and performing a second processing operation on the intermediate data by a neural processing unit to obtain a reconstructed image. This video processing method can significantly reduce power consumption and latency by jointly constructing the reconstructed image through the neural processing unit and the first processing unit, with the neural processing unit sharing the data processing pressure of the first processing unit.

[0049] For example, the video data processing method provided by the embodiments of the present disclosure may be applicable to terminals including a neural processing unit (e.g., mobile phones, tablet computers, laptop computers, wearable devices, etc.), and may be applicable to the above-mentioned various exemplary video encoding and decoding schemes or similar hybrid encoding video encoding and decoding schemes, and the embodiments of the present disclosure are not limited thereto. The embodiments of the present disclosure are applied, for example, to scenarios such as mobile live broadcasting, real-time video conferencing, and AR / VR interaction, where mobile devices need to achieve low-latency video encoding and decoding processing under limited power consumption.

[0050] Figure 4 A flowchart of a video data processing method provided by at least one embodiment of the present disclosure is shown.

[0051] like Figure 4 As shown, the video data processing method includes step S401 and step S402.

[0052] Step S401: in response to conversion between a video frame and a bit stream of the video, in the process of constructing a reconstructed image corresponding to the video frame, intermediate data obtained by a first processing operation performed by a first processing unit is obtained.

[0053] Step S402: The neural processing unit performs a second processing operation on the intermediate data to obtain a reconstructed image.

[0054] For step S401, the processed video can be a video shot in real time by the camera of the terminal device, a video downloaded from the Internet (such as streaming media), or a locally stored video, etc. For example, it can be a low dynamic range (LDR) video, a standard dynamic range (SDR) video, etc. The embodiments of the present disclosure do not impose any restrictions on this.

[0055] The conversion between a video frame and a video bitstream includes encoding the video into a bitstream, or decoding the bitstream to obtain a video frame. For example, the conversion between a video frame and a bitstream may include encoding the current block into a bitstream, or may include decoding the current block from the bitstream. That is, in different cases, the conversion process may include an encoding process or a decoding process, which is not limited in the embodiments of the present disclosure.

[0056] The encoding process may include the process of constructing a reconstructed image corresponding to a video frame. Figure 1 In the encoder shown, the process of constructing a reconstructed image corresponding to a video frame includes inverse quantization, inverse transformation, filtering, etc. For example, in the encoder, the constructed reconstructed image corresponding to the video frame is stored in a decoded image buffer as a historical reference frame.

[0057] During decoding, the process of constructing a reconstructed image corresponding to a video frame includes, for example, Figure 2 The process described.

[0058] In some embodiments of the present disclosure, the first processing unit may be a general-purpose processor, such as a central processing unit (CPU). A general-purpose processor refers to a processor designed to perform a wide range of tasks, rather than a processor optimized for a specific application or computing type. For example, a general-purpose processor is commonly used in personal computers, servers, and other computing devices that need to process a variety of different types of tasks. A general-purpose processor has a powerful computing core that can perform complex logical operations, system management, and other tasks, and can efficiently process a single or a small number of complex tasks. On the other hand, its performance is relatively low when processing large-scale data parallel computing.

[0059] The embodiments of the present disclosure do not impose any restrictions on the instruction set (such as X86, ARM, RISC-V, MIPS, etc.) and microarchitecture adopted by the CPU.

[0060] The first processing operation in step S401 includes, for example, at least one of entropy decoding, inverse quantization, inverse transformation, inter-frame prediction, or intra-frame prediction.

[0061] For the example of the first processing operation mentioned above, the general-purpose processor can rely on its more powerful logical operations and complex instruction processing capabilities to accurately handle complex tasks such as syntax parsing and entropy decoding in video decoding, ensuring the accuracy and stability of decoding. In addition, the general-purpose processor is highly versatile and can support decoding of various different video encoding formats through software algorithms.

[0062] In some embodiments of the present disclosure, the intermediate data includes reconstructed frames of a code stream, where the code stream is a bit stream obtained by encoding video frames.

[0063] For example, using Figure 1 The encoder described encodes the video frame to obtain a code stream of the video frame, which is, for example, a bit stream. Figure 2 During the decoding process of the code stream by the described decoder, the reconstructed frame of the code stream is, for example, an image frame obtained by combining a prediction signal obtained by intra-frame prediction or inter-frame prediction with a reconstructed residual signal.

[0064] In some other embodiments of the present disclosure, the code stream may include a bit stream obtained by transforming and quantizing the residual signal. Figure 1 In the example, the output of the transform / quantization module can be used as a code stream, and the reconstructed frame of the code stream is an image frame obtained by combining the prediction signal obtained by intra-frame prediction or inter-frame prediction with the reconstructed residual signal, for example, the sum of the prediction signal and the reconstructed residual signal.

[0065] For example, the first processing unit performs entropy decoding, inverse quantization, inverse transform, inter-frame prediction, or intra-frame prediction to obtain a reconstructed frame of the code stream. Alternatively, the first processing unit may only perform inverse quantization without inverse transform, that is, the first processing unit performs entropy decoding, inverse quantization, inter-frame prediction, or intra-frame prediction to obtain a reconstructed frame of the code stream. That is, the first processing operation may include one or more of entropy decoding, inverse quantization, inverse transform, inter-frame prediction, or intra-frame prediction to obtain the reconstructed frame.

[0066] In some embodiments of the present disclosure, the first processing operation further includes deriving filter parameters based on the reconstructed frame, and correspondingly, the intermediate data also includes filter parameters. Since deriving filter parameters requires relatively complex processing logic and processes, the process of deriving filter parameters based on the reconstructed frame can also be performed by the first processing unit. For example, Figure 1 In the example, the program of the filter control analysis module can be executed by the first processing unit, and the program of the filter control analysis module is executed to obtain the parameters of the deblocking filter and / or SAO, and the intermediate data includes the parameters of the deblocking filter and / or SAO. Figure 2 In the example, the part of deriving the filter parameters in the deblocking filter can be performed by the first processing unit.

[0067] Regarding step S402, the Neural Processing Unit (NPU) differs from the GPU in terms of architecture, function, and purpose. It is a processor designed to accelerate deep learning, especially neural network computing, and is used to achieve low-power accelerated artificial intelligence (AI) reasoning. Its architecture can adopt different specific architectures, instruction sets, etc. as AI algorithms, models, and use cases evolve. The NPU has a large number of parallel computing units (or computing cores) that can simultaneously process multiple neural network computing tasks (such as convolution calculations, activation functions, etc.) in parallel. Therefore, it can process large amounts of data in parallel, greatly improving computing efficiency. In addition, the NPU may also include control logic (controller) for scheduling computing tasks, managing resource allocation, and coordinating workflows between various components, including various interfaces for communicating with other system components (such as the CPU or GPU) via the system bus. These interfaces may include, for example, high-speed bus interfaces such as PCIe and NVLink, or network interfaces such as Ethernet and InfiniBand. For example, the NPU uses a series of low-power design technologies, such as optimized circuit design and dynamic voltage and frequency adjustment, to reduce energy consumption while ensuring computing performance.

[0068] Compared to general-purpose processors, the NPU can complete the same amount of computation with lower power consumption when processing neural network tasks. For example, in at least one embodiment, the first processing unit and the neural processing unit can be integrated into the same system-on-chip (SoC).

[0069] For example, the intermediate data includes a reconstructed frame and filtering parameters, and the second processing operation includes loop filtering the reconstructed frame using the filtering parameters to obtain filtered data; and post-processing the filtered data to obtain a reconstructed image. This embodiment uses a general-purpose processor to implement entropy decoding, inverse quantization, inverse transform, inter-frame prediction, intra-frame prediction, and filtering parameter derivation of the decoder. After the general-purpose processor completes frame reconstruction, it transmits the filtering parameters and reconstructed frame to the NPU, which then implements loop filtering and post-processing of the decoder.

[0070] For example, in at least one embodiment, the NPU includes multiple operation units, such as scalar operation subunits, matrix operation subunits, tensor operation subunits, etc., including one or more levels of cache (such as global cache, private cache, etc.), and may also include modules for accelerating specific types of computing tasks as needed, such as pooling units, normalization units, activation function units, etc.

[0071] For example, the scalar operator unit, the matrix operator unit, and the tensor operator unit can be programmed to implement the algorithm of video encoding and decoding. For example, a software development kit (SDK) is used to directly control hardware units such as the scalar operator unit, the matrix operator unit, and the tensor operator unit. For example, the SDK is used to load parameters directly into the tensor operator unit so that the tensor operator unit performs tensor operations. For another example, a model is written through a deep learning framework, and the compiler is relied upon to automatically map to hardware units such as the scalar operator unit, the matrix operator unit, and the tensor operator unit. In at least one embodiment, the NPU uses a scalar subunit to process part of the second processing operation, and uses a matrix operator unit to process another part of the second processing operation. The second processing operation, for example, includes a branch structure processing operation and a multiplication and addition operation operation. The scalar subunit is configured to perform the branch structure processing operation, and the matrix operator unit is configured to perform matrix operations, such as convolution calculations. The operation unit, for example, also includes a vector operator unit for performing vector operations.

[0072] Compared with GPU, NPU is suitable for computing relatively repetitive and simple tasks (for example, the task usually does not include too many branches, such as for specific image rendering and other functions). GPU relies on a large amount of parallelism to accelerate calculations, but the video encoding and decoding process is a serial and small-area parallel structure with a large number of branch structures. NPU can process branches through the scalar operation subunit inside NPU, and perform small-area parallel acceleration through vector operation (vector) subunit and matrix operation (cube) subunit. Therefore, in the embodiment of the present disclosure, a general-purpose processor is used for software decoding, and then the NPU is used to implement filtering and post-processing on the hardware. This combination reduces power consumption, improves the efficiency of encoding and decoding operations, and reduces the communication delay between software and hardware.

[0073] As mentioned above, video codecs typically operate in parallel on a block-by-block basis, encoding and decoding the residuals. Therefore, the data corresponding to these blocks is typically matrices or sparse matrices, which is compatible with the NPU's computational core. Furthermore, filtering operations such as deblocking, SAO, and ALF performed in video codecs are also suitable for high-density computation using the NPU. For example, the basic algorithm structure for deblocking filtering involves performing a linear weighted operation on pixels on both sides of a block boundary, for example, Y = a0*X0+a1*X1+a2*X2+..., where Y is the filtered pixel value, Xi is the original pixel value on both sides of the boundary, and ai is a weighting parameter, i = 0, 1, 2, etc. The weighting parameters vary for different boundaries, allowing for branching using scalar operators (for example, conditionally determining whether filtering should be skipped for certain areas (such as texture detail areas in I-frames) or selecting different filter coefficients based on boundary type). Vector and matrix operators are then used to implement high-density multiplication-addition computations. SAO and ALF also involve class lookups to retrieve specific parameters, and the NPU's tensor operation subunit supports multi-vector parallel lookups. Therefore, the NPU is compatible with the loop filtering and post-processing of video codecs. Using the NPU for loop filtering and post-processing can improve processing efficiency and reduce power consumption and latency.

[0074] Figure 5 A schematic diagram of a process for decoding a bit stream into a video image provided by at least one embodiment of the present disclosure is shown. Figure 5 To illustrate some embodiments of the present disclosure. Figure 5 In the figure, the solid-line box portion can be executed by the first processing unit, and the dotted-line box portion can be executed by the neural processing unit.

[0075] like Figure 5 As shown, a bitstream (also referred to as a "codestream") undergoes entropy decoding, inverse quantization, and inverse transformation to obtain a reconstructed residual signal. A predicted signal obtained by intra-frame prediction or inter-frame prediction is added to the reconstructed residual signal, for example, to obtain a reconstructed frame. The process of obtaining the reconstructed frame can be performed by a first processing unit. For example, the first processing unit is a CPU. In this example, the intermediate data includes the reconstructed frame, and the first processing operation includes entropy decoding, inverse quantization and inverse transformation, and intra-frame prediction or inter-frame prediction.

[0076] In some embodiments of the present disclosure, after obtaining a reconstructed frame, the reconstructed frame can be analyzed to derive filtering parameters required for subsequent filtering. Subsequent filtering may include, for example, deblocking filtering and sample adaptive offset filtering. Filtering parameters may include, for example, deblocking filtering parameters required for deblocking filtering and / or compensation parameters required for sample adaptive offset filtering.

[0077] In some embodiments of the present disclosure, the derivation of filter parameters may also be performed by the first processing unit. In this embodiment, the intermediate data includes not only the reconstructed frame but also the filter parameters. The first processing operation may include entropy decoding, inverse quantization and inverse transformation, intra-frame prediction or inter-frame prediction, and may also include derivation of filter parameters.

[0078] In some other embodiments of the present disclosure, the derivation of filtering parameters may also be performed by a neural processing unit.

[0079] After the reconstructed frame is obtained, a second processing operation is performed. The second processing operation, for example, includes loop filtering the reconstructed frame using the filtering parameters to obtain filtered data (ie, Figure 3 In some embodiments of the present disclosure, loop filtering of the reconstructed frame and post-processing of the filtered data may be performed by a neural processing unit.

[0080] In some embodiments of the present disclosure, the neural processing unit includes not only a scalar operation unit, a matrix operation unit, etc., but may also further include a programmable portion (e.g., on-chip memory). In some embodiments of the present disclosure, for example, an instruction sequence for loop filtering and post-processing is written to the on-chip memory of the NPU, and then the NPU reads the compiled instruction sequence from the on-chip memory and distributes the instruction sequence to the corresponding operation unit for execution.

[0081] The reconstructed frame is subjected to loop filtering using the filtering parameters to obtain filtered data, including: performing deblocking filtering on the reconstructed frame using the filtering parameters to obtain object data; and sequentially performing sample adaptive offset (SAO) and adaptive loop filtering (ALF) on the object data to obtain filtered data.

[0082] For example, deblocking filtering is performed on the reconstructed frame using deblocking filtering parameters of a deblocking filter to obtain object data, and then sample adaptive offset (SAO) and adaptive loop filtering are used to filter the object data to obtain filtered data.

[0083] In some embodiments of the present disclosure, post-processing includes, for example, at least one of image scaling, cropping, noise reduction, sharpening, and color space conversion to improve the image quality and visual effects of the video.

[0084] Figure 6 A schematic diagram showing an example of data flow between a CPU and an NPU provided by at least one embodiment of the present disclosure is shown.

[0085] like Figure 6As shown, in this example, CPU 610 uses buffer 620 to transfer intermediate data to NPU 630. CPU 610 is an example of a first processing unit. For example, buffer 630 can be a cache or hardware device independent of NPU and CPU. Buffer 630 is, for example, a ring buffer.

[0086] A ring buffer is essentially a fixed-size buffer whose data storage area is logically connected in a ring. A ring buffer consists of an array (for storing data) and two pointers: a read pointer and a write pointer. The read pointer points to the location where data can currently be read, and the write pointer points to the location where data can currently be written. When the write pointer reaches the end of the buffer, it automatically wraps around to the beginning and continues writing. The read pointer also follows a similar circular pattern when reading data. A ring buffer effectively utilizes fixed-size storage space, avoiding the space waste associated with reading and writing data in a standard buffer. Once data is read, the space occupied by it is immediately available for new data, eliminating the need for data movement or reallocation required by a standard buffer. In a multitasking or multithreaded environment, a ring buffer conveniently supports concurrent read and write operations. Multiple threads can read and write to the buffer simultaneously. By properly controlling the movement of the read and write pointers, data consistency and correctness can be ensured, reducing inter-thread synchronization overhead.

[0087] The CPU 610 transfers the intermediate data to the NPU 630 using the buffer 620 , for example, between step S401 and step S402 .

[0088] In some embodiments of the present disclosure, the ring buffer is divided into multiple cache sub-areas, and the multiple cache sub-areas correspond one-to-one to the multiple tasks. The second processing operation on the intermediate data is divided into multiple tasks, and the parameters required by the multiple tasks (for example, the target virtual address below) are cached in the multiple cache sub-areas. In some embodiments of the present disclosure, the cache sub-areas can be divided according to the size of the parameters required by the tasks, and the size of each cache sub-area can be equal or unequal.

[0089] Each cache sub-area includes status information, and the status information includes a processing status and a completion status. The processing status indicates that the task corresponding to the cache sub-area is being processed and has not yet been completed. The completion status indicates that the task corresponding to the cache sub-area has been processed and completed. The method also includes: for each of the multiple cache sub-areas, in response to the neural processing unit completing the processing of the task corresponding to the cache sub-area, updating the status information of the cache sub-area from the processing status to the completion status, the first processing unit reads the status information, and in response to the status information being updated from the processing status to the completion status, executes the subsequent task of the task.

[0090] For example, the status information of the cache sub-area is set to the completion state, which is equivalent to a notification that tells the CPU that the processing is completed, so that the CPU can handle subsequent operations. For example, after receiving this notification, the display hardware can be notified to display this frame on the display screen (i.e., on-screen operation).

[0091] The NPU executes tasks or commands by checking the ring buffer, reducing the NPU driver communication delay. The CPU does not need to synchronize the ring buffer status and can directly write tasks to be executed, avoiding the CPU waiting for NPU execution and reducing CPU waiting.

[0092] The video processing method also includes configuring a virtual address space for access by the neural processing unit by the CPU 610. The storage space for the intermediate data that the NPU needs to access to perform a task is configured by the first processing unit (such as the CPU). That is, the first processing unit allocates a virtual address space to the NPU and uses a virtual address to manage it for access by the NPU, so that the first processing unit only needs to pass the target virtual address in the virtual address space to the NPU to enable the NPU to access the intermediate data. The virtual address space allocated by the first processing unit to the NPU is, for example, a portion of the space in the memory of a computing device (such as a terminal) including the first processing unit and the NPU, which allows the NPU to access it.

[0093] For example, the first processing unit uses the buffer to transfer intermediate data to the neural processing unit, including: the first processing unit writes the intermediate data to the target virtual address in the virtual address space in response to obtaining the intermediate data; writes the target virtual address to the buffer; and the neural processing unit reads the target virtual address from the buffer and accesses the target virtual address to obtain the intermediate data. This embodiment unifies the NPU virtual address space, and the first processing unit can uniformly configure the NPU virtual address for access by the NPU. Through the unified address space, the data and parameters required for the task to be processed by the NPU can be flexibly transferred, reducing the frequency of communication with the NPU required for decoding and multiple post-processing tasks.

[0094] For example, in at least one example, the CPU allocates a virtual address space of 0000-FFFF to the NPU. If intermediate data is written to the target virtual address 00FF by the CPU, the CPU writes 00FF to the ring buffer. The NPU reads the target virtual addresses in the ring buffer in the order in which the ring buffer was read. If the target virtual address 00FF is read, the CPU accesses the address 00FF in the CPU memory to retrieve the intermediate data.

[0095] In some embodiments of the present disclosure, the NPU includes, in addition to the arithmetic unit and the programmable portion, a dedicated cache for a second processing operation. Process data generated during the processing of intermediate data by the second processing operation is cached in the dedicated cache via the neural processing unit; and in response to subsequent processing using the process data, the neural processing unit reads the process data from the dedicated cache for subsequent processing.

[0096] For example, the object data, filtered data and reconstructed image obtained by deblocking the reconstructed frame can all be stored in a dedicated cache. It should be noted that process data not only includes object data, filtered data, etc., but also includes any other data generated during the filtering process (for example, data generated or used by the post-processing process). When performing the second processing operation, data that needs to be accessed frequently can be stored in a dedicated cache. Dedicated caches usually use high-speed storage media and optimized storage architectures, which can temporarily store these frequently used data close to the processing unit, greatly shortening the data access path and reducing the time delay of data transmission. For example, when the NPU performs a convolution operation, the dedicated cache can quickly provide the convolution kernel weight data and input feature map data, so that the operation can be performed efficiently. Compared with reading data from the main memory, the speed is greatly improved. Using the NPU dedicated cache can significantly improve the cache hit rate.

[0097] In some embodiments of the present disclosure, after obtaining the reconstructed image, the format of the reconstructed image may be converted to obtain a display image that can be directly displayed on a display unit. The display unit includes, for example, display hardware such as a display screen.

[0098] In some embodiments of the present disclosure, the method further includes using a second processing unit to convert the format of the reconstructed image to obtain a display image, and the display image is directly used for display on the display unit, and the second processing unit is different from the neural processing unit.

[0099] For example, the data format of the reconstructed image is YUV (Y represents luminance, U and V represent chrominance) format, and the data format of the display image supported by the display unit is, for example, red, green, and blue (RGB) format, then the YUV format is converted into RGB format by the second processing unit.

[0100] In some embodiments of the present disclosure, the second processing unit is, for example, a display processor (Display Processing Unit, DPU, or "display processing unit"), for example, including a timing controller (Tcon), for example, implemented as a timing controller chip or module. The display processor is, for example, a chip or processor specifically used to process display-related tasks, and is used to convert received image signals into a variety of signals that drive the display screen to perform display operations, including but not limited to pixel data signals (such as specific numerical values in color spaces such as RGB or YUV), horizontal synchronization signals (Hsync), vertical synchronization signals (Vsync), data enable signals (DE), backlight control signals, power management signals, etc. The display processor is a dedicated hardware module or integrated circuit, mainly used to convert, optimize and control the transmission of input image signals.

[0101] The embodiments of the present disclosure do not impose any restrictions on the specifications, implementation methods, etc. of the display screen. For example, it can be a liquid crystal display (LCD), an organic light-emitting display (OLED), an inorganic light-emitting display (LED), etc. It can be a small display screen for wearable devices (such as smart glasses, smart watches, head-mounted display devices, etc.), mobile phones, etc., it can also be a medium-sized display screen for tablets, laptops, etc., and it can also be a large display screen for user televisions, monitors, etc.

[0102] Figure 7 An architectural diagram of displaying video on a mobile terminal provided by at least one embodiment of the present disclosure is shown.

[0103] like Figure 7 As shown, the mobile terminal may include a CPU, an NPU, and a DPU, etc. The CPU is an example of a first processing unit, and the DPU is an example of a second processing unit. This embodiment has no limitation on the type of mobile terminal, and may be, for example, a mobile phone, a tablet computer, a wearable device, etc.

[0104] The CPU implements entropy decoding, inverse quantization, inverse transformation, inter-frame prediction, intra-frame prediction, and filter parameter derivation in the decoding process. After the CPU completes the reconstruction of the frame, it passes the filter parameters and the reconstructed frame to the NPU. The NPU completes the subsequent decoding step, namely loop filtering, and also performs post-processing. The DPU is then used to process the YUV data generated by the post-processing into RGB data and directly display it on the display.

[0105] In the CPU soft-decoding and GPU post-processing solution, due to differences in the CPU and GPU hardware architectures, the GPU must copy the decoded data of a frame of image from the CPU after the CPU has processed it, reducing communication overhead between the CPU and GPU. In some embodiments of the present disclosure, a one-way data flow is adopted, with decoded data flowing from the CPU -> NPU -> DPU (data is fully processed in a single hardware unit before entering the next-level coprocessor). This can reduce communication synchronization overhead and share decoded data across coprocessors. Data is transmitted between all hardware units, and no data is copied in the intermediate process.

[0106] In the current CPU soft decoding GPU post-processing solution, the communication overhead between the CPU and GPU is large. Therefore, after a frame of image is decoded, the GPU usually copies the decoded data of the frame from the CPU (for example, the filtered data obtained by the ALF). In some embodiments of the present disclosure, the loop filtering and post-processing are migrated to the NPU, and the communication overhead between the CPU and NPU is reduced by using a cache area. Therefore, there is no need to wait for a frame of image to be decoded before copying data from the CPU. For example, post-processing can be performed after a stripe is decoded. This allows decoding and post-processing to be completed simultaneously, significantly reducing rendering latency. Strips are one or more independent rectangular areas into which a video image is divided during encoding. A frame of image can contain one or more strips, and each stripe contains an integer number of basic coding units such as CTUs. Strips are independent of each other during the encoding process, that is, the coding information of each stripe does not depend on other stripes and can be decoded independently. On the other hand, since decoding and post-processing are completely separate in the CPU soft decoding GPU post-processing solution, the data volume is too large and the cache cannot be reused. In some embodiments of the present disclosure, the use of the NPU's dedicated cache can improve the cache hit rate.

[0107] Figure 8 A block diagram of a video data processing device 800 provided by at least one embodiment of the present disclosure is shown.

[0108] like Figure 8 As shown, the video data processing device 800 includes a first processing unit 801 and a neural processing unit 802. For example, the video data processing device can be implemented as a video signal encoder, such as can be deployed on a server providing video services, or can be implemented as a video signal decoder, such as can be deployed on a terminal. The video data processing device can reduce power consumption and latency of video processing.

[0109] The first processing unit 801 is configured to obtain intermediate data obtained by performing the first processing operation in response to the conversion between the video frame and the video bit stream in the process of constructing the reconstructed image corresponding to the video frame. Figure 4 Step S401.

[0110] The neural processing unit 802 is configured to perform a second processing operation on the intermediate data to obtain a reconstructed image. The neural processing unit 802 performs, for example, Figure 4 Step S402.

[0111] For example, in the device / system of the video data processing apparatus 800, the first processing unit 801 can serve as a master, and the neural processing unit 802 can serve as a slave; for example, the first processing unit 801 is a CPU, and the neural processing unit 802 is a coprocessor. The master is used to control the operation of the slave, for example, providing the slave with data and code to be executed (for example, via direct memory access (DMA) or shared memory area) to the slave, so that the slave executes the corresponding code to process the data and then returns the processed results.

[0112] In some embodiments of the present disclosure, the device 800 also includes a buffer, and the first processing unit is further configured to: configure a virtual address space for access by the neural processing unit; in response to obtaining intermediate data, write the intermediate data to a target virtual address in the virtual address space; and write the target virtual address to the buffer, and the neural processing unit is further configured to: read the target virtual address from the buffer, and access the target virtual address to obtain the intermediate data.

[0113] In some embodiments of the present disclosure, the video data processing apparatus 800 further includes a second processing unit configured to perform format conversion on the reconstructed image to obtain a display image, and the display image is directly used for display on a display unit.

[0114] For example, the first processing unit 801 and the second processing unit can be implemented by hardware (e.g., circuit) modules or chips, etc. The following embodiments are the same and will not be described in detail. For example, the first processing unit can be implemented by a central processing unit (CPU), a field programmable gate array (FPGA), or other processing units with data processing capabilities and / or instruction execution capabilities, as well as corresponding computer instructions to implement these units or modules.

[0115] It should be noted that for the sake of clarity and brevity, the present embodiment does not illustrate all components of the video data processing device 800. To implement the necessary functions of the video data processing device, those skilled in the art may provide or configure other components not shown as needed, and the present embodiment does not limit this.

[0116] It should be noted that in the embodiment of the present disclosure, the various units of the device 800 correspond to the various steps of the aforementioned video data processing method. For the specific functions of the device 800, reference can be made to the relevant description of the video data processing method, which will not be repeated here. Figure 8 The components and structures of the device 800 shown are merely exemplary and non-limiting. The device 800 may further include other components and structures as needed.

[0117] At least one embodiment of the present disclosure further provides an electronic device, comprising a processor and a memory. The processor comprises a first processing unit and a neural processing unit. The memory comprises one or more computer program instructions; when the one or more computer program instructions are executed by the processor, the video data processing method provided by any embodiment of the present disclosure is executed. For example, the video data processing method comprises: in response to the conversion between the video frame of the video and the bit stream of the video, in the process of constructing a reconstructed image corresponding to the video frame, obtaining intermediate data obtained by performing a first processing operation by the first processing unit; and performing a second processing operation on the intermediate data by the neural processing unit to obtain a reconstructed image. The electronic device can reduce the power consumption and delay of video processing.

[0118] Figure 9 A schematic block diagram of an electronic device provided for some embodiments of the present disclosure.

[0119] like Figure 9 As shown, the electronic device 1200 includes at least one processor 1210 and a memory 1220. The memory 1220 is used to store non-transitory computer-readable instructions (e.g., one or more computer program modules). The processor 1210 is used to execute the non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are executed by the processor 1210, one or more steps in the video data processing method described above can be performed. The memory 1220 and the processor 1210 can be interconnected via a bus system and / or other form of connection mechanism (not shown).

[0120] For example, in addition to the first processing unit (e.g., CPU) and the neural processing unit, the processor may also include a digital signal processor (DSP), a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), or other processing units with data processing capabilities and / or instruction execution capabilities, as needed. The first processing unit may be a general-purpose processor, and the neural processing unit may be a dedicated processor, and may control other components in the electronic device to perform desired functions.

[0121] For example, the memory 1220 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, a flash memory, etc. One or more computer program modules may be stored on the computer-readable storage medium, and the processor 1210 may execute one or more computer program modules to implement various functions of the electronic device 1200. Various applications and various data, as well as various data used and / or generated by the applications, may also be stored in the computer-readable storage medium.

[0122] It should be noted that, in the embodiment of the present disclosure, the specific functions and technical effects of the electronic device 1200 can be referred to the above description of the video data processing method, which will not be repeated here.

[0123] Figure 10 A schematic structural diagram of another electronic device provided by at least one embodiment of the present disclosure is shown.

[0124] Reference below Figure 10 , which shows a schematic structural diagram of an electronic device (e.g., a terminal device or a server) 1500 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 10 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0125] like Figure 10 As shown, the electronic device 1500 may include a processing device (such as a central processing unit, an NPU, a DPU, etc.) 1501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1502 or a program loaded from a storage device 1508 into a random access memory (RAM) 1503, such as executing the video data processing method of any embodiment of the present disclosure, encoding or decoding a video signal, and displaying the obtained video signal on a display after decoding. The central processing unit is an example of a first processing unit, the DPU is an example of a second processing unit, and the NPU serves as a neural processing unit.

[0126] Various programs and data required for the operation of the electronic device 1500 are also stored in the RAM 1503. The processing device 1501, the ROM 1502, and the RAM 1503 are connected to each other via a bus 1504. An input / output (I / O) interface 1505 is also connected to the bus 1504.

[0127] Typically, the following devices may be connected to the I / O interface 1505: an input device 1506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1509. The communication device 1509 may allow the electronic device 1500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 10 The electronic device 1500 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0128] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1509, or installed from the storage device 1508, or installed from the ROM 1502. When the computer program is executed by the processing device 1501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0129] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0130] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0131] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0132] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the video data processing method provided by any embodiment of the present disclosure.

[0133] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0134] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0135] The units or modules involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit or module does not necessarily limit the unit or module itself.

[0136] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0137] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0138] According to one or more embodiments of the present disclosure, Example 1 provides a video data processing method, including:

[0139] In response to conversion between a video frame of a video and a bitstream of the video, obtaining intermediate data obtained by a first processing unit performing a first processing operation during construction of a reconstructed image corresponding to the video frame; and

[0140] The neural processing unit performs a second processing operation on the intermediate data to obtain the reconstructed image.

[0141] According to one or more embodiments of the present disclosure, Example 2 provides that the first processing operation in Example 1 includes at least one of entropy decoding, inverse quantization, inverse transformation, inter-frame prediction, or intra-frame prediction,

[0142] The intermediate data includes reconstructed frames of a code stream, wherein the code stream is obtained by encoding the video frames.

[0143] According to one or more embodiments of the present disclosure, Example 3 provides that the first processing operation in Example 2 further includes deriving filtering parameters based on the reconstructed frame, and the intermediate data further includes the filtering parameters.

[0144] The second processing operation includes:

[0145] Using the filtering parameters, loop filtering is performed on the reconstructed frame to obtain filtered data; and

[0146] The filtered data is post-processed to obtain the reconstructed image.

[0147] According to one or more embodiments of the present disclosure, Example 4 provides the method of performing the loop filtering on the reconstructed frame using the filtering parameters in Example 3 to obtain the filtered data, including:

[0148] Using the filtering parameters, performing deblocking filtering on the reconstructed frame to obtain object data; and

[0149] The object data is sequentially subjected to sample adaptive compensation and adaptive loop filtering to obtain the filtered data.

[0150] According to one or more embodiments of the present disclosure, Example 5 provides the method in Example 1, further comprising:

[0151] using a second processing unit to convert the format of the reconstructed image to obtain a display image,

[0152] The display image is directly used for display on a display unit, and the second processing unit is different from the neural processing unit.

[0153] According to one or more embodiments of the present disclosure, Example 6 provides that the data format of the reconstructed image in Example 5 is a YUV format, and the data format of the displayed image is an RGB format.

[0154] According to one or more embodiments of the present disclosure, Example 7 provides the method in Example 5, wherein the first processing unit includes a general-purpose processor, and the second processing unit includes a display processor.

[0155] According to one or more embodiments of the present disclosure, Example 8 provides the method of any one of Examples 1 to 7, further comprising:

[0156] The first processing unit transfers the intermediate data to the neural processing unit using a buffer.

[0157] According to one or more embodiments of the present disclosure, Example 9 provides the method of Example 8, further comprising:

[0158] configuring, by the first processing unit, a virtual address space for access by the neural processing unit;

[0159] The first processing unit transmits the intermediate data to the neural processing unit using the buffer, including:

[0160] In response to obtaining the intermediate data, the first processing unit writes the intermediate data into a target virtual address in the virtual address space;

[0161] Writing the target virtual address into the buffer; and

[0162] The neural processing unit reads the target virtual address from the buffer and accesses the target virtual address to obtain the intermediate data.

[0163] According to one or more embodiments of the present disclosure, Example 10 provides that the buffer in Example 8 is a ring buffer, and the ring buffer is divided into a plurality of cache sub-areas, the plurality of cache sub-areas correspond one-to-one to a plurality of tasks, each cache sub-area includes status information, and the status information includes a processing status and a completion status, and the method further includes:

[0164] For each of the plurality of cache sub-regions, in response to the neural processing unit completing processing of a task corresponding to the cache sub-region, updating the state information of the cache sub-region from the processing state to the completion state,

[0165] The first processing unit reads the status information and, in response to the status information being updated from the processing status to the completion status, executes a subsequent task of the task.

[0166] According to one or more embodiments of the present disclosure, Example 11 provides that the neural processing unit in any one of Examples 1 to 7 further includes a dedicated cache, and process data generated during the processing of the intermediate data by the second processing operation is cached in the dedicated cache via the neural processing unit; and

[0167] In response to subsequent processing using the process data, the neural processing unit reads the process data from the dedicated cache to perform the subsequent processing.

[0168] According to one or more embodiments of the present disclosure, Example 12 provides that the conversion in any one of Examples 1 to 7 includes encoding the video into the bitstream, or the conversion includes decoding the video from the bitstream.

[0169] According to one or more embodiments of the present disclosure, Example 13 provides a video data processing apparatus, including:

[0170] a first processing unit configured to, in response to conversion between a video frame of a video and a bitstream of the video, obtain intermediate data obtained by performing a first processing operation in a process of constructing a reconstructed image corresponding to the video frame; and

[0171] A neural processing unit is configured to perform a second processing operation on the intermediate data to obtain a reconstructed image.

[0172] According to one or more embodiments of the present disclosure, Example 14 provides that the apparatus in Example 13 further includes a buffer, and the first processing unit is further configured to:

[0173] configuring a virtual address space for access by the neural processing unit;

[0174] In response to obtaining the intermediate data, writing the intermediate data into a target virtual address in the virtual address space; and

[0175] Writing the target virtual address into the buffer,

[0176] The neural processing unit is further configured to: read the target virtual address from the buffer, and access the target virtual address to obtain the intermediate data.

[0177] According to one or more embodiments of the present disclosure, Example 15 provides the apparatus of Example 13 or Example 14, further comprising:

[0178] The second processing unit is configured to convert the format of the reconstructed image to obtain a display image,

[0179] The display image is directly used for display on the display unit.

[0180] According to one or more embodiments of the present disclosure, Example 16 provides an electronic device, including:

[0181] a processor, wherein the processor comprises a first processing unit and a neural processing unit; and

[0182] a memory including one or more computer program instructions,

[0183] The one or more computer program instructions are executed by the processor when it is executed:

[0184] In response to conversion between a video frame of a video and a bitstream of the video, obtaining intermediate data obtained by a first processing unit performing a first processing operation during construction of a reconstructed image corresponding to the video frame; and

[0185] The neural processing unit performs a second processing operation on the intermediate data to obtain the reconstructed image.

[0186] According to one or more embodiments of the present disclosure, Example 17 provides a computer-readable storage medium, which non-transitorily stores computer-readable instructions. When the computer-readable instructions are executed by a processor, the following are implemented:

[0187] In response to conversion between a video frame of a video and a bitstream of the video, obtaining intermediate data obtained by a first processing unit performing a first processing operation during construction of a reconstructed image corresponding to the video frame; and

[0188] The neural processing unit performs a second processing operation on the intermediate data to obtain the reconstructed image.

[0189] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0190] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0191] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A video data processing method, comprising: In response to conversion between a video frame of a video and a bitstream of the video, obtaining intermediate data obtained by a first processing unit performing a first processing operation during construction of a reconstructed image corresponding to the video frame; as well as The neural processing unit performs a second processing operation on the intermediate data to obtain the reconstructed image.

2. The method according to claim 1, wherein The first processing operation includes at least one of entropy decoding, inverse quantization, inverse transformation, inter-frame prediction or intra-frame prediction, The intermediate data includes reconstructed frames of a code stream, wherein the code stream is obtained by encoding the video frames.

3. The method according to claim 2, wherein: The first processing operation further comprises deriving filtering parameters based on the reconstructed frame, and the intermediate data further comprises the filtering parameters. The second processing operation includes: Using the filtering parameters, loop filtering is performed on the reconstructed frame to obtain filtered data; and The filtered data is post-processed to obtain the reconstructed image.

4. The method according to claim 3, wherein: Performing the loop filtering on the reconstructed frame using the filtering parameters to obtain the filtered data includes: Using the filtering parameters, performing deblocking filtering on the reconstructed frame to obtain object data; and The object data is sequentially subjected to sample adaptive compensation and adaptive loop filtering to obtain the filtered data.

5. The method according to claim 1, further comprising: using a second processing unit to convert the format of the reconstructed image to obtain a display image, The display image is directly used for display on a display unit, and the second processing unit is different from the neural processing unit.

6. The method according to claim 5, wherein: The data format of the reconstructed image is YUV format, and the data format of the displayed image is RGB format.

7. The method according to claim 5, wherein: The first processing unit includes a general-purpose processor, and the second processing unit includes a display processor.

8. The method according to any one of claims 1 to 7, further comprising: The first processing unit transfers the intermediate data to the neural processing unit using a buffer.

9. The method according to claim 8, further comprising: configuring, by the first processing unit, a virtual address space for access by the neural processing unit; The first processing unit transmits the intermediate data to the neural processing unit using the buffer, including: In response to obtaining the intermediate data, the first processing unit writes the intermediate data into a target virtual address in the virtual address space; Writing the target virtual address into the buffer; and The neural processing unit reads the target virtual address from the buffer and accesses the target virtual address to obtain the intermediate data.

10. The method according to claim 8, wherein The buffer is a ring buffer and the ring buffer is divided into a plurality of cache sub-areas, the plurality of cache sub-areas correspond one-to-one to a plurality of tasks, each cache sub-area includes status information, the status information includes a processing status and a completion status, and the method further includes: For each of the plurality of cache sub-regions, in response to the neural processing unit completing processing of a task corresponding to the cache sub-region, updating the state information of the cache sub-region from the processing state to the completion state, The first processing unit reads the status information and, in response to the status information being updated from the processing status to the completion status, executes a subsequent task of the task.

11. The method according to any one of claims 1 to 7, wherein: The neural processing unit further includes a dedicated cache, and process data generated during the processing of the intermediate data by the second processing operation is cached in the dedicated cache via the neural processing unit; as well as In response to subsequent processing using the process data, the neural processing unit reads the process data from the dedicated cache to perform the subsequent processing.

12. The method according to any one of claims 1 to 7, wherein: The converting includes encoding the video into the bitstream, or the converting includes decoding the video from the bitstream.

13. A video data processing device, comprising: a first processing unit configured to, in response to conversion between a video frame of a video and a bitstream of the video, obtain intermediate data obtained by performing a first processing operation in a process of constructing a reconstructed image corresponding to the video frame; and A neural processing unit is configured to perform a second processing operation on the intermediate data to obtain the reconstructed image.

14. The device according to claim 13, wherein The device further includes a buffer, and the first processing unit is further configured to: configuring a virtual address space for access by the neural processing unit; In response to obtaining the intermediate data, writing the intermediate data into a target virtual address in the virtual address space; and Writing the target virtual address into the buffer, The neural processing unit is further configured to: read the target virtual address from the buffer, and access the target virtual address to obtain the intermediate data.

15. The apparatus according to claim 13 or 14, further comprising: The second processing unit is configured to convert the format of the reconstructed image to obtain a display image, The display image is directly used for display on the display unit.

16. An electronic device comprising: a processor, wherein the processor comprises a first processing unit and a neural processing unit; and a memory including one or more computer program instructions, The one or more computer program instructions are executed by the processor when it is executed: In response to conversion between a video frame of a video and a bitstream of the video, obtaining intermediate data obtained by a first processing unit performing a first processing operation during construction of a reconstructed image corresponding to the video frame; and The neural processing unit performs a second processing operation on the intermediate data to obtain the reconstructed image.

17. A computer-readable storage medium non-transitorily storing computer-readable instructions, wherein: When executed by a processor, the computer-readable instructions: In response to conversion between a video frame of a video and a bitstream of the video, obtaining intermediate data obtained by a first processing unit performing a first processing operation during construction of a reconstructed image corresponding to the video frame; as well as The neural processing unit performs a second processing operation on the intermediate data to obtain the reconstructed image.