Audio and video processing method and system, storage medium and electronic equipment
By deploying multiple processing steps on the GPU in the audio and video processing system, a full-link GPU data processing pipeline is built, which solves the efficiency bottleneck brought about by data interaction between the CPU and the GPU, and realizes efficient audio and video processing, which is suitable for ultra-high-definition video and intelligent media applications.
Patent Information
- Application Number
- CN202510220965.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, frequent interaction between the CPU and the GPU leads to bandwidth bottlenecks, some key steps have not achieved GPU acceleration, cannot fully utilize the advantages of parallel computing, and lacks scalability in high-resolution scenarios, and the system efficiency is difficult to meet the needs.
By deploying the entire process of demultiplexing, decoding, format conversion, machine learning reasoning, encoding and multiplexing on the GPU side, an end-to-end GPU data processing pipeline is built to reduce data transmission between the GPU and the CPU, and improve system throughput and real-time.
Significantly reduce PCIe communication overhead, improve system throughput and real-time performance, and provide efficient underlying support for ultra-high-definition video processing and intelligent media applications.
Smart Images

Figure CN119996761A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimedia technology, and in particular to an audio and video processing method, system, storage medium and electronic equipment. Background Art
[0002] In the field of multimedia technology, when it comes to audio and video processing that requires machine learning algorithms, due to the large order of magnitude of audio and video data itself and the complexity of reasoning calculations, the reasoning part is generally accelerated using a GPU (Graphics Processing Unit), while other parts of the processing flow, such as audio and video transpackaging, transcoding, format conversion, and other filter processing, are generally processed on the CPU (Central Processing Unit). In a typical x86 machine environment, since the GPU can only communicate with the CPU through PCIe, data exchange efficiency (for example, insufficient bandwidth and latency) issues will arise.
[0003] In the existing technology, although there are audio and video processing solutions based on GPU acceleration, there are still problems with insufficient full-link optimization. For example, the frequent interaction between the CPU and GPU in the data processing link leads to bandwidth bottlenecks, some key steps are not accelerated by GPU, and the advantages of parallel computing cannot be fully utilized. In addition, the scalability is insufficient in high-resolution scenarios, and the system efficiency is difficult to meet the needs.
[0004] Therefore, how to implement a technical solution for GPU acceleration of the entire link of audio and video processing and break through the efficiency bottleneck under the existing heterogeneous architecture has become a technical problem that technical personnel in this field urgently need to solve. Summary of the invention
[0005] In view of the above problems, the present invention provides an audio and video processing method, system, storage medium and electronic device that overcome the above problems or at least partially solve the above problems. The technical solution is as follows:
[0006] An audio and video processing method, comprising:
[0007] The hardware encoding and decoding process is performed by GPU acceleration to demultiplex and decode the input audio and video data, and output the first audio and video stream data;
[0008] Convert the first audio and video stream data from a video pixel format to a machine learning input format on the GPU, and convert the color space format of the machine learning output to an encoding input format;
[0009] Performing video processing on the audio and video stream data converted into the machine learning input format using a deep learning model to obtain second audio and video stream data, wherein the deep learning model supports a variety of image processing algorithms and machine learning frameworks;
[0010] Converting the second audio and video stream data from the machine learning input format to the video pixel format on the GPU;
[0011] The encoding process is performed through GPU acceleration, and the second audio and video stream data converted into the video pixel format is multiplexed, packaged and output.
[0012] Optionally, the demultiplexing and decoding the input audio and video data to output the first audio and video stream data includes:
[0013] Read input audio and video data;
[0014] Separating the audio and video elementary streams and media parameters of the audio and video data;
[0015] The audio and video basic stream is acceleratedly decoded on the GPU to generate first audio and video stream data in a video pixel format.
[0016] Optionally, the method further includes:
[0017] Dynamically adjust the model input preprocessing parameters of the deep learning model based on the media parameters.
[0018] Optionally, the using a deep learning model to perform video processing on the audio and video stream data converted into the machine learning input format to obtain second audio and video stream data includes:
[0019] According to the resolution and frame rate of the audio and video stream data converted into the machine learning input format, a pre-trained deep learning model is dynamically loaded, and the video frames of the audio and video stream data are converted into a tensor format and then an image processing algorithm is executed to obtain a second audio and video stream data, wherein the image processing algorithm includes at least one of image super-resolution reconstruction, video noise reduction and restoration, frame rate improvement interpolation and HDR effect enhancement.
[0020] Optionally, dynamically loading a pre-trained deep learning model according to the resolution and frame rate of the audio and video stream data converted into the machine learning input format includes:
[0021] According to the resolution and frame rate of the audio and video stream data converted into the machine learning input format, the model input size is automatically matched, and a deep learning model that supports at least one machine learning framework including Torch, Caffe, Tensorflow, TensorRT and OpenVINO is loaded.
[0022] Optionally, the performing encoding processing by GPU acceleration to multiplex, package and output the second audio and video stream data converted into the video pixel format includes:
[0023] Calling the GPU hardware encoder to compress and encode the second audio and video stream data converted into the video pixel format;
[0024] The encoded second audio and video stream data is multiplexed, packaged and output according to the target encapsulation format.
[0025] Optionally, the video pixel format is YUV format or NV12 format, and the machine learning input format is RGB format or BGR format.
[0026] An audio and video processing system, comprising: a demultiplexing decoding module, a format conversion module, a machine learning filter module and a coding multiplexing module, wherein the system realizes the coordinated work of hardware coding and decoding and machine learning processing through GPU acceleration, wherein:
[0027] The demultiplexing and decoding module is used for audio and video separation and hardware accelerated decoding, and outputs original audio and video stream data;
[0028] The format conversion module is used to implement bidirectional data format conversion on the GPU, converting the decoded video pixel format into the machine learning input format, and converting the processed machine learning output color space format into the encoding input format;
[0029] The machine learning filter module is used to provide a pluggable deep learning model interface and support a variety of image processing algorithms and machine learning frameworks;
[0030] The encoding multiplexing module is used to implement GPU accelerated encoding and audio and video multiplexing packaging.
[0031] A computer-readable storage medium stores a program, and when the program is executed by a processor, the audio and video processing method is implemented.
[0032] An electronic device comprises at least one processor, and at least one memory and a bus connected to the processor; wherein the processor and the memory communicate with each other via the bus; and the processor is used to call program instructions in the memory to execute the audio and video processing method.
[0033] By means of the above technical scheme, the present invention provides an audio and video processing method, system, storage medium and electronic device, the method comprising: performing hardware encoding and decoding processing through GPU acceleration, demultiplexing and decoding the input audio and video data, and outputting the first audio and video stream data; converting the first audio and video stream data from the video pixel format to the machine learning input format on the GPU, and converting the color space format of the machine learning output to the encoding input format; using a deep learning model to perform video processing on the audio and video stream data converted to the machine learning input format to obtain the second audio and video stream data, wherein the deep learning model supports a variety of image processing algorithms and machine learning frameworks; converting the second audio and video stream data from the machine learning input format to the video pixel format on the GPU; performing encoding processing through GPU acceleration, and multiplexing and packaging the second audio and video stream data converted to the video pixel format for output. The present invention innovatively deploys the entire process of demultiplexing, decoding, format conversion, machine learning reasoning, encoding and multiplexing on the GPU side, and constructs an end-to-end GPU data processing pipeline, thereby significantly reducing PCIe communication overhead, improving system throughput and real-time performance, and providing efficient underlying support for ultra-high-definition video processing and intelligent media applications.
[0034] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0036] Figure 1 A schematic diagram showing a flow chart of an implementation of an audio and video processing method according to an embodiment of the present invention;
[0037] Figure 2 A schematic diagram of the architecture of an audio and video processing system provided by an embodiment of the present invention is shown;
[0038] Figure 3 A schematic diagram showing a logical framework of audio and video processing provided by an embodiment of the present invention is shown;
[0039] Figure 4 A logic block diagram of using a filter in a Python version provided by an embodiment of the present invention is shown;
[0040] Figure 5A schematic structural diagram of an electronic device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0041] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to enable the scope of the present invention to be fully communicated to those skilled in the art.
[0042] In the field of multimedia processing, with the explosive growth of demand for ultra-high-definition video, real-time streaming and intelligent audio and video processing, machine learning algorithms (such as image super-resolution, denoising, interpolation, HDR, etc.) are widely used in audio and video enhancement and content analysis. However, audio and video data is massive (such as the high resolution and high frame rate characteristics of 4K / 8K video) and has processing complexity (such as encoding and decoding, format conversion, algorithm reasoning, etc.), which places extremely high demands on computing resources. In traditional solutions, heterogeneous computing architecture is usually adopted to deploy machine learning reasoning tasks on GPU (Graphics Processing Unit) for acceleration, while audio and video transpackaging, transcoding, format conversion and other processes still rely on CPU (Central Processing Unit) for processing. In a typical x86 architecture, data is exchanged between GPU and CPU through the PCIe bus. When processing high-resolution video or multi-task concurrency, uncompressed raw audio and video data (such as RGB / YUV format) needs to be frequently transmitted between CPU and GPU, resulting in PCIe bandwidth bottlenecks, communication delays and other problems, which seriously restricts the overall throughput of the system and becomes a significant efficiency bottleneck in large-scale machine learning applications.
[0043] Based on this, in an embodiment of the present invention, an audio and video processing method and system are provided. First, secondary development is performed based on the FFmpeg framework. The I / O reading, writing and transpackaging operations in the multimedia processing flow are handed over to the CPU for processing, while high-computational tasks such as decoding, encoding, format conversion, filter processing and machine learning reasoning are handed over to the GPU for execution, so as to improve the computing efficiency. Secondly, only one data interaction between the GPU and the CPU is required in the processing flow, and the interactive data is compressed audio and video data, which significantly reduces the amount of communication data, thereby effectively alleviating the performance bottleneck problem caused by insufficient PCIe bandwidth. It can be seen that the present invention achieves a significant improvement in multimedia processing speed through full-link GPU acceleration optimization, and provides efficient support for audio and video processing and intelligent applications.
[0044] like Figure 1As shown, a flowchart of an implementation of an audio and video processing method provided by an embodiment of the present invention is provided. The method may include:
[0045] S100, performing hardware encoding and decoding processing through GPU acceleration, demultiplexing and decoding the input audio and video data, and outputting first audio and video stream data.
[0046] Specifically, the embodiment of the present invention can call the GPU hardware decoder (such as NVIDIA or NVENC) based on the FFmpeg framework to directly demultiplex and decode the input audio and video streams, output the original NV12 format data and keep it in the GPU video memory throughout the process. Compared with the traditional CPU demultiplexing + GPU decoding solution, the embodiment of the present invention saves the first transmission of the original data from the CPU to the GPU, thereby improving the decoding speed.
[0047] S110. Convert the first audio and video stream data from a video pixel format to a machine learning input format on the GPU, and convert the color space format of the machine learning output to an encoding input format.
[0048] Specifically, the embodiment of the present invention can convert NV12 into RGB / BGR format through the CUDA-accelerated NPP library to meet the tensor input requirements of models such as TensorFlow and PyTorch.
[0049] S120. Perform video processing on the audio and video stream data converted into a machine learning input format using a deep learning model to obtain second audio and video stream data, wherein the deep learning model supports a variety of image processing algorithms and machine learning frameworks.
[0050] Specifically, the embodiment of the present invention can directly process RGB or BGR data on the GPU based on the dynamically loaded pre-trained model using the TensorRT optimization engine to reduce the time consumption of single-frame reasoning. The embodiment of the present invention supports multi-model cascading (such as super-resolution followed by HDR), realizes pipeline parallelism through CUDA streams, and reduces model switching delays.
[0051] S130. Convert the second audio and video stream data from a machine learning input format to a video pixel format on the GPU.
[0052] Specifically, the embodiment of the present invention can reversely convert the RGB or BGR results output by the model into NV12 format to provide a compatible format for GPU encoding. It can be understood that both format conversions are completed in the GPU video memory, avoiding the two PCIe transmissions required to return the data to the CPU for conversion in the traditional solution.
[0053] Specifically, the embodiment of the present invention can use FFmpeg command line parameters to specify whether the task is executed on the CPU or GPU. The following is a complete command based on GPU accelerated reasoning:
[0054] "CUDA_VISIBLE_DEVICES=3 / home / zhoujs / mgtvML_FFmpeg / ffmpeg -hwaccelcuvid -c:v h264_cuvid -hwaccel_output_format cuda -i
[0055] / home / zhoujs / mnt_source_code / video_test / 540_test.mp4-vf
[0056] format_cuda=pix_fmt=rgbpf32,mgtvsr_tensorrt= / home / zhoujs / mgtvML_FFmpeg / ISR / isr_model_540p_s.trt,format_cuda=pix_fmt=nv12 -c:v h264_nvenc -b:v4096k -y / home / zhoujs / mnt_source_code / video_test / 540_test-nvenc.ts".
[0057] Among them, CUDA_VISIBLE_DEVICES=3 means that the specified task is executed on GPU card numbered 3, ensuring that all processing is completed on this GPU. Next, in the input decoding stage, use the "-hwaccel cuvid -c:v h264_cuvid-hwaccel_output_format cuda" parameter to specify that the I / O file is read and decapsulated by the CPU, and then the data is handed over to the GPU for decoding. The decoded video data is stored in the GPU video memory in NV12 format. Then, in the GPU filter processing stage, you can call the format_cuda filter to convert the NV12 data to the RGB or BGR data format required for machine learning model processing on the GPU through the CUDA interface without returning the data to the CPU. Call the mgtvsr_tensorrt filter to perform super-resolution processing on the GPU (based on the specified TensorRT model file " / home / zhoujs / mgtvML_FFmpeg / ISR / isr_model_540p_s.trt"). After the inference processing is completed, the format_cuda filter is called again to convert the RGB or BGR data back to the NV12 data format that the hardware encoder can accept. The entire process is completed in the GPU memory without going through the CPU. In the output encoding stage, the "-c:v h264_nvenc -b:v 4096k" parameter is used to call the GPU hardware encoder to encode the NV12 format video data in H.264. Finally, after the hardware encoding is completed, the audio and video streams are multiplexed and packaged on the CPU to generate output files (such as TS files) that conform to the target encapsulation format and write them to the storage. In this way, the filter processing and the bidirectional conversion of data formats are completed in the GPU memory, which minimizes the data transmission between the CPU and the GPU and improves the execution efficiency of inference and encoding.
[0058] S140, performing encoding processing through GPU acceleration, and multiplexing and packaging the second audio and video stream data converted into a video pixel format for output.
[0059] Specifically, the embodiment of the present invention can call the GPU hardware encoder to compress NV12 data in real time, and complete audio and video synchronization and encapsulation (MP4 / MKV) through the GPU memory pass-through interface of FFmpeg. The final data is directly written from the GPU video memory to the storage, avoiding transmission back to the CPU, thereby reducing end-to-end processing delay.
[0060] The present invention provides an audio and video processing method, which includes: performing hardware encoding and decoding processing through GPU acceleration, demultiplexing and decoding processing on input audio and video data, and outputting first audio and video stream data; converting the first audio and video stream data from a video pixel format to a machine learning input format on the GPU, and converting the color space format of the machine learning output to a coding input format; using a deep learning model to perform video processing on the audio and video stream data converted to the machine learning input format to obtain second audio and video stream data, wherein the deep learning model supports a variety of image processing algorithms and machine learning frameworks; converting the second audio and video stream data from the machine learning input format to a video pixel format on the GPU; performing encoding processing through GPU acceleration, and multiplexing and packaging the second audio and video stream data converted to the video pixel format for output. The present invention innovatively deploys the entire process of demultiplexing, decoding, format conversion, machine learning reasoning, encoding and multiplexing on the GPU side, and constructs an end-to-end GPU data processing pipeline, thereby significantly reducing PCIe communication overhead, improving system throughput and real-time performance, and providing efficient underlying support for ultra-high-definition video processing and intelligent media applications.
[0061] Optional, same as above Figure 1 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present invention, demultiplexing and decoding the input audio and video data to output the first audio and video stream data may specifically include:
[0062] Read the input audio and video data. Separate the audio and video basic stream and media parameters of the audio and video data. Accelerate the decoding of the audio and video basic stream on the GPU to generate the first audio and video stream data in the video pixel format.
[0063] Specifically, the embodiment of the present invention can call the libavformat protocol reading module in the FFmpeg framework to read the media file data in the network streaming media or the local to obtain audio and video data. Then, the libavformat demultiplexing module in the FFmpeg framework is used to demultiplex the read audio and video data stream to achieve the separation of audio and video streams. At the same time, the media parameters of the audio and video files will be extracted and saved, including the resolution and frame rate of the video, as well as the sampling rate, number of channels and bit depth of the audio, etc. The media parameters provide the necessary parameters for the subsequent model initialization. The above process is executed on the CPU, and the overhead is also small. Finally, the libavcodec decoding module in the FFmpeg framework is called to decode the separated audio and video streams and convert them into original uncompressed audio and video stream data, wherein the audio and video decoding process is placed on the GPU for accelerated processing, especially when using NVIDIA hardware decoding, the data format after video decoding will be NV12.
[0064] The embodiment of the present invention efficiently demultiplexes and decodes the input audio and video data, and finally outputs the first audio and video stream data, thereby realizing GPU-accelerated hardware encoding and decoding processing.
[0065] Optional, same as above Figure 1 On the basis of one or more corresponding embodiments, another optional embodiment provided by the embodiment of the present invention may further include:
[0066] Dynamically adjust the model input preprocessing parameters of the deep learning model based on the media parameters.
[0067] Specifically, the embodiment of the present invention can dynamically load the offline pre-trained machine learning model according to the media parameters. The adapted model file is initialized by parsing the properties of the audio and video stream (such as resolution, frame rate, etc.) and the corresponding configuration file is loaded to adjust the model input preprocessing parameters of the deep learning model.
[0068] The embodiment of the present invention can significantly improve the adaptability of the deep learning model to audio and video data with different resolutions and frame rates by dynamically adjusting the input preprocessing parameters of the deep learning model according to the media parameters, thereby optimizing the processing effect and improving the processing efficiency.
[0069] Optional, same as above Figure 1 On the basis of one or more corresponding embodiments, in another optional embodiment provided by an embodiment of the present invention, using a deep learning model to perform video processing on the audio and video stream data converted into a machine learning input format to obtain the second audio and video stream data may specifically include:
[0070] According to the resolution and frame rate of the audio and video stream data converted into the machine learning input format, a pre-trained deep learning model is dynamically loaded, and the video frames of the audio and video stream data are converted into a tensor format and then an image processing algorithm is executed to obtain a second audio and video stream data, wherein the image processing algorithm includes at least one of image super-resolution reconstruction, video noise reduction and restoration, frame rate improvement interpolation and HDR effect enhancement.
[0071] Specifically, the embodiment of the present invention can determine whether the deep learning model has been initialized. If not, the step of dynamically adjusting the model input preprocessing parameters of the deep learning model according to the media parameters is executed. After initialization, the video frames of the audio and video stream data are converted into tensor data in an adapted model input format. Then, an image processing algorithm is executed, including but not limited to image super-resolution reconstruction, video noise reduction and repair, frame rate enhancement interpolation and HDR effect enhancement, to obtain the second audio and video stream data.
[0072] The embodiment of the present invention dynamically loads a pre-trained deep learning model according to the resolution and frame rate of the audio and video stream data, and executes an image processing algorithm. It can automatically optimize the processing process according to different media features, improve video quality and visual effects, and improve processing efficiency.
[0073] Optionally, an embodiment of the present invention can automatically match the model input size according to the resolution and frame rate of the audio and video stream data converted into the machine learning input format, and load a deep learning model that supports at least one machine learning framework including Torch, Caffe, Tensorflow, TensorRT and OpenVINO.
[0074] Optional, same as above Figure 1 On the basis of one or more corresponding embodiments, another optional embodiment provided by the embodiment of the present invention is to accelerate the encoding process by GPU, and multiplex and package the second audio and video stream data converted into the video pixel format for output, which may specifically include:
[0075] The GPU hardware encoder is called to compress and encode the second audio and video stream data converted into a video pixel format; and the encoded second audio and video stream data is multiplexed, packaged and output according to a target encapsulation format.
[0076] Specifically, the embodiment of the present invention can call the libavcodec encoding module inside FFmpeg to encode and compress the converted second audio and video stream data, and then call the libavformat multiplexing module inside FFmpeg to multiplex the encoded second audio and video stream according to the target encapsulation format and package it into a final playable audio and video file.
[0077] The embodiment of the present invention can significantly improve the video processing efficiency and quality by calling the GPU hardware encoder to perform compression encoding of the video pixel format, and multiplexing and packaging the output according to the target encapsulation format, thereby ensuring that the final generated audio and video files have better playback performance and compatibility.
[0078] Optional, same as above Figure 1 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present invention, the video pixel format is YUV format or NV12 format, and the machine learning input format is RGB format or BGR format.
[0079] Although operations are depicted in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in a sequential order.Multitasking and parallel processing may be advantageous under certain circumstances.
[0080] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0081] like Figure 2 As shown, an architecture diagram of the audio and video processing system provided by an embodiment of the present invention, the system may include: a demultiplexing decoding module M1, a format conversion module M2, a machine learning filter module M3 and an encoding multiplexing module M4, and the system realizes the collaborative work of hardware encoding and decoding and machine learning processing through GPU acceleration.
[0082] The demultiplexing and decoding module M1 is used for audio and video separation and hardware accelerated decoding, and outputs original audio and video stream data.
[0083] Specifically, the demultiplexing and decoding module M1 is based on the FFmpeg framework, calls the libavformat demultiplexer to separate the audio and video streams, and uses GPU hardware (such as NVIDIA or NVENC) to accelerate decoding and decode compressed video (such as H.264) into raw data (NV12 format). The demultiplexing and decoding module M1 is deeply integrated with GPU hardware decoding through FFmpeg, and directly retains the decoded data in the GPU memory to avoid bandwidth consumption for transmission to the CPU.
[0084] Optionally, the system dynamically adjusts the model input preprocessing parameters of the machine learning by using the original audio and video stream data parsed by the demultiplexing decoding module M1.
[0085] The present invention extracts the media parameters (such as resolution, frame rate, and color space) of the audio and video stream in real time through the demultiplexing decoding module M1, and dynamically adjusts the input preprocessing strategy of the machine learning model: adaptively scale the input size based on the resolution to avoid image distortion; combine the color space matching data normalization parameters to improve the model reasoning accuracy; intelligently set the timing algorithm interval according to the frame rate to enhance the accuracy of motion compensation. This mechanism realizes zero-copy parameter transfer through GPU video memory sharing, thereby improving the model initialization speed, supporting the automatic adaptation of multi-source heterogeneous videos, and optimizing the video memory allocation efficiency, improving the image quality and reducing the time consumption of reasoning.
[0086] The format conversion module M2 is used to realize bidirectional conversion of data formats on the GPU, convert the decoded video pixel format into the machine learning input format, and convert the color space format of the processed machine learning output into the encoding input format.
[0087] Specifically, the format conversion module M2 implements bidirectional conversion between decoded data (NV12) and machine learning input format (RGB / BGR) on the GPU side, as well as reverse conversion of the format after the inference result is returned (RGB / BGR→NV12). The format conversion module M2 accelerates the conversion through CUDA kernel functions or NPP libraries (NVIDIA Performance Primitives), eliminating the intermediate link in traditional solutions where data needs to be transferred back to the CPU for format processing.
[0088] The machine learning filter module M3 is used to provide a pluggable deep learning model interface and support a variety of image processing algorithms and machine learning frameworks.
[0089] Among them, the machine learning filter module M3 can be implemented based on different mainstream machine learning frameworks, such as Torch, Caffe, TensorFlow, TensorRT and OpenVINO, providing a flexible interface for FFmpeg's dnn filter to call, thereby realizing efficient algorithm processing of audio and video data.
[0090] Among them, the dnn filter is a new filter module introduced in FFmpeg that focuses on machine learning processing. By defining supported data formats and processing methods (such as CPU processing, GPU acceleration or calling third-party libraries, etc.), it realizes intelligent audio and video processing functions such as super-resolution, interpolation, HDR, etc. It fully utilizes GPU acceleration capabilities by flexibly calling interfaces such as TensorRT or Python. While ensuring efficient performance, it supports users to specify filters through simple command lines to complete specified tasks, with high flexibility and scalability.
[0091] Specifically, the machine learning filter module M3 provides a pluggable interface, supports dynamic loading of models in frameworks such as TensorFlow and PyTorch, and executes super-resolution, denoising and other algorithms based on GPU inference engines (such as TensorRT). The machine learning filter module M3 directly receives RGB / BGR data in the GPU memory through a collaborative mechanism, and the processed frames still reside on the GPU to avoid cross-device data transmission. It supports multi-model parallel processing and achieves pipelining through CUDA streams.
[0092] The encoding multiplexing module M4 is used to implement GPU accelerated encoding and audio and video multiplexing and packaging.
[0093] Specifically, the encoding multiplexing module M4 compresses the processed NV12 data by calling the GPU hardware encoder (such as NVIDIA NVDEC), and completes audio and video multiplexing and packaging (such as MP4 / MKV) in the GPU memory through FFmpeg's libavformat. The entire data stream from decoding to encoding in the encoding multiplexing module M4 is closed-loop processed in the GPU memory, completely avoiding the PCIe communication bottleneck between the CPU and GPU.
[0094] The full-link GPU pipeline provided by the audio and video processing system in an embodiment of the present invention is: input audio and video → GPU demultiplexing and decoding → GPU format conversion → GPU machine learning reasoning → GPU format inverse conversion → GPU encoding multiplexing → output file.
[0095] An audio and video processing system provided by an embodiment of the present invention comprises: a demultiplexing and decoding module M1, a format conversion module M2, a machine learning filter module M3 and a coding multiplexing module M4. The system realizes the coordinated work of hardware coding and decoding and machine learning processing through GPU acceleration, wherein: the demultiplexing and decoding module M1 is used for audio and video separation and hardware accelerated decoding, and outputs original audio and video stream data; the format conversion module M2 is used to realize the bidirectional conversion of data formats on the GPU, convert the decoded video pixel format into the machine learning input format, and convert the processed machine learning output color space format into the coding input format; the machine learning filter module M3 is used to provide a pluggable deep learning model interface, supporting a variety of image processing algorithms and machine learning frameworks; the coding multiplexing module M4 is used to realize GPU accelerated coding and audio and video multiplexing packaging. Based on the FFmpeg framework, the present invention innovatively deploys the whole process of demultiplexing, decoding, format conversion, machine learning reasoning, coding and multiplexing on the GPU side, constructs an end-to-end GPU data processing pipeline, thereby significantly reducing PCIe communication overhead, improving system throughput and real-time performance, and providing efficient underlying support for ultra-high-definition video processing and intelligent media applications.
[0096] Optionally, the demultiplexing and decoding module M1 may include: a protocol reading unit, a demultiplexing unit and a hardware decoding unit.
[0097] The protocol reading unit is used to read input audio and video data.
[0098] Specifically, the protocol reading unit is used to call the libavformat module inside the FFmpeg framework to read the network streaming media data or local file data to the CPU for processing. This operation is completed in the CPU, which occupies less resources and has low processing overhead.
[0099] The demultiplexing unit is used to separate the audio and video streams and extract media parameters.
[0100] Specifically, the demultiplexing unit demultiplexes the read audio and video data streams through the libavformat module within the FFmpeg framework, separates the audio stream and the video stream, and extracts the relevant parameter information of the media stream, such as the resolution and frame rate of the video, the sampling rate, number of channels and bit depth of the audio, so as to provide support for the initialization of the subsequent model file. This operation is completed in the CPU, with low resource overhead.
[0101] A hardware decoding unit that performs video decoding on the GPU and outputs data in a video pixel format.
[0102] Specifically, the hardware decoding unit is used to decode the separated audio and video streams through the libavcodec module within the FFmpeg framework and convert them into original uncompressed audio and video stream data. The hardware decoding unit uses the GPU for accelerated processing. For example, when using NVIDIA hardware decoding, the decoded video pixel data format is NV12.
[0103] The demultiplexing decoding module M1 provided in the embodiment of the present invention significantly improves the audio and video processing efficiency through multi-level optimization: the protocol reading unit reads the input data based on FFmpeg's libavformat in a lightweight manner, releasing GPU computing resources; the demultiplexing unit quickly separates the audio and video streams and accurately extracts key media parameters such as resolution and frame rate, providing a pre-configuration basis for subsequent modules and shortening the system initialization time; the hardware decoding unit relies on the GPU parallel architecture to improve the video decoding speed, and the decoded NV12 data directly resides in the GPU memory, eliminating the PCIe transmission link of the decoded data returned to the CPU in the traditional solution, effectively reducing bandwidth occupancy, reducing end-to-end delay, and improving system throughput, effectively supporting real-time processing requirements in high-resolution and multi-concurrency scenarios.
[0104] Optionally, the format conversion module M2 may include: a first conversion unit and a second conversion unit.
[0105] The first conversion unit is used to convert the video pixel format into a color space format for machine learning processing.
[0106] The first conversion unit is used to convert the original video pixel format after hardware decoding into the color space format (such as RGB or BGR) required by the machine learning model to meet the requirements of machine learning inference processing.
[0107] The second conversion unit is used to inversely convert the processed color space format into a video pixel format for hardware encoding.
[0108] The second conversion unit is used to inversely convert the color space format (such as RGB or BGR) after the machine learning inference processing into the video pixel format required by the hardware encoding, so as to be encoded by the hardware encoder;
[0109] Optionally, the format conversion process is accelerated by CUDA.
[0110] The format conversion module M2 provided in an embodiment of the present invention achieves significant efficiency gains through full GPU-side collaborative processing: the first conversion unit efficiently converts the decoded NV12 pixel format into the RGB / BGR format required for machine learning based on CUDA parallel computing, and the second conversion unit reversely quickly restores the RGB / BGR results after algorithm processing to the NV12 encoding format. Both stages of conversion are completed in a closed loop in the GPU memory, eliminating the cross-device transmission of data returned to the CPU for format processing in traditional solutions, and combining CUDA acceleration to reduce the overall conversion time, providing seamless technical guarantees for real-time machine learning processing and encoding of high-resolution videos.
[0111] Optionally, the machine learning filter module M3 may include: a model initialization unit and a data processing unit.
[0112] Model initialization unit, used to dynamically load pre-trained models based on media parameters.
[0113] Specifically, the model initialization unit is used to dynamically load the offline pre-trained machine learning model according to the parameter information of the media stream. By parsing the properties of the audio and video stream (such as resolution, frame rate, etc.), the adapted model file is initialized and the corresponding configuration file is loaded to meet specific business needs.
[0114] Optionally, a model initialization unit is specifically used to automatically adapt the model input size according to the resolution and frame rate parameters of the input video, and support model deployment of at least one inference framework.
[0115] Specifically, the model initialization unit can determine whether the machine learning model has been initialized. If it has not been initialized, the offline trained model file is loaded, and the dynamic initialization of the model is completed according to the parsed media parameters. At the same time, the configuration file of the machine learning model is loaded, and the specific algorithm model to be processed is determined according to business needs.
[0116] The data processing unit is used to convert the video frame into a tensor format and perform specified algorithm processing, wherein the specified algorithm includes at least one of image super-resolution, denoising and restoration, frame interpolation, and HDR processing.
[0117] Specifically, the data processing unit is used to receive the decoded audio and video raw data, convert it into a tensor format acceptable to the model, and execute the specified machine learning algorithm processing. These algorithms may include but are not limited to image super-resolution, denoising and restoration, interpolation, HDR processing or image quality assessment. After the processing is completed, the results are returned to the subsequent module for further processing or encoding operations.
[0118] Furthermore, the data processing unit can receive audio and video data, convert it into a format supported by the model, and then call the specified algorithm model for processing. After the processing is completed, the result is returned to the filter or encoding multiplexing module M4 of FFmpeg to support subsequent audio and video workflows.
[0119] The machine learning filter module M3 provided in the embodiment of the present invention realizes efficient processing through intelligent model management and collaborative computing on the GPU side: the model initialization unit dynamically adapts and loads the pre-trained model based on media parameters (such as resolution and frame rate), avoiding resource mismatch caused by manual configuration of the model in traditional solutions, and improving model loading efficiency; the data processing unit directly converts the video frames in the GPU memory into tensor format through CUDA acceleration, and calls the optimized inference engine of TensorRT / OpenVINO to execute super-resolution, denoising and other algorithms, thereby reducing the time consumption of single-frame processing, and supporting multi-algorithm parallel pipelining (such as interpolation + HDR composite processing), combined with the GPU memory zero-copy mechanism, eliminating the transmission overhead of data to and from the CPU, improving the overall system throughput, and meeting the low-latency and high-precision requirements of real-time ultra-high-definition video enhancement and intelligent processing.
[0120] Optionally, the encoding multiplexing module M4 may include: a format inverse conversion unit, a hardware encoding unit and a multiplexing packaging unit.
[0121] The format inverse conversion unit is used to realize data format conversion.
[0122] Specifically, the format inverse conversion unit calls the format_cuda filter module to receive the original audio and video data (such as video data in RGB or BGR format) processed by the machine learning filter module M3, and converts it into a data format acceptable to the hardware encoder (such as NV12), that is, inversely converts the data format to adapt to the hardware encoding requirements.
[0123] Hardware encoding unit for GPU-accelerated video encoding.
[0124] Specifically, the hardware encoding unit calls the libavcodec module inside FFmpeg to encode and compress the converted audio and video streams based on GPU acceleration. The video stream is compressed using hardware encoding acceleration, and the audio stream is encoded according to specific needs.
[0125] The multiplexing packaging unit is used to generate output files that conform to the target packaging format.
[0126] Specifically, the multiplexing and packaging unit calls the libavformat module inside FFmpeg to multiplex and package the encoded audio and video streams, generate output files that conform to the target encapsulation format, form playable audio and video finished files, and complete the entire processing flow.
[0127] The encoding multiplexing module M4 provided in the embodiment of the present invention realizes efficient output through closed-loop processing on the full GPU side: the format inverse conversion unit completes the lossless conversion of the processing result to the encoding format (such as NV12) in the GPU memory based on the format_cuda filter, eliminating the cross-device transmission of data back to the CPU in the traditional solution; the hardware encoding unit calls the GPU accelerated encoder to realize real-time compression of the video and improve the encoding speed; the multiplexing and packaging unit completes the synchronization and encapsulation of the audio and video streams (such as MP4 / MKV) directly in the GPU memory, avoiding the return copy of the final data to the CPU memory, reducing the overall encoding multiplexing delay, saving video processing PCIe bandwidth, supporting the concurrent output of multiple high-code streams, and providing end-to-end high-throughput, low-latency encoding and packaging capabilities for ultra-high-definition video processing.
[0128] The audio and video processing system provided by the embodiment of the present invention adds two GPU-based high-efficiency filter modules in the FFmpeg framework to optimize the processing link of audio and video data. One is a data format conversion filter, which is used to efficiently convert between the NV12 format generated by hardware encoding and decoding and the RGB or BGR format required for machine learning reasoning; the second is a reasoning acceleration filter, which is used to complete the data interaction and reasoning process between the audio and video data and the machine learning algorithm module, and realize GPU accelerated processing at the same time. Except for a small amount of I / O reading and writing and transpackaging work completed by the CPU, all other core processing links are executed on the GPU, from data decoding to reasoning, and then to encoding, forming a complete GPU acceleration link. By reducing the amount of data transmission between the GPU and the CPU, the system performance bottleneck problem caused by insufficient PCIe bandwidth is effectively solved, and the throughput and efficiency of audio and video processing tasks are greatly improved. In addition, the present invention supports a variety of mainstream deep learning frameworks (such as Torch, TensorFlow, TensorRT and OpenVINO, etc.), provides a flexible pluggable interface, and easily supports complex audio and video intelligent processing tasks such as super-resolution, interpolation, HDR, etc., which significantly reduces the engineering landing cost of machine learning algorithms.
[0129] Figure 3The figure shows a logical framework diagram of audio and video processing provided by an embodiment of the present invention. First, the demultiplexing decoding module M1 is used to parse the audio and video file into an undecoded media stream, and then the undecoded media stream is parsed into the original uncompressed audio and video data (such as video frame data and audio PCM data) for use in subsequent processing steps. Next, the format conversion module M2 is used to call the format_cuda filter: the storage format of the original uncompressed data is converted into a data format suitable for machine learning model processing. For example, the data is converted from NV12 format to RGB or BGR format through the CUDA interface. The format conversion module M2 can complete the data format conversion in the GPU video memory through the coded encapsulated CUDA interface (libformat_cuda_kernel.so), while avoiding data backhaul between the CPU and the GPU to improve efficiency. Then, the machine learning filter module M3 judges the model file type. If the input machine learning model file suffix is .trt, it means that the model has been an accelerated model optimized by TensorRT. At this time, the C language version of the mgtvsr_tensorrt filter is called to directly use the CUDA hardware interface and the TensorRT interface for reasoning. If the input model file has the suffix .pt, it is a PyTorch model. In this case, the Python version of the mgtvsr_python filter is called to perform machine learning algorithm reasoning through the Python interface, relying on the CUDA library's support for Python. The Python version of the filter uses the following logic: Figure 4 As shown. Whether it is the C version or the Python version filter, two global static storage spaces need to be opened in the GPU video memory for data transfer between the filter and the inference algorithm. At the same time, the core logic of video memory allocation and data copying is implemented by the encapsulated CUDA interface (libformat_cuda_kernel.so library compiled by .cu file). Before the AI algorithm is processed, first check whether the model has been initialized. If not, load the model file and complete the initialization. After the initialization is completed, use the AI algorithm to perform specific processing tasks, such as super resolution, interpolation, repair, HDR processing, image quality assessment or image enhancement. After the AI algorithm is processed, call the format_cuda filter again to convert the processed data from RGB or BGR format back to a format that the hardware encoder can accept (such as NV12). This step is also completed in the GPU video memory, without the need to return data to the CPU, thus avoiding the performance bottleneck of PCIe data transmission. Finally, the encoding multiplexing module M4 is used to encode the format-converted audio and video raw data into the target standard video stream format through the hardware encoder. Encapsulate the encoded audio and video stream into an audio and video file in the target encapsulation format and write it to the storage device.
[0130] By using the format_cuda filter and related CUDA interfaces, the embodiment of the present invention completes all data format conversion and processing in the GPU video memory, avoiding multiple data transmissions between the CPU and the GPU in the traditional method, and significantly improving the processing efficiency.
[0131] The present invention provides an audio and video processing system and method for realizing full-link GPU inference acceleration based on the FFmpeg framework, focusing on improving the collaborative working efficiency of the GPU and the CPU under a heterogeneous computing architecture. By handing over all large-scale data computing tasks to the GPU for execution, the amount of data interaction between the GPU and the CPU is reduced, the processing speed is significantly improved, and the overall production efficiency of audio and video machine learning processing is optimized. At the same time, the method focuses on the versatility of the model, supports the application requirements of the C language and Python versions of the machine learning algorithm in different scenarios, thereby effectively reducing the cost of algorithm engineering implementation. This solution is widely used in the field of audio and video processing based on machine learning algorithms, and provides an efficient solution to the problem of reduced system efficiency due to insufficient PCIe bandwidth in the x86 architecture, greatly improving the throughput and performance of the system.
[0132] An embodiment of the present invention provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the audio and video processing method is implemented.
[0133] An embodiment of the present invention provides a processor, which is used to run a program, wherein the audio and video processing method is executed when the program is running.
[0134] like Figure 5 As shown, an embodiment of the present invention provides an electronic device 1000, which includes at least one processor 1001, at least one memory 1002 connected to the processor 1001, and a bus 1003; wherein the processor 1001 and the memory 1002 communicate with each other through the bus 1003; the processor 1001 is used to call the program instructions in the memory 1002 to execute the above-mentioned audio and video processing method. The electronic device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0135] The present invention also provides a computer program product, which, when executed on an electronic device, is suitable for executing a program that initializes the steps of the audio and video processing method.
[0136] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatuses, electronic devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable device to generate a machine, so that the instructions executed by the processor of the computer or other programmable device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0137] In a typical configuration, an electronic device includes one or more processors (CPU), a memory, and a bus. The electronic device may also include an input / output interface, a network interface, and the like.
[0138] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip. The memory is an example of a computer-readable medium.
[0139] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0140] In the description of the present invention, it should be understood that the terms "up", "down", "front", "back", "left" and "right" etc. indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description. They do not indicate or imply that the positions or elements referred to must have specific directions, be constructed and operate in specific directions. Therefore, they should not be understood as limitations of the present invention.
[0141] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. It should also be noted that the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0142] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0143] The above are only embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of the present invention.
Claims
1. An audio and video processing method, characterized in that: include: The hardware encoding and decoding process is performed by GPU acceleration to demultiplex and decode the input audio and video data, and output the first audio and video stream data; Convert the first audio and video stream data from a video pixel format to a machine learning input format on the GPU, and convert the color space format of the machine learning output to an encoding input format; Performing video processing on the audio and video stream data converted into the machine learning input format using a deep learning model to obtain second audio and video stream data, wherein the deep learning model supports a variety of image processing algorithms and machine learning frameworks; Converting the second audio and video stream data from the machine learning input format to the video pixel format on the GPU; The encoding process is performed through GPU acceleration, and the second audio and video stream data converted into the video pixel format is multiplexed, packaged and output.
2. The method according to claim 1, characterized in that The step of demultiplexing and decoding the input audio and video data and outputting the first audio and video stream data includes: Read input audio and video data; Separating the audio and video elementary streams and media parameters of the audio and video data; The audio and video basic stream is acceleratedly decoded on the GPU to generate first audio and video stream data in a video pixel format.
3. The method according to claim 2, characterized in that Also includes: Dynamically adjust the model input preprocessing parameters of the deep learning model based on the media parameters.
4. The method according to claim 1, characterized in that: The step of using a deep learning model to perform video processing on the audio and video stream data converted into the machine learning input format to obtain second audio and video stream data includes: According to the resolution and frame rate of the audio and video stream data converted into the machine learning input format, a pre-trained deep learning model is dynamically loaded, and the video frames of the audio and video stream data are converted into a tensor format and then an image processing algorithm is executed to obtain a second audio and video stream data, wherein the image processing algorithm includes at least one of image super-resolution reconstruction, video noise reduction and restoration, frame rate improvement interpolation and HDR effect enhancement.
5. The method according to claim 4, characterized in that The method of dynamically loading a pre-trained deep learning model according to the resolution and frame rate of the audio and video stream data converted into the machine learning input format includes: According to the resolution and frame rate of the audio and video stream data converted into the machine learning input format, the model input size is automatically matched, and a deep learning model that supports at least one machine learning framework including Torch, Caffe, Tensorflow, TensorRT and OpenVINO is loaded.
6. The method according to claim 1, characterized in that The encoding process is performed by GPU acceleration to multiplex, package and output the second audio and video stream data converted into the video pixel format, including: Calling the GPU hardware encoder to compress and encode the second audio and video stream data converted into the video pixel format; The encoded second audio and video stream data is multiplexed, packaged and output according to the target encapsulation format.
7. The method according to any one of claims 1 to 6, characterized in that The video pixel format is YUV format or NV12 format, and the machine learning input format is RGB format or BGR format.
8. An audio and video processing system, characterized in that: include: Demultiplexing decoding module, format conversion module, machine learning filter module and encoding multiplexing module, the system realizes the collaborative work of hardware encoding and decoding and machine learning processing through GPU acceleration, wherein: The demultiplexing and decoding module is used for audio and video separation and hardware accelerated decoding, and outputs original audio and video stream data; The format conversion module is used to implement bidirectional data format conversion on the GPU, converting the decoded video pixel format into the machine learning input format, and converting the processed machine learning output color space format into the encoding input format; The machine learning filter module is used to provide a pluggable deep learning model interface and support a variety of image processing algorithms and machine learning frameworks; The encoding multiplexing module is used to implement GPU accelerated encoding and audio and video multiplexing packaging.
9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the audio and video processing method according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: The electronic device includes at least one processor, and at least one memory and a bus connected to the processor; wherein the processor and the memory communicate with each other through the bus; the processor is used to call the program instructions in the memory to execute the audio and video processing method as described in any one of 1 to 7.
Citation Information
Cited By
Multi-source audio and video access gateway system based on machine learning
CN121173981A