System architecture for video processing
Through a software-defined VPU architecture based on RISC-V architecture, the integration of high-performance computing cores and ASIC-specific acceleration units is solved, and the problem of inefficiency of traditional video processing systems is achieved, and efficient and flexible video processing and storage solutions are achieved.
Patent Information
- Application Number
- CN202510622887.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-14
AI Technical Summary
Traditional CPU computing methods are inefficient, GPU costs are high and flexibility are limited, and ASIC VPUs are difficult to meet diversified and flexible needs in cloud computing scenarios. Existing video processing systems have challenges in efficient transmission and storage.
It adopts a software-defined VPU architecture (SD-VPU) based on RISC-V architecture, integrates high-performance computing cores and ASIC-specific acceleration units, supports video transcoding, compression, monitoring and other functions, processes video streams in parallel, and combines AI computing power units for image quality enhancement and coding optimization.
Significantly improve the programmability of VPUs, reduce host communication frequency, improve video processing efficiency, reduce storage costs, and adapt to diversified video processing needs.
Smart Images

Figure CN120499342A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of video processing, and in particular to a system architecture for video processing. Background Art
[0002] With the rapid development of video surveillance and visual networking services, the popularization of ultra-high-definition video (such as 4K / 8K) and the expansion of intelligent application scenarios, the amount of video data has exploded, bringing unprecedented challenges to video transmission and storage systems.
[0003] Traditional CPU computing methods are inefficient, and although GPUs have certain processing capabilities, their high cost, low resource utilization, and limited flexibility make it difficult to meet the needs of large-scale video services. Traditional ASIC-based VPUs (Video Processing Units), although they have certain performance advantages, their high cost and customization limitations make it difficult to meet the diverse and elastic needs of cloud computing scenarios.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention
[0005] Currently, there is a lack of VPU products based on the RISC-V (Reduced Instruction Set Computer V) architecture. This paper provides a system architecture for video processing based on the open nature and modular design of the RISC-V architecture. At the hardware level, it adopts software-defined VPU architecture (SD-VPU) technology, integrates a high-performance computing core (RISC-V) and an ASIC-specific acceleration unit, significantly enhances the programmability of the VPU, effectively reduces the frequency of communication with the host, and thus improves efficiency.
[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.
[0007] According to one aspect of the present disclosure, a system architecture for video processing is provided, comprising a computing power base, a driver layer, a service layer, and an application layer arranged in sequence;
[0008] The computing power base hardware includes multiple video processing cards, each of which includes multiple parallel video processing chips. Each video processing chip includes a RISC-V computing core and a video processing unit. The multiple parallel video processing chips are all connected to a PCIe switch.
[0009] The application layer is used to initiate video processing tasks;
[0010] The video processing service engine in the service layer allocates video processing tasks to target video processing chips based on the task requirements of the video processing tasks and the load of the RISC-V computing core in each video processing card;
[0011] The driver layer is used to perform standardized preprocessing on the input video corresponding to the video processing task, and push the processed video to the target video processing chip through the PCIe switch. The RISC-V computing core and video processing unit of the video processing chip work together to process the video.
[0012] In one embodiment of the present disclosure, each video processing chip further includes an AI computing unit that uses algorithms to analyze video content to achieve image quality enhancement, intelligent adjustment of encoding parameters, and / or scene recognition;
[0013] The processed video is pushed to the RISC-V computing core of the target video processing chip through the PCIe switch. The RISC-V computing core calls the video processing unit and the AI computing power unit to collaboratively process the video according to the task requirements of the video processing task.
[0014] In one embodiment of the present disclosure, the video processing service engine in the service layer distributes the video processing tasks to multiple target video processing chips based on the task requirements of the video processing tasks and the load of the RISC-V computing core in each video processing card. The multiple target processing chips are used to simultaneously process multiple video streams or simultaneously process different parts of a video.
[0015] In one embodiment of the present disclosure, the multiple target processing chips belong to the same video processing card, or the multiple target processing chips belong to multiple video processing cards.
[0016] In one embodiment of the present disclosure, the video processing service engine in the service layer allocates video processing tasks to one or more target video processing chips based on the task requirements of the video processing tasks and hardware resource status information. The hardware resource status information includes the load and hardware temperature of the RISC-V computing core in each video processing card.
[0017] In one embodiment of the present disclosure, the application layer initiates multiple video processing tasks; the video processing service engine of the service layer is further configured to adjust the execution order of the multiple video processing tasks in combination with task priorities and hardware resource status information.
[0018] In one embodiment of the present disclosure, the driver layer also collects hardware resource status information in real time and feeds the hardware resource status information back to the video processing service engine of the service layer.
[0019] In one embodiment of the present disclosure, each video processing card is provided with a separate baseboard management controller, which is used to monitor the hardware status of the transcoding card in real time; the driver layer is provided with a status monitoring interface, and the driver layer obtains hardware resource status information from each baseboard management controller through the status monitoring interface.
[0020] In one embodiment of the present disclosure, the video processing task includes one or more of the following tasks:
[0021] Video transcoding, video compression, video monitoring, hyperparameter adaptive optimization, splicing, scaling, and AI bitrate control.
[0022] In one embodiment of the present disclosure, the video processing service engine of the service layer also dynamically adjusts the transcoding parameters of the input video corresponding to the video processing task for pre-processing at the driver layer in combination with the task load and hardware resource status information.
[0023] A system architecture for video processing provided by the embodiments of the present disclosure adopts software-defined VPU architecture (SD-VPU) technology at the hardware level. Each video processing card includes multiple parallel video processing chips, and each video processing chip integrates a RISC-V computing core and a video processing unit, which significantly improves the programmability of the VPU; multiple parallel video processing chips can be processed in parallel to improve efficiency; the video processing service engine in the service layer allocates video processing tasks to target video processing chips based on the task requirements of the video processing tasks and the load of the RISC-V computing core in each video processing card. The RISC-V computing core and video processing unit of the video processing chip collaborate to process the video, and through unified allocation and scheduling of tasks, VPU cluster resource optimization and performance improvement are achieved.
[0024] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0026] Obviously, the drawings described below are only some embodiments of the present disclosure. A person skilled in the art can obtain other drawings based on these drawings without any creative effort.
[0027] Figure 1 A schematic diagram of a system architecture for video processing according to an embodiment of the present disclosure is shown;
[0028] Figure 2 A schematic diagram showing the architecture of a video processing card in an embodiment of the present disclosure is shown;
[0029] Figure 3 A schematic diagram of a pure ASIC architecture according to an embodiment of the present disclosure is shown;
[0030] Figure 4 A schematic diagram of a control core + ASIC dedicated unit architecture is shown in an embodiment of the present disclosure;
[0031] Figure 5 A schematic diagram illustrating an architecture of a high-performance computing core + ASIC dedicated unit in an embodiment of the present disclosure is shown;
[0032] Figure 6 A schematic diagram of another system architecture for video processing according to an embodiment of the present disclosure is shown;
[0033] Figure 7 A schematic diagram showing the architecture of another video processing card in an embodiment of the present disclosure is shown;
[0034] Figure 8 A schematic diagram showing the effect of high bit rate video transcoding in an embodiment of the present disclosure is shown;
[0035] Figure 9 A schematic diagram showing the effect of low-bitrate video transcoding in an embodiment of the present disclosure is shown;
[0036] Figure 10a and 10b A schematic diagram showing the effect of low-bitrate video transcoding in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. The components of the embodiments of the present disclosure generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the disclosure for which protection is sought, but merely represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present disclosure.
[0038] For ease of understanding, the following first explains the relevant concepts involved in this disclosure as follows:
[0039] RISC-V (Reduced Instruction Set Computer V): is an open source instruction set architecture developed by the University of California, Berkeley. It adopts the principles of reduced instruction set computing and has a simple, efficient and modular design. It supports multiple data widths (such as 32 bits, 64 bits, and 128 bits), does not require instruction set authorization, and has low cost and low power consumption.
[0040] VPU (Video Processing Unit): Video processing unit is a new core engine of the video processing platform. It has hard decoding function and the ability to reduce CPU load, which can reduce server load and network bandwidth consumption. It is used to distinguish it from the traditional GPU (Graph Process Unit).
[0041] H.264: A highly compressed digital video codec standard developed by the Joint Video Team (JVT), a joint effort of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG). This standard is often referred to as H.264 / AVC (or AVC / H.264, H.264 / MPEG-4 AVC, or MPEG-4 / H.264 AVC), specifically identifying its two developers.
[0042] H.265: High Efficiency Video Coding (HEVC), also known as H.265 and MPEG-H part 2, is a video compression standard and one of several potential successors to the widely used AVC (H.264 or MPEG-4 part 10). Compared to AVC, HEVC provides approximately twice the data compression ratio at the same video quality level, or significantly improves video quality at the same bit rate. It supports resolutions up to 8192×4320, including 8K UHD.
[0043] With the rapid development of cloud computing and video applications, the demand for cloud storage and computing resources for video encoding and decoding technologies continues to rise. On the one hand, the demand for network bandwidth for high-definition video transmission has skyrocketed, especially in multi-channel concurrent and real-time monitoring scenarios, where traditional solutions struggle to meet efficient transmission requirements. On the other hand, the long-term storage of massive amounts of video data has led to increased storage pressure, and existing storage architectures are gradually experiencing bottlenecks in capacity, cost, and efficiency. Furthermore, the heterogeneity of encoding formats (such as H.264, H.265, and AV1) in different business scenarios has further increased system complexity and processing costs. At the same time, with intensifying market competition, customers are placing higher demands on video storage services, demanding not only storage capacity that can quickly adapt to data growth but also services with higher cost-performance and enhanced video processing capabilities. This poses dual challenges to cost control and service quality for visual networking services, and also raises expectations for technological innovation in the underlying hardware and system architecture.
[0044] Traditional CPU computing methods are inefficient, and while GPUs offer some processing power, their high cost, low resource utilization, and limited flexibility make them inadequate for large-scale video services. Against this backdrop, the VPU (Video Processing Unit) emerged as a specialized processor for accelerating image and video processing. It offers hardware-based encoding and decoding, image stitching, and post-processing acceleration, effectively reducing the burden on server CPUs and network bandwidth consumption, significantly reducing the number of racks and operating costs in data centers.
[0045] Furthermore, raw video data contains a significant amount of redundant information. While video transcoding and compression technologies can remove redundant data to reduce video size, data with weak scene relevance and low user interest may still exist. With the continuous advancement of AI technology, AI-based video analysis can identify key details in the image, achieving deeper compression and redundancy removal. However, using separate hardware accelerator cards to handle video transcoding and AI recognition tasks will undoubtedly increase costs.
[0046] Driven by market demand, ASIC-based VPUs have emerged, but VPUs based on the RISC-V architecture are still a niche. While traditional ASIC VPUs offer certain performance advantages, their high cost and customization limitations make them difficult to meet the diverse and flexible demands of cloud computing. Leveraging the open nature and modular design of the RISC-V architecture, developing adaptive video transcoding solutions for cloud computing scenarios has become a breakthrough.
[0047] In this context, for high-concurrency video transcoding application scenarios such as video cloud and security, the disclosed embodiments propose a system architecture for video processing. At the hardware level, it adopts software-defined VPU architecture (SD-VPU) technology, integrates high-performance computing cores (RISC-V) and ASIC dedicated acceleration units, significantly improving the programmability of VPU. TeleVPU supports flexible super-resolution virtualization and is equipped with larger-capacity video memory. It can cache multiple AI models and a large number of video files at the same time, effectively reducing the frequency of communication with the host, thereby improving efficiency. At the software level, it provides full-stack driver service capabilities, supporting key functions including video transcoding, video compression, video monitoring, hyperparameter adaptive optimization, splicing, scaling, AI bit rate control, and multi-card and multi-core scheduling, providing users with an innovative and cost-effective video processing solution.
[0048] The defects of the above solutions and the proposed solutions are the results obtained by the inventor after practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed in the present disclosure for the above problems below should be the contributions made by the inventor to the present disclosure during the disclosure process.
[0049] This exemplary implementation is described in detail below with reference to the accompanying drawings and examples.
[0050] Figure 1 A system architecture for video processing in an embodiment of the present disclosure is shown. Figure 1 As shown, the system architecture for video processing provided in the embodiment of the present disclosure includes a computing power base, a driver layer, a service layer and an application layer arranged in sequence.
[0051] In some embodiments, the hardware of the computing power base includes multiple video processing cards (TeleVPU cards), each of which includes multiple video processing chips connected in parallel. Figure 2 As shown, each video processing chip includes a RISC-V computing core and a video processing unit (NPU), and multiple parallel video processing chips are connected to a PCIe switch. In some embodiments, the video processing unit is an ASIC dedicated acceleration unit.
[0052] In some embodiments, the application layer is used to initiate video processing tasks; the video processing service engine in the service layer allocates video processing tasks to the target video processing chip based on the task requirements of the video processing task and the load of the RISC-V computing core in each video processing card; the driver layer is used to perform standardized pre-processing on the input video corresponding to the video processing task and push the processed video to the target video processing chip via a PCIe switch. The RISC-V computing core and video processing unit of the video processing chip collaborate to process the video. In some embodiments, the video processing tasks include one or more of the following tasks: video transcoding, video compression, video monitoring, hyperparameter adaptive optimization, splicing, scaling, and AI bitrate control.
[0053] In some embodiments, each video processing chip also includes an AI computing unit, which uses algorithms to analyze video content to achieve image quality enhancement, intelligent adjustment of encoding parameters and / or scene recognition; the processed video is pushed to the RISC-V computing core of the target video processing chip through a PCIe switch, and the RISC-V computing core calls the video processing unit and the AI computing unit to collaboratively process the video according to the task requirements of the video processing task.
[0054] In some embodiments, the video processing card (TeleVPU card) hardware may be composed of five high-performance video processing chips connected in parallel, such as Figure 2 As shown, each chip features a powerful video processing unit (video processing accelerator) and a quad-core RISC-V processor. The five computing chips exchange data with the server host via PCIE ports. Intensive computations such as video data decoding, encoding, and post-processing are all performed internally on the chip, leaving the host only to perform light-load computations such as scheduling.
[0055] In some embodiments, the video processing card (TeleVPU) features independent board-level BMC management to integrate seamlessly into the server cluster's operations and maintenance system. The five accelerator chips operate independently, ensuring that a video task blockage or anomaly on one chip does not affect the normal operation of the other chips. Furthermore, under the control of the VPU card service software, transcoding tasks and video streams can be switched between the chips in real time, ensuring disaster recovery and robustness for massive transcoding tasks.
[0056] In some embodiments, video data is accessed to the video processing card (TeleVPU card) at high speed through a PCIe Switch. As a data exchange hub, the PCIe Switch can distribute data to each RISC-V computing core to achieve parallel processing and improve efficiency.
[0057] In some embodiments, the video processing service engine in the service layer distributes video processing tasks to multiple target video processing chips based on the task requirements and the load of the RISC-V compute cores in each video processing card. These target processing chips are used to simultaneously process multiple video streams or different portions of a single video. Each RISC-V compute core acts as a core control unit, leveraging its flexible and efficient architecture to dispatch the associated video processing unit (NPU) and AI computing unit. Based on task requirements, the RISC-V compute core assigns video transcoding tasks to the NPU and can also call upon the AI unit for auxiliary processing. The NPU provides hardware acceleration for video codecs, supporting conversion to multiple formats such as H.264 and H.265, rapidly processing video encoding, compression, and decompression to ensure transcoding speed and compatibility. The AI computing unit uses algorithms to analyze video content and implement functions such as image quality enhancement, intelligent adjustment of encoding parameters, and scene recognition, improving transcoding quality and adaptability to meet diverse needs. Multiple RISC-V compute cores and supporting units working in parallel can simultaneously process multiple video streams or different portions of high-resolution video, significantly improving overall transcoding performance and adapting to high-load scenarios.
[0058] Traditional VPU is mainly divided into two routes, such as Figure 3 As shown, it has a pure ASIC dedicated unit, or as Figure 4 As shown, it has a control core + ASIC dedicated unit, lacks general computing power, and has poor programmability. With the emergence of AI-driven video compression technology, applications have gradually increased their demands for VPU programmability and cache, and traditional VPUs are difficult to meet emerging needs. The video processing card (TeleVPU card) of the RISC-V instruction set proposed in the present embodiment adopts the software-defined VPU architecture (SD-VPU) technology route, with built-in high-performance computing core (RISC-V) and ASIC dedicated unit. The VPU has stronger programmability and supports more flexible super-resolution virtualization, etc., and the VPU side memory is larger, supporting the simultaneous caching of multiple AI models and a large number of video files, avoiding frequent communication with the host, such as Figure 5 shown.
[0059] The software of the video processing card (TeleVPU card) is divided into the host side and the VPU side. The host side software runs on the server host, and its function is to integrate the video computing resources of up to 8 TeleVPU cards installed on a server, and map them into an easy-to-use software interface that can be directly called on the host. At the same time, it provides multi-card task overall scheduling, video stream distribution, hardware status monitoring, and real-time load statistics services. The VPU side software runs inside the card and performs specific computing tasks. The host side software provides a general Ffmpeg interface that can be seamlessly integrated into the existing soft transcoding system. At the same time, it provides whole-machine background service software at the streaming protocol level such as RTSP and RTMP, supports direct management of video streaming tasks through the HTTP command interface outside the server, and quickly integrates the multi-card server as a whole product into the monitoring and live broadcast system without having to pay attention to the hardware resource allocation management and load balancing inside the server. The overall architecture of RISC-V software is as follows Figure 6 shown.
[0060] The underlying hardware layer is based on the TeleVPU video transcoding computing power of the SD-VPU architecture, enabling efficient execution of key functions such as video transcoding, compression, and splicing. This provides high-density, low-cost, and customizable video transcoding and compression capabilities for video processing scenarios such as the Internet of Vision. The driver layer constructs the advanced encapsulation SDK OpenVPU, which covers key capabilities such as video transcoding, compression, splicing, AI bitrate control, and hyperparameter adjustment, shielding application complexity and facilitating agile development for users. The service layer uses the VPU Server to empower upper-layer applications, responsible for single-card multi-core, multi-level multi-card, distributed server, and multi-task scheduling. By uniformly allocating and scheduling tasks, it optimizes VPU cluster resources and improves performance. The application layer handles key user services such as video transcoding, video compression, video splicing, and video confluence.
[0061] In some embodiments, the application layer initiates a video processing task, such as a video transcoding request, and calls the video transcoding function of the service layer VPUServer engine.
[0062] In some embodiments, the service layer is responsible for task parsing and resource scheduling, including single-card multi-core management, multi-machine multi-card management, and multi-task scheduling.
[0063] In some embodiments, the application layer is used to initiate video processing tasks; the video processing service engine in the service layer allocates the video processing tasks to one or more target video processing chips based on the task requirements of the video processing tasks and the load of the RISC-V computing core in each video processing card. In some embodiments, multiple target processing chips belong to the same video processing card, or multiple target processing chips belong to multiple video processing cards.
[0064] Multiple target processing chips belong to the same video processing card, which is the aforementioned single-card multi-core management. If the task is processed on a single card, the VPU Server engine allocates the transcoding task to the appropriate computing core based on the transcoding task requirements and the load of each core on the single card.
[0065] Multiple target processing chips belong to multiple video processing cards, which means multi-machine multi-card management (if cross-card is required). If the task load is high or the resources of a single card are insufficient, the multi-machine multi-card management scans the global resource pool to determine the distribution of multiple cards involved in transcoding, and uses distributed services to call resources across cards.
[0066] In some embodiments, the video processing service engine in the service layer allocates video processing tasks to one or more target video processing chips based on the task requirements of the video processing tasks and hardware resource status information. The hardware resource status information includes the load and hardware temperature of the RISC-V computing core in each video processing card.
[0067] In some embodiments, the application layer initiates multiple video processing tasks; the video processing service engine of the service layer is further configured to adjust the execution order of the multiple video processing tasks by combining task priorities with hardware resource status information.
[0068] In some embodiments, the driver layer also collects hardware resource status information in real time and feeds the hardware resource status information back to the video processing service engine of the service layer.
[0069] In some embodiments, each video processing card is provided with a separate baseboard management controller, which is used to monitor the hardware status of the transcoding card in real time; the driver layer is provided with a status monitoring interface, and the driver layer obtains hardware resource status information from each baseboard management controller through the status monitoring interface, such as Figure 7 The BMC (Baseboard Management Controller) monitors the transcoding card hardware status (temperature, power consumption, etc.) in real time to ensure stable device operation and supports remote management for easy maintenance and troubleshooting.
[0070] In some embodiments, the video processing service engine of the service layer also combines the task load and hardware resource status information to dynamically adjust the transcoding parameters of the input video corresponding to the video processing task for pre-processing at the driver layer. The service layer calls the video transcoding interface of OpenVPUSDK. After receiving the instruction, the driver layer performs standardized pre-processing (such as format conversion and decoding) on the input video based on the FFmpeg video processing library. Hyperparameter adaptive optimization is used to analyze the task load and hardware status, and dynamically adjust transcoding parameters (such as resolution and frame rate) to improve efficiency. Call AI bitrate control to optimize the encoding process, reducing bitrate waste while ensuring image quality.
[0071] In some embodiments, the driver layer pushes the processed video data to the RISC-V VPU computing core and supporting units (such as the NPU) for transcoding. The OpenVPU SDK's status monitoring interface collects hardware status (such as computing power usage and temperature) in real time and feeds it back to the VPU Server engine in the service layer. Based on the monitoring feedback, the VPU Server engine dynamically adjusts resource allocation through multi-tasking scheduling. For example, if a core overheats, some tasks will be migrated to other idle cores.
[0072] After transcoding is completed, the driver layer integrates the transcoded data through the FFmpeg library, and the service layer returns the results to the application layer, completing a complete resource management and calling process.
[0073] The architecture of the disclosed embodiment achieves efficient, flexible and high-quality video transcoding processing through a combination of hardware acceleration, parallel processing and intelligent optimization.
[0074] The disclosed embodiments adopt software-defined VPU architecture (SD-VPU) technology at the hardware level, integrating high-performance computing cores (RISC-V) and ASIC-specific acceleration units, significantly improving the programmability of the VPU. TeleVPU supports flexible super-resolution virtualization and is equipped with larger-capacity video memory. It can cache multiple AI models and a large number of video files at the same time, effectively reducing the frequency of communication with the host, thereby improving efficiency. At the software level, it provides full-stack driver service capabilities, supporting key functions including video transcoding, video compression, video monitoring, hyperparameter adaptive optimization, splicing, scaling, AI bit rate control, and multi-card and multi-core scheduling.
[0075] Figure 8 In order to apply the RISC-V VPU high bit rate video transcoding effect of the embodiment of the present disclosure, Figure 8 It can be seen that for high-bitrate original videos, while ensuring basic clarity, the use of VPU hardware and video compression can reduce the video size by more than 90%, significantly reducing video storage costs and promoting the rapid implementation and application of this product in video cloud services.
[0076] Figure 9 For the low bit rate video transcoding effect of the disclosed embodiment, Figure 9 As can be seen from the figure, for low-bitrate original videos, the original video is no longer clear. By reducing the video bit rate and frame rate, VPU can still achieve a certain compression multiple, and the video size can be reduced by 50%-80%.
[0077] Table 1 shows the number of transcoding channels tested in the disclosed embodiment. It shows that when the VPU transcodes without reducing the video frame rate (25 fps), it can transcode up to 40 channels of 1080P video. When the frame rate is reduced, the overall transcoding time is significantly reduced because fewer video frames need to be encoded, allowing the transcoding of more video files in the same amount of time.
[0078] Table 1
[0079]
[0080] Figure 10a and Figure 10b To achieve the low-bitrate video transcoding effect of the disclosed embodiment, when the VPU does not reduce the video frame rate (25fps) for transcoding, the overall transcoding time does not change much because the video frames being encoded remain unchanged. When the video frame rate is reduced (10fps), the overall transcoding time is significantly reduced because the number of video frames being encoded decreases.
[0081] In some embodiments, the video processing card (TeleVPU card) adopts the form of a single-width, full-length, full-height PCIe card, which can be easily integrated into existing 2U and 4U rack-mounted servers to provide high-efficiency support for video processing. A single Xinchuang server has 8 slots, allowing up to 8 transcoding cards to be inserted at the same time, providing users with greater processing power. In terms of software form, the transcoding card provides a rich selection of software interfaces. Users can choose to use the standard FFmpeg interface or their own interface. These interfaces cover a variety of video processing functions such as encoding, decoding, scaling, splicing, video compression, AI, etc. to meet the needs of different users. This flexibility enables users to choose the most suitable software interface according to their own needs to achieve customized video processing solutions.
[0082] In the embodiments of the present disclosure, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Concepts such as "first" and "second" mentioned in this disclosure are used solely to distinguish different devices, modules, or units and are not intended to limit the order or interdependence of the functions performed by these devices, modules, or units.
[0083] In this disclosure, the term "and / or" simply describes an association relationship between related objects, indicating that three possible relationships exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the related objects are in an "or" relationship.
[0084] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such division is not mandatory.
[0085] In fact, according to the embodiment of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0086] Some of the blocks shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0087] Those skilled in the art will appreciate that all or part of the steps for implementing the above embodiments may be implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software, which may be collectively referred to herein as a "circuit," "module," or "system."
[0088] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein.
[0089] This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
Claims
1. A system architecture for video processing, characterized in that: It includes the computing power base, driver layer, service layer and application layer set up in sequence; The hardware of the computing power base includes multiple video processing cards, each video processing card includes multiple parallel video processing chips, each video processing chip includes a RISC-V computing core and a video processing unit; the multiple parallel video processing chips are all connected to a PCIe switch; The application layer is used to initiate video processing tasks; The video processing service engine in the service layer allocates the video processing task to the target video processing chip according to the task requirements of the video processing task and the load of the RISC-V computing core in each of the video processing cards; The driver layer is used to perform standardized preprocessing on the input video corresponding to the video processing task, and push the processed video to the target video processing chip through the PCIe switch. The RISC-V computing core and video processing unit of the video processing chip collaboratively process the video.
2. The method according to claim 1, characterized in that Each video processing chip also includes an AI computing unit that uses algorithms to analyze video content to achieve image quality enhancement, intelligent adjustment of encoding parameters and / or scene recognition; The processed video is pushed to the RISC-V computing core of the target video processing chip through the PCIe switch. The RISC-V computing core calls the video processing unit and the AI computing power unit to collaboratively process the video according to the task requirements of the video processing task.
3. The method according to claim 1, characterized in that The video processing service engine in the service layer distributes the video processing task to multiple target video processing chips based on the task requirements of the video processing task and the load of the RISC-V computing core in each of the video processing cards. The multiple target processing chips are used to simultaneously process multiple video streams or simultaneously process different parts of a video.
4. The method according to claim 3, characterized in that The multiple target processing chips belong to the same video processing card, or the multiple target processing chips belong to multiple video processing cards.
5. The method according to claim 1, wherein The video processing service engine in the service layer allocates the video processing task to one or more target video processing chips based on the task requirements and hardware resource status information of the video processing task. The hardware resource status information includes the load and hardware temperature of the RISC-V computing core in each of the video processing cards.
6. The method according to claim 1, characterized in that The application layer initiates multiple video processing tasks; the video processing service engine of the service layer is further used to adjust the execution order of the multiple video processing tasks in combination with task priorities and hardware resource status information.
7. The method according to claim 5 or 6, characterized in that The driver layer also collects hardware resource status information in real time and feeds the hardware resource status information back to the video processing service engine of the service layer.
8. The method according to claim 7, characterized in that Each video processing card is provided with a separate baseboard management controller, which is used to monitor the hardware status of the transcoding card in real time; the driver layer is provided with a status monitoring interface, and the driver layer obtains hardware resource status information from each baseboard management controller through the status monitoring interface.
9. The method according to claim 1, characterized in that The video processing tasks include one or more of the following tasks: Video transcoding, video compression, video monitoring, hyperparameter adaptive optimization, splicing, scaling, and AI bitrate control.
10. The method according to claim 1, characterized in that The video processing service engine of the service layer also dynamically adjusts the transcoding parameters of the input video corresponding to the video processing task for pre-processing at the driver layer in combination with the task load and hardware resource status information.
Citation Information
Patent Citations
Parallel computing accelerator and embedded system
CN112306663A
Multi-scene data processing acceleration system and method based on FPGA
CN118860653A
Multi-core heterogeneous design method and system for RISC-V chip applied to power grid
CN119885987A
Configurable functional multi-processing architecture for video processing
US20080170611A1
Ai video processing method and apparatus
WO2021139173A1
Cited By
IPC analysis and restoration system and method based on non-inductive shunting
CN121691755A