Video transcoding method and device, electronic equipment and storage medium

CN122824941APending Publication Date: 2026-09-25SHANGHAI HONGJUN RUITONG MICROELECTRONICS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611232665.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-14
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]本发明提供了一种视频转码方法、装置、电子设备、存储介质及程序产品,以解决计算资源浪费严重的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824941A_ABST
    Figure CN122824941A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of video processing, and discloses a video transcoding method and device, electronic equipment and storage medium, wherein after obtaining a decoded object and position information thereof, and a plurality of nodes respectively corresponding image regions, the target node responsible for the decoded object can be determined according to the position information and the plurality of nodes respectively corresponding image regions, and the decoded object is output to the target output buffer of the target node. The target node can read the decoded object from the target output buffer and directly perform filtering and encoding processing without reading image data across nodes. After completing the encoding operation, each node can output the code stream packet to the first node, and the first node generates the final image encoding frame according to the encoded code stream packet. The encoding and decoding operation is distributed to multiple nodes for parallel execution, which can greatly improve the utilization rate of the computing resources of the nodes and avoid wasting the computing resources of the nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing technology, and more specifically to video transcoding methods, apparatus, electronic devices, and storage media. Background Technology

[0002] In video decoding, the video transcoding process is typically forcibly bound to a single node. However, in a multi-node computing system, this forced binding of video transcoding to a particular node leads to significant waste of computing resources on other nodes. Summary of the Invention

[0003] This invention provides a video transcoding method, apparatus, electronic device, storage medium, and program product to solve the problem of serious waste of computing resources.

[0004] In a first aspect, the present invention provides a video transcoding method, which is applied to a target computing system comprising multiple nodes. The method is executed by a first node, which is any one of the multiple nodes. The method includes: Obtain the decoded object, the position information of the decoded object in the target frame image, and the image regions corresponding to multiple nodes respectively. The decoded object includes one or more pixels in the decoded target frame image, and the target frame image is any frame image in the video stream. Based on the location information and the image regions corresponding to multiple nodes, the target node corresponding to the decoded object is determined among the multiple nodes; The decoded object is output to the target output buffer in the target node, so that the target node can perform filtering and encoding operations on the decoded object to obtain the bit stream packet corresponding to the decoded object, and output it to the first node; Obtain the bitstream packets output by multiple nodes; The final image encoded frame is generated based on the bitstream packets output by multiple nodes.

[0005] The video transcoding method provided in this invention, after obtaining the decoded object and its location information, as well as the image regions corresponding to multiple nodes, can first calculate and determine the target node responsible for the decoded object based on the location information and the image regions corresponding to multiple nodes, and then output the decoded object to the target output buffer of the target node. The target node can read the decoded object from the target output buffer and directly perform filtering and encoding processing without needing to read image data across nodes. After each node completes the encoding operation, it can output the bitstream packet back to the first node, which generates the final transcoded image of the target frame image based on the encoded bitstream packet.

[0006] By distributing encoding and decoding operations across multiple nodes for parallel execution, the utilization rate of computing resources on each node in the target computing system can be greatly improved, avoiding waste of node computing resources. Furthermore, encoding operations consume a significant load during video transcoding, and the video stream contains a large number of images. In this scheme, the encoding operation for each frame no longer relies on a single node but is completed collaboratively by multiple nodes, which can greatly improve the utilization rate of node computing resources and the efficiency of video transcoding.

[0007] In one optional implementation, the location information is the target row number, and the image region is the range of row numbers; Based on location information and the image regions corresponding to multiple nodes, the target node corresponding to the decoded object is determined among the multiple nodes, including: Within the range of row numbers corresponding to multiple nodes, determine the range of target row numbers to which the target row number belongs; The node corresponding to the target row number range is identified as the target node.

[0008] Thus, since the horizontal correlation in video content is generally higher than the vertical correlation, when segmenting by line, the motion estimation operation in the encoding process can still be fully searched in the horizontal direction, minimizing the efficiency loss in motion estimation. Furthermore, calculating positional information is more convenient with line segmentation. For example, when a decoder completes the decoding of each image unit, its row number is known, eliminating the need to re-analyze and calculate the pixel positional information.

[0009] In one alternative implementation, outputting the decoded object to the target output buffer in the target node includes: Get the target pointer corresponding to the target node; Based on the target pointer, output the decoded object to the target output buffer.

[0010] In this way, by setting the target pointer for the target node, the decoded object can be accurately output to the corresponding target output buffer, avoiding the memory bandwidth occupation caused by reading data across nodes during the encoding process.

[0011] In one alternative implementation, the method further includes: Obtain the load information for each of the multiple nodes; The load information corresponding to multiple nodes is compared to obtain the comparison results; Based on the comparison results, adjust the image regions corresponding to multiple nodes respectively.

[0012] In this way, by comparing the load information between nodes, the load difference between nodes can be determined. Based on the load difference, the image area can be adjusted so that nodes with low load are responsible for encoding more pixels, and nodes with high load are responsible for encoding fewer pixels. This allows for load balancing among nodes. For encoding operations of the same frame of image, the encoding operations can be completed as simultaneously as possible, avoiding situations where one node completes the encoding operation too early and another node completes the encoding operation too late, requiring a long wait before the transcoding operation of that frame of image can be completed.

[0013] In one optional implementation, load information corresponding to multiple nodes is obtained, including: Obtain the current load information of multiple nodes during the transcoding of the target frame image; According to the preset adjustment strategy, the historical load information corresponding to multiple nodes is obtained. The historical load information is the load information of the corresponding node in the historical images before the transcoded target frame image. The load information of the second node is determined based on the current load information and historical load information corresponding to the second node, where the second node is any one of multiple nodes.

[0014] Thus, if the image features of different regions of the target frame image differ greatly and the load information changes significantly, adjusting directly based on the load information of the transcoded target frame image may result in abrupt changes, leading to inaccurate adjustment operations of the image regions and consequently, poor stability of the adjustment operations.

[0015] In an optional implementation, before obtaining the decoded object, the position information of the decoded object in the target frame image, and the image regions corresponding to the multiple nodes, the method further includes: After receiving the video transcoding start command, allocate a preset size of memory space in the target node as the target output buffer corresponding to the target node. Obtain the target number of nodes in the target computing system; Enable the encoder's segmentation mode; Set the encoder's encoding threads to the target number; Bind an encoding thread to a node so that any one of the multiple nodes can perform encoding operations only on the fragments stored in its own output buffer.

[0016] By directly binding an encoding instance to a specific node, cross-node access operations can be avoided during the encoding process.

[0017] In one optional implementation, the decoded object includes all pixels in the decoded target frame image; the method further includes: The decoded object is completely output to the decoder's total buffer so that the target device can split the decoded object according to the target number of nodes in the target computing system, and obtain image slices corresponding to multiple nodes. Each image slice is output to the output buffer of the corresponding node so that the corresponding node can perform filtering and encoding operations on the image slices in its own output buffer to obtain a bit stream packet. The target device is a memory controller or the central processing unit in any of the multiple nodes.

[0018] Since decoding operations are generally performed by the decoder, implementing the aforementioned fragmented output usually requires modifying the decoder's source code. Therefore, in cases where the decoder's source code cannot be modified, this solution can also achieve the same result by outputting the complete decoded target frame image, then segmenting and storing it, allowing each encoding thread to perform encoding operations on its own node without cross-node access.

[0019] In a second aspect, the present invention provides a video transcoding apparatus applied to a target computing system, the target computing system including multiple nodes, the apparatus being configured on a first node, the first node being any one of the multiple nodes, the apparatus comprising: The acquisition module is used to acquire the decoded object, the position information of the decoded object in the target frame image, and the image regions corresponding to multiple nodes respectively. The decoded object includes one or more pixels in the decoded target frame image, and the target frame image is any frame image in the video stream. The determination module is used to determine the target node corresponding to the decoded object among multiple nodes based on the location information and the image regions corresponding to multiple nodes respectively; The output module is used to output the decoded object to the target output buffer in the target node, so that the target node can perform filtering and encoding operations on the decoded object to obtain the bit stream packet corresponding to the decoded object, and output it to the first node; The acquisition module is also used to acquire the bitstream packets output by multiple nodes. The determination module is also used to generate the final image encoded frame based on the bitstream packets output by multiple nodes.

[0020] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the video transcoding method described in the first aspect or any corresponding embodiment thereof.

[0021] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to perform the video transcoding method described in the first aspect or any of its corresponding embodiments.

[0022] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the video transcoding method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description

[0023] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0024] Figure 1 This is a schematic diagram of an application scenario according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the first type of video transcoding method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a second process for a video transcoding method according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the third process of the video transcoding method according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the fourth process of the video transcoding method according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the fifth process of the video transcoding method according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the sixth process of the video transcoding method according to an embodiment of the present invention; Figure 8 This is a structural block diagram of a video transcoding apparatus according to an embodiment of the present invention; Figure 9 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0027] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0028] The video transcoding method provided in this invention can be applied to target computing systems, such as... Figure 1 As shown, the target computing system includes multiple nodes, which can be called Non-Uniform Memory Access (NUMA) nodes. Each node may include a central processing unit (CPU) and memory.

[0029] refer to Figure 2 The process illustrated here typically involves binding the video transcoding process to a specific NUMA node, such as node 0. Accordingly, the decoder runs on node 0, and its decoded image frames are stored in the memory of node 0. The filter runs on node 0, and its filtered image frames are stored in the memory of node 0. The encoder runs on node 0, and after encoding, it outputs the bitstream to a file or a URL.

[0030] In this process, all computation and memory access are completed within a single node. Although there is no cross-node access, it leads to a significant waste of computing resources in the target computing system. For example, in a two-node server, where each node includes 128 processing cores and the server has 256 processing cores, if only one node is performing a video transcoding task, the remaining 128 processing cores will be completely idle, resulting in an extremely low return on hardware investment.

[0031] Furthermore, video transcoding is a computationally intensive task, and a single node would severely hinder the feasibility of real-time transcoding, especially for high-resolution (e.g., 4K, 8K) video transcoding. Although multiple transcoding processes can be started when multiple processing streams exist, with each node handling a separate video transcoding task, for a single real-time stream, the resource limitations of a single node still apply.

[0032] In addition, some encoders (e.g., libx265) detect all NUMA nodes by default and create cross-node thread pools. Even if nodes and processes are bound, the encoder may still allocate auxiliary data structures on other nodes, leading to accidental remote memory accesses and affecting performance stability.

[0033] refer to Figure 3 The illustrated process, in another related technology, implements a random load balancing scheduling method. That is, the operating system schedules the decoding, filtering, and encoding threads based on the load of each node. Specifically, the decoding thread may be scheduled to node 0 or node 1, and the decoded image frames are stored in the memory of the node performing the decoding operation. The filtering thread may also be scheduled to node 0 or node 1; if it is scheduled to a different node than the decoding thread, cross-node memory read operations are required. During the encoding phase, the encoder's thread pool utilizes the processing cores of all nodes by default; that is, the encoding thread may reside on node 0 or node 1. Therefore, the bitstream packet may be output from either node.

[0034] This approach has the following drawbacks: First, transmitting image frames across nodes consumes a significant amount of memory bandwidth. For example, a 4K YUV420p image is 12MB in size. If the decoding operation is performed on node 0 and the encoding operation on node 1, then each frame requires transmitting 12MB of original image data across nodes. Here, YUV420 represents a color encoding format, where Y represents luminance, U and V represent chrominance, and p indicates that the three YUV components are stored separately. For a 30 frame per second (FPS) video, the bandwidth required is 360MB / s.

[0035] In addition, the cross-node access to reconstructed frames within the encoder can easily fill up the cross-interconnect bus, which can be a performance bottleneck caused by the Cache Coherent Interconnect for Accelerators (CCIX).

[0036] Second, during the encoder's encoding process, internal motion estimation is required. This internal motion estimation needs to reference the reconstructed frame, which is the same size as a single image frame. If the current encoding thread is on node 1, while the reconstructed frame is in the memory of node 0, then every time the reference frame is needed, a cross-node access will occur. This access operation happens frequently within the encoder, potentially causing the actual cross-node bandwidth to far exceed the original transmission volume, and developers may find it difficult to detect this bandwidth consumption.

[0037] Third, the scheduling decisions of this solution are affected by the load, and the video transcoding speed may fluctuate significantly at different times and with different input videos, making it unsuitable for scenarios requiring stable latency, such as real-time communication and live streaming. For example, if the encoding thread originally runs on node 0 and the relevant data is stored in the memory of node 0, the encoding thread may be migrated to node 1 when node 0 is under high load, resulting in a large number of cross-node accesses during the encoding process and a sharp decrease in encoding speed.

[0038] Fourth, although the encoder can associate nodes with thread pools by setting parameters (e.g., the numa pool parameter), its actual behavior depends on whether the relevant code libraries (e.g., libnuma) are linked at compile time, and the kernel version it depends on at runtime. Generally, by default, the encoder may attempt to balance threads across all nodes, resulting in extremely complex cross-node memory access patterns that are difficult for ordinary users to analyze and optimize.

[0039] To address the aforementioned technical problems, an embodiment of a video transcoding method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0040] This embodiment provides a video transcoding method, which can be executed by a first node, and more specifically, by the central processing unit (CPU) within the first node. The first node can be any one of multiple nodes included in the target computing system. Figure 4 This is a flowchart of a video transcoding method according to an embodiment of the present invention, such as... Figure 4 As shown, the process includes the following steps: Step 401: Obtain the decoded object, the position information of the decoded object in the target frame image, and the image regions corresponding to multiple nodes respectively.

[0041] The decoded object can include one or more pixels from the decoded target frame image. The target frame image can be any frame image in the video stream, which can be a live video stream, an offline video stream, a game video stream, a real-time communication video stream, etc. The video stream can conform to any video standard such as High Efficiency Video Coding (HEVC) tile, AV1 tile, or H.264 slice.

[0042] The decoded object includes multiple pixels that can be one or more of the following: a row of pixels, a column of pixels, or pixels in a segmented region of the decoded target frame image. During the calculation of position information, the first node can calculate position information for each pixel individually, or it can calculate position information for multiple pixels.

[0043] Location information may include one or more elements such as row number, column number, and cell number. A cell number is a unique number assigned to each block in the image after the image is divided into blocks, indicating the block's location in the image.

[0044] The image region is a range of numbers, and the format of the image region is the same as that of the location information. For example, when the location information is a row number, the image region is a range of row numbers; when the location information is a column number, the image region is a range of column numbers; and when the location information is a cell number, the image region is a range of cell numbers. Each image region corresponds to a segment number.

[0045] Specifically, in the video transcoding process, the decoding operation can be performed on each frame of the image according to preset units. Accordingly, after completing the decoding operation of the target frame image, the decoded target frame image can be obtained, and the pixels in the decoded target frame image are all decoded objects. The preset unit can be a macroblock, a coding tree unit (CTU), etc.

[0046] To facilitate subsequent encoding and avoid cross-node memory access, pixels at different locations can be handled by different nodes. Therefore, after obtaining a decoded object, the position information of the decoded object in the target frame image can be determined first. For example, for a certain pixel, the row number or column number of the pixel in the target frame image can be determined.

[0047] When transcoding the first frame of a video stream during decoding, the first node can divide the image region of the target frame image into multiple image regions according to the target number and a preset segmentation method (e.g., horizontal segmentation, vertical segmentation, etc.), and assign each image region to a node. Alternatively, the first node can first obtain the type of the video stream and select the corresponding segmentation method based on the type. For example, if the video stream type is a static scene (with small differences between adjacent image frames) or a vertical motion scene, vertical segmentation can be selected, and the position information can be column numbers, with the image region being a range of column numbers. If the video stream type is a dynamic scene or a horizontal motion scene, horizontal segmentation can be selected, and the position information can be row numbers, with the image region being a range of row numbers. In the processing of subsequent image frames, there is no need to re-perform the segmentation operation; the pre-configured image regions can be read directly.

[0048] Step 402: Based on the location information and the image regions corresponding to the multiple nodes, determine the target node corresponding to the decoded object among the multiple nodes.

[0049] Specifically, the first node can compare the location information of the decoded object with the image regions of each node to determine the image region that matches the location information. Then, the node corresponding to the matching image region can be determined as the target node.

[0050] Step 403: Output the decoded object to the target output buffer in the target node.

[0051] The target node performs filtering and encoding operations on the decoded objects in the target output buffer to obtain the bit stream packet corresponding to the decoded object, and outputs it to the first node.

[0052] Specifically, the first node can output the decoded object to the target output buffer of the target node. In this way, for decoded images at different positions in the target frame image, the first node can output the decoded image to the output buffer of the corresponding node.

[0053] Taking the target node as an example, the target node can read the decoded object from the target output buffer and use a filter to filter the target output buffer. After the filtering operation is completed, the encoder can continue to perform the encoding operation to obtain the final bit stream packet, which is then output to the first node. Both the filter and the encoder are software encoders, running in a multi-threaded manner on each node.

[0054] In this way, there is no data exchange between multiple nodes during the filtering and encoding process, and no large amount of data is transmitted across nodes.

[0055] Step 404: Obtain the bitstream packets output by multiple nodes.

[0056] Specifically, multiple nodes execute the above encoding operations in parallel, and each node is only responsible for encoding a portion of the pixels in the target frame image. Therefore, before performing subsequent operations, the first node needs to obtain the bitstream packets output by each node. The bitstream packets can be Network Abstraction Layer (NAL) units, slices, etc., where a slice is an independent encoded region in a frame image.

[0057] Since video transcoding involves transcoding a series of images, each node can have an encoder output queue. After generating a bitstream packet, each node can output the bitstream packet and the frame number of the target frame image to its corresponding output queue and send a merge notification to the first node. The output queue can be a lock-free queue or a message channel. The merge notification includes at least the frame number, and may also include one or more of the following: node identification information, the location information of the decoded image, and the fragment number, to indicate the image region corresponding to the bitstream. After receiving merge notifications for the same frame image from all nodes, the first node can read the bitstream packet from the output queues of each node.

[0058] Step 405: Generate the final image encoded frame based on the bitstream packets output by multiple nodes.

[0059] Specifically, after obtaining all the bitstream packets corresponding to the target frame image, the first node can write the bitstream packets output by multiple nodes to the output file sequentially according to their positions in the target frame image, or push the stream. For example, in the case of HEVC video stream format, the bitstream packets can be written to the output file sequentially from smallest to largest according to their corresponding segment numbers. Simultaneously, corresponding sequence parameter sets (VPS), image parameter sets (SPS), and picture parameter sets (PPS) can be generated, and the VPS, SPS, and PPS are added to the beginning of the bitstream packets in the output file.

[0060] In this way, cross-node transmission of the bitstream packet is only required during the merging phase. Compared to the original pre-data, the bitstream packet data size is smaller, generally tens of KB. The amount of data transmitted across nodes is small, and the impact on the total memory bandwidth is small.

[0061] The video transcoding method provided in this embodiment, after obtaining the decoded object and its location information, as well as the image regions corresponding to multiple nodes, can first calculate the target node responsible for the decoded object based on the location information and the image regions corresponding to multiple nodes, and output the decoded object to the target output buffer of the target node. The target node can read the decoded object from the target output buffer and directly perform filtering and encoding processing without reading image data across nodes. After each node completes the encoding operation, it can output the bitstream packet back to the first node. The first node generates the final transcoded image of the target frame image based on the encoded bitstream packet. In this way, the encoding and decoding operations are distributed to multiple nodes for parallel execution, which can greatly improve the utilization rate of the computing resources of each node in the target computing system and avoid wasting the computing resources of the nodes. In addition, the encoding operation occupies a large load in the video transcoding process, and the video stream contains a large number of images. In this solution, the encoding operation of each frame image no longer depends on a single node, but is completed by multiple nodes together, which can greatly improve the utilization rate of node computing resources and video transcoding efficiency.

[0062] In some optional implementations, based on any of the above embodiments, before performing the video transcoding task, the first node can perform a pre-configuration operation. Accordingly, before step 401 above, the first node can execute the following specific steps, the process of which is as follows: Figure 5 As shown: Step 501: After receiving the video transcoding start command, allocate a preset size of memory space in the target node as the target output buffer corresponding to the target node.

[0063] Step 502: Obtain the target number of nodes in the target computing system.

[0064] Step 503: Start the encoder's segmentation mode.

[0065] Step 504: Set the encoder's encoding threads to the target number.

[0066] Step 505: Bind an encoding thread to a node so that any one of the multiple nodes performs encoding operations only on the fragments stored in its own output buffer.

[0067] The preset size can be the ratio between the size of a frame image and the number of targets.

[0068] Specifically, when the transcoding task starts, the first node can call the operating system's NUMA interface (e.g., the libnuma interface on Linux) to obtain a list of available (i.e., free) nodes in the target computing system. This list can include identification information for multiple nodes, and the number of node identifications in the list is determined as the target number. Then, based on the node identification information in the list, memory allocation instructions can be used to create an independent memory pool for each node. This memory pool is physically located within the node; for example, memory allocation instructions could be `numa_alloc_onnode`, `mbind`, etc.

[0069] Furthermore, for each node, a memory resource of a preset size in the memory pool corresponding to that node can be used as an output buffer. For example, the preset size can be the size of half a frame of image. A memory resource of the size of a slice is allocated in node 0 as the output buffer of node 0, and a memory resource of the size of a slice is allocated in node 1 as the output buffer of node 1.

[0070] Simultaneously, the first node can configure the encoder's encoding threads based on the target number of nodes in the target computing system. Specifically, it can use an encoder that supports independent encoding and enable the encoder's tile / slice mode (used to instruct the encoder to perform tile-based independent encoding operations). For example, the software encoder could be x265 or SVT-AV1, and tile mode can be enabled via command-line parameters such as `tiles 2x1` or `no-deblock`. Furthermore, the first node can set the number of tile encoding threads within the encoder to be the same as the target number, for example, setting two tile encoding threads.

[0071] Finally, the first node can bind a fragmentation coding thread to the CPU of a node through the thread binding interface provided by the encoder or the operating system affinity settings. For example, fragmentation coding thread 0 can be bound to the CPU of node 0, and fragmentation coding thread 1 can be bound to the CPU of node 1.

[0072] In this way, this scheme forcibly binds each encoder instance (i.e., the encoding thread) to a unit node, and the data required for encoding is all within this node. This naturally restricts the thread and memory access within the encoder to the node, avoiding implicit cross-node access and resulting in stable and predictable performance. Specifically, during encoding, the central processing unit in each node only needs to read the slice from the output buffer of its own node to perform the corresponding encoding operations. Furthermore, during the encoding operation, there is no cross-slice reference; that is, motion estimation, intra-frame prediction, and loop filtering involved in the encoding operation are all limited to the slices it owns.

[0073] In some alternative implementations, in step 403 above, the decoded object may include a row of pixels, and correspondingly, the position information of the decoded object may be a target row number, and the image region may be a range of row numbers. Alternatively, the decoded object may be a column of pixels, and correspondingly, the position information of the decoded object may be a target column number, and the image region may be a range of column numbers. Or, the decoded object may be a unit number; for example, the target frame image may be divided into four equal parts to obtain four partitions, each partition being a unit with a unit number.

[0074] The following example, using the first scenario, further illustrates the process of determining the target node corresponding to the decoded object among multiple nodes based on location information and the image regions corresponding to each node: Step 1: Determine the target row number range to which the target row number belongs from the row number ranges corresponding to the multiple nodes.

[0075] Step 2: Identify the nodes corresponding to the target row number range as target nodes.

[0076] Specifically, the first node can compare the target row number with the target row number range of each node to determine the target row number range into which the target row number falls. Then, the node corresponding to the target row number range into which the target row number falls is determined as the target node. Other schemes are similar and will not be described in detail here.

[0077] Thus, since the horizontal correlation in video content is generally higher than the vertical correlation, when segmenting by line, the motion estimation operation in the encoding process can still be fully searched in the horizontal direction, minimizing the efficiency loss in motion estimation. Furthermore, calculating positional information is more convenient with line segmentation. For example, when a decoder completes the decoding of each image unit, its row number is known, eliminating the need to re-analyze and calculate the pixel positional information.

[0078] In some optional implementations, in addition to configuring the encoder, this solution can also modify the decoder's source code. The decoder can be an FFmpeg built-in decoder or a VLC decoder. VLC stands for Video LAN Client. Accordingly, the first node can also perform the following configuration operations on the decoder: Configure pointers, node identifiers, and image regions for each of the multiple output buffers of the decoder, so that image data in the corresponding image region can be written to the corresponding output buffer when the decoding operation is completed.

[0079] Furthermore, in step 403 above, the first node can output the decoded object to the target output buffer in the target node in the following manner: Step 1: Obtain the target pointer corresponding to the target node.

[0080] Step 2: Output the decoded object to the target output buffer according to the target pointer.

[0081] Specifically, the first node can obtain the target pointer based on the target node's identification information, and then determine the storage address of the decoded object based on the target pointer. Based on this storage address, the decoded object is written to the target output buffer in the target node's memory.

[0082] Accordingly, when the target node performs the encoding operation, it first determines the target output buffer based on the target pointer corresponding to itself. Then, it can read the decoded object that matches the image region it is responsible for from the target output buffer, and perform filtering and encoding operations on the read decoded object to obtain the bit stream packet, and output it to the target output buffer according to the target pointer.

[0083] In this way, by setting the target pointer for the target node, the decoded object can be accurately output to the corresponding target output buffer, avoiding the memory bandwidth occupation caused by reading data across nodes during the encoding process.

[0084] In some optional implementations, based on any of the above embodiments, since the subsequent merging operation requires waiting for all nodes to complete the encoding operation of the same frame, in order to enable all nodes to complete the encoding operation at the same time as much as possible and minimize synchronization waiting, the first node may also perform the following image region adjustment operation. The specific adjustment process can be found in [reference needed]. Figure 6 : Step 601: Obtain the load information corresponding to each of the multiple nodes.

[0085] Step 602: Compare the load information corresponding to multiple nodes respectively, and obtain the comparison results.

[0086] Step 603: Based on the comparison results, adjust the image regions corresponding to the multiple nodes respectively.

[0087] The load information can include encoding time, encoding queue depth, CPU utilization, etc. For example, it can be the load information of a node in the current cycle, or the load information of a node across multiple cycles. The encoding queue can be used to store unencoded pixels, i.e., decoded objects.

[0088] Specifically, the adjustment operation can be performed periodically. In each period, each node can report its own load information to the first node. The first node compares the load information of each node to obtain the load ratio between the nodes, i.e., obtains the comparison result. Then, the reciprocal of the load ratio is used to determine the image region ratio. Based on the image region ratio and the region information of the target frame image, the image region corresponding to each node is determined, and the corresponding adjustment operation is performed.

[0089] For example, the target computing system includes node 1, node 2, and node 3. Originally, the image regions of the three nodes were rows 1 to 4, rows 5 to 8, and rows 9 to 12, respectively. The ratio of the image regions of the three nodes was 3:2:1, and the ratio of the image regions of the three nodes was 1:2:3. The region information of the target frame image is 12 rows. Therefore, based on the ratio of the image regions, the 12 rows can be divided into 6 equal parts. Rows 1 to 2 are determined as the image region of node 1, rows 3 to 6 are determined as the image region of node 2, and rows 7 to 12 are determined as the image region of node 3.

[0090] For example, suppose we initially divide the workload equally (each node 50%). After a period of time, if we detect that the average encoding time of node 0, or the CPU utilization, is 20% longer than that of node 1, it indicates that node 0 is overloaded. Therefore, we can adjust the partitioning ratio of the next frame to 45% for node 0 and 55% for node 1, allowing node 1 to handle a larger portion of the load. Subsequent monitoring and fine-tuning over several frames will bring the completion times of the two nodes closer together, thereby reducing the waiting time of the merging thread and improving overall parallel efficiency.

[0091] In this way, since the video content complexity varies in different regions of different images, and each node may also perform other tasks besides video transcoding, dynamic adjustment can make the amount of encoding tasks among the nodes more balanced, that is, make the encoding processing time of each node more consistent, and minimize synchronization waiting.

[0092] In some optional implementations, there can be a variety of strategies for adjusting the size of the image region, such as sliding smoothing, exponential smoothing, and amplitude limiting adjustment (the adjustment amplitude is less than or equal to a preset amplitude threshold). The required load information is different for different strategies. For example, sliding average requires N frames of images before the target frame image, exponential smoothing requires the previous frame image before the target frame image, and amplitude limiting adjustment does not require historical images. It only needs to use the preset amplitude threshold for adjustment operation when it is determined that the adjustment amplitude is greater than the preset amplitude threshold. For example, it can be 5%.

[0093] Accordingly, in step 601 above, the load information of multiple nodes can be obtained through the following specific steps: Step 1: Obtain the current load information of multiple nodes during the transcoding of the target frame image.

[0094] Step 2: Based on the preset adjustment strategy, obtain the historical load information corresponding to multiple nodes.

[0095] Step 3: Determine the load information of the second node based on the current load information and historical load information corresponding to the second node.

[0096] The current load information can be the load information of the corresponding node during the transcoding of the target frame image. The historical load information can be the load information of the corresponding node during the transcoding of historical images before the target frame image. For example, the load information of the corresponding node when transcoding N frames of images, or the load information of the corresponding node when transcoding the previous frame image. The second node can be any one of multiple nodes.

[0097] Specifically, each node records its load information when encoding a frame of image and reports it to the master controller of the first node. During image region adjustment operations in each cycle, for each node, the first node can determine the historical load information required to execute the preset adjustment strategy from the historically reported load information. Furthermore, the load information of the node can be determined based on its current and historical load information. For example, for the moving average algorithm, when determining the load information of the second node, the first node can average the N load information points of the second node with the current load information to obtain the load information of the second node. For the exponential average algorithm, the load information of the second node can be determined using the expression (a × current load information of the second node + (a-1) × historical load information of the second node).

[0098] In this way, by limiting the amplitude, moving average, or exponentially smoothing the changes in the fragmentation ratio, and avoiding boundary abrupt changes that could cause fluctuations in coding quality or oscillations in bitrate allocation, the stability of dynamic adjustments can be guaranteed, and subjective quality will not be affected.

[0099] In some alternative implementations, during the video transcoding process described above, the decoding operation can be performed by one of the multiple nodes or by multiple nodes in parallel.

[0100] When the decoding operation is performed by a certain node (e.g., the first node), the aforementioned pre-configuration operation, bitstream packet merging operation, and image region adjustment operation can all be performed by that node.

[0101] Since decoding, configuration, merging, and adjustment operations consume fewer resources than encoding, they can be processed on a single node for easier management.

[0102] In some optional implementations, when the decoding operation is performed by a single node, before performing the decoding operation, the computational resource utilization and / or memory resource utilization of multiple nodes can be obtained based on the identification information of multiple nodes included in the available node list. Then, a node can be selected as the first node based on the computational resource utilization and / or memory resource utilization. For example, the node with the lowest computational resource utilization can be selected as the first node, or the node with the lowest memory resource utilization can be selected as the first node. Alternatively, for each node, the weighted sum of the computational resource utilization and memory resource utilization can be determined first, and then the node with the smallest sum can be selected as the first node.

[0103] This ensures load balancing across nodes, preventing the first node from bearing too many tasks, which would lead to slow encoding operations and low efficiency in transcoding the target frame image.

[0104] In some optional implementations, the decoded object may include all pixels in the decoded target frame image. Accordingly, after completing the decoding operation, the first node outputs the decoded object completely to the decoder's total buffer. This allows the target device to split the decoded object according to the target number of nodes in the target computing system, obtaining image slices corresponding to multiple nodes. Each image slice is then output to the output buffer of its corresponding node, allowing the corresponding node to perform filtering and encoding operations on the image slices in its own output buffer to obtain a bitstream packet. The target device is either a memory controller or the central processing unit (CPU) of any of the multiple nodes. The memory controller may be a Direct Memory Access (DMA) controller.

[0105] Specifically, after the first node completes the decoding operation, or after multiple nodes complete the decoding operation, they can directly output the decoded object to the total buffer and send a fragmentation notification to the target device.

[0106] Upon receiving the fragmentation notification, the target device can read the decoded target frame image from the total buffer, split the decoded image frame image into multiple image fragments according to the target quantity, and determine the image regions corresponding to each image fragment. For each image fragment, a matching operation is performed based on the image region corresponding to the image fragment and the image regions corresponding to multiple nodes to determine the node that matches the image fragment, and obtain the pointer corresponding to the node. Then, based on the pointer, the image fragment can be output to the output buffer of the corresponding node, where the corresponding node performs filtering and encoding operations to obtain the bitstream packet, which is then output to the first node.

[0107] Therefore, this method does not require modification of the decoder source code, meaning that this solution is suitable for decoder scenarios where the source code cannot be modified, making it quite convenient.

[0108] In some alternative implementations, the encoder can also be a hardware encoder, such as an NVENC or QSV encoder, which can provide a slice-mode application programming interface (API). Accordingly, the number of hardware encoders can be the same as the number of nodes in the target computing system, and one hardware encoder can be physically connected to one node. After completing the decoding operation, any node can output the decoded object to the corresponding output buffer according to the aforementioned processing method. Each node then filters the decoded image and re-outputs it to the corresponding output buffer, sending a completion notification to the hardware encoder corresponding to that node. Upon receiving the completion notification from any node, the hardware encoder connected to the memory of the output buffer can directly read the decoded image and encode it. After completing the encoding process, the hardware encoder can write the obtained bitstream packet back to the output buffer and notify the first node so that the first node can merge the bitstream packets.

[0109] In this way, having multiple hardware encoders perform the encoding operations in parallel can improve encoding efficiency and reduce the computational load on each node.

[0110] The following is a detailed explanation of the process of the video transcoding method described above, using a specific example. Figure 7 As shown.

[0111] The first node mentioned above can be node 0. Node 0 performs configuration operations on the encoder, decoder, and filters, and monitors its own load information. Node 1 can monitor its own load information and send it to node 0 periodically. The decoder and stream combiner can run on node 0. Node 0 and node 1 can run the encoder and filter in parallel.

[0112] Node 0 performs image decoding, and then segments the decoded image and outputs the segments to different nodes.

[0113] When the decoder outputs image slices, it writes image slice 0 into the memory of node 0 and image slice 1 into the memory of node 1. Then, node 0 can perform filtering and encoding operations on image slice 0 to obtain the corresponding bitstream packet. Node 1 can perform filtering and encoding operations on image slice 1 to obtain the corresponding bitstream packet and send it to node 0.

[0114] After completing the filtering and encoding operations for image segment 0, and after receiving the bitstream packet sent by node 1, node 0 performs a merging operation to complete the output bitstream.

[0115] In this way, if both nodes include 128 processing cores, the two nodes can process half a frame of image at the same time, halving the computational load, which is suitable for real-time transcoding scenarios with high resolution or high frame rate.

[0116] This embodiment also provides a video transcoding apparatus for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0117] This embodiment provides a video transcoding device, such as... Figure 8 As shown, it includes: The acquisition module 810 is used to acquire the decoded object, the position information of the decoded object in the target frame image, and the image regions corresponding to multiple nodes respectively. The decoded object includes one or more pixels in the decoded target frame image, and the target frame image is any frame image in the video stream. The determination module 820 is used to determine the target node corresponding to the decoded object among multiple nodes based on the location information and the image regions corresponding to multiple nodes respectively; The output module 830 is used to output the decoded object to the target output buffer in the target node, so that the target node can perform filtering and encoding operations on the decoded object to obtain the bit stream packet corresponding to the decoded object and output it to the first node; The acquisition module 810 is also used to acquire the bit stream packets output by multiple nodes respectively; The determination module 820 is also used to generate the final image encoded frame based on the bit stream packets output by multiple nodes.

[0118] In some optional implementations, the location information is the target row number, and the image region is the range of row numbers; Module 820 is specifically used for: Within the range of row numbers corresponding to multiple nodes, determine the range of target row numbers to which the target row number belongs; The node corresponding to the target row number range is identified as the target node.

[0119] In some alternative implementations, the output module 830 is specifically used for: Get the target pointer corresponding to the target node; Based on the target pointer, output the decoded object to the target output buffer.

[0120] In some alternative embodiments, the device further includes an adjustment module 840 for: Obtain the load information for each of the multiple nodes; The load information corresponding to multiple nodes is compared to obtain the comparison results; Based on the comparison results, adjust the image regions corresponding to multiple nodes respectively.

[0121] In some alternative implementations, the adjustment module 840 is specifically used for: Obtain the current load information of multiple nodes during the transcoding of the target frame image; According to the preset adjustment strategy, the historical load information corresponding to multiple nodes is obtained. The historical load information is the load information of the corresponding node in the historical images before the transcoded target frame image. The load information of the second node is determined based on the current load information and historical load information corresponding to the second node, where the second node is any one of multiple nodes.

[0122] In some alternative implementations, the device further includes a configuration module 850 for: After receiving the video transcoding start command, allocate a preset size of memory space in the target node as the target output buffer corresponding to the target node. Obtain the target number of nodes in the target computing system; Enable the encoder's segmentation mode; Set the encoder's encoding threads to the target number; Bind an encoding thread to a node so that any one of the multiple nodes can perform encoding operations only on the fragments stored in its own output buffer.

[0123] In some optional implementations, the decoded object includes all pixels in the decoded target frame image; the output module 830 is further configured to: The decoded object is completely output to the decoder's total buffer so that the target device can split the decoded object according to the target number of nodes in the target computing system, and obtain image slices corresponding to multiple nodes. Each image slice is output to the output buffer of the corresponding node so that the corresponding node can perform filtering and encoding operations on the image slices in its own output buffer to obtain a bit stream packet. The target device is a memory controller or the central processing unit in any of the multiple nodes.

[0124] The video transcoding apparatus provided in this embodiment of the invention can execute the video transcoding method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.

[0125] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0126] The following is a detailed reference. Figure 9 This diagram illustrates a suitable structural schematic for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 901, which can perform various appropriate actions and processes according to a program stored in ROM 902 or a program loaded from memory 908 into RAM 903. RAM 903 also stores various programs and data required for the operation of the electronic device. The processor 901, ROM 902, and RAM 903 are interconnected via bus 904. An input / output (I / O) interface 905 is also connected to bus 904; wherein ROM is a read-only memory and RAM is a random access memory.

[0127] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays, speakers, vibrators, etc.; memory devices 908 including, for example, magnetic tape, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0128] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a memory 908, or installed from a ROM 902. When the computer program is executed by the processor 901, it performs the functions defined in the video transcoding method of the embodiments of the present invention.

[0129] Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0130] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the video transcoding method shown in the above embodiments is implemented.

[0131] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0132] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A video transcoding method, characterized in that, The method is applied to a target computing system, which includes multiple nodes. The method is executed by a first node, which is any one of the multiple nodes. The method includes: Obtain the decoded object, the position information of the decoded object in the target frame image, and the image regions corresponding to the multiple nodes respectively, wherein the decoded object includes one or more pixels in the decoded target frame image, and the target frame image is any frame image in the video stream; Based on the location information and the image regions corresponding to the plurality of nodes respectively, the target node corresponding to the decoded object is determined among the plurality of nodes; The decoded object is output to the target output buffer in the target node, so that the target node can perform filtering and encoding operations on the decoded object to obtain the bit stream packet corresponding to the decoded object, and output it to the first node; Obtain the bitstream packets output by the multiple nodes respectively; The final image encoded frame is generated based on the bitstream packets output by the multiple nodes.

2. The video transcoding method according to claim 1, characterized in that, The location information is the target row number, and the image region is the range of row numbers; The step of determining the target node corresponding to the decoded object among the plurality of nodes based on the location information and the image regions corresponding to the plurality of nodes respectively includes: Among the row number ranges corresponding to the multiple nodes, the target row number range to which the target row number belongs is determined; The node corresponding to the target row number range is determined as the target node.

3. The video transcoding method according to claim 1 or 2, characterized in that, The step of outputting the decoded object to the target output buffer in the target node includes: Obtain the target pointer corresponding to the target node; The decoded object is output to the target output buffer according to the target pointer.

4. The video transcoding method according to claim 1 or 2, characterized in that, The method further includes: Obtain the load information corresponding to each of the multiple nodes; The load information corresponding to each of the multiple nodes is compared to obtain the comparison results; Based on the comparison results, the image regions corresponding to the multiple nodes are adjusted respectively.

5. The video transcoding method according to claim 4, characterized in that, The step of obtaining the load information corresponding to each of the multiple nodes includes: Obtain the current load information of each of the multiple nodes during the transcoding of the target frame image; According to a preset adjustment strategy, historical load information corresponding to multiple nodes is obtained, wherein the historical load information is the load information of the corresponding node in the process of transcoding the target frame image in the historical image. The load information of the second node is determined based on the current load information and historical load information corresponding to the second node, wherein the second node is any one of the plurality of nodes.

6. The video transcoding method according to claim 1 or 2, characterized in that, Before obtaining the decoded object, the position information of the decoded object in the target frame image, and the image regions corresponding to the multiple nodes, the method further includes: After receiving the video transcoding start command, the memory space of the target node is allocated to a preset size as the target output buffer corresponding to the target node. Obtain the target number of nodes in the target computing system; Enable the encoder's segmentation mode; Set the number of encoding threads of the encoder to the target number; Bind an encoding thread to a node so that any one of the multiple nodes performs encoding operations only on the fragments stored in its own output buffer.

7. The video transcoding method according to claim 1 or 2, characterized in that, The decoded object includes all pixels in the decoded target frame image; the method further includes: The decoded object is completely output to the decoder's total buffer so that the target device can split the decoded object according to the target number of nodes in the target computing system to obtain image slices corresponding to the multiple nodes respectively. Each image slice is output to the output buffer of the corresponding node so that the corresponding node can perform filtering and encoding operations on the image slices in its own output buffer to obtain a bitstream packet. The target device is a memory controller or the central processing unit of any one of the multiple nodes.

8. A video decoding device, characterized in that, The device is applied to a target computing system, the target computing system including multiple nodes, the device being configured on a first node, the first node being any one of the multiple nodes, the device comprising: The acquisition module is used to acquire a decoded object, the position information of the decoded object in the target frame image, and the image regions corresponding to the multiple nodes respectively, wherein the decoded object includes one or more pixels in the decoded target frame image, and the target frame image is any frame image in the video stream; The determining module is used to determine the target node corresponding to the decoded object among the multiple nodes based on the location information and the image regions corresponding to the multiple nodes respectively; The output module is used to output the decoded object to the target output buffer in the target node, so that the target node can perform filtering and encoding operations on the decoded object to obtain the bit stream packet corresponding to the decoded object, and output it to the first node; The acquisition module is also used to acquire the code stream packets output by the multiple nodes respectively; The determining module is further configured to generate a final image encoded frame based on the bitstream packets output by the multiple nodes respectively.

9. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the video transcoding method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the video transcoding method according to any one of claims 1 to 7.