Video transcoding method, device and equipment and computer readable storage medium

By determining and adjusting the encoding information of the region of interest in the video frame during the video transcoding process, and multiplexing the encoding information of the non-interested area, the problems of large amount of video transcoding and high cost in the prior art are solved, and the transcoding calculation amount and cost are reduced while ensuring the video quality.

CN120017881APending Publication Date: 2025-05-16BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311532666.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

While ensuring the quality of video, the existing video transcoding method has a large amount of calculation, resulting in an increase in CDN cost.

Method used

By decoding the original video stream, the original resolution data and low resolution data are obtained, the region of interest is determined for each frame, the encoding information of the region of interest is adjusted, and the encoding information of the non-area of ​​interest is multiplexed to reduce the amount of calculation during the transcoding process.

Benefits of technology

Under the condition of the overall code rate reduction, the video quality is kept from decreasing, significantly reducing the amount of calculation during the transcoding process, thereby reducing CDN costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017881A_ABST
    Figure CN120017881A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video transcoding method, device and equipment and a computer readable storage medium. The method comprises the following steps: acquiring an original video stream; decoding the original video stream to obtain original resolution data, low resolution data and initial coding information corresponding to the original video stream; obtaining a region of interest of each frame in the original resolution data; adjusting initial coding information corresponding to a region of interest of each frame in the original resolution data and retaining initial coding information corresponding to a non-region of interest of each frame to obtain final coding information of each frame; and coding each frame in the original resolution data by using the final coding information of each frame. In this way, under the condition that the overall code rate is reduced, the video quality is not obviously reduced, and the calculated amount in the transcoding process is reduced to reduce the CDN cost while the video quality is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video, and in particular to the field of video transcoding technology. Background Art

[0002] With the popularity of short videos, the proportion of video data on the Internet is getting higher and higher. Statistics show that the proportion of video data on the Internet has reached more than 80%. Especially now with the advent of 5G, users have higher requirements for image quality, and there are more and more 4k, 8k, HDR and other videos. As a result, the video quality has increased exponentially, which has led to the CDN (Content Delivery Network) cost (such as transcoding cost) of the video operation platform becoming higher and higher. For example: the existing video transcoding method mainly uses more advanced encoders such as HEVC, AV1, etc. for transcoding, which is accompanied by a significant increase in computing power.

[0003] Therefore, how to reduce the amount of computation in the transcoding process (where the essence of transcoding is decoding before encoding) while ensuring video quality in order to reduce CDN costs has become an urgent problem to be solved. Summary of the invention

[0004] The present disclosure provides a video transcoding method, apparatus, device, storage medium and vehicle.

[0005] According to a first aspect of the present disclosure, a video transcoding method is provided. The method comprises:

[0006] Get the original video stream;

[0007] Decoding the original video stream to obtain original resolution data, low resolution data, and initial coding information corresponding to the original video stream, wherein the resolution of the low resolution data is lower than the resolution of the original resolution data;

[0008] Acquire a region of interest of each frame in the original resolution data, wherein the region of interest of each frame is determined based on the low resolution data;

[0009] Adjusting the initial coding information corresponding to the region of interest of each frame in the original resolution data, and retaining the initial coding information corresponding to the region of non-interest of each frame, so as to obtain the final coding information of each frame;

[0010] Each frame in the original resolution data is encoded using the final encoding information of each frame to obtain a transcoded video stream.

[0011] According to the above aspect and any possible implementation, an implementation is further provided, wherein obtaining the region of interest of each frame in the original resolution data includes:

[0012] Determining a current resolution and a current scene of each frame in the low-resolution data;

[0013] Get the preset correspondence between resolution, scene and region of interest;

[0014] Matching the current resolution and the current scene with the preset corresponding relationship to determine a region of interest for each frame in the low-resolution data;

[0015] The region of interest of each frame in the low-resolution data is scaled according to a resolution ratio relationship between the low-resolution data and the original-resolution data to determine the region of interest of each frame in the original-resolution data.

[0016] According to the above aspects and any possible implementation manner, an implementation manner is further provided, wherein decoding the original video stream to obtain original resolution data, low resolution data, and initial coding information corresponding to the original video stream includes:

[0017] Acquire a preset resolution corresponding to the low-resolution data;

[0018] Based on the preset resolution, the original video stream is decoded to obtain original resolution data and initial coding information corresponding to the original video stream and the low-resolution data, wherein the initial coding information includes: an initial macroblock corresponding to the original resolution data, a motion vector of the initial macroblock, and a type of slice to which the initial macroblock belongs.

[0019] According to the above aspects and any possible implementation manner, an implementation manner is further provided, wherein adjusting the initial coding information corresponding to the region of interest of each frame in the original resolution data comprises:

[0020] Determine at least one macroblock division method corresponding to a preset coding standard;

[0021] Calculating the distortion amount corresponding to each macroblock division method in the at least one macroblock division method;

[0022] According to the magnitude of the distortion amount corresponding to each macroblock division method, the initial coding information corresponding to the region of interest of each frame in the original resolution data is adjusted.

[0023] According to the above aspects and any possible implementation manner, an implementation manner is further provided, wherein adjusting the initial coding information corresponding to the region of interest of each frame in the original resolution data according to the size of the distortion amount corresponding to each macroblock division manner includes:

[0024] Selecting a macroblock partitioning method with the smallest distortion from the at least one macroblock partitioning method;

[0025] Determining the macroblock division mode with the smallest distortion as the optimal macroblock division mode;

[0026] At least one new macroblock corresponding to the region of interest of each frame in the macroblock optimal partitioning mode and the type of slice to which each of the at least one new macroblock belongs are determined, and a motion vector of each of the at least one new macroblock is calculated.

[0027] According to the above aspect and any possible implementation manner, an implementation manner is further provided, wherein obtaining the final encoding information of each frame includes:

[0028] Multiplexing and storing the initial macroblocks of the non-interested area of ​​each frame in the original resolution data, the motion vectors of the initial macroblocks of the non-interested area of ​​each frame, and the types of slices to which the initial macroblocks of the non-interested area of ​​each frame belong;

[0029] At least one new macroblock corresponding to the area of ​​interest of each frame in the original resolution data, the type of slice to which the at least one new macroblock belongs, the motion vector of the at least one new macroblock, the initial macroblock of the non-interest area of ​​each frame, the motion vector of the initial macroblock of the non-interest area of ​​each frame, and the type of slice to which the initial macroblock of the non-interest area of ​​each frame belongs are determined as the final encoding information of each frame.

[0030] According to the above aspects and any possible implementation manner, an implementation manner is further provided, wherein encoding each frame in the original resolution data using the final encoding information of each frame to obtain a transcoded video stream includes:

[0031] reducing the quantization parameter of the region of interest of each frame in the original resolution data and increasing the quantization parameter of the region of non-interest of each frame in the original resolution data;

[0032] Each frame in the original resolution data after the quantization parameter is adjusted is encoded using the final encoding information of each frame to obtain a transcoded video stream.

[0033] According to a second aspect of the present disclosure, a video transcoding device is provided. The device comprises:

[0034] A first acquisition module, used for acquiring an original video stream;

[0035] A decoding module, used for decoding the original video stream to obtain original resolution data, low resolution data and initial coding information corresponding to the original video stream, wherein the resolution of the low resolution data is lower than the resolution of the original resolution data;

[0036] A second acquisition module, used to acquire a region of interest of each frame in the original resolution data, wherein the region of interest of each frame is determined based on the low resolution data;

[0037] A processing module, used for adjusting the initial coding information corresponding to the region of interest of each frame in the original resolution data, and retaining the initial coding information corresponding to the non-region of interest of each frame, so as to obtain the final coding information of each frame;

[0038] The encoding module is used to encode each frame in the original resolution data using the final encoding information of each frame to obtain a transcoded video stream.

[0039] According to a third aspect of the present disclosure, an electronic device is provided, which includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the program, the method described above is implemented.

[0040] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method according to the first aspect of the present disclosure is implemented.

[0041] According to a fifth aspect of the present disclosure, a vehicle is provided, the vehicle comprising the video transcoding device as described in the second aspect and / or the electronic device as described in the third aspect.

[0042] In the present disclosure, after obtaining the original video stream, the original resolution data, low-resolution data and initial coding information corresponding to the original video stream can be obtained by decoding the original video stream, and then the region of interest of each frame in the original resolution data is obtained, and then the initial coding information corresponding to the region of interest of each frame in the original resolution data is adjusted and the initial coding information corresponding to the non-interest region of each frame is retained, so as to obtain the final coding information of each frame, so as to encode each frame in the original resolution data using the final coding information of each frame to obtain a transcoded video stream. In this way, during the transcoding process, only the coding information of the region of interest needs to be adjusted, and the coding information of the non-interest region can be reused and saved, which obviously reduces the computational complexity of the transcoding process, and the human eye generally pays more attention to the region of interest (such as the face) and pays less attention to the non-interest region (such as the background). Therefore, under the condition of reducing the overall bit rate, the quality of the video is not significantly reduced, which also achieves the goal of reducing the computational complexity in the transcoding process while ensuring the video quality to reduce the CDN cost.

[0043] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:

[0045] Figure 1 A flowchart of a video transcoding method according to an embodiment of the present disclosure is shown;

[0046] Figure 2 A flowchart of another video transcoding method according to an embodiment of the present disclosure is shown;

[0047] Figure 3 A flowchart of another video transcoding method according to an embodiment of the present disclosure is shown;

[0048] Figure 4 A block diagram of a video transcoding device according to an embodiment of the present disclosure is shown;

[0049] Figure 5 A block diagram of an exemplary electronic device capable of implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0051] In addition, the term "and / or" in this article is only a description of the association relationship between the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0052] Figure 1 The flowchart of the video transcoding method 100 according to the embodiment of the present disclosure is shown. The execution subject of the method 100 may be a video server, and the function of the video server is to transcode, that is, decode and then re-encode the encoded original video stream sent by the video sending end, and then send the obtained transcoded video stream to the video receiving end. The method 100 may include:

[0053] Step 110, obtaining the original video stream;

[0054] Step 120, decoding the original video stream to obtain original resolution data, low resolution data, and initial coding information corresponding to the original video stream, wherein the resolution of the low resolution data is lower than the resolution of the original resolution data;

[0055] The original resolution data may be 1280*1280, and the low resolution data may be 640*780 or 640*640.

[0056] The initial coding information includes, but is not limited to: an initial macroblock corresponding to the original resolution data, a motion vector of the initial macroblock, and a type of slice to which the initial macroblock belongs.

[0057] In video coding, a coded image is usually divided into several macroblocks. A macroblock consists of a luminance pixel block and two additional chrominance pixel blocks. Generally speaking, the luminance block is a 16x16 pixel block, and the size of the two chrominance image pixel blocks depends on the sampling format of the image. For example, for a YUV420 sampled image, the chrominance block is an 8x8 pixel block. In each image, several macroblocks are arranged in the form of slices. The video coding algorithm encodes each macroblock one by one in units of macroblocks and organizes them into a continuous video code stream.

[0058] There may be multiple macroblocks in the original resolution data. The specific number depends on the macroblock division method. No matter how many macroblocks there are in the original resolution data, they are collectively referred to as initial macroblocks. The macroblock division method can be 16*16 pixel, 16*8 pixel macroblock division, etc.

[0059] Slice: A frame of video image can be encoded into one or more slices, each slice contains an integer number of macroblocks, that is, each slice contains at least one macroblock and at most contains the macroblocks of the entire image.

[0060] The purpose of slices: To limit the spread and transmission of bit errors, the coded slices are kept independent of each other. The slice types (i.e., Slice_type) include I slices (containing only I macroblocks), P slices (containing P and I macroblocks), B slices (containing B and I macroblocks), SP slices (used for switching between different coded streams), and SI slices (special types of coded macroblocks)

[0061] That is, the types of slices to which the initial macroblock belongs include I slice, P slice, B slice, SP slice and SI slice.

[0062] The motion vector MV of a macroblock usually adopts a block matching algorithm, whose main idea is to divide a frame of image into NxN macroblocks, and then each block searches for the best matching macroblock within the search range of the previous frame according to a certain matching criterion. The obtained displacement difference is called a motion vector; or the x-axis and y-axis displacement of the macroblock in the previous and next frames of image is called a motion vector.

[0063] Step 130, obtaining a region of interest of each frame in the original resolution data, wherein the region of interest of each frame is determined based on the low resolution data;

[0064] Region of interest (ROI) is a term used to describe the region of interest in machine vision and image processing. In machine vision and image processing, the region to be processed is outlined in the form of a box, circle, ellipse, irregular polygon, etc., which is called a region of interest.

[0065] Step 140, adjusting the initial coding information corresponding to the region of interest of each frame in the original resolution data, and retaining the initial coding information corresponding to the non-region of interest of each frame, so as to obtain the final coding information of each frame;

[0066] The non-interest region of each frame is other regions in each frame except the region of interest. For example, if a frame is a person image, the region of interest is the face, and the non-interest region is other regions in the person image except the face.

[0067] Step 150: Encode each frame in the original resolution data using the final encoding information of each frame to obtain a transcoded video stream.

[0068] After obtaining the original video stream, the original resolution data, low-resolution data and initial coding information corresponding to the original video stream can be obtained by decoding the original video stream, and then the region of interest of each frame in the original resolution data is obtained, and then the initial coding information corresponding to the region of interest of each frame in the original resolution data is adjusted and the initial coding information corresponding to the non-interest region of each frame is retained, so as to obtain the final coding information of each frame, so as to encode each frame in the original resolution data using the final coding information of each frame to obtain a transcoded video stream. In this way, during the transcoding process, only the coding information of the region of interest needs to be adjusted, and the coding information of the non-interest region can be reused and saved, which obviously reduces the computational complexity of the transcoding process, and the human eye generally pays more attention to the region of interest (such as the face) and pays less attention to the non-interest region (such as the background). Therefore, under the condition of reducing the overall bit rate, the quality of the video is not significantly reduced, which also achieves the goal of reducing the computational complexity in the transcoding process while ensuring the video quality to reduce the CDN cost.

[0069] In some embodiments, obtaining the region of interest of each frame in the original resolution data includes:

[0070] Determining a current resolution and a current scene of each frame in the low-resolution data;

[0071] Get the preset correspondence between resolution, scene and region of interest;

[0072] For example, for the original video stream of a live broadcast scene, the region of interest is usually the face area or the cargo area;

[0073] For a video call scenario, the region of interest is usually a human face region, while the region of interest for some other scenarios may be an animal region or other object region.

[0074] Matching the current resolution and the current scene with the preset corresponding relationship to determine a region of interest for each frame in the low-resolution data;

[0075] Due to differences in resolution and scene, the region of interest will be different. For example, a frame with a large resolution usually has a large region of interest. Therefore, by matching the current resolution and the current scene with the preset correspondence, the region of interest of each frame in the low-resolution data can be determined, that is, the position of the region of interest of each frame can be determined.

[0076] Preferably, the region of interest may be a rectangular region, and the region of interest may be characterized by the coordinate position of the upper left corner of the rectangular region and the length and width of the rectangular region or by the coordinate positions of the four corners of the rectangular region.

[0077] The region of interest of each frame in the low-resolution data is scaled according to a resolution ratio relationship between the low-resolution data and the original-resolution data to determine the region of interest of each frame in the original-resolution data.

[0078] For example: the size of each frame in the low-resolution data is 640*640, and the size of each frame in the original resolution data is 1280*1280, then the resolution ratio between the low-resolution data and the original resolution data is 1:4. In this way, after determining the region of interest of each frame in the low-resolution data, the coordinates of the region of interest of each frame in the low-resolution data can be multiplied by 4 (i.e., enlarged 4 times) to obtain the coordinates of the region of interest of each frame in the original resolution data, and the box formed by the coordinates of the region of interest of each frame in the original resolution data is the region of interest of each frame in the original resolution data.

[0079] After obtaining the preset correspondence between the resolution, scene and region of interest, the current resolution and the current scene can be matched with the preset correspondence to determine the region of interest of each frame in the low-resolution data, and then the region of interest of each frame in the low-resolution data can be scaled according to the resolution ratio between the low-resolution data and the original resolution data, so as to accurately determine the region of interest of each frame in the original resolution data.

[0080] In some embodiments, decoding the original video stream to obtain original resolution data, low resolution data, and initial coding information corresponding to the original video stream includes:

[0081] Acquire a preset resolution corresponding to the low-resolution data;

[0082] Based on the preset resolution, the original video stream is decoded to obtain original resolution data and initial coding information corresponding to the original video stream and the low-resolution data, wherein the initial coding information includes: an initial macroblock corresponding to the original resolution data, a motion vector of the initial macroblock, and a type of slice to which the initial macroblock belongs.

[0083] By obtaining the preset resolution corresponding to the low-resolution data, the original video stream can be decoded based on the preset resolution to accurately obtain the original resolution data and the initial encoding information corresponding to the original video stream and the low-resolution data.

[0084] In some embodiments, adjusting the initial coding information corresponding to the region of interest of each frame in the original resolution data includes:

[0085] Determine at least one macroblock division method corresponding to a preset coding standard;

[0086] The preset encoding standards include but are not limited to H.264, H.265 and H.266.

[0087] The macroblock division method is explained as follows:

[0088] I macroblock supports 16x16, 4 8x8 blocks, and 16 4x4 blocks.

[0089] P macroblocks support 16x16, 2 16x8 blocks, 2 8x16 blocks, and 4 8x8 blocks (8x8 blocks need to be divided again);

[0090] B macroblock supports 16x16, 2 16x8 blocks, 2 8x16 blocks, and 16 8x8 blocks (8x8 blocks need to be divided again).

[0091] Calculating the distortion amount corresponding to each macroblock division method in the at least one macroblock division method;

[0092] The distortion amount is the distortion amount calculated based on rate-distortion optimization (RDO).

[0093] According to the magnitude of the distortion amount corresponding to each macroblock division method, the initial coding information corresponding to the region of interest of each frame in the original resolution data is adjusted.

[0094] After calculating the distortion amount corresponding to each macroblock division method in the at least one macroblock division method, the initial encoding information corresponding to the region of interest of each frame in the original resolution data can be reasonably adjusted according to the size of the distortion amount corresponding to each macroblock division method, so as to obtain accurate encoding information of the region of interest of each frame, thereby ensuring that the quality of the transcoded video stream will not be reduced.

[0095] In some embodiments, adjusting the initial coding information corresponding to the region of interest of each frame in the original resolution data according to the magnitude of the distortion amount corresponding to each macroblock division method includes:

[0096] Selecting a macroblock partitioning method with the smallest distortion from the at least one macroblock partitioning method;

[0097] Determining the macroblock division mode with the smallest distortion as the optimal macroblock division mode;

[0098] At least one new macroblock corresponding to the region of interest of each frame in the macroblock optimal partitioning mode and the type of slice to which each of the at least one new macroblock belongs are determined, and a motion vector of each of the at least one new macroblock is calculated.

[0099] By selecting a macroblock division method with the smallest distortion from the at least one macroblock division method, and determining the macroblock division method with the smallest distortion as the optimal macroblock division mode, at least one new macroblock corresponding to the area of ​​interest of each frame and the type of slice to which the at least one new macroblock belongs can be determined according to the optimal macroblock division mode, and the motion vector of the at least one new macroblock is calculated, thereby obtaining new encoding information of the macroblocks in the area of ​​interest of each frame in the original resolution data.

[0100] In some embodiments, obtaining the final encoding information of each frame includes:

[0101] Multiplexing and storing the initial macroblocks of the non-interested area of ​​each frame in the original resolution data, the motion vectors of the initial macroblocks of the non-interested area of ​​each frame, and the types of slices to which the initial macroblocks of the non-interested area of ​​each frame belong;

[0102] At least one new macroblock corresponding to the area of ​​interest of each frame in the original resolution data, the type of slice to which the at least one new macroblock belongs, the motion vector of the at least one new macroblock, the initial macroblock of the non-interest area of ​​each frame, the motion vector of the initial macroblock of the non-interest area of ​​each frame, and the type of slice to which the initial macroblock of the non-interest area of ​​each frame belongs are determined as the final encoding information of each frame.

[0103] By further optimizing the original coding information for the region of interest and directly reusing the coding information for the region of no interest, the final coding information can be obtained, which can not only reduce the amount of transcoding calculations but also ensure that the overall visual subjective quality is not reduced.

[0104] In some embodiments, encoding each frame in the original resolution data using the final encoding information of each frame to obtain a transcoded video stream includes:

[0105] reducing the quantization parameter of the region of interest of each frame in the original resolution data and increasing the quantization parameter of the region of non-interest of each frame in the original resolution data;

[0106] The quantization parameter is the QP parameter, which is mainly used to adjust the details of the image and ultimately adjust the picture quality. The QP value is inversely proportional to the bit rate. The smaller the QP value and the higher the bit rate, the higher the picture quality; conversely, the larger the QP value and the higher the bit rate, the lower the picture quality.

[0107] There can be a target correspondence between the quantization parameter and the bit rate, so that the first bit rate and the second bit rate required by the region of interest and the non-region of interest of each frame in the original resolution data can be obtained respectively, and then the first bit rate is matched with the target correspondence to obtain the first target quantization parameter required by the region of interest of each frame in the original resolution data, and then when it is reduced, the quantization parameter of the region of interest of each frame in the original resolution data is reduced to the first target quantization parameter; similarly, the second bit rate is matched with the target correspondence to obtain the second target quantization parameter required by the non-region of interest of each frame in the original resolution data, and then when it is increased, the quantization parameter of the region of interest of each frame in the original resolution data is increased to the second target quantization parameter.

[0108] Each frame in the original resolution data after the quantization parameter is adjusted is encoded using the final encoding information of each frame to obtain a transcoded video stream.

[0109] By reducing the quantization parameter of the region of interest of each frame in the original resolution data and increasing the quantization parameter of the non-interest region of each frame in the original resolution data, the bit rate of the original resolution data is reduced as a whole, which can reduce the bandwidth cost in the CDN cost. Then, the final encoding information of each frame is used to encode each frame in the original resolution data after the quantization parameter is adjusted, which can achieve the guarantee of the video quality of the transcoded video stream on the basis of reducing the overall bit rate and reducing the transcoding workload.

[0110] The following will be combined Figure 2 The technical solution of the present disclosure is further described as follows:

[0111] like Figure 2 As shown, the original video stream is sent to the VPU decoder for decoding. The decoder will output two videos, one with original resolution data and one with low resolution video, i.e., low resolution data (such as 640x480). The low resolution data is sent to the NPU (a device independent of the transcoder, i.e., the network processor) to obtain ROI information (i.e., initial encoding information). The ROI information is usually rectangular frame information. The ROI information is then sent to the VPU encoder together with the original resolution data. At the same time, the encoding information of each frame will be saved during video decoding, mainly block division information, MV, slice type, etc. During encoding, the saved encoding information is also sent to the encoder. The encoded video is then output, which is the final transcoded video stream. Among them, the VPU decoder and the VPU encoder are transcoders, and the transcoder is a separate device.

[0112] The following will be combined Figure 3 The technical solution of the present disclosure is further described as follows:

[0113] S1. Input original video stream.

[0114] S2. The original video stream is input into the decoder, and the decoder outputs two data streams, one is the original resolution data, and the other is the downscaled small resolution data. At the same time, the initial decoding information is saved, such as the macroblock division of this frame, MV data, slice type and other information.

[0115] S3. Send the low-resolution data to the NPU to obtain ROI information (i.e., initial coding information). Usually, the ROI is a rectangular frame, and the selected ROI information can be determined according to the actual video scene. For example, for live video, the region of interest is usually a face or a product; for video call scenes, the region of interest is usually only a face; for other scenes, the region of interest may be an animal or other object.

[0116] S4. Send the original resolution data and ROI information, as well as the saved initial encoding information, to the encoder for encoding.

[0117] S5. For the ROI area, in order to improve its encoding quality, first reduce its QP to increase the bit rate. Secondly, for the macroblocks in the ROI area, the division is further divided on the basis of the original division to select the optimal division mode of the macroblocks, and further calculate MV and other information based on the optimal division mode of the macroblocks.

[0118] S6. For the non-ROI area, in order to reduce the bit rate, its QP is appropriately increased. At the same time, in order to reduce the overall calculation amount of transcoding, the macroblocks in the non-ROI area directly reuse the saved initial encoding information.

[0119] S7. Output the final coding number information.

[0120] S8. Encode each frame in the original resolution data using the final encoding information to obtain a transcoded video stream.

[0121] In this way, the present disclosure combines ROI information, and reuses the coded information once based on the ROI, further optimizes the coded information in the ROI area, and directly reuses the coded information once in the non-ROI area, firstly, the amount of transcoding calculation is reduced, and the transcoding cost is reduced, and secondly, the bit rate is reduced without reducing the visual quality, thereby reducing the bandwidth cost.

[0122] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present disclosure.

[0123] The above is an introduction to the method embodiment. The following is a further explanation of the scheme disclosed in the present invention through an apparatus embodiment.

[0124] Figure 4 FIG. 5 shows a block diagram of a video transcoding device 500 according to an embodiment of the present disclosure. Figure 4 As shown, the device 400 includes:

[0125] A first acquisition module 410, used to acquire an original video stream;

[0126] A decoding module 420 is used to decode the original video stream to obtain original resolution data, low resolution data and initial coding information corresponding to the original video stream, wherein the resolution of the low resolution data is lower than the resolution of the original resolution data;

[0127] A second acquisition module 430 is used to acquire a region of interest of each frame in the original resolution data, wherein the region of interest of each frame is determined based on the low resolution data;

[0128] The processing module 440 is used to adjust the initial coding information corresponding to the region of interest of each frame in the original resolution data, and retain the initial coding information corresponding to the non-region of interest of each frame, so as to obtain the final coding information of each frame;

[0129] The encoding module 450 is used to encode each frame in the original resolution data using the final encoding information of each frame to obtain a transcoded video stream.

[0130] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0131] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, including:

[0132] at least one processor; and

[0133] a memory communicatively connected to the at least one processor; wherein,

[0134] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any one of the above method embodiments.

[0135] According to an embodiment of the present disclosure, the present disclosure further provides a vehicle, comprising: the video transcoding device as described in the above embodiment or the electronic device as described in the above embodiment.

[0136] According to an embodiment of the present disclosure, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute any one of the above method embodiments.

[0137] Figure 5 A schematic block diagram of an electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0138] The device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0139] A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0140] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as method 100. For example, in some embodiments, the method 100 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the method 100 described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the method 100 in any other appropriate manner (e.g., by means of firmware).

[0141] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0142] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0143] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0144] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0145] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0146] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0147] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0148] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A video transcoding method, characterized in that: include: Get the original video stream; Decoding the original video stream to obtain original resolution data, low resolution data, and initial coding information corresponding to the original video stream, wherein the resolution of the low resolution data is lower than the resolution of the original resolution data; Acquire a region of interest of each frame in the original resolution data, wherein the region of interest of each frame is determined based on the low resolution data; Adjusting the initial coding information corresponding to the region of interest of each frame in the original resolution data, and retaining the initial coding information corresponding to the region of non-interest of each frame, so as to obtain the final coding information of each frame; Each frame in the original resolution data is encoded using the final encoding information of each frame to obtain a transcoded video stream.

2. The method according to claim 1, characterized in that The step of obtaining the region of interest of each frame in the original resolution data comprises: Determining a current resolution and a current scene of each frame in the low-resolution data; Get the preset correspondence between resolution, scene and region of interest; Matching the current resolution and the current scene with the preset corresponding relationship to determine a region of interest for each frame in the low-resolution data; According to the resolution ratio relationship between the low-resolution data and the original-resolution data, the region of interest of each frame in the low-resolution data is scaled to determine the region of interest of each frame in the original-resolution data.

3. The method according to claim 1, characterized in that The decoding of the original video stream to obtain original resolution data, low resolution data and initial coding information corresponding to the original video stream includes: Acquire a preset resolution corresponding to the low-resolution data; Based on the preset resolution, the original video stream is decoded to obtain original resolution data and initial coding information corresponding to the original video stream and the low-resolution data, wherein the initial coding information includes: an initial macroblock corresponding to the original resolution data, a motion vector of the initial macroblock, and a type of slice to which the initial macroblock belongs.

4. The method according to claim 3, characterized in that: The adjusting the initial coding information corresponding to the region of interest of each frame in the original resolution data includes: Determine at least one macroblock division method corresponding to a preset coding standard; Calculating the distortion amount corresponding to each macroblock division method in the at least one macroblock division method; According to the magnitude of the distortion amount corresponding to each macroblock division method, the initial coding information corresponding to the region of interest of each frame in the original resolution data is adjusted.

5. The method according to claim 4, characterized in that The adjusting the initial coding information corresponding to the region of interest of each frame in the original resolution data according to the magnitude of the distortion amount corresponding to each macroblock division method includes: Selecting a macroblock partitioning method with the smallest distortion from the at least one macroblock partitioning method; Determining the macroblock division mode with the smallest distortion as the optimal macroblock division mode; At least one new macroblock corresponding to the region of interest of each frame in the macroblock optimal partitioning mode and the type of slice to which each of the at least one new macroblock belongs are determined, and a motion vector of each of the at least one new macroblock is calculated.

6. The method according to claim 5, characterized in that The obtaining of the final encoding information of each frame comprises: Multiplexing and storing the initial macroblocks of the non-interested area of ​​each frame in the original resolution data, the motion vectors of the initial macroblocks of the non-interested area of ​​each frame, and the types of slices to which the initial macroblocks of the non-interested area of ​​each frame belong; At least one new macroblock corresponding to the area of ​​interest of each frame in the original resolution data, the type of slice to which the at least one new macroblock belongs, the motion vector of the at least one new macroblock, the initial macroblock of the non-interest area of ​​each frame, the motion vector of the initial macroblock of the non-interest area of ​​each frame, and the type of slice to which the initial macroblock of the non-interest area of ​​each frame belongs are determined as the final encoding information of each frame.

7. The method according to claim 5, characterized in that The encoding of each frame in the original resolution data using the final encoding information of each frame to obtain a transcoded video stream includes: reducing the quantization parameter of the region of interest of each frame in the original resolution data, and increasing the quantization parameter of the region of no interest of each frame in the original resolution data; Each frame in the original resolution data after the quantization parameter is adjusted is encoded using the final encoding information of each frame to obtain a transcoded video stream.

8. A video transcoding device, characterized in that: include: A first acquisition module, used for acquiring an original video stream; A decoding module, used for decoding the original video stream to obtain original resolution data, low resolution data and initial coding information corresponding to the original video stream, wherein the resolution of the low resolution data is lower than the resolution of the original resolution data; A second acquisition module, used to acquire a region of interest of each frame in the original resolution data, wherein the region of interest of each frame is determined based on the low resolution data; A processing module, used for adjusting the initial coding information corresponding to the region of interest of each frame in the original resolution data, and retaining the initial coding information corresponding to the non-region of interest of each frame, so as to obtain the final coding information of each frame; The encoding module is used to encode each frame in the original resolution data using the final encoding information of each frame to obtain a transcoded video stream.

9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

11. A vehicle, characterized in that: include: The apparatus as claimed in claim 8, and / or the electronic device as claimed in claim 9, and / or the readable storage medium as claimed in claim 10.