Video processing method and device and related equipment

By downsampling and filling data processing of video frame images, a video code stream comparable to the original image data is generated, which solves the problems of low video transmission efficiency and waste of bandwidth resources, and realizes efficient video interaction and complete mask information transmission.

CN120583253APending Publication Date: 2025-09-02HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410240101.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-01
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

During video transmission, the efficiency of interacting video between the sending end and the receiving end is low, and it occupies a lot of bandwidth resources, especially due to the additional transmission of mask information.

Method used

By downsampling the video frame image, a third image with a size comparable to the original image is generated, and complete mask information is included during the transmission process, the data amount of the video code stream is reduced, and the image size is adjusted to match the original size using the fill data, and the video code stream is coded to generate.

Benefits of technology

The bandwidth resources required for video transmission are reduced, the video interaction efficiency between the sending and the receiving end is improved, and the visual presentation effect of video playback and the integrity of mask information is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120583253A_ABST
    Figure CN120583253A_ABST
Patent Text Reader

Abstract

The invention discloses a video processing method and device and related equipment, and relates to the technical field of video processing. The video processing device obtains a first image in a to-be-processed video and mask information corresponding to the first image, wherein the size of the first image is a first size; and then, the video processing device performs down-sampling on the first image to obtain a second image, and generates a third image according to the second image and the mask information, the size of the third image is the first size, so that the video processing device performs coding based on the third image to obtain a video code stream corresponding to the third image. Therefore, the data volume of the third image is equivalent to the data volume of the first image in the video to be processed, so that the data volume of the video code stream required to be transmitted subsequently can be reduced, and bandwidth resources required to be occupied by video transmission can be reduced. Meanwhile, the video code stream comprises complete mask information, so that the requirement of a receiving end on the mask information can be met, and the visual presentation effect during video playing can be guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video processing technology, and in particular to a video processing method, apparatus, and related equipment. Background Art

[0002] In scenarios like video conferencing and live streaming, video frames are often processed based on the corresponding mask information. Mask information refers to information that can be used to identify pixel regions within a video frame. For example, mask information can be used to identify the pixel region representing a person's background within a video frame. By replacing the pixel values ​​within this pixel region with specified pixel values, the background of the person in the video can be replaced.

[0003] Currently, before sending a video to a receiver, the sender typically converts each frame into YUV420 format (or YUV422 format, etc.). Y represents the brightness of the image, U represents the chroma of the image, and V represents the density of the image. The sender then mixes and sorts the YUV420 image data with the mask information corresponding to the frame into a fixed YUV444 format. The sender then encodes this YUV444 data into a video stream, which is then sent to the receiver.

[0004] However, in addition to sending the video to the receiving end, the sending end also sends additional mask information, which will result in low efficiency of video interaction between the sending end and the receiving end and more bandwidth resources required. Summary of the Invention

[0005] This application provides a video processing method to improve the efficiency of interactive video between a transmitter and a receiver and reduce the bandwidth resources required for video transmission. In addition, this application also provides a video processing device, a computing device, a computer-readable storage medium, and a computer program product.

[0006] In the first aspect, the present application provides a video processing method, which can be executed by a corresponding video processing device. Specifically, the video processing device obtains a first image in a video to be processed and mask information corresponding to the first image. The mask information can be, for example, a mask map, which can be used to indicate one or more pixel areas in the first image, and the size of the first image is a first size; then, the video processing device downsamples the first image to obtain a second image, and generates a third image based on the second image and the mask information, and the size of the third image is the first size, so that the video processing device encodes the third image based on the third image to obtain a video code stream corresponding to the third image.

[0007] In this way, during video processing, the video processing device first downsamples the first image in the video and then generates a third image of the first size based on the mask information corresponding to the first image and the downsampled second image. This ensures that the data volume of the generated third image is comparable to that of the first image in the video to be processed. This reduces the data volume of the video stream required for subsequent transmission, thereby reducing the bandwidth resources required for video transmission and improving the efficiency of video exchange between the transmitter and receiver. Furthermore, the video stream corresponding to the generated third image includes complete mask information, which satisfies the receiver's mask information requirements while reducing the bandwidth resources required for video transmission. Furthermore, the complete mask information ensures that the receiver's visual presentation quality when playing video based on the received video stream. Furthermore, the color information in the first image is typically rich and redundant. This means that the color information difference between the fourth image obtained by the subsequent upsampling of the third image by the receiver and the first image is typically small. This ensures that the video content played by the receiver based on the fourth image is as close to the video content in the first image as possible, thereby ensuring the video playback quality at the receiver.

[0008] In one possible embodiment, the size of the second image is a second size, and the second size is smaller than the first size. When downsampling the first image, the video processing device may first determine the second size based on the data volume of the mask information corresponding to the first image, and then downsample the first image based on the second size to obtain a second image, where the sum of the data volume of the second image of the second size and the data volume of the mask information does not exceed the data volume of the first image of the first size. In this way, determining the size of the second image based on the data volume of the mask information allows, after concatenating the second image and the mask information, to obtain an image with a data volume no larger than the size of a single frame. This reduces the data volume of the video stream required for subsequent transmission of the image, thereby reducing the bandwidth resources required for video transmission and improving the efficiency of video interaction between the transmitter and receiver.

[0009] In one possible embodiment, the third image includes padding data, and the value of the padding data is a preset value. When the video processing device generates the third image based on the second image and the mask information, specifically when the sum of the data volume of the second image and the data volume of the mask information is less than the data volume of the first image, the second image, the mask information, and the padding data are concatenated to generate the third image. In this way, by adding padding data of a specified value, an image with a data volume corresponding to the size of a single frame can be concatenated. This reduces the amount of video data required for subsequent transmission of the image, thereby reducing the bandwidth resources required for video transmission and improving the efficiency of video interaction between the transmitter and receiver.

[0010] In a possible implementation, the mask information corresponding to the first image includes a mask image, and the size of the mask image is a first size, or the size of the mask image is smaller than the first size.

[0011] In one possible embodiment, when the video processing device downsamples the first image, it may specifically downsample the first image in a single directional dimension, such as downsampling the first image in the height direction, or downsampling the first image in the width direction, to obtain a second image; in this way, the width of the obtained second image is equal to the width of the first image (downsampling in the height direction), or the height of the second image is equal to the height of the first image (downsampling in the width direction).

[0012] In the second aspect, the present application provides a video processing method, which can be executed by a corresponding video processing device. Specifically, the video processing device obtains a video code stream to be processed, the video code stream includes a third image, and the third image includes mask information, and the size of the third image is a first size; then, the video processing device decodes the video code stream to obtain the third image, and splits the third image into a second image and mask information, so that the video processing device upsamples the second image to obtain a fourth image, the size of the fourth image is the first size, and performs preset video processing operations based on the fourth image and the mask information, such as performing an operation of playing a video frame image.

[0013] In this way, the data volume of the third image is equivalent to the data volume of the first image in the video to be processed, thereby reducing the data volume of the video stream transmission, thereby reducing the bandwidth resources required for video transmission and improving the efficiency of interactive video between the sender and the receiver. At the same time, the video stream corresponding to the third image includes complete mask information, which satisfies the receiver's demand for mask information while reducing the bandwidth resources required for video transmission. Moreover, the complete mask information can also guarantee the visual presentation effect of the receiver when playing the video according to the received video stream. In addition, the color information in the first image is usually rich and redundant, which makes the color information difference between the fourth image obtained by the video processing device by upsampling the third image and the original video frame image usually small, thereby ensuring that the video content played by the video processing device according to the fourth image can restore the video content in the video frame image as much as possible, thereby ensuring the subsequent video playback effect.

[0014] In a third aspect, the present application provides a video processing device, which includes various modules for executing the video processing method in the first aspect or any possible implementation of the first aspect.

[0015] In a fourth aspect, the present application provides a video processing device, which includes various modules for executing the video processing method in the second aspect or any possible implementation of the second aspect.

[0016] In a fifth aspect, the present application provides a computing device, comprising a processor and a memory. The processor and the memory communicate with each other. The processor is used to execute instructions stored in the memory, so that the computing device executes the video processing method as described in the first aspect or any one of the implementations of the first aspect, or executes the video processing method as described in the second aspect. It should be noted that the memory can be integrated into the processor or can be independent of the processor. The computing device may further include a bus. The processor is connected to the memory via the bus. The memory may include a readable memory and a random access memory.

[0017] In a sixth aspect, the present application provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computing device, the computing device executes the operating steps of the video processing method described in the first aspect or any one of the implementations of the first aspect, or the computing device executes the operating steps of the video processing method described in the second aspect.

[0018] In the seventh aspect, the present application provides a computer program product comprising instructions, which, when run on a computing device, enables the computing device to execute the operating steps of the video processing method described in the first aspect or any one of the implementations of the first aspect, or enables the computing device to execute the operating steps of the video processing method described in the second aspect.

[0019] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is a schematic diagram of an exemplary application scenario provided by this application;

[0021] Figure 2 A flowchart of a video processing method provided in this application;

[0022] Figure 3a A schematic diagram of obtaining a third image by stitching the second image and the mask image together;

[0023] Figure 3b A schematic diagram of obtaining a third image by left-right splicing of the second image and the mask image;

[0024] Figure 4a Schematic diagram of stitching the second image and the mask image together to obtain the third image;

[0025] Figure 4b A schematic diagram of obtaining a third image by left-right splicing of the second image and the mask image;

[0026] Figure 5 A schematic diagram of the structure of a video processing device provided in this application;

[0027] Figure 6 A schematic structural diagram of another video processing device provided in this application;

[0028] Figure 7 A schematic diagram of the hardware structure of a computing device provided in this application. DETAILED DESCRIPTION

[0029] To make the above-mentioned objects, features, and advantages of the present application more clearly understood, various non-limiting embodiments of the embodiments of the present application will be exemplified below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, not all of them. Based on the embodiments in this application, all other embodiments obtained based on the above content are within the scope of protection of this application.

[0030] See also Figure 1 , is a schematic diagram of an exemplary application scenario provided in the embodiment of this application. Figure 1 As shown, user 1 and user 2 can interact via video through terminal 100 and terminal 200, such as holding a video conference. The video stream sent by terminal 100 can be forwarded to terminal 200 via server 300; at the same time, terminal 100 can also receive the video stream sent by terminal 200 via server 200.

[0031] The terminal 100 may be configured with a video processing device 101, and the terminal 200 may be configured with a video processing device 201. Furthermore, the video processing device in each terminal can be used to encode video frames to obtain a video stream, and to decode the received video stream. Each terminal may be, for example, a mobile phone, a tablet computer (Pad), a virtual reality (VR) terminal, an augmented reality (AR) terminal, an in-vehicle terminal, etc., without limitation.

[0032] Taking the example of terminal 100 sending a video to terminal 200, user 1 can import the video on terminal 100; alternatively, terminal 100 can be equipped with a camera, allowing user 1 to capture the video. Furthermore, user 1 can semantically annotate the image content in the video on terminal 100. For example, user 1 can select and annotate the image area in the video frame containing a face, and the remaining image area as the background area. In this case, the size of each frame image in the video can be 4032×3024 pixels, that is, the image is 4032 pixels wide and 3024 pixels high. Generally, the size of the video frame image determines the resolution of the video.

[0033] Then, the video processing device 101 in the terminal 100 can begin processing the video. Specifically, taking the processing of a frame of image in the video (hereinafter referred to as the first image) as an example, the video processing device 101 can determine a mask image for the first image based on the semantic annotation results of user 1. The mask image can identify the image area in the first image where the content is a human face. The mask image can be a grayscale image, and thus the data volume of the mask image can be 1 / 3 of the data volume of the first image. Furthermore, the video processing device 101 can downsample the first image (i.e., reduce the size of the first image) to obtain a second image. The data volume of the second image can be, for example, 2 / 3 of the data volume of the first image, for example, the size of the second image can be 4032×2016 pixels. Then, based on the second image and the mask image, the video processing device 101 can generate a third image with a size of 4032×3024 pixels. For example, the video processing device 101 can splice the pixel data of the second image with the pixel data of the mask image to obtain a third image with a size of 4032×3024 pixels. Finally, the video processing device 101 may obtain a video code stream corresponding to the third image based on the encoding of the third image, so that the terminal 100 may send the video code stream to the terminal 200 through the server 300 .

[0034] Accordingly, after receiving the video stream, the video processing device 201 on the terminal 200 can decode the video stream to obtain a third image. The video processing device 201 can then decompose the third image to obtain a second image and a mask image, and upsample the second image to obtain a fourth image of size 4032×3024 pixels. The video processing device 201 can then present the fourth image based on the mask image. For example, the video processing device 201 can determine the background area in the fourth image based on the mask image and, when presenting the fourth image, replace the pixel values ​​of the background area in the fourth image with the pixel values ​​corresponding to the new background, thereby achieving background replacement.

[0035] In actual application, for each frame image in the video, the terminal 100 and the terminal 200 can refer to the above-mentioned processing method for the first image to process it, so as to achieve video interaction between the terminal 100 and the terminal 200.

[0036] Thus, during video interaction between terminal 100 and terminal 200, video processing device 101 first downsamples the first image in the video, and then generates a third image of size 4032×3024 pixels based on the mask image corresponding to the first image and the downsampled second image. This ensures that the data volume of the generated third image is equivalent to that of the first image in the video. Thus, for a frame of image in the video, the data volume of the video stream sent by terminal 100 to terminal 200 is the data volume of the frame, rather than the data volume of the frame plus the data volume of the mask image. This effectively reduces the data volume of the video stream required for subsequent transmission, thereby reducing the bandwidth resources required for video transmission and improving the efficiency of video interaction between the sender and receiver.

[0037] At the same time, the video stream sent by terminal 100 includes a complete mask image, which satisfies terminal 200's mask image requirements while reducing the bandwidth resources required for video transmission. Furthermore, the complete mask image also ensures the visual presentation effect when terminal 200 plays the video based on the received video stream, such as accurately implementing background replacement (avoiding replacing image content in non-background areas).

[0038] Furthermore, the color information in the first image is usually rich and redundant, which makes the color information difference between the fourth image obtained by upsampling the third image by the terminal 200 and the original first image in the video usually small. This can ensure that the video content played by the terminal 200 based on the fourth image can restore the video content in the first image as much as possible, thereby ensuring the video playback effect at the receiving end.

[0039] Similarly, when terminal 200 needs to send a video to terminal 100, the video processing device 201 on terminal 200 can also use a similar method to encode the transmitted video to obtain a video code stream, and the video processing device 100 on terminal 100 can use a similar method to decode the received video code stream and play the video, which will not be elaborated on.

[0040] Exemplarily, the video processing device 101 (and the video processing device 201 ) may be implemented by software or hardware.

[0041] In the first example, when implemented by software, the video processing device 101 may include code running on an instance, which may be at least one of a host, a virtual machine, a container, a thread, and a process. In actual application, the video processing device 101 may be, for example, a video codec on the terminal 100.

[0042] In the second example, when implemented by hardware, the video processing device 101 can be implemented by at least one physical device including a processor, such as a server. The processor can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a system on chip (SoC), a software-defined infrastructure (SDI) chip, an artificial intelligence (AI) chip, a data processing unit (DPU), or any combination thereof. In addition, the number of processors included in the video processing device 101 can be one or more, and the type of processors included can be one or more. The number and type of processors can be set according to the business requirements of the actual application, and this embodiment does not limit this.

[0043] It is worth noting that the above Figure 1 The application scenario shown is only an example and is not intended to be limiting. For example, in actual application, the number of terminals accessing the server 300 can be any number, so that the server 300 can forward the video stream generated by one terminal to multiple other terminals in parallel. For example, in a live video broadcast scenario, the server 300 can forward the video stream generated by the terminal 100 to multiple different terminals, and the multiple terminals can respectively present the corresponding live video content based on the received video stream, so that the users of each terminal can see the live screen. Alternatively, in a multi-person video conferencing scenario, the video streams generated by multiple terminals can be sent in parallel to the same terminal, so that the terminal can present multiple video contents based on the fusion of the received multi-channel video streams.

[0044] For ease of understanding, an embodiment of the video processing method provided in this application is described below with reference to the accompanying drawings.

[0045] See also Figure 2 , Figure 2 A flow chart of a video processing method provided in an embodiment of the present application, which can be applied to Figure 1 The application scenario described above may be applied to other applicable application scenarios. Figure 1 The video processing device 101 in the application scenario shown is used as an example for illustrative description. Figure 2 In the illustrated embodiment, terminal 100 transmits a video to terminal 200 .

[0046] in, Figure 2 The video processing method shown may specifically include:

[0047] S201: The video processing device 101 obtains a first image in a video and mask information corresponding to the first image, wherein the size of the first image is a first size.

[0048] The video to be processed may be provided by user 1, such as by user 1 importing the video into terminal 100. Alternatively, the video to be processed may be acquired by terminal 100, such as by capturing the video using a camera configured thereon (e.g., a camera), though this is not a limitation. For ease of understanding and explanation, this embodiment uses the processing of one frame of image (i.e., the first image) in the video as an example for explanation. The processing of other frames of image in the video may be similar to the processing of the first image.

[0049] In addition, the video processing device 101 will also obtain mask information corresponding to the first image. Exemplarily, the mask information can be a mask map, in which each pixel can have the pixel value of only one channel (each pixel in the first image can have the pixel values ​​of three channels, R, G, and B), and the multiple pixel values ​​in the mask map can be used to indicate the pixel area in the first image. For example, when the first image is specifically a face image in a video conferencing scene, the value of the pixel point corresponding to the face image in the mask map can be 255, and the value of the remaining pixel points in the mask map can be 0, so that the face image area and the background area in the first image can be distinguished according to the value of the pixel points in the mask map. In actual application, the mask information can also be other types of information, which is not limited to this.

[0050] When the mask information corresponding to the first image is specifically a mask map, this embodiment provides the following implementation examples for obtaining the mask map.

[0051] In a first implementation example, the video processing device 101 may use a semantic segmentation algorithm to perform semantic recognition on the image content in the first image, and segment the pixel regions in the first image based on the semantic recognition result to obtain a mask image (or semantic mask image) corresponding to the first image. The semantic segmentation algorithm may be, for example, a normalized cut (N-cut) algorithm.

[0052] In a second implementation example, the video processing device 101 can be configured with an artificial intelligence (AI) model for image segmentation, and the video processing device 101 can input the first image into the AI ​​model, and the AI ​​model performs semantic recognition on the picture content in the first image and outputs a corresponding mask image, in which different pixel areas in the mask image correspond to different semantics recognized by the AI ​​model (the pixels in each pixel area have the same value).

[0053] In the third implementation example, user 1 can frame the picture content in the first image on the terminal 100. For example, when the first image is an image including a face, user 1 can frame the face area in the first image, so that the video processing device 101 can generate a mask image corresponding to the first image based on the framing operation performed by user 1. The pixel area indicated by the mask image can be used to indicate the pixel area framed by user 1 in the first image.

[0054] In addition to the above implementation examples, the video processing device 101 may also adopt other methods to obtain the mask information corresponding to the first image, which is not limited to this.

[0055] S202: The video processing device 101 downsamples the first image to obtain a second image.

[0056] Each pixel in the first image may include pixel values ​​for three channels: red (R), green (G), and blue (B). Typically, the color information in the first image is rich and redundant, that is, the difference in pixel values ​​between two adjacent pixels in the first image is small. Therefore, the video processing device 100 may compress the first image, specifically, downsample the first image to obtain a second image smaller in size than the first image.

[0057] In a specific implementation, the video processing device 101 may first determine the size of the second image to be generated (hereinafter referred to as the second size). Then, based on the size, the video processing device 101 may downsample the first image using a bilinear interpolation algorithm to obtain a second image of the second size. The sum of the data volume of the second image of the second size and the data volume of the mask information does not exceed the data volume of the first image of the first size. In actual applications, the video processing device 101 may also use other algorithms to downsample the first image, and this is not limited to this.

[0058] Exemplarily, the video processing apparatus 101 may determine the second size of the second image to be generated based on the data volume of the mask information. Assuming that the mask information is a mask map, the video processing apparatus 101 may determine the second size based on the following various methods.

[0059] In a first implementation, the height and width of the first image and the mask image are both H and W (H represents height and W represents width). Therefore, when each pixel in the mask image has only one pixel value, the data size of the mask image is typically 1 / 3 of the data size of the first image. In this case, the video processing device 101 can determine that the second size of the second image to be generated is 2 / 3 of the first image. Furthermore, the video processing device 101 can downsample the first image in a single dimension to obtain the second image.

[0060] Exemplarily, the video processing device 101 may calculate the height H″ and width W′ of the second image to be generated based on the following formula (1).

[0061]

[0062] in, Indicates rounding down.

[0063] At this time, the width of the second image obtained by downsampling is equal to the width of the first image, but the height of the second image is 2 / 3 of the height of the first image.

[0064] Alternatively, the video processing device 101 may calculate the height H′ and width W′ of the second image to be generated based on the following formula (2).

[0065]

[0066] At this time, the height of the second image obtained by downsampling is equal to the height of the first image, but the width of the second image is 2 / 3 of the width of the first image.

[0067] In a second implementation, the height and width of the first image are H1 and W1, respectively, and the height and width of the mask image are H2 and W2, respectively. Furthermore, the height of the mask image is less than or equal to the height of the first image, and the width of the mask image is less than or equal to the width of the first image. Therefore, when each pixel in the mask image has only one pixel value, the video processing device 101 can calculate the height H′ and width W′ of the second image to be generated based on the following formula (3).

[0068]

[0069] At this time, the video processing device 101 may downsample the first image in the height direction of the first image, that is, the video processing device 101 may sample multiple rows of pixels in the first image to obtain the second image.

[0070] Alternatively, the video processing device 101 may also calculate the height H′ and width W′ of the second image to be generated based on the following formula (4).

[0071]

[0072] At this time, the video processing apparatus 101 may downsample the first image in the width direction of the first image, that is, the video processing apparatus 101 may sample multiple columns of pixels in the first image to obtain the second image.

[0073] It should be noted that the above description is based on an example in which the mask information is specifically a mask image. When the mask information is specifically other types of information, the video processing device 101 can also determine the second size of the second image in a similar manner to the above, and downsample the first image according to the second size to obtain the second image.

[0074] S203: The video processing device 101 generates a third image according to the down-sampled second image and the acquired mask information. The size of the third image is the first size.

[0075] It can be understood that the size of the second image obtained by downsampling is smaller than the size of the first image. Therefore, the video processing device 101 can generate a new image of the first size, that is, the third image mentioned above, based on the smaller second image and the mask information.

[0076] In a first implementation example, the video processing device 101 can directly splice the first image with the mask information to obtain a third image of the first size. For example, when the mask information is specifically a mask map, the size of the mask map is {height: H, width: W}, and the data volume of the mask map is H*W; when the values ​​of each pixel in the image are stored row by row, the size of the second image is {height: H, width: W*2 / 3}, and the data volume of the second image is H*W*2 (that is, 3*H*W*2 / 3, each pixel includes the pixel value of 3 channels). Among them, the W is divisible by 3. Then, the video processing device 101 can store the values ​​of each pixel in the mask map continuously after the second image to form a third image of the first size, and the data volume of the third image is 3*H*W (that is, H*W*2+H*W). At this time, the generated third image can be as follows Figure 3a As shown, a top-bottom splicing method is formed. For another example, when the values ​​of each pixel in the image are stored in columns, the size of the second image is {height: H*2 / 3, width: W}, and the video processing device 101 can store the values ​​of each pixel in the mask image continuously after the second image to form a third image of the first size. The data volume of the third image is 3*H*W (i.e., H*W*2+H*W). Among them, H can be divided by 3. At this time, the generated third image can be as follows Figure 3b As shown, a left-right splicing method is formed. Wherein, the color format of the third image is RGB format, that is, the value of each pixel includes color information of three channels.

[0077] In a second implementation, when determining the size of the second image to be generated based on the above formulas (1) to (4), since the height or width of the first image is determined by rounding down, if the video processing device 101 can directly splice the first image with the mask information, the size of the generated third image may be smaller than the first size. For example, assuming that the first size is 4000×3000 pixels, where the width is 4000 pixels and the height is 3000 pixels, and the sizes of the first image and the mask image are both the first size, the height of the second image calculated by the video processing device 101 according to formula (2) is 3000 pixels and the width is 2666 pixels (i.e., 4000*2 / 3, with decimals rounded down). In this way, after the video processing device 101 splices the second image with the mask image, the size of the third image obtained is {height: 3000, width: 3999}, which is smaller than the first size {height: 3000, width: 4000}.

[0078] Therefore, the video processing device 101 can splice the second image, the mask information and the padding data to obtain a third image of the first size. The padding data is a pixel value, and the value of the padding data is a fixed value between 0 and 255, such as the values ​​of the padding data are all 0 or all 255. Still taking the first size of {height: 3000, width: 4000} as an example, after the video processing device 101 splices the second image with the mask image (the height of the splicing result is 3000 and the width is 3999), it can continue to splice with the padding data of the size of {height: 3000, width: 1}, so that the size of the final third image can be {height: 3000, width: 4000}.

[0079] Exemplarily, when the size of the mask image is the first size, the amount S of padding data can be calculated using the following formula (5).

[0080] S=H*W*2-H′*W′*3 Formula (5)

[0081] When the size of the mask image is smaller than the first size, the amount S of padding data can be calculated using the following formula (6).

[0082] S=3(H1W1-H′W′)-(H2W2) Formula (6)

[0083] The order in which the video processing device 101 splices the second image, the mask information, and the filling data can be any order and is not limited. For example, the video processing device 101 can store the values ​​of each pixel in the mask image continuously after the second image, and then store the filling data after the mask image to form a third image of the first size. Then, when the values ​​of each pixel in the image are stored row by row, the generated third image can be as follows: Figure 4a As shown, a top-bottom splicing method is formed. When the values ​​of each pixel in the image are stored in columns, the generated third image can be as follows Figure 4b In actual applications, the data splicing order in the third image may also be mask information - second image - padding data, mask information - padding data - second image, padding data - mask information - second image, or padding data - second image - mask data, etc., and this is not limited to this.

[0084] S204: The video processing device 101 encodes the third image to obtain a video stream corresponding to the third image.

[0085] In one possible embodiment, the third image is in RGB color format. Therefore, the video processing device 101 first converts the third image into an image in a specified format, such as YUV420 format, to convert the third image into a specified color space. The video processing device 101 then performs video encoding on the converted image in the specified format, such as encoding the image based on the H.264 standard, to obtain a corresponding video stream corresponding to the third image. In actual applications, the video processing device 101 can use any video encoding method to encode the video stream, and this embodiment is not limited to this.

[0086] It is worth noting that in this embodiment, video encoding is performed on one frame of the video, and the resulting video stream is the video stream corresponding to the third image. For other frames of the video, the video processing device 101 can use the method described in steps S201 to S204 above to process and encode the frame to obtain the corresponding video stream, thereby obtaining the video stream corresponding to the entire video.

[0087] S205 : The video processing device 101 forwards the video stream corresponding to the third image to the terminal 200 via the server 300 .

[0088] For example, the video processing device 101 may forward the video stream to the terminal 200 via the server 300 based on the Real-time Transport Protocol (RTP) or other types of protocols, which is not limited to this.

[0089] In this embodiment, the video processing device 101 sending a video stream is used as an example for description. In actual application, the video processing device 101 can send the video stream to the terminal 200 by calling a network card or other device on the terminal 100.

[0090] S206: The video processing device 201 in the terminal 200 decodes the received video code stream to obtain a third image of the first size.

[0091] The specific implementation process of decoding the video code stream has relevant applications in actual scenarios and will not be described in detail here.

[0092] S207: The video processing device 201 separates the third image to obtain a second image and mask information.

[0093] It can be understood that the video processing device 101 in the terminal 100 generates a third image by splicing the second image and the mask information (and filling data). Therefore, the video processing device 101 in the terminal 200 can split the third image according to the data splicing order in the third image to obtain the second image and complete mask information from the third image.

[0094] In a specific implementation, when image data is stored by row, if the video processing device 101 generates a third image by sequentially splicing the second image, mask information, and padding data, the video processing device 201 can extract the first N data from the third image based on the data volume N of the second image and use the first N data as the second image. Then, based on the data volume M of the mask information, after extracting the second image, the video processing device 201 can continue to extract M data from the third image, where the M data serves as the mask data.

[0095] During the video exchange process between terminal 100 and terminal 200, the size (or data volume) of the second image and the data volume of the mask information (or the size of the mask image) can be fixed values, so that the video processing device 201 extracts the second image from the third image based on the data volume N and extracts the mask information from the third image based on the data volume M. Alternatively, when the video processing device 101 sends the video stream to terminal 200, it can also send the size of the second image and the data volume of the mask information to terminal 200, such as by adding the size of the second image and the data volume of the mask information to the corresponding packet header of the video stream, so that the size and data volume are transmitted to terminal 200 along with the video stream.

[0096] S208: The video processing device 201 upsamples the second image to obtain a fourth image, where the size of the fourth image is the first size.

[0097] It can be understood that since the size of the second image is the second size (smaller than the first size), and the size of the video frame image played by the terminal 200 is the first size, the video processing device 201 can enlarge the second image after obtaining the second image, specifically by upsampling the second image to obtain an image of the first size (i.e., the fourth image). The video processing device 201 can use a bilinear interpolation algorithm or other algorithm to achieve upsampling of the second image, and this is not limited; and the direction dimension in which the video processing device 201 upsamples the second image matches the direction dimension in which the video processing device 101 downsamples the first image. In addition, the video processing device 101 can specifically upsample the second image according to the first size of the third image.

[0098] Because the first image has rich and redundant color information, after downsampling the first image to obtain the second image and then upsampling the second image to obtain the fourth image, the color information of the fourth image is generally less different from the color information of the first image. Accordingly, the image content of the fourth image generated by the video processing device 201 is also highly similar to the image content of the first image, thereby achieving restoration of the original video frame image.

[0099] S209: The video processing device 201 performs a preset video processing operation according to the fourth image and the mask information.

[0100] Exemplarily, the video processing operation may be, for example, an operation of playing a video.

[0101] In a specific implementation, when playing the fourth image, the video processing device 201 can replace the background of the fourth image according to the mask information, or add special effects to the fourth image. For example, the video processing device 201 can determine the mask area in the fourth image according to the mask information, so that the video processing device 201 can set the values ​​of pixels in other areas outside the mask area to the pixel values ​​corresponding to the new background, thereby replacing the background of the fourth image. For another example, the video processing device 201 can determine the mask area in the fourth image according to the mask information, and set the values ​​of some or all pixels in the mask area to the pixel values ​​corresponding to the special effect, thereby adding special effects to the fourth image.

[0102] In other embodiments, the video processing operation performed by the video processing device 201 according to the fourth image and the mask information may also be other types of operations, which may be specifically set according to the requirements of the actual application scenario and are not limited thereto.

[0103] In actual application, the video processing device 201 usually continuously receives video code streams corresponding to multiple images. Therefore, for the video code stream corresponding to each image, the video processing device 201 can use the above-mentioned steps S206 to S209 to process the video code stream corresponding to the image, so that the terminal 200 can continuously play multiple frames of complete video frame images to ensure the smoothness of video frame playback.

[0104] It should be noted that in this embodiment, the video processing device 101 downsamples the first image in one directional dimension, and the video processing device 201 upsamples the first image in the directional dimension. This ensures that signal aliasing is not likely to occur after the second image and the mask information are spliced, thereby avoiding affecting the compression ratio of subsequent video intra-frame encoding. However, in other embodiments, the video processing device 101 and the video processing device 201 may also downsample and upsample the image in multiple directional dimensions respectively. In addition, the size of the image when the video processing device 101 downsamples can be determined based on the data volume of the mask information or can be a preset value, and there is no limitation on this.

[0105] Furthermore, when terminal 200 transmits a video to terminal 100, video processing device 201 may process the original video frame image in the manner described in steps S201 to S205 above to generate a video stream including complete mask information, and transmit the video stream to terminal 100 via server 300. Accordingly, video processing device 101 in terminal 100 may decode the received video stream in the manner described in steps S206 to S209 above to obtain the original video frame image and complete mask information, which will not be described in detail here.

[0106] In this embodiment, the video code stream sent by terminal 100 to terminal 200 can not only carry complete mask information, thereby ensuring the video presentation effect when terminal 200 plays the video according to the complete mask information, but also the data amount of the video code stream corresponding to the third image sent by terminal 100 to terminal 200 is the data amount of one frame of image, rather than the data amount of one frame of image + the data amount of the mask image. This can effectively reduce the bandwidth resources required for video transmission and improve the video interaction efficiency between terminal 100 and terminal 200.

[0107] At the same time, during the video interaction process, the first image is compressed and transmitted by utilizing the rich and redundant color information in the first image, so that the color information difference between the fourth image obtained by the terminal 200 upsampling the third image and the original first image in the video is usually small. This can ensure that the video content played by the terminal 200 based on the fourth image can restore the video content in the first image as much as possible, thereby ensuring the video playback effect at the receiving end.

[0108] It is worth noting that in this embodiment, the Figure 1The illustrated application scenario of interactive video between multiple terminals is used as an example. In other embodiments and application scenarios, after processing each frame image in the video and generating a corresponding video stream, the video processing device 101 can encapsulate the video stream, such as into an MP4 format video. Finally, the video processing device 101 can save the video carrying the mask information locally or in the cloud. Accordingly, the video processing device 201 can obtain the video by accessing the cloud and decode it through the above steps S206 to S208 to obtain the original video frame image and the corresponding mask information.

[0109] It is worth noting that other reasonable step combinations that can be thought of by those skilled in the art based on the above description also fall within the scope of protection of this application. Secondly, those skilled in the art should also be familiar with that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.

[0110] Combination of the above Figures 1 to 4b The video processing method provided in the embodiment of the present application is introduced, and then the structure of the video processing device and computing equipment provided in the embodiment of the present application is introduced with reference to the accompanying drawings.

[0111] See also Figure 5 , shows a schematic structural diagram of a video processing device, the video processing device 500 includes:

[0112] An acquisition module 501 is configured to acquire a first image in a video to be processed and mask information corresponding to the first image, where the size of the first image is a first size;

[0113] A downsampling module 502 is configured to downsample the first image to obtain a second image;

[0114] An image generating module 503 is configured to generate a third image according to the second image and the mask information, where the size of the third image is the first size;

[0115] The encoding module 504 is configured to encode the third image to obtain a video stream corresponding to the third image.

[0116] In a possible implementation, the size of the second image is a second size, and the second size is smaller than the first size;

[0117] The downsampling module 502 is configured to:

[0118] determining a second size according to a data amount of the mask information corresponding to the first image;

[0119] The first image is downsampled according to the second size to obtain a second image, and the sum of the data volume of the second image of the second size and the data volume of the mask information does not exceed the data volume of the first image of the first size.

[0120] In a possible implementation, the third image includes filling data, and a value of the filling data is a preset value;

[0121] Then, the image generation module 503 is configured to, when the sum of the data volume of the second image and the data volume of the mask information is less than the data volume of the first image, concatenate the second image, the mask information and the padding data to obtain a third image.

[0122] In a possible implementation, the mask information corresponding to the first image includes a mask map;

[0123] The size of the mask image is the first size, or the size of the mask image is smaller than the first size.

[0124] In one possible implementation, the downsampling module 502 is configured to:

[0125] Downsampling the first image in a single direction dimension to obtain a second image;

[0126] The width of the second image is equal to the width of the first image, or the height of the second image is equal to the height of the first image.

[0127] because Figure 5 The video processing device 500 shown corresponds to the above Figure 2 The video processing device 101 in the embodiment shown, Figure 5 For the specific implementation of the video processing device 500 and its technical effects, see the above Figure 2 The description of the relevant parts in the illustrated embodiment will not be repeated here.

[0128] See also Figure 6 , shows a schematic structural diagram of a video processing device, the video processing device 600 includes:

[0129] An acquisition module 601 is configured to acquire a video stream to be processed, where the video stream includes a third image, the third image includes mask information, and the size of the third image is a first size;

[0130] A decoding module 602 is configured to decode the video stream to obtain a third image;

[0131] A splitting module 603 is configured to split the third image to obtain a second image and mask information;

[0132] an upsampling module 604, configured to upsample the second image to obtain a fourth image, where the size of the fourth image is the first size;

[0133] The operation execution module 605 is configured to execute a preset video processing operation according to the fourth image and the mask information.

[0134] because Figure 6 The video processing device 600 shown corresponds to the above Figure 2 The video processing device 201 in the embodiment shown, Figure 6 For the specific implementation of the video processing device 600 and its technical effects, see the above Figure 2 The description of the relevant parts in the illustrated embodiment will not be repeated here.

[0135] Figure 7 This is a hardware structure diagram of a computing device 700 provided in this application. The computing device 700 can, for example, implement the above Figure 2 The video processing device 101 in the embodiment shown, or the video processing device 101 that implements the above Figure 2 The video processing device 201 in the illustrated embodiment, etc.

[0136] like Figure 7 As shown, the computing device 700 includes a processor 701, a memory 702, and a communication interface 703. The processor 701, the memory 702, and the communication interface 703 communicate via a bus 704, and may also communicate via other means such as wireless transmission. The memory 702 is used to store instructions, and the processor 701 is used to execute the instructions stored in the memory 702. Furthermore, the computing device 700 may also include a memory unit 705, and the memory unit 705 may be connected to the processor 701, the storage medium 702, and the communication interface 703 via a bus 704. The memory 702 stores program code, and the processor 701 may call the program code stored in the memory 702 to perform the following operations:

[0137] Obtaining a first image in a video to be processed and mask information corresponding to the first image, where the size of the first image is a first size;

[0138] downsampling the first image to obtain a second image;

[0139] generating a third image according to the second image and the mask information, wherein the size of the third image is the first size;

[0140] Based on the third image, encoding is performed to obtain a video stream corresponding to the third image.

[0141] Alternatively, the processor 701 may call the program code stored in the memory 702 to perform the following operations:

[0142] Acquire a video stream to be processed, where the video stream includes a third image, the third image includes mask information, and the size of the third image is the first size;

[0143] Decoding the video code stream to obtain the third image;

[0144] Splitting the third image to obtain a second image and the mask information;

[0145] Upsampling the second image to obtain a fourth image, where the size of the fourth image is the first size;

[0146] A preset video processing operation is performed according to the fourth image and the mask information.

[0147] It should be understood that in this embodiment, the processor 701 may be a CPU, or may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete device components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0148] The memory 702 may include a read-only memory and a random access memory, and provides instructions and data to the processor 701. The memory 702 may also include a nonvolatile random access memory.

[0149] The memory 702 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DRRAM).

[0150] The communication interface 703 is used to communicate with other devices connected to the computing device 700. The bus 704 includes not only a data bus, but also a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as bus 704 in the figure.

[0151] It should be understood that the computing device 700 according to the embodiment of the present application may correspond to the video processing device 500 according to the embodiment of the present application, and may correspond to the device that executes the video processing according to the embodiment of the present application. Figure 2 The method performed by the video processing apparatus 101 or the video processing apparatus 201 in the method shown, and the above and other operations and / or functions implemented by the computing device 700 are respectively to achieve Figure 2 For the sake of brevity, the process of the corresponding method in will not be repeated here.

[0152] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned video processing method.

[0153] The present application also provides a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the computer program product fully or partially generates the process or function described in the present application.

[0154] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0155] The computer program product may be a software installation package. When any of the aforementioned video processing methods needs to be used, the computer program product may be downloaded and executed on a computing device.

[0156] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0157] The terms used in the above embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and claims of this application, the singular expressions "one", "a kind of", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the embodiments of the present application, "one or more" refers to one, two or more; the character " / " generally indicates that the objects associated with each other are in an "or" relationship. In the embodiments of the present application. "Simultaneously" means within the same time period, including situations at the same time. The terms "first", "second", etc. in the specification, claims and drawings of this application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable where appropriate, and this is merely a way of distinguishing objects with the same properties when describing them in the embodiments of the present application.

[0158] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0159] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A video processing method, characterized in that: The method comprises: Obtaining a first image in a video to be processed and mask information corresponding to the first image, where the size of the first image is a first size; downsampling the first image to obtain a second image; generating a third image according to the second image and the mask information, wherein the size of the third image is the first size; Based on the third image, encoding is performed to obtain a video stream corresponding to the third image.

2. The method according to claim 1, characterized in that The size of the second image is a second size, and the second size is smaller than the first size; The downsampling the first image includes: determining the second size according to the data amount of the mask information corresponding to the first image; The first image is downsampled according to the second size to obtain the second image, and the sum of the data volume of the second image of the second size and the data volume of the mask information does not exceed the data volume of the first image of the first size.

3. The method according to claim 1 or 2, characterized in that The third image includes filling data, and a value of the filling data is a preset value; Then, generating a third image according to the second image and the mask information includes: When the sum of the data volume of the second image and the data volume of the mask information is less than the data volume of the first image, the second image, the mask information and the padding data are spliced ​​to obtain the third image.

4. The method according to any one of claims 1 to 3, characterized in that The mask information corresponding to the first image includes a mask map; The size of the mask image is the first size, or the size of the mask image is smaller than the first size.

5. The method according to any one of claims 1 to 4, characterized in that The downsampling the first image includes: Downsampling the first image in a single direction dimension to obtain the second image; The width of the second image is equal to that of the first image, or the height of the second image is equal to that of the first image.

6. A video processing method, characterized in that: The method comprises: Acquire a video stream to be processed, where the video stream includes a third image, the third image includes mask information, and the size of the third image is the first size; Decoding the video code stream to obtain the third image; Splitting the third image to obtain a second image and the mask information; Upsampling the second image to obtain a fourth image, where the size of the fourth image is the first size; A preset video processing operation is performed according to the fourth image and the mask information.

7. A video processing device, characterized in that: The video processing device comprises: An acquisition module, configured to acquire a first image in a video to be processed and mask information corresponding to the first image, wherein the size of the first image is a first size; a downsampling module, configured to downsample the first image to obtain a second image; an image generation module, configured to generate a third image according to the second image and the mask information, wherein the size of the third image is the first size; The encoding module is used to encode the third image based on the third image to obtain a video stream corresponding to the third image.

8. The device according to claim 7, characterized in that The size of the second image is a second size, and the second size is smaller than the first size; The downsampling module is used to: determining the second size according to the data amount of the mask information corresponding to the first image; The first image is downsampled according to the second size to obtain the second image, and the sum of the data volume of the second image of the second size and the data volume of the mask information does not exceed the data volume of the first image of the first size.

9. The device according to claim 7 or 8, characterized in that The third image includes filling data, and a value of the filling data is a preset value; Then, the image generation module is used to splice the second image, the mask information and the filling data to obtain the third image when the sum of the data volume of the second image and the data volume of the mask information is less than the data volume of the first image.

10. The device according to any one of claims 7 to 9, characterized in that The mask information corresponding to the first image includes a mask map; The size of the mask image is the first size, or the size of the mask image is smaller than the first size.

11. The device according to any one of claims 7 to 10, characterized in that The downsampling module is used to: Downsampling the first image in a single direction dimension to obtain the second image; The width of the second image is equal to that of the first image, or the height of the second image is equal to that of the first image.

12. A video processing device, characterized in that: The device comprises: an acquisition module, configured to acquire a video stream to be processed, wherein the video stream includes a third image, the third image includes mask information, and the size of the third image is the first size; A decoding module, configured to decode the video code stream to obtain the third image; a splitting module, configured to split the third image to obtain the second image and the mask information; an upsampling module, configured to upsample the second image to obtain a fourth image, where the size of the fourth image is the first size; An operation execution module is configured to execute a preset video processing operation according to the fourth image and the mask information.

13. A computing device, characterized in that The computing device includes a processor and a memory; The processor is configured to execute instructions stored in the memory, so as to enable the computing device to perform the steps of the method according to any one of claims 1 to 6.

14. A computer-readable storage medium, characterized in that The method comprises instructions which, when executed on a computing device, cause the computing device to perform the steps of the method according to any one of claims 1 to 6.

15. A computer program product comprising instructions, characterized in that When the method is executed on at least one computing device, the at least one computing device is enabled to perform the steps of the method according to any one of claims 1 to 6.