Video processing method, device and system, terminal equipment and readable storage medium

By segmenting the video image into target objects and background, encoding only the target object image, and using the uncompressed complete background image as the background for the reconstructed image, the problem of reducing the video bitrate while ensuring video quality is solved, achieving the effect of significantly reducing encoding resources and bandwidth consumption.

CN120676200APending Publication Date: 2025-09-19HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410317257.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

How to reduce the bit rate of the video stream after encoding while ensuring video quality, so as to reduce storage and transmission costs and improve the user's playback experience under poor network conditions.

Method used

By segmenting the target object and background of the video image, only the target object image is encoded, and the uncompressed complete background image is used as the background of the reconstructed image, and the image code stream of the target object image and the complete background image are sent to the terminal device for display.

Benefits of technology

It achieves the goal of significantly reducing the video bit rate while ensuring the video quality, reducing encoding resources and bandwidth consumption, while maintaining high video definition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676200A_ABST
    Figure CN120676200A_ABST
Patent Text Reader

Abstract

The invention provides a video processing method, device and system, terminal equipment and a readable storage medium, and aims to reduce the code rate of a code stream of a video by only encoding a target object image segmented from an image in the video; when a reconstructed image of the image is determined, a complete background image without video compression is adopted as a background image, and the definition of a background part in the reconstructed image is ensured; therefore, the code rate of the code stream of the video is reduced on the basis of ensuring the image quality of the video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video technology, and in particular to a video processing method, apparatus, system, terminal device, and readable storage medium. Background Art

[0002] Video encoding, also known as video compression, refers to converting video signals into digital signals and using compression algorithms to reduce the storage space and transmission bandwidth required for the video. Video decoding, on the other hand, converts compressed video data back into a playable video stream. Video encoding and decoding are crucial steps in the video processing process. During video processing, the bitrate of the encoded video stream is carefully controlled to balance video quality and file size. If the bitrate of the encoded video stream is too low, video quality will degrade, affecting the viewing experience. If the bitrate is too high, the video file will be too large, increasing storage and transmission costs. For example, in live streaming scenarios such as conferencing, online education, video chat, and surveillance, a lower bitrate, while maintaining consistent image quality, reduces costs for operators and improves the user experience even in poor network conditions.

[0003] Therefore, how to reduce the bit rate of the video stream after encoding while ensuring the video quality has become a technical problem that needs to be solved urgently. Summary of the Invention

[0004] The present application provides a video processing method, apparatus, system, terminal device and readable storage medium, which reduces the bit rate of the video stream by encoding only the target object image segmented from the image in the video; when determining the reconstructed image of the image, a complete background image without video compression is used as the background image to ensure the clarity of the background part in the reconstructed image; thereby achieving the reduction of the video stream bit rate while ensuring the video quality.

[0005] In a first aspect, a video processing method is provided, which is applied to a first terminal device, and includes: acquiring a first image of a video, where the first image is a frame of image in the video; segmenting the first image into a target object and a background to obtain image segmentation information, where the image segmentation information includes a first target object image and a first background image, where the first target object image is an image of the target object in the first image, and the first background image is an image other than the first target object image in the first image; when it is determined that there is a first complete background image, encoding the first target object image to obtain an image code stream corresponding to the first target object image, where the first complete background image is a complete background image corresponding to the first background image, and the complete background image is a complete image of the background where the target object is located; sending the first complete background image and the image code stream corresponding to the first target object image to a second terminal device, where the image code stream corresponding to the first target object image and the first complete background image are used by the second terminal device to determine and display a reconstructed image of the first image, where the background image of the reconstructed image of the first image is the first complete background image.

[0006] The video processing method provided in the first aspect determines a first complete background image and encodes a first target object image obtained by segmenting the first image to obtain an image code stream of the first target object image; the image code streams of the first complete background image and the first target object image are sent to a second terminal device, so that the second terminal device determines and displays a reconstructed image of the first image based on the image code streams corresponding to the first complete background image and the first target object image, wherein the background image of the reconstructed image of the first image is the first complete background image. Since the first terminal device only needs to encode a portion of the image in the first image (i.e., the first target object image), the bit rate of the image code stream of the obtained first target object image is much lower than the bit rate of the image code stream of the complete first image. For the entire video, when most of the image frames in the video are encoded by only encoding a portion of the image, the bit rate of the code stream corresponding to the entire video will be significantly reduced. Therefore, this method can effectively reduce the bit rate of the video code stream. In addition, because the first complete background image has not been compressed, it can maintain a high degree of clarity. Using the first complete background image as the background image of the reconstructed image of the first image allows the background of the displayed reconstructed image of the first image to have a high degree of clarity, thereby ensuring clear video quality. In summary, the video processing method of this application can reduce the video bit rate while ensuring video quality. In addition, since the video bit rate is reduced, the consumption of encoding resources and bandwidth resources of the first terminal device is also significantly reduced.

[0007] Exemplarily, the video processing method may further include: if it is determined that the first complete background image does not exist, encoding the first image to obtain an image stream corresponding to the first image; and sending the image stream corresponding to the first image to a second terminal device, where the image stream corresponding to the first image is used by the second terminal device to determine and display a reconstructed image of the first image. In this implementation, if it is determined that the first complete background image does not exist, it means that a complete background image has not yet been generated, or that the first background image segmented from the first image has significantly changed from an existing complete background image, and therefore the existing complete background image cannot be used as the first complete background image. In this case, the first terminal device encodes the complete first image to obtain an image stream corresponding to the first image, so that the second terminal device can determine and display the reconstructed image of the first image based on the image stream corresponding to the first image, thereby ensuring that the second terminal can display images that do not have corresponding complete background images. This design allows the second terminal device to display reconstructed images of both images in the video that have corresponding complete background images and images that do not have corresponding complete background images, thereby ensuring the continuity of the video display on the second terminal device.

[0008] In a possible implementation of the first aspect, each frame image in a video corresponds to a time point, with different frames corresponding to different time points. Each complete background image is obtained based on at least two background images. The at least two background images corresponding to the first complete background image include the first background image. The second background image is any background image other than the first background image from the at least two background images corresponding to the first complete background image. The second background image and the second target object image are obtained by segmenting the second image in the video. The time point corresponding to the second image is before the time point corresponding to the first image. The position of the first target object image in the first image is different from the position of the second target object image in the second image. In this implementation, each complete background image is obtained based on at least two background images. The first background image is the last background image of the at least two background images used to form the first complete background image. That is, after determining the first complete background image based on the at least two background images, the portion of the first image corresponding to the last background image of the at least two background images (i.e., the first target object image) is encoded, ensuring that as many images in the video as possible are encoded using a method of encoding only partial images, thereby minimizing the bitrate of the video stream. In addition, since the position of the first target object image in the first image is different from the position of the second target object image in the second image, the position of the area from which the first target object image is segmented in the first background image is different from the position of the area from which the second target object image is segmented in the second background image. Therefore, the first complete background image can be obtained by complementing the first background image and the second background image. The method is simple and easy to implement.

[0009] In a possible implementation of the first aspect, each frame image in a video corresponds to a time point, and different frames of image correspond to different time points. Each complete background image is obtained based on at least two background images. The second background image and the third background image are any two background images from the at least two background images corresponding to the first complete background image. The second background image and the second target object image are obtained by segmenting the second image in the video. The third background image and the third target object image are obtained by segmenting the third image in the video. The time point corresponding to the second image and the time point corresponding to the third image are both before the time point corresponding to the first image. The position of the second target object image in the second image is different from the position of the third target object image in the third image. In this implementation, the second background image and the third background image are any two background images from the at least two background images corresponding to the first complete background image. Since the position of the second target object image in the second image is different from the position of the third target object image in the third image, the region from which the second target object image is segmented in the second background image is different from the region from which the third target object image is segmented in the third background image. Therefore, the first complete background image can be obtained by complementing the second and third background images. The method is simple and easy to implement.

[0010] In a possible implementation of the first aspect, in at least two background images corresponding to the first complete background image, the time interval between any two background images is greater than the time interval between any two adjacent frame images in the video. Generally speaking, the difference between two consecutive frame images in a video is relatively small. In this implementation, the time interval between any two background images used to form the first complete background image is greater than the time interval between any two adjacent frame images in the video, which can avoid the large amount of data processing caused by frequently acquiring background images; for example, when the position of the target object in the background changes little, frequently acquiring background images increases the amount of data processing but cannot obtain the first complete background image in a short time. When acquiring background images used to form a complete background image in this implementation, the number of background images used can be controlled, and segmentation of the target object and the background on too many images can be avoided, thereby reducing the cost of calculation and reducing computing power consumption.

[0011] Exemplarily, among the at least two background images corresponding to the first complete background image, the time interval between any two adjacent background images is a first preset duration. In this implementation, a background image for forming the first complete background image is obtained at intervals of the first preset duration.

[0012] In a possible implementation of the first aspect, the method further includes: determining whether there is an overlapping portion between a blank area of ​​the i-th background image and a blank area of ​​the first stitched image, the first stitched image is obtained by stitching together the first i-1 background images, the i-th background image is the i-th background image among the background images used to form the first complete background image, the i-th background image is obtained by segmenting a target object and a background in a frame image before the first image in the video, the blank area of ​​the i-th background image is the area after the target object image is segmented out of the i-th background image, the blank area of ​​the first stitched image is the blank area after the first i-1 background images are stitched together, and i is an integer greater than or equal to 2; if there is no overlapping portion between the blank area of ​​the i-th background image and the blank area of ​​the first stitched image, stitching the image portion of the i-th background image corresponding to the blank area of ​​the first stitched image into the blank area of ​​the first stitched image, or stitching the image portion of the first stitched image corresponding to the blank area of ​​the i-th background image into the blank area of ​​the i-th background image, to obtain the first complete background image. In this implementation, there is no overlap between the blank areas of the i-th background image and the blank areas of the first stitched image, indicating that the blank areas of the i-th background image and the blank areas of the first stitched image are located in completely different locations. Therefore, the first complete background image can be obtained by splicing the image portion of the i-th background image corresponding to the blank area of ​​the first stitched image into the blank area of ​​the first stitched image, or by splicing the image portion of the first stitched image corresponding to the blank area of ​​the i-th background image into the blank area of ​​the i-th background image. Obtaining the first complete background image by filling the blank areas consumes less computation and power, facilitating implementation of the method.

[0013] In a possible implementation of the first aspect, the method further includes: if there is an overlapping portion between the blank area of ​​the i-th background image and the blank area of ​​the first stitched image, stitching the image portion of the i-th background image corresponding to the blank area of ​​the first stitched image into the blank area of ​​the first stitched image, or stitching the image portion of the first stitched image corresponding to the blank area of ​​the i-th background image into the blank area of ​​the i-th background image, to obtain a second stitched image; if there is no overlapping portion between the blank area of ​​the i+1-th background image and the blank area of ​​the second stitched image, stitching the image portion of the i-th background image corresponding to the blank area of ​​the second stitched image into the blank area of ​​the second stitched image, or stitching the image portion of the second stitched image corresponding to the blank area of ​​the i+1-th background image into the blank area of ​​the i+1-th background image, to obtain a first complete background image. In this implementation, there is an overlapping part between the blank area of ​​the i-th background image and the blank area of ​​the first stitched image, indicating that the blank area of ​​the i-th background image and the blank area of ​​the first stitched image have the same part at their locations. Therefore, the image part corresponding to the blank area of ​​the first stitched image in the i-th background image is spliced ​​into the blank area of ​​the first stitched image, or the image part corresponding to the blank area of ​​the i-th background image in the first stitched image is spliced ​​into the blank area of ​​the i-th background image. What is obtained is a second stitched image that still has a blank area. Therefore, it is still necessary to compare the blank areas of the second stitched image with those of the i+1-th background image until the first complete background image can be spliced.

[0014] In one possible implementation of the first aspect, when i is equal to 2, the first stitched image is the first background image among the background images used to form the first complete background image. In this implementation, i being equal to 2 indicates that the first complete background image is obtained based on two background images. In this case, the first stitched image is the first background image among the two background images, and the first background image is the background image corresponding to the earlier time point of the two background images.

[0015] In a possible implementation of the first aspect, a fourth target object image and a fourth background image are obtained by segmenting a fourth image in a video. The first complete background image is obtained based on at least two background images. The fourth background image is the background image corresponding to the latest time point of the at least two background images corresponding to the first complete background image. The image stream corresponding to the fourth target object image includes a uniform resource locator (URL) of the first complete background image on a remote server. The URL is used by a second terminal device to obtain the first complete background image. In this implementation, the first complete background image is stored on a remote server, and the image stream corresponding to the fourth target object image carries the URL of the first complete background image on the remote server. After obtaining the image stream corresponding to the fourth target object image, the second terminal device downloads the first complete background image based on the URL. The first terminal device sends the first complete background image to the second terminal device via the URL. This transmission method reduces bandwidth resource consumption. The first URL is carried in the image stream, eliminating the need for the first terminal device to separately send the URL to the second terminal device, further reducing bandwidth resource consumption.

[0016] It's understandable that using an image stream to carry the uniform locator of the first complete background image on a remote server is particularly suitable for live broadcast scenarios with a single host and multiple players. In this scenario, each player can obtain the image stream from the host or remote server by pulling the stream, then decode the image stream to obtain the uniform resource locator (URL), and then download the first complete background image based on the URL. This approach facilitates the player's access to the first complete image. For example, the host or remote server is the first terminal device, and the player is the second terminal device.

[0017] In one possible implementation of the first aspect, the first terminal device may directly send the first complete background image to the second terminal device; alternatively, the first terminal device may directly send the unified locator of the first complete background image on a remote server to the second terminal device. This implementation is primarily applicable to point-to-point communication scenarios, such as video calls, home monitoring, and video sharing.

[0018] In one possible implementation of the first aspect, the uniform resource locator (URL) of the first complete background image on the remote server is located in supplemental enhancement information (SEI) of the image stream corresponding to the fourth target object image. In this implementation, the URL is located in the supplemental enhancement information (SEI), which meets standard requirements and facilitates implementation.

[0019] In one possible implementation of the first aspect, the image segmentation information further includes: position information of the first target object image within the first image, and the image code stream corresponding to the first target object image includes: position information of the first target object image within the first image. In this implementation, the image code stream corresponding to the first target object image carries encoded data of the position information of the first target object image within the first image, thereby transmitting the position information of the first target object image within the first image to a second terminal device. The second terminal device can use this position information when determining and displaying a reconstructed image of the first image, thereby ensuring that the position of the target object in the reconstructed image of the first image is consistent with the position of the target object in the first image.

[0020] In one possible implementation of the first aspect, the supplemental enhancement information of the image codestream corresponding to the first target object image includes location information of the first target object image within the first image. In this implementation, the uniform resource locator (URI) is located in the supplemental enhancement information (SII), which meets standard requirements and facilitates implementation and widespread application.

[0021] In one possible implementation of the first aspect, before encoding the first target object image, the method further includes: adjusting the first target object image to a preset target size. In this implementation, adjusting the target object image to be encoded to the target size normalizes the target object image size, thereby meeting encoding requirements of the first terminal device and decoding requirements of the second terminal device, and improving data processing efficiency and stability.

[0022] It is understood that since the size of the target object images in different frames of a video may vary, converting the target object images to a preset target size ensures that all encoded target object images have the same size (or resolution), ensuring normal encoding of the target object images and avoiding the high bitrate caused by the presence of a large number of I-frames after encoding. If the target object images have different sizes, then all the target object images will be I-frames after encoding, and the bitrate of I-frames is relatively high, resulting in a high bitrate of the final bitstream.

[0023] In one possible implementation of the first aspect, a sequence parameter set of an image code stream corresponding to the first target object image includes a preset target size. In this implementation, the preset target size is carried in the sequence parameter set, which meets standard requirements and is easy to implement.

[0024] In a possible implementation of the first aspect, each frame image in the video corresponds to a time point, and different frames of image correspond to different time points. The method further includes: determining a first degree of change of the first background image relative to the second complete background image, where the second complete background image is the complete background image closest to the time point corresponding to the first image at the moment of determination; if the first degree of change is less than or equal to a preset threshold, then determining the second complete background image as the first complete background image. In this implementation, when the degree of change of the first background image relative to the most recently formed complete background image (i.e., the second complete background image) is less than a preset threshold, the second complete background image is determined as the first complete background image to ensure that the first complete background image and the first background image have a high degree of similarity.

[0025] In a possible implementation of the first aspect, the degree of change of the first background image relative to the second complete background image is determined at a first moment, and the method further includes: determining a second degree of change of the fifth background image relative to the second complete background image at a second moment, and if the second degree of change is less than or equal to a preset threshold, determining the second complete background image as the complete background image corresponding to the fifth complete background image, where the fifth background image is obtained by segmenting the fifth image in the video; the time interval between the first moment and the second moment is greater than the time interval between any two adjacent frame images in the video. In this implementation, the time point corresponding to the fifth background image is after the time point corresponding to the first background image, and there is at least one background image between the first background image and the fifth background image, that is, in this embodiment, the background image of each frame image in the video will not be compared with the complete background image, thereby reducing the amount of data processing and improving data processing efficiency.

[0026] For example, the time interval between the first moment and the second moment may be a second preset duration, that is, a comparison of the degree of change of a background image and a second complete background image is performed once every second preset duration.

[0027] In a second aspect, a video processing method is provided, which is applied to a second terminal device, and includes: receiving a first complete background image and an image code stream corresponding to a first target object image; decoding the image code stream corresponding to the first target object image to obtain a reconstructed image of the first target object image; determining and displaying the reconstructed image of the first image based on the first complete background image and the reconstructed image of the first target object image, wherein the first image is an image from which the first target object image is segmented, and the background image of the reconstructed image of the first image is the first complete background image.

[0028] In the video processing method provided in the second aspect, the image code stream received by the second terminal device is the image code stream of the local image (i.e., the first target image) in the first image. The bit rate of the image code stream is much smaller than the bit rate of the image code stream of the complete first image. This can effectively reduce the bit rate of the video code stream. The reduction in the bit rate of the video code stream can significantly reduce the computing power consumption of the second terminal device; in addition, the first complete background image has not undergone video compression, and can maintain a high clarity. The first complete background image is used as the background image of the reconstructed image of the first image, so that the reconstructed image of the first image has clear picture quality.

[0029] In one possible implementation of the second aspect, receiving the first complete background image includes: receiving an image stream corresponding to a fourth target object image, the image stream corresponding to the fourth target object image including a uniform resource locator (URL) for the first complete background image on a remote server; and obtaining the first complete background image based on the URL for the first complete background image on the remote server. In this implementation, obtaining the first complete background image based on the URL for the first complete background image on the remote server involves downloading the first complete background image from the remote server based on the URL. Obtaining the first complete background image using the URL in the received image stream corresponding to the fourth target object image allows the second terminal device to quickly and easily obtain the first complete background image.

[0030] In one possible implementation of the second aspect, the uniform resource locator (URL) of the first complete background image on the remote server is located in supplemental enhancement information (SEMI) of the image stream corresponding to the fourth target object image. In this implementation, the URL is located in the supplemental enhancement information (SEMI), which meets standard requirements and facilitates implementation.

[0031] In one possible implementation of the second aspect, the image code stream corresponding to the first target object image is decoded to obtain position information of the first target object image within the first image; the position of the reconstructed image of the first target object image within the first image is determined based on the position information of the first target object image within the first image. In this implementation, the position information of the first target object image within the first image is taken into account when determining the reconstructed image of the first image, thereby ensuring that the reconstructed image of the first image has a high degree of fidelity to the first image.

[0032] In one possible implementation of the second aspect, before displaying the reconstructed image of the first image based on the first complete background image and the reconstructed image of the first target object image, the method further includes: determining the original size of the first target object image, where the original size is the size of the first target object image in the first image; and adjusting the reconstructed image of the first target object image to the original size. In this implementation, before superimposing the reconstructed image of the first target object image on the first complete background image, the reconstructed image of the first target object image is adjusted to the original size, thereby ensuring that the reconstructed image of the first target object image is the same size as the first target object image, further ensuring that the reconstructed image of the first image has a high degree of fidelity to the first image.

[0033] In a possible implementation of the second aspect, a reconstructed image of the first image is displayed based on the first complete background image and the reconstructed image of the first target object image, including: rendering the first complete background image to the display area of ​​the video; and superimposing the first target object image on the first complete background image to display the reconstructed image of the first image.

[0034] In a third aspect, a video processing device is provided, which includes units for performing each step of the method in the above first aspect or any possible implementation of the first aspect.

[0035] In a fourth aspect, a video processing device is provided, which includes units for each step of the method in the above second aspect or any possible implementation of the second aspect.

[0036] In a fifth aspect, a communication device is provided, which includes units for each step of the method in any one of the above aspects or any possible implementation of any one of the aspects.

[0037] In a sixth aspect, a communication device is provided, which includes at least one processor and a memory, the processor and the memory are coupled, the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method in any one of the above aspects or any possible implementation of any one of the aspects is executed.

[0038] In a seventh aspect, a communication device is provided, which includes at least one processor and an interface circuit, and the at least one processor is used to execute: the method in any one of the above aspects or any possible implementation of any one of the aspects.

[0039] In an eighth aspect, a terminal device is provided, comprising a processor and a memory, the memory being used to store instructions, the processor being used to read the instructions to execute the method in the above first aspect or any possible implementation of any first aspect, or the processor being used to read the instructions to execute the method in the above second aspect or any possible implementation of any second aspect.

[0040] In the ninth aspect, a video processing system is provided, which includes a first terminal device and a second terminal device that are communicatively connected, the first terminal device being used to execute the method in the above first aspect or any possible implementation of any first aspect, the second terminal device being used to execute the method in the above second aspect or any possible implementation of any second aspect, the first terminal device sending the first complete background image and the image code stream corresponding to the first target object image to the second terminal device.

[0041] In a possible implementation of the ninth aspect, the first terminal device is a remote server, or the system also includes a remote server; the first complete background image is stored in the remote server, and the second terminal device obtains the first complete background image from the remote server through the uniform resource locator of the first complete background image in the remote server.

[0042] In a tenth aspect, a computer program product is provided, which includes a computer program, and when the computer program is executed by a processor, it is used to execute the method in any one of the above aspects or any possible implementation of any one of the aspects.

[0043] In the eleventh aspect, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed, it is used to execute the method in any one of the above aspects or any possible implementation of any one of the aspects.

[0044] In the twelfth aspect, a chip is provided, which includes: a processor for calling and running a computer program from a memory, so that a terminal device equipped with the chip executes a method in any one of the above aspects or any possible implementation of any one of the aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a schematic diagram of a video processing system provided in one embodiment of the present application.

[0046] Figure 2 is a schematic diagram of a video processing system provided in another embodiment of the present application.

[0047] Figure 3is a schematic diagram of a video processing system provided in another embodiment of the present application.

[0048] Figure 4 This is a flow chart of an example provided in an embodiment of the present application.

[0049] Figure 5 This is a schematic diagram of image segmentation in a video processing method provided in an embodiment of the present application.

[0050] Figure 6 This is a schematic diagram of a complete background image in a video processing method provided in an embodiment of the present application, as well as a schematic diagram of the image overlay process.

[0051] Figure 7 This is a flow chart of an example provided in an embodiment of the present application.

[0052] Figure 8 This is a schematic diagram of a process of determining a complete background image in a video processing method provided in an embodiment of the present application.

[0053] Figure 9 This is a schematic diagram of determining whether there are overlapping parts in a blank area in a video processing method provided in one embodiment of the present application.

[0054] Figure 10 This is a flowchart of determining whether a first complete background image exists in a video processing method provided in an embodiment of the present application.

[0055] Figure 11 This is a schematic diagram of the application of a video processing method provided in an embodiment of the present application in a video processing process.

[0056] Figure 12 This is a flow chart of an example provided in an embodiment of the present application.

[0057] Figure 13 This is a hardware structure block diagram of an example of a terminal device provided in this application.

[0058] Figure 14 This is a schematic diagram of a chip system provided in this application. DETAILED DESCRIPTION

[0059] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.

[0060] The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to be limiting of the present application. As used in the specification and appended claims of this application, the singular expressions "one", "a kind of", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless there is a clear contrary indication in the context. It should also be understood that in the embodiments of the present application, "one or more" refers to one or more (including two); "and / or" describes the association relationship of associated objects, indicating that three relationships may exist; for example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship.

[0061] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0062] The "multiple" involved in the embodiments of the present application means greater than or equal to two. It should be noted that in the description of the embodiments of the present application, the words "first" and "second" are only used for the purpose of distinguishing the description and cannot be understood as indicating or implying relative importance or order.

[0063] In addition, various aspects or features of the present application can be implemented as methods, devices or products using standard programming and / or engineering technology. The term "product" used in the present application embodiment covers a computer program that can be accessed from any computer-readable device, carrier or medium. For example, computer-readable media can include, but are not limited to: magnetic storage devices (for example, hard disks, floppy disks or magnetic tapes, etc.), optical disks (for example, compact disks (compact disks, CDs), digital versatile disks (digital versatile disks, DVDs), etc.), smart cards and flash memory devices (for example, erasable programmable read-only memory (EPROM), cards, sticks or key drives, etc.). In addition, various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" can include, but is not limited to, wireless channels and various other media that can store, contain and / or carry instructions and / or data.

[0064] Video is an audiovisual media format that presents dynamic scenes through a series of static images. Its multimedia presentation is more engaging and appealing than simple text or images, and therefore plays an irreplaceable role in people's daily lives, learning, and work. For example, live streaming, a form of real-time online video communication, has rapidly gained popularity worldwide in recent years. Live streaming includes online conferencing, online education, video chat, e-commerce live streaming, and surveillance.

[0065] To provide users with a superior video experience, videos generally undergo a series of processing steps, including video data acquisition, generation, editing, compression (encoding), and transmission to playback (decoding). Video encoding and decoding are crucial steps in this process. Video encoding converts video signals into digital signals and uses compression algorithms to reduce the storage space and transmission bandwidth required. Video decoding restores the compressed video data into a playable stream. In the video field, bitrate, also known as bitrate, refers to the amount of video data per unit time. Kbps (thousand bits per second) is a common unit, meaning thousands of bits of data are used per second. During video processing, the bitrate of the encoded video stream is optimally controlled to balance video quality and file size. If the bitrate of the encoded video stream is too low, video quality will degrade, affecting viewing quality. If the bitrate is too high, the video file will be too large, increasing storage and transmission costs. For example, in live streaming scenarios like conferencing, online education, video chat, and surveillance, the lower the bitrate of the encoded video stream while maintaining image quality, the lower the operator's costs. Furthermore, even in poor network conditions, the user experience is better. Therefore, how to reduce the bitrate of the encoded video stream while maintaining image quality has become a key issue in the field of video technology.

[0066] In view of this, the present application provides a video processing method, wherein a first terminal device obtains a target object image and a background image by segmenting an image in a video (i.e., a frame of image in the video); determines a complete background image corresponding to the image; encodes the target object image to obtain an image code stream corresponding to the target object image; and sends the image code stream corresponding to the target object image and the complete background image to a second terminal device, so that the second terminal device determines and displays a reconstructed image of the image based on the complete background image and the image code stream corresponding to the target object image, wherein the background image of the reconstructed image of the image is the complete background image. In the video processing method of the present application, since the first terminal device only needs to encode a portion of the image (i.e., the target object image) in the image, for example, assuming that the target object image occupies a quarter of the area of ​​the corresponding complete image, the number of pixels actually encoded is only a quarter of the number of pixels of the entire image. The smaller the number of encoded pixels, the smaller the bit rate of the image code stream formed. Therefore, the bit rate of the image code stream corresponding to the target object image is much smaller than the bit rate of the image code stream of the complete image. For the entire video, when most of the image frames in the video are encoded using a partial image encoding method, the bit rate of the code stream corresponding to the entire video will be significantly reduced. And because the complete background image has not been compressed, the complete background image can maintain a high degree of clarity. The second terminal uses the complete background image as the background image of the reconstructed image, so that the reconstructed image can maintain a high degree of clarity. In summary, the video processing method in this application can reduce the bit rate of the video while ensuring the video quality.

[0067] The video processing method provided in this application will be exemplarily described below. Those skilled in the art will appreciate that the following content is merely an example and is not intended to limit the scope of protection of this application.

[0068] It should be noted that in the embodiments of the present application, "terminal device", "electronic device" and "terminal" all have the same meaning and can be interchanged.

[0069] For ease of understanding, the application scenario of the video processing method in the embodiment of the present application is first described below.

[0070] refer to Figure 1 , Figure 1 The video processing method in the embodiment of the present application can be applied to the following examples: Figure 1 The video processing system shown.

[0071] like Figure 1 As shown, in Figure 1In the illustrated embodiment, the video processing system includes a live broadcast end 110, a remote server 120, and a playback end 130. The live broadcast end 110, the remote server 120, and the playback end 130 can all be terminal devices. In addition, in this application, the terminal device can also be referred to as an electronic device.

[0072] It is understandable that in the video processing system, the number of live broadcast terminals 110 can be one or more, the number of remote servers 120 can be one or more, and the number of playback terminals 130 can also be one or more. This application does not impose any restrictions on this.

[0073] It should be understood that Figure 1 The illustrated video processing system is used in live streaming scenarios, such as online conferencing, online education, video chat, e-commerce live streaming, and live show streaming. Taking e-commerce live streaming as an example, the live streaming end 110 is the device used by the host to broadcast live, responsible for collecting, processing, and transmitting audio and video data. It typically includes devices such as a camera, microphone, and corresponding audio and video processing software. The playback end 130 is the device used by users to watch the live broadcast, such as a smartphone, computer, or smart TV. Users can receive and play the live content on the playback end 130.

[0074] Exemplarily, the playback end 130 can be a terminal device such as a mobile phone, laptop computer, desktop PC, mobile tablet, car tablet, etc., which serves as a human-computer interaction interface for users to use live broadcast applications. Users interact with the terminal device by clicking, touching, pressing buttons, etc. to watch live broadcasts, video conferences, and other functions.

[0075] In some embodiments, based on Figure 1 In the video processing system shown in FIG. 1 , the first terminal device in the embodiment of the video processing method of the present application may be the remote server 120, and the second terminal device may be the playback terminal 130. The following is an exemplary description of the composition and function of the live broadcast terminal 110, the remote server 120, and the playback terminal 130 in this embodiment:

[0076] Exemplarily, the live broadcast end 110 may include: a terminal device including an acquisition module and an encoding and streaming module, wherein: the acquisition module is used for video acquisition, for example, the acquisition module may include a camera in the terminal device, and the camera acquires images according to a set frame rate and resolution to form original video data; the encoding and streaming module encodes the image sequence in the video data in the order of acquisition time to obtain a live stream, and streams the live stream to the remote server 120.

[0077] Exemplarily, the remote server 120 may include: an image segmentation module, an encoding module and a distribution module, wherein: the image segmentation module is used to decode the received live stream to obtain an image sequence corresponding to the live stream, the image segmentation module can also segment the image in the image sequence into a target object and a background to obtain a target object image, and the image segmentation module can also be used to generate a complete background image; the encoding module is used to encode the segmented target object image to obtain an image stream of the target object image; the distribution module is used to distribute the image stream of the target object image to the playback end 130.

[0078] In addition, the remote server 120 can also be used to transmit the complete background image to the playback terminal 130. For example, the remote server 120 can distribute the complete background image to the playback terminal 130 through a distribution module, or the image code stream of the target object image can carry the URL (Uniform Resource Locator) of the complete background image on the remote server 120. The playback terminal 130 decodes the image code stream to obtain the URL and downloads the complete background image from the remote server 120 based on the URL. This application does not limit the specific method by which the remote server 120 transmits the complete background image to the playback terminal 130.

[0079] For example, the remote server 120 may provide the playback terminal 130 with an image stream of the target object image and a complete background image via a network request stream interface. Alternatively, the remote server 120 may provide the playback terminal 130 with an image stream encoding the complete image in the image sequence, which is not limited in this application.

[0080] It is understandable that the remote server 120 can send the image stream to the playback terminal 130 via a CDN (Content Delivery Network).

[0081] Exemplarily, the playback end 130 can be a terminal device including a stream receiving and decoding module and a rendering module, wherein: the stream receiving and decoding module is used to obtain the image code stream of the complete background image and the target object image, and decode the obtained target object image code stream frame by frame in chronological order to obtain an image sequence of the target object image; the rendering module is used to restore the image sequence in the video, including rendering the complete background image, and overlaying the reconstructed image of the decoded target object image on the complete background image.

[0082] In other embodiments, based on Figure 1In the video processing system shown, the first terminal device in the embodiment of the video processing method of the present application can be the live broadcast terminal 110, and the second terminal device can be the playback terminal 130. The following is an exemplary description of the composition and function of the live broadcast terminal 110, the remote server 120 and the playback terminal 130 in this embodiment:

[0083] Exemplarily, the live broadcast end 110 may include: a terminal device including an acquisition module, an image segmentation module, an encoding module and a streaming module, wherein: the acquisition module is used for video acquisition, for example, the acquisition module may include a camera in the terminal device, and the camera acquires images according to a set frame rate and resolution to form original video data; the image segmentation module is used to segment the images in the image sequence in the original video data into target objects and backgrounds to obtain target object images, and the image segmentation module can also be used to generate a complete background image and send the complete background image to the remote server 120; the encoding module is used to encode the segmented target object image to obtain an image code stream of the target object image; the streaming module is used to stream the image code stream of the target object image to the remote server 120; the streaming module can also be used to send the complete background image to the remote server 120.

[0084] Exemplarily, the remote server 120 includes a distribution module configured to distribute the received image stream of the target object image to the playback terminal 130. Furthermore, the remote server 120 may also be configured to transmit the complete background image to the playback terminal 130. The specific manner in which the remote server 120 transmits the complete background image to the playback terminal 130 can be found in the foregoing description and is not further elaborated here.

[0085] Exemplarily, the playback terminal 130 may be a terminal device including a stream receiving and decoding module and a rendering module. The specific functions of the stream receiving and decoding module and the rendering module in the playback terminal 130 may be referred to as described above and will not be elaborated here.

[0086] In some embodiments, since both the live broadcast end 110 and the playback end 130 can be terminal devices, the terminal device serving as the live broadcast end 110 can also serve as the playback end, and the terminal device serving as the playback end 130 can also serve as the live broadcast end. This application does not impose any restrictions on this.

[0087] refer to Figure 2 , Figure 2 FIG. 1 is a schematic diagram of a video processing system provided in another embodiment of the present application. The video processing method in the embodiment of the present application can be applied to the following examples: Figure 2 The video processing system shown.

[0088] like Figure 2 As shown, in Figure 2In the illustrated embodiment, the video processing system includes a remote server 210 and a playback terminal 220. Both the remote server 210 and the playback terminal 220 may be terminal devices, which may also be referred to as electronic devices in this application.

[0089] It is understandable that in the video processing system, the number of remote servers 210 can be one or more, and the number of playback terminals 220 can also be one or more, and this application does not impose any restrictions on this.

[0090] It should be understood that Figure 2 The video processing system shown is used in a recording and broadcasting scenario. This scenario involves pre-recording audio and video content and then playing it back when needed. For example, recording and broadcasting video can be used in scenarios such as online education, corporate training, product demonstrations, and event recording.

[0091] It is understood that the remote server 210 stores a video file to be recorded and played, and the remote server 210 is in communication with the playback terminal 220. The interaction process between the remote server 210 and the playback terminal 220 may include: the playback terminal 220 sends a request to the remote server 210 to obtain the video file; the remote server 210 provides the corresponding video file to the playback terminal 220 according to the request of the playback terminal 220; the playback terminal 220 decodes the received video file and then plays (or displays) it.

[0092] It should be understood that the video file provided by the remote server 210 to the playback end is an encoded file, specifically, it can be encoded using the processing method in this application; the playback end 220 can use the video processing method in this application for decoding and display.

[0093] In some embodiments, based on Figure 2 In the video processing system shown, the first terminal device in the embodiment of the video processing method of the present application can be a remote server 210, and the second terminal device can be a playback terminal 220. The composition and function of the remote server 210 and the playback terminal 220 are described below as follows:

[0094] Exemplarily, the remote server 210 may include: an image segmentation module, an encoding module and a distribution module, wherein: the image segmentation module is used to segment the target object and background of the image in the image sequence in the received original video file to obtain the target object image, and the image segmentation module can also be used to generate a complete background image; the encoding module is used to encode the segmented target object image to obtain the image code stream of the target object image; the distribution module is used to distribute the image code stream of the target object image to the playback end 220.

[0095] In addition, the remote server 210 can also be used to transmit the complete background image to the playback terminal 220. For example, the remote server 210 can distribute the complete background image to the playback terminal 220 via a distribution module, or the image code stream of the target object image can carry the URL (Uniform Resource Locator) of the complete background image on the remote server 210. The playback terminal 220 decodes the image code stream to obtain the URL and downloads the complete background image from the remote server 210 based on the URL. This application does not limit the specific method by which the remote server 210 transmits the complete background image to the playback terminal 220.

[0096] Exemplarily, the playback end 220 can be a terminal device including a decoding module and a rendering module, wherein: the decoding module is used to obtain the image code stream of the complete background image and the target object image, and decode the obtained image code stream of the target object frame by frame in chronological order to obtain an image sequence of the target object image; the rendering module renders the complete background image, and overlays (Over l ay) the reconstructed image of the decoded target object image on the complete background image for display.

[0097] refer to Figure 3 , Figure 3 FIG. 1 is a schematic diagram of a video processing system provided in another embodiment of the present application. The video processing method in the embodiment of the present application can be applied to the following examples: Figure 3 The video processing system shown.

[0098] like Figure 3 As shown, in Figure 3 In the illustrated embodiment, the video processing system includes a live broadcast end 310 and a playback end 320. Both the live broadcast end 310 and the playback end 320 can be terminal devices. In addition, in this application, the terminal device can also be referred to as an electronic device.

[0099] It is understandable that Figure 3 The application scenario of the video processing system shown is a point-to-point communication scenario, for example: a scenario in which a user at the live broadcast end 310 and a user at the playback end 320 make a video call, or it can also be a scenario such as home monitoring and video sharing, which is not limited in this application.

[0100] It should be understood that since the communication between the live broadcast end 310 and the playback end 320 in the video processing system is point-to-point, the number of the live broadcast end 310 and the playback end 320 in the video processing system is only one.

[0101] In some embodiments, based on Figure 3In the video processing system shown, the first terminal device in the embodiment of the video processing method of the present application can be a live broadcast terminal 310, and the second terminal device can be a playback terminal 320. The composition and function of the live broadcast terminal 310 and the playback terminal 320 are described below in an exemplary manner.

[0102] Exemplarily, the live broadcast terminal 310 may include: a terminal device comprising an acquisition module, an image segmentation module, an encoding module, and a transmission module, wherein: the acquisition module is used to acquire video. For example, the acquisition module may include a camera in the terminal device, which acquires images at a set frame rate and resolution to form original video data; the image segmentation module is used to segment the images in the image sequence in the received original video data into target objects and backgrounds to obtain target object images. The image segmentation module may also be used to generate a complete background image; the encoding module is used to encode the segmented target object images to obtain an image stream of the target object images; and the transmission module is used to distribute the image stream of the target object images to the playback terminal 320. In addition, the transmission module may also be used to transmit the complete background image to the playback terminal 320.

[0103] Exemplarily, the playback end 320 can be a terminal device including a decoding module and a rendering module, wherein: the decoding module is used to obtain the image code stream of the complete background image and the target object image, and decode the obtained image code stream of the target object frame by frame in chronological order to obtain an image sequence of the target object image; the rendering module renders the complete background image, and overlays (overlay) the reconstructed image of the decoded target object image on the complete background image for display.

[0104] In some embodiments, since both the live broadcast end 310 and the playback end 320 can be terminal devices, the terminal device serving as the live broadcast end 310 can also serve as the playback end, and the terminal device serving as the playback end 320 can also serve as the acquisition end. This application does not impose any restrictions on this.

[0105] It should be understood that in the examples of the above-mentioned video processing systems, live broadcast terminal 110 and live broadcast terminal 310 serve as the human-computer interaction interface of the user's live broadcast application. The user interacts with the terminal device by clicking, touching, pressing keys, etc. to realize functions such as video capture and interaction with the playback terminal in various application scenarios; while playback terminal 130, playback terminal 220, and playback terminal 320 serve as the human-computer interaction interface of the user's live broadcast application. The user interacts with the terminal device by clicking, touching, pressing keys, etc. to realize functions such as watching live broadcasts, video chatting, and video conferencing. The live broadcast terminal and playback terminal in each of the above-mentioned video processing systems can also realize various different functions, which are not detailed in this application.

[0106] It can be understood that in the examples of the above-mentioned video processing systems, the terminal device serving as the live broadcast end or the playback end can be: a smart phone, a tablet computer, a laptop computer, a desktop PC, a foldable screen mobile phone, a wearable device, a mobile tablet computer, a car tablet computer, a large-screen terminal, etc. This application does not limit the specific form of the terminal device.

[0107] Exemplarily, the video processing method in the embodiment of the present application mainly relies on the terminal device (i.e., the first terminal device and the second terminal device), and is widely used in the live broadcast APP on the terminal device, for example, it can be applied to the streaming APP, playback APP, live broadcast encoding service, etc.

[0108] It should be understood that in the video processing system, there are communication connections between the live broadcast end and the playback end, between the live broadcast end and the remote server, or between the playback end and the remote server to achieve mutual interaction. The communication connection can be a wired communication connection or a wireless communication connection, and this application does not impose any restrictions on this.

[0109] In the embodiment of the present application, the wireless communication can be one or more of wireless local area network (WLAN), wireless local area network (Wi-Fi), Bluetooth, Bluetooth Low Energy (BLE), mobile communication, radio frequency identification (RFID), infrared, and ultra-wideband (UWB).

[0110] It should be understood that in the present application, the terminal devices (such as live broadcast ends, remote services and playback ends) in the video processing systems in the above-mentioned examples, and the functional modules included in the terminal devices (such as live broadcast ends, remote services and playback ends) are merely illustrative and do not constitute specific limitations on the embodiments of the present application.

[0111] For example, in other embodiments of the present application, the video processing system may include more or fewer terminal devices than in the above-described video processing system. For another example, in other embodiments of the present application, each terminal device may include more or fewer functional modules than in the above-described terminal devices, or may combine certain functional modules, split certain functional modules, or include different functional modules. This is not a limitation of the present application.

[0112] The following is an exemplary introduction to the steps of the video processing method provided in this application based on the video processing system in the above application scenario.

[0113] An embodiment of the present application provides a video processing method, which is applied to a first terminal device, and the video processing method includes: obtaining a first image of a video, where the first image is a frame of image in the video; segmenting the first image into a target object and a background to obtain image segmentation information, where the image segmentation information includes a first target object image and a first background image, where the first target object image is an image of the target object in the first image, and the first background image is an image other than the first target object image in the first image; when it is determined that there is a first complete background image, encoding the first target object image to obtain an image code stream corresponding to the first target object image, where the first complete background image is a complete background image corresponding to the first background image, and the complete background image is a complete image of the background where the target object is located; sending the first complete background image and the image code stream corresponding to the first target object image to a second terminal device; the image code stream corresponding to the first complete background image and the first target object image is used by the second terminal device to determine and display a reconstructed image of the first image, where the background image of the reconstructed image of the first image is the first complete background image.

[0114] It can be understood that the first terminal device obtains the first target object image and the first background image by segmenting the first image into a target object and a background, and when it is determined that there is a first complete background image, encodes the first target object image to obtain an image code stream of the first target object image; and sends the image code stream of the first complete background image and the first target object image to the second terminal device, so that the second terminal device determines and displays the reconstructed image of the first image based on the image code stream corresponding to the first complete background image and the first target object image, wherein the background image of the reconstructed image of the first image is the first complete background image. Since the first terminal device only needs to encode a portion of the image in the first image (i.e., the first target object image), the bit rate of the image code stream of the obtained first target object image is much lower than the bit rate of the image code stream of the complete first image. For the entire video, when most of the image frames in the video are encoded by only encoding a portion of the image, the bit rate of the code stream corresponding to the entire video will be significantly reduced. Therefore, this method can effectively reduce the bit rate of the video code stream. In addition, since the first complete background image has not undergone video compression, it can maintain a high clarity. Therefore, the first complete background image is used as the background image of the reconstructed image of the first image determined by the second terminal device, so that the background of the reconstructed image of the first image displayed in the second terminal device has a high clarity, thereby ensuring clear video quality. In summary, the video processing method in this application can reduce the video bit rate while ensuring clear video quality.

[0115] Exemplarily, when it is determined that the first complete background image does not exist, the first image is encoded to obtain an image code stream corresponding to the first image; the image code stream corresponding to the first image is sent to the second terminal device, and the image code stream corresponding to the first image is used by the second terminal device to determine and display a reconstructed image of the first image. In this implementation, it is determined that the first complete background image does not exist, which means that a complete background image does not exist, or the first background image segmented from the first image has changed significantly relative to the existing complete background image, and therefore the existing complete background image cannot be used as the first complete background image; in this case, the first terminal device encodes the complete first image to obtain an image code stream of the first image and sends it to the second terminal device, so that the second terminal device can determine and display the reconstructed image of the first image based on the image code stream of the first image, thereby ensuring that the second terminal can display the reconstructed image of each image in the video, and ensuring the continuity of the video display on the second terminal device.

[0116] It should be understood that the first terminal device can be Figure 1 The live broadcast terminal 110 or the remote server 120 in the video processing system shown, or the first terminal device can be Figure 2 The remote server 210 in the video processing system shown, or the first terminal device can be Figure 3 The live broadcast end 310 in the video processing system shown.

[0117] An embodiment of the present application also provides a video processing method, which is applied to a second terminal device. The video processing method includes: receiving a first complete background image and an image code stream corresponding to a first target object image; decoding the image code stream corresponding to the first target object image to obtain a reconstructed image of the first target object image; determining and displaying the reconstructed image of the first image based on the first complete background image and the reconstructed image of the first target object image, wherein the first image is an image from which the first target object image is segmented, and the background image of the reconstructed image of the first image is the first complete background image.

[0118] It can be understood that the image code stream received by the second terminal device is the image code stream of the local image in the first image (i.e., the first target image). The bit rate of this image code stream is much smaller than the bit rate of the image code stream of the complete first image. This can effectively reduce the bit rate of the video code stream. The reduction in the bit rate of the video code stream can significantly reduce the computing power consumption of the second terminal device. In addition, the first complete background image has not undergone video compression and can maintain a high clarity. The first complete background image is used as the background image of the reconstructed image of the first image, so that the reconstructed image of the first image has clear picture quality.

[0119] It should be understood that the second terminal device can be Figure 1The playback terminal 130 in the video processing system shown, or the second terminal device can be Figure 2 The playback terminal 220 in the video processing system shown, or the second terminal device can be Figure 3 The playback end 320 in the video processing system is shown.

[0120] For ease of understanding, the specific processes of the video processing methods in different scenarios are exemplarily described below with reference to the accompanying drawings.

[0121] Figure 4 The flowchart is a flow chart of an embodiment of the present application, wherein the devices involved in the flow chart include a first terminal device and a second terminal device.

[0122] like Figure 4 As shown, the process includes the following steps: S410-S470.

[0123] S410, the first terminal device obtains a first image of the video, where the first image is a frame of image in the video.

[0124] It should be understood that in different application scenarios, the specific form of the first terminal device is different, and the method for the first terminal device to obtain the first image in the video is also different. An exemplary explanation is given below.

[0125] In some application scenarios, based on Figure 1 In the video processing system shown, the first terminal device in the video processing method is the remote server 120, and the second terminal device is the playback end 130. In this case, the first terminal device obtains an image sequence corresponding to the live stream by decoding the live stream sent by the live end 110, and the first image is a frame image in the image sequence.

[0126] In other application scenarios, based on Figure 1 In the video processing system shown, the first terminal device in the video processing method is the live broadcast end 110, and the second terminal device is the playback end 130. In this case, the first terminal device captures images according to the set frame rate and resolution to form original video data, and the first image is a frame image in the image sequence included in the video data.

[0127] In some application scenarios, based on Figure 2 In the video processing system shown, the first terminal device in the video processing method is the remote server 210, and the second terminal device is the playback end 220. The application scenario is mainly a recording and broadcasting scenario. The pre-recorded video file is stored in the first terminal device, and the first image is a frame image in the recorded and broadcasted video file.

[0128] In other application scenarios, based on Figure 3In the video processing system shown, the first terminal device in the video processing method is the live broadcast end 310, and the second terminal device is the playback end 320. In this case, the first terminal device captures images according to the set frame rate and resolution to form original video data. The first image is a frame image in the image sequence included in the video data.

[0129] S420: The first terminal device segments the first image into a target object and a background to obtain image segmentation information, where the image segmentation information includes a first target object image and a first background image.

[0130] In the embodiment, the first target object image is an image of the target object in the first image, and the first background image is an image other than the first target object image in the first image.

[0131] It can be understood that the target object refers to the main object that is focused on or displayed in the image corresponding to the video. The target object can be anything, and the specific type of the target object depends on the content and purpose of the video; for example, the products, anchors or teams that need to be displayed in e-commerce live broadcast videos, teachers explaining courses in online education videos, individuals during video chats, etc. This application does not limit the specific type of target objects; and the background is the part of the image corresponding to the video other than the target object.

[0132] In one embodiment, the target object and background segmentation process can be implemented using a target detection algorithm. The target detection algorithm automatically identifies and labels objects in the image, that is, generates a bounding box that tightly surrounds the target object. For example, in the first image, the image within the bounding box is the first target object image, and the image outside the bounding box is the first background image.

[0133] Figure 5 This is a schematic diagram of an example of image segmentation in a video processing method provided in an embodiment of the present application. Figure 5 As shown, by segmenting the target object and background in image 500, a target object image 510 and a background image 520 are obtained. Target object image 510 includes all pixel information within a bounding box in image 500, and the location where the target image is segmented from background image 520 is a blank area 521 in background image 520. In this embodiment, blank area 521 is represented by black. Of course, blank area 521 can also be represented by any other color that can make blank area 521 significantly different from other parts of background image 520, which is not described in detail in this application.

[0134] For example, the target detection algorithm can be a deep learning-based target detection algorithm, such as YOLO (You Only Look Once), RetinaNet, Faster R-CNN, and SSD (Single Shot MultiBox Detector). Through training and learning, these algorithms can automatically identify and extract the bounding box of the target object from the image. Of course, any other feasible target detection algorithm can also be used to segment the target object and background in the first image, and this application does not limit this.

[0135] In some other embodiments, the target object and background segmentation process can be implemented by using a cutout method, where the target object is cut out from the first image, the cutout image is used as the first target object image, and then the image after the target object is cut out of the first image is used as the first background image.

[0136] It should be understood that the embodiments of the present application can use any feasible cutout algorithm to segment the target object and background of the first image, and the present application does not limit the specific cutout algorithm.

[0137] In some embodiments, the image segmentation information obtained by the first terminal device by segmenting the target object and the background of the first image may also include position information of the first target image in the first image.

[0138] Exemplarily, when the image within the bounding box is used as the first target object image, the position information of the first target object image in the first image can be represented by border data, for example, it can be represented by a (x, y, w, h) quadruple, where (x, y) is the coordinate of the upper left corner of the bounding box in the first image, and (w, h) is the width and height of the bounding box.

[0139] For another example, the position information of the first target object image in the first image can also be represented by (x, y, x1, y1), where (x, y) is the coordinate of the point in the upper left corner of the bounding box in the first image, and (x1, y1) is the coordinate of the point in the lower right corner of the bounding box in the first image. The width and height (w, h) of the bounding box can be calculated based on the coordinates (x, y) and (x1, y1), which will not be elaborated in this application.

[0140] It can be understood that if the image of the target object extracted from the first image is used as the first target object image, the position information of the first target object image in the first image can be represented by the coordinates of the center of mass of the first target object image. The method for calculating the center of mass can be determined according to the existing technology and will not be elaborated here.

[0141] Of course, the first terminal device can use any possible method to segment the target object and background of the first image, and the position information of the first target object image in the first image can also be represented in any possible way, which is not enumerated in this application.

[0142] S430: The first terminal device determines that there is a first complete background image, where the first complete background image is a complete background image corresponding to the first background image, and the complete background image is a complete image of the background where the target object is located.

[0143] For ease of understanding, let's first combine Figure 5 and Figure 6 Figure a illustrates the complete background image.

[0144] like Figure 5 As shown, Figure 5 There is a blank area in the background image 520, so Figure 5 The background image 520 in is not a complete background image; and Figure 6 The image a in the figure is the background image with complete pixel information. Figure 6 As shown in FIG. a, the difference between the complete background image 610 and the background image 520 is that the position in the complete background image 610 corresponding to the blank area 521 of the background image 520 is no longer a blank area, but displays complete pixel information. Figure 6 The image in Figure a can be used as a complete background image corresponding to the background image 520.

[0145] It can be understood that determining the existence of a first complete background image indicates that the first background image segmented from the first image has little change relative to the existing complete background image, so the existing complete background image can be used as the first complete background image corresponding to the first background image.

[0146] Exemplarily, the second background image and the third background image are any two background images from the at least two background images corresponding to the first complete background image. The second background image and the second target object image are segmented from the second image in the video, and the third background image and the third target object image are segmented from the third image in the video. The time points corresponding to the second image and the third image are both before the time points corresponding to the first image. The position of the second target object image in the second image is different from the position of the third target object image in the third image. This example shows that the first background image segmented from the first image has little change relative to the existing complete background image. In this example, the second background image and the third background image are any two background images from the at least two background images corresponding to the first complete background image. Since the position of the second target object image in the second image is different from the position of the third target object image in the third image, the region from which the second target object image is segmented in the second background image is different from the region from which the third target object image is segmented in the third background image. Therefore, the first complete background image can be obtained by complementing the second background image with the third background image.

[0147] In some embodiments, in at least two background images corresponding to the first complete background image: the time interval between any two background images is greater than the time interval between any two adjacent frame images in the video. Generally speaking, the difference between two consecutive frame images in a video is relatively small. In this example, the time interval between any two background images used to form the first complete background image is greater than the time interval between any two adjacent frame images in the video, which can avoid the frequent acquisition of background images that brings about a large amount of data processing; for example, when the position of the target object in the background changes little, frequently acquiring background images increases the amount of data processing but cannot obtain the first complete background image in a short time. When acquiring background images used to form a complete background image in this way, the number of background images used can be controlled, and segmentation of the target object and the background on too many images can be avoided, thereby reducing the cost of calculation and reducing computing power consumption.

[0148] Exemplarily, in the at least two background images corresponding to the first complete background image, the time interval between any two adjacent background images is a first preset duration. In this example, when determining the first complete background image based on the at least two background images, a background image is acquired at intervals of the first preset duration. By setting the time interval for acquiring background images to a fixed first preset duration, the operation of acquiring background images becomes more regular and convenient.

[0149] S440: The first terminal device encodes the first target object image to obtain an image code stream corresponding to the first target object image.

[0150] It should be understood that the first terminal device may encode the first target object image using any possible video coding standard, such as H.264, H.265, VP8, VP9, ​​and other video coding standards, and this application does not impose any restrictions on this. With the development of science and technology, subsequent video coding standards may be applied to the encoding process of the first target object image in the embodiments of this application.

[0151] For ease of understanding, the following uses H.264 as an example to illustrate the code stream structure of an image code stream. The H.264 code stream structure may include: Supplemental Enhancement Information (SII), Coding Parameter Set (SPS / PPS, Sequence Parameter Set / Picture Parameter Set), and image coding data (Slice). Among them:

[0152] Supplemental Enhancement Information (SEI) is a key technology in the H.264 standard, primarily serving as a supplement and enhancement. SEI messages provide additional information about the video stream. This information isn't directly involved in the decoding process, but is very helpful for understanding the structure and content of the video stream. For example, SEI can include information such as timecode, scene change markers, and resume points, helping the decoder better process the video stream.

[0153] Coding Parameter Sets (SPS / PPS): The SPS (Sequence Parameter Set) and PPS (Picture Parameter Set) are key components of the H.264 bitstream. The SPS contains the coding parameters for the entire video sequence, such as resolution, frame rate, and color space. The PPS contains the coding parameters for each picture or group of pictures, such as quantization parameters and entropy coding method. These parameter sets are generated during the encoding process and transmitted along with the video stream, enabling decoders to correctly decode the video data.

[0154] Image Coded Data (Slice): Slices are the basic coding units in the H.264 bitstream and contain the actual image coded data. Each Slice contains a portion of the video frame data. By decoding the Slice, the original video frame can be restored. The Slice division method can be adjusted as needed to adapt to different encoding requirements and network conditions.

[0155] In some embodiments, the image segmentation information obtained by the first terminal device by segmenting the target object and background of the first image also includes: position information of the first target object image in the first image, and the image code stream corresponding to the first target object image includes: position information of the first target object image in the first image.

[0156] Exemplarily, if the first terminal device uses H.264 to encode the first target object image, the position information of the first target object image in the first image is the relevant data of the bounding box (x, y, w, h). When encoding the first target object image, the position information (x, y, w, h) can be encoded into the supplementary enhancement information (SE I) of the image code stream corresponding to the first target object image.

[0157] It is understandable that when the first terminal device encodes the first target object image according to other video coding standards, the position information of the first target object image in the first image is encoded into a position in the image code stream that has a similar function to the supplemental enhancement information (SE I), and this application does not elaborate on this.

[0158] In some embodiments, before encoding the first target object image, the first target object image may be resized to a preset target size. In this embodiment, the size of the target object image is normalized. Assuming the preset target size is (w', h'), the first target object image needs to be scaled from its original size (w, h) to (w', h'), and then the scaled first target object image is encoded.

[0159] Exemplarily, if the first terminal device uses H.264 to encode the first target object image, the sequence parameter set (SPS) of the image code stream corresponding to the first target object image includes a preset target size, that is, the preset target size is carried by the sequence parameter set of the image code stream corresponding to the first target object image.

[0160] It is understandable that when the first terminal device encodes the first target object image according to other video coding standards, the preset target size is encoded into a position in the image code stream that has a similar function to the sequence parameter set (SPS), and this application does not elaborate on this.

[0161] S450: The first terminal device sends the first complete background image and the image code stream corresponding to the first target object image to the second terminal device.

[0162] In an embodiment of the present application, in order to ensure the clarity of the first complete background image, the first terminal device does not perform video compression on the first complete background image.

[0163] In some embodiments, before sending the first complete background image to the second terminal device, the first complete background image may also be subjected to lossless compression, such as PNG (Portable Network Graphics) compression, to ensure the clarity of the first complete background image. PNG compression is a technology for reducing image file size using a lossless compression algorithm that can reduce file size without sacrificing image quality.

[0164] It is understandable that in different scenarios, the way in which the first terminal device sends the first complete background image to the second terminal device may be different. The following is an illustrative description of the way in which the first terminal device sends the first complete background image to the second terminal device in different application scenarios.

[0165] In some application scenarios, based on Figure 1 In the video processing system shown, the first terminal device is the remote server 120, and the second terminal device is the playback end 130. In this case, after the first terminal device (i.e., the remote server 120) obtains the first complete background image based on at least two background images, it distributes the uniform resource locator (URL) of the first complete background image in the remote server 120 to the second terminal device (i.e., the playback end 130), and the second terminal device downloads the first complete background image from the remote server 120 according to the uniform resource locator. Alternatively, after obtaining the first complete background image based on at least two background images, the first terminal device (i.e., the remote server 120) encodes the uniform resource locator (URL) of the first complete background image in the remote server 120 into the image code stream of the fourth target object image, where the fourth background image is the background image corresponding to the latest time point among the at least two background images, and the fourth background image and the fourth target object image are obtained by segmenting the fourth image; the first terminal device distributes the image code stream of the fourth target object image to the second terminal device, and the second terminal device can obtain the uniform resource locator of the first complete background image in the remote server 120 after decoding the image code stream of the fourth target object image, and the second terminal device downloads the first complete background image from the remote server 120 according to the uniform resource locator.

[0166] In other application scenarios, based on Figure 1In the video processing system shown, the first terminal device is the live broadcast end 110, and the second terminal device is the playback end 130. In this case, after the first terminal device (i.e., the live broadcast end 110) obtains a first complete background image based on at least two background images, the first complete background image is sent to the remote server 120. The remote server 120 distributes the uniform resource locator (URL) of the first complete background image in the remote server 120 to the second terminal device (i.e., the playback end 130). The second terminal device downloads the first complete background image from the remote server 120 according to the uniform resource locator. Alternatively, after obtaining the first complete background image based on at least two background images, the first terminal device (i.e., the live broadcast end 110) sends the first complete background image to the remote server 120; the first terminal device encodes the uniform resource locator (URL) of the first complete background image in the remote server 120 into the image code stream of the fourth target object image, where the fourth background image is the background image corresponding to the latest time point among the at least two background images, and the fourth background image and the fourth target object image are obtained by segmenting the fourth image; the first terminal device sends the image code stream of the fourth target object image to the remote server 120, and the remote server 120 distributes the image code stream of the fourth target object image to the second terminal device (i.e., the playback end 130), and the second terminal device can obtain the uniform resource locator of the first complete background image in the remote server 120 after decoding the image code stream of the fourth target object image, and the second terminal device downloads the first complete background image from the remote server 120 according to the uniform resource locator.

[0167] In some application scenarios, based on Figure 2In the video processing system shown, the first terminal device is the remote server 210, and the second terminal device is the playback end 220. The application scenario is mainly a recording and playback scenario. In this scenario, after the first terminal device (i.e., the remote server 210) obtains the first complete background image based on at least two background images, it distributes the uniform resource locator (URL) of the first complete background image in the remote server 210 to the second terminal device (i.e., the playback end 130). The second terminal device downloads the first complete background image from the remote server 210 according to the uniform resource locator. Alternatively, after obtaining the first complete background image based on at least two background images, the first terminal device encodes the uniform resource locator (URL) of the first complete background image in the remote server 210 into the image code stream of the fourth target object image, where the fourth background image is the background image corresponding to the latest time point among the at least two background images, and the fourth background image and the fourth target object image are obtained by segmenting the fourth image; the first terminal device distributes the image code stream of the fourth target object image to the second terminal device, and the second terminal device can obtain the uniform resource locator of the first complete background image in the remote server 210 after decoding the image code stream of the fourth target object image, and the second terminal device downloads the first complete background image from the remote server 210 according to the uniform resource locator.

[0168] In other application scenarios, based on Figure 3 In the video processing system shown, the first terminal device is the live broadcast end 310, and the second terminal device is the playback end 320. In this case, after the first terminal device (i.e., the live broadcast end 310) obtains a first complete background image based on at least two background images, it sends the first complete background image to the second terminal device (i.e., the playback end 320).

[0169] In some embodiments, the uniform resource locator of the first complete background image in the remote server is located in the supplemental enhancement information (SE I) of the image code stream corresponding to the fourth target object image. In this case, the first terminal device encodes the fourth target object image according to H.264, and the code stream structure of the image code stream includes the supplemental enhancement information.

[0170] It is understandable that when the first terminal device encodes the fourth target object image according to other video coding standards, the uniform resource locator of the first complete background image in the remote server is encoded into the image code stream corresponding to the fourth target object image in a position similar to the supplemental enhancement information (SE I). This application does not elaborate on this.

[0171] It should be understood that the first terminal device may use conventional technology in the field to send the image code stream corresponding to the first target object image to the second terminal device, and this application will not elaborate on this.

[0172] S460: The second terminal device decodes the image code stream corresponding to the first target object image to obtain a reconstructed image of the first target object image.

[0173] It should be understood that the reconstructed image of the first target object image can also be called the restored image of the first target object image. The specific decoding process only needs to be adapted to the video coding standard adopted by the image code stream. This application does not limit or elaborate on this.

[0174] S470: The second terminal device displays a reconstructed image of the first image based on the first complete background image and the reconstructed image of the first target object image, where the background image of the reconstructed image of the first image is the first complete background image.

[0175] It can be understood that the reconstructed image of the first image can also be called the restored image of the first image.

[0176] In some embodiments, step S470 includes: the second terminal device rendering the first complete background image to the display area of ​​the video; and the second terminal device overlaying the first target object image on the first complete background image to display a reconstructed image of the first image. The reconstructed image of the first target object image is overlaid on the first complete background image, so that the first target object image is displayed on top of the first complete background image, visually representing the reconstructed image of the first image. In addition, by overlaying the first target object image on the first complete background image, the digital signal of the image is allowed to be directly output to the display screen through the video memory without being processed by the display chip, thereby improving and reducing the display chip utilization of the second terminal device.

[0177] Figure 6 Figure b in FIG is a schematic diagram of the process of superimposing the target object image onto the complete background image in one embodiment of the present application. Figure 6 As shown in Figure b, by overlaying the target object image 510 onto the complete background image 610, an image 620 is obtained. The image 620 is Figure 5 The reconstructed image of image 500 in FIG.

[0178] In some embodiments, image 620 is a reconstructed image of image 500, and image 620 is basically identical to image 500. In this case, image 620 has the highest degree of restoration of image 500. Therefore, the video watched by the user in the second terminal device is basically identical to the video before processing in the first terminal device based on the video processing method in this application, so as to ensure that the original video is displayed to the user and the user's viewing experience is guaranteed.

[0179] In some other embodiments, there may be slight differences between image 620 and image 500. As long as the differences are within an acceptable range, the user's viewing experience can be guaranteed.

[0180] In some other embodiments, step S470 includes: the second terminal device overlaying the first target object image onto the first complete background image to obtain a reconstructed image of the first image; and the second terminal device rendering and displaying the reconstructed image of the first image. In this embodiment, the reconstructed image of the first image is first obtained, and then the reconstructed image of the first image is rendered and displayed, which is simple and easy to implement.

[0181] In some embodiments, the second terminal device decodes the image code stream corresponding to the first target object image and obtains position information of the first target object image within the first image. The position of the reconstructed image of the first target object image within the reconstructed image of the first image is determined based on the position information of the first target object image within the first image. In this manner, the position of the target object in the reconstructed image of the first image is closer to the position of the target object in the first image, ensuring that the reconstructed image of the first image is faithful to the first image.

[0182] In some embodiments, before the second terminal device displays the reconstructed image of the first image based on the first complete background image and the reconstructed image of the first target object image, the video processing method further includes: the second terminal device determining the original size of the first target object image, where the original size is the size of the first target object image in the first image; and adjusting the reconstructed image of the first target object image to the original size. In this manner, the size of the target object in the reconstructed image of the first image is the same as the size of the target object in the first image, ensuring that the reconstructed image of the first image accurately reproduces the first image.

[0183] exist Figure 4 In the embodiment shown, the first terminal device determines that the first complete background image exists. In other exemplary embodiments, the first terminal device determines that the first complete background image does not exist. Figure 7 This situation will be described.

[0184] Figure 7 The figure shows a flow chart of an example provided by an embodiment of the present application, wherein the devices involved in the flow chart include a first terminal device, a second terminal device and a remote server.

[0185] like Figure 7 As shown, the process includes the following steps: S710-S770.

[0186] S710, a first terminal device obtains a first image of a video, where the first image is a frame of image in the video.

[0187] It should be understood that step S710 is substantially the same as step S410 , and therefore the specific content of step S710 may refer to the description of step S410 in the foregoing text, and will not be elaborated here.

[0188] S720: The first terminal device segments the first image into a target object and a background to obtain image segmentation information, where the image segmentation information includes a first target object image and a first background image.

[0189] In the embodiment, the first target object image is an image of the target object in the first image, and the first background image is an image other than the first target object image in the first image.

[0190] It should be understood that step S720 is substantially the same as step S420, and therefore the specific content of step S720 may refer to the description of step S420 in the foregoing text and will not be elaborated here.

[0191] S730: The first terminal device determines that the first complete background image does not exist.

[0192] In the embodiment, the first complete background image refers to the complete background image corresponding to the first background image. The definition of the complete background image can be found in the description of step S430 and will not be elaborated here.

[0193] It should be understood that the first terminal device determines that the first complete background image does not exist, which means that the first background image segmented from the first image has changed significantly compared to the existing complete background image, and therefore the existing complete background image cannot be used as the first complete background image; or no complete background image has yet to be formed.

[0194] S740: The first terminal device encodes the first image to obtain an image code stream corresponding to the first image.

[0195] It can be understood that when it is determined that the first complete background image does not exist, the first terminal device encodes the entire first image. Any available video coding standard can be used to encode the first image, for example, H.264, H.265, VP8, VP9 and other video coding standards can be used. This application does not impose any restrictions on this.

[0196] S750: The first terminal device sends the image code stream corresponding to the first image to the second terminal device.

[0197] It should be understood that the first terminal device may use conventional technology in the field to send the image code stream corresponding to the first image to the second terminal device, and this application will not elaborate on this.

[0198] S760: The second terminal device decodes the image code stream corresponding to the first image to obtain a reconstructed image of the first image.

[0199] S770: The second terminal device displays the reconstructed image of the first image.

[0200] It can be understood that the reconstructed image of the first image can also be called the restored image of the first image. The specific decoding process only needs to be adapted to the video coding standard adopted by the image code stream. This application does not limit or elaborate on this.

[0201] exist Figure 7 In the illustrated embodiment, when it is determined that the first complete background image does not exist, the complete image is encoded to ensure that each frame of the video can be displayed in the second terminal device, thereby ensuring the continuity of the video playback.

[0202] For ease of understanding, the following Figure 1 Taking the live broadcast scene in which the remote server 120 in the video processing system shown is used as the first terminal device as an example, the process of the first terminal device determining a complete background image based on at least two background images is exemplarily described.

[0203] For example, at the start of a live broadcast, the live broadcast terminal 110 captures images through a camera at a set frame rate and resolution to form raw video data. The raw video data includes a sequence of images captured in a time sequence. The live broadcast terminal 110 encodes the sequence of images captured in a time sequence to obtain a live stream, and pushes the live stream to the remote server 120. The remote server 120 receives the live stream pushed from the live broadcast terminal 110 and decodes the live stream to obtain the original image sequence of the video.

[0204] When the live broadcast begins, there is no complete background image in the remote server 120, so the remote server 120 needs to determine the complete background image based on at least two background images. Figure 8 1 is a flow chart of determining a complete background image in one embodiment of the present application, wherein the remote server 120 is the execution entity.

[0205] like Figure 8 As shown, the process of determining the complete background image includes the following steps: S810-S880.

[0206] S810. Obtain the first background image: At time t0, the remote server 120 separates the target object and the background of image A1 in the original image sequence to obtain the first target object image and the first background image. The blank area in the first background image is the area after the first target object image is separated from the first background image.

[0207] S820. Obtain the second background image: At time t0+t, separate the target object and the background of image A2 in the original image sequence to obtain the second target object image and the second background image. The blank area in the second background image is the area after the second target object image is segmented out of the second background image.

[0208] It should be understood that the time interval between the acquisition of the first and second background images is t, that is, a background image is acquired every time interval t to form a complete background image, and the time interval t is greater than the time interval between two adjacent frames in the video. Therefore, when generating a complete background image, it is not necessary to segment each frame in the video. For example, a background image can be acquired every 1 second to reduce computational cost.

[0209] S830: Determine whether there is an overlapping portion between the blank area of ​​the first background image and the blank area of ​​the second background image.

[0210] It is understandable that the blank area in each background image may be represented by the position information of the target object image corresponding to the background image, for example, by the bounding box of the target object image.

[0211] Figure 9 This is a schematic diagram of determining whether there are overlapping parts in a blank area in one embodiment of the present application. Figure 9 As shown, the coordinate origin is O, the blank area of ​​the first background image is represented by a bounding box as (x1, y1, w1, h1); the blank area of ​​the second background image is represented by a bounding box as (x2, y2, w2, h2), where x2>x1>0, y1>y2>0, and w1, h1, w2 and h2 are all greater than 0.

[0212] exist Figure 9 In the embodiment shown, if x1+w1>x2 and y2+h2>y1 are both satisfied, it is determined that the blank area of ​​the first background image overlaps with the blank area of ​​the second background image; if x1+w1>x2 and y2+h2>y1 are not both satisfied, it is determined that the blank area of ​​the first background image does not overlap with the blank area of ​​the second background image.

[0213] It should be understood that in other embodiments, the relative position relationship between the blank area of ​​the first background image and the blank area of ​​the second background image is different from Figure 9 The ones shown may be different, but a similar method can also be used to calculate whether there is an overlapping area, which is not enumerated in this application.

[0214] S840: If there is no overlapping portion between the blank area of ​​the first background image and the blank area of ​​the second background image, the first background image and the second background image are spliced ​​together to obtain a complete background image.

[0215] Exemplarily, splicing the first background image and the second background image can be accomplished by splicing the image portion of the first background image corresponding to the blank area of ​​the second background image into the blank area of ​​the second background image, or by splicing the image portion of the second background image corresponding to the blank area of ​​the first background image into the blank area of ​​the first background image, thereby obtaining a complete background image and ending the process of determining the complete background image.

[0216] It is understandable that if there is no overlapping portion between the blank area of ​​the first background image and the blank area of ​​the second background image, a complete background image can be formed by using the two background images.

[0217] S850: If there is an overlap between the blank area of ​​the first background image and the blank area of ​​the second background image, the first background image and the second background image are spliced ​​to obtain a spliced ​​image B. 1至2 , and jump to S860.

[0218] Exemplarily, the first background image and the second background image can be spliced ​​together by splicing the image portion of the first background image corresponding to the blank area of ​​the second background image into the blank area of ​​the second background image, or by splicing the image portion of the second background image corresponding to the blank area of ​​the first background image into the blank area of ​​the first background image, to obtain a spliced ​​image B1 of the first background image and the second background image.

[0219] S860, obtain the jth background image, where j is an integer greater than or equal to 3: at time t0+(j-1)×t, the remote server 120 performs a search on image A in the original image sequence. j The target object and the background are separated to obtain the j-th target object image and the j-th background image. The blank area in the j-th background image is the area after the j-th target object image is segmented from the j-th background image.

[0220] S870, determine whether the blank area of ​​the j-th background image is consistent with the spliced ​​image B 1至(j-1) Check whether there is any overlap between the blank areas.

[0221] S880: If the blank area of ​​the jth background image is consistent with the spliced ​​image B 1至(j-1) If there is no overlap between the blank areas of the j-th background image and the spliced ​​image B 1至(j-1)The image part corresponding to the blank area is stitched into the stitched image B 1至(j-1) In the blank area of ​​​​the stitching image B 1至(j-1) The image portion corresponding to the blank area of ​​the j-th background image is spliced ​​into the blank area of ​​the j-th background image to obtain a complete background image.

[0222] In some other embodiments, the process of determining the complete background image includes the following step S890, wherein:

[0223] S890, if the blank area of ​​the j-th background image is consistent with the spliced ​​image B 1至(j-1) If there is an overlap between the blank areas of the j-th background image and the spliced ​​image B 1至(j-1) The image part corresponding to the blank area is stitched into the stitched image B 1至(j-1) In the blank area of ​​​​the stitched image B 1至(j-1) The image portion corresponding to the blank area of ​​the j-th background image is spliced ​​into the blank area of ​​the j-th background image to obtain the spliced ​​image B 1至j , jump to step S860 and set j=j+1.

[0224] It should be understood that since the first j background images cannot form a complete background image in step S890, it is necessary to obtain the j+1th background image to continue to form the complete background image, that is, jump to step S860 to obtain the j+1th background image, and repeat this cycle until the complete background image is obtained.

[0225] Assume that Figure 8 In the embodiment shown, the kth background image is the last background image that forms the complete background image (where k is an integer greater than or equal to 2), and the image corresponding to the kth background image in the image sequence is image A. k , then the image sequence is located at image A k All previous images are encoded using full image encoding; Image A k The encoding is performed using local image encoding; and the image sequence is located at image A k For subsequent images, it is necessary to compare the background image with the complete background image to determine whether to use the local image coding method or the complete image coding method for coding.

[0226] It can be understood that the following is an exemplary description of the process of comparing the first background image with the complete background image to determine whether the first complete background image exists.

[0227] Figure 10This is a flow chart of determining whether a first complete background image exists in one embodiment of the present application. The execution subject of this embodiment is the first terminal device. Figure 10 As shown, the process of determining whether a first complete background image exists includes the following steps: S1010-S1030, wherein:

[0228] S1010: Determine a first change degree of a first background image relative to a second complete background image, where the second complete background image is the complete background image closest to a time point corresponding to the first image at the determination moment.

[0229] It is understood that when the background of a video changes, a new complete background image is generally required to encode the partial image of the image after the background change. Therefore, during the entire video encoding process, multiple complete background images may be generated. When determining whether a complete background image corresponding to a first background image exists, the degree of change of the first background image needs to be compared with the newly generated second complete background image.

[0230] In some embodiments, when determining the degree of change between the first background image and the second complete background image, the areas outside the blank areas of the second complete background image and the first background image are compared. Specifically, local areas can be randomly sampled and brightness comparison can be performed, and the degree of change can be determined based on the brightness comparison results.

[0231] Of course, any other possible method may be used to determine the degree of change between the first background image and the second complete background image, which will not be elaborated in this application.

[0232] In some embodiments, as Figure 11 As shown, when the background of the video is constantly changing, the normal encoding mode (i.e., the complete image encoding mode) can be switched to. When the background tends to be stable, the mode of encoding only the target object image (i.e., the partial image encoding mode) can be switched to in the embodiment of the present application.

[0233] Exemplarily, the first terminal device may notify the second terminal device that the encoding mode has been switched through supplemental enhancement information (SEI) in the image code stream, and the first terminal device also needs to re-encode the SPS / PPS parameter set.

[0234] S1020: If the first change degree is less than or equal to a preset threshold, determine the second complete background image as the first complete background image.

[0235] S1030: If the first change degree is greater than a preset threshold, determine that the first complete background image does not exist.

[0236] It should be understood that after the complete background image is determined, there is no need to compare the degree of change of the background image of each frame. It is only necessary to compare the degree of change once every certain time interval.

[0237] For example, the comparison of the degree of change may be performed at intervals of a second preset time length, wherein the second preset time length is greater than the time interval between any two adjacent frame images in the video.

[0238] Exemplarily, the second preset duration may be 0.5 seconds, 1 second, or 2 seconds, etc., and this application does not impose any limitation on this.

[0239] For ease of understanding, the following uses H.264 to encode the first target object image, and the target object is a human body as an example to illustrate the video processing method in the embodiment of the present application. Figure 1 In the video processing system shown, the remote server 120 serves as the first terminal device, and the playback terminal 130 serves as the second terminal device. Figure 12 Figure a in FIG is a schematic diagram of the process in the first terminal device, Figure 12 Figure b is a schematic diagram of the process in the second terminal device.

[0240] like Figure 12 As shown in Figure a, the process in the first terminal device is divided into two parts: background generation and human body image encoding.

[0241] The detailed steps for background generation are as follows:

[0242] S101: First, an original live stream is received and decoded to obtain an original image sequence.

[0243] S102: Use a target detection algorithm (including but not limited to the Yo-lo series, RetinaNet, etc.) to perform human body detection on the images in the original image sequence to obtain the bounding box data of the human body. The bounding box data can be expressed as a (x, y, w, h) four-tuple, where (x, y) is the coordinate of the upper left corner of the bounding box in the original image, and (w, h) is the width and height of the bounding box. Separate the human body image from the background image: Crop the original image according to the (x, y, w, h) four-tuple to obtain the human body image; the remaining image is the background image with the human body image removed. At this time, the complete background image has not yet been formed, so the image encoding method is the normal encoding of the complete image (i.e., complete image encoding).

[0244] S103: After obtaining the background image, a black hole is formed at the original position of the human body image, thereby obtaining the first separated background image (ie, the first background image), and the position of the black hole is recorded as (x1, y1, w1, h1).

[0245] S104: Set the background acquisition time interval t. After t time, use the method in S102 to obtain the second separated background image (i.e., the second background image). The black hole position is (x2, y2, w2, h2). The position where the first black hole differs from the second black hole is the background area that can be used for restoration (inpainting). Save the image after inpainting as the target background image to be generated.

[0246] S105: Similarly, the Nth separated background image (ie, the Nth background image) can be obtained, and the black hole position is (xN, yN, wN, hN). Assume that the black hole disappears in the target background image to be generated at this time.

[0247] S106: Detect that the black hole in the third background image completely disappears, and the complete background generation process is completed. Human body image coding starts at this position of the Nth image (compared to the complete image coding in S102, only the human body image portion needs to be coded at this time).

[0248] S107: Put the third background image on a cloud server (such as a CDN device), and encode the URL of the third background image into the SE I.

[0249] Among them, the detailed steps of human body image coding are as follows:

[0250] S108: Obtain the human body image and the bounding box data of the human body from step S102. The data format is a (x, y, w, h) quadruple, where (x, y) is the coordinate of the upper left corner of the bounding box in the original image, and (w, h) is the width and height of the bounding box.

[0251] S109: Encode the (x, y, w, h) quadruple into supplementary enhancement information (SEI).

[0252] S110: Set the normalized size of the human body image to (w', h'), scale the human body image to (w', h'), and then encode the scaled image (for example, regenerate the encoding parameter set (i.e., SPS and PPS in H.264), with w' and h' carried by SPS).

[0253] S111: encapsulating the supplementary enhancement information (SEI) and the human body image coding data to generate a human body image live stream, and sending the generated human body image live stream to the CDN device.

[0254] After the complete background is generated and the human body image is encoded, the second terminal device obtains the human body image live stream and the complete background image from the cloud server (such as a CDN device). The second terminal device (i.e., the playback end, or the player) first renders the complete background image during playback. After receiving the human body image live stream, it parses SE I to obtain the Bounding Box data. After decoding the human body image live stream, it overlays (Overlay) the bounding box data onto the background image. In this way, the second terminal device completes the image restoration. The specific process is as follows: Figure 12 As shown in Figure b.

[0255] like Figure 12 As shown in Figure b, the detailed steps for the second terminal device to restore the image are as follows:

[0256] S101: The second terminal device obtains a live stream of a human body image from a cloud server (such as a CDN device).

[0257] S102: The second terminal device parses SE I (the first terminal device will encode the URL of the complete background image into SE I after the complete background image is generated, and will also encode the Bounding Box data into SEI when encoding the human body image) to obtain the URL of the complete background image and the human body Bounding Box data.

[0258] S103: The second terminal device downloads the complete background image from a cloud server (such as a CDN device) according to the URL of the complete background image.

[0259] S104: The second terminal device obtains the current video display area.

[0260] S105: The second terminal device renders the complete background image into the video display area.

[0261] S106: The second terminal device obtains the (x, y, w, h) quadruple according to the Bounding Box data in step S102.

[0262] S107: The second terminal device parses the coding parameter set SPS / PPS to obtain the coded human body image size (w', h'), and decodes the Silice coded data to obtain the human body image.

[0263] S108: The second terminal device scales the size of the decoded human body image from (w', h') to (w, h), thereby restoring the size of the original human body image.

[0264] S109: The second terminal device overlays the human body image onto the (x, y) coordinate position of the complete background image to restore the entire image.

[0265] In this embodiment, after human image encoding is enabled, assuming that the human body area occupies an average of one-quarter of the entire image (that is, one-half the width and one-half the height), the number of pixels actually encoded is only one-quarter of the original image. This significantly reduces the CPU or GPU consumption, memory consumption, and bit rate of the live stream on the second terminal device. Furthermore, because the background image is losslessly compressed, the entire background is very clear. Furthermore, due to the reduced bit rate, the lag rate is also greatly reduced in weak network conditions.

[0266] It is understandable that, in some embodiments, for the H.264 / H.265 video coding standard, the URL of the complete background image and / or the Bounding Box data (ie, the location information of the target object image) can be encoded into a custom NAL type.

[0267] It should be understood that the above is only intended to help those skilled in the art better understand the embodiments of the present application, and is not intended to limit the scope of the embodiments of the present application. Based on the above examples given, those skilled in the art can obviously make various equivalent modifications or changes. For example, certain steps in each embodiment of the above method may be unnecessary, or certain new steps may be added. Or a combination of any two or any multiple embodiments described above. Such modifications, changes, or combined solutions also fall within the scope of the embodiments of the present application.

[0268] It should also be understood that the above description of the embodiments of the present application focuses on emphasizing the differences between the various embodiments. The same or similar points that are not mentioned can be referenced with each other. For the sake of brevity, they will not be repeated here.

[0269] It should also be understood that the division of the modes, situations, categories and embodiments in the embodiments of the present application is only for the convenience of description and should not constitute a special limitation. The features of various modes, categories, situations and embodiments can be combined without contradiction.

[0270] It should also be understood that the various numerical numbers involved in the embodiments of this application are only for the convenience of description and are not intended to limit the scope of the embodiments of this application. The order of the sequence numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0271] It should also be understood that in the various embodiments of the present application, unless otherwise specified or there is a logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other, and the technical features in different embodiments can be combined to form new embodiments according to their internal logical relationships.

[0272] The above combination Figures 1 to 12 The embodiments of the video processing system and the video processing method provided by the embodiments of the present application are described, and the terminal device provided by the embodiments of the present application is described below.

[0273] In this embodiment, each device (including each terminal device described above) can be divided into functional modules based on the above method embodiment. For example, each function can be divided into functional modules, or two or more functions can be integrated into a single processing module. The above integrated modules can be implemented in the form of hardware. It should be noted that the module division in this embodiment is schematic and is only a logical functional division. In actual implementation, other division methods may be used.

[0274] It should be noted that the relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.

[0275] The terminal device provided in the embodiments of the present application is used to execute the process provided in any of the above embodiments, and thus can achieve the same effect as the above implementation method. In the case of an integrated unit, the terminal device may include a processing module, a storage module, and a communication module. Among them, the processing module can be used to control and manage the actions of the terminal device. For example, it can be used to support the terminal device to execute the steps performed by the processing unit. The storage module can be used to support the storage of program code and data, etc. The communication module can be used to support communication between the terminal device and other devices.

[0276] The processing module may be a processor or a controller. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, and the like. The storage module may be a memory. The communication module may specifically be a device that interacts with other terminal devices, such as a radio frequency circuit, a Bluetooth chip, or a Wi-Fi chip.

[0277] An embodiment of the present application provides a terminal device, which is used to execute the steps executed by the first terminal device in the video processing method provided by the present application, or to execute the steps executed by the second terminal device in the video processing method provided by the present application; the terminal device in the embodiment of the present application can be a handheld device (such as a mobile phone terminal), various portable notebooks, various tablet computers, smart cameras, wearable devices, etc., and the embodiment of the present application is not limited to this.

[0278] For example, Figure 13 1300 is a schematic diagram showing the hardware structure of the terminal device 1300. Figure 13 As shown, the terminal device 1300 may include a processor 1310, an external memory interface 1320, an internal memory 1330, a universal serial bus (USB) interface 1340, a charging management module 1350, a power management module 1351, a battery 1352, an antenna 1, an antenna 2, a wireless communication module 1360, a video processing module 1370, and a display screen 1380. The display screen 1380 is used to display pictures, photos, text, videos, etc. to the user.

[0279] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the terminal device 1300. In other embodiments of the present application, the terminal device 1300 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0280] The processor 1310 may include one or more processing units. For example, the processor 1310 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent components or integrated into one or more processors. In some embodiments, the terminal device 1300 may also include one or more processors 1310. The controller may generate an operation control signal based on the instruction opcode and the timing signal to complete the control of instruction fetching and execution.

[0281] In some embodiments, the processor 1310 may include one or more interfaces. These interfaces may include an inter-integrated circuit (I2C) interface, an integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a USB interface. The USB interface 1340 is an interface that complies with USB standards and may be, for example, a Mini USB interface, a Micro USB interface, or a USB Type-C interface. The USB interface 1340 may be used to transmit data between the terminal device 1300 and peripheral devices.

[0282] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely illustrative and does not constitute a structural limitation on the terminal device 1300. In other embodiments of the present application, the terminal device 1300 may also adopt a different interface connection method from the above embodiments, or a combination of multiple interface connection methods.

[0283] The wireless communication function of the terminal device 1300 can be implemented through antenna 1, antenna 2 and wireless communication module 1360, etc.

[0284] The wireless communication module 1360 can provide wireless communication solutions including Wi-Fi (including Wi-Fi sensing and Wi-Fi AP), Bluetooth (BT), wireless data transmission modules (for example, 1333MHz, 868MHz, 915MHz) applied to the terminal device 1300. The wireless communication module 1360 can be one or more devices integrating at least one communication processing module. The wireless communication module 1360 receives electromagnetic waves via antenna 1 or antenna 2 (or antenna 1 and antenna 2), filters and frequency modulates the electromagnetic wave signals, and sends the processed signals to the processor 1310. The wireless communication module 1360 can also receive the signal to be transmitted from the processor 1310, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through antenna 1 or antenna 2. For example, the terminal device 1300 can communicate with other terminal devices through the wireless communication module 1360.

[0285] The external memory interface 1320 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal device 1300. The external memory card communicates with the processor 1310 via the external memory interface 1320 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0286] The internal memory 1330 can be used to store one or more computer programs, which include instructions. The processor 1310 can execute the above instructions stored in the internal memory 1330, thereby causing the terminal device 1300 to perform the methods provided in some embodiments of the present application, as well as various applications and data processing. The internal memory 1330 may include a code storage area and a data storage area. Among them, the code storage area can store an operating system. The data storage area can store data created during the use of the terminal device 1300, etc. In addition, the internal memory 1330 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more disk storage components, flash memory components, universal flash storage (UFS), etc. In some embodiments, the processor 1310 can execute the instructions stored in the internal memory 1330 and / or the instructions stored in the memory provided in the processor 1310, thereby causing the terminal device 1300 to execute the methods provided in the embodiments of the present application, as well as other applications and data processing.

[0287] The present application also provides a chip system. Figure 14 As shown, the chip system includes at least one processor 1401 and at least one interface circuit 1402. The processor 1401 and the interface circuit 1402 can be interconnected via lines. For example, the interface circuit 1402 can be used to receive signals from other devices (such as the memory of any of the above-mentioned terminal devices). For another example, the interface circuit 1402 can be used to send signals to other devices (such as the processor 1401). Exemplarily, the interface circuit 1402 can read the instructions stored in the memory and send the instructions to the processor 1401. When the instructions are executed by the processor 1401, the terminal device can execute the various steps executed by any terminal device in the above-mentioned embodiments (for example, the terminal device can be a handheld device (such as a mobile phone terminal), various portable notebooks, various tablet computers, smart cameras, wearable devices, etc.). Of course, the chip system can also include other discrete devices, which are not specifically limited in the embodiments of the present application.

[0288] The present application also provides an apparatus, included in a terminal device, that implements the terminal device behavior described in any of the above embodiments. This functionality can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes at least one module or unit corresponding to the above functionality.

[0289] It should also be understood that the division of units in the above device is merely a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. Moreover, the units in the device can all be implemented in the form of software called through processing elements; or all be implemented in the form of hardware; or some units can be implemented in the form of software called through processing elements, and some units can be implemented in the form of hardware. For example, each unit can be a separately established processing element, or it can be integrated into a certain chip of the device. In addition, it can also be stored in a memory in the form of a program, and called by a certain processing element of the device to execute the function of the unit. Here, the processing element can also be called a processor, which can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each unit above can be implemented by the integrated logic circuit of the hardware in the processor element or in the form of software called through the processing element. In one example, the unit in any of the above devices can be one or more integrated circuits configured to implement the above method, such as one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms. For another example, when the unit in the device can be implemented in the form of a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call a program. For another example, these units can be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0290] The present application also provides a computer-readable storage medium for storing computer program code, wherein the computer program includes steps for executing the steps of executing or displaying an interface on a terminal device in any of the embodiments provided above. The readable medium may be a read-only memory (ROM) or a random access memory (RAM), which is not limited in the present application.

[0291] The present application also provides a computer program product, which includes instructions. When the instructions are executed, the terminal device executes the steps of executing or displaying the interface of the terminal device in any of the above embodiments.

[0292] An embodiment of the present application also provides a graphical user interface on a terminal device, wherein the terminal device has a display screen, a camera, a memory, and one or more processors, wherein the one or more processors are used to execute one or more computer programs stored in the memory, and the graphical user interface includes a graphical user interface displayed when the terminal device executes the steps performed by the terminal device in any of the above embodiments.

[0293] Among them, the terminal equipment, device, computer-readable storage medium, computer program product or chip system provided in this embodiment is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.

[0294] It is understandable that, in order to realize the above functions, the above-mentioned terminal devices etc. include hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0295] The embodiment of the present application can divide the functional modules of the above-mentioned terminal device etc. according to the above-mentioned method example. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. In actual implementation, there may be other division methods.

[0296] Through the description of the above embodiments, those skilled in the art will clearly understand that for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0297] The functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0298] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.

[0299] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A video processing method, characterized in that: Applied to a first terminal device, the method includes: Acquire a first image of a video, where the first image is a frame of image in the video; Segmenting the first image into a target object and a background to obtain image segmentation information, the image segmentation information including a first target object image and a first background image, wherein the first target object image is an image of the target object in the first image, and the first background image is an image of the first image excluding the first target object image; When it is determined that a first complete background image exists, encoding the first target object image to obtain an image code stream corresponding to the first target object image, wherein the first complete background image is a complete background image corresponding to the first background image, and the complete background image is a complete image of the background where the target object is located; The first complete background image and the image code stream corresponding to the first target object image are sent to a second terminal device. The image code stream corresponding to the first target object image and the first complete background image are used by the second terminal device to display a reconstructed image of the first image. The background image of the reconstructed image of the first image is the first complete background image.

2. The method according to claim 1, characterized in that Each frame image in the video corresponds to a time point, and different frame images correspond to different time points. Each complete background image is obtained based on at least two background images. The second background image and the third background image are any two background images among the at least two background images corresponding to the first complete background image. The second background image and the second target object image are obtained by segmenting the second image in the video, and the third background image and the third target object image are obtained by segmenting the third image in the video. The time point corresponding to the second image and the time point corresponding to the third image are both before the time point corresponding to the first image; the position of the second target object image in the second image is different from the position of the third target object image in the third image.

3. The method according to claim 2, characterized in that In the at least two background images corresponding to the first complete background image: the time interval between any two background images is greater than the time interval between any two adjacent frame images in the video.

4. The method according to claim 3, characterized in that In the at least two background images corresponding to the first complete background image: the time interval between any two adjacent background images is a first preset time length.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Determine whether there is an overlap between a blank area of ​​an i-th background image and a blank area of ​​a first stitched image, where the first stitched image is obtained by stitching together preceding i-1 background images, the i-th background image is the i-th background image among the background images used to form the first complete background image, the i-th background image is obtained by segmenting a target object and background from a frame of image preceding the first image in a video, the blank area of ​​the i-th background image is the area after the target object image is segmented from the i-th background image, and the blank area of ​​the first stitched image is the blank area after stitching together the preceding i-1 background images, where i is an integer greater than or equal to 2; If there is no overlapping part between the blank area of ​​the i-th background image and the blank area of ​​the first stitched image, the image part of the i-th background image corresponding to the blank area of ​​the first stitched image is stitched into the blank area of ​​the first stitched image, or the image part of the first stitched image corresponding to the blank area of ​​the i-th background image is stitched into the blank area of ​​the i-th background image to obtain the first complete background image.

6. The method according to claim 5, characterized in that The method further comprises: If there is an overlap between the blank area of ​​the i-th background image and the blank area of ​​the first stitched image, stitching the image portion of the i-th background image corresponding to the blank area of ​​the first stitched image into the blank area of ​​the first stitched image, or stitching the image portion of the first stitched image corresponding to the blank area of ​​the i-th background image into the blank area of ​​the i-th background image, to obtain a second stitched image; When there is no overlapping part between the blank area of ​​the i+1th background image and the blank area of ​​the second stitched image, the image part in the i-th background image corresponding to the blank area of ​​the second stitched image is stitched into the blank area of ​​the second stitched image, or the image part in the second stitched image corresponding to the blank area of ​​the i+1th background image is stitched into the blank area of ​​the i+1th background image to obtain the first complete background image.

7. The method according to claim 5 or 6, characterized in that When i is equal to 2, the first spliced ​​image is the first background image among the background images used to form the first complete background image.

8. The method according to any one of claims 1 to 7, characterized in that The fourth target object image and the fourth background image are obtained by segmenting the fourth image in the video, the first complete background image is obtained based on at least two background images, the fourth background image is the background image corresponding to the latest time point among the at least two background images corresponding to the first complete background image, and the image code stream corresponding to the fourth target object image includes the uniform resource locator of the first complete background image in the remote server, and the uniform resource locator is used by the second terminal device to obtain the first complete background image.

9. The method according to claim 8, characterized in that The uniform resource locator of the first complete background image in the remote server is located in the supplementary enhancement information of the image code stream corresponding to the fourth target object image.

10. The method according to any one of claims 1 to 9, characterized in that The image segmentation information further includes: position information of the first target object image in the first image, and the image code stream corresponding to the first target object image includes: position information of the first target object image in the first image.

11. The method according to claim 10, characterized in that The supplementary enhancement information of the image code stream corresponding to the first target object image includes: position information of the first target object image in the first image.

12. The method according to any one of claims 1 to 11, characterized in that Before encoding the first target object image, the method further includes: adjusting the first target object image to a preset target size.

13. The method according to claim 12, characterized in that The sequence parameter set of the image code stream corresponding to the first target object image includes the preset target size.

14. The method according to any one of claims 1 to 13, characterized in that Each frame image in the video corresponds to a time point, and different frames of image correspond to different time points. The method further includes: Determining a first degree of change of the first background image relative to a second complete background image, where the second complete background image is the complete background image closest to a time point corresponding to the first image at the determination moment; If the first change degree is less than or equal to a preset threshold, the second complete background image is determined as the first complete background image.

15. The method according to claim 14, characterized in that The degree of change of the first background image relative to the second complete background image is determined at a first moment, and the method further includes: determining, at a second moment, a second degree of change of the fifth background image relative to the second complete background image; and if the second degree of change is less than or equal to the preset threshold, determining the second complete background image as the complete background image corresponding to the fifth background image, where the fifth background image is obtained by segmenting the fifth image in the video; The time interval between the first moment and the second moment is greater than the time interval between any two adjacent frame images in the video.

16. A video processing method, characterized in that: Applied to a second terminal device, the method includes: Receive a first complete background image and an image stream corresponding to a first target object image; Decoding an image code stream corresponding to the first target object image to obtain a reconstructed image of the first target object image; Based on the first complete background image and the reconstructed image of the first target object image, a reconstructed image of the first image is displayed, where the first image is an image obtained by segmenting the first target object image, and the background image of the reconstructed image of the first image is the first complete background image.

17. The method according to claim 16, characterized in that The receiving of the first complete background image includes: receiving an image code stream corresponding to a fourth target object image, wherein the image code stream corresponding to the fourth target object image includes a uniform resource locator of the first complete background image in a remote server; The first complete background image is acquired according to the uniform resource locator of the first complete background image in the remote server.

18. The method according to claim 17, characterized in that The uniform resource locator of the first complete background image in the remote server is located in the supplementary enhancement information of the image code stream corresponding to the fourth target object image.

19. The method according to any one of claims 16 to 18, characterized in that The image code stream corresponding to the first target object image is decoded to obtain position information of the first target object image in the first image; the position of the reconstructed image of the first target object image in the reconstructed image of the first image is determined based on the position information of the first target object image in the first image.

20. The method according to claim 19, characterized in that Before displaying the reconstructed image of the first image based on the first complete background image and the reconstructed image of the first target object image, the method further includes: Determine an original size of the first target object image, where the original size is the size of the first target object image in the first image; The reconstructed image of the first target object image is adjusted to the original size.

21. The method according to any one of claims 16 to 20, characterized in that The step of displaying the reconstructed image of the first image based on the first complete background image and the reconstructed image of the first target object image includes: Rendering the first complete background image to a display area of ​​the video; The first target object image is superimposed on the first complete background image to display a reconstructed image of the first image.

22. A video processing device, characterized in that: The apparatus comprises means for performing the steps of the method according to any one of claims 1 to 15 , or the apparatus comprises means for performing the steps of the method according to any one of claims 16 to 21 .

23. A terminal device, characterized in that: The terminal device includes a processor and a memory, the memory is used to store instructions, the processor is used to read the instructions to execute the method according to any one of claims 1 to 15, or the processor is used to read the instructions to execute the method according to any one of claims 16 to 21.

24. A video processing system, characterized in that: The invention comprises a first terminal device and a second terminal device which are communicatively connected, wherein the first terminal device is used to execute the method according to any one of claims 1 to 15, and the second terminal device is used to execute the method according to any one of claims 16 to 21, and the first terminal device sends a first complete background image and an image code stream corresponding to a first target object image to the second terminal device.

25. The system according to claim 24, wherein: The first terminal device is a remote server, or the system also includes the remote server; the first complete background image is stored in the remote server, and the second terminal device obtains the first complete background image from the remote server through the uniform resource locator of the first complete background image in the remote server.

26. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the processor performs the method according to any one of claims 1 to 15, or when the program instructions are executed by the processor, the processor performs the method according to any one of claims 16 to 21.

27. A chip, characterized in that: include: A processor, configured to call and run a computer program from a memory, so that a communication device equipped with the chip performs the method according to any one of claims 1 to 15, or a communication device equipped with the chip performs the method according to any one of claims 16 to 21.