Video coding method and device

By utilizing the prediction information and detection box information of video frames to infer the outline of the Region of Interest (ROI), the problem of inaccurate ROI location in existing technologies is solved, improving the effect and efficiency of video encoding, especially in resource-constrained devices.

CN121967702APending Publication Date: 2026-05-01HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2024-10-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately determine the location of the region of interest (ROI), resulting in poor video encoding quality. This makes it particularly difficult to apply high-computing AI-based solutions in resource-constrained devices such as security cameras.

Method used

By utilizing the prediction information of the frame to be encoded and the contour information of the detection boxes or target regions of other frames, the contour of the target region of the frame to be encoded can be inferred, thereby improving the positional accuracy of the ROI and reducing computational power consumption.

Benefits of technology

It improves the accuracy and efficiency of video encoding without increasing computing power consumption, making it suitable for resource-limited devices such as security cameras, and enhancing the effectiveness of ROI encoding and privacy protection encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967702A_ABST
    Figure CN121967702A_ABST
Patent Text Reader

Abstract

The invention provides a video coding method and device, and the method comprises the steps: obtaining a detection frame of a first target frame in a video; a detection frame of a second target frame and / or the contour of a target area in the second target frame are / is obtained, the target area in the second target frame is located in the detection frame of the second target frame, and the second target frame is related to the video; determining the contour of the target area in the first target frame in the detection frame of the first target frame according to at least two of the following items: the prediction information of the first target frame, the detection frame of the second target frame or the contour of the target area in the second target frame; and coding the first target frame according to the contour of the target area in the first target frame. According to the scheme of the embodiment of the invention, the more accurate ROI position can be provided, so that the video coding effect can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding methods and apparatus Technical Field

[0001] This application relates to the field of video processing technology, and more specifically, to a method and apparatus for video encoding. Background Technology

[0002] To reduce the resources required for transmission and storage, video data is typically encoded using an encoder, for example, into a standard format such as H.265, before being transmitted.

[0003] Regions of interest (ROIs) play a crucial role in video coding. For example, ROI coding schemes can effectively guarantee the coding quality and stability of ROIs, reducing the bitrate without compromising ROI coding quality, thus improving the efficiency of video transmission and storage. Furthermore, ROIs can also be used to implement privacy protection features in video coding, such as occlusion, to conceal private areas. The accuracy of the ROI is critical to the overall effectiveness of video coding.

[0004] In related solutions, artificial intelligence (AI)-based methods, such as image segmentation models, can be used to obtain relatively accurate regions of interest (ROIs). However, this approach incurs significant computational costs, making it difficult to apply in real-world scenarios. Typically, non-AI methods are used for foreground detection, or other intelligent business models, such as perimeter detection models, are reused to determine ROIs. However, this approach only yields bounding boxes containing the ROI, not the precise location of the ROI itself.

[0005] Therefore, how to obtain a more accurate location of ROI has become an urgent problem to be solved. Summary of the Invention

[0006] This application provides a video encoding method and apparatus, which helps to provide more accurate ROI locations, thereby improving video encoding performance.

[0007] In a first aspect, embodiments of this application provide a video encoding method, the method comprising: acquiring a video to be encoded; acquiring a detection box of a first target frame in the video; acquiring a detection box of a second target frame and / or the outline of a target region in the second target frame, wherein the target region in the second target frame is located within the detection box of the second target frame, and the second target frame is related to the video; determining the outline of the target region in the first target frame based on at least two of the following: prediction information of the first target frame, the detection box of the second target frame, or the outline of the target region in the second target frame, wherein the prediction information of the first target frame includes inter-frame prediction information and / or intra-frame prediction information of the first target frame; and encoding the first target frame based on the outline of the target region in the first target frame.

[0008] According to the above scheme, the contour of the target region of the frame to be encoded can be inferred from the detection box of the frame to be encoded (such as the first target frame), the detection box information of other frames (such as the second target frame), and / or the contour information of the target region in other frames. This is beneficial to provide a more accurate range of the target region, thereby improving the encoding effect, while not introducing a large amount of computing power consumption, which is beneficial to ensuring inference efficiency.

[0009] The second target frame may include one or more video frames.

[0010] For example, the target region can be used as the ROI.

[0011] For example, the first target frame can be an I-frame, a P-frame, or a B-frame.

[0012] Optionally, the first target frame can be any frame in the video other than the first frame.

[0013] For example, the outline of the target region in the first target frame can be used for ROI encoding.

[0014] For example, the outline of the target region in the first target frame can be used for privacy-preserving coding.

[0015] In conjunction with the first aspect, in some implementations of the first aspect, the second target frame includes one or more frames that are located before the first target frame in the encoding order.

[0016] According to the above scheme, the second target frame is located before the first target frame in the encoding order, which can reduce the impact on video encoding and help ensure encoding efficiency.

[0017] In conjunction with the first aspect, in some implementations of the first aspect, the second target frame includes one or more reference frames of the first target frame.

[0018] For example, when the first target frame is a P-frame, the second target frame may include one or more reference frames of the first target frame.

[0019] For example, when the first target frame is B, the second target frame may include one or more forward reference frames of the first target frame.

[0020] In conjunction with the first aspect, in some implementations of the first aspect, the first target frame is a video frame other than the first frame in the video, and the method further includes: obtaining the detection box of the first frame in the video; and encoding the first frame based on the detection box of the first frame.

[0021] In conjunction with the first aspect, in some implementations of the first aspect, the contour of the target region in the first target frame is determined based on at least two of the following: prediction information of the first target frame, the detection box of the second target frame, or the contour of the target region in the second target frame, including: offsetting the contour of the target region in the second target frame according to the offset between the detection box of the first target frame and the detection box of the second target frame to obtain the contour of the target region in the first target frame.

[0022] Optionally, when the second target frame includes multiple video frames, the contour of the target region in the first target frame is determined based on at least two of the following: prediction information of the first target frame, the detection box of the second target frame, or the contour of the target region in the second target frame, including: offsetting the contour of the target region in each video frame according to the offset between the detection box of the first target frame and the detection box of each of the multiple video frames, to obtain multiple offset contours; and determining the contour of the target region in the first target frame based on the multiple offset contours.

[0023] Optionally, the first target frame is an I-frame.

[0024] According to the above scheme, the contour of the target region in the frame to be encoded (such as the first target frame) is obtained by inferring the contour of the target region in other frames (such as the second target frame) and the offset between detection boxes. This is beneficial to obtain a more accurate range of the target region. At the same time, this method is relatively simple and efficient, which helps to ensure that the encoding efficiency is not affected.

[0025] In conjunction with the first aspect, in certain implementations of the first aspect, the contour of the target region in the first target frame is determined based on at least two of the following in the detection box of the first target frame: prediction information of the first target frame, the detection box of the second target frame, or the contour of the target region in the second target frame, including: performing outer contour detection based on a first target initial point set to obtain the contour of the target region in the first target frame, wherein the first target initial point set includes a first prediction block and / or a second prediction block, the inter-frame prediction information of the first target frame is used to indicate the prediction information of the inter-frame prediction block of the first target frame, the first prediction block belongs to the inter-frame prediction block of the first target frame, the first prediction block is located within the detection box of the first target frame, the first prediction block is the target point of the first motion vector in the motion vector corresponding to the first target frame, the starting point of the first motion vector is located within the detection box of the second target frame, the intra-frame prediction information of the first target frame is used to indicate the prediction information of the intra-frame prediction block of the first target frame, the second prediction block belongs to the intra-frame prediction block of the first target frame, and the second prediction block is located within the detection box of the first target frame.

[0026] Optionally, when the second target frame includes multiple video frames, the contour of the target region in the first target frame is determined based on at least two of the following in the detection box of the first target frame: prediction information of the first target frame, detection box of the second target frame, or the contour of the target region in the second target frame, including: performing outer contour detection based on multiple first target initial point sets respectively to obtain multiple candidate contours of the target region in the first target frame, wherein the multiple first target initial point sets respectively correspond to the multiple video frames, and determining the contour of the target region in the first target frame based on the multiple candidate contours.

[0027] Optionally, the first target frame is a P-frame or a B-frame.

[0028] According to the above scheme, the target initial point set (i.e. the first target initial point set) is determined by using prediction information. For example, the appropriate motion vector is determined by using inter-frame prediction information, thereby determining the appropriate inter-frame prediction block. The appropriate intra-frame prediction block is determined by using intra-frame prediction information, and then the outer contour detection is performed on the target initial point set. This is beneficial to obtain a more accurate range of the target area, while not introducing a large amount of computing power consumption.

[0029] In conjunction with the first aspect, in certain implementations of the first aspect, the contour of the target region in the first target frame is determined based on at least two of the following in the detection box of the first target frame: prediction information of the first target frame, the detection box of the second target frame, or the contour of the target region in the second target frame, including: performing outer contour detection based on a second target initial point set to obtain the contour of the target region in the first target frame, wherein the second target initial point set includes a third prediction block and / or a second prediction block, the inter-frame prediction information of the first target frame is used to indicate the prediction information of the inter-frame prediction block of the first target frame, the third prediction block belongs to the inter-frame prediction block of the first target frame, the third prediction block is located within the detection box of the first target frame, the third prediction block is the target point of the second motion vector in the motion vector corresponding to the first target frame, the starting point of the second motion vector is located within the contour of the target region in the second target frame, the intra-frame prediction information of the first target frame is used to indicate the prediction information of the intra-frame prediction block of the first target frame, the second prediction block belongs to the intra-frame prediction block of the first target frame, and the second prediction block is located within the detection box of the first target frame.

[0030] Optionally, when the second target frame includes multiple video frames, the contour of the target region in the first target frame is determined based on at least two of the following in the detection box of the first target frame: prediction information of the first target frame, detection box of the second target frame, or the contour of the target region in the second target frame, including: performing outer contour detection based on multiple initial point sets of second targets respectively to obtain multiple candidate contours of the target region in the first target frame, wherein the multiple initial point sets of second targets respectively correspond to the multiple video frames, and determining the contour of the target region in the first target frame based on the multiple candidate contours.

[0031] Optionally, the first target frame is a P-frame or a B-frame.

[0032] According to the above scheme, the initial target point set (i.e., the second initial target point set) is determined using prediction information. For example, suitable motion vectors are determined using inter-frame prediction information, thereby determining suitable inter-frame prediction blocks. Suitable intra-frame prediction blocks are determined using intra-frame prediction information. Then, the outer contour of the initial target point set is detected, which helps to obtain a more accurate target region range without introducing large computational costs. Furthermore, the starting point of the selected motion vectors comes from the contour of the target region in the reference frame, which helps to further improve the accuracy of the selected motion vectors. Consequently, the accuracy of the initial target point set is further improved, thus improving the accuracy of the target region contour.

[0033] Secondly, embodiments of this application provide a video encoding apparatus, which includes units for implementing the first aspect or any possible implementation of the first aspect.

[0034] Thirdly, embodiments of this application provide a computer device including a processor for coupling with a memory to read and execute instructions and / or program code in the memory to perform the first aspect or any possible implementation thereof.

[0035] Fourthly, embodiments of this application provide a chip system including logic circuitry for coupling with an input / output interface to transmit data via the input / output interface, thereby executing the first aspect or any possible implementation thereof.

[0036] Fifthly, embodiments of this application provide a computer-readable storage medium storing program code that, when executed on a computer, causes the computer to perform the first aspect or any possible implementation thereof.

[0037] In a sixth aspect, embodiments of this application provide a computer program product comprising: computer program code, which, when run on a computer, causes the computer to perform the first aspect or any possible implementation thereof. Attached Figure Description

[0038] Figure 1 is a schematic flowchart of foreground detection used in video coding.

[0039] Figure 2 is a schematic diagram of the encoding effect based on the detection box.

[0040] Figure 3 is a schematic flowchart of a video encoding method according to an embodiment of this application.

[0041] Figure 4 is a schematic diagram of an ROI encoding process according to an embodiment of this application.

[0042] Figure 5 is a schematic diagram of a privacy-preserving coding process according to an embodiment of this application.

[0043] Figure 6 is a schematic flowchart of a method for inferring the contour of a target region in a video frame according to an embodiment of this application.

[0044] Figure 7 is a schematic diagram of the reasoning result of the contour of a target region according to an embodiment of this application.

[0045] Figure 8 is a schematic diagram of the contour reasoning process according to an embodiment of this application.

[0046] Figure 9 is a schematic diagram of the intra-prediction block and intra-prediction information according to an embodiment of this application.

[0047] Figure 10 is a schematic diagram of the reasoning result of the contour of a target region according to an embodiment of this application.

[0048] Figure 11 is a schematic diagram of the reasoning result of the contour of a target region according to an embodiment of this application.

[0049] Figure 12 is a schematic diagram of a video encoding method according to an embodiment of this application.

[0050] Figure 13 is a schematic diagram of a reasoning profile process according to an embodiment of this application.

[0051] Figure 14 is a schematic diagram of the range of ROI determined by different schemes.

[0052] Figure 15 is a schematic structural block diagram of a computer device provided according to an embodiment of this application.

[0053] Figure 16 is a schematic diagram of another computer device provided in an embodiment of this application.

[0054] Figure 17 is a schematic diagram of a chip system provided in an embodiment of this application. Detailed Implementation

[0055] The technical solutions in this application will now be described with reference to the accompanying drawings.

[0056] In the embodiments of this application, the words "exemplary," "for example," etc., are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design that is described as "exemplary" in this application should not be construed as being more preferred or advantageous than other embodiments or design options. Specifically, the use of the term "exemplary" is intended to present the concept in a concrete manner.

[0057] In the embodiments of this application, "corresponding" and "corresponding" can sometimes be used interchangeably. It should be noted that when the distinction is not emphasized, their intended meanings are consistent.

[0058] The network architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0059] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0060] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0061] To better illustrate aspects of the embodiments of this application, the terms that may be involved in the embodiments of this application will be explained below.

[0062] (1) ROI;

[0063] A Region of Interest (ROI) is a region in an image or video frame that is defined as a rectangle, circle, ellipse, irregular polygon, or other shape to be processed.

[0064] (2) Video coding generally refers to the processing of a sequence of images that form a video or video sequence. In the field of video coding, the terms "picture," "frame," or "image" can be used synonymously. Video coding is performed on the source side and typically involves processing (e.g., by compression) the raw video images to reduce the amount of data required to represent them, thereby enabling more efficient storage and / or transmission. Video decoding is performed on the destination side and typically involves inverse processing relative to the encoder to reconstruct the video images.

[0065] (3) ROI encoding;

[0066] ROI coding is a technique in video coding used to improve the coding quality and priority of specific regions (such as faces, text, or moving objects). The goal of ROI coding is to ensure that the most important visual information is best protected and presented with limited bitrate resources.

[0067] (4) Intra-coded frame, i.e., I-frame;

[0068] An I-frame stands for keyframe. It is a complete image frame that contains all the information of the frame and can be decoded independently without relying on other frames.

[0069] (5) Forward predictive frame, i.e. P-frame;

[0070] A P-frame is constructed by referencing a previously decoded I-frame or P-frame. It only stores change information relative to the reference frame, making a P-frame smaller than an I-frame. When decoding a P-frame, the decoder needs to apply the change information to one or more previously decoded reference frames to reconstruct the P-frame.

[0071] (6) Bi-directional interpolated prediction frame, i.e., B-frame;

[0072] B-frames represent bidirectional prediction frames, which take into account both the encoded frames preceding the original image sequence and the temporal redundancy information between the encoded frames following the original image sequence, in order to compress the amount of data transmitted in the encoded image.

[0073] (7) Bitrate;

[0074] Bitrate, also known as video transmission bitrate, bandwidth consumption, or throughput, is the number of bits transmitted per unit of time. Bitrate is usually expressed in bits per second (bit / s or bps).

[0075] (8) YUV;

[0076] YUV is a color encoding method where Y represents luminance, and U and V represent chrominance. Compared to RGB, YUV is more suitable for encoding and decoding.

[0077] (9) Raw image;

[0078] Raw images are the original data captured by complementary metal-oxide-semiconductor (CMOS) or charge-coupled device (CCD) image sensors and converted into digital signals. A raw file records the raw information from the image sensor, along with metadata generated by the sensor, such as International Organization for Standardization (ISO) settings, shutter speed, aperture value, and white balance.

[0079] (10)H.265;

[0080] H.265 is a new video coding standard developed by the Video Coding Experts Group (VCEG) of the International Telecommunication Union (ITU-T) Telecommunication Standardization Sector, following H.264. H.265 defines the process steps for video encoding and decoding, as well as the characteristic format that the encoded bitstream should conform to.

[0081] Object detection is the process of finding all objects of interest (bounding boxes) in an image and determining their category and location.

[0082] (12) Target segmentation;

[0083] Object segmentation is the accurate separation of specific targets or objects from an image or video. Unlike object detection, which focuses on object location and bounding boxes, object segmentation requires the precise identification and labeling of each pixel of the target.

[0084] To reduce the resources required for transmission and storage, video data is typically encoded using an encoder, for example, into a standard format like H.265, before transmission. To further reduce the bitrate after compression, Region of Interest (ROI) encoding can be used during video encoding. This approach effectively ensures the encoding quality and stability of the region of interest, reducing the bitrate without compromising the ROI's encoding quality, thus improving video transmission and storage efficiency. However, ROI encoding will reduce the image quality of non-ROI regions.

[0085] For example, a camera acquires raw data through its lens and sensor, then corrects the raw data using image signal processing (ISP) to obtain YUV data. The YUV data is then encoded (compressed) into a bitstream in standard formats such as H.265 by an encoder. To further reduce the bitrate after compression, ROI encoding is used when converting the YUV data to a bitstream in formats such as H.265.

[0086] In addition, Region of Interest (ROI) is generally used when implementing privacy protection features in video encoding. By performing operations such as occlusion on the ROI, the purpose of obscuring private areas can be achieved.

[0087] In related solutions, artificial intelligence (AI)-based approaches, such as image segmentation models, can be used to obtain relatively accurate regions of interest. However, this approach consumes a great deal of computing power and is difficult to apply in real-world scenarios.

[0088] To reduce computational consumption, non-AI methods are typically used for foreground detection, or other intelligent business models, such as perimeter detection models, are reused to determine the region of interest. However, this approach can only obtain bounding boxes containing the region of interest, not the precise location of the region of interest.

[0089] Figure 1 shows a schematic flowchart of foreground detection used for video encoding.

[0090] As shown in Figure 1, a detection box is output through foreground detection, which includes the Region of Interest (ROI). This detection box can be provided to the encoder. If the ROI is used for privacy protection, the encoder can perform operations such as occlusion on the image within the detection box during encoding. If the ROI is used for ROI encoding, the encoder can use specific techniques to implement ROI encoding, such as setting different quantization parameters (QP) for the areas inside and outside the detection box.

[0091] Compared to the actual ROI, an excessively large detection bounding box will, on the one hand, affect the bitrate reduction effect of ROI encoding, and on the other hand, result in an excessively large occlusion area for privacy protection, affecting the visual effect.

[0092] Figure 2 shows the effect of encoding using detection boxes.

[0093] Figure 2(a) illustrates the bitrate usage of video frames encoded with actual ROIs and those encoded with detection boxes. As shown in Figure 2(a), the white dots represent bitrate usage. The bitrate usage of ROI encoding based on detection boxes is significantly higher than that of ROI encoding based on actual ROIs, resulting in wasted bitrate. Figure 2(b) illustrates the bitrate usage of video frames and the bitrate usage when privacy protection is achieved based on detection boxes.

[0094] In view of this, embodiments of this application provide a video encoding method that facilitates obtaining more accurate ROI locations, thereby improving the video encoding effect.

[0095] The method described in this application can be applied to scenarios requiring video encoding, such as ROI encoding or privacy-preserving encoding. For example, the solution described in this application can be applied to scenarios such as video encoding for security cameras, multimedia entertainment video encoding, live streaming video encoding, or game video encoding.

[0096] Taking security cameras as an example, security scenarios generally focus on motor vehicles, non-motor vehicles, and people, with lower image quality requirements for other non-focused objects. Therefore, ROI encoding is becoming increasingly common in security cameras. Furthermore, privacy-preserving encoding is also becoming more prevalent in security cameras. Typically, the neural processing unit (NPU) of security cameras has limited computing power, making it difficult to run an additional segmentation model for ROI encoding or privacy-preserving encoding. The solution adopted in this application can obtain more accurate ROI locations, thereby improving the effectiveness of both ROI encoding and privacy-preserving encoding.

[0097] Figure 3 shows a schematic flowchart of a video encoding method according to an embodiment of this application.

[0098] For example, the method 300 shown in FIG3 can be performed by a device with video encoding capabilities, such as a camera or a server.

[0099] As shown in Figure 3, method 300 may include the following steps.

[0100] 310, Get the video to be encoded.

[0101] 320, Obtain the detection bounding box of the first target frame in the video.

[0102] 330. Obtain the detection bounding box of the second target frame and / or the outline of the target region in the second target frame. The target region in the second target frame is located within the detection bounding box of the second target frame. The second target frame is relevant to the video.

[0103] 340. The outline of the target region in the first target frame is determined based on at least two of the following: prediction information of the first target frame, the detection box of the second target frame, or the outline of the target region in the second target frame. The prediction information of the first target frame includes inter-frame prediction information and / or intra-frame prediction information of the first target frame.

[0104] 350, Encode the first target frame according to the contour of the target region in the first target frame.

[0105] The video to be encoded can be obtained in a variety of ways.

[0106] For example, the video to be encoded can be Raw data collected by a local device through a camera or sensor, and then YUV data obtained after correction by an ISP. Alternatively, the video to be encoded can be pre-stored on a local device. Or, the video to be encoded can be obtained through other means; this application does not limit this method.

[0107] The video to be encoded can include multiple video frames arranged in chronological order. Taking YUV data as an example, one video frame can be a YUV image.

[0108] Optionally, the first target frame can be any frame in the video other than the first frame.

[0109] The terms "first" and "second" in "first target frame" and "second target frame" are used only to distinguish different frames and have no other limiting function. For example, the first target frame is the Nth frame in the video, where N is an integer greater than 1, meaning the Nth frame can be any frame in the video other than the first frame.

[0110] The first target frame can be an I-frame, P-frame, or B-frame.

[0111] The first target frame can be the frame to be encoded.

[0112] The second target frame may include one or more video frames. A second target frame comprising one video frame can also be understood as the first target frame corresponding to one second target video frame. When the second target frame comprises multiple video frames, these multiple video frames can also be referred to as multiple second target frames, or in other words, the first target frame corresponds to multiple second target frames.

[0113] For different first target frames, the number of corresponding second target frames can be the same or different.

[0114] Optionally, the second target frame may include one or more frames that are located before the first target frame in the encoding order, and are simply referred to as frames before the first target frame.

[0115] The second target frame can be understood as either the original frame to be encoded in the video, or the frame obtained after encoding and decoding the original frame to be encoded in the video. This application does not distinguish between the two. For example, if the second target frame is understood as the original frame to be encoded in the video, then the second target frame can be the frame to be encoded that precedes the first target frame in the encoding order. As another example, if the second target frame is understood as the frame obtained after encoding and decoding the original frame to be encoded in the video, then the second target frame can be the frame decoded before encoding the first target frame.

[0116] For example, the second target frame can be any frame or multiple frames preceding the first target frame. For instance, if the first target frame is the Nth frame, the second target frame can be the (N-1)th frame.

[0117] Furthermore, optionally, if a reference frame exists for the first target frame, the second target frame may include one or more reference frames of the first target frame.

[0118] For example, when the first target frame is a P-frame, the second target frame may include one or more reference frames of the first target frame. For instance, if the first target frame is the Nth frame, the second target frame, i.e., the reference frame, may be the (N-1)th frame.

[0119] For example, when the first target frame is B, the second target frame may include one or more forward reference frames of the first target frame.

[0120] It should be understood that the above are merely examples and not limitations. For instance, if a reference frame exists for the first target frame, the second target frame can also be any frame or multiple frames preceding the first target frame.

[0121] For ease of description, this application embodiment only uses the encoding process of one frame (such as the Nth frame) as an example for illustration. Other video frames in the video, except for the first frame, can also be encoded by referring to steps 320 to 350.

[0122] The detection boxes in each frame can also be referred to as the detection boxes of the target regions in each frame.

[0123] A detection box can also be called a bounding box. For example, a detection box can be a rectangular detection box, or simply a rectangular box.

[0124] For example, the detection box can be obtained by other modules and provided to the encoder.

[0125] For example, the detection box can be obtained by the foreground detection module by performing foreground detection on the video frame. Alternatively, the detection box can be obtained by reusing models from other intelligent services (such as a perimeter detection model) to process the video frame.

[0126] Alternatively, the detection frame can also be determined by the encoder.

[0127] The detection boxes for different video frames may be obtained by the same method or by different methods, and this application does not limit this.

[0128] The target region of each frame can be considered as the ROI of that frame. For example, the target region of each frame can be the foreground region of that frame.

[0129] The outline of the target region in the second target frame can be obtained in a variety of ways.

[0130] For example, the outline of the target region in the second target frame can be obtained by the method of the embodiments of this application. When the second target frame is not the first frame, the first target frame in method 300 can be replaced with the second target frame, and the second target frame in method 300 can be replaced with other frames to obtain the outline of the target region in the second target frame.

[0131] Alternatively, the contour of the target region in the second target frame can also be obtained through other methods. For example, the contour of the target region in the second target frame can be obtained through a segmentation model.

[0132] The prediction information of the first target frame can be obtained through an encoder. For example, the prediction information of the first target frame can be obtained through an encoder kernel. For instance, the first target frame is input into the encoder kernel, and intra-frame prediction information of the first target frame is obtained through intra-frame prediction, and / or inter-frame prediction information of the first target frame is obtained through inter-frame prediction.

[0133] For example, the first target frame is an I-frame, and the prediction information of the first target frame may include the intra-frame prediction information of the first target frame.

[0134] For example, the first target frame is a P-frame or a B-frame, and the prediction information of the first target frame may include inter-frame prediction information of the first target frame. Further, the prediction information of the first target frame may also include intra-frame prediction information of the first target frame.

[0135] For example, intra-frame prediction information can correspond to a prediction block (PB). A prediction block can also be replaced by a prediction unit, etc. A prediction block can be the basic unit for performing intra-frame prediction. This prediction block can be called an intra-frame prediction block.

[0136] Intra-frame prediction information can be used to indicate the prediction information of the corresponding intra-frame prediction block. For example, the prediction information may include at least one of the following: prediction mode, partitioning mode, prediction direction, or context information.

[0137] For example, inter-frame prediction information can correspond to prediction blocks. A prediction block can be the basic unit for performing inter-frame prediction. This prediction block can be called an inter-frame prediction block.

[0138] Inter-frame prediction information can be used to indicate the prediction information of the corresponding inter-frame prediction block. For example, the prediction information may include at least one of the following: motion vector (MV), reference frame index, prediction mode, partition mode, or motion vector precision, etc.

[0139] For example, step 340 can be performed by the encoder, for example, by the encoder kernel. Alternatively, step 340 can also be performed by other modules.

[0140] For a detailed description of step 340, please refer to method 600 below.

[0141] As previously mentioned, the first target frame can be any video frame other than the first frame in the video. Optionally, method 300 may also include steps 360 to 370 (not shown in the figure).

[0142] 360, to obtain the detection box of the first frame in the video.

[0143] 370. Encode the first frame based on the detection box of the first frame.

[0144] The first frame in the video can be encoded using steps 360 and 370.

[0145] As an example, when encoding the Region of Interest (ROI) of a video, the ROI of the first frame is encoded based on the detection bounding box of the first frame, and the ROI of the other frames is encoded based on the contours of the target regions in the other frames. In other words, the shape of the ROI range in the first frame is a rectangle, and the shape of the ROI range in the other frames is the contour of the target region. The contours of the target regions in the other frames can be obtained using the method described in the embodiments of this application.

[0146] As another example, when performing privacy-preserving encoding on a video, the first frame is coded based on the detection bounding box of the first frame, and the other frames are coded based on the contours of the target regions in other frames. In other words, the shape of the ROI range in the first frame is a rectangle, and the shape of the ROI range in the other frames is a contour. Accordingly, the privacy-preserving region of the first frame in the decoded video is a rectangle, and the privacy-preserving region of the other frames is a contour. The contours of the target regions in the other frames can be obtained using the method in the embodiments of this application.

[0147] The contour of the target region of the first target frame obtained through reasoning in step 340 can be used for encoding the first target frame, and finally the bitstream of the first target frame is obtained.

[0148] For example, the outline of the target region in the first target frame can be used for ROI encoding.

[0149] Figure 4 illustrates a schematic diagram of an ROI encoding process according to an embodiment of this application. As shown in Figure 4, the inference of the target region can be performed by the encoder kernel.

[0150] For example, as shown in Figure 4, the first target frame is input into the encoder kernel, which performs inter-frame prediction and / or intra-frame prediction to obtain the prediction information of the first target frame. Based on the prediction information of the first target frame, the detection box of the second target frame, and / or the contour of the target region in the second target frame, the encoder kernel infers the contour of the target region in the first target frame from the detection box of the first target frame. As shown in Figure 4, the contour of the target region in the first target frame can be used in the quantization step, such as setting different QPs for the inner and outer regions of the contour of the target region in the first target frame.

[0151] The above are just examples. The outline of the target region can also be applied in other steps of ROI encoding, such as replacing the role of the detection box in ROI encoding. This application does not limit the specific way in which the outline of the target region is applied to ROI encoding.

[0152] For example, the outline of the target region in the first target frame can be used for privacy-preserving coding.

[0153] Figure 5 shows a schematic diagram of the privacy protection coding process according to an embodiment of this application.

[0154] For example, as shown in Figure 5(a), the first target frame is occluded based on the contour of the target region of the first target frame. For example, the image within the contour of the target region in the first target frame is mosaicked or Gaussian blurred. The occluded first target frame is then input into the encoder kernel for encoding.

[0155] For example, as shown in Figure 5(b), after the first target frame is encoded by the encoder kernel, the bitstream of the first target frame is obtained. Based on the contour of the target region of the first target frame, the relevant bitstream elements in the bitstream of the first target frame are processed so that the image within the contour of the target region in the first target frame obtained after decoding the processed bitstream is occluded.

[0156] The above are merely examples. The outline of the target region can also be protected for privacy in other ways, such as replacing the role of the detection box in privacy protection coding. This application does not limit the specific way in which the outline of the target region is applied for privacy protection.

[0157] Figure 6 illustrates a method for inferring the contour of a target region in a video frame according to an embodiment of this application. Exemplarily, method 600 can be executed by an encoder, such as by an encoder kernel. The contour inferred by method 600 can be used in method 300.

[0158] As shown in Figure 6, method 600 may include the following steps.

[0159] 610, Obtain the detection box of the first target frame in the video to be encoded.

[0160] 620, Obtain the detection bounding box of the second target frame and / or the outline of the target region in the second target frame. The target region in the second target frame is located within the detection bounding box of the second target frame. The second target frame is relevant to the video.

[0161] 630. The outline of the target region in the first target frame is determined based on at least two of the following: prediction information of the first target frame, the detection box of the second target frame, or the outline of the target region in the second target frame. The prediction information of the first target frame includes inter-frame prediction information and / or intra-frame prediction information of the first target frame.

[0162] Steps 610 and 620 correspond to steps 320 and 330 in method 300, respectively. For a detailed description, please refer to steps 320 and 330. To avoid repetition, they will not be repeated here.

[0163] Step 630 corresponds to step 340.

[0164] Step 630 (i.e., step 340) will be explained below.

[0165] Step 630 can be implemented in several ways. The following describes the implementation of step 340 using four methods (method 1, method 2, method 3, and method 4) as examples.

[0166] Method 1:

[0167] As one possible implementation, the contour of the target region in the first target frame is determined based on the detection frame of the second target frame and the contour of the target region in the second target frame.

[0168] Optionally, step 630 may include: determining the contour of the target region in the first target frame based on the offset between the detection box of the first target frame and the detection box of the second target frame, and the contour of the target region in the second target frame.

[0169] That is, the outline of the target region in the second target frame is offset based on the offset amount. The position of the offset outline can be used to determine the position of the outline of the target region in the first target frame. In other words, the offset outline can be used to determine the outline of the target region in the first target frame.

[0170] The following explanation uses the second target frame as an example of a video frame.

[0171] When the second target frame is a video frame, in step 630, the outline of the target region in the second target frame can be offset according to the offset between the detection box of the first target frame and the detection box of the second target frame. The position of the offset outline can be used as the position of the outline of the target region in the first target frame, or in other words, the offset outline can be used as the outline of the target region in the first target frame.

[0172] For example, the first target frame is frame N, the second target frame is frame (N-1), and the detection box S of frame N... N And the detection box S in the (N-1)th frame N-1 The offset between them is represented by S N -S N-1 Then R N =R N-1 +Offset.

[0173] For example, the offset between the detection box of the first target frame and the detection box of the second target frame can be the offset of the center of the detection box of the first target frame relative to the center of the detection box of the second target frame.

[0174] For example, calculate the difference (dx, dy) between the coordinates of the center of the detection box in the first target frame and the coordinates of the center of the detection box in the second target frame, offset the contour of the target region in the second target frame by (dx, dy), and obtain the contour of the target region in the first target frame.

[0175] Figure 7 illustrates a schematic diagram of the inference result of the contour of a target region according to an embodiment of this application. As shown in Figure 7, the contour of the target region in the Nth frame is inferred from the detection box of the (N-1)th frame and the contour of the target region in the (N-1)th frame.

[0176] When the second target frame comprises multiple video frames, in step 630, the contour of the target region in each of the multiple second target frames can be offset according to the offset between the detection box of the first target frame and the detection box of each of the multiple second target frames. In this case, the positions of multiple offset contours can be obtained, and the position of the target region contour in the first target frame can be determined based on the positions of these multiple offset contours. For example, the target region in the first target frame can be the union of the regions indicated by the multiple offset contours, and the contour of this union is the contour of the target region in the first target frame.

[0177] For example, the first target frame can be an I-frame, that is, the outline of the target region can be obtained by reasoning from other I-frames in the video besides the first frame through method 1.

[0178] For example, the first target frame can be a P-frame or a B-frame, that is, the contour of the target region can be obtained by reasoning from the P-frame or B-frame in the video through method 1.

[0179] In Method 1, the contour of the target region in the frame to be encoded (such as the first target frame) is obtained by inferring the contour of the target region in other frames (such as the second target frame) and the offset between detection boxes. This is beneficial to obtain a more accurate range of the target region. At the same time, this method is relatively simple and efficient, which helps to ensure that the encoding efficiency is not affected.

[0180] Method 2:

[0181] As one possible implementation, the outline of the target region in the first target frame is determined in the detection box of the first target frame based on the prediction information of the first target frame and the detection box of the second target frame.

[0182] Method 2 can be understood as inferring the contour from the detection box, as shown in Figure 8(a), inferring the contour of the target region in the frame to be encoded based on the detection boxes of other frames.

[0183] Optionally, the prediction information of the first target frame includes inter-frame prediction information of the first target frame. The inter-frame prediction information of the first target frame is used to indicate the prediction information of the inter-frame prediction blocks in the first target frame. Step 630 may include: performing outer contour detection based on the first target initial point set to obtain the contour of the target region of the first target frame. The first target initial point set includes a first prediction block. The first prediction block belongs to the inter-frame prediction block of the first target frame. The first prediction block is located within the detection box of the first target frame. The first prediction block is the target point of the first motion vector in the motion vector corresponding to the first target frame. The starting point of the first motion vector is located within the detection box of the second target frame.

[0184] In this case, the second target frame may include one or more reference frames of the first target frame.

[0185] As mentioned earlier, the second target frame may include one or more reference frames, and correspondingly, the first target initial point set may include one or more sets. These multiple sets can also be referred to as multiple first target initial point sets. For ease of description, the following explanation will use the second target frame as a reference frame of the first target frame as an example to illustrate Method 2.

[0186] The term "first" in "first motion vector" is for descriptive convenience only and has no limiting effect. Any motion vector whose starting point is within the detection box of the second target frame and whose target point is within the detection box of the first target frame can be considered as the first motion vector. The term "first" in "first prediction block" is for descriptive convenience only and has no limiting effect. The target point of the first motion vector is the first prediction block.

[0187] For example, motion vectors whose starting point is located within the detection box of the second target frame and whose target point is located within the detection box of the first target frame are selected based on the inter-frame prediction information of the first target frame (i.e., the first motion vector). The first target initial point set may include some or all of the target points (i.e., the first prediction block) of the selected motion vectors.

[0188] The target point is a prediction block in the frame to be encoded, which is the object to be predicted, and is predicted based on the corresponding block in the reference frame.

[0189] The starting point is a block or region in the reference frame, which is considered the best matching block for the "target point" in the reference frame.

[0190] Table 1 shows a set of inter-frame prediction information, including relevant data for multiple motion vectors.

[0191] Table 1

[0192]

[0193] `source` indicates the position of the reference frame. For example, "-1" indicates that the reference frame is the frame before the current frame, and "+1" indicates that the reference frame is the frame after the current frame. Taking Table 1 as an example, the current frame in Table 1 is frame 2, and its reference frame is frame 1. `blockw` indicates the width of the current block, `blockh` indicates the height of the current block, `srcx` indicates the starting x-coordinate of the reference block in the reference frame, `srcy` indicates the starting y-coordinate of the current block in the current frame, `dstx` indicates the starting x-coordinate of the current block in the current frame, and `dsty` indicates the starting y-coordinate of the current block in the current frame. `flags` are flags used to mark some special attributes or behaviors of the current block, such as whether the current block uses bidirectional prediction or intra-frame prediction. `motion_x` represents the horizontal component of the motion vector, `motion_y` represents the vertical component of the motion vector, and `motion_scale` represents the scaling factor of the motion vector.

[0194] The criteria for determining whether a predicted block is inside or outside the detection box can be set as needed. For example, if all pixels in a predicted block are inside the detection box, then the predicted block can be considered to be inside the detection box; otherwise, the predicted block can be considered to be outside the detection box. Similarly, if some pixels in a predicted block are inside the detection box, then the predicted block can be considered to be inside the detection box; otherwise, the predicted block can be considered to be outside the detection box.

[0195] External contour detection can be achieved using external contour detection algorithms, such as contour detection in OpenCV.

[0196] Optionally, the prediction information of the first target frame may include intra-frame prediction information of the first target frame, which is used to indicate the prediction information of the intra-frame prediction block of the first target frame. Step 630 may include: performing outer contour detection based on the first target initial point set to obtain the contour of the target region of the first target frame. The first target initial point set includes a second prediction block. The second prediction block belongs to the intra-frame prediction block of the first target frame. The second prediction block is located within the detection box of the first target frame.

[0197] Intra-prediction blocks can also be replaced with other description methods such as intra-pixel blocks.

[0198] The term "second" in "second prediction block" is for descriptive convenience only and has no limiting effect. Any intra-frame prediction block located within the detection box of the first target frame can be regarded as the second prediction block.

[0199] For example, intra-prediction blocks located within the detection box of the first target frame can be filtered based on the intra-prediction information of the first target frame, and the target initial point set may include some or all of the filtered intra-prediction blocks.

[0200] Figure 9 shows a schematic diagram of an intra-prediction block located within the detection frame of the first target frame, as well as a schematic diagram of intra-prediction information. As shown in Figure 9, for this prediction block, i.e., the prediction unit in Figure 9, the prediction information indicates the following: the location and index (idx) of the prediction block are 608x960 and 2419 respectively, the type of the prediction block is an intra-prediction block, and the dimensions of the prediction block are 32x32, etc. It should be understood that Figure 9 is only an example, and the prediction information may also include other content or be represented in other forms. This application embodiment does not limit this.

[0201] As mentioned above, the prediction information of the first target frame may include intra-frame prediction information and inter-frame prediction information of the first target frame. Accordingly, the initial point set of the first target may include a first prediction block and a second prediction block.

[0202] For example, taking the first target frame as a P-frame, the contour of the target region in the video can be obtained through reasoning using method 2. For instance, the prediction information of the P-frame can include intra-frame prediction information and inter-frame prediction information. Based on the prediction information of the P-frame, motion vectors whose starting point is located within the detection box of the previous frame and whose target point is located within the detection box of the current frame can be selected, as well as intra-frame prediction blocks located within the detection box of the current frame. The target points of all selected motion vectors and intra-frame prediction blocks constitute the first target initial point set. Then, based on the first target initial point set, outer contour detection is performed to obtain the contour of the target region of the frame.

[0203] Figure 10 shows a schematic diagram of the inference result of the contour of a target region according to an embodiment of this application. As shown in Figure 10, the contour of the target region in the Nth frame is inferred from the detection box of the (N-1)th frame and the prediction information of the Nth frame.

[0204] The second target frame may also include multiple reference frames of the first target frame, meaning the first target frame corresponds to multiple second target frames. In this case, candidate contours of the target region in the first target frame can be determined within the detection frames of the first target frame based on the prediction information of the first target frame and the detection boxes of each second target frame. Then, the contour of the target region in the first target frame is determined based on the multiple candidate contours. For example, the target region in the first target frame can be the union of the regions indicated by the multiple candidate contours, and the contour of this union is the contour of the target region in the first target frame.

[0205] The method for determining the candidate contour of the target region in the first target frame can refer to the method for determining the contour of the target region in the first target frame when the second target frame is a reference frame.

[0206] For example, the second target frame consists of M reference frames, where M is an integer greater than 1. Correspondingly, the first target initial point set may include M sets, i.e., M first target initial point sets, each set corresponding to one second target frame. Step 630 can be understood as: performing outer contour detection based on the M first target initial point sets respectively to obtain the contour of the target region of the first target frame.

[0207] For each second target frame, an initial set of points for the first target can be determined using the method described above, and outer contour detection can be performed based on this initial set of points. The resulting contour can be used as a candidate contour. For M reference frames, M candidate contours can be obtained.

[0208] For example, the second target frame includes reference frame #1 and reference frame #2. The first target initial point set corresponding to reference frame #1 includes a first prediction block, which belongs to the inter-frame prediction block of the first target frame. The first prediction block is located within the detection box of the first target frame. The first prediction block is the target point of the first motion vector in the motion vector corresponding to the first target frame, and the starting point of the first motion vector is located within the detection box of reference frame #1. The first target initial point set corresponding to reference frame #2 includes a first prediction block, which belongs to the inter-frame prediction block of the first target frame. The first prediction block is located within the detection box of the first target frame. The first prediction block is the target point of the first motion vector in the motion vector corresponding to the first target frame, and the starting point of the first motion vector is located within the detection box of reference frame #2.

[0209] In Method 2, the target initial point set (i.e., the first target initial point set) is determined using prediction information. For example, appropriate motion vectors are determined using inter-frame prediction information, thereby determining appropriate inter-frame prediction blocks. Appropriate intra-frame prediction blocks are determined using intra-frame prediction information, and then the outer contour of the target initial point set is detected. This is beneficial to obtain a more accurate range of the target area without introducing a large computational cost.

[0210] Method 3:

[0211] As one possible implementation, step 630 may include: determining the contour of the target region in the first target frame within the detection frame of the first target frame based on the prediction information of the first target frame and the contour of the target region in the second target frame.

[0212] Method 3 can be understood as inferring the contour from the contour, as shown in Figure 8(b), inferring the contour of the target region in the frame to be encoded based on the contour of the target region in other frames.

[0213] Optionally, the prediction information of the first target frame includes inter-frame prediction information of the first target frame. The inter-frame prediction information of the first target frame is used to indicate the prediction information of the inter-frame prediction blocks in the first target frame. Step 630 may include: performing outer contour detection based on the second target initial point set to obtain the contour of the target region of the first target frame. The second target initial point set includes a third prediction block. The third prediction block belongs to the inter-frame prediction block of the first target frame. The third prediction block is located within the detection box of the first target frame. The third prediction block is the target point of the second motion vector in the motion vector corresponding to the first target frame. The starting point of the second motion vector is located within the contour of the target region in the second target frame.

[0214] In this case, the second target frame may include one or more reference frames of the first target frame.

[0215] As mentioned earlier, the second target frame may include one or more reference frames, and correspondingly, the second target initial point set may include one or more sets. These multiple sets can also be referred to as multiple second target initial point sets. For ease of description, the following explanation will use the second target frame as a reference frame of the first target frame as an example to illustrate method 3.

[0216] The term "second" in "second motion vector" is for descriptive convenience only and has no limiting effect. Any motion vector whose starting point is located within the outline of the target region in the second target frame and whose target point is located within the detection box of the first target frame can be considered as a second motion vector. The term "third" in "third prediction block" is for descriptive convenience only and has no limiting effect. The target point of the second motion vector is the third prediction block.

[0217] For example, based on the inter-frame prediction information of the first target frame, motion vectors whose starting point is located within the outline of the target region in the second target frame and whose target points are located within the detection box of the first target frame (i.e., second motion vectors) can be selected. The initial set of second target points may include some or all of the target points (i.e., third prediction blocks) of the selected motion vectors.

[0218] The criteria for determining whether a predicted block is inside or outside the target region's outline can be set as needed. For example, if all pixels in the predicted block are inside the target region's outline, then the predicted block can be considered to be inside the target region's outline; otherwise, the predicted block can be considered to be outside the target region's outline. Similarly, if some pixels in the predicted block are inside the target region's outline, then the predicted block can be considered to be inside the target region's outline; otherwise, the predicted block can be considered to be outside the target region's outline.

[0219] The criteria for determining whether the predicted block is inside or outside the detection box can be found in the description in Method 2. To avoid repetition, they will not be repeated here.

[0220] External contour detection can be achieved using external contour detection algorithms, such as contour detection in OpenCV.

[0221] Optionally, the prediction information of the first target frame may include intra-frame prediction information of the first target frame, which is used to indicate the prediction information of the intra-frame prediction block of the first target frame. Step 630 may include: performing outer contour detection based on the second target initial point set to obtain the contour of the target region of the first target frame. The second target initial point set includes a second prediction block. The second prediction block belongs to the intra-frame prediction block of the first target frame. The second prediction block is located within the detection box of the first target frame.

[0222] The description of the second prediction block can be found in the relevant description in Method 2. To avoid repetition, it will not be repeated here.

[0223] As mentioned above, the prediction information of the first target frame may include intra-frame prediction information and inter-frame prediction information of the first target frame. Correspondingly, the initial point set of the second target may include the third prediction block and the second prediction block.

[0224] For example, taking the first target frame as a P-frame, the contour of the target region in the video can be obtained through reasoning using method 3. For instance, the prediction information of the P-frame can include intra-frame prediction information and inter-frame prediction information. Based on the prediction information of the P-frame, motion vectors whose starting point is located within the contour of the target region in the previous frame and whose target point is located within the detection box of the current frame can be selected, as well as intra-frame prediction blocks located within the detection box of the current frame. The target points of all selected motion vectors and intra-frame prediction blocks constitute the second target initial point set. Then, based on the second target initial point set, outer contour detection is performed to obtain the contour of the target region of the current frame.

[0225] Figure 11 shows a schematic diagram of the inference result of the contour of a target region according to an embodiment of this application. As shown in Figure 11, the contour of the target region in the Nth frame is inferred in the detection box of the Nth frame based on the contour of the target region in the (N-1)th frame and the prediction information of the Nth frame.

[0226] The second target frame may also include multiple reference frames of the first target frame, meaning the first target frame corresponds to multiple second target frames. In this case, candidate contours of the target region in the first target frame can be determined within the detection bounding box of the first target frame based on the prediction information of the first target frame and the contour of the target region in each of the second target frames. Then, the contour of the target region in the first target frame is determined based on the multiple candidate contours. For example, the target region in the first target frame can be the union of the regions indicated by the multiple candidate contours, and the contour of this union is the contour of the target region in the first target frame.

[0227] The method for determining the candidate contour of the target region in the first target frame can refer to the method for determining the contour of the target region in the first target frame when the second target frame is a reference frame.

[0228] For example, the second target frame consists of M reference frames, where M is an integer greater than 1. Correspondingly, the second target initial point set may include M sets, i.e., M second target initial point sets, each set corresponding to one second target frame. Step 630 can be understood as: performing outer contour detection based on the M second target initial point sets to obtain the contour of the target region in the first target frame.

[0229] For each second target frame, an initial set of points for the second target can be determined using the method described above, and outer contour detection can be performed based on this initial set of points. The resulting contour can be used as a candidate contour. For M reference frames, M candidate contours can be obtained.

[0230] For example, the second target frame includes reference frame #1 and reference frame #2. The initial point set of the second target corresponding to reference frame #1 includes a third prediction block. The third prediction block belongs to the inter-frame prediction block of the first target frame. The third prediction block is located within the detection box of the first target frame. The third prediction block is the target point of the second motion vector in the motion vector corresponding to the first target frame. The starting point of the second motion vector is located within the contour of the target region in reference frame #1. The initial point set of the second target corresponding to reference frame #2 includes a third prediction block. The third prediction block belongs to the inter-frame prediction block of the first target frame. The third prediction block is located within the detection box of the first target frame. The third prediction block is the target point of the second motion vector in the motion vector corresponding to the first target frame. The starting point of the second motion vector is located within the contour of the target region in reference frame #2.

[0231] In Method 3, the initial target point set (i.e., the second initial target point set) is determined using prediction information. For example, suitable motion vectors are determined using inter-frame prediction information, thereby determining suitable inter-frame prediction blocks. Suitable intra-frame prediction blocks are determined using intra-frame prediction information. Then, the outer contour of the initial target point set is detected, which helps to obtain a more accurate target region range without introducing significant computational overhead. Furthermore, the motion vectors selected in Method 3 originate from the contour of the target region in the reference frame. This helps to further improve the accuracy of the selected motion vectors, thereby further improving the accuracy of the initial target point set and ultimately enhancing the accuracy of the target region's contour.

[0232] Method 4:

[0233] As one possible implementation, step 630 may include: determining the contour of the target region in the first target frame based on the prediction information of the first target frame, the detection box of the second target frame, and the contour of the target region in the second target frame within the detection box of the first target frame.

[0234] For example, method 4 can be obtained by combining the above three methods. For instance, two results can be obtained based on any two of the above three methods, and these two results can be used as two initial contours of the target region in the first target frame. Then, the final contour of the target region in the first target frame can be obtained based on these two initial contours. For example, the average value of the coordinates of the two initial contours can be calculated and used as the contour of the target region in the first target frame.

[0235] Different frames in a video can use the same method to determine the outline of the target region, or they can use different methods to determine the outline of the target region.

[0236] As an example, for an I-frame, method 1 can be used to determine the outline of the target region. Further, this I-frame is any I-frame other than the first frame of the video. For P-frames and / or B-frames, method 2 or method 3 can be used to determine the outline of the target region. For instance, if the outline of the target region in the reference frame of the P-frame has been determined, method 3 can be used to determine the outline of the target region in the P-frame; if the outline of the target region in the reference frame of the P-frame has not yet been determined, method 2 can be used to determine the outline of the target region in the P-frame.

[0237] In the embodiments of this application, the contour of the target region can be inferred in a variety of ways. For example, method 1 can be used to determine the contour of the target region in the I-frame, and method 2 or method 3 can be used to determine the contour of the target region in the P-frame and / or B-frame. This is beneficial to provide a more accurate contour of the target region while ensuring inference efficiency.

[0238] In the scheme of this application embodiment, the contour of the target region of the frame to be encoded is inferred from the detection box of the frame to be encoded by using the prediction information of the frame to be encoded, the detection box information of other frames and / or the contour information of the target region in other frames. This is beneficial to provide a more accurate range of the target region, thereby improving the encoding effect, while not introducing a large computational power consumption, which is beneficial to ensuring inference efficiency.

[0239] For example, by adopting the scheme of the embodiments of this application, when encoding a non-compact detection box, the contour of the target region can be inferred within the detection box by a small amount of additional computation in the encoder, so that the encoding effect is close to the encoding effect achieved when encoding the contour of the input target region.

[0240] For example, the outline of the target region can be used for ROI encoding, which helps reduce the bitrate. For example, the outline of the target region can be used for privacy-preserving encoding, which helps provide more granular protection.

[0241] Figure 12 shows a schematic diagram of a video encoding method according to an embodiment of this application. Method 1100 can be regarded as a specific implementation of method 300. For a detailed description, please refer to method 300 or method 600. To avoid repetition, some descriptions are omitted appropriately when describing method 1100. In method 1100, the Nth frame is an example of a first target frame, and the (N-1)th frame is an example of a second target frame.

[0242] For example, the method 1100 shown in FIG12 can be performed by an encoder.

[0243] As shown in Figure 12, method 1100 may include the following steps:

[0244] 1110, Receive the Nth frame. N is an integer greater than 1.

[0245] 1120, The Nth frame is predicted by the encoder to obtain the prediction information of the Nth frame. The prediction information of the Nth frame includes the inter-frame prediction information and / or the intra-frame prediction information of the Nth frame.

[0246] 1130, Obtain the detection box of frame N and the detection box of frame N-1 and / or the outline of the target region in frame N-1.

[0247] 1140, the contour of the target region in the Nth frame is inferred from the detection box of the Nth frame by the encoder based on the prediction information of the Nth frame, the detection box of the N-1th frame and / or the contour of the target region in the N-1th frame.

[0248] 1150, The Nth frame is encoded by the encoder based on the contour of the target region in the Nth frame to obtain the bitstream of the Nth frame.

[0249] Step 1140 is described below as an example.

[0250] For example, when the Nth frame is an I-frame, an I-frame inference algorithm can be used to obtain the outline of the target region in the Nth frame. This I-frame can be any I-frame in the video other than the first frame. The I-frame inference algorithm can be implemented using method 1 in method 600.

[0251] For example, when the Nth frame is a P-frame, if the contour of the target region in the N-1th frame is not determined, a bounding box inference contour algorithm can be used; if the contour of the target region in the N-1th frame is determined, a contour inference contour algorithm can be used to obtain the contour of the target region in the Nth frame. The bounding box inference contour algorithm can be implemented using method 3 in method 600. The contour inference contour algorithm can be implemented using method 2 in method 600.

[0252] It should be understood that method 1100 is only described using the example of the Nth frame being an I-frame or a P-frame, and does not constitute a limitation on the scheme of the embodiments of this application. In other implementations, the Nth frame can also be a B-frame, and its encoding process can refer to that of a P-frame.

[0253] Figure 13 shows a schematic diagram of a reasoning profile process.

[0254] The registered variables include: the detection box S (one frame) and the contour R of the target region (one frame).

[0255] For the Nth frame, the input is the detection box S of the Nth frame. N The output is the outline R of the target region in the Nth frame. N .

[0256] For the first frame, buffers S1 to S and R. S1 is the detection bounding box of the first frame. That is, the shape of the target region in the first frame is the detection bounding box.

[0257] If the Nth frame is an I-frame, match S N And the detection box S in the (N-1)th frame N-1 For each detection box S N The contour R of the target region is obtained through the I-frame inference algorithm. N Accordingly, the shape of the target region in the Nth frame is the outline of the target region.

[0258] If the Nth frame is a P frame, match S N and S N-1 For each detection box S N If R is empty (i.e., the outline R of the target region in frame N-1 was not obtained), N-1 Then, the contour R of the target region is obtained through the bounding box inference contour algorithm. N Accordingly, the shape of the target region in the Nth frame is the outline of the target region.

[0259] If the Nth frame is a P frame, match S N and S N-1 The detection boxes, for each detection box S N If R is not empty (i.e., R has been obtained) N-1 Then, the contour R of the target region is obtained through the contour reasoning algorithm.N Accordingly, the shape of the target region in the Nth frame is the outline of the target region.

[0260] S N Cache to S, and R N Cache in R.

[0261] For example, the I-frame inference algorithm can be implemented through the following steps. The Nth frame is an I-frame that is not the first frame.

[0262] 1) Calculate the detection box S in the (N-1)th frame. N-1 And the detection box S in the Nth frame N The difference between the center coordinates (dx, dy).

[0263] 2) The contour R of the target region in the (N-1)th frame N-1 Offset (dx, dy) to obtain the contour R of the target region in the Nth frame. N .

[0264] For example, the bounding box inference contour algorithm can be implemented through the following steps. The Nth frame is the P frame.

[0265] 1) Filter MVs based on the inter-frame prediction information of the Nth frame.

[0266] Filter the starting point from frame S in the MV of frame N. N-1 In the middle, the target point is at S N MV (an example of the first motion vector) in the example.

[0267] 2) Filter the intra prediction block based on the intra prediction information of the Nth frame.

[0268] Filter the intra prediction block located at S from the Nth frame. N The intra prediction block (an example of the second prediction block).

[0269] 3) Based on the initial point set of the first target, the outer contour is obtained using the outer contour detection algorithm as R. N The initial target point set can include all the selected target points of the MV (an example of the first prediction block) and the selected intra prediction blocks.

[0270] For example, the contour reasoning algorithm can be implemented through the following steps. The Nth frame is the P frame.

[0271] 1) Filter MVs based on the inter-frame prediction information of the Nth frame.

[0272] Filter out the starting point R from the MV of frame N. N-1 In the middle, the target point is at S N MV (an example of the second motion vector) in the example.

[0273] 2) Filter the intra prediction block based on the intra prediction information of the Nth frame.

[0274] Filter the intra prediction block located at S from the Nth frame. N The intra prediction block in the text.

[0275] 3) Based on the initial point set of the second target, the outer contour is obtained using the outer contour detection algorithm as R. N The second initial target point set may include all the target points of the selected MV (an example of the third prediction block) and the selected intra prediction blocks.

[0276] In the scheme of this application embodiment, the contour information of the ROI can be inferred within the encoder using detection box information and prediction information, which is beneficial for providing a more accurate ROI range, for example, through a more accurate range of the foreground region. The inferred contour information can be used for video encoding, such as ROI encoding or privacy-preserving encoding.

[0277] Figure 14 shows a schematic diagram of the range of ROI determined by the three schemes.

[0278] As shown in Figure 14(a), when performing ROI encoding based on detection boxes, the shape of the ROI range in the first frame and subsequent frames is a rectangular box.

[0279] As shown in Figure 14(b), when using the ROI obtained by the segmentation scheme for ROI encoding, the shape of the ROI range in the first frame and subsequent frames is the outline of the ROI.

[0280] As shown in Figure 14(c), when using the scheme of this application for ROI encoding, the shape of the ROI range in the first frame can be a rectangle, and the shape of the ROI range in subsequent frames can be the outline of the ROI.

[0281] As shown in Figure 14(d), when using the privacy protection coding scheme of this application, the shape of the ROI range in the first frame can be a rectangle, and the shape of the ROI range in subsequent frames can be the outline of the ROI.

[0282] The methods of the embodiments of this application have been described in detail above. The apparatus of the embodiments of this application will now be described with reference to Figures 15 to 17. The apparatus of the embodiments of this application is capable of performing the above methods. To avoid unnecessary repetition, some descriptions have been appropriately omitted when describing the apparatus of the embodiments of this application.

[0283] Figure 15 is a schematic structural block diagram of an electronic device according to an embodiment of this application. The electronic device shown in Figure 15 can be used to perform the methods shown in Figure 3, Figure 6, or Figure 12.

[0284] For example, the electronic device 800 shown in FIG15 can be used to perform the method shown in FIG3.

[0285] The electronic device 800 shown in Figure 15 includes a first acquisition unit 801, a second acquisition unit 802, and a processing unit 803.

[0286] The first acquisition unit 801 is used to acquire the video to be encoded.

[0287] The second acquisition unit 802 is used for:

[0288] Obtain the detection bounding box of the first target frame in the video;

[0289] Obtain the detection box of the second target frame and / or the outline of the target region in the second target frame, wherein the target region in the second target frame is located within the detection box of the second target frame, and the second target frame is related to the video.

[0290] The processing unit 803 is configured to determine the outline of a target region in a first target frame based on at least two of the following: prediction information of the first target frame, the detection frame of the second target frame, or the outline of the target region in the second target frame, wherein the prediction information of the first target frame includes inter-frame prediction information and / or intra-frame prediction information of the first target frame.

[0291] The first target frame is encoded based on the contour of the target region in the first target frame.

[0292] Optionally, the second target frame includes one or more frames that are located before the first target frame in the encoding order.

[0293] Optionally, the second target frame includes one or more reference frames of the first target frame.

[0294] Optionally, the first target frame is a video frame other than the first frame in the video, and the second acquisition unit 802 is further configured to: acquire the detection box of the first frame in the video; the processing unit 803 is further configured to: encode the first frame according to the detection box of the first frame.

[0295] Optionally, the processing unit 803 is specifically used to: offset the contour of the target region in the second target frame according to the offset between the detection box of the first target frame and the detection box of the second target frame, so as to obtain the contour of the target region in the first target frame.

[0296] Optionally, the processing unit 803 is specifically configured to: perform outer contour detection based on a first target initial point set to obtain the contour of the target region in the first target frame, wherein the first target initial point set includes a first prediction block and / or a second prediction block, the inter-frame prediction information of the first target frame is used to indicate the prediction information of the inter-frame prediction block of the first target frame, the first prediction block belongs to the inter-frame prediction block of the first target frame, the first prediction block is located within the detection box of the first target frame, the first prediction block is the target point of the first motion vector in the motion vector corresponding to the first target frame, the starting point of the first motion vector is located within the detection box of the second target frame, the intra-frame prediction information of the first target frame is used to indicate the prediction information of the intra-frame prediction block of the first target frame, the second prediction block belongs to the intra-frame prediction block of the first target frame, and the second prediction block is located within the detection box of the first target frame.

[0297] Optionally, the processing unit 803 is specifically configured to: perform outer contour detection based on the second target initial point set to obtain the contour of the target region in the first target frame, wherein the second target initial point set includes a third prediction block and / or a second prediction block, the inter-frame prediction information of the first target frame is used to indicate the prediction information of the inter-frame prediction block of the first target frame, the third prediction block belongs to the inter-frame prediction block of the first target frame, the third prediction block is located within the detection box of the first target frame, the third prediction block is the target point of the second motion vector in the motion vector corresponding to the first target frame, the starting point of the second motion vector is located within the contour of the target region in the second target frame, the intra-frame prediction information of the first target frame is used to indicate the prediction information of the intra-frame prediction block of the first target frame, the second prediction block belongs to the intra-frame prediction block of the first target frame, and the second prediction block is located within the detection box of the first target frame.

[0298] It should be understood that the electronic device 800 is embodied in the form of a functional unit. The term "unit" here can refer to an application-specific integrated circuit (ASIC), electronic circuitry, a processor (e.g., a shared processor, a proprietary processor, or a group processor, etc.) and memory for executing one or more software or firmware programs, integrated logic circuitry, and / or other suitable components supporting the described functions. In an alternative example, those skilled in the art will understand that the electronic device 800 can be used to perform various processes and / or steps corresponding to the video encoding method in the above method embodiments; to avoid repetition, these will not be described further here.

[0299] Figure 16 is a schematic diagram of a computer device provided in an embodiment of this application. As shown in Figure 16, the computer device 900 includes a processor 901, which executes computer programs or instructions stored in a memory 902, or reads data / signaling stored in the memory 902, to perform the methods in the above-described method embodiments. Optionally, there may be one or more processors 901.

[0300] Optionally, as shown in FIG16, the computer device 900 further includes a memory 902 for storing computer programs or instructions and / or data. The memory 902 may be integrated with the processor 901 or may be separately configured. Optionally, there may be one or more memories 902.

[0301] Optionally, as shown in FIG16, the computer device 900 further includes a transceiver 903 for receiving and / or transmitting signals. For example, the processor 901 is used to control the transceiver 903 to receive and / or transmit signals.

[0302] It should be understood that the processor mentioned in the embodiments of this application can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0303] It should also be understood that the memory mentioned in the embodiments of this application can be volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM). For example, RAM can be used as an external cache. By way of example and not limitation, RAM includes the following forms: static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).

[0304] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, the memory (storage module) can be integrated into the processor.

[0305] It should also be noted that the memory described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0306] Figure 17 is a schematic diagram of a chip system 1000 provided in an embodiment of this application. The chip system 1000 (or may also be called a processing system) includes logic circuits 1001 and input / output interface 1002.

[0307] The logic circuit 1001 can be a processing circuit in the chip system 1000. The logic circuit 1001 can be coupled to a memory unit, calling instructions from the memory unit, enabling the chip system 1000 to implement the methods and functions of the embodiments of this application. The input / output interface 1002 can be an input / output circuit in the chip system 1000, outputting processed information from the chip system 1000, or inputting data or signaling information to be processed into the chip system 1000 for processing.

[0308] As one approach, the chip system 1000 is used to implement the operations described in the various method embodiments above.

[0309] For example, logic circuit 1001 is used to implement the relevant operations in the above method embodiments; input / output interface 1002 is used to implement the sending and / or receiving related operations in the above method embodiments.

[0310] This application also provides a computer-readable storage medium storing computer instructions for implementing the methods in the above-described method embodiments.

[0311] For example, when the computer program is executed by a computer, it enables the computer to implement the methods in the various embodiments described above.

[0312] This application also provides a computer program product comprising instructions which, when executed by a computer, implement the methods described in the above-described method embodiments.

[0313] The explanations and beneficial effects of the relevant contents in any of the devices provided above can be referred to the corresponding method embodiments provided above, and will not be repeated here.

[0314] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of apparatus or units may be electrical, mechanical, or other forms.

[0315] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. For example, the computer can be a personal computer, a server, or a network device, etc. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks, SSDs). For example, the aforementioned available media include, but are not limited to, USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, and other media capable of storing program code.

[0316] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A video encoding method, characterized in that, include: Obtain the video to be encoded; Obtain the detection bounding box of the first target frame in the video; Obtain the detection box of the second target frame and / or the outline of the target region in the second target frame, wherein the target region in the second target frame is located within the detection box of the second target frame, and the second target frame is related to the video; determine the outline of the target region in the first target frame in the detection box of the first target frame based on at least two of the following: prediction information of the first target frame, the detection box of the second target frame, or the outline of the target region in the second target frame, wherein the prediction information of the first target frame includes inter-frame prediction information and / or intra-frame prediction information of the first target frame; encode the first target frame based on the outline of the target region in the first target frame.

2. The method according to claim 1, characterized in that, The second target frame includes one or more frames that are encoded before the first target frame.

3. The method according to claim 2, characterized in that, The second target frame includes one or more reference frames of the first target frame.

4. The method according to any one of claims 1 to 3, characterized in that, The first target frame is a video frame other than the first frame in the video, and the method further includes: obtaining a detection box of the first frame in the video; and encoding the first frame according to the detection box of the first frame.

5. The method according to any one of claims 1 to 4, characterized in that, The step of determining the contour of the target region in the first target frame in the detection box of the first target frame based on at least two of the following: prediction information of the first target frame, the detection box of the second target frame, or the contour of the target region in the second target frame, includes: offsetting the contour of the target region in the second target frame according to the offset between the detection box of the first target frame and the detection box of the second target frame to obtain the contour of the target region in the first target frame.

6. The method according to any one of claims 1 to 4, characterized in that, The step of determining the contour of the target region in the first target frame within the detection frame of the first target frame based on at least two of the following: prediction information of the first target frame, the detection frame of the second target frame, or the contour of the target region in the second target frame, includes: performing outer contour detection based on a first target initial point set to obtain the contour of the target region in the first target frame, wherein the first target initial point set includes a first prediction block and / or a second prediction block, the inter-frame prediction information of the first target frame is used to indicate the prediction information of the inter-frame prediction block of the first target frame, the first prediction block belongs to the inter-frame prediction block of the first target frame, the first prediction block is located within the detection frame of the first target frame, the first prediction block is the target point of a first motion vector in the motion vector corresponding to the first target frame, the starting point of the first motion vector is located within the detection frame of the second target frame, the intra-frame prediction information of the first target frame is used to indicate the prediction information of the intra-frame prediction block of the first target frame, the second prediction block belongs to the intra-frame prediction block of the first target frame, and the second prediction block is located within the detection frame of the first target frame.

7. The method according to any one of claims 1 to 4, characterized in that, The step of determining the contour of the target region in the first target frame within the detection frame of the first target frame based on at least two of the following: prediction information of the first target frame, the detection frame of the second target frame, or the contour of the target region in the second target frame, includes: performing outer contour detection based on a second target initial point set to obtain the contour of the target region in the first target frame, wherein the second target initial point set includes a third prediction block and / or a second prediction block, the inter-frame prediction information of the first target frame is used to indicate the prediction information of the inter-frame prediction block of the first target frame, the third prediction block belongs to the inter-frame prediction block of the first target frame, the third prediction block is located within the detection frame of the first target frame, the third prediction block is the target point of the second motion vector in the motion vector corresponding to the first target frame, the starting point of the second motion vector is located within the contour of the target region in the second target frame, the intra-frame prediction information of the first target frame is used to indicate the prediction information of the intra-frame prediction block of the first target frame, the second prediction block belongs to the intra-frame prediction block of the first target frame, and the second prediction block is located within the detection frame of the first target frame.

8. A video encoding apparatus, characterized in that, include: The first acquisition unit is used to acquire the video to be encoded. The second acquisition unit is used to: acquire the detection box of the first target frame in the video; A processing unit is configured to: obtain a detection box of a second target frame and / or the outline of a target region in the second target frame, wherein the target region in the second target frame is located within the detection box of the second target frame, and the second target frame is related to the video; and a processing unit is configured to: determine the outline of a target region in the first target frame in the detection box of the first target frame based on at least two of the following: prediction information of the first target frame, the detection box of the second target frame, or the outline of the target region in the second target frame, wherein the prediction information of the first target frame includes inter-frame prediction information and / or intra-frame prediction information of the first target frame; and encode the first target frame based on the outline of the target region in the first target frame.

9. The apparatus according to claim 8, characterized in that, The second target frame includes one or more frames that are encoded before the first target frame.

10. The apparatus according to claim 9, characterized in that, The second target frame includes one or more reference frames of the first target frame.

11. The apparatus according to any one of claims 8 to 10, characterized in that, The first target frame is a video frame other than the first frame in the video, and the second acquisition unit is further configured to: acquire the detection box of the first frame in the video; the processing unit is further configured to: encode the first frame according to the detection box of the first frame.

12. The apparatus according to any one of claims 8 to 11, characterized in that, The processing unit is specifically used to: offset the contour of the target region in the second target frame according to the offset between the detection box of the first target frame and the detection box of the second target frame, so as to obtain the contour of the target region in the first target frame.

13. The apparatus according to any one of claims 8 to 11, characterized in that, The processing unit is specifically used to: perform outer contour detection based on a first target initial point set to obtain the contour of the target region in the first target frame, wherein the first target initial point set includes a first prediction block and / or a second prediction block, the inter-frame prediction information of the first target frame is used to indicate the prediction information of the inter-frame prediction block of the first target frame, the first prediction block belongs to the inter-frame prediction block of the first target frame, the first prediction block is located within the detection box of the first target frame, the first prediction block is the target point of the first motion vector in the motion vector corresponding to the first target frame, the starting point of the first motion vector is located within the detection box of the second target frame, the intra-frame prediction information of the first target frame is used to indicate the prediction information of the intra-frame prediction block of the first target frame, the second prediction block belongs to the intra-frame prediction block of the first target frame, and the second prediction block is located within the detection box of the first target frame.

14. The apparatus according to any one of claims 8 to 11, characterized in that, The processing unit is specifically used to: perform outer contour detection based on a second target initial point set to obtain the contour of the target region in the first target frame, wherein the second target initial point set includes a third prediction block and / or a second prediction block, the inter-frame prediction information of the first target frame is used to indicate the prediction information of the inter-frame prediction block of the first target frame, the third prediction block belongs to the inter-frame prediction block of the first target frame, the third prediction block is located within the detection box of the first target frame, the third prediction block is the target point of the second motion vector in the motion vector corresponding to the first target frame, the starting point of the second motion vector is located within the contour of the target region in the second target frame, the intra-frame prediction information of the first target frame is used to indicate the prediction information of the intra-frame prediction block of the first target frame, the second prediction block belongs to the intra-frame prediction block of the first target frame, and the second prediction block is located within the detection box of the first target frame.

15. A computer device, characterized in that, include: A processor configured to be coupled to a memory, read and execute instructions and / or program code in the memory to perform the method as described in any one of claims 1 to 7.

16. A chip system, characterized in that, include: A logic circuit for coupling with an input / output interface, through which data is transmitted to perform the method as described in any one of claims 1 to 7.

17. A computer-readable medium, characterized in that, The computer-readable medium stores program code that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 7.