Video Encoding Method, Apparatus, Electronic Device, and Computer-Readable Medium

By obtaining the image sequence of the target video, determining the structured description information and generating background frame images, combining video intelligent analysis technology, selecting the macroblock with the smallest pixel value for encoding, solving the problem of low unit encoding efficiency of macroblocks and improving the video compression coding efficiency.

CN116016934BActive Publication Date: 2025-07-29GUANGDONG VIMICRO +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310018730.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-06
Publication Date
2025-07-29
Estimated Expiration
2043-01-06

AI Technical Summary

Technical Problem

When existing video encoding is performed in macroblocks, it is difficult to accurately describe the background and macromotion of the target object, resulting in low video encoding efficiency and thus requiring more network resources.

Method used

By obtaining the image sequence of the target video, determining the structured description information, generating background frame images and structured reconstruction frame images, combining video intelligent analysis technology, selecting the macroblock with the smallest pixel value as the prediction macroblock for encoding, and generating a video code stream.

Benefits of technology

It improves the video compression and coding efficiency in complex scenarios, and reduces the computing volume and network resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116016934B_ABST
    Figure CN116016934B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a video encoding method, apparatus, electronic device, and computer-readable medium. A specific implementation of the method includes: obtaining a target image sequence; for each target image, performing the following first encoding steps: determining structured description information; generating a background frame image according to at least one historical target image; generating a background reconstructed frame image; generating a structured reconstructed frame image according to the background reconstructed frame image, at least one historical reconstructed frame image, and the structured description information; for each target macroblock, performing the following second encoding steps: determining a first macroblock; generating a first macroblock bitstream; determining a second macroblock; generating a second macroblock bitstream; determining the first macroblock as a prediction macroblock; encoding the target image according to the prediction macroblock set to obtain a first target image bitstream; generating a video bitstream according to the first target image bitstream set. This implementation can improve the video compression encoding efficiency in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, and more particularly, to video encoding methods, apparatuses, electronic devices, and computer-readable media. Background Art

[0002] Due to the large amount of information, video signals have certain requirements for the transmission network bandwidth, which brings great pressure to both transmission and storage. Therefore, video encoding is required in practical applications. Video encoding technology is a technology that reduces the volume or bit rate of video data without significantly affecting video quality. For video encoding, the commonly used method is to perform video compression encoding by removing spatial redundancy, temporal redundancy, and other redundant information in the video frame sequence in units of macroblocks.

[0003] However, the inventors found that when the above method is used to encode videos, the following technical problems often exist:

[0004] Encoding videos in units of macroblocks is difficult to accurately describe the macroscopic motion of the background and target objects, resulting in low video encoding efficiency, and thus more network resources are required.

[0005] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention

[0006] This content part of the present disclosure is used to briefly introduce concepts that will be described in detail in the following detailed implementation part. This content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0007] Some embodiments of the present disclosure propose video encoding methods, apparatuses, electronic devices, and computer-readable media to solve one or more of the technical problems mentioned in the above background art section.

[0008] In a first aspect, some embodiments of the present disclosure provide a video encoding method, including: obtaining a target image sequence for a target video; for each target image in the target image sequence, performing the following first encoding step: in response to determining that the target image is not the first-frame target image, determining structured description information corresponding to the target image according to the target image and at least one historical target image, where the structured description information includes: background change description information of the target image and foreground target object description information of the target image, and the historical target image in the at least one historical target image is an image located before the target image; generating a background frame image according to the at least one historical target image; generating a background reconstructed frame image according to the background frame image; generating a structured reconstructed frame image corresponding to the target image according to the background reconstructed frame image, at least one historical reconstructed frame image, and the structured description information, where the historical reconstructed frame image in the at least one historical reconstructed frame image is an image obtained by encoding and decoding and reconstructing the historical target image; for each target macroblock in the target image, performing the following second encoding step: determining a macroblock in the structured reconstructed frame image that has a position correspondence with the target macroblock as a first macroblock; generating a first macroblock bitstream according to the first macroblock and the target macroblock; determining a macroblock in a pre-obtained basic prediction image set that has a position correspondence with the target macroblock as a second macroblock; generating a second macroblock bitstream according to the second macroblock and the target macroblock; in response to determining that the length corresponding to the first macroblock bitstream is less than or equal to the length corresponding to the second macroblock bitstream, determining the first macroblock as a prediction macroblock; encoding the target image according to the obtained prediction macroblock set to obtain a first target image bitstream, where the image parameter set in the first target image bitstream includes parameter information related to structured information; generating a video bitstream according to the obtained first target image bitstream set.

[0009] Second aspect, some embodiments of the present disclosure provide a video encoding device, including: an acquisition unit configured to acquire a target image sequence for a target video; an execution unit configured to, for each target image in the target image sequence, perform the following first encoding steps: in response to determining that the target image is not the first-frame target image, determine structured description information corresponding to the target image according to the target image and at least one historical target image, where the structured description information includes: background change description information of the target image and foreground target object description information of the target image, and the historical target images in the at least one historical target image are images located before the target image; generate a background frame image according to the at least one historical target image; generate a background reconstruction frame image according to the background frame image; generate a structured reconstruction frame image corresponding to the target image according to the background reconstruction frame image, at least one historical reconstruction frame image, and the structured description information, where the historical reconstruction frame images in the at least one historical reconstruction frame image are images obtained by encoding and decoding and reconstructing historical target images; for each target macroblock in the target image, perform the following second encoding steps: determine a macroblock in the structured reconstruction frame image that has a position correspondence with the target macroblock as the first macroblock; generate a first macroblock bitstream according to the first macroblock and the target macroblock; determine a macroblock in a pre-obtained basic prediction image set that has a position correspondence with the target macroblock as the second macroblock; generate a second macroblock bitstream according to the second macroblock and the target macroblock; in response to determining that the length of the first macroblock bitstream is less than or equal to the length of the second macroblock bitstream, determine the first macroblock as the prediction macroblock; encode the target image according to the obtained prediction macroblock set to obtain a first target image bitstream, where the image parameter set in the first target image bitstream includes parameter information related to structured information; a generation unit configured to generate a video bitstream according to the obtained first target image bitstream set.

[0010] Third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any implementation manner in the first aspect.

[0011] Fourth aspect, some embodiments of the present disclosure provide a computer-readable medium storing a computer program thereon, where when the program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.

[0012] The above-described various embodiments of the present disclosure have the following beneficial effects: The video encoding methods of some embodiments of the present disclosure can combine video intelligent analysis technology and macroblock-based video encoding technology, improving the video compression encoding efficiency in complex scenarios. Specifically, the reason for the low video encoding efficiency is that video encoding in units of macroblocks is difficult to accurately describe the macroscopic motion of the background and target objects, resulting in low video encoding efficiency and thus requiring more network resources. Based on this, the video encoding methods of some embodiments of the present disclosure can, first, obtain a target image sequence for the target video. Here, the obtained target encoding image sequence is used for subsequent encoding of the target image. Second, for each target image in the above target image sequence, perform the following first encoding step: In response to determining that the above target image is not the first frame target image, determine the structured description information corresponding to the above target image according to the above target image and at least one historical target image, where the above structured description information includes: background change description information of the target image and foreground target object description information of the target image, and the historical target image in the above at least one historical target image is an image located before the above target image. Here, the obtained structured description information is used for subsequent determination of the structured reconstruction frame image corresponding to the target image. Then, generate a background frame image according to the above at least one historical target image. Generate a background reconstruction frame image according to the above background frame image. Here, the obtained background reconstruction frame is used for subsequent determination of the structured reconstruction frame image. Generate the structured reconstruction frame image corresponding to the above target image according to the above background reconstruction frame image, at least one historical reconstruction frame image, and the above structured description information, where the historical reconstruction frame image in the above at least one historical reconstruction frame image is an image obtained by encoding and decoding reconstruction of the historical target image. Here, the obtained structured reconstruction frame image is used for subsequent determination of the predicted macroblock corresponding to the position of the macroblock in the above target image. Since the structured reconstruction frame image is obtained by filling the pixels of the background reconstruction image and at least one historical reconstruction frame image through the above structured description information, when determining the pixel value of the target macroblock corresponding to the position of the above target image, the amount of calculation is reduced, and the speed of image encoding is improved. Then, for each target macroblock in the above target image, perform the following second encoding step: Determine the macroblock in the above structured reconstruction frame image that has a position correspondence with the above target macroblock as the first macroblock. Generate a first macroblock bitstream according to the above first macroblock and the above target macroblock. Here, using the first macroblock as the predicted macroblock reduces the amount of calculation for determining the pixels of the target macroblock and improves the rate of generating the first macroblock bitstream. Determine the macroblock in the pre-obtained basic prediction image set that has a position correspondence with the above target macroblock as the second macroblock. Generate a second macroblock bitstream according to the above second macroblock and the above target macroblock.In response to determining that the length of the above-mentioned first macroblock bitstream is less than or equal to the length of the above-mentioned second macroblock bitstream, the above-mentioned first macroblock is determined as the predicted macroblock. Here, through macroblock-level pixel comparison, the macroblock with the smallest pixel value is selected as the predicted macroblock to achieve local optimality in video coding, which is beneficial to improving the speed of image coding. Finally, according to the obtained set of predicted macroblocks, the above-mentioned target image is encoded to obtain a first target image bitstream, where the image parameter set in the above-mentioned first target image bitstream includes parameter information related to structured information. According to the obtained set of first target image bitstreams, a video bitstream is generated. Here, the macroblocks obtained from the structured reconstruction frame are combined with the macroblocks obtained from the basic predicted image, and the macroblock with the smallest corresponding macroblock bitstream is selected as the predicted macroblock to improve the coding efficiency. Thus, it can be seen that this video coding method can combine video intelligent analysis technology and macroblock-based video coding technology, improving the video compression coding efficiency in complex scenarios. Description of the Drawings

[0013] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.

[0014] Figure 1 is a flowchart of some embodiments of a video coding method according to the present disclosure;

[0015] Figure 2 is a schematic structural diagram of some embodiments of a video coding apparatus according to the present disclosure;

[0016] Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Description of the Embodiments

[0017] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not used to limit the protection scope of the present disclosure.

[0018] It should also be noted that, for the sake of convenience of description, only the parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0019] It should be noted that the concepts such as "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or mutual dependence relationship of the functions performed by these devices, modules or units.

[0020] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".

[0021] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0022] The following will detail this disclosure with reference to the accompanying drawings and in conjunction with embodiments.

[0023] Figure 1 Flow 100 of some embodiments of a video encoding method according to this disclosure is shown. The video encoding method includes the following steps:

[0024] Step 101, obtain a target image sequence for a target video.

[0025] In some embodiments, the execution subject (e.g., an electronic device) of the above video encoding method can obtain the target image sequence for the target video through a wired connection method or a wireless connection method. Among them, the above target video can be a video that needs to be encoded. The above target image sequence can be the image sequence in the above target video.

[0026] Step 102, for each target image in the target image sequence, perform the following first encoding step:

[0027] Step 1021, in response to determining that the target image is not the first-frame target image, determine the structured description information corresponding to the target image according to the target image and at least one historical target image.

[0028] In some embodiments, the above-mentioned execution entity may, in response to determining that the above-mentioned target image is not the first-frame target image, determine the structured description information corresponding to the above-mentioned target image according to the above-mentioned target image and at least one historical target image. Among them, the above-mentioned structured description information includes: background change description information of the target image and foreground target object description information of the target image. The historical target images in the above-mentioned at least one historical target image are images located before the above-mentioned target image. The background change description information of the target image may be position parameter information characterizing the background change in the target image. For example, the background change description information of the target image may include: {Type: global motion; Moving direction: moving 1.023 pixels in the positive direction of the abscissa, moving 3.231 pixels in the negative direction of the ordinate, and rotating 5.1 degrees in the positive direction with the pixel at (56, 78) as the center}. The foreground target description information of the above-mentioned target image may be position change information characterizing the foreground target object in the above-mentioned target image. For example, the foreground target description information of the above-mentioned target image may include: {Target object 1: vehicle; Dataset of boundary coordinates of the previous frame image: [(26, 46), (+1, 0), (+1, 0), (+1, 0)...(0, -1), (0, -1), (0, -1)]; Moving direction: moving 1.5 pixels in the positive direction of the abscissa, moving 3.2 pixels in the negative direction of the ordinate}.

[0029] As an example, the above-mentioned execution entity may determine the structured description information corresponding to the above-mentioned target image according to the above-mentioned target image and at least one historical target image through a video intelligent analysis algorithm. Among them, the above-mentioned video intelligent analysis algorithm may be an algorithm that divides an image into a background and a foreground. For example, the above-mentioned video intelligent analysis algorithm may be an optical flow method. The above-mentioned video intelligent analysis algorithm may also be an artificial intelligence algorithm.

[0030] Optionally, before the above-mentioned step of, in response to determining that the above-mentioned target image is not the first-frame target image, determining the structured description information corresponding to the above-mentioned target image according to the above-mentioned target image and at least one historical target image, the above-mentioned execution entity may further perform the following steps:

[0031] The above-mentioned execution entity may, in response to determining that the above-mentioned target image is the first-frame target image, encode the above-mentioned first-frame target image to obtain a first-frame target image bitstream.

[0032] As an example, the above-mentioned execution entity may perform the following encoding steps for each macroblock to be encoded in the above-mentioned first-frame target image: First step, determine the predicted macroblock of the macroblock to be encoded. Among them, the above-mentioned predicted macroblock may be the encoded macroblock directly above the macroblock to be encoded. The above-mentioned predicted macroblock may also be the encoded macroblock to the left of the macroblock to be encoded. Second step, determine the pixel difference between the above-mentioned macroblock to be encoded and the predicted macroblock. Fourth step, perform transformation, quantization, and entropy encoding on the above-mentioned pixel difference to obtain the bitstream of the first-frame target image.

[0033] Step 1022, generate a background frame image according to at least one historical target image.

[0034] In some embodiments, the above-mentioned execution entity may determine a background frame image according to the above-mentioned at least one historical target image. Among them, the above-mentioned background frame image may be an image representing the background pixel information of at least one frame of images in the target image sequence. After the key frame image is encoded and output, the background frame image is encoded and output again. The at least one frame of images may be multiple images from the first frame of image to the next key frame image. The key frame image may be an image obtained by performing intra-frame encoding on the target image in the target image sequence. In the present disclosure, by extending the nal_unit_type syntax of the NAL (Network Abstract Layer) unit, when the value of nal_unit_type is 16, the target image is determined as the background frame image.

[0035] As an example, the above-mentioned execution entity may use a video intelligent analysis algorithm to perform background extraction on the above-mentioned at least one historical target image to obtain a background frame image.

[0036] Step 1023, generate a background reconstructed frame image according to the background frame image.

[0037] In some embodiments, the above-mentioned execution entity may generate a background reconstructed frame image according to the background frame image. Among them, the above-mentioned background reconstructed frame image may be an image obtained by encoding, decoding, and reconstructing the background frame image.

[0038] As an example, the above-mentioned execution entity may encode the above-mentioned background frame image to obtain an encoded background frame image. Secondly, decode the above-mentioned encoded background frame image to obtain a decoded background frame image as the background reconstructed frame image.

[0039] Step 1024, generate a structured reconstructed frame image corresponding to the target image according to the background reconstructed frame image, at least one historical reconstructed frame image, and structured description information.

[0040] In some embodiments, the above-mentioned execution entity may generate a structured reconstruction frame image corresponding to the above-mentioned target image according to the above-mentioned background reconstruction frame image, at least one historical reconstruction frame image, and the above-mentioned structured description information. Among them, the above-mentioned structured reconstruction frame image may be an image obtained by filling the original pixels of the target object in the historical reconstruction frame image after the target object moves in the historical reconstruction frame image with the background reconstruction frame. The above-mentioned at least one historical reconstruction frame image is an image obtained by lossy encoding and decoding reconstruction of a historical target image.

[0041] In some alternative implementation manners of some embodiments, the generating, according to the above-mentioned background reconstruction frame image, at least one historical reconstruction frame image, and the above-mentioned structured description information, a structured reconstruction frame image corresponding to the above-mentioned target image may include the following steps:

[0042] First step, the above-mentioned execution entity may determine the background change information in the above-mentioned target image and the position change information of the foreground target object in the above-mentioned target image according to the above-mentioned structured description information.

[0043] As an example, the above-mentioned execution entity may determine the description information of the foreground target object from the historical reconstruction frame image to the above-mentioned target image and the description information of the background change in the above-mentioned target image through the above-mentioned structured description information, so as to determine the position change information of the foreground target object in the above-mentioned reconstruction frame image and the background change information in the above-mentioned target image.

[0044] Second step, the above-mentioned execution entity may generate a structured reconstruction frame image corresponding to the above-mentioned target image according to the above-mentioned background reconstruction frame image, the above-mentioned at least one historical reconstruction frame image, the above-mentioned position change information, and the above-mentioned background change information.

[0045] As an example, the above-mentioned execution entity may first determine the position information of the pixels that need to be filled after the background change and the movement of the foreground target object in the above-mentioned at least one historical reconstruction frame image through the above-mentioned position change information. Finally, fill the pixels after the movement of the foreground target object in the above-mentioned historical reconstruction frame image with the pixels of the above-mentioned background reconstruction frame image to obtain the structured reconstruction frame image corresponding to the above-mentioned target image.

[0046] Step 1025, for each target macroblock in the above-mentioned target image, perform the following second encoding step:

[0047] Step 10251, determine that the macroblock in the structured reconstruction frame image that has a position correspondence relationship with the target macroblock is the first macroblock.

[0048] In some embodiments, the aforementioned execution entity may determine that the macroblock in the aforementioned structured reconstructed frame image that has a positional correspondence with the aforementioned target macroblock is the first macroblock. The first macroblock has a one-to-one positional correspondence with the target macroblock. For example, if the first macroblock is the macroblock located in the upper left of the aforementioned structured reconstructed frame image, then the target macroblock is also the macroblock located in the upper left of the aforementioned target image.

[0049] Step 10252: Generate a first macroblock bitstream based on the first macroblock and the target macroblock.

[0050] In some embodiments, the aforementioned execution entity may generate a first macroblock bitstream based on the aforementioned first macroblock and the aforementioned target macroblock. The first macroblock bitstream may be the bitstream obtained by encoding the target macroblock through the first macroblock.

[0051] As an example, the aforementioned execution entity may first determine the pixel difference between the aforementioned first macroblock and the aforementioned target macroblock. Then, perform transform quantization and entropy coding on the aforementioned pixel difference to obtain the first macroblock bitstream.

[0052] Step 10253: Determine that the macroblock in the pre-obtained basic prediction image set that has a positional correspondence with the target macroblock is the second macroblock.

[0053] In some embodiments, the above-mentioned execution entity may determine that the macroblock in the pre-obtained basic prediction image set that has a positional correspondence with the above-mentioned target macroblock is the second macroblock. The pre-obtained basic prediction image set includes an intra-prediction image and an inter-prediction image set obtained by a search algorithm. The intra-prediction image may be an image obtained by encoding the target macroblock using an adjacent macroblock set that has been encoded in the target reconstruction image. The adjacent macroblock set that has been encoded may be an encoded macroblock that has an adjacent relationship with the above-mentioned target macroblock in terms of position. The above-mentioned adjacent relationship in terms of position may be the macroblocks located above, to the left, top-left, top-left and top-right of the target macroblock. The second macroblock may be obtained through the following steps: First step, determine the length of the code stream corresponding to the macroblock adjacent to the target macroblock in the intra-prediction image and obtained by encoding the target macroblock, as the first length value. Second step, determine the length of the code stream corresponding to the macroblock in the above-mentioned inter-prediction image that has a positional correspondence with the above-mentioned second macroblock and obtained by encoding the target macroblock, as the second length value. Third step, when the above-mentioned first length value is less than or equal to the above-mentioned second length value, determine the macroblock with the smallest length of the macroblock code stream corresponding to the macroblock code stream set obtained by encoding the target macroblock using the adjacent macroblock set in the intra-prediction image as the second macroblock. Fourth step, when the above-mentioned first pixel difference is greater than the above-mentioned second pixel difference, determine the macroblock with the smallest length of the macroblock code stream corresponding to the macroblock code stream set obtained by encoding the target macroblock using the macroblocks with positional correspondence in the inter-prediction image set as the second macroblock. Among them, the macroblocks with positional correspondence may be the macroblocks located in the inter-prediction image set and having the same coordinates as the target macroblock.

[0054] Step 10254, generate a first macroblock code stream according to the first macroblock and the target macroblock.

[0055] In some embodiments, the above-mentioned execution entity may generate a second macroblock code stream according to the above-mentioned second macroblock and the above-mentioned target macroblock. Among them, the second macroblock code stream may be a code stream obtained by encoding the target macroblock through the second macroblock.

[0056] As an example, the above-mentioned execution entity may first determine the pixel difference between the above-mentioned second macroblock and the above-mentioned target macroblock. Then, perform transform quantization and entropy coding on the above-mentioned pixel difference to obtain the second macroblock code stream.

[0057] Step 10255, in response to determining that the length of the first macroblock code stream is less than or equal to the length of the second macroblock code stream, determine the first macroblock as the predicted macroblock.

[0058] In some embodiments, the execution entity may determine the first macroblock as the predicted macroblock in response to determining that the corresponding length of the first macroblock bitstream is less than or equal to the corresponding length of the second macroblock bitstream. Wherein, the corresponding length of the first macroblock bitstream may be the amount of data transmitted per unit time in the first macroblock bitstream. The corresponding length of the second macroblock bitstream may be the amount of data transmitted per unit time in the second macroblock bitstream.

[0059] Optionally, after determining the first macroblock as the predicted macroblock in response to determining that the corresponding length of the first macroblock bitstream is less than or equal to the corresponding length of the second macroblock bitstream, the method may further include the following steps:

[0060] The execution entity may determine the second macroblock as the predicted macroblock in response to determining that the corresponding length of the first macroblock bitstream is greater than the corresponding length of the second macroblock bitstream.

[0061] Step 1026, encode the target image according to the obtained set of predicted macroblocks to obtain a first target image bitstream.

[0062] In some embodiments, the execution entity may encode the target image according to the obtained set of predicted macroblocks to obtain a first target image bitstream. Wherein, the image parameter set in the first target image bitstream includes parameter information related to the structured information. The obtained set of predicted macroblocks may include: macroblocks located in the structured reconstruction frame image and having a position correspondence with the target macroblock, and macroblocks located in the basic prediction image set and having a position correspondence with the target macroblock. The first target image bitstream may be a bitstream obtained by encoding each macroblock in the target image with the pixels of each predicted macroblock in the set of predicted macroblocks as reference pixels. The parameter information related to the structured description information may include at least one of the following: the writing language type of the structured description information, the data length of the structured description information, and the data of the structured description information. The description type of the structured description information may be a type written in XML (Extensible Markup Language) language when ai_data_type is 0. The description type of the structured description information may also be a type without writing language restrictions when ai_data_type is other values.

[0063] As an example, the execution entity may first determine the predicted macroblocks having a position correspondence with each macroblock in the target image. Secondly, determine a third pixel difference set between each macroblock in the target image and the set of predicted macroblocks. Finally, perform transform quantization and entropy coding on the third pixel difference set to obtain a first target image bitstream.

[0064] In some alternative implementations of some embodiments, encoding the target image according to the obtained predicted macroblock set to obtain a first target image bitstream may include the following steps:

[0065] First, the execution subject may determine a plurality of predicted pixel values corresponding to the target image according to the obtained predicted macroblock set.

[0066] As an example, the execution subject may predict each target macroblock corresponding to the predicted macroblock in the predicted macroblock set to obtain a plurality of predicted pixel values corresponding to each macroblock in the target image.

[0067] Second, the execution subject may determine a plurality of differences between the plurality of predicted pixel values and the plurality of pixel values of the target image.

[0068] Third, the execution subject may encode the plurality of differences to obtain a first target image bitstream. The encoding may be transform quantization and entropy encoding of the plurality of differences.

[0069] Step 103, generating a video bitstream according to the obtained first target image bitstream set.

[0070] In some embodiments, the execution subject may generate a video image bitstream according to the obtained image bitstream set. Among them, the video bitstream may be a bitstream obtained by encoding each image in the target video.

[0071] As an example, the execution subject may splice the obtained image bitstream set to obtain a video bitstream.

[0072] Optionally, after encoding the target image according to the obtained predicted macroblock set to obtain a first target image bitstream, it includes:

[0073] First, the execution subject may encode the target image according to the obtained second macroblock set to obtain a second target image bitstream.

[0074] As an example, the execution subject may first determine the second macroblock corresponding to each macroblock in the target image. Secondly, determine a fourth pixel difference set between each macroblock in the target image and the second macroblock. Finally, perform transform quantization and entropy encoding on the fourth pixel difference set to obtain a second target image bitstream.

[0075] Second step, in response to determining that the length of the first target image bitstream is less than or equal to the length of the second target image bitstream, the execution entity may determine the first target image bitstream as the bitstream corresponding to the target image. The length of the first target image bitstream may be the amount of data transmitted per unit time of the first target image bitstream. The length of the second target image bitstream may be the amount of data transmitted per unit time of the second target image bitstream.

[0076] Third step, in response to determining that the length of the first target image bitstream is greater than the length of the second target image bitstream, the execution entity may determine the second target image bitstream as the bitstream corresponding to the target image.

[0077] Optionally, after step 103, the execution entity may further perform the following steps:

[0078] First step, obtain the image parameter set in the video bitstream. The image parameter set may be the relevant parameter information of the image sequence of the target video. For example, the image parameter set may include at least one of the following: image type, serial number of the image, and entropy coding type.

[0079] Second step, in response to determining that the image parameter set includes parameter information related to the background frame image, perform syntax parsing on the background frame image to obtain a background reconstructed frame image. The parameter information related to the background frame image may include at least one of the following: image type of the background frame and coding method of the background frame. The image type of the background frame may be an I-frame. The coding method of the background frame may be macroblock-level coding. Third step, in response to determining that the image parameter set includes parameter information related to the structured description information, generate the structured reconstructed frame image according to the structured description information, the background reconstructed frame image, and the at least one historical reconstructed frame image.

[0080] As an example, the execution entity may first determine, through the position change information, the position information of the pixels that need to be filled after the foreground target object moves in the historical reconstructed frame image. Finally, fill the pixels of the foreground target object after moving in the historical reconstructed frame image with the pixels of the background reconstructed frame image to obtain the structured reconstructed frame image corresponding to the target image.

[0081] Fourth step, decode the video bitstream according to the structured reconstructed frame image to obtain a target decoded output image sequence. The target decoded output image in the target decoded output image sequence may be an image obtained by decoding the target image with the image corresponding to the predicted macroblock as a reference image.

[0082] As an example, the above-mentioned execution entity may use the image corresponding to the predicted macroblock as a reference image to perform entropy decoding and inverse transform quantization on the above-mentioned image bitstream, so as to obtain the target decoded output image sequence.

[0083] The above embodiments of the present disclosure have the following beneficial effects: The video encoding method of some embodiments of the present disclosure can combine video intelligent analysis technology and macroblock-based video encoding technology, improving the video compression encoding efficiency in complex scenarios. Specifically, the reason for the low video encoding efficiency is that video encoding in units of macroblocks is difficult to accurately describe the macroscopic motion of the background and target objects, resulting in low video encoding efficiency and thus requiring more network resources. Based on this, the video encoding method of some embodiments of the present disclosure can, first, obtain a target image sequence for the target video. Here, the obtained target encoding image sequence is used for subsequent encoding of the target image. Secondly, for each target image in the above target image sequence, perform the following first encoding step: In response to determining that the above target image is not the first frame target image, determine the structured description information corresponding to the above target image according to the above target image and at least one historical target image, where the above structured description information includes: background change description information of the target image and foreground target object description information of the target image, and the historical target image in the above at least one historical target image is an image located before the above target image. Here, the obtained structured description information is used for subsequent determination of the structured reconstruction frame image corresponding to the target image. Then, generate a background frame image according to the above at least one historical target image. Generate a background reconstruction frame image according to the above background frame image. Generate the structured reconstruction frame image corresponding to the above target image according to the above background reconstruction frame image, at least one historical reconstruction frame image, and the above structured description information, where the historical reconstruction frame image in the above at least one historical reconstruction frame image is an image obtained by encoding and decoding reconstruction of the historical target image. Here, the obtained structured reconstruction frame image is used for subsequent determination of the prediction macroblock corresponding to the position of the macroblock in the above target image. Then, for each target macroblock in the above target image, perform the following second encoding step: Determine the macroblock in the above structured reconstruction frame image that has a position correspondence with the above target macroblock as the first macroblock. Generate a first macroblock bitstream according to the above first macroblock and the above target macroblock. Determine the macroblock in the pre-obtained basic prediction image set that has a position correspondence with the above target macroblock as the second macroblock. Generate a second macroblock bitstream according to the above second macroblock and the above target macroblock. In response to determining that the length of the above first macroblock bitstream is less than or equal to the length of the above second macroblock bitstream, determine the above first macroblock as the prediction macroblock. Here, macroblock-level pixel comparison is performed, and the macroblock with the smallest pixel value is selected as the prediction macroblock to achieve local optimality in video encoding. Finally, encode the above target image according to the obtained prediction macroblock set to obtain a first target image bitstream, where the image parameter set in the above first target image bitstream includes parameter information related to structured information. Generate a video bitstream according to the obtained first target image bitstream set.It can be seen that the video encoding method can combine video intelligent analysis technology and macroblock-based video encoding technology, improving the video compression and encoding efficiency in complex scenarios.

[0084] For further reference Figure 2 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a video encoding apparatus. These apparatus embodiments correspond to Figure 1 the method embodiments shown, and the apparatus can be specifically applied to various electronic devices.

[0085] As Figure 2 shown, a video encoding apparatus 200 includes: an acquisition unit 201, an execution unit 202, and a generation unit 203. Among them, the acquisition unit 201 is configured to: acquire a target image sequence for a target video. The execution unit 202 is configured to: for each target image in the above target image sequence, perform the following first encoding step: in response to determining that the above target image is not the first-frame target image, determine structured description information corresponding to the above target image according to the above target image and at least one historical target image, where the above structured description information includes: background change description information of the target image and foreground target object description information of the target image, and the historical target images in the above at least one historical target image are images located before the above target image; generate a background frame image according to the above at least one historical target image; generate a background reconstructed frame image according to the above background frame image; generate a structured reconstructed frame image corresponding to the above target image according to the above background reconstructed frame image, at least one historical reconstructed frame image, and the above structured description information, where the historical reconstructed frame images in the above at least one historical reconstructed frame image are images obtained by encoding and decoding and reconstructing historical target images; for each target macroblock in the above target image, perform the following second encoding step: determine the macroblock in the above structured reconstructed frame image that has a position correspondence with the above target macroblock as the first macroblock; generate a first macroblock bitstream according to the above first macroblock and the above target macroblock; determine the macroblock in the pre-obtained basic prediction image set that has a position correspondence with the above target macroblock as the second macroblock; generate a second macroblock bitstream according to the above second macroblock and the above target macroblock; in response to determining that the length of the above first macroblock bitstream is less than or equal to the length of the above second macroblock bitstream, determine the above first macroblock as the prediction macroblock; encode the above target image according to the obtained prediction macroblock set to obtain a first target image bitstream, where the image parameter set in the above first target image bitstream includes parameter information related to structured information. The generation unit 203 is configured to: generate a video bitstream according to the obtained first target image bitstream set.

[0086] It can be understood that the various units described in the apparatus 200 are related to the referenceFigure 1 corresponds to each step in the described method. Thus, the operations, features, and beneficial effects described above for the method also apply to the apparatus 200 and the units included therein, and will not be elaborated here.

[0087] Refer to the following Figure 3 , which shows a schematic structural diagram of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0088] As Figure 3 shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0089] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 shows an electronic device 300 having various devices, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be alternatively implemented or included. Figure 3 Each block shown in

[0090] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such some embodiments, the computer program can be downloaded and installed from the network via the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the methods of some embodiments of the present disclosure are performed.

[0091] It should be noted that the computer-readable medium in some embodiments of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in some embodiments of the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0092] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0093] The above computer-readable medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device is caused to: obtain a target image sequence for a target video; for each target image in the target image sequence, perform the following first encoding step: in response to determining that the target image is not the first-frame target image, determine structured description information corresponding to the target image according to the target image and at least one historical target image, where the structured description information includes: background change description information of the target image and foreground target object description information of the target image, and the historical target image in the at least one historical target image is an image located before the target image; generate a background frame image according to the at least one historical target image; generate a background reconstruction frame image according to the background frame image; generate a structured reconstruction frame image corresponding to the target image according to the background reconstruction frame image, at least one historical reconstruction frame image, and the structured description information, where the historical reconstruction frame image in the at least one historical reconstruction frame image is an image obtained by encoding and decoding and reconstructing a historical target image; for each target macroblock in the target image, perform the following second encoding step: determine the macroblock in the structured reconstruction frame image that has a position correspondence with the target macroblock as the first macroblock; generate a first macroblock bitstream according to the first macroblock and the target macroblock; determine the macroblock in the pre-obtained basic prediction image set that has a position correspondence with the target macroblock as the second macroblock; generate a second macroblock bitstream according to the second macroblock and the target macroblock; in response to determining that the length of the first macroblock bitstream is less than or equal to the length of the second macroblock bitstream, determine the first macroblock as the prediction macroblock; encode the target image according to the obtained set of prediction macroblocks to obtain a first target image bitstream, where the image parameter set in the first target image bitstream includes parameter information related to structured information; generate a video bitstream according to the obtained set of first target image bitstreams.

[0094] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0095] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0096] The units described in some embodiments of the present disclosure may be implemented in software or in hardware. The described units may also be provided in a processor. For example, it may be described as: a processor includes an acquisition unit, an execution unit, and a generation unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the acquisition unit may also be described as "the unit for acquiring the target image sequence for the target video".

[0097] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0098] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the embodiments of the present disclosure that have similar functions.

Claims

1. A video encoding method, comprising: Obtaining a target image sequence for a target video; For each target image in the target image sequence, performing the following first encoding step: In response to determining that the target image is not the first-frame target image, determining structured description information corresponding to the target image according to the target image and at least one historical target image, where the structured description information includes: background change description information of the target image and foreground target object description information of the target image, and the historical target images in the at least one historical target image are images located before the target image; Generating a background frame image according to the at least one historical target image; Generating a background reconstructed frame image according to the background frame image; Generating a structured reconstructed frame image corresponding to the target image according to the background reconstructed frame image, at least one historical reconstructed frame image, and the structured description information, where the historical reconstructed frame images in the at least one historical reconstructed frame image are images obtained by encoding and decoding and reconstructing historical target images; For each target macroblock in the target image, performing the following second encoding step: Determining a macroblock in the structured reconstructed frame image that has a position correspondence with the target macroblock as the first macroblock; Generating a first macroblock bitstream according to the first macroblock and the target macroblock; Determining a macroblock in a pre-obtained basic prediction image set that has a position correspondence with the target macroblock as the second macroblock; Generating a second macroblock bitstream according to the second macroblock and the target macroblock; In response to determining that the length corresponding to the first macroblock bitstream is less than or equal to the length corresponding to the second macroblock bitstream, determining the first macroblock as the prediction macroblock; Encoding the target image according to the obtained prediction macroblock set to obtain a first target image bitstream, where the image parameter set in the first target image bitstream includes parameter information related to structured information; Generating a video bitstream according to the obtained first target image bitstream set.

2. The method according to claim 1, wherein, After the step of, in response to determining that the length corresponding to the first macroblock bitstream is less than or equal to the length corresponding to the second macroblock bitstream, determining the first macroblock as the prediction macroblock, the method further includes: In response to determining that the length corresponding to the first macroblock bitstream is greater than the length corresponding to the second macroblock bitstream, determining the second macroblock as the prediction macroblock.

3. The method according to claim 1, wherein After the step of encoding the target image according to the obtained prediction macroblock set to obtain a first target image bitstream, it includes: Encoding the target image according to the obtained second macroblock set to obtain a second target image bitstream; In response to determining that the length corresponding to the first target image bitstream is less than or equal to the length corresponding to the second target image bitstream, determining the first target image bitstream as the bitstream corresponding to the target image; In response to determining that the length corresponding to the first target image bitstream is greater than the length corresponding to the second target image bitstream, determining the second target image bitstream as the bitstream corresponding to the target image.

4. The method according to claim 1, wherein Before determining, in response to determining that the target image is not the first-frame target image, the structured description information corresponding to the target image according to the target image and at least one historical target image, the method further includes: In response to determining that the target image is the first-frame target image, encoding the first-frame target image to obtain a first-frame target image bitstream.

5. The method according to claim 1, wherein, The generating the structured reconstruction frame image corresponding to the target image according to the background reconstruction frame image, at least one historical reconstruction frame image, and the structured description information includes: Determining, according to the structured description information, background change information in the target image and position change information of foreground target objects in the target image; Generating the structured reconstruction frame image corresponding to the target image according to the background reconstruction frame image, the at least one historical reconstruction frame image, the position change information, and the background change information.

6. The method according to claim 1, wherein The encoding the target image according to the obtained set of predicted macroblocks to obtain a first target image bitstream includes: Determining, according to the obtained set of predicted macroblocks, a plurality of predicted pixel values corresponding to the target image; Determining a plurality of differences between the plurality of predicted pixel values and a plurality of pixel values of the target image; Encoding the plurality of differences to obtain a first target image bitstream.

7. The method according to claim 1, wherein The method further includes: Obtaining a set of image parameters in the video bitstream; In response to determining that the set of image parameters includes parameter information related to the background frame image, performing syntax parsing on the background frame image to obtain a background reconstruction frame image; In response to determining that the set of image parameters includes parameter information related to the structured description information, generating the structured reconstruction frame image according to the structured description information, the background reconstruction frame image, and the at least one historical reconstruction frame image; Decoding the video bitstream according to the structured reconstruction frame image to obtain a target decoded output image sequence.

8. A video encoding apparatus, including: An obtaining unit configured to obtain a target image sequence for a target video; An execution unit is configured to, for each target image in the target image sequence, perform the following first encoding steps: in response to determining that the target image is not the first-frame target image, determine structured description information corresponding to the target image according to the target image and at least one historical target image, where the structured description information includes: background change description information of the target image and foreground target object description information of the target image, and the historical target images in the at least one historical target image are images located before the target image; generate a background frame image according to the at least one historical target image; generate a background reconstructed frame image according to the background frame image; generate a structured reconstructed frame image corresponding to the target image according to the background reconstructed frame image, at least one historical reconstructed frame image, and the structured description information, where the historical reconstructed frame images in the at least one historical reconstructed frame image are images obtained by encoding, decoding, and reconstructing historical target images; for each target macroblock in the target image, perform the following second encoding steps: determine the macroblock in the structured reconstructed frame image that has a position correspondence with the target macroblock as the first macroblock; generate a first macroblock bitstream according to the first macroblock and the target macroblock; determine the macroblock in the pre-obtained basic prediction image set that has a position correspondence with the target macroblock as the second macroblock; generate a second macroblock bitstream according to the second macroblock and the target macroblock; in response to determining that the length corresponding to the first macroblock bitstream is less than or equal to the length corresponding to the second macroblock bitstream, determine the first macroblock as the prediction macroblock; encode the target image according to the obtained prediction macroblock set to obtain a first target image bitstream, where the image parameter set in the first target image bitstream includes parameter information related to structured information; A generation unit is configured to generate a video bitstream according to the obtained first target image bitstream set.

9. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method according to any one of claims 1-7.

10. A computer-readable medium having a computer program stored thereon, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Video transcoding method and device, electronic equipment and storage medium

    CN113014926A

  • Video frame error hiding method and device, electronic equipment and medium

    CN114827632A