Video encoding method and video decoding method and device
By using texture edge segmentation and filtering processing based on the target image block in video encoding, the problem of predicting pixels of block boundaries is solved, and a smoother reconstruction of image boundaries and improved encoding performance is achieved.
Patent Information
- Application Number
- CN202311658808.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-06
AI Technical Summary
In video encoding, the pixels of the synthetic prediction block boundary are not smooth, resulting in discontinuous, non-smooth, or obvious boundaries in the boundary area of the reconstructed image.
By acquiring the target image block and dividing it into sub-image blocks based on its texture edges, the prediction blocks of the sub-image blocks are combined to obtain the initial prediction block, the minimum distance between the pixel points in the initial prediction block to calculate the filtering intensity, and the initial prediction block is filtered to obtain the fusion prediction block, and the residuals of the target image block and the fusion prediction block are finally calculated to generate encoded data.
This method can effectively filter unsmooth pixel points in the prediction block, avoid the problem of unsmooth pixels in the prediction block boundary, thereby improving video encoding performance and improving boundary continuity of the reconstructed image.
Smart Images

Figure CN120111232A_ABST
Abstract
Description
Technical Field
[0001] Some embodiments of the present application relate to the technical field of video encoding and decoding, and more specifically, to a video encoding method, a video decoding method and a device. Background Art
[0002] A complete image in a video is usually called a video frame. The content within a video frame and between video frames is often highly correlated, so there is a lot of redundant information in the video. The purpose of video encoding is to remove as much redundant information as possible in the video to reduce the amount of data representing the video content.
[0003] At present, when encoding video frames, mainstream video coding standards first divide video frames into coding tree units (CTUs) of different sizes, and use the resulting CTUs as the minimum unit for encoding and decoding. The encoding and decoding tasks for CTUs can be further refined to divide the CTUs into multiple coding units (CUs), a process called block partitioning. The coding performance of a video is closely related to the block partitioning. In order to improve the video coding performance, the related art proposes a block partitioning scheme that obtains the segmentation curve of an image block based on the internal texture edge of the image and uses the segmentation curve to partition the image block. The block partitioning scheme supports partitioning blocks into asymmetric, irregular, and irregular shaped coding blocks, which can make the division of coding blocks as close to the real image edge as possible, thereby improving the coding performance of the video. After performing motion estimation and motion compensation on the shaped coding blocks obtained by image block segmentation to obtain the prediction blocks of each shaped coding block, these prediction blocks need to be merged to perform residual operations on the image block. However, since the prediction blocks of different shaped coding blocks may be obtained based on reference blocks in different video frames, the boundary pixels of the synthesized prediction blocks will be uneven, which will cause irregular jumping of the residual during compensation. If not properly processed, the boundary area of the reconstructed image may appear discontinuous, non-smooth or obvious. Summary of the invention
[0004] The exemplary embodiments of the present application provide a video encoding method, a video decoding method and an apparatus for solving the problem of uneven boundary pixels of a synthesized prediction block.
[0005] Some embodiments of the present application provide the following technical solutions:
[0006] In a first aspect, some embodiments of the present application provide a video encoding method, including:
[0007] Obtaining a target image block, wherein the target image block is a rectangular image block obtained by dividing the video frame into blocks;
[0008] dividing the target image block into a first sub-image block and a second sub-image block based on a texture edge in the target image block;
[0009] Acquire a first prediction block of the first sub-image block and a second prediction block of the second sub-image block;
[0010] Combining the first prediction block and the second prediction block to obtain an initial prediction block, and obtaining a boundary line between the first prediction block and the second prediction block in the initial prediction block;
[0011] Determining the filtering strength of each pixel point in the initial prediction block according to the minimum distance from each pixel point in the initial prediction block to the boundary line;
[0012] Filtering each pixel in the initial prediction block based on the filtering strength of each pixel in the initial prediction block to obtain a fused prediction block;
[0013] Calculating the residual between the target image block and the fused prediction block to obtain residual data of the target image block;
[0014] Generate encoding data of the target image block according to the residual data of the target image block.
[0015] In a second aspect, some embodiments of the present application provide a video decoding method, including:
[0016] Obtaining encoded data of a target image block, wherein the target image block is a rectangular image block obtained by dividing a video frame into blocks;
[0017] Obtain residual data, a first prediction block, and a second prediction block of the target image block according to the encoded data of the target image block, wherein the first prediction block and the second prediction block are prediction blocks of a first sub-image block and a second sub-image block, respectively, and the first sub-image block and the second sub-image block are two sub-image blocks obtained by segmenting the target image block based on a texture edge in the target image block;
[0018] Combining the first prediction block and the second prediction block to obtain an initial prediction block, and obtaining a boundary line between the first prediction block and the second prediction block in the initial prediction block;
[0019] Determining the filtering strength of each pixel point in the initial prediction block according to the minimum distance from each pixel point in the initial prediction block to the boundary line;
[0020] Filtering each pixel in the initial prediction block based on the filtering strength of each pixel in the initial prediction block to obtain a fused prediction block;
[0021] A reconstructed image block of the target image block is obtained according to the residual data of the fused prediction block and the target image block.
[0022] In a third aspect, some embodiments of the present application provide a video encoding device, including:
[0023] a memory configured to store a computer program;
[0024] The processor is configured to, when calling a computer program, enable the video encoding device to implement the video encoding method described in the first aspect.
[0025] In a fourth aspect, some embodiments of the present application provide a video decoding device, including:
[0026] a memory configured to store a computer program;
[0027] The processor is configured to, when calling a computer program, enable the video decoding device to implement the video decoding method described in the second aspect.
[0028] In a fifth aspect, some embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computing device, the computing device implements the video encoding method described in the first aspect or the video decoding method described in the second aspect.
[0029] In a sixth aspect, some embodiments of the present application provide a computer program product, which, when executed on a computer, enables the computer to implement the video encoding method described in the first aspect or the video decoding method described in the second aspect.
[0030] It can be seen from the above technical scheme that when encoding the target image block obtained by block division of the video frame, the video encoding method provided by the embodiment of the present application first divides the target image block into a first sub-image block and a second sub-image block based on the texture edge in the target image block, and then obtains the first prediction block of the first sub-image block and the second prediction block of the second sub-image block, and combines the first prediction block and the second prediction block to obtain the initial prediction block, and obtains the boundary line between the first prediction block and the second prediction block in the initial prediction block, and determines the filtering strength of each pixel point in the initial prediction block according to the minimum distance from each pixel point in the initial prediction block to the boundary line, and then filters each pixel point in the initial prediction block based on the filtering strength of each pixel point in the initial prediction block to obtain a fused prediction block, and finally calculates the residual between the target image block and the fused prediction block to obtain the residual data of the target image block; and generates the encoding data of the target image block according to the residual data of the target image block. Since the video encoding method provided in the embodiment of the present application obtains the boundary line between the first prediction block and the second prediction block in the initial prediction block after combining the first prediction block and the second prediction block to obtain the initial prediction block, determines the filtering strength of each pixel in the initial prediction block according to the minimum distance from each pixel in the initial prediction block to the boundary line, and filters each pixel in the initial prediction block based on the filtering strength of each pixel in the initial prediction block, the video encoding method provided in the embodiment of the present application can filter the uneven pixels in the prediction block, thereby avoiding uneven pixels at the boundary of the prediction block. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the implementation methods of some embodiments of the present application or related technologies, the following is a brief introduction to the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings described below are some embodiments of the present application, and a person skilled in the art can also obtain other drawings based on these drawings.
[0032] Figure 1 A block diagram of a video decoding system in some embodiments of the present application is shown;
[0033] Figure 2 A schematic diagram showing the structure of a video encoder in some embodiments of the present application is shown;
[0034] Figure 3 A schematic diagram showing the structure of a video decoder in some embodiments of the present application is shown;
[0035] Figure 4 One of the flow charts of the steps of the video encoding method in some embodiments of the present application is shown;
[0036] Figure 5 A schematic diagram showing an initial prediction block and a boundary line in some embodiments of the present application is shown;
[0037] Figure 6 A schematic diagram showing the distance from a pixel point to a boundary line in some embodiments of the present application;
[0038] Figure 7 The second flowchart of the video encoding method in some embodiments of the present application is shown;
[0039] Figure 8 A schematic diagram showing a first area of an initial prediction block in some embodiments of the present application;
[0040] Fig. 9 A flowchart showing the steps of a video decoding method in some embodiments of the present application is shown. DETAILED DESCRIPTION
[0041] In order to make the purpose and implementation method of the present application clearer, the exemplary implementation method of the present application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only part of the embodiments of the present application, rather than all the embodiments.
[0042] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and common meanings.
[0043] The terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.
[0044] References to "some implementations", "some embodiments", etc. in the specification indicate that the described implementations or embodiments may include specific features, structures, or characteristics, but not every embodiment may include the specific features, structures, or characteristics. In addition, such phrases do not necessarily refer to the same implementation. In addition, when a specific feature, structure, or characteristic is described in connection with an embodiment, it is considered to be within the knowledge of those skilled in the art to implement such feature, structure, or characteristic in connection with other implementations (whether or not explicitly described herein).
[0045] The embodiments of the present application relate to the field of video coding and decoding technology. The following first describes a video coding and decoding framework for executing the video encoding method and video decoding method provided by the embodiments of the present application.
[0046] Video can be regarded as a sequence of multiple video frames (images). Video playback can be regarded as the display of video frames at a preset speed (for example, 24 frames per second, 30 frames per second, and 60 frames per second) in the order of the sequence. In theory, the amount of video data is positively correlated with the resolution of the video frame. The higher the resolution of the video frame, the larger the amount of video data. If the pixel data of each pixel of all video frames is directly saved in the video file, the amount of video data will be very large, which will make the video difficult to store and transmit. Video encoding and decoding is proposed to solve this problem to a certain extent. Video decoding mainly includes: video encoding and video decoding. Among them, video encoding can be understood as the process of compressing the original video frame, and video decoding can be understood as the process of reconstructing the video frame based on the compressed video data.
[0047] Reference Figure 1 The block diagram of the video decoding system in some embodiments of the present application is shown in FIG. Figure 1 As shown, the video decoding system 100 includes: a source device 10 and a destination device 20. The source device 10 can obtain original video data through a video source 101, and encode the original video frame through a video encoder 102 to obtain video encoding data, and provide the video encoder 102 output video encoding data to the destination device 20 through an output interface 103. The destination device 20 can obtain the video encoding data provided by the source device 10 through an input interface 201, and decode the video encoding data through a video decoder 202 to obtain video decoding data, and input the video decoding data into a player 203 to play the video. The source device 10 and the destination device 20 may include any of a wide range of devices, such as: a personal computer (Program Counter), a notebook computer, a tablet computer, a set-top box, a mobile phone, a television, a camera, a display, a digital media player, a video game console, a video streaming device, etc.
[0048] In some embodiments, the video source 101 of the source device 10 may be a video shooting device, such as a camera. In other embodiments, the video source 101 may be a component capable of generating video based on computer graphics, such as a screen recording component, an animation generation component, etc.
[0049] In some embodiments, the destination device 20 may receive the video encoding data provided by the source device 10 via a computer-readable medium. The computer-readable medium may include any type of medium or device capable of moving the video encoding data from the source device 10 to the destination device 20. In one example, the computer-readable medium may include a communication medium. The communication medium may modulate the video encoding data according to a communication standard (e.g., a wireless communication protocol) and transmit it to the destination device 20. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet network (e.g., a local area network, a wide area network, or a global network, such as the Internet). The communication medium may include a router, a switch, a base station, or any other device that can be used to facilitate communication from the source device 10 to the destination device 20.
[0050] In some examples, the video encoding data may be output from the output interface 103 of the source device 10 to a storage device. Accordingly, the video encoding data may be accessed from the storage device by the input interface 203 of the destination device 20. The storage device may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing video encoding data. In another example, the storage device may be a server or an intermediate storage device for storing the video encoding data generated by the source device 10. The destination device 20 may obtain the stored video encoding data from the storage device via streaming or downloading. The file server may be any type of server capable of storing encoded video data and transmitting the encoded video data to the destination device 20. In some embodiments. The file server includes a network server (e.g., for a website), an FTP server, a network attached storage device, or a local disk drive. The destination device 20 may access the encoded video data through any standard data connection (including an Internet connection). This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both suitable for accessing encoded video data stored on a file server. The transmission of the encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.
[0051] At present, the video coding standard has gradually evolved from ISO / IECMPEG-1 (International StandardizationOrganization / International Electrotechnical CommissionMoving Picture ExpertsGroup-1), through ISO / IECMPEG-2, ISO / IECMPEG-4, Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), etc. to Versatile Video Coding (VVC). The video encoding method and video decoding method provided in the embodiments of the present application can be applied to any suitable video decoding standard. For example: HEVC standard, VVC standard, etc.
[0052] Reference Figure 2 As shown, Figure 2The schematic diagram of the structure of the video encoder in some embodiments of the present application. The video encoder 200 provided in the embodiment of the present application includes: a segmentation module 21, a motion estimation module 22, a motion compensation module 23, a fusion filter module 24, a residual calculation module 25, a quantization transformation module 26 and an entropy coding module 27. The process of video encoding by the video encoder 200 includes: 1. The segmentation module 21 divides the video frame to be encoded into blocks. For the rectangular image blocks obtained by block division, the segmentation module 21 can decide whether to further segment according to the actual encoding situation. Specifically: Since selecting a larger coding block is more conducive to improving the encoding efficiency, for rectangular image blocks with relatively flat pixels, fewer segmentation times can be selected, and the segmentation depth is shallow. On the contrary, for rectangular image blocks with more complex textures or located at the edge of the image, more segmentation times can be selected, and the segmentation depth is deeper. The segmentation module 21 will also divide at least one rectangular image block into two heterogeneous sub-image blocks based on a segmentation curve that fits the real texture edge. 2. The motion estimation module 22 performs motion estimation (ME) on each coding block (including: rectangular image blocks and irregular sub-image blocks that are not further divided) to obtain the reference block (RB) corresponding to each coding block. 3. The motion compensation module 23 performs motion compensation (MC) on the reference block to obtain the prediction block and motion vector (MV) corresponding to each coding block. 4. For the rectangular image block divided into two irregular sub-image blocks, the fusion filter module 24 combines the prediction blocks of the two irregular sub-image blocks into the initial prediction block of the rectangular image block, and filters the initial prediction block to obtain the fused prediction block. 5. The residual calculation module 25 calculates the residual between the prediction block and the corresponding coding block to obtain the prediction residual. 6. The transform quantization module 26 transforms and quantizes the prediction residual to obtain the transform coefficient. 7. The entropy coding module 27 performs entropy coding on the segmentation curve, motion vector, transform coefficient and other information to obtain the encoded data of the video frame to be encoded.
[0053] It should be noted that the above encoding process only describes some steps in the video encoding process. In addition to the above steps, the video encoding process also includes other steps. For example, inputting information such as the prediction mode and the transformation mode into the entropy encoding module 27, and adding the entropy encoding result of the prediction mode and the transformation mode to the encoded data of the video frame to be encoded.
[0054] Reference Figure 3 As shown, Figure 3The schematic diagram of the structure of the video decoder in some embodiments of the present application. The video encoder 300 provided in the embodiment of the present application includes: an entropy decoding module 31, a motion estimation module 32, a motion compensation module 33, a fusion filter module 34, an inverse quantization transformation module 35, a fusion module 36 and a reconstruction module 37. The process of video decoding by the video encoder 300 includes: 1. The entropy decoding module 31 performs entropy decoding on the coded data of the video frame to be coded, and obtains at least one of the information such as a segmentation curve for segmenting the rectangular image block of the video frame to be reconstructed into two heterogeneous sub-image blocks, motion vectors of the two heterogeneous sub-image blocks, transformation coefficients of the coding block, and motion vectors according to the entropy decoding result. 2. The motion estimation module 32 determines the reference block corresponding to each image block to be reconstructed and the two heterogeneous sub-image blocks of the image block to be reconstructed. 3. The motion compensation module 33 performs motion compensation on the reference block according to the motion vector to obtain the prediction block corresponding to the image block to be reconstructed or the prediction block of the two heterogeneous sub-image blocks of the image block to be reconstructed. 4. The fusion filter module 34 combines the prediction blocks of the two heterogeneous sub-image blocks into the initial prediction block of the image block to be reconstructed, and filters the initial prediction block to obtain the fused prediction block. 5. The inverse quantization transformation module 35 performs inverse quantization and inverse transformation operations on the transformation coefficients to obtain the prediction residual, and the fusion module 36 adds and fuses the prediction residual with the prediction block or the fused prediction block to obtain the data to be reconstructed. 7. The reconstruction module 37 reconstructs each block to be reconstructed according to the reconstructed data to obtain a reconstructed video frame.
[0055] The present application embodiment provides a video encoding method, referring to Figure 4 As shown, the video encoding method includes the following steps S41 to S48:
[0056] S41, obtaining a target image block.
[0057] The target image block is a rectangular image block obtained by dividing the video frame into blocks.
[0058] Under video coding standards such as HEVC and VVC, after receiving the video frame to be encoded, the encoder will first divide the video frame to be encoded into multiple image blocks, and encode the image blocks obtained by dividing the video frame to be encoded as the minimum coding unit. In addition, the encoding task for the image block can be further refined. For example: In the HEVC standard, the image block obtained by dividing the video frame to be encoded is called a coding tree unit (Coding Tree Unit, CTU), and the coding tree unit can be further divided into smaller square coding units (Coding unit, CU) by quadtree division. The maximum size of the coding tree unit can be supported to 64×64 and the minimum size can be supported to 16×16. In order to meet the needs of ultra-high-definition video encoding such as 4K and 8K, the VVC standard expands the maximum size of CTU to 128×128, and further division methods are no longer limited to quadtree division, but also support binary tree and ternary tree division. The target image block in the embodiment of the present application can be a coding tree unit, or it can be a coding unit obtained by further dividing the coding tree unit, and the embodiment of the present application does not limit this.
[0059] S42: Divide the target image block into a first sub-image block and a second sub-image block based on a texture edge in the target image block.
[0060] Exemplarily, texture edge detection may be performed on the target image block using edge detection algorithms such as Sobel, Prewitt, Roberts, Canny, and Marr-Hildreth to obtain texture edges in the target image block.
[0061] Since the texture edge in the target image block is generally not a horizontal or vertical straight line, the first sub-image block and the second sub-image block are probably not rectangular image blocks, but irregularly shaped image blocks.
[0062] In some embodiments, the above step S42 (segmenting the target image block into a first sub-image block and a second sub-image block based on a texture edge in the target image block) includes the following steps 1 to 4:
[0063] Step 1: Segment the target image block into two image areas based on a texture edge in the target image block, and obtain a boundary line between the two image areas.
[0064] It should be noted that in some cases, the brightness of the pixels in the image block is relatively flat, and texture edge detection on the image block cannot obtain a texture edge that can divide the image block into two image regions. For such image blocks, the embodiment of the present application may not further divide the sub-image blocks, but directly encode the image block as a coding unit. In other cases, the texture edge obtained by texture edge detection on the image block will divide the image block into image regions greater than 2. For such image blocks, multiple image regions can be first merged into two image regions, and then the subsequent video encoding steps can be performed.
[0065] In some embodiments, dividing the target image block into two image areas based on the at least one texture edge includes: performing image area segmentation on the target image block based on the at least one texture edge to obtain an image area set corresponding to the target image block; obtaining the number of image areas in the image area set; if the number of image areas in the image area set is greater than 2, merging the image areas in the image area set into two image areas to divide the target image block into two image areas.
[0066] In some embodiments, merging the image areas in the image area set into two image areas includes: obtaining at least two merging schemes for the image area set, the at least two merging schemes including various schemes for merging the image areas in the image area set into two image areas; respectively obtaining the difference between the at least two merging schemes, the difference between any merging scheme being the absolute difference between a first similarity and a second similarity of the merging scheme; the first similarity and the second similarity of any merging scheme being the sum of the absolute differences between the grayscale values of each pixel point of the two image areas and the corresponding co-located blocks under the merging scheme; and merging the image areas in the image area set into two image areas by using the merging scheme with the largest difference among the at least two merging schemes.
[0067] Step 2: Perform curve fitting on the boundary line to obtain a segmentation curve of the target image block.
[0068] Generally, the boundary line between two image areas has no function expression or the function expression is very complex. Directly representing the boundary line between the two image areas in the bitstream data will greatly increase the overhead of the bitstream data. Therefore, the embodiment of the present application performs curve fitting on the boundary line between the two image areas to simplify the function expression of the boundary line, thereby reducing the overhead caused by the boundary line.
[0069] In some embodiments, performing curve fitting on the dividing line to obtain the segmentation curve of the target image block includes: performing second-order or third-order Bezier curve fitting on the dividing line, and determining the Bezier curve obtained by performing second-order or third-order Bezier curve fitting on the dividing line as the segmentation curve of the target image block.
[0070] Step 3: Sampling the segmentation curve with integer pixel accuracy to obtain a sampling result of the segmentation curve.
[0071] Since the segmentation curve is a continuous curve, the continuous segmentation curve may pass through the smallest unit (pixel point) of the digital image, and the smallest unit of the digital image cannot be segmented and encoded, so the segmentation curve cannot be directly applied to the segmentation of the digital image. Based on this, when the target image block is segmented into two sub-image blocks based on the segmentation curve of the target image block, the embodiment of the present application first samples the segmentation curve with an integer pixel accuracy to sample the segmentation curve as a discrete curve, and then segments the target image block into two sub-image blocks through the sampled curve (the sampling result of the segmentation curve), thereby avoiding segmenting a pixel point into different sub-image blocks.
[0072] Step 4: Segment the target image block into the first sub-image block and the second sub-image block based on the sampling result of the segmentation curve.
[0073] In some embodiments, the above step S42 (segmenting the target image block into a first sub-image block and a second sub-image block based on a texture edge in the target image block) includes the following steps a to e:
[0074] Step a: dividing the target image block into two image areas based on a texture edge in the target image block, and obtaining a boundary line between the two image areas.
[0075] The implementation of step a may refer to the above step 1, and will not be described in detail here to avoid redundancy.
[0076] Step b: obtaining two endpoints of the segmentation curve according to the encoded video content.
[0077] In some embodiments, obtaining two endpoints of a segmentation curve according to video content that has been encoded includes: determining whether spatially adjacent pixel points of four edges of the target image block can be obtained in the video content that has been encoded; if so, obtaining the segmentation start point and the segmentation end point according to the spatially adjacent pixel points of the four edges of the target image block; if not, obtaining the best matching block of the target image block in the video content that has been encoded, and obtaining the segmentation start point and the segmentation end point according to the edge pixel points of the best matching block of the target image block.
[0078] In some embodiments, obtaining the segmentation starting point and the segmentation end point based on the spatially adjacent pixel points at the four edges of the target image block includes: obtaining the gradient amplitude of the spatially adjacent pixel points at the four edges of the target image block in the direction of the corresponding edge; obtaining at least two pixel value step points based on the gradient amplitude of the spatially adjacent pixel points at the four edges of the target image block in the direction of the corresponding edge, the pixel value step point being a pixel point with the largest gradient amplitude and a gradient amplitude greater than a threshold gradient amplitude among the spatially adjacent pixel points at any edge of the target image block; and respectively determining the two pixel value step points with the largest gradient amplitude among the at least two pixel value step points as the segmentation starting point and the segmentation end point.
[0079] In some embodiments, the acquiring the segmentation start point and the segmentation end point according to the edge pixel points of the best matching block of the target image block comprises: acquiring the gradient amplitudes of the pixel points on four edges of the best matching block of the target image block in the direction of the corresponding edges; acquiring at least two pixel value step points according to the gradient amplitudes of the pixel points on four edges of the best matching block of the target image block in the direction of the corresponding edges, wherein the pixel value step points are pixel points with the largest gradient amplitude and a gradient amplitude greater than a threshold gradient amplitude among the pixel points on any edge of the best matching block of the target image block; acquiring two pixel value step points with the largest gradient amplitude among the at least two pixel value step points; and determining the segmentation start point and the segmentation end point according to the two pixel value step points with the largest gradient amplitude and the motion vector of the best matching block of the target image block, respectively.
[0080] Step c: determining the segmentation curve according to the two endpoints of the segmentation curve and the dividing line.
[0081] In some embodiments, the segmentation curve is determined according to the two endpoints of the segmentation curve and the dividing line, including: fitting the dividing line with a third-order Bezier curve with the segmentation starting point and the segmentation end point as the two endpoints of the third-order Bezier curve to obtain a first control point and a second control point; determining whether the first control point and the second control point are located on the same side of a straight line connecting the segmentation starting point and the segmentation end point; if not, obtaining the segmentation curve according to the segmentation starting point, the segmentation end point, the first control point and the second control point; if yes, fitting the dividing line with a second-order Bezier curve with the segmentation starting point and the segmentation end point as the two endpoints of the second-order Bezier curve to obtain control points, and obtaining the segmentation curve according to the segmentation starting point, the segmentation end point and the control points.
[0082] Step d: sampling the segmentation curve with integer pixel accuracy to obtain a sampling result of the segmentation curve.
[0083] Step e: dividing the target image block into the first sub-image block and the second sub-image block based on the sampling result of the segmentation curve.
[0084] Since the above-mentioned embodiment obtains the segmentation starting point and the segmentation end point of the segmentation curve for segmenting the target image block into two sub-image blocks according to the video content that has been encoded, the decoding end can also obtain the segmentation starting point and the segmentation end point according to the video content that has been encoded. Therefore, the encoding end only needs to add the control information of the segmentation curve to the bitstream data corresponding to the target image block, and the decoding end can obtain the complete segmentation curve information. Therefore, the embodiment of the present application can avoid adding the segmentation starting point and the segmentation end point of the segmentation curve to the bitstream data corresponding to the target image block, thereby reducing the bit rate overhead brought by representing the segmentation curve.
[0085] S43: Obtain a first prediction block of the first sub-image block and a second prediction block of the second sub-image block.
[0086] In some embodiments, obtaining a first prediction block of the first sub-image block and a second prediction block of the second sub-image block includes: performing motion estimation on the first sub-image block to obtain a first reference block and a first motion vector; performing motion compensation on the first reference block based on the first motion vector to obtain the first prediction block; performing motion estimation on the second sub-image block to obtain a second reference block and a second motion vector; and performing motion compensation on the second reference block based on the second motion vector to obtain the second prediction block.
[0087] S44. Combine the first prediction block and the second prediction block to obtain an initial prediction block, and obtain a boundary line between the first prediction block and the second prediction block in the initial prediction block.
[0088] For example, refer to Figure 5 As shown, the first prediction block 51 and the second prediction block 52 are combined to obtain an initial prediction block 53 , and a boundary line 500 between the first prediction block 51 and the second prediction block 52 in the initial prediction block 53 is obtained.
[0089] S45 . Determine the filtering strength of each pixel in the initial prediction block according to the minimum distance from each pixel in the initial prediction block to the boundary line.
[0090] In the embodiment of the present application, the minimum distance from the pixel point to the boundary line refers to the minimum value of the distances from the pixel point to each point on the boundary line. For example, as shown in reference 6, the distances from the pixel point 600 to each point on the boundary line 500 include: d1, d2, d3, ..., dn, and the minimum value of d1, d2, d3, ..., dn is dx, then the minimum distance dx from the pixel point 600 to the boundary line 500 is determined.
[0091] S46. Filter each pixel in the initial prediction block based on the filtering strength of each pixel in the initial prediction block to obtain a fused prediction block.
[0092] In the embodiment of the present application, filtering the pixel points means that the convolution kernel slides within the range of the pixel point set to be filtered, and performs a convolution operation with the original pixel value of each pixel point to obtain the filtered pixel value. The pixel value in the embodiment of the present application can be the grayscale value of the pixel point.
[0093] In some embodiments, the filtering strength of each pixel in the initial prediction block is negatively correlated with the minimum distance from each pixel in the initial prediction block to the boundary line. That is, for any pixel in the initial prediction block, if the minimum distance from the pixel to the boundary line is smaller, the filtering strength of the pixel is greater; conversely, if the minimum distance from the pixel to the boundary line is larger, the filtering strength of the pixel is smaller.
[0094] In the embodiment of the present application, the filtering strength of each pixel point in the initial prediction block is negatively correlated with the minimum distance from each pixel point in the initial prediction block to the boundary line. Therefore, the embodiment of the present application can adaptively filter the image according to the content and needs of the image, so that for pixels with stronger noise, stronger filtering is applied to remove noise, and for cleaner pixels, weaker filtering is applied to retain authenticity.
[0095] S47. Calculate the residual between the target image block and the fused prediction block to obtain residual data of the target image block.
[0096] In some embodiments, calculating the residual between the target image block and the fused prediction block includes: calculating the difference between the pixel values of each pixel point of the target image block and the co-located pixel point in the fused prediction block to obtain the residual data of the target image block.
[0097] S48. Generate encoding data of the target image block according to the residual data of the target image block.
[0098] In some embodiments, the encoding data of the target image block is generated according to the residual data of the target image block, including performing operations such as transformation, quantization, and entropy encoding on the residual data of the target image block to obtain the encoding data of the target image block.
[0099] The video encoding method provided in the embodiment of the present application encodes the target image block obtained by block division of the video frame. First, based on the texture edge in the target image block, the target image block is divided into a first sub-image block and a second sub-image block. Then, the first prediction block of the first sub-image block and the second prediction block of the second sub-image block are obtained, and the first prediction block and the second prediction block are combined to obtain an initial prediction block. The boundary line between the first prediction block and the second prediction block in the initial prediction block is obtained, and the filtering strength of each pixel point in the initial prediction block is determined according to the minimum distance from each pixel point in the initial prediction block to the boundary line. Then, based on the filtering strength of each pixel point in the initial prediction block, each pixel point in the initial prediction block is filtered to obtain a fused prediction block. Finally, the residual between the target image block and the fused prediction block is calculated to obtain residual data of the target image block. The encoding data of the target image block is generated according to the residual data of the target image block. Since the video encoding method provided in the embodiment of the present application obtains the boundary line between the first prediction block and the second prediction block in the initial prediction block after combining the first prediction block and the second prediction block to obtain the initial prediction block, determines the filtering strength of each pixel in the initial prediction block according to the minimum distance from each pixel in the initial prediction block to the boundary line, and filters each pixel in the initial prediction block based on the filtering strength of each pixel in the initial prediction block, the video encoding method provided in the embodiment of the present application can filter the uneven pixels in the prediction block, thereby avoiding uneven pixels at the boundary of the prediction block.
[0100] As an extension and refinement of the above embodiment, the present application embodiment also provides another video encoding method, referring to Figure 7 As shown, the video encoding method includes the following steps:
[0101] S701, obtaining a target image block.
[0102] The target image block is a rectangular image block obtained by dividing the video frame into blocks.
[0103] S702: Divide the target image block into a first sub-image block and a second sub-image block based on a texture edge in the target image block.
[0104] S703: Perform motion estimation on the first sub-image block to obtain a first reference block and a first motion vector.
[0105] S704: Perform motion compensation on the first reference block based on the first motion vector to obtain the first prediction block.
[0106] S705: Perform motion estimation on the second sub-image block to obtain a second reference block and a second motion vector.
[0107] S706: Perform motion compensation on the first reference block based on the first motion vector to obtain the first prediction block.
[0108] S707 . Combine the first prediction block and the second prediction block to obtain an initial prediction block, and obtain a boundary line between the first prediction block and the second prediction block in the initial prediction block.
[0109] S708: Set the filtering strength of each pixel point in the first area of the initial prediction block to a preset filtering strength.
[0110] The first area is an area composed of pixel points adjacent to the boundary line.
[0111] For example, refer to Figure 8 As shown, the pixels of the initial prediction block 53 that are adjacent to the boundary line 500 include pixels with pixel coordinates: (4,13), (4,14), (4,15), (4,16), (5,13), (5,14), (5,15), (5,16), (6,13), (6,14), (7,7), (7,8), (7,9)..., so the area composed of these pixels is determined as the first area 531 of the initial prediction block 53.
[0112] In some embodiments, the preset filtering strength is 1.
[0113] S709. For each pixel point in the second area according to the initial prediction block, calculate the pixel distance between the pixel point and each pixel point in the first area, and determine the minimum value of the pixel distance between the pixel point and each pixel point in the first area as the minimum distance from the pixel point to the boundary line.
[0114] The second area is an area composed of pixel points that are not adjacent to the boundary line.
[0115] Since the first region is a region composed of pixels adjacent to the boundary line, and the second region is a region composed of pixels not adjacent to the boundary line, the set of pixels in the first region and the set of pixels in the second region are complementary to each other. Figure 8As shown, the second region includes other regions in the initial prediction block 53 except the first region 531 .
[0116] In some embodiments, the pixel distance between the pixel points in the second area and the pixel points in the first area can be calculated by the following formula (1):
[0117]
[0118] Among them, i and j are the pixel coordinates of the pixel points in the second area respectively, i_target and j_target are the pixel coordinates of the pixel points in the first area respectively, and d is the pixel distance between the pixel point (i, j) in the second area and the pixel point (i_target, j_target) in the first area.
[0119] S710: Calculate the filtering strength of each pixel in the second area of the initial prediction block according to the minimum distance from each pixel in the second area to the boundary line.
[0120] In some embodiments, the above step S709 (calculating the filtering strength of each pixel in the second area of the initial prediction block according to the minimum distance from each pixel in the second area to the boundary line) includes:
[0121] Calculate the inverse of the square value of the minimum distance from each pixel point in the second area to the boundary line to obtain a first calculated value for each pixel point in the second area; calculate the inverse of the first calculated value for each pixel point in the second area to obtain a second calculated value for each pixel point in the second area; determine the second calculated value for each pixel point in the second area as the filtering strength for each pixel point in the second area.
[0122] The minimum distance from the pixel point (i, j) in the second area to the boundary line is expressed as d i,j , the filtering intensity of the pixel point (i, j) in the second area is expressed as w i,j , then the filter strength w i,j The calculation formula can be shown as formula (2):
[0123] w i,j =1 / (d i,j ) 2 (2)
[0124] At this point, the filtering strength of each pixel point in the first area and the filtering strength of each pixel point in the second area are determined through the above steps, and the first area and the second area can be combined into an initial prediction block, so the filtering strength of each pixel point in the initial prediction block is determined.
[0125] That is, the filtering strength of each pixel in the initial prediction block can be expressed as the following formula (3):
[0126]
[0127] Among them, d i,j is the minimum distance from the pixel point (i, j) in the initial prediction block to the boundary line, w i,j is the filtering strength of pixel (i, j).
[0128] S711. Obtain a set of pixels to be filtered according to the filtering strength of each pixel in the initial prediction block.
[0129] The to-be-filtered pixel point set is a set of pixel points in the initial prediction block whose filtering strength is greater than a strength threshold.
[0130] In the embodiment of the present application, the closer the intensity threshold is to 0, the larger the pixel area to be filtered, the more pixels are included in the set of pixel points to be filtered, the greater the computing power cost, but the better the fusion effect; the closer the intensity threshold is to 1, the smaller the pixel area to be filtered, the fewer pixels are included in the set of pixel points to be filtered, the smaller the computing power cost, but the worse the fusion effect. Therefore, the intensity threshold value needs to be adjusted according to the specific application scenario, video content and encoder characteristics. In some embodiments, the value range of the intensity threshold is [0.16, 0.8].
[0131] In the embodiment of the present application, the filter strength w of each pixel in the initial prediction block is i,j Compare with the intensity threshold. i,j When the intensity is greater than or equal to the intensity threshold, the pixel is classified into the set of pixel points to be filtered, and the pixels not classified into the set of pixel points to be filtered are not processed. Therefore, the above embodiment can reduce the number of pixels that need to be processed, thereby reducing the calculation complexity and saving computing power consumption.
[0132] S712: Convert the filtering intensity of each pixel point in the set of pixel points to be filtered into a filtering parameter of a target filter to obtain the filtering parameter of each pixel point in the set of pixel points to be filtered.
[0133] The target filter in the embodiment of the present application can be a weighted mean filter, a Gaussian two-dimensional filter, an adaptive median filter, a bilateral filter, a Laplace filter, or other filters.
[0134] Different filters have different applications and characteristics: weighted mean filter uses different weights to calculate the average value, while smoothing the image, retaining some detailed features in the image; adaptive filter uses dynamic weights, adjusted according to local pixel values, used to process uneven noise, retain image details; edge enhancement filter weights depend on pixel position, focusing on processing edge features; Laplacian filter center pixel weight is positive, surrounding pixel weight is negative, used to enhance the high-frequency components in the image, enhance the details and edges of the image; frequency domain filter filters the signal according to the frequency, commonly used in image restoration and frequency domain analysis, remove periodic noise, frequency domain specific frequency analysis. These filters have different effects and uses in different applications. In actual application, the appropriate filter type is selected according to specific needs. The following is an example of the target filter being a Gaussian two-dimensional filter. Gaussian two-dimensional filter is a linear filter commonly used in image processing, used to smooth images to reduce noise and details, and it calculates the weight of the filter based on the Gaussian function to control the strength and smoothness of the filter. The kernel of the Gaussian two-dimensional filter (also known as the Gaussian kernel) is usually an odd-sized two-dimensional square matrix. The variance (also called standard deviation) of a Gaussian 2D filter determines the smoothness and response characteristics of the filter. The range of standard deviation values affects the performance and effect of the filter. When the standard deviation is small (usually less than 1.0), the Gaussian 2D filter has a small smoothing effect and is suitable for tasks that need to retain image details and are sensitive to noise, such as noise removal before edge detection. When the standard is medium (usually between 1.0 and 4.0), it is usually used for general smoothing operations, which can remove noise to a certain extent and retain the main features of the image. When the standard deviation is large (usually greater than 4.0), the Gaussian 2D filter has a stronger smoothing effect, removes more details, and is suitable for tasks that require strong smoothing of images to remove a lot of noise.
[0135] In some embodiments, when the target filter is a Gaussian two-dimensional filter, the filter strength of each pixel in the set of pixels to be filtered is converted into the filter parameters of the target filter to obtain the filter parameters of each pixel in the set of pixels to be filtered, including: determining the tangent value of the filter strength of each pixel in the set of pixels to be filtered as the standard deviation of each pixel in the set of pixels to be filtered. That is, when the target filter is a Gaussian two-dimensional filter, the standard deviation of the pixel point (i, j) is expressed as σ i,j , the filtering intensity of pixel (i, j) is expressed as w i,j , then the standard deviation σ i,j The calculation formula can be shown as formula (4):
[0136] σ i,j =tan(w i,j ) (4)
[0137] S713. Obtain a filter matrix for each pixel in the set of pixel points to be filtered according to the filter parameters of each pixel in the set of pixel points to be filtered and the response characteristics of the target filter.
[0138] Two-dimensional Gaussian kernel GK i,j The calculation formula is shown in formula (5):
[0139]
[0140] Substitute the standard deviation of each pixel in the set of pixels to be filtered into the two-dimensional Gaussian kernel GK i,j The calculation formula can be used to calculate the two-dimensional Gaussian kernel corresponding to each pixel point in the set of pixel points to be filtered.
[0141] In some embodiments, before obtaining the filter matrix of each pixel point in the initial prediction block based on the filter parameters of each pixel point in the set of pixel points to be filtered and the response characteristics of the target filter, the method also includes: determining the size of the filter matrix of each pixel point in the set of pixel points to be filtered based on the filter parameters of each pixel point in the set of pixel points to be filtered.
[0142] In some embodiments, the size of the filter matrix of each pixel in the set of pixels to be filtered is positively correlated with the filter parameters of each pixel in the set of pixels to be filtered. That is, for any pixel, if the filter parameter of the pixel is larger, the size of the filter matrix of the pixel is larger; conversely, if the filter parameter of the pixel is smaller, the size of the filter matrix of the pixel is smaller.
[0143] For example: when the target filter is a Gaussian two-dimensional filter, if the standard deviation of the pixel points (filter parameters) is less than 1.0, a Gaussian kernel (filter matrix) of size 3x3 will be selected; if the standard deviation of the pixel points is between 1.0 and 4.0, a Gaussian kernel of size 5x5 will be selected; if the standard deviation of the pixel points is greater than 4.0, a Gaussian kernel of size 7x7 or larger will be selected.
[0144] S714: Filter each pixel point in the set of pixel points to be filtered according to the filter matrix of each pixel point in the set of pixel points to be filtered, so as to obtain a fused prediction block.
[0145] S715. Calculate the residual between the target image block and the fused prediction block to obtain residual data of the target image block.
[0146] S716: Generate encoding data of the target image block according to the residual data of the target image block.
[0147] In some embodiments, encoding data of the target image block is generated based on the residual data of the target image block, including: transforming, entropy encoding, and other operations on the residual data of the target image block, information of a segmentation curve for segmenting the target image block into the first sub-image block and the second sub-image block, the motion vector of the first sub-image block, and the motion vector of the second sub-image block, so as to obtain the encoding data of the target image block.
[0148] In some embodiments, based on any of the above embodiments, the video encoding method provided in the embodiments of the present application further includes performing the following steps ① to ③ before calculating the residual between the target image block and the fused prediction block to obtain the residual data of the target image block:
[0149] Step ①: perform noise point detection on the fused prediction block based on a preset noise detection algorithm.
[0150] Exemplarily, the preset noise detection algorithm may be a square root algorithm, an inter-frame difference detection algorithm, an autocorrelation algorithm, a spectrum analysis algorithm, a wavelet transform algorithm, a threshold processing algorithm, or other noise detection algorithms.
[0151] In some embodiments, noise point detection is performed on the fused prediction block based on threshold processing, including: calculating the gradient amplitude of each pixel in the fused prediction block, and determining the pixel whose gradient amplitude is higher than the threshold gradient amplitude as a noise point.
[0152] In some embodiments, the gradient amplitude of each pixel point in the fused prediction block may be calculated by using a Prewitt operator, a Roberts operator, a Sobel operator, a Laplacian operator, a Scharr operator, or the like.
[0153] For example, the Prewitt operator is used to calculate the gradient amplitude of each pixel in the horizontal and vertical directions, including:
[0154] First, create the horizontal Prewitt kernel P x and the vertical Prewitt kernel P y .
[0155]
[0156]
[0157] Secondly, align the center point of the convolution kernel with each pixel of the fusion prediction block and perform convolution operation to obtain the gradient amplitude P of the pixel in the horizontal direction. x and the vertical gradient amplitude P y .
[0158] That is, using the horizontal Prewitt kernel P x and the vertical Prewitt kernel P y Separate convolution fusion prediction blocks Get the gradient amplitude G of the image in the horizontal direction x and the vertical gradient magnitude G y . It can be specifically expressed as the following formula (6) and formula (7):
[0159]
[0160]
[0161] in Represents a convolution operation.
[0162] Finally, according to the gradient amplitude P of each pixel point in the horizontal direction of the fused prediction block x and the vertical gradient amplitude P y Calculate the gradient magnitude of each pixel of the fused prediction block.
[0163] Exemplarily, according to the gradient amplitude P of each pixel point in the fusion prediction block in the horizontal direction x and the vertical gradient amplitude P y Calculating the gradient magnitude of each pixel point of the fused prediction block may include: calculating the gradient magnitude GM of each pixel point of the fused prediction block according to the following formula (8):
[0164]
[0165] Step ②: If the number of noise points in the fused prediction block is greater than the threshold number, the filtering strength of each pixel point in the fused prediction block is obtained.
[0166] In some embodiments, the method for obtaining the filtering strength of each pixel point in the fused prediction block can be similar to the method for obtaining the filtering strength of each pixel point in the initial prediction block in the above embodiment. To avoid redundancy, it will not be described in detail here.
[0167] Step ③: filter each pixel in the fused prediction block based on the filtering strength of each pixel in the fused prediction block.
[0168] It should be noted that the filter used to filter each pixel in the fused prediction block may be the same as the filter (target filter) used to filter the pixel in the initial prediction block, or may be different from the filter used to filter the pixel in the initial prediction block.
[0169] The present application also provides a video decoding method, referring to Fig. 9 As shown, the video decoding method includes the following steps S91 to S96:
[0170] S91. Obtain encoding data of a target image block.
[0171] The target image block is a rectangular image block obtained by dividing the video frame into blocks.
[0172] In some embodiments, obtaining the encoding data of the target image block includes: receiving video stream data of a to-be-played video sent by a media resource server, and extracting the encoding data of the target image block from the video stream data.
[0173] In some embodiments, the encoding data of the target image block includes: first encoding data, second encoding data, third encoding data and fourth encoding data; the first encoding data is encoding data obtained by encoding information of a segmentation curve used to segment the target image block into a first sub-image block and a second sub-image block; the second encoding data is encoding data obtained by encoding residual data of the target image block, and the third encoding data and the fourth encoding data are encoding data obtained by encoding motion vectors of the first sub-image block and the second sub-image block, respectively.
[0174] In some embodiments, the information of the segmentation curve for segmenting the target image block into the first sub-image block and the second sub-image block includes: position information of the starting point, position information of the end point and position information of the control point of the second-order Bezier curve.
[0175] In some embodiments, the information of the segmentation curve for segmenting the target image block into the first sub-image block and the second sub-image block includes: the position information of the starting point, the end point, the first control point and the second control point of the third-order Bezier curve.
[0176] In some embodiments, the information of the segmentation curve for segmenting the target image block into the first sub-image block and the second sub-image block includes: position information of control points of the second-order Bezier curve.
[0177] In some embodiments, the information of the segmentation curve for segmenting the target image block into the first sub-image block and the second sub-image block includes: position information of the first control point and the position information of the second control point of the third-order Bezier curve.
[0178] S92. Obtain residual data, a first prediction block, and a second prediction block of the target image block according to the encoded data of the target image block.
[0179] The first prediction block and the second prediction block are prediction blocks of a first sub-image block and a second sub-image block respectively, and the first sub-image block and the second sub-image block are two sub-image blocks obtained by segmenting the target image block based on a texture edge in the target image block.
[0180] In some embodiments, the encoded data corresponding to the target image block includes: second encoded data obtained by performing transformation, quantization, entropy encoding and other operations on the residual data of the target image block, and the residual data of the target image block can be obtained by performing inverse operations such as entropy decoding, inverse transformation, and inverse quantization on the second encoded data.
[0181] In some embodiments, the encoding data corresponding to the target image block includes: first encoding data obtained by encoding information of a segmentation curve for segmenting the target image block into two sub-image blocks, third encoding data obtained by encoding a motion vector of the first sub-image block, and fourth encoding data obtained by encoding a motion vector of the second sub-image block. Acquiring the first prediction block and the second prediction block according to the encoding data of the target image block includes: performing entropy decoding, inverse transformation, inverse quantization and other operations on the first encoding data, the third encoding data and the fourth encoding data to obtain the segmentation curve, the motion vector of the first sub-image block and the motion vector of the second sub-image block, segmenting the target image block based on the segmentation curve, and determining the reference blocks corresponding to the two sub-image blocks according to the motion vector of the first sub-image block and the motion vector of the second sub-image block, and performing motion compensation on the reference blocks of the two sub-image blocks according to the motion vectors of the two sub-image blocks, so as to obtain the first prediction block and the second prediction block respectively.
[0182] S93. Combine the first prediction block and the second prediction block to obtain an initial prediction block, and obtain a boundary line between the first prediction block and the second prediction block in the initial prediction block.
[0183] S94: Determine the filtering strength of each pixel in the initial prediction block according to the minimum distance from each pixel in the initial prediction block to the boundary line.
[0184] S95. Filter each pixel in the initial prediction block based on the filtering strength of each pixel in the initial prediction block to obtain a fused prediction block.
[0185] The implementation method of steps S93 to S95 can refer to the above steps S44 to S46, and to avoid redundancy, they will not be described in detail here.
[0186] S96. Obtain a reconstructed image block of the target image block according to the residual data of the fused prediction block and the target image block.
[0187] The video decoding method provided in the embodiment of the present application can reconstruct the target image block based on the encoded data of the target image block obtained by the video encoding provided in the above embodiment to obtain the reconstructed image block of the target image block. Therefore, the embodiment of the present application can achieve normal decoding of the video while improving the performance of video encoding.
[0188] In some embodiments, some embodiments of the present application provide a video encoding device, the video encoding device comprising:
[0189] a memory configured to store a computer program;
[0190] The processor is configured to enable the video encoding device to implement the video encoding method described in any of the above embodiments when calling the computer program.
[0191] In some embodiments, some embodiments of the present application provide a video decoding device, the video decoding device comprising:
[0192] a memory configured to store a computer program;
[0193] The processor is configured to enable the video decoding device to implement the video decoding method described in any of the above embodiments when calling the computer program.
[0194] In some embodiments, some embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computing device, the computing device implements the video encoding method described in any of the above embodiments or the video decoding method described in any of the above embodiments.
[0195] In some embodiments, some embodiments of the present application provide a computer program product, which, when executed on a computer, enables the computer to implement the video encoding method described in any of the above embodiments or the video decoding method described in any of the above embodiments.
[0196] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0197] For the convenience of explanation, the above description has been made in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and various different variations of the embodiments suitable for specific use considerations.
Claims
1. A video encoding method, It is characterized in that include: Obtaining a target image block, wherein the target image block is a rectangular image block obtained by dividing the video frame into blocks; dividing the target image block into a first sub-image block and a second sub-image block based on a texture edge in the target image block; Acquire a first prediction block of the first sub-image block and a second prediction block of the second sub-image block; Combining the first prediction block and the second prediction block to obtain an initial prediction block, and obtaining a boundary line between the first prediction block and the second prediction block in the initial prediction block; Determining the filtering strength of each pixel point in the initial prediction block according to the minimum distance from each pixel point in the initial prediction block to the boundary line; Filtering each pixel in the initial prediction block based on the filtering strength of each pixel in the initial prediction block to obtain a fused prediction block; Calculating the residual between the target image block and the fused prediction block to obtain residual data of the target image block; Generate encoding data of the target image block according to the residual data of the target image block.
2. The method according to claim 1, It is characterized in that The filtering strength of each pixel point in the initial prediction block is negatively correlated with the minimum distance from each pixel point in the initial prediction block to the boundary line.
3. The method according to claim 2, It is characterized in that The step of determining the filtering strength of each pixel in the initial prediction block according to the minimum distance from each pixel in the initial prediction block to the boundary line comprises: Setting the filtering strength of each pixel point in a first area of the initial prediction block to a preset filtering strength, wherein the first area is an area composed of pixel points adjacent to the boundary line; The filtering strength of each pixel in the second area of the initial prediction block is calculated according to the minimum distance from each pixel in the second area to the boundary line, and the second area is an area composed of pixels not adjacent to the boundary line.
4. The method according to claim 3, It is characterized in that Before calculating the filtering strength of each pixel point in the second area according to the minimum distance from each pixel point in the second area to the boundary line, the method further includes: For each pixel point in the second area, the pixel distance between the pixel point and each pixel point in the first area is calculated, and the minimum value of the pixel distance between the pixel point and each pixel point in the first area is determined as the minimum distance from the pixel point to the boundary line.
5. The method according to claim 4, It is characterized in that The obtaining, according to the minimum distance from each pixel point in the second area to the boundary line, the filtering strength of each pixel point in the second area includes: Calculating the reciprocal of the square value of the minimum distance from each pixel point in the second area to the boundary line to obtain a first calculated value of each pixel point in the second area; Calculating the reciprocal of the first calculated value of each pixel point in the second area to obtain a second calculated value of each pixel point in the second area; The second calculated value of each pixel in the second area is determined as the filtering strength of each pixel in the second area.
6. The method according to claim 1, It is characterized in that The filtering each pixel in the initial prediction block based on the filtering strength of each pixel in the initial prediction block to obtain a fused prediction block includes: Converting the filtering strength of each pixel in the initial prediction block into a filtering parameter of a target filter to obtain the filtering parameter of each pixel in the initial prediction block; Acquire a filter matrix for each pixel in the initial prediction block according to a filter parameter of each pixel in the initial prediction block and a response characteristic of the target filter; Each pixel point in the initial prediction block is filtered according to the filter matrix of each pixel point in the initial prediction block to obtain a fused prediction block.
7. The method according to claim 1, It is characterized in that The filtering each pixel in the initial prediction block based on the filtering strength of each pixel in the initial prediction block to obtain a fused prediction block includes: Acquire a set of pixels to be filtered according to the filtering strength of each pixel in the initial prediction block, wherein the set of pixels to be filtered is a set of pixels in the initial prediction block whose filtering strength is greater than a strength threshold; Each pixel point in the set of pixel points to be filtered is filtered according to the filtering strength of each pixel point in the set of pixel points to be filtered to obtain a fused prediction block.
8. The method according to any one of claims 1 to 7, It is characterized in that The step of dividing the target image block into a first sub-image block and a second sub-image block based on a texture edge in the target image block comprises: Dividing the target image block into two image areas based on a texture edge in the target image block, and acquiring a boundary line between the two image areas; Performing curve fitting on the dividing line to obtain a segmentation curve of the target image block; Sampling the segmentation curve with integer pixel accuracy to obtain a sampling result of the segmentation curve; The target image block is segmented into the first sub-image block and the second sub-image block based on the sampling result of the segmentation curve.
9. The method according to any one of claims 1 to 7, It is characterized in that The step of dividing the target image block into a first sub-image block and a second sub-image block based on a texture edge in the target image block comprises: Dividing the target image block into two image areas based on a texture edge in the target image block, and acquiring a boundary line between the two image areas; Obtain two endpoints of the segmentation curve according to the encoded video content; Determine the segmentation curve according to two endpoints of the segmentation curve and the dividing line; Sampling the segmentation curve with integer pixel accuracy to obtain a sampling result of the segmentation curve; The target image block is segmented into the first sub-image block and the second sub-image block based on the sampling result of the segmentation curve.
10. The method according to any one of claims 1 to 7, It is characterized in that The obtaining a first prediction block of the first sub-image block and a second prediction block of the second sub-image block comprises: Performing motion estimation on the first sub-image block to obtain a first reference block and a first motion vector; Performing motion compensation on the first reference block based on the first motion vector to obtain the first prediction block; Performing motion estimation on the second sub-image block to obtain a second reference block and a second motion vector; Perform motion compensation on the first reference block based on the first motion vector to obtain the first prediction block.
11. The method according to any one of claims 1 to 7, It is characterized in that Before calculating the residual between the target image block and the fused prediction block to obtain residual data of the target image block, the method further includes: Performing noise point detection on the fused prediction block based on a preset noise detection algorithm; If the number of noise points in the fused prediction block is greater than a threshold number, obtaining the filtering strength of each pixel point in the fused prediction block; Each pixel in the fused prediction block is filtered based on the filtering strength of each pixel in the fused prediction block.
12. A video decoding method, It is characterized in that include: Obtaining encoded data of a target image block, wherein the target image block is a rectangular image block obtained by dividing a video frame into blocks; Obtain residual data, a first prediction block, and a second prediction block of the target image block according to the encoded data of the target image block, wherein the first prediction block and the second prediction block are prediction blocks of a first sub-image block and a second sub-image block, respectively, and the first sub-image block and the second sub-image block are two sub-image blocks obtained by segmenting the target image block based on a texture edge in the target image block; Combining the first prediction block and the second prediction block to obtain an initial prediction block, and obtaining a boundary line between the first prediction block and the second prediction block in the initial prediction block; Determining the filtering strength of each pixel point in the initial prediction block according to the minimum distance from each pixel point in the initial prediction block to the boundary line; Filtering each pixel in the initial prediction block based on the filtering strength of each pixel in the initial prediction block to obtain a fused prediction block; A reconstructed image block of the target image block is obtained according to the residual data of the fused prediction block and the target image block.
13. A video encoding device, It is characterized in that include: a memory configured to store a computer program; The processor is configured to enable the video encoding device to implement the video encoding method according to any one of claims 1 to 11 when calling a computer program.
14. A video decoding device, It is characterized in that include: a memory configured to store a computer program; The processor is configured to enable the video decoding device to implement the video decoding method according to claim 12 when calling the computer program.
Citation Information
Cited By
Intelligent traffic video data coding method and system
CN120529080A
Image transmission optimization method
CN120602648A
A method for optimizing image transmission
CN120602648B