Video coding method, decoding method, electronic terminal and storage medium

By determining compression parameters and bitrate control parameters based on the feature information of the target object in the video frame, the problem of difficulty in balancing subjective quality and compression rate in existing technologies is solved, and a balance between improving subjective quality and compression rate is achieved without increasing the encoding bitrate is realized.

CN121985126APending Publication Date: 2026-05-05ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG DAHUA TECH CO LTD
Filing Date
2025-12-12
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing video coding methods struggle to achieve optimal compression rates while ensuring the target's subjective quality.

Method used

Compression parameters, including compression level and level parameters, are determined based on the feature information of each target object in the current frame. Based on these parameters, code control parameters are determined for encoding to obtain bitstream data.

Benefits of technology

While ensuring the subjective quality of the image, the optimal compression ratio was achieved, improving both the objective effect and subjective quality of the encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121985126A_ABST
    Figure CN121985126A_ABST
Patent Text Reader

Abstract

The invention provides a video coding method, a video decoding method, an electronic terminal and a storage medium. The video coding method provided by the invention comprises the following steps: determining a compression parameter corresponding to a current frame according to feature information of each target object in the current frame; wherein the compression parameter corresponding to the current frame comprises a compression level corresponding to the current frame and / or a level parameter corresponding to the target object; determining a code control parameter of the current frame based on the compression parameter; and coding the current frame based on the code control parameter to obtain code stream data. Therefore, the compression parameter of the image frame is mapped based on the feature of the target object in the image, and the compression parameter is associated with the code control parameter, so that the subjective quality of the target object in the image can be ensured, and the compression rate can be optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image technology, and in particular to a video encoding method, a decoding method, an electronic terminal, and a storage medium. Background Technology

[0002] In recent years, with the development of video and media transmission technologies, higher demands have been placed on video quality. To provide a better video playback experience, internet video platforms have seen a surge in video bitrates. However, some existing video encoding methods cannot simultaneously achieve both target subjective quality and compression rate. Summary of the Invention

[0003] This application provides a video encoding method, a decoding method, an electronic terminal, and a storage medium. The method of this application can achieve the optimal compression rate while ensuring the subjective quality of the target.

[0004] To solve the above-mentioned technical problems, the first technical solution adopted in this application is: to provide a video coding method, including: The compression parameters corresponding to the current frame are determined based on the feature information of each target object in the current frame; wherein, the compression parameters corresponding to the current frame include the compression level corresponding to the current frame and / or the level parameter corresponding to the target object; The code control parameters of the current frame are determined based on the compression parameters; The current frame is encoded based on the code control parameters to obtain the bitstream data.

[0005] In one embodiment, determining the compression parameters corresponding to the current frame based on the feature information of each target object in the current frame includes: The compression level corresponding to the current frame is determined based on the preset quality level and the level parameters of each target object, thereby determining the compression parameters corresponding to the current frame.

[0006] In one embodiment, determining the compression level corresponding to the current frame based on a preset quality level and a level parameter for each target object includes: The region level of the current frame is determined based on the level parameters of each target object, the screen proportion of each target object, and the resolution of the current frame. The first complexity of the current frame is determined based on the first temporal texture information and the first spatial texture information of the current frame. The compression level of the current frame is determined based on the first complexity of the current frame, the regional level of the current frame, and the preset quality level.

[0007] In one embodiment, determining the region level of the current frame based on the level parameter of each target object, the screen proportion of each target object, and the resolution of the current frame includes: The total screen percentage of all target objects is calculated based on the level parameters of each target object, the screen percentage of each target object, and the resolution of the current frame. The region level of the current frame is determined based on the total screen occupancy of all target objects. Among them, the region level of the current frame is positively correlated with the total screen area.

[0008] In one embodiment, determining the compression parameters corresponding to the current frame based on the feature information of each target object in the current frame includes: Determine the key parameters of the target object based on its attribute information; The second complexity of the target object is determined based on the second temporal texture information and the second spatial texture information of the target object; Based on the key parameters of the target object, its second degree of complexity, and the preset quality level, the corresponding level parameters of the target object are determined, thereby determining the compression parameters corresponding to the current frame.

[0009] In one embodiment, determining the compression parameters corresponding to the current frame based on the feature information of each target object in the current frame includes: Feature extraction is performed on the current frame to obtain the region of interest corresponding to the current frame; Masking is performed on the region of interest to determine the level parameters of the target object, thereby determining the compression parameters corresponding to the current frame.

[0010] In one embodiment, determining the code control parameters of the current frame based on compression parameters includes: Based on the compression level corresponding to the current frame in the compression parameters, at least one of the following is determined: the global quantization parameters of the current frame, the quantization parameter offset value of the target object, the reference frame interval length, and the frame rate of the current frame; thereby determining the code control parameters of the current frame; and / or Based on the level parameters corresponding to the target object in the compression parameters, at least one of the following is determined: the target quantization parameter of the target object, the quantization parameter offset value of the target object, the reference frame interval length, and the frame rate of the current frame, thereby determining the code control parameters of the current frame. In one embodiment, determining the quantization parameter offset value of the target object based on the compression parameters includes: The quantization parameter offset value of the target object is determined based on the target object's level parameter in the compression parameters.

[0011] In one embodiment, before determining the compression parameters corresponding to the current frame based on the feature information of each target object in the current frame, the method further includes: Image enhancement is performed on the foreground region of the current frame and image filtering is performed on the background region of the current frame according to the preset quality level; Among them, the intensity of image enhancement is positively correlated with the preset quality level, while the intensity of image filtering is negatively correlated with the preset quality level.

[0012] In one embodiment, after encoding the current frame based on code control parameters to obtain bitstream data, the method further includes: The location information of the target object, the preset quality level, and the bitstream data are packaged together to obtain the encoded data.

[0013] To solve the above-mentioned technical problems, the second technical solution adopted in this application is: to provide a video decoding method, including: Obtain bitstream data; The bitrate control parameters are obtained from the bitrate data, and the bitrate data is decoded based on the bitrate control parameters to obtain the video data; The code control parameters are determined based on the compression parameters, which are determined according to the feature information of each target object in the image frame. The compression parameters include the compression level corresponding to the current frame and / or the level parameter corresponding to the target object.

[0014] In one embodiment, acquiring bitstream data includes: Obtain encoded data. The encoded data is decompressed to obtain the bitstream data, and the location information and preset quality level of the target object corresponding to each image frame in the video data corresponding to the bitstream data are obtained. Video decoding methods also include: Image processing is performed on the image frame based on the location information of each target object in the image frame and the preset quality level.

[0015] To solve the above-mentioned technical problems, the third technical solution adopted in this application is: to provide an electronic terminal, which includes a memory and a processor coupled to each other, the processor being used to execute program instructions stored in the memory, and the processor being used to execute program data to implement the steps in the video encoding method or the video decoding method described above.

[0016] To solve the above-mentioned technical problems, the fourth technical solution adopted in this application is: to provide a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, to implement the steps in the video encoding method or the video decoding method described above.

[0017] The beneficial effects of this application are as follows: Unlike existing technologies, the video coding method provided in this application determines the compression parameters corresponding to the current frame based on the feature information of each target object in the current frame; wherein, the compression parameters corresponding to the current frame include the compression level corresponding to the current frame and / or the level parameter corresponding to the target object; the code control parameters of the current frame are determined based on the compression parameters; and the current frame is encoded based on the code control parameters to obtain the bitstream data. In this way, by mapping the compression parameters of the image frame based on the features of the target object in the image and associating the compression parameters with the code control parameters, the subjective quality of the target object in the image can be guaranteed, and the compression ratio can be optimized. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating an embodiment of the video encoding method provided in this application; Figure 2 This is a flowchart illustrating the first embodiment of the encoding method for determining the level parameter corresponding to the target object. Figure 3a This is a schematic diagram of the frame difference method for calculating temporal texture information provided in this application; Figure 3b This is a schematic diagram illustrating the calculation of temporal texture information for gradient values ​​provided in this application; Figure 4 This is a flowchart illustrating the second embodiment of the encoding method for determining the level parameter corresponding to the target object. Figure 5 This application Figure 2 A flowchart illustrating another embodiment of step S21; Figure 6 This is a flowchart illustrating an embodiment of the encoding method of this application for determining the compression level corresponding to the current frame; Figure 7 This is a flowchart illustrating an embodiment of the video decoding method provided in this application; Figure 8 This is a schematic diagram of the framework of an embodiment of the electronic terminal provided in this application; Figure 9 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium provided in this application. Detailed Implementation

[0020] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0021] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0022] In this article, the term "and / or" simply describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "more" in this article means two or more objects.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0024] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0025] The video encoding method provided in this application can be implemented by a server or terminal alone, or by a server and terminal working together. In some embodiments, the terminal or server can implement the video encoding method provided in this application by running a computer program. For example, the computer program can be a native program or software module in an operating system; it can be a native application (APP), i.e., a program that needs to be installed in the operating system to run, such as a client that supports virtual scenes, such as a game APP; it can also be a mini-program, i.e., a program that only needs to be downloaded to a browser environment to run; or it can be a mini-program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of application, module, or plugin.

[0026] To enable those skilled in the art to better understand the technical solution of this application, a video coding method provided by this application will be described in further detail below with reference to the accompanying drawings and specific embodiments.

[0027] Please see Figure 1 This is a flowchart illustrating the first embodiment of the video encoding method of this application, specifically including: Step S11: Determine the compression parameters corresponding to the current frame based on the feature information of each target object in the current frame.

[0028] Understandably, each image frame will contain some target objects, such as stationary or moving targets. The level of attention given to different target objects varies in different applications. Taking the currently encoded image frame as the current frame, this application determines the compression parameters of the current frame based on the feature information of each target object in the current frame. It should be noted that the compression parameters of the current frame include the compression level corresponding to the current frame and / or the level parameters corresponding to the target objects. Understandably, the feature information of the target objects may be, for example, the attribute information of the target objects (attribute information is used to characterize the category of the target objects), or the feature information of the target objects may also be, for example, the degree of attention given to the target objects (used to characterize the importance of the target objects). In this way, the category and importance of the target objects can be associated with the compression parameters.

[0029] In one specific embodiment, the compression level corresponding to the current frame is the global compression level for the entire image frame. The compression level parameter corresponding to the target object refers to the compression level parameter for each target object in the image.

[0030] In one embodiment, before determining the compression level of the current frame, image enhancement is performed on the foreground region of the current frame according to a preset quality level, and image filtering is performed on the background region of the current frame; wherein, the intensity of image enhancement is positively correlated with the preset quality level, and the intensity of image filtering is negatively correlated with the preset quality level.

[0031] Specifically, a quality management interface can be provided in the device, allowing users to set preset quality levels for image quality management according to their needs. For example, a quality level selection interface can be provided, with quality level values ​​ranging from 1 to N, which users can customize. In one embodiment, the quality level value range is, for example, 1 to 10.

[0032] Image enhancement is performed on the foreground region of the current frame based on a preset quality. The foreground region is the area where the target object is located in the current frame. In some embodiments, an intelligent algorithm is used to determine all target objects contained in the current frame, and the area where the target object is located is designated as the foreground region; the remaining areas in the current frame excluding the target object are designated as the background region.

[0033] Image enhancement of the foreground region of the current frame according to a preset quality level specifically includes: if the preset quality level is greater than or equal to a first preset value, then secondary image enhancement is performed on the foreground region; if the preset quality level is less than the first preset value but greater than or equal to a second preset value, then primary image enhancement is performed on the foreground region; if the preset quality level is less than the second preset value, then no image enhancement is performed. Wherein, if the first preset value is greater than the second preset value, the intensity of secondary image enhancement is greater than the intensity of primary image enhancement. For example, the first preset value is 7, and the second preset value is 5; of course, in other embodiments, the first and second preset values ​​can also be other values, and are not specifically limited. It is understood that the higher the preset quality level, the greater the image enhancement intensity, thus resulting in better subjective quality of the final compressed image.

[0034] Image filtering of the background region of the current frame according to a preset quality level specifically includes: if the preset quality level is less than or equal to a third preset value, then a first-level image filtering is performed on the background region; if the preset quality level is less than or equal to a fourth preset value, then a second-level image filtering is performed on the background region; if the preset quality level falls within other ranges, then no image filtering is performed. The third preset value is greater than the fourth preset value, and the third preset value is less than the second preset value; the filtering intensity of the first-level image filtering is less than the filtering intensity of the second-level image filtering. For example, the third preset value is 4, and the fourth preset value is 2. Of course, in other embodiments, the third and fourth preset values ​​can also be other values, and there is no specific limitation. It is understood that the higher the preset quality level, the lower the intensity of image filtering, thus allowing the image to retain more details.

[0035] Specifically, the level parameter corresponding to the target object represents the importance of the target object.

[0036] In one embodiment, a management interface can be provided in the device, allowing users to customize the level parameters corresponding to the target objects. For example, the level parameter for the first type of target object can be set to 10, the level parameter for the second type of target object can be set to 4, and the level parameter for the third type of target object can be set to 1. This indicates that the first type of target object is more important than the second type of target object, and the second type of target object is more important than the third type of target object.

[0037] In another embodiment, the level parameter corresponding to the target object can be calculated and determined. Specifically, in conjunction with... Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the coding method of this application for determining the level parameter corresponding to the target object.

[0038] Step S31: Determine the important parameters of the target object based on its attribute information.

[0039] In some embodiments, intelligent algorithms are used to determine all target objects contained in the current frame (which can be an image after image enhancement and filtering). For example, a target recognition algorithm is used to identify first-type, second-type, and third-type target objects in the current frame. Specifically, the position information of each type of target object is determined, the coordinates of each type of target object are recorded, a bounding box of each target object is formed, and the attribute information, i.e., type information, of each target object is recorded. Important parameters of the target objects are determined based on their attribute information. Taking the first-type target object as an example, assuming that the first-type target object belongs to the type of focus, important parameters of the first-type target object are set. .

[0040] Step S32: Determine the second complexity of the target object based on the second temporal texture information and the second spatial texture information of the target object.

[0041] In one embodiment, the second temporal texture information and the second spatial texture of the target object can be calculated on a pixel-by-pixel basis. In other real-time examples, the second temporal texture information and the second spatial texture of the target object can also be calculated on a macroblock basis in the current frame.

[0042] In one embodiment, temporal texture information can be obtained through motion detection, including but not limited to frame difference methods, Gaussian models, etc. Combined with... Figure 3a Taking the frame difference method as an example to calculate temporal texture information, the frame difference method determines whether the magnitude of pixel change exceeds a threshold by analyzing the pixel differences between consecutive frames. Figure 3a As shown, with a threshold of 2, the top-left pixel changes from 5 to 1, and the top-right pixel changes from 1 to 5, with a pixel change value of 4, both exceeding the threshold of 2. Therefore, these two pixels are considered moving pixels. The bottom-left pixel changes from 2 to 3, while the bottom-right pixel remains unchanged, with pixel differences of 1 and 0, respectively, which are less than the threshold of 2. Therefore, these two pixels are considered stationary pixels. The second temporal texture information corresponding to the target object can be obtained by averaging the frame differences of the pixels within the bounding rectangle of the target object across consecutive frames. It can be understood that if the second temporal texture information of the target object is calculated using macroblocks of the current frame as units, then based on the macroblock size, the frame differences of each pixel within the macroblock are averaged to obtain the temporal texture information corresponding to that macroblock. Further averaging the temporal texture information of the macroblocks within the bounding rectangle of the target object yields the second temporal texture information corresponding to the target object.

[0043] In one embodiment, spatial texture information can be obtained by calculating image complexity, gradient, Sobel operator, etc. Figure 3bTaking the gradient algorithm as an example to calculate spatial texture information, the gradient calculation method is to average the differences between the current pixel and its surrounding pixels to obtain the gradient value of the current pixel. Figure 3b As shown, the gradient value of the center pixel (i.e., the current pixel) is calculated as: (|3-1|+|3-5|+|3-5|+|3-3|+|3-3|+|3-4|+|3-5|+|3-6|) / 8=1.5. The gradient values ​​of the pixels within the bounding rectangle of the target object are calculated using the above method, and then the gradient values ​​of the pixels within the bounding rectangle are averaged to obtain the second spatial texture information corresponding to the target object. It can be understood that if the second spatial texture information of the target object is calculated using macroblocks of the current frame as units, then based on the macroblock size, the gradient values ​​of each pixel within the macroblock are averaged to obtain the spatial texture information corresponding to that macroblock. Further averaging the spatial texture information of the macroblocks within the bounding rectangle of the target object yields the second spatial texture information corresponding to the target object.

[0044] After calculating the second temporal texture information and the second spatial texture information of the target object in the above manner, the second complexity of the target object is determined based on the second temporal texture information and the second spatial texture information of the target object.

[0045] In one embodiment, the second temporal texture information and the second spatial texture information of the target object are mapped using the following formula (1) to determine the second complexity of the target object: Formula (1); In formula (1), This indicates the second level of complexity of the target object. This represents the second temporal texture information of the target object. f3 represents the second spatial texture information of the target object, and f3 represents the first mapping algorithm.

[0046] In one embodiment, the first mapping algorithm is, for example, a weighted summation algorithm. Assuming the second temporal texture information ranges from 1 to 10, the second spatial texture information ranges from 1 to 10, and assuming the current target object is not complex but is a moving target, then the second temporal texture information is valued at 10, the second spatial texture information is valued at 2, and the weights of the second temporal and second spatial texture information are each set to 0.5. The second complexity of the current target object is then calculated as follows: =0.5* +0.5* =10*0.5+2*0.5=6.

[0047] The weighted summation algorithm described above is only one example of the first mapping algorithm. In other embodiments, the first mapping algorithm may also be other types of algorithms, and there is no specific limitation.

[0048] Step S33: Determine the level parameters corresponding to the target object based on the important parameters of the target object, the second complexity, and the preset quality level, thereby determining the compression parameters corresponding to the current frame.

[0049] In step S31, the key parameters of the target object are determined. And in step S32, the second complexity of the target object is determined. Furthermore, based on the key parameters of the target object, its second level of complexity, and the preset quality level, a mapping is performed to determine the corresponding level parameters for the target object.

[0050] In one embodiment, the level parameter corresponding to the target object is calculated using the following formula (2). : Formula (2); In formula (2), quality represents the preset quality level, and f4 represents the second mapping algorithm.

[0051] In one embodiment, the second mapping algorithm f4 is, for example: ; Among them, quality max This indicates the maximum value within the preset quality level range. For example, if the preset quality level range is 1-10, then quality... max It is 10.

[0052] In other embodiments, the second mapping algorithm f4 can also be other algorithms, and there is no specific limitation. The level parameter of the target object can be calculated in the above manner. .

[0053] In this embodiment, the target object's level parameters are mapped based on a combination of the target object's attribute information, temporal texture information, and spatial texture information. This is used to determine the code control parameters, thereby improving subjective quality after encoding and ensuring subsequent intelligent analysis of the target.

[0054] In another embodiment of this application, a deep learning network can also be used to process the current frame to obtain the level parameters of the target object. Combined with... Figure 4 , Figure 4 This is a flowchart illustrating the second embodiment of the coding method for determining the level parameter corresponding to the target object in this application, specifically including: Step S51: Extract features from the current frame to obtain the region of interest corresponding to the current frame.

[0055] Specifically, in combination Figure 5 The brightness channel image of the current frame is input into the first convolutional layer. The first convolutional layer outputs two branches: the first branch outputs the feature map, and the second branch is used to obtain the region of interest (ROI). The first convolutional layer is, for example, a 3×3 convolutional layer.

[0056] The ROI (Region of Interest) branch is further processed using a second convolutional layer (3×3). The output of the second convolutional layer is connected to a third convolutional layer (1×1) and a fourth convolutional layer (1×1). The fourth convolutional layer is used to calculate the feature offset; the output of the third convolutional layer is used for dimensionality transformation. After dimensionality transformation, the softmax function is used for processing, followed by an inverse dimensionality transformation. This, along with the feature offset, determines the proposed coordinate information. The proposed coordinate information and the feature map together determine the division of the ROI, thus identifying the Region of Interest (ROI) corresponding to the current frame.

[0057] Step S52: Perform mask calculation on the region of interest to determine the level parameters of the target object, thereby determining the compression parameters corresponding to the current frame.

[0058] Sampling points are selected within the region of interest to obtain multiple sampling points. Then, bilinear interpolation is performed on each sampling point to calculate its feature value from neighboring pixels in the feature map. This process is then sequentially passed through a fifth and sixth convolutional layer to obtain a mask. This mask image represents the level parameters corresponding to the target object. The fifth and sixth convolutional layers are 3×3 convolutional layers.

[0059] The level parameters of each target object can be determined by the above method, and then the compression level corresponding to the current frame can be determined according to the preset quality level and the level parameters of each target object.

[0060] Specifically, in combination Figure 6 , Figure 6 This is a flowchart illustrating an embodiment of the encoding method of this application for determining the compression level corresponding to the current frame, specifically including: Step S71: Determine the region level of the current frame based on the level parameters of each target object, the screen proportion of each target object, and the resolution of the current frame.

[0061] In the steps described above, the bounding rectangle of each target object is determined during target recognition, and the frame percentage of the target object is determined based on this bounding rectangle. It can be understood that the frame percentage of a target object is, for example, the proportion of pixels within the target object's bounding rectangle to the total number of pixels in the current frame.

[0062] Current frame region level The calculation method is as follows: Formula (3); Where resolution represents the resolution of the current frame. This indicates the percentage of the image frame occupied by the target object, and f5 indicates the third mapping algorithm. Understandably, the higher the calculated region level, the higher the image quality desired by the user.

[0063] In one embodiment, the total screen percentage of all target objects is calculated in advance based on the grade parameter of each target object, the screen percentage of each target object, and the resolution of the current frame. Specifically, the total screen percentage of the target objects is calculated using the following formula (4). : Formula (4); Where n represents the total number of target objects. This represents the level parameter of the i-th target object. This indicates the maximum value of the target object's grade parameter. It's a custom setting, for example, 10. It can be the same as the maximum value within the preset quality grade range. The i-th target object represents the percentage of the frame, and resolution represents the resolution of the current frame.

[0064] The region level of the current frame is determined based on the total screen share of all target objects; the region level of the current frame is positively correlated with the total screen share.

[0065] In one embodiment, if the total screen area occupied by the target object is less than a first threshold, it indicates that the target object area in the current frame is relatively small, and the region level of the current frame is 0; if the total screen area occupied by the target object is greater than or equal to the first threshold but less than a second threshold, it indicates that the target object in the current frame is relatively large, and the region level of the current frame is 1; if the total screen area occupied by the target object is greater than or equal to the second threshold, it indicates that the target object area in the current frame is large, and the region level of the current frame is 2. It can be understood that the larger the area occupied by the target object in the current frame, the higher the region level of the current frame.

[0066] Step S72: Determine the first complexity of the current frame based on the first temporal texture information and the first spatial texture information of the current frame.

[0067] It is understandable that the calculation method of the first temporal texture information of the current frame is the same as that of the second temporal texture information of the target object, and the calculation method of the first spatial texture information of the current frame is the same as that of the second spatial texture information of the target object, which will not be repeated here.

[0068] Specifically, after determining the first temporal texture information and the first spatial texture information of the current frame, the first complexity of the current frame can be calculated using the following formula (5): Formula (5); in, Indicates the first level of complexity of the current frame. This represents the first temporal texture information of the current frame. This represents the first spatial texture information of the current frame, and f6 represents the fourth mapping algorithm. The fourth mapping algorithm can be the same as the first mapping algorithm.

[0069] Step S73: Determine the compression level corresponding to the current frame based on the first complexity of the current frame, the regional level of the current frame, and the preset quality level.

[0070] In one embodiment, the compression level corresponding to the current frame is calculated using the following formula (6). : Formula (6); in, Indicates the region level of the current frame. f7 indicates the first level of complexity for the current frame, and f7 indicates the fifth mapping algorithm.

[0071] In one embodiment, the fifth mapping algorithm f7 is: ; Among them, quality max This indicates the maximum value within the preset quality level range.

[0072] Step S12: Determine the code control parameters of the current frame based on the compression parameters.

[0073] Specifically, after calculating the compression parameters corresponding to the current frame (i.e., the compression level corresponding to the current frame and / or the level parameters corresponding to the target object) through the above process, the code control parameters of the current frame are determined based on the compression parameters.

[0074] In one specific embodiment, at least one of the following is determined based on the compression level corresponding to the current frame in the compression parameters: global quantization parameters of the current frame, quantization parameter offset value of the target object, reference frame interval length, and frame rate of the current frame, thereby determining the code control parameters of the current frame.

[0075] In one embodiment, the global quantization parameters of the current frame are determined based on the compression level corresponding to the current frame in the compression parameters. Specifically, the global quantization parameters of the current frame are calculated using the following formula (7): Formula (7); in, This represents the minimum value of the global quantization parameter. f8 represents the maximum value of the global quantization parameter and the sixth mapping algorithm.

[0076] In one embodiment, the compression level is assumed. =6, assuming the minimum value of the initial global quantization parameter is 10 and the maximum value is 29, then the calculated minimum value of the global quantization parameter is... =10+6=16, the maximum value of the global quantization parameter. =29+6=35, and These are used to limit the maximum and minimum quantization parameters of the current frame, respectively. The above embodiment uses addition to determine the global quantization parameters; in another embodiment, subtraction can also be used. Specifically, the minimum value of the calculated global quantization parameter... =10-6=4, the maximum value of the global quantization parameter. =29-6=23. In other embodiments, the global quantization parameters of the current frame can also be determined based on the compression level using other calculation methods, without any specific limitations.

[0077] In another embodiment, the reference frame interval length is determined based on the compression level. The reference frame is an I-frame, and the reference frame interval length is denoted as GOP. The GOP is calculated as follows: f8 represents the seventh mapping algorithm. In a specific embodiment, assuming the initial reference frame interval length GOPori = 50, the calculated reference frame interval length GOP = GOPori + =50 + 6 = 56. This embodiment uses the seventh mapping algorithm as an addition operation to determine the reference frame interval length based on the compression level. In other embodiments, the seventh mapping algorithm can also be used in other ways, such as subtraction, to determine the reference frame interval length based on the compression level; no specific limitation is made. Determining the reference frame interval length based on the compression level combines the compression level with the compression ratio, improving the compression ratio while maintaining image quality.

[0078] In another embodiment, the frame rate of the current frame is determined based on the compression level. The frame rate of the current frame is denoted as FPS. In one embodiment, the frame rate of the current frame is calculated as follows: f 10 This represents the eighth mapping algorithm. In one embodiment, the eighth mapping algorithm is: ;in, This represents the initial frame rate. The frame rate of the current frame is calculated based on the eighth mapping algorithm, combining the initial frame rate and the compression level. The compression level is then combined with the frame rate during compression of the current frame to improve the compression ratio while maintaining image quality.

[0079] In another embodiment, the quantization parameter offset value of the target object is determined based on the compression level. Specifically, the quantization parameter offset value of the target object is determined based on the compression level and the level parameter of the target object. In one embodiment, the quantization parameter offset value of the target object is calculated using the following formula (8). : Formula (8); in, Indicates the compression level. f represents the level parameter of the target object. 11 This represents the ninth mapping algorithm. In one specific embodiment, the ninth mapping algorithm is: ; in, The maximum value of the compression level is determined by the above formula (6) and the fifth mapping algorithm f7. The quantization parameter offset of the target object and the global quantization parameter determine the quantization parameters within the bounding rectangle of the target object.

[0080] In another embodiment, at least one of the following is determined based on the level parameter corresponding to the target object in the compression parameters: the target quantization parameter of the target object, the quantization parameter offset value of the target object, the reference frame interval length, and the frame rate of the current frame, thereby determining the code control parameters of the current frame. The calculation method in this embodiment is the same as that in the above embodiments; only the variable in the formula needs to be replaced from compression level to level parameter, and the specific method is not limited.

[0081] In another embodiment, a set of code control parameters can be calculated based on compression parameters and grade parameters respectively, and then the code control parameters with the best quality and compression ratio can be selected as the final code control parameters for encoding.

[0082] The video coding method of this application dynamically adjusts the global quantization parameters and quantization parameter offset values ​​according to the compression level and / or the level parameters of the target object. This allows for a more reasonable allocation of the encoded bits in the current frame, improving the bit allocation in areas of human visual focus and thus enhancing the subjective effect. It also increases the compression ratio while reducing the bits in areas of non-human visual focus, making the decrease in subjective effect virtually imperceptible. This method improves subjective quality and compression ratio without increasing the bitrate, achieving a balance between objective results, subjective quality, and compression ratio gains.

[0083] Step S13: Encode the current frame based on the code control parameters to obtain the bitstream data.

[0084] Specifically, the global quantization parameters of the current frame, the quantization parameter offset of the target object, the reference frame interval length, and the frame rate of the current frame are calculated as above and used as the code control parameters of the current frame to encode the current frame and obtain the bitstream data.

[0085] The video coding method proposed in this application determines information such as the importance, location, and screen proportion of the target object. It determines the region level of the current frame based on the level parameter representing the importance of the target object and the screen proportion of the target object. It then maps the compression level to the complexity, region level, and quality level of the current frame. Finally, it maps the compression level and / or level parameter to code control parameters and sends them to the encoder for encoding. This method can achieve the best results in terms of objective coding, subjective quality, and post-intelligent analysis.

[0086] In one embodiment, the target object's location information, a preset quality level, and the bitstream data are packaged to obtain encoded data. The target object's location information and the preset quality level are used for image enhancement at the decoding end. It should be noted that at the encoding end, the preset quality level is used to enhance the foreground region (i.e., the region where the target object is located) and the background region (other regions besides the target object) of the image. Therefore, at the decoding end, to improve image quality, the preset quality level is used again to enhance the foreground and background regions of the image. Thus, the target object's location information, the preset quality level, and the data transmitted with the bitstream data need to be transmitted to the decoding end.

[0087] The video encoding method of this application can improve subjective quality without increasing the encoding bitrate, thus achieving a balance between subjective quality, compression rate, and user needs.

[0088] Please see Figure 7 , Figure 7 A flowchart illustrating an embodiment of the video decoding method provided in this application specifically includes: Step S81: Obtain the bitstream data.

[0089] In one embodiment, since the code control parameters used in the video encoding of this application are determined based on the above method, during encoding, the location information of the target object, the level parameters corresponding to the target object, the compression level, and the bitstream data are packaged together to obtain encoded data. When the decoding end interprets the video, it first obtains the encoded data, decompresses the encoded data to obtain bitstream data, and obtains the location information of the target object and the preset quality level corresponding to each image frame in the video data corresponding to the bitstream data.

[0090] Step S82: Obtain the bitrate control parameters from the bitrate stream data, and decode the bitrate stream data based on the bitrate control parameters to obtain video data.

[0091] After obtaining the bitstream data, the decoding end extracts the bitstream control parameters from the bitstream data, and decodes the bitstream data based on the bitstream control parameters to obtain the video data.

[0092] The code control parameters are determined based on compression parameters, which are determined according to the feature information of each target object in the image frame. The compression parameters include the compression level corresponding to the current frame and / or the level parameter corresponding to the target object. In one specific embodiment, the global quantization parameter, target quantization parameter, quantization parameter offset value of the target object, reference frame interval, and frame rate of the image frame in the code control parameters are at least partially determined based on the compression level and / or the level parameter.

[0093] Specifically, since image enhancement and filtering are performed before video data encoding, further image processing is applied to the image frames based on the location information of each target object and a preset quality level. Specifically, given the location information of each target object, image processing is performed on the target object at the corresponding location based on the preset quality level.

[0094] In one specific embodiment, if the preset quality level is greater than or equal to a first preset value, then secondary image enhancement is performed on the foreground region, i.e., the target object; if the preset quality level is less than the first preset value but greater than or equal to a second preset value, then primary image enhancement is performed on the foreground region; if the preset quality level is less than the second preset value, then no image enhancement is performed. Wherein, if the first preset value is greater than the second preset value, the intensity of the secondary image enhancement is greater than the intensity of the primary image enhancement. For example, the first preset value is 7, and the second preset value is 5; of course, in other embodiments, the first and second preset values ​​can also be other values, and there is no specific limitation. It is understood that the higher the preset quality level, the greater the image enhancement intensity, thus resulting in better subjective quality of the final compressed image.

[0095] In another embodiment, if the preset quality level is less than or equal to a third preset value, then first-level image filtering is performed on the background region (i.e., the region in the image frame other than the target object); if the preset quality level is less than or equal to a fourth preset value, then second-level image filtering is performed on the background region; if the preset quality level falls within other ranges, then no image filtering is performed. The third preset value is greater than the fourth preset value, and the third preset value is less than the second preset value, meaning the filtering intensity of the first-level image filtering is less than the filtering intensity of the second-level image filtering. For example, the third preset value is 4, and the fourth preset value is 2; however, in other embodiments, the third and fourth preset values ​​can be other values, without limitation. It is understood that a higher preset quality level results in a lower image filtering intensity, thus allowing the image to retain more detail.

[0096] Please see Figure 8 , Figure 8This is a schematic diagram of a framework of an embodiment of the electronic terminal provided in this application. The electronic terminal 80 includes a memory 81 and a processor 82 coupled to each other. The processor 82 is used to execute program instructions stored in the memory 81 to implement the steps of any of the above-described video encoding method embodiments. In a specific implementation scenario, the terminal 80 may include, but is not limited to, a microcomputer or a server. In addition, the terminal 80 may also include mobile devices such as laptops and tablets, which are not limited here.

[0097] Specifically, processor 82 controls itself and memory 81 to implement the steps of any of the video encoding method embodiments described above. Processor 82 may also be referred to as a CPU (Central Processing Unit). Processor 82 may be an integrated circuit chip with signal processing capabilities. Processor 82 may also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. A general-purpose processor may be a microprocessor or any conventional processor. Furthermore, processor 82 may be implemented using integrated circuit chips.

[0098] Please see Figure 9 , Figure 9 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium provided in this application. The computer-readable storage medium 90 stores program instructions 901 that can be executed by a processor. The program instructions 901 are used to implement the steps of any of the above-described video encoding method embodiments.

[0099] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0100] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0101] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0102] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0103] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0104] The above are merely embodiments of this application and do not limit the scope of patent protection of this application. Any equivalent structural or procedural changes made using the content of this application’s specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of this application.

Claims

1. A video encoding method, characterized in that, include: The compression parameters corresponding to the current frame are determined based on the feature information of each target object in the current frame; wherein, the compression parameters corresponding to the current frame include the compression level corresponding to the current frame and / or the level parameter corresponding to the target object; The code control parameters of the current frame are determined based on the compression parameters; The current frame is encoded based on the code control parameters to obtain the bitstream data.

2. The video encoding method according to claim 1, characterized in that, The compression parameters corresponding to the current frame are determined based on the feature information of each target object in the current frame, including: The compression level corresponding to the current frame is determined based on the preset quality level and the level parameter of each target object, thereby determining the compression parameter corresponding to the current frame.

3. The video encoding method according to claim 2, characterized in that, Determining the compression level corresponding to the current frame based on a preset quality level and the level parameter of each target object includes: The region level of the current frame is determined based on the level parameter of each target object, the screen proportion of each target object, and the resolution of the current frame; The first complexity of the current frame is determined based on the first temporal texture information and the first spatial texture information of the current frame. The compression level of the current frame is determined based on the first complexity of the current frame, the regional level of the current frame, and the preset quality level.

4. The video encoding method according to claim 3, characterized in that, The region level of the current frame is determined based on the level parameter of each target object, the screen proportion of each target object, and the resolution of the current frame, including: The total screen percentage of all target objects is calculated based on the level parameter of each target object, the screen percentage of each target object, and the resolution of the current frame. The region level of the current frame is determined based on the total screen proportion of all target objects; The region level of the current frame is positively correlated with the total screen percentage.

5. The video encoding method according to claim 1, characterized in that, The compression parameters corresponding to the current frame are determined based on the feature information of each target object in the current frame, including: Determine the key parameters of the target object based on its attribute information; The second complexity of the target object is determined based on the second temporal texture information and the second spatial texture information of the target object; Based on the important parameters of the target object, the second complexity, and the preset quality level, the level parameter corresponding to the target object is determined, thereby determining the compression parameter corresponding to the current frame.

6. The video encoding method according to claim 1, characterized in that, The compression parameters corresponding to the current frame are determined based on the feature information of each target object in the current frame, including: Feature extraction is performed on the current frame to obtain the region of interest corresponding to the current frame; A mask calculation is performed on the region of interest to determine the level parameters of the target object, thereby determining the compression parameters corresponding to the current frame.

7. The video coding method according to any one of claims 1 to 6, characterized in that, Determining the code control parameters of the current frame based on the compression parameters includes: Based on the compression level corresponding to the current frame in the compression parameters, at least one of the following is determined: the global quantization parameter of the current frame, the quantization parameter offset value of the target object, the reference frame interval length, and the frame rate of the current frame, thereby determining the code control parameters of the current frame; and / or Based on the level parameter corresponding to the target object in the compression parameters, at least one of the following is determined: the target quantization parameter of the target object, the quantization parameter offset value of the target object, the reference frame interval length, and the frame rate of the current frame, thereby determining the code control parameters of the current frame.

8. The video encoding method according to claim 7, characterized in that, Determining the quantization parameter offset value of the target object based on the compression parameters includes: The quantization parameter offset value of the target object is determined based on the level parameter of the target object in the compression parameters.

9. The video encoding method according to claim 2, characterized in that, Before determining the compression parameters corresponding to the current frame based on the feature information of each target object in the current frame, the method further includes: Image enhancement is performed on the foreground region of the current frame and image filtering is performed on the background region of the current frame according to the preset quality level; The intensity of image enhancement is positively correlated with the preset quality level, while the intensity of image filtering is negatively correlated with the preset quality level.

10. The video encoding method according to claim 2, characterized in that, After encoding the current frame based on the code control parameters to obtain the bitstream data, the method further includes: The location information of the target object, the preset quality level, and the bitstream data are packaged together to obtain encoded data.

11. A video decoding method, characterized in that, include: Obtain bitstream data; The bitstream data is processed to obtain bitrate control parameters, and the bitstream data is decoded based on the bitrate control parameters to obtain video data. The code control parameters are determined based on compression parameters, which are determined according to the feature information of each target object in the image frame, and the compression parameters include the compression level corresponding to the current frame and / or the level parameter corresponding to the target object.

12. The video decoding method according to claim 11, characterized in that, Obtain the bitstream data, including: Obtain encoded data. The encoded data is decompressed to obtain the bitstream data, and the location information and preset quality level of the target object corresponding to each image frame in the video data corresponding to the bitstream data are obtained. The video decoding method further includes: Image processing is performed on the image frame based on the location information of each target object in the image frame and a preset quality level.

13. An electronic terminal, characterized in that, The electronic terminal includes a memory and a processor coupled to each other. The processor is used to execute program instructions stored in the memory and to execute program data to implement the steps in the video encoding method as claimed in any one of claims 1 to 10 or the steps in the video decoding method as claimed in any one of claims 11 to 12.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the video encoding method as described in any one of claims 1 to 10 or the steps of the video decoding method as described in any one of claims 11 to 12.