Static region detection method, video coding method, storage medium and electronic equipment
By detecting the static area of the encoding block in the video image and calculating the average value by sliding window sliding to judge the static area, the problem of inaccurate detection of static area in the prior art is solved, and the effective reduction of code rate in video encoding and the maintenance of visual quality is achieved.
Patent Information
- Application Number
- CN202510374546.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art is difficult to accurately detect static areas in video images, which makes it difficult to effectively reduce the code rate of video encoding.
By obtaining the difference between the brightness and chrominance components of the original pixel of the coded block and the predicted pixel, the static area detection point is determined, and the sliding window is used to slide in the detection point area to calculate the average value of the data points in the sliding window to determine whether the coded block is a static area.
It improves the accuracy and reliability of static area detection, saves the code rate in video encoding, and ensures subjective perception of visual quality.
Smart Images

Figure CN120238654A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of video image processing, and relates to a method for detecting static regions, in particular to a method for detecting static regions, a video encoding method, a storage medium, and an electronic device. Background Art
[0002] With the wide application of intelligent devices and the increasing popularity of social networks, the amount of massive video and digital image data generated in daily life is constantly increasing, which poses higher and higher requirements for video image data compression technology. One of the core challenges in the field of video encoding is how to effectively reduce the bit rate of video compression on the premise that the subjective perception of visual quality is not affected. For static regions in video images, usually these regions have less impact on the user's visual attention, so more efficient compression methods can be used to reduce the required bit rate. Therefore, how to accurately detect static regions in video images has become one of the technical problems that need to be solved urgently by relevant technical personnel. Summary of the Invention
[0003] Embodiments of this application provide a method for detecting static regions, a video encoding method, a storage medium, and an electronic device, which are used to accurately detect static regions in video images.
[0004] In a first aspect, embodiments of this application provide a method for detecting static regions. The method for detecting static regions includes: obtaining the difference between the luminance component and the difference between the chrominance components of the original pixels and the predicted pixels of the coding block; determining static region detection points based on the difference between the luminance components and its absolute value and the difference between the chrominance components and its absolute value; sliding a window within the region including the static region detection points, and obtaining the average value of the data points within the window each time it slides as the pre-sliding foreground mean; determining whether the coding block is a static region according to each of the pre-sliding foreground means, the foreground mean threshold, and the foreground pixel point threshold.
[0005] In some implementation manners of the first aspect, the method for detecting static regions further includes: determining whether color cast occurs in the coding block according to the difference between the luminance component and the difference between the chrominance components; if color cast occurs in the coding block, then amplify the absolute value of the difference of the component corresponding to the color cast.
[0006] In some implementation manners of the first aspect, determining whether color cast occurs in the coding block according to the difference between the luminance component and the difference between the chrominance components includes: for the original pixels of the coding block, respectively counting the number of positive and negative values of the difference between the luminance component and the difference between the chrominance components; if the number of positive or negative values of the difference of a certain component is greater than the quantity threshold, it is determined that color cast occurs in the coding block.
[0007] In some implementations of the first aspect, determining whether the coding block is a static region according to the foreground point means of the sliding windows, the foreground mean threshold, and the foreground pixel point threshold includes: obtaining the quotient and difference between the foreground point mean of each sliding window and the foreground mean threshold; summing the quotients corresponding to all eligible sliding windows, and if the result of the summation is less than or equal to the foreground pixel point threshold, determining that the coding block is the static region, otherwise determining that the coding block is not the static region; wherein, if the foreground detection point value of a sliding window is greater than the difference of the sliding window and the difference of the sliding window is greater than 0, it is determined that the sliding window is eligible, otherwise it is determined that the sliding window is ineligible.
[0008] In some implementations of the first aspect, determining the static region detection points based on the difference of the luminance component and its absolute value and the difference of the chrominance component and its absolute value includes: performing downsampling processing on the difference of the luminance component so that the difference of the luminance component has the same data volume as the difference of the chrominance component; selecting the largest value from the absolute value of the difference of the luminance component and the absolute value of the difference of the chrominance component as the foreground point detection data source; filling the data points including a plurality of the foreground point detection data sources to obtain the static region detection points.
[0009] In some implementations of the first aspect, the static region detection method further includes: obtaining a corresponding reconstruction error threshold according to the complexity of the original pixels of the coding block; if the motion vector of the coding block is not 0, reducing a preset threshold to obtain the foreground pixel point threshold, reducing the noise threshold corresponding to the previous frame, and obtaining the foreground mean threshold according to the reduced noise threshold and the reconstruction error threshold.
[0010] In some implementations of the first aspect, the static region detection method further includes: obtaining the mean value of the luminance component of the previous frame, where the mean value of the luminance component of the previous frame is the average value of the luminance components of all pixels in the previous frame; determining whether each coding block in the previous frame belongs to the brightest region according to the mean value of the luminance component of each coding block in the previous frame and the mean value of the luminance component of the previous frame; obtaining the noise threshold corresponding to the previous frame according to the sum of the prediction errors and the sum of the quantities of the coding blocks belonging to the brightest region in the previous frame.
[0011] Second aspect, an embodiment of the present application provides a video encoding method. The video encoding method includes: obtaining the difference between the original pixels and the predicted pixels of the encoding block in the luminance component and the difference in the chrominance component; determining static region detection points based on the difference in the luminance component and its absolute value and the difference in the chrominance component and its absolute value; using a window to slide within the region including the static region detection points, and obtaining the average value of the data points within the window at each slide as the pre-sliding foreground mean value; determining whether the encoding block is a static region according to each of the pre-sliding foreground mean values, the foreground mean value threshold, and the foreground pixel point threshold; if the encoding block is a static region, when encoding the encoding block, the intra prediction mode is not used, and cleaning residual processing is performed on the encoding block.
[0012] Third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the static region detection method or the video encoding method provided by the embodiment of the present application.
[0013] Fourth aspect, an embodiment of the present application provides an electronic device. The electronic device includes: a memory storing a computer program; a processor communicatively connected to the memory, and when calling the computer program, it executes the static region detection method or the video encoding method provided by the embodiment of the present application.
[0014] As described above, an embodiment of the present application provides a static region detection method, and the static region detection method has a high accuracy.
[0015] In addition, in the embodiment of the present application, static region detection is performed based on the encoding block, and factors such as video picture light change, noise introduction, and texture distortion are considered simultaneously during the detection, which is beneficial to further improving the accuracy and reliability of static region detection.
[0016] An embodiment of the present application also provides a video encoding method, which combines static region detection and video encoding strategies, which is beneficial to saving the encoding bit rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1A It shows a flowchart of the static region detection method provided by the embodiment of the present application.
[0018] Figure 1B It shows a sampling ratio schematic diagram of the Y component, U component, and V component in the embodiment of the present application.
[0019] Figure 2 It shows a flowchart of obtaining DiffY, DiffU, DiffV, and AbsDiffY, AbsDiffU, and AbsDiffV in the embodiment of the present application.
[0020] Figure 3 It shows a flowchart for determining whether a coding block has a color cast phenomenon in an embodiment of the present application.
[0021] Figure 4 It shows a flowchart for obtaining static region detection points in an embodiment of the present application.
[0022] Figure 5 It shows a schematic diagram of a window sliding in an area including static region detection points in an embodiment of the present application.
[0023] Figure 6 It shows a flowchart for determining whether a coding block is a static region in an embodiment of the present application.
[0024] Figure 7A It shows a flowchart for obtaining a reconstruction error threshold and a foreground mean threshold in an embodiment of the present application.
[0025] Figure 7B It shows a schematic diagram for calculating the complexity of a coding block in an embodiment of the present application.
[0026] Figure 7C It shows a schematic diagram for calculating the reconstruction error of a coding block in an embodiment of the present application.
[0027] Figure 8 It shows a flowchart for obtaining a noise threshold corresponding to the previous frame in an embodiment of the present application.
[0028] Figure 9 It shows a flowchart for a video coding method provided by an embodiment of the present application.
[0029] Figure 10 It shows a schematic diagram of the structure of an electronic device provided by an embodiment of the present application.
[0030] Description of component numbers
[0031] 1000 Electronic device
[0032] 1001 Processor
[0033] 1002 Memory
[0034] 10021 Operating system
[0035] 10022 Application program
[0036] 1003 Network interface
[0037] 1004 Bus system
[0038] 1005 User interface
[0039] Steps S11 to S14
[0040] Steps S111 to S116
[0041] Steps S31 to S32
[0042] Steps S41 to S43
[0043] Steps S61 to S62
[0044] Steps S71 to S72
[0045] Steps S81 to S83
[0046] Steps S91 to S95 DETAILED DESCRIPTION
[0047] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0048] It should be noted that the illustrations provided in the following embodiments are only used to illustrate the basic concept of the present application in a schematic manner, and therefore the illustrations only show components related to the present application rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0049] In the embodiments of the present application, the words "exemplary" or "for example" represent examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0050] Some technical solutions use the accumulation of foreground masking to detect static areas. Specifically, when the cumulative value of foreground masking in a region exceeds a certain threshold, the region is determined to be a static region. However, in the video encoding process, due to changes in scene illumination and noise introduced by image acquisition sensors, the accurate detection of static areas is often affected, and the video quality is negatively affected.
[0051] At least to address the above-mentioned problem, an embodiment of the present application provides a static area detection method. Figure 1A The flowchart of the static area detection method provided by the embodiment of the present application is shown.Figure 1A As shown in Figure 1A , the static region detection method includes the following steps S11 to S14.
[0052] S11, obtain the difference DiffY between the original pixel and the predicted pixel of the coding block in the luminance component (i.e., Y component) and the differences DiffU and DiffV in the chrominance components (i.e., U / V components). The predicted pixel is obtained by predicting using the encoded image pixels, i.e., the reconstructed pixels, through video prediction coding technology. In video coding, operations such as transformation, quantization, and entropy coding can be performed on the difference between the original pixel and the predicted pixel, that is, the residual, so as to achieve the compression of the video image.
[0053] Among them, in the YUV color space, the Y component represents the luminance information of the coding block, that is, the brightness and darkness of the image. The U / V components represent the color information of the image, that is, the color and hue of the image. There are various different formats in the YUV color space, such as 4:4:4, 4:2:2, 4:2:0, etc., and these formats represent the sampling ratios of the Y, U, and V components. For example, the 4:2:0 format means that for every four Y components, there is one U component and one V component, as Figure 1B shown.
[0054] S12, determine the static region detection points based on the difference of the Y component and its absolute value and the differences of the U / V components and their absolute values.
[0055] S13, use a window to slide within the region including the static region detection points, and obtain the average value of the data points within the window each time it slides as the foreground mean before sliding.
[0056] S14, judge whether the coding block is a static region according to each foreground mean before sliding, the foreground mean threshold, and the foreground pixel point threshold.
[0057] Please refer to Figure 2 , in some implementation manners, obtaining the difference between the original pixel and the predicted pixel of the coding block in the Y component and the U / V components includes:
[0058] S111, calculate the mean of the Y component and the means of the U / V components of the original pixels in the coding block.
[0059] S112, subtract the corresponding means from the Y component and the U / V components of the original pixels in the coding block to obtain the first luminance component and the first chrominance component. By step S112, the influence brought by the change of bright and dark light can be removed.
[0060] S113, calculate the mean of the Y component and the means of the U / V components of the predicted pixels in the coding block.
[0061] S114. Subtract the mean values of the Y component and the U / V components of the predicted pixels in the coding block to obtain a second luminance component and a second chrominance component.
[0062] S115. Obtain the difference between the first luminance component and the second luminance component as DiffY, and obtain the differences between the first chrominance component and the second chrominance component as DiffU and DiffV.
[0063] S116. Obtain the absolute values of DiffY, DiffU, and DiffV, denoted as AbsDiffY, AbsDiffU, and AbsDiffV respectively.
[0064] In some implementation manners, the static region detection method provided by the embodiments of the present application may further include: judging whether a color cast phenomenon occurs in the coding block according to the difference of the Y component and the difference of the U / V components. If a color cast phenomenon occurs in the coding block, amplify the absolute value of the difference of the component corresponding to the color cast phenomenon, and do not amplify the absolute values of the differences of the remaining components.
[0065] Exemplarily, if the component corresponding to the color cast phenomenon is the Y component, in step S12, the static region detection points may be determined based on the difference of the Y component, the absolute value of the amplified difference of the Y component, the difference of the U / V components, and their absolute values. When the component corresponding to the color cast phenomenon is other components, the method for determining the static region detection points is similar to the above, and will not be elaborated here.
[0066] Please refer to Figure 3 , in some implementation manners, judging whether a color cast phenomenon occurs in the coding block according to the difference of the Y component and the difference of the U / V components includes the following steps S31 and S32.
[0067] S31. For the original pixels of the coding block, respectively count the number of positive and negative values of the difference of the Y component and the difference of the U / V components.
[0068] S32. If the number of positive or negative values of the difference of a certain component is greater than the quantity threshold, it is judged that a color cast phenomenon occurs in the coding block. Among them, the quantity threshold can be set according to actual requirements or experience. For example, the quantity threshold can be set to α*N*N, where α is a value greater than 0 and less than 1, and N*N is the total number of original pixels included in the coding block.
[0069] Taking the difference DiffY of the Y component as an example, count the number of original pixels YPosi with positive DiffY in the coded block, and count the number of original pixels YNega with negative DiffY in the coded block. If the value of YPosi or YNega is greater than the quantity threshold, it is determined that the coded block has a color cast phenomenon. The process of determining whether the coded block has a color cast phenomenon based on the differences DiffU and DiffV of the U / V components is similar to that of DiffY, which will not be elaborated here. When it is determined that the coded block does not have a color cast phenomenon according to none of DiffY, DiffU, and DiffV, it is determined that the coded block does not have a color cast phenomenon.
[0070] Exemplarily, if the value of YPosi is greater than the quantity threshold, the absolute value AbsDiffY of the difference of the component corresponding to the color cast phenomenon (in this case, the Y component) is amplified, including: amplifying AbsDiffY to (10 + YLevel) / 10 times, where YLevel is YPosi - α * N * N.
[0071] Exemplarily, if the value of UNega is greater than the quantity threshold, the absolute value AbsDiffU of the difference of the component corresponding to the color cast phenomenon (in this case, the U component) is amplified, including: amplifying AbsDiffU to (10 + ULevel) / 10 times, where ULevel is UNega - α * N * N.
[0072] The implementation method of amplifying the absolute value of the difference of the component corresponding to the color cast phenomenon in other cases is similar to the above, which will not be elaborated here.
[0073] Please refer to Figure 4 , in some implementation manners, determining the static region detection points based on the difference of the Y component and its absolute value and the differences of the U / V components and their absolute values includes the following steps S41 to S43.
[0074] S41, perform downsampling processing on the difference of the Y component so that the difference of the Y component has the same data volume as the differences of the U / V components.
[0075] Exemplarily, the implementation method of performing downsampling processing on the difference of the Y component may include: starting from the upper left corner, rounding the average value of every adjacent 4 differences of the Y component as the new DiffY, that is, rounding the result of (DiffY[i][j] + DiffY[i + 1][j] + DiffY[i][j + 1] + DiffY[i + 1][j + 1]) / 4 as the new DiffY, so that DiffY, DiffU, and DiffV have the same data volume.
[0076] S42. Select the maximum value among the absolute value of the difference in the Y component, AbsDiffY, and the absolute values of the differences in the U / V components, AbsDiffU and AbsDiffV, as the data source for foreground point detection, denoted as Diff[i][j].
[0077] S43. Fill the data points including multiple data sources for foreground point detection to obtain static region detection points.
[0078] Exemplarily, for N / 2*N / 2 data points including the data source for foreground point detection, the following methods shown in S431 and S432 can be used for filling:
[0079] S431. Copy the data on the leftmost and rightmost sides of the data points and fill 3 columns outwards respectively to form (3 + N / 2 + 3)*N / 2 data points.
[0080] S432. For the data points obtained in S431, copy the data on the uppermost and lowermost sides and fill 3 rows outwards respectively to finally form (3 + N / 2 + 3)*(3 + N / 2 + 3) static region detection points, and each static region detection point is denoted as P[i][j].
[0081] In some implementation manners, a sliding window of size M*M can be used to obtain the mean value MeanP[i] of N / 2*N / 2 points among the above (3 + N / 2 + 3)*(3 + N / 2 + 3) static region detection points, and the calculation method of this mean value is as Figure 5 shown. In Figure 5 , the shaded part is the original data, and the white part is the filled data. The sliding window moves one grid to the right one by one, and then one grid downwards one by one. Each time it moves one grid, the average value of all data points within the sliding window can be obtained, denoted as MeanP[i]. The sliding window can move 8 grids to the right in sequence and 8 grids downwards in sequence, and a total of 64 average values MeanP[i] can be obtained.
[0082] Please refer to Figure 6 , in some implementation manners, determining whether a coding block is a static region according to the foreground point mean value of each sliding window, the foreground mean value threshold, and the foreground pixel point threshold includes the following steps S61 and S62.
[0083] S61. Obtain the quotient value PixDelta[i] and the difference value ith[i] between the foreground point mean value MeanP[i] of each sliding window and the foreground mean value threshold Meanth. Wherein, PixDelta[i] = MeanP[i] / Meanth, and ith[i] = MeanP[i] - Meanth.
[0084] S62. Sum the quotient values PixDelta[i] corresponding to all eligible sliding windows. If the result of the summation is less than or equal to the foreground pixel threshold PixNumTh, determine that the coding block is a static region; otherwise, determine that the coding block is not a static region.
[0085] Among them, if the foreground detection point value of a certain sliding window is greater than the difference value ith[i] of this sliding window, and the difference value ith[i] of this sliding window is greater than 0, then determine that this sliding window meets the conditions; otherwise, determine that this sliding window does not meet the conditions. (i, j) corresponds to the point at the upper left corner of the sliding window. For example, for Figure 5 the sliding window at the upper left corner shown, i = 0 and j = 0, and its foreground detection point value is the value corresponding to the static region detection point P[3 + i, 3 + j] (this value can be obtained through step S42). When this sliding window slides one step to the right, i = 1 and j = 0, and its foreground detection point value is the value corresponding to the static region detection point P[3 + i, 3 + j], and so on.
[0086] Please refer to Figure 7A , in some implementation manners, the static region detection method provided by the embodiments of the present application may further include the following steps S71 and S72.
[0087] S71. Obtain the corresponding reconstruction error threshold according to the complexity of the original pixels of the coding block.
[0088] Exemplarily, according to the complexity of the original pixels of the coding block, the corresponding reconstruction error can be obtained by querying the complexity-error mapping table, and this reconstruction error is used as the reconstruction error threshold.
[0089] Exemplarily, after each frame ends, the complexity Comp[i] of the original pixels in the image can be calculated respectively according to the size of the coding block, and the reconstruction error ReconSad[i] between the corresponding reconstructed pixels and the original pixels can be calculated, so as to obtain the complexity Comp[i] and the reconstruction error ReconSad[i] of each coding block in this frame. Traverse each coding block. For the complexity Comp[j] of a certain coding block j, if there is no other coding block with the same complexity as Comp[j], then determine that Comp[j] corresponds to the reconstruction error ReconSad[j] of this coding block j, and add this corresponding relationship to the complexity-error mapping table; if there are other coding blocks k1 to km with the same complexity as Comp[j], then determine that Comp[j] corresponds to the average value of ReconSad[j] + ReconSad[k1] +... + ReconSad[km], and add this corresponding relationship to the complexity-error mapping table. Where m is an integer greater than or equal to 1, representing the number of coding blocks with the same complexity as Comp[j].
[0090] For example, if a certain frame contains 3 coded blocks, by traversing, the reconstruction error ReconSad[1] of the first coded block is 10, and its complexity Comp[1] is 15; the reconstruction error ReconSad[2] of the second coded block is 12, and its complexity Comp[2] is 10; the reconstruction error ReconSad[3] of the third coded block is 14, and its complexity Comp[3] is 10. Based on this, it can be determined that the reconstruction error corresponding to a complexity of 15 is 10, and the reconstruction error corresponding to a complexity of 10 is (12 + 14) / 2 = 13.
[0091] Figure 7B It shows the calculation method of complexity in the embodiment of the present application. As Figure 7B shown, the gradient of a pixel is the average of the absolute value of the difference between the pixel and the right pixel and the absolute value of the difference between the pixel and the bottom pixel. For example, the horizontal gradient of pixel A is the absolute value of the difference between pixel A and pixel B, the vertical gradient of pixel A is the absolute value of the difference between pixel A and pixel C, and the gradient of pixel A is the average of the horizontal gradient and the vertical gradient of pixel A. For a coded block, its complexity is the average of the gradients of the remaining pixels except the rightmost 1 column and the bottommost 1 row in the coded block. For example, Figure 7B the complexity of the coded block shown is the average of the gradients of the 3*3 pixels in the upper left corner.
[0092] Figure 7C It shows the calculation method of the reconstruction error in the embodiment of the present application. The reconstruction error of a coded block is the average of the absolute values of the differences between each pixel point in the reconstructed coded block and the corresponding original pixel point. As Figure 7C shown, the absolute values of the differences between pixel point A and A', pixel point B and B', …, pixel point J and J' are respectively obtained, and the average of these absolute values is used as the reconstruction error of the coded block.
[0093] S72. If the motion vector of the coded block is not 0, then reduce the preset threshold to obtain the foreground pixel point threshold, reduce the noise threshold corresponding to the previous frame, and obtain the foreground mean threshold according to the reduced noise threshold and the reconstruction error threshold.
[0094] Exemplarily, if the motion vector of the coded block is not 0, 1 / 2 of the preset threshold can be used as the foreground pixel point threshold, the noise threshold corresponding to the previous frame is reduced to 1 / 2, and a larger value is selected from the reconstruction error threshold obtained in step S71 and the reduced noise threshold as the foreground mean threshold.
[0095] If the motion vector of a coding block is 0, use a preset threshold as the foreground pixel threshold, and obtain the foreground mean threshold according to the noise threshold and reconstruction error threshold corresponding to the previous frame. For example, a larger value can be selected from the reconstruction error threshold obtained in step S71 and the noise threshold corresponding to the previous frame as the foreground mean threshold.
[0096] Please refer to Figure 8 , in some implementation manners, the static region detection method provided by the embodiments of the present application may further include the following steps S81 to S83.
[0097] S81, obtain the mean value Yth of the Y component of the previous frame, and the mean value Yth of the Y component of the previous frame is the average value of the Y components of all pixels in the previous frame.
[0098] S82, determine whether each coding block in the previous frame belongs to the brightest region according to the mean value of the Y component of each coding block in the previous frame and the mean value Yth of the Y component of the previous frame.
[0099] Exemplarily, if the mean value of the Y component of a certain coding block is greater than or equal to MIN(Yth * 0.8, 50), it can be considered that the coding block belongs to the brightest region.
[0100] S83, obtain the noise threshold corresponding to the previous frame according to the sum of the prediction errors and the sum of the numbers of the coding blocks belonging to the brightest region in the previous frame. For example, if the sum of the numbers of the coding blocks belonging to the brightest region is Cnt, and the sum of the prediction errors of these coding blocks is SADSum, then the noise threshold noiseSadTh corresponding to the previous frame = SADSum / Cnt. Among them, the prediction error of a coding block is the average value of the absolute values of the differences between the original pixels and the predicted pixels in the coding block.
[0101] The embodiments of the present application also provide a video coding method. Figure 9 Shown is a flowchart of the video coding method provided by the embodiments of the present application. As Figure 9 shown, the video coding method includes the following steps S91 to S95.
[0102] S91, obtain the differences between the Y components and the differences between the U / V components of the original pixels and the predicted pixels of the coding block.
[0103] S92, determine static region detection points based on the differences between the Y components and their absolute values and the differences between the U / V components and their absolute values.
[0104] S93, use a window to slide within the region including the static region detection points, and obtain the average value of the data points within the window each time it slides as the foreground mean before sliding.
[0105] S94. Determine whether the coding block is a static region according to the average value of the foreground points in front of each sliding window, the foreground average value threshold, and the foreground pixel point threshold.
[0106] S95. If the coding block is a static region, when encoding the coding block, do not use the intra prediction mode, and perform residual cleaning processing on the coding block.
[0107] Specifically, during video encoding, the intra prediction mode generates prediction pixels by using the spatial domain correlation of the video image and the already encoded pixels in the current image, thereby effectively removing the spatial redundancy of the video. The inter prediction mode utilizes the temporal domain correlation of the video image and generates prediction pixels by using the reconstructed pixels of the already encoded images before the current image, thereby effectively removing the temporal redundancy of the video. During the video image encoding process, it is necessary to perform related processing such as transformation and quantization on the difference between the input pixels and the prediction pixels, that is, the pixel residual data. Generally, the larger the residual, the greater the coding bit rate.
[0108] In some implementation manners, if the current coding block is a static region, it can be considered that the current coding block has obvious temporal domain correlation. During the rate-distortion optimization processing of the coding block, the current block can be made not to be encoded as the intra prediction mode, thereby eliminating the obvious temporal redundancy.
[0109] In some implementation manners, if the current coding block is a static region, the distortion between the current coding block and the reconstructed block is small, and it can be considered that the residual of the current coding block is also small. During encoding, the residual cleaning processing can be performed on the current coding block, that is, the residual of the current coding block is set to zero for cleaning processing, and the residual data is no longer encoded, thereby saving the bit rate.
[0110] Among them, steps S91 to S95 are the same as the static region detection method provided in the embodiments of the present application, and will not be elaborated here.
[0111] The protection scope of the static region detection method and the video encoding method provided in the embodiments of the present application is not limited to the execution order of the steps listed in this embodiment. Any solution achieved by adding or subtracting steps of the prior art and replacing steps according to the principle of the present application is included in the protection scope of the present application.
[0112] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the static area detection method or the video encoding method provided by the embodiments of the present application. Those of ordinary skill in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing a processor through a program. The program can be stored in a computer-readable storage medium, and the storage medium is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid state drive, magnetic tape, floppy disk, optical disc, and any combination thereof. The above storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center integrating one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)).
[0113] The embodiments of the present application further provide an electronic device. Figure 10 is a schematic block diagram of the electronic device provided by the embodiments of the present application. As Figure 10 shown, the electronic device 1000 includes: at least one processor 1001, a memory 1002, at least one network interface 1003, and a user interface 1005. Each component in the electronic device 1000 is coupled together through a bus system 1004. It can be understood that the bus system 1004 is used to implement the connection and communication between these components. In addition to the data bus, the bus system 1004 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 10 all kinds of buses are labeled as the bus system.
[0114] Among them, the user interface 1005 may include a display, a keyboard, a mouse, a trackball, a click gun, a key, a button, a touchpad, or a touch screen, etc.
[0115] It can be understood that the memory 1002 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM, Static Random Access Memory), synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory). The memory described in the embodiments of the present application is intended to include but not limited to these and any other suitable categories of memories.
[0116] The memory 1002 in the embodiments of the present application is used to store various categories of data to support the operation of the electronic device 1000. Examples of such data include: any executable programs for operating on the electronic device 1000, such as the operating system 10021 and application programs 10022; the operating system 10021 contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application programs 10022 can include various application programs, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. Implementing the static area detection method or video encoding method provided in the embodiments of the present application can be included in the application programs 10022.
[0117] The method disclosed in the embodiments of the present application can be applied to or implemented by the processor 1001. The processor 1001 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in hardware or instructions in software form in the processor 1001. The above-mentioned processor 1001 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 1001 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor 1001 may be a microprocessor or any conventional processor, etc. Combining the steps of the accessory optimization method provided in the embodiments of the present application can be directly reflected as being completed by the hardware decoding processor, or completed by a combination of hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the memory. The processor reads the information in the memory and combines its hardware to complete the steps of the foregoing method.
[0118] In an exemplary embodiment, the electronic device 1000 may be an application-specific integrated circuit (ASIC, Application Specific Integrated Circuit), a DSP, a programmable logic device (PLD, ProgrammableLogic Device), or a complex programmable logic device (CPLD, Complex Programmable Logic Device) for executing the foregoing method.
[0119] In summary, the embodiments of the present application provide a static region detection method. The static region detection method performs static region detection based on coding blocks, and simultaneously considers factors such as video frame light changes, noise introduction, and texture distortion during the detection, which is beneficial to further improving the accuracy and reliability of static region detection. In addition, the embodiments of the present application also provide a video coding method. This method first performs pixel-domain sampling on the image to be processed at the block level, then calculates the frame-level and block-level detection thresholds, adaptively configures the static region detection model parameters, then performs static region detection, and finally encodes the video image based on the detection results. In this way, while the subjective visual quality of the video-coded image does not decrease or slightly decreases, the bit rate is effectively saved, and it is easy to expand and adapt to different scenarios. Therefore, the present application effectively overcomes various shortcomings in the prior art and has high industrial utilization value.
[0120] The above embodiments are only illustrative of the principles and effects of the present application and are not intended to limit the present application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present application should still be covered by the claims of the present application.
Claims
1. A static area detection method, characterized in that: The static area detection method comprises: Obtain the difference between the brightness component and the chrominance component of the original pixel and the predicted pixel of the coding block; Determine a static area detection point based on the difference value and the absolute value of the luminance component and the difference value and the absolute value of the chrominance component; Slide the window in the area including the static area detection points, and obtain the average value of the data points in the window during each sliding as the sliding window foreground point average value; Whether the coding block is a static area is determined according to the foreground pixel point mean value of each sliding window, the foreground mean threshold value, and the foreground pixel point threshold value.
2. The static area detection method according to claim 1, characterized in that: The static area detection method further includes: Determining whether a color cast occurs in the coding block according to a difference value of the luminance component and a difference value of the chrominance component; If a color cast occurs in the coding block, the absolute value of the difference of the component corresponding to the color cast is amplified.
3. The static area detection method according to claim 2, characterized in that: Judging whether a color cast occurs in the coding block according to a difference value of the brightness component and a difference value of the chrominance component, comprising: For the original pixels of the coding block, respectively counting the number of positive values and negative values of the difference of the luminance component and the difference of the chrominance component; If the difference of a certain component is a positive value or the number of negative values is greater than the quantity threshold, it is determined that the coding block has a color cast phenomenon.
4. The static area detection method according to claim 1, characterized in that: Judging whether the coding block is a static area according to the foreground point mean value of each sliding window, the foreground mean threshold value, and the foreground pixel point threshold value includes: Obtaining the quotient and difference between the foreground point mean value of each sliding window and the foreground mean value threshold; The quotient values corresponding to all sliding windows that meet the conditions are summed, and if the sum result is less than or equal to the foreground pixel threshold, the coding block is judged to be the static area, otherwise, the coding block is judged not to be the static area; If the foreground detection point value of a certain sliding window is greater than the difference value of the sliding window, and the difference value of the sliding window is greater than 0, then it is determined that the sliding window meets the condition; otherwise, it is determined that the sliding window does not meet the condition.
5. The static area detection method according to claim 1, characterized in that: Determining a static area detection point based on the difference value and the absolute value of the luminance component and the difference value and the absolute value of the chrominance component, including: Downsampling the difference of the luminance component so that the difference of the luminance component and the difference of the chrominance component have the same data amount; Selecting the largest value among the absolute value of the difference of the brightness component and the absolute value of the difference of the chrominance component as the foreground point detection data source; The data points including the plurality of foreground point detection data sources are filled to obtain the static area detection points.
6. The static area detection method according to claim 1, characterized in that: The static area detection method further includes: Obtaining a corresponding reconstruction error threshold according to the complexity of the original pixels of the coding block; If the motion vector of the coding block is not 0, the preset threshold is reduced to obtain the foreground pixel threshold, the noise threshold corresponding to the previous frame is reduced, and the foreground mean threshold is obtained according to the reduced noise threshold and the reconstruction error threshold.
7. The static area detection method according to claim 6, characterized in that: The static area detection method further includes: Obtaining a mean value of the brightness component of the previous frame, wherein the mean value of the brightness component of the previous frame is an average value of the brightness components of all pixels in the previous frame; Determining whether each coding block in the previous frame belongs to the brightest area according to the average value of the brightness components of each coding block in the previous frame and the average value of the brightness components of the previous frame; The noise threshold corresponding to the previous frame is obtained according to the sum of the prediction errors and the sum of the numbers of the coding blocks belonging to the brightest area in the previous frame.
8. A video encoding method, characterized in that: The video encoding method comprises: Obtain the difference between the brightness component and the chrominance component of the original pixel and the predicted pixel of the coding block; Determine a static area detection point based on the difference value and the absolute value of the luminance component and the difference value and the absolute value of the chrominance component; Slide the window in the area including the static area detection points, and obtain the average value of the data points in the window during each sliding as the sliding window foreground point average value; Determine whether the coding block is a static area according to the foreground pixel mean value of each sliding window, the foreground mean threshold value and the foreground pixel threshold value; If the coding block is a static area, the intra-frame prediction mode is not used when encoding the coding block, and residual cleaning processing is performed on the coding block.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the static area detection method according to any one of claims 1 to 7 or the video encoding method according to claim 8 is implemented.
10. An electronic device, characterized in that: The electronic device comprises: A memory storing a computer program; A processor is communicatively connected to the memory, and executes the static area detection method according to any one of claims 1 to 7 or the video encoding method according to claim 8 when calling the computer program.