Data processing method and device, equipment and storage medium
By downsampling each image unit in the image frame and building the Vinahof equation, the problem of low coding efficiency of image frames is solved and a more efficient coding effect is achieved.
Patent Information
- Application Number
- CN202410114525.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art leads to a low overall coding efficiency of the image frame when constructing the Wienerhof equation of the image frame.
By acquiring the types of each image unit in the image frame, downsampling the image units that meet the downsampling conditions are performed, the Vinahof equation associated with the image unit is constructed, the filter coefficient group of the image frame is determined, and the code stream data of the image frame is generated.
The number of pixel points in the image unit is reduced, and the number of construction of the Vinahof equation is reduced, thereby improving the overall encoding efficiency of the image frame.
Smart Images

Figure CN120378635A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly relates to a data processing method, a data processing device, a computer device, and a computer-readable storage medium. Background Art
[0002] With the progress of scientific research, a vast amount of multimedia resources have emerged in the network. During the encoding process of multimedia resources, it is necessary to solve the filter coefficient group corresponding to the multimedia resource (such as an image frame) through one or more Wiener-Hopf equations. It has been found that the Wiener-Hopf equation for solving the filter coefficient group corresponding to an image frame is obtained based on the Wiener-Hopf equations associated with each pixel point in the image frame. Individually constructing the Wiener-Hopf equations associated with each pixel point in the image frame will result in a low overall encoding efficiency of the image frame. Summary of the Invention
[0003] Embodiments of this application provide a data processing method, device, equipment, and computer-readable storage medium, which can improve the overall encoding efficiency of an image frame.
[0004] On the one hand, embodiments of this application provide a data processing method, including:
[0005] Obtaining the types of each image unit in the image frame to be encoded, where each image unit includes at least two pixel points in the image frame;
[0006] Performing downsampling processing on the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of each type of image unit;
[0007] Based on the downsampling results of each type of image unit, constructing the Wiener-Hopf equation associated with this type of image unit;
[0008] Determining the filter coefficient group corresponding to the image frame through the Wiener-Hopf equations associated with each type of image unit, where the filter coefficient group is used to generate the bitstream data of the image frame.
[0009] On the one hand, embodiments of this application provide a data processing device, and this data processing device includes:
[0010] An obtaining unit, configured to obtain the types of each image unit in the image frame to be encoded, where each image unit includes at least two pixel points in the image frame;
[0011] A processing unit, configured to perform downsampling processing on the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of each type of image unit;
[0012] and constructing a Wiener-Hopf equation associated with the image units of each type based on the downsampling results of the image units of each type;
[0013] and determining a set of filtering coefficients corresponding to the image frame through the Wiener-Hopf equations associated with the image units of each type, where the set of filtering coefficients is used to generate the bitstream data of the image frame.
[0014] In one embodiment, the image frame includes M types of image units, where M is a positive integer; the target type is any one of the M types; the processing unit is configured to perform downsampling processing on the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of the image units of each type, specifically:
[0015] detecting whether there is a first image unit in the image units of the target type, where the image region associated with the first image unit does not coincide with the edge region of the image frame, or the first image unit is an image unit encoded by using a non-screen content encoding technique;
[0016] if there is a first image unit in the image units of the target type, performing downsampling processing on the first image unit to obtain the downsampling result of the image units of the target type;
[0017] wherein the downsampling processing includes at least one of horizontal downsampling processing and vertical downsampling processing.
[0018] In one embodiment, the image frame includes M types of image units, where M is a positive integer; the target type is any one of the M types; the processing unit is configured to perform downsampling processing on the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of the image units of each type, specifically:
[0019] detecting whether there is a second image unit in the image units of the target type, where there is an overlapping part between the image region associated with the second image unit and the edge region of the image frame in the target direction;
[0020] if there is a second image unit in the image units of the target type, performing downsampling processing on the second image unit along the target direction to obtain the downsampling result of the image units of the target type.
[0021] In one embodiment, the image frame includes M types of image units, where M is a positive integer; the target type is any one of the M types; the processing unit is configured to perform downsampling processing on the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of the image units of each type, specifically:
[0022] Detect whether there are target pixel points in the image units of the target type. The target pixel points are the pixel points in the image units of the target type with a difference value less than the difference threshold. The difference value of each pixel point is used to indicate the difference between the first color information and the second color information of the pixel point. The first color information of any pixel point is the color information of the pixel point in the image frame, and the second color information of any pixel point is the color information obtained after performing sample adaptive compensation on the compression result of the pixel point;
[0023] If there are target pixel points in the image units of the target type, then remove the target pixel points from the image units of the target type to obtain the downsampling result of the image units of the target type.
[0024] In one implementation, the image frame includes M types of image units, where M is a positive integer; the target type is any one of the M types, and the downsampling result of the image units of the target type includes N pixel points, where N is a positive integer; the processing unit is configured to construct the Wiener-Hopf equation associated with the image units of this type based on the downsampling results of the image units of each type, specifically:
[0025] Construct the Wiener-Hopf equation associated with N pixel points based on the color information of K pixel points associated with each pixel point and the color information of this pixel point. The Wiener-Hopf equation associated with the i-th pixel point is constructed based on the color information of K pixel points associated with the i-th pixel point and the color information of the i-th pixel point, where i is a positive integer less than or equal to N;
[0026] Perform an accumulation process on the Wiener-Hopf equations associated with N pixel points to obtain the Wiener-Hopf equation associated with the image units of the target type.
[0027] In one implementation, the processing unit is configured to determine the filter coefficient group corresponding to the image frame through the Wiener-Hopf equations associated with the image units of each type, specifically:
[0028] Group the image units of each type according to a preset grouping rule to obtain P groups of image units. There is a target group among the P groups of image units. The target group includes at least two types of image units, where P is a positive integer;
[0029] Merge the Wiener-Hopf equations associated with different types of image units in the target group, and use the merged Wiener-Hopf equation as the Wiener-Hopf equation associated with the target group;
[0030] Determine the filter coefficient group corresponding to the image frame based on the Wiener-Hopf equations associated with the P groups of image units.
[0031] In one embodiment, the processing unit is configured to determine a set of filtering coefficients corresponding to an image frame based on the Wiener-Hopf equations associated with P sets of image units, specifically:
[0032] Generate P combined results based on the P sets of image units, where the h-th combined result includes P - h + 1 sets of image units, and h is a positive integer less than or equal to P;
[0033] Calculate the rate-distortion optimization corresponding to the P combined results through the Wiener-Hopf equations associated with the P sets of image units;
[0034] Determine the set of filtering coefficients associated with the target combined result as the set of filtering coefficients corresponding to the image frame, where the target combined result is the combined result with the minimum rate-distortion optimization among the P combined results.
[0035] In one embodiment, an image frame includes M types of image units, where M is a positive integer; the processing unit is configured to determine a set of filtering coefficients corresponding to the image frame through the Wiener-Hopf equations associated with each type of image unit, specifically:
[0036] Generate M combined results based on the M types of image units, where the g-th combined result includes M - g + 1 types of image units, and g is a positive integer less than or equal to M;
[0037] Calculate the rate-distortion optimization corresponding to the M combined results through the Wiener-Hopf equations associated with the M types of image units;
[0038] Determine the set of filtering coefficients associated with the target combined result as the set of filtering coefficients corresponding to the image frame, where the target combined result is the combined result with the minimum rate-distortion optimization among the M combined results.
[0039] In one embodiment, the process by which the processing unit generates M combined results based on the M types of image units includes:
[0040] Merge the x-th type of image unit and the y-th type of image unit to obtain the merged x-th type of image unit, and the combined rate-distortion optimization associated with the x-th type of image unit and the y-th type of image unit is the minimum among the combined rate-distortion optimizations associated with any two types of image units among the M types of image units;
[0041] Merge the Wiener-Hopf equation associated with the x-th type of image unit and the Wiener-Hopf equation associated with the y-th type of image unit to obtain the merged Wiener-Hopf equation associated with the x-th type of image unit.
[0042] In one embodiment, the processing unit is configured to calculate the rate-distortion optimization corresponding to the M combined results through the Wiener-Hopf equations associated with the M types of image units, specifically:
[0043] Solve the Wiener - Hoff equation associated with M - g + 1 types of image units to obtain the filtering coefficients associated with each type of image unit;
[0044] Based on the filtering coefficients associated with each type of image unit, calculate the rate - distortion optimization of this type of image unit;
[0045] Through the rate - distortion optimization of M - g + 1 types of image units, calculate the rate - distortion optimization corresponding to the g - th combination result.
[0046] In one implementation, the processing unit is used to obtain the types of each image unit in the image frame to be encoded, specifically:
[0047] Obtain the features of each image unit in the image frame to be encoded, where the features of each image unit include a direction factor and a quantization activity factor;
[0048] Based on the features of each image unit, determine the type of this image unit.
[0049] In one implementation, the processing unit is also used for:
[0050] Generate the bitstream data of the image frame based on the filtering coefficient group corresponding to the image frame;
[0051] Send the bitstream data to the decoding end so that the decoding end can present the image frame based on the bitstream data.
[0052] Correspondingly, the present application provides a computer device, which includes:
[0053] A memory, in which a computer program is stored;
[0054] A processor, configured to load the computer program to implement the above - mentioned data processing method.
[0055] Correspondingly, the present application provides a computer - readable storage medium, which stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the above - mentioned data processing method.
[0056] Correspondingly, the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer - readable storage medium. The processor of the computer device reads the computer instructions from the computer - readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above - mentioned data processing method.
[0057] In the embodiments of the present application, the types of each image unit in the image frame to be encoded are obtained, the image units that meet the downsampling conditions in each type of image unit are downsampled to obtain the downsampling results of each type of image unit, a Wiener-Hopf equation associated with each type of image unit is constructed based on the downsampling results of each type of image unit, and a filter coefficient group corresponding to the image frame is determined through the Wiener-Hopf equations associated with each type of image unit. The filter coefficient group is used to generate the bitstream data of the image frame. It can be seen that by downsampling each type of image unit, the number of pixel points included in each type of image unit can be reduced, thereby reducing the number of Wiener-Hopf equations (associated with pixel points) involved in constructing the Wiener-Hopf equation associated with each type of image unit, and further improving the overall encoding efficiency of the image frame. Description of the Drawings
[0058] Figure 1a Schematic diagram of a general video encoder architecture provided by an embodiment of the present application;
[0059] Figure 1b Data processing scenario diagram provided by an embodiment of the present application;
[0060] Figure 2 Flowchart of a data processing method provided by an embodiment of the present application;
[0061] Figure 3a Schematic diagram of a positional relationship provided by an embodiment of the present application;
[0062] Figure 3b Another schematic diagram of a positional relationship provided by an embodiment of the present application;
[0063] Figure 3c Schematic diagram of a filter for a luminance component provided by an embodiment of the present application;
[0064] Figure 3d Schematic diagram of a filter for a chrominance component provided by an embodiment of the present application;
[0065] Figure 3e Schematic diagram of an image unit provided by an embodiment of the present application;
[0066] Figure 4 Flowchart of another data processing method provided by an embodiment of the present application;
[0067] Figure 5 Schematic diagram of the structure of a data processing device provided by an embodiment of the present application;
[0068] Figure 6 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed Embodiments
[0069] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.
[0070] This application relates to technologies related to artificial intelligence and encoding / decoding. The related technologies involved are briefly introduced below:
[0071] Artificial Intelligence (AI): So-called AI is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machine to have the functions of perception, reasoning, and decision-making. The embodiments of the present application mainly involve analyzing the attribute information of image units (including at least one of the encoding technology adopted by the image unit, the position of the image unit in the view frame, and the color information of the pixel points in the image unit) through a downsampling analysis model to obtain an analysis result.
[0072] AI technology is an interdisciplinary subject involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operating / interactive systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the basic model, can be widely applied to downstream tasks in various major directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0073] Machine Learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how a computer simulates or implements human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make a computer intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pre-trained model is the latest development result of deep learning, integrating the above technologies. The embodiments of the present application mainly involve training the downsampling analysis model based on a sample data set to further improve the accuracy of the analysis result of the downsampling analysis model.
[0074] Loop filtering is one of the core technologies in video coding. The Versatile Video Coding (VVC) encoder can support three loop filters, including the Deblocking Filter (DF), the Sample Adaptive Offset (SAO), and the Adaptive Loop Filter (ALF). Among them, ALF is an adaptive filter based on the Wiener Filter, whose function is to optimize the output image signal to minimize the Mean Squared Error (MSE) between it and the original image, thereby achieving the effect of reducing distortion. Figure 1a This is a schematic diagram of a general video encoder architecture provided by an embodiment of this application. As Figure 1a shown, ALF in VVC is executed after the DF filter and the SAO filter. ALF derives the filter coefficients by solving the Wiener-Hopf equation of the Wiener filter, and it is determined by the encoder-side Rate-Distortion Optimization (RDO) whether to enable ALF. If ALF is enabled, the encoder needs to transfer the filter coefficient group to the decoder.
[0075] During the encoding process, the filter parameter signals of the Adaptive Loop Filter (ALF) and the Cross Component Adaptive Loop Filter (CCALF) can be included in the Adaptive Parameter Set (APS). One APS can contain up to 25 sets of luminance ALF filter coefficients and clipping values, and up to 8 sets of chrominance ALF filter coefficients and clipping values. Each chrominance component of CCALF can have up to 4 sets of filter coefficients in one APS.
[0076] For the purpose of saving bitrate, a Merge operation can be performed on the signals between different classifications of the luminance filter coefficients. The index of the APS used by the current Slice can be sent in the Slice Header.
[0077] It should be noted that to limit the computational complexity, the luminance and chrominance ALF filter coefficients can be quantized to integers in the range of [-2 7 , 2 7 -1] with a normalization constant of 128, and the coefficient at the center position is fixed at 128.
[0078] The clipping value index decoded from the APS can be used to confirm the clipping values of luminance and chrominance by looking up a table. These clipping values are related to the internal bit depth. For the specific correspondence between the clipping values and the internal bit depth, see Table 1 (the correspondence table of clipping values and internal bit depth):
[0079] Table 1
[0080]
[0081] For the luminance ALF filter, up to 7 APS index signals can be sent in the slice header to describe the luminance filter parameter set used for the current slice. Whether to use ALF for each coding tree block (CTB) can be controlled by sending a switch signal at the CTB level. Each CTB can select filter parameters from 16 sets of fixed ALF parameters and the APS of the current slice through the filter parameter set index. The 16 sets of fixed ALF parameters are predefined and stored in the encoder and decoder.
[0082] For the chrominance filter, an APS index can be sent in the slice header to specify the chrominance filter parameter set used for the current slice. If there are multiple filter parameters in the APS, the filter used for the current chrominance CTB can also be determined by sending the filter parameter set index at the CTB level.
[0083] Based on the above technologies related to artificial intelligence and coding / decoding, the embodiments of this application provide a data processing solution, which can improve the overall coding efficiency of image frames. Figure 1b A data processing scenario diagram provided by the embodiments of this application, as Figure 1bAs shown in the figure, the data processing scenario provided by this application includes a terminal device 101 and a server 102, and the data processing solution provided by this application can be executed by the server 102. Among them, the terminal device may include, but is not limited to: smart phones (such as Android phones, IOS phones, etc.), tablet computers, portable personal computers, mobile Internet devices (Mobile Internet Devices, abbreviated as MID), intelligent voice interaction devices, intelligent household appliances, vehicle-mounted terminals, aircraft, wearable devices, etc. The embodiments of this application do not make limitations in this regard; the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The embodiments of this application do not make limitations in this regard.
[0084] It should be noted that Figure 1b the numbers of the terminal device 101 and the server 102 are only for illustration and do not constitute an actual limitation of this application. The terminal device 101 and the server 102 can be connected by wired or wireless means, and this application does not make restrictions in this regard.
[0085] The general process of the data processing solution provided by this application is as follows:
[0086] (1) The server 102 obtains the types of each image unit in the image frame to be encoded, and each image unit includes at least two pixel points in the image frame. The image unit may include, but is not limited to, a pixel block composed of 4*4 pixel points, a coding tree unit (Coding Tree Unit, CTU). In one implementation manner, the server 102 obtains the features of each image unit in the image frame to be encoded, and determines the type of the image unit based on the features of each image unit; among them, the features of each image unit include a directionality factor (Directionality) and an activity factor (Activity).
[0087] (2) The server 102 performs downsampling on the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of each type of image unit. The downsampling conditions include at least one of the following: the image region associated with the image unit (such as the region with a distance less than z pixel points from the center point of the image unit, where z is a positive integer) does not overlap with the edge region of the image frame (that is, there is no overlapping part), the image unit is an image unit encoded using non-ScreenContent Coding (SCC) technology, the image region associated with the image unit has an overlapping part with the edge region of the image frame in the target direction (any one or more directions of the image frame), and the edge region can be determined by a boundary detection algorithm), there are target pixel points in the image frame, and the difference between the color information (including at least one of the luminance component and the chrominance component) of the target pixel points in the image frame and the color information obtained after sample adaptive compensation of the compression result of the target pixel points is less than the difference threshold. Downsampling refers to compressing the pixel points in the image unit; for example, assuming the image unit is a 4*4 pixel block, it can be downsampled horizontally to a 2*4 pixel block, or vertically downsampled to a 4*2 pixel block, or downsampled both horizontally and vertically to a 2*2 pixel block.
[0088] (3) The server 102 constructs the Wiener-Hopf equation associated with each type of image unit based on the downsampling results of each type of image unit. In one implementation, the image frame includes M types of image units, where M is a positive integer. The target type is any one of the M types, and the downsampling result of the image units of the target type includes N pixel points, where N is a positive integer. The server 102 constructs the Wiener-Hopf equation associated with N pixel points based on the color information of K pixel points associated with each pixel point and the color information of this pixel point; among them, the Wiener-Hopf equation associated with the i-th pixel point is constructed based on the color information of K pixel points associated with the i-th pixel point and the color information of the i-th pixel point, and i is a positive integer less than or equal to N. After obtaining the Wiener-Hopf equations associated with N pixel points, the server 102 performs an accumulation process on the Wiener-Hopf equations associated with N pixel points to obtain the Wiener-Hopf equation associated with the image units of the target type.
[0089] (4) The server 102 determines a set of filtering coefficients corresponding to the image frame through the Wiener - Hoff equations associated with each type of image unit. The set of filtering coefficients is used to generate the bit - stream data of the image frame. In one implementation, the server 102 groups each type of image unit according to a preset grouping rule, obtaining P groups of image units. There is a target group among the P groups of image units. The target group includes at least two types of image units, and P is an integer greater than 1. Then the server 102 merges the Wiener - Hoff equations associated with different types of image units in the target group, and takes the merged Wiener - Hoff equation as the Wiener - Hoff equation associated with the target group. After determining the Wiener - Hoff equation associated with the target group, the server 102 determines the set of filtering coefficients corresponding to the image frame based on the Wiener - Hoff equations associated with the P groups of image units. In another implementation, the image frame includes M types of image units, and M is a positive integer. The server 102 generates M combined results based on the M types of image units. Among them, the g - th combined result includes M - g + 1 types of image units, and g is a positive integer less than or equal to M. After obtaining the M combined results, the server 102 calculates the rate - distortion optimization corresponding to the M combined results through the Wiener - Hoff equations associated with the M types of image units, and determines the set of filtering coefficients associated with the target combined result as the set of filtering coefficients corresponding to the image frame. The target combined result is the combined result with the minimum rate - distortion optimization among the M combined results.
[0090] Further, the server 102 generates the bit - stream data of the image frame based on the set of filtering coefficients corresponding to the image frame, and sends the bit - stream data to the terminal device 101 (decoding end). After receiving the bit - stream data of the image frame, the terminal device 101 decodes the bit - stream data and presents the image frame based on the decoding result of the bit - stream data.
[0091] In the embodiments of the present application, the types of each image unit in the image frame to be encoded are obtained, and the image units that meet the down - sampling condition in each type of image unit are down - sampled to obtain the down - sampling results of each type of image unit. Based on the down - sampling results of each type of image unit, the Wiener - Hoff equation associated with each type of image unit is constructed. Through the Wiener - Hoff equations associated with each type of image unit, the set of filtering coefficients corresponding to the image frame is determined. The set of filtering coefficients is used to generate the bit - stream data of the image frame. It can be seen that by down - sampling each type of image unit, the number of pixel points included in each type of image unit can be reduced, thereby reducing the number of Wiener - Hoff equations (associated with pixel points) involved in constructing the Wiener - Hoff equation associated with each type of image unit, and further improving the overall encoding efficiency of the image frame.
[0092] Based on the above data processing solution, embodiments of the present application propose a more detailed data processing method, and the data processing method proposed in the embodiments of the present application will be introduced in detail below with reference to the accompanying drawings.
[0093] Please refer to Figure 2 , Figure 2 which is a flowchart of a data processing method provided by an embodiment of the present application. This data processing method can be executed by a computer device; specifically, the computer device can be Figure 1b the server 102 shown in Figure 2 As shown, this data processing method may include but is not limited to S201 - S204:
[0094] S201. Obtain the types of each image unit in the image frame to be encoded.
[0095] The image frame can specifically be an independent image or an image in a video, and the present application does not limit this. Each image unit includes at least two pixel points in the image frame. The image unit can include but is not limited to pixel blocks of various scales (such as composed of 4 * 4 pixel points), Coding Tree Units (CTUs).
[0096] In one implementation, the computer device obtains the features of each image unit in the image frame to be encoded, and classifies the image units according to the features of the image units (that is, based on the features of each image unit, determines the type of the image unit). In one embodiment, the features of each image unit include the directionality and activity of the image unit. The features of each image unit can be derived from the pixel gradient values of the image unit horizontally, vertically, diagonally, and anti - diagonally. The color information of different image units can be the same or different. The color information of the image unit includes the luminance component and chrominance component of the image unit, where the classification index C of the luminance component of the image unit can be determined by the directionality and activity of the image unit, and the specific formula can be expressed as:
[0097]
[0098] where C is the classification index of the image unit of the luminance component, D is the directionality of the image unit, is the activity of the image unit.
[0099] Both the directionality and activity of the image unit can be divided into 5 levels. The 5 levels of the directionality of the image unit can be expressed as D0 - D4, and the 5 levels of the activity of the image unit can be expressed as Based on the direction factor and quantization activity factor of an image unit, a computer device can divide image units into 25 types. For the specific division method of the 25 types of image units, refer to Table 2 (Image Unit Classification Relationship Table):
[0100] Table 2
[0101]
[0102] As shown in Table 2, a computer device can determine the type of an image unit based on the direction factor level and quantization activity factor level of each image unit.
[0103] S202. Downsample the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of image units of each type.
[0104] The downsampling conditions include at least one of the following: the image area associated with the image unit (such as the area within a distance less than z pixel points from the center point of the image unit, where z is a positive integer) does not coincide with the edge area of the image frame; the image unit is an image unit encoded using non-screen content coding (SCC) technology; the image area associated with the image unit has an overlapping part with the edge area of the image frame in the target direction (any one or more directions of the image frame); there are target pixel points in the image frame, and the difference between the color information of the target pixel points in the image frame (including at least one of the luminance component and the chrominance component) and the color information obtained after sample adaptive compensation of the compression result of the target pixel points is less than the difference threshold. Downsampling refers to compressing the pixel points in the image unit. For example, if the image unit is a 4*4 pixel block, it can be downsampled horizontally to a 2*4 pixel block, or vertically to a 4*2 pixel block, or both horizontally and vertically to a 2*2 pixel block.
[0105] In one implementation, the image frame includes M types of image units, where M is a positive integer. Let the target type be any one of the M types. The computer device detects whether there is a first image unit in the image units of the target type. The image area associated with the first image unit (the area within a distance less than z pixel points from the center point of the first image unit, where z is a positive integer) does not coincide with the edge area of the image frame (the area within a distance less than the distance threshold from the image frame boundary), or the first image unit is an image unit encoded using non-screen content coding (SCC) technology. If there is a first image unit in the image units of the target type, the computer device performs downsampling processing on the first image unit to obtain the downsampling result of the image units of the target type; where the downsampling processing includes at least one of horizontal downsampling processing and vertical downsampling processing.
[0106] Specifically, if there is a non-SCC region in the image unit of the target type (i.e., the image unit encoded by using a non-screen content coding technology), the computer device performs downsampling processing on the non-SCC region to obtain the downsampling result of the image unit of the target type. If there is an image unit in the image unit of the target type that does not coincide with the edge region of the image frame (i.e., the image region associated with the image unit does not coincide with the edge region of the image frame), the computer device performs downsampling processing on the image unit to obtain the downsampling result of the image unit of the target type. Optionally, if the image frame is encoded by using the non-SCC technology, the computer device may directly perform downsampling processing on the image frame.
[0107] Figure 3a This is a schematic diagram of a positional relationship provided by an embodiment of the present application. As Figure 3a shown, the image region 302 is the image region associated with the image unit 301. It should be noted that the shape of the image region 302 is only for illustration. In practical applications, the shape of the image region may also be a rectangle, a polygon, etc. The present application does not limit this. The image frame edge region 303 is the region whose distance from the image frame boundary is less than the distance threshold. Since the image region 302 does not coincide with the image frame edge region 303, the computer device performs downsampling processing on the image unit 301 to obtain the downsampling result of the image unit 301.
[0108] In another implementation manner, the image frame includes M types of image units, where M is a positive integer. Let the target type be any one of the M types. The computer device detects whether there is a second image unit in the image unit of the target type, and there is an overlapping part between the image region associated with the second image unit and the edge region of the image frame in the target direction. If there is a second image unit in the image unit of the target type, the computer device performs downsampling processing on the second image unit along the target direction to obtain the downsampling result of the image unit of the target type.
[0109] Figure 3b This is another schematic diagram of a positional relationship provided by an embodiment of the present application. As Figure 3b shown, the image region 305 is the image region associated with the image unit 304. It should be noted that the shape of the image region 305 is only for illustration. In practical applications, the shape of the image region may also be a rectangle, a polygon, etc. The present application does not limit this. The image frame edge region 306 is the region whose distance from the image frame boundary is less than the distance threshold. Since there is an overlapping part between the image region 305 and the image frame edge region 306, the computer device performs downsampling processing on the image unit 304 along the right direction (i.e., the target direction) (for example, the computer device performs downsampling processing on 8 pixel points on the right side of the middle line of the image unit 304) to obtain the downsampling result of the image unit 304.
[0110] In yet another embodiment, an image frame includes M types of image units, where M is a positive integer. Let the target type be any one of the M types. The computer device detects whether there is a target pixel point in the image units of the target type. The target pixel point is a pixel point in the image units of the target type whose difference value is less than the difference threshold. The difference value of each pixel point is used to indicate the difference between the first color information and the second color information of the pixel point. The first color information of any pixel point is the color information of the pixel point in the image frame (which can be a video frame obtained by filtering a video using a Temporal Filter). The second color information of any pixel point is the color information obtained by performing sample adaptive compensation on the compression result of the pixel point. The color information may include at least one of a luminance component and a chrominance component. The determination method of the target pixel point can be expressed as:
[0111] abs(src(P)-rec(P))<Thre
[0112] where abs(x) represents the absolute value of x, src(P) represents the color information of pixel point P in the image frame (i.e., the first color information), rec(P) is the color information obtained by performing sample adaptive compensation on the compression result of pixel point P (i.e., the second color information), and Thre is the difference threshold; that is, when the difference between the first color information and the second color information of pixel point P in the image frame is less than the difference threshold, the computer device determines pixel point P as the target pixel point.
[0113] It can be understood that in the specific implementation process, the computer device can determine whether multiple pixel points are target pixel points in parallel. The specific quantity can be determined based on the performance of the computer device, and this application does not limit it.
[0114] In one embodiment, the color information includes a luminance component. The difference value of each pixel point includes the difference between the luminance component of the pixel point in the image frame (the first luminance component) and the luminance component obtained by performing sample adaptive compensation on the compression result of the pixel point (the second luminance component). In this case, when the difference between the first luminance component and the second luminance component of a pixel point is less than the luminance threshold, the computer device determines the pixel point as the target pixel point.
[0115] In another embodiment, the color information includes a chrominance component, and the difference value of each pixel includes the difference between the chrominance component of the pixel in the image frame (the first chrominance component) and the chrominance component obtained after sample adaptive compensation of the compression result of the pixel (the second chrominance component). In this case, when the difference between the first chrominance component and the second chrominance component of a pixel is less than the luminance threshold, the computer device determines that the pixel is a target pixel.
[0116] In yet another embodiment, the color information includes a luminance component and a chrominance component, and the difference value of each pixel is used to indicate the difference between the first luminance component and the first chrominance component of the pixel in the image frame and the second luminance component and the second chrominance component obtained after sample adaptive compensation of the compression result of the pixel; that is, the difference value of each pixel is the sum of a first difference (the difference between the first luminance component and the second luminance component of the pixel) and a second difference (the difference between the first chrominance component and the second chrominance component of the pixel). In this case, when the difference between the first luminance component and the first chrominance component of a pixel in the image frame and the second luminance component and the second chrominance component obtained after sample adaptive compensation of the compression result of the pixel is less than the difference threshold, the computer device determines that the pixel is a target pixel.
[0117] Further, if there are target pixels in the image unit of the target type, the computer device removes the target pixels from the image unit of the target type to obtain the downsampling result of the image unit of the target type.
[0118] It should be noted that S202 can be executed not only after S201, but also before S201 or in parallel with S201 (when S202 is executed before S201 or in parallel with S201, the computer device can directly perform downsampling processing on the image units that meet the downsampling conditions according to the above downsampling conditions without distinguishing the types of image units), and this application does not limit this. When an image unit meets multiple downsampling conditions, the image unit can be downsampled multiple times; for example, assume that image unit 1 is encoded using non-SCC technology, and the computer device performs downsampling processing on image unit 1 to obtain image unit 2 (the number of pixels included in image unit 2 is less than that of image unit 1). If there are target pixels in image unit 2, the computer device removes the target pixels in image unit 2 to obtain image unit 3 (the number of pixels included in image unit 3 is less than that of image unit 2).
[0119] S203. Based on the downsampling results of the image units of each type, construct the Wiener-Hopf equation associated with the image units of this type.
[0120] In one implementation, an image frame includes M types of image units, where M is a positive integer. The target type is any one of the M types, and the downsampling result of the image units of the target type includes N pixel points, where N is a positive integer. The computer device constructs a Wiener-Hopf equation associated with the N pixel points based on the color information of K pixel points associated with each pixel point and the color information of the pixel point; among them, the Wiener-Hopf equation associated with the i-th pixel point is constructed based on the color information of K pixel points associated with the i-th pixel point and the color information of the i-th pixel point. The color information includes at least one of a luminance component and a chrominance component, and i is a positive integer less than or equal to N. After obtaining the Wiener-Hopf equations associated with the N pixel points, the computer device performs an accumulation process on the Wiener-Hopf equations associated with the N pixel points to obtain the Wiener-Hopf equation associated with the image units of the target type.
[0121] The pixel points associated with each pixel point are determined based on the size of the filter. The computer device can use filters of different sizes for the luminance component and the chrominance component of the pixel point respectively.
[0122] In one embodiment, the computer device uses a 7x7 diamond filter for the luminance component. The position corresponding to the center of the filter is the position of the currently filtered pixel point, and the pixel points symmetric to the center of the currently filtered pixel point use the same filtering coefficient. Figure 3c Schematic diagram of a filter for the luminance component provided by an embodiment of the present application. As Figure 3c shown, the size of the diamond filter is 7x7. When this diamond filter is used, the number of pixel points associated with the currently filtered pixel point is 24 (i.e., K = 24). The Wiener-Hopf equation of the luminance component associated with the i-th pixel point is constructed based on the luminance components of 24 pixel points associated with the i-th pixel point and the luminance of the i-th pixel point, where i is a positive integer less than or equal to N.
[0123] In another embodiment, the computer device uses a 5x5 diamond filter for the chrominance component. The position corresponding to the center of the filter is the position of the currently filtered pixel point, and the pixel points symmetric to the center of the currently filtered pixel point use the same filtering coefficient. Figure 3d Schematic diagram of a filter for the chrominance component provided by an embodiment of the present application. As Figure 3d shown, the size of the diamond filter is 5x5. When this diamond filter is used, the number of pixel points associated with the currently filtered pixel point is 12 (i.e., K = 12). The Wiener-Hopf equation of the chrominance component associated with the i-th pixel point is constructed based on the chrominance components of 12 pixel points associated with the i-th pixel point and the luminance of the i-th pixel point, where i is a positive integer less than or equal to N.
[0124] In another embodiment, the Wiener - Hoff equation associated with the image unit includes the Wiener - Hoff equation for the luminance component and the Wiener - Hoff equation for the chrominance component. The image frame includes M types of image units, where M is a positive integer. The target type is any one of the M types, and the down - sampling result of the image units of the target type includes N pixel points, where N is a positive integer. The Wiener - Hoff equation for the luminance component associated with the i - th pixel point is constructed based on the luminance components of K pixel points associated with the i - th pixel point and the luminance component of the i - th pixel point. The Wiener - Hoff equation for the luminance component associated with the image units of the target type is obtained by accumulating the Wiener - Hoff equations for the luminance components of the N pixel points, where i is a positive integer less than or equal to N; similarly, the Wiener - Hoff equation for the chrominance component associated with the i - th pixel point is constructed based on the chrominance components of K pixel points associated with the i - th pixel point and the chrominance component of the i - th pixel point. The Wiener - Hoff equation for the chrominance component associated with the image units of the target type is obtained by accumulating the Wiener - Hoff equations for the chrominance components of the N pixel points.
[0125] Figure 3e FIG. is a schematic diagram of an image unit provided by an embodiment of the present application. As Figure 3e shown, the gray - scale part (consisting of 4 * 4 pixel points) is the image unit. The following will describe the embodiment of constructing the Wiener - Hoff equation for the luminance component of the image unit in conjunction with Figure 3c and Figure 3e Let the currently filtered pixel point be pixel point P, and pixel point P corresponds to C12 in Figure 3c , pixel points A and B correspond to C0 in Figure 3c (centrally symmetric with pixel point P), pixel points C and D correspond to C9 in Figure 3c (centrally symmetric with pixel point P), and pixel points E and F correspond to C1 in Figure 3c (centrally symmetric with pixel point P). Calculate the non - linear clipping values corresponding to the positions of C0 - C12 in Figure 3c respectively, which can be expressed as M0 - M12; among them, M0 is calculated based on the luminance reconstruction values of pixel point A, pixel point B, and pixel point P. The luminance reconstruction value of pixel point A is the luminance component obtained by performing sample - point adaptive compensation on the compression result of pixel point A, the luminance reconstruction value of pixel point B is the luminance component obtained by performing sample - point adaptive compensation on the compression result of pixel point B, and the luminance reconstruction value of pixel point P is the luminance component obtained by performing sample - point adaptive compensation on the compression result of pixel point P. The specific calculation formula of M0 can be expressed as:
[0126] clip(rec(A)-rec(P))+clip(rec(B)-rec(P))
[0127] Among them, clip is the non-linear limit value of the adaptive loop filter, rec(A) is the luminance component obtained by performing sample adaptive compensation on the compression result of pixel point A, rec(B) is the luminance component obtained by performing sample adaptive compensation on the compression result of pixel point B, and rec(P) is the luminance component obtained by performing sample adaptive compensation on the compression result of pixel point P. Similarly, the specific calculation formula of M1 can be expressed as:
[0128] clip(rec(E)-rec(P))+clip(rec(F)-rec(P))
[0129] According to the above implementation manner, M0 - M12 can be calculated. After obtaining M0 - M12, the computer device can construct the Wiener - Hoff equation of pixel point P in the luminance component based on M0 - M12, the luminance component of pixel point P in the image frame, and the luminance reconstruction value of pixel point P. The Wiener - Hoff equation of pixel point P in the luminance component includes matrix X
[13]
[13] and matrix Y
[13] ; among them, the value of X[i][j] (both i and j are integers from 0 to 12) can be expressed as:
[0130] X[i][j] = Mi * Mj
[0131] That is to say, X
[13]
[13] can be obtained through M0 - M12. The value of Y[i] (i is an integer from 0 to 12) can be expressed as:
[0132] Y[i] = Mi * (src(P)-rec(P))
[0133] Among them, src(P) represents the luminance component of pixel point P in the image frame, and rec(P) is the luminance component obtained by performing sample adaptive compensation on the compression result of pixel point P.
[0134] According to the above implementation manner, the computer device can construct Figure 3e the Wiener - Hoff equations of each pixel point in the luminance component in the image unit (a total of 16). After obtaining the Wiener - Hoff equations of each pixel point in the luminance component, the computer device accumulates the 16 Wiener - Hoff equations of the pixel points in the luminance component to obtain the Wiener - Hoff equation of the image unit in the luminance component; among them, X
[13]
[13] in the Wiener - Hoff equation of the image unit in the luminance component is obtained by accumulating X
[13]
[13] in the Wiener - Hoff equations corresponding to the 16 pixel points, and Y
[13] in the Wiener - Hoff equation of the image unit in the luminance component is obtained by accumulating Y
[13] in the Wiener - Hoff equations corresponding to the 16 pixel points. Similarly, based on the Wiener - Hoff equations of the luminance components of the same type of image units, the Wiener - Hoff equation of the luminance component associated with this type of image unit can be constructed.
[0135] The following combines Figure 3d and Figure 3e to illustrate the implementation of the Wiener - Hoff equation for constructing image units in the chrominance component. Let the currently filtered pixel be pixel P, and pixel P corresponds to Figure 3d C6 in Figure 3d Pixels G and H correspond to Figure 3d C0 in (centrally symmetric to pixel P), pixels R and S correspond to Figure 3d C4 in (centrally symmetric to pixel P), and pixels Q and T correspond to Figure 3d C1 in (centrally symmetric to pixel P). Calculate the non - linear clipping values corresponding to the positions of C0 - C6 in
[0136] clip(rec(G)-rec(P))+clip(rec(H)-rec(P))
[0137] where clip is the non - linear clipping value of the adaptive loop filter, rec(G) is the chrominance component obtained by performing sample - adaptive compensation on the compression result of pixel G, rec(H) is the chrominance component obtained by performing sample - adaptive compensation on the compression result of pixel H, and rec(P) is the chrominance component obtained by performing sample - adaptive compensation on the compression result of pixel P. Similarly, the specific calculation formula for M1 can be expressed as:
[0138] clip(rec(Q)-rec(P))+clip(rec(T)-rec(P))
[0139] According to the above implementation, M0 - M6 can be calculated. After obtaining M0 - M6, the computer device can construct the Wiener - Hoff equation of pixel P in the chrominance component based on M0 - M6, the chrominance component of pixel P in the image frame, and the chrominance reconstruction value of pixel P. The Wiener - Hoff equation of pixel P in the chrominance component includes matrix X[7][7] and matrix Y[7]; where the value of X[i][j] (i, j are both integers from 0 to 6) can be expressed as:
[0140] X[i][j] = Mi * Mj
[0141] That is to say, X[7][7] can be obtained through M0 - M6. The value of Y[i] (where i is an integer from 0 to 6) can be expressed as:
[0142] Y[i] = Mi * (src(P) - rec(P))
[0143] Among them, src(P) represents the chrominance component of pixel point P in the image frame, and rec(P) is the chrominance component obtained after sample - adaptive compensation of the compression result of pixel point P.
[0144] According to the above - mentioned implementation manner, the computer device can construct Figure 3e the Wiener - Hoff equations (a total of 16) of each pixel point in the chrominance component in the image unit. After obtaining the Wiener - Hoff equations of each pixel point in the chrominance component, the computer device accumulates the 16 Wiener - Hoff equations of the pixel points in the chrominance component to obtain the Wiener - Hoff equation of the image unit in the chrominance component; among them, X[7][7] in the Wiener - Hoff equation of the image unit in the chrominance component is obtained by accumulating X[7][7] in the Wiener - Hoff equations corresponding to the 16 pixel points, and Y[7] in the Wiener - Hoff equation of the image unit in the chrominance component is obtained by accumulating Y[7] in the Wiener - Hoff equations corresponding to the 16 pixel points. Similarly, based on the Wiener - Hoff equations of the same - type image units in the chrominance component, the Wiener - Hoff equation associated with the chrominance component of this type of image unit can be constructed.
[0145] S204. Determine the filter coefficient group corresponding to the image frame through the Wiener - Hoff equations associated with each type of image unit.
[0146] The filter coefficient group is used to generate the bit - stream data of the image frame.
[0147] In one implementation manner, the computer device groups each type of image unit according to a preset grouping rule to obtain P groups of image units. There is a target group among the P groups of image units, and the target group includes at least two types of image units, where P is an integer greater than 1. Then the computer device merges the Wiener - Hoff equations associated with different types of image units in the target group and uses the merged Wiener - Hoff equation as the Wiener - Hoff equation associated with the target group. After determining the Wiener - Hoff equation associated with the target group, the computer device determines the filter coefficient group corresponding to the image frame based on the Wiener - Hoff equations associated with the P groups of image units.
[0148] In another embodiment, an image frame includes M types of image units, where M is a positive integer. The computer device generates M merging results based on the M types of image units; among them, the g-th merging result includes M - g + 1 types of image units, and g is a positive integer less than or equal to M. After obtaining the M merging results, the computer device calculates the rate-distortion optimization corresponding to the M merging results through the Wiener-Hopf equations associated with the M types of image units, and determines the filter coefficient group associated with the target merging result as the filter coefficient group corresponding to the image frame; where the target merging result is the merging result with the minimum rate-distortion optimization among the M merging results.
[0149] Further, the computer device may generate bitstream data of the image frame based on the filter coefficient group corresponding to the image frame, and send the bitstream data to the decoding end, so that after receiving the bitstream data of the image frame, the decoding end decodes the bitstream data and presents the image frame based on the decoding result of the bitstream data.
[0150] In the embodiments of the present application, the types of each image unit in the image frame to be encoded are obtained, the image units that meet the downsampling conditions in each type of image unit are downsampled to obtain the downsampling results of each type of image unit, and based on the downsampling results of each type of image unit, the Wiener-Hopf equation associated with this type of image unit is constructed. Through the Wiener-Hopf equations associated with each type of image unit, the filter coefficient group corresponding to the image frame is determined, and the filter coefficient group is used to generate the bitstream data of the image frame. It can be seen that by downsampling each type of image unit, the number of pixel points included in each type of image unit can be reduced, thereby reducing the number of Wiener-Hopf equations (associated with pixel points) involved in constructing the Wiener-Hopf equation associated with each type of image unit, and further improving the overall encoding efficiency of the image frame.
[0151] Please refer to Figure 4 , Figure 4 which is a flowchart of another data processing method provided by the embodiments of the present application. This data processing method can be executed by a computer device; specifically, the computer device can be Figure 1b the server 102 shown in Figure 4 As shown in
[0152] S401. Obtain the types of each image unit in the image frame to be encoded.
[0153] The specific implementation manner of S401 can refer to Figure 2 the implementation manner of S201 therein, which will not be elaborated here.
[0154] S402. Downsample the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of each type of image unit.
[0155] In one implementation, meeting the downsampling conditions means that the analysis result of the image unit indicates that the image unit needs to be downsampled. The computer device analyzes the attribute information of each type of image unit (including at least one of the coding technology adopted by the image unit, the position of the image unit in the view frame, and the color information of the pixel points in the image unit) through a downsampling analysis model to obtain an analysis result (the analysis result is used to indicate the corresponding downsampling strategy of the image unit); among them, the downsampling analysis model is obtained by training a model to be trained using a sample data set. Specifically, the sample data set includes sample data and verification data corresponding to the sample data. The computer device calls the model to be trained to analyze the sample data to obtain the analysis result of the sample data, and optimizes the model to be trained based on the difference between the analysis result of the sample data and the verification data corresponding to the sample data to obtain the downsampling analysis model. After obtaining the analysis result of each type of image unit, the computer device downsamples the image unit based on the corresponding downsampling strategy of each image unit to obtain the downsampled image unit.
[0156] S403. Based on the downsampling results of each type of image unit, construct the Wiener-Hopf equation associated with the image unit of this type.
[0157] For the specific implementation of S403, reference can be made to Figure 2 the implementation of S203 in, which will not be elaborated here.
[0158] S404. Generate P merging results based on M types of image units.
[0159] The g-th merging result includes P - g + 1 groups of image units, where g is a positive integer less than or equal to P. It can be understood that when M = P (that is, the M types of image units are not grouped), the g-th merging result includes M - g + 1 types of image units.
[0160] In one implementation, P is a positive integer less than M. The computer device groups each type of image unit according to a preset grouping rule to obtain P groups of image units, and there is a target grouping among the P groups of image units, and the target grouping includes at least two types of image units. The preset grouping rule may include: grouping the image units based on the directionality factor of the image unit, grouping the image units based on the activity factor of the image unit, and grouping the image units based on the directionality factor and the activity factor of the image unit.
[0161] After obtaining P groups of image units, the computer device merges the Wiener-Hopf equations associated with different types of image units in the target group, and uses the merged Wiener-Hopf equation as the Wiener-Hopf equation associated with the target group. In one embodiment, on the one hand, the computer device performs a fusion process on the left equalities of the Wiener-Hopf equations associated with different types of image units in the target group to obtain a first fusion result; on the other hand, the computer device performs a fusion process on the right equalities of the Wiener-Hopf equations associated with different types of image units in the target group to obtain a second fusion result. After obtaining the first fusion result and the second fusion result, the computer device constructs a target Wiener-Hopf equation based on the first fusion result and the second fusion result, and uses the target Wiener-Hopf equation as the Wiener-Hopf equation associated with the target group; wherein, the left equality of the target Wiener-Hopf equation is the first fusion result, and the right equality of the target Wiener-Hopf equation is the second fusion result.
[0162] Further, the computer device generates P combined results based on the P groups of image units. In one embodiment, the first combined result includes the P groups of image units, the g-th combined result is obtained by merging g groups of the P groups of image units, the g-th combined result includes P - g + 1 groups of image units, and g is an integer greater than 1 and less than or equal to P. In one embodiment, the process by which the computer device generates P combined results based on the P groups of image units includes: on the one hand, the computer device merges the i-th group of image units and the j-th group of image units to obtain the merged i-th group of image units (i.e., the merged group of image units); on the other hand, the computer device merges the Wiener-Hopf equation associated with the i-th group of image units and the Wiener-Hopf equation associated with the j-th group of image units to obtain the Wiener-Hopf equation associated with the merged i-th group of image units (i.e., the merged group of image units); wherein, the combined rate-distortion optimization associated with the i-th group of image units and the j-th group of image units is the smallest among the combined rate-distortion optimizations associated with any two groups of the P groups of image units.
[0163] The combined rate - distortion optimization is obtained by summing the rate - distortion optimization of the $i$-th group of picture elements and the rate - distortion optimization of the $j$-th group of picture elements. The rate - distortion optimization of the $i$-th group of picture elements is calculated based on the bit - rate overhead of the filter coefficients associated with the $i$-th group of picture elements and the distortion of each type of picture element in the $i$-th group of picture elements. Among them, the filter coefficients associated with the $i$-th group of picture elements are obtained by solving the Wiener - Hopf equation associated with the combined group (i.e., obtained by combining the Wiener - Hopf equation associated with the $i$-th group of picture elements and the Wiener - Hopf equation associated with the $j$-th group of picture elements). The distortion of each type of picture element in the $i$-th group of picture elements is calculated based on the difference between the reconstruction result of each type of picture element in the $i$-th group of picture elements and the original image (i.e., the picture frame). The reconstruction result of each type of picture element in the $i$-th group of picture elements is generated by the filter coefficients associated with the $i$-th group of picture elements (e.g., by filtering the compression result of each type of picture element in the $i$-th group of picture elements with the filter coefficients associated with the $i$-th group of picture elements to obtain the reconstruction result of each type of picture element in the $i$-th group of picture elements). Similarly, the rate - distortion optimization of the $j$-th group of picture elements is calculated based on the bit - rate overhead of the filter coefficients associated with the $j$-th group of picture elements and the distortion of each type of picture element in the $j$-th group of picture elements. Among them, the filter coefficients associated with the $j$-th group of picture elements are obtained by solving the Wiener - Hopf equation associated with the combined group (i.e., obtained by combining the Wiener - Hopf equation associated with the $i$-th group of picture elements and the Wiener - Hopf equation associated with the $j$-th group of picture elements). The distortion of each type of picture element in the $j$-th group of picture elements is calculated based on the difference between the reconstruction result of each type of picture element in the $j$-th group of picture elements and the original image (i.e., the picture frame). The reconstruction result of each type of picture element in the $j$-th group of picture elements is generated by the filter coefficients associated with the $i$-th group of picture elements. According to the above - mentioned implementation manner, the computer device can combine the $P$ groups of picture elements in pairs, solve the corresponding combined rate - distortion optimization, and combine the two groups of picture elements with the smallest combined rate - distortion optimization.
[0164] According to the above method, the computer device combines the two groups of picture elements with the smallest combined rate - distortion optimization in the first combined result (the first combined result includes $P$ groups of picture elements) to obtain the second combined result (the second combined result includes $P - 1$ groups of picture elements). Then the computer device combines the two groups of picture elements with the smallest combined rate - distortion optimization in the second combined result to obtain the third combined result (the third combined result includes $P - 2$ groups of picture elements). Repeat the above steps until the $P$-th combined result (the $P$-th combined result includes 1 group of picture elements) is obtained.
[0165] In another embodiment, P = M, and the computer device generates M combined results based on M types of image units. That is, the computer device may not group the M types of image units, but directly generate M combined results based on the M types of image units. In this case, the g-th combined result includes M - g + 1 types of image units.
[0166] In one embodiment, the process by which the computer device generates M combined results based on M types of image units includes: on the one hand, the computer device combines the x-th type of image unit and the y-th type of image unit to obtain the combined x-th type of image unit. The combined rate-distortion optimization associated with the x-th type of image unit and the y-th type of image unit is the smallest among the combined rate-distortion optimizations associated with any two types of image units among the M types of image units. On the other hand, the computer device combines the Wiener-Hopf equation associated with the x-th type of image unit and the Wiener-Hopf equation associated with the y-th type of image unit to obtain the combined Wiener-Hopf equation associated with the x-th type of image unit.
[0167] According to the above method, the computer device combines the two sets of image units with the smallest combined rate-distortion optimization in the first combined result (the first combined result includes M types of image units) to obtain the second combined result (the second combined result includes M - 1 types of image units). Then, the computer device combines the two sets of image units with the smallest combined rate-distortion optimization in the second combined result to obtain the third combined result (the third combined result includes M - 2 types of image units). Repeat the above steps until the M-th combined result is obtained (the M-th combined result includes 1 type of image unit, and this type of image unit is obtained by combining the M types of image units).
[0168] S405. Calculate the rate-distortion optimization corresponding to the P combined results through the Wiener-Hopf equations associated with the M types of image units.
[0169] When P = 1, the computer device combines the Wiener-Hopf equations associated with the M types of image units to obtain the Wiener-Hopf equation associated with the combined result, and solves the Wiener-Hopf equation associated with the combined result to obtain the filter coefficient group associated with the combined result. The computer device directly determines the filter coefficient group associated with the combined result as the filter coefficient group corresponding to the image frame.
[0170] When P is greater than 1 and less than or equal to M, the computer device solves the Wiener-Hopf equation associated with P groups of image units, obtains the filtering coefficients associated with each group of image units, and calculates the rate-distortion optimization of each group of image units based on the filtering coefficients associated with each group of image units. When P = M, the implementation manner for the computer device to calculate the rate-distortion optimization corresponding to the g-th combined result is as follows: The computer device solves the Wiener-Hopf equation associated with M - g + 1 types of image units, and obtains the filtering coefficients associated with each type of image unit. After obtaining the filtering coefficients associated with each type of image unit, the computer device calculates the rate-distortion optimization of each type of image unit based on the filtering coefficients associated with each type of image unit, and calculates the rate-distortion optimization corresponding to the g-th combined result through the rate-distortion optimizations of M - g + 1 types of image units.
[0171] In one embodiment, the computer device calculates the distortion of the h-th type of image unit based on the filtering coefficients associated with the h-th type of image unit, where h is a positive integer less than or equal to P - g + 1. Specifically, the computer device performs filtering processing on the image unit to be filtered based on the filtering coefficients associated with the h-th type of image unit, obtains the reconstruction result of the h-th type of image unit, and calculates the distortion of this type of image unit through the difference between the reconstruction result of the h-th type of image unit and the h-th type of image unit in the original image. Further, the computer device may calculate the rate-distortion optimization of the h-th type of image unit based on the distortion of the h-th type of image unit and the bitrate overhead of the filtering coefficients associated with the h-th type of image unit.
[0172] It can be understood that when a group of image units (P < M) includes multiple types of image units, the computer device can calculate the distortion of each type of image unit in this group of image units according to the above implementation manner. After obtaining the distortion of each type of image unit in each group of image units, the computer device calculates the rate-distortion optimization of this group of image units through the bitrate overhead of the filtering coefficients associated with each group of image units and the distortion of each type of image unit in this group of image units. Specifically, the computer device calculates the rate-distortion optimization of the h-th group of image units through the bitrate overhead of the filtering coefficients associated with the h-th group of image units and the distortion of each type of image unit in the h-th group of image units. In one implementation manner, the h-th group of image units includes Q types of image units, where Q is a positive integer, and the computer device calculates the rate-distortion optimization of the t-th type of image unit based on the distortion of the t-th type of image unit and the bitrate overhead of the filtering coefficients associated with the h-th group of image units, where t is a positive integer less than or equal to Q. After obtaining the rate-distortion optimizations of Q types of image units, the computer device performs a summation process on the rate-distortion optimizations of Q types of image units to obtain the rate-distortion optimization of the h-th group of image units.
[0173] According to the above embodiment, after obtaining the rate-distortion optimization of P-g+1 groups of image units, the computer device calculates the rate-distortion optimization corresponding to the g-th combined result through the rate-distortion optimization of the P-g+1 groups of image units. In one embodiment, the computer device accumulates the rate-distortion optimizations of the P-g+1 groups of image units to obtain the rate-distortion optimization corresponding to the g-th combined result.
[0174] Optionally, before filtering the image units to be filtered by the filter coefficients, the computer device may also perform a geometric transformation on the filter coefficients and the (clipping value) according to the gradient value of the image units to be filtered (including at least one of the gradient values of horizontal, vertical, diagonal, and skew diagonal). The geometric transformation includes 4 types of transformations: identity, diagonal transformation (Diagonal), vertical flip (Vertical Flip), and rotation (Rotation). Applying the geometric transformation to the filter parameters is equivalent to performing the corresponding geometric transformation on the image units to be filtered and then filtering them when the parameters remain unchanged. By performing a geometric transformation on the filter coefficients, the directionality of the filtering operation can be made closer, enabling more pixel points to share the same filter parameters, reducing distortion without encoding more filter parameters, and thus improving the overall coding efficiency.
[0175] S406: Determine the filter coefficient group associated with the target combined result as the filter coefficient group corresponding to the image frame.
[0176] The target combined result is the combined result with the minimum rate-distortion optimization among the P combined results. The filter coefficient group associated with the target combined result is obtained by solving the Wiener-Hopf equations associated with each group of image units in the target combined result. In one embodiment, the target combined result is the g-th combined result, and the g-th combined result includes P-g+1 groups of image units. The computer device separately solves the Wiener-Hopf equations associated with the P-g+1 groups of image units to obtain P-g+1 groups of filter coefficients, and determines these P-g+1 groups of filter coefficients as the filter coefficient group corresponding to the image frame.
[0177] S407: Generate the bitstream data of the image frame based on the filter coefficient group corresponding to the image frame.
[0178] In one implementation, the computer device generates a strip header based on the filter coefficient group corresponding to the image frame, and generates the bitstream data of the image frame through this strip header. After obtaining the bitstream data of the image frame, the computer device sends the bitstream data of the image frame to the decoding end (such as a terminal device). Correspondingly, after the decoding end obtains the bitstream data of the image frame, it decodes the bitstream data and presents the image frame based on the decoding result of the bitstream data. The decoding process includes filtering the reconstructed pixel points based on the filter coefficient group to obtain the filtered pixel points, which can be specifically expressed as:
[0179]
[0180] Among them, R(I,j) represents the reconstructed pixel point, R′(I,j) represents the pixel point after filtering, f(k,l) represents the filtering parameter corresponding to the pixel point to be filtered, K(x,y) is the clipping value formula, and K(x,y) = min(y, max(-y, x)) = Clip3(-y, y, x). c(k,l) is the solved clipping value parameter, and the variables k and l take values in the range of [-L / 2, L / 2], where L is the filter length. It should be noted that the clipping value operation introduces non-linear characteristics, which can effectively reduce the interference of the surrounding pixel values on the filtering output of the current pixel when the surrounding pixel values are quite different from the current pixel.
[0181] In the embodiments of the present application, the types of each image unit in the image frame to be encoded are obtained, the image units that meet the downsampling conditions in each type of image unit are downsampled to obtain the downsampling results of each type of image unit, and based on the downsampling results of each type of image unit, a Wiener-Hopf equation associated with the image unit of this type is constructed. Through the Wiener-Hopf equations associated with each type of image unit, a filter coefficient group corresponding to the image frame is determined, and the filter coefficient group is used to generate the bitstream data of the image frame. It can be seen that by downsampling each type of image unit, the number of pixel points included in each type of image unit can be reduced, thereby reducing the number of (pixel-point-associated) Wiener-Hopf equations participating in the construction of the Wiener-Hopf equation associated with each type of image unit, and further improving the overall encoding efficiency of the image frame. In addition, by grouping each type of image unit and merging the Wiener-Hopf equations associated with different types of image units in the same group, the number of Wiener-Hopf equations to be solved in the encoding process can be reduced (there is no need to solve the Wiener-Hopf equation associated with each type of image unit), and the computational amount of rate-distortion optimization of the filter coefficient group (the number of groups is less than the number of types of image units) can be reduced, further improving the encoding efficiency.
[0182] The method of the embodiments of the present application is elaborated in detail above. To facilitate better implementation of the above solutions of the embodiments of the present application, correspondingly, the device of the embodiments of the present application is provided below.
[0183] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a data processing device provided by the embodiments of the present application. This device can be mounted on a computer device, and the computer device can specifically be Figure 1b the server 102 shown. Figure 5 The data processing device shown can be used to execute the above Figure 2 and Figure 4Some or all of the functions in the described method embodiments. Please refer to Figure 5 , and the detailed descriptions of each unit are as follows:
[0184] An acquisition unit 501, configured to acquire the types of each image unit in an image frame to be encoded, and each image unit includes at least two pixel points in the image frame;
[0185] A processing unit 502, configured to perform downsampling processing on the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of each type of image unit;
[0186] and configured to construct a Wiener-Hopf equation associated with the image unit of each type based on the downsampling results of the image unit of each type;
[0187] and configured to determine a set of filter coefficients corresponding to the image frame through the Wiener-Hopf equations associated with the image units of each type, and the set of filter coefficients is used to generate the bitstream data of the image frame.
[0188] In one implementation, the image frame includes M types of image units, where M is a positive integer; the target type is any one of the M types; the processing unit 502 is configured to perform downsampling processing on the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of each type of image unit, and specifically:
[0189] detect whether there is a first image unit in the image units of the target type, where the image area associated with the first image unit does not coincide with the edge area of the image frame, or the first image unit is an image unit encoded by using a non-screen content encoding technology;
[0190] If there is a first image unit in the image units of the target type, perform downsampling processing on the first image unit to obtain the downsampling result of the image unit of the target type;
[0191] wherein, the downsampling processing includes at least one of horizontal downsampling processing and vertical downsampling processing.
[0192] In one implementation, the image frame includes M types of image units, where M is a positive integer; the target type is any one of the M types; the processing unit 502 is configured to perform downsampling processing on the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of each type of image unit, and specifically:
[0193] detect whether there is a second image unit in the image units of the target type, where there is an overlapping part between the image area associated with the second image unit and the edge area of the image frame in the target direction;
[0194] If there is a second image unit in the image units of the target type, downsample the second image unit along the target direction to obtain the downsampling result of the image units of the target type.
[0195] In one implementation, the image frame includes M types of image units, where M is a positive integer; the target type is any one of the M types; the processing unit 502 is configured to downsample the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of the image units of each type. Specifically, it is configured to:
[0196] Detect whether there is a target pixel point in the image units of the target type. The target pixel point is a pixel point in the image units of the target type whose difference value is less than the difference threshold. The difference value of each pixel point is used to indicate the difference between the first color information and the second color information of the pixel point. The first color information of any pixel point is the color information of the pixel point in the image frame, and the second color information of any pixel point is the color information obtained after performing sample adaptive compensation on the compression result of the pixel point.
[0197] If there is a target pixel point in the image units of the target type, remove the target pixel point from the image units of the target type to obtain the downsampling result of the image units of the target type.
[0198] In one implementation, the image frame includes M types of image units, where M is a positive integer; the target type is any one of the M types, and the downsampling result of the image units of the target type includes N pixel points, where N is a positive integer; the processing unit 502 is configured to construct the Wiener-Hopf equation associated with the image units of this type based on the downsampling results of the image units of each type. Specifically, it is configured to:
[0199] Based on the color information of K pixel points associated with each pixel point and the color information of this pixel point, construct the Wiener-Hopf equation associated with N pixel points. The Wiener-Hopf equation associated with the i-th pixel point is constructed based on the color information of K pixel points associated with the i-th pixel point and the color information of the i-th pixel point, where i is a positive integer less than or equal to N.
[0200] Perform an accumulation process on the Wiener-Hopf equations associated with N pixel points to obtain the Wiener-Hopf equation associated with the image units of the target type.
[0201] In one implementation, the processing unit 502 is configured to determine the filter coefficient group corresponding to the image frame through the Wiener-Hopf equations associated with the image units of each type. Specifically, it is configured to:
[0202] Group the image units of each type according to the preset grouping rules to obtain P groups of image units. There is a target group among the P groups of image units, and the target group includes at least two types of image units. P is a positive integer;
[0203] Merge the Wiener-Hopf equations associated with the image units of different types in the target group, and use the merged Wiener-Hopf equation as the Wiener-Hopf equation associated with the target group;
[0204] Based on the Wiener-Hopf equations associated with the P groups of image units, determine the filter coefficient group corresponding to the image frame.
[0205] In one embodiment, the processing unit 502 is configured to determine the filter coefficient group corresponding to the image frame based on the Wiener-Hopf equations associated with the P groups of image units, and specifically:
[0206] Generate P combined results based on the P groups of image units. The h-th combined result includes P - h + 1 groups of image units, and h is a positive integer less than or equal to P;
[0207] Calculate the rate-distortion optimization corresponding to the P combined results through the Wiener-Hopf equations associated with the P groups of image units;
[0208] Determine the filter coefficient group associated with the target combined result as the filter coefficient group corresponding to the image frame. The target combined result is the combined result with the minimum rate-distortion optimization among the P combined results.
[0209] In one embodiment, the image frame includes M types of image units, and M is a positive integer; the processing unit 502 is configured to determine the filter coefficient group corresponding to the image frame through the Wiener-Hopf equations associated with each type of image unit, and specifically:
[0210] Generate M combined results based on the M types of image units. The g-th combined result includes M - g + 1 types of image units, and g is a positive integer less than or equal to M;
[0211] Calculate the rate-distortion optimization corresponding to the M combined results through the Wiener-Hopf equations associated with the M types of image units;
[0212] Determine the filter coefficient group associated with the target combined result as the filter coefficient group corresponding to the image frame. The target combined result is the combined result with the minimum rate-distortion optimization among the M combined results.
[0213] In one embodiment, the process of the processing unit 502 generating the M combined results based on the M types of image units includes:
[0214] Merge the x-th type of image units and the y-th type of image units to obtain the merged x-th type of image units. The rate-distortion optimization associated with the x-th type of image units and the y-th type of image units is the smallest among the rate-distortion optimizations associated with any two types of image units among the M types of image units.
[0215] Merge the Wiener-Hopf equation associated with the x-th type of image units and the Wiener-Hopf equation associated with the y-th type of image units to obtain the merged Wiener-Hopf equation associated with the x-th type of image units.
[0216] In one embodiment, the processing unit 502 is configured to calculate the rate-distortion optimization corresponding to M combinations of merge results through the Wiener-Hopf equations associated with the M types of image units. Specifically, it is configured to:
[0217] Solve the Wiener-Hopf equations associated with M-g+1 types of image units to obtain the filtering coefficients associated with each type of image unit;
[0218] Based on the filtering coefficients associated with each type of image unit, calculate the rate-distortion optimization of this type of image unit;
[0219] Calculate the rate-distortion optimization corresponding to the g-th combination of merge results through the rate-distortion optimizations of M-g+1 types of image units.
[0220] In one embodiment, the processing unit 502 is configured to obtain the types of each image unit in the image frame to be encoded. Specifically, it is configured to:
[0221] Obtain the features of each image unit in the image frame to be encoded. The features of each image unit include a direction factor and a quantization activity factor;
[0222] Based on the features of each image unit, determine the type of this image unit.
[0223] In one embodiment, the processing unit 502 is further configured to:
[0224] Generate the bitstream data of the image frame based on the filtering coefficient group corresponding to the image frame;
[0225] Send the bitstream data to the decoding end so that the decoding end can present the image frame based on the bitstream data.
[0226] According to an embodiment of the present application, Figure 2 and Figure 4 Some of the steps involved in the data processing method shown can be executed by each unit in the Figure 5 shown data processing device. For example, Figure 2 S201 shown in Figure 5The obtaining unit 501 shown executes, and S202 - S204 can be executed by Figure 5 the processing unit 502 shown. Figure 4 S401 shown in Figure 5 can be executed by the obtaining unit 501 shown, and S402 - S407 can be executed by Figure 5 the processing unit 502 shown. Figure 5 Each unit in the data processing device shown can be respectively or all combined into one or several other units to form, or a certain (some) unit can also be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, the data processing device can also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.
[0227] According to another embodiment of this application, it can be achieved by running a computer program (including program code) that can execute the respective steps involved in the corresponding methods shown in Figure 2 and Figure 4 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct a data processing device as shown in Figure 5 and to implement the data processing method of the embodiments of this application. The computer program can be recorded on, for example, a computer-readable recording medium, and be loaded into the above computing device through the computer-readable recording medium and run therein.
[0228] Based on the same inventive concept, the principle and beneficial effects of the data processing device provided in the embodiments of this application for solving problems are similar to the principle and beneficial effects of the data processing method in the method embodiments of this application. The principle and beneficial effects of the method implementation can be referred to. For the sake of brevity of description, they will not be elaborated here.
[0229] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a computer device provided by the embodiments of this application. As shown in Figure 6As shown in the figure, the computer device at least includes a processor 601, a communication interface 602, and a memory 603. Among them, the processor 601, the communication interface 602, and the memory 603 can be connected through a bus or other means. The processor 601 (or the Central Processing Unit (CPU)) is the computing core and control core of the computer device. It can parse various instructions in the computer device and process various data of the computer device. For example, the CPU can be used to parse the power-on and power-off instructions sent by the user to the computer device and control the computer device to perform power-on and power-off operations. Another example is that the CPU can transmit various interactive data between the internal structures of the computer device, and so on. The communication interface 602 can optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.). Controlled by the processor 601, it can be used to send and receive data. The communication interface 602 can also be used for the transmission and interaction of internal data of the computer device. The memory 603 (Memory) is the memory device in the computer device and is used to store programs and data. It can be understood that the memory 603 here can include both the built-in memory of the computer device and, of course, the extended memory supported by the computer device. The memory 603 provides a storage space, and this storage space stores the operating system of the computer device, which can include but is not limited to: Android system, iOS system, Windows Phone system, etc. This application does not make any limitations in this regard.
[0230] An embodiment of this application also provides a computer-readable storage medium (Memory). The computer-readable storage medium is the memory device in the computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the computer device. And, in this storage space, there are also stored one or more instructions suitable for being loaded and executed by the processor 601. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.
[0231] In one embodiment, the computer device can specifically be Figure 1b the server 102 shown in the figure. The processor 601 executes the following operations by running the executable program code in the memory 603:
[0232] Obtain the types of each picture element in the image frame to be encoded, where each picture element includes at least two pixel points in the image frame;
[0233] Perform downsampling on the picture elements that meet the downsampling conditions in each type of picture element to obtain the downsampling results of each type of picture element;
[0234] Based on the downsampling results of each type of picture element, construct the Wiener-Hopf equation associated with the picture element of this type;
[0235] Determine the filter coefficient group corresponding to the image frame through the Wiener-Hopf equations associated with each type of picture element, and the filter coefficient group is used to generate the bitstream data of the image frame.
[0236] As an optional embodiment, the image frame includes M types of picture elements, where M is a positive integer; the target type is any one of the M types; the specific embodiment in which the processor 601 performs downsampling on the picture elements that meet the downsampling conditions in each type of picture element to obtain the downsampling results of each type of picture element is:
[0237] Detect whether there is a first picture element in the picture elements of the target type, where the image area associated with the first picture element does not coincide with the edge area of the image frame, or the first picture element is a picture element encoded by using a non-screen content encoding technique;
[0238] If there is a first picture element in the picture elements of the target type, perform downsampling on the first picture element to obtain the downsampling result of the picture elements of the target type;
[0239] Among them, the downsampling process includes at least one of horizontal downsampling and vertical downsampling.
[0240] As an optional embodiment, the image frame includes M types of picture elements, where M is a positive integer; the target type is any one of the M types; the specific embodiment in which the processor 601 performs downsampling on the picture elements that meet the downsampling conditions in each type of picture element to obtain the downsampling results of each type of picture element is:
[0241] Detect whether there is a second picture element in the picture elements of the target type, where there is an overlapping part between the image area associated with the second picture element and the edge area of the image frame in the target direction;
[0242] If there is a second picture element in the picture elements of the target type, perform downsampling on the second picture element along the target direction to obtain the downsampling result of the picture elements of the target type.
[0243] As an alternative embodiment, the image frame includes M types of image units, where M is a positive integer; the target type is any one of the M types; a specific embodiment in which the processor 601 performs downsampling processing on the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of each type of image unit is as follows:
[0244] Detect whether there are target pixel points in the image units of the target type. The target pixel points are pixel points in the image units of the target type with a difference value less than the difference threshold. The difference value of each pixel point is used to indicate the difference between the first color information and the second color information of the pixel point. The first color information of any pixel point is the color information of the pixel point in the image frame, and the second color information of any pixel point is the color information obtained after performing sample adaptive compensation on the compression result of the pixel point;
[0245] If there are target pixel points in the image units of the target type, then remove the target pixel points from the image units of the target type to obtain the downsampling result of the image units of the target type.
[0246] As an alternative embodiment, the image frame includes M types of image units, where M is a positive integer; the target type is any one of the M types, and the downsampling result of the image units of the target type includes N pixel points, where N is a positive integer; a specific embodiment in which the processor 601 constructs the Wiener-Hopf equation associated with the image units of this type based on the downsampling results of each type of image unit is as follows:
[0247] Based on the color information of K pixel points associated with each pixel point and the color information of this pixel point, construct the Wiener-Hopf equation associated with N pixel points. The Wiener-Hopf equation associated with the i-th pixel point is constructed based on the color information of K pixel points associated with the i-th pixel point and the color information of the i-th pixel point, where i is a positive integer less than or equal to N;
[0248] Perform an accumulation process on the Wiener-Hopf equations associated with N pixel points to obtain the Wiener-Hopf equation associated with the image units of the target type.
[0249] As an alternative embodiment, a specific embodiment in which the processor 601 determines the filter coefficient group corresponding to the image frame through the Wiener-Hopf equations associated with each type of image unit is as follows:
[0250] Group each type of image unit according to a preset grouping rule to obtain P groups of image units. There is a target group among the P groups of image units. The target group includes at least two types of image units, where P is a positive integer;
[0251] Merge the Wiener-Hopf equations associated with different types of image units in the target group, and use the merged Wiener-Hopf equation as the Wiener-Hopf equation associated with the target group;
[0252] Based on the Wiener-Hopf equations associated with P groups of image units, determine the filter coefficient group corresponding to the image frame.
[0253] As an optional embodiment, the specific embodiment in which the processor 601 determines the filter coefficient group corresponding to the image frame based on the Wiener-Hopf equations associated with P groups of image units is as follows:
[0254] Generate P combined results based on P groups of image units. The h-th combined result includes P - h + 1 groups of image units, where h is a positive integer less than or equal to P;
[0255] Calculate the rate-distortion optimization corresponding to the P combined results through the Wiener-Hopf equations associated with the P groups of image units.
[0256] Determine the filter coefficient group associated with the target combined result as the filter coefficient group corresponding to the image frame, where the target combined result is the combined result with the minimum rate-distortion optimization among the P combined results.
[0257] As an optional embodiment, the image frame includes M types of image units, where M is a positive integer; the specific embodiment in which the processor 601 determines the filter coefficient group corresponding to the image frame through the Wiener-Hopf equations associated with each type of image unit is as follows:
[0258] Generate M combined results based on M types of image units. The g-th combined result includes M - g + 1 types of image units, where g is a positive integer less than or equal to M;
[0259] Calculate the rate-distortion optimization corresponding to the M combined results through the Wiener-Hopf equations associated with the M types of image units.
[0260] Determine the filter coefficient group associated with the target combined result as the filter coefficient group corresponding to the image frame, where the target combined result is the combined result with the minimum rate-distortion optimization among the M combined results.
[0261] As an optional embodiment, the process by which the processor 601 generates M combined results based on M types of image units includes:
[0262] Merge the x-th type of image unit and the y-th type of image unit to obtain the merged x-th type of image unit. The combined rate-distortion optimization associated with the x-th type of image unit and the y-th type of image unit is the minimum among the combined rate-distortion optimizations associated with any two types of image units among the M types of image units;
[0263] Combine the Wiener - Hoff equation associated with the x - th type of image unit and the Wiener - Hoff equation associated with the y - th type of image unit to obtain the combined Wiener - Hoff equation associated with the x - th type of image unit.
[0264] As an alternative embodiment, a specific embodiment of the processor 601 calculating the rate - distortion optimization corresponding to M combined results through the Wiener - Hoff equations associated with M types of image units is as follows:
[0265] Solve the Wiener - Hoff equations associated with M - g + 1 types of image units to obtain the filtering coefficients associated with each type of image unit;
[0266] Based on the filtering coefficients associated with each type of image unit, calculate the rate - distortion optimization of the image unit of this type;
[0267] Calculate the rate - distortion optimization corresponding to the g - th combined result through the rate - distortion optimizations of M - g + 1 types of image units.
[0268] As an alternative embodiment, a specific embodiment of the processor 601 obtaining the types of each image unit in the image frame to be encoded is as follows:
[0269] Obtain the features of each image unit in the image frame to be encoded, where the features of each image unit include a direction factor and a quantization activity factor;
[0270] Based on the features of each image unit, determine the type of this image unit.
[0271] As an alternative embodiment, the processor 601 also performs the following operations by running the executable program code in the memory 603:
[0272] Generate the bit - stream data of the image frame based on the filtering coefficient group corresponding to the image frame;
[0273] Send the bit - stream data to the decoding end so that the decoding end presents the image frame based on the bit - stream data.
[0274] Based on the same inventive concept, the principle of problem - solving and the beneficial effects of the computer device provided in the embodiments of the present application are similar to those of the data - processing method in the method embodiments of the present application. One can refer to the principle of problem - solving and the beneficial effects of the method. For the sake of brevity, it will not be elaborated here.
[0275] The embodiments of the present application also provide a computer - readable storage medium. The computer - readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by a processor to perform the data - processing method in the above - mentioned method embodiments.
[0276] An embodiment of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above data processing method.
[0277] The steps in the method of the embodiment of the present application can be adjusted, combined, and deleted according to actual needs.
[0278] The modules in the device of the embodiment of the present application can be combined, divided, and deleted according to actual needs.
[0279] In the embodiment of the present application, the "module" or "unit" involved refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0280] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The computer-readable storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0281] The above-disclosed is only a preferred embodiment of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the application.
Claims
1. A data processing method, characterized in that, The method includes: obtaining the types of each image unit in the image frame to be encoded, where each image unit includes at least two pixel points in the image frame; performing downsampling processing on the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of each type of image unit; constructing a Wiener-Hopf equation associated with the image unit of each type based on the downsampling results of the image unit of each type; determining a filter coefficient group corresponding to the image frame through the Wiener-Hopf equations associated with the image units of each type, where the filter coefficient group is used to generate the bitstream data of the image frame.
2. The method according to claim 1, wherein The image frame includes M types of image units, where M is a positive integer; the target type is any one of the M types; The performing downsampling processing on the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of each type of image unit includes: detecting whether there is a first image unit in the image units of the target type, where the image area associated with the first image unit does not coincide with the edge area of the image frame, or the first image unit is an image unit encoded using a non-screen content encoding technique; if there is the first image unit in the image units of the target type, performing downsampling processing on the first image unit to obtain the downsampling result of the image units of the target type; wherein the downsampling processing includes at least one of horizontal downsampling processing and vertical downsampling processing.
3. The method according to claim 1, characterized in that, The image frame includes M types of image units, where M is a positive integer; the target type is any one of the M types; The performing downsampling processing on the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of each type of image unit includes: detecting whether there is a second image unit in the image units of the target type, where there is an overlapping part between the image area associated with the second image unit and the edge area of the image frame in the target direction; if there is the second image unit in the image units of the target type, performing downsampling processing on the second image unit along the target direction to obtain the downsampling result of the image units of the target type.
4. The method according to claim 1, wherein The image frame includes M types of image units, where M is a positive integer; the target type is any one of the M types; The performing downsampling processing on the image units that meet the downsampling conditions in each type of image unit to obtain the downsampling results of each type of image unit includes: detecting whether there is a target pixel point in the image units of the target type, where the target pixel point is a pixel point in the image units of the target type with a difference value less than a difference threshold, and the difference value of each pixel point is used to indicate the difference between the first color information and the second color information of the pixel point, the first color information of any pixel point is the color information of the pixel point in the image frame, and the second color information of any pixel point is the color information obtained after performing sample adaptive compensation on the compression result of the pixel point; If the target pixel point exists in the image units of the target type, remove the target pixel point from the image units of the target type to obtain the downsampling result of the image units of the target type.
5. The method according to claim 1, characterized in that, The image frame includes M types of image units, where M is a positive integer; the target type is any one of the M types, and the downsampling result of the image units of the target type includes N pixel points, where N is a positive integer; Constructing the Wiener-Hopf equation associated with the image units of each type based on the downsampling results of the image units of each type includes: Construct the Wiener-Hopf equation associated with the N pixel points based on the color information of K pixel points associated with each pixel point and the color information of this pixel point. The Wiener-Hopf equation associated with the i-th pixel point is constructed based on the color information of K pixel points associated with the i-th pixel point and the color information of the i-th pixel point, where i is a positive integer less than or equal to N; Perform an accumulation process on the Wiener-Hopf equations associated with the N pixel points to obtain the Wiener-Hopf equation associated with the image units of the target type.
6. The method according to claim 1, characterized in that, Determining the filter coefficient group corresponding to the image frame through the Wiener-Hopf equations associated with the image units of each type includes: Group the image units of each type according to a preset grouping rule to obtain P groups of image units. There is a target group in the P groups of image units. The target group includes at least two types of image units, where P is a positive integer; Merge the Wiener-Hopf equations associated with the different types of image units in the target group, and use the merged Wiener-Hopf equation as the Wiener-Hopf equation associated with the target group; Determine the filter coefficient group corresponding to the image frame based on the Wiener-Hopf equations associated with the P groups of image units.
7. The method according to claim 6, wherein Determining the filter coefficient group corresponding to the image frame based on the Wiener-Hopf equations associated with the P groups of image units includes: Generate P combined results based on the P groups of image units. The h-th combined result includes P - h + 1 groups of image units, where h is a positive integer less than or equal to P; Calculate the rate-distortion optimization corresponding to the P combined results through the Wiener-Hopf equations associated with the P groups of image units; Determine the filter coefficient group associated with the target combined result as the filter coefficient group corresponding to the image frame. The target combined result is the combined result with the minimum rate-distortion optimization among the P combined results.
8. The method according to claim 1, wherein The image frame includes M types of image units, where M is a positive integer; determining the filter coefficient group corresponding to the image frame through the Wiener-Hopf equations associated with the image units of each type includes: Generate M combined results based on the M types of image units. The g-th combined result includes M - g + 1 types of image units, where g is a positive integer less than or equal to M; Calculate the rate-distortion optimization corresponding to the M combined results through the Wiener-Hopf equations associated with the M types of image units; Determine the filter coefficient group associated with the target combined result as the filter coefficient group corresponding to the image frame. The target combined result is the combined result with the minimum rate-distortion optimization among the M combined results.
9. The method according to claim 8, wherein The process of generating M combined results based on the M types of image units includes: Combining the x-th type of image unit and the y-th type of image unit to obtain the combined x-th type of image unit, and the combined rate-distortion optimization associated with the x-th type of image unit and the y-th type of image unit is the smallest among the combined rate-distortion optimizations associated with any two types of image units in the M types of image units; Combining the Wiener-Hopf equation associated with the x-th type of image unit and the Wiener-Hopf equation associated with the y-th type of image unit to obtain the Wiener-Hopf equation associated with the combined x-th type of image unit.
10. The method according to claim 8, characterized in that, The calculation of the rate-distortion optimization corresponding to the M combined results through the Wiener-Hopf equations associated with the M types of image units includes: Solving the Wiener-Hopf equations associated with the M - g + 1 types of image units to obtain the filtering coefficients associated with each type of image unit; Based on the filtering coefficients associated with each type of image unit, calculating the rate-distortion optimization of this type of image unit; Calculating the rate-distortion optimization corresponding to the g-th combined result through the rate-distortion optimizations of the M - g + 1 types of image units.
11. The method according to claim 1, characterized in that, The obtaining of the types of the respective image units in the image frame to be encoded includes: Obtaining the features of the respective image units in the image frame to be encoded, where the feature of each image unit includes a direction factor and a quantization activity factor; Based on the features of each image unit, determining the type of this image unit.
12. The method according to claim 1, wherein The method further includes: Generating the bitstream data of the image frame based on the filtering coefficient group corresponding to the image frame; Sending the bitstream data to the decoding end so that the decoding end presents the image frame based on the bitstream data.
13. A data processing device, characterized in that, The data processing device includes: An obtaining unit, configured to obtain the types of the respective image units in the image frame to be encoded, and each image unit includes at least two pixel points in the image frame; A processing unit, configured to perform downsampling processing on the image units that meet the downsampling condition in each type of image unit to obtain the downsampling results of the respective types of image units; And for constructing the Wiener-Hopf equation associated with each type of image unit based on the downsampling results of each type of image unit; And for determining the filtering coefficient group corresponding to the image frame through the Wiener-Hopf equations associated with the respective types of image units, and the filtering coefficient group is used to generate the bitstream data of the image frame.
14. A computer device, characterized in that, Includes: A memory, in which a computer program is stored; A processor, configured to load the computer program to implement the data processing method according to any one of claims 1 - 12.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the data processing method according to any one of claims 1 - 12.