An image processing method and related device
By segmenting the image into multiple tiles in image encoding and using adaptive data and neural networks for feature extraction and compensation, the problem of image quality loss in deep learning image encoding is solved, and higher image reconstruction quality is achieved.
Patent Information
- Application Number
- CN202010754333.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-30
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-07-30
AI Technical Summary
Existing deep learning image encoding methods are difficult to effectively reduce image quality loss during the encoding process.
By segmenting the image into multiple tiles at the encoding end, the adaptive data of each tiles are extracted for pre-processing, and the tiles are featured and quantized by encoding neural network, and finally the encoded representation is obtained through entropy coding. The decoding end reconstructs tiles by entropy decoding and decoding the neural network, and uses adaptive data to compensate to improve image quality.
Compensating multiple reconstructed tiles through multiple adaptive data highlights the local characteristics of each tiles, significantly improving the reconstruction quality of the image and reducing image quality loss during the encoding process.
Smart Images

Figure CN114066914B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and particularly to an image processing method and related devices. Background Art
[0002] Nowadays, multimedia data occupies the vast majority of the traffic on the Internet. The compression of image data plays a crucial role in the storage and efficient transmission of multimedia data. Therefore, image coding is a technology with great practical value.
[0003] The research on image coding has a long history. Researchers have proposed a large number of methods and formulated a variety of international standards, such as image coding standards like JPEG, JPEG2000, WebP, BPG, etc. Although these coding methods are widely used at present, for the continuously increasing amount of image data and the emerging new media types, these traditional methods show some limitations. In recent years, some researchers have started to conduct research on image coding methods based on deep learning. Some researchers have achieved good results. For example, Ballé et al. proposed an end-to-end optimized image coding method, achieving better image coding performance than the current best, and even surpassing the current best traditional coding standard BPG. Deep learning image coding is a lossy image coding technology. The general process of deep learning image coding is as follows: at the encoding end, extract the adaptive data of the image, preprocess the image using the adaptive data, encode the preprocessed image through an encoding neural network to obtain compressed data, and at the decoding end, decode the compressed data to obtain an image similar to the original image.
[0004] Although the above deep learning image coding has made great progress compared with traditional coding methods, how to reduce the loss of image quality during the coding process is a problem that lossy image coding technology has always needed to solve. Summary of the Invention
[0005] This application provides an image processing method and related devices for improving image quality.
[0006] In the first aspect of this application, an image processing method is provided. The method includes:
[0007] The encoding end obtains a first image, then segments the first image to obtain N first image patches, where N is an integer greater than 1. The encoding end obtains N first adaptive data from the N first image patches, and the N first adaptive data correspond to the N first image patches one by one. The encoding end preprocesses the N first image patches through the N first adaptive data. After preprocessing, the encoding end processes the preprocessed N first image patches through an encoding neural network to obtain N groups of first feature maps. The encoding end quantizes and entropy-codes the N groups of first feature maps to obtain N first encoded representations. Among them, by extracting multiple adaptive information, the multiple adaptive information can be used to compensate for multiple reconstructed image patches, thereby highlighting local characteristics and improving the image quality of the second image.
[0008] In an alternative design of the first aspect, if the N first encoded representations are entropy-decoded, N groups of second feature maps can be obtained. If the N groups of second feature maps are processed through a decoding neural network, N first reconstructed image patches can be obtained, and the N first adaptive data are used to compensate for the N first reconstructed image patches. If the N first reconstructed image patches after compensation are combined, a second image can be obtained.
[0009] In an alternative design of the first aspect, the method further includes: the encoding end sends the N first encoded representations, the N first adaptive data, and the correspondence relationship to the decoding end, where the correspondence relationship includes the correspondence relationship between the N first adaptive data and the N first encoded representations.
[0010] In an alternative design of the first aspect, the method further includes: the encoding end quantizes the N first adaptive data to obtain N first adaptive quantization data. The encoding end sends the N first adaptive quantization data to the decoding end, and the N first adaptive quantization data are used to compensate for the N first reconstructed image patches. Among them, N is an integer greater than 1, and the encoding end needs to obtain multiple first adaptive data. Compared with obtaining only one adaptive data from the first image, when obtaining multiple first adaptive data, quantizing the first adaptive data can reduce the data volume of the first adaptive data.
[0011] In an alternative design of the first aspect, the larger N is, the smaller the information entropy of a single first adaptive quantization data is. Among them, the larger N is, the more the number of first image patches, and the more the number of first adaptive quantization data. In this case, by reducing the information entropy of the first adaptive quantization data, the quantization degree of the first adaptive data can be further improved, and the data volume of the first adaptive data can be reduced.
[0012] In an alternative design of the first aspect, the arrangement order of the N first coding representations is the same as the arrangement order of the N first tiles, and the arrangement order of the N first tiles is the arrangement order of the N first tiles in the first image. The corresponding relationship includes the arrangement order of the N first coding representations and the arrangement order of the N first tiles. Compared with only obtaining one piece of adaptive data from the first image, there are multiple pieces of adaptive data and multiple first tiles in this application. Therefore, this application needs to ensure the corresponding relationship between the multiple pieces of adaptive data and the multiple first tiles. By using the arrangement order to ensure the above corresponding relationship, the data volume can be reduced.
[0013] In an alternative design of the first aspect, if the second image is processed by a fusion neural network to obtain a third image, the fusion neural network is used to reduce the difference between the second image and the first image, and the difference includes blocking artifacts. By highlighting the local characteristics of each image block, the performance of each tile is enhanced in this application, but this also easily causes blocking artifacts between tiles. By processing the second image with a fusion neural network, the influence caused by blocking artifacts can be reduced, and the image quality can be improved.
[0014] In an alternative design of the first aspect, the size of each of the N first tiles is the same. Among them, if the size of each first tile is the same, then in the operations of the feature map and the convolutional layer in the encoding neural network, the number of multiplications and additions of each tile participating in the convolutional operation is the same, thereby improving the operation efficiency.
[0015] In an alternative design of the first aspect, when the method is used to segment the first images of different sizes, the size of the first tile is a fixed value. Among them, when the encoding end processes images of different sizes, the image is segmented into tiles of the same size. By fixing the size of the first tile, it is possible to match the convolutional operation unit with the height of the tile, thereby reducing the cost of the convolutional operation unit or improving the usage efficiency of the convolutional operation unit.
[0016] In an alternative design of the first aspect, the pixels of the first tile are a×b, where a and b are obtained according to the target pixels, the target pixels are c×d, equal to an integer, equal to an integer, a and c are the number of pixel points in the width direction, b and d are the number of pixel points in the height direction, the target pixels are obtained according to the target resolution of the terminal device, the terminal device includes a camera component, and the pixels of the image obtained by the camera component under the setting of the target resolution are the target pixels. The first image is obtained by the camera component. The encoding end and / or the decoding end can be the terminal device or not. Among them, under the setting of the target resolution, the image obtained by the encoding end can be just divided into different tiles, thereby avoiding filling useless data and improving the image quality.
[0017] In an alternative design of the first aspect, the target resolution is obtained according to the resolution setting of the imaging component in the imaging application interface. Among them, the resolution of the imaging component obtained by shooting can be set in the imaging application interface. The selected resolution in the setting interface is used as the target resolution, which improves the acquisition efficiency of the target resolution.
[0018] In an alternative design of the first aspect, the target resolution is obtained according to the target image group in the image library obtained by the imaging component. The pixels of the target image group are target pixels. Among the image groups with different pixels, the ratio of the target image group in the image library is the largest. Among them, the image library obtained by the encoding end through the imaging component includes image groups with different pixels. Determining the target pixels through the target image group can ensure that most images can just be divided into different tiles, thereby avoiding filling useless data and improving the image quality.
[0019] In an alternative design of the first aspect, it is characterized in that images with multiple pixels are obtained through the imaging component, and the multiple pixels are e×f, is equal to an integer, is equal to an integer, e includes c, and f includes d. Among them, the terminal device can obtain images with different pixels through the imaging component, and e×f is the pixel set of the images with different pixels. This application stipulates that the images with different pixels obtained by the imaging component can all be just divided into different tiles, avoiding filling useless data, thereby improving the image quality.
[0020] In an alternative design of the first aspect, the multiple pixels are obtained according to the resolution setting of the imaging component in the imaging application interface.
[0021] In an alternative design of the first aspect, the pixels of the first tile are a×b, a is the number of pixel points in the width direction, and b is the number of pixel points in the height direction. The pixels of the first image are r×t. After obtaining the first image and before splitting the first image, the method further includes: if is not equal to an integer, and / or is not equal to an integer, then fill the edge of the first image with the pixel median value so that is equal to an integer, is equal to an integer, and the pixels of the filled first image are r1×t1. Among them, the size of the first tile is fixed, and the encoding end may need to face images with different pixels, that is, some pixel images may not be just divisible. In the case where the image cannot be just divided, filling the edge of the image with the image median value can improve the compatibility of the model while reducing the impact on the image quality. The image median value is the median value of the pixel points.
[0022] In an alternative design of the first aspect, after obtaining the first image and before padding the edge of the first image, the method further includes: If is not an integer, then proportionally enlarge r and t to obtain a first image with pixels of r2×t2, is an integer. If is not an integer, then pad the edge of the first image with the pixel median value. Among them, the number of tiles filled with the pixel median value will affect the image quality. By proportionally enlarging the image, the number of tiles filled with the pixel median value is reduced, improving the image quality.
[0023] In an alternative design of the first aspect, after proportionally enlarging r and t, if is not an integer, then obtain remainder. If the remainder is greater than then only pad the pixel median value on one side of the width direction of the first image. Among them, only padding the pixel median value on one side of the image further reduces the number of tiles filled with the pixel median value while reducing the impact of padding on the image blocks, improving the image quality.
[0024] In an alternative design of the first aspect, if the remainder is less than then pad the pixel median value on both sides of the width direction of the first image, so that the width of the pixel median value padded on each side is where g is the remainder. Among them, the impact of padding on the image blocks is reduced, improving the image quality.
[0025] In an alternative design of the first aspect, the N first tiles include a first target tile, and the range of pixel values of the first target tile is smaller than the range of pixel values of the first image. Before obtaining N first adaptive data from the N first tiles, the method further includes: The encoding end dequantizes the pixel values of the first target tile. The encoding end obtains a first adaptive data from the dequantized first target tile. Dequantizing the pixel values of the first target tile further highlights the local characteristics of the image.
[0026] A second aspect of the present application provides an image processing method, the method includes:
[0027] The decoding end obtains N first encoded representations, N first adaptive data, and a correspondence relationship. The correspondence relationship includes the correspondence between the N first adaptive data and the N first encoded representations. The N first adaptive data correspond one-to-one with the N first encodings, and N is an integer greater than 1. The decoding end performs entropy decoding on the N first encoded representations to obtain N groups of second feature maps. The decoding end processes the N groups of second feature maps through a decoding neural network to obtain N first reconstructed blocks. The decoding end compensates the N first reconstructed blocks with the N first adaptive data. The decoding end combines the compensated N first reconstructed blocks to obtain a second image. Among them, compensating multiple reconstructed blocks with multiple adaptive data highlights the local characteristics of each block, thereby improving the image quality of the second image.
[0028] In an alternative design of the second aspect, the N first encoded representations are obtained by quantizing and entropy encoding N groups of first feature maps. The N groups of first feature maps are obtained by processing N first blocks after preprocessing through an encoding neural network. The N first blocks after preprocessing are obtained by preprocessing the N first blocks with the N first adaptive data. The N first adaptive data are obtained from the N first blocks. The N first blocks are obtained by splitting a first image.
[0029] In an alternative design of the second aspect, the N first adaptive data are N first adaptive quantization data, and the N first adaptive quantization data are obtained by quantizing the N first adaptive data. Specifically, the decoding end compensates the N first reconstructed blocks with the N first adaptive quantization data.
[0030] In an alternative design of the second aspect, the larger N is, the smaller the information entropy of a single first adaptive quantization data is.
[0031] In an alternative design of the second aspect, the arrangement order of the N first encoded representations is the same as the arrangement order of the N first blocks. The arrangement order of the N first blocks is the arrangement order of the N first blocks in the first image. The correspondence relationship includes the arrangement order of the N first encoded representations and the arrangement order of the N first blocks.
[0032] In an alternative design of the second aspect, the method further includes: the decoding end processes the second image through a fusion neural network to obtain a third image. Processing the second image through a fusion neural network to reduce the difference between the second image and the first image. The difference includes block effects.
[0033] In an alternative design of the second aspect, the size of each of the N first blocks is the same.
[0034] In an alternative design of the second aspect, when the method is used to combinatorially generate second images of different sizes, the size of the first tile is a fixed value. In an alternative design of the second aspect, the pixels of the first tile are a×b. a and b are obtained according to the target pixels, and the target pixels are c×d, equal to an integer, equal to an integer, where a and c are the number of pixel points in the width direction, and b and d are the number of pixel points in the height direction. The target pixels are obtained according to the target resolution of the terminal device. The terminal device includes a camera component, and the pixels of the image obtained by the camera component under the setting of the target resolution are the target pixels. The first image is obtained by the camera component.
[0035] In an alternative design of the second aspect, the target resolution is obtained according to the resolution setting of the camera component in the setting interface of the camera application.
[0036] In an alternative design of the second aspect, the target resolution is obtained according to the target image group in the image library obtained by the camera component, and the pixels of the target image group are the target pixels. Among the image groups with different pixels, the ratio of the target image group in the image library is the largest.
[0037] In an alternative design of the second aspect, images with multiple pixels are obtained by the camera component, and the multiple pixels are e×f, equal to an integer, equal to an integer, where e includes c and f includes d.
[0038] In an alternative design of the second aspect, the multiple pixels can be obtained according to the resolution setting of the camera component in the setting interface of the camera application.
[0039] In an alternative design of the second aspect, the pixels of the first tile are a×b, where a is the number of pixel points in the width direction and b is the number of pixel points in the height direction, and the pixels of the first image are r×t. In not equal to an integer, and / or in the case of not being equal to an integer, the edges of the first image are filled with the pixel median, so that equal to an integer, equal to an integer, and the pixels of the filled first image are r1×t1.
[0040] In an alternative design of the second aspect, in the case of not being equal to an integer, r and t are scaled up proportionally by the encoding end to obtain a first image with pixels of r2×t2, equal to an integer.
[0041] In an alternative design of the second aspect, if not equal to an integer, the remainder of is greater than The first image is filled with the pixel median only on one side in the width direction.
[0042] In an alternative design of the second aspect, if the remainder is less than The first image is filled with the pixel median on both sides in the width direction, and the width of the pixel median filled on each side is where g is the remainder.
[0043] In an alternative design of the second aspect, the N first tiles include a first target tile, the range of the pixel values of the first target tile is less than the range of the pixel values of the first image, and at least one of the first adaptive data is obtained from the dequantized first target tile, and the dequantized first target tile is obtained by dequantizing the pixel values of the first target tile.
[0044] A third aspect of the present application provides a model training method, the method comprising:
[0045] Obtain a first image;
[0046] Segment the first image to obtain N first tiles, where N is an integer greater than 1;
[0047] Obtain N first adaptive data from the N first tiles, and the N first adaptive data correspond to the N first tiles one by one;
[0048] Preprocess the N first tiles according to the N first adaptive data;
[0049] Process the preprocessed N first tiles through a first encoding neural network to obtain N groups of first feature maps;
[0050] Quantize and entropy-encode the N groups of first feature maps to obtain N first encoded representations;
[0051] Entropy-decode the N first encoded representations to obtain N groups of second feature maps;
[0052] Process the N groups of second feature maps through a first decoding neural network to obtain N first reconstructed tiles;
[0053] Compensate the N first reconstructed tiles through the N first adaptive data;
[0054] Combine the compensated N first reconstructed tiles to obtain a second image;
[0055] Obtain the distortion loss of the second image relative to the first image;
[0056] The model is jointly trained using a loss function until the image distortion value between the first image and the second image reaches a first preset level. The model includes the first encoding neural network, quantization network, entropy encoding network, entropy decoding network, and the first decoding neural network. Optionally, the model further includes a segmentation network, and the trainable parameters in the segmentation network are the size of the first tile.
[0057] Output a second encoding neural network and a second decoding neural network. The second encoding neural network is the model obtained after the first encoding neural network has undergone iterative training, and the second decoding neural network is the model obtained after the first decoding neural network has undergone iterative training.
[0058] In an alternative design of the third aspect, the method further includes:
[0059] Quantize the N first adaptive data to obtain N first adaptive quantization data, which are used to compensate for the N first reconstructed tiles.
[0060] In an alternative design of the third aspect, the larger the N, the smaller the information entropy of a single first adaptive quantization data.
[0061] In an alternative design of the third aspect, the arrangement order of the N first encoded representations is the same as the arrangement order of the N first tiles, and the arrangement order of the N first tiles is the arrangement order of the N first tiles in the first image.
[0062] In an alternative design of the third aspect, the second image is processed by a fusion neural network to obtain a third image. The fusion neural network is used to reduce the difference between the second image and the first image, and the difference includes blocking artifacts.
[0063] Obtaining the distortion loss of the second image relative to the first image includes:
[0064] Obtaining the distortion loss of the third image relative to the first image;
[0065] The model includes a fusion neural network.
[0066] In an alternative design of the third aspect, each of the N first tiles has the same size.
[0067] In an alternative design of the third aspect, in two iterative trainings, the size of the first image used for training is different, and the size of the first tile is a fixed value.
[0068] In an alternative design of the third aspect, the pixels of the first tile are a×b, where a and b are obtained based on target pixels, and the target pixels are c×d. Equal to an integer. Equal to an integer. Here, a and c are the number of pixel points in the width direction, and b and d are the number of pixel points in the height direction. The target pixels are obtained based on the target resolution of the terminal device. The terminal device includes a camera component, and the pixels of the image obtained by the camera component under the setting of the target resolution are the target pixels. The first image is obtained by the camera component.
[0069] In an alternative design of the third aspect, the target resolution is obtained by setting the resolution of the camera component through the setting interface in the camera application.
[0070] In an alternative design of the third aspect, the target resolution is obtained based on a target image group in the image library obtained by the camera component. The pixels of the target image group are the target pixels, and among the image groups with different pixels, the ratio of the target image group in the image library is the largest.
[0071] In an alternative design of the third aspect, images with multiple pixels are obtained by the camera component, and the multiple pixels are e×f. Equal to an integer. Equal to an integer. Here, e includes c, and f includes d.
[0072] In an alternative design of the third aspect, the multiple pixels are obtained by setting the resolution of the camera component through the setting interface in the camera application.
[0073] In an alternative design of the third aspect, the pixels of the first tile are a×b, where a is the number of pixel points in the width direction and b is the number of pixel points in the height direction. The pixels of the first image are r×t.
[0074] After obtaining the first image and before segmenting the first image, the method further includes:
[0075] If Is not equal to an integer, and / or Is not equal to an integer, then fill the edge of the first image with the pixel median so that Equal to an integer. Equal to an integer. The pixels of the filled first image are r1×t1.
[0076] In an alternative design of the third aspect, after obtaining the first image and before filling the edge of the first image, the method further includes:
[0077] If the is not an integer, then scale r and t proportionally to obtain the first image with pixels of r2×t2, and the is an integer;
[0078] If the is not an integer, and / or is not an integer, then filling the edges of the first image with the pixel median includes:
[0079] If is not an integer, then fill the edges of the first image with the pixel median.
[0080] In an alternative design of the third aspect, after proportionally scaling r and t, if is not an integer, then obtain the remainder of. If the remainder is greater than then fill the pixel median only on one side in the width direction of the first image.
[0081] In an alternative design of the third aspect, if the remainder is less than then fill the pixel median on both sides in the width direction of the first image, so that the width of the pixel median filled on each side is where g is the remainder.
[0082] In an alternative design of the third aspect, the N first tiles include a first target tile, and the range of the pixel values of the first target tile is smaller than the range of the pixel values of the first image;
[0083] Before obtaining N first adaptive data from the N first tiles, the method further includes:
[0084] Dequantize the pixel values of the first target tile;
[0085] Obtaining N first adaptive data from the N first tiles includes:
[0086] Obtain the one first adaptive data from the dequantized first target tile.
[0087] A fourth aspect of the present application provides an encoding device, and the device includes:
[0088] A first acquisition module, configured to acquire a first image;
[0089] A segmentation module, configured to segment the first image to obtain N first tiles, where N is an integer greater than 1;
[0090] A second acquisition module, configured to acquire N first adaptive data from N first tiles, where the N first adaptive data corresponds to the N first tiles one by one;
[0091] A preprocessing module, configured to preprocess the N first tiles according to the N first adaptive data;
[0092] An encoding neural network module, configured to process the preprocessed N first tiles through an encoding neural network to obtain N groups of first feature maps;
[0093] A quantization and entropy encoding module, configured to perform quantization and entropy encoding on the N groups of first feature maps to obtain N first encoded representations.
[0094] In an alternative design of the fourth aspect, the N first encoded representations are used for entropy decoding to obtain N groups of second feature maps, the N groups of second feature maps are used to be processed by a decoding neural network to obtain N first reconstructed tiles, the N first adaptive data are used to compensate the N first reconstructed tiles, and the compensated N first reconstructed tiles are used to form a second image.
[0095] In an alternative design of the fourth aspect, the apparatus further includes:
[0096] A sending module, configured to send the N first encoded representations, the N first adaptive data, and the correspondence relationship to a decoding end, where the correspondence relationship includes the correspondence relationship between the N first adaptive data and the N first encoded representations.
[0097] In an alternative design of the fourth aspect, the apparatus further includes:
[0098] A quantization module, configured to quantize the N first adaptive data to obtain N first adaptive quantization data, and the N first adaptive quantization data are used to compensate the N first reconstructed tiles;
[0099] The sending module is specifically configured to send the N first adaptive quantization data to a decoding end.
[0100] In an alternative design of the fourth aspect, the larger N is, the smaller the information entropy of a single first adaptive quantization data is.
[0101] In an alternative design of the fourth aspect, the arrangement order of the N first encoded representations is the same as the arrangement order of the N first tiles, the arrangement order of the N first tiles is the arrangement order of the N first tiles in the first image, and the correspondence relationship includes the arrangement order of the N first encoded representations and the arrangement order of the N first tiles.
[0102] In an alternative design of the fourth aspect, the second image is used to be processed by a fusion neural network to obtain a third image, and the fusion neural network is used to reduce the difference between the second image and the first image, where the difference includes blocking artifacts.
[0103] In an alternative design of the fourth aspect, each of the N first tiles has the same size.
[0104] In an alternative design of the fourth aspect, when the device is used to process the first images of different sizes, the size of the first tile is a fixed value. In an alternative design of the fourth aspect, the pixels of the first tile are a×b, where a and b are obtained according to the target pixels. The target pixels are c×d, equal to an integer, equal to an integer. a and c are the number of pixel points in the width direction, and b and d are the number of pixel points in the height direction. The target pixels are obtained according to the target resolution of the terminal device. The terminal device includes a camera component. The pixels of the image obtained by the camera component under the setting of the target resolution are the target pixels. The first image is obtained by the camera component.
[0105] In an alternative design of the fourth aspect, the target resolution is obtained by setting the resolution of the camera component according to the setting interface in the camera application.
[0106] In an alternative design of the fourth aspect, the target resolution is obtained according to the target image group in the image library obtained by the camera component. The pixels of the target image group are the target pixels. Among the image groups with different pixels, the ratio of the target image group in the image library is the largest.
[0107] In an alternative design of the fourth aspect, images with multiple pixels are obtained by the camera component. The multiple pixels are e×f, equal to an integer, equal to an integer, where e includes c and f includes d.
[0108] In an alternative design of the fourth aspect, the multiple pixels are obtained by setting the resolution of the camera component according to the setting interface in the camera application.
[0109] In an alternative design of the fourth aspect, the pixels of the first tile are a×b, where a is the number of pixel points in the width direction and b is the number of pixel points in the height direction. The pixels of the first image are r×t. The device further includes:
[0110] A filling module, used for if is not equal to an integer, and / or is not equal to an integer, then fill the edge of the first image with the pixel median value so that is equal to an integer, is equal to an integer, and the pixels of the filled first image are r1×t1.
[0111] In an alternative design of the fourth aspect, the device further includes:
[0112] An amplification module, configured to, if the is not an integer, amplify the r and the t in equal proportion to obtain the first image with pixels of r2×t2, and the is an integer;
[0113] The filling module is specifically configured to, if is not an integer, fill the edge of the first image with the median pixel value.
[0114] In an alternative design of the fourth aspect, the second obtaining unit is further configured to, if is not an integer, obtain the remainder of;
[0115] The filling module is specifically configured to, if the remainder is greater than then only fill the median pixel value on one side in the width direction of the first image.
[0116] In an alternative design of the fourth aspect, the filling module is specifically configured to, if the remainder is less than then fill the median pixel value on both sides in the width direction of the first image, so that the width of the median pixel value filled on each side is where g is the remainder.
[0117] In an alternative design of the fourth aspect, the N first tiles include a first target tile, and the range of the pixel values of the first target tile is smaller than the range of the pixel values of the first image. The apparatus further includes:
[0118] An inverse quantization module, configured to inverse quantize the pixel values of the first target tile;
[0119] The second obtaining module is specifically configured to obtain a first adaptive data from the inversely quantized first target tile.
[0120] A fifth aspect of the present application provides a decoding apparatus, the decoding apparatus includes:
[0121] An obtaining module, configured to obtain N first encoded representations, N first adaptive data, and a correspondence relationship, where the correspondence relationship includes the correspondence relationship between the N first adaptive data and the N first encoded representations, the N first adaptive data and the N first encodings correspond one by one, and N is an integer greater than 1;
[0122] An entropy decoding module, configured to perform entropy decoding on the N first encoded representations to obtain N groups of second feature maps;
[0123] A decoding neural network module, configured to process the N groups of second feature maps to obtain N first reconstructed tiles;
[0124] A compensation module, configured to compensate N first reconstructed patches with N first adaptive data;
[0125] A combination module, configured to combine the compensated N first reconstructed patches to obtain a second image.
[0126] In an alternative design of the fifth aspect, the N first encoded representations are obtained by quantizing and entropy encoding N groups of first feature maps, the N groups of first feature maps are obtained by processing N first patches after preprocessing through an encoding neural network, the N first patches after preprocessing are obtained by preprocessing N first patches with the N first adaptive data, the N first adaptive data are obtained from the N first patches, and the N first patches are obtained by splitting a first image.
[0127] In an alternative design of the fifth aspect, the N first adaptive data are N first adaptive quantization data, and the N first adaptive quantization data are obtained by quantizing the N first adaptive data;
[0128] The compensation module is specifically configured to compensate the N first reconstructed patches with the N first adaptive quantization data.
[0129] In an alternative design of the fifth aspect, the larger N is, the smaller the information entropy of a single first adaptive quantization data is.
[0130] In an alternative design of the fifth aspect, the arrangement order of the N first encoded representations is the same as the arrangement order of the N first patches, and the arrangement order of the N first patches is the arrangement order of the N first patches in the first image. The corresponding relationship includes the arrangement order of the N first encoded representations and the arrangement order of the N first patches.
[0131] In an alternative design of the fifth aspect, the apparatus further includes:
[0132] A fusion neural network module, configured to process the second image to obtain a third image, so as to reduce the difference between the second image and the first image, and the difference includes blocking artifacts.
[0133] In an alternative design of the fifth aspect, each of the N first patches has the same size.
[0134] In an alternative design of the fifth aspect, when the apparatus is used to combine and generate second images of different sizes, the size of the first patch is a fixed value. In an alternative design of the fifth aspect, the pixels of the first patch are a×b, and a and b are obtained according to the target pixels, and the target pixels are c×d, equal to an integer, Equal to an integer, where a and c are the number of pixels in the width direction, and b and d are the number of pixels in the height direction. The target pixel is obtained according to the target resolution of the terminal device. The terminal device includes a camera component. The pixels of the image obtained by the camera component under the setting of the target resolution are the target pixels, and the first image is obtained by the camera component.
[0135] In an alternative design of the fifth aspect, the target resolution is obtained according to the resolution setting of the camera component in the setting interface of the camera application.
[0136] In an alternative design of the fifth aspect, the target resolution is obtained according to the target image group in the image library obtained by the camera component. The pixels of the target image group are the target pixels. Among the image groups with different pixels, the ratio of the target image group in the image library is the largest.
[0137] In an alternative design of the fifth aspect, images with multiple pixels are obtained by the camera component. The multiple pixels are e×f, Equal to an integer, Equal to an integer, e includes c, and f includes d.
[0138] In an alternative design of the fifth aspect, the multiple pixels are obtained according to the resolution setting of the camera component in the setting interface of the camera application.
[0139] In an alternative design of the fifth aspect, the pixels of the first tile are a×b, where a is the number of pixels in the width direction and b is the number of pixels in the height direction. The pixels of the first image are r×t. In Not equal to an integer, and / or In the case of not being equal to an integer, the edges of the first image are filled with the pixel median value, so that Equal to an integer, Equal to an integer, and the pixels of the filled first image are r1×t1.
[0140] In an alternative design of the fifth aspect, in In the case of not being equal to an integer, r and t are enlarged proportionally by the encoding end, and a first image with pixels of r2×t2 is obtained, Equal to an integer.
[0141] In an alternative design of the fifth aspect, if Not equal to an integer, The remainder of is greater than The first image is only filled with the pixel median value on one side in the width direction.
[0142] In an alternative design of the fifth aspect, if the remainder is less than The first image is filled with the pixel median value on both sides in the width direction, and the width of the pixel median value filled on each side is where g is the remainder.
[0143] In an alternative design of the fifth aspect, the N first tiles include a first target tile, the range of pixel values of the first target tile is smaller than the range of pixel values of the first image, and at least one of the first adaptive data is obtained from the dequantized first target tile, and the dequantized first target tile is obtained by dequantizing the pixel values of the first target tile.
[0144] The sixth aspect of the present application provides a training device, the device includes:
[0145] A first acquisition module, configured to acquire a first image;
[0146] A segmentation module, configured to segment the first image to obtain N first tiles, where N is an integer greater than 1;
[0147] A second acquisition module, configured to acquire N first adaptive data from the N first tiles, and the N first adaptive data correspond to the N first tiles one by one;
[0148] A preprocessing module, configured to preprocess the N first tiles according to the N first adaptive data;
[0149] A first encoding neural network module, configured to process the N preprocessed first tiles to obtain N groups of first feature maps;
[0150] A quantization and entropy encoding module, configured to perform quantization and entropy encoding on the N groups of first feature maps to obtain N first encoded representations;
[0151] An entropy decoding module, configured to perform entropy decoding on the N first encoded representations to obtain N groups of second feature maps;
[0152] A first decoding neural network module, configured to process the N groups of second feature maps to obtain N first reconstructed tiles;
[0153] A compensation module, configured to compensate the N first reconstructed tiles through the N first adaptive data;
[0154] A combination module, configured to combine the N compensated first reconstructed tiles to obtain a second image;
[0155] A third acquisition module, configured to acquire the distortion loss of the second image relative to the first image;
[0156] A training module, configured to jointly train a model using a loss function until the image distortion value between the first image and the second image reaches a first preset level. The model includes the first encoding neural network, a quantization network, an entropy encoding network, an entropy decoding network, and the first decoding neural network. Optionally, the model further includes a segmentation network, and the trainable parameter in the segmentation network is the size of the first tile. Optionally, the model further includes a segmentation network, and the trainable parameter in the segmentation network is the size of the first tile;
[0157] An output module, configured to output a second encoding neural network and a second decoding neural network. The second encoding neural network is a model obtained after the first encoding neural network has undergone iterative training, and the second decoding neural network is a model obtained after the first decoding neural network has undergone iterative training.
[0158] In an alternative design of the sixth aspect, the apparatus further includes:
[0159] A quantization module, configured to quantize the N first adaptive data to obtain N first adaptive quantization data, and the N first adaptive quantization data are used to compensate the N first reconstructed tiles;
[0160] In an alternative design of the sixth aspect, the larger the N, the smaller the information entropy of a single first adaptive quantization data.
[0161] In an alternative design of the sixth aspect, the arrangement order of the N first encoded representations is the same as the arrangement order of the N first tiles, and the arrangement order of the N first tiles is the arrangement order of the N first tiles in the first image.
[0162] In an alternative design of the sixth aspect, the second image is processed by a fusion neural network to obtain a third image. The fusion neural network is used to reduce the difference between the second image and the first image, and the difference includes blocking artifacts;
[0163] The third acquisition module is specifically configured to acquire the distortion loss of the third image relative to the first image;
[0164] The model includes a fusion neural network.
[0165] In an alternative design of the sixth aspect, the size of each of the N first tiles is the same.
[0166] In an alternative design of the sixth aspect, in two iterative trainings, the size of the first image used for training is different, and the size of the first tile is a fixed value.
[0167] In an alternative design of the sixth aspect, the pixels of the first tile are a×b, where a and b are obtained based on the target pixels, and the target pixels are c×d. Equal to an integer. Equal to an integer. a and c are the number of pixel points in the width direction, and b and d are the number of pixel points in the height direction. The target pixels are obtained based on the target resolution of the terminal device. The terminal device includes a camera component. The pixels of the image obtained by the camera component under the setting of the target resolution are the target pixels. The first image is obtained by the camera component.
[0168] In an alternative design of the sixth aspect, the target resolution is obtained by setting the resolution of the camera component according to the setting interface in the camera application.
[0169] In an alternative design of the sixth aspect, the target resolution is obtained based on the target image group in the image library obtained by the camera component. The pixels of the target image group are the target pixels. Among the image groups with different pixels, the ratio of the target image group in the image library is the largest.
[0170] In an alternative design of the sixth aspect, images with multiple pixels are obtained by the camera component. The multiple pixels are e×f. Equal to an integer. Equal to an integer. e includes c, and f includes d.
[0171] In an alternative design of the sixth aspect, the multiple pixels are obtained by setting the resolution of the camera component according to the setting interface in the camera application.
[0172] In an alternative design of the sixth aspect, the pixels of the first tile are a×b, where a is the number of pixel points in the width direction and b is the number of pixel points in the height direction. The pixels of the first image are r×t.
[0173] The apparatus further includes:
[0174] A filling module, configured to, if Is not equal to an integer, and / or Is not equal to an integer, then fill the edge of the first image with the pixel median value, so that Is equal to an integer. Is equal to an integer. The pixels of the filled first image are r1×t1.
[0175] In an alternative design of the sixth aspect, the apparatus further includes:
[0176] An amplification module, configured to, if the If it is not equal to an integer, then scale the r and the t proportionally to obtain the first image with pixels of r2×t2, the is equal to an integer;
[0177] The filling module is specifically configured to if is not equal to an integer, then fill the edge of the first image with the pixel median value.
[0178] In an alternative design of the sixth aspect, the second acquisition module is further configured to, after scaling the r and the t proportionally, if is not equal to an integer, then obtain the remainder of;
[0179] The filling module is specifically configured to if the remainder is greater than then fill the pixel median value only on one side in the width direction of the first image.
[0180] In an alternative design of the sixth aspect, if the remainder is less than then fill the pixel median value on both sides in the width direction of the first image, so that the width of the pixel median value filled on each side is wherein, the g is the remainder.
[0181] In an alternative design of the sixth aspect, the N first tiles include a first target tile, and the range of the pixel values of the first target tile is smaller than the range of the pixel values of the first image;
[0182] The device further includes:
[0183] An inverse quantization module, configured to inverse quantize the pixel values of the first target tile;
[0184] The second acquisition module is specifically configured to obtain a first adaptive data from the inverse quantized first target tile.
[0185] A seventh aspect of the present application provides an encoding device, which may include a memory, a processor, and a bus system. Among them, the memory is used to store a program, and the processor is used to execute the program in the memory, including the following steps:
[0186] Obtain a first image;
[0187] Segment the first image to obtain N first tiles, where N is an integer greater than 1;
[0188] Obtain N first adaptive data from the N first tiles, and the N first adaptive data correspond to the N first tiles one by one;
[0189] Preprocess the N first tiles through the N first adaptive data;
[0190] Process the preprocessed N first tiles through an encoding neural network to obtain N groups of first feature maps;
[0191] Quantize and entropy-encode the N groups of first feature maps to obtain N first encoded representations.
[0192] In an alternative design of the seventh aspect, the encoding device is a virtual reality (VR) device, a mobile phone, a tablet, a laptop computer, a server, or a smart wearable device.
[0193] In the seventh aspect of the present application, the processor can also be used to execute the steps performed by the encoding end in each possible implementation manner of the first aspect. Specifically, reference can be made to the first aspect, and details are not elaborated here.
[0194] The eighth aspect of the present application provides a decoding device, which may include a memory, a processor, and a bus system. Among them, the memory is used to store programs, and the processor is used to execute the programs in the memory, including the following steps:
[0195] Obtain N first encoded representations, N first adaptive data, and a correspondence relationship. The correspondence relationship includes the correspondence between the N first adaptive data and the N first encoded representations. The N first adaptive data correspond one-to-one with the N first encodings, and N is an integer greater than 1;
[0196] Perform entropy decoding on the N first encoded representations to obtain N groups of second feature maps;
[0197] Process the N groups of second feature maps through a decoding neural network to obtain N first reconstructed tiles;
[0198] Compensate the N first reconstructed tiles with the N first adaptive data;
[0199] Combine the compensated N first reconstructed tiles to obtain a second image.
[0200] In an alternative design of the eighth aspect, the decoding device is a virtual reality (VR) device, a mobile phone, a tablet, a laptop computer, a server, or a smart wearable device.
[0201] In the eighth aspect of the present application, the processor can also be used to execute the steps performed by the decoding end in each possible implementation manner of the second aspect. Specifically, reference can be made to the second aspect, and details are not elaborated here.
[0202] The ninth aspect of the present application provides a training device, which may include a memory, a processor, and a bus system. Among them, the memory is used to store programs, and the processor is used to execute the programs in the memory, including the following steps:
[0203] Obtain a first image;
[0204] Divide the first image to obtain N first image patches, where N is an integer greater than 1;
[0205] Obtain N first adaptive data from the N first image patches, where the N first adaptive data correspond one-to-one to the N first image patches;
[0206] Preprocess the N first image patches according to the N first adaptive data;
[0207] Process the preprocessed N first image patches through a first encoding neural network to obtain N groups of first feature maps;
[0208] Quantize and entropy-encode the N groups of first feature maps to obtain N first encoded representations;
[0209] Entropy-decode the N first encoded representations to obtain N groups of second feature maps;
[0210] Process the N groups of second feature maps through a first decoding neural network to obtain N first reconstructed patches;
[0211] Compensate the N first reconstructed patches with the N first adaptive data;
[0212] Combine the compensated N first reconstructed patches to obtain a second image;
[0213] Obtain the distortion loss of the second image relative to the first image;
[0214] Jointly train the first encoding neural network, quantization network, entropy encoding network, entropy decoding network, and the first decoding neural network using a loss function until the image distortion value between the first image and the second image reaches a first preset level;
[0215] Output a second encoding neural network and a second decoding neural network, where the second encoding neural network is the model obtained after the first encoding neural network has undergone iterative training, and the second decoding neural network is the model obtained after the first decoding neural network has undergone iterative training.
[0216] In the ninth aspect of the present application, the processor can also be used to execute the steps performed by the decoding end in each possible implementation manner of the third aspect. Specifically, reference can be made to the third aspect, and details are not elaborated here.
[0217] In a tenth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When it runs on a computer, it causes the computer to execute the image processing method described in any one of the first aspect to the third aspect above.
[0218] In an eleventh aspect, an embodiment of the present application provides a computer program which, when running on a computer, causes the computer to execute the image processing method described in any one of the first to third aspects above.
[0219] In a twelfth aspect, the present application provides a chip system. The chip system includes a processor for supporting an execution device or a training device to implement the functions involved in the above aspects, for example, sending or processing the data and / or information involved in the above method. In a possible design, the chip system further includes a memory for storing the necessary program instructions and data of the execution device or the training device. The chip system may be composed of chips or may include chips and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0220] Figure 1 It is a schematic structural diagram of a framework of an artificial intelligence entity;
[0221] Figure 2a It is a schematic diagram of an application scenario of an embodiment of the present application;
[0222] Figure 2b It is another schematic diagram of an application scenario of an embodiment of the present application;
[0223] Figure 2c It is another schematic diagram of an application scenario of an embodiment of the present application;
[0224] Figure 3a It is a schematic flowchart of an image processing method provided by an embodiment of the present application;
[0225] Figure 3b It is another schematic flowchart of an image processing method provided by an embodiment of the present application;
[0226] Figure 4 It is a schematic diagram of splitting and combining images in an embodiment of the present application;
[0227] Figure 5 It is a schematic diagram of an image encoding processing process based on CNN in an embodiment of the present application;
[0228] Figure 6 It is a schematic diagram of an image decoding process based on CNN in an embodiment of the present application;
[0229] Figure 7 It is a schematic diagram of a setting interface for setting the resolution of a camera of a terminal device in an embodiment of the present application;
[0230] Figure 8 It is a schematic flowchart of image filling in an embodiment of the present application;
[0231] Figure 9Another schematic diagram of the image filling process in the embodiments of the present application;
[0232] Figure 10 A comparison schematic diagram of the image compression quality in the embodiments of the present application;
[0233] Figure 11 A system architecture diagram of the image processing system provided in the embodiments of the present application;
[0234] Figure 12 A schematic diagram of a process of the model training method provided in the embodiments of the present application;
[0235] Figure 13 A schematic diagram of a training process provided in the embodiments of the present application;
[0236] Figure 14 A schematic diagram of the structure of an encoding device provided in the embodiments of the present application;
[0237] Figure 15 A schematic diagram of the structure of a decoding device provided in the embodiments of the present application;
[0238] Figure 16 A schematic diagram of the structure of a training device provided in the embodiments of the present application;
[0239] Figure 17 A schematic diagram of the structure of an execution device provided in the embodiments of the present application;
[0240] Figure 18 A schematic diagram of the structure of a training device provided in the embodiments of the present application;
[0241] Figure 19 A schematic diagram of the structure of a chip provided in the embodiments of the present application. Detailed implementation manners
[0242] The embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention. The terms used in the embodiments of the present invention are only used to explain the specific embodiments of the present invention, rather than to limit the present invention.
[0243] The embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0244] In the description and claims of this application and the above-mentioned drawings, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing embodiments of this application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.
[0245] First, the overall workflow of the artificial intelligence system will be described. Please refer to Figure 1 , Figure 1 FIG. is a schematic structural diagram of an artificial intelligence main framework. The above artificial intelligence theme framework will be elaborated from two dimensions: "intelligent information chain" (horizontal axis) and "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.
[0246] (1) Infrastructure
[0247] The infrastructure provides computing power support for the artificial intelligence system, realizes communication with the external world, and is supported through the basic platform. Communicate with the outside through sensors; the computing power is provided by intelligent chips (such as hardware acceleration chips like CPU, NPU, GPU, ASIC, FPGA, etc.); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for calculation.
[0248] (2) Data
[0249] The data on the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves Internet of Things data of traditional devices, including business data of existing systems and perception data such as force, displacement, liquid level, temperature, humidity, etc.
[0250] (3) Data processing
[0251] Data processing generally includes data training, machine learning, deep learning, search, inference, decision-making, etc.
[0252] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.
[0253] Inference refers to the process of simulating the intelligent reasoning mode of humans in a computer or intelligent system, based on an inference control strategy, and using formalized information for machine thinking and problem-solving. The typical functions are search and matching.
[0254] Decision-making refers to the process of making decisions after intelligent information is inferred, and usually provides functions such as classification, ranking, prediction, etc.
[0255] (4) General capabilities
[0256] After the data is processed through the above-mentioned data processing, some general capabilities can be further formed based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0257] (5) Intelligent products and industry applications
[0258] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields, which are the encapsulation of the overall artificial intelligence solution, productize intelligent information decision-making, and realize practical applications. Its application fields mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, safe cities, etc.
[0259] This application can be applied to the field of image processing in the field of artificial intelligence. Multiple application scenarios implemented in products will be introduced below.
[0260] First, for image compression and decompression processes applied to terminal devices, both the encoding end and the decoding end are terminal devices.
[0261] The image compression method provided by the embodiments of this application can be applied to the image compression process in terminal devices. Specifically, it can be applied to the photo albums, video surveillance, etc. on terminal devices. Specifically, reference can be made to Figure 2a , Figure 2a which is a schematic diagram of an application scenario of the embodiments of this application. As shown in Figure 2aAs shown in [FIGURE], the terminal device can obtain the image to be compressed, where the image to be compressed can be a photo taken by a camera component or a frame captured from a video. The camera component is generally a camera. The terminal device divides and extracts the obtained image through a central processing unit (CPU) to obtain multiple tiles. After obtaining multiple tiles, the terminal device can perform feature extraction on the obtained multiple tiles through an artificial intelligence (AI) encoding neural network (referred to as the encoding neural network) in a neural-network processing unit (NPU), transform the tile data into output features with lower redundancy, and generate probability estimates for each feature point in the output features. The central processing unit (CPU) performs entropy encoding on the extracted output features through the probability estimates of each point in the output features, reduces the encoding redundancy of the output features, further reduces the data transmission volume during the tile compression process, and saves the encoded data in the corresponding storage location in the form of a data file. When the user needs to obtain the file saved in the above storage location, the CPU can obtain and load the saved file at the corresponding storage location, and obtain the decoded feature map through entropy decoding. The AI decoding neural network (referred to as the decoding neural network) in the NPU reconstructs the feature map to obtain multiple reconstructed tiles. After obtaining multiple tiles, the terminal device combines the multiple reconstructed tiles through the CPU to obtain a reconstructed image.
[0262] Specifically, in this scenario, the terminal device can save the encoded data on the cloud device. When the user needs to obtain the above encoded data, the encoded data can be obtained from the cloud device.
[0263] Second, for image compression and decompression applied to the cloud, both the encoding end and the decoding end are cloud devices.
[0264] The image compression method provided by the embodiments of this application can be applied to the image compression process on the cloud. Specifically, it can be applied to functions such as cloud albums on cloud devices. The cloud device can be a cloud server. Specifically, reference can be made to Figure 2b , Figure 2b which is another application scenario schematic diagram of the embodiments of this application. As shown in Figure 2bAs shown, the terminal device can obtain the image to be compressed, where the image to be compressed can be a photo taken by a camera component or a frame intercepted from a video. The terminal device can perform entropy encoding on the image to be compressed through the CPU to obtain encoded data. In addition to using entropy encoding, any lossless compression method in the prior art can also be used. The terminal device can transmit the encoded data to the cloud device, and the cloud device can perform corresponding entropy decoding on the received encoded data to obtain the image to be compressed. The terminal device extracts the obtained image through the CPU to obtain multiple tiles. After obtaining multiple tiles, the server can perform feature extraction on the obtained multiple tiles through the encoding neural network in the graphics processing unit (GPU), transform the tile data into output features with lower redundancy, and generate probability estimates for each point in the output features. The CPU performs entropy encoding on the extracted output features through the probability estimates for each point in the output features, reduces the encoding redundancy of the output features, further reduces the data transmission volume during the tile compression process, and saves the encoded data in the corresponding storage location in the form of a data file. When the user needs to obtain the file saved in the above storage location, the CPU can obtain and load the above saved file in the corresponding storage location, and based on entropy decoding, obtain the decoded feature map. The decoding neural network in the NPU reconstructs the feature map to obtain multiple reconstructed tiles. After obtaining multiple tiles, the cloud device combines the multiple reconstructed tiles through the CPU to obtain a reconstructed image. The cloud device can perform entropy encoding on the image to be compressed through the CPU to obtain encoded data, and the encoding method can also be any other lossless compression method in the prior art. The cloud device can transmit the encoded data to the terminal device, and the terminal device can perform corresponding entropy decoding on the received encoded data to obtain the decoded image.
[0265] III. Applied to the image decompression of the terminal device, the image compression process of the cloud device, the encoding end is the cloud device, and the decoding end is the terminal device.
[0266] The image compression method provided in the embodiments of the present application can be applied to the image compression of the terminal device and the image decompression process of the cloud device. Specifically, it can be applied to functions such as cloud albums on the cloud device, and the cloud device can be a cloud server. Specifically, reference can be made to Figure 2c , Figure 2c which is another application scenario schematic diagram of the embodiments of the present application. As shown in Figure 2cAs shown, the terminal device can obtain the image to be compressed, where the image to be compressed can be a photo taken by the imaging component or a frame intercepted from a video. The terminal device can perform entropy encoding on the image to be compressed through the CPU to obtain encoded data. In addition to using entropy encoding, any lossless compression method in the prior art can also be used. The terminal device can transmit the encoded data to the cloud device, and the cloud device can perform corresponding entropy decoding on the received encoded data to obtain the image to be compressed. The terminal device extracts the obtained image through the CPU to obtain multiple tiles. After obtaining multiple tiles, the server can perform feature extraction on the obtained multiple tiles through the encoding neural network in the GPU, transform the tile data into output features with lower redundancy, and generate probability estimates for each point in the output features. The CPU performs entropy encoding on the extracted output features through the probability estimates for each point in the output features, reduces the encoding redundancy of the output features, further reduces the data transmission volume during the tile compression process, and saves the encoded data in the form of a data file at the corresponding storage location. When the terminal device needs to obtain the above image, the terminal device receives the encoded data sent by the cloud device and obtains the decoded feature map based on entropy decoding. The terminal device reconstructs the feature map through the decoding neural network in the NPU to obtain multiple reconstructed tiles. After obtaining multiple tiles, the terminal device combines the multiple reconstructed tiles through the CPU to obtain the reconstructed image.
[0267] Since the embodiments of the present application involve the application of a large number of neural networks, for the convenience of understanding, the relevant terms and concepts of the neural networks that the embodiments of the present application may involve are introduced below.
[0268] (1) Neural network
[0269] A neural network can be composed of neural units. A neural unit can refer to an arithmetic unit with xs and intercept 1 as inputs, and the output of this arithmetic unit can be:
[0270]
[0271] Among them, s = 1, 2, ……, n, where n is a natural number greater than 1, Ws is the weight of Xs, and b is the bias of the neuron. f is the activation function of the neuron, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer, and the activation function can be the sigmoid function. A neural network is a network formed by connecting multiple such single neurons together, that is, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neurons.
[0272] (2) Deep neural network
[0273] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with multiple hidden layers. Dividing the DNN according to the positions of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i + 1-th layer.
[0274] Although the DNN looks very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression: Among them, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also known as the coefficient), and α() is the activation function. Each layer is just a simple operation on the input vector to obtain the output vector Due to the large number of layers in the DNN, the number of coefficients W and offset vectors is also relatively large. The definitions of these parameters in the DNN are as follows: Taking the coefficient W as an example: Suppose in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscript corresponds to the index 2 of the output third layer and the index 4 of the input second layer.
[0275] In summary, the coefficient from the k-th neuron in the L - 1-th layer to the j-th neuron in the L-th layer is defined as
[0276] Note that there is no W parameter in the input layer. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically speaking, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is a process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).
[0277] (3) Convolutional Neural Network
[0278] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers, and this feature extractor can be regarded as a filter. A convolutional layer refers to the neuron layer in a convolutional neural network that performs convolution processing on the input signal. In the convolutional layer of a convolutional neural network, a neuron can be connected to only some adjacent layer neurons. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weight here is the convolutional kernel. Sharing weights can be understood as a way of extracting image information that is independent of position. The convolutional kernel can be initialized in the form of a matrix of random size, and during the training process of the convolutional neural network, the convolutional kernel can learn to obtain reasonable weights. Additionally, the direct benefit of sharing weights is to reduce the connections between layers of the convolutional neural network while also reducing the risk of overfitting.
[0279] (4) Loss Function
[0280] During the process of training a deep neural network, since it is desired that the output of the deep neural network is as close as possible to the value that is truly wanted to be predicted, the weight vectors of each layer of the neural network can be updated by comparing the predicted value of the current network with the truly desired target value and then according to the difference between the two. (Of course, there is usually an initialization process before the first update, that is, parameters are pre-configured for each layer in the deep neural network). For example, if the predicted value of the network is high, the weight vector is adjusted to make it predict lower, and continuously adjusted until the deep neural network can predict the truly desired target value or a value very close to the truly desired target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function. They are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the deep neural network becomes a process of minimizing this loss as much as possible.
[0281] (5) Backpropagation algorithm
[0282] The neural network can use the backpropagation (BP) algorithm to correct the magnitudes of the parameters in the initial neural network model during the training process, so that the reconstruction error loss of the neural network model becomes smaller and smaller. Specifically, the forward propagation of the input signal until the output will generate an error loss, and the initial neural network model parameters are updated by backpropagating the error loss information, so as to converge the error loss. The backpropagation algorithm is a backpropagation movement dominated by the error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix.
[0283] In the embodiments of the present application, not only the operation of image segmentation is performed, but also a step of adaptively extracting data for multiple tiles is added between the image segmentation and the encoding neural network. The adaptive data can be the mean, the mean square error, etc. The extracted adaptive data is used for preprocessing the tiles. The encoding neural network extracts features from the preprocessed multiple tiles. The adaptive data is also used for compensating the reconstructed tiles in addition to preprocessing the tiles. For the convenience of description, the following will take the adaptive data as the mean and the application scenario as the third application scenario described above as an example to describe in detail the image processing method in the embodiments of the present application. In the third application scenario, the cloud device is the encoding end and the terminal device is the decoding end.
[0284] As an example, the terminal device can be a mobile phone, a tablet computer, a laptop computer, a smart wearable device, etc. As another example, the terminal device can be a virtual reality (VR) device. As another example, the embodiments of the present application can also be applied to intelligent monitoring. A camera can be configured in the intelligent monitoring, and the intelligent monitoring can obtain pictures to be compressed through the camera. It should be understood that the embodiments of the present application can also be applied to other scenarios that require image compression, and other scenarios will not be listed one by one here.
[0285] Please refer to Figure 3a , Figure 3a which is a schematic flowchart of an image processing method provided by an embodiment of the present application.
[0286] In step 301, the terminal device obtains a first image.
[0287] The terminal device can obtain a first image, where the first image can be a photo taken by a camera component or a frame captured from a video being shot. The terminal device includes the camera component, and the camera component is generally a camera. The first image can also be an image obtained by the terminal device from the network or an image obtained by the terminal device using a screenshot tool.
[0288] Specifically, regarding the image processing method in the embodiments of the present application, reference can also be made to Figure 3b , Figure 3b which is another schematic flowchart of the image processing method provided by an embodiment of the present application. Figure 3b which illustrates the entire process of outputting a third image from the first image.
[0289] In step 302, the terminal device sends the first image to the cloud device.
[0290] Before the terminal device sends the first image to the cloud device, the terminal device can perform lossless encoding on the first image to obtain encoded data. The encoding method can be entropy encoding or other lossless compression methods.
[0291] In step 303, the cloud device divides the first image to obtain N first image blocks.
[0292] The cloud device can receive the first image sent by the terminal device. If the first image has been losslessly encoded by the terminal device, the cloud device also needs to perform lossless decoding on it. The cloud device divides the first image to obtain N first image blocks, where N is an integer greater than 1. Figure 4 which is a schematic diagram of dividing and combining images in an embodiment of the present application. As shown in Figure 4As shown, the first image 401 is divided into 12 first tiles. Among them, when the size of the first image is determined, the size of the first tile determines the value of N. The N described here as 12 is only an example. In the subsequent description, the size of the first tile will be described in detail.
[0293] Optionally, each of the N first tiles has the same size.
[0294] In step 304, the cloud device obtains M first means from the N first tiles.
[0295] If the first image is a three-channel image, the first tile includes data of three channels, and the number M of the first means obtained by the cloud device is equal to 3N. If the first image is a grayscale image, that is, a one-channel image, the first tile includes data of one channel, and the number M of the first means obtained by the cloud device is equal to N. Since the processing method for each channel is similar, for the convenience of description, in the embodiments of the present application, only one channel is taken as an example for description. The mean refers to the mean of the pixel values of all pixel points in the first tile.
[0296] In step 305, the cloud device preprocesses the N first tiles with the N first means.
[0297] The preprocessing may be to subtract the mean from the pixel value of each pixel point in the first tile to obtain the N preprocessed first tiles.
[0298] In step 306, the N preprocessed first tiles are processed by an encoding neural network to obtain N groups of first feature maps.
[0299] In the embodiments of the present application, optionally, the encoding neural network is a CNN, and the terminal device may perform feature extraction on the N preprocessed first tiles based on the CNN to obtain N groups of first feature maps. Each group of first feature maps corresponds to a first tile, and each group of first feature maps includes at least one feature map. Hereinafter, the first feature map may also be referred to as a channel feature map image, and each semantic channel corresponds to a first feature map.
[0300] In the embodiments of the present application, refer to Figure 5 , Figure 5 is a schematic diagram of an image encoding processing process based on a CNN in the embodiments of the present application. Figure 5 It shows the first tile 501, the CNN 502, the channel feature map 503, and a group of first feature maps 504, where the CNN 502 may include multiple CNN layers.
[0301] For example, CNN502 can multiply the upper-left 3×3 pixels of the input data (the first tile) by weights and map them to the neurons at the upper-left end of the first feature map. The weights to be multiplied will also be 3×3. Thereafter, in the same process, CNN502 scans the input data (the first tile) one by one from left to right and from top to bottom, and multiplies by the weights to map the neurons of the feature map. Here, the 3×3 weights used are called filters or filter kernels. That is to say, the process of applying the filter in CNN502 is the process of performing a convolution operation using the filter kernel, and the extracted result is called a "channel feature map". Among them, the channel feature map can also be called a multi-channel feature map image. The term "multi-channel feature map image" can refer to a set of feature map images corresponding to multiple channels. According to an embodiment, the channel feature map can be generated by CNN502, and CNN502 is also called the "feature extraction layer" or "convolution layer" of the CNN. The layer of the CNN can define the mapping from the output to the input. The mapping defined by the layer is performed as one or more filter kernels (convolution kernels) to be applied to the input data to generate the channel feature map to be output to the next layer. The input data can be the first tile or the channel feature map output by CNN502.
[0302] Referring to Figure 5 , during forward execution, CNN502 receives the first tile 501 as input and generates the channel feature map 503. Additionally, during forward execution, the next layer CNN receives the channel feature map 503 as input and generates the channel feature map 503 as output. Then, each subsequent layer will receive the channel feature map generated in the previous layer and generate the channel feature map of the next layer as input. Finally, a set of first feature maps 504 generated in the (X1)th layer is received. Among them, X1 is an integer greater than 1, that is, the channel feature map of each layer mentioned above may all be a set of first feature maps 504.
[0303] The cloud device repeats the above operations for each first tile, and N sets of first feature maps can be obtained.
[0304] Optionally, as the level of CNN502 increases, the length and width of each feature map in the multi-channel feature map image gradually decrease, and the number of semantic channels of the multi-channel feature map image gradually increases, so as to achieve data compression of the first tile.
[0305] Meanwhile, in addition to the operation of applying the convolution kernel that maps the input feature map to the output feature map, other processing operations can also be performed. Examples of other processing operations can include but are not limited to the application of activation functions, pooling, resampling, etc.
[0306] For example, as Figure 3bAs shown, optionally, after each layer of convolutional kernels, a GDN (generalized divisive normalization) activation function is further included. The expression form of GDN is as follows:
[0307]
[0308] Among them, u represents the j-th channel of the output of the i-th convolutional layer. v represents the output result of the corresponding activation function. β and γ are respectively the trainable parameters of the activation function, used to enhance the non-linear expression ability of the neural network.
[0309] It should be noted that the above is only one implementation manner for feature extraction of the first tile. In practical applications, the specific implementation manner of feature extraction is not limited.
[0310] In the embodiments of the present application, through the above method, the first tile is transformed into another space (at least one first feature map) through a CNN convolutional neural network. Optionally, the number of first feature maps is 192, that is, the number of semantic channels is 192, and each semantic channel corresponds to a first feature map. In the embodiments of the present application, at least one first feature map may be in the form of a three-dimensional tensor, and its size may be 192×w×h, where w×h is the width and length of the matrix corresponding to the first feature map of a single channel.
[0311] In step 307, the N groups of first feature maps are quantized and entropy encoded to obtain N first encoded representations.
[0312] In the embodiments of the present application, after obtaining N groups of first feature maps by processing the preprocessed N first tiles through an encoding neural network, the processed N groups of first feature maps can be quantized and entropy encoded to obtain N first encoded representations.
[0313] In the embodiments of the present application, the N groups of first feature maps are converted to quantization centers according to specified rules for subsequent entropy encoding. The quantization operation can convert the N groups of first feature maps from floating-point numbers to bitstreams (for example, bitstreams of specific-bit integers such as 8-bit integers or 4-bit integers). In some embodiments, the quantization operation can be performed on the N groups of first feature maps by rounding, but is not limited thereto.
[0314] In the embodiments of the present application, the probability estimates of each point in the output feature can be obtained by using an entropy estimation network, and the output feature is entropy encoded using the probability estimate to obtain a binary bitstream. It should be noted that the entropy encoding process mentioned in the present application can adopt existing entropy encoding technologies, and the present application will not elaborate on this.
[0315] In step 308, the cloud device sends N first encoded representations, N first means, and corresponding relationships to the terminal device.
[0316] In step 302 above, the terminal device stores the first image in the cloud device. If the terminal device needs to obtain the first image, it can send a request to the cloud device. After receiving the request sent by the terminal device, the cloud device sends N first encoded representations, N first means, and a correspondence relationship to the terminal device. The correspondence relationship refers to the correspondence relationship between the N first encoded representations and the N first means.
[0317] Optionally, the arrangement order of the N first encoded representations is the same as the arrangement order of the N first tiles, and the arrangement order of the N first tiles is the arrangement order of the N first tiles in the first image. The correspondence relationship includes the arrangement order of the N first encoded representations and the arrangement order of the N first tiles.
[0318] Optionally, before the cloud device sends the N first means to the terminal device, the cloud device quantizes the N first means to obtain N first quantized means. For example, the pixel value of each pixel point in the first tile is represented by 8 bits, and the first mean of the first tile is a 32-bit floating-point number. The cloud device quantizes the first mean, and the number of bits of the obtained first quantized mean is less than 32. The smaller the number of bits of the first quantized mean, the smaller the information entropy of the first quantized mean. Further, the number of bits of the first quantized mean is equal to the number of bits of the pixel value of each pixel point in the first tile, that is, when the pixel value of each pixel point in the first tile is represented by 8 bits, the first quantized mean is also represented by 8 bits.
[0319] Optionally, based on the cloud device quantizing the N first means, the larger the value of N, the smaller the information entropy of a single first quantized mean. Information entropy is used to describe the quantization degree of the N first means by the cloud device. If the information entropy of a single first quantized mean is smaller, it means that the quantization degree of the N first means by the cloud device is higher. When processing the first image with the same number of pixels, the smaller the pixels of each first image, the larger N is. The larger N is, the larger the data volume of the N first means is. For example, assume that the pixels of the first image are 640×480, the pixels of the first tile are 320×480, N is 2, and each first quantized mean is represented by 8 bits, then the data volume of the N first quantized means is 2×8 bits. Assume that the pixels of the first tile are 1×1, N is 640×480, and each first quantized mean is represented by 8 bits, then the data volume of the N first quantized means is 640×480×8 bits. The data volume of the first image is also 640×480×8 bits. It can be seen that the larger the value of N, the larger the data volume of the N first means. When N is equal to the pixel size of the first image, even the data volume of the quantized first means reaches the data volume of the first image. Therefore, in the embodiments of the present application, the larger the value of N, the smaller the information entropy of a single first quantized mean.
[0320] In step 309, the terminal device performs entropy decoding on N first encoded representations to obtain N sets of second feature maps.
[0321] After the terminal device receives the N first encoded representations sent by the cloud device, the terminal device performs entropy decoding on the N first encoded representations to obtain N sets of second feature maps.
[0322] In step 310, the terminal device processes the N sets of second feature maps through a decoding neural network to obtain N first reconstructed blocks.
[0323] In an embodiment of the present application, optionally, the decoding neural network is a CNN. The terminal device can reconstruct the N sets of second feature maps based on the CNN to obtain N sets of first blocks. Each set of second feature maps corresponds to a first block, and each set of second feature maps includes at least one feature map. Hereinafter, the second feature map can also be referred to as a reconstructed feature map image, and each semantic channel corresponds to a second feature map.
[0324] In an embodiment of the present application, referring to Figure 6 , Figure 6 is a schematic diagram of an image decoding process based on a CNN in an embodiment of the present application, Figure 6 showing a set of second feature maps 601, a transposed CNN 602, a reconstructed feature map 603, and a first reconstructed block 604. The transposed CNN 602 may include multiple transposed CNN layers.
[0325] For example, the transposed CNN 602 can multiply the upper left pixel of the input data (a set of second feature maps 601) by a weight and map it to the neuron at the upper left end of the reconstructed feature map 603. The weight to be multiplied will be 3×3. Thereafter, in the same process, the transposed CNN 602 scans the input data (a set of second feature maps 601) one by one from left to right and from top to bottom, and multiplies by the weight to map the neurons of the feature map. After passing through the transposed CNN 602 with a weight of 3×3, the length and width of the obtained reconstructed feature map 603 become 3 times that of the second feature map. Here, the 3×3 weight used is called an inverse filter or an inverse filter kernel. That is, the process of applying the inverse filter in the transposed CNN 602 is a process of performing deconvolution operation using the inverse filter kernel, and the extracted result is called a "reconstructed feature map". According to the embodiment, the reconstructed feature map can be generated by the transposed CNN 602, and the transposed CNN 602 is also called the transposed convolutional layer of the CNN. The layer of the CNN can define the mapping from the output to the input. The mapping defined by the layer is used as one or more inverse filter kernels (transposed convolutional layers) to be applied to the input data to generate the reconstructed feature map to be output to the next layer. The input data can be a set of second feature maps or the reconstructed feature image of a specific layer.
[0326] Referring toFigure 6 , the transposed CNN602 receives a set of second feature maps 601 and generates a reconstructed feature map 603 as output. Additionally, the next layer of the transposed CNN receives the reconstructed feature map 603 as input and generates the reconstructed feature map of the next layer as output. Then, each subsequent transposed CNN layer will receive the reconstructed feature map generated in the previous layer and generate the next reconstructed feature map as output. Finally, the first reconstructed patch 604 generated in the (X2) layer is received, where X2 is an integer greater than 1, that is, the reconstructed feature map of each layer above may serve as the first reconstructed patch 604. The cloud device repeats the above operations for each set of second feature maps to obtain N first reconstructed patches.
[0327] Optionally, as the number of layers of the transposed CNN increases, the length and width of each feature map in the reconstructed feature map gradually increase until the size of the first patch before inputting into the encoding neural network is restored. The number of semantic channels of the reconstructed feature map gradually decreases until the semantic channels of the first patch before inputting into the encoding neural network are restored. When the first patch is a single-channel image, the semantic channel of the first reconstructed patch 604 is 1, and when the first patch is a three-channel image, the semantic channel of the first reconstructed patch 604 is 3. Through the above reconstruction, data decoding of the first patch is achieved.
[0328] Meanwhile, in addition to applying the operation of the transposed convolutional kernel that maps each set of second feature maps to the reconstructed feature map, other processing operations can also be performed. Examples of other processing operations can include but are not limited to applications such as activation functions, pooling, resampling, etc.
[0329] For example, optionally, as Figure 3b shown, after each layer of the transposed convolutional kernel in the decoding neural network, it also includes inverse generalized divisive normalization (iGDN). iGDN is an approximate inverse form of the GDN activation function in the encoding segment. The expression form of iGDN is:
[0330]
[0331] where v represents the j-th channel of the output of the i-th convolutional layer. u represents the output result of the corresponding activation function, and β and γ are the trainable parameters of the activation function, used to enhance the non-linear expression ability of the neural network.
[0332] In step 311, the terminal device compensates for the N first reconstructed patches with N first means.
[0333] In the above step 308, the cloud device sends N first means and corresponding relationships to the terminal device. After obtaining N first reconstructed patches through the decoding neural network, the terminal device compensates the N first reconstructed patches using the corresponding relationships and the N first means. Compensation means adding the pixel value of each pixel point in the first reconstructed patch with the first mean to obtain the compensated first reconstructed patch. After the terminal device repeats compensating the N first reconstructed patches, it can obtain the compensated N first reconstructed patches.
[0334] Optionally, when the terminal device receives N first quantization means from the cloud device, the terminal device compensates the N first reconstructed patches using the N first quantization means. It should be determined that when the terminal device compensates the N first reconstructed patches using the N first quantization means, the cloud device will also preprocess the N first patches using the N first quantization means.
[0335] In step 312, the terminal device combines the compensated N first reconstructed patches to obtain a second image.
[0336] Please refer to Figure 4 , combination is the inverse process of segmentation. Replace the N first patches with the N first reconstructed patches, and then combine the N first reconstructed patches.
[0337] In step 313, the terminal device processes the second image through a fusion neural network to obtain a third image.
[0338] In the embodiment of the present application, by highlighting the local characteristics of each first patch, the performance of each patch is enhanced, but it is also easy to cause block effects between the first reconstructed patches. Block effect means that discontinuity will occur at the boundary between the first reconstructed patches, forming defects in the reconstructed image. By processing the second image through a fusion neural network, the influence caused by block effects can be reduced and the image quality can be improved.
[0339] Optionally, the fusion neural network is a CNN. Please refer to Figure 5 and Figure 6 , in terms of the structure of the CNN, the fusion neural network can be a combination of an encoding neural network and a decoding neural network. By using Figure 5 's output 504 as Figure 6 's input 601, using the second image as Figure 5 's input 501, Figure 6The output is the third image. By using a fusion neural network, the blocking effect in the second image can be eliminated. It should be noted that this is simply an example to illustrate the framework of the fusion neural network. In practical applications, the framework of the fusion neural network, such as the number of layers of the CNN, the number of layers of the transposed CNN, the size of the matrix of each CNN layer, etc., may have no relation to the encoding neural network and the decoding neural network.
[0340] Optionally, as Figure 3b shown, after the convolutional kernel of the fusion neural network, there is also a rectified linear unit layer ReLU, and ReLU is used to correct negative numbers in the feature map output by the convolutional kernel to zero.
[0341] The above has described the process of processing an image using the image processing method in the embodiments of the present application. Optionally, the image processing method in the embodiments of the present application can process images of different sizes, such as the fourth image, and the pixels of the fourth image are different from those of the first image. The process of processing the fourth image using the image processing method in the embodiments of the present application is similar to the process of processing the first image described above, and will not be elaborated here specifically. In particular, when using the image processing method in the embodiments of the present application to process the fourth image, the cloud device segments the fourth image, and M second tiles can be obtained. Among them, the size of the second tile is the same as that of the first tile. When the sizes of the first tile and the second tile are the same, the same encoding neural network and decoding neural network are used to process the first tile and the second tile, and the number of convolution operations and the number of data participating in each convolution operation in the processing process are the same. In this case, a corresponding convolution operation unit can be designed according to the above number of convolution operations and / or the number of data participating in each convolution operation, so that the convolution operation unit matches the processing process. Since the number of convolution operations and the number of data participating in each convolution operation in the processing process are determined by the size of the first tile and the CNN, it can also be considered that the convolution operation unit matches the first tile, or the convolution operation unit matches the encoding neural network and / or the decoding neural network. The higher the matching degree between the convolution operation unit and the first tile, the smaller the number of idle multipliers and adders in the convolution operation unit in the processing process, that is, the higher the usage efficiency of the convolution operation unit.
[0342] The image processing method in the embodiments of the present application has been described above. In the above process, the size of the first tile not only affects the size of N, but also affects whether the image is exactly divided into integer tiles. Generally, the determinants of the size of the first tile are as follows. On the first hand, it is the influence of the model on the size of the first tile. The model includes an encoding neural network and a decoding neural network, and may also include a fusion neural network. The influence of the model on the size of the first tile generally includes the influence on the size of the first tile during model training and the influence on the size of the first tile when using the model. The influence on the size of the first tile during model training includes training the model with tiles of different sizes to determine in which interval or at which value of the tile size the model has a faster convergence speed, or the model outputs high-quality images, or the model has high compression performance. Different models are for different-sized image blocks, and when using different models, the performance of different models may vary in different scenarios, that is, the generalization problem of the model. The influence on the size of the first tile when using the model includes this generalization problem. On the second hand, it is the influence of whether the image is exactly divided into integer tiles on the size of the first tile. If the image cannot be exactly divided into integer tiles, there will be some incomplete tiles, which affect the model's reconstruction of the tile and reduce the image quality. To reduce the influence of the second aspect on the image quality, some related technical solutions are proposed below.
[0343] In step 302 above, the terminal device sends a first image to the cloud device. In this scenario, the encoding neural network in the cloud device can specifically serve this terminal device or this type of terminal device. If the terminal device includes a camera component, such as a camera, it is desired that the first image obtained by the terminal device through the camera can be divided into integer tiles by the cloud device. Assume that the pixels of the first tile are a×b, where a is the number of pixel points in the width direction and b is the number of pixel points in the height direction. a and b are obtained according to the target pixels, and the target pixels are c×d, is equal to an integer, is equal to an integer. The target pixels are obtained according to the target resolution of the terminal device, and the target resolution is the default resolution of the shooting component of the terminal device or the resolution set by the terminal device. The pixels of the image obtained by the shooting component of the terminal device under the setting of the target resolution are the target pixels, and the first image is obtained according to the shooting component. In particular, when the terminal device is an encoding end, such as the first application scenario described above, it is more meaningful to determine the size of the first tile through the target resolution. Because the target resolution indicates the pixels of the image that the terminal device may obtain in the future, that is, the pixels of the image that the encoding end will use the image processing method in the embodiments of the present application to process in the future. Therefore, during model training, the training can be carried out for the target pixels.
[0344] Optionally, the target resolution is obtained according to the resolution setting of the imaging component in the imaging application. Among them, the resolution obtained by the imaging component during imaging can be set in the setting interface of the imaging application. The selected resolution in the setting interface is used as the target resolution. Please refer to Figure 7 , Figure 7 which is a schematic diagram of the setting interface for setting the resolution of the camera of the terminal device in the embodiment of the present application. In the schematic diagram of the setting interface, the option 701 with a resolution of [4:3] 10MP is selected. Although the specific value of the target resolution is not specifically described for this option here. However, according to the first image obtained by shooting, it can be known that the pixels of the first image are 2736×3648, that is, the target pixels are 2736×3648. By determining the target pixels, the size of the first tile is determined, so that is equal to an integer, is equal to an integer.
[0345] Optionally, the target resolution is obtained according to the target image group in the image library obtained by the imaging component. The pixels of the target image group are the target pixels, and among the image groups with different pixels, the ratio of the target image group in the image library is the largest. Among them, the image library obtained by the encoding end through the imaging component includes image groups with different pixels. For example, as Figure 7 shown, the terminal corresponding to this schematic diagram of the setting interface can obtain images with 4 kinds of pixels through the camera. In the image library of the camera of the terminal device, determining which kind of pixel image has the largest proportion can ensure that the image of this pixel can be just divided into integer blocks.
[0346] Optionally, images with multiple pixels are obtained through the imaging component, and the multiple pixels are e×f, is equal to an integer, is equal to an integer, e includes c, and f includes d. Among them, the terminal device can obtain images with different pixels through the imaging component, and e×f is the pixel set of the images with different pixels. For example, as Figure 7 shown, the terminal corresponding to this schematic diagram of the setting interface can obtain images with 4 kinds of pixels through the camera. If is equal to an integer, is equal to an integer, it means that the images of these 4 kinds of pixels can all be divided into integer blocks.
[0347] Optionally, e×f also includes the pixels obtained by the terminal device through screenshot.
[0348] Optionally, the multiple pixels are obtained by setting the resolution of the imaging component in the setting interface of the imaging application.
[0349] The above describes the solution of trying to ensure that the first image is divided into integer blocks. However, in practical applications, there are always images that cannot be divided into integer blocks. In this case, to improve the compatibility of the model, it is necessary to pad the first image, and the following is the relevant description.
[0350] Optionally, the pixels of the first tile are a×b, where a is the number of pixel points in the width direction and b is the number of pixel points in the height direction, and the pixels of the first image are r×t. After obtaining the first image and before dividing the first image, the method further includes: if is not an integer, and / or is not an integer, then pad the edges of the first image with the pixel median value, so that is an integer, is an integer, and the pixels of the padded first image are r1×t1. The pixel median value refers to the median value of the pixel values of a single pixel point of the first image. For example, when an 8-bit representation is used for a pixel point of the first image, the pixel median value is 128. By padding the edges of the image with the image median value, the compatibility of the model can be improved while reducing the impact on the image quality. The image median value is the median value of the pixel points.
[0351] Optionally, before padding the edges of the first image, it further includes: if the is not an integer, then scale r and t proportionally to obtain the first image with pixels r2×t2, and the is an integer. Among them, the number of tiles filled with the pixel median value will affect the image quality. By proportionally scaling the image, the number of tiles filled with the image median value is reduced, improving the image quality. As Figure 8 shown, Figure 8 is a schematic flowchart of image padding in an embodiment of the present application. In Figure 8 8a, the pixels of the first tile are a×b, and the pixels of the first image are r×t. After proportionally scaling r and t, as Figure 8 shown in 8b, the pixels of the first image are r2×t2. Before magnification, as Figure 8 shown in 8a, the number of tiles to be filled is 6. After magnification, as Figure 8 shown in 8b, the number of tiles to be filled is 4. Therefore, the number of tiles to be filled is reduced. In particular, if is not an integer, scale r and t proportionally so that then the number of tiles to be filled is reduced to 2.
[0352] Optionally, after proportionally scaling r and t, if is not an integer, then obtain the remainder of . If the remainder is greater than Then, only the pixel median is filled on one side in the width direction of the first image. Among them, only filling the pixel median on one side of the image can further reduce the number of tiles filled with the image median while reducing the impact of filling on the image block, thereby improving the image quality. As Figure 8 shown in 8b, the remainder g is greater than As Figure 8 shown in 8c, the pixel median is filled on one side of the first image.
[0353] Optionally, if the remainder is less than then the pixel median is filled on both sides in the width direction of the first image, and the width of the pixel median filled on each side is where g is the remainder. Among them, the impact of filling on the image block is reduced, and the image quality is improved. As Figure 9 shown, Figure 9 is another schematic flowchart of image filling in the embodiment of the present application. If the remainder g is less than then the image median is filled on both sides of the first image, and the width of the filled image median is
[0354] In this application, the first image is segmented to obtain N first tiles, the respective means are obtained from different first tiles, and then the first reconstructed tiles are compensated using the means to achieve the purpose of highlighting the local characteristics of the first image. In particular, the N first tiles include a first target tile, and the range of pixel values of the first target tile is smaller than the range of pixel values of the first image. Before obtaining N first adaptive data from the N first tiles, the method further includes: the cloud device dequantizing the pixel values of the first target tile. The cloud device obtains N first adaptive data from the dequantized first target tile. By dequantizing the pixel values of the first target tile, the local characteristics of the first image are further highlighted. Any first tile can be understood as the local characteristics of the first image. By highlighting the local characteristics of the first image, the reconstruction quality of the image can be improved, that is, the compression quality of the image can be improved. As Figure 10 shown, Figure 10This is a comparison schematic diagram of the image compression quality in the embodiments of the present application. The abscissa represents the number of bits per pixel (bit-per-pixel, BPP), which is used to measure the bit rate. The ordinate represents the peak signal-to-noise ratio (PSNR), which is used to measure the quality. The compression algorithms compared with the image processing method in the embodiments of the present application include different implementations of JPEG2000, HEVC (high efficiency video coding), and VVC (versatile video coding) standards. For JPEG2000, the reference software OpenJPEG is used to represent its compression performance. At the same time, the implementation integrated in Matlab is used as a supplement to the compression performance of JPEG2000. For HEVC, the reference software HM-16.15 is used to reflect the rate-distortion (RD) performance. The performance of the VVC standard is represented by the VVC standard reference software VTM-6.2. It should be noted that in the encoding configuration of VTM-6.2, the input image bit depth and the internal calculation bit depth are set to 8 to be compatible with the format of the input image, and the test image is encoded using the all intra (AI) configuration. The rate-distortion performance of various compression algorithms is as Figure 10 shown. The rate-distortion performance curve of OpenJPEG is 1001, the rate-distortion performance curve of the Matlab implementation of JPEG2000 is 1002, the performance curve of the 420 image format compression of the reference software HM-16.15 is 1003, the performance curve of the convolutional neural network image compression algorithm without block division is 1004, the performance curve of the present invention is 1005, and the performance curve of the 420 image format compression of the reference software VTM-6.2 is 1006.
[0355] The image processing method in the embodiments of the present application has been described above. Next, the image processing system in the embodiments of the present application will be described.
[0356] Please refer to Figure 11 , Figure 11 which is a system architecture diagram of the image processing system provided in the embodiments of the present application. In Figure 11 it, the image processing system 200 includes an execution device 210, a training device 220, a database 230, a client device 240, and a data storage system 250. The execution device 210 includes a calculation module 211.
[0357] Among them, a first image set is stored in the database 230. Optionally, a fourth image set is also included in the database 230. The training device 220 generates a target model / rule 201 for processing the first image and / or the fourth image, and iteratively trains the target model / rule 201 using the first image and / or the fourth image in the database to obtain a mature target model / rule 201. In the embodiments of the present application, the target model / rule 201 includes an encoding neural network and a decoding neural network. Optionally, the target model / rule 201 further includes a fusion neural network.
[0358] The encoding neural network and the decoding neural network obtained by the training device 220 can be applied to different systems or devices, such as mobile phones, tablets, laptops, VR devices, monitoring systems, and so on. Among them, the execution device 210 can call data, code, etc. in the data storage system 250, and can also store data, instructions, etc. in the data storage system 250. The data storage system 250 can be placed in the execution device 210, or the data storage system 250 can be an external memory relative to the execution device 210.
[0359] The calculation module 211 receives the first image sent by the client device 240, segments the first image to obtain N first tiles, extracts N first adaptive data from the N first tiles, preprocesses the N first tiles using the N first adaptive data, and then extracts features from the preprocessed N first tiles through the encoding neural network to obtain N groups of first feature maps. The obtained N groups of first feature maps are quantized and entropy encoded to obtain N encoding tables, where N is an integer greater than 1.
[0360] The calculation module 211 can also perform entropy decoding on the N encoding representations to obtain N groups of second feature maps, and then process the N groups of second feature groups through the decoding neural network to obtain N first reconstructed tiles. After obtaining the N first reconstructed tiles, the N first reconstructed tiles are compensated using the N first adaptive data. The calculation module 211 combines the N first reconstructed tiles to obtain a second image. Optionally, when the target model / rule 201 further includes a fusion neural network, the calculation module 211 can also use the fusion neural network to process the second image to obtain a third image. Among them, the fusion neural network is used to reduce the difference between the second image and the first image, and the difference includes blocking artifacts.
[0361] In some embodiments of the present application, please refer to Figure 11, the execution device 210 and the terminal device 240 can be separate and independent devices. The execution device 210 is configured with an I / O interface 212 to interact with the terminal device 240. The "user" can input a first image to the I / O interface 212 through the terminal device 240, and the execution device 210 returns a second image to the terminal device 240 through the I / O interface 212 for the user. In addition, the relationship between the terminal device 240 and the execution device 210 can be described by the relationship between the terminal device and the encoding end and the decoding end. The encoding end is a device using an encoding neural network, and the decoding end is a device using a decoding neural network. The encoding end and the decoding end can be the same device or independent devices. The terminal device is similar to the terminal device in the above image processing method, and the terminal device can be the encoding end and / or the decoding end. To facilitate understanding of the relationship between the terminal device 240 and the execution device 210, reference can be made to the relevant description in the foregoing Figure 2a - Figure 2c description.
[0362] It should be noted that Figure 11 is only a schematic diagram of the architecture of the image processing system provided by the embodiments of the present invention. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in some other embodiments of the present application, the execution device 210 can be configured in the terminal device 240. As an example, when the terminal device is a mobile phone or a tablet, the execution device 210 can be a module for performing array image processing in the main processor (Host CPU) of the mobile phone or the tablet, or the execution device 210 can also be a graphics processing unit (GPU) or a neural network processor (NPU) in the mobile phone or the tablet. The GPU or NPU is mounted on the main processor as a coprocessor, and tasks are allocated by the main processor.
[0363] Combined with the above description, the specific implementation process of the training stage of the image processing method provided by the embodiments of the present application will be described below.
[0364] Specifically, please refer to Figure 12 , Figure 12 which is a schematic flowchart of a model training method provided by the embodiments of the present application. The model training method provided by the embodiments of the present application may include:
[0365] In step 1201, the training device acquires a first image.
[0366] In step 1202, the training device segments the first image to obtain N first image patches, where N is an integer greater than 1.
[0367] In step 1203, the training device acquires N first adaptive data from the N first image patches, and the N first adaptive data correspond to the N first image patches one by one.
[0368] In step 1204, the training device preprocesses the N first tiles according to the N first adaptive data.
[0369] In step 1205, the training device processes the N preprocessed first tiles through a first encoding neural network to obtain N groups of first feature maps.
[0370] In step 1206, the training device quantizes and entropy-encodes the N groups of first feature maps to obtain N first encoded representations.
[0371] In step 1207, the training device entropy-decodes the N first encoded representations to obtain N groups of second feature maps.
[0372] In step 1208, the training device processes the N groups of second feature maps through a first decoding neural network to obtain N first reconstructed tiles.
[0373] In step 1209, the training device compensates the N first reconstructed tiles with the N first adaptive data.
[0374] In step 1210, the training device combines the N compensated first reconstructed tiles to obtain a second image.
[0375] In step 1211, the training device obtains the distortion loss of the second image relative to the first image.
[0376] In step 1212, the training device jointly trains the model using a loss function until the image distortion value between the first image and the second image reaches a first preset level. The model includes the first encoding neural network, quantization network, entropy encoding network, entropy decoding network, and the first decoding neural network.
[0377] Please refer to Figure 13 , Figure 13 which is a schematic diagram of a training process provided by an embodiment of the present application. The loss function of the model in the embodiment is:
[0378] loss = l d + P × l r
[0379] In the above loss function, l d represents the information entropy of the first encoded representation. P × l r is used to represent the distortion metric between the first image and the second image, and l r represents the distortion loss between the first image and the second image. P represents the balance factor between the two loss functions, which is used to characterize the relative relationship between the first encoded representation and the quality of the reconstructed image.
[0380] Optionally, to obtain a suitable block size, the training process includes, in multiple iterative trainings, dividing the first image into tiles of different sizes, i.e., with different values of N. Comparing the loss functions obtained from multiple iterations to optimize the size of the first tiles.
[0381] In step 1213, the training device outputs a second encoding neural network and a second decoding neural network. The second encoding neural network is a model obtained after the first encoding neural network has undergone iterative training, and the second decoding neural network is a model obtained after the first decoding neural network has undergone iterative training.
[0382] The specific descriptions of steps 1201 to 1211 can be referred to the descriptions in the above image processing method.
[0383] Optionally, the method further includes:
[0384] The training device quantizes the N first adaptive data to obtain N first adaptive quantization data, which are used to compensate the N first reconstructed tiles.
[0385] Optionally, the larger the N, the smaller the information entropy of a single first adaptive quantization data.
[0386] Optionally, the arrangement order of the N first encoding representations is the same as the arrangement order of the N first tiles, and the arrangement order of the N first tiles is the arrangement order of the N first tiles in the first image.
[0387] Optionally, the training device processes the second image through a fusion neural network to obtain a third image to reduce the difference between the second image and the first image, where the difference includes block artifacts;
[0388] The training device is specifically configured to obtain the distortion loss of the third image relative to the first image;
[0389] The model includes a fusion neural network.
[0390] Optionally, each of the N first tiles has the same size.
[0391] Optionally, in two iterative trainings, the size of the first image for training is different, and the size of the first tiles is a fixed value.
[0392] Optionally, the pixels of the first tile are a×b, where a and b are obtained according to the target pixels, and the target pixels are c×d, equal to an integer, Equal to an integer, where a and c are the number of pixel points in the width direction, b and d are the number of pixel points in the height direction, the target pixel is obtained according to the target resolution of the terminal device, the terminal device includes a camera component, and the pixels of the image obtained by the camera component under the setting of the target resolution are the target pixels. The first image is obtained by the camera component.
[0393] Optionally, the target resolution is obtained by setting the resolution of the camera component according to the setting interface in the camera application.
[0394] Optionally, the target resolution is obtained from the target image group in the image library obtained by the camera component. The pixels of the target image group are the target pixels. Among the image groups with different pixels, the ratio of the target image group in the image library is the largest.
[0395] Optionally, images with multiple pixels are obtained by the camera component. The multiple pixels are e×f, Equal to an integer, Equal to an integer, where e includes c and f includes d.
[0396] Optionally, the multiple pixels are obtained by setting the resolution of the camera component according to the setting interface in the camera application.
[0397] Optionally, the pixels of the first tile are a×b, where a is the number of pixel points in the width direction and b is the number of pixel points in the height direction. The pixels of the first image are r×t;
[0398] After obtaining the first image and before splitting the first image, the method further includes:
[0399] If Is not equal to an integer, and / or Is not equal to an integer, then fill the edge of the first image with the pixel median value so that Is equal to an integer, Is equal to an integer, and the pixels of the filled first image are r1×t1.
[0400] Optionally, after obtaining the first image and before filling the edge of the first image, the method further includes:
[0401] If the Is not equal to an integer, then magnify r and t proportionally to obtain the first image with pixels r2×t2, and the Is equal to an integer;
[0402] The if Is not equal to an integer, and / or If it is not equal to an integer, filling the edge of the first image with the median pixel value includes:
[0403] If is not equal to an integer, fill the edge of the first image with the median pixel value.
[0404] Optionally, after scaling r and t by a ratio, if is not equal to an integer, obtain the remainder. If the remainder is greater than the training device fills the median pixel value only on one side of the width direction of the first image.
[0405] Optionally, if the remainder is less than then fill the median pixel value on both sides of the width direction of the first image, so that the width of the median pixel value filled on each side is where g is the remainder.
[0406] Optionally, the N first tiles include a first target tile, and the range of pixel values of the first target tile is smaller than the range of pixel values of the first image;
[0407] Before obtaining N first adaptive data from the N first tiles, the method further includes:
[0408] The training device dequantizes the pixel values of the first target tile;
[0409] The training device is specifically configured to obtain a first adaptive data from the dequantized first target tile.
[0410] Based on the corresponding embodiments of Figures 1 to 13 in order to better implement the above solutions of the embodiments of the present application, the following also provides related devices for implementing the above solutions. Specifically, refer to Figure 14 , Figure 14 FIG. 1400 is a schematic structural diagram of an encoding device 1400 provided by an embodiment of the present application. The encoding device 1400 corresponds to an encoding end. The encoding device 1400 may be a terminal device or a cloud device. The encoding device 1400 includes:
[0411] A first obtaining module 1401, configured to obtain a first image;
[0412] A splitting module 1402, configured to split the first image to obtain N first tiles, where N is an integer greater than 1;
[0413] A second obtaining module 1403, configured to obtain N first adaptive data from the N first tiles, and the N first adaptive data correspond to the N first tiles one by one;
[0414] A preprocessing module 1404, configured to preprocess N first tiles according to N first adaptive data;
[0415] An encoding neural network module 1405, configured to process the N preprocessed first tiles through an encoding neural network to obtain N groups of first feature maps;
[0416] A quantization and entropy encoding module 1406, configured to perform quantization and entropy encoding on the N groups of first feature maps to obtain N first encoded representations. Optionally, the encoding device is further configured to perform all or part of the operations performed by the cloud device in the corresponding foregoing Figure 3a embodiment.
[0417] The encoding device in the embodiments of the present application has been described above. On the basis of the corresponding Figures 1 to 13 embodiment, in order to better implement the foregoing solution of the embodiments of the present application, the decoding device in the embodiments of the present application is further described below. Specifically, refer to Figure 15 , Figure 15 which is a schematic structural diagram of a decoding device 1500 provided by an embodiment of the present application. The decoding device 1500 corresponds to a decoding end. The decoding device 1500 may be a terminal device or a cloud device. The decoding device 1500 includes:
[0418] An acquisition module 1501, configured to acquire N first encoded representations, N first adaptive data, and a correspondence relationship. The correspondence relationship includes the correspondence relationship between the N first adaptive data and the N first encoded representations. The N first adaptive data and the N first encodings are in one-to-one correspondence, and N is an integer greater than 1;
[0419] An entropy decoding module 1502, configured to perform entropy decoding on the N first encoded representations to obtain N groups of second feature maps;
[0420] A decoding neural network module 1503, configured to process the N groups of second feature maps to obtain N first reconstructed tiles;
[0421] A compensation module 1504, configured to compensate the N first reconstructed tiles through the N first adaptive data;
[0422] A combination module 1505, configured to combine the N first reconstructed tiles after compensation to obtain a second image.
[0423] Optionally, the decoding device is further configured to perform all or part of the operations performed by the terminal device in the corresponding Figure 3a embodiment.
[0424] The decoding device in the embodiments of the present application has been described above. On the basis of the corresponding Figures 1 to 13Based on the corresponding embodiments, in order to better implement the above solutions of the embodiments of the present application, the following also provides a description of the training device in the embodiments of the present application. Specifically, refer to Figure 16 , Figure 16 FIG. Figure 16 is a schematic structural diagram of a training device 1600 provided by an embodiment of the present application. The training device 1600 includes:
[0425] A first acquisition module 1601, configured to acquire a first image.
[0426] A segmentation module 1602, configured to segment the first image to obtain N first tiles, where N is an integer greater than 1.
[0427] A second acquisition module 1603, configured to acquire N first adaptive data from the N first tiles, and the N first adaptive data correspond to the N first tiles one by one.
[0428] A preprocessing module 1604, configured to preprocess the N first tiles according to the N first adaptive data;
[0429] A first encoding neural network module 1605, configured to process the N preprocessed first tiles to obtain N groups of first feature maps.
[0430] A quantization and entropy encoding module 1606, configured to perform quantization and entropy encoding on the N groups of first feature maps to obtain N first encoded representations.
[0431] An entropy decoding module 1607, configured to perform entropy decoding on the N first encoded representations to obtain N groups of second feature maps.
[0432] A first decoding neural network module 1608, configured to process the N groups of second feature maps to obtain N first reconstructed tiles.
[0433] A compensation module 1609, configured to compensate the N first reconstructed tiles through the N first adaptive data.
[0434] A combination module 1610, configured to combine the N compensated first reconstructed tiles to obtain a second image.
[0435] A third acquisition module 1611, configured to acquire the distortion loss of the second image relative to the first image.
[0436] A training module 1612, configured to jointly train the model by using a loss function until the image distortion value between the first image and the second image reaches a first preset level. The model includes the first encoding neural network, the quantization network, the entropy encoding network, the entropy decoding network, and the first decoding neural network. Optionally, the model further includes a segmentation network, and the trainable parameter in the segmentation network is the size of the first tile. Optionally, the model further includes a segmentation network, and the trainable parameter in the segmentation network is the size of the first tile.
[0437] An output module 1613, configured to output a second encoding neural network and a second decoding neural network. The second encoding neural network is a model obtained after the first encoding neural network has undergone iterative training, and the second decoding neural network is a model obtained after the first decoding neural network has undergone iterative training.
[0438] Optionally, the training device is further configured to perform all or part of the operations performed by the terminal device and / or the cloud device in the foregoing Figure 3a corresponding embodiments.
[0439] In an alternative design of the sixth aspect, the N first tiles include a first target tile, and the range of pixel values of the first target tile is smaller than the range of pixel values of the first image;
[0440] The device further includes:
[0441] An inverse quantization module, configured to inverse-quantize the pixel values of the first target tile;
[0442] The second acquisition module 1603 is specifically configured to obtain a first adaptive data from the inverse-quantized first target tile.
[0443] Next, an execution device provided in an embodiment of the present application is introduced. Please refer to Figure 17 , Figure 17 which is a schematic structural diagram of an execution device provided in an embodiment of the present application. The execution device 1700 may specifically be a virtual reality (VR) device, a mobile phone, a tablet computer, a laptop computer, a smart wearable device, a monitoring data processing device, a server, etc., which is not limited herein. Among them, the encoding device described in the corresponding embodiment and / or Figure 14 the decoding device described in the corresponding embodiment may be deployed on the execution device 1700 to implement Figure 15 and / or Figure 14 and / or Figure 15 the functions of the device in the corresponding embodiment. Specifically, the execution device 1700 includes: a receiver 1701, a transmitter 1702, a processor 1703, and a memory 1704 (where the number of processors 1703 in the execution device 1700 may be one or more,Figure 17 Taking a processor as an example, the processor 1703 may include an application processor 17031 and a communication processor 17032. In some embodiments of the present application, the receiver 1701, the transmitter 1702, the processor 1703, and the memory 1704 may be connected by a bus or other means.
[0444] The memory 1704 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1703. A part of the memory 1704 may also include a non-volatile random access memory (NVRAM). The memory 1704 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, where the operation instructions may include various operation instructions for implementing various operations.
[0445] The processor 1703 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system, where the bus system may include a power bus, a control bus, a status signal bus, etc. in addition to the data bus. However, for the sake of clear illustration, all kinds of buses are referred to as the bus system in the figure.
[0446] The method disclosed in the embodiments of the present application can be applied to or implemented by the processor 1703. The processor 1703 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 1703 or by instructions in the form of software. The above-mentioned processor 1703 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 1703 can implement or execute various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1704, and the processor 1703 reads the information in the memory 1704 and combines its hardware to complete the steps of the above method.
[0447] The receiver 1701 can be used to receive input digital or character information, and generate signal inputs related to the relevant settings and function controls of the execution device. The transmitter 1702 can be used to output digital or character information through the first interface; the transmitter 1702 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1702 can also include a display device such as a display screen.
[0448] In an embodiment of the present application, in one case, the processor 1703 is used to execute Figure 3a the operations performed by the terminal device and / or the cloud device in the corresponding embodiment.
[0449] Optionally, the application processor 17031 is used to obtain a first image;
[0450] Segment the first image to obtain N first tiles, where N is an integer greater than 1;
[0451] Obtain N first adaptive data from the N first tiles, and the N first adaptive data correspond to the N first tiles one by one;
[0452] Preprocess N first tiles according to N first adaptive data;
[0453] Process the preprocessed N first tiles through an encoding neural network to obtain N groups of first feature maps;
[0454] Quantize and entropy-encode the N groups of first feature maps to obtain N first encoded representations;
[0455] In addition, the application processor 17031 can also be used to execute all or part of the operations that the cloud device in the corresponding embodiment above can execute. Figure 3a The corresponding embodiment of the terminal device can execute all or part of the operations.
[0456] Optionally, the application processor 17031 is used to obtain N first encoded representations;
[0457] Entropy-decode the N first encoded representations to obtain N groups of second feature maps;
[0458] Process the N groups of second feature maps through a decoding neural network to obtain N first reconstructed tiles;
[0459] Compensate the N first reconstructed tiles with N first adaptive data;
[0460] Combine the compensated N first reconstructed tiles to obtain a second image;
[0461] In addition, the application processor 17031 can also be used to execute all or part of the operations that the terminal device in the corresponding embodiment above can execute. Figure 3a The corresponding embodiment of the terminal device can execute all or part of the operations.
[0462] The embodiment of the present application also provides a training device. Please refer to Figure 18 , Figure 18 which is a schematic structural diagram of a training device provided by the embodiment of the present application. The training device 1800 can be deployed with Figure 16 the training device described in the corresponding embodiment, used to implement Figure 16Corresponding to the functions of the training device in the embodiments, specifically, the training device 1800 is implemented by one or more servers. The training device 1800 may vary significantly due to configuration or performance differences and may include one or more central processing units (CPUs) 1822 (e.g., one or more processors) and a memory 1832, and one or more storage media 1830 (e.g., one or more mass storage devices) for storing application programs 1842 or data 1844. Among them, the memory 1832 and the storage media 1830 may be transient storage or persistent storage. The program stored in the storage media 1830 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the training device. Further, the central processing unit 1822 may be configured to communicate with the storage media 1830 and execute a series of instruction operations in the storage media 1830 on the training device 1800.
[0463] The training device 1800 may further include one or more power supplies 1826, one or more wired or wireless network interfaces 1850, one or more input / output interfaces 1858, and / or one or more operating systems 1841, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0464] In the embodiments of the present application, the central processing unit 1822 is used to execute Figure 16 all or part of the operations executed by the training device in the corresponding embodiments.
[0465] The embodiments of the present application also provide a computer program product. When it runs on a computer, it causes the computer to execute the steps executed by the execution device in the method described in the foregoing Figure 17 illustrated embodiments, or causes the computer to execute the steps executed by the training device in the method described in the foregoing Figure 18 illustrated embodiments.
[0466] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a program for signal processing. When it runs on a computer, it causes the computer to execute the steps executed by the execution device in the method described in the foregoing Figure 17 illustrated embodiments, or causes the computer to execute the steps executed by the training device in the method described in the foregoing Figure 18 illustrated embodiments.
[0467] The execution device and training device provided by the embodiments of the present application may specifically be chips, and the chips include: a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin, a circuit, etc. The processing unit may execute the computer execution instructions stored in the storage unit to enable the chips in the execution device to execute the operations performed by the terminal device and / or the cloud device described in the above Figure 3a embodiments, or to enable the chips in the training device to execute the model training method described in the above Figure 13 embodiments. Optionally, the storage unit is a storage unit within the chip, such as a register, a cache, etc. The storage unit may also be a storage unit outside the chip within the radio access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0468] Specifically, please refer to Figure 19 , Figure 19 which is a schematic structural diagram of a chip provided by the embodiments of the present application. The chip may be embodied as a neural network processor NPU2000, and the NPU2000 is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are allocated by the Host CPU. The core part of the NPU is the arithmetic circuit 2003, and the arithmetic circuit 2003 is controlled by the controller 2004 to extract matrix data from the memory and perform multiplication operations.
[0469] In some implementations, the arithmetic circuit 2003 includes multiple processing units (Process Engine, PE) inside. In some implementations, the arithmetic circuit 2003 is a two-dimensional systolic array. The arithmetic circuit 2003 may also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 2003 is a general matrix processor.
[0470] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 2002 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 2001 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are saved in the accumulator 2008.
[0471] The unified memory 2006 is used to store input data and output data. The weight data directly passes through the Direct Memory Access Controller (DMAC) 2005 and is transferred to the weight memory 2002 by the DMAC. The input data is also transferred to the unified memory 2006 by the DMAC.
[0472] The BIU is the Bus Interface Unit, i.e., the bus interface unit 2010, which is used for the interaction between the AXI bus, the DMAC, and the Instruction Fetch Buffer (IFB) 2009.
[0473] The bus interface unit 2010 (Bus Interface Unit, abbreviated as BIU) is used for the instruction fetch buffer 2009 to obtain instructions from the external memory, and is also used for the storage unit access controller 2005 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0474] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 2006, transfer the weight data to the weight memory 2002, or transfer the input data to the input memory 2001.
[0475] The vector calculation unit 2007 includes multiple arithmetic processing units, which, if necessary, further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as Batch Normalization, pixel-level summation, upsampling of the feature plane, etc.
[0476] In some implementations, the vector calculation unit 2007 can store the processed output vector in the unified memory 2006. For example, the vector calculation unit 2007 can apply a linear function and / or a non-linear function to the output of the arithmetic circuit 2003, such as linear interpolation of the feature plane extracted by the convolutional layer, or a vector of accumulated values, to generate activation values. In some implementations, the vector calculation unit 2007 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 2003, for example, for use in subsequent layers in the neural network.
[0477] The instruction fetch buffer 2009 connected to the controller 2004 is used to store the instructions used by the controller 2004;
[0478] The unified memory 2006, the input memory 2001, the weight memory 2002, and the fetch memory 2009 are all On-Chip memories. The external memory is private to the NPU hardware architecture.
[0479] Among them, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the method in the first aspect above.
[0480] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0481] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by means of dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in various embodiments of this application.
[0482] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0483] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
Claims
1. An image processing method, characterized in that, comprising: obtaining a first image; segmenting the first image to obtain N first tiles, where N is an integer greater than 1; obtaining N first adaptive data from the N first tiles, and the N first adaptive data correspond to the N first tiles one by one; preprocessing the N first tiles according to the N first adaptive data; processing the preprocessed N first tiles through an encoding neural network to obtain N groups of first feature maps; quantizing and entropy encoding the N groups of first feature maps to obtain N first encoded representations.
2. The method according to claim 1, characterized in that, the N first encoded representations are used for entropy decoding to obtain N groups of second feature maps, the N groups of second feature maps are used for processing through a decoding neural network to obtain N first reconstructed tiles, and the N first adaptive data are used for compensating the N first reconstructed tiles, and the compensated N first reconstructed tiles are used for combining into a second image.
3. The method according to claim 2, characterized in that, the method further comprises: sending the N first encoded representations, the N first adaptive data and the correspondence relationship to a decoding end, where the correspondence relationship includes the correspondence relationship between the N first adaptive data and the N first encoded representations.
4. The method according to claim 3, characterized in that, the method further comprises: quantizing the N first adaptive data to obtain N first adaptive quantization data, and the N first adaptive quantization data are used for compensating the N first reconstructed tiles; the sending the N first adaptive data to the decoding end includes: sending the N first adaptive quantization data to the decoding end.
5. The method according to claim 4, characterized in that, the larger the N, the smaller the information entropy of a single first adaptive quantization data.
6. The method according to any one of claims 3 to 5, characterized in that, the arrangement order of the N first encoded representations is the same as the arrangement order of the N first tiles, the arrangement order of the N first tiles is the arrangement order of the N first tiles in the first image, and the correspondence relationship includes the arrangement order of the N first encoded representations and the arrangement order of the N first tiles.
7. The method according to any one of claims 1 to 5, characterized in that, each of the N first tiles has the same size.
8. The method according to claim 7, characterized in that, when the method is used for segmenting the first images of different sizes, the size of the first tiles is a fixed value.
9. The method according to any one of claims 1 to 5, characterized in that, The pixels of the first tile are a×b, where a and b are obtained based on a target pixel, and the target pixel is c×d. Equal to an integer. Equal to an integer. a and c are the number of pixel points in the width direction, and b and d are the number of pixel points in the height direction. The target pixel is obtained based on the target resolution of the terminal device. The terminal device includes a camera component. The pixels of the image obtained by the camera component under the setting of the target resolution are the target pixel. The first image is obtained by the camera component.
10. The method according to claim 9, characterized in that, the target resolution is set according to the resolution of the imaging component in the imaging application through the setting interface.
11. The method according to claim 9, characterized in that, The target resolution is obtained based on a target image group in a picture library obtained by the imaging component. The pixels of the target image group are the target pixels, and among image groups with different pixels, the ratio of the target image group in the picture library is the largest.
12. The method according to any one of claims 1 to 5, wherein, the pixels of the first tile are a×b, a is the number of pixel points in the width direction, b is the number of pixel points in the height direction, and the pixels of the first image are r×t; after obtaining the first image and before segmenting the first image, the method further includes: If is not equal to an integer, and / or is not equal to an integer, then fill the edge of the first image with the median value of the pixels such that is equal to an integer, is equal to an integer, and the pixels of the filled first image are r1×t1.
13. The method according to claim 12, wherein, after obtaining the first image and before filling the edge of the first image, the method further includes: If the is not an integer, then proportionally enlarge the r and the t to obtain the first image with pixels of r2×t2, and the is an integer; If the following is not an integer, and / or is not an integer, filling the edge of the first image with the pixel median value includes: If is not equal to an integer, fill the edge of the first image with the median value of the pixels.
14. The method according to any one of claims 1 to 5, wherein, the N first tiles include a first target tile, and the range of the pixel values of the first target tile is smaller than the range of the pixel values of the first image; before obtaining N first adaptive data from the N first tiles, the method further includes: inverse quantizing the pixel values of the first target tile; obtaining N first adaptive data from the N first tiles includes: obtaining one first adaptive data from the inverse quantized first target tile.
15. An image processing method, wherein, it includes: obtaining N first coded representations, N first adaptive data, and a correspondence relationship, where the correspondence relationship includes the correspondence relationship between the N first adaptive data and the N first coded representations, the N first adaptive data correspond one-to-one with N first tiles, and N is an integer greater than 1; performing entropy decoding on the N first coded representations to obtain N groups of second feature maps; processing the N groups of second feature maps through a decoding neural network to obtain N first reconstructed tiles; compensating the N first reconstructed tiles with the N first adaptive data; combining the compensated N first reconstructed tiles to obtain a second image.
16. The method according to claim 15, wherein, the N first coded representations are obtained through quantization and entropy coding of N groups of first feature maps, the N groups of first feature maps are obtained by processing N preprocessed first tiles through an encoding neural network, the preprocessed N first tiles are obtained by preprocessing N first tiles with the N first adaptive data, the N first adaptive data are obtained from the N first tiles, and the N first tiles are obtained by segmenting a first image.
17. The method according to claim 15 or 16, wherein, the larger N is, the smaller the information entropy of a single first adaptive quantization data.
18. The method according to claim 16, wherein, the method further includes: processing the second image through a fusion neural network to obtain a third image to reduce the difference between the second image and the first image, where the difference includes block effect.
19. An encoding device, wherein, it includes: A first acquisition module, configured to acquire a first image; A segmentation module, configured to segment the first image to obtain N first tiles, where N is an integer greater than 1; A second acquisition module, configured to acquire N first adaptive data from the N first tiles, where the N first adaptive data correspond to the N first tiles one by one; A preprocessing module, configured to preprocess the N first tiles according to the N first adaptive data; An encoding neural network module, configured to process the N preprocessed first tiles to obtain N groups of first feature maps; A quantization and entropy encoding module, configured to perform quantization and entropy encoding on the N groups of first feature maps to obtain N first encoded representations.
20. The apparatus according to claim 19, wherein, the N first encoded representations are used for entropy decoding to obtain N groups of second feature maps, the N groups of second feature maps are used to be processed by a decoding neural network to obtain N first reconstructed tiles, the N first adaptive data are used to compensate the N first reconstructed tiles, and the compensated N first reconstructed tiles are used to be combined into a second image.
21. The apparatus according to claim 20, wherein, the apparatus further comprises: A sending module, configured to send the N first encoded representations, the N first adaptive data, and a correspondence relationship to a decoding end, where the correspondence relationship includes the correspondence relationship between the N first adaptive data and the N first encoded representations.
22. The apparatus according to claim 21, wherein, the apparatus further comprises: A quantization module, configured to quantize the N first adaptive data to obtain N first adaptive quantization data, and the N first adaptive quantization data are used to compensate the N first reconstructed tiles; The sending module is specifically configured to send the N first adaptive quantization data to the decoding end.
23. The apparatus according to claim 22, wherein, the larger the N is, the smaller the information entropy of a single first adaptive quantization data is.
24. The apparatus according to any one of claims 21 to 23, wherein, the arrangement order of the N first encoded representations is the same as the arrangement order of the N first tiles, the arrangement order of the N first tiles is the arrangement order of the N first tiles in the first image, and the correspondence relationship includes the arrangement order of the N first encoded representations and the arrangement order of the N first tiles.
25. The apparatus according to any one of claims 19 to 23, wherein, each of the N first tiles has the same size.
26. The apparatus according to claim 25, wherein, when the apparatus is used to process the first images of different sizes, the size of the first tile is a fixed value.
27. A decoding apparatus, wherein, comprises: An acquisition module, configured to acquire N first encoded representations, N first adaptive data, and a correspondence relationship, where the correspondence relationship includes the correspondence relationship between the N first adaptive data and the N first encoded representations, the N first adaptive data are in one-to-one correspondence with N first tiles, and N is an integer greater than 1; An entropy decoding module, configured to perform entropy decoding on the N first encoded representations to obtain N groups of second feature maps; A decoding neural network module, configured to process the N groups of second feature maps to obtain N first reconstructed tiles; A compensation module, configured to compensate the N first reconstructed tiles by using the N first adaptive data; A combination module, configured to combine the N first reconstructed tiles after compensation to obtain a second image.
28. The apparatus according to claim 27, wherein, the N first encoded representations are obtained by quantizing and entropy encoding N groups of first feature maps, the N groups of first feature maps are obtained by processing N first tiles after preprocessing through an encoding neural network, the N first tiles after preprocessing are obtained by preprocessing N first tiles by using the N first adaptive data, the N first adaptive data are obtained from the N first tiles, and the N first tiles are obtained by splitting a first image.
29. The apparatus according to claim 27 or 28, wherein, the larger the N, the smaller the information entropy of a single first adaptive quantization data.
30. The apparatus according to claim 28, wherein, the apparatus further includes: A fusion neural network module, configured to process the second image to obtain a third image, so as to reduce the difference between the second image and the first image, where the difference includes blocking artifacts.
31. An image processing device, including: A non-volatile memory and a processor coupled to each other, and the processor calls program code stored in the memory to execute the method described in any one of claims 1-18.
Citation Information
Patent Citations
Method and apparatus for encoding and decoding image
CN101978698A
Adaptive sampling point compensation coding method and device, and method and device for decoding video code stream
CN105635732A
Method and device for encoding or decoding image
CN111052740A