Image Processing Method, Apparatus and Electronic Device
By determining the target pixel area in the super-segment display device for super-segment processing, and updating the image using other pixel areas of the previous frame image, the problems of large amount of super-segment processing and high cost in the prior art are solved, and efficient image processing and display effects are guaranteed.
Patent Information
- Application Number
- CN201911008119.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-10-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2039-10-22
AI Technical Summary
In the image processing of the existing super-score display device, the super-score processing has a large amount of computing and a high cost.
By acquiring the inter-frame residuals between the image and the previous frame image, the target pixel area is determined, and the area is super-segmented, and other pixel areas in the image are updated with other pixel areas in the previous frame image after the super-segmented process, reducing unnecessary super-segmented processing.
It effectively reduces the calculation amount and calculation cost of super-score processing, improves the efficiency of image processing, and ensures the display effect of image.
Smart Images

Figure CN112700368B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and particularly to an image processing method, apparatus, and electronic device. Background Art
[0002] With the development of technology, super-resolution display devices, such as smart TVs, are increasingly widely used. Super-resolution, that is, super-resolution. A super-resolution display device is a display device capable of performing super-resolution processing on an image, and super-resolution processing is a technology for reconstructing a low-resolution image into a high-resolution image.
[0003] Currently, a super-resolution display device inputs a decoded image into a super-resolution model, and the super-resolution model performs super-resolution processing on the image. However, in such an image processing method, the computational amount of super-resolution processing is large, and the computational cost is high. Summary of the Invention
[0004] Embodiments of the present application provide an image processing method, apparatus, and electronic device, which can reduce the computational amount of current super-resolution processing and lower the computational cost. The following introduces the present application through different aspects. It should be understood that the implementation manners and beneficial effects of the following different aspects can be mutually referred to.
[0005] The "first" and "second" that appear in the present application are only used to distinguish two objects, and do not have the meaning of a sequence.
[0006] Embodiments of the present application provide an image processing method, and the method includes:
[0007] Obtain the inter-frame residual between a first image and the adjacent previous frame image to obtain a residual block. The residual block includes a plurality of residual points corresponding one-to-one to a plurality of pixel positions of the first image, and each residual point has a residual value; based on the residual block, determine a target pixel region in the first image; perform super-resolution processing on the target pixel region in the first image to obtain a super-resolution processed target pixel region; update other pixel regions in the first image with other pixel regions in the super-resolution processed previous frame image, where the other pixel regions include pixel regions in the first image except the target pixel region. In this way, the same effect as performing super-resolution processing on other pixel regions is achieved without performing super-resolution processing on other pixel regions.
[0008] Wherein, the super-resolution processed first image includes the super-resolution processed target pixel region and the updated other pixel regions (equivalent to the super-resolution processed other pixel regions).
[0009] In the embodiment of the present application, by determining the target pixel region in the first image and performing super-resolution processing on the target pixel region, super-resolution processing is achieved for the region where the pixel points with differences between the first image and the previous frame image are located. Moreover, the other pixel regions in the first image are updated with the other pixel regions in the super-resolved previous frame image, achieving the same effect of super-resolution processing for the other pixel regions, making full use of the characteristics of video temporal redundancy. Therefore, by performing super-resolution processing on a partial region of the first image, the effect of performing super-resolution processing on the entire first image is achieved, reducing the computational amount of super-resolution processing and lowering the computational cost.
[0010] Since the other pixel regions include the pixel regions in the first image except the target pixel region, the sizes of the target pixel region and the other pixel regions may or may not match. Correspondingly, the method for obtaining the super-resolved first image is also different. The embodiment of the present application will be described by taking the following two examples:
[0011] In one example, the sizes of the target pixel region and the other pixel regions match, that is, the other pixel regions are the pixel regions in the first image except the target pixel region; correspondingly, the sizes of the super-resolved target pixel region and the updated other pixel regions also match. Then, the super-resolved first image can be formed by splicing the super-resolved target pixel region and the updated other pixel regions.
[0012] In another example, the sizes of the target pixel region and the other pixel regions do not match, and there is an overlapping region at their edges, that is, the other pixel regions include other pixel regions in addition to the pixel regions in the first image except the target pixel region. Correspondingly, the sizes of the super-resolved target pixel region and the updated other pixel regions also do not match, and there is an overlapping region at their edges. Since the other pixel regions in the first image are updated from the other pixel regions in the super-resolved second image, the pixel data of the included pixel points are usually more accurate. Therefore, the pixel data of the overlapping region of the super-resolved first image usually takes the pixel data of the updated other pixel regions in the first image as the standard. Then, the super-resolved first image can be formed by splicing the updated target pixel region and the updated other pixel regions. The updated target pixel region is obtained by subtracting (also known as removing) its overlapping region with the other pixel regions from the super-resolved target pixel region. The updated target pixel region is shrunk inward relative to the target pixel region before updating, and the size of the updated target pixel region matches the size of the updated other pixel regions.
[0013] Optionally, the target pixel region includes the region of the pixels corresponding to the position of the first target residual point in the first image, and the first target residual point is the point in the residual block where the residual value is greater than a specified threshold. Optionally, the specified threshold is 0.
[0014] Exemplarily, the target pixel region is the region of the pixels corresponding to the positions of the first target residual point and the second target residual point in the first image, and the second target residual point is the residual point around the first target residual point in the residual block. The residual points around the first target residual point refer to the residual points set around the first target residual point, which are the peripheral points of the first target residual point that meet the specified conditions. For example, the residual points above, below, left, and right of the first target residual point; or, the residual points above, below, left, right, top - left, bottom - left, top - right, and bottom - right of the first target residual point. The aforementioned specified conditions are determined based on the requirements of the super - resolution process, such as being set based on the receptive field (e.g., the receptive field of the last convolutional layer) in the super - resolution model.
[0015] Exemplarily, the second target residual point is the residual point in the residual block where the residual value around the first target residual point is not greater than the specified threshold. That is, the second target residual point is the point among the peripheral points of the first target residual point that meet the specified conditions and where the residual value is not greater than the specified threshold.
[0016] When the residual values of all the residual points in the residual block are 0, it indicates that the content of the first image and the previous frame image has not changed, and the super - resolved images of both should also not change. The previous frame image after super - resolution processing can be used to update the first image to achieve the same effect as super - resolving the first image, without the need to perform super - resolution processing on the first image again, that is, without performing the action of determining the target pixel region, which can effectively reduce the computational cost. Correspondingly, the aforementioned determination of the target pixel region in the first image based on the residual block may include: when at least one of the residual points included in the residual block has a non - zero residual value, determining the target pixel region in the first image based on the residual block. That is, when there are residual points with non - zero residual values in the residual block, the action of determining the target pixel region is then performed.
[0017] In some implementation manners, the determination of the target pixel region in the first image based on the residual block includes:
[0018] Generating a mask pattern based on the residual block, the mask pattern including a plurality of first mask points, and the plurality of first mask points corresponding one - to - one to the positions of a plurality of target residual points in the residual block; inputting the mask pattern and the first image into a super - resolution model, and determining, through the super - resolution model, the region of the pixels corresponding to each mask point position among the plurality of first mask points in the first image as the target pixel region.
[0019] Correspondingly, the super-resolution processing of the target pixel points in the first image to obtain the super-resolution processed target pixel points includes:
[0020] Performing super-resolution processing on the target pixel region in the first image through the super-resolution model to obtain a super-resolution processed target pixel region.
[0021] In some other implementation manners, the determining of the target pixel region in the first image based on the residual block includes:
[0022] Generating a mask pattern based on the residual block, the mask pattern including a plurality of first mask points, and the positions of the plurality of first mask points corresponding one-to-one to the positions of a plurality of target residual points in the residual block; determining the region where the pixel points corresponding to each mask point in the plurality of first mask points are located in the first image as the target pixel region.
[0023] Correspondingly, the super-resolution processing of the target pixel points in the first image to obtain the super-resolution processed target pixel points includes:
[0024] Inputting the target pixel points in the first image into the super-resolution model, and performing super-resolution processing on the target pixel region in the first image through the super-resolution model to obtain a super-resolution processed target pixel region.
[0025] In some implementation manners, the generating of the mask pattern based on the residual block includes:
[0026] Generating an initial mask pattern including a plurality of mask points based on the residual block, the plurality of mask points corresponding one-to-one to the positions of a plurality of pixel points in the first image, and the plurality of mask points including the plurality of first mask points and a plurality of second mask points; assigning a first value to the mask value of the first mask points in the initial mask pattern, and assigning a second value to the mask value of the second mask points in the mask pattern to obtain the mask pattern, where the first value and the second value are different;
[0027] The determining of the pixel points corresponding to the plurality of first mask points in the first image as the target pixel region includes: traversing the mask points in the mask pattern, and in the first image, determining the pixel points corresponding to the mask points with the mask value of the first value as the target pixel region.
[0028] In some implementation manners, the generating of the mask pattern based on the residual block includes:
[0029] Perform morphological transformation processing on the residual block to obtain the mask pattern. The morphological transformation processing includes binarization processing and dilation processing on the first mask points in the binarized residual block. Exemplarily, the binarization processing and the dilation processing can be performed sequentially.
[0030] Among them, binarization processing is a processing method that sets the pixel value of each pixel point in the image to one of two pixel values, a first value and a second value, and the first value and the second value are different. After the image undergoes binarization processing, it only includes pixel points with two pixel values. This binarization processing can reduce the interference of various elements in the image on the subsequent image processing process.
[0031] The binarization processing in the embodiments of the present application can adopt any method among the global binarization threshold method, the local binarization threshold method, the maximum inter-class variance method, and the iterative binarization threshold method. The embodiments of the present application do not limit this.
[0032] Dilation processing is a processing for finding local maxima. Convolve the image to be processed with a preset kernel (also called a core). During each convolution process, the maximum value in the kernel coverage area is assigned to the specified pixel point, so that the bright ones become brighter. The resulting effect is that the bright areas of the image to be processed expand. Among them, the kernel has a definable anchor point, which is usually the center point of the kernel, and the aforementioned specified pixel point is this anchor point.
[0033] In the embodiments of the present application, the super-resolution model includes at least one convolution kernel, and the kernel of the dilation processing has the same size as the receptive field of the last convolution layer of the super-resolution model. This last convolution layer is the output layer of the super-resolution model, and the image after the super-resolution processing of the super-resolution model is output from this layer. The receptive field of this layer is the largest receptive field among the receptive fields corresponding to each convolution layer in the super-resolution model. In this way, the obtained mask pattern is adapted to the size of the largest receptive field of the super-resolution model, so as to play a good guiding role, avoid the situation that the convolvable area of the image input into the super-resolution model subsequently is too small, resulting in the image not being convolvable by the convolution layer, and ensure that the image input into the super-resolution model subsequently can be effectively super-resolved. For example, the size of this kernel can be 3×3 or 5×5 pixel points.
[0034] In some implementation manners, generating a mask pattern based on the residual block includes:
[0035] Divide the residual block into multiple sub-residual blocks, and perform block processing on each divided sub-residual block. The block processing includes:
[0036] When the residual values of at least one residual point included in the sub-residual block are not 0, divide the sub-residual block into multiple sub-residual blocks, and perform the block processing on each divided sub-residual block until the residual values of the residual points included in the divided sub-residual block are all 0, or the total number of residual points in the divided sub-residual block is less than the point threshold, or the total number of times of dividing the residual block reaches the number threshold; generate a sub-mask pattern corresponding to each target residual block, where at least one residual point included in the target residual block has a non-zero residual value; wherein, the mask pattern includes the generated sub-mask patterns.
[0037] In some implementation manners, the performing super-resolution processing on the target pixel points in the first image to obtain the super-resolution processed target pixel points includes:
[0038] Obtain a target image block corresponding to each sub-mask pattern in the first image; perform super-resolution processing on sub-regions of the target pixel region included in each target image block respectively to obtain the super-resolution processed target pixel region, and the super-resolution processed target pixel region is composed of sub-regions of the target pixel region included in each super-resolution processed target image block.
[0039] In the embodiments of the present application, multiple target image blocks can be screened out from the first image, and super-resolution processing is performed on each target image block. Since the super-resolution processing is performed on the target image blocks, and the size of each target image block is smaller than that of the first image, the computational complexity of the super-resolution processing can be reduced, and the computational cost can be lowered. Especially when the super-resolution processing is performed by a super-resolution model, the complexity of the super-resolution model can be effectively reduced, and the efficiency of the super-resolution operation can be improved.
[0040] In some implementation manners, the dividing the residual block into multiple sub-residual blocks includes: dividing the residual block into multiple sub-residual blocks in a quadtree partitioning manner;
[0041] The dividing the sub-residual block into multiple sub-residual blocks includes: dividing the sub-residual block into multiple sub-residual blocks in a quadtree partitioning manner.
[0042] Since the traditional video decoding process requires image chunking, which is usually performed in a quadtree partitioning manner, in the embodiments of the present application, when the residual blocks and sub-residual blocks are partitioned in a quadtree partitioning manner, traditional image processing methods can be compatible. For example, in practical applications, the residual blocks and sub-residual blocks can also be partitioned by the image partitioning module used in the aforementioned video decoding process, so as to achieve module reuse and save computational costs. And by using the quadtree partitioning method, each time of partitioning, a residual block or a sub-residual block can be divided into four sub-residual blocks with equal sizes, making the sizes of the obtained residual blocks uniform and facilitating subsequent processing.
[0043] In some implementation manners, before updating the other pixel regions in the first image with the other pixel regions in the previous frame image after super-resolution processing, the method further includes:
[0044] Performing erosion processing on multiple first mask points in the mask pattern to obtain an updated mask pattern, where the kernel of the erosion processing has the same size as the receptive field of the last convolutional layer of the super-resolution model; determining the pixel points corresponding to the positions of the multiple first mask points after erosion processing in the first image as auxiliary pixel points; and determining the region where the pixel points other than the auxiliary pixel points in the first image are located as the other pixel regions.
[0045] Erosion processing can eliminate the edge noise of the image. The other pixel regions determined by the updated mask pattern obtained through erosion processing have clearer edges and less noise compared to the other pixel regions obtained by the aforementioned first and second optional manners. When performing the subsequent pixel update process, negative effects such as detail blurring, edge blunting, graininess, and noise enhancement can be reduced, ensuring the display effect of the finally reconstructed first image.
[0046] In some implementation manners, before determining the target pixel region in the first image, determining the target pixel region in the first image based on the residual block includes: counting the first proportion of the number of residual points with a residual value of 0 in the residual block in the total number of residual points in the residual block; when the first proportion is greater than the first super-resolution trigger proportion threshold, determining the target pixel region in the first image based on the residual block.
[0047] In this way, it is possible to determine whether to execute the partial super-resolution algorithm based on the content difference between the first image and the previous frame image, thereby improving the flexibility of image processing.
[0048] In some implementation manners, the super-resolution model is a CNN model, such as models like SRCNN or ESPCN; the super-resolution model can also be a GAN, such as models like SRGAN or ESRGAN.
[0049] In a second aspect, an exemplary embodiment of the present application provides an image processing apparatus, which includes one or more modules for implementing any one of the image processing methods in the foregoing first aspect.
[0050] In a third aspect, an embodiment of the present application provides an electronic device, such as a terminal. The electronic device includes a processor and a memory. The processor generally includes a CPU and / or a GPU. The memory is used to store a computer program; the processor is used to implement any one of the image processing methods in the foregoing first aspect when executing the computer program stored in the memory. Among them, the CPU and the GPU can be two chips or integrated on the same chip.
[0051] In a fourth aspect, an embodiment of the present application provides a storage medium, which can be non-volatile. The storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement any one of the image processing methods in the foregoing first aspect.
[0052] In a fifth aspect, an embodiment of the present application provides a computer program or a computer program product including computer-readable instructions. When the computer program or the computer program product runs on a computer, the computer is caused to execute any one of the image processing methods in the foregoing first aspect. The computer program product may include one or more program units for implementing the foregoing method.
[0053] In a sixth aspect, the present application provides a chip, such as a CPU. The chip includes a logic circuit, and the logic circuit can be a programmable logic circuit. When the chip runs, it is used to implement any one of the image processing methods in the foregoing first aspect.
[0054] In a seventh aspect, the present application provides a chip, such as a CPU. The chip includes one or more physical cores and a storage medium. After the one or more physical cores read the computer instructions in the storage medium, any one of the image processing methods in the foregoing first aspect is implemented.
[0055] In an eighth aspect, the present application provides a chip, such as a GPU. The chip includes one or more physical cores and a storage medium. After the one or more physical cores read the computer instructions in the storage medium, any one of the image processing methods in the foregoing first aspect is implemented.
[0056] In a ninth aspect, the present application provides a chip, such as a GPU. The chip includes a logic circuit, and the logic circuit can be a programmable logic circuit. When the chip runs, it is used to implement any one of the image processing methods in the foregoing first aspect.
[0057] In summary, in the embodiment of the present application, by determining the target pixel region in the first image and performing super-resolution processing on the target pixel region, super-resolution processing is achieved for the region where the pixel points different between the first image and the previous frame image are located. Moreover, other pixel regions in the first image are updated with other pixel regions in the super-resolution processed previous frame image, achieving the same effect of super-resolution processing for other pixel regions, making full use of the characteristics of video temporal redundancy. Therefore, by performing super-resolution processing on a partial region of the first image, the effect of performing super-resolution processing on the entire first image is achieved, reducing the computational amount of super-resolution processing and the computational cost.
[0058] For the test video processed by the embodiment of the present application, compared with directly performing full super-resolution processing on the video in the traditional technology, it can save approximately 45% of the super-resolution calculation amount. The significant reduction in the super-resolution calculation amount, on the one hand, is beneficial to accelerating the video processing speed, ensuring that the video can meet the basic frame rate requirements, thereby guaranteeing the real-time nature of the video and preventing situations such as playback delay and stuttering; on the other hand, the reduction in the calculation amount means less processing tasks and consumption of the calculation unit in the super-resolution display device, resulting in a decrease in the overall power consumption and saving the power consumption of the device.
[0059] Moreover, the partial super-resolution algorithm proposed in the embodiment of the present application is not a method of sacrificing the effect in exchange for efficiency by only super-resolving partial image regions and processing other parts by non-super-resolution means. Instead, it avoids repeated super-resolution of the unchanged regions and redundant temporal information between the front and back frames of the video. Essentially, it is a method that pursues the maximization of information utilization rate. For the first image using the partial super-resolution algorithm, by setting a mask pattern, it guides the super-resolution model to perform super-resolution at the pixel level. In the finally processed video, in essence, all pixel values of each frame image come from the super-resolution calculation results, which is the same as the display effect of the traditional full super-resolution algorithm, avoiding the sacrifice of the display effect.
[0060] Moreover, if the method of inputting a sub-mask pattern and a target image block is adopted and the super-resolution model performs super-resolution processing, the complexity of each super-resolution processing by the super-resolution model is relatively low, and the requirement for the structural complexity of the super-resolution model is relatively low. Thus, the super-resolution model can be simplified, the requirement for the processor performance can be reduced, and the super-resolution processing efficiency can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 It is a schematic structural diagram of a super-resolution display device involved in an image processing method provided by an embodiment of the present application;
[0062] Figure 2 It is a flowchart of an image processing method provided by an embodiment of the present application;
[0063] Figure 3 It is a schematic diagram of the pixel values of the first image provided by an embodiment of the present application;
[0064] Figure 4 It is a schematic diagram of pixel values of the second image provided by an embodiment of the present application;
[0065] Figure 5 It is a schematic diagram of a residual block provided by an embodiment of the present application;
[0066] Figure 6 is Figure 5 A schematic diagram for explaining the principle of the residual block shown;
[0067] Figure 7 It is a schematic diagram of a process for determining a target pixel region in a first image provided by an embodiment of the present application;
[0068] Figure 8 It is a schematic diagram of the principle of dilation processing provided by an embodiment of the present application;
[0069] Figure 9 It is another schematic diagram of the principle of dilation processing provided by an embodiment of the present application;
[0070] Figure 10 It is another schematic diagram of a process for determining a target pixel region in a first image provided by an embodiment of the present application;
[0071] Figure 11 It is a schematic diagram of the principle of erosion processing provided by an embodiment of the present application;
[0072] Figure 12 It is a schematic diagram of the principle of updating another pixel region K2 in the first image with another pixel region K1 in the second image after super-resolution processing provided by an embodiment of the present application;
[0073] Figure 13 It is a schematic diagram of the principle of an image processing method provided by an embodiment of the present application;
[0074] Figure 14 It is a block diagram of an image processing device provided by an embodiment of the present application;
[0075] Figure 15 It is a block diagram of a first determination module provided by an embodiment of the present application;
[0076] Figure 16 It is a block diagram of another image processing device provided by an embodiment of the present application;
[0077] Figure 17 It is a block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0078] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.
[0079] For the convenience of readers' understanding, the following will first explain the terms involved in the embodiments of this application.
[0080] In this application, "a plurality of" refers to "two or more" or "at least two" without special explanation. "A and / or B" in this application includes at least three cases: "A", "B", and "A and B".
[0081] Image resolution: Used to reflect the amount of information stored in an image, referring to the total number of pixels in the image. Image resolution is usually expressed as the number of horizontal pixels × the number of vertical pixels.
[0082] 1080p: A display format, where P represents progressive scan. The image resolution of 1080p is usually 1920×1080.
[0083] 2k resolution: A display format, and the corresponding image resolution is usually 2048×1152.
[0084] 4k resolution: A display format, and the corresponding image resolution is usually 3840×2160.
[0085] 480p resolution: A display format, and the corresponding image resolution is usually 640×480.
[0086] 360p resolution: A display format, and the corresponding image resolution is usually 480×360.
[0087] Super resolution, that is, Super Resolution. Super resolution processing is a technology that reconstructs a low-resolution image into a high-resolution image, that is, the image resolution of the reconstructed image is greater than that of the image before reconstruction. The reconstructed image is also called a super resolution image. For example, super resolution processing can reconstruct an image with 360p resolution into an image with 480p resolution, or reconstruct an image with 2k resolution into an image with 4K resolution.
[0088] Color space, also known as color model, color space, or color system, is used to reflect the colors involved in an image. Different color spaces correspond to different color coding formats. Currently, the two relatively commonly used color spaces are the YUV color space and the RGB color space (each color in the color space is also called a color channel), and the corresponding color coding formats are the YUV format and the RGB format.
[0089] Among them, when the color coding format is the YUV format, the pixel values of a pixel point include: the value of the luminance component Y, the value of the chrominance component U, and the value of the chrominance component V; when the color coding format is the RGB format, the pixel values of a pixel point include the value of the transparency component and the values of multiple color components, and the multiple color components may include the red component R, the green component G, and the blue component B.
[0090] A Convolutional Neural Network (CNN) is a feedforward neural network. Its artificial neurons can respond to the surrounding units within a part of the coverage range and can process images according to image features.
[0091] Generally, the basic structure of a convolutional neural network includes two layers. One is the feature extraction layer, where the input of each neuron is connected to the local receptive field of the previous layer and extracts the features of the local receptive field. The other is the feature mapping layer. Each feature mapping layer of the network consists of multiple feature maps, and each feature map is a plane. The feature mapping layer is provided with an activation function, and the common activation function is a non-linear mapping function, which can be a sigmoid function or a Rectified linear unit (ReLU) function. A convolutional neural network is composed of a large number of nodes (also called "neurons" or "units") connected to each other, and each node represents a specific output function. The connection between every two nodes represents a weighted value, which is called a weight. Different weights and activation functions will result in different outputs of the convolutional neural network.
[0092] Usually, a convolutional neural network includes at least one convolutional layer. Each convolutional layer includes a feature extraction layer and a feature mapping layer. When the convolutional neural network includes multiple convolutional layers, the multiple convolutional layers are connected in sequence. The receptive field is the size of the area on the original image (referring to the image input into the convolutional neural network) mapped by each pixel point on the feature map output by each convolutional layer of the convolutional neural network.
[0093] One of the advantages of convolutional neural networks compared with traditional image processing algorithms is that it avoids the complex preprocessing process of images (such as extracting artificial features, etc.), and can directly input the original image for end-to-end learning. One of the advantages of convolutional neural networks compared with traditional neural networks is that traditional neural networks all adopt a fully connected method, that is, the neurons from the input layer to the hidden layer are all fully connected, which will lead to a huge number of parameters, making the network training time-consuming or even difficult to train, while convolutional neural networks avoid this problem through methods such as local connection and weight sharing.
[0094] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of a super-resolution display device 10 involved in an image processing method provided by an embodiment of the present application. The super-resolution display device 10 may be a product or component with a display function and a super-resolution processing function, such as a smart TV, a smart screen, a smart phone, a tablet computer, an electronic paper, a display, a notebook computer, a digital photo frame, or a navigator. The super-resolution display device 10 includes: a processor 101, a display control module 102, and a memory 103. Among them, the processor 101 is configured to process an image in a video obtained from a video source and transmit the processed image to the display control module 102, and the processed image is adapted to the format requirements of the display control module 102; the display control module 102 is configured to process the received processed image to obtain a driving signal adapted to a display module ( Figure 1 not labeled) and drive the display module to display an image based on the driving signal; the memory 103 is configured to store video data.
[0095] Exemplarily, the processor 101 may include a Central Processing Unit (CPU) and / or a Graphics Processing Unit (GPU), and the processor 101 may be integrated on a graphics card; the display control module 102 may be a Timing Controller (TCON) or a Microcontroller Unit (MCU); the display module may be a display screen; the memory 103 may be a Double Data Rate (DDR) dynamic random access memory. In the embodiment of the present application, a super-resolution model 1031 is stored in the memory 103, and the process of the processor 101 processing the images in the video obtained from the video source may include: decoding the video obtained from the video source, preprocessing the decoded image (such as subsequent step 201 or step 202, etc.), inputting the preprocessed image into the super-resolution model 1031, and performing super-resolution processing on the preprocessed image through the super-resolution model 1031. Exemplarily, the super-resolution model may be a CNN model, such as a Super-Resolution Convolutional Neural Network (SRCNN) or an Efficient sub-pixel Convolutional Neural Network (ESPCN), etc.; the super-resolution model may also be a Generative Adversarial Network (GAN), such as a Super-Resolution Generative Adversarial Network (SRGAN) or an Enhanced Super-Resolution Generative Adversarial Networks (ESRGAN), etc.
[0096] Currently, the super-resolution display device directly inputs the decoded image into the super-resolution model, and the super-resolution model performs super-resolution processing on the image. However, for such an image processing method, the computational workload of the super-resolution processing is large and the computational cost is high.
[0097] The embodiment of the present application provides an image processing method, which proposes a partial super-resolution algorithm, which can reduce the computational workload of the super-resolution processing and reduce the computational cost. This image processing method can be applied to Figure 1The super-resolution display device shown, since a video may include multiple images, in the embodiments of the present application, the first image is taken as an example to illustrate the image processing method. The first image is a frame image in the video, and the first image is a non-first frame image in the video (i.e., not the first frame image). The processing methods for other non-first frame images can refer to the processing method of the first image. Assume that the previous frame image adjacent to the first image is the second image, as Figure 2 shown, the method includes:
[0098] Step 201, the super-resolution display device obtains the inter-frame residual between the first image and the second image to obtain a residual block.
[0099] The inter-frame residual refers to the absolute value difference of the pixel values of two adjacent frame images in the video, which can reflect the content change situation (i.e., the change of pixel values) between two adjacent frame images. The residual block is the result of obtaining the inter-frame residual. The size of the residual block is the same as the size of the first image and the size of the second image. The residual block includes multiple residual points corresponding one by one to the positions of multiple pixel points of the first image. Each residual point has a residual value, and each residual value is the absolute value of the difference between the pixel values of the pixel points at the corresponding positions of the first image and the second image.
[0100] In the embodiments of the present application, there are multiple ways to obtain the residual block. In the embodiments of the present application, the following two ways are taken as examples to illustrate the ways to obtain the residual block:
[0101] The first implementable way is to calculate the inter-frame residual between the first image and the second image to obtain a residual block. The inter-frame residual is the absolute value of the difference between the pixel values of the pixel points at the corresponding positions of the first image and the second image. Assume that the first image is the t-th frame image in the video, t > 1, then the second image is the (t - 1)-th frame image.
[0102] In the first example, when the color spaces of the first image and the second image are RGB color spaces, the color coding format involved is RGB coding format. In this case, the transparency component values in the pixel values of the pixel points of the first image and the second image are usually ignored in the inter-frame residual. The pixel values of the pixel points of the first image and the second image include the values of the red component R, the green component G, and the blue component B. Then the inter-frame residual Residual includes Residual[R], Residual[G], and Residual[B]. Then the inter-frame residual Residual satisfies:
[0103] Residual[R] = Absdiff(R Frame(t-1) ,R Frame(t) ); (Formula 1)
[0104] Residual[G] = Absdiff(G Frame(t-1) ,GFrame(t) )); (Formula 2)
[0105] Residual[B]=Absdiff(B Frame(t-1) , B Frame(t) ). (Formula 3)
[0106] Where, Absdiff represents calculating the absolute value of the difference in pixel values (R value, G value, or B value) of corresponding pixels in two images; Frame(t) is the first image; Frame(t - 1) is the second image; R Frame(t-1) represents the R value of the red component of the second image, R Frame(t) represents the R value of the red component of the first image; G Frame(t-1) represents the G value of the green component of the second image, G Frame(t) represents the G value of the green component of the first image; B Frame(t-1) represents the B value of the blue component of the second image, B Frame(t) represents the B value of the blue component of the first image.
[0107] In the second example, when the color spaces of the first image and the second image are YUV color spaces, the color coding format involved is YUV coding format. In this case, the inter-frame residual between the first image and the second image is characterized by the inter-frame residual of the luminance component Y of the first image and the second image, then the inter-frame residual Residual = Residual[Y].
[0108] In an alternative way, the inter-frame residual Residual satisfies:[[]]
[0109] Residual = Residual[Y]=0.299·Residual[R]+0.587·Residual[G]+0.144·Residual[B]. (Formula 4)
[0110] Where, the obtaining methods of Residual[R], Residual[G], and Residual[B] respectively refer to the aforementioned Formulas 1 to 3, and the inter-frame residual Residual is obtained by converting the inter-frame residuals of the aforementioned three RGB color channels according to a certain ratio.
[0111] In another alternative way, the inter-frame residual Residual satisfies:[[]]
[0112] Residual = |Residual[Y1]-Residual[Y2]|; (Formula 5)
[0113] Among them, Residual[Y1] is the value of the luminance component of the first image, which is obtained by converting the RGB values of the first image according to a certain ratio; Residual[Y2] is the value of the luminance component of the second image, which is obtained by converting the RGB values of the second image according to a certain ratio. Among them, the certain ratio can be the ratio in Formula 4, that is, the ratios of the R value, G value, and B value are 0.299, 0.587, and 0.144 respectively.
[0114] Exemplarily, assume that the first image and the second image each include 5·5 pixel points, and the pixel value is represented by the luminance component Y. The pixel values of the pixel points of the first image are as Figure 3 shown, and the pixel values of the pixel points of the second image are as Figure 4 shown. Then the finally obtained residual block is as Figure 5 shown, including 5·5 residual points, and the residual value of each residual point is the absolute value of the difference between the pixel values of the pixel points at the corresponding positions in the first image and the second image.
[0115] It should be noted that the protection scope of the embodiments of the present application is not limited thereto. When the color coding format of the image is other formats, any person skilled in the art in the technical field disclosed in the embodiments of the present application can also easily think of transformation or replacement to calculate the inter-frame residual within the technical scope disclosed in the embodiments of the present application. Therefore, these easily conceivable changes or replacements are also covered by the protection scope of the embodiments of the present application.
[0116] The second implementable manner is to obtain the pre-stored inter-frame residual to obtain a residual block.
[0117] As described above, the first image obtained by the super-resolution display device is a decoded image. During the video decoding process, the calculation process of the inter-frame residual between two adjacent frames of images is involved. Therefore, during the video decoding process, the calculated inter-frame residual of each two adjacent frames of images can be stored, and when a residual block needs to be obtained, the pre-stored inter-frame residual can be directly extracted to obtain the residual block.
[0118] In the embodiments of the present application, the standard adopted by the processor for video decoding can be any one of H.261 to H.265, and MPEG-4V1 to MPEG-4V3, etc. Among them, H.264, also known as Advanced Video Coding (AVC), and H.265, also known as High Efficiency Video Coding (HEVC), both adopt the motion compensation hybrid coding algorithm.
[0119] Taking H.265 as an example, the encoding architecture of H.265 is generally similar to that of H.264, mainly including: entropy coding module, intra prediction module, inter prediction module, inverse transform module, inverse quantization module, loop filter module and other modules. The loop filter module includes deblocking module and Sample Adaptive Offset (SAO), etc. Among them, the entropy decoding module is used to process the bitstream provided by the video source to obtain mode information and inter-frame residual. After the entropy decoding module processes and obtains the inter-frame residual, the inter-frame residual can be stored to extract the inter-frame residual when step 201 is executed.
[0120] By obtaining the residual block through this second implementation method, the repeated calculation of the inter-frame residual can be reduced, the operation cost can be reduced, and the overall duration of image processing can be saved. Especially in the scenario where the image resolution of the video provided by the video source is relatively high, the image processing delay can be effectively reduced.
[0121] Step 202: The super-resolution display device detects whether the residual value of the residual point included in the residual block is 0. If it is determined that the residual value of at least one residual point included in the residual block is not 0, step 203 is executed. If it is determined that the residual values of all residual points included in the residual block are 0, step 207 is executed.
[0122] The super-resolution display device can traverse each residual point in the residual block to detect whether the residual value of each residual point is 0. When the residual value of at least one residual point in the residual block is not 0, step 203 can be executed. The subsequent steps 203 to 206 correspond to part of the super-resolution algorithm. When the residual values of all residual points in the residual block are 0, step 207 can be executed. For example, the super-resolution display device can traverse the residual points in the scanning order from left to right and from top to bottom.
[0123] Step 203: The super-resolution display device determines the target pixel area in the first image based on the residual block.
[0124] In the embodiments of the present application, the target pixel region is the region where the pixel points corresponding to the positions of the target residual points in the residual block are located in the first image, and the points corresponding to the positions of the target residual points in the first image are called target pixel points. Generally, the target residual points include two types of residual points: the first target residual point and the second target residual point. Optionally, the target pixel region includes the region where the pixel points corresponding to the position of the first target residual point are located in the first image. For example, the target pixel region is the region where the pixel points corresponding to the positions of the first target residual point and the second target residual point are located in the first image. Among them, the first target residual point is the point in the residual block where the residual value is greater than the specified threshold, and the second target residual point is the residual point around the first target residual point in the residual block. The residual points around the first target residual point refer to the residual points set around the first target residual point, which are the surrounding points that meet the specified conditions for the first target residual point. For example, the residual points above, below, left, and right of the first target residual point; or, the residual points above, below, left, right, top left, bottom left, top right, and bottom right of the first target residual point. Optionally, the aforementioned specified threshold is 0.
[0125] Exemplarily, the second target residual point is the residual point in the residual block where the residual value around the first target residual point is not greater than (i.e., less than or equal to) the specified threshold. That is, the second target residual point is the point among the surrounding points that meet the specified conditions for the first target residual point and whose residual value is not greater than the specified threshold. When the aforementioned specified threshold is 0, the second target residual point is the residual point with a residual value of 0 around the first target residual point in the residual block.
[0126] The region where the first target residual point is located is the region where the contents of the first image and the second image are different. Based on this region, the region in the first image that actually needs to be super-resolved can be found. Since the pixel values of the surrounding regions of a region in the first image are usually required to be referred to when performing super-resolution processing on a region of the first image, the surrounding region of the region in the first image that needs to be super-resolved also needs to be found. And the region where the second target residual point is located is the surrounding region of the region where the first target residual point is located. By determining the second target residual point, the surrounding region of the region in the first image that actually needs to be super-resolved can be determined, so as to adapt to the requirements of super-resolution processing and ensure effective super-resolution processing in the subsequent process. The aforementioned specified conditions are determined based on the requirements of super-resolution processing, for example, set based on the receptive field (such as the receptive field of the last convolutional layer) in the super-resolution model.
[0127] Assume that the first target residual point is the point in the residual block where the residual value is greater than the specified threshold, the second target residual point is the residual point in the residual block where the residual value around the first target residual point is not greater than the specified threshold, and the specified threshold is 0. Take Figure 6 the shown residual block as an example. Figure 6 It is Figure 5Schematic diagram for explaining the principle of the residual block shown. The area where the target residual points of the residual block are located is the area composed of area K and area M. Among them, area K includes the first target residual point P1 and the second target residual points P7 to P14 around it; area M includes the first target residual points P2 to P6 and the second target residual points P15 to P19 around them. Therefore, the finally determined first target residual points are P1 to P6, and the second target residual points include P7 to P19. Then the target pixel points determined in the first image are the pixel points with the same positions as the residual points P1 to P19. Figure 6 Taking the surrounding points that meet the specified conditions as the upper, lower, left, right, upper left, lower left, upper right, and lower right residual points of each first target residual point as an example for illustration, but it is not limited thereto.
[0128] In the embodiments of the present application, multiple methods can be used to determine the target pixel area. The embodiments of the present application will be described by taking the following two determination methods as examples:
[0129] In the first determination method, the target pixel area is determined inside the super-resolution model. As Figure 7 shown, the process of determining the target pixel area in the first image based on the residual block includes:
[0130] Step 2031: The super-resolution display device generates a mask graph based on the residual block.
[0131] This mask graph is used to indicate the position of the target pixel area and plays a guiding role in screening the target pixel area in the first image. The mask graph includes multiple first mask points, and the positions of the multiple first mask points correspond one by one to the positions of the multiple target residual points in the residual block. The multiple target residual points at least include the first target residual points. Usually, the multiple target residual points include the first target residual points and the second target residual points; that is, the multiple first mask points are used to identify the positions of the multiple target residual points in the residual block. Since the positions of the multiple target residual points of the residual block correspond one by one to the positions of the multiple target pixel points in the target pixel area of the first image, the multiple first mask points are used to identify the positions of the multiple target pixel areas in the first image. The target pixel area where the target pixel points are located can be found through the first mask points.
[0132] In the embodiments of the present application, the following two optional implementation methods are used to schematically illustrate step 2031:
[0133] In the first optional implementation, morphological transformations can be performed on the residual block to obtain a mask pattern. The morphological transformation includes binarization and dilation. For example, first perform binarization on the residual block to obtain a binarized residual block; then perform dilation on the binarized residual block to obtain a dilated residual block, and use the dilated residual block as the mask pattern.
[0134] Among them, binarization is a processing method that sets the pixel value of each pixel point in the image to one of two pixel values, a first value and a second value, and the first value and the second value are different. After the image undergoes binarization processing, it only includes pixel points with two pixel values. This binarization processing can reduce the interference of various elements in the image on the subsequent image processing process.
[0135] The residual value of each residual point in the residual block is the absolute value of the difference between the pixel values of two pixel points. The residual block is equivalent to the difference image between two images. Therefore, the residual block can also be regarded as an image, the residual points it contains are equivalent to the pixel points of the image, and the residual value of the residual point is equivalent to the pixel value of the pixel point.
[0136] For example, in the RGB color space, since the residual value of each residual point includes the values of three color components, R, G, and B, that is, the aforementioned Residual[R], Residual[G], and Residual[B], the residual value of the residual point can be characterized by its grayscale value to simplify the calculation process. The grayscale value of the residual point is used to reflect the brightness and darkness of the residual point, and it can be obtained by converting the R, G, and B values of the residual point, that is, converting the residual value of the residual point, such as the aforementioned Residual[R], Residual[G], and Residual[B]. This conversion process can refer to the traditional process of converting R, G, and B values to grayscale values, and this application will not elaborate on it. The grayscale value range of the residual point is generally from 0 to 255, the grayscale value of the white residual point is 255, and the grayscale value of the black residual point is 0. When performing binarization on the residual block, it can be determined whether the residual value of each residual point in the residual block is greater than the binarization threshold (for example, this binarization threshold can be a fixed value or a variable value. When it is a variable value, the binarization threshold can be determined by using the method of local adaptive binarization). When the residual value of a certain residual point is greater than the binarization threshold, set the residual value of this residual point to the first value; when the residual value of a certain residual point is less than or equal to the binarization threshold, set the residual value of this residual point to the second value.
[0137] Exemplarily, in the YUV color space, since the residual value of each residual point includes a value, i.e., the aforementioned Residual[Y], when performing binarization processing on the residual block, it can be determined whether the residual value of each residual point in the residual block is greater than the binarization threshold (for example, the binarization threshold can be a fixed value or a variable value. When it is a variable value, a local adaptive binarization method can be used to determine the binarization threshold). When the residual value of a certain residual point is greater than the binarization threshold, the residual value of this residual point is set to a first value. When the residual value of a certain residual point is less than or equal to the binarization threshold, the residual value of this residual point is set to a second value.
[0138] Exemplarily, the aforementioned first value is a non-zero value, and the second value is 0. For example, the aforementioned first value is 255 and the second value is 0; or the first value is 1 and the second value is 0. In actual implementation, the smaller the value, the less storage space is occupied. Therefore, usually the first value is set to 1 and the second value is set to 0 to save storage space.
[0139] The binarization threshold can be determined based on the aforementioned specified threshold. For example, in the RGB color space, since the grayscale value of the residual point compared with the binarization threshold is converted from the residual value, the binarization threshold is obtained by converting the specified threshold using the same conversion rule; in the YUV color space, since the residual value of the residual point is compared with the binarization threshold, the binarization threshold is equal to the specified threshold. Exemplarily, the specified threshold can be 0, and correspondingly, the binarization threshold can be 0.
[0140] In the binarized residual block obtained by the aforementioned method, there are only two values, which reduces the subsequent calculation complexity; and the residual point corresponding to the first value is the aforementioned first target residual point, and the first target residual point can be simply and quickly located in the subsequent processing.
[0141] The binarization processing in the embodiments of the present application can adopt any method among the global binarization threshold method, the local binarization threshold method, the maximum between-class variance method, and the iterative binarization threshold method, and the embodiments of the present application do not limit this.
[0142] Dilation processing is a process of finding local maxima. The image to be processed is convolved with a preset kernel (also called a core). In each convolution process, the maximum value in the kernel coverage area is assigned to the specified pixel point, so that the bright ones become brighter, and the obtained effect is that the bright area of the image to be processed expands. Among them, the kernel has a definable anchor point, which is usually the center point of the kernel, and the aforementioned specified pixel point is this anchor point.
[0143] As Figure 8 shown, Figure 8Assume the image to be processed is F1. The image to be processed includes 5·5 pixels, where the shaded areas represent bright points. The kernel is the shaded part in F2, which consists of 5 pixels in total. The anchor point is the center point B of these 5 pixels. Then the image after the final dilation processing is F3. Figure 8 In this, "*" represents convolution operation.
[0144] In the embodiments of this application, the image to be processed is a residual block. Performing dilation processing on the residual block means performing dilation processing on the first target residual point of the binarized residual block. To meet the requirements of the aforementioned dilation processing, if the aforementioned first value is non-zero and the second value is 0, then the first target residual point is a bright point, and other residual points (i.e., points other than the first target residual point) are dark points, and the aforementioned dilation processing can be directly performed to achieve the dilation of the area where the first target residual point is located; if the aforementioned first value is 0 and the second value is non-zero, then the first value can be updated to a non-zero value and the second value can be updated to 0 through a specified algorithm. In this way, the first target residual point is a bright point, and other residual points are dark points, and then the aforementioned dilation processing is performed. It should be noted that, please refer to Figure 8 and Figure 9 , if the residual block is F1 and the residual block after dilation processing is F3, then the residual points corresponding to the diagonal shaded areas in F1 and F3 are the first target residual points, and the residual points corresponding to the "×" shaped shaded areas in F3 are the second target residual points. Figure 9 The kernel in Figure 8 is F4, which is different from F2 in
[0145] In the embodiments of this application, the super-resolution model includes at least one convolution kernel, and the kernel for dilation processing has the same size as the receptive field of the last convolution layer of the super-resolution model. This last convolution layer is the output layer of the super-resolution model, and the image after the super-resolution processing of the super-resolution model is output from this layer. The receptive field of this layer is the largest receptive field among the receptive fields corresponding to each convolution layer in the super-resolution model. In this way, the obtained mask pattern is adapted to the size of the largest receptive field of the super-resolution model, thus playing a good guiding role and avoiding the situation where the convolvable area of the image input to the super-resolution model later is too small, resulting in the image not being convolvable by the convolution layer, and ensuring that the image input to the super-resolution model later can be effectively super-resolved. For example, the size of this kernel can be 3·3 or 5·5 pixels.
[0146] In the second optional implementation manner, due to the principle of the aforementioned dilation processing, which is similar to the algorithm for finding the m-neighborhood, the neighborhood refers to an open interval centered on the target point. The algorithm for finding the m-neighborhood means obtaining an open interval of m points adjacent to the target point centered on the target point. m is the size of the receptive field of the last convolution layer of the super-resolution model minus 1. For example, if the receptive field of the last convolution layer of the super-resolution model is 5·5 pixels, then m = 8. As described aboveFigure 8 and Figure 9 as shown Figure 8 The dilation process of [[ ]] is equivalent to finding the 4-neighborhood for each first target residual point in F1. Figure 9 The dilation process of [[ ]] is equivalent to finding the 8-neighborhood for each first target residual point in F1. In the embodiments of the present application, since the receptive field of the last convolutional layer of the super-resolution model is usually 3·3 or 5·5 pixel points, therefore, m = 8 or 24. Among them, when m = 8, the 8-neighborhood refers to 8 points including the upper, lower, left, right, upper left, lower left, upper right, and lower right centered on the target point; when m = 24, the 24-neighborhood refers to 8 points including the upper, lower, left, right, upper left, lower left, upper right, and lower right centered on the target point, and 16 points in a circle around these 8 points, a total of 24 points.
[0147] Then the above step 2031 includes: first performing binarization processing on the residual block to obtain the binarized residual block; then finding the m-neighborhood for each first target residual point in the binarized residual block, and filling the residual value (i.e., the aforementioned first value) of each target residual point into the corresponding m-neighborhood to obtain the mask pattern. The process of binarization processing can refer to the foregoing implementation manner, and the method for obtaining the m-neighborhood can refer to the related art, which will not be elaborated in the embodiments of the present application.
[0148] It should be noted that in the embodiments of the present application, the process of generating the mask pattern based on the residual block in the foregoing step 2031 can also be implemented in other ways. For example, it includes: generating an initial mask pattern including a plurality of mask points based on the residual block, the plurality of mask points corresponding one-to-one to the positions of a plurality of pixel points of the first image, and the plurality of mask points including a plurality of first mask points and a plurality of second mask points. That is, the initial mask pattern is a pattern in which the first mask points and the second mask points are determined; assigning the mask value of the first mask points in the initial mask pattern to the first value, and assigning the mask value of the second mask points in the mask pattern to the second value to obtain the mask pattern, where the first value and the second value are different. The mask pattern obtained in this way is a binary image. Subsequently, when locating the target pixel points, the first mask points can be located by searching for the first value, and then the target pixel points can be indexed, which can realize the rapid determination of the target pixel points.
[0149] In the embodiments of the present application, the residual block can be directly processed as a whole to generate the foregoing mask pattern, or the residual block can be divided into multiple sub-residual blocks, and each sub-residual block is processed to obtain a sub-mask pattern, and the generated sub-mask patterns form the mask pattern. Since the mask pattern is generated in blocks, the computational complexity can be reduced and the computational cost can be lowered. When the computing power of the processor is strong, the generation processes of multiple sub-mask patterns can be executed simultaneously to save the image processing time.
[0150] Exemplarily, the steps of generating a mask pattern based on a residual block may include:
[0151] Step A1: Divide the residual block into multiple sub-residual blocks, and perform a block processing on each obtained sub-residual block. The block processing includes: when at least one residual point in the sub-residual block has a non-zero residual value (i.e., the sub-residual block includes a residual point with a non-zero residual value), divide the sub-residual block into multiple sub-residual blocks, and perform a block processing on each obtained sub-residual block, until the residual values of the residual points included in the obtained sub-residual block are all 0 (i.e., the sub-residual block does not include a residual point with a non-zero residual value), or the total number of residual points in the obtained sub-residual block is less than the point threshold, or the total number of divisions of the residual block reaches the number threshold.
[0152] Among them, when the residual values of the residual points included in the obtained sub-residual block are all 0, it indicates that there is no content change in the regions of the first image and the second image corresponding to the sub-residual block, and there is no need to continue the block division; when the total number of residual points in the obtained sub-residual block is less than the point threshold, it indicates that the size of the obtained sub-residual block is small enough. If further division is performed, on the one hand, if the corresponding target image block is input to the super-resolution model in the subsequent process (refer to subsequent step B2), it is likely to cause the super-resolution model to be unable to perform effective super-resolution and affect the super-resolution effect of the super-resolution model. On the other hand, the size of the sub-residual block is too small, which is likely to cause too large an operation cost. Therefore, there is no need to continue the division. When the total number of divisions of the residual block reaches the number threshold, if further division is continued, on the one hand, it may cause too large an operation cost due to too many divisions, and on the other hand, it may cause the size of the obtained sub-residual block to be too small, affecting the super-resolution effect. Therefore, there is no need to continue the division, thereby reducing the operation cost and ensuring the super-resolution effect. Exemplarily, the foregoing point threshold and number threshold may be determined based on the image resolution of the first image, for example, being positively correlated with the image resolution of the first image. That is, the larger the image resolution, the larger the point threshold and the number threshold. Usually, the number threshold may be 2 or 3 times.
[0153] By performing the foregoing step A1, cyclic division of the residual block can be achieved, and finally multiple sub-residual blocks are obtained. When the computing power of the processor is strong, when multiple sub-residual blocks all need to be divided, the division processes of multiple sub-residuals can be executed simultaneously to save the image processing time.
[0154] Among them, both the residual block and the sub-residual block can be divided into blocks in the way of binary tree division or quadtree division. The residual block or sub-residual block divided by binary tree is divided into 2 sub-residual blocks with equal or unequal sizes each time; the residual block or sub-residual block divided by quadtree is divided into 4 sub-residual blocks with equal or unequal sizes each time. It should be noted that the residual block and the sub-residual block can also have other division methods, as long as effective block division can be achieved, and the embodiments of the present application do not limit this.
[0155] Since the traditional video decoding process requires image block division, and this block division usually adopts the quadtree division method, in the embodiments of the present application, when the residual block and the sub-residual block adopt the quadtree division method, they can be compatible with the traditional image processing method. For example, in practical applications, the residual block and the sub-residual block can also be divided through the image division module used in the foregoing video decoding process, so as to achieve the reuse of the module and save the operation cost. As mentioned above, generally, the image resolution of images in a video is 360p, 480p, 720p, 1080p, 2k, 4k, etc., all of which are integer multiples of 4. Therefore, by adopting the quadtree division method, each time of division can divide the residual block or sub-residual block into four sub-residual blocks with equal sizes, that is, achieve the four-equal division of the residual block or sub-residual block, making the sizes of the divided residual blocks uniform and facilitating subsequent processing. Of course, after the residual block is divided multiple times, there may be a situation where the sub-residual block cannot be four-equal divided. It can be divided as evenly as possible so that the size difference between any two of the four sub-residual blocks after each division is less than or equal to the specified difference threshold, which will not affect the function of the finally obtained sub-mask pattern either.
[0156] Step A2, generate a sub-mask pattern corresponding to each target residual block, and the foregoing mask pattern includes the generated sub-mask pattern. Among them, the target residual block includes at least one residual point with a non-zero residual value.
[0157] In the embodiments of the present application, since the residual values of the residual points included in other residual blocks except the target residual block are all 0, it means that there is no content change in the regions of the first image and the second image corresponding to this sub-residual block, and there is no need to perform super-resolution on the corresponding region of this sub-residual block of the first image. Therefore, there is also no need to obtain a sub-mask pattern to guide the target pixel region. In this way, compared with generating the mask pattern as a whole, the method of generating the mask pattern in blocks can reduce the processing of other residual blocks except the target residual block, and only generate a sub-mask pattern for the target residual block, reducing the operation cost.
[0158] Among them, the generation method of each sub-mask pattern can refer to the two optional implementation methods in the foregoing step 2031. For example, morphological transformation processing is performed on each target sub-residual block to obtain a sub-mask pattern corresponding to each target sub-residual block. Specifically, first perform binarization processing on the target sub-residual block to obtain the binarized target sub-residual block; then perform dilation processing on the binarized target sub-residual block to obtain the dilated target sub-residual block, and use the dilated target sub-residual block as the sub-mask pattern. Another example is to first perform binarization processing on each target sub-residual block to obtain the binarized target sub-residual block; then find the m-neighborhood of each first target residual point in each binarized target sub-residual block, and fill the residual value of each target residual point into the corresponding m-neighborhood to obtain the corresponding sub-mask pattern. The processing process of each target sub-residual block can refer to the foregoing Figure 8 and Figure 9 , which will not be elaborated in this embodiment of the present application.
[0159] Optionally, the size of each of the foregoing sub-mask patterns is the same as the size of the corresponding target sub-residual block.
[0160] Step 2032: The super-resolution display device inputs the mask pattern and the first image into the super-resolution model, and determines, through the super-resolution model, the area where the pixel points corresponding to the positions of multiple first mask points in the first image are located as the target pixel area.
[0161] In step 2032, the super-resolution display device determines the target pixel area through the super-resolution model. The traditional super-resolution model only performs super-resolution processing on the received image. In this embodiment of the present application, code for determining the target pixel area can be added to the front end (i.e., the input end) of the traditional super-resolution model to implement the determination of the target pixel area. In this way, the super-resolution display device only needs to input the mask pattern and the first image into the super-resolution model, reducing the computational complexity of the modules outside the super-resolution model in the super-resolution display device.
[0162] Referring to the foregoing embodiments, the mask pattern includes multiple mask points, and the multiple mask points correspond one-to-one to the positions of multiple pixel points of the first image, and each mask point has a mask value. In an optional example, the mask pattern is a binarized image, and the multiple mask points include multiple first mask points and multiple second mask points. The mask value of the first mask point is the first value, and the mask value of the second mask point is the second value, and the first value and the second value are different. As described above, the first value and the second value can be one of a non-zero value and 0 respectively. In this embodiment of the present application, it is assumed that the first value is a non-zero value (such as 1) and the second value is 0. In another optional example, the mask pattern is a monochromatic image, and the multiple mask points only include multiple first mask points.
[0163] In step 2032, the super-resolution model traverses the mask points in the mask pattern. In the first image, the area where the pixel points corresponding to the mask points with the mask value being the first value (i.e., the target pixel points) are located is determined as the target pixel area.
[0164] In the second determination method, the target pixel area is determined outside the super-resolution model. As Figure 10 shown, the process of determining the target pixel area in the first image includes:
[0165] Step 2033: The super-resolution display device generates a mask pattern based on the residual block.
[0166] The mask pattern includes a plurality of first mask points, and the positions of the plurality of first mask points correspond one-to-one to the positions of the plurality of target residual points in the residual block. The process of step 2033 can refer to the process of the foregoing step 2031, and the embodiments of the present application will not elaborate on this.
[0167] Step 2034: The super-resolution display device determines the area where the pixel points corresponding to the positions of the plurality of first mask points (i.e., the target pixel points) in the first image are located as the target pixel area.
[0168] Referring to the foregoing step 2032, the super-resolution display device can traverse the mask points in the mask pattern, and in the first image, determine the pixel points corresponding to the mask points with the mask value being the first value as the target pixel area.
[0169] In the foregoing two determination methods, the super-resolution display device determines the target pixel area from the first image under the guidance of the mask pattern, shields the pixel points other than the target pixel area, and realizes the rapid positioning of the target pixel area, thereby effectively saving the processing time of the image.
[0170] In the embodiments of the present application, the mask pattern can have various shapes. In the first example, the mask pattern only includes a plurality of first mask points, that is, the mask pattern is composed of a plurality of first mask points, and the obtained mask pattern is usually an irregular pattern; in the second example, the mask pattern only includes a plurality of sub-mask patterns, that is, the mask pattern is composed of a plurality of sub-mask patterns, and the size of each sub-mask pattern is the same as that of the corresponding target residual block, that is, it includes both first mask points and second mask points, and the obtained mask pattern is usually an irregular pattern formed by splicing sub-mask patterns; in the third example, the size of the mask pattern is the same as that of the first pattern, and it includes both first mask points and second mask points. Usually, the memory in the super-resolution display device stores graphic data in the form of a one-dimensional or multi-dimensional array, etc. If the mask pattern is the mask pattern in the first example above, the data granularity to be stored is the pixel-level data granularity, and the storage complexity is relatively high; if the mask pattern is the mask pattern in the second example above, the data granularity to be stored is the pixel block-level (the size of one pixel block is the size of one sub-mask pattern above) data granularity, and the stored graphic is more regular than that in the first example, and the storage complexity is relatively low; if the mask pattern is the mask pattern in the third example above, the stored graphic is a rectangle, which is more regular than the first example and the second example, and the storage complexity is relatively low. Therefore, usually the mask pattern is usually in the shapes of the second and third examples above.
[0171] Step 204, the super-resolution display device performs super-resolution processing on the target pixel region in the first image to obtain the super-resolution processed target pixel region.
[0172] In the embodiments of the present application, the super-resolution display device can perform super-resolution processing on the target pixel region in the first image through a super-resolution model to obtain the super-resolution processed target pixel region. Referring to the foregoing step 203, since there are various ways to determine the target pixel region, correspondingly, there are various ways to perform super-resolution processing. The embodiments of the present application will be described by taking the following two processing methods as examples:
[0173] The first processing method, corresponding to the first determination method in the foregoing step 203, the process of the super-resolution display device performing super-resolution processing on the target pixel region in the first image to obtain the super-resolution processed target pixel region can be: the super-resolution display device performs super-resolution processing on the target pixel region in the first image through a super-resolution model to obtain the super-resolution processed target pixel region.
[0174] Exemplarily, referring to step 2032, since the super-resolution display device inputs the mask pattern and the first image into the super-resolution model and determines the target pixel region through the super-resolution model, the super-resolution model can continue to perform super-resolution processing on the target pixel region in the first image to obtain the super-resolution processed target pixel region. The process of the super-resolution model performing super-resolution processing can refer to related technologies, and the embodiments of the present application will not elaborate on this.
[0175] For the second processing method, corresponding to the second determination method in the foregoing step 203, the process of the super-resolution display device performing super-resolution processing on the target pixel region in the first image to obtain the super-resolution processed target pixel region can be: the super-resolution display device inputs the target pixel region in the first image into the super-resolution model, and performs super-resolution processing on the target pixel region in the first image through the super-resolution model to obtain the super-resolution processed target pixel region.
[0176] In the embodiments of the present application, multiple target image blocks can be screened out from the first image to perform super-resolution processing on each target image block. Since the super-resolution processing is performed on the target image blocks, and the size of each target image block is smaller than that of the first image, the computational complexity of the super-resolution processing can be reduced, and the computational cost can be reduced. Especially when the super-resolution processing is performed by the super-resolution model, the complexity of the super-resolution model can be effectively reduced, and the efficiency of the super-resolution operation can be improved.
[0177] Exemplarily, step 204 may include:
[0178] Step B1: The super-resolution display device obtains the target image block corresponding to each sub-mask pattern in the first image.
[0179] In the first image, the image block corresponding to the position of each sub-mask pattern is determined as the target image block.
[0180] Step B2: The super-resolution display device respectively performs super-resolution processing on the sub-regions of the target pixel regions included in each target image block to obtain the super-resolution processed target pixel regions.
[0181] Referring to the foregoing content, it can be known that the mask pattern is used to indicate the position of the target pixel region. Since the mask pattern is divided into multiple sub-mask patterns, the target pixel region can also be divided into multiple sub-regions corresponding to the multiple sub-mask patterns. Also, since each sub-mask pattern corresponds to a target image block, each target image block contains a sub-region of the target pixel region, that is, the multiple target image blocks obtained above correspond one-to-one to the multiple sub-regions. Therefore, the super-resolution processed target pixel region is composed of the sub-regions of the target pixel regions included in each super-resolution processed target image block.
[0182] Exemplarily, the super-resolution process can be performed by a super-resolution model. Corresponding to the first processing method in step 204, referring to the foregoing step 2032, the super-resolution model can first determine the target pixel region and the corresponding region of the target image block, that is, the sub-region of the target pixel region included in each target image block, and then perform the super-resolution process. Then the foregoing step 2032 can specifically include: the super-resolution display device inputs each sub-mask pattern and the corresponding target image block into the super-resolution model respectively (each input is a target image block and a corresponding sub-mask pattern), and through the super-resolution model, the region where the pixel points corresponding to the multiple first mask point positions of the corresponding sub-mask pattern in the target image block is determined as the sub-region of the target pixel region included in the target image block; step B2 can specifically include: through the super-resolution model, perform super-resolution processing on the sub-regions of the target pixel regions included in each target image block respectively, to obtain the sub-regions of the target pixel regions included in each target image block after super-resolution processing. The target pixel region after super-resolution processing is composed of the sub-regions of the target pixel regions included in each target image block after super-resolution processing. In this way, the size of the image input into the super-resolution model each time is small, which can effectively reduce the computational complexity of the super-resolution model. Thus, a super-resolution operation can be achieved by using a super-resolution model with a relatively simple structure, reducing the complexity of the super-resolution model and improving the efficiency of the super-resolution operation.
[0183] Corresponding to the second processing method in step 204, step B2 can specifically include: the super-resolution display device inputs the sub-regions of the target pixel regions included in each target image block into the super-resolution model respectively, and through the super-resolution model, perform super-resolution processing on the sub-regions of the target pixel regions included in each target image block respectively, to obtain the sub-regions of the target pixel regions included in each target image block after super-resolution processing. The target pixel region after super-resolution processing is composed of the sub-regions of the target pixel regions included in each target image block after super-resolution processing.
[0184] After receiving the sub-regions of the target pixel regions included in each target image block after super-resolution processing output by the super-resolution model, the super-resolution display device can splice (also referred to as combine) them according to the positions of the respective target image blocks in the first image to obtain the sub-region of the target pixel region after super-resolution processing. Then the finally obtained target pixel region after super-resolution processing is composed of the sub-regions of the target pixel regions included in each target image block after super-resolution processing.
[0185] It should be noted that the number of pixel points in the target pixel region after super-resolution processing is greater than the number of pixels in the target pixel region before super-resolution processing, that is, the super-resolution processing realizes an increase in the pixel density in the target pixel region, thereby realizing the super-resolution effect. For example, the target pixel region before super-resolution processing includes 5·5 pixel points, and the target pixel region after super-resolution processing includes 9·9 pixel points.
[0186] Step 205: The super-resolution display device determines other pixel regions in the first image.
[0187] Among them, the other pixel regions include pixel regions other than the target pixel region. In the first optional method, the other pixel region is the region where the pixel points in the first image other than the target pixel region are located (it can also be understood as directly performing step 206 without performing the so-called "determination" step); in the second optional method, the other pixel region is the region where the pixel points in the first image other than the pixel points corresponding to the first mask points are located; in the third optional method, the other pixel region is the region where the pixel points in the first image other than the auxiliary pixel points are located. The region where the auxiliary pixel points are located includes the target pixel region, but the number of auxiliary pixel points in the first image is greater than the number of pixel points (i.e., target pixel points) in the target pixel region, that is, the area of the region where the auxiliary pixel points are located is greater than the area of the target pixel region.
[0188] When adopting the third optional method, for example, step 205 may include: performing erosion processing on multiple first mask points in the mask pattern to obtain an updated mask pattern; determining the pixel points in the first image corresponding to the positions of the multiple first mask points after the erosion processing as auxiliary pixel points; and determining the region where the pixel points in the first image other than the auxiliary pixel points are located as the other pixel region.
[0189] Among them, the erosion processing is a process of finding local minimum values. Convolving the image to be processed with a preset kernel (also called a core). In each convolution process, the minimum value in the region covered by the kernel is assigned to the specified pixel point, and the effect obtained is that the bright region of the image to be processed shrinks. The kernel has a definable anchor point, which is usually the center point of the kernel, and the aforementioned specified pixel point is the anchor point. It should be noted that the aforementioned dilation processing and this erosion processing do not have a reciprocal relationship.
[0190] As Figure 11 shown, Figure 11 Suppose the image to be processed is F5, the image to be processed includes 5·5 pixel points, the shaded part therein represents the bright points, the kernel is the shaded part in F6, a total of 5 pixel points, and the anchor point is the center point D of these 5 pixel points. Then the finally eroded image is F7. Figure 11 In "*" represents convolution.
[0191] It is worth noting that, please refer to Figure 11 , if the mask pattern is F5 and the eroded mask pattern is F7, then the mask points corresponding to the diagonal shaded areas in F5 and F7 are the first mask points.
[0192] As described above, the super-resolution model includes at least one convolutional kernel. The kernel for erosion processing has the same size as the receptive field of the last convolutional layer of the super-resolution model, that is, the kernel for erosion processing has the same size as the kernel for the aforementioned dilation processing. The region corresponding to the first mask point in the updated mask pattern obtained in this way is actually the region corresponding to the first mask point in the mask pattern obtained by the aforementioned dilation processing after removing the outermost layer of first mask points (that is, the inner contraction of the region corresponding to the first mask point is achieved). Compared with the mask pattern obtained after the aforementioned binarization processing, the edge of the updated mask pattern is smoother and has less noise.
[0193] In the embodiments of the present application, edge noise of an image can be eliminated through erosion processing. For other pixel regions determined by using the updated mask pattern obtained through erosion processing, compared with the other pixel regions obtained by the aforementioned first and second optional methods, the edges are clearer and the noise is less. When performing the pixel update process in step 206 subsequently, negative effects such as detail blurring, edge blunting, graininess, and noise enhancement can be reduced, ensuring the display effect of the first image after the final super-resolution processing.
[0194] In actual implementation of the embodiments of the present application, other processing methods can also be used to update the mask pattern, as long as the region corresponding to the first mask point in the updated mask pattern is the region corresponding to the first mask point in the mask pattern obtained by the aforementioned dilation processing after removing the outermost layer of first mask points, that is, the same effect as the erosion processing can be achieved. The embodiments of the present application do not limit this.
[0195] Step 206: The super-resolution display device updates other pixel regions in the first image by using other pixel regions in the second image after super-resolution processing.
[0196] Since the number of pixel points in other pixel regions in the second image after super-resolution processing is greater than the number of pixel points in other pixel regions in the first image, the number of pixel points in the updated other pixel regions is greater than the number of pixel points in the other pixel regions before the update, that is, the update process realizes an increase in the pixel density in the other pixel regions. The display effect of the updated other pixel regions is the same as the display effect of the other pixel regions after super-resolution processing. Therefore, the updated other pixel regions are equivalent to the other pixel regions after super-resolution processing.
[0197] Such as Figure 15As shown, assume that the other pixel region K1 in the second image after super-resolution processing includes 12·12 pixel points, and the other pixel region K2 in the first image includes 6·6 pixel points. Each square represents a pixel point. The sizes of the other pixel region K1 and the other pixel region K2, as well as their positions in their respective figures, are the same. Updating the other pixel region K2 in the first image with the other pixel region K1 in the second image after super-resolution processing means updating the pixel value, pixel position, and other pixel data of the pixels in the other pixel region K2 with the pixel value, pixel position, and other pixel data of the corresponding pixels in the other pixel region K1. Then referring to Figure 12 , the number of pixel points, pixel values, and pixel positions in the other pixel region K2 in the updated first image are correspondingly the same as those in the other pixel region K1 in the second image.
[0198] The first image after super-resolution processing, that is, the reconstructed first image, includes the target pixel region after super-resolution processing obtained through step 204 and the updated other pixel region obtained through step 206. The display effect of this first image after super-resolution processing is the same as that of the traditional first image after super-resolution processing.
[0199] Since this other pixel region includes the pixel region in the first image except the target pixel region, the sizes of the target pixel region and the other pixel region may or may not match. Correspondingly, the methods for obtaining the first image after super-resolution processing are also different. The embodiments of the present application will be described by taking the following several optional methods as examples:
[0200] In an alternative approach, the size of the target image region matches that of other pixel regions, that is, the other pixel regions are the pixel regions in the first image except the target pixel region; correspondingly, the sizes of the super-resolved target image region and the updated other pixel regions also match. Then, the super-resolved first image can be formed by stitching together the super-resolved target pixel region and the updated other pixel regions; in another alternative approach, the size of the target image region does not match that of other pixel regions, and there is an overlapping region at their edges, that is, the other pixel regions include additional pixel regions on the basis of the pixel regions in the first image except the target pixel region. Correspondingly, the sizes of the super-resolved target image region and the updated other pixel regions also do not match, and there is an overlapping region at their edges. Since the other pixel regions of the first image are updated from the other pixel regions of the super-resolved second image, the pixel data of the included pixels are usually more accurate. Therefore, the pixel data of the overlapping region of the super-resolved first image is based on the pixel data of the updated other pixel regions of the first image. Then, the super-resolved first image can be formed by stitching together the updated target pixel region and the updated other pixel regions. For example, the updated target pixel region is obtained by subtracting (also known as removing) its overlapping region with the other pixel regions from the super-resolved target pixel region. The updated target pixel region is shrunk relative to the target pixel region before the update, and the size of the updated target pixel region matches the size of the updated other pixel regions.
[0201] It should be noted that in practical applications, it is also possible not to consider whether the size of the target image region matches that of other pixel regions, and directly obtain the super-resolved first image in the following way:
[0202] In an alternative approach, the super-resolved second image can be used as the background image, and the super-resolved target pixel region obtained in step 204 is used to cover the corresponding region of the second image to obtain the super-resolved first image; in another alternative approach, the first image or a blank image can be used as the background image, the super-resolved target pixel region obtained in step 204 is used to cover the corresponding region of the first image, and the other pixel regions of the super-resolved second image are used to cover the corresponding region of the first image to obtain the super-resolved first image. As long as it is ensured that the final super-resolved first image is consistent with the image obtained by performing overall super-resolution processing on the first image or the difference is within an acceptable range, the embodiments of the present application do not make any limitations in this regard.
[0203] In the embodiments of the present application, the process of generating the super-resolved first image based on the obtained super-resolved target pixel region and the obtained updated other pixel regions can also be implemented by the aforementioned super-resolution model.
[0204] Exemplarily, the process of generating the first image after super-resolution processing can satisfy:
[0205] R(F(t)) = R(F(t - 1)(w, h))[Mask(w, h) = R2] + SR(F(t)(w, h))[Mask(w, h) = R1];
[0206] Wherein, R(F(t)) represents the first image after super-resolution processing, (w, h) represents any point in the image; Mask(w, h) represents the mask point (w, h) in the mask pattern (such as the updated mask pattern in step 205 above), and SR represents super-resolution processing. R(F(t - 1)(w, h))[Mask(w, h) = R2] represents the pixel region where the pixel point (w, h) corresponding to the mask point (w, h) with the second value in the mask pattern in the second image after super-resolution processing is located, that is, the other pixel region determined in step 205 above; SR(F(t)(w, h))[Mask(w, h) = R1] represents the pixel region where the pixel point corresponding to the mask point (w, h) with the first value in the mask pattern in the region after super-resolution processing of the first image (i.e., the target pixel region determined in step 203) is located, that is, all or part of the region in the target pixel region after super-resolution processing above (for example, if the mask pattern update operation described in step 205 is not performed, it is the entire region; if the mask pattern update operation described in step 205 is performed, it is a partial region. This situation is equivalent to the target pixel region determined in step 203 being updated with the update of the mask pattern, and this partial region is the updated target pixel region above, and this partial region can be obtained by subtracting the overlapping region between it and the other pixel region from the target pixel region determined by super-resolution processing in step 203). R1 represents the first value, and R2 represents the second value. Exemplarily, the first value is 1 and the second value is 0; or, the first value is 255 and the second value is 0.
[0207] Step 207, the super-resolution display device updates the first image with the second image after super-resolution processing.
[0208] Since the inter-frame residuals corresponding to the residual blocks are used to reflect the content change situation between two adjacent frames of images, when the residual values of all the residual points in the residual blocks obtained based on the first image and the second image are 0, it indicates that the content of the first image and the second image has not changed. Then, similarly, the images after super-resolution processing of the two should also not change. Therefore, the first image is updated with the second image after super-resolution processing to obtain the first image after super-resolution processing.
[0209] It should be noted that the second image after super-resolution processing is an image determined by using the image processing method provided in the embodiments of the present application, or a traditional image processing method, or other super-resolution processing methods. Updating the first image with the second image after super-resolution processing can make the number of pixel points in the updated first image greater than the number of pixel points in the first image before updating, that is, the update process realizes an increase in the pixel density in the first image, thereby achieving the effect of super-resolution. Therefore, the updated first image is also a super-resolution image, that is, the updated first image is equivalent to the first image after super-resolution processing.
[0210] It should be noted that when performing step 203, the super-resolution display device can also count the first proportion of the number of residual points with a residual value of 0 in the total number of residual points in the residual block; when the first proportion is greater than the first super-resolution trigger proportion threshold, based on the residual block, a target pixel area is determined in the first image. When the first proportion is not greater than the first super-resolution trigger proportion threshold, other methods are used to perform super-resolution processing on the entire first image, for example, using a traditional method to perform super-resolution processing on the first image. (Or, when the first proportion is greater than or equal to the first super-resolution trigger proportion threshold, based on the residual block, a target pixel area is determined in the first image. When the first proportion is less than the first super-resolution trigger proportion threshold, other methods are used to perform super-resolution processing on the entire first image.)
[0211] By determining whether the first proportion of the number of residual points with a residual value of 0 in the total number of residual points in the residual block is greater than the first super-resolution trigger proportion threshold, it can be detected whether the difference in the content of two consecutive frames of images is large. When the first proportion is not greater than the first super-resolution trigger proportion threshold, it indicates that the difference in the content of the two frames of images is large and the temporal correlation is not strong. If the computational cost of directly performing super-resolution processing on the entire first image (for example, directly inputting the first image into a super-resolution model) is less than or equal to the computational cost of using the aforementioned steps 203 to 206, then super-resolution processing can be directly performed on the entire first image, that is, all of the first image is super-resolved; when the first proportion is greater than the first super-resolution trigger proportion threshold, it indicates that the difference in the content of the two frames of images is small and the computational cost of directly performing super-resolution processing on the entire first image is greater than the computational cost of using the aforementioned steps 203 to 206, then the aforementioned steps 203 to 206 can be executed. This can determine whether to execute a partial super-resolution algorithm based on the content difference between the first image and the second image, thereby improving the flexibility of image processing.
[0212] Similarly, when performing step 203, the super-resolution display device may further count the second proportion of the number of residual points with non-zero residual values in the residual block in the total number of residual points in the residual block; when the second proportion is not greater than the second super-resolution trigger proportion threshold, based on the residual block, determine a target pixel region in the first image; when the second proportion is greater than the second super-resolution trigger proportion threshold, perform super-resolution processing on the entire first image in other ways, such as performing super-resolution processing on the first image in a traditional way (or, when the second proportion is greater than the second super-resolution trigger proportion threshold, based on the residual block, determine a target pixel region in the first image; when the second proportion is not greater than the second super-resolution trigger proportion threshold, perform super-resolution processing on the entire first image in other ways). By determining whether the second proportion is greater than the second super-resolution trigger proportion threshold, it is possible to detect whether the content difference between two consecutive frames of images is large. When the second proportion is greater than the second super-resolution trigger proportion threshold, it indicates that the content difference between the two frames of images is large and the temporal correlation is not strong. If the computational cost of directly performing super-resolution processing on the entire first image (for example, directly inputting the first image into the super-resolution model) is less than or equal to the computational cost of using the aforementioned steps 203 to 206, then the super-resolution processing can be directly performed on the first image, that is, full super-resolution of the first image is performed; when the second proportion is not greater than the second super-resolution trigger proportion threshold, it indicates that the content difference between the two frames of images is small, and the computational cost of directly performing super-resolution processing on the entire first image is greater than the computational cost of using the aforementioned steps 203 to 206, then the aforementioned steps 203 to 206 can be executed. This can determine whether to execute the partial super-resolution algorithm based on the content difference between the first image and the second image, thereby improving the flexibility of image processing.
[0213] The aforementioned first super-resolution trigger proportion threshold and the second super-resolution trigger proportion threshold may be the same or different. By way of example, both are 50%.
[0214] As mentioned above, in the embodiments of the present application, the image resolution of the images in the video may be 360p, 480p, 720p, 1080p, 2k, 4k, etc. The image resolutions exemplified in the foregoing embodiments are relatively small. For example, it is assumed that the first image and the second image each include 5·5 pixel points. This is only for the convenience of the reader's understanding and does not limit the actual resolution of the image to the resolution in the foregoing examples.
[0215] In the foregoing embodiments, inputting an image or a region of an image into the super-resolution model means inputting the pixel data of the pixel points in the image or the region of the image into the super-resolution model.
[0216] The sequence of steps of the image processing method provided by the embodiments of the present application can be appropriately adjusted, and the steps can also be increased or decreased accordingly according to the situation. Any person skilled in the art in the technical field disclosed in the present application can easily think of a changed method, which should be covered within the protection scope of the present application, so it will not be elaborated here. The foregoing steps 201 to 207 can all be controlled and executed by a processor as shown in Figure 1 shown.
[0217] Referring to Figure 13 , for some super-resolution methods provided by the embodiments of the present application, in fact, the first image is divided into a target pixel region H1 and other pixel regions H2 for separate processing (in actual implementation, there may be an overlapping region between the boundary of the target pixel region H1 determined in step 203 and the other pixel region H2 determined in step 205. Figure 13 Taking the case where their shapes match and there is no overlapping region as an example for illustration), by determining the target pixel region H1 in the first image and performing super-resolution processing on the target pixel region, super-resolution processing of the region where the pixel points different between the first image and the previous frame image is achieved, and the other pixel region in the previous frame image after super-resolution processing is used to update the other pixel region H2 in the first image, achieving the same effect of super-resolution processing of the other pixel region, making full use of the characteristics of video temporal redundancy. Therefore, by performing super-resolution processing on a partial region of the first image, the effect of performing super-resolution processing on the entire first image is achieved, reducing the computational amount of the actual super-resolution processing and reducing the computational cost.
[0218] For the test video processed by the embodiments of the present application, compared with directly performing full super-resolution processing on the video in the traditional technology, it can save about 45% of the super-resolution calculation amount. The significant reduction of the super-resolution calculation amount, on the one hand, is beneficial to accelerating the video processing speed, ensuring that the video can reach the basic frame rate requirement, thereby ensuring the real-time performance of the video and preventing situations such as playback delay and stuttering; on the other hand, the reduction of the calculation amount means less processing tasks and consumption of the calculation unit in the super-resolution display device, resulting in a decrease in the overall power consumption and saving the power consumption of the device.
[0219] Moreover, some of the partial super-resolution algorithms proposed in the embodiments of the present application are not methods that sacrifice the effect for efficiency by only super-resolving some image regions and processing other parts by non-super-resolution means, but avoid redundant super-resolution of unchanged regions and redundant temporal information between consecutive frames of the video. Essentially, it is a method that pursues the maximization of information utilization rate. For the first image using the partial super-resolution algorithm, by setting a mask pattern, it guides the super-resolution model to perform super-resolution at the pixel level. In the finally processed video, in fact, all pixel values of each frame image come from the super-resolution calculation results, which is the same as the display effect of the traditional full super-resolution algorithm, avoiding the sacrifice of the display effect.
[0220] Moreover, if the input sub-mask pattern and the target image block are used for super-resolution processing by the super-resolution model, the complexity of each super-resolution processing by the super-resolution model is relatively low, and the requirement for the structural complexity of the super-resolution model is relatively low. Thus, the super-resolution model can be simplified, the requirement for the processor performance can be reduced, and the super-resolution processing efficiency can be improved.
[0221] The following is the device embodiment of the present application, which can be used to execute the method embodiment of the present application. For the details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.
[0222] Please refer to Figure 14 , Figure 14 which is a block diagram of an image processing device 300. The device includes:
[0223] An acquisition module 301, configured to acquire the inter-frame residual between the first image and the adjacent previous frame image to obtain a residual block, where the residual block includes a plurality of residual points corresponding one-to-one to a plurality of pixel point positions of the first image, and each residual point has a residual value;
[0224] A first determination module 302, configured to determine a target pixel region in the first image based on the residual block;
[0225] A partial super-resolution module 303, configured to perform super-resolution processing on the target pixel region in the first image to obtain a super-resolution processed target pixel region;
[0226] An update module 304, configured to update the other pixel region in the first image with the other pixel region in the super-resolution processed previous frame image, where the other pixel region includes the pixel region in the first image except the target pixel region;
[0227] Wherein, the super-resolution processed first image includes the super-resolution processed target pixel region and the updated other pixel region.
[0228] Optionally, the target pixel region is the region where the pixel points corresponding to the positions of the first target residual point and the second target residual point in the first image are located. The first target residual point is the point in the residual block where the residual value is greater than the specified threshold, and the second target residual point is the residual point around the first target residual point in the residual block.
[0229] In the embodiment of the present application, the first determination module determines target pixel points in the first image, and the partial super-resolution module performs super-resolution processing on the target pixel points, so as to implement super-resolution processing on the area where the pixel points different between the first image and the previous frame image are located. And the update module updates the other pixel areas in the first image with the other pixel areas in the previous frame image after super-resolution processing, making full use of the characteristics of video temporal redundancy. Therefore, by performing super-resolution processing on a partial area of the first image, the effect of performing super-resolution processing on the entire first image is achieved, reducing the computational amount of super-resolution processing and lowering the computational cost.
[0230] As Figure 15 shown, in an optional manner, the first determination module 302 includes:
[0231] A generation sub-module 3021, configured to generate a mask pattern based on the residual block, where the mask pattern includes a plurality of first mask points, and the positions of the plurality of first mask points correspond one-to-one to the positions of a plurality of target residual points in the residual block;
[0232] A determination sub-module 3022, configured to input the mask pattern and the first image into a super-resolution model, and determine, through the super-resolution model, the area where the pixel points corresponding to each mask point in the plurality of first mask points are located in the first image as the target pixel area.
[0233] Correspondingly, the partial super-resolution module 303 is configured to: perform super-resolution processing on the target pixel area in the first image through the super-resolution model to obtain a super-resolution processed target pixel area.
[0234] In another optional manner, as Figure 15 shown, the first determination module 302 includes:
[0235] A generation sub-module 3021, configured to generate a mask pattern based on the residual block, where the mask pattern includes a plurality of first mask points, and the positions of the plurality of first mask points correspond one-to-one to the positions of a plurality of target residual points in the residual block;
[0236] A determination sub-module 3022, configured to determine the area where the pixel points corresponding to each mask point in the plurality of first mask points are located in the first image as the target pixel area.
[0237] Correspondingly, the partial super-resolution module 303 is configured to:
[0238] Input the target pixel points in the first image into a super-resolution model, and perform super-resolution processing on the target pixel area in the first image through the super-resolution model to obtain a super-resolution processed target pixel area.
[0239] Optionally, the mask pattern includes a plurality of mask points, the plurality of mask points corresponding one-to-one to the positions of the plurality of pixel points of the first image. Each of the mask points has a mask value. The plurality of mask points include the plurality of first mask points and a plurality of second mask points. The mask value of the first mask point is a first value, and the mask value of the second mask point is a second value. The first value and the second value are different;
[0240] In the above two optional manners, the determining sub-module 3022 can be used to:
[0241] Traverse the mask points in the mask pattern, and in the first image, determine the pixel points corresponding to the mask points with the first value as the mask value as the target pixel points.
[0242] Optionally, the generating sub-module 3021 is used to:
[0243] Perform morphological transformation processing on the residual block to obtain the mask pattern. The morphological transformation processing includes binarization processing and dilation processing on the first mask points in the binarized residual block. The super-resolution model includes at least one convolutional layer, and the kernel of the dilation processing has the same size as the receptive field of the last convolutional layer of the super-resolution model.
[0244] Optionally, the generating sub-module 3021 is used to:
[0245] Divide the residual block into a plurality of sub-residual blocks, and perform block processing on each divided sub-residual block. The block processing includes:
[0246] When the residual value of at least one residual point included in the sub-residual block is not 0, divide the sub-residual block into a plurality of sub-residual blocks, and perform the block processing on each divided sub-residual block until the residual values of the residual points included in the divided sub-residual block are all 0, or the total number of residual points in the divided sub-residual block is less than the point threshold, or the total number of divisions of the residual block reaches the number threshold;
[0247] Generate a sub-mask pattern corresponding to each target residual block, where at least one residual point included in the target residual block has a non-zero residual value;
[0248] Wherein, the mask pattern includes the generated sub-mask patterns.
[0249] Optionally, the partial super-resolution module 303 is used to:
[0250] Obtain a target image block corresponding to each sub-mask pattern in the first image;
[0251] Super-resolution processing is performed on sub-regions of the target pixel regions included in each of the target image patches respectively to obtain the target pixel regions after super-resolution processing, and the target pixel regions after super-resolution processing are composed of sub-regions of the target pixel regions included in each of the target image patches after super-resolution processing.
[0252] Optionally, both the residual block and the sub-residual block are divided into blocks in a quadtree partitioning manner.
[0253] Optionally, as Figure 16 shown, the apparatus 300 further includes:
[0254] An erosion module 305, configured to perform erosion processing on a plurality of first mask points in the mask pattern to obtain an updated mask pattern before updating the pixel values of the other pixel regions in the first image with the pixel values of the pixel points corresponding to the positions of the other pixel regions in the previous frame image, where the kernel of the erosion processing has the same size as the receptive field of the last convolutional layer of the super-resolution model;
[0255] A second determination module 306, configured to determine, in the first image, the pixel points corresponding to the positions of the plurality of first mask points after erosion processing as auxiliary pixel points;
[0256] A third determination module 307, configured to determine the region of the pixel points other than the auxiliary pixel points in the first image as the other pixel regions.
[0257] Optionally, the first determination module 302 is configured to:
[0258] Statistically calculate a first proportion of the number of residual points with a residual value of 0 in the residual block to the total number of residual points in the residual block;
[0259] When the first proportion is greater than a first super-resolution trigger proportion threshold, determine the target pixel region in the first image based on the residual block.
[0260] Optionally, the super-resolution model can be a CNN model, such as the SRCNN or ESPCN model; the super-resolution model can also be a GAN, such as the SRGAN or ESRGAN model.
[0261] In an embodiment of the present application, the first determination module realizes super-resolution processing on the area where the pixels different between the first image and the previous frame image are located by determining the target pixels in the first image and performing super-resolution processing on the target pixels by a partial super-resolution module, and the update module updates the other pixel areas in the first image with the other pixel areas in the previous frame image after super-resolution processing, making full use of the characteristics of video temporal redundancy. Therefore, by performing super-resolution processing on a partial area of the first image, the effect of performing super-resolution processing on the entire first image is achieved, reducing the computational amount of super-resolution processing and the computational cost.
[0262] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0263] Moreover, each module in the above device can be implemented by software or a combination of software and hardware. When at least one module is hardware, the hardware can be a logic integrated circuit module, which can specifically include transistors, logic gate arrays, or algorithmic logic circuits, etc. When at least one module is software, the software exists in the form of a computer program product and is stored in a computer-readable storage medium. The software can be executed by a processor. Therefore, alternatively, the image rendering device can be implemented by a processor executing a software program, and this embodiment is not limited thereto.
[0264] An embodiment of the present application provides an electronic device, including: a processor and a memory;
[0265] The memory is used to store a computer program;
[0266] When the processor executes the computer program stored in the memory, it implements any one of the image processing methods of the present application.
[0267] Figure 17 The structural schematic diagram of the electronic device 400 involved in the image processing method is shown. The electronic device 400 can be, but is not limited to, a laptop computer, a desktop computer, a mobile phone, a smart phone, a tablet computer, a multimedia player, an e-reader, a smart vehicle device, a smart home appliance (such as a smart TV), an artificial intelligence device, a wearable device, an Internet of Things device, or a virtual reality / augmented reality / mixed reality device, etc. For example, the electronic device 400 can include the structure of the super-resolution display device 100 shown above. Figure 1 The structure of the super-resolution display device 100 shown above.
[0268] The electronic device 400 may include a processor 410, an external memory interface 420, an internal memory 421, a universal serial bus (USB) interface 430, a charging management module 440, a power management module 441, a battery 442, an antenna 4, an antenna 2, a mobile communication module 450, a wireless communication module 460, an audio module 470, a speaker 470A, a receiver 470B, a microphone 470C, a headphone interface 470D, a sensor module 480, a button 490, a motor 491, an indicator 492, a camera 493, a display screen 494, and a subscriber identification module (SIM) card interface 495, etc. Among them, the sensor module 480 may include one or more of a pressure sensor 480A, a gyroscope sensor 480B, a barometric pressure sensor 480C, a magnetic sensor 480D, an acceleration sensor 480E, a distance sensor 480F, a proximity light sensor 480G, a fingerprint sensor 480H, a temperature sensor 480J, a touch sensor 480K, an ambient light sensor 480L, a bone conduction sensor 480M, etc.
[0269] It can be understood that the structure schematically shown in the embodiments of this application does not constitute a specific limitation on the electronic device 400. In other embodiments of this application, the electronic device 400 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0270] It can be understood that the interface connection relationship between the modules schematically shown in the embodiments of this application is only for illustrative purposes and does not constitute a structural limitation on the electronic device 400. In other embodiments of this application, the electronic device 400 may also adopt different interface connection methods (such as a bus connection method) or a combination of multiple interface connection methods as described in the above embodiments.
[0271] The processor 410 may include one or more processing units, such as a central processing unit CPU (e.g., an application processor (AP)), a graphics processing unit (GPU). Further, it may also include a modem processor, an image signal processor (ISP), an MCU, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0272] A memory may also be provided in the processor 410 for storing instructions and data. In some embodiments, the memory in the processor 410 is a cache memory. This memory can save the instructions or data that the processor 410 has just used or recycled. If the processor 410 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 410, and thus improves the efficiency of the system.
[0273] In some embodiments, the processor 410 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0274] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 410 may include multiple groups of I2C buses. The processor 410 may be respectively coupled to the touch sensor 480K, the charger, the flash, the camera 493, etc. through different I2C bus interfaces. For example, the processor 410 may be coupled to the touch sensor 480K through the I2C interface, enabling the processor 410 and the touch sensor 480K to communicate through the I2C bus interface, thereby implementing the touch function of the electronic device 400.
[0275] The I2S interface can be used for audio communication. In some embodiments, the processor 410 may include multiple groups of I2S buses. The processor 410 may be coupled to the audio module 470 through the I2S bus to implement communication between the processor 410 and the audio module 470. In some embodiments, the audio module 470 may transmit audio signals to the wireless communication module 460 through the I2S interface, thereby implementing the function of answering a call through a Bluetooth headset.
[0276] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some embodiments, the audio module 470 and the wireless communication module 460 may be coupled through the PCM bus interface. In some embodiments, the audio module 470 may also transmit audio signals to the wireless communication module 460 through the PCM interface, thereby implementing the function of answering a call through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0277] The UART interface is a general-purpose serial data bus for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 410 and the wireless communication module 460. For example, the processor 410 communicates with the Bluetooth module in the wireless communication module 460 through the UART interface to implement the Bluetooth function. In some embodiments, the audio module 470 may transmit audio signals to the wireless communication module 460 through the UART interface, thereby implementing the function of playing music through a Bluetooth headset.
[0278] The MIPI interface can be used to connect the processor 410 to peripheral devices such as the display screen 494 and the camera 493. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 410 and the camera 493 communicate through the CSI interface to implement the shooting function of the electronic device 400. The processor 410 and the display screen 494 communicate through the DSI interface to implement the display function of the electronic device 400.
[0279] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 410 to the camera 493, the display screen 494, the wireless communication module 460, the audio module 470, the sensor module 480, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0280] The USB interface 430 is an interface that complies with the USB standard specification, and can specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 430 can be used to connect a charger to charge the electronic device 400, and can also be used for data transmission between the electronic device 400 and peripheral devices. It can also be used to connect headphones to play audio through the headphones. This interface can also be used to connect other electronic devices, such as AR devices, etc.
[0281] The charging management module 440 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 440 can receive the charging input from a wired charger through the USB interface 430. In some embodiments of wireless charging, the charging management module 440 can receive the wireless charging input through the wireless charging coil of the electronic device 400. While the charging management module 440 charges the battery 442, it can also supply power to the electronic device through the power management module 441.
[0282] The power management module 441 is used to connect to the battery 442, the charging management module 440, and the processor 410. The power management module 441 receives inputs from the battery 442 and / or the charging management module 440 and supplies power to the processor 410, the internal memory 421, the display screen 494, the camera 493, the wireless communication module 460, etc. The power management module 441 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 441 can also be disposed in the processor 410. In some other embodiments, the power management module 441 and the charging management module 440 can also be disposed in the same device.
[0283] Optionally, the wireless communication function of the electronic device 400 can be implemented by the antenna 4, the antenna 2, the mobile communication module 450, the wireless communication module 460, the modulation and demodulation processor, and the baseband processor, etc.
[0284] The antenna 4 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 400 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 4 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0285] The mobile communication module 450 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device 400. The mobile communication module 450 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 450 can receive electromagnetic waves by the antenna 4, filter and amplify the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 450 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 4 for radiation. In some embodiments, at least some functional modules of the mobile communication module 450 can be disposed in the processor 410. In some embodiments, at least some functional modules of the mobile communication module 450 and at least some modules of the processor 410 can be disposed in the same device.
[0286] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 470A, the receiver 470B, etc.), or displays an image or video through the display screen 494. In some embodiments, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 410 and be provided in the same device as the mobile communication module 450 or other functional modules.
[0287] The wireless communication module 460 may provide solutions for wireless communications applied to the electronic device 400, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 460 may be one or more devices integrating at least one communication processing module. The wireless communication module 460 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signal, and transmits the processed signal to the processor 410. The wireless communication module 460 may also receive the signal to be transmitted from the processor 410, perform frequency modulation on it, amplify it, and convert it into electromagnetic waves through the antenna 2 and radiate it out.
[0288] In some embodiments, the antenna 4 of the electronic device 400 is coupled to the mobile communication module 450, and the antenna 2 is coupled to the wireless communication module 460, such that the electronic device 400 can communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).
[0289] The electronic device 400 implements the display function through the GPU, the display screen 494, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 494 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 410 may include one or more GPUs, which execute program instructions to generate or change the display information.
[0290] The display screen 494 is used to display images, videos, etc. The display screen 494 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 400 may include 4 or N display screens 494, where N is a positive integer greater than 4.
[0291] The electronic device 400 can implement the shooting function through the ISP, the camera 493, the video codec, the GPU, the display screen 494, and the application processor, etc.
[0292] The ISP is used to process the data fed back by the camera 493. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 493.
[0293] The camera 493 is used to capture static images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, the electronic device 400 may include 4 or N cameras 493, where N is a positive integer greater than 4.
[0294] The digital signal processor is used to process digital signals. In addition to being able to process digital image signals, it can also process other digital signals. For example, when the electronic device 400 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0295] The video codec is used to compress or decompress digital videos. The electronic device 400 can support one or more video codecs. In this way, the electronic device 400 can play or record videos in multiple encoding formats, such as: Moving Picture Experts Group (MPEG) 4, MPEG2, MPEG3, MPEG4, etc.
[0296] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 400 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.
[0297] The external memory interface 420 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 400. The external memory card communicates with the processor 410 through the external memory interface 420 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0298] The internal memory 421 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 421 can include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.). The data storage area can store the data created during the use of the electronic device 400 (such as audio data, phone book, etc.). In addition, the internal memory 421 can include high-speed random access memory, such as double data rate synchronous dynamic random access memory (DDR), and can also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 410 executes various functional applications and data processing of the electronic device 400 by running the instructions stored in the internal memory 421 and / or the instructions stored in the memory provided in the processor.
[0299] The electronic device 400 can implement audio functions through the audio module 470, the speaker 470A, the receiver 470B, the microphone 470C, the headphone jack 470D, and the application processor, etc. For example, music playback, recording, etc.
[0300] The audio module 470 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 470 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 470 can be disposed in the processor 410, or some functional modules of the audio module 470 can be disposed in the processor 410.
[0301] The speaker 470A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 400 can listen to music or hands-free calls through the speaker 470A.
[0302] The receiver 470B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 400 answers a call or a voice message, the voice can be listened to by bringing the receiver 470B close to the human ear.
[0303] The microphone 470C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone 470C to input the sound signal into the microphone 470C. The electronic device 400 can be provided with at least one microphone 470C. In some other embodiments, the electronic device 400 can be provided with two microphones 470C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 400 can also be provided with three, four or more microphones 470C to implement functions such as collecting sound signals, noise reduction, identifying the sound source, and implementing a directional recording function.
[0304] The headphone jack 470D is used to connect a wired headphone. The headphone jack 470D can be a USB interface 430, or a 3.5-mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0305] The pressure sensor 480A is used to sense pressure signals and can convert pressure signals into electrical signals. In some embodiments, the pressure sensor 480A may be disposed on the display screen 494. There are many types of pressure sensors 480A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor may include at least two parallel plates having conductive materials. When a force acts on the pressure sensor 480A, the capacitance between the electrodes changes. The electronic device 400 determines the intensity of the pressure based on the change in capacitance. When a touch operation acts on the display screen 494, the electronic device 400 detects the intensity of the touch operation according to the pressure sensor 480A. The electronic device 400 can also calculate the position of the touch based on the detection signal of the pressure sensor 480A. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities may correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, the instruction to view short messages is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.
[0306] The gyroscope sensor 480B can be used to determine the motion posture of the electronic device 400. In some embodiments, the angular velocity of the electronic device 400 around three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 480B. The gyroscope sensor 480B can be used for anti-shake during shooting. Exemplarily, when the shutter is pressed, the gyroscope sensor 480B detects the angle of jitter of the electronic device 400, calculates the distance that the lens module needs to compensate based on the angle, and enables the lens to counteract the jitter of the electronic device 400 through reverse movement to achieve anti-shake. The gyroscope sensor 480B can also be used for navigation and somatosensory game scenarios.
[0307] The barometric pressure sensor 480C is used to measure barometric pressure. In some embodiments, the electronic device 400 calculates the altitude based on the barometric pressure value measured by the barometric pressure sensor 480C to assist in positioning and navigation.
[0308] The magnetic sensor 480D includes a Hall sensor. The electronic device 400 can use the magnetic sensor 480D to detect the opening and closing of the flip leather case. In some embodiments, when the electronic device 400 is a flip phone, the electronic device 400 can detect the opening and closing of the flip according to the magnetic sensor 480D. Furthermore, according to the detected opening and closing state of the leather case or the opening and closing state of the flip, features such as automatic flip unlocking are set.
[0309] The acceleration sensor 480E can detect the magnitude of the acceleration of the electronic device 400 in various directions (generally three axes). When the electronic device 400 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.
[0310] A distance sensor 480F for measuring distance. The electronic device 400 can measure distance through infrared or laser. In some embodiments, when shooting a scene, the electronic device 400 can use the distance sensor 480F to measure distance to achieve fast focusing.
[0311] The proximity light sensor 480G can include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The light-emitting diode can be an infrared light-emitting diode. The electronic device 400 emits infrared light outward through the light-emitting diode. The electronic device 400 uses the photodiode to detect the infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 400. When insufficient reflected light is detected, the electronic device 400 can determine that there is no object near the electronic device 400. The electronic device 400 can use the proximity light sensor 480G to detect when the user holds the electronic device 400 close to the ear for a call, so as to automatically turn off the screen to save power. The proximity light sensor 480G can also be used for automatic unlocking and locking of the holster mode and pocket mode.
[0312] The ambient light sensor 480L is used to sense the ambient light brightness. The electronic device 400 can adaptively adjust the brightness of the display screen 494 according to the sensed ambient light brightness. The ambient light sensor 480L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 480L can also cooperate with the proximity light sensor 480G to detect whether the electronic device 400 is in the pocket to prevent accidental touch.
[0313] The fingerprint sensor 480H is used to collect fingerprints. The electronic device 400 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application locks, fingerprint photography, fingerprint answering of incoming calls, etc.
[0314] The temperature sensor 480J is used to detect temperature. In some embodiments, the electronic device 400 executes a temperature processing strategy using the temperature detected by the temperature sensor 480J. For example, when the temperature reported by the temperature sensor 480J exceeds a threshold, the electronic device 400 reduces the performance of the processor located near the temperature sensor 480J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the electronic device 400 heats the battery 442 to avoid abnormal shutdown of the electronic device 400 caused by low temperature. In some other embodiments, when the temperature is lower than yet another threshold, the electronic device 400 boosts the output voltage of the battery 442 to avoid abnormal shutdown caused by low temperature.
[0315] The touch sensor 480K, also known as the "touch control device". The touch sensor 480K can be disposed on the display screen 494, and the touch sensor 480K and the display screen 494 form a touch screen, also known as the "touch control screen". The touch sensor 480K is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 494. In some other embodiments, the touch sensor 480K can also be disposed on the surface of the electronic device 400, at a different position from the display screen 494.
[0316] The bone conduction sensor 480M can acquire vibration signals. In some embodiments, the bone conduction sensor 480M can acquire vibration signals of the vibrating bone mass of the human vocal tract. The bone conduction sensor 480M can also contact the human pulse to receive blood pressure pulsation signals. In some embodiments, the bone conduction sensor 480M can also be disposed in the earphone to form a bone conduction earphone. The audio module 470 can parse out voice signals based on the vibration signals of the vibrating bone mass acquired by the bone conduction sensor 480M to implement the voice function. The application processor can parse out heart rate information based on the blood pressure pulsation signals acquired by the bone conduction sensor 480M to implement the heart rate detection function.
[0317] In some other embodiments of the present application, the electronic device 400 can also adopt different interface connection methods in the above embodiments. For example, some or all of the above multiple sensors are connected to the MCU, and then connected to the AP through the MCU.
[0318] The keys 490 include a power-on key, volume keys, etc. The keys 490 can be mechanical keys or touch keys. The electronic device 400 can receive key inputs and generate key signal inputs related to the user settings and function controls of the electronic device 400.
[0319] The motor 491 can generate vibration prompts. The motor 491 can be used for incoming call vibration prompts and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playing, etc.) can correspond to different vibration feedback effects. For touch operations acting on different regions of the display screen 494, the motor 491 can also correspond to different vibration feedback effects. Different application scenarios (such as time reminder, receiving information, alarm clock, game, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0320] The indicator 492 can be an indicator light and can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc.
[0321] The SIM card interface 495 is used to connect to a SIM card. The SIM card can be inserted into or removed from the SIM card interface 495 to achieve contact with and separation from the electronic device 400. The electronic device 400 can support 4 or N SIM card interfaces, where N is a positive integer greater than 4. The SIM card interface 495 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 495 simultaneously. The types of the multiple cards can be the same or different. The SIM card interface 495 can also be compatible with different types of SIM cards. The SIM card interface 495 can also be compatible with external memory cards. The electronic device 400 interacts with the network through the SIM card to implement functions such as calls and data communication. In some embodiments, the electronic device 400 uses an eSIM, that is, an embedded SIM card. The eSIM card can be embedded in the electronic device 400 and cannot be separated from the electronic device 400.
[0322] The software system of the electronic device 400 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of this application, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the electronic device 400.
[0323] The embodiments of this application also provide an image processing device, including a processor and a memory; when the processor executes the computer program stored in the memory, the image processing device executes the image processing method provided by the embodiments of this application. Optionally, the image processing device can be deployed in a smart TV.
[0324] The embodiments of this application also provide a storage medium. The storage medium can be a non-volatile computer-readable storage medium. A computer program is stored in the storage medium, and the computer program instructs the terminal to execute any of the image processing methods provided by the embodiments of this application. The storage medium can include various media that can store program codes, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0325] The embodiments of the present application also provide a computer program product containing instructions. When the computer program product runs on a computer, it causes the computer to execute the image processing method provided by the embodiments of the present application. The computer program product may include one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer-readable storage medium may be any available medium that the computer can access or a data storage device such as a server or a data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.
[0326] The embodiments of the present application also provide a chip, such as a CPU chip. The chip includes one or more physical cores and a storage medium. After reading the computer instructions in the storage medium, the one or more physical cores implement the foregoing image processing method. In other embodiments, the chip may implement the foregoing image processing method in a pure hardware or a hardware-software combination manner, that is, the chip includes a logic circuit. When the chip runs, the logic circuit is used to implement any one of the image processing methods in the foregoing first aspect, and the logic circuit may be a programmable logic circuit. Similarly, a GPU may also be implemented like a CPU.
[0327] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware or by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, etc.
[0328] In the embodiments of the present application, "A refers to B" means that A is the same as B or A is simply deformed on the basis of B.
[0329] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. An image processing method, characterized in that, the method includes: obtaining an inter-frame residual between a first image and an adjacent previous frame image to obtain a residual block, where the inter-frame residual is the absolute value difference between the pixel values of the first image and the pixel values of the previous frame image, and the residual block includes a plurality of residual points corresponding one-to-one to the positions of a plurality of pixel points of the first image, and each residual point has a residual value; based on the residual block, determining a target pixel region in the first image, where the target pixel region is the region of pixel points in the first image corresponding to the positions of a first target residual point and a second target residual point, the first target residual point is a point in the residual block whose residual value is greater than a specified threshold, and the second target residual point is a residual point around the first target residual point in the residual block; performing super-resolution processing on the target pixel region in the first image to obtain a super-resolution processed target pixel region; updating other pixel regions in the first image with other pixel regions in the super-resolution processed previous frame image, where the other pixel regions include pixel regions in the first image other than the target pixel region; wherein, the super-resolution processed first image includes the super-resolution processed target pixel region and the updated other pixel regions.
2. The method according to claim 1, characterized in that, the determining a target pixel region in the first image based on the residual block includes: generating a mask pattern based on the residual block, where the mask pattern includes a plurality of first mask points, and the plurality of first mask points correspond one-to-one to the positions of a plurality of target residual points in the residual block; inputting the mask pattern and the first image into a super-resolution model, and determining, by the super-resolution model, the region of pixel points in the first image corresponding to the position of each mask point in the plurality of first mask points as the target pixel region.
3. The method according to claim 2, characterized in that, the performing super-resolution processing on the target pixel region in the first image to obtain a super-resolution processed target pixel region includes: performing super-resolution processing on the target pixel region in the first image by the super-resolution model to obtain a super-resolution processed target pixel region.
4. The method according to claim 1, characterized in that, the determining a target pixel region in the first image based on the residual block includes: generating a mask pattern based on the residual block, where the mask pattern includes a plurality of first mask points, and the plurality of first mask points correspond one-to-one to the positions of a plurality of target residual points in the residual block; determining the region of pixel points in the first image corresponding to the position of each mask point in the plurality of first mask points as the target pixel region.
5. The method according to claim 4, characterized in that, the performing super-resolution processing on the target pixel region in the first image to obtain a super-resolution processed target pixel region includes: Input the target pixel region in the first image into the super-resolution model, and perform super-resolution processing on the target pixel region in the first image through the super-resolution model to obtain the super-resolved target pixel region.
6. The method according to claim 4 or 5, wherein, generating a mask pattern based on the residual block includes: generating an initial mask pattern including a plurality of mask points based on the residual block, the plurality of mask points corresponding one-to-one to the positions of a plurality of pixel points in the first image, and the plurality of mask points including the plurality of first mask points and a plurality of second mask points; assigning a first value to the mask value of the first mask points in the initial mask pattern, and assigning a second value to the mask value of the second mask points in the mask pattern to obtain the mask pattern, where the first value and the second value are different; determining the region where the pixel points corresponding to each mask point among the plurality of first mask points in the first image are located as the target pixel region includes: traversing the mask points in the mask pattern, and in the first image, determining the pixel points corresponding to the mask points with a mask value of the first value as the target pixel region.
7. The method according to any one of claims 2, 3, and 5, wherein, generating a mask pattern based on the residual block includes: performing morphological transformation processing on the residual block to obtain the mask pattern, where the morphological transformation processing includes binarization processing and dilation processing on the first mask points in the binarized residual block, the super-resolution model includes at least one convolutional layer, and the kernel of the dilation processing has the same size as the receptive field of the last convolutional layer of the super-resolution model.
8. The method according to any one of claims 2 to 5, wherein, generating a mask pattern based on the residual block includes: dividing the residual block into a plurality of sub-residual blocks, and performing block processing on each divided sub-residual block, where the block processing includes: when the residual values of at least one residual point included in the sub-residual block are not 0, dividing the sub-residual block into a plurality of sub-residual blocks, and performing the block processing on each divided sub-residual block until the residual values of the residual points included in the divided sub-residual block are all 0, or the total number of residual points in the divided sub-residual block is less than the point threshold, or the total number of divisions of the residual block reaches the number threshold; generating a sub-mask pattern corresponding to each target residual block, where the target residual block includes at least one residual point with a non-zero residual value; wherein, the mask pattern includes the generated sub-mask patterns.
9. The method according to claim 8, wherein, performing super-resolution processing on the target pixel region in the first image to obtain the super-resolved target pixel region includes: acquiring a target image block corresponding to each sub-mask pattern in the first image; Super-resolution processing is respectively performed on sub-regions of the target pixel regions included in each of the target image patches to obtain the target pixel regions after the super-resolution processing, and the target pixel regions after the super-resolution processing are composed of the sub-regions of the target pixel regions included in each of the target image patches after the super-resolution processing.
10. The method according to claim 9, wherein, both the residual block and the sub-residual block are divided into blocks in a quadtree partitioning manner.
11. The method according to claim 7, wherein, before updating the other pixel regions in the first image with the other pixel regions in the previous frame image after the super-resolution processing, the method further includes: performing erosion processing on a plurality of first mask points in the mask pattern to obtain an updated mask pattern, and the kernel of the erosion processing has the same size as the receptive field of the last convolutional layer of the super-resolution model; determining pixel points corresponding to the positions of the plurality of first mask points after the erosion processing in the first image as auxiliary pixel points; determining the region where the pixel points other than the auxiliary pixel points are located in the first image as the other pixel regions.
12. The method according to any one of claims 1 to 5 and 9 to 11, wherein, determining a target pixel region in the first image based on the residual block includes: counting a first proportion of the number of residual points with a residual value of 0 in the residual block in the total number of residual points in the residual block; when the first proportion is greater than a first super-resolution trigger proportion threshold, determining the target pixel region in the first image based on the residual block.
13. An image processing apparatus, wherein, the apparatus includes: an acquisition module, configured to acquire an inter-frame residual between a first image and an adjacent previous frame image to obtain a residual block, where the inter-frame residual is the absolute value difference between the pixel values of the first image and the previous frame image, and the residual block includes a plurality of residual points corresponding one-to-one to the positions of a plurality of pixel points in the first image, and each residual point has a residual value; a first determination module, configured to determine a target pixel region in the first image based on the residual block, where the target pixel region is the region where the pixel points corresponding to the positions of a first target residual point and a second target residual point in the first image are located, the first target residual point is a point in the residual block with a residual value greater than a specified threshold, and the second target residual point is a residual point around the first target residual point in the residual block; a partial super-resolution module, configured to perform super-resolution processing on the target pixel region in the first image to obtain a target pixel region after the super-resolution processing; an update module, configured to update the other pixel regions in the first image with the other pixel regions in the previous frame image after the super-resolution processing, where the other pixel regions include the pixel regions other than the target pixel region in the first image; wherein, the first image after the super-resolution processing includes the target pixel region after the super-resolution processing and the updated other pixel regions.
14. The apparatus according to claim 13, wherein, The first determination module includes: A generation sub-module, configured to generate a mask pattern based on the residual block, where the mask pattern includes a plurality of first mask points, and the positions of the plurality of first mask points correspond one-to-one to the positions of a plurality of target residual points in the residual block; A determination sub-module, configured to input the mask pattern and the first image into a super-resolution model, and determine, through the super-resolution model, a region where a pixel point corresponding to each mask point in the plurality of first mask points in the first image is located as the target pixel region.
15. The apparatus according to claim 14, wherein, The partial super-resolution module is configured to: Perform super-resolution processing on the target pixel region in the first image through the super-resolution model to obtain a super-resolution processed target pixel region.
16. The apparatus according to claim 13, wherein, The first determination module includes: A generation sub-module, configured to generate a mask pattern based on the residual block, where the mask pattern includes a plurality of first mask points, and the positions of the plurality of first mask points correspond one-to-one to the positions of a plurality of target residual points in the residual block; A determination sub-module, configured to determine, in the first image, a region where a pixel point corresponding to each mask point in the plurality of first mask points is located as the target pixel region.
17. The apparatus according to claim 16, wherein, The partial super-resolution module is configured to: Input a target pixel point in the first image into the super-resolution model, and perform super-resolution processing on the target pixel region in the first image through the super-resolution model to obtain a super-resolution processed target pixel region.
18. The apparatus according to claim 16 or 17, wherein, The generation sub-module is configured to: Generate an initial mask pattern including a plurality of mask points based on the residual block, where the plurality of mask points correspond one-to-one to the positions of a plurality of pixel points in the first image, and the plurality of mask points include the plurality of first mask points and a plurality of second mask points; Assign a first value to the mask value of the first mask point in the initial mask pattern, and assign a second value to the mask value of the second mask point in the mask pattern, to obtain the mask pattern, where the first value and the second value are different; The determination sub-module is configured to: Traverse the mask points in the mask pattern, and in the first image, determine a pixel point corresponding to a mask point with a mask value of the first value as the target pixel point.
19. The apparatus according to any one of claims 14, 15, and 17, wherein, The generation sub-module is configured to: Perform morphological transformation processing on the residual block to obtain the mask pattern, where the morphological transformation processing includes binarization processing and dilation processing on the first mask points in the binarized residual block, the super-resolution model includes at least one convolutional layer, and the kernel of the dilation processing has the same size as the receptive field of the last convolutional layer of the super-resolution model.
20. The apparatus according to any one of claims 14 to 17, wherein, The generation sub-module is configured to: Divide the residual block into multiple sub - residual blocks, and perform block - processing on each divided sub - residual block. The block - processing includes: When the residual values of at least one residual point included in the sub - residual block are not 0, divide the sub - residual block into multiple sub - residual blocks, and perform the block - processing on each divided sub - residual block until the residual values of the residual points included in the divided sub - residual block are all 0, or the total number of residual points in the divided sub - residual block is less than the point threshold, or the total number of divisions of the residual block reaches the number threshold; Generate a sub - mask pattern corresponding to each target residual block, where at least one residual point included in the target residual block has a non - zero residual value; Wherein, the mask pattern includes the generated sub - mask patterns.
21. The apparatus according to claim 20, characterized in that, The partial super - resolution module is configured to: Obtain a target image block corresponding to each sub - mask pattern in the first image; Perform super - resolution processing on sub - regions of the target pixel region included in each target image block respectively to obtain the super - resolution processed target pixel region, and the super - resolution processed target pixel region is composed of sub - regions of the target pixel region included in each super - resolution processed target image block.
22. The apparatus according to claim 21, characterized in that, Both the residual block and the sub - residual block are divided into blocks in a quadtree partitioning manner.
23. The apparatus according to claim 19, characterized in that, The apparatus further includes: An erosion module, configured to perform erosion processing on multiple first mask points in the mask pattern to obtain an updated mask pattern before updating other pixel regions in the first image with other pixel regions in the previous frame image after super - resolution processing. The kernel of the erosion processing has the same size as the receptive field of the last convolutional layer of the super - resolution model; A second determination module, configured to determine pixel points corresponding to the positions of the multiple first mask points after erosion processing in the first image as auxiliary pixel points; A third determination module, configured to determine the region of pixel points other than the auxiliary pixel points in the first image as the other pixel region.
24. The apparatus according to any one of claims 13 to 17, 21 to 23, characterized in that, The first determination module is configured to: Statistically calculate a first proportion of the number of residual points with a residual value of 0 in the residual block in the total number of residual points in the residual block; When the first proportion is greater than the first super - resolution trigger proportion threshold, determine the target pixel region in the first image based on the residual block.
25. An electronic device, characterized in that, comprises: A processor and a memory; The memory is used to store a computer program; The processor is configured to implement the image - processing method according to any one of claims 1 to 12 when executing the computer program stored in the memory.
Citation Information
Patent Citations
Escalator handrail boundary area cross-border detection method and system
CN110009650A