Image processing method and device, electronic equipment and storage medium
By performing resolution compression and restoration on image data, and combining it with a trained depth estimation network, the problems of high computational cost and low accuracy in depth estimation algorithms are solved, achieving efficient depth estimation and image blurring effects.
Patent Information
- Application Number
- CN202111416602.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-23
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2041-11-23
AI Technical Summary
Existing depth estimation algorithms require a large amount of computation to achieve high accuracy, making it difficult to maintain accuracy while reducing computational load.
By converting the image data from a first data arrangement to a second data arrangement, compressing the resolution, and using a trained depth estimation network to perform depth estimation, the resolution is then restored to obtain an accurate depth-estimated image.
It improves the accuracy of depth estimation while reducing computational load and supports real-time preview of image blurring effects.
Smart Images

Figure CN116168068B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image technology, in particular to an image processing method and device, electronic equipment and storage medium. BACKGROUND
[0002] When implementing the image background blurring function on a small electronic device such as a mobile phone, different blurring degrees need to be designed according to the distance between the photographed object and the electronic device: the farther the distance between the photographed object and the electronic device, the higher the blurring degree of the background, and the more blurred the background; the closer the distance between the photographed object and the electronic device, the lower the blurring degree of the background, and the clearer the background. It can be seen that one of the prerequisites for achieving a better background blurring function includes accurately identifying the distance between the photographed object and the electronic device, i.e., depth estimation of the captured image.
[0003] However, in practice, it is found that if the existing depth estimation algorithm wants to obtain a high-precision depth estimation result, the amount of calculation required to run the algorithm will increase; if the amount of calculation of the depth estimation algorithm is reduced, the accuracy of the depth estimation will decrease. SUMMARY
[0004] The image processing method, device, electronic equipment and storage medium disclosed by the embodiments of the present application can reduce the amount of calculation while improving the accuracy of depth estimation.
[0005] The first aspect of the embodiments of the present application discloses an image processing method, which comprises: converting image data of a first to-be-processed image from a first data arrangement mode to a second data arrangement mode based on a first operator to obtain a second to-be-processed image with compressed resolution; the components in the depth dimension of the first data arrangement mode and the second data arrangement mode are the number of channels, and the number of channels corresponding to the first data arrangement mode is less than the number of channels corresponding to the second data arrangement mode; performing depth estimation on the second to-be-processed image through a trained depth estimation network to obtain a first depth estimation image corresponding to the second to-be-processed image, the resolution of the first depth estimation image being consistent with the resolution of the second to-be-processed image; converting image data of the first depth estimation image from the second data arrangement mode to the first data arrangement mode based on a second operator to obtain a second depth estimation image with a resolution consistent with the first to-be-processed image.
[0006] The embodiment of the present application discloses an image processing device, characterized in that comprising: a first conversion module, configured to convert image data of a first to-be-processed image from a first data arrangement mode to a second data arrangement mode based on a first operator, to obtain a second to-be-processed image with compressed resolution; the components in the depth dimension of the first data arrangement mode and the second data arrangement mode are the number of channels, and the number of channels corresponding to the first data arrangement mode is less than the number of channels corresponding to the second data arrangement mode; a depth estimation module, configured to perform depth estimation on the second to-be-processed image through a trained depth estimation network, to obtain a first depth estimation image corresponding to the second to-be-processed image, the resolution of the first depth estimation image is consistent with the resolution of the second to-be-processed image; a second conversion module, configured to convert image data of the first depth estimation image from the second data arrangement mode to the first data arrangement mode based on a second operator, to obtain a second depth estimation image with the same resolution as the first to-be-processed image.
[0007] The third aspect of the embodiment of the present application discloses an electronic device, comprising a memory and a processor, the memory stores a computer program, and the computer program is executed by the processor to make the processor implement any one of the image processing methods disclosed by the embodiment of the present application.
[0008] The fourth aspect of the present application discloses a computer readable storage medium storing a computer program, wherein the computer program is executed by a processor to implement any one of the image processing methods disclosed by the embodiment of the present application.
[0009] Compared with the related art, the embodiment of the present application has the following beneficial effects:
[0010] Based on the conversion of the arrangement mode of the image data of the first to-be-processed image by the first operator, the resolution of the image is compressed while the original image information is retained, and the second to-be-processed image is obtained. Based on the trained depth estimation network, the depth estimation of the second to-be-processed image is performed, and the first depth estimation image is obtained. The resolution of the first depth estimation image output by the depth estimation network is consistent with the resolution of the second to-be-processed image input into the depth estimation network, so that the arrangement mode conversion of the image data of the first depth estimation image can be performed based on the second operator, to obtain the second depth estimation image. In the embodiment of the present application, the conversion of the arrangement mode of the image data does not lose the original image information, and the image with compressed resolution is input into the depth estimation network for depth estimation, which can reduce the calculation amount while improving the depth estimation accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description only represent some of the embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.
[0012] Figure 1 FIG. 1 is a structural schematic diagram of an image processing circuit disclosed by an embodiment;
[0013] Figure 2 FIG. 2 is a method flow schematic diagram of an image processing method disclosed by an embodiment;
[0014] Figure 3 FIG. 3 is an example diagram of a data arrangement mode of a first to-be-processed image based on a sapce-to-depth function conversion disclosed by an embodiment;
[0015] Figure 4 FIG. 4 is an example diagram of a data arrangement mode of a first depth estimation image based on a depth-to-space function conversion disclosed by an embodiment;
[0016] Figure 5 FIG. 5 is a structural schematic diagram of a depth estimation network disclosed by an embodiment;
[0017] Figure 6 FIG. 6 is a method flow schematic diagram of an image processing method disclosed by another embodiment;
[0018] Figure 7 FIG. 7 is a method flow schematic diagram of a depth estimation network training method disclosed by an embodiment;
[0019] Figure 8 FIG. 8 is a method flow schematic diagram of a depth estimation network training method disclosed by another embodiment;
[0020] Figure 9 FIG. 9 is a structural schematic diagram of an image processing device disclosed by an embodiment;
[0021] Figure 10 FIG. 10 is a structural schematic diagram of an electronic device disclosed by an embodiment of the present application. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only represent some of the embodiments of the present application, and not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0023] It should be noted that the terms "comprising" and "having" and any variations thereof in the embodiments of the present application and the accompanying drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally further include steps or units not listed, or can optionally further include other steps or units inherent to the process, method, product or device.
[0024] Embodiments of the present application disclose image processing methods and devices, electronic devices and storage media, which can reduce the amount of calculation while improving the accuracy of depth estimation. The following are described in detail.
[0025] Please refer to Figure 1 , Figure 1 is a structural schematic diagram of an image processing circuit disclosed by an embodiment. The image processing circuit can be applied to electronic devices such as smart phones, smart tablets, smart watches, but is not limited thereto. As shown in Figure 1 , the image processing circuit can include an imaging device (camera) 110, a posture sensor 120, an image memory 130, an image signal processing (ISP) processor 140, a logic controller 150, and a display 160.
[0026] The image processing circuit includes the ISP processor 140 and the control logic 150. The image data captured by the imaging device 110 is first processed by the ISP processor 140, which analyzes the image data to capture image statistics that can be used to determine one or more control parameters of the imaging device 110. The imaging device 110 can include one or more lenses 112 and an image sensor 114. The image sensor 114 can include a color filter array (such as a Bayer filter), and the image sensor 114 can obtain light intensity and wavelength information captured by each imaging pixel and provide a set of raw image data (RAW image data) that can be processed by the ISP processor 140. The posture sensor 120 (such as a three-axis gyroscope, a Hall sensor, an accelerometer, etc.) can provide image processing parameters (such as anti-shake parameters) collected by the posture sensor 120 to the ISP processor 140 based on the posture sensor 120 interface type. The posture sensor 120 interface can use an SMIA (Standard Mobile Imaging Architecture) interface, other serial or parallel camera interfaces, or a combination of the above interfaces.
[0027] In addition, the image sensor 114 can also send raw image data to the pose sensor 120, which can provide the raw image data to the ISP processor 140 based on the pose sensor 120 interface type or store the raw image data into the image memory 130.
[0028] The ISP processor 140 processes the raw image data pixel by pixel in multiple formats. For example, each image pixel can have a bit depth of 8, 10, 12, or 14 bits, and the ISP processor 140 can perform one or more image processing operations on the raw image data, collect statistics about the image data, etc. The image processing operations can be performed at the same or different bit depth precisions.
[0029] The ISP processor 140 can also receive image data from the image memory 130. For example, the pose sensor 120 interface sends raw image data to the image memory 130, which provides the raw image data to the ISP processor 140 for processing. The image memory 130 can be part of a memory device, a storage device, or a separate dedicated memory within the electronic device, and can include a DMA (Direct Memory Access) feature.
[0030] When receiving raw image data from the image sensor 114 interface or from the pose sensor 120 interface or from the image memory 130, the ISP processor 140 can perform one or more image processing operations, such as temporal filtering. The processed image data can be sent to the image memory 130 for additional processing before being displayed. The ISP processor 140 receives the processed data from the image memory 130 and performs image data processing in the raw domain and in one or more color spaces, such as YUV, RGB, YCbCr, etc. The processed image data from the ISP processor 140 can be output to the display 160 for viewing by a user and / or further processing by a graphics engine or GPU (Graphics Processing Unit). In addition, the output from the ISP processor 140 can also be sent to the image memory 130, and the display 160 can read the image data from the image memory 130. In one embodiment, the image memory 130 can be configured to implement one or more frame buffers.
[0031] The statistics determined by the ISP processor 140 can be sent to the control logic 150. For example, the statistics can include image sensor 114 statistics such as vibration frequency of the gyroscope, auto exposure, auto white balance, auto focus, flicker detection, black level compensation, lens 112 shading correction, etc. The control logic 150 can include a processor and / or microcontroller that executes one or more routines (e.g., firmware) that can determine control parameters for the imaging device 110 and control parameters for the ISP processor 140 based on the received statistics. For example, the control parameters for the imaging device 110 can include attitude sensor 120 control parameters (e.g., gain, integration time for exposure control, anti-shake parameters, etc.), camera flash control parameters, camera anti-shake displacement parameters, lens 112 control parameters (e.g., focus or zoom focal length), or a combination of these parameters. The ISP control parameters can include gain levels and color correction matrices for auto white balance and color adjustment (e.g., during YUV processing), and lens 112 shading correction parameters.
[0032] In one embodiment, the ISP processor 140 can obtain a first to-be-processed image from the imaging device 110 or the image storage 130, and convert image data of the first to-be-processed image from a first data arrangement mode to a second data arrangement mode to obtain a second to-be-processed image with a resolution compressed. The first to-be-processed image can be captured by the imaging device 110 or captured by another device and temporarily stored in the image storage 130.
[0033] The trained depth estimation network in the ISP processor 140 can also be invoked to perform depth estimation on the second to-be-processed image to obtain a first depth estimation image corresponding to the second to-be-processed image; wherein the resolution of the first depth estimation image is consistent with the resolution of the second to-be-processed image, and the first depth estimation image includes a depth estimation value corresponding to each pixel point in the second to-be-processed image, which can be predicted by the trained depth estimation network.
[0034] After obtaining the first depth estimation image, the ISP processor 140 can also convert the image data of the first depth estimation image from the second data arrangement mode to the first data arrangement mode based on a second operator to restore the resolution of the first depth estimation image to be consistent with the resolution of the first to-be-processed image, and obtain a second depth estimation image.
[0035] The ISP processor 140 can perform a blurring process on the first to-be-processed image based on the depth estimation value corresponding to each pixel point in the second depth estimation image to achieve a blurring effect of a clear foreground and a blurred background.
[0036] Please refer to Figure 2 , Figure 2Fig. 1 is a schematic diagram of a method flow of an image processing method according to an embodiment. The image processing method can be applied to any of the electronic devices disclosed in the foregoing embodiments. As shown in Fig. 1, the image processing method can include the following steps. Figure 2
[0037] 210. Convert the image data of the first to-be-processed image from the first data arrangement manner to a second data arrangement manner based on a first operator, to obtain a second to-be-processed image with compressed resolution.
[0038] The image data can be represented in the form of an array. In a neural network, the array used to represent the image data can also be referred to as a tensor. The shape of the tensor can be determined based on three dimensions of height, width, and depth. Among them, the height dimension and the width dimension can be referred to as the space dimension. The shape of the tensor can determine the arrangement manner of the image data.
[0039] Therefore, the first data arrangement manner and the second data arrangement manner described above can be two tensors with different shapes. The number of components of the first data arrangement manner in the height dimension is greater than that of the second data arrangement manner in the height dimension, and the number of components of the first data arrangement manner in the width dimension can be greater than that of the second data arrangement manner in the width dimension. The number of components of the first data arrangement manner and the second data arrangement manner in the depth dimension can be the number of channels, and the number of channels corresponding to the first data arrangement manner can be less than the number of channels corresponding to the second data arrangement manner. In a popular way, compared with the first data arrangement manner, the second data arrangement manner is smaller in the height and width dimensions and longer in the depth dimension.
[0040] In the field of imaging, the image resolution can be represented as "number of horizontal pixels x number of vertical pixels", that is, the image resolution can be used to represent the size of the image data in the width and height dimensions. Therefore, if the image data is converted from the first data arrangement manner to the second data arrangement manner, the image resolution after conversion is reduced, that is, the image resolution is compressed.
[0041] In the embodiments of the present application, the electronic device can convert the image data of the first to-be-processed image from the first data arrangement manner to the second arrangement manner through the first operator, to compress the resolution of the first to-be-processed image.
[0042] Among them, the first operator can convert the image data arranged in the space dimension to the depth dimension. For example, the first operator can be a space-to-depth function in the algorithm platform Tensor Flow.
[0043] Please refer to Figure 3 , Figure 3 This is an example diagram of a method for transforming the data arrangement of a first image to be processed based on a space-to-depth function, as disclosed in one embodiment.
[0044] In deep learning frameworks, image data can be represented in NCHW or NHWC format. Here, N represents the batch, C represents the channel, H represents the height, and W represents the width.
[0045] like Figure 3 As shown, the first image to be processed is represented in NHWC format, and the first data layout 310 is NHWC = [1, 4, 6, 3]. That is, when the image data of the first image to be processed is arranged in the first data layout, the batch size is 1, the height is 4, the width is 6, and the number of channels is 3. Furthermore, in the length dimension of the first data layout 310, the first and second columns of data are image data of the red (Red, R) channel, the third and fourth columns of data are image data of the green (Green, G) channel, and the fifth and sixth columns are image data of the blue (Blue, B) channel.
[0046] The image data of the first image to be processed is transformed from the first data arrangement method 310 to the second data arrangement method 320 based on the space-to-depth function. For example... Figure 3 As shown, the second data arrangement 320 is NHWC = [1, 2, 3, 12]. That is, when the image data of the first image to be processed is arranged in the second data arrangement, the batch size is 1, the height is 2, the width is 3, and the number of channels is 12.
[0047] In other words, if the data arrangement of the first image to be processed is transformed using the space-to-depth function, and the image data is represented in NHWC format, the resulting second image to be processed can be represented as follows:
[0048] [in_batch_1, in_height_1 / block_size_1, in_width_1 / block_size_1, in_channel_1*(block_size_1*block_size_1)].
[0049] Here, `in_batch_1` represents the batch number of the second batch of images to be processed, which is the same as the batch number of the first batch of images to be processed. `in_height_1` represents the height of the first batch of images to be processed, and `in_width_1` represents the width of the first batch of images to be processed. `block_size_1` can be a custom parameter that can be set according to actual business needs. The larger the value of `block_size_1`, the greater the degree of image resolution compression, and the lower the resolution of the compressed image.
[0050] Optionally, the electronic device can compress the resolution of the first image to be processed to half based on the aforementioned first operator to obtain the second image to be processed. For example, the value of the aforementioned block_size_1 can be set to half the height or width of the first image to be processed, thereby compressing the resolution of the image to half of its original value based on the space-to-depth function.
[0051] like Figure 3 As shown, when the image data of the first image to be processed is converted from the first data arrangement to the second data arrangement, the image resolution decreases, but the information contained in the image data is not lost; the information in the image data is completely preserved. Since the conversion of the data arrangement can reduce the image resolution, the operation of converting the image data from the first data arrangement to the second data arrangement can be considered a generalized downsampling operation.
[0052] While interpolation and other downsampling operations can compress image resolution, they can easily lead to the loss of image information. For example, a set of image data [1, 512, 512, 3] in NHWC format, after interpolation and downsampling, yields image data [1, 256, 256, 3]. Although the image resolution decreases from 512×512 to 256×256, some image information is clearly lost. However, the same set of image data, after being processed by the aforementioned first operator, yields image data [1, 256, 256, 12], which retains the original image information.
[0053] 220. The depth of the second image to be processed is estimated by using the trained depth estimation network to obtain the first depth estimation map corresponding to the second image to be processed.
[0054] The depth estimation network can be any type of neural network model, such as Convolutional Neural Network (CNN), U-Net, etc., but is not limited to these.
[0055] The depth estimation network can be trained by using sample data. The sample data can include a plurality of first sample images and first reference depth maps corresponding to the first sample images. The resolution of the first sample image is consistent with that of the first reference depth map. The first reference depth map can include a depth reference value corresponding to each pixel point in the first sample image. The depth reference value is an accurate true value, which can be manually marked or measured by a ranging device such as a laser radar.
[0056] In the design of the depth estimation network, the resolution of the input image of the depth estimation network is generally fixed. The smaller the resolution of the input image, the smaller the amount of calculation required for training and applying the depth estimation network.
[0057] In the embodiments of the present application, the second to-be-processed image input into the depth estimation network is an image with a compressed resolution. Therefore, when the image features of the second to-be-processed image are transmitted in each network layer of the depth estimation network, the resolution of the image features is reduced as a whole. Moreover, since the image information in the first to-be-processed image is not lost in the second to-be-processed image, the image information contained in the image features transmitted in the depth estimation network is not reduced. The reduction of the resolution can improve the calculation efficiency of the depth estimation network, and the retained image information can ensure the depth estimation accuracy of the depth estimation network.
[0058] In the embodiments of the present application, after each network layer included in the depth estimation network processes the second to-be-processed image in turn, the first depth estimation image can be output. The first depth estimation image can include a depth estimation value corresponding to each pixel point in the second to-be-processed image, and the resolution of the first depth estimation image is consistent with that of the second to-be-processed image, i.e., the resolution of the first depth estimation image is smaller than that of the first to-be-processed image.
[0059] The purpose of the depth estimation by the electronic device through the depth estimation network is to determine the blurring strength required for the blurring operation by using the depth estimation value. The processing of the image by the depth estimation network does not change the resolution of the image. The resolution of the second depth estimation image is consistent with that of the second to-be-processed image, which is different from the resolution of the first to-be-processed image obtained originally. Therefore, after obtaining the first depth estimation image output by the depth estimation network, the electronic device can further convert the data arrangement mode of the first depth estimation image, so that the resolution of the first depth estimation image is restored to be consistent with that of the first to-be-processed image.
[0060] 230、convert the image data of the first depth estimation image from the second data arrangement mode to the first data arrangement mode based on the second operator, to obtain a second depth estimation image with a resolution consistent with that of the first to-be-processed image.
[0061] The second operator can convert the image data arranged in the depth dimension to the spatial dimension. The conversion operation performed by the second operator can be regarded as the reverse operation of the conversion operation performed by the first operator. Similar to the first operator, the conversion operation of the arrangement of the image data based on the second operator can restore the resolution of the image without losing the image information. That is, compared with the first depth estimation image, the resolution of the second depth estimation image is restored to be consistent with the first to-be-processed image, and the depth estimation values included in the second depth estimation image do not change and are still the more accurate depth estimation values output by the depth estimation network.
[0062] For example, the second operator can be a depth-to-space function of an algorithm platform Tensor Flow. Please refer to Figure 4 , Figure 4 Figure 1 is an example diagram of converting the arrangement of the first depth estimation image based on the depth-to-space function according to an embodiment.
[0063] As shown in Figure 4 , the first depth estimation image is represented in the NHWC format, and the second data arrangement 410 is NHWC = [1, 2, 2, 4]. That is, when the image data of the first depth estimation image is arranged in the second data arrangement, the batch is 1, the height is 2, the width is 2, and the number of channels is 4.
[0064] The image data of the first depth estimation image is converted from the second data arrangement 410 to the first data arrangement 420 based on the depth-to-space function. As shown in Figure 4 , the second data arrangement 420 is NHWC = [1, 4, 4, 1]. That is, when the image data of the first depth estimation image is arranged in the first data arrangement, the batch is 1, the height is 4, the width is 4, and the number of channels is 1.
[0065] That is, if the arrangement of the first depth estimation image is converted based on the depth-to-space function, and the restored second depth estimation image is represented in the data format of NHWC, the converted second depth estimation image can be represented as follows:
[0066] [in_batch_2, in_height_2 * block_size_2, in_width_2 * block_size_2, in_channel_2 / (block_size_2 * block_size_2)].
[0067] in_batch_2 can be used to represent the batch number of the second depth estimation image, which is the same as the batch number of the first depth estimation image, in_height_2 can be used to represent the height of the first to-be-processed image, in_width_2 can be used to represent the width of the first to-be-processed image. block_size_2 can be a self-defined parameter, which can be set according to actual business requirements. Optionally, if the first operator is a space-to-depth function, the value of block_size_2 can be the same as the value of block_size_1 described above, so as to restore the resolution of the first depth estimation image to be consistent with the resolution of the first to-be-processed image.
[0068] It should be noted that, Figure 4 The first depth estimation image and the second depth estimation image shown are only examples, Figure 4 The first depth estimation image shown is obtained by depth estimation of a second to-be-processed image input into the network by the depth estimation network, Figure 4 The first depth estimation image shown is not necessarily obtained by depth estimation of the second to-be-processed image. Figure 3 The second to-be-processed image shown is not necessarily input into the network.
[0069] As can be seen in the foregoing embodiments, the electronic device can first compress the resolution of the first to-be-processed image by the first operator to reduce the resolution of the image features that need to be processed by the depth estimation network, so as to reduce the amount of calculation required by the depth estimation network to perform depth estimation on the second to-be-processed image input into the network, and improve the calculation efficiency of the depth estimation network.
[0070] At the same time, since the operation of the first operator to compress the resolution of the first to-be-processed image does not lose image information in the image, the depth estimation network can still output a first depth estimation image with high accuracy, and the accuracy of the depth estimation will not decrease due to the decrease in the resolution of the input image.
[0071] After the electronic device obtains the first depth estimation image output by the depth estimation network, the resolution of the first depth estimation image can be restored by using a second operator that is a reverse operation of the first operator, so that the resolution of the second depth estimation image obtained after restoration is consistent with the resolution of the original input first to-be-processed image, and further so that the electronic device can directly calculate the blurring degree required for the blurring operation using the second depth estimation image, thereby achieving a good image blurring effect.
[0072] In addition, due to the decrease in the amount of calculation required for depth estimation, the electronic device can support real-time preview of the image blurring effect.
[0073] In an embodiment, in order to further improve the depth estimation accuracy of the depth estimation network, an attention mechanism can also be introduced to improve the depth estimation network. Please refer to Figure 5 , Figure 5 is a structural schematic diagram of a depth estimation network disclosed in an embodiment. As shown in Figure 5 , the depth estimation network 500 can include an encoder 510 and a decoder 520.
[0074] The encoder 510 can be any backbone network with encoding capability, for example, any network in the Mobile Net series, and can include one or more network layers, which are not specifically limited. The encoder 510 can obtain a second to-be-processed image input to the depth estimation network, and perform encoding processing on the second to-be-processed image to obtain an encoded feature image.
[0075] The decoder 520 can include one or more attention modules 521. The attention module 521 can be any neural network module for processing image features based on an attention mechanism. For example, the attention module 521 can include a network layer, a first mapping function, and a multiplication operator. The network layer can be a convolution layer, the first mapping function can be a Sigmoid function, and the multiplication operator is used to perform a multiplication operation. It should be noted that in other embodiments, the attention module 521 can also be a channel attention module (Squeeze-and-Excitation Networks, SENet) or a convolutional network attention module (Convolutional Block Attention Module, CBAM), which are not specifically limited.
[0076] In the depth estimation network disclosed in the embodiments of the present application, the attention module can be used to perform depth estimation on the encoded feature image output by the encoder based on an attention mechanism, thereby obtaining a first depth estimation image corresponding to the second to-be-processed image.
[0077] In an embodiment, please continue to refer to Figure 5 , as shown in Figure 5 , the decoder 520 can include two or more attention modules 521 connected in sequence. Each attention module 521 can include a network layer 21a, a first mapping function 21b, and a multiplication operator 21c.
[0078] In an embodiment, the implementation of the attention module in the depth estimation network decoder for performing depth estimation on the encoded feature image output by the encoder based on an attention mechanism to obtain a first depth estimation image output by the attention module can include:
[0079] The input image input to the Mth attention module is decoded by the network layer of the Mth attention module to obtain a first original feature image, the first original feature image is mapped by a first mapping function in the Mth attention module, and the first correlation feature image obtained after mapping is multiplied by the first original feature image by a multiplication operator in the Mth attention module to obtain a decoding feature image output by the Mth attention module.
[0080] Wherein, M is a positive integer greater than or equal to 1. When M is equal to 1, the input image is the encoding feature image output by the encoder; when M is greater than 1, the input image is the decoding feature image output by the M-1th attention module; and the decoding feature image output by the last attention module of the decoder can be the first depth estimation image finally output by the depth estimation network, and the resolutions of the decoding feature images output by different attention modules are different.
[0081] As shown in Figure 5 The network layer 21a of the first attention module 521 can receive the encoding feature image output by the encoder, and the network layer 21a can decode the encoding feature image to obtain a first original feature image. The first mapping function 21b can map the first original feature image to obtain a first correlation feature image. The multiplication operator 21c can multiply the first correlation feature image output by the first mapping function 21b and the first original feature image output by the network layer 21a to obtain a decoding feature image output by the first attention module 521.
[0082] The decoding feature image output by the first attention module 521 can be input to the network layer of the second attention module, and the second attention module can process the decoding feature image output by the first attention module. The process is similar to that of the first attention module 521, and the following content will not be repeated.
[0083] As shown in Figure 5 The decoding feature image output by the last attention module of the decoder 520 can be used as the first depth estimation image finally output by the depth estimation network 500.
[0084] As can be seen, the attention module can multiply the first relevant feature image obtained after mapping based on the first mapping function with the first original feature image output by the network layer of the attention module. This enhances the expressive power of image features based on the attention mechanism, thereby improving the accuracy and precision of depth estimation. Furthermore, the decoded feature images output by different attention modules have different resolutions. Therefore, the decoded feature images output by each attention module can be considered as depth estimation results at different resolutions corresponding to the second image to be processed. That is, the decoder can calculate attention for image features at different scales, thereby improving the accuracy of depth estimation for objects at different scales.
[0085] In this embodiment of the application, the electronic device can perform depth estimation on the second image to be processed with resolution compression based on any of the aforementioned deep neural networks, and after obtaining the first depth estimation image output by the deep neural network, restore the resolution of the first depth estimation image to be consistent with the resolution of the first image to be processed based on the aforementioned second operator, and then use the second depth estimation image after resolution restoration to perform blurring processing on the first image to be processed.
[0086] In one embodiment, the electronic device can use a single-channel depth estimation image to blur the first image to be processed. Therefore, after obtaining the first depth estimation image output by the depth estimation network, and converting the image data of the first depth estimation image from the second data arrangement to the first data arrangement based on the second operator to restore the resolution of the first depth estimation image to be consistent with the resolution of the first image to be processed, the electronic device can further map the resolution-restored first depth estimation image based on the second mapping function, and determine the mapped image as the second depth estimation image. The second mapping function can be any normalization function, such as the Sigmoid function, the Softmax function, etc., without specific limitations.
[0087] To more clearly illustrate the image processing method disclosed in the embodiments of this application, please refer to the exemplary examples. Figure 6 , Figure 6 This is a schematic flowchart of an image processing method disclosed in another embodiment.
[0088] like Figure 6 As shown, after the first image to be processed is processed by the first operator 610 (sapce-to-depth function), the first data arrangement of the first image to be processed is converted to the second data arrangement, resulting in a second image to be processed with reduced resolution.
[0089] The second to-be-processed image is input to the depth estimation network 620, and the encoder 621 of the depth estimation network can include at least two network layers connected in sequence. The encoder 621 can perform encoding processing on the second to-be-processed image based on the foregoing network layers, and the last network layer included in the encoder 621 can input the encoded feature image of the second to-be-processed image to the decoder 622.
[0090] The decoder 622 can include at least three attention modules connected in sequence.
[0091] Taking the first attention module 6221 as an example, the first attention module 6221 can include a network layer, a first mapping function (Sigmoid function), and a multiplication operator.
[0092] The encoded feature image output by the encoder 621 is subjected to decoding processing of the network layer in the first attention module 6221, to obtain a first original feature image. After the first original feature image is mapped by the first mapping function, a first correlation feature image can be obtained. The first correlation feature image and the first original feature image are multiplied based on the multiplication operator, to obtain a decoding feature map output by the first attention module.
[0093] The decoding feature map output by the first attention module is input to the second attention module, and the operations performed by each attention module included in the decoder 620 can refer to the operations performed by the first attention module, and the following content will not be described again.
[0094] The decoding feature map output by the last attention module included in the decoder 620 can be used as a first depth estimation image output by the depth estimation network.
[0095] After the first depth estimation image is processed by the second operator 630 (depth-to-sapce function), the image data of the first depth estimation image is converted from the second data arrangement mode to the first data arrangement mode, to obtain a resolution-restored first depth estimation image. The resolution-restored first depth estimation image is further mapped by the second mapping function 640 (Sigmoid function), to obtain a second depth estimation image. The electronic device can use the second depth estimation image to perform blurring processing on the first to-be-processed image.
[0096] It can be seen that in the foregoing embodiments, the electronic device can first compress the resolution of the first to-be-processed image based on the first operator, and then perform depth estimation on the second to-be-processed image with the resolution compressed by using the depth estimation network. The compression of the resolution does not lose image information, reduces the calculation amount of depth estimation, and improves the accuracy of depth estimation. The depth estimation network can perform depth estimation on the second to-be-processed image based on the attention mechanism, which can further improve the accuracy and precision of depth estimation. The first depth estimation image output by the depth estimation network is restored in resolution based on the second operator, and is normalized based on the second mapping function, so that the second depth estimation image obtained finally can be directly applied to the calculation of the blurring processing.
[0097] In one embodiment, the specific implementation in which the electronic device performs blurring processing on the first to-be-processed image by using the second depth estimation image can include the following steps:
[0098] The electronic device determines the background region of the first to-be-processed image, and determines the depth estimation value corresponding to the background region according to the second depth estimation image; and determines the blurring level of the background region according to the depth estimation value corresponding to the background region, and performs blurring processing on the background region according to the blurring level.
[0099] In the foregoing embodiments, the electronic device can detect the foreground region or the background region selected by the user from the first to-be-processed image through clicking, long pressing, or other user operations. If the foreground region is selected from the first to-be-processed image by the user operation, the electronic device can determine the remaining image region in the first to-be-processed image except the foreground region selected by the user operation as the background region. Alternatively, the electronic device can also identify the portrait region from the first to-be-processed image by using a portrait recognition algorithm such as portrait segmentation or portrait matting, and determine the other image region in the first to-be-processed image except the portrait region as the background region. Alternatively, the electronic device can also determine the background region of the first to-be-processed image by using the depth estimation value included in the second depth estimation image, for example, the pixel points in the background region are pixel points with a depth estimation value less than a depth threshold in the second depth estimation image. The above are some embodiments of how the electronic device determines the background region of the first to-be-processed image, and the specific implementation is not limited.
[0100] After the electronic device determines the background region of the first to-be-processed image, the depth estimation value corresponding to the background region can be obtained from the second depth estimation image accordingly, that is, the depth estimation value corresponding to each pixel point included in the background region is obtained, so as to determine the blurring level of the background region according to the depth estimation value corresponding to each pixel point included in the background region.
[0101] Optionally, the blurring level of the background region can be positively correlated with the depth estimation value of the background region, that is, the greater the depth estimation value of the background region, the higher the blurring level, and the more blurred the background.
[0102] Further optionally, the background region can be divided into two or more background sub-regions according to the depth estimation value corresponding to each pixel point, and each background sub-region can correspond to a respective blur level. The blur level corresponding to the background sub-region can be in a positive correlation with the depth estimation value corresponding to the background sub-region.
[0103] For example, if the electronic device performs the blur processing on the background region in the manner of Gaussian blur, the blur level can be represented by the size of the Gaussian kernel. The Gaussian kernel can be represented by the following formula:
[0104]
[0105] where G(u,v) can be used to represent the Gaussian kernel, (u,v) can be used to represent the image coordinates of the pixel point, and σ can be a preset standard deviation.
[0106] The higher the blur level, the larger the size of the Gaussian kernel, and the more blurred the background region after the blur processing; the lower the blur level, the smaller the size of the Gaussian kernel, and the clearer the background region after the blur processing.
[0107] It can be seen that, based on the image processing method disclosed in the foregoing embodiments, the electronic device can obtain an accurate and high-precision second depth estimation image, and further perform the blur processing on the background region of the first to-be-processed image using the second depth estimation image, so that the blur effect of the background region is real and natural, and the shooting effect of the foreground scene can be highlighted.
[0108] The depth estimation network disclosed in the foregoing embodiments needs to be trained by sample data to accurately estimate the depth of the first to-be-processed image.
[0109] In one embodiment, the depth estimation network is obtained by adjusting the parameters in the depth estimation network to be trained with reference to the logarithmic loss between the second correlation feature map and the first reference depth image; the second correlation feature map is generated by the depth estimation network to be trained in the process of estimating the depth of the resolution-compressed second sample image, the depth estimation network to be trained estimates the depth of the second sample image based on the attention mechanism through the attention module of the decoder, and the attention module includes the first mapping function, and the second correlation feature image is output by the first mapping function.
[0110] In order to more clearly illustrate the process of training the depth estimation network in the embodiments of the present application, the following describes the training method of the depth estimation network.
[0111] Please refer to Figure 7 , Figure 7Fig. 1 is a schematic diagram of a method flow of a depth estimation network training method according to an embodiment. The method can be applied to a cloud server, a personal computer or other computing platform with strong computing capability, and the trained depth estimation network is transmitted to any of the aforementioned electronic devices for storage. Alternatively, the method can also be applied to the aforementioned electronic device, which is not limited in particular, and the training method of the depth estimation network is described below with the electronic device as an example. As shown in Fig. 1, the method can include the following steps: Figure 7
[0112] 710, converting the image data of the first sample image from the first data arrangement to the second data arrangement based on the first operator to obtain a second sample image with compressed resolution.
[0113] The first sample image can be any one of the multiple images included in the sample data. The first sample image can correspond to a first reference depth image, the resolution of the first reference depth image is consistent with the resolution of the first sample image, and the first reference depth image can include depth reference values corresponding to each pixel point in the first sample image.
[0114] The implementation of the electronic device converting the image data of the first sample image based on the first operator can refer to the foregoing embodiments, and the following content will not be repeated.
[0115] 720, performing depth estimation on the second sample image by the depth estimation network to be trained, and obtaining a mapping result generated by the depth estimation network to be trained in the depth estimation process.
[0116] As shown in the foregoing embodiments, the depth estimation network can include an encoder and a decoder, and the decoder can include one or more attention modules, each of which can include a network layer, a first mapping function and a multiplication operator. In the process of the depth estimation network performing depth estimation on the second sample image, the image features of the second sample image can be transmitted in each attention module included in the decoder. Each attention module can decode the image input into the attention module through the network layer, and map the original feature image obtained after decoding by the network layer through the first mapping function, and the second correlation feature image obtained after mapping can be used as the mapping result generated by the depth estimation network in the depth estimation process.
[0117] The process of the depth estimation network to be trained performing depth estimation on the second sample image is similar to the process of the trained depth estimation network performing depth estimation on the second sample image in the foregoing embodiments, and the following content will not be repeated.
[0118] It can be seen that each attention module included in the decoder of the depth estimation network can generate a second correlation feature image based on the first mapping function. Therefore, the mapping result obtained by the electronic device can include one or more second correlation feature images, and each second correlation feature image included in the mapping result is generated by an attention module included in the decoder based on the first mapping function.
[0119] It should be noted that, as shown in the foregoing embodiments, when the decoder includes multiple attention modules, the resolutions of the image features processed by each attention module are different. Therefore, when the mapping result includes multiple second correlation feature images, the resolutions of the second correlation feature images are also different.
[0120] 730. Determine the training loss according to the logarithmic loss between the mapping result and the first reference depth image.
[0121] The resolution of the first reference depth image is consistent with the resolution of the first sample image before resolution compression. The resolutions of the second correlation feature images included in the mapping result are not the same as the resolution of the first reference depth image, and therefore, when calculating the training loss, the problem of scale inconsistency caused by the different resolutions needs to be considered. In order to eliminate the problem of scale inconsistency caused by the different resolutions, the electronic device can determine the training loss according to the logarithmic loss between the mapping result and the first reference depth image.
[0122] In one embodiment, when the mapping result includes one second correlation feature image, the electronic device can calculate the logarithmic loss between the second correlation feature image and the first reference image as the training loss.
[0123] In one embodiment, when the mapping result includes two or more second correlation feature images, the electronic device can calculate the logarithmic loss between each second correlation feature image and the first reference image, respectively, and sum the logarithmic loss corresponding to each second correlation feature image to obtain the training loss.
[0124] That is, when the decoder includes two or more attention modules, the electronic device can calculate the training loss using the second correlation feature images with different resolutions output by different attention modules, and adjust the parameters in the depth estimation network using the training loss, so as to improve the depth estimation accuracy of the depth estimation network for objects of different scales.
[0125] For example, the logarithmic loss between each second correlation feature image and the first reference depth image can be calculated by the following formula:
[0126]
[0127]
[0128] wherein D(g) can be used to represent a logarithmic loss, T can be used to represent a total number of pixel points, i can be used to represent an i-th pixel point, and l can be used to represent a custom weight of a hyperparameter in the depth estimation network, can be used to represent a depth prediction value of the i-th pixel point, and d i can be used to represent a depth reference value of the i-th pixel point.
[0129] 740. Adjusting the parameters in the depth estimation network by using the training loss.
[0130] The electronic device can adjust the parameters in the depth estimation network by combining the training loss based on Stochastic Gradient Descent (SGD), Momentum, etc. Wherein the parameters in the encoder and / or decoder of the depth estimation network can be adjusted.
[0131] In one embodiment, the depth estimation network to be trained can output a first depth training image corresponding to the second sample image in addition to the aforementioned second correlation feature image in the process of depth estimation on the second sample image. The first depth training image includes a depth prediction value corresponding to each pixel point in the second sample image, which is predicted by the depth estimation network to be trained.
[0132] As shown in the foregoing embodiment, the first depth training image can be output by the last attention module in the decoder of the depth estimation network.
[0133] In order to train a depth estimation network with higher depth estimation accuracy, the electronic device can also use the first depth training image to participate in the calculation of the training loss. Specifically, it can include the following steps:
[0134] The electronic device can convert the image data of the first depth training image from the second data arrangement to the first data arrangement based on the aforementioned second operator to obtain a second depth training image whose resolution is restored to be consistent with that of the first sample image; and map the second depth training image by using the aforementioned second mapping function to obtain a third depth training image.
[0135] The electronic device calculates the training loss according to one or more second correlation feature images included in the mapping result, the third depth training image, and the first reference depth image.
[0136] The electronic device calculates a logarithmic loss between each second correlation feature image included in the mapping result and the first reference depth image, and calculates a logarithmic loss or a non-logarithmic loss between the third depth training image and the first reference depth image. The non-logarithmic loss can include an LI regularization loss, a cross-entropy loss, etc., and is not limited in particular. Since the resolution of the third depth training image is restored to be consistent with that of the first sample image, there is no problem of scale inconsistency between the third depth training image and the first reference depth image, and the non-logarithmic loss between the third depth training image and the first reference depth image can be calculated.
[0137] The training loss can be the sum of the logarithmic loss corresponding to each second correlation feature image and the logarithmic or non-logarithmic loss between the third depth training image and the first reference depth image.
[0138] In order to more clearly describe the depth estimation network training method disclosed in the embodiments of the present application, please refer to Figure 8 , Figure 8 is a method flow diagram of a depth estimation network training method disclosed in another embodiment.
[0139] As shown in Figure 8 , after the first sample image is processed by the first operator 810 (sapce-to-depth function), the first data arrangement mode of the first sample image is converted to the second data arrangement mode, and a second sample image with compressed resolution is obtained.
[0140] The second sample image is input into the depth estimation network 820, and the structure of the depth estimation network 820 is similar to that of the depth estimation network 620 in the foregoing embodiments. The process of depth estimation of the second sample image by the depth estimation network 820 is similar to the process of depth estimation of the second sample image by the depth estimation network 620 in the foregoing embodiments, and the following content will not be described in detail.
[0141] As shown in Figure 8 , the decoder 822 included in the depth estimation network 820 can include at least three attention modules connected in sequence. Taking the first attention module 8221 as an example, the first mapping function 22b (Sigmoid function) included in the first attention module 8221 can output a second correlation feature image. On the one hand, the second correlation feature image output by the first attention module 8221 can be used to participate in the calculation of the training loss, and on the other hand, the second correlation feature image can be multiplied with the third original feature image output by the network layer 22a through the multiplication operator 22c, and the multiplication result is sent to the next attention module for processing.
[0142] As shown in Figure 8As shown, the multiplication operator of the last attention module included in the decoder 822 can multiply the second correlation feature map output by the first mapping function with the third original feature map obtained after decoding the network layer input to the last attention module, to obtain the first depth training image.
[0143] The first depth training image can be converted into a second depth training image after being processed by a second operator 830 (depth-to-sapce function), and the second depth training image can be converted into a third depth training image after being mapped by a second mapping function 840 (Sigmoid function). The third depth training image can be used to calculate the training loss.
[0144] As can be seen, in the foregoing embodiments, the training loss of the depth estimation network can be calculated based on the attention mechanism, and the training loss can be calculated using the second correlation feature map output by the first mapping function included in the attention module, so as to train the depth estimation network to obtain a depth estimation network with improved depth estimation accuracy. Moreover, the training loss can be calculated using the second correlation feature maps of different resolutions output by different attention modules, so as to train the neural network to be able to accurately estimate the depth of objects of different sizes.
[0145] Please refer to Figure 9 , Figure 9 FIG. 1 is a structural schematic diagram of an image processing apparatus disclosed in an embodiment, which can be applied to any one of the electronic devices described above. As shown in Figure 9 The image processing apparatus 900 can include a first conversion module 910, a depth estimation module 920, and a second conversion module 930.
[0146] The first conversion module 910 is configured to convert image data of a first to-be-processed image from a first data arrangement manner to a second data arrangement manner based on a first operator, to obtain a second to-be-processed image with compressed resolution; the components in the depth dimension of the first data arrangement manner and the second data arrangement manner are channel numbers, and the channel number corresponding to the first data arrangement manner is less than the channel number corresponding to the second data arrangement manner.
[0147] The depth estimation module 920 is configured to perform depth estimation on the second to-be-processed image by using the trained depth estimation network, to obtain a first depth estimation image corresponding to the second to-be-processed image, and the resolution of the first depth estimation image is consistent with the resolution of the second to-be-processed image.
[0148] Optionally, the resolution of the second to-be-processed image can be one half of the resolution of the first to-be-processed image.
[0149] The second conversion module 930 is configured to convert the image data of the first depth estimation image from the second data arrangement to the first data arrangement based on a second operator, to obtain a second depth estimation image with a resolution consistent with the first to-be-processed image.
[0150] In an embodiment, the depth estimation network can include an encoder and a decoder, and the decoder includes an attention module.
[0151] The encoding unit is configured to perform encoding processing on the second to-be-processed image by the encoder, to obtain an encoded feature image.
[0152] The decoding unit is configured to perform depth estimation on the encoded feature image based on an attention mechanism by the attention module in the decoder, to obtain the first depth estimation image corresponding to the second to-be-processed image.
[0153] In an embodiment, the decoder includes at least two attention modules connected in sequence, and each attention module includes a network layer, a first mapping function, and a multiplication operator.
[0154] The decoding unit is further configured to perform decoding processing on the encoded feature image output by the encoder by the network layer in the first attention module, to obtain a first original feature image, perform mapping on the first original feature image by the first mapping function in the first attention module, multiply the first correlation feature image obtained after the mapping by the encoded feature image by the multiplication operator in the first attention module, and obtain a decoding feature image output by the first attention module.
[0155] In addition, perform decoding processing on the decoding feature image output by the M-1th attention module by the network layer in the Mth attention module, to obtain a second original feature image, perform mapping on the second original feature image by the first mapping function in the Mth attention module, multiply the second correlation feature image obtained after the mapping by the decoding feature image output by the M-1th attention module by the multiplication operator in the Mth attention module, and obtain a decoding feature image output by the Mth attention module; M is a positive integer greater than or equal to 2, the first depth estimation image is a decoding feature image output by the last attention module included in the decoder, and the resolutions of the decoding feature images output by different attention modules are different.
[0156] In an embodiment, the second conversion module 930 can also be configured to convert the image data of the first depth estimation image from the second data arrangement to the first data arrangement based on a second operator, so as to restore the resolution of the first depth estimation image to be consistent with the resolution of the first to-be-processed image; and map the first depth estimation image after the resolution is restored by using a second mapping function, and determine the image obtained after the mapping as the second depth estimation image.
[0157] In an embodiment, the image processing apparatus 900 can further include a blurring module.
[0158] The blurring module can be configured to, after obtaining the second depth estimation image output by the second conversion module 930, determine a background region of the first to-be-processed image, and determine a depth estimation value corresponding to the background region according to the second depth estimation image; and determine a blurring level of the background region according to the depth estimation value corresponding to the background region, and perform blurring processing on the background region according to the blurring level.
[0159] In an embodiment, the image processing apparatus 900 can further include a training module.
[0160] The first conversion module 910 can also be configured to, before the first conversion module 910 converts the image data of the first to-be-processed image from the first data arrangement to the second data arrangement based on the first operator to obtain the second to-be-processed image with the compressed resolution, convert the image data of the first sample image from the first data arrangement to the second data arrangement based on the first operator to obtain the second sample image with the compressed resolution.
[0161] The training module can be configured to obtain a mapping result generated in a process in which a to-be-trained depth estimation network performs depth estimation on the second sample image, the mapping result including a second correlation feature image, the second correlation feature image being output by a first mapping function in the to-be-trained depth estimation network; the to-be-trained depth estimation network including a decoder, and the first mapping function belonging to an attention module of the decoder.
[0162] The training module can also be configured to determine a training loss according to a logarithmic loss between the mapping result and a first reference depth image, the first reference depth image including a depth reference value corresponding to each pixel point in the first sample image, and the resolution of the first reference depth image being consistent with the resolution of the first sample image.
[0163] The training module can also be configured to adjust parameters in the depth estimation network by using the training loss.
[0164] In one embodiment, the decoder comprises at least two attention modules connected in sequence; the mapping result comprises second correlation feature images output by the first mapping functions included in each attention module respectively, and the resolutions of the second correlation feature images output by different attention modules are different.
[0165] The training module is further configured to calculate a log loss of each second correlation feature image and the first reference depth image respectively, and sum the log losses corresponding to each second correlation feature image to obtain the training loss.
[0166] In one embodiment, the second conversion module 930 is further configured to obtain a first depth training image output by the depth estimation network to be trained, the first depth training image comprising a depth prediction value corresponding to each pixel in the second sample image; and convert image data of the first depth training image from the second data arrangement mode to the first data arrangement mode based on the second operator to obtain a second depth training image whose resolution is restored to be consistent with that of the first sample image.
[0167] The second depth training image is mapped by the second mapping function to obtain a third depth training image.
[0168] The training module is further configured to calculate a training loss based on the mapping result generated in the process of depth estimation of the second sample image by the depth estimation network to be trained, the third depth training image and the first reference depth image.
[0169] It can be seen that, by implementing the image processing apparatus disclosed in the foregoing embodiments, the resolution of the first to-be-processed image can be compressed by the first operator first, which can not only reduce the calculation amount of the depth estimation network in depth estimation, but also ensure the accuracy of the depth estimation. Meanwhile, after obtaining the first depth estimation image output by the depth estimation network, the resolution of the first depth estimation image can be restored by the second operator which is the inverse operation of the first operator, so that the electronic device can directly calculate the blurring degree required for the blurring operation based on the second depth estimation image whose resolution is restored, thereby achieving a good image blurring effect. In addition, the calculation amount required for depth estimation is reduced, so that the electronic device can support real-time preview of the image blurring effect.
[0170] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of an electronic device disclosed in an embodiment of the present application.
[0171] As Figure 10 shown, the electronic device can comprise:
[0172] a memory 1010 storing executable program codes;
[0173] A processor 1020 coupled to the memory 1010;
[0174] The processor 1020 invokes the executable program code stored in the memory 1010 to execute any one of the image processing methods disclosed in the embodiments of the present application.
[0175] It should be noted that, Figure 10 The electronic device shown can also include a power supply, input keys, a camera, a speaker, a screen, RF circuitry, a Wi-Fi module, a Bluetooth module, sensors, and other components not shown, and the embodiments of the present application will not be described here.
[0176] The embodiments of the present application disclose a computer readable storage medium storing a computer program, wherein the computer program is executed by a processor to enable the processor to implement any one of the image processing methods disclosed in the embodiments of the present application.
[0177] The embodiments of the present application disclose a computer program product including a non-transitory computer readable storage medium storing a computer program, and the computer program is operable to enable a computer to implement any one of the image processing methods disclosed in the embodiments of the present application.
[0178] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in any suitable manner in one or more embodiments. Those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application.
[0179] In various embodiments of the present application, it should be understood that the size of the serial number of the above processes does not mean the inevitable sequence of execution, and the execution sequence of the processes should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0180] The units described above as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e. can be located in one place, or can be distributed to multiple network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiments of the present application.
[0181] In addition, each of the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware, or in the form of a software functional unit.
[0182] When the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solutions of the present application, essentially or in part, or all or part of the technical solutions, can be embodied in the form of a software product. The computer software product is stored in a memory and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc., and specifically can be a processor in the computer device) to perform all or part of the steps of the methods according to the embodiments of the present application.
[0183] Those of ordinary skill in the art can understand that all or part of the steps of the various methods in the above embodiments can be completed by a program instructing related hardware, and the program can be stored in a computer-readable storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disk storage, a magnetic disk storage, a magnetic tape storage, or any other medium of a computer-readable type.
[0184] The image processing method and device, the electronic device, and the computer storage medium disclosed by the embodiments of the present application are described in detail above, and the principles and implementation manners of the present application are described by applying specific examples. The above embodiment description is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manner and application range will be changed, and the above description should not be understood as a limitation on the present application.
Claims
1. An image processing method, characterized by, The method comprises: Converting image data of a first to-be-processed image from a first data arrangement mode to a second data arrangement mode based on a first operator to obtain a second to-be-processed image with compressed resolution; the number of components in the depth dimension of the first data arrangement mode and the second data arrangement mode is the number of channels, and the number of channels corresponding to the first data arrangement mode is less than the number of channels corresponding to the second data arrangement mode; Performing depth estimation on the second to-be-processed image through a depth estimation network to obtain a first depth estimation image corresponding to the second to-be-processed image, wherein the resolution of the first depth estimation image is consistent with the resolution of the second to-be-processed image; wherein the depth estimation network is obtained by adjusting parameters in a to-be-trained depth estimation network with reference to a logarithmic loss between a second correlation feature map and a first reference depth image; the second correlation feature map is generated by the to-be-trained depth estimation network in the process of performing depth estimation on a second sample image with compressed resolution, the to-be-trained depth estimation network performs depth estimation on the second sample image based on an attention mechanism through an attention module of a decoder, the attention module comprises a first mapping function, and the second correlation feature image is output by the first mapping function; and the second sample image is obtained after converting image data of a first sample image from the first data arrangement mode to the second data arrangement mode based on the first operator; the first reference depth image comprises a depth reference value corresponding to each pixel point in the first sample image, and the resolution of the first reference depth image is consistent with the resolution of the first sample image; Converting image data of the first depth estimation image from the second data arrangement mode to the first data arrangement mode based on a second operator to obtain a second depth estimation image with a resolution consistent with the first to-be-processed image.
2. The method of claim 1, wherein, The resolution of the second to-be-processed image is half of the resolution of the first to-be-processed image.
3. The method of claim 1, wherein, The depth estimation network comprises an encoder and a decoder, the decoder comprises an attention module, and the depth estimation network trained performs depth estimation on the second to-be-processed image to obtain a first depth estimation image corresponding to the second to-be-processed image, which comprises: Performing encoding processing on the second to-be-processed image through the encoder to obtain an encoded feature image; Performing depth estimation on the encoded feature image based on an attention mechanism through the attention module in the decoder to obtain a first depth estimation image corresponding to the second to-be-processed image.
4. The method of claim 3, wherein, The decoder comprises at least two attention modules connected in sequence, each attention module comprises a network layer, a first mapping function and a multiplication operator, and the depth estimation on the encoded feature image based on an attention mechanism through the attention module in the decoder to obtain a first depth estimation image output by the attention module comprises: The input image input to the Mth attention module is decoded by a network layer of the Mth attention module to obtain a first original feature image, the first original feature image is mapped by a first mapping function in the Mth attention module, and the first original feature image is multiplied by a first correlation feature image obtained after mapping by a multiplication operator in the Mth attention module to obtain a decoding feature image output by the Mth attention module; Wherein, M is a positive integer greater than or equal to 1; when M is equal to 1, the input image is the encoding feature image output by the encoder; when M is greater than 1, the input image is a decoding feature image output by the M-1th attention module; the first depth estimation image is a decoding feature image output by the last attention module included in the decoder, and the resolutions of the decoding feature images output by different attention modules are different.
5. The method according to any one of claims 1 to 4, characterized in that, The image data of the first depth estimation image is converted from the second data arrangement mode to the first data arrangement mode based on the second operator to restore the resolution of the first depth estimation image to be consistent with the resolution of the first to-be-processed image to obtain a second depth estimation image, including: The image data of the first depth estimation image is converted from the second data arrangement mode to the first data arrangement mode based on the second operator to restore the resolution of the first depth estimation image to be consistent with the resolution of the first to-be-processed image; The first depth estimation image after resolution restoration is mapped by a second mapping function, and the image obtained after mapping is determined as a second depth estimation image.
6. The method of claim 1, wherein, After obtaining the second depth estimation image, the method further includes: determining the background region of the first to-be-processed image, and determining the depth estimation value corresponding to the background region according to the second depth estimation image; determining the blurring level of the background region according to the depth estimation value corresponding to the background region, and blurring the background region according to the blurring level.
7. An image processing apparatus characterized by comprising: including: a first conversion module configured to convert image data of a first to-be-processed image from a first data arrangement mode to a second data arrangement mode based on a first operator to obtain a second to-be-processed image with compressed resolution; The components in the depth dimension of the first data arrangement mode and the second data arrangement mode are the number of channels, and the number of channels corresponding to the first data arrangement mode is less than the number of channels corresponding to the second data arrangement mode; The depth estimation module is configured to perform depth estimation on the second to-be-processed image by using a depth estimation network to obtain a first depth estimation image corresponding to the second to-be-processed image, wherein the resolution of the first depth estimation image is consistent with the resolution of the second to-be-processed image; wherein the depth estimation network is obtained by adjusting parameters in a to-be-trained depth estimation network based on a logarithmic loss between a second correlation feature map and a first reference depth image; the second correlation feature map is generated by the to-be-trained depth estimation network during depth estimation on a second sample image with compressed resolution, and the to-be-trained depth estimation network performs depth estimation on the second sample image based on an attention mechanism through an attention module of a decoder, the attention module includes a first mapping function, and the second correlation feature image is output by the first mapping function; the second sample image is obtained by converting image data of a first sample image from the first data arrangement mode to the second data arrangement mode based on a second operator; the first reference depth image includes a depth reference value corresponding to each pixel point in the first sample image, and the resolution of the first reference depth image is consistent with the resolution of the first sample image; The second conversion module is configured to convert image data of the first depth estimation image from the second data arrangement mode to the first data arrangement mode based on a second operator to obtain a second depth estimation image with a resolution consistent with the first to-be-processed image.
8. An electronic device, comprising: The computer program is executed by the processor to implement the method of any one of claims 1 to 6.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Monocular depth estimation method
CN110060286A
Multi-modal information-based light field depth estimation method
CN112767466A
Image processing method and device, electronic equipment and storage medium
CN112991171A