A depth estimation network training method, depth estimation method and electronic device
Through the depth estimation network training method that integrates human semantic information and depth characteristics, the problem of inaccurate depth estimation in monocular depth estimation in complex scenarios is solved, which significantly improves the portrait blur effect.
Patent Information
- Application Number
- CN202311718735.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-12-13
AI Technical Summary
The existing monocular depth estimation method is difficult to accurately estimate image depth information in complex scenes such as night scenes and movements, resulting in unsatisfactory portrait blur effect.
The depth estimation network training method including human analytical subnet and monocular depth estimation subnet is used to train the monocular depth estimation subnet by fusing human semantic information features and depth features to improve the accuracy of depth estimation.
Improve the quality of portrait details and continuity in the image, especially in complex scenes such as night scenes and movements, reducing the error and leakage of portrait blur.
Smart Images

Figure CN118447288B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of network training technology, and in particular, relates to a depth estimation network training method, a depth estimation method and an electronic device. Background Art
[0002] Portrait blur processing refers to obtaining the depth relationship between the portrait and the background in the image based on the depth information of the image. Then, based on the depth relationship between the portrait and the background, the portrait is focused and the background is blurred. This portrait blur effect can highlight the human body and enhance the artistic conception.
[0003] Among them, the key to achieving the portrait blur effect is to correctly estimate the depth information of the image. Commonly used depth estimation methods mainly include binocular depth estimation methods and monocular depth estimation methods.
[0004] However, for images taken in complex scenes such as night scenes and sports, both the binocular depth estimation method and the monocular depth estimation method are unable to correctly estimate the depth information of the image, resulting in unsatisfactory portrait blur effect and affecting the user experience. Summary of the invention
[0005] The present application provides a depth estimation network training method, a depth estimation method and an electronic device, which can improve the portrait blur effect of an image based on a monocular depth estimation method.
[0006] In a first aspect, the present application provides a depth estimation network training method, wherein the depth estimation network includes a human body parsing subnetwork and a monocular depth estimation subnetwork, and the method includes: obtaining training data, wherein the training data includes a training image and real depth information corresponding to the training image; using the human body parsing subnetwork to perform feature extraction on the training image to obtain a first intermediate feature containing human body semantic information; using the monocular depth estimation subnetwork to perform feature extraction on the training image to obtain a second intermediate feature containing depth information; using the monocular depth estimation subnetwork to perform feature fusion on the first intermediate feature and the second intermediate feature to obtain a fused feature; using the monocular depth estimation subnetwork to perform depth estimation on the fused feature pair to obtain predicted depth information; based on the predicted depth information and the real depth information, training the monocular depth estimation subnetwork to obtain a trained monocular depth estimation subnetwork.
[0007] In this way, when training the monocular depth estimation subnetwork, the present application fuses the first intermediate features containing human semantic information, and predicts the depth information of the training image based on the fused features. In this way, the first intermediate features containing human semantic information can help the monocular depth estimation subnetwork better learn the subject information in the training image and the correlation between the various parts of the human body, so as to distinguish the various parts of the human body from the background in the image, so that the depth estimation result is consistent with the human body analysis result, thereby improving the accuracy of the subject depth and the depth details of various parts of the human body.
[0008] In one implementable manner, the monocular depth estimation subnetwork is trained based on the predicted depth information and the real depth information to obtain a trained monocular depth estimation subnetwork, including: determining the loss weights of different regions in the training image; wherein the loss weight of the portrait region in the training image is greater than the loss weight of the background region; calculating the depth loss based on the loss weights of different regions, the predicted depth information and the real depth information; when the depth loss is less than a depth loss threshold, terminating the training of the monocular depth estimation subnetwork to obtain a trained monocular depth estimation network.
[0009] In this way, when back-propagating to update the network parameters, since the loss gradient of the portrait area is larger, the network update will focus more on improving the quality of the depth map of the portrait area, so that the portrait area in the output depth map appears clearer and more coherent with smaller errors.
[0010] In one implementable manner, determining the loss weights of different areas in the training image includes: determining the key parts of the portrait in the training image based on the motion amplitude of each part of the human body and / or the color of each part of the human body in the training image; and determining that the loss weight of the key part is greater than the loss weights of other parts of the portrait.
[0011] In this way, when back-propagating to update network parameters, the network update will focus more on improving the quality of the depth map of key parts due to the larger loss gradient of key parts. In this way, the monocular depth sub-network can generate more accurate depth results in the areas where key parts such as hair, limbs, and dark clothes are located, making the human body area in the output depth map appear clearer and more coherent with smaller errors.
[0012] In one implementable manner, based on the motion amplitude of each human body part and / or the color of each human body part in the training image, determining the key parts of the portrait in the training image includes: determining the brightness of each human body part in the training image; determining the human body parts in the training image whose brightness is less than a brightness threshold as the key parts; and / or determining the motion amplitude of each human body part in the training image; determining the human body parts in the training image whose motion amplitude is greater than a motion amplitude threshold as the key parts.
[0013] Lightness can represent the brightness of a color from white to black. The higher the lightness, the lighter the color, and the lower the lightness, the darker the color. In this way, based on the lightness of each part of the human body, the dark and light areas in each part of the human body can be determined. The dark areas can be further determined as key parts.
[0014] In one implementable manner, the method further includes: acquiring an original training image and real human semantic information corresponding to the original training image; performing motion blur enhancement processing on the portrait in the original training image based on the real human semantic information to obtain a motion scene training image; performing dark light enhancement processing on the original training image based on the real human semantic information to obtain a night scene training image; wherein the motion scene training image and the night scene training image constitute the training image.
[0015] In this way, considering that there are few real scene images captured in complex scenes such as night scenes and sports in the sample library, we can obtain sports scene training images by performing motion blur enhancement processing on the images collected in ordinary scenes, and obtain night scene training images by performing dark light enhancement processing on the images collected in ordinary scenes.
[0016] In one implementable manner, motion blur enhancement processing is performed on the portrait in the original training image based on the real human semantic information to obtain a motion scene training image, including: determining the parts of the human body in the training image based on the real human semantic information; determining blur parameters of different parts of the human body, the blur parameters including motion direction and motion speed; wherein, a human body part with a larger motion amplitude has a greater motion speed; based on the blur parameters, different degrees of blur processing is performed on different parts of the human body in the training image to obtain the motion scene training image.
[0017] In real sports scenes, different parts of the human body often have different movement amplitudes. For example, in a waving motion scene, the movement amplitude of the human body's limbs is significantly greater than that of the trunk. Therefore, in order to make the obtained motion blurred image more consistent with the actual sports scene, the application can perform different degrees of blurring on different parts of the human body based on the movement amplitude of different parts of the human body to obtain a sports scene training image.
[0018] In one implementable manner, the training data also includes real human body semantic information corresponding to the training image; the method also includes: using the human body parsing subnetwork to perform human body parsing processing on the training image to obtain predicted human body semantic information; based on the predicted human body semantic information and the real human body semantic information, training the human body parsing subnetwork to obtain a trained human body parsing subnetwork.
[0019] When training the depth estimation network, the present application can use a pre-trained human parsing sub-network to further train the depth estimation network, or can use a non-pre-trained human parsing sub-network to train synchronously with the monocular depth estimation sub-network.
[0020] In one implementable manner, the human body parsing subnetwork is used to perform feature extraction on the training image to obtain a first intermediate feature containing human body semantic information, including: using the trained human body parsing subnetwork to perform human body parsing processing on the training image to obtain the first intermediate feature.
[0021] In this way, the first intermediate feature extracted based on the trained human body analysis subnetwork is more accurate and can better guide the monocular depth estimation subnetwork to learn the subject information in the training image and the correlation between various parts of the human body.
[0022] In a second aspect, the present application also provides a depth estimation method, which is applied to an electronic device, wherein the electronic device includes a depth estimation network, and the depth estimation network includes a human body parsing subnetwork and a monocular depth estimation subnetwork; the method includes: obtaining an image to be processed; using the human body parsing subnetwork to perform feature extraction on the image to be processed to obtain a first intermediate feature containing human body semantic information; using the monocular depth estimation subnetwork to perform feature extraction on the image to be processed to obtain a second intermediate feature containing depth information; using the monocular depth estimation subnetwork to perform feature fusion on the first intermediate feature and the second intermediate feature to obtain a fused feature; using the monocular depth estimation subnetwork to perform depth estimation on the fused feature pair to obtain depth information corresponding to the image to be processed.
[0023] In one achievable manner, the method further includes: based on the depth information, blurring the background of the image to be processed to obtain a blurred portrait image.
[0024] In this way, the trained depth estimation network can be applied to the portrait blur shooting function, so that the portrait details and continuity of the portrait in the portrait blur photos or videos are better, and the probability of incorrect blurring of the portrait is reduced. Especially for complex portrait blur scenes such as night scenes and sports, the portrait blur effect obtained by using the depth estimation network trained in this application is particularly outstanding.
[0025] In a third aspect, the present application also provides an electronic device, comprising a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, the computer program code comprises computer instructions, and when the processor executes the computer instructions, the electronic device executes a method as described in any one of the first aspect or the second aspect.
[0026] In a fourth aspect, the present application also provides a chip system, which includes a processor; the processor is coupled to a memory, the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes a method as described in any one of the first aspect or the second aspect.
[0027] In a fifth aspect, the present application also provides a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction is executed on a computer, the computer executes the method as described in any one of the first aspect or the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 A flowchart of a depth estimation network training method provided in an embodiment of the present application;
[0029] Figure 2 An algorithm architecture diagram of a depth estimation network provided in an embodiment of the present application;
[0030] Figure 3 A network architecture diagram of a depth estimation network provided in an embodiment of the present application;
[0031] Figure 4 An example diagram of a depth estimation network training method provided in an embodiment of the present application;
[0032] Figure 5 A flowchart of a depth estimation method provided in an embodiment of the present application;
[0033] Figure 6 An algorithm architecture diagram of a depth estimation method provided in an embodiment of the present application;
[0034] Figure 7 A structural block diagram of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0035] To facilitate the understanding of the technical solution of the application, the application scenario of the application is first described below.
[0036] Portrait blur processing refers to obtaining the depth relationship between the portrait and the background in the image based on the depth information of the image. Then, based on the depth relationship between the portrait and the background, the portrait is focused and the background is blurred. This portrait blur effect can highlight the human body and enhance the artistic conception.
[0037] Among them, the key to achieving the portrait blur effect is to correctly estimate the depth information of the image. Commonly used depth estimation methods mainly include binocular depth estimation methods and monocular depth estimation methods.
[0038] Binocular depth estimation method: It means taking two images with a binocular camera. Then, through the stereo matching algorithm, find the corresponding points of the same scene point in the two images and calculate their disparity. After that, according to the camera parameters (such as focal length, baseline distance, etc.) and disparity, the depth of the point is calculated through the principle of triangulation. Applying the above process to all points in the image, the depth information of the entire scene can be obtained.
[0039] Monocular depth estimation method: refers to using a single camera to predict the depth distribution of the entire scene image through a deep learning network without the need for special hardware support.
[0040] It can be seen that the binocular depth estimation method requires the support of the stereo matching algorithm. However, in some complex photo scenes, stereo matching is prone to errors, resulting in inaccurate estimated depth information, thus affecting the portrait blur effect. For example, in night scene portrait blur photography scenes, stereo matching is prone to errors due to the poor lighting conditions in the night scene environment, insufficient texture in the target area, and the easy appearance of noise in the image. For another example, in the motion portrait blur photography scene, stereo matching is prone to errors due to the motion blur of the human body in the motion scene and the deformation of the image under different viewing angles.
[0041] Since the monocular depth estimation method does not involve the stereo matching process, it is better than the binocular depth estimation method for complex portrait blurring scenes such as night scenes and sports. However, due to the complex depth changes in complex scenes such as night scenes and sports, and the limited content of a single image, the current monocular depth estimation network has a weak generalization ability for these complex scenes. Therefore, the portrait blurring photos obtained based on the monocular depth estimation method need to be further improved in terms of portrait details and continuity.
[0042] For example, in a night scene portrait blur photography scene, since the color of the human hair and the background are both black, the monocular depth estimation network cannot accurately distinguish between the hair and the background. Therefore, the hair in the portrait may be blurred incorrectly.
[0043] In order to solve the above technical problems, an embodiment of the present application provides a depth estimation network training method, which uses two branches, a portrait parsing subnetwork and a monocular depth estimation subnetwork, to train the image. When training the monocular depth estimation subnetwork, the human body semantic information features extracted by the portrait parsing subnetwork are fused with the depth features extracted by the monocular depth estimation subnetwork to train the monocular depth estimation subnetwork. In this way, the human body semantic information features can help the monocular depth estimation subnetwork better learn the subject information in the image and the correlation between various parts of the human body, so that the depth estimation result is consistent with the human body parsing result, thereby improving the accuracy of the subject depth and the depth details of various parts of the human body. In this way, the depth estimation network trained in this application acts on the portrait blur shooting scene, which can improve the portrait details and continuity in the image. Especially in complex portrait blur shooting scenes such as night scenes and sports, the situation of false false and leaked false can be significantly reduced.
[0044] The depth estimation network training method provided in the embodiment of the present application can be implemented by deploying a neural network model and computer program code in software form in a hardware computing environment. Available hardware computing environments include: personal computers, servers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, programmable consumer electronic devices, cloud servers, server instances, supercomputers, etc.
[0045] The depth estimation network training method provided in the embodiment of the present application is described in detail below with reference to the accompanying drawings.
[0046] Figure 1 A flowchart of a depth estimation network training method provided in an embodiment of the present application, Figure 2 The following is a schematic diagram of the algorithm structure of a depth estimation network provided in an embodiment of the present application. Figure 1 and 2 As shown, the depth estimation network training method provided in the embodiment of the present application may include the following steps:
[0047] S101, obtaining training data.
[0048] In the embodiment of the present application, the training data may include a training image and real depth information and real human semantic information corresponding to the training image. In other words, the training data includes paired training images, depth information, and human semantic information.
[0049] The real depth information may include the depth value corresponding to each pixel in the training image. The depth value refers to the distance information from each point in the scene in the image to the camera.
[0050] The real human semantic information may include the semantic labels corresponding to each pixel of the human figure in the training image, for example, the semantic labels may include the semantic labels of the head, body, hand, leg, top, etc.
[0051] In some embodiments, the real depth information may be a depth map, and the real human semantic information may be a human semantic segmentation map. The depth map is presented as a grayscale gradient change from white to black to represent the depth difference of each pixel in the image. The human semantic segmentation map may be a color map, in which different colors may be used to represent different semantic labels. For example, in the human semantic segmentation map, blue represents the face, red represents the top, green represents the pants, yellow represents the hands, and so on. In some examples, each semantic label may also be annotated in the human semantic segmentation map.
[0052] It should be noted that, in some embodiments, the training images may be images taken in real scenes. For example, the training images may include images of real scenes taken in complex scenes such as night scenes and sports. Among them, the training images are images that have not been processed with portrait blur.
[0053] Considering that there are few real scene images captured in complex scenes such as night scenes and sports in the sample library, the embodiment of the present application also provides a method for simulating real scene images captured in complex scenes such as night scenes and sports based on images captured in ordinary scenes.
[0054] The images captured in ordinary scenes may be sufficient images including portraits in the sample library. For example, the images captured in ordinary scenes may be open source and directly acquired images. The images may be images captured in static and well-lit scenes.
[0055] In one possible implementation, the simulation of night scene images can be implemented in the following manner: obtaining an original image corresponding to a normal scene, and then performing dark light enhancement processing on the original image. In this way, the original image after dark light enhancement processing becomes darker in overall color, similar to a night scene image.
[0056] In one possible implementation, the simulated motion scene image can be implemented in the following manner: obtaining an original image corresponding to a normal scene and real human semantic information corresponding to the original image, and then performing motion blur processing on the original image based on the real human semantic information to obtain a motion blurred image.
[0057] In some embodiments, the portrait in the original image can be blurred to the same degree based on the semantic information of the real human body. However, in real motion scenes, different parts of the human body often have different motion amplitudes. For example, in a waving motion scene, the motion amplitude of the human body's limbs is significantly greater than that of the torso. Therefore, this blurring method does not fit the actual motion scene.
[0058] In order to make the obtained motion blurred image more closely fit the actual motion scene, the embodiment of the present application can perform different degrees of blurring on different parts of the human body based on the motion amplitude of different parts of the human body to obtain a motion scene training image.
[0059] Exemplarily, the embodiments of the present application can first determine the parts of the human body in the image based on the semantic information of the real human body. Then, the movement direction and movement speed parameters of different parts of the human body can be set based on the prior knowledge of body structure biology. For example, the movement amplitude of the hands and feet is larger, and the head and body remain relatively still. Afterwards, the original image can be motion blurred based on the above-set parameters to obtain a motion scene training image. The scene training image obtained in this way is more realistic and natural.
[0060] In this way, the embodiment of the present application can first obtain the original image corresponding to the ordinary scene, and then obtain the training image by performing dark light enhancement processing and / or motion blur enhancement processing on the original image. The training image may include the image after dark light enhancement processing and / or the image after motion blur enhancement processing.
[0061] It should be noted that after the original image is processed by dark light enhancement and / or motion blur enhancement, the depth information and human semantic information of the image will not change. In other words, the depth information and human semantic information of the training image and the original image are the same.
[0062] S102, input the training image into the human parsing subnetwork, and output the predicted human semantic information.
[0063] The human body analysis sub-network is a deep learning network model for semantic analysis of human body images. The embodiment of the present application does not limit the specific network architecture of the human body analysis sub-network.
[0064] Exemplarily, the specific network architecture of the human body parsing subnetwork can be Attention-to-Scale, JPPNet, UNet and other network architectures. The Attention-to-Scale network architecture is a human body parsing network architecture that uses pyramid pooling to fuse features of different scales. The JPPNet network architecture is a human body parsing network architecture that jointly utilizes global human body information and local part information. The UNet network architecture is an encoder-decoder network architecture.
[0065] In some embodiments, the predicted human semantic information output by the human body parsing subnetwork may include one or more of a human body semantic segmentation map and a semantic label.
[0066] S103, determining a first loss based on the predicted human body semantic information and the actual human body semantic information.
[0067] In some embodiments, the first loss may be a cross entropy loss. By minimizing the cross entropy loss, the predicted human body semantic information is made as close as possible to the real human body semantic information, so as to achieve the parsing effect of various parts of the human body.
[0068] Exemplarily, step S102 may output a predicted human semantic label for each pixel in the training image. Then, the loss corresponding to each pixel may be determined based on the predicted human semantic label and the true human semantic label. Finally, the losses of all pixels may be summed or averaged to obtain a first loss.
[0069] S104: Based on the first loss, train the human body parsing sub-network.
[0070] The training process of the human body parsing sub-network is an iterative process. Through multiple iterations, the network parameters inside the human body parsing sub-network are continuously optimized and updated to make the network loss converge continuously. When the first loss is less than the preset threshold, the training ends and a trained human body parsing sub-network is obtained.
[0071] It should be noted that the embodiment of the present application is only an example of training the human body parsing sub-network based on the first loss, and does not limit the training method of the human body parsing sub-network. For example, when the number of training times of the human body parsing sub-network training reaches a preset number, the training can be terminated to obtain a trained human body parsing sub-network.
[0072] The following uses the UNet network architecture of the human body parsing sub-network as an example to illustrate the training process of the human body parsing sub-network.
[0073] like Figure 3 As shown in the figure, the UNet network architecture includes an encoder and a decoder. The encoder can use convolution blocks to gradually downsample the input training image to obtain multiple layers of feature maps containing human semantic information. The decoder can gradually upsample the spatial dimensions based on the encoder to obtain detailed information. After that, the feature maps containing human semantic information in each layer of the encoder can be symmetrically skipped and directly spliced to the decoder for fusion. Then, in the last upsampling layer, a 1x1 convolution can be used to map to the target number of categories. Finally, the Softmax activation function is used to predict the category probability of each pixel (that is, the probability of each part of the human body).
[0074] Furthermore, the human body parsing subnetwork can be optimized through cross entropy loss until the human body parsing subnetwork converges. In this way, the UNet network architecture can fully combine features at different scales through encoding-decoding, retaining the overall picture without losing details, and better achieve human body parsing.
[0075] S105, using the human body analysis sub-network to perform feature extraction on the input training image to obtain a first intermediate feature.
[0076] Among them, the first intermediate feature can be a plurality of feature maps containing human semantic information extracted by the encoder of the human body parsing subnetwork.
[0077] For example, Figure 4 As shown, the encoder of the human body parsing subnetwork extracts features from the input training image and obtains four layers of feature maps containing human body semantic information, namely feature map M1, feature map M2, feature map M3 and feature map M4.
[0078] S106, input the training image and the first intermediate feature into a monocular depth estimation subnetwork, and output predicted depth information.
[0079] In the embodiment of the present application, when training the monocular depth estimation subnetwork, the monocular depth estimation subnetwork can perform feature extraction on the training image to obtain a second intermediate feature containing depth information. Then, the monocular depth estimation subnetwork is used to perform feature fusion on the first intermediate feature and the second intermediate feature to obtain a fused feature. Furthermore, the monocular depth estimation subnetwork can predict the depth information of the training image based on the fused feature. In this way, the first intermediate feature containing human semantic information can help the monocular depth estimation subnetwork better learn the subject information in the training image and the correlation between various parts of the human body, so as to distinguish various parts of the human body from the background in the image, so that the depth estimation result is consistent with the human body analysis result, thereby improving the accuracy of the subject depth and the depth details of various parts of the human body.
[0080] The monocular depth estimation subnetwork is a network model that predicts the depth information of an input image. The embodiment of the present application does not limit the specific network architecture of the monocular depth estimation subnetwork. For example, the monocular depth estimation subnetwork can adopt an encoder-decoder network architecture. Among them, the specific network architecture of the monocular depth estimation subnetwork and the human body analysis subnetwork can be the same or different.
[0081] The following takes the monocular depth estimation subnetwork and the human body parsing subnetwork using the same network architecture as an example to illustrate the training method of the monocular depth estimation subnetwork. For example, both the monocular depth estimation subnetwork and the human body parsing subnetwork use the encoder-decoder UNet network architecture.
[0082] Exemplary, combined Figure 3 and Figure 4As shown in FIG. 1 , the original training image is subjected to motion blur enhancement processing to obtain a motion blurred training image. Afterwards, the motion blurred training image is input into the human body parsing subnetwork and the monocular depth estimation subnetwork respectively.
[0083] The encoder of the human body parsing subnetwork extracts features from the input training image and obtains four layers of feature maps containing human body semantic information, namely feature map M1, feature map M2, feature map M3 and feature map M4. Correspondingly, the encoder of the monocular depth estimation subnetwork extracts features from the input training image and obtains four layers of feature maps containing depth information, namely feature map N1, feature map N2, feature map N3 and feature map N4.
[0084] After the encoder of the human body parsing subnetwork obtains four layers of feature maps containing human body semantic information, on the one hand, the decoder of the human body parsing subnetwork can perform deconvolution and other calculations based on the feature map M4 to output a predicted human body semantic segmentation map.
[0085] For example, Figure 4 As shown, the output predicted human semantic segmentation map only includes the portrait area, and different parts of the portrait can be represented by different colors.
[0086] It should be noted that Figure 4 The human body semantic segmentation map is an image after grayscale processing, and the actual output human body semantic segmentation map can be a color image, where different colors can represent different parts of the human body.
[0087] On the other hand, the decoder of the monocular depth estimation subnetwork fuses the feature map N4 with the feature map M1 to obtain the fused feature map K1. Furthermore, the fused feature map K1 is fused with the feature map M2 to obtain the fused feature map K2. The fused feature map K2 is then fused with the feature map M3 to obtain the fused feature map K3. Finally, the fused feature map K3 is fused with the feature map M4 to obtain the fused feature map K4. Subsequently, the predicted depth map can be output based on the feature map K4. Among them, the fused feature map K1, the fused feature map K2, the fused feature map K3, and the fused feature map K3 all contain the human features and depth features of the training image.
[0088] For example, Figure 4 As shown, the output predicted depth map includes all the information in the training image, and the depth value corresponding to each area can be represented by different grayscales.
[0089] It should be understood that the above-mentioned feature maps M1, M2, M3 and M4 are first intermediate features, feature maps N1, N2, N3 and N4 are second intermediate features, and fused feature maps K1, K2, K3 and K4 are fused features.
[0090] In some embodiments, in order to better fuse the first intermediate feature with the second intermediate feature, the first intermediate feature may be subjected to a convolution process before being fused with the second intermediate feature. In this way, after the feature map M1, the feature map M2, the feature map M3 and the feature map M4 are subjected to convolution process, feature maps M1', M2', M3' and M4' are obtained. Then, the decoder of the monocular depth estimation subnetwork may fuse the feature map N4 with the feature map M1' to obtain a fused feature map K1. Furthermore, the fused feature map K1 is fused with the feature map M2' to obtain a fused feature map K2. The fused feature map K2 is then fused with the feature map M3' to obtain a fused feature map K3. Finally, the fused feature map K3 is fused with the feature map M4' to obtain a feature map K4. In this way, the first intermediate feature after convolution process can be better fused with the second intermediate feature, thereby improving the training efficiency.
[0091] It should be noted that the embodiments of the present application do not limit the specific implementation method of extracting features from training images using an encoder. For example, it can be implemented by convolution calculation, downsampling, etc. For details, please refer to the relevant technology, which will not be repeated here. The embodiments of the present application do not limit the specific implementation method of feature fusion using a decoder. For example, the fused feature can be obtained by performing pixel-level addition operations, deconvolution calculations, etc. on the first intermediate feature and the second intermediate feature.
[0092] In the embodiment of the present application, the human body analysis subnetwork and the monocular depth estimation subnetwork can be trained synchronously or step by step.
[0093] If the human body parsing subnetwork and the monocular depth estimation subnetwork are trained synchronously, the first intermediate feature can be extracted during each iterative training of the human body parsing subnetwork. Then, the extracted first intermediate feature is input into the monocular depth estimation subnetwork. After that, the monocular depth estimation subnetwork fuses the first intermediate feature with the second intermediate feature. That is to say, if the human body parsing subnetwork and the monocular depth estimation subnetwork are trained synchronously, the first intermediate feature input into the monocular depth estimation subnetwork each time is based on the feature extracted by the untrained human body parsing subnetwork.
[0094] If the human body parsing subnetwork and the monocular depth estimation subnetwork are trained in steps, that is, the human body parsing subnetwork is trained first, and after the human body parsing subnetwork is trained to convergence, the monocular depth estimation subnetwork is trained. In this way, when training the monocular depth estimation subnetwork, the trained human body parsing subnetwork can be used to extract the first intermediate feature of the training image. Then, the extracted first intermediate feature is input into the monocular depth estimation subnetwork. After that, the monocular depth estimation subnetwork fuses the first intermediate feature with the second intermediate feature. In other words, if the human body parsing subnetwork and the monocular depth estimation subnetwork are trained in steps, the first intermediate feature input into the monocular depth estimation subnetwork each time is based on the feature extracted by the trained human body parsing subnetwork. In this way, the first intermediate feature extracted based on the trained human body parsing subnetwork is more accurate, and can better guide the monocular depth estimation subnetwork to learn the subject information in the training image and the correlation between various parts of the human body.
[0095] S107: Determine a second loss (also referred to as depth loss) based on the predicted depth information and the actual depth information.
[0096] In some embodiments, the second loss may traverse each training sample (training data), calculate the cross entropy loss of the predicted probability of the training sample for each depth category and the true depth label, and then average them.
[0097] Exemplarily, the second loss can be calculated using the following formula (1):
[0098]
[0099] Wherein, L2 represents the second loss, N represents the total number of training samples, i represents any training sample among all training samples, M represents the total number of depth categories, and both N and M are positive integers greater than or equal to 1. ic represents the true label of the i-th training sample for the c-th class, p ic Represents the predicted probability that the i-th sample belongs to the c-th class. ic log(p ic ) represents the cross entropy of each class.
[0100] In some embodiments, in order to make the portrait area in the output depth map appear clearer, more coherent, and with smaller errors. Different loss weights can be set for different areas of the training image. For example, an embodiment of the present application can set the loss weight of the portrait area in the training image to be greater than the loss weight of the background area. In this way, when calculating the second loss, the second loss can be calculated based on the loss weights of different areas, the predicted depth information, and the actual depth information. In this way, when backpropagation is used to update the network parameters, since the loss gradient of the portrait area is larger, the network update will pay more attention to improving the quality of the depth map of the portrait area, thereby making the portrait area in the output depth map appear clearer, more coherent, and with smaller errors.
[0101] In some embodiments, considering that in motion scenes, areas with large movement such as limbs are prone to depth estimation errors, and in night scenes, dark areas such as hair and dark clothes are prone to depth estimation errors. Therefore, the loss weights of key parts such as hair, limbs, and dark clothes can be increased when calculating the second loss. In this way, when backpropagating to update network parameters, since the loss gradient of key parts is larger, the network update will pay more attention to improving the quality of the depth map of key parts. In this way, the monocular depth subnetwork can generate more accurate depth results in the areas where key parts such as hair, limbs, and dark clothes are located, making the human body area in the output depth map appear clearer and more coherent, with smaller errors.
[0102] In one achievable manner, the second loss can be calculated in the following manner: the loss weights of each region in the training image can be preconfigured. Among them, the loss weights of key parts such as hair, limbs, dark clothes, etc. are higher than the loss weights of other regions in the image. Then, based on the loss weights of each region in the training image, the predicted depth information, and the true depth information, the second loss is calculated. In this way, the quality of dark areas such as hair and dark clothes in the output depth map for night scenes can be improved, and the quality of areas with large movement such as limbs in the output depth map for motion scenes can be improved.
[0103] In one implementable manner, the second loss can also be calculated in the following manner: first, based on the training image and the human semantic information corresponding to the training image, determine the brightness of each part of the human body in the image. Brightness can represent the brightness degree of color from white to black. The higher the brightness, the lighter the color, and the lower the brightness, the darker the color. Then, the area where the brightness is less than the brightness threshold is determined as the dark area, and the area where the brightness is greater than or equal to the brightness threshold is determined as the light area. Among them, the dark area is the key part. In this way, the loss weight of the dark area can be set greater than the loss weight of the light area. Finally, the determined loss weight is introduced into the calculation of the second loss to obtain the second loss. In this way, the quality of dark areas such as hair and dark clothes in the output depth map for night scenes can be improved.
[0104] Exemplarily, the second loss can be calculated using the following formula (2), where formula (2) is:
[0105]
[0106] Among them, w c Represents the loss weight corresponding to category c.
[0107] For example, c represents the loss weight corresponding to the depth region. For another example, w c Represents the loss weights corresponding to key parts such as hair, limbs, dark clothes, etc.
[0108] S108, training a monocular depth estimation subnetwork based on the second loss.
[0109] The training process of the monocular depth estimation subnetwork is an iterative process. Through multiple iterations, the network parameters inside the monocular depth estimation subnetwork are continuously optimized and updated to make the network loss continue to converge. When the second loss is less than the preset threshold, the training ends and a trained monocular depth estimation subnetwork is obtained.
[0110] It should be noted that the embodiment of the present application is only illustrative of the training of the monocular depth estimation subnetwork based on the second loss, and does not limit the training method of the monocular depth estimation subnetwork. For example, when the number of training times of the monocular depth estimation subnetwork training reaches a preset number, the training can be terminated to obtain a trained monocular depth estimation subnetwork.
[0111] In summary, the depth estimation network provided in the embodiment of the present application includes two network branches, namely a human body analysis subnetwork and a monocular depth estimation subnetwork. In the training stage of the depth estimation network, the first intermediate feature containing human body semantic information can be extracted through the human body analysis subnetwork, and the first intermediate feature can be fused with the second intermediate feature containing depth information of the monocular depth estimation subnetwork. In this way, by training the monocular depth estimation subnetwork with fusion features containing human body features and depth features, the monocular depth estimation subnetwork can better learn the subject information in the image and the correlation between various parts of the human body, so that the depth estimation result is consistent with the human body analysis result, thereby improving the accuracy of the subject depth and the depth details of various parts of the human body.
[0112] The depth estimation network after training in the embodiment of the present application is a lightweight structure and can be deployed in mobile electronic devices such as mobile phones and tablet computers. In this way, the electronic device deployed with the above-mentioned depth estimation network can use the trained depth estimation network for image processing. For example, the trained depth estimation network can be applied to the portrait blur camera function, so that the details and continuity of the portraits in the obtained portrait blur photos or videos are better, and the probability of incorrect blurring of the portraits is reduced. Especially for complex portrait blur scenes such as night scenes and sports, the portrait blur effect obtained by applying the depth estimation network trained in the embodiment of the present application is particularly prominent.
[0113] The following describes a method for blurring a portrait based on a depth estimation network trained according to an embodiment of the present application.
[0114] Figure 5 The following is a schematic diagram of the workflow of a method for blurring a portrait provided in an embodiment of the present application. Figure 5 As shown, the method may include the following steps:
[0115] S201, obtaining an image to be processed.
[0116] The image to be processed may be an original, clear image acquired by an image acquirer of the electronic device.
[0117] S202, using the human body parsing sub-network to perform feature extraction on the image to be processed, to obtain a first intermediate feature containing human body semantic information.
[0118] S203, using a monocular depth estimation subnetwork to perform feature extraction on the image to be processed to obtain a second intermediate feature containing depth information.
[0119] S204, using a monocular depth estimation sub-network to perform feature fusion on the first intermediate feature and the second intermediate feature to obtain a fused feature.
[0120] S205, using a monocular depth estimation subnetwork to perform depth estimation on the fused feature pair to obtain depth information corresponding to the image to be processed.
[0121] Among them, steps S202 to S205 can refer to the description of the above training process and will not be repeated here.
[0122] S206, based on the depth information, filtering the background of the image to be processed to obtain a blurred portrait image.
[0123] In one possible implementation, filtering the background of the image to be processed to obtain a portrait blurred image can be implemented in the following manner: first, based on the depth information, determine the distance between the portrait and other objects other than the portrait in the image to be processed. The distance information between the portrait and other objects other than the portrait in the image to be processed can be represented in the form of a sigma graph. Then, based on the sigma graph, filtering is performed to different degrees on other objects other than the portrait to obtain a portrait blurred image.
[0124] Through filtering, objects other than the portrait can achieve different degrees of blurring effects, wherein the closer the object is to the portrait, the smaller the blurring degree is, and the farther the object is from the portrait, the larger the blurring degree is.
[0125] The embodiments of the present application do not limit the specific filtering processing method. For example, filtering processing methods such as circular filtering and infinite impulse response filtering (IIR, Infinite Impulse Response) can be used.
[0126] The various method embodiments described in this document may be independent solutions or may be combined according to internal logic, and these solutions all fall within the protection scope of this application.
[0127] It can be understood that, in the above-mentioned various method embodiments, the methods and operations implemented by the electronic device can also be implemented by components (such as chips or circuits) that can be used in the electronic device.
[0128] The above embodiments introduce the depth estimation network training method and the depth estimation method provided by the present application. It is understandable that, in order to realize the above functions, the electronic device includes a hardware structure and / or software module corresponding to each function. It should be easily appreciated by those skilled in the art that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0129] The embodiment of the present application further provides a processing device, which includes at least one processor and a communication interface. The communication interface is used to provide information input and / or output to the at least one processor, and the at least one processor is used to execute the method in the above method embodiment.
[0130] It should be understood that the above-mentioned processing device can be a chip. Figure 7 , Figure 7A structural block diagram of a chip provided in an embodiment of the present application. Figure 7 The chip shown may be a general-purpose processor or a dedicated processor. The chip 300 may include at least one processor 301. The at least one processor 301 may be used to support execution Figures 1 to 6 The technical solution shown in any one of the embodiments.
[0131] Optionally, the chip 300 may further include a transceiver 302, which is used to accept the control of the processor 301 and to support the execution Figures 1 to 6 The technical solution shown in any one of the embodiments. Optionally, Figure 7 The chip 300 shown may further include a storage medium 303. Specifically, the transceiver 302 may be replaced by a communication interface, and the communication interface provides information input and / or output for the at least one processor 301.
[0132] It should be noted that Figure 7 The chip 300 shown can be implemented using the following circuits or devices: one or more field programmable gate arrays (FPGA), programmable logic devices (PLD), application specific integrated circuits (ASIC), system on chip (SoC), central processor unit (CPU), network processor (NP), digital signal processor (DSP), micro controller unit (MCU), controller, state machine, gate logic, discrete hardware components, any other suitable circuits, or any combination of circuits capable of performing the various functions described throughout the present application.
[0133] In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in a processor or an instruction in the form of software. The steps of the method disclosed in conjunction with the embodiment of the present application can be directly embodied as a hardware processor for execution, or a combination of hardware and software modules in a processor for execution. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it is not described in detail here.
[0134] It should be noted that the processor in the embodiment of the present application can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiment can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to perform, or the hardware and software modules in the decoding processor can be combined and performed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0135] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0136] According to the method provided in the embodiments of the present application, the embodiments of the present application also provide a computer program product, which includes: a computer program or instructions, when the computer program or instructions are run on a computer, the computer executes the method of any one of the embodiments of the method.
[0137] According to the method provided by the embodiments of the present application, the embodiments of the present application also provide a computer storage medium, which stores a computer program or instructions. When the computer program or instructions are run on a computer, the computer executes the method of any one of the embodiments of the method.
[0138] According to the method provided by an embodiment of the present application, an embodiment of the present application also provides an electronic device, including a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the method of any one of the embodiments of the method embodiment.
[0139] According to the method provided in the embodiment of the present application, the embodiment of the present application also provides a chip system, which includes a processor, the processor is coupled to a memory, and is used to execute a computer program or instruction stored in the memory. When the computer program or instruction is executed, the chip system can implement all or part of the steps in the method embodiment. The chip system can be composed of a chip, or it can include a chip and other discrete devices.
[0140] Those of ordinary skill in the art will appreciate that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0141] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage media include: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.
[0142] The computer storage medium, computer program product, and electronic device provided in the above-mentioned embodiments of the present application are all used to execute the method provided above. Therefore, the beneficial effects that can be achieved can refer to the corresponding beneficial effects of the method provided above, and will not be repeated here.
[0143] It should be understood that in each embodiment of the present application, the execution order of each step should be determined by its function and internal logic, and the size of the sequence number of each step does not mean the order of execution and does not limit the implementation process of the embodiment.
[0144] Each part of this specification is described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, computer storage medium, computer program product, and electronic device, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiments.
[0145] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0146] The above-described embodiments of the present application do not constitute a limitation on the protection scope of the present application.
Claims
1. A depth estimation network training method, characterized in that: The depth estimation network includes a human body analysis subnetwork and a monocular depth estimation subnetwork, and the method includes: Acquire training data, the training data comprising a training image and real depth information corresponding to the training image; Using the human body parsing subnetwork to perform feature extraction on the training image to obtain a first intermediate feature containing human body semantic information; Using the monocular depth estimation subnetwork to perform feature extraction on the training image to obtain a second intermediate feature containing depth information; Using the monocular depth estimation subnetwork to perform feature fusion on the first intermediate feature and the second intermediate feature to obtain a fused feature; Using the monocular depth estimation subnetwork to perform depth estimation on the fused feature pair to obtain predicted depth information; Based on the predicted depth information and the real depth information, the monocular depth estimation subnetwork is trained to obtain a trained monocular depth estimation subnetwork; The monocular depth estimation subnetwork is trained based on the predicted depth information and the real depth information to obtain a trained monocular depth estimation subnetwork, including: Determining loss weights of different regions in the training image; wherein the loss weight of the portrait region in the training image is greater than the loss weight of the background region; Calculating a depth loss based on the loss weights of the different regions, the predicted depth information, and the actual depth information; When the depth loss is less than the depth loss threshold, the training of the monocular depth estimation subnetwork is terminated to obtain a trained monocular depth estimation network.
2. The method according to claim 1, characterized in that The determining the loss weights of different regions in the training image comprises: Determining key parts of the portrait in the training image based on the movement amplitude of each part of the human body and / or the color of each part of the human body in the training image; It is determined that the loss weight of the key part is greater than the loss weight of other parts of the portrait.
3. The method according to claim 2, characterized in that Determining key parts of the portrait in the training image based on the movement amplitude of each part of the human body and / or the color of each part of the human body in the training image includes: Determining the brightness of each part of the human body in the training image; Determine the human body parts whose brightness in the training image is less than the brightness threshold as key parts; and / or, Determining the movement amplitude of each part of the human body in the training image; The human body parts in the training image whose movement amplitude is greater than the movement amplitude threshold are determined as key parts.
4. The method according to claim 1, characterized in that The method further comprises: Acquire an original training image and real human semantic information corresponding to the original training image; Based on the real human body semantic information, a motion blur enhancement process is performed on the human figure in the original training image to obtain a motion scene training image; Performing dark light enhancement processing on the original training image based on the real human body semantic information to obtain a night scene training image; The motion scene training image and the night scene training image constitute the training image.
5. The method according to claim 4, characterized in that Performing motion blur enhancement processing on the human figure in the original training image based on the real human semantic information to obtain a motion scene training image includes: Based on the real human body semantic information, determining various parts of the human body in the training image; Determine fuzzy parameters of different parts of the human body, wherein the fuzzy parameters include movement direction and movement speed; wherein the movement speed of the human body part with a larger movement amplitude is greater; Based on the blur parameters, different parts of the human body in the training image are blurred to different degrees to obtain the motion scene training image.
6. The method according to any one of claims 1 to 5, characterized in that The training data also includes real human semantic information corresponding to the training image; the method also includes: Using the human body parsing subnetwork to perform human body parsing processing on the training image to obtain predicted human body semantic information; Based on the predicted human body semantic information and the real human body semantic information, the human body parsing sub-network is trained to obtain a trained human body parsing sub-network.
7. The method according to claim 6, characterized in that The human body parsing subnetwork is used to perform feature extraction on the training image to obtain a first intermediate feature containing human body semantic information, including: The trained human body parsing subnetwork is used to perform human body parsing processing on the training image to obtain a first intermediate feature.
8. A depth estimation method, characterized in that: Applied to an electronic device, the electronic device includes a depth estimation network, the depth estimation network includes a human body analysis subnetwork and a monocular depth estimation subnetwork; the monocular depth estimation subnetwork is a subnetwork obtained by determining the loss weights of different regions in a training image during a training process; and calculating the depth loss based on the loss weights of different regions, predicted depth information and real depth information; wherein the loss weight of the portrait region in the training image is greater than the loss weight of the background region; the method includes: Get the image to be processed; Using the human body parsing subnetwork to perform feature extraction on the image to be processed, to obtain a first intermediate feature containing human body semantic information; Using the monocular depth estimation subnetwork to perform feature extraction on the image to be processed to obtain a second intermediate feature containing depth information; Using the monocular depth estimation subnetwork to perform feature fusion on the first intermediate feature and the second intermediate feature to obtain a fused feature; The monocular depth estimation subnetwork is used to perform depth estimation on the fused feature pair to obtain depth information corresponding to the image to be processed.
9. The method according to claim 8, characterized in that The method further comprises: Based on the depth information, the background of the image to be processed is blurred to obtain a blurred portrait image.
10. An electronic device, characterized in that: It comprises a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, the computer program code comprises computer instructions, and when the processor executes the computer instructions, the electronic device executes the method as described in any one of claims 1-9.
11. A chip system, characterized in that: The chip system includes a processor; the processor is coupled to a memory, the memory is used to store computer program code, the computer program code includes computer instructions, and when the processor executes the computer instructions, the method as described in any one of claims 1-9 is executed.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program or instruction. When the computer program or instruction is executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Scene structure learning method and device and electronic equipment
CN109658418A
Depth estimation method and device, electronic equipment and computer readable storage medium
CN114359361A