Image processing method and device

By semantic segmentation and depth feature fusion of the initial RGB image, the target depth map is generated, which solves the problem of missing and noise in the depth map collected by the depth sensor in the indoor environment, improves the density and accuracy of the depth information, and enhances the execution ability of computer vision tasks.

CN114511778BActive Publication Date: 2025-05-06MIDEA GRP (SHANGHAI) CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210062375.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-19
Publication Date
2025-05-06
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

The depth maps collected by existing depth sensors in indoor environments have problems such as a large number of continuous depth information missing and depth data being noisy, which affects the execution of computer vision tasks.

Method used

An image processing method is proposed, which extracts semantic feature maps by semantic segmentation of the initial RGB image; combines the depth feature vector with the initial RGB image to generate a fusion confidence map and a fusion depth map; finally, based on the local and fusion confidence map depth map, a target depth map is generated to improve the density and accuracy of the depth information.

Benefits of technology

By extracting semantic information and depth features fusion, the global depth map information is optimized, which significantly improves the density and accuracy of the depth information of the target depth map, and enhances the execution ability of computer vision tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114511778B_ABST
    Figure CN114511778B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and provides an image processing method and device, the method comprising: performing semantic segmentation on an initial RGB image to obtain a semantic feature map; performing depth feature extraction on an initial depth image corresponding to the initial RGB image based on the foreground probability and semantic features included in the semantic feature map to obtain a depth feature vector, the depth feature vector is used to determine a local confidence map and a local depth map; performing feature fusion on the depth feature vector and the initial RGB image to obtain a fused confidence map and a fused depth map; and obtaining a target depth map based on the local confidence map, the local depth map, the fused confidence map, and the fused depth map. The method extracts the semantic information of the initial RGB image, guides the accurate depiction of the local depth information of the image, and optimizes the global depth map information at the same time, improves the density and accuracy of the depth information of the target depth map, and provides a guarantee for subsequent computer vision tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an image processing method and device. Background Art

[0002] Scene modeling technology has flourished due to the combination of depth maps collected by depth sensors and RGB images collected by color cameras, and is widely used in computer vision fields such as autonomous driving, robot recognition, map navigation, motion planning, and augmented reality.

[0003] However, due to the limitations of the devices themselves, the reliability of the depth maps collected by various depth sensors, including structured light-based sensors and lidar, is affected. In particular, in indoor environment scenes, there are often continuous and large amounts of partial depth information missing and noisy depth data, which in turn has adverse effects when performing computer vision tasks. Summary of the invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes an image processing method to fill in the depth map collected by the sensor and improve the density and accuracy of the depth information in the depth map.

[0005] The image processing method according to the first embodiment of the present invention includes:

[0006] Perform semantic segmentation on the initial RGB image to obtain a semantic feature map;

[0007] Based on the foreground probability and the semantic features included in the semantic feature map, a depth feature extraction is performed on the initial depth image corresponding to the initial RGB image to obtain a depth feature vector, wherein the depth feature vector is used to determine a local confidence map and a local depth map;

[0008] Performing feature fusion on the depth feature vector and the initial RGB image to obtain a fused confidence map and a fused depth map;

[0009] A target depth map is obtained based on the local confidence map, the local depth map, the fused confidence map and the fused depth map.

[0010] According to the image processing method of the embodiment of the present invention, by extracting the semantic information of the initial RGB image, the initial depth image is guided to accurately characterize the local depth information of the image. At the same time, the global depth map information is optimized by fusing the depth feature vector and the initial RGB image. By introducing the confidence map corresponding to the local depth map and the fused depth map, a reliable depth prediction area is retained, and the density and accuracy of the depth information of the target depth map are improved, thereby providing a guarantee for subsequent computer vision tasks.

[0011] According to an embodiment of the present invention, performing semantic segmentation on the initial RGB image to obtain a semantic feature map includes:

[0012] Inputting the initial RGB image into the local guidance module of the depth completion model for semantic segmentation, and obtaining the semantic feature map output by the local guidance module;

[0013] The step of extracting depth features from the initial depth image corresponding to the initial RGB image based on the foreground probability and the semantic features included in the semantic feature map to obtain a depth feature vector includes:

[0014] Inputting the semantic feature map and the initial depth image into the constraint network of the depth completion model to extract depth features, so as to obtain the depth feature vector;

[0015] The step of fusing the depth feature vector and the initial RGB image to obtain a fused confidence map and a fused depth map includes:

[0016] Inputting the initial RGB image and the depth feature vector into the adversarial generative network of the depth completion model to obtain the fused depth map and the fused confidence map output by the adversarial generative network;

[0017] The obtaining a target depth map based on the local confidence map, the local depth map, the fused confidence map and the fused depth map includes:

[0018] Inputting the local depth map, the local confidence map, the fused depth map and the fused confidence map into an output network of the depth completion model to obtain the target depth map output by the output network;

[0019] Wherein, the depth completion model is trained based on a sample training set.

[0020] According to an embodiment of the present invention, inputting the initial RGB image and the depth feature vector into the adversarial generative network of the depth completion model to obtain the fused depth map and the fused confidence map output by the adversarial generative network includes:

[0021] Inputting the depth feature vector and the initial RGB image into the generator of the generative adversarial network, and having the generator reconstruct the depth information of the initial RGB image based on the depth feature vector, to obtain the fused depth map and the fused confidence map output by the generator;

[0022] The generator is obtained by conducting adversarial training based on sample RGB images and sample dense depth maps in conjunction with the discriminator of the adversarial generative network.

[0023] According to an embodiment of the present invention, the step of inputting the semantic feature map and the initial depth image into the constraint network of the depth completion model to extract depth features to obtain the depth feature vector includes:

[0024] Inputting the semantic feature map and the initial depth image into the encoder of the constraint network for downsampling processing to obtain a first depth feature vector;

[0025] The first depth feature vector is input into the decoder of the constraint network for upsampling to obtain a second depth feature vector.

[0026] According to an embodiment of the present invention, inputting the initial RGB image and the depth feature vector into the adversarial generative network of the depth completion model to obtain the fused depth map and the fused confidence map output by the adversarial generative network includes:

[0027] Inputting the initial RGB image into the generative adversarial network for downsampling processing to obtain a first feature vector;

[0028] Performing feature fusion on the first feature vector and the first depth feature vector through the instance normalization module of the depth completion model to obtain a target fused feature vector;

[0029] Performing upsampling processing on the target fused feature vector to obtain a second feature vector;

[0030] The second feature vector and the second depth feature vector are subjected to feature fusion through the instance normalization module to obtain the fused depth map and the fused confidence map output by the generative adversarial network.

[0031] According to an embodiment of the present invention, the local depth map and the local confidence map are determined by the following steps:

[0032] The depth feature vector in the encoder of the constraint network is input into a symmetrical position in the decoder of the constraint network by means of a jump connection, so as to obtain the local depth map and the local confidence map output by the decoder.

[0033] According to an embodiment of the present invention, obtaining a target depth map based on the local confidence map, the local depth map, the fused confidence map and the fused depth map includes:

[0034] Based on the local depth map, the local confidence map, the fused depth map and the fused confidence map, obtaining a local attention weight and a fused attention weight;

[0035] The local depth map and the fused depth map are weighted based on the local attention weight and the fused attention weight to obtain the target depth map.

[0036] According to one embodiment of the present invention, the sample training set includes a plurality of sample RGB images and their corresponding simulated depth images;

[0037] The simulated depth image is determined by the following steps:

[0038] Performing a highlight area mask on the sample depth image corresponding to the sample RGB image to obtain a first simulated depth map;

[0039] Performing a target pixel area mask based on the sample depth image to obtain a second simulated depth map;

[0040] Perform noise point masking based on the sample depth image to obtain a third simulated depth map;

[0041] Performing random masking of semantic labels based on the sample depth image to obtain a fourth simulated depth map;

[0042] Perform semantic segmentation training based on all sample depth images in the sample training set, determine a mask matrix, and obtain a fifth simulated depth map;

[0043] The simulated depth image is determined based on the depth information of the sample depth image and at least one of the first to fifth simulated depth maps.

[0044] According to one embodiment of the present invention, the target loss function of the depth completion model training is determined based on the following steps:

[0045] Obtaining a first loss function of a generator and a second loss function of a discriminator of the generative adversarial network;

[0046] Based on the local depth map and the target depth map, obtaining a third loss function and a fourth loss function;

[0047] The target loss function is obtained based on the first loss function, the second loss function, the third loss function and the fourth loss function.

[0048] An image processing device according to an embodiment of a first aspect of the present invention includes:

[0049] The first processing module is used to perform semantic segmentation on the initial RGB image to obtain a semantic feature map;

[0050] A second processing module is used to extract depth features from an initial depth image corresponding to the initial RGB image based on the foreground probability and the semantic features included in the semantic feature map to obtain a depth feature vector, wherein the depth feature vector is used to determine a local confidence map and a local depth map;

[0051] A third processing module is used to perform feature fusion on the depth feature vector and the initial RGB image to obtain a fused confidence map and a fused depth map;

[0052] The fourth processing module is used to obtain a target depth map based on the local confidence map, the local depth map, the fused confidence map and the fused depth map.

[0053] According to an embodiment of the third aspect of the present invention, an electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of any one of the above-mentioned image processing methods are implemented.

[0054] According to the non-transitory computer-readable storage medium of the fourth aspect of the present invention, a computer program is stored thereon, and when the computer program is executed by a processor, the steps of any one of the above-mentioned image processing methods are implemented.

[0055] A computer program product according to an embodiment of the fifth aspect of the present invention comprises a computer program, which, when executed by a processor, implements the steps of any one of the above-mentioned image processing methods.

[0056] The above one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects:

[0057] By extracting the semantic information of the target RGB image, the target depth image is guided to accurately characterize the local depth information of the image. At the same time, the global depth map information is optimized by fusing the depth feature vector and the target RGB image. By introducing the confidence map corresponding to the local depth map and the fused depth map, the reliable depth prediction area is retained, and the density and accuracy of the depth information of the target dense depth map are improved, providing guarantee for subsequent computer vision tasks.

[0058] Furthermore, the generator of the generative adversarial network performs style transfer-like processing on the depth feature vector and the target RGB image, fusing the depth information represented by the depth feature vector with the information of the target RGB image, so that the generative adversarial network can output a dense fused depth map.

[0059] Furthermore, the attention weight is used to make the target dense depth map evenly retain the depth prediction areas corresponding to the local depth map and the fused depth map. Compared with the target depth image collected by the depth sensor, the depth information of the target dense depth map is denser and more accurate. The point cloud information converted from the target dense depth map contains more points than the target depth image and can better cover the shape of objects in the scene.

[0060] Furthermore, by generating different types of realistic depth maps through the five-mask method and fusing the depth information of different depth maps, we can generate realistic depth images with a more reasonable distribution of missing depth information, and use this to train a more robust depth completion model.

[0061] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0063] Figure 1 is a flowchart of an image processing method provided by an embodiment of the present invention;

[0064] Figure 2 is a schematic diagram of the structure of a depth completion model provided by an embodiment of the present invention;

[0065] Figure 3 is a schematic diagram of the structure of a constraint network provided by an embodiment of the present invention;

[0066] Figure 4 is an example diagram of a sample depth map provided by an embodiment of the present invention;

[0067] Figure 5 is an example diagram of a first simulated depth map provided by an embodiment of the present invention;

[0068] Figure 6 is an example diagram of a second simulated depth map provided by an embodiment of the present invention;

[0069] Figure 7 is an example diagram of a third simulated depth map provided by an embodiment of the present invention;

[0070] Figure 8 is a fourth simulated depth image example diagram provided by an embodiment of the present invention;

[0071] Fig. 9 is a fifth simulated depth image example diagram provided by an embodiment of the present invention;

[0072] Fig.10 is an example diagram of a simulated depth image provided by an embodiment of the present invention;

[0073] Fig.11 is a structural schematic diagram of an image processing device provided by an embodiment of the present invention;

[0074] Fig.12 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0075] The following embodiments of the present invention are described in further detail in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0076] In the description of the embodiments of the present invention, it should be noted that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the embodiments of the present invention and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limitations on the embodiments of the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance.

[0077] In the description of the embodiments of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "connected" and "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms in the embodiments of the present invention can be understood according to specific circumstances.

[0078] In the embodiments of the present invention, unless otherwise clearly specified and limited, the first feature being "above" or "below" the second feature may mean that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature being "above", "above" and "above" the second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. The first feature being "below", "below" and "below" the second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.

[0079] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiment of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0080] The depth map information collected by the depth sensor in indoor environment scenes often has continuous and large amounts of partial depth information missing and noisy depth data. The reliability of the depth map is low. Such raw depth maps will have adverse effects on computer vision tasks such as 3D target detection, three-dimensional reconstruction and indoor navigation.

[0081] Depth Completion refers to the process of completing and repairing the depth map information by combining the depth map and RGB image collected at the same time. Depth Completion is a common method to improve the reliability of the depth map and obtain accurate depth information.

[0082] At present, depth map completion methods are mainly based on neural networks and convolutional neural networks, which can be roughly divided into three categories: the first category is to directly use the encoder-decoder structure to complete the depth map, but the depth information generated by this method is seriously blurred, and the completed depth map is not effective; the second category optimizes depth completion by learning global point relationships and iterative propagation, which not only requires relatively rich global depth information as input, but also the completion process is time-consuming and inefficient; the third category modifies the neural network parameter structure by analyzing and modeling the physical characteristics of the depth map. This method has insufficient generalization ability and is not effective in handling the continuous missing of depth information in indoor scenes.

[0083] Combine the following Figures 1 to 10 The image processing method of the embodiment of the present invention is described. By constructing a depth completion model, the semantic and structural information of the RGB image is extracted and deeply fused, the global depth information is optimized, the accuracy of the depth information is improved, and the reliability of the depth map is improved, which is beneficial to the execution of computer vision tasks.

[0084] like Figure 1 As shown, the image processing method of the embodiment of the present invention includes steps 110 to 140, and the execution subject of the method can be a controller, or a cloud, or an edge server.

[0085] Step 110: perform semantic segmentation on the initial RGB image to obtain a semantic feature map.

[0086] RGB images refer to three-channel color images. Various colors are obtained by changing and superimposing the three color channels of red (R), green (G) and blue (B), so that RGB images can be prepared to reflect visual color characteristics. Among them, the initial RGB image is an RGB image obtained in a certain scene, and the scene may include multiple objects. The corresponding initial RGB image includes pixels corresponding to the multiple objects.

[0087] For example, the initial RGB image may be an RGB image of an indoor scene, and the indoor scene may include objects such as tables and chairs, household appliances, and furniture in the room.

[0088] In this step, the initial RGB image is processed by semantic segmentation, which performs dense prediction and inference labels for each pixel in the initial RGB image to achieve fine-grained reasoning, so that each pixel can be labeled as the category of the closed area.

[0089] In this embodiment, the initial RGB image is processed by semantic segmentation to obtain a corresponding semantic feature map, which includes relevant information of foreground probability and semantic features, that is, the semantic feature map is a feature map that marks the labels and positions corresponding to multiple objects.

[0090] In actual implementation, the initial RGB image can be semantically segmented through the U-Net network.

[0091] For example, the U-Net network is used to perform semantic segmentation on the initial RGB image with a dimension of H×W×3 to obtain a semantic feature map with a dimension of H×W×2. The first channel and the second channel of the semantic feature map represent the foreground probability and the semantic feature, respectively, where H is the height of the image and W is the width of the image.

[0092] Step 120 : performing feature extraction processing on the initial depth image based on the foreground probability and semantic features included in the semantic feature map to obtain a depth feature vector.

[0093] A depth image, also called a distance image, refers to an image that uses the depth value from the image collector to each point in the scene as the pixel value, which can reflect the geometric shape of the visible surface in the scene.

[0094] The depth image can be calculated into point cloud data after coordinate transformation, and the point cloud data with rules and necessary information can also be inversely calculated into depth image data.

[0095] The initial depth image is a depth image of a scene corresponding to the acquired initial RGB image, and the initial depth image also includes depth information corresponding to multiple objects.

[0096] For example, the initial depth image may be a depth image of an indoor scene, and the indoor scene may include objects such as tables, chairs, household appliances, and furniture in the room.

[0097] It is understandable that due to the diversity of objects in indoor scenes, scene lighting and the limitations of the acquisition equipment itself, the initial depth image has continuous and large amounts of partial depth information missing or noisy depth data.

[0098] In this step, before feature extraction, the initial depth image and the semantic feature map are feature spliced, and a depth feature vector can be obtained by extracting features from the image spliced ​​with the initial depth image and the semantic feature map.

[0099] It can be understood that the initial depth image and the semantic feature map have the same length and width dimensions, only the number of channels is different, and they can be directly spliced.

[0100] For example, the dimension of the semantic feature map is H×W×2, the dimension of the initial depth image is H×W×1, and the dimension of the image obtained by direct splicing is H×W×3.

[0101] By utilizing the foreground probability and related information of the semantic features included in the semantic feature map, the feature extraction of the initial depth image can be guided to focus on the local depth information, so that the depth feature vector can better represent the depth information of the local objects in the image.

[0102] Feature extraction of the image obtained by splicing the initial depth image and the semantic feature map can include downsampling processing to reduce the size of the spliced ​​image and upsampling processing to increase the resolution of the feature. The corresponding depth feature vector includes the feature vectors in the downsampling processing and upsampling processing processes.

[0103] In this embodiment, the initial depth image and the semantic feature map are spliced, feature extraction is performed on the spliced ​​image, and based on the obtained depth feature vector, a local confidence map and a local depth map are obtained.

[0104] Step 130: perform feature fusion on the depth feature vector and the initial RGB image to obtain a fused confidence map and a fused depth map.

[0105] In this step, a style transfer-like process is performed based on the depth feature vector and the initial RGB image. The depth feature vector is used as an injection to fuse the depth information represented by the depth feature vector with the information of the initial RGB image.

[0106] In this embodiment, the depth feature vector is used as a mapping in the initial RGB image to construct the depth information in the initial RGB image, and a fused depth map fused with the depth feature vector and a fused confidence map representing the pixel confidence of the fused depth map are obtained.

[0107] It can be understood that the depth information corresponding to the fused depth map obtained by fusing the initial RGB image with the depth feature vector is denser than the depth information of the initial depth image.

[0108] Step 140: Obtain a target depth map based on the local confidence map, the local depth map, the fused confidence map, and the fused depth map.

[0109] In this embodiment, the reliable depth prediction area in the local depth map can be better focused through the local confidence map, and the reliable depth prediction area in the fused depth map can be better focused through the fused confidence map.

[0110] In actual implementation, by introducing the attention mechanism and utilizing the local confidence map and the fused confidence map, a weighted summation is performed on the local depth map and the fused depth map, so that the obtained target depth map can more accurately retain the reliable depth prediction area, thereby improving the density and accuracy of the depth information of the target depth map.

[0111] It can be understood that the target depth map is a predicted depth map of the scene corresponding to the initial depth image and the initial RGB image, and is used to represent the depth information of the scene.

[0112] Compared with the initial depth image, the target depth map has denser and more accurate depth information. The target depth map can effectively improve the accuracy and efficiency of subsequent computer vision tasks such as autonomous driving, robot recognition, map navigation, motion planning and augmented reality.

[0113] According to the image processing method provided by the present invention, by extracting the semantic information of the initial RGB image, the initial depth image is guided to accurately characterize the local depth information of the image. At the same time, the global depth map information is optimized by fusing the depth feature vector and the initial RGB image. By introducing the confidence map corresponding to the local depth map and the fused depth map, a reliable depth prediction area is retained, the density and accuracy of the depth information of the target depth map are improved, and a guarantee is provided for subsequent computer vision tasks.

[0114] In some embodiments, a depth completion model is constructed and trained, and the initial depth image and the initial RGB image are processed to obtain a target depth map.

[0115] It should be noted that the initial RGB image may be an RGB image directly captured by the image capture device, or may be an RGB image obtained by converting images in other color spaces.

[0116] In actual implementation, Figure 2 As shown, the initial depth image (shown as d raw ) and the initial RGB image (shown as r) are input into the depth completion model to obtain the target depth map (shown as d pred ).

[0117] It can be understood that after the depth completion model is trained with the sample training set, the initial depth image and the initial RGB image are input for processing.

[0118] The depth image and RGB image in the sample training set can be input into the depth completion model. The loss function of the depth completion model is calculated through the output dense depth image, and the parameters of the depth completion model are adjusted. When the obtained loss function meets the conditions for stopping training, the training of the depth completion model is terminated.

[0119] In this embodiment, if Figure 2 As shown, the deep completion model can include a local guidance module (shown as LocalGuidance Module), a constraint network (shown as Constraint Network), an adversarial generation network (shown as RDF-GAN) and an output network (shown as Confidence Fusion Head).

[0120] The initial RGB image is input into the local guidance module for semantic segmentation to obtain the semantic feature map output by the local guidance module.

[0121] The input of the local guidance module is the initial RGB image. The local guidance module performs semantic segmentation on the initial RGB image and extracts relevant information of foreground probability and semantic features. The output of the local guidance module is a semantic feature map.

[0122] In this step, the initial RGB image is processed by semantic segmentation, which performs dense prediction and inference labels for each pixel in the initial RGB image to achieve fine-grained reasoning, so that each pixel can be labeled as the category of the closed area.

[0123] That is, semantic segmentation is to segment different objects in the initial RGB image from the perspective of pixels and label the pixels in the initial RGB image.

[0124] In actual implementation, the local guidance module can perform semantic segmentation on the initial RGB image through the U-Net network.

[0125] For example, the U-Net network is used to perform semantic segmentation on the initial RGB image with a dimension of H×W×3 to obtain a semantic feature map with a dimension of H×W×2. The first channel and the second channel of the semantic feature map represent the foreground probability and the semantic feature, respectively, where H is the height of the image and W is the width of the image.

[0126] The constraint network performs feature extraction based on the initial depth image and semantic feature map to obtain a depth feature vector.

[0127] In actual implementation, an initial depth image may be collected through a depth sensor. The initial depth image corresponds to the same scene as the initial RGB image, and the two images also correspond to the same viewing angle.

[0128] In this step, before the constraint network performs feature extraction, the initial depth image and the semantic feature map are feature spliced, and a depth feature vector can be obtained by extracting features from the image spliced ​​with the initial depth image and the semantic feature map.

[0129] It can be understood that the initial depth image and the semantic feature map have the same length and width dimensions, only the number of channels is different, and they can be directly spliced.

[0130] For example, the dimension of the semantic feature map is H×W×2, the dimension of the initial depth image is H×W×1, and the dimension of the image obtained by direct splicing is H×W×3.

[0131] In the process of feature extraction by the constraint network, the foreground probability and related information of the semantic features included in the semantic feature map can be used to guide the feature extraction of the initial depth image to focus on local depth information, and obtain a depth feature vector that can characterize the depth information of local objects in the image.

[0132] The constraint network further outputs a local confidence map (shown as c l ) and the local depth map (shown as d l ).

[0133] The feature extraction of the image obtained by splicing the initial depth image and the semantic feature map by the constraint network can include downsampling processing to reduce the size of the spliced ​​image and upsampling processing to increase the resolution of the feature. The corresponding depth feature vector includes the feature vectors in the downsampling processing and upsampling processing processes.

[0134] In this embodiment, the constraint network outputs a local confidence map and a local depth map based on the obtained depth feature vector, where the local depth map refers to a depth image focusing on local depth information, and the local confidence map is a confidence map representing the confidence values ​​of pixels in the local depth map.

[0135] The confidence map is composed of the confidence value of each pixel. The confidence value indicates the possibility that each pixel belongs to the target. The larger the confidence value, the greater the possibility that the pixel belongs to the target.

[0136] Among them, the local depth map refers to a depth image focusing on local depth information, and the local confidence map is a confidence map representing the confidence values ​​of pixels in the local depth map.

[0137] The adversarial generative network performs feature fusion on the deep feature vector and the initial RGB image to obtain a fused confidence map and a fused depth map.

[0138] The input of the generative adversarial network is the depth feature vector and the initial RGB image. The generative adversarial network performs style transfer-like processing based on the depth feature vector and the initial RGB image, injects the depth feature vector as a condition, and fuses the depth information represented by the depth feature vector with the information of the initial RGB image.

[0139] The output of the adversarial generation network is a fused depth map (shown as d f ) and the corresponding fusion confidence map (shown as c f ).

[0140] It should be noted that in the generative adversarial network, before feature extraction, the initial RGB image is first fused with the semantic feature map extracted by the local guidance module, and then during the feature extraction process, the initial RGB image fused with the semantic feature map is fused with the depth feature vector, and finally the fused confidence map and the fused depth map are output.

[0141] The dimension of the initial RGB image fused with the semantic feature map is HxWx5, where 3 channels are R, G, and B, and 2 channels are foreground probability and semantic features.

[0142] In this embodiment, the depth feature vector is used as a mapping in the initial RGB image to construct the depth information in the initial RGB image, and a fused depth map fused with the depth feature vector and a fused confidence map representing the pixel confidence of the fused depth map are obtained.

[0143] The depth feature vector is fused into the initial RGB image. The texture properties of the image corresponding to the image area described by the initial RGB image are fused with the depth information represented by the depth feature vector. The adversarial generative network obtains a fused depth map that integrates visual information and depth information.

[0144] It can be understood that the depth information corresponding to the fused depth map obtained by fusing the initial RGB image with the depth feature vector is denser than the depth information of the initial depth image.

[0145] The output network obtains the target depth map based on the local confidence map, local depth map, fused confidence map and fused depth map.

[0146] The input of the output network is the local confidence map, the local depth map, the fused confidence map and the fused depth map. By introducing the attention mechanism, the local confidence map and the fused confidence map can be used to perform weighted summation on the local depth map and the fused depth map.

[0147] Among them, the local confidence map represents the confidence value of the pixel in the local depth map, and the fused confidence map represents the confidence value of the pixel in the fused depth map.

[0148] It can be understood that the local depth map is obtained by performing depth prediction based on the initial depth image, and the fused depth map is obtained by performing depth prediction based on the initial RGB image.

[0149] In this embodiment, the reliable depth prediction area in the local depth map can be better focused through the local confidence map, and the reliable depth prediction area in the fused depth map can be better focused through the fused confidence map.

[0150] In actual implementation, by introducing the attention mechanism and utilizing the local confidence map and the fused confidence map, a weighted summation is performed on the local depth map and the fused depth map, so that the obtained target depth map can more accurately retain the reliable depth prediction area, thereby improving the density and accuracy of the depth information of the target depth map.

[0151] It can be understood that the target depth map is a predicted depth map of the scene corresponding to the initial depth image and the initial RGB image, and is used to represent the depth information of the scene.

[0152] For example, the initial depth image and the initial RGB image correspond to a certain indoor scene, and the target depth map corresponds to the predicted depth information of the indoor scene.

[0153] Compared with the initial depth image, the target depth map has denser and more accurate depth information. The target depth map can effectively improve the accuracy and efficiency of subsequent computer vision tasks such as autonomous driving, robot recognition, map navigation, motion planning and augmented reality.

[0154] In some embodiments, the generative adversarial network includes a generator and a discriminator.

[0155] The generator reconstructs the depth information of the sample RGB image as much as possible with the depth feature vector extracted from the sample RGB image, and the discriminator (shown as Discriminator) distinguishes the depth image output by the generator from the input sample dense depth map (shown as d gt ) to see if they are consistent, and output the corresponding authenticity result (Fake / Real in the figure).

[0156] During the training process, the generator and discriminator in the adversarial generative network compete with each other to achieve more accurate and reliable output results.

[0157] Among them, the sample dense depth map can be a depth image acquired by using a more accurate depth sensor, and the sample dense depth map can also be an accurate and dense depth map generated based on the size layout of objects in the scene.

[0158] It is understandable that the discriminator and the sample dense depth map are only used as training labels in the training process of the adversarial generation network, and are not needed in the application process of the generator inputting the initial RGB image to predict the output fused depth map.

[0159] In actual implementation, the discriminator can use the design structure of PatchGAN to output a matrix in the form of full convolution to evaluate the depth map generated by the generator, where the authenticity judgment of each point in the matrix represents the evaluation value of each small area in the depth map generated by the generator.

[0160] In the application process of the adversarial generative network, the generator reconstructs the depth information of the initial RGB image with the depth feature vector extracted from the initial depth image to obtain the fused depth map output by the generator and the corresponding fused confidence map.

[0161] It should be noted that the generator of the generative adversarial network performs style transfer-like processing on the depth feature vector and the initial RGB image. The initial RGB image is used as the conditional input of the generator, and the depth feature vector is used as the injection of the generator to fuse the depth information represented by the depth feature vector with the information of the initial RGB image.

[0162] In some embodiments, the loss function of the entire adversarial generative network includes a first loss function corresponding to the generator and a second loss function corresponding to the discriminator.

[0163] In this embodiment, the adversarial generative network can be trained through the network training method of WGAN, which can effectively balance the training degree between the generator and the discriminator. During the training process, a certain numerical value can be generated to indicate the progress of the adversarial generative network training, while also ensuring the diversity of images generated by the generator.

[0164] In the training process of the adversarial generative network, the second loss function corresponding to the discriminator can be determined by the following formula:

[0165] L S =E draw~Draw [S(G(M(d raw ))|r)]-E dgt~Dgt [S(d gt |r)]

[0166] Among them, L S is the second loss function corresponding to the discriminator, r is the initial RGB image, d raw is the initial depth image, D raw is the distribution of the initial depth image, d gt is a sample dense depth image, D gt is the distribution of sample dense depth images, S is the discriminator of the adversarial generation network, G is the generator of the adversarial generation network, M is the encoder of the constraint network, E draw~Draw (*) indicates distribution based on D raw The sampled d raw The expected value of the case*, E dgt~Dgt (*) indicates distribution based on D gt The sampled d gt The expected value of the case*.

[0167] Correspondingly, the first loss function corresponding to the generator can be determined by the following formula:

[0168] L G =-E draw~Draw [S(G(M(d raw ))|r)]+λ g L1[G(M(d raw ))]

[0169] Among them, L G is the second loss function corresponding to the generator, r is the initial RGB image, d raw is the initial depth image, D raw is the distribution of the initial depth image, S is the discriminator of the adversarial generation network, G is the generator of the adversarial generation network, M is the encoder of the constraint network, and E draw~Draw (*) indicates distribution based on D raw The sampled d raw The expected value of the case*.

[0170] Among them, L1 is the loss function of the local depth map and the target depth map, λ g is the weight factor of L1. In actual implementation, λ g It can be taken as 0.5.

[0171] In some embodiments, the constraint network extracts features from the initial depth image and the semantic feature map to obtain the corresponding depth feature vector including:

[0172] like Figure 2As shown, the encoder of the constraint network (shown as Constraint Encoder) downsamples the semantic feature map and the initial depth image to reduce the image size and obtain the first depth feature vector (shown as Z1) output by the encoder.

[0173] The decoder of the constraint network (shown as Upsample1) upsamples the first depth feature vector to restore the resolution and obtains the second depth feature vector (shown as Z2) output by the decoder.

[0174] In this embodiment, the constraint network includes an encoder and a decoder. The encoder learns the feature map of the image spliced ​​by the input initial depth image and the semantic feature map through a neural network, and the decoder labels and segments the feature map that the encoder can provide.

[0175] In actual implementation, the encoder-decoder structure of the constraint network convolution can be constructed using the ResNet-18 network. Before the encoder and decoder of the constraint network are used, the constraint network can be preliminarily trained using the ImageNet image dataset.

[0176] In some embodiments, Figure 2 As shown, the depth completion model also includes an instance normalization module (shown as W-AdaIN), which fuses the deep feature vector of the constraint network and the features of the initial RGB image input into the adversarial generation network.

[0177] Among them, the instance normalization module is a module that performs feature integration based on AdaIN and attention mechanism, and AdaIN is applied to the style migration processing of images.

[0178] The fusion operation of the instance normalization module can be defined as:

[0179]

[0180] Among them, A and B represent the corresponding weights of the depth feature vector z and the initial RGB image r obtained by the self-attention mechanism, respectively, and f r is the feature vector extracted from the initial RGB image, μ and σ are the mean operation and variance operation respectively, y s and b They are respectively identified as the spatial scale factor and bias factor obtained by performing an affine transformation on the depth feature vector z.

[0181] In this embodiment, the instance normalization module uses the self-attention mechanism to assign similar weight values ​​to regions on the same depth plane to appropriately scale the spatial scale factor, and to smooth the depth blocks by assigning similar weight values ​​to locally similar color levels.

[0182] The instance normalization module performs a feature fusion operation on the depth feature vector and the initial RGB image, highlighting the depth information represented by the depth feature vector on the basis of the initial RGB image, so that the adversarial generative network can output a fused depth map and a corresponding fused confidence map based on the features fused by the instance normalization module.

[0183] In this embodiment, if Figure 2 As shown, the generator of the generative adversarial network (shown as Generated Encoder and Upsample2) is similar to the network structure of the constraint network, both of which are encoding and decoding structures. The first deep feature vector output by the constraint network encoder and the second deep feature vector output by the decoder are input and fused to the corresponding positions in the generator of the generative adversarial network through the instance normalization module.

[0184] In actual implementation, the adversarial generative network downsamples the input initial RGB image to obtain the corresponding first eigenvector. At the corresponding position in the constraint network, the encoder of the constraint network outputs the first depth feature vector obtained by the downsampling process.

[0185] The first feature vector and the first depth feature vector are input into an instance normalization module, and the instance normalization module fuses them to obtain a target fused feature vector that fuses the visual information of the initial RGB image and the depth information of the depth feature vector.

[0186] The target fused feature vector is then upsampled in the adversarial generative network to obtain the corresponding second feature vector. At the corresponding position in the constraint network, the decoder of the constraint network outputs the second deep feature vector obtained by the upsampling process.

[0187] The second feature vector and the second depth feature vector are input into the instance normalization module, and the instance normalization module fuses them so that the adversarial generation network outputs a fused depth map and a corresponding fused confidence map.

[0188] It should be noted that the features fused by the instance normalization module are only output to the generative adversarial network for processing and will not be output to the constraint network.

[0189] In some embodiments, the encoder of the constraint network continuously downsamples the input image of the initial depth image and the spliced ​​image of the semantic feature map to obtain a series of depth feature vectors of different dimensional sizes, and the corresponding position of the decoder continuously upsamples.

[0190] like Figure 3 As shown, between the encoder and the decoder, the depth feature vector obtained by downsampling the encoder is input to the corresponding upsampling position in the decoder through a jump connection, so that the decoder outputs a local depth map and its corresponding local confidence map.

[0191] For example, after the encoder inputs the initial depth image and the image concatenated with the semantic feature map, it first performs a convolution with a convolution kernel of 7x7, a step size of 2, and an output channel of 64, and then performs a 3x3 maximum pooling (Max pool) to obtain a depth feature vector with a dimension of H / 4xW / 4,64, which is further downsampled to obtain a depth feature vector of H / 8xW / 8,128 and a depth feature vector of H / 16xW / 16,256.

[0192] At symmetrical positions in the decoder, deep feature vectors with input dimensions of H / 4xW / 4, 64, H / 8xW / 8, 128, and H / 16xW / 16, 256 are jump-connected so that the decoder outputs a local depth map and its corresponding local confidence map.

[0193] Among them, the dimension size of the local depth map and its corresponding local confidence map is HxW.

[0194] The depth feature vector in the encoder of the constraint network is input to the symmetrical position in the decoder of the constraint network by means of jump connection, and the local depth map and local confidence map output by the decoder are obtained.

[0195] In some embodiments, after obtaining the local confidence map, local depth map, fused confidence map and fused depth map, by introducing the attention mechanism, the local confidence map and the fused confidence map are used to obtain the local attention weight corresponding to the local depth map and the fused attention weight corresponding to the fused depth map.

[0196] The fused depth map and the local depth map are weighted by fusion attention weights and local attention weights to calculate the target depth map.

[0197] Through the local confidence map and the fused confidence map, the attention weights corresponding to the local depth map and the fused depth map are obtained, so that when the attention weights are used to calculate the target depth map, the target depth map balances the depth prediction areas corresponding to the local depth map and the fused depth map, so that the depth information of the target depth map is closer to the real depth information.

[0198] It can be understood that compared with the initial depth image collected by the depth sensor, the depth information of the target depth map is denser and more accurate, and the point cloud information converted from the target depth map contains more points than the initial depth image and can better cover the shape of the objects in the scene.

[0199] In actual implementation, based on the local confidence map, the local depth map, the fused confidence map, and the fused depth map, the target depth map can be determined by the following formula:

[0200]

[0201] Among them, d pred represents the target depth image, d l represents the local depth map, d f represents the fused depth map, represents the local attention weight, represents the fused attention weight, and (i, j) represents the coordinates of the pixel with x coordinate i and y coordinate j in the input image.

[0202] In some embodiments, the sample training set includes a plurality of sample RGB images and corresponding simulated depth images.

[0203] In this embodiment, the sample training set used for depth completion mode training includes multiple sample training groups, each of which includes a sample RGB image and at least one simulated depth image corresponding to the sample RGB image.

[0204] The simulated depth image is an image generated by simulating the sample depth image corresponding to the sample RGB image.

[0205] It should be noted that the sample depth image is a depth image of a scene corresponding to the acquired sample RGB image, and the sample RGB image and the sample depth image include the same object.

[0206] In actual implementation, the depth completion model can be trained accordingly by selecting sample training sets corresponding to sample RGB images in different scenes to improve the depth completion effect of the depth completion model in specific scenes.

[0207] For example, the depth completion model may be trained by selecting a sample training set in an indoor scene, where the sample training set includes a sample RGB image in the indoor scene and a simulated depth image corresponding to the sample RGB image.

[0208] It can be understood that by generating realistic depth images through simulation, we can fully simulate the continuous and large amount of partial depth information missing or noisy depth data that appear in the depth images collected by the depth sensor, and then fully train the depth completion model, effectively improving the accuracy and efficiency of the depth completion model training.

[0209] In some embodiments, five different simulated depth maps can be obtained through the following five different implementation methods, and then the simulated depth image corresponding to the sample RGB image is obtained to form a sample training set.

[0210] 1. Mask the highlight area.

[0211] In this embodiment, the sample RGB image is segmented to obtain a highlight area with reflective properties in the sample RGB image, and the area corresponding to the highlight area in the sample depth image is masked, that is, the highlight area in the sample depth image is covered to obtain a first simulated depth map.

[0212] For example, Figure 4 The sample depth image shown is masked with a highlight area to obtain Figure 5 The first realistic depth map is shown.

[0213] 2. Black mask.

[0214] In this embodiment, based on the pixel value of each pixel in the sample RGB image, the pixel values ​​in the sample depth image are divided into target pixel areas in the target pixel range, and the area corresponding to the target pixel area in the sample depth image is masked to obtain a second simulated depth map.

[0215] In actual implementation, the pixels in the target pixel area are pixels whose R, G, and B values ​​are all in [0, 5].

[0216] For example, Figure 4 The sample depth image shown in the figure is masked to the target pixel area, and the result is as follows Figure 6 The second realistic depth map is shown.

[0217] 3. Noise point mask.

[0218] In this embodiment, the sample depth image is segmented to obtain noise points of depth information of the sample depth image, and then the noise points in the sample depth image are masked to obtain a third simulated depth map.

[0219] For example, Figure 4 The sample depth image shown in the figure is subjected to noise point masking to obtain the following Figure 7 The third realistic depth map is shown.

[0220] 4. Semantic label random mask.

[0221] In this embodiment, semantic segmentation is performed on the sample depth image to divide various semantic labels in the sample depth image, and semantic labels in the sample depth image are randomly selected for masking to obtain a fourth simulated depth map.

[0222] For example, Figure 4 The sample depth image shown in the figure is randomly masked with semantic labels to obtain Figure 8 The fourth realistic depth map is shown.

[0223] 5. Mask matrix mask.

[0224] In this embodiment, a U-Net network can be used to perform semantic segmentation task training on a portion of sample depth images randomly divided in a sample training set, and then the results of the semantic segmentation training of the portion of sample depth images are used to perform semantic segmentation on the remaining sample depth images in the sample training set, and the sample depth images are ORed together with the real dense depth images corresponding to the sample depth images to obtain the corresponding mask matrix, and then the mask matrix is ​​used to cover the mask to obtain the fifth simulated depth map.

[0225] Among them, the real dense depth image is acquired by using a more accurate depth sensor, or it can be an accurate and dense depth map generated based on the size layout of objects in the scene. The real dense depth image corresponds to the same scene as the sample depth image.

[0226] For example, Figure 4 The sample depth image shown is masked by the mask matrix to obtain Fig. 9 The fifth realistic depth map is shown.

[0227] After obtaining the first simulated depth map, the second simulated depth map, the third simulated depth map, the fourth simulated depth map and the fifth simulated depth map, at least one image is selected from the five simulated depth maps and the original sample depth image, and the depth information in the image is retained to determine the simulated depth image in the sample training set.

[0228] For example, keep Fig. 9 The fifth realistic depth map and Figure 7 The depth information of the third simulated depth map shown is obtained as follows Fig.10 The simulated depth image.

[0229] In this embodiment, the depth information of at least one of the five simulated depth maps and the six sample depth images is selected to be retained. When two images are selected, the depth information that exists in both images is retained. When three images are selected, the depth information that exists in the three images is retained, and so on.

[0230] In actual implementation, one image retained from the five simulated depth images and the six images of the sample depth image may be the sample depth image. In this case, the simulated depth image is the sample depth image collected by the depth sensor.

[0231] In the related art, a random sampling mask method is used to generate a sparse depth map. The depth information distribution of this type of depth map is very different from that of the real scene. The pixels of the random sampling mask cover all areas in the scene, but in the real scene, the missing depth information usually forms a continuous area.

[0232] In this embodiment, different types of realistic depth maps are generated by a five-mask method, and the depth information of different depth maps is fused to generate a realistic depth image with a more reasonable distribution of missing depth information, which can be used to train a more robust depth completion model.

[0233] In some embodiments, the target loss function used to guide the depth completion model training process includes the loss functions of the generator and the discriminator in the adversarial generation network, the loss function of the local depth map output by the constraint network, and the loss function of the target depth map output by the output network.

[0234] Among them, the loss functions of the generator and the discriminator in the adversarial generative network are the first loss function and the second loss function respectively, the loss function of the local depth map output by the constraint network is the third loss function, and the loss function of the target depth map output by the output network is the fourth loss function.

[0235] In actual implementation, the target loss function of the depth completion model is the sum of the first to fourth loss functions, which can be determined by applying the following formula:

[0236] L overall =L S +L G +λ l L1(d l )+λ pred L1(d pred )

[0237] Among them, L overall is the target loss function, L1(*) is the depth map * and the sample dense depth image d gt The loss function obtained by calculating the absolute value loss, L S is the second loss function, L G is the first loss function, λ l L1(d l ) is the third loss function, λ pred L1(d pred ) is the fourth loss function.

[0238] λ l and λ pred are the loss function weights corresponding to the local depth map and the target depth map, respectively. In actual implementation, λ l and λ pred The values ​​can be 1 and 10.

[0239] The image processing device provided by the embodiment of the present invention is described below. The image processing device described below and the image processing method described above can be referred to each other.

[0240] like Fig.11As shown, the image processing device provided by the embodiment of the present invention includes:

[0241] The first processing module 1110 is used to perform semantic segmentation on the initial RGB image to obtain a semantic feature map;

[0242] A second processing module 1120 is used to extract depth features from an initial depth image corresponding to the initial RGB image based on the foreground probability and the semantic features included in the semantic feature map to obtain a depth feature vector, where the depth feature vector is used to determine a local confidence map and a local depth map;

[0243] The third processing module 1130 is used to perform feature fusion on the depth feature vector and the initial RGB image to obtain a fused confidence map and a fused depth map;

[0244] The fourth processing module 1140 is configured to obtain a target depth map based on the local confidence map, the local depth map, the fused confidence map, and the fused depth map.

[0245] According to the image processing device provided by the present invention, by extracting the semantic information of the target RGB image, the target depth image is guided to accurately characterize the local depth information of the image. At the same time, the global depth map information is optimized by fusing the depth feature vector and the target RGB image. By introducing the confidence map corresponding to the local depth map and the fused depth map, the reliable depth prediction area is retained, the density and accuracy of the depth information of the target dense depth map are improved, and the guarantee for subsequent computer vision tasks is provided.

[0246] In some embodiments, the first processing module 1110 is further used to input the initial RGB image into the local guidance module of the depth completion model for semantic segmentation, and obtain a semantic feature map output by the local guidance module;

[0247] The second processing module 1120 is further used to input the semantic feature map and the initial depth image into the constraint network of the depth completion model to extract the depth feature and obtain the depth feature vector;

[0248] The third processing module 1130 is further used to input the initial RGB image and the depth feature vector into the adversarial generation network of the depth completion model to obtain a fused depth map and a fused confidence map output by the adversarial generation network;

[0249] The fourth processing module 1140 is further used to input the local depth map, the local confidence map, the fused depth map and the fused confidence map into the output network of the depth completion model to obtain a target depth map output by the output network;

[0250] Among them, the depth completion model is trained based on the sample training set.

[0251] In some embodiments, the third processing module 1130 is further used to input the depth feature vector and the initial RGB image into the generator of the adversarial generation network, and the generator reconstructs the depth information of the initial RGB image based on the depth feature vector to obtain the fused depth map and fused confidence map output by the generator;

[0252] The generator is based on sample RGB images and sample dense depth maps, and is trained adversarially against the discriminator of the joint adversarial generative network.

[0253] In some embodiments, the loss function of the adversarial generative network includes a first loss function of the generator and a second loss function of the discriminator.

[0254] In some embodiments, the second processing module 1120 is further used to input the semantic feature map and the initial depth image into the encoder of the constraint network for downsampling processing to obtain a first depth feature vector;

[0255] The first deep feature vector is input into the decoder of the constraint network for upsampling to obtain a second deep feature vector.

[0256] In some embodiments, the third processing module 1130 is further used to input the initial RGB image into the adversarial generation network for downsampling processing to obtain a first feature vector;

[0257] Performing feature fusion on the first feature vector and the first deep feature vector through an instance normalization module of the depth completion model to obtain a target fused feature vector;

[0258] Upsampling the target fused feature vector to obtain a second feature vector;

[0259] The second feature vector and the second depth feature vector are feature fused through the instance normalization module to obtain a fused depth map and a fused confidence map output by the adversarial generation network.

[0260] In some embodiments, the second processing module 1120 is further used to input the depth feature vector in the encoder of the constraint network to a symmetrical position in the decoder of the constraint network through a jump connection to obtain a local depth map and a local confidence map output by the decoder.

[0261] In some embodiments, the fourth processing module 1140 is further configured to obtain a local attention weight and a fused attention weight based on the local depth map, the local confidence map, the fused depth map, and the fused confidence map;

[0262] The local depth map and the fused depth map are weighted based on the local attention weight and the fused attention weight to obtain the target depth map.

[0263] In some embodiments, the sample training set includes a plurality of sample RGB images and their corresponding simulated depth images; the simulated depth images are determined by the following steps:

[0264] Perform a highlight area mask on the sample depth image corresponding to the sample RGB image to obtain a first simulated depth map;

[0265] Perform a target pixel area mask based on the sample depth image to obtain a second simulated depth map;

[0266] Perform noise point masking based on the sample depth image to obtain a third simulated depth map;

[0267] Performing random masking of semantic labels based on the sample depth image to obtain a fourth realistic depth map;

[0268] Perform semantic segmentation training based on all sample depth images in the sample training set, determine the mask matrix, and obtain a fifth simulated depth map;

[0269] Based on the depth information of the sample depth image and at least one of the first to fifth simulated depth maps, a simulated depth image is determined.

[0270] In some embodiments, the target loss function for depth completion model training is determined based on the following steps:

[0271] Obtain a first loss function of the generator of the adversarial generation network and a second loss function of the discriminator;

[0272] Based on the local depth map and the target depth map, a third loss function and a fourth loss function are obtained;

[0273] Based on the first loss function, the second loss function, the third loss function and the fourth loss function, a target loss function is obtained.

[0274] Fig.12 An example of a physical structure diagram of an electronic device is shown in FIG. Fig.12As shown, the electronic device may include: a processor 1210, a communication interface 1220, a memory 1230 and a communication bus 1240, wherein the processor 1210, the communication interface 1220 and the memory 1230 communicate with each other through the communication bus 1240. The processor 1210 may call the logic instructions in the memory 1230 to execute the image processing method, which includes: performing semantic segmentation on the initial RGB image to obtain a semantic feature map; performing depth feature extraction on the initial depth image corresponding to the initial RGB image based on the foreground probability and the semantic feature included in the semantic feature map to obtain a depth feature vector, and the depth feature vector is used to determine a local confidence map and a local depth map; performing feature fusion on the depth feature vector and the initial RGB image to obtain a fused confidence map and a fused depth map; and obtaining a target depth map based on the local confidence map, the local depth map, the fused confidence map and the fused depth map.

[0275] In addition, the logic instructions in the above-mentioned memory 1230 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.

[0276] Furthermore, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the image processing method provided by the above-mentioned method embodiments, which includes: performing semantic segmentation on the initial RGB image to obtain a semantic feature map; based on the foreground probability and semantic features included in the semantic feature map, performing depth feature extraction on an initial depth image corresponding to the initial RGB image to obtain a depth feature vector, and the depth feature vector is used to determine a local confidence map and a local depth map; performing feature fusion on the depth feature vector and the initial RGB image to obtain a fused confidence map and a fused depth map; and obtaining a target depth map based on the local confidence map, the local depth map, the fused confidence map and the fused depth map. On the other hand, an embodiment of the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute the image processing method provided by the above embodiments, the method comprising: performing semantic segmentation on the initial RGB image to obtain a semantic feature map; performing depth feature extraction on an initial depth image corresponding to the initial RGB image based on the foreground probability and semantic features included in the semantic feature map to obtain a depth feature vector, and the depth feature vector is used to determine a local confidence map and a local depth map; performing feature fusion on the depth feature vector and the initial RGB image to obtain a fused confidence map and a fused depth map; and obtaining a target depth map based on the local confidence map, the local depth map, the fused confidence map and the fused depth map.

[0277] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0278] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0279] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

[0280] The above embodiments are only used to illustrate the present invention, but not to limit the present invention. Although the present invention is described in detail with reference to the embodiments, it should be understood by those skilled in the art that various combinations, modifications or equivalent substitutions of the technical solutions of the present invention do not depart from the spirit and scope of the technical solutions of the present invention, and should be included in the scope of the claims of the present invention.

Claims

1. An image processing method, characterized in that: The method further comprises: Perform semantic segmentation on the initial RGB image to obtain a semantic feature map; Based on the foreground probability and the semantic features included in the semantic feature map, a depth feature extraction is performed on the initial depth image corresponding to the initial RGB image to obtain a depth feature vector, wherein the depth feature vector is used to determine a local confidence map and a local depth map; Performing feature fusion on the depth feature vector and the initial RGB image to obtain a fused confidence map and a fused depth map; A target depth map is obtained based on the local confidence map, the local depth map, the fused confidence map and the fused depth map.

2. The image processing method according to claim 1, characterized in that: The semantic segmentation of the initial RGB image to obtain a semantic feature map includes: Inputting the initial RGB image into the local guidance module of the depth completion model for semantic segmentation, and obtaining the semantic feature map output by the local guidance module; The step of extracting depth features from the initial depth image corresponding to the initial RGB image based on the foreground probability and the semantic features included in the semantic feature map to obtain a depth feature vector includes: Inputting the semantic feature map and the initial depth image into the constraint network of the depth completion model to extract depth features, so as to obtain the depth feature vector; The step of fusing the depth feature vector and the initial RGB image to obtain a fused confidence map and a fused depth map includes: Inputting the initial RGB image and the depth feature vector into the adversarial generative network of the depth completion model to obtain the fused depth map and the fused confidence map output by the adversarial generative network; The obtaining a target depth map based on the local confidence map, the local depth map, the fused confidence map and the fused depth map includes: Inputting the local depth map, the local confidence map, the fused depth map and the fused confidence map into the output network of the depth completion model to obtain the target depth map output by the output network; Wherein, the depth completion model is trained based on a sample training set.

3. The image processing method according to claim 2, characterized in that: The step of inputting the initial RGB image and the depth feature vector into the adversarial generative network of the depth completion model to obtain the fused depth map and the fused confidence map output by the adversarial generative network comprises: Inputting the depth feature vector and the initial RGB image into the generator of the generative adversarial network, and having the generator reconstruct the depth information of the initial RGB image based on the depth feature vector, to obtain the fused depth map and the fused confidence map output by the generator; The generator is obtained by conducting adversarial training based on sample RGB images and sample dense depth maps in conjunction with the discriminator of the adversarial generative network.

4. The image processing method according to claim 2, characterized in that: The step of inputting the semantic feature map and the initial depth image into the constraint network of the depth completion model to extract depth features and obtain the depth feature vector includes: Inputting the semantic feature map and the initial depth image into the encoder of the constraint network for downsampling processing to obtain a first depth feature vector; The first depth feature vector is input into the decoder of the constraint network for upsampling to obtain a second depth feature vector.

5. The image processing method according to claim 4, characterized in that: The step of inputting the initial RGB image and the depth feature vector into the adversarial generative network of the depth completion model to obtain the fused depth map and the fused confidence map output by the adversarial generative network comprises: Inputting the initial RGB image into the generative adversarial network for downsampling processing to obtain a first feature vector; Performing feature fusion on the first feature vector and the first depth feature vector through the instance normalization module of the depth completion model to obtain a target fused feature vector; Performing upsampling processing on the target fused feature vector to obtain a second feature vector; The second feature vector and the second depth feature vector are subjected to feature fusion through the instance normalization module to obtain the fused depth map and the fused confidence map output by the generative adversarial network.

6. The image processing method according to claim 2, characterized in that: The local depth map and the local confidence map are determined by the following steps: The depth feature vector in the encoder of the constraint network is input into a symmetrical position in the decoder of the constraint network by means of a jump connection, so as to obtain the local depth map and the local confidence map output by the decoder.

7. The image processing method according to any one of claims 1 to 6, characterized in that: The obtaining a target depth map based on the local confidence map, the local depth map, the fused confidence map and the fused depth map includes: Based on the local depth map, the local confidence map, the fused depth map and the fused confidence map, obtaining a local attention weight and a fused attention weight; The local depth map and the fused depth map are weighted based on the local attention weight and the fused attention weight to obtain the target depth map.

8. The image processing method according to any one of claims 2 to 6, characterized in that: The sample training set includes a plurality of sample RGB images and their corresponding simulated depth images; The simulated depth image is determined by the following steps: Performing a highlight area mask on the sample depth image corresponding to the sample RGB image to obtain a first simulated depth map; Performing a target pixel area mask based on the sample depth image to obtain a second simulated depth map; Perform noise point masking based on the sample depth image to obtain a third simulated depth map; Performing random masking of semantic labels based on the sample depth image to obtain a fourth simulated depth map; Perform semantic segmentation training based on all sample depth images in the sample training set, determine a mask matrix, and obtain a fifth simulated depth map; The simulated depth image is determined based on the depth information of the sample depth image and at least one of the first to fifth simulated depth maps.

9. The image processing method according to any one of claims 2 to 6, characterized in that: The target loss function of the depth completion model training is determined based on the following steps: Obtaining a first loss function of a generator and a second loss function of a discriminator of the generative adversarial network; Based on the local depth map and the target depth map, obtaining a third loss function and a fourth loss function; The target loss function is obtained based on the first loss function, the second loss function, the third loss function and the fourth loss function.

10. An image processing device, characterized in that: include: The first processing module is used to perform semantic segmentation on the initial RGB image to obtain a semantic feature map; A second processing module is used to extract depth features from an initial depth image corresponding to the initial RGB image based on the foreground probability and the semantic features included in the semantic feature map to obtain a depth feature vector, wherein the depth feature vector is used to determine a local confidence map and a local depth map; A third processing module is used to perform feature fusion on the depth feature vector and the initial RGB image to obtain a fused confidence map and a fused depth map; The fourth processing module is used to obtain a target depth map based on the local confidence map, the local depth map, the fused confidence map and the fused depth map.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the image processing method according to any one of claims 1 to 9 are implemented.

12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image processing method according to any one of claims 1 to 9 are implemented.

13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the image processing method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Unsupervised cross-domain self-adaptive medical image segmentation method based on deep adversarial learning

    AU2020103905A4

  • Remote sensing image semi-supervised semantic segmentation method based on generative adversarial network

    CN111080645A