Image processing method and device, electronic equipment and computer program product
Patent Information
- Application Number
- CN202410545134.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2025-10-31
Smart Images

Figure CN120877031A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing, and in particular to an image processing method, apparatus, electronic device, and computer program product. Background Technology
[0002] In the field of computer vision, inferring 3D information from a single 2D image has always been a classic topic. Using only a single ordinary camera combined with algorithms to perceive the depth of objects (i.e., the distance of objects from the lens) is a highly promising application, and this research area is known as monocular depth estimation. In real life, there are numerous scenarios requiring the extraction of depth information from 2D images, such as augmented reality applications on mobile phones and indoor positioning. Furthermore, monocular camera-based depth perception solutions are cheaper and more portable than expensive depth perception solutions like LiDAR, which helps to reduce reliance on costly radar equipment in robot navigation and mapping. Summary of the Invention
[0003] This disclosure provides an image processing method, apparatus, electronic device, and computer program product.
[0004] In a first aspect, embodiments of this disclosure provide an image processing method, the method comprising:
[0005] Acquire the raw images captured by the monocular camera;
[0006] The original image is subjected to preset processing to obtain a processed image with a different resolution than the original image, and the processed image and the original image constitute an image set;
[0007] Multi-scale fusion and multi-level feature extraction are performed on the image information contained in the image set to obtain a hierarchical perceptual sampling depth map for representing depth information at different levels;
[0008] Based on the hierarchical sensing sampling depth map, a fine-grained sensing depth map that aggregates depth information from different layers is obtained, thus obtaining the depth map corresponding to the original image.
[0009] Secondly, embodiments of this disclosure also provide an image processing apparatus, the apparatus comprising:
[0010] The acquisition module is configured to acquire the raw images captured by the monocular camera;
[0011] The first processing module is configured to perform preset processing on the original image to obtain a processed image with a different resolution than the original image, and the processed image and the original image constitute an image set;
[0012] The second processing module is configured to perform multi-scale fusion and multi-level feature extraction on the image information contained in the image set to obtain a hierarchical perceptual sampling depth map for representing depth information at different levels.
[0013] The acquisition module is configured to acquire a fine-grained sensing depth map that aggregates depth information from different layers based on the hierarchical sensing sampling depth map, thereby obtaining the depth map corresponding to the original image.
[0014] Thirdly, embodiments of this disclosure provide an electronic device, the electronic device comprising:
[0015] One or more processors;
[0016] A memory having stored one or more programs that, when executed by one or more processors, enable the one or more processors to implement the image processing method.
[0017] One or more input / output (I / O) interfaces are connected between the processor and the memory and configured to enable information exchange between the processor and the memory.
[0018] Fourthly, embodiments of this disclosure provide a computer program product, which includes a computer program that, when executed by a processor, implements the image processing method.
[0019] This embodiment acquires the original image captured by a monocular camera, performs pre-processing on the original image to obtain a processed image with a different resolution, which enriches the image information of the original image and facilitates subsequent perception of the image's depth information. An image set is constructed from the processed image and the original image. Multi-scale fusion and multi-level feature extraction are performed on the image information contained in this image set, enabling analysis of the image at different scales, thereby obtaining rich multi-scale information. This allows the scheme of this embodiment to capture scene features at different scales, thus providing a more comprehensive understanding of the image content. Feature extraction at different levels effectively captures local details and global structural information in the image, balancing the contributions of local and global information, and fully utilizing depth information at different levels, making the depth estimation results more accurate and stable. Furthermore, due to the adoption of a multi-scale, hierarchical perception sampling strategy, the scheme of this embodiment has a certain degree of anti-interference capability against noise in the image. During the depth estimation process, the fusion of multi-scale information can reduce the impact of noise and improve the reliability of the algorithm. Information at different levels reflects the spatial structure at different scales, effectively capturing the geometric shape and depth distribution of the scene, thereby improving the accuracy and generalization ability of depth estimation. The embodiments disclosed herein enable a monocular vision system to acquire rich depth information from a single image, further assisting in the perception and understanding of three-dimensional space. This is of great significance for applications such as intelligent driving, augmented reality, and intelligent monitoring, and improves the intelligence level of the monocular vision system, thereby achieving a safer and more convenient human-computer interaction experience. Attached Figure Description
[0020] In the accompanying drawings of the embodiments disclosed herein:
[0021] Figure 1 This is a schematic flowchart of an image processing method provided in an embodiment of the present disclosure;
[0022] Figure 2 This is a schematic diagram of the image processing model composition provided in the embodiments of this disclosure;
[0023] Figure 3 This is a schematic diagram of an image processing method based on an image processing model provided in an embodiment of the present disclosure;
[0024] Figure 4 This is a schematic flowchart of a method for obtaining a hierarchical sensing sampling depth map provided in an embodiment of the present disclosure;
[0025] Figure 5 A schematic diagram of a method for obtaining a predicted depth map corresponding to an original image, provided in an embodiment of this disclosure;
[0026] Figure 6 A schematic flowchart of a method for obtaining a predicted depth map corresponding to an original image, provided in an embodiment of this disclosure;
[0027] Figure 7 A schematic diagram of a hierarchical depth sensing sampling coefficient acquisition method provided in this embodiment of the disclosure;
[0028] Figure 8 This is a schematic diagram of the hierarchical depth sensing sampling coefficient acquisition process provided in an embodiment of the present disclosure;
[0029] Figure 9 This is a schematic flowchart of a hierarchical sensing sampling depth map acquisition method provided in an embodiment of this disclosure;
[0030] Figure 10 This is a schematic flowchart of a method for obtaining depth maps after sampling of the corresponding layer, provided in an embodiment of this disclosure.
[0031] Figure 11 This is a schematic diagram of the composition of a deep fine-grained sensing network provided in an embodiment of the present disclosure;
[0032] Figure 12 A schematic diagram of a method for obtaining a fine-grained sensing depth map provided in an embodiment of this disclosure;
[0033] Figure 13 This is a schematic flowchart of a method for obtaining a fine-grained depth map provided in an embodiment of the present disclosure;
[0034] Figure 14 This is a schematic diagram of the method for obtaining the output features of each layer of perceptual cascaded convolutional block provided in the embodiments of this disclosure;
[0035] Figure 15 This is a schematic block diagram of the image processing apparatus provided in the embodiments of this disclosure;
[0036] Figure 16 This is a schematic block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0037] To enable those skilled in the art to better understand the technical solutions of this disclosure, the communication-sensing data processing method and computer-readable storage medium provided in the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings.
[0038] The present disclosure will be described more fully below with reference to the accompanying drawings; however, the embodiments shown may be embodied in different forms, and the present disclosure should not be construed as limited to the embodiments set forth below. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will enable those skilled in the art to fully understand the scope of the disclosure.
[0039] The accompanying drawings of the embodiments disclosed herein are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the detailed embodiments to explain this disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the description of the detailed embodiments with reference to the accompanying drawings.
[0040] This disclosure may be described with reference to plan and / or cross-sectional views using the ideal schematic diagrams of this disclosure. Therefore, the example illustrations may be modified according to manufacturing techniques and / or tolerances.
[0041] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.
[0042] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the disclosure. The term "and / or" as used in this disclosure includes any and all combinations of one or more of the associated enumerated entries. The singular forms "a" and "the" as used in this disclosure are also intended to include the plural forms, unless the context clearly indicates otherwise. The terms "comprising," "made of," etc., as used in this disclosure specify the presence of the stated feature, integral, step, operation, element, and / or component, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof.
[0043] Unless otherwise specified, all terms used in this disclosure (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined in this disclosure.
[0044] In the field of computer vision, inferring 3D information from a single 2D image has always been a classic topic. Using only a single ordinary camera combined with algorithms to perceive the depth of objects (i.e., the distance of objects from the lens) is a highly promising application, and this research area is known as monocular depth estimation. In real life, there are numerous scenarios requiring the extraction of depth information from 2D images, such as augmented reality applications on mobile phones and indoor positioning. Furthermore, monocular camera-based depth perception solutions are cheaper and more portable than expensive depth perception solutions like LiDAR, which helps to reduce reliance on costly radar equipment in robot navigation and mapping.
[0045] Currently, convolutional neural network (CNN) technology is sufficient to predict a certain depth information from a single image. Monocular depth estimation based on neural networks mainly falls into two categories: supervised algorithms and self-supervised algorithms. Supervised algorithms require real depth labels during model training. The need for large amounts of high-quality manually labeled information increases the training cost of traditional supervised algorithms and limits their application. Self-supervised algorithms do not require real depth labels for training, which alleviates the pressure of data labeling. However, self-supervised strategies face challenges such as insufficient perception of the original image during training, insufficient granularity of the depth map, and relatively low depth prediction accuracy. Improvements to monocular depth estimation are expected to have applications in various fields, such as improving human-computer interaction in autonomous driving, virtual reality, and bokeh rendering.
[0046] This embodiment acquires the original image captured by a monocular camera, performs pre-processing on the original image to obtain a processed image with a different resolution, which enriches the image information of the original image and facilitates subsequent perception of the image's depth information. An image set is constructed from the processed image and the original image. Multi-scale fusion and multi-level feature extraction are performed on the image information contained in this image set, enabling analysis of the image at different scales, thereby obtaining rich multi-scale information. This allows the scheme of this embodiment to capture scene features at different scales, thus providing a more comprehensive understanding of the image content. Feature extraction at different levels effectively captures local details and global structural information in the image, balancing the contributions of local and global information, and fully utilizing depth information at different levels, making the depth estimation results more accurate and stable. Furthermore, due to the adoption of a multi-scale, hierarchical perception sampling strategy, the scheme of this embodiment has a certain degree of anti-interference capability against noise in the image. During the depth estimation process, the fusion of multi-scale information can reduce the impact of noise and improve the reliability of the algorithm. Depth information at different levels reflects the spatial structure at different scales, effectively capturing the geometric shape and depth distribution of the scene, thereby improving the accuracy and generalization ability of depth estimation. The embodiments disclosed herein enable a monocular vision system to acquire rich depth information from a single image, thereby achieving perception and understanding of three-dimensional space. This is of great significance for applications such as intelligent driving, augmented reality, and intelligent monitoring, and improves the intelligence level of the monocular vision system, thereby achieving a safer and more convenient human-computer interaction experience.
[0047] The image processing method of this disclosure can be executed by any electronic device, such as a terminal device or a server. The terminal device may include, but is not limited to, in-vehicle devices, user equipment (UE), mobile devices, computing devices, wearable devices, etc., such as, but not limited to, cellular phones, cordless phones, personal digital assistants (PDAs), and portable computers. This image processing method can be implemented by a processor calling computer-readable program instructions stored in memory, or it can be implemented by a server.
[0048] The image processing method of this disclosure can be applied to, but is not limited to, the fields of 3D reconstruction and map building, which require recovering the 3D structure and geometric information of a scene from monocular images. The monocular image depth estimation method based on hierarchical sensing sampling in this disclosure can help 3D reconstruction and map building systems obtain the depth information of a scene, thereby achieving accurate 3D reconstruction and map building, and providing support for applications such as indoor navigation and unmanned aerial vehicle path planning.
[0049] Furthermore, the image processing method of this disclosure can be applied to, but is not limited to, computer vision, autonomous driving, robotics, medical care, virtual reality, game development, training, virtual tourism, remote collaboration, smart home, and Internet of Things fields. For example, the solution of this disclosure can help vehicles perceive the distance information of the road and obstacles ahead, thereby realizing functions such as intelligent obstacle avoidance and driving assistance, and improving the safety of intelligent transportation.
[0050] The solutions disclosed herein have broad market potential, particularly in fields such as computer vision, autonomous driving, and virtual reality. Currently, an increasing number of users are seeking to improve the visual quality of photos taken by cameras through bokeh rendering, creating a demand for good depth estimation. Autonomous driving and intelligent transportation systems are also major applications of depth estimation technology. There is significant market demand, with automakers, technology suppliers, and traffic management departments all seeking reliable monocular depth estimation techniques. The depth estimation methods disclosed herein can provide more accurate and robust environmental perception, potentially improving the safety and feasibility of autonomous vehicles. The virtual reality and augmented reality markets are growing rapidly, and depth estimation methods are crucial for realizing immersive virtual and augmented reality experiences. These technologies can be applied to areas such as gaming, training, virtual tourism, and remote collaboration. The smart home and IoT markets require depth estimation for indoor navigation, security monitoring, and device interconnection. These applications offer growth opportunities for smart home devices and IoT solutions.
[0051] The embodiments of this disclosure will be described in detail below.
[0052] This disclosure provides an image processing method, such as... Figure 1 As shown, the method may include steps S11-S14:
[0053] S11. Acquire the raw image captured by the monocular camera.
[0054] In this embodiment of the disclosure, the original image may include, but is not limited to, an optical RGB (red, green, blue) image.
[0055] In this embodiment of the disclosure, acquiring the raw image captured by a monocular camera may include: acquiring the t-th frame optical RGB image I captured by the monocular camera. t .
[0056] S12. Perform preset processing on the original image to obtain a processed image with a different resolution than the original image. The processed image and the original image constitute an image set.
[0057] In this embodiment of the disclosure, the preset processing may include, but is not limited to, multi-layer downsampling preprocessing to obtain a processed image. The processed image may be one or more images, and the number of processed images is not limited here; it can be defined according to requirements. For example, it may be two images.
[0058] In this embodiment, the resolution of each processed image is different from that of the original image, in order to enhance the image information of the original image. The processed images, together with the original images, can constitute an image set, which may be simply referred to as an image set. For example, an optical RGB image I t The two processed images contain a total of three resolutions, which can form the corresponding image set S(I). t ).
[0059] In this embodiment of the disclosure, when there are multiple original images, a corresponding image set can be obtained for each original image.
[0060] In this embodiment of the disclosure, the resolution of the processed image is not limited in detail and can be defined according to the requirements. For example, for two post-processed images, the resolution can be one-half and one-quarter of the original image size, respectively.
[0061] In this embodiment of the disclosure, there is no limitation on the downsampling method of the original image. Any downsampling method that can be implemented is acceptable, such as including but not limited to: bilinear interpolation, nearest neighbor interpolation, bicubic interpolation, etc.
[0062] S13. Perform multi-scale fusion and multi-level feature extraction on the image information contained in the image set to obtain a hierarchical perceptual sampling depth map for representing depth information at different levels.
[0063] In this embodiment of the disclosure, the method may further include:
[0064] Obtain the preset image processing model;
[0065] The image processing model is used to perform multi-scale fusion and multi-level feature extraction on the image information contained in the image set to obtain a hierarchical perceptual sampling depth map for representing depth information at different levels.
[0066] In this embodiment, the image processing model can be a deep neural network model built and trained using the deep learning framework PyTorch. PyTorch is an open-source deep learning framework for building and training artificial neural networks, designed to provide researchers and engineers with a flexible and dynamic deep learning platform. PyTorch employs a dynamic computation graph approach, meaning that it allows programs to build, modify, and debug neural networks at runtime without needing to predefine the entire computation graph. This makes it highly flexible and intuitive for experimentation and research. The core data structure of PyTorch is the tensor, which is similar to multidimensional arrays in NumPy (an extension library for Python that supports large-dimensional arrays and matrix operations, and also provides a large library of mathematical functions for array operations), but with additional features, enabling accelerated computation on GPUs (Graphics Processing Units). Tensors are the primary way to represent input data, weights, gradients, etc., in PyTorch. PyTorch's Autograd module can automatically compute the gradient of tensors. This is very useful for training neural networks and performing optimization algorithms such as gradient descent, as it allows the model to automatically backpropagate gradients and update weights. PyTorch's design allows users to break down neural network models into multiple modules, each of which can be defined and tuned independently. This modular nature makes it easier to create and reuse complex models.
[0067] In the embodiments disclosed herein, such as Figure 2 As shown, the image processing model 10 may include: a depth prediction network 11, a depth sensing sampling coefficient prediction module 12, and a depth fine-grained sensing network 13.
[0068] In this embodiment, the image processing model 10 includes three deep learning networks. The depth prediction network 11 extracts and summarizes preliminary depth features from the image to obtain a predicted depth map. These coarse preliminary depth features are used to roughly describe the depth distribution of the image. The main function of the depth-aware sampling coefficient prediction module 12 is to obtain depth-aware sampling coefficients at different levels, preparing for the depth-aware adaptive sampling step. Depth-aware adaptive sampling is used to further process the predicted depth map to obtain a hierarchical perception sampling depth map, which represents depth information at different levels. The fine-grained depth perception network 13 comprehensively judges the depth distribution of the scene based on the depth information at different levels contained in the multi-layer hierarchical perception sampling depth map and outputs a refined depth prediction result.
[0069] In this embodiment of the disclosure, obtaining the preset image processing model includes:
[0070] Acquire the target sample image;
[0071] The parameters of the preset neural network are continuously modified based on the principle of minimizing the difference between the target sample image and the preset original sample image. The image processing model is obtained when the difference between the target sample image and the original sample image meets the preset requirements.
[0072] In this embodiment of the disclosure, the preset neural network may include the depth prediction network 11, the depth sensing sampling coefficient prediction module 12, and the depth fine-grained sensing network 13 described above.
[0073] In this embodiment of the disclosure, the target sample image can be obtained by acquiring two-dimensional image coordinates based on the sample estimated depth map and sample fine-grained perception depth map of the original sample image, and then performing coordinate transformation and sampling based on the two-dimensional image coordinates in the coordinate system of the pose transformation sample original image corresponding to the original sample image.
[0074] In this embodiment of the disclosure, acquiring the target sample image may include:
[0075] Based on the preset sample original image, the sample estimated depth map and the fine-grained perception depth map, the first two-dimensional image coordinates are generated on the sample original image.
[0076] The coordinates of the first two-dimensional image are transformed to the coordinate system of the original pose transformation sample image corresponding to the original sample image to obtain the coordinates of the second two-dimensional image.
[0077] The original image of the pose transformation sample is sampled based on the coordinates of the second two-dimensional image to obtain the synthesized target sample image.
[0078] In this embodiment of the disclosure, when training the image processing model of this embodiment, the pose transformation parameters T of the monocular camera in different frames can be used. t→t′ The intrinsic parameter matrix K of the monocular camera itself, and the image I1 at time t. t (This can be used as the original image of the sample mentioned above) corresponds to image I1 at time t′. t′ (This can be used as the original image for the pose transformation sample mentioned above, i.e., image I1) t The image after the attitude change (image I1 at time t′) t′ and the image I1 at time t t The depth map D1 can be synthesized into an image I1′ at time t. t =I1 t′→t The target sample image mentioned above is shown in the following formula:
[0079] I1 t′→t =I1 t′ {proj(D,T t→t′ ,K)};
[0080] In the above formula, D1 represents the image I1 at time t. t The initial estimated depth map D1 c And fine-grained perception depth map D1 x The set, proj() represents the set of depth maps D 1 Image I1 at time t t The generated two-dimensional image coordinates (i.e., the first two-dimensional image coordinates mentioned above) are represented by {}, where {} denotes the sampler. This process utilizes the depth map D to sample image I1. t Projected onto the point cloud, and image I1 t The image is transformed to the coordinate system at time t′, and then the corresponding two-dimensional coordinates (i.e., the second two-dimensional image coordinates mentioned above) are obtained using a bilinear sampler. Based on the obtained coordinate information (i.e., the second two-dimensional image coordinates mentioned above), the image I1 is processed. t′ Sampling yields a composite image I1′ at time t. t =I1 t′→t (i.e., the target sample image mentioned above), and then the synthesized image I1′ can be generated. t =I1 t′→t With image I1 t A comparison is performed to minimize the differences in the comparison, thereby correcting the network parameters of the image processing model.
[0081] In this embodiment of the disclosure, training the image processing model may include:
[0082] The depth estimation network 11 is trained separately by fixing the weight coefficients of the depth sensing sampling coefficient prediction module 12 and the depth fine-grained sensing network 13.
[0083] When the number of training steps reaches the first preset value, the weight coefficients of the depth prediction network 11 are fixed, the fixed weight coefficients of the depth perception sampling coefficient prediction module 12 and the depth fine-grained perception network 13 are canceled, and the training of the depth perception sampling coefficient prediction module 12 and the depth fine-grained perception network 13 begins.
[0084] When the number of training steps reaches the second preset value, the fixation of the weight coefficients of the depth prediction network 11 is cancelled, and the depth prediction network 11, the depth perception sampling coefficient prediction module 12 and the depth fine-grained perception network 13 are trained together until the training ends; wherein, the second preset value is greater than the first preset value.
[0085] In this embodiment of the disclosure, the detailed values of the first preset value and the second preset value are not limited, and can be defined according to requirements.
[0086] In this embodiment of the disclosure, for example, when training an image processing model, the weight coefficients of the depth-sensing sampling coefficient prediction module 12 and the depth fine-grained sensing network 13 can be fixed first, and the depth prediction network 11 can be trained separately. When the number of training steps reaches half of the total number of steps, the weight coefficients of the depth prediction network 11 can be fixed, the fixing operation on the depth-sensing sampling coefficient prediction module 12 and the depth fine-grained sensing network 13 can be canceled, and the training of the depth-sensing sampling coefficient prediction module 12 and the depth fine-grained sensing network 13 can begin. When the number of training steps reaches three-quarters of the total number of steps, the fixing operation on the depth prediction network 11 can be canceled, and the depth prediction network 11, the depth-sensing sampling coefficient prediction module 12, and the depth fine-grained sensing network 13 can be trained together until the training ends.
[0087] In the embodiments disclosed herein, such as Figure 3 , Figure 4 As shown, a preset image processing model is used to perform multi-scale fusion and multi-level feature extraction on the image information contained in the image set to obtain a hierarchical perceptual sampling depth map for representing depth information at different levels. This may include steps S21-S23:
[0088] S21. Input the original image into the depth prediction network of the preset image processing model to obtain the predicted depth map corresponding to the original image.
[0089] In this embodiment of the disclosure, the original image I with meta-resolution can be... t Input the depth prediction network to obtain the predicted depth map D c .
[0090] In the embodiments disclosed herein, such as Figure 5As shown, the depth prediction network 11 may include: shallow depth concatenated convolutional block 111, deep depth concatenated convolutional block 112, upsampling decoding convolutional block 113, and depth prediction convolutional block 114.
[0091] In the embodiments disclosed herein, such as Figure 6 As shown, inputting the original image into a preset depth estimation network to obtain the estimated depth map corresponding to the original image may include steps S31-S34:
[0092] S31. Input the original image into a shallow depth concatenated convolutional block to obtain shallow depth features.
[0093] In this embodiment of the disclosure, the shallow depth-cascaded convolutional block 111 may include: a plurality of first convolutional layers having a first kernel size and a plurality of second convolutional layers having a second kernel size; the plurality of first convolutional layers and the plurality of second convolutional layers are stacked sequentially;
[0094] In this configuration, the first kernel size is larger than the second kernel size; the stride of the first convolutional layer at the front of the multiple first convolutional layers is the first step size, and the stride of the other first convolutional layers and second convolutional layers excluding the first part of the first convolutional layers is the second step size, with the first step size being larger than the second step size.
[0095] In this embodiment of the disclosure, the exact number of the first convolutional layer and the second convolutional layer is not limited and can be defined according to requirements.
[0096] In this embodiment, the detailed dimensions of the first core size and the second core size are not limited, and the values of the first core size and the second core size can be defined according to requirements. For example, 1×1, 3×3, 5×5, etc.
[0097] In this embodiment of the disclosure, the detailed dimensions of the first step length and the second step length are not limited, and the values of the first step length and the second step length can be defined according to the requirements.
[0098] In this embodiment of the disclosure, for example, the shallow depth convolutional block 111 may include: six first convolutional layers with a first kernel size of 3×3 stacked sequentially, and four second convolutional layers with a second kernel size of 1×1. The convolution stride (i.e., the first step size) of the first four first convolutional layers may be 2, and the convolution stride of the remaining convolutional layers may be 1. That is, the first step size of the last two first convolutional layers and the second step size of all the second convolutional layers may be 1.
[0099] In this embodiment of the disclosure, when performing a convolution operation with a stride of 2 (the first stride of the first four first convolutional layers), the resolution of the original input image will change. When the original image resolution changes, the output features of the convolutional layer (e.g., the first four first convolutional layers) will be stored in a cosine positional encoding manner to store coordinate information. This coordinate information will be passed through a convolutional layer with a kernel size of 1×1 (i.e., the second convolutional layer) and concatenated with the original features (i.e., the features before cosine positional encoding, such as the output features of the first four first convolutional layers mentioned above). The above calculation process can be written in the following form:
[0100]
[0101] in, This represents the obtained shallow depth features. This represents a shallow depth-cascaded convolutional block, which incorporates the aforementioned convolution and coordinate encoding operations. Shallow depth features are features initially extracted from the original image, including information such as texture. These shallow depth features will be further refined into deep features in subsequent processes to extract the hidden depth information in the image.
[0102] The aforementioned cosine position encoding can be represented as: p k,2u =sin(k / 1000) 2u / d1 ), p k,2u+1 =cos(k / 1000) 2u / d1 );p k,2u and p k,2u+1 These represent the encoded coordinates, where K refers to the position coordinates in a single dimension (such as the X-axis or Y-axis in a coordinate system); u is the dimension of the position encoding; and d1 is a preset constant term used to constrain u.
[0103] S32. Input the shallow depth features into the deep depth concatenated convolutional block to obtain the deep depth features.
[0104] In this embodiment of the disclosure, the deep convolutional block 112 may include: a plurality of third convolutional layers with a third kernel size and a first activation function; the plurality of third convolutional layers are stacked sequentially, and the last third convolutional layer is connected to the first activation function;
[0105] The last third convolutional layer has the first dilation rate.
[0106] In this embodiment, the exact number of the third convolutional layer is not limited and can be defined according to requirements.
[0107] In this embodiment, the detailed dimensions of the third core are not limited, and the value of the third core size can be defined according to requirements. For example, 1×1, 3×3, 5×5, etc.
[0108] In this embodiment of the disclosure, the detailed value of the first expansion rate is not limited, and the value of the first expansion rate can be defined according to the requirements.
[0109] In this embodiment of the disclosure, for example, the deep concatenated convolutional block 112 may include four stacked third convolutional layers with a kernel size of 3×3 and a first activation function of ReLU1. The convolution stride (i.e., the third stride) of each third convolutional layer can be 1, and the dilation rate (i.e., the first dilation rate) of the last third convolutional layer can be 2. The remaining third convolutional layers have no dilation operation. The output adopts a residual addition structure, and the addition method is element-wise addition. The convolution process can be represented as follows:
[0110]
[0111] in, This represents the obtained deep depth features. This represents the convolution operation of the deep concatenated convolution block 112. The deep depth features have already encoded the implicit depth information in the image, and the deep depth features will be combined with the shallow depth features in subsequent processes to gradually decode the depth map.
[0112] S33. Input the deep depth features and shallow depth features into the upsampled decoding convolutional block to obtain the decoding depth features.
[0113] In this embodiment of the disclosure, the upsampling decoding convolutional block 113 may include: a plurality of fourth convolutional layers having a fourth kernel size; the plurality of fourth convolutional layers are stacked sequentially.
[0114] In this embodiment, the exact number of the fourth convolutional layer is not limited and can be defined according to requirements.
[0115] In this embodiment, the detailed dimensions of the fourth core are not limited, and the value of the fourth core size can be defined according to requirements. For example, 1×1, 3×3, 5×5, etc.
[0116] In this embodiment of the disclosure, for example, the upsampling decoding convolutional block 113 may include: eight fourth convolutional layers with a kernel size of 3×3 stacked sequentially. The first four of the eight fourth convolutional layers can be used to process the aggregation operation of shallow and deep features. The input of shallow and deep features can be concatenated. During concatenation, the features are first processed by convolution and global pooling operations by the fourth convolutional layer with a kernel size of 3×3 to obtain the channel weight coefficient (or weight coefficient). The channel weight coefficient is used to measure the ratio between shallow and deep features.
[0117] In this embodiment of the disclosure, the deep depth features can be upsampled to the same resolution as the shallow depth features before the stitching operation. Subsequently, the shallow depth features are multiplied by the weight coefficient corresponding to the shallow depth features to avoid the feature values of the shallow depth features suppressing the deep depth features. The calculation process can be represented as follows:
[0118]
[0119] in, This represents the obtained decoding depth features. This represents the computation operation of the upsampling decoding convolutional block 113, where U represents the nearest neighbor upsampling operation and Con represents the concatenation operation along the channel dimension. The decoded depth features contain information fused from deep and shallow depth features. This fused information will be used as the input to the last layer of the depth prediction network (i.e., depth prediction convolutional block 114) to produce the predicted depth map.
[0120] S34. Input the decoded depth features into the depth prediction convolution block to obtain the estimated depth map corresponding to the original image.
[0121] In this embodiment of the disclosure, the depth prediction convolutional block 114 may include at least one fifth convolutional layer having a fifth kernel size.
[0122] In this embodiment, the exact number of the fifth convolutional layer is not limited and can be defined according to requirements.
[0123] In this embodiment, the detailed dimensions of the fifth core are not limited, and the value of the fifth core size can be defined according to requirements. For example, 1×1, 3×3, 5×5, etc.
[0124] In this embodiment of the disclosure, for example, the depth prediction convolutional block 114 may include a fifth convolutional layer with a kernel size of 3×3. The convolutional stride of the fifth convolutional layer may be 1, and the convolution process can be represented as follows:
[0125]
[0126] Among them, D c This represents the estimated depth map generated by depth prediction convolutional block 114. This represents the convolution operation of depth prediction convolution block 114.
[0127] In this embodiment, steps S31-S34 above can obtain a predicted depth map corresponding to the original image based on the original image and the depth prediction network. This depth prediction operation performs preliminary processing and analysis on the original image, thereby reducing the computational load of subsequent depth estimation processes and improving overall computational efficiency. Through steps S31-S34, the depth prediction operation generates relatively coarse but fast depth estimation results. These results can serve as a preprocessing step, providing preliminary estimation information for subsequent deep learning models (e.g., the depth-aware sampling coefficient prediction module 12 and the depth fine-grained perception network 13), thereby accelerating the training and inference processes of subsequent deep learning models and the entire image processing model. The depth prediction network can estimate the global structure of the image, extract basic scene information, etc., thereby improving the robustness and stability of subsequent deep learning models and reducing depth estimation errors caused by data quality or environmental changes.
[0128] S22. Input the estimated depth map and image set into the depth perception sampling coefficient prediction module of the preset image processing model to obtain the depth perception sampling coefficients of different levels, which are used as hierarchical depth perception sampling coefficients.
[0129] In this embodiment of the disclosure, this step is to prepare for the depth-sensing adaptive sampling in the next step S23. The depth-sensing adaptive sampling is to further process the estimated depth map to obtain a hierarchical sensing sampling depth map, which is used to represent the depth information of different layers.
[0130] In the embodiments disclosed herein, such as Figure 7 As shown, the depth-sensing sampling coefficient prediction module 12 may include, but is not limited to: shallow sampling concatenated convolutional block 121, deep sampling concatenated convolutional block 122, and sampling coefficient output module 123.
[0131] In the embodiments disclosed herein, such as Figure 8 As shown, the estimated depth map and image set are input into a preset depth-sensing sampling coefficient prediction module to obtain depth-sensing sampling coefficients at different levels, which are used as hierarchical depth-sensing sampling coefficients. This process may include steps S41-S44:
[0132] S41. Input the estimated depth map and the processed image into the shallow sampling convolutional block to obtain the depth-aware sampling shallow features.
[0133] In this embodiment of the disclosure, the estimated depth map obtained in the aforementioned steps and the multi-layer downsampling preprocessed image (i.e. the aforementioned processed image) can be input into the shallow sampling concatenated convolution block 121 in the depth-aware sampling coefficient prediction module 12 to obtain the depth-aware sampling shallow features.
[0134] In this embodiment of the disclosure, the shallow sampling convolutional block 121 may include: a plurality of sixth convolutional layers with a sixth kernel size and a second activation function; the plurality of sixth convolutional layers are stacked sequentially, and the last sixth convolutional layer is connected to the second activation function;
[0135] In this case, the stride of the last sixth convolutional layer is smaller than the stride of the other sixth convolutional layers among the multiple sixth convolutional layers except for the last sixth convolutional layer.
[0136] In this embodiment, the exact number of the sixth convolutional layer is not limited and can be defined according to requirements.
[0137] In this embodiment, the detailed dimensions of the sixth core are not limited, and the value of the sixth core size can be defined according to requirements. For example, 1×1, 3×3, 5×5, etc.
[0138] In this embodiment, the stride of each sixth convolutional layer is not limited and can be defined according to requirements.
[0139] In this embodiment, for example, the shallow sampling convolutional block 121 may include six stacked sixth convolutional layers with a kernel size of 3×3 and a second activation function of ReLU2. The stride of the first five sixth convolutional layers may be 2, and the stride of the last sixth convolutional layer may be 1. The lower-resolution images in the processed images obtained after multi-layer downsampling preprocessing in the image set will be concatenated with the low-resolution feature map generated by the sixth convolutional layer, rather than with the estimated depth map. It should be noted that since the RGB image (a three-channel image) and the depth map (a single-channel image) belong to different modalities, the images in the multi-layer image set (such as the processed image) will be processed by convolutional layers with a preset kernel size (such as 1×1) and GELU activation units (such as the second activation function mentioned above) before the concatenation operation to reduce fusion conflicts between different modalities. That is, the processed image (low-resolution image) obtained after multi-layer downsampling preprocessing is first processed by a 1×1 convolutional layer and GELU activation unit to achieve modality fusion. The modality-fused image is then stitched together with the feature image (also low-resolution) output by the estimated depth map after passing through six sixth convolutional layers to obtain the shallow features of depth-aware sampling.
[0140] In this embodiment of the disclosure, the above calculation process can be represented as follows:
[0141]
[0142] in, This represents the obtained shallow features sampled from depth-sensing (such as shallow features sampled from depth-sensing). D represents the convolution operation of shallow sampling concatenated convolution block 121 (such as a shallow downsampling concatenated convolution block). c To predict the depth map, S(I) t () refers to an image set.
[0143] In this embodiment of the disclosure, the predicted sampling coefficients are used to further optimize the estimated depth map, based on the previously obtained estimated depth map D. c In the initial training stage, due to the small number of training iterations and the local linear sampler during self-supervised training, the sampling range of the high-resolution depth map in the view synthesis process is too small. The depth-aware sampling method of this disclosure can expand the sampling range and improve the accuracy of the depth map.
[0144] S42. Input the shallow features of the depth-sensing sampling into the deep sampling concatenated convolutional block to obtain the deep features of the depth-sensing sampling.
[0145] In this embodiment of the disclosure, the deep sampling convolutional block 122 may include: multiple seventh convolutional layers with a seventh kernel size and a third activation function; the multiple seventh convolutional layers are stacked sequentially, and the last seventh convolutional layer is connected to the third activation function;
[0146] The last seventh convolutional layer has a second dilation rate.
[0147] In this embodiment, the exact number of the seventh convolutional layer is not limited and can be defined according to requirements.
[0148] In this embodiment, the detailed dimensions of the seventh core are not limited, and the value of the seventh core size can be defined according to requirements. For example, 1×1, 3×3, 5×5, etc.
[0149] In this embodiment, the stride and dilation rate of each seventh convolutional layer are not limited and can be defined according to requirements.
[0150] In this embodiment of the disclosure, for example, the deep sampling cascaded convolutional block 122 may include: four sequentially stacked seventh convolutional layers with a seventh kernel size of 3×3 and a third activation function of ReLU3. The convolution stride of the seventh convolutional layer can be 1, the dilation rate of the last seventh convolutional layer can be 2, and the remaining seventh convolutional layers other than the last seventh convolutional layer have no dilation operation. The output can adopt a residual addition structure, and the addition method is element-wise addition.
[0151] In this embodiment of the disclosure, the convolution process described above can be represented as follows:
[0152]
[0153] in, This represents the deep features obtained from the depth-sensing sampling. This represents the convolution operation of the deep sampling cascaded convolution block 122.
[0154] In this embodiment of the disclosure, the deep features of the depth-sensing sampling encode the sampling cues hidden in the image, which will be decoded into sampling coefficients at different levels in a subsequent process.
[0155] S43. Input the deep features of the depth perception sampling into the sampling coefficient output module to obtain the depth perception sampling weight map.
[0156] In this embodiment of the disclosure, the sampling coefficient output module 123 may include at least one eighth convolutional layer having an eighth kernel size.
[0157] In this embodiment, the exact number of the eighth convolutional layer is not limited and can be defined according to requirements.
[0158] In this embodiment, the detailed dimensions of the eighth core are not limited, and the value of the eighth core size can be defined according to requirements. For example, 1×1, 3×3, 5×5, etc.
[0159] In this embodiment of the disclosure, the stride of each eighth convolutional layer is not limited and can be defined according to requirements.
[0160] In this embodiment of the disclosure, for example, the sampling coefficient output module 123 may include: an eighth convolutional layer with an eighth kernel size of 3×3, the convolution stride of which can be 1, and the convolution process can be represented as follows:
[0161]
[0162] Among them, e w The obtained depth-sensing sampling weight map is used as the output of the depth-sensing sampling coefficient prediction module 12. This indicates the convolution operation of the sampling coefficient output module 123.
[0163] In this embodiment, the depth-aware sampling weight map e is obtained based on the convolution operations of the 3×3 and 1×1 convolutional layers in all the foregoing embodiments. w It can be a four-channel weighted graph, which can be represented sequentially according to the channel order. and That is, for each depth-sensing sampling weight map e w This allows us to view the depth-sensing sampling weight map e. w Each channel is processed to obtain the sampling weight value of the corresponding channel, i.e., the weight value of the first channel. The weight value of the second channel The weight value of the third channel The weight value of the fourth channel Among them, the depth-sensing sampling weight map e w Processing each channel may include: averaging all weight values contained in the space corresponding to each channel to obtain the weight value of each channel.
[0164] S44. Based on the numerical values of multiple channels along the spatial dimension of the depth-sensing sampling weight map, the depth-sensing sampling coefficients of different levels of the depth map corresponding to the original image are determined, and used as hierarchical depth-sensing sampling coefficients.
[0165] In this embodiment of the disclosure, depth-sensing sampling coefficients for different levels of the depth map corresponding to the original image are determined based on the numerical values of multiple channels along the spatial dimension of the depth-sensing sampling weight map. These coefficients, used as hierarchical depth-sensing sampling coefficients, may include:
[0166] Calculate the average value of each channel along the spatial dimension in all depth-sensing sampling weight maps containing the multiple channels respectively, and take two average values from the average values corresponding to the multiple channels each time, and use them as the depth smoothing coefficient and position smoothing coefficient of the sampling layer respectively.
[0167] The depth smoothing coefficient and position smoothing coefficient obtained from the multi-layer sampling are used as the hierarchical depth-sensing sampling coefficients.
[0168] In this embodiment of the disclosure, the multiple channels may include four channels; calculating the average value along the spatial dimension of each channel of the total depth-aware sampling weight map containing the multiple channels, and extracting two average values from the average values corresponding to the multiple channels each time, which are respectively used as the depth smoothing coefficient and position smoothing coefficient of a layer of sampling; including:
[0169] The average value of the first channel along the spatial dimension is used as the first layer sampling depth smoothing coefficient of the depth map.
[0170] The average value of the second channel along the spatial dimension is used as the first layer sampling position smoothing coefficient of the depth map;
[0171] The average value of the third channel along the spatial dimension is used as the second-layer sampling depth smoothing coefficient of the depth map.
[0172] The average value of the fourth channel along the spatial dimension is used as the second-layer sampling position smoothing coefficient of the depth map.
[0173] In this embodiment of the disclosure, for example, for the four-channel depth-sensing sampling weight map e described above... w The weight values of each channel are expressed sequentially according to the channel order. and The depth-aware sampling weight map e w The average value along the spatial dimension of each channel from the first to the fourth channel is used as the first-layer sampling depth smoothing coefficient, the first-layer sampling position smoothing coefficient, the second-layer sampling depth smoothing coefficient, and the second-layer sampling position smoothing coefficient of the depth map, respectively. The calculation process can be represented as follows:
[0174]
[0175]
[0176] in, This represents the smoothing coefficient of the first layer sampling depth. This represents the smoothing coefficient for the first layer sampling position. This represents the smoothing coefficient for the second layer sampling depth. represents the smoothing coefficient of the second-layer sampling position, |·| represents the number of elements in the specified depth-aware sampling weight map, and ∑ represents the summation in the spatial dimension.
[0177] In this embodiment of the disclosure, and These are used to control the degree of blurring in the depth domain during the first and second layer sampling processes, respectively. and These parameters are used to control the degree of blurring in the position domain during the first and second layer sampling processes. By obtaining these control parameters (i.e., hierarchical depth-sensing sampling coefficients) in a learnable manner, adaptive sampling operations can be performed, preventing excessive blurring during sampling.
[0178] In this embodiment, steps S41-S44 input the estimated depth map and the image set obtained after multi-level downsampling (this step mainly uses the processed images in the image set) into the depth-aware sampling coefficient prediction module 12 to obtain depth-aware sampling coefficients at different levels. The estimated depth map contains depth information for each pixel in the image, while the image set after multi-level downsampling contains feature information at different scales. Combining these two allows for full utilization of the depth information and multi-scale features in the image. Hierarchical depth-aware sampling coefficients can help the image processing model better adapt to image features at different depths and scales, improving the accuracy and robustness of depth perception, thereby improving the performance of subsequent depth estimation.
[0179] In this embodiment, the hierarchical depth-sensing sampling coefficients take into account both the two-dimensional image plane coordinates and the three-dimensional depth coordinates of the depth map, and obtain sampling coefficients at different levels through a learnable method. This improves the utilization of multi-dimensional information by the image processing model and makes the acquisition of hierarchical depth-sensing sampling coefficients more adaptive. In contrast, existing common solutions only calculate filtering parameters in the two-dimensional plane or use fixed filter configurations, which is not conducive to fully utilizing the multi-dimensional cues in the depth map.
[0180] S23. Use the hierarchical depth sensing sampling coefficients to perform adaptive sampling processing on the estimated depth map to obtain hierarchical sensing sampling depth maps that represent depth information at different levels, and obtain a set of hierarchical sensing sampling depth maps.
[0181] In this embodiment of the disclosure, the estimated depth map is subjected to adaptive sampling processing using hierarchical depth sensing sampling coefficients to obtain a set of hierarchical sensing sampling depth maps. This is to ensure that each pixel in the depth map can well integrate the depth information of its neighboring pixels during sampling, and to ensure that the pixel's own characteristics are not obscured by a large amount of pixel information in its neighborhood, so as to well represent the depth information of different layers.
[0182] In this embodiment of the disclosure, the hierarchical depth-sensing sampling coefficients may include, but are not limited to, the depth smoothing coefficients and position smoothing coefficients corresponding to each layer of sampling.
[0183] In the embodiments disclosed herein, such as Figure 9 As shown, the estimated depth map is subjected to adaptive sampling processing using hierarchical depth-sensing sampling coefficients to obtain hierarchical sensing sampling depth maps that represent depth information at different levels, resulting in a set of hierarchical sensing sampling depth maps. This process may include steps S51-S52:
[0184] S51. Using the depth smoothing coefficient and position smoothing coefficient corresponding to each layer of sampling, perform depth-aware adaptive sampling processing on the estimated depth map to obtain the depth map after sampling of the corresponding layer.
[0185] In this embodiment of the disclosure, depth-aware adaptive sampling processing is performed on the estimated depth map using the depth smoothing coefficient and position smoothing coefficient of each sampling layer to obtain the depth map after sampling of the corresponding layer, which may include:
[0186] Based on the depth smoothing coefficient and position smoothing coefficient of each layer of sampling and the preset sampling formula, the depth value of each pixel in the estimated depth map is sampled based on the depth values of the neighboring pixels of the pixel to obtain the depth map after sampling of the corresponding layer.
[0187] In this embodiment of the disclosure, the depth smoothing coefficient may include: a first layer sampling depth smoothing coefficient and a second layer sampling depth smoothing coefficient; the position smoothing coefficient may include: a first layer sampling position smoothing coefficient and a second layer sampling position smoothing coefficient.
[0188] In the embodiments disclosed herein, such as Figure 10 As shown, the depth-aware adaptive sampling process is performed on the estimated depth map using the depth smoothing coefficient and position smoothing coefficient of each sampling layer to obtain the depth map after sampling of the corresponding layer, which may include steps S61-S62:
[0189] S61. Using the first-layer sampling depth smoothing coefficient and the first-layer sampling position smoothing coefficient, the estimated depth map is subjected to first-layer depth perception adaptive sampling processing to obtain the first-layer sampled depth map.
[0190] In this embodiment of the disclosure, the sampling kernel function weights (i.e., and Equal weight values represent the different sampling coefficients in the hierarchical depth-sensing sampling coefficients, such as the depth smoothing coefficient of the first layer sampling. Smoothing coefficient of the first layer sampling position The allocation method (as described in the following calculation formula) can be implemented using bilateral filtering. The specific sampling calculation process is as follows:
[0191]
[0192]
[0193] In the above formula, o represents the pixel p within the first preset neighborhood (e.g., 3×3) (i.e., the neighboring pixel), and the first term of the equation... |po| 2 This represents the proximity of the current pixel p to its neighboring pixel o in two-dimensional image coordinates. Pixel o in the image that is closer to pixel p in coordinates is... The larger the value obtained, the greater the weight that pixel 'o' has in the filtering operation; the second term of the equation |D c (p)-D c (o)| 2 This indicates how similar the depth values of the current pixel p are to those of its neighboring pixel o. Pixels with more similar depth values are passed through... The larger the obtained value, the greater the weight that pixel 'o' has in the filtering operation. The above two steps prevent abrupt changes in depth in edge regions from interfering with the filtering operation, resulting in more consistent depth within the same object, while maintaining clear depth boundaries at edges. D c (p) represents the depth value at pixel location p in the estimated depth map, D c(o) represents the depth value at pixel position o in the estimated depth map, H1 indicates that the downsampling operation of the first layer will reduce the output resolution to half of the input resolution, and D1 represents the depth map after the first layer sampling.
[0194] S62. Using the second-layer sampling depth smoothing coefficient and the second-layer sampling position smoothing coefficient, perform adaptive sampling processing of the estimated depth map for second-layer depth perception to obtain the depth map after second-layer sampling.
[0195] In this embodiment of the disclosure, the sampling kernel function weights (i.e., and Equal weight values represent the different sampling coefficients in the hierarchical depth-sensing sampling coefficients, such as the second-layer sampling depth smoothing coefficient. Second layer sampling position smoothing coefficient The allocation method (as described in the following calculation formula) can be implemented using bilateral filtering. The specific sampling calculation process is as follows:
[0196]
[0197]
[0198] In the above formula, o represents the pixel p within the second preset neighborhood (e.g., 5×5), H2 indicates that the downsampling operation of the second layer reduces the output resolution to one-quarter of the input resolution, and D2 represents the depth map after the second layer sampling. The operation and principle of this part are similar to the process of generating D1, but the resulting resolution is lower, in order to capture multi-scale information at different resolutions.
[0199] S52. A set of hierarchical sensing sampling depth maps is composed of the estimated depth map and the depth map after multi-layer sampling.
[0200] In this embodiment of the disclosure, it can be provided by D c D1 and D2 together constitute a set of hierarchical sensing sampling depth maps.
[0201] In this embodiment, the hierarchical sensing sampling depth map can dynamically adjust the sampling strategy to better adapt to the depth distribution of different scenes and objects, thereby improving the efficiency and accuracy of depth estimation. The hierarchical sensing sampling depth map allows image processing models to focus more on regions with significant depth changes at different scales, thus enhancing the image processing model's ability to perceive depth information. This helps the image processing model more accurately capture the depth structure in the image, improving the accuracy and robustness of depth estimation.
[0202] In this embodiment, the filtering operation is based on two-dimensional image planar coordinates and three-dimensional depth coordinates. Traditional bilateral filtering operations can only filter based on planar coordinates and image color in a two-dimensional image, failing to utilize depth information. The filtering operation in this embodiment takes into account the abrupt changes in depth of different objects in three-dimensional space. This helps maintain the clear boundaries of objects at the foreground-background interface, preventing depth blurring between the foreground and background.
[0203] S14. Based on the hierarchical sensing sampling depth map, obtain a fine-grained sensing depth map that aggregates depth information from different levels, and obtain the depth map corresponding to the original image.
[0204] In this embodiment of the disclosure, the method may further include:
[0205] Obtain the preset image processing model;
[0206] The image processing model is used to obtain a fine-grained sensing depth map that aggregates depth information from different layers based on the hierarchical sensing sampling depth map, thus obtaining the depth map corresponding to the original image.
[0207] In this embodiment of the disclosure, obtaining a fine-grained sensing depth map that aggregates depth information from different layers based on a hierarchical sensing sampling depth map may include:
[0208] The set of hierarchical sensing sampling depth maps and the estimated depth map are input into a preset depth fine-grained sensing network. The depth fine-grained sensing network aggregates depth information from different layers in the hierarchical sensing sampling depth maps to obtain a fine-grained sensing depth map.
[0209] In this embodiment of the disclosure, this step is to aggregate depth information from different layers in different hierarchical sensing sampling depth maps to obtain fine-grained sensing depth.
[0210] In the embodiments disclosed herein, such as Figure 11 As shown, the deep fine-grained perception network 13 may include: a multi-layer perception cascaded convolutional block 131, a fine upsampling decoding convolutional block 132, and a fine-grained depth prediction convolutional block 133.
[0211] In the embodiments disclosed herein, such as Figure 12 , Figure 13 As shown, the set of hierarchical sensing sampling depth maps and the estimated depth map are input into a preset depth fine-grained sensing network. This fine-grained sensing network aggregates depth information from different layers in the hierarchical sensing sampling depth maps to obtain a fine-grained sensing depth map, which may include steps S71-S73:
[0212] S71. Input the predicted depth map Dc and the multi-layer sampled depth maps from the set of hierarchical sensing sampling depth maps into the multi-layer sensing convolutional block respectively, and use the output of each layer sensing convolutional block as the input of the next layer sensing convolutional block to obtain the features output by each layer sensing convolutional block.
[0213] In this embodiment of the disclosure, the input of each layer of the perceptual convolutional block, except for the first layer of the perceptual convolutional block in the multilayer perceptual convolutional block 131, also includes the output of the previous layer of the perceptual convolutional block 131.
[0214] In this embodiment of the disclosure, the multilayer sensing convolutional block 131 may include: a shallow sensing convolutional block 1311, a deep sensing convolutional block 1312, and a thinning layer sensing convolutional block 1313; the depth map after multilayer sampling may include: a depth map D1 after the first layer sampling and a depth map D2 after the second layer sampling.
[0215] In the embodiments disclosed herein, such as Figure 14 As shown, the multi-layered sampled depth maps from the set of predicted depth maps and hierarchical perceptual sampled depth maps are respectively input into the multi-layer perceptual convolutional block, and the output of each perceptual convolutional block is used as the input of the next perceptual convolutional block to obtain the features output by each perceptual convolutional block. This can include steps S81-S83:
[0216] S81. Input the estimated depth map into the shallow perceptual cascaded convolutional block to obtain shallow fine-grained depth features.
[0217] In this embodiment of the disclosure, the shallow-sensing cascaded convolutional block 1311 may include: multiple ninth convolutional layers with a ninth kernel size and a fourth activation function; the multiple ninth convolutional layers are stacked sequentially, and the last ninth convolutional layer is connected to the fourth activation function;
[0218] Among the multiple ninth convolutional layers, the stride of the first part of the ninth convolutional layers is the third stride, and the stride of the other ninth convolutional layers (excluding the first part of the ninth convolutional layers) is the fourth stride, with the third stride being greater than the fourth stride.
[0219] In this embodiment, the exact number of the ninth convolutional layer is not limited and can be defined according to requirements.
[0220] In this embodiment, the detailed dimensions of the ninth core are not limited, and the value of the ninth core size can be defined according to requirements. For example, 1×1, 3×3, 5×5, etc.
[0221] In this embodiment, the detailed stride of each ninth convolutional layer is not limited and can be defined according to requirements.
[0222] In this embodiment, for example, the shallow-sensing cascaded convolutional block 1311 (or shallow-fine-grained sensing cascaded convolutional block) may include two stacked ninth convolutional layers with a kernel size of 3×3 and a fourth activation function ReLU4. The convolution stride of the first ninth convolutional layer (e.g., the third stride) can be 2, and the convolution stride of the second ninth convolutional layer (e.g., the fourth stride) can be 1. When performing a convolution operation with a stride of 2, it will cause a change in resolution. In the case of a change in resolution, the output features of the ninth convolutional layer (e.g., the first ninth convolutional layer) will be stored in the form of cosine position encoding to store coordinate information. This coordinate information will pass through a convolutional layer with a kernel size of 1×1. The output of the 1×1 convolutional layer is concatenated with the original features (the output features after the estimated depth map is input into the first ninth convolutional layer). The above process can be written in the following form:
[0223]
[0224] in, This represents the shallow, fine-grained depth characteristics obtained. D represents the convolution and coordinate calculation operations of shallow perceptual cascaded convolution block 1311. c This refers to the estimated depth map with the original resolution, which has the highest resolution. Therefore, the estimated depth map D can be used first here. c The input is a depth fine-grained perception network 13, and other lower-resolution depth maps will be gradually added to the depth fine-grained perception network 13 in subsequent processes.
[0225] S82. Input the depth map after sampling of the first layer and the shallow fine-grained depth features into the deep perception cascaded convolution block to obtain the deep fine-grained depth features.
[0226] In this embodiment of the disclosure, the deep-perception cascaded convolutional block 1312 may include: multiple tenth convolutional layers with a tenth kernel size and a fifth activation function; the multiple tenth convolutional layers are stacked sequentially, and the last tenth convolutional layer is connected to the fifth activation function;
[0227] Among them, the stride of the first part of the 10th convolutional layers is the fifth stride, and the stride of the other 10th convolutional layers besides the first part is the sixth stride, with the fifth stride being greater than the sixth stride.
[0228] The last tenth convolutional layer has a third dilation rate.
[0229] In this embodiment, the exact number of the tenth convolutional layer is not limited and can be defined according to requirements.
[0230] In this embodiment, the detailed dimensions of the tenth core are not limited, and the value of the tenth core size can be defined according to requirements. For example, 1×1, 3×3, 5×5, etc.
[0231] In this embodiment, the detailed stride (fifth stride and sixth stride) and dilation rate (third dilation rate) of each tenth convolutional layer are not limited and can be defined according to requirements.
[0232] In this embodiment of the disclosure, for example, the deep-perception cascaded convolutional block 1312 may include two stacked 10th convolutional layers with a 3×3 kernel size and a fifth activation function ReLU5; wherein, the convolution stride of the first 10th convolutional layer (e.g., the fifth stride) can be 2, the convolution stride of the second 10th convolutional layer (e.g., the sixth stride) can be 1, the third dilation rate of the last 10th convolutional layer can be 2, and the remaining 10th convolutional layers may have no dilation operation, providing shallow, fine-grained depth features. The input to the depth map D1 after sampling from the first layer can be concatenated. The convolution process can be represented as follows:
[0233]
[0234] in, This represents the obtained deep fine-grained depth features. This represents the convolution operation of the deep perception cascaded convolution block 1312, where Con represents the convolution operation on the shallow fine-grained depth features. The operation of stitching together the channel dimension features of the depth map D1 after the first layer sampling.
[0235] S83. Input the depth map after sampling of the second layer and the deep fine-grained depth features into the refining layer perceptual cascaded convolution block to obtain the fine-grained depth features of the refining layer.
[0236] In this embodiment of the disclosure, the thinning layer sensing cascaded convolutional block 1313 (which may be referred to as the thinning layer fine-grained sensing cascaded convolutional block) may include: a plurality of eleventh convolutional layers having an eleventh kernel size; the plurality of eleventh convolutional layers are stacked sequentially, and the last eleventh convolutional layer has a fourth dilation rate.
[0237] In this embodiment, the exact number of the eleventh convolutional layer is not limited and can be defined according to requirements.
[0238] In this embodiment, the detailed dimensions of the eleventh core are not limited, and the value of the eleventh core size can be defined according to requirements. For example, 1×1, 3×3, 5×5, etc.
[0239] In this embodiment, the detailed stride and dilation rate (fourth dilation rate) of each eleventh convolutional layer are not limited and can be defined according to requirements.
[0240] In this embodiment of the disclosure, for example, the refined layer perceptual cascaded convolutional block 1313 may include: two sequentially stacked eleventh convolutional layers with a kernel size of 3×3, a convolution stride of 2, and a dilation rate (fourth dilation rate) of the last eleventh convolutional layer of 2. The remaining eleventh convolutional layers may have no dilation operation, providing deep fine-grained depth features. The input to the depth map D2 after the second layer sampling can be stitched together, and the convolution process can be represented as follows:
[0241]
[0242] in, This represents the fine-grained depth characteristics of the refined layer. This represents the convolution operation of the refined layer perceptual cascaded convolution block 1313.
[0243] S72. Obtain all features output by the multi-layer perceptron convolutional block, and input all features into the refined upsampling decoding convolutional block to obtain decoded fine-grained depth features.
[0244] In this embodiment of the disclosure, when the multilayer perceptual cascaded convolutional block 131 includes a shallow perceptual cascaded convolutional block 1311, a deep perceptual cascaded convolutional block 1312, and a thinning layer perceptual cascaded convolutional block 1313, all features output by the multilayer perceptual cascaded convolutional block 131 may include the aforementioned shallow fine-grained depth features, deep fine-grained depth features, and thinning layer fine-grained depth features.
[0245] In this embodiment of the disclosure, obtaining all features output by the multilayer perceptron concatenated convolutional block and inputting all features into the refined upsampling decoding convolutional block to obtain decoded fine-grained depth features may include:
[0246] The shallow fine-grained depth features, deep fine-grained depth features, and fine-grained depth features of the thinning layer are input into the thinning upsampling decoding convolutional block 132 to obtain the decoding fine-grained depth features.
[0247] In this embodiment of the disclosure, the refined upsampling decoding convolutional block 132 may include: a plurality of twelfth convolutional layers having a twelfth kernel size; the plurality of twelfth convolutional layers are stacked sequentially.
[0248] In this embodiment, the exact number of the twelfth convolutional layer is not limited and can be defined according to requirements.
[0249] In this embodiment, the detailed dimensions of the twelfth core are not limited, and the value of the twelfth core size can be defined according to requirements. For example, 1×1, 3×3, 5×5, etc.
[0250] In this embodiment, the detailed stride and dilation rate of each twelfth convolutional layer are not limited and can be defined according to requirements.
[0251] In this embodiment of the disclosure, for example, the upsampling decoding convolutional block 132 may include eight sequentially stacked twelfth convolutional layers with a kernel size of 3×3. All features output from the input multilayer perceptron cascaded convolutional block (shallow fine-grained depth features, deep fine-grained depth features, and refinement layer fine-grained depth features) can be concatenated. The first four twelfth convolutional layers can be used to process the aggregation operation of fine-grained depth features (shallow fine-grained depth features, deep fine-grained depth features, and refinement layer fine-grained depth features) from different layers. This aggregation operation may include concatenating the fine-grained depth features from different layers and inputting the concatenated features into the first four twelfth convolutional layers of the refinement upsampling decoding convolutional block 132.
[0252] In this embodiment of the disclosure, the above-mentioned aggregation operation process may include: fine-grained depth features from different layers are first convolved by eight twelfth convolutional layers with a kernel size of 3×3 during the convolution operation. After the convolution operation, a global pooling operation (averaging) is performed to obtain a channel weight coefficient to measure the proportional relationship of fine-grained features from different layers. Before the convolution operation, the deep fine-grained depth features and the refined layer fine-grained depth features are upsampled to the same resolution as the shallow fine-grained depth features. Then, the shallow fine-grained depth features are multiplied by the weight coefficient corresponding to the shallow fine-grained depth features to avoid the feature values of the shallow fine-grained depth features suppressing the feature values of the deep fine-grained depth features and the refined layer fine-grained depth features. The upsampled deep fine-grained depth features, the refined layer fine-grained depth features, and the shallow fine-grained depth features multiplied by the weight coefficient are multiplied by the above-mentioned channel weight coefficient. The three features obtained are then convolved, and the convolved features are input into the first four twelfth convolutional layers for convolution operation, thereby completing the above-mentioned aggregation operation and obtaining the decoded fine-grained depth features. The convolution process of the above aggregation operation can be represented as follows:
[0253]
[0254] in, This represents the obtained decoded fine-grained depth features. This represents the computational operation of the refined upsampled decoded convolutional block 132, where U represents the nearest neighbor upsampling operation and Con represents the fine-grained depth features of the obtained refined layer. The splicing operation at the channel level. Here. It incorporates information from a large sampling area in the original depth map (i.e., the estimated depth map).
[0255] S73. Input the decoded fine-grained depth features into the fine-grained depth prediction convolutional block to obtain the fine-grained perception depth map.
[0256] In this embodiment of the disclosure, the fine-grained depth prediction convolutional block 133 may include at least one thirteenth convolutional layer having a thirteenth kernel size.
[0257] In this embodiment, the exact number of the thirteenth convolutional layer is not limited and can be defined according to requirements.
[0258] In this embodiment, the detailed dimensions of the thirteenth core are not limited, and the value of the thirteenth core size can be defined according to requirements. For example, 1×1, 3×3, 5×5, etc.
[0259] In this embodiment, the detailed stride and dilation rate of each thirteenth convolutional layer are not limited and can be defined according to requirements.
[0260] In this embodiment of the disclosure, for example, the fine-grained depth prediction convolutional block 133 may include a thirteenth convolutional layer with a thirteenth kernel size of 3×3, and the convolution stride may be 1. The convolution process can be represented as follows:
[0261]
[0262] Among them, D x This represents a fine-grained depth map. This represents the convolution operation of fine-grained depth prediction convolution block 133, and the fine-grained perception depth map D. x Compared to the original estimated depth map D c The depth of each pixel also contains information about other neighboring pixels in a larger neighborhood.
[0263] In this embodiment of the disclosure, the above scheme can be used to obtain a depth map corresponding to a two-dimensional image (such as the original image), namely the fine-grained perception depth map D mentioned above. x .
[0264] In this embodiment, the deep fine-grained perceptual network 13 can further refine and optimize the initial depth map (estimated depth map) by utilizing information from the hierarchical perceptual sampling depth map. This means that the image processing model can more accurately estimate the depth value of each pixel in the image, thereby improving the precision and accuracy of depth estimation and enhancing the image processing model's ability to understand the scene. The deep fine-grained perceptual network 13 can not only consider the overall depth distribution and global consistency, but also perform local refinement based on local features in the hierarchical perceptual sampling depth map. This ensures both the overall consistency and stability of the depth map and enables more refined depth estimation at details, achieving good depth estimation results at different scales and granularities.
[0265] In this embodiment, the multi-level feature (such as all features output by a multilayer perceptron convolutional block) convergence operation utilizes depth information from different levels, rather than simply stitching together features of different resolutions. In this embodiment, the image processing model accepts pre-processed multilayer depth maps as priors, which helps reduce the learning difficulty of the image processing model.
[0266] In this embodiment of the disclosure, the solution includes at least the following advantages:
[0267] 1. This disclosure enables image analysis at different scales, thereby obtaining rich multi-scale information. This allows the algorithm to capture scene features at different scales, thus providing a more comprehensive understanding of image content. Feature extraction at different levels effectively captures local details and global structural information in the image. This balances the contributions of local and global information, making the depth estimation results more accurate and stable. Due to the adoption of a multi-scale hierarchical perceptual sampling strategy, this disclosure provides a certain degree of resistance to noise in the image. During the depth estimation process, fusing multi-scale information reduces the impact of noise and improves the reliability of the algorithm. Information at different levels reflects the spatial structure at different scales. This design effectively captures the geometry and depth distribution of the scene, thereby improving the accuracy and generalization ability of depth estimation.
[0268] 2. The embodiments of this disclosure improve the environmental understanding capability of a monocular perception system. Through the solutions of the embodiments of this disclosure, a monocular vision system can acquire rich depth information from a single image, achieving perception and understanding of three-dimensional space. This is of great significance for applications such as intelligent driving, augmented reality, and intelligent monitoring, and can improve the intelligence level of the monocular perception system, thereby achieving a safer and more convenient human-computer interaction experience.
[0269] 3. This embodiment of the present disclosure utilizes a learnable method to predict hierarchical depth-sensing sampling coefficients and performs adaptive depth-sensing sampling processing on the depth map to obtain a hierarchical sensing sampling depth map. This allows each pixel in the depth map to effectively integrate the depth information of its neighboring pixels and ensures that the characteristics of each pixel are not obscured by a large amount of pixel information in its neighborhood, thus effectively representing depth information at different levels. This embodiment of the present disclosure fully utilizes the depth-sensing features of multi-layer images by aggregating hierarchical sensing sampling depth maps. It combines sampling strategies at different levels to perform adaptive depth sampling feature aggregation. Through a learnable method, it adaptively adjusts the sampling degree of the depth domain and the position domain, allowing pixels of different depths to aggregate neighborhood information without obscuring their own information. This extracts the rich depth information hidden in the image and achieves multi-layer aggregation of depth information to generate a high-accuracy depth map, which is beneficial for the application of self-supervised depth estimation in practical production and life.
[0270] This disclosure also provides an image processing apparatus 100, such as... Figure 15 As shown, it includes:
[0271] The acquisition module 101 is configured to acquire the raw images captured by the monocular camera.
[0272] The first processing module 102 is configured to perform preset processing on the original image to obtain a processed image with a different resolution than the original image, and the processed image and the original image constitute an image set.
[0273] The second processing module 103 is configured to perform multi-scale fusion and multi-level feature extraction on the image information contained in the image set to obtain a hierarchical perceptual sampling depth map for representing depth information at different levels.
[0274] The module 104 is configured to obtain a fine-grained sensing depth map that aggregates depth information from different layers based on the hierarchical sensing sampling depth map, thereby obtaining the depth map corresponding to the original image.
[0275] In this embodiment of the disclosure, the second processing module 103 includes: a depth prediction network 11 of a preset image processing model 10 and a depth perception sampling coefficient prediction module 12;
[0276] The output of the depth prediction network 11 is connected to the input of the depth perception sampling coefficient prediction module 12.
[0277] In this embodiment of the disclosure, the depth prediction network 11 includes: a shallow depth concatenated convolutional block 111, a deep depth concatenated convolutional block 112, an upsampling decoding convolutional block 113, and a depth prediction convolutional block 114.
[0278] Shallow depth concatenated convolutional block 111, deep depth concatenated convolutional block 112, upsampling decoding convolutional block 113 and depth prediction convolutional block 114 are connected in sequence; the input of shallow depth concatenated convolutional block 111 is used as the input of depth prediction network 11, and the output of depth prediction convolutional block 114 is used as the output of depth prediction network 11.
[0279] In this embodiment of the disclosure, the shallow depth-cascaded convolutional block 111 includes: a plurality of first convolutional layers having a first kernel size and a plurality of second convolutional layers having a second kernel size; the plurality of first convolutional layers and the plurality of second convolutional layers are stacked sequentially;
[0280] In this configuration, the first kernel size is larger than the second kernel size; the stride of the first convolutional layer at the front of the multiple first convolutional layers is the first step size, and the stride of the other first convolutional layers and second convolutional layers excluding the first part of the first convolutional layers is the second step size, with the first step size being larger than the second step size.
[0281] In this embodiment of the disclosure, the deep convolutional block 112 includes: a plurality of third convolutional layers with a third kernel size and a first activation function; the plurality of third convolutional layers are stacked sequentially, and the last third convolutional layer is connected to the first activation function;
[0282] The last third convolutional layer has the first dilation rate.
[0283] In this embodiment of the disclosure, the upsampling decoding convolutional block 113 may include: a plurality of fourth convolutional layers having a fourth kernel size; the plurality of fourth convolutional layers are stacked sequentially;
[0284] And / or,
[0285] The depth prediction convolutional block 114 includes at least one fifth convolutional layer with a fifth kernel size.
[0286] In this embodiment of the disclosure, the depth-sensing sampling coefficient prediction module 12 includes: a shallow sampling concatenated convolutional block 121, a deep sampling concatenated convolutional block 122, and a sampling coefficient output module 123;
[0287] Shallow sampling concatenated convolutional block 121, deep sampling concatenated convolutional block 122, and sampling coefficient output module 123 are connected in sequence;
[0288] The output of the sampling coefficient output module 123 is used as the output of the depth perception sampling coefficient prediction module 12.
[0289] In this embodiment of the disclosure, the shallow sampling convolutional block 121 includes: a plurality of sixth convolutional layers with a sixth kernel size and a second activation function; the plurality of sixth convolutional layers are stacked sequentially, and the last sixth convolutional layer is connected to the second activation function;
[0290] In this case, the stride of the last sixth convolutional layer is smaller than the stride of the other sixth convolutional layers among the multiple sixth convolutional layers except for the last sixth convolutional layer.
[0291] In this embodiment of the disclosure, the deep sampling convolutional block 122 includes: multiple seventh convolutional layers with a seventh kernel size and a third activation function; the multiple seventh convolutional layers are stacked sequentially, and the last seventh convolutional layer is connected to the third activation function;
[0292] The last seventh convolutional layer has a second dilation rate.
[0293] In this embodiment of the disclosure, the sampling coefficient output module 123 includes at least one eighth convolutional layer having an eighth kernel size.
[0294] In this embodiment of the disclosure, the obtaining module 104 includes: a deep fine-grained perception network 13 of a preset image processing model 10;
[0295] The outputs of the depth prediction network 11 and the depth sensing sampling coefficient prediction module 12 are both connected to the input of the depth fine-grained sensing network 13.
[0296] In this embodiment of the disclosure, the deep fine-grained sensing network 13 includes: a multi-layer sensing cascaded convolutional block 131, a fine upsampling decoding convolutional block 132, and a fine-grained depth prediction convolutional block 133;
[0297] The multilayer sensing convolutional block 131, the fine-grained upsampling decoding convolutional block 132, and the fine-grained depth prediction convolutional block 133 are connected in sequence.
[0298] In this embodiment of the disclosure, the multilayer sensing concatenated convolutional block 131 may include: a shallow sensing concatenated convolutional block 1311, a deep sensing concatenated convolutional block 1312, and a thinning layer sensing concatenated convolutional block 1313.
[0299] Shallow perception concatenated convolutional block 1311, deep perception concatenated convolutional block 1312, and thinning layer perception concatenated convolutional block 1313 are connected in sequence; the thinning layer perception concatenated convolutional block 1313 is connected to the thinning upsampling decoding convolutional block 132.
[0300] The input of the shallow-sensory cascaded convolutional block 1311 is used as the input of the multi-sensory cascaded convolutional block 131.
[0301] In this embodiment of the disclosure, the shallow-sensing cascaded convolutional block 1311 includes: multiple ninth convolutional layers with a ninth kernel size and a fourth activation function; the multiple ninth convolutional layers are stacked sequentially, and the last ninth convolutional layer is connected to the fourth activation function;
[0302] Among the multiple ninth convolutional layers, the stride of the first part of the ninth convolutional layers is the third stride, and the stride of the other ninth convolutional layers (excluding the first part of the ninth convolutional layers) is the fourth stride, with the third stride being greater than the fourth stride.
[0303] In this embodiment of the disclosure, the deep-perception cascaded convolutional block 1312 includes: multiple tenth convolutional layers with a tenth kernel size and a fifth activation function; the multiple tenth convolutional layers are stacked sequentially, and the last tenth convolutional layer is connected to the fifth activation function;
[0304] Among them, the stride of the first part of the 10th convolutional layers is the fifth stride, and the stride of the other 10th convolutional layers besides the first part is the sixth stride, with the fifth stride being greater than the sixth stride.
[0305] The last tenth convolutional layer has a third dilation rate.
[0306] In this embodiment of the disclosure, the thinning layer sensing cascaded convolutional block 1313 includes: a plurality of eleventh convolutional layers with an eleventh kernel size; the plurality of eleventh convolutional layers are stacked sequentially, and the last eleventh convolutional layer has a fourth dilation rate.
[0307] In this embodiment of the disclosure, the refined upsampling decoding convolutional block 132 includes: a plurality of twelfth convolutional layers having a twelfth kernel size; the plurality of twelfth convolutional layers are stacked sequentially;
[0308] And / or,
[0309] The fine-grained depth prediction convolutional block 133 includes: at least one thirteenth convolutional layer with a thirteenth kernel size.
[0310] In this disclosure, any of the embodiments in the foregoing method embodiments are applicable to this device embodiment, and will not be described in detail here.
[0311] This disclosure also provides an electronic device 200, such as... Figure 16 As shown, the electronic device 200 includes:
[0312] One or more processors 201;
[0313] The memory 202 stores one or more programs, which, when executed by the one or more processors 201, enable the one or more processors to implement the image processing method.
[0314] One or more input / output I / O interfaces 203 are connected between the processor 201 and the memory 202 and configured to enable information exchange between the processor 201 and the memory 202.
[0315] Among them, processor 201 is a device with data processing capabilities, including but not limited to central processing unit (CPU); memory 202 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH); I / O interface (read-write interface) 203 is connected between processor 201 and memory 202, and can realize information interaction between processor 201 and memory 202, including but not limited to data bus (Bus).
[0316] In some embodiments, the processor 201, memory 202, and I / O interface 203 are interconnected via bus 204, and thus connected to other components of the computing device.
[0317] This disclosure also provides a computer program product, which includes a computer program that, when executed by a processor, implements the image processing method described above.
[0318] Those skilled in the art will understand that all or some of the functional modules / units disclosed above can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0319] In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be executed by several physical components working together.
[0320] Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit (CPU), digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technique for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH) or other disk storage; read-only optical disc (CD-ROM), digital versatile disc (DVD) or other optical disc storage; magnetic cartridges, magnetic tapes, disk storage or other magnetic storage; and any other media that can be used to store desired information and can be accessed by a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0321] This disclosure has disclosed exemplary embodiments, and although specific terminology has been used, it is for general illustrative purposes only and should not be construed as limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. An image processing method, characterized in that, The method includes: Acquire the raw images captured by the monocular camera; The original image is subjected to preset processing to obtain a processed image with a different resolution than the original image, and the processed image and the original image constitute an image set; Multi-scale fusion and multi-level feature extraction are performed on the image information contained in the image set to obtain a hierarchical perceptual sampling depth map for representing depth information at different levels; Based on the hierarchical sensing sampling depth map, a fine-grained sensing depth map that aggregates depth information from different layers is obtained, thus obtaining the depth map corresponding to the original image.
2. The image processing method according to claim 1, characterized in that, The step of performing multi-scale fusion and multi-level feature extraction on the image information contained in the image set to obtain a hierarchical perceptual sampling depth map for representing depth information at different levels includes: The original image is input into the depth estimation network of a preset image processing model to obtain the estimated depth map corresponding to the original image. The estimated depth map and the image set are input into the depth perception sampling coefficient prediction module of the preset image processing model to obtain depth perception sampling coefficients at different levels, which are used as hierarchical depth perception sampling coefficients. The estimated depth map is subjected to adaptive sampling processing using the hierarchical depth sensing sampling coefficients to obtain hierarchical sensing sampling depth maps that represent depth information at different levels, thus obtaining a set of hierarchical sensing sampling depth maps.
3. The image processing method according to claim 2, characterized in that, The process of obtaining a fine-grained sensing depth map that aggregates depth information from different layers based on the hierarchical sensing sampling depth map includes: The set of hierarchical sensing sampling depth maps and the estimated depth map are input into the depth fine-grained sensing network of the preset image processing model. The depth fine-grained sensing network aggregates depth information from different layers of the hierarchical sensing sampling depth maps to obtain the fine-grained sensing depth map.
4. The image processing method according to claim 2, characterized in that, The depth prediction network includes: shallow depth concatenated convolutional blocks, deep depth concatenated convolutional blocks, upsampling decoding convolutional blocks, and depth prediction convolutional blocks; The step of inputting the original image into a preset depth estimation network to obtain a predicted depth map corresponding to the original image includes: The original image is input into the shallow depth concatenated convolutional block to obtain shallow depth features; The shallow depth features are input into the deep depth concatenated convolutional block to obtain deep depth features; The deep depth features and shallow depth features are input into the upsampled decoding convolutional block to obtain the decoding depth features; The decoded depth features are input into the depth prediction convolution block to obtain the estimated depth map corresponding to the original image.
5. The image processing method according to claim 2, characterized in that, The depth-sensing sampling coefficient prediction module includes: a shallow sampling concatenated convolutional block, a deep sampling concatenated convolutional block, and a sampling coefficient output module; The step of inputting the estimated depth map and the image set into a preset depth-sensing sampling coefficient prediction module to obtain depth-sensing sampling coefficients at different levels, as hierarchical depth-sensing sampling coefficients, includes: The estimated depth map and the processed image are input into the shallow sampling concatenated convolutional block to obtain depth-aware sampling shallow features. The shallow features of the depth-sensing sampling are input into the concatenated convolutional block of the deep sampling to obtain the deep features of the depth-sensing sampling. The deep features of the depth-sensing sampling are input into the sampling coefficient output module to obtain the depth-sensing sampling weight map; Based on the numerical values of multiple channels along the spatial dimension of the depth-sensing sampling weight map, the depth-sensing sampling coefficients of different levels of the depth map corresponding to the original image are determined, and these coefficients are used as the hierarchical depth-sensing sampling coefficients.
6. The image processing method according to claim 5, characterized in that, The step of determining the depth-sensing sampling coefficients for different levels of the depth map corresponding to the original image based on the numerical values of multiple channels along the spatial dimension of the depth-sensing sampling weight map, as the hierarchical depth-sensing sampling coefficients, includes: Calculate the average value of each channel along the spatial dimension of all depth-sensing sampling weight maps containing the multiple channels respectively, and take two average values from the average values corresponding to the multiple channels each time, and use them as the depth smoothing coefficient and position smoothing coefficient of a layer of sampling respectively. The depth smoothing coefficient and position smoothing coefficient obtained from the multi-layer sampling are used as the hierarchical depth-sensing sampling coefficients.
7. The image processing method according to claim 6, characterized in that, The plurality of channels includes four channels; the calculation of the numerical average value of each of the plurality of channels along the spatial dimension involves taking two average values from all the average values corresponding to the plurality of channels each time, and using them as the depth smoothing coefficient and position smoothing coefficient of a layer of sampling, respectively, including: The average value of the first channel among the four channels along the spatial dimension is used as the first layer sampling depth smoothing coefficient of the depth map; The average value of the second channel along the spatial dimension among the four channels is used as the first layer sampling position smoothing coefficient of the depth map; The average value of the third channel among the four channels along the spatial dimension is used as the second layer sampling depth smoothing coefficient of the depth map; The average value of the fourth channel along the spatial dimension is used as the second layer sampling position smoothing coefficient of the depth map.
8. The image processing method according to claim 2, characterized in that, The hierarchical depth-sensing sampling coefficients include: a depth smoothing coefficient and a position smoothing coefficient corresponding to each layer of sampling; the step of using the hierarchical depth-sensing sampling coefficients to perform adaptive depth-sensing sampling processing on the estimated depth map to obtain hierarchical sensing sampling depth maps for representing depth information at different levels, and obtaining a set of hierarchical sensing sampling depth maps, includes: using the depth smoothing coefficient and the position smoothing coefficient corresponding to each layer of sampling to perform adaptive depth-sensing sampling processing on the estimated depth map to obtain a depth map after sampling at the corresponding layer; The estimated depth map and the multi-layer sampled depth map constitute the hierarchical sensing sampling depth map, and a set of hierarchical sensing sampling depth maps is obtained.
9. The image processing method according to claim 8, characterized in that, The depth smoothing coefficient includes: a first-layer sampling depth smoothing coefficient and a second-layer sampling depth smoothing coefficient; the position smoothing coefficient includes: a first-layer sampling position smoothing coefficient and a second-layer sampling position smoothing coefficient; The step of performing depth-aware adaptive sampling processing on the estimated depth map using the depth smoothing coefficient and the position smoothing coefficient corresponding to each sampling layer to obtain the depth map after sampling at the corresponding layer includes: The estimated depth map is subjected to adaptive sampling processing for first-layer depth perception using the first-layer sampling depth smoothing coefficient and the first-layer sampling position smoothing coefficient to obtain the depth map after first-layer sampling. The estimated depth map is subjected to adaptive sampling processing for second-layer depth perception using the second-layer sampling depth smoothing coefficient and the second-layer sampling position smoothing coefficient to obtain the second-layer sampled depth map.
10. The image processing method according to claim 3, characterized in that, The deep fine-grained sensing network includes: a multi-layer sensing cascaded convolutional block, a fine upsampling decoding convolutional block, and a fine-grained depth prediction convolutional block; The step of inputting the set of hierarchical sensing sampling depth maps and the estimated depth map into a preset fine-grained depth sensing network, and having the fine-grained depth sensing network aggregate depth information from different layers of the hierarchical sensing sampling depth maps to obtain the fine-grained sensing depth map, includes: The estimated depth map and the multi-layer sampled depth map from the set of hierarchical sensing sampling depth maps are respectively input into the multi-layer sensing convolutional block, and the output of each layer sensing convolutional block is used as the input of the next layer sensing convolutional block to obtain the features output by each layer sensing convolutional block. Obtain all features output by the multilayer perceptron convolutional block, and input all features into the refined upsampling decoding convolutional block to obtain decoded fine-grained depth features; The decoded fine-grained depth features are input into the fine-grained depth prediction convolutional block to obtain the fine-grained sensing depth map.
11. The image processing method according to claim 10, characterized in that, The multi-layer perceptual convolutional block includes: a shallow perceptual convolutional block, a deep perceptual convolutional block, and a thinning layer perceptual convolutional block; the multi-layer sampled depth map includes: a depth map after the first layer sampling and a depth map after the second layer sampling. The step of inputting the estimated depth map and the multi-layer sampled depth maps from the set of hierarchical sensing sampled depth maps into the multi-layer sensing convolutional block, and using the output of each layer of the sensing convolutional block as the input of the next layer of the sensing convolutional block, to obtain the features output by each layer of the sensing convolutional block includes: The estimated depth map is input into the shallow perception cascaded convolutional block to obtain shallow fine-grained depth features. The depth map sampled from the first layer and the shallow fine-grained depth features are input into the deep-sensing cascaded convolutional block to obtain the deep fine-grained depth features; The depth map sampled from the second layer and the deep fine-grained depth features are input into the perceptual cascaded convolutional block of the refinement layer to obtain the fine-grained depth features of the refinement layer.
12. The image processing method according to claim 1, characterized in that, The method further includes: Obtain the preset image processing model; The image processing model is used to perform multi-scale fusion and multi-level feature extraction on the image information contained in the image set to obtain a hierarchical perceptual sampling depth map that represents depth information at different levels; based on the hierarchical perceptual sampling depth map, a fine-grained perceptual depth map that aggregates depth information at different levels is obtained to obtain the depth map corresponding to the original image. The process of obtaining the preset image processing model includes: Acquire the target sample image; The parameters of the preset neural network are continuously modified based on the principle of minimizing the difference between the target sample image and the preset original sample image. The image processing model is obtained when the difference between the target sample image and the preset original sample image meets the preset requirements.
13. The image processing method according to claim 12, characterized in that, The acquisition of the target sample image includes: Based on the preset sample original image, the sample estimated depth map and the fine-grained sensing depth map, a first two-dimensional image coordinate is generated on the sample original image. The first two-dimensional image coordinates are transformed to the coordinate system of the original pose transformation sample image corresponding to the original sample image to obtain the second two-dimensional image coordinates; The original image of the pose transformation sample is sampled based on the second two-dimensional image coordinates to obtain the synthesized target sample image.
14. The image processing method according to claim 12, characterized in that, The preset image processing model includes: a depth prediction network, a depth-aware sampling coefficient prediction module, and a depth fine-grained perception network; training the image processing model includes: The depth prediction network is trained separately by fixing the weight coefficients of the depth sensing sampling coefficient prediction module and the depth fine-grained sensing network. When the number of training steps reaches the first preset value, the weight coefficients of the depth prediction network are fixed, the fixed weight coefficients of the depth perception sampling coefficient prediction module and the depth fine-grained perception network are canceled, and the training of the depth perception sampling coefficient prediction module and the depth fine-grained perception network begins. When the number of training steps reaches the second preset value, the fixation of the weight coefficients of the depth prediction network is cancelled, and the depth prediction network, the depth perception sampling coefficient prediction module, and the depth fine-grained perception network are trained together until the training ends; wherein, the second preset value is greater than the first preset value.
15. An image processing apparatus, characterized in that, include: The acquisition module is configured to acquire the raw images captured by the monocular camera; The first processing module is configured to perform preset processing on the original image to obtain a processed image with a different resolution than the original image, and the processed image and the original image constitute an image set; The second processing module is configured to perform multi-scale fusion and multi-level feature extraction on the image information contained in the image set to obtain a hierarchical perceptual sampling depth map for representing depth information at different levels. The acquisition module is configured to acquire a fine-grained sensing depth map that aggregates depth information from different layers based on the hierarchical sensing sampling depth map, thereby obtaining the depth map corresponding to the original image.
16. The image processing apparatus according to claim 15, characterized in that, The second processing module includes: a depth prediction network for a preset image processing model and a depth-aware sampling coefficient prediction module; The output of the depth prediction network is connected to the input of the depth perception sampling coefficient prediction module.
17. The image processing apparatus according to claim 16, characterized in that, The acquisition module includes: a deep fine-grained perception network of the preset image processing model; The output of the depth prediction network and the output of the depth sensing sampling coefficient prediction module are both connected to the input of the depth fine-grained sensing network.
18. The image processing apparatus according to claim 16, characterized in that, The depth prediction network includes: shallow depth concatenated convolutional blocks, deep depth concatenated convolutional blocks, upsampling decoding convolutional blocks, and depth prediction convolutional blocks; The shallow depth concatenated convolutional block, the deep depth concatenated convolutional block, the upsampling decoding convolutional block, and the depth prediction convolutional block are connected in sequence; The output of the depth prediction convolutional block is used as the output of the depth estimation network.
19. The image processing apparatus according to claim 16, characterized in that, The depth-sensing sampling coefficient prediction module includes: a shallow sampling concatenated convolutional block, a deep sampling concatenated convolutional block, and a sampling coefficient output module; The shallow sampling concatenated convolutional block, the deep sampling concatenated convolutional block, and the sampling coefficient output module are connected in sequence; The output of the sampling coefficient output module is used as the output of the depth perception sampling coefficient prediction module.
20. The image processing apparatus according to claim 17, characterized in that, The deep fine-grained sensing network includes: a multi-layer sensing cascaded convolutional block, a fine upsampling decoding convolutional block, and a fine-grained depth prediction convolutional block; The multilayer sensing concatenated convolutional block, the refined upsampling decoding convolutional block, and the fine-grained depth prediction convolutional block are connected in sequence; The input of the multilayer perceptron convolutional block is used as the input of the deep fine-grained perceptron network.
21. The image processing apparatus according to claim 20, characterized in that, The multilayer sensing concatenated convolutional block includes: shallow sensing concatenated convolutional block, deep sensing concatenated convolutional block, and refined layer sensing concatenated convolutional block; The shallow sensing concatenated convolutional block, the deep sensing concatenated convolutional block, and the refined sensing concatenated convolutional block are connected in sequence; the refined sensing concatenated convolutional block is connected to the refined upsampling decoding convolutional block. The input of the shallow-sensory cascaded convolutional block is used as the input of the multi-sensory cascaded convolutional block.
22. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory having stored one or more programs thereon, which, when executed by the one or more processors, cause the one or more processors to implement the image processing method according to any one of claims 1-14; One or more input / output (I / O) interfaces are connected between the processor and the memory and configured to enable information exchange between the processor and the memory.
23. A computer program product comprising a computer program that, when executed by a processor, implements the image processing method according to any one of claims 1-14.