Image depth estimation method, apparatus, device, medium, and computer program product

By employing an image depth estimation method, utilizing a spatiotemporal information extraction network and a depth evaluation network, combined with a spatial attention network and a loss function, the accuracy problem of binocular depth estimation in complex lighting environments is solved, achieving higher depth estimation accuracy.

CN121120740BActive Publication Date: 2026-07-24HUBEI JINGCHU HUMANOID ROBOT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUBEI JINGCHU HUMANOID ROBOT CO LTD
Filing Date
2025-08-07
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing methods for binocular depth estimation have low depth estimation accuracy in complex lighting environments.

Method used

An image depth estimation method is adopted, which uses a spatiotemporal information extraction network, a binocular stereo matching network, and a depth evaluation network. Multiple skip-connected encoder-decoder modules are used to extract and evaluate features from the left and right eye images. By combining a spatial attention network and a loss function, illumination differences are eliminated, thereby improving the depth estimation accuracy.

Benefits of technology

By learning the dynamic changes between binocular image frames, and utilizing motion information and temporal consistency in dynamic scenes, the accuracy of image depth estimation can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120740B_ABST
    Figure CN121120740B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image processing, and provides a kind of image depth estimation method, device, equipment, medium and computer program product, left and right eye images including left eye adjacent frame image and correction binocular image are input into depth estimation model including space-time information extraction network, binocular stereo matching network and depth evaluation network to obtain depth map;Space-time information extraction network is used to carry out space-time feature extraction to left eye adjacent frame image, obtain binocular pose transformation information and image change information between adjacent frame image;Binocular stereo matching network is used to carry out multi-layer feature extraction to correction binocular image, obtain multi-level parallax map;Depth evaluation network is used to evaluate binocular camera pose transformation information, image target change information and multi-level parallax map, obtain the depth map of left and right eye images.The present application improves the precision of image depth evaluation by learning the dynamic change rule between binocular image frames, using motion information and time sequence consistency in dynamic scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, and more particularly to an image depth estimation method, apparatus, device, medium, and computer program product. Background Technology

[0002] With the development of automation technology, grasping robots have been widely used in automated production scenarios (such as material handling and component assembly). However, the lighting environment in natural production scenarios is quite complex, and existing solutions for binocular depth estimation only consider stereo matching of a single frame image. This limitation leads to low depth estimation accuracy in existing binocular depth estimation solutions. Summary of the Invention

[0003] This invention provides an image depth estimation method, apparatus, device, medium, and computer program product to solve the problem of low depth estimation accuracy in existing binocular depth estimation schemes, thereby improving the accuracy of image depth estimation.

[0004] This invention provides an image depth estimation method, comprising the following steps: The left and right eye images are input into the depth estimation model to obtain the depth map output by the depth estimation model; the left and right eye images include the left eye's adjacent frame image and the corrected binocular image; The depth estimation model includes a spatiotemporal information extraction network, a binocular stereo matching network, and a depth evaluation network. The spatiotemporal information extraction network includes multiple skip-connected encoder-decoder modules, each of which is used for: Spatiotemporal features are extracted from the adjacent frames of the left eye to obtain the binocular camera pose transformation information and image target change information between adjacent frames; The binocular stereo matching network is used to perform multi-level feature extraction on the corrected binocular image to obtain a multi-level disparity map; The depth evaluation network is used to evaluate the pose transformation information of the binocular camera, the image target change information, and the multi-level disparity map to obtain the depth maps of the left and right eye images.

[0005] According to an image depth estimation method provided by the present invention, the encoder-decoder module includes an encoder, a spatial attention network, and a decoder; The encoder is used to obtain the initial encoding features of the left-eye adjacent frame image; The spatial attention network is used to extract texture features, enhance edge perception, and suppress reflective areas from the initial encoded features to obtain multi-scale attention features. The decoder is used to decode the multi-scale attention features to obtain the binocular camera pose transformation information and image target change information between adjacent frame images.

[0006] According to an image depth estimation method provided by the present invention, the encoder-decoder module includes a pose transformation encoder-decoder, an optical flow encoder-decoder, and an appearance flow encoder-decoder; the stereo camera pose transformation information is a stereo pose transformation matrix; the image target change information includes an optical flow map and an appearance flow map; wherein: The decoder of the pose transformation codec is used to perform dimensionality compression and feature refinement on the features output by the encoder of the pose transformation codec to obtain the binocular pose transformation matrix. The decoder of the optical flow codec is used to decode the multi-scale attention features output by the spatial attention network of the optical flow codec to obtain the optical flow map; The appearance stream codec is used to extract features from the adjacent frame images of the left eye and the optical flow map to obtain the appearance flow map.

[0007] According to an image depth estimation method provided by the present invention, the spatial attention network includes a texture feature extraction network, an edge awareness enhancement network, a reflective region suppression network, a channel stitching module, and a residual connection network, wherein: The texture feature extraction network is used to extract texture features from the initial encoded features to obtain image texture features; The edge-aware enhancement network is used to extract edge features from the initial encoded features to obtain image edge features; The reflective region suppression network is used to extract reflective suppression features from the initial encoded features to obtain image reflective suppression features; The channel stitching module is used to perform channel-dimensional stitching processing on the image texture features, the image edge features, and the image reflection suppression features to obtain an attention map; The residual connection network is used to perform residual connections between the attention map and the initial encoded features to obtain multi-scale attention features.

[0008] According to an image depth estimation method provided by the present invention, the step of extracting features from the adjacent frame images of the left eye and the optical flow map to obtain the appearance flow map includes: Obtain the brightness change value of each pixel in the adjacent frame image of the left eye; Based on the brightness change value and the geometric deformation field, the luminance change value of each pixel is determined; the geometric deformation field is determined based on the optical flow map. The appearance flow map is determined based on the luminance variation values ​​of each pixel.

[0009] According to an image depth estimation method provided by the present invention, the binocular stereo matching network includes a multi-layer feature extraction module, a stereo matching module, and a self-attention module, wherein: The multi-layer feature extraction module is used to extract multi-layer features from the corrected binocular image; The stereo matching module is used to determine the feature similarity matrix based on the multi-layer features; The self-attention module is used to perform region propagation processing on the multi-level disparity map to obtain a multi-level disparity map; the multi-level disparity map is generated based on the feature similarity matrix.

[0010] The present invention also provides an image depth estimation device, comprising the following modules: The image depth evaluation module is used to input the left and right eye images into the depth estimation model to obtain the depth map output by the depth estimation model; the left and right eye images include the left eye adjacent frame image and the corrected binocular image; The depth estimation model includes a spatiotemporal information extraction network, a binocular stereo matching network, and a depth evaluation network. The spatiotemporal information extraction network includes multiple skip-connected encoder-decoder modules, each of which is used for: Spatiotemporal features are extracted from the adjacent frames of the left eye to obtain the binocular camera pose transformation information and image target change information between adjacent frames; The binocular stereo matching network is used to perform multi-level feature extraction on the corrected binocular image to obtain a multi-level disparity map; The depth evaluation network is used to evaluate the pose transformation information of the binocular camera, the image target change information, and the multi-level disparity map to obtain the depth maps of the left and right eye images.

[0011] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the image depth estimation method as described above.

[0012] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image depth estimation method as described above.

[0013] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the image depth estimation method as described above.

[0014] The image depth estimation method, apparatus, device, medium, and computer program product provided by this invention extracts spatiotemporal features from adjacent frames of the left-eye image through multiple skip-connected codec modules to obtain spatiotemporal information between adjacent frames, including binocular camera pose transformation information and image target change information. Furthermore, by extracting and correcting multi-layer features from the binocular images, and utilizing the structural similarity between the left and right eye images, the accurate prediction results of reliable image regions are propagated to other regions, resulting in a multi-level disparity map. A depth evaluation network extracts the spatiotemporal dynamic information learned by the network and evaluates the multi-level disparity map to obtain the depth maps of the left and right eye images. This invention improves the accuracy of image depth estimation by learning the dynamic change patterns between binocular image frames and utilizing motion information and temporal consistency in dynamic scenes. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0016] Figure 1 This is a flowchart illustrating the image depth estimation method provided by the present invention.

[0017] Figure 2 This is a schematic diagram of the network framework of the depth estimation model provided by the present invention.

[0018] Figure 3 This is a schematic diagram of the framework of the spatial attention network provided by the present invention.

[0019] Figure 4 This is a schematic diagram of the image depth estimation device provided by the present invention.

[0020] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0022] The following is combined Figures 1-5 This invention describes the image depth estimation method, apparatus, device, medium, and computer program product.

[0023] Figure 1 This is one of the flowcharts illustrating the image depth estimation method provided by the present invention, such as... Figure 1 As shown, the method includes the following: Step 100: Input the left and right eye images into the depth estimation model to obtain the depth map output by the depth estimation model; the left and right eye images include the left eye adjacent frame image and the corrected binocular image; The depth estimation model includes a spatiotemporal information extraction network, a binocular stereo matching network, and a depth evaluation network. The spatiotemporal information extraction network includes multiple skip-connected encoder-decoder modules, each of which is used for: Spatiotemporal features are extracted from the adjacent frames of the left eye to obtain the binocular camera pose transformation information and image target change information between adjacent frames; The binocular stereo matching network is used to perform multi-level feature extraction on the corrected binocular image to obtain a multi-level disparity map; The depth evaluation network is used to evaluate the pose transformation information of the binocular camera, the image target change information, and the multi-level disparity map to obtain the depth maps of the left and right eye images.

[0024] The image depth estimation method provided by this invention estimates the depth information of the left and right eye images by training a depth estimation model, and obtains the depth map of the left and right eye images (a single-channel image that records the distance information from the object corresponding to each pixel in the application scenario to the camera).

[0025] like Figure 2 As shown, the depth estimation model provided by this invention includes at least a spatiotemporal information extraction network, a stereo matching network, and a depth evaluation network. The spatiotemporal information extraction network consists of multiple skip-connected codecs (i.e., codec modules). Taking three codec modules as an example, adjacent frames of the left camera are input into the spatiotemporal information extraction network. Through one of the codec modules in the spatiotemporal information extraction network, the stereo camera pose transformation information between adjacent frames is calculated and output (used to describe the relative pose relationship between the left and right camera coordinate systems).

[0026] By extracting two other codec modules in the spatiotemporal information network, target change information (i.e. image target change information) between adjacent frames is calculated and output. The image target change information reflects the dynamic changes of pixels between adjacent frames and the differences of the target between adjacent frames.

[0027] To eliminate illumination differences (e.g., brightness and contrast) between adjacent frames during depth estimation and prevent the depth estimation model from being interfered with by non-geometric information, a loss function resistant to inconsistencies in light intensity is constructed to ensure that the depth estimation model focuses on geometric relationships (e.g., disparity and motion) rather than illumination changes. Then, the right-eye image is warped to the left-eye image to generate a pseudo-left-eye image, and the previous frame is warped to the current frame to generate a pseudo-current-frame image. These warped composite images (pseudo-left-eye and pseudo-current-frame images) are used as self-supervised signals. The image dynamic change patterns (implicitly) learned by the spatiotemporal information extraction network are distilled into the depth evaluation network. This allows the depth evaluation network to evaluate the stereo camera pose transformation information, image target change information, and multi-level disparity maps, resulting in depth maps for both the left and right-eye images.

[0028] This embodiment extracts spatiotemporal features from adjacent frames of the left-eye image using multiple skip-connected codec modules to obtain spatiotemporal information between adjacent frames, including binocular camera pose transformation information and image target change information. Furthermore, by extracting multi-layer features from the corrected binocular images and utilizing the structural similarity between the left and right eye images, accurate prediction results for reliable image regions are propagated to other regions, resulting in a multi-level disparity map. A depth evaluation network extracts the spatiotemporal dynamic information learned by the network and evaluates the multi-level disparity map to obtain depth maps for both the left and right eye images. This invention improves the accuracy of image depth evaluation by learning the dynamic change patterns between binocular image frames and utilizing motion information and temporal consistency in dynamic scenes.

[0029] In one embodiment, the codec module includes an encoder, a spatial attention network, and a decoder; The encoder is used to obtain the initial encoding features of the left-eye adjacent frame image; The spatial attention network is used to extract texture features, enhance edge perception, and suppress reflective areas from the initial encoded features to obtain multi-scale attention features. The decoder is used to decode the multi-scale attention features to obtain the binocular camera pose transformation information and image target change information between adjacent frame images.

[0030] Specifically, such as Figure 2As shown, the spatiotemporal information extraction network consists of an encoder-decoder module with multiple skip connections. The encoder can be a standard ResNet18 (a lightweight model in Residual Networks, which addresses the degradation problem in deep convolutional neural networks (CNNs) by removing the last connected layer and introducing residual blocks and skip connections) with the input channels expanded to N×3 (N being the number of adjacent input frames, which can be set to 2). The weights of the first convolutional kernel of the ResNet18 are initialized using a weighted average to ensure effective fusion of temporal information. Adding a spatial attention mechanism (i.e., a spatial attention network) after the ResNet18 allows the spatial attention network to better extract features from adjacent frames.

[0031] The initial encoded features output by the encoder and the three-level features (texture features, edge features, and reflection suppression features) extracted by the spatial attention network are combined to form a four-level feature map, thus forming a multi-scale feature pyramid (i.e., multi-scale attention features). This provides more spatial context information for the subsequent decoder to decode the multi-scale attention features and obtain the stereo camera pose transformation information and image target change information between adjacent frames.

[0032] This embodiment extracts multi-level features from adjacent frame images through a skip-connected encoding / decoding module to provide more spatiotemporal information.

[0033] In one embodiment, the codec module includes a pose transformation codec, an optical flow codec, and an appearance flow codec; the binocular camera pose transformation information is a binocular pose transformation matrix; the image target change information includes an optical flow map and an appearance flow map; wherein: The decoder of the pose transformation codec is used to perform dimensionality compression and feature refinement on the features output by the encoder of the pose transformation codec to obtain the binocular pose transformation matrix. The decoder of the optical flow codec is used to decode the multi-scale attention features output by the spatial attention network of the optical flow codec to obtain the optical flow map; The appearance stream codec is used to extract features from the adjacent frame images of the left eye and the optical flow map to obtain the appearance flow map.

[0034] Specifically, the codec module in the spatiotemporal information extraction network includes at least a codec for generating stereo camera pose transformation information between adjacent frames (i.e., a pose transformation codec, including a pose transformation encoder and a pose transformation decoder), a codec for generating optical flow graphs between adjacent frames (i.e., an optical flow codec, including an optical flow encoder and an optical flow decoder), and a codec for generating appearance flow graphs between adjacent frames (i.e., an appearance flow codec, including an appearance flow encoder and an appearance flow decoder). The stereo camera pose transformation information can be a stereo pose transformation matrix; the image target change information includes optical flow graphs and appearance flow graphs.

[0035] The features output by the pose transformation encoder are further processed by a spatial attention network to extract multi-scale features. The resulting multi-scale attention features are then input into the pose transformation decoder. The pose decoder first compresses the dimensionality of the multi-scale attention features, and then refines the features step by step through a three-level convolutional network to obtain multi-degree-of-freedom pose parameters. The stereo pose transformation matrix is ​​then obtained from the multi-degree-of-freedom pose parameters.

[0036] The encoding and decoding process of the optical flow encoder and decoder is similar to that of the pose transformation encoder described above. The optical flow decoder decodes the multi-scale attention features (generated by the spatial attention network in the optical flow encoder and decoder module) to obtain the optical flow map.

[0037] Depend on Figure 2 It can be seen that the optical flow graph output by the optical flow encoder-decoder module and the left adjacent frame image (from the input spatiotemporal information extraction network) are jointly input into the appearance flow encoder so that the appearance flow decoder outputs the appearance flow graph.

[0038] This embodiment uses a skip-connected codec module to fuse multi-scale coding features of adjacent frame images, thereby obtaining more spatiotemporal information from adjacent frame images.

[0039] In one embodiment, the spatial attention network includes a texture feature extraction network, an edge awareness enhancement network, a reflective region suppression network, a channel stitching module, and a residual connection network, wherein: The texture feature extraction network is used to extract texture features from the initial encoded features to obtain image texture features; The edge-aware enhancement network is used to extract edge features from the initial encoded features to obtain image edge features; The reflective region suppression network is used to extract reflective suppression features from the initial encoded features to obtain image reflective suppression features; The channel stitching module is used to perform channel-dimensional stitching processing on the image texture features, the image edge features, and the image reflection suppression features to obtain an attention map; The residual connection network is used to perform residual connections between the attention map and the initial encoded features to obtain multi-scale attention features.

[0040] Specifically, such as Figure 3 As shown, the spatial attention network added after ResNet18 includes at least a texture feature extraction network, an edge awareness enhancement network, a reflective region suppression network, a channel concatenation module, and a residual connection network. Among them, conv1×1 is a regular 1×1 convolution (using a 1×1 kernel / filter to perform a full-channel linear weighted summation on each pixel / spatial location of the input feature map); conv3×3 is a depthwise separable convolution (including depthwise convolution for processing spatial information and pointwise convolution for processing channel information. Depthwise separable convolution decouples spatial and channel information by performing a step-by-step operation of spatial information first and then channel information, thereby reducing the computational cost); Rectified Linear Unit (ReLU) is one type of activation function; Sigmoid is another type of activation function.

[0041] The process by which spatial attention networks obtain multi-scale attention features is shown in Equation 1, where, Multi-scale attention features; The kernel size is Depth-separable convolution; for Figure 3 The original features of the input spatial attention network, i.e., the initial encoded features; For attention graphs; For element-wise multiplication; This is an element-wise addition.

[0042] (1) (2) in, The generation incorporates multi-source feature information (image texture features, image edge features, and image reflection suppression features), as shown in Equation 2, where, Use the Sigmoid activation function; Image texture features; Image edge features; Features for suppressing image reflection; This is a concatenation operation at the channel level. Among them, Figure 2 Sobel Kernel is a tool used in computer vision for edge detection. It calculates the image gradient (grayscale change rate) through convolution operations to quickly locate regions in the image where grayscale changes abruptly, i.e., edges.

[0043] This embodiment uses multi-scale texture feature extraction, edge awareness enhancement, and reflective area suppression to accurately allocate attention to key regions in adjacent frame images.

[0044] In one embodiment, the step of extracting features from the adjacent frame images of the left eye and the optical flow map to obtain the appearance flow map includes: Obtain the brightness change value of each pixel in the adjacent frame image of the left eye; Based on the brightness change value and the geometric deformation field, the luminance change value of each pixel is determined; the geometric deformation field is determined based on the optical flow map. The appearance flow map is determined based on the luminance variation values ​​of each pixel.

[0045] Specifically, the process of the appearance stream codec generating the appearance stream graph is shown in Equation 3, where, for Pixels in adjacent frames at different times Brightness value; For geometric deformation fields, describing the pixel from arrive The displacement is determined by the optical flow map. The resulting appearance flow map specifically includes pixels. The change in luminance from the previous frame to the current frame is caused by changes in illumination and non-Lambertian reflection (the phenomenon of light reflection from the surface of a real object, where the intensity of reflected light changes with the viewing direction / angle of view or the angle of incidence). .

[0046] (3) like Figure 2 As shown, the inputs to the pose transformation codec and the optical flow codec are... and ,as well as and Two sets of adjacent left-eye frames, among which, and The time frame number of the image; The left eye is represented by the optical flow map output by the optical flow codec. The pixel displacement is represented by a two-dimensional vector and superimposed on the coordinates of adjacent frames (pixels) to form a sampling grid. Then, bilinear interpolation resampling is used to synthesize a registration image (the image of the adjacent frame of the left eye and the optical flow map). The registration image is input into the appearance flow codec to generate the appearance flow map.

[0047] This embodiment generates an appearance flow graph by analyzing the changes in pixels in adjacent frame images.

[0048] In one embodiment, the binocular stereo matching network includes a multi-layer feature extraction module, a stereo matching module, and a self-attention module, wherein: The multi-layer feature extraction module is used to extract multi-layer features from the corrected binocular image; The stereo matching module is used to determine the feature similarity matrix based on the multi-layer features; The self-attention module is used to perform region propagation processing on the multi-level disparity map to obtain a multi-level disparity map; the multi-level disparity map is generated based on the feature similarity matrix.

[0049] Specifically, the stereo-corrected left and right eye images are adjusted to a specified size and used as input to the binocular stereo matching network. First, a convolutional neural network (CNN) is used to extract global features. The extracted global features are then downsampled multiple times, and positional encoding is added to the compressed pixel features after downsampling.

[0050] Then, the features with added positional encoding are input into a multi-layered Transformer (feature extraction unit, i.e., multi-layer feature extraction module) to extract multi-layer features. A parallel softmax stereo matching layer (i.e., stereo matching module) is directly connected after each Transformer to calculate the feature similarity matrix, establishing pixel-level correspondences to avoid limitations imposed by a preset disparity range. Based on the feature similarity matrix between the calibrated binocular images, three disparity maps with different hierarchical characteristics are generated.

[0051] The three disparity maps are input into a weight-sharing self-attention module to jointly optimize the predictions at multiple levels. By utilizing the structural similarity of the binocular images themselves, the more accurate prediction results of reliable regions (regions with little change) are propagated to occluded regions (regions with significant change).

[0052] This embodiment analyzes the structural similarity of the binocular images at the same time using a binocular stereo matching network to obtain a multi-level disparity map for generating a depth map.

[0053] Specifically, the depth estimation model provided by this invention also includes a loss function part, such as... Figure 2 As shown, the loss function specifically includes reprojection perception loss. Photometric calibration loss 3D matching loss and parallax smoothing loss .

[0054] The role of reprojection perception loss is to improve the accuracy of depth estimation and camera pose estimation through multi-view geometric consistency constraints. The method of warping image frames from previous and subsequent time steps to the current image frame is shown in Equation 4, where... The current image frame; These are image frames from previous and subsequent moments; To deeply evaluate the depth map output by the network; for arrive Relative camera pose between them; This is the camera intrinsic parameter matrix; This is a bilinear sampling operation.

[0055] (4) The accuracy of depth estimation and camera pose estimation is improved by using reprojection perceptual loss (adjusting the parameters of the depth estimation model). A feature space-based perceptual loss is employed, which does not directly calculate the pixel intensity difference between adjacent image frames, but instead reprojects the distorted image frames... and corresponding The feature map is obtained by inputting it into the Transformer.

[0056] The effect of photometric inconsistency is reduced by photometric calibration loss. The left image is synthesized from the right image by calculation. Corresponding to the actual left image The pixel intensity loss and structural similarity index Measure (SSIM) loss are used to calculate the stereo matching loss. The disparity smoothing loss is used to maintain the structural integrity of the disparity map and resolve ill-conditioned regions in the image.

[0057] (5) The total loss function obtained through the above various loss functions is shown in Formula 5. Wherein, , , and These are the weights (hyperparameters) of the corresponding loss function.

[0058] The image depth estimation apparatus provided by the present invention is described below. The image depth estimation apparatus described below can be referred to in correspondence with the image depth estimation method described above.

[0059] Please refer to Figure 4 The present invention also provides an image depth estimation apparatus, comprising: The image depth evaluation module 401 is used to input the left and right eye images into the depth estimation model to obtain the depth map output by the depth estimation model; the left and right eye images include the left eye adjacent frame image and the corrected binocular image; The depth estimation model includes a spatiotemporal information extraction network, a binocular stereo matching network, and a depth evaluation network. The spatiotemporal information extraction network includes multiple skip-connected encoder-decoder modules, each of which is used for: Spatiotemporal features are extracted from the adjacent frames of the left eye to obtain the binocular camera pose transformation information and image target change information between adjacent frames; The binocular stereo matching network is used to extract multi-level features from the corrected binocular image to obtain a multi-level disparity map; The depth evaluation network is used to evaluate the pose transformation information of the binocular camera, the image target change information, and the multi-level disparity map to obtain the depth maps of the left and right eye images.

[0060] Optionally, the codec module includes an encoder, a spatial attention network, and a decoder; The encoder is used to obtain the initial encoding features of the left-eye adjacent frame image; The spatial attention network is used to extract texture features, enhance edge perception, and suppress reflective areas from the initial encoded features to obtain multi-scale attention features. The decoder is used to decode the multi-scale attention features to obtain the binocular camera pose transformation information and image target change information between adjacent frame images.

[0061] Optionally, the codec module includes a pose transformation codec, an optical flow codec, and an appearance flow codec; the stereo camera pose transformation information is a stereo pose transformation matrix; the image target change information includes an optical flow graph and an appearance flow graph; wherein: The decoder of the pose transformation codec is used to perform dimensionality compression and feature refinement on the features output by the encoder of the pose transformation codec to obtain the binocular pose transformation matrix. The decoder of the optical flow codec is used to decode the multi-scale attention features output by the spatial attention network of the optical flow codec to obtain the optical flow map; The appearance stream codec is used to extract features from the adjacent frame images of the left eye and the optical flow map to obtain the appearance flow map.

[0062] Optionally, the spatial attention network includes a texture feature extraction network, an edge awareness enhancement network, a reflective region suppression network, a channel stitching module, and a residual connection network, wherein: The texture feature extraction network is used to extract texture features from the initial encoded features to obtain image texture features; The edge-aware enhancement network is used to extract edge features from the initial encoded features to obtain image edge features; The reflective region suppression network is used to extract reflective suppression features from the initial encoded features to obtain image reflective suppression features; The channel stitching module is used to perform channel-dimensional stitching processing on the image texture features, the image edge features, and the image reflection suppression features to obtain an attention map; The residual connection network is used to perform residual connections between the attention map and the initial encoded features to obtain multi-scale attention features.

[0063] Optionally, the appearance stream codec is further configured to: Obtain the brightness change value of each pixel in the adjacent frame image of the left eye; Based on the brightness change value and the geometric deformation field, the luminance change value of each pixel is determined; the geometric deformation field is determined based on the optical flow map. The appearance flow map is determined based on the luminance variation values ​​of each pixel.

[0064] Optionally, the binocular stereo matching network includes a multi-layer feature extraction module, a stereo matching module, and a self-attention module, wherein: The multi-layer feature extraction module is used to extract multi-layer features from the corrected binocular image; The stereo matching module is used to determine the feature similarity matrix based on the multi-layer features; The self-attention module is used to perform region propagation processing on the multi-level disparity map to obtain a multi-level disparity map; the multi-level disparity map is generated based on the feature similarity matrix.

[0065] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communications bus 540. The processor 510 can call logical instructions in the memory 530 to execute an image depth estimation method, which includes: inputting left and right eye images into a depth estimation model to obtain a depth map output by the depth estimation model; the left and right eye images include adjacent frame images of the left eye and a calibrated binocular image; wherein, the depth estimation model includes a spatiotemporal information extraction network, a binocular stereo matching network, and a depth evaluation network; the spatiotemporal information extraction network includes multiple skip-connected encoder-decoder modules, each encoder-decoder module being used to: extract spatiotemporal features from the adjacent frame images of the left eye to obtain binocular camera pose transformation information and image target change information between adjacent frame images; the binocular stereo matching network is used to perform multi-layer feature extraction on the calibrated binocular image to obtain a multi-level disparity map; the depth evaluation network is used to evaluate the binocular camera pose transformation information, the image target change information, and the multi-level disparity map to obtain the depth maps of the left and right eye images.

[0066] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0067] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the image depth estimation method provided by the above methods. The method includes: inputting left and right eye images into a depth estimation model to obtain a depth map output by the depth estimation model; the left and right eye images include a left eye adjacent frame image and a corrected binocular image; wherein, the depth estimation model includes a spatiotemporal information extraction network, a binocular stereo matching network, and a depth evaluation network; the spatiotemporal information extraction network includes multiple skip-connected encoder-decoder modules, each encoder-decoder module being used to: extract spatiotemporal features from the left eye adjacent frame image to obtain binocular camera pose transformation information and image target change information between adjacent frame images; the binocular stereo matching network is used to extract multi-level features from the corrected binocular image to obtain a multi-level disparity map; the depth evaluation network is used to evaluate the binocular camera pose transformation information, the image target change information, and the multi-level disparity map to obtain the depth map of the left and right eye images.

[0068] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the image depth estimation method provided by the methods described above. The method includes: inputting left and right eye images into a depth estimation model to obtain a depth map output by the depth estimation model; the left and right eye images include adjacent frame images of the left eye and a calibrated binocular image; wherein the depth estimation model includes a spatiotemporal information extraction network, a binocular stereo matching network, and a depth evaluation network; the spatiotemporal information extraction network includes multiple skip-connected encoder-decoder modules, each encoder-decoder module being used to: extract spatiotemporal features from the adjacent frame images of the left eye to obtain binocular camera pose transformation information and image target change information between adjacent frame images; the binocular stereo matching network is used to perform multi-level feature extraction on the calibrated binocular image to obtain a multi-level disparity map; the depth evaluation network is used to evaluate the binocular camera pose transformation information, the image target change information, and the multi-level disparity map to obtain the depth maps of the left and right eye images.

[0069] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0070] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An image depth estimation method, characterized in that, include: The left and right eye images are input into the depth estimation model to obtain the depth map output by the depth estimation model; the left and right eye images include the left eye's adjacent frame image and the corrected binocular image; The depth estimation model includes a spatiotemporal information extraction network, a binocular stereo matching network, and a depth evaluation network. The spatiotemporal information extraction network includes multiple skip-connected encoder-decoder modules, each of which is used for: Spatiotemporal features are extracted from the adjacent frames of the left eye to obtain the binocular camera pose transformation information and image target change information between adjacent frames; The binocular stereo matching network is used to extract multi-level features from the corrected binocular image to obtain a multi-level disparity map; The depth evaluation network is used to evaluate the pose transformation information of the binocular camera, the image target change information, and the multi-level disparity map to obtain the depth maps of the left and right eye images; The encoder-decoder module includes an encoder, a spatial attention network, and a decoder; The encoder is used to obtain the initial encoding features of the left-eye adjacent frame image; The spatial attention network is used to extract texture features, enhance edge perception, and suppress reflective areas from the initial encoded features to obtain multi-scale attention features. The decoder is used to decode the multi-scale attention features to obtain binocular camera pose transformation information and image target change information between adjacent frame images; The codec module includes a pose transformation codec, an optical flow codec, and an appearance flow codec; the binocular camera pose transformation information is a binocular pose transformation matrix; the image target change information includes an optical flow map and an appearance flow map; wherein: The decoder of the pose transformation codec is used to perform dimensionality compression and feature refinement on the features output by the encoder of the pose transformation codec to obtain the binocular pose transformation matrix. The decoder of the optical flow codec is used to decode the multi-scale attention features output by the spatial attention network of the optical flow codec to obtain the optical flow map; The appearance stream codec is used to extract features from the adjacent frame images of the left eye and the optical flow map to obtain the appearance flow map.

2. The image depth estimation method according to claim 1, characterized in that, The spatial attention network includes a texture feature extraction network, an edge awareness enhancement network, a reflective region suppression network, a channel stitching module, and a residual connection network, wherein: The texture feature extraction network is used to extract texture features from the initial encoded features to obtain image texture features; The edge-aware enhancement network is used to extract edge features from the initial encoded features to obtain image edge features; The reflective region suppression network is used to extract reflective suppression features from the initial encoded features to obtain image reflective suppression features; The channel stitching module is used to perform channel-dimensional stitching processing on the image texture features, the image edge features, and the image reflection suppression features to obtain an attention map; The residual connection network is used to perform residual connections between the attention map and the initial encoded features to obtain multi-scale attention features.

3. The image depth estimation method according to claim 1, characterized in that, The step of extracting features from the adjacent frame images of the left eye and the optical flow map to obtain the appearance flow map includes: Obtain the brightness change value of each pixel in the adjacent frame image of the left eye; Based on the brightness change value and the geometric deformation field, the luminance change value of each pixel is determined; the geometric deformation field is determined based on the optical flow map. The appearance flow map is determined based on the luminance variation values ​​of each pixel.

4. The image depth estimation method according to claim 1, characterized in that, The binocular stereo matching network includes a multi-layer feature extraction module, a stereo matching module, and a self-attention module, wherein: The multi-layer feature extraction module is used to extract multi-layer features from the corrected binocular image; The stereo matching module is used to determine the feature similarity matrix based on the multi-layer features; The self-attention module is used to perform region propagation processing on the multi-level disparity map to obtain a multi-level disparity map; the multi-level disparity map is generated based on the feature similarity matrix.

5. An image depth estimation device, characterized in that, include: The image depth evaluation module is used to input the left and right eye images into the depth estimation model to obtain the depth map output by the depth estimation model; the left and right eye images include the left eye adjacent frame image and the corrected binocular image; The depth estimation model includes a spatiotemporal information extraction network, a binocular stereo matching network, and a depth evaluation network. The spatiotemporal information extraction network includes multiple skip-connected encoder-decoder modules, each of which is used for: Spatiotemporal features are extracted from the adjacent frames of the left eye to obtain the binocular camera pose transformation information and image target change information between adjacent frames; The binocular stereo matching network is used to extract multi-level features from the corrected binocular image to obtain a multi-level disparity map; The depth evaluation network is used to evaluate the pose transformation information of the binocular camera, the image target change information, and the multi-level disparity map to obtain the depth maps of the left and right eye images; The encoder-decoder module includes an encoder, a spatial attention network, and a decoder; The encoder is used to obtain the initial encoding features of the left-eye adjacent frame image; The spatial attention network is used to extract texture features, enhance edge perception, and suppress reflective areas from the initial encoded features to obtain multi-scale attention features. The decoder is used to decode the multi-scale attention features to obtain binocular camera pose transformation information and image target change information between adjacent frame images; The codec module includes a pose transformation codec, an optical flow codec, and an appearance flow codec; the binocular camera pose transformation information is a binocular pose transformation matrix; the image target change information includes an optical flow map and an appearance flow map; wherein: The decoder of the pose transformation codec is used to perform dimensionality compression and feature refinement on the features output by the encoder of the pose transformation codec to obtain the binocular pose transformation matrix. The decoder of the optical flow codec is used to decode the multi-scale attention features output by the spatial attention network of the optical flow codec to obtain the optical flow map; The appearance stream codec is used to extract features from the adjacent frame images of the left eye and the optical flow map to obtain the appearance flow map.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the image depth estimation method as described in any one of claims 1 to 4.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the image depth estimation method as described in any one of claims 1 to 4.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the image depth estimation method as described in any one of claims 1 to 4.