Depth estimation method, electronic device, storage medium and program product
By performing position encoding of image data and feature fusion of echo data, the accuracy and robustness of existing depth estimation methods in noise and complex environments are solved, and a more efficient depth estimation effect is achieved.
Patent Information
- Application Number
- CN202510095027.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
Existing depth estimation methods are sensitive to image noise, ambient light changes and sensor noise, which lead to unstable and inaccurate depth estimation results. The echo waveform-based method handles insufficient multipath reflection in complex environments, resulting in poor accuracy and robustness of depth estimation.
By acquiring the echo data and image data of the detection scene, the pixel points in the image data are positionally encoded, and the feature information of the echo data and position encoding are fused, and the depth estimation process is performed using the fused feature information.
It effectively improves the accuracy and robustness of depth estimation, and can capture depth information more accurately in complex environments and reduce noise impact.
Smart Images

Figure CN120014011A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a depth estimation method, an electronic device, a storage medium, and a program product. Background Art
[0002] In the fields of autonomous driving and robot navigation, depth estimation is often required to achieve the mapping of two-dimensional images to three-dimensional space. Among the current relevant depth estimation methods, pixel-based depth estimation methods are highly sensitive to image noise. Factors such as ambient light changes, sensor noise, and image compression will introduce noise, resulting in unstable and inaccurate depth estimation results. Other depth estimation methods based on echo waveforms have the defect of insufficient processing capabilities for multipath reflections when facing complex environments, resulting in poor accuracy and robustness of depth estimation. Summary of the invention
[0003] The embodiments of the present application provide a depth estimation method, an electronic device, a storage medium, and a program product, which can effectively improve the accuracy and robustness of depth estimation.
[0004] In a first aspect, an embodiment of the present application provides a depth estimation method, the method comprising:
[0005] Acquire echo data and image data of the detection scene;
[0006] Performing position coding processing on each pixel in the image data to obtain the position code corresponding to each pixel;
[0007] Perform feature fusion processing on echo data and position code to obtain fused feature information;
[0008] Depth estimation is performed based on the fused feature information to obtain the depth estimation result corresponding to the detection scene.
[0009] In this embodiment, the electronic device can obtain image data of the detection scene and perform position encoding on each pixel in the image data, thereby capturing the spatial position of the pixel in the two-dimensional image data; then, the echo data and position encoding can be used to perform feature fusion, so that the fused feature information not only contains the depth information provided by the echo data, but also can be embedded with the spatial position information of the pixel, thereby using the fused feature information to perform depth estimation processing, which can effectively improve the accuracy and robustness of the depth estimation.
[0010] Furthermore, in some embodiments, the echo data includes echo information corresponding to each laser point; the echo data and the position code are subjected to feature fusion processing to obtain fused feature information, including:
[0011] Preprocess the echo data to obtain the preprocessed echo information corresponding to each laser point;
[0012] In the position codes corresponding to the respective pixel points, the target position codes of the target pixel points corresponding to the respective laser points are determined;
[0013] The pre-processed echo information and target position code corresponding to each laser point are spliced to obtain the spliced feature information corresponding to each laser point;
[0014] Linear transformation is performed on the spliced feature information corresponding to each laser point to obtain fused feature information.
[0015] In this embodiment, the electronic device can pre-process the echo data to obtain pre-processed echo information corresponding to each laser point, and then determine the target position code of the target pixel point corresponding to each laser point in the position code corresponding to each pixel point, so as to splice the pre-processed echo information and target position code corresponding to each laser point respectively to obtain the spliced feature information corresponding to each laser point, and then perform linear transformation processing on the spliced feature information to improve the semantic expression of the fused feature information, thereby improving the accuracy of subsequent depth estimation.
[0016] Furthermore, in some embodiments, depth estimation processing is performed based on the fused feature information to obtain a depth estimation result corresponding to the detection scene, including:
[0017] Performing feature enhancement processing on the fused feature information based on a preset attention model to obtain first feature information;
[0018] Performing depth estimation processing on the first feature information based on a preset depth estimation model to obtain depth probability distribution information;
[0019] A depth estimation result is determined based on the depth probability distribution information.
[0020] In this embodiment, the electronic device can perform feature enhancement processing on the fused feature information based on the attention mechanism to improve the expressiveness of the first feature information, and then use a preset depth estimation model to perform depth estimation processing on the first feature information to determine the depth estimation result based on the obtained depth probability distribution information, thereby improving the accuracy of the depth estimation.
[0021] Furthermore, in some embodiments, feature enhancement processing is performed on the fused feature information based on a preset attention model to obtain first feature information, including:
[0022] Based on the preset attention model, the fused feature information is globally averaged pooled to obtain the channel description vector corresponding to the fused feature information;
[0023] Determine a channel weight vector according to the channel description vector;
[0024] Determine a channel enhanced feature vector according to the channel weight vector and the fused feature information;
[0025] The first feature information is determined based on the channel-enhanced feature vector.
[0026] In this embodiment, when feature enhancement processing is performed on the fused feature information based on the attention mechanism, the fused feature information can first be globally average pooled based on the preset attention model to obtain a channel description vector, and then a channel weight vector corresponding to the channel description vector is generated using the preset fully connected neural network model. The channel-enhanced feature vector is then determined based on the channel weight vector and the fused feature information, and the first feature information is determined using the channel-enhanced feature vector, which can enhance the important features in the fused feature information and thus improve the feature expression.
[0027] Further, in some embodiments, determining the first feature information based on the channel enhanced feature vector includes:
[0028] Perform channel dimension pooling on the channel-enhanced feature vector to obtain a spatial description vector;
[0029] Generate a spatial weight map according to the spatial description vector;
[0030] The first feature information is determined according to the spatial weight map and the channel-enhanced feature vector.
[0031] In this embodiment, when determining the first feature information based on the channel-enhanced feature vector, the channel-enhanced feature vector can be pooled in the channel dimension, and then a spatial weight map is generated based on the obtained spatial description vector. The first feature information is determined based on the spatial weight map and the channel-enhanced feature vector, which can effectively enhance the expressiveness of the first feature information and thus improve the robustness of subsequent depth estimation.
[0032] Further, in some embodiments, performing depth estimation processing on the first feature information based on a preset depth estimation model to obtain depth probability distribution information includes:
[0033] Perform peak detection processing based on a preset depth estimation model and the first feature information to obtain peak points and depth values of the peak points corresponding to each laser point;
[0034] Determine the depth probability information corresponding to the peak point of each laser point based on the depth value;
[0035] The depth probability distribution information corresponding to each laser point is determined according to the depth probability information corresponding to each laser point.
[0036] In this embodiment, when performing depth estimation processing on the first feature information based on a preset depth estimation model, peak detection processing can be performed first to obtain the peak points and depth values of the peak points corresponding to each laser point, and then the depth probability information is determined based on the depth value. Furthermore, the depth probability distribution information is determined based on the depth probability information corresponding to each laser point, which can improve the accuracy of the depth probability distribution information.
[0037] Further, in some embodiments, determining a depth estimation result based on the depth probability distribution information includes:
[0038] Determine an expected value of the depth probability distribution according to the depth probability distribution information corresponding to each laser point, and determine a depth estimation result according to the expected value of the depth probability distribution; or,
[0039] The depth estimation result is determined according to the maximum likelihood estimation value of the depth probability distribution information corresponding to each laser point.
[0040] In this embodiment, when determining the depth estimation result based on the depth probability distribution information, the expected value of the depth probability distribution can be calculated according to the depth probability distribution information and the depth value corresponding to each laser point, so that the expected value of the depth probability distribution corresponding to each laser point is determined as the depth estimation result. Alternatively, the depth estimation result can also be determined according to the maximum likelihood estimation value of the depth probability distribution information corresponding to each laser point, which can effectively improve the accuracy of the depth estimation result.
[0041] Furthermore, in some embodiments, position coding is performed on each pixel in the image data to obtain position codes corresponding to each pixel, including:
[0042] Performing horizontal position coding processing and vertical position coding processing on each pixel point in the image data respectively to obtain a horizontal position code and a vertical position code of each pixel point;
[0043] The horizontal position code and the vertical position code of each pixel point are spliced to obtain the position code corresponding to each pixel point.
[0044] In this embodiment, when determining the position code corresponding to each pixel point, the horizontal position code and the vertical position code of each pixel point are determined separately, so as to splice the horizontal position code and the vertical position code to obtain the position code of each pixel point, which can effectively capture the spatial position of the pixel point without adding additional parameter burden, so as to improve the accuracy of depth estimation.
[0045] In a second aspect, an embodiment of the present application provides an electronic device, the electronic device comprising an acquisition unit, an encoding unit, a fusion unit, and a depth estimation unit;
[0046] An acquisition unit, used for acquiring echo data and image data of a detection scene;
[0047] The encoding unit is used to perform position encoding processing on each pixel point in the image data to obtain the position code corresponding to each pixel point;
[0048] A fusion unit is used to perform feature fusion processing on the echo data and the position code to obtain fused feature information;
[0049] The depth estimation unit is used to perform depth estimation processing based on the fused feature information to obtain a depth estimation result corresponding to the detection scene.
[0050] In this embodiment, the electronic device can obtain image data of the detection scene and perform position encoding on each pixel in the image data, thereby capturing the spatial position of the pixel in the two-dimensional image data; then, the echo data and position encoding can be used to perform feature fusion, so that the fused feature information not only contains the depth information provided by the echo data, but also can be embedded with the spatial position information of the pixel, thereby using the fused feature information to perform depth estimation processing, which can effectively improve the accuracy and robustness of the depth estimation.
[0051] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory storing instructions executable by the processor; when the instructions are executed by the processor, the above-mentioned depth estimation method is implemented.
[0052] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned depth estimation method.
[0053] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the steps in the above-mentioned depth estimation method. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 Schematic diagram of the implementation process of the depth estimation method provided in the embodiment of the present application Figure 1 ;
[0055] Figure 2 Schematic diagram of the implementation process of the depth estimation method provided in the embodiment of the present application Figure 2 ;
[0056] Figure 3Schematic diagram of the composition structure of the electronic device provided in the embodiment of the present application Figure 1 ;
[0057] Figure 4 Schematic diagram of the composition structure of the electronic device provided in the embodiment of the present application Figure 2 . DETAILED DESCRIPTION
[0058] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, not for limiting the scope of protection of the present invention.
[0059] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0060] In the field of modern sensors and image processing, depth estimation is a key step in mapping two-dimensional images to three-dimensional space, and is widely used in many fields such as autonomous driving, robot navigation, and augmented reality. Traditional pixel-based depth estimation methods mainly rely on monocular or binocular vision to infer depth by analyzing image texture, edges, and stereo matching information. However, these methods face many challenges in practical applications.
[0061] First, pixel-based depth estimation methods are highly sensitive to image noise. Factors such as ambient light changes, sensor noise, and image compression can introduce noise, resulting in unstable and inaccurate depth estimation results. Second, such methods are usually computationally complex, especially when processing high-resolution images or real-time applications, the consumption of computing resources increases significantly, limiting their application on resource-constrained devices. In addition, bandwidth limitation is also an important bottleneck, especially in scenarios where a large amount of image data needs to be transmitted, where data transmission efficiency directly affects the overall performance of the system. In order to solve the above problems, in recent years, researchers have begun to explore the use of hardware-detected echo waveform characteristics for depth estimation. The echo waveform can provide more direct and accurate distance information through the time difference and shape characteristics of the transmitted and received signals. Compared with traditional pixel-based methods, the use of echo waveforms for depth estimation can not only significantly improve the accuracy of depth measurement, but also effectively optimize the calculation and transmission efficiency through feature extraction and data normalization techniques. This method reduces the dependence on high-complexity image processing algorithms and reduces the system's demand for computing resources and bandwidth, thereby improving performance while expanding the scope of application of depth estimation technology.
[0062] However, existing depth estimation methods based on echo waveforms still have certain limitations, such as insufficient processing of multipath reflections in complex environments, which leads to poor accuracy and robustness of depth estimation.
[0063] In order to solve the above problems, in an embodiment of the present application, a depth estimation method, an electronic device, a storage medium and a program product are proposed, wherein the electronic device obtains echo data and image data of a detection scene; performs position encoding processing on each pixel in the image data to obtain a position code corresponding to each pixel; performs feature fusion processing on the echo data and the position code to obtain fused feature information; performs depth estimation processing based on the fused feature information to obtain a depth estimation result corresponding to the detection scene, which can effectively improve the accuracy and robustness of the depth estimation.
[0064] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0065] Figure 1 Schematic diagram of the implementation process of the depth estimation method provided in the embodiment of the present application Figure 1 ,like Figure 1 As shown, in an embodiment of the present application, a depth estimation method of an electronic device may include the following steps:
[0066] Step 101: Acquire echo data and image data of a detection scene.
[0067] In an embodiment of the present application, the electronic device may first acquire echo data and image data of the detection scene.
[0068] It should be noted that in the embodiments of the present application, the electronic device can be any device with communication and storage functions, and the present application does not limit it; for example, the electronic device can be an in-vehicle electronic device, a car machine, a personal computer (PC), a laptop computer and other electronic devices.
[0069] In the embodiments of the present application, the detection scene may be any scene. For example, when the electronic device is an in-vehicle electronic device, the detection scene may be a real-time environment or real-time scene in which the vehicle is located during driving.
[0070] In an embodiment of the present application, the echo data may be data acquired based on a laser radar; the laser radar may scan and detect a scene based on a set signal type, signal transmission frequency, and bandwidth to obtain echo data; for example, the signal type may include pulsed laser, modulated continuous wave (CW), or frequency modulated continuous wave (Frequency Modulated Continuous Wave, FMCW), the frequency modulation range may be set to 77 GHz to 78 GHz, and the bandwidth may be set according to actual application requirements.
[0071] In an embodiment of the present application, the echo data may include echo information corresponding to each laser point.
[0072] In an embodiment of the present application, the electronic device may include an echo signal receiving module for receiving echo data; the echo signal receiving module may be configured with a high-sensitivity low-noise amplifier and a directional antenna, and improve the overall performance of the receiver through a low-noise amplifier (Low Noise Amplifier, LNA) and efficient signal modulation technology.
[0073] In an embodiment of the present application, the image data may be data acquired based on a visual sensor. For example, the visual sensor may be a camera.
[0074] Step 102: Perform position coding processing on each pixel in the image data to obtain a position code corresponding to each pixel.
[0075] In an embodiment of the present application, after acquiring the echo data and image data of the detection scene, the electronic device may perform position coding processing on each pixel in the image data to obtain the position code corresponding to each pixel.
[0076] In an embodiment of the present application, position encoding provides spatial context information by embedding the two-dimensional position information of each pixel into the feature vector, which helps to understand the geometric structure of the image and can effectively distinguish pixels at different positions in the feature fusion stage, thereby improving the accuracy of depth estimation.
[0077] In some embodiments of the present application, when the electronic device performs position coding processing on each pixel point in the image data to obtain the position code corresponding to each pixel point, it can perform horizontal position coding processing and vertical position coding processing on each pixel point in the image data to obtain the horizontal position code and vertical position code of each pixel point; and splice the horizontal position code and vertical position code of each pixel point to obtain the position code corresponding to each pixel point.
[0078] In some embodiments of the present application, when the electronic device performs horizontal position coding and vertical position coding processing on each pixel point in the image data respectively to obtain the horizontal position coding and vertical position coding of each pixel point, it can use sine function and cosine function to perform horizontal position coding and vertical position coding processing on each pixel point in the image data to obtain the horizontal position coding and vertical position coding of each pixel point.
[0079] Exemplarily, for each pixel point (x, y) in the image data, its horizontal position code and vertical position code can be determined using a sine function and a cosine function; the determination method of the horizontal position code can be expressed as the following formula:
[0080]
[0081] Among them, PE(x,y,2i) and PE(x,y,2i+1) together constitute the horizontal position code, i represents the i-th dimension in the horizontal position code vector; the determination method of the vertical position code can be expressed as the following formula:
[0082]
[0083]
[0084] Among them, PE(y,y,2i) and PE(y,y,2i+1) together constitute the vertical position code.
[0085] In some embodiments of the present application, the horizontal position code and the vertical position code of each pixel point can be spliced or added to obtain the position code of the pixel point, which is also a two-dimensional position code vector of the pixel point.
[0086] Exemplarily, the position encoding can be expressed as the following formula:
[0087] PE(x,y)=[PE(x),PE(y)] (5)
[0088] Among them, PE(x) represents the horizontal position encoding, and PE(y) represents the vertical position encoding.
[0089] Step 103: Perform feature fusion processing on the echo data and the position code to obtain fused feature information.
[0090] In an embodiment of the present application, the electronic device may perform position coding processing on each pixel point in the image data to obtain the position code corresponding to each pixel point, and then perform feature fusion processing on the echo data and the position code to obtain fused feature information.
[0091] In some embodiments of the present application, when the electronic device performs feature fusion processing on the echo data and the position code to obtain fused feature information, it can first pre-process the echo data to obtain pre-processed echo information corresponding to each laser point; determine the target position code of the target pixel point corresponding to each laser point in the position code corresponding to each pixel point; respectively splice the pre-processed echo information and the target position code corresponding to each laser point to obtain the spliced feature information corresponding to each laser point; perform linear transformation processing on the spliced feature information corresponding to each laser point to obtain fused feature information.
[0092] In an embodiment of the present application, preprocessing may include filtering processing and compression processing, wherein the filtering processing may include at least one of Fourier transform filtering, low-pass filtering, high-pass filtering, band-pass filtering, adaptive filtering and moving average filtering; through filtering processing, the signal-to-noise ratio of the echo data can be improved; then the filtered echo data is compressed using a compression algorithm to extract the coordinates of key points therein, thereby obtaining preprocessed echo data, and the preprocessed echo data includes the preprocessed echo information corresponding to each laser point.
[0093] In the embodiments of the present application, the compression algorithm is not limited in the present application. For example, a discrete wavelet transform (DWT) compression algorithm can be used to compress the filtered echo data. The compression algorithm can detect significant change points in the filtered echo data, such as zero crossing points and local extreme points, and record the time and amplitude information of these key points. Through compression processing, the amount of data can be significantly reduced, the demand for storage space and transmission bandwidth can be reduced, and key feature information can be extracted, that is, the main features of the signal are retained to ensure that the compressed data has sufficient descriptive capabilities. The pre-processed echo data may include key point coordinate information, and the key point coordinate information may include the time t and amplitude information f(t) of the key point. This information can be used to determine the main features and changing trends of the signal.
[0094] In an embodiment of the present application, since the echo data is relatively sparse relative to the image data, that is, the laser points in the echo data are fewer than the pixel points in the image data, the present application selects target pixel points corresponding to each laser point from each pixel point, thereby splicing the preprocessed echo information of each laser point with the target position code of the target pixel point corresponding to the laser point, thereby obtaining the spliced feature information of the laser point.
[0095] In an embodiment of the present application, feature fusion processing can fuse the preprocessed echo information and position coding, so that the fused feature information not only contains the depth information provided by the echo waveform, but also embeds the spatial position information of the pixel points, laying a solid foundation for subsequent feature enhancement and depth estimation.
[0096] Exemplarily, the concatenated feature information can be expressed as concat(f,g), where concat represents a feature concatenation operation; f represents the echo information after preprocessing, R represents a real number; g represents the target position code, but
[0097] Exemplarily, linear transformation is performed on the concatenated feature information to obtain fused feature information which can be expressed as the following formula:
[0098] FusionFeature=ReLU(W×concat(f,g)+b) (6)
[0099] Among them, FusionFeature represents the fused feature information after linear transformation processing, W represents weight, b represents bias, and ReLU represents Rectified Linear Unit, which is a commonly used activation function in neural networks.
[0100] Step 104: Perform depth estimation processing based on the fused feature information to obtain a depth estimation result corresponding to the detection scene.
[0101] In an embodiment of the present application, the electronic device may perform feature fusion processing on the echo data and the position code to obtain fused feature information, and then perform depth estimation processing based on the fused feature information to obtain a depth estimation result corresponding to the detection scene.
[0102] In some embodiments of the present application, when the electronic device performs depth estimation processing based on the fused feature information to obtain a depth estimation result corresponding to the detection scene, it can perform feature enhancement processing on the fused feature information based on a preset attention model to obtain first feature information; perform depth estimation processing on the first feature information based on the preset depth estimation model to obtain depth probability distribution information; and determine the depth estimation result based on the depth probability distribution information.
[0103] In an embodiment of the present application, the preset attention model is a network model constructed based on the attention mechanism. Through the preset attention model, the ability of the fused feature information in feature expression can be enhanced. This is because the fused feature information contains rich echo waveform depth information and spatial position information of pixel points. However, not all features are equally important in depth estimation. Some features may contain key information, while other features may be noise or redundant information. The attention mechanism can automatically identify and strengthen important features and suppress unimportant features by assigning different weights to different features, thereby improving the quality of the overall feature representation and the performance of depth estimation. The preset attention model combines channel attention and spatial attention to form a hybrid attention mechanism to make full use of the multi-dimensional information of the fused features.
[0104] In some embodiments of the present application, the fused feature information can be dynamically weighted based on the hybrid attention mechanism in the preset attention model to generate an enhanced feature vector, i.e., the first feature information; the preset attention model can include a neural network with a pooling layer and two fully connected layers.
[0105] In some embodiments of the present application, when the electronic device performs feature enhancement processing on the fused feature information based on a preset attention model to obtain the first feature information, it can perform global average pooling processing on the fused feature information based on the preset attention model to obtain a channel description vector corresponding to the fused feature information; determine a channel weight vector based on the channel description vector; determine a channel-enhanced feature vector based on the channel weight vector and the fused feature information; and determine the first feature information based on the channel-enhanced feature vector.
[0106] Exemplarily, the fused feature information is globally averaged pooled based on the preset attention model, and the channel description vector obtained can be expressed as the following formula:
[0107]
[0108] Among them, c represents the channel index, H and W represent the height and width of the feature map corresponding to the fused feature information, i and j represent the row and column respectively; then, the channel description vector z can be c Input into a neural network containing two fully connected layers to generate a channel weight vector, which can be expressed as the following formula:
[0109] s = σ(W2×ReLU(W1×z c )) (8)
[0110] Among them, W1 and W2 are the weight matrices of the two fully connected layers, and σ represents the Sigmoid activation function; then, we can calculate the weight matrix of the channel weight vector s and the fused feature information [f c ,g c ] c,i,j Determine the channel enhancement feature vector, which can be expressed as the following formula:
[0111] [f c ,g c ] c,i,j =s×[f,g] c,i,j (9)
[0112] In some embodiments of the present application, when the electronic device determines the first feature information based on the channel-enhanced feature vector, it can perform channel dimension pooling processing on the channel-enhanced feature vector to obtain a spatial description vector; generate a spatial weight map based on the spatial description vector; and determine the first feature information based on the spatial weight map and the channel-enhanced feature vector.
[0113] In some embodiments of the present application, the pooling process may include average pooling and maximum pooling; for example, the pooling process of the channel dimension of the channel-enhanced feature vector may be expressed as the following formula:
[0114]
[0115] Where C represents the number of channels, m avg represents the vector obtained by averaging the channel dimension of the channel-enhanced feature vector, m max represents the vector obtained by performing the maximum pooling of the channel dimension on the channel-enhanced feature vector; the vector m obtained by performing the average pooling of the channel dimension avg The vector m obtained after the maximum pooling of the channel dimension max By performing feature concatenation, we can obtain a spatial description vector m. Then, we can input the spatial description vector m into a 7×7 convolutional layer to generate a spatial weight map p. This process can be expressed as p=σ(Con 7×7 (m)), σ represents the Sigmoid activation function; finally, the spatial weight map p can be applied to the channel-enhanced feature vector [f c ,g c ] c,i,j , thereby obtaining the first characteristic information, which can be expressed as the following formula:
[0116] [f′,g′] c,i,j =p×[f c ,g c ] c,i,j (13)
[0117] In some embodiments of the present application, when the electronic device performs depth estimation processing on the first feature information based on a preset depth estimation model to obtain depth probability distribution information, it can perform peak detection processing based on the preset depth estimation model and the first feature information to obtain the peak point and the depth value of the peak point corresponding to each laser point; determine the depth probability information corresponding to the peak point of each laser point based on the depth value; and determine the depth probability distribution information corresponding to each laser point according to the depth probability information corresponding to each laser point.
[0118] In an embodiment of the present application, the peak point corresponding to each laser point in the first feature information may be analyzed based on a preset depth estimation model, thereby reflecting the possibility of depth position distribution through the peak point.
[0119] Exemplarily, for a certain peak point, the preset depth estimation model can predict or estimate its depth at different depth positions to obtain multiple depth estimation values, and the depth estimation value can be understood as the probability corresponding to each depth position; for example, when calculating the depth probability information corresponding to the k-th peak point, it can be calculated by the following formula:
[0120]
[0121] Among them, d represents the actual depth value of the peak point, μ k represents the depth mean corresponding to the peak point, σ k It represents the standard deviation of the depth corresponding to the peak point, reflecting the uncertainty of the depth position of the peak. The depth mean and the depth standard deviation are calculated through multiple depth estimation values corresponding to the peak point. The depth mean is obtained by averaging multiple depth estimation values, and the standard deviation is obtained by calculating the standard deviation of multiple depth estimation values.
[0122] In some embodiments of the present application, one laser point may correspond to multiple peak points. When determining the depth probability distribution information corresponding to the laser point, the depth probability information corresponding to the multiple peak points corresponding to the laser point can be added together to obtain the depth probability distribution information corresponding to the laser point. Similarly, this operation is performed on each laser point to obtain the depth probability distribution information corresponding to each laser point.
[0123] Exemplarily, the depth probability information of the K peak points corresponding to a certain laser point is added together, and the obtained depth probability distribution information can be expressed as the following formula:
[0124]
[0125] Among them, w k Represents the weight of the k-th peak point, which can be used to reflect the importance of this peak point in depth estimation.
[0126] In some embodiments of the present application, when the electronic device determines the depth estimation result based on the depth probability distribution information, it can determine the expected value of the depth probability distribution according to the depth probability distribution information corresponding to each laser point, and determine the depth estimation result according to the expected value of the depth probability distribution; or, determine the depth estimation result according to the maximum likelihood estimation value of the depth probability distribution information corresponding to each laser point.
[0127] Exemplarily, the expected value of the depth probability distribution information is used as the final depth result of the laser point, and the expected value D1 can be expressed as the following formula:
[0128]
[0129] Among them, d max Represents the maximum value of the depth estimate, d min Indicates the minimum value among the depth estimates.
[0130] Exemplarily, the maximum likelihood estimation value D2 of the depth probability distribution information is used as the depth result of the laser point, which can be expressed as the following formula:
[0131]
[0132] It can be understood that the depth results of all laser points can constitute the final depth estimation result of the detection scene.
[0133] The embodiment of the present application provides a depth estimation method, in which an electronic device obtains echo data and image data of a detection scene; performs position encoding processing on each pixel in the image data to obtain the position encoding corresponding to each pixel; performs feature fusion processing on the echo data and the position encoding to obtain fused feature information; performs depth estimation processing based on the fused feature information to obtain a depth estimation result corresponding to the detection scene. It can be seen that the electronic device can obtain the image data of the detection scene, and perform position encoding on each pixel in the image data, thereby being able to capture the spatial position of the pixel in the two-dimensional image data; and then can use the echo data and the position encoding to perform feature fusion, so that the fused feature information contains both the depth information provided by the echo data and the spatial position information of the pixel, thereby using the fused feature information to perform depth estimation processing, which can effectively improve the accuracy and robustness of the depth estimation.
[0134] Based on the above embodiment, in another embodiment of the present application, illustratively, as Figure 2 As shown, when performing depth estimation, echo data can be obtained (step 201), and then the echo data is preprocessed (step 202), and then the image data is position encoded (step 203), and then the position encoding and echo data are feature fused (step 204), and then the fused feature information is feature enhanced using the attention mechanism (step 205), and finally depth estimation is performed (step 206), that is, depth estimation is performed based on the first feature information obtained after feature enhancement, which can effectively improve the accuracy and robustness of depth estimation.
[0135] In an embodiment of the present application, a laser radar or other sensor can be used as a signal source, and the appropriate signal type can be selected according to the needs; common signals include pulsed laser, modulated continuous wave or frequency modulated continuous wave, and the appropriate bandwidth is set according to actual needs, and the frequency modulation range is set to 77GHz to 78GHz. The transmission frequency and bandwidth of the signal directly affect the accuracy and range of depth estimation. High-frequency signals can provide higher distance resolution, but may be limited by transmission distance and environmental factors. The design of bandwidth requires a balance between resolution and system complexity. In order to avoid interference with the surrounding environment and equipment, the transmission power needs to be strictly controlled. This application adopts two technologies, pulse amplitude modulation and automatic power control, to dynamically adjust the transmission power according to actual measurement requirements.
[0136] In some embodiments of the present application, the electronic device may be configured with an echo signal receiving module, and the echo data is received through the echo signal receiving module. The echo signal receiving module is configured with a high-sensitivity low-noise amplifier and a directional antenna, and a low-noise amplifier and an efficient signal modulation technology are used to improve the overall performance of the receiver. In order to improve the anti-interference ability and measurement accuracy of the system, the present invention also adopts a multi-channel receiving architecture. The echo signal is synchronously received by multiple antennas or multiple detectors, and the reliability of the signal is enhanced by combining spatial diversity technology. The electronic device can also be configured with a data acquisition module and a synchronization control module. The data acquisition module is provided with an analog-to-digital converter (ADC) with a sampling rate of 2Gs / s to achieve high-speed sampling of the echo signal; the synchronization control module can ensure the time alignment of the transmission and reception, and use digital filtering technology to remove stray signals and noise; finally, the echo data can be stored in the high-speed memory of the electronic device for use by subsequent depth estimation algorithms.
[0137] In some embodiments of the present application, the electronic device can pre-process the acquired echo data, including Fourier transform filtering denoising, low-pass filtering and other operations to improve the signal-to-noise ratio of the echo waveform. Subsequently, a compression algorithm is used to compress the filtered echo waveform, extract key point coordinate information, and obtain pre-processed echo data.
[0138] In some embodiments of the present application, echo data is often interfered by various noises during the acquisition process, such as environmental noise, electromagnetic interference, etc.; the purpose of filtering and denoising is to eliminate or reduce the impact of these noises on the signal, thereby improving the clarity and accuracy of the signal; commonly used denoising methods include: moving average filtering and Fourier transform filtering. Filtering is a key step in the preprocessing process, and its main purpose is to further improve the quality of the signal and remove residual noise and unnecessary frequency components.
[0139] In some embodiments of the present application, the filtering method may include low-pass filtering, high-pass filtering, band-pass filtering, adaptive filtering, etc. Dynamically adjusting the filtering parameters according to the statistical characteristics of the signal can maintain a good filtering effect when the signal and noise characteristics change. After filtering the echo data, the echo data often still has a high redundancy. The present application uses a compression algorithm of discrete wavelet transform to compress the data, which significantly reduces the demand for storage space and transmission bandwidth, and also extracts key feature information.
[0140] In some embodiments of the present application, in the compressed pre-processed echo waveform data, the key point coordinate information contains the main features and change trends of the signal. These key points usually correspond to the peaks, valleys or inflection points of the signal, reflecting the important characteristics of the signal. The compression algorithm first detects significant change points in the signal, such as zero crossing points, local extreme points, etc., and then records the time and amplitude information of the key points to form a coordinate pair [t, f(t)], where t represents time and f(t) represents signal amplitude. Finally, by selecting the most representative key points, the amount of data is reduced while retaining the main features of the signal to ensure that the compressed information has sufficient descriptive power. After the above preprocessing steps, the compressed echo information f finally obtained contains the main features and key information of the signal. These compressed data not only take up less storage space, but also speed up subsequent data transmission and processing.
[0141] In some embodiments of the present application, in the process of depth estimation, it is crucial to accurately capture the spatial position information of each pixel in the two-dimensional image; position coding, as an effective means, can provide the necessary spatial reference for subsequent feature fusion, thereby improving the accuracy and robustness of depth estimation. When mapping two-dimensional image information to three-dimensional space, relying solely on echo waveforms for depth estimation may ignore the spatial distribution characteristics of pixels. Position coding provides spatial contextual information by embedding the two-dimensional position information of each pixel into the feature vector, which not only helps to understand the geometric structure of the image, but also effectively distinguishes pixels at different positions in the feature fusion stage, thereby improving the accuracy of depth estimation.
[0142] In some embodiments of the present application, for each pixel in the two-dimensional image data, its position index in the horizontal and vertical directions can be defined. Assuming that the size of the image data is H×W, the horizontal position index can be x∈{0,1,…,W-1}, and the vertical position index can be y∈{0,1,…,H-1}. Assuming that the dimension of the position code is d, an even dimension is usually selected so that the horizontal and vertical position codes can be embedded in different dimensions respectively. For each position index (x, y), its position code vector can be calculated according to the above formula (1) and formula (2) to obtain the horizontal position code, and according to formula (3) and formula (4) to obtain the vertical position code. Finally, the horizontal position code and the vertical position code are concatenated or added to obtain the final two-dimensional position code.
[0143] In some embodiments of the present application, feature fusion is a key step to integrate feature information from different sources to form a richer and more semantically rich information representation. The preprocessed echo information f can be fused with the position code g to generate a fused feature vector [f, g]. This fused feature vector contains both the depth information provided by the echo waveform information and the embedded spatial position information of the pixel points, laying a solid foundation for subsequent feature enhancement and depth estimation.
[0144] In some embodiments of the present application, since feature information from a single source is often insufficient to fully describe the spatial structure and depth information in an image, the preprocessed echo information provides direct depth measurement data, and the position code contains the position information of each pixel in the two-dimensional image. Using echo information alone may not be able to fully utilize the spatial structure of the image, and relying solely on position coding cannot reflect the depth information. Therefore, by effectively integrating the two, it is possible to comprehensively utilize the depth information of the echo waveform and the spatial information of the position code to improve the accuracy and robustness of depth estimation.
[0145] In some embodiments of the present application, feature fusion is mainly achieved by combining feature splicing and linear transformation to ensure that the fused feature vector not only maintains the independence of each feature, but also achieves information complementarity and enhancement through linear transformation; in the present application, a feature splicing and fusion method is first adopted to splice two feature vectors in a specific dimension to form a new high-dimensional feature vector. In order to further enhance the expressive power of the fused features, a linear transformation layer is introduced after feature splicing to perform linear transformation on the spliced features, so as to learn a more semantic information representation and obtain fused feature information.
[0146] In some embodiments of the present application, applying an attention mechanism to fused feature information can enhance the expressiveness of feature representation and improve the accuracy and robustness of depth estimation; the fused feature information contains rich echo waveform depth information and spatial position information of pixel points. However, not all features are equally important in depth estimation. Some features may contain key information, while other features may be noise or redundant information. The attention mechanism can automatically identify and strengthen important features and suppress unimportant features by assigning different weights to different features, thereby improving the quality of the overall feature representation and the performance of depth estimation.
[0147] In some embodiments of the present application, a hybrid attention mechanism is formed by combining channel attention and spatial attention to make full use of the multi-dimensional information of the fused features, and the enhanced first feature information is generated by dynamically weighting the fused feature information.
[0148] In some embodiments of the present application, the fused feature information can first be globally average pooled to generate a channel description vector, and then the channel description vector can be input into a neural network containing two fully connected layers to generate a channel weight vector, and the channel weight vector can be applied to the fused feature information to generate a channel-enhanced feature vector, and then the channel-enhanced feature vector can be average pooled and maximum pooled in the channel dimension to obtain a spatial description vector, and the spatial description vector can be input into a 7×7 convolutional layer to generate a spatial weight map, and finally the spatial weight map can be applied to the channel-enhanced feature vector to generate the first feature information that is finally enhanced by the attention mechanism.
[0149] In some embodiments of the present application, when performing depth estimation based on the first feature information, the peak points in the first feature information can be analyzed based on a constructed preset depth estimation model to reflect the possible distribution of the object at different depth positions, which can more comprehensively capture the complex shape and multi-level structure of the object, thereby significantly improving the detail expression capability of depth perception; the preset depth estimation model can be a model based on a convolutional neural network.
[0150] In some embodiments of the present application, a depth probability distribution curve in the first feature information can be generated based on a preset depth estimation model. The depth probability distribution curve includes a peak point and a depth estimation value corresponding to the peak point. The depth estimation value can be understood as the probability value of the depth position of the Zener corresponding to this peak point. Each peak point can have multiple depth estimation values. The depth probability information of each peak point can be calculated according to the Gaussian distribution function. For example, for the kth peak point, when calculating its depth probability information, the aforementioned formula (14) can be used for calculation, and then the depth probability information of all the peak points corresponding to the laser point can be superimposed to obtain the depth probability distribution information of the laser point.
[0151] In an embodiment of the present application, the final depth result of each laser point can be extracted based on a preset depth estimation model.
[0152] In some embodiments of the present application, the expected value of the depth probability distribution information may be used as the final depth result.
[0153] In some embodiments of the present application, the maximum likelihood estimate of the depth probability distribution information may be used as the final depth result; it may also be understood that the peak position of the depth probability distribution information is used as the final depth result.
[0154] In some embodiments of the present application, the above-mentioned depth estimation method can be applied to the generation of a bird's eye view (BEV) and tasks such as 3D target detection based on the bird's eye view, which can improve the depth accuracy of the bird's eye view and the reliability of related tasks.
[0155] In summary, this application can effectively convert one-dimensional waveform data into a two-dimensional probability distribution curve of depth information by generating a depth probability distribution curve, which can improve the depth estimation accuracy and reliability of the bird's-eye view related algorithm; secondly, by using multiple peak points of the echo waveform, the probability distribution form is used to reflect the possibility of different depth positions. Compared with the traditional single peak detection method, it can more comprehensively capture the complex shape and multi-level structure of the object, and significantly improve the detail expression ability of depth perception; in addition, while improving the depth estimation performance, it can also optimize the utilization efficiency of computing resources, so that the depth estimation method can be more widely used in high-tech fields such as autonomous driving and robot navigation, especially in resource-constrained and real-time application scenarios. This application utilizes the characteristics of the echo waveform and optimizes it in the feature extraction and perspective transformation stages. It not only solves the shortcomings of the traditional pixel-based depth estimation method in terms of accuracy, efficiency and resource consumption, but also significantly improves the robustness and detail expression ability of the depth estimation, which has important theoretical significance and broad practical application prospects.
[0156] The embodiment of the present application provides a depth estimation method, in which an electronic device obtains echo data and image data of a detection scene; performs position encoding processing on each pixel in the image data to obtain the position encoding corresponding to each pixel; performs feature fusion processing on the echo data and the position encoding to obtain fused feature information; performs depth estimation processing based on the fused feature information to obtain a depth estimation result corresponding to the detection scene. It can be seen that the electronic device can obtain the image data of the detection scene, and perform position encoding on each pixel in the image data, thereby being able to capture the spatial position of the pixel in the two-dimensional image data; and then can use the echo data and the position encoding to perform feature fusion, so that the fused feature information contains both the depth information provided by the echo data and the spatial position information of the pixel, thereby using the fused feature information to perform depth estimation processing, which can effectively improve the accuracy and robustness of the depth estimation.
[0157] Based on the above embodiment, in another embodiment of the present application, Figure 3 Schematic diagram of the composition structure of the electronic device provided in the embodiment of the present application Figure 1 ,like Figure 3 As shown, the electronic device 1 provided in the embodiment of the present application may include an acquisition unit 11, an encoding unit 12, a fusion unit 13 and a depth estimation unit 14.
[0158] The acquisition unit 11 may be used to acquire echo data and image data of the detection scene.
[0159] The encoding unit 12 may be used to perform position encoding processing on each pixel point in the image data to obtain a position code corresponding to each pixel point.
[0160] The fusion unit 13 can be used to perform feature fusion processing on the echo data and the position code to obtain fused feature information.
[0161] The depth estimation unit 14 may be used to perform depth estimation processing based on the fused feature information to obtain a depth estimation result corresponding to the detection scene.
[0162] In some embodiments of the present application, the echo data includes echo information corresponding to each laser point; the fusion unit 13 can also be used to pre-process the echo data to obtain pre-processed echo information corresponding to each laser point; and determine the target position code of the target pixel point corresponding to each laser point in the position code corresponding to each pixel point; and splice the pre-processed echo information and target position code corresponding to each laser point to obtain the spliced feature information corresponding to each laser point; and perform linear transformation on the spliced feature information corresponding to each laser point to obtain fused feature information.
[0163] In some embodiments of the present application, the depth estimation unit 14 can also be used to perform feature enhancement processing on the fused feature information based on a preset attention model to obtain first feature information; and perform depth estimation processing on the first feature information based on a preset depth estimation model to obtain depth probability distribution information; and determine the depth estimation result based on the depth probability distribution information.
[0164] In some embodiments of the present application, the depth estimation unit 14 can also be used to perform global average pooling processing on the fused feature information based on a preset attention model to obtain a channel description vector corresponding to the fused feature information; and determine a channel weight vector based on the channel description vector; and determine a channel-enhanced feature vector based on the channel weight vector and the fused feature information; and determine the first feature information based on the channel-enhanced feature vector.
[0165] In some embodiments of the present application, the depth estimation unit 14 can also be used to perform channel dimension pooling processing on the channel-enhanced feature vector to obtain a spatial description vector; and generate a spatial weight map based on the spatial description vector; and determine the first feature information based on the spatial weight map and the channel-enhanced feature vector.
[0166] In some embodiments of the present application, the depth estimation unit 14 can also be used to perform peak detection processing based on a preset depth estimation model and the first feature information to obtain the peak point and the depth value of the peak point corresponding to each laser point; and determine the depth probability information corresponding to the peak point of each laser point based on the depth value; and determine the depth probability distribution information corresponding to each laser point according to the depth probability information corresponding to each laser point.
[0167] In some embodiments of the present application, the depth estimation unit 14 can also be used to determine an expected value of the depth probability distribution based on the depth probability distribution information corresponding to each laser point, and determine the depth estimation result based on the expected value of the depth probability distribution; or, determine the depth estimation result based on the maximum likelihood estimate value of the depth probability distribution information corresponding to each laser point.
[0168] In some embodiments of the present application, the encoding unit 12 can also be used to perform horizontal position encoding processing and vertical position encoding processing on each pixel point in the image data, respectively, to obtain the horizontal position code and vertical position code of each pixel point; and to splice the horizontal position code and vertical position code of each pixel point, respectively, to obtain the position code corresponding to each pixel point.
[0169] In the embodiments of the present application, further, Figure 4 Schematic diagram of the composition structure of the electronic device provided in the embodiment of the present application Figure 2 ,like Figure 4 As shown, the electronic device 1 provided in the embodiment of the present application may further include a processor 15 and a memory 16 storing executable instructions of the processor 15 ; further, the electronic device 1 may further include a communication interface 17 and a bus 18 for connecting the processor 15 , the memory 16 and the communication interface 17 .
[0170] In the embodiment of the present application, the processor 15 can be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic device used to implement the function of the processor can also be other, and the embodiment of the present application is not specifically limited. The electronic device 10 can also include a memory 16, which can be connected to the processor 15, wherein the memory 16 is used to store executable program code, the program code includes computer operation instructions, and the memory 16 may include a high-speed RAM memory, and may also include a non-volatile memory, for example, at least two disk memories.
[0171] In the embodiment of the present application, the bus 18 is used to connect the communication interface 17, the processor 15 and the memory 16, and the mutual communication between these devices.
[0172] In the embodiment of the present application, the memory 16 is used to store instructions and data.
[0173] Further, in an embodiment of the present application, the processor 15 is used to obtain echo data and image data of the detection scene;
[0174] Performing position coding processing on each pixel in the image data to obtain the position code corresponding to each pixel;
[0175] Perform feature fusion processing on echo data and position code to obtain fused feature information;
[0176] Depth estimation is performed based on the fused feature information to obtain the depth estimation result corresponding to the detection scene.
[0177] In practical applications, the memory 16 may be a volatile memory, such as a random access memory (RAM); or a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk (HDD) or a solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 15.
[0178] In addition, each functional module in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or software functional modules.
[0179] If the integrated unit is implemented in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method of this embodiment. The aforementioned storage medium includes: U disk, mobile hard disk, read only memory (ROM), random access memory (RAM), disk or optical disk, etc. Various media that can store program codes.
[0180] The embodiment of the present application provides an electronic device, which includes an acquisition unit, an encoding unit, a fusion unit and a depth estimation unit; the acquisition unit is used to acquire echo data and image data of the detection scene; the encoding unit is used to perform position encoding processing on each pixel in the image data to obtain the position encoding corresponding to each pixel; the fusion unit is used to perform feature fusion processing on the echo data and the position encoding to obtain fused feature information; the depth estimation unit is used to perform depth estimation processing based on the fused feature information to obtain the depth estimation result corresponding to the detection scene. It can be seen that the electronic device can acquire the image data of the detection scene and perform position encoding on each pixel in the image data, thereby capturing the spatial position of the pixel in the two-dimensional image data; and then the echo data and the position encoding can be used to perform feature fusion, so that the fused feature information contains both the depth information provided by the echo data and the spatial position information of the pixel, so that the fused feature information is used to perform depth estimation processing, which can effectively improve the accuracy and robustness of the depth estimation.
[0181] Specifically, a program instruction corresponding to a depth estimation method in this embodiment may be stored on a storage medium such as a CD, a hard disk, or a USB flash drive. When a program instruction corresponding to a depth estimation method in the storage medium is read or executed by an electronic device, the following steps are included:
[0182] Acquire echo data and image data of the detection scene;
[0183] Performing position coding processing on each pixel in the image data to obtain the position code corresponding to each pixel;
[0184] Perform feature fusion processing on echo data and position code to obtain fused feature information;
[0185] Depth estimation is performed based on the fused feature information to obtain the depth estimation result corresponding to the detection scene.
[0186] An embodiment of the present application provides a computer program product, including a computer program or instructions. When the computer program or instructions are executed by a processor, the steps in the above-mentioned depth estimation method are implemented.
[0187] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.
[0188] The present application is described with reference to implementation flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0189] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which is implemented in the implementation flow diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0190] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing the steps in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0191] The above embodiments are only preferred embodiments for fully illustrating the present invention, and the protection scope of the present invention is not limited thereto. Any equivalent substitution or change made by a person skilled in the art based on the present invention is within the protection scope of the present invention.
Claims
1. A depth estimation method, characterized in that: The method comprises: Acquire echo data and image data of the detection scene; Performing position coding processing on each pixel point in the image data to obtain a position code corresponding to each pixel point; Performing feature fusion processing on the echo data and the position code to obtain fused feature information; Depth estimation processing is performed based on the fused feature information to obtain a depth estimation result corresponding to the detection scene.
2. The method according to claim 1, characterized in that The echo data includes echo information corresponding to each laser point; the echo data and the position code are subjected to feature fusion processing to obtain fused feature information, including: Preprocessing the echo data to obtain preprocessed echo information corresponding to each laser point; Determine the target position codes of the target pixel points respectively corresponding to the respective laser points in the position codes respectively corresponding to the respective pixel points; Respectively performing splicing processing on the pre-processed echo information and the target position code corresponding to each of the laser points to obtain the spliced feature information corresponding to each of the laser points; Linear transformation is performed on the spliced feature information corresponding to each of the laser points to obtain the fused feature information.
3. The method according to claim 1 or 2, characterized in that: The performing depth estimation processing based on the fused feature information to obtain a depth estimation result corresponding to the detection scene includes: Performing feature enhancement processing on the fused feature information based on a preset attention model to obtain first feature information; Performing depth estimation processing on the first feature information based on a preset depth estimation model to obtain depth probability distribution information; The depth estimation result is determined based on the depth probability distribution information.
4. The method according to claim 3, characterized in that The performing feature enhancement processing on the fused feature information based on the preset attention model to obtain the first feature information includes: Performing global average pooling processing on the fused feature information based on the preset attention model to obtain a channel description vector corresponding to the fused feature information; Determine a channel weight vector according to the channel description vector; Determine a channel enhanced feature vector according to the channel weight vector and the fused feature information; The first feature information is determined based on the channel-enhanced feature vector.
5. The method according to claim 4, characterized in that The determining the first feature information based on the feature vector enhanced by the channel includes: Performing channel dimension pooling processing on the channel enhanced feature vector to obtain a spatial description vector; Generate a spatial weight map according to the spatial description vector; The first feature information is determined according to the spatial weight map and the channel-enhanced feature vector.
6. The method according to claim 3, characterized in that: The performing depth estimation processing on the first feature information based on a preset depth estimation model to obtain depth probability distribution information includes: Perform peak detection processing based on a preset depth estimation model and the first feature information to obtain peak points corresponding to each laser point and depth values of the peak points; Determine depth probability information corresponding to the peak points of each laser point based on the depth value; The depth probability distribution information corresponding to each laser point is determined according to the depth probability information corresponding to each laser point.
7. The method according to claim 6, characterized in that The determining the depth estimation result based on the depth probability distribution information includes: Determine an expected value of depth probability distribution according to the depth probability distribution information corresponding to each of the laser points, and determine the depth estimation result according to the expected value of depth probability distribution; or, The depth estimation result is determined according to the maximum likelihood estimation value of the depth probability distribution information corresponding to each of the laser points.
8. The method according to claim 1 or 2, characterized in that: The performing position coding processing on each pixel point in the image data to obtain the position code corresponding to each pixel point includes: Performing horizontal position coding processing and vertical position coding processing on each pixel point in the image data respectively to obtain a horizontal position code and a vertical position code of each pixel point; The horizontal position codes and the vertical position codes of the respective pixel points are spliced to obtain the position codes corresponding to the respective pixel points.
9. An electronic device, characterized in that: The electronic device includes an acquisition unit, an encoding unit, a fusion unit and a depth estimation unit; The acquisition unit is used to acquire echo data and image data of the detection scene; The encoding unit is used to perform position encoding processing on each pixel point in the image data to obtain the position code corresponding to each pixel point; The fusion unit is used to perform feature fusion processing on the echo data and the position code to obtain fused feature information; The depth estimation unit is used to perform depth estimation processing based on the fused feature information to obtain a depth estimation result corresponding to the detection scene.
10. An electronic device, characterized in that: The electronic device includes a processor and a memory storing instructions executable by the processor; when the instructions are executed by the processor, the method according to any one of claims 1 to 8 is implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
12. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps in the depth estimation method according to any one of claims 1 to 8 are implemented.