Image depth prediction method, device and equipment based on multi-feature fusion
Through the feature fusion of global encoder and local encoder network combined with cross attention network, the problem of insufficient processing of global context and local details in drone image depth estimation is solved, and the depth estimation capability of drone in complex environments is improved.
Patent Information
- Application Number
- CN202510260226.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-04
AI Technical Summary
The existing drone image depth estimation methods have limited capabilities when processing global context information, making it difficult to capture local details, resulting in errors in the judgment of obstacle distances in complex terrain environments, affecting the safety of the drone and mapping accuracy.
The global encoder network and the local encoder network are used to extract the global and local features of the image, and feature fusion is performed through the cross attention network, combining Fourier transform and inverse Fourier transform, and finally input the trained depth prediction network to obtain the image depth parameters.
It improves the accuracy of image depth prediction, enhances the depth estimation performance of the drone in complex scenarios, and ensures the reliability and accuracy of depth prediction.
Smart Images

Figure CN120259395A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to an image depth prediction method, device and equipment based on multi-feature fusion. Background Art
[0002] In recent years, UAV technology has been widely used in fields such as aerial surveying and mapping, environmental monitoring, and disaster assessment. The image data collected by the cameras carried by UAVs provides important visual information support for these applications. However, traditional UAV images record three-dimensional scenes in a two-dimensional form, and depth information is inevitably lost during the imaging process. This lack of depth information not only limits our understanding of the spatial structure of the scene, but also affects the precise positioning and size measurement of target objects. Especially in complex terrain environments, the lack of depth information may lead to incorrect judgments of the distance to obstacles, increasing the flight risk of UAVs. Therefore, carrying out research on depth estimation of UAV images is of great practical significance for improving the autonomous navigation ability of UAVs, enhancing mapping accuracy, and strengthening environmental perception performance. With the rapid iteration of computing hardware and the continuous breakthrough of deep learning theory, image depth estimation methods based on deep learning have gradually become the mainstream of research. However, current image depth estimation methods have limited capabilities in processing global context information and are also slightly insufficient in capturing local details, making it difficult to meet the actual application requirements. Summary of the Invention
[0003] In view of this, the present invention proposes an image depth prediction method, device and equipment based on multi-feature fusion.
[0004] The technical solution of the present invention is implemented as follows: In the first aspect of the present invention, an image depth prediction method based on multi-feature fusion is provided, including:
[0005] Inputting the image to be processed into a global encoder network and a local encoder network respectively to obtain a first global feature and a first local feature of the image to be processed;
[0006] Inputting the first global feature and the first local feature into a cross-attention network, respectively obtaining a first intermediate feature through convolution processing and a second intermediate feature through Fourier transform and inverse Fourier transform, and splicing the first intermediate feature and the second intermediate feature to obtain a target fusion feature;
[0007] Inputting the target fusion feature into a trained depth prediction network to obtain image depth parameters.
[0008] On the basis of the above technical solution, preferably, the global encoder network includes a Transformer network, and the local encoder network includes a convolutional neural network.
[0009] Based on the above technical solutions, preferably, before inputting the image to be processed into the global encoder network and the local encoder network respectively to obtain the first global feature and the first local feature of the image to be processed, it includes:
[0010] Perform preprocessing operations of randomly cropping, randomly rotating, horizontally flipping, adjusting brightness and contrast on the image to be processed, and replace the random area of the image to be processed with a mask.
[0011] Based on the above technical solutions, preferably, the step of inputting the image to be processed into the global encoder network and the local encoder network respectively to obtain the global feature and the local feature of the image to be processed includes:
[0012] Input the preprocessed image to be processed into the local encoder network to obtain the first local feature of the image to be processed;
[0013] Perform patch embedding processing and position encoding processing on the preprocessed image to be processed, and input it into the global encoder network to obtain the first global feature.
[0014] Based on the above technical solutions, preferably, the cross-attention network includes a first attention network and a second attention network: the step of inputting the first global feature and the first local feature into the cross-attention network, respectively obtaining a first intermediate feature through convolutional processing and a second intermediate feature through Fourier transform and inverse Fourier transform, and splicing the first intermediate feature and the second intermediate feature to obtain the target fusion feature includes:
[0015] Input the first global feature and the first local feature into the first attention network, use the first global feature as the Query vector of cross-attention, and use the first local feature as the Key vector and Value vector of cross-attention to obtain the first fusion feature;
[0016] Input the first fusion feature and the second global feature into the second attention network, use the second global feature as the Query vector of cross-attention, and use the first fusion feature as the Key vector and Value vector of cross-attention to obtain the target fusion feature; wherein, the second global feature is the feature obtained by processing the first global feature and the first fusion feature through the global encoder network.
[0017] Based on the above technical solutions, preferably, the steps of inputting the first global feature and the first local feature into the first attention network and inputting the first fusion feature and the second global feature into the second attention network include the following steps:
[0018] F Atten1 = Atten1(F Conv , F Trans1 )
[0019] F Trans2 = Transblock(F Trans1 + F Atten1 )
[0020] F Atten2 = Atten2(F Atten1 , F Trans2 );
[0021] Among them, F Conv is the first local feature, F Trans1 is the first global feature, F Trans2 is the second global feature, Atten1(·, ·) and Atten2(·, ·) are the first attention network and the second attention network respectively, F Atten1 and F Atten2 respectively represent the image features processed by the first attention network and the second attention network.
[0022] Based on the above technical solution, preferably, inputting the target fusion feature into a trained depth prediction network to obtain an image depth parameter includes: determining the image depth parameter d pred :
[0023]
[0024] Among them, d ni , d max and d min respectively represent the normalized inverse depth, the maximum true depth, and the minimum true depth.
[0025] Even more preferably, the second aspect of the present invention provides an image depth prediction device based on multi-feature fusion, including: a feature extraction module, a feature fusion module, and a depth prediction module; among them,
[0026] The feature extraction module is configured to input the image to be processed into a global encoder network and a local encoder network respectively to obtain the first global feature and the first local feature of the image to be processed;
[0027] The feature fusion module is configured to input the first global feature and the first local feature into a cross-attention network, and respectively obtain a first intermediate feature through convolutional processing and a second intermediate feature through Fourier transform and inverse Fourier transform, and splice the first intermediate feature and the second intermediate feature to obtain a target fusion feature;
[0028] The depth prediction module is configured to input the target fusion feature into a trained depth prediction network to obtain image depth parameters.
[0029] More preferably, a third aspect of the present invention provides an electronic device, including a processor and a memory; the memory stores a computer program, wherein the computer program, when executed by the processor, implements the multi-feature fusion-based image depth prediction method described in the first aspect.
[0030] More preferably, a fourth aspect of the present invention provides a computer storage medium, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the multi-feature fusion-based image depth prediction method described in the first aspect.
[0031] The multi-feature fusion-based image depth prediction method of the present invention has the following beneficial effects compared with the prior art:
[0032] 1. By using a global encoder network and a local encoder network to extract the global feature and local feature of the image to be processed respectively, and then using cross-attention to fuse the global feature and local feature, it effectively takes into account both local detail preservation and global consistency constraints, realizes the collaborative utilization of local details and global information, thereby improving the prediction accuracy of image depth and providing reliable technical support for depth estimation in the UAV scenario.
[0033] 2. By replacing the random region of the image to be processed with a mask during the model training process, it simulates the image information loss situation that may occur in the actual scenario, prompts the model to pay more attention to the global context information, and thus improves its depth estimation performance in complex scenarios.
[0034] 3. By predicting the normalized inverse depth and then converting the inverse depth into the actual depth, it avoids the occurrence of too large numerical values during the calculation process and reduces the calculation error. In addition, by outputting the inverse depth, the gradient propagation during model training will be smoother, ensuring the reliability of the prediction result of the depth prediction model. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0036] Figure 1Schematic flowchart of an image depth prediction method based on multi - feature fusion provided by an embodiment of the present invention;
[0037] Figure 2 Schematic architecture diagram of the global encoder network, local encoder network and cross - attention network provided by an embodiment of the present invention;
[0038] Figure 3 Schematic structure diagram of the cross - attention network provided by an embodiment of the present invention;
[0039] Figure 4 Schematic structure diagram of an image depth prediction device based on multi - feature fusion provided by an embodiment of the present invention;
[0040] Figure 5 Schematic structure diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0041] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0042] In some embodiments, as Figure 1 shown, Figure 1 Schematic flowchart of an image depth prediction method based on multi - feature fusion provided by an embodiment of the present invention; The image depth prediction method based on multi - feature fusion provided by the present invention includes:
[0043] S110: Input the image to be processed into the global encoder network and the local encoder network respectively to obtain the first global feature and the first local feature of the image to be processed.
[0044] In this actual example, the image to be processed can be an environmental image obtained by a drone through a lightweight monocular vision system. The global encoder network can capture the overall structure and information of the image to be processed. Inputting the image to be processed into the global encoder network can obtain a first global feature in the form of a vector. The first global feature can be regarded as a "summary" or "overview" of the image to be processed. Similarly, the local encoder network can capture the local details and features of the image, and the first local feature reflects the specific details of different regions in the image to be processed.
[0045] In some embodiments, the global encoder network includes a Transformer network, and the local encoder network includes a convolutional neural network.
[0046] In this embodiment, the global encoder network can be a Transformer network. Through the attention mechanism, the Transformer network can directly capture the dependencies at any position in the sequence without being limited by the sequence length, enabling it to process longer context features. By segmenting the image into patches or extracting the feature map as the input, the Transformer network can capture the global features in the image to be processed. The local encoder network can be a convolutional neural network (CNN). The CNN has good local perception characteristics and can focus on the local regions of the input data to capture local features. It should be noted that the global encoder network can also be Vision Transformer, Swin Transformer, etc., and the local encoder network can also be a residual neural network (Resnet), SWSL_ResNext, etc.
[0047] In some embodiments, before S110, where the image to be processed is respectively input into the global encoder network and the local encoder network to obtain the first global feature and the first local feature of the image to be processed, it includes:
[0048] Perform preprocessing operations on the image to be processed, such as random cropping, random rotation, horizontal flipping, adjusting brightness and contrast, and replace a random region of the image to be processed with a mask.
[0049] Random cropping is to randomly select a region from the image to be processed for cropping to ensure the generalization ability of subsequent features. Random rotation is to randomly rotate the image to be processed by a certain angle to increase the diversity of the image to be processed. Horizontal flipping is to flip the image to be processed in the horizontal direction. Adjusting brightness and contrast is to change the brightness level and the contrast of different regions of the image by changing the pixel values of the image to be processed. Replacing a random region of the image to be processed with a mask simulates the situation of missing image information that may occur in the actual scene, which can prompt the model to pay more attention to the global context information, thereby improving its depth estimation performance in complex scenes. Specifically:
[0050]
[0051] Among them, RGB(x, y) represents the RGB value at pixel (x, y), and random(x, y) represents the random region.
[0052] In some embodiments, S110, where the image to be processed is respectively input into the global encoder network and the local encoder network to obtain the global feature and the local feature of the image to be processed, includes:
[0053] Input the preprocessed image to be processed into the local encoder network to obtain the first local feature of the image to be processed;
[0054] Perform patch embedding processing and positional encoding processing on the preprocessed image to be processed, and input it into the global encoder network to obtain the first global feature.
[0055] Patch embedding processing is the process of splitting the feature map into a series of small, non-overlapping patches and converting these patches into vector representations. Patch embedding can convert image data into a patch sequence, which is more adaptable to the global encoder network. Positional encoding enables the model to understand the relative positions and spatial relationships of different patches by adding positional information to each patch embedding. By combining positional information and patch embedding, the model can more accurately capture and represent features in the image.
[0056] In this embodiment, the preprocessed image is input into the local encoder network and the global encoder network in parallel, and the first local feature F Conv and the first global feature F Trans1 :
[0057] F Cow = Convnet(image);
[0058] F Trans1 = Transnet(image);
[0059] where Convnet(·) represents the local encoder network, and Transnet(·) represents the global encoder network. The global encoder network needs to perform patch embedding and positional encoding processing on the image to be processed:
[0060] Transnet = Transblock(patch(·)+pos_encoding(·));
[0061] where Transblock(·) represents the global encoder module, which is a partial functional module of the global encoder network, patch(·) represents patch embedding, and pos_encoding(·) represents positional encoding.
[0062] S120, input the first global feature and the first local feature into the cross-attention network, and obtain the first intermediate feature through convolution processing and the second intermediate feature through Fourier transform and inverse Fourier transform, and splice the first intermediate feature and the second intermediate feature to obtain the target fusion feature.
[0063] Here, a convolution operation can be performed on at least one of the global features and local features to extract a higher-level feature representation. The global feature or local feature (or a combination of both) is subjected to a Fourier transform to convert it from the spatial domain to the frequency domain. In the frequency domain, various operations (such as filtering, feature enhancement, etc.) can be performed, and then it is converted back to the spatial domain through an inverse Fourier transform, thereby achieving feature fusion.
[0064] In some embodiments, the cross-attention network includes a first attention network and a second attention network: S120, inputting the first global feature and the first local feature into the cross-attention network, respectively obtaining a first intermediate feature through convolution processing and a second intermediate feature through Fourier transform and inverse Fourier transform, and splicing the first intermediate feature and the second intermediate feature to obtain a target fusion feature, including:
[0065] Input the first global feature and the first local feature into the first attention network. The first global feature serves as the Query vector of the cross-attention, and the first local feature serves as the Key vector and Value vector of the cross-attention to obtain a first fusion feature;
[0066] Input the first fusion feature and the second global feature into the second attention network. The second global feature serves as the Query vector of the cross-attention, and the first fusion feature serves as the Key vector and Value vector of the cross-attention to obtain a target fusion feature; wherein, the second global feature is the feature obtained by processing the first global feature and the first fusion feature through a global encoder network.
[0067] In this exemplary embodiment, two attention modules are used to fuse the first global feature and the first local feature. In the first attention network, the first global feature serves as the Query vector of the cross-attention, and its global context information can be utilized to guide the local feature selection of the first local feature. The first local feature serves as the Key vector and Value vector of the cross-attention, providing a specific local feature response for the first global feature, so that the fused feature contains both global semantic information and retains local details. Similarly, in the second attention network, the second global feature serves as the Query vector of the cross-attention, and the first fusion feature serves as the Key vector and Value vector of the cross-attention, further realizing the complementary advantages of different image features. The attention mechanism allows the network model to compare each element in the input sequence with other elements when processing the sequence, so as to correctly process each element in different contexts. The attention mechanism generates attention weights by calculating the correlation between the query vector (Query), key vector (Key), and value vector (Value), and applies the weights to the value vector to obtain a weighted sum representation.
[0068] In an alternative embodiment, refer toFigure 2 , Figure 2 This is a schematic diagram of the architectures of the global encoder network, local encoder network, and cross-attention network provided by the embodiments of the present invention; among them, the global encoder network is a Transformer network, and the local encoder network is a convolutional network. In the specific implementation process, the image to be processed is preprocessed and then sent to parallel convolutional branches and Transformer branches for feature extraction respectively. The convolutional branches focus on capturing the local detailed features of the image, while the Transformer branches are responsible for modeling the global spatial relationships of the image. The features extracted by the two branches are subjected to information interaction and integration through a carefully designed fusion module, and finally an accurate depth estimation map is generated through a depth prediction network. It should be noted that the fusion module includes attention module 1 and attention module 2, and the two cooperate with the Transformer branch to form a cross-attention network. The collaborative working mechanism of the convolutional branch and the Transformer branch effectively takes into account local detail preservation and global consistency constraints, and the cross-attention network enables the global features and local features to "pay attention" to each other, that is, enables the network to learn how to enhance or adjust local features according to global context information, or local features optimize global context information, providing reliable technical support for depth estimation in the UAV scenario.
[0069] In some embodiments, inputting the first global feature and the first local feature into the first attention network and inputting the first fusion feature and the second global feature into the second attention network includes the following steps:
[0070] F Atten1 = Atten1(F Conv , F Trans1 )
[0071] F Trans2 = Transblock(F Trans1 + F Atten1 )
[0072] F Atten2 = Atten2(F Atten1 , F Trans2 );
[0073] Among them, F Conv is the first local feature, F Trans1 is the first global feature, F Trans2 is the second global feature, Atten1(·, ·) and Atten2(·, ·) are the first attention network and the second attention network respectively, F Atten1 and F Atten2 respectively represent the image features after being processed by the first attention network and the second attention network.
[0074] In an alternative embodiment, please refer to Figure 3 , Figure 3 which is a schematic structural diagram of the cross-attention network provided by the embodiment of the present invention; in the cross-attention network, a convolution operation and a Fourier transform operation are respectively performed on the image features, and the first intermediate feature and the second intermediate feature output by the two are spliced together. The convolution operation extracts local features in the spatial domain and can capture the texture, edges and detail information of the image. The Fourier transform converts the image from the spatial domain to the frequency domain and can capture the global frequency information of the image. By splicing the two types of features, the model can utilize both local and global information simultaneously, improving the richness of feature representation.
[0075] Exemplarily, the processes that the first global feature and the first local feature experience in the cross-attention network are as follows:
[0076] F1 = crossatten(·, ·);
[0077] F2 = Convblock(F1);
[0078] F3 = ifft(fft(F1));
[0079]
[0080] where crossatten(·, ·) represents the cross-attention module, Convblock(·) represents the convolution module, fft(·) and ifft(·) respectively represent the Fourier transform and the inverse Fourier transform, denotes splicing the features, and finally obtaining the target fusion feature F to be output fusion .
[0081] S130, input the target fusion feature into the trained depth prediction network to obtain the image depth parameter.
[0082] Here, sample data covering various scenarios and depth changes can be obtained in advance. After preprocessing, image annotation and image segmentation of the sample data, the data is divided into a training set, a validation set and a test set to train the depth prediction model. The depth prediction model can be a convolutional neural network CNN and its variants, such as ResNet, VGG, etc.
[0083] In some embodiments, S130, input the target fusion feature into the trained depth prediction network to obtain the image depth parameter, including: determining the image depth parameter d based on the intermediate depth parameter output by the depth prediction network pred :
[0084]
[0085] where dni , d max and d min respectively represent the normalized inverse depth, the maximum value of the true depth, and the minimum value of the true depth.
[0086] In this actual example, the normalized inverse depth can be predicted through the depth prediction head, and then the inverse depth can be converted into the actual depth. The inverse depth value has some regularization effects mathematically. Since the depth of an object is usually a positive number, and the range of the inverse depth is usually easier to control (usually between 0 and 1), this helps to avoid excessive numerical values in the calculation process. For example, an object with a very large depth may cause overflow or instability in numerical calculations. By outputting the inverse depth, the gradient propagation during training will be smoother, and the predicted values of the model are usually more concentrated and easier to optimize.
[0087] In some embodiments, please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an image depth prediction device based on multi-feature fusion provided by an embodiment of the present invention. The present invention provides an image depth prediction device 400 based on multi-feature fusion, including: a feature extraction module 410, a feature fusion module 420, and a depth prediction module 430; wherein,
[0088] The feature extraction module 410 is configured to input the image to be processed into the global encoder network and the local encoder network respectively, and obtain the first global feature and the first local feature of the image to be processed;
[0089] The feature fusion module 420 is configured to input the first global feature and the first local feature into the cross-attention network, and respectively obtain the first intermediate feature through convolutional processing and the second intermediate feature through Fourier transform and inverse Fourier transform, and splice the first intermediate feature and the second intermediate feature to obtain the target fusion feature;
[0090] The depth prediction module 430 is configured to input the target fusion feature into the trained depth prediction network to obtain the image depth parameter.
[0091] In some embodiments, the global encoder network includes a Transformer network, and the local encoder network includes a convolutional neural network.
[0092] In some embodiments, the image depth prediction device 400 based on multi-feature fusion includes a preprocessing module; the preprocessing module is specifically configured to:
[0093] Perform preprocessing operations of random cropping, random rotation, horizontal flipping, adjusting brightness and contrast on the image to be processed, and replace the random area of the image to be processed with a mask.
[0094] In some embodiments, the feature extraction module 410 is specifically configured to:
[0095] Input the pre - processed image to be processed into the local encoder network to obtain the first local feature of the image to be processed;
[0096] Perform patch embedding processing and position encoding processing on the pre - processed image to be processed, and input it into the global encoder network to obtain the first global feature.
[0097] In some embodiments, the cross - attention network includes a first attention network and a second attention network: The feature fusion module 420 is specifically configured to:
[0098] Input the first global feature and the first local feature into the first attention network, where the first global feature serves as the Query vector of the cross - attention, and the first local feature serves as the Key vector and Value vector of the cross - attention, to obtain the first fusion feature;
[0099] Input the first fusion feature and the second global feature into the second attention network, where the second global feature serves as the Query vector of the cross - attention, and the first fusion feature serves as the Key vector and Value vector of the cross - attention, to obtain the target fusion feature; wherein, the second global feature is the feature obtained by processing the first global feature and the first fusion feature through the global encoder network.
[0100] In some embodiments, the feature fusion module 420 is specifically configured to perform the following steps:
[0101] F Atten1 = Atten1(F Conv , F Trans1 )
[0102] F Trans2 = Transblock(F Trans1 + F Atten1 )
[0103] F Atten2 = Atten2(F Atten1 , F Trans2 );
[0104] Wherein, F Conv is the first local feature, F Trans1 is the first global feature, F Trans2 is the second global feature, Atten1(·, ·) and Atten2(·, ·) are the first attention network and the second attention network respectively, and F Atten1 and F Atten2 represent the image features after being processed by the first attention network and the second attention network respectively.
[0105] In some embodiments, the depth prediction module 430 is specifically configured to determine the image depth parameter d based on the intermediate depth parameters output by the depth prediction network pred :
[0106]
[0107] where d ni 、d max and d min respectively represent the normalized inverse depth, the maximum true depth, and the minimum true depth.
[0108] It should be noted that the image depth prediction device based on multi-feature fusion provided in the embodiments of the present application and the image depth prediction method based on multi-feature fusion provided in the embodiments of the present application are based on the same application concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned image depth prediction method based on multi-feature fusion, and the repeated parts will not be elaborated.
[0109] In some embodiments, please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device provided in the embodiments of the present application. An electronic device 500 provided in the embodiments of the present application includes a processor 510 and a memory 520; the memory 520 stores a computer program, wherein the computer program, when executed by the processor, implements the above-mentioned image depth prediction method based on multi-feature fusion.
[0110] Specifically, the processor 510 may include, for example, a general microprocessor, an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), and so on. The processor 510 may also include on-board memory for caching purposes. The processor 510 may be a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiments of the present application.
[0111] The memory 520 may be, for example, any medium capable of containing, storing, transmitting, propagating, or transporting instructions. For example, the memory 520 may include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, components, or propagation media. Specific examples of the memory 520 include: magnetic storage devices, such as magnetic tapes or hard disk drives (HDDs); optical storage devices, such as compact discs (CD-ROMs); may also be, such as random access memory (RAM) or flash memory; and / or wired / wireless communication links.
[0112] The present application also provides a computer-readable medium, on which a computer program is stored. When the program is executed by a processor, it implements the above-described method for image depth prediction based on multi-feature fusion. The computer-readable medium may be included in the device / device / system described in the above embodiments; or it may exist separately without being assembled into the device / device / system. The above computer-readable medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present application is implemented.
[0113] According to an embodiment of the present application, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may, for example, be but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted by any suitable medium, including but not limited to: wireless, wired, optical cable, radio frequency signal, etc., or any suitable combination of the above.
[0114] Those skilled in the art will understand that the features recited in the various embodiments and / or claims of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present application. In particular, without departing from the spirit and teachings of the present application, the features recited in the various embodiments and / or claims of the present application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application. Therefore, the scope of the present application should not be limited to the above embodiments, but should be determined not only by the appended claims, but also by the equivalents of the appended claims. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. An image depth prediction method based on multi-feature fusion, characterized in that, Including: Input the image to be processed into the global encoder network and the local encoder network respectively to obtain the first global feature and the first local feature of the image to be processed; Input the first global feature and the first local feature into the cross-attention network, and respectively obtain the first intermediate feature through convolutional processing and the second intermediate feature through Fourier transform and inverse Fourier transform, and splice the first intermediate feature and the second intermediate feature to obtain the target fusion feature; Input the target fusion feature into the trained depth prediction network to obtain the image depth parameter.
2. The image depth prediction method based on multi-feature fusion according to claim 1, wherein The global encoder network includes a Transformer network, and the local encoder network includes a convolutional neural network.
3. The image depth prediction method based on multi-feature fusion according to claim 1, characterized in that Before inputting the image to be processed into the global encoder network and the local encoder network respectively to obtain the first global feature and the first local feature of the image to be processed, it includes: Perform preprocessing operations such as random cropping, random rotation, horizontal flipping, adjusting brightness and contrast on the image to be processed, and replace the random area of the image to be processed with a mask.
4. The method for predicting image depth based on multi-feature fusion according to claim 1, wherein, The step of inputting the image to be processed into the global encoder network and the local encoder network respectively to obtain the global feature and the local feature of the image to be processed includes: Input the preprocessed image to be processed into the local encoder network to obtain the first local feature of the image to be processed; Perform patch embedding processing and position encoding processing on the preprocessed image to be processed, and input it into the global encoder network to obtain the first global feature.
5. The method for predicting image depth based on multi-feature fusion according to claim 1, wherein The cross-attention network includes a first attention network and a second attention network: The step of inputting the first global feature and the first local feature into the cross-attention network, and respectively obtaining the first intermediate feature through convolutional processing and the second intermediate feature through Fourier transform and inverse Fourier transform, and splicing the first intermediate feature and the second intermediate feature to obtain the target fusion feature includes: Input the first global feature and the first local feature into the first attention network, use the first global feature as the Query vector of cross-attention, and use the first local feature as the Key vector and Value vector of cross-attention to obtain the first fusion feature; Input the first fusion feature and the second global feature into the second attention network, use the second global feature as the Query vector of cross-attention, and use the first fusion feature as the Key vector and Value vector of cross-attention to obtain the target fusion feature; wherein, the second global feature is the feature obtained by processing the first global feature and the first fusion feature through the global encoder network.
6. The method for predicting image depth based on multi-feature fusion according to claim 5, wherein The step of inputting the first global feature and the first local feature into the first attention network and inputting the first fusion feature and the second global feature into the second attention network includes the following steps: F Atten1 = Atten1(F Conv , F Trans1 ) F Trans2 = Transblock(F Trans1 + F Atten1 ) F Atten2 = Atten2(F Atten1 , F Trans2 ); Among them, F Conv is the first local feature, F Trans1 is the first global feature, F Trans2 is the second global feature, Atten1(·, ·) and Atten2(·, ·) are the first attention network and the second attention network respectively, F Atten1 and F Atten2 respectively represent the image features processed by the first attention network and the second attention network.
7. The method for predicting image depth based on multi-feature fusion according to claim 1, wherein Inputting the target fusion feature into the trained depth prediction network to obtain the image depth parameter, including: determining the image depth parameter d based on the intermediate depth parameter output by the depth prediction network pred : Among them, d ni , d max and d min respectively represent the normalized inverse depth, the maximum true depth, and the minimum true depth.
8. An image depth prediction device based on multi-feature fusion, characterized in that, Including: A feature extraction module, a feature fusion module and a depth prediction module; wherein, The feature extraction module is configured to input the image to be processed into the global encoder network and the local encoder network respectively, and obtain the first global feature and the first local feature of the image to be processed; The feature fusion module is configured to input the first global feature and the first local feature into the cross-attention network, and respectively obtain the first intermediate feature through convolutional processing and the second intermediate feature through Fourier transform and inverse Fourier transform, and splice the first intermediate feature and the second intermediate feature to obtain the target fusion feature; The depth prediction module is configured to input the target fusion feature into the trained depth prediction network to obtain the image depth parameter.
9. An electronic device, comprising a processor and a memory; the memory stores a computer program, wherein, When the computer program is executed by the processor, it implements the multi-feature fusion-based image depth prediction method according to any one of claims 1 to 7.
10. A computer storage medium, characterized in that, A computer program is stored thereon, wherein when the computer program is executed by the processor, it implements the multi-feature fusion-based image depth prediction method according to any one of claims 1 to 7.