Lightweight student network model training method and unmanned aerial vehicle monocular depth estimation method
Through self-supervised learning and the knowledge distillation method of the teacher-student network paradigm, combined with edge detection and guided filter fusion technology, a lightweight student network model is trained to solve the deployment problem of monocular depth estimation algorithm on UAV platform, and realize efficient and real-time depth perception and visual navigation.
Patent Information
- Application Number
- CN202511139978.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-08-14
AI Technical Summary
Existing monocular depth estimation algorithm models have a large number of parameters and consume high computing resources. They are difficult to be effectively deployed on UAV platforms with limited computing resources and cannot meet the real-time depth perception requirements of lightweight UAVs.
A self-supervised learning method is used to train a lightweight student network model. The knowledge of the teacher network is transferred to the student network through knowledge distillation in the teacher-student network paradigm. The edge detection image and guided filter fusion technology are combined to improve the depth estimation accuracy. The image is reconstructed through the posture network, and the overall loss function is optimized to achieve lightweight model.
Efficient and real-time depth perception is achieved on UAV platforms with limited computing resources, which significantly reduces the computational complexity of the algorithm model, improves the inference speed, and enhances the UAV's monocular depth estimation capability, providing a technical foundation for efficient and real-time visual navigation.
Smart Images

Figure CN120656095A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aircraft technology, and in particular to a training method for a lightweight student network model and a monocular depth estimation method for an unmanned aerial vehicle (UAV). Background Art
[0002] With the rapid development of autonomous driving, augmented reality, robotic navigation, and other fields, real-time and accurate depth estimation technology has become one of the key technologies in these application areas. Depth estimation technology provides key data support for drones in depth perception, path planning, obstacle detection, and obstacle avoidance decisions. It is a core component for drones to perceive and understand their surroundings, and plays a vital role in drones' autonomous navigation and environmental interaction. Although depth sensors such as lidar, time-of-flight cameras, and binocular cameras can provide relatively accurate depth information, these depth sensors are typically bulky, costly, and power-hungry, making them difficult to meet the strict weight, volume, size, and power requirements of drones.
[0003] For example, in high-risk post-disaster search and rescue scenarios involving drones, the risks of close-range human exploration are extremely high due to the characteristics of environmental structural damage and frequent secondary disasters. Therefore, micro-drones are prioritized for rapid environmental perception and positioning of living entities to improve search and rescue efficiency and reduce rescue risks. Accurate and real-time depth information is crucial in such missions, as drones rely on depth perception to identify obstacles, analyze terrain undulations, and determine the spatial location of trapped individuals. This allows them to plan safe flight paths and assist in rescue decision-making, providing critical data support for the rapid implementation of search and rescue missions.
[0004] However, due to limitations in payload capacity, flight time, and onboard computing power, micro-UAVs are unable to carry traditional depth sensors such as lidar. These traditional devices are often bulky, heavy, power-hungry, and expensive, making them unsuitable for lightweight UAV platforms.
[0005] In contrast, monocular cameras are low-cost, lightweight, small, low-power, highly flexible, adaptable, and easy-to-install and configure sensors. They only require a single RGB camera to estimate the depth of the surrounding environment and infer the depth information between each object in the scene and the camera. They are particularly suitable for efficient environmental perception in resource-constrained devices.
[0006] There have been some studies on monocular depth estimation algorithms, such as Monodepth2, GeoNet, CADepth, etc., and these algorithms have performed well in depth estimation accuracy. However, due to the large number of algorithm model parameters and excessive consumption of computing resources, the model size is large, the inference speed is slow, and the computing power requirements of the hardware platform are high. This makes it difficult to effectively deploy these algorithms on drone edge computing platforms with limited computing resources, thereby limiting the widespread application of this technology on resource-constrained platforms. Summary of the Invention
[0007] The purpose of this application is to provide a lightweight student network model training method and a monocular depth estimation method for drones to at least solve the above technical problems. The various technical effects that can be produced by the optional technical solutions provided by the present invention are described in detail below.
[0008] To achieve the above objectives, in a first aspect, the present application provides a training method for a lightweight student network model, comprising: Extracting an edge detection image of the original image, inputting the edge detection image and the original image into a pre-trained teacher network for inference, and obtaining an edge depth map and an initial depth map; Using the edge depth map as a guide map, fusing the initial depth map and the edge depth map to obtain a fused depth map, and using the fused depth map as a soft label to guide student network training; Inputting the original image into the student network, so that the student network performs self-supervised training under the guidance of the soft label using a teacher-student network paradigm distillation method, and outputs a predicted depth map; After splicing the original image with adjacent frame images, the spliced images are input into a posture network to obtain a posture transformation matrix, and image reconstruction is performed based on the posture transformation matrix and the predicted depth map to obtain a reconstructed image; Calculating the overall loss function of the student network; when the overall loss function is within a preset range, the student network converges to obtain a lightweight student network model; Among them, the overall loss function is the sum of the pairwise loss function, the Laplace distillation loss function, the pixel-by-pixel distillation loss function, and the basic loss function. The basic loss function is the sum of the photometric reprojection loss function and the smoothing loss function between the reconstructed image and the original image. The pixel-by-pixel distillation loss function is the pixel-by-pixel distillation loss function between the fused depth map and the predicted depth map. The pairwise loss function and the Laplace distillation loss function are determined based on the teacher intermediate feature map of the teacher network and the student intermediate feature map of the student network.
[0009] In some embodiments, the edge depth map is used as a guide map to fuse the initial depth map and the edge depth map to obtain a fused depth map, which is expressed by the following formula: ; in, represents the fused depth map, GF represents the guided filtering operation, Representing the initial depth map, including the overall depth information and spatial structure information of the original image; represents the edge depth map, including edge detail information, and serves as a guide map; r represents the radius of the guided filter window, which is used to control the local filtering range; Represents the regularization parameter, which is used to prevent overfitting and ensure the smoothness of the filter.
[0010] In some embodiments, the guided filtering operation includes: Using the initial depth map as an input image and the edge depth map as a guide image, performing guided filtering, and outputting an initial fused image that satisfies a local linear model; When the filter window is located in the edge area of the guide image, the gradient change of the initial fused image in the edge direction is enhanced to highlight the edge details; When the filtering window is located in a flat area of the guide image, local mean filtering is performed on the corresponding flat area.
[0011] In some embodiments, after splicing the original image with adjacent frame images, inputting the spliced images into a posture network to obtain a posture transformation matrix, and performing image reconstruction based on the posture transformation matrix and the predicted depth map to obtain a reconstructed image, including: After splicing the original image with adjacent frame images, a spliced image is obtained; Inputting the stitched image into the posture network to obtain a posture transformation matrix; Based on the camera intrinsic parameter matrix, the predicted depth map and the pose transformation matrix, mapping the pixel points on the original image to the corresponding pixel points in the coordinate system of the adjacent frame image through spatial geometric transformation; The pixel values on the adjacent frame images are sampled to the viewing angle of the original image to obtain a reconstructed image.
[0012] In some embodiments, calculating the overall loss function of the student network includes: Calculating the basic loss function, where the basic loss function is a photometric reprojection loss function and a smoothing loss function between the reconstructed image and the original image; Calculating a pixel-by-pixel distillation loss function between the fused depth map and the predicted depth map; Calculating a pairwise loss function and the Laplace distillation loss function between the teacher intermediate feature map of the teacher network and the student intermediate feature map of the student network; Based on the photometric reprojection loss function, the smoothing loss function, the pixel-by-pixel distillation loss function, the pairwise loss function, and the Laplace distillation loss function, the overall loss function is calculated using the following formula; ; in, represents the overall loss function, represents the basic loss function, which is the sum of the photometric reprojection loss function and the smoothing loss function. represents the affinity pairwise loss function between the teacher intermediate feature map of the teacher network and the student intermediate feature map of the student network, represents the Laplace distillation loss function, represents the pixel-wise distillation loss function.
[0013] In some embodiments, calculating a pairwise loss function between the teacher intermediate feature map of the teacher network and the student intermediate feature map of the student network includes: Based on the teacher's intermediate feature map and the student's intermediate feature map, an affinity map is constructed; wherein the nodes in the affinity map represent different spatial positions, and the connection relationship between two nodes represents similarity; the size of the teacher's intermediate feature map or the student's intermediate feature map is , the affinity map includes nodes, and the number of edges connected to each node is ; Indicates the connection scope; Indicates granularity; Based on the affinity graph, the formula for calculating the pairwise loss function is: ; in, , ; in, and Represents the width and height of the teacher's intermediate feature map or the student's intermediate feature map, respectively, Represents all nodes, represents the similarity between node i and node j in the affinity graph corresponding to the student network, represents the similarity between node i and node j in the affinity graph corresponding to the teacher network, and They represent the aggregated features obtained by average pooling. The aggregation operation means averaging the features of all channels in the node and compressing them into a feature vector of fixed length, thereby extracting the global representation information of the node. The Laplace distillation loss function between the teacher intermediate feature map of the teacher network and the student intermediate feature map of the student network is calculated and expressed as follows: ; Among them, C i is the total number of units contained in the features obtained by the decoder of the teacher network or the i-th decoder layer of the student network, N represents the total number of decoder layers of the teacher network or the student network, x is the spatial position index on the teacher's intermediate feature map, Represents the absolute difference of the residual of the first layer of the Laplacian pyramid feature, Represents the absolute difference of the residual of the second layer of the Laplacian pyramid feature, It is the squared difference between the features of the teacher network and the student network after downsampling to 1 / 4 scale.
[0014] In some embodiments, the pixel-by-pixel distillation loss function between the fused depth map and the predicted depth map is calculated using the following formula: ; in, represents the pixel-wise distillation loss function, C represents the number of scales predicted by the decoder of the teacher network or the decoder of the student network, and N i represents the number of pixels at the i-th scale, represents the depth prediction value of the teacher network at the i-th scale and the j-th pixel point, Represents the depth prediction value of the student network at the jth pixel at the i-th scale.
[0015] In some embodiments, the calculating the basic loss function includes: The photometric reprojection loss function is calculated using the photometric reprojection loss formula, which is: ; in, represents the photometric reprojection loss function, pe represents the photometric reprojection error, and SSIM represents the structural similarity index; Represents the first-order norm of the image and calculates pixel-level differences; represents the weight coefficient, represents the original image, represents the reconstructed image; The smooth loss function is calculated using the smooth loss formula, which is: ; in, represents the smooth loss function, Represents the original image The predicted depth map The depth value is normalized by the mean, Indicates that the original image Input the predicted depth map predicted by the student network, Represents the predicted depth map The mean of and Represents the gradient calculation in the x and y directions respectively, and represents the image gradient, as a weighting factor; Based on the photometric reprojection loss function and the smoothing loss function, a basic loss function is calculated, which is expressed by the following formula: ; in, represents the base loss function, C represents the number of scales predicted by the decoder of the teacher network or the decoder of the student network, represents the weight coefficient of smoothing loss, Represents a predicted depth map A binary matrix with consistent scale. When a pixel point is blocked or out of view during the reprojection process, the corresponding position The value is 0, otherwise the value is 1.
[0016] In some embodiments, after obtaining the lightweight student network model, the training method of the lightweight student network model further includes: Verify the monocular depth estimation performance of the lightweight student network model; If the monocular depth estimation performance is verified to be successful, the lightweight student network model is deployed to the UAV platform; If the monocular depth estimation performance verification fails, adjusting the training parameters of the student network until the monocular depth estimation performance verification passes.
[0017] In a second aspect, the present application provides a monocular depth estimation method for a drone, comprising: The image acquired by the monocular camera of the drone at the current moment is input into the trained lightweight student network model to obtain a depth estimation value of the drone; the lightweight student network model is trained by the training method of the lightweight student network model described in any one of the first aspects.
[0018] In a third aspect, the present application also provides a processing device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors are used to execute the one or more computer programs stored in the memory, so that the one or more processors execute a training method for a lightweight student network model as described in any one of the first aspects and a monocular depth estimation method for a drone as described in the second aspect.
[0019] In a fourth aspect, the present application provides an aircraft, comprising the processing device described in the third aspect.
[0020] Implementing one of the above technical solutions of this application has the following advantages or beneficial effects: The lightweight student network model training method and the monocular depth estimation method for drones in this application, first, use a self-supervised learning method to train the student network model, without relying on labeled data, thereby effectively reducing the cost and time of manual labeling and being able to better adapt to the actual application environment. Secondly, a knowledge distillation method based on the teacher-student network paradigm is adopted to transfer the knowledge of a teacher network with a large number of parameters and strong performance to a student network model with a smaller number of parameters, ensuring that the student network model reduces the number of parameters, computational complexity and storage requirements while maintaining high performance, and also significantly reduces the dependence on hardware computing resources, thereby achieving lightweight model. In addition, an edge extraction algorithm based on deep learning is introduced into the teacher network. By obtaining the edge detection map corresponding to the original image, the teacher network receives the original image and its corresponding edge detection image as input, and predicts the edge depth map and the initial depth map. Subsequently, the edge depth map and the initial depth map are fused using guided filter fusion technology to fully combine the overall depth layout and edge detail information, thereby improving the overall depth estimation accuracy. Subsequently, the fused depth map is used as a guidance signal for the student network to further enhance the depth estimation capability of the student network. Ultimately, efficient and real-time depth perception was achieved on a UAV platform with limited computing resources, which has significant engineering application value and application prospects.
[0021] This application proposes a monocular depth estimation method for drones. While ensuring the accuracy of monocular depth estimation, it effectively reduces the computational complexity of the algorithm model, improves the inference speed, and significantly reduces the dependence on hardware computing resources. The present invention can successfully deploy the algorithm on edge devices with limited computing resources, such as embedded platforms such as drones, thereby enhancing the application potential of monocular depth estimation algorithms in real-time perception and environmental understanding. Compared with the existing technology, the present invention further enhances the monocular depth estimation capability of drones, provides a solid technical foundation for achieving efficient and real-time visual navigation, and has significant engineering application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive work. In the drawings: Figure 1 This is a flow chart of a training method for a lightweight student network model according to an embodiment of the present application; Figure 2 This is a schematic block diagram of a training method for a lightweight student network model according to an embodiment of the present application; Figure 3 This is a schematic diagram of the structure of the teacher network of this application; Figure 4 This is a schematic diagram of the structure of the student network of this application; Figure 5 This is a picture of a post-disaster search and rescue scene taken by a monocular camera of a drone in an embodiment of the present application; Figure 6 This is a monocular depth estimation image of a drone post-disaster search and rescue scenario before the lightweight student network model is adopted in the embodiment of the present application; Figure 7 This is a monocular depth estimation image of a drone post-disaster search and rescue scenario after adopting a lightweight student network model in an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solutions and advantages of the present application clearer, the various exemplary embodiments to be described below may refer to the corresponding drawings, which constitute a part of the exemplary embodiments, which describe various exemplary embodiments that may be used to implement the present application. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation methods described in the following exemplary embodiments do not represent all implementation methods consistent with the present disclosure. It should be understood that they are only examples of processes, methods and devices that are consistent with some aspects disclosed in the present application as detailed in the attached claims, and other embodiments may also be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and essence of the present application.
[0024] In the description of this application, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. The term "plurality" means two or more. The terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, an integral connection, a mechanical connection, an electrical connection, a communication connection, a direct connection, an indirect connection through an intermediate medium, and can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to the specific circumstances.
[0025] In order to illustrate the technical solution described in this application, a specific embodiment is provided below, and only the parts related to the embodiment of this application are shown.
[0026] like Figure 1-2 As shown, the present application provides a training method for a lightweight student network model, and the training method for the lightweight student network model includes the following steps S10 to S50.
[0027] S10: Extracting an edge detection image of the original image, and inputting the edge detection image and the original image into a pre-trained teacher network for inference, respectively, to obtain an edge depth map and an initial depth map.
[0028] The pre-trained teacher network is a teacher network model that has been trained in advance and is a monocular depth estimation teacher model. In some embodiments, the training method of the lightweight student network model may further include: Based on a dataset containing deep information, a teacher network with a large number of parameters is trained to obtain a pre-trained teacher network.
[0029] Specifically, the dataset is one of the most important public datasets in the field of computer vision and autonomous driving. It contains image data from actual driving environments and is widely used to verify depth estimation algorithms. In the monocular depth estimation task, the images of the KITTI dataset are taken from the front camera of a car, covering a variety of complex urban streets, suburbs, and highway scenes. For each image, the KITTI dataset provides a corresponding depth map, which is acquired by a laser scanner and therefore has high-precision real depth values. The resolution of the dataset is 1242×375 and contains 44,931 accurately labeled high-resolution images, including 39,810 training images, 4,424 verification images, and 697 test images. Therefore, the dataset of the embodiment of the present application can be the KITTI dataset, and the dataset is a public dataset of sequential images containing depth information.
[0030] The collected dataset is used to verify the accuracy and real-time performance of the monocular depth estimation algorithm.
[0031] Then, a teacher network T with a large parameter scale is constructed and trained based on the above-collected dataset to obtain a pre-trained teacher network, which has the characteristics of high depth and precision.
[0032] like Figure 3 As shown, Figure 3 This is a schematic diagram of the module structure of the teacher network ResNeXt in the embodiment of the present application. In the embodiment of the present application, ResNeXt101 is used as the feature extraction backbone network during the training of the teacher network T. The feature decoder consists of 5 layers. Each layer fuses the encoder features through progressive upsampling and jump connection, and restores the depth information layer by layer, thereby constructing a teacher network with a large parameter scale. Figure 3 As shown in the figure, the teacher network uses grouped convolution instead of standard convolution, which not only effectively reduces the amount of computation but also improves the feature representation capability. At the same time, ResNeXt introduces a new hyperparameter called cardinality, which represents the number of grouped convolutions in the network. A larger cardinality can increase the learning ability of the network without adding additional computational cost. In ResNeXt, each module does not simply stack multiple convolutional layers, but processes different features by using multiple parallel paths. Each path processes different features through grouped convolution, thereby increasing the network's expressive power. Therefore, increasing the cardinality is a more effective way to improve accuracy. Compared with simply deepening or widening the network, it is more conducive to optimizing network performance.
[0033] The embodiment of the present application adopts a self-supervised approach to train the ResNeXt101 teacher network. During the training process, the Adam optimizer is used to optimize the parameters of the teacher model, where the initial learning rate is 0.0001, the first-order moment estimation coefficient is 0.9, and the second-order moment estimation coefficient is 0.999. By calculating the gradient and updating the weights of the convolution kernel and the fully connected layer in the network, the loss function is gradually minimized. In order to improve the training efficiency and model convergence speed, the embodiment of the present application introduces a learning rate decay strategy to dynamically reduce the learning rate in the later stage of training, thereby avoiding gradient oscillation. At the same time, batch normalization technology is used to normalize the data distribution of each batch, effectively enhancing the stability of the training process. After 20 rounds of iterative training, a teacher network with optimal depth estimation accuracy is obtained, which lays a solid foundation for the subsequent knowledge distillation process.
[0034] After obtaining the pre-trained teacher network, the pre-trained teacher network is used to infer the original image to obtain the initial depth map. The original image can be any image taken by a monocular camera, for example, Figure 2 As shown, Figure 2 The original image is a scene image taken by a monocular camera. For example, Figure 5 As shown, Figure 5 Post-disaster search and rescue scene captured by a drone's monocular camera.
[0035] Moreover, edge detection can be performed on the original image, and the edge detection image can be input into the pre-trained teacher network for inference to obtain the edge depth map.
[0036] In some embodiments, a deep learning-based pixel difference network (PiDiNet) can be used to implement efficient edge detection of the original image, thereby obtaining an edge detection image.
[0037] The network structure of PiDiNet mainly includes an efficient backbone network (Efficient Backbone) and an efficient side structure (Efficient Side Structure). Among them, the backbone network adopts a method of connecting depth-separable convolution and residual blocks in series to gradually extract multi-scale features through 4 stages. Each stage contains 4 residual blocks, and downsampling is performed through the maximum pooling layer between each stage, and the number of channels is doubled step by step. In the embodiment of the present application, the number of channels corresponding to these 4 stages are 60, 120, 240 and 240 respectively. In addition, each residual block contains a depth convolution layer, a ReLU activation layer and a point-by-point convolution layer to ensure that multi-scale edge features are effectively extracted while keeping the model lightweight.
[0038] The edge feature extraction architecture extracts edge feature maps from each stage and calculates an edge loss for deep supervision. To further optimize edge feature extraction, the pixel difference network introduces a compact dilated convolution module after each stage's output to enrich multi-scale edge information, and a compact spatial attention module suppresses background noise. Subsequently, a 1×1 convolutional layer compresses the multi-channel feature map into a single-channel edge map, which is then interpolated to restore the original image size. Finally, a sigmoid activation function is applied to generate an edge detection image.
[0039] The edge map not only provides structural information and object boundary features in the image, but also plays a key role in the depth estimation task. In this embodiment, the edge detection map is input into the pre-trained teacher network. The edge depth map generated after the pre-trained teacher network inference not only contains global depth information, but also retains rich object edge depth details. With the help of the edge information in the edge depth map, the depth estimation network can be further assisted to better understand the edge contours and significantly changed areas in the image, thereby improving the depth estimation accuracy of object boundaries and edge areas, and alleviating the boundary blurring problem common in traditional monocular depth estimation.
[0040] S20: Using the edge depth map as a guide map, fusing the initial depth map and the edge depth map to obtain a fused depth map, and using the fused depth map as a soft label to guide student network training.
[0041] In some embodiments, in step S20, the edge depth map is used as a guide map to fuse the initial depth map and the edge depth map to obtain a fused depth map, which can be expressed by the following formula: ; in, represents the fused depth map, GF represents the guided filtering operation, Representing the initial depth map, including the overall depth information and spatial structure information of the original image; represents the edge depth map, including edge detail information, and serves as a guide map; r represents the radius of the guide filter window, which is used to control the local filtering range; Represents the regularization parameter, which is used to prevent overfitting and ensure the smoothness of the filter.
[0042] Specifically, the embodiment of the present application adopts a depth map fusion method based on guided filtering, which uses the initial depth map obtained by the pre-trained teacher network to predict the original image. Edge depth map predicted by pre-trained teacher network with edge detection image Perform adaptive fusion.
[0043] Different from traditional weighted fusion, guided filtering can be based on edge depth map The gradient change dynamically adjusts the fusion result, maintains the smoothness of the depth in the flat area, and highlights the detail information in the edge area. While enhancing the edge, the guided filter ensures the continuity of the overall depth structure, so that the obtained fused depth map maintains spatial consistency in both edge and non-edge areas.
[0044] By using the initial depth map corresponding to the original image As the input image, the edge depth map corresponding to the edge detection image As a guide map, it can give full play to the guiding role of edge information, guide the edge optimization direction of the input image, and enhance the gradient change of the edge area while ensuring the consistency of the overall depth structure, so that the edges of the fused depth map are clearer and the details are more accurate. The output of the pre-trained teacher network is used as soft labels to guide the distillation training of the student network.
[0045] Furthermore, when using the depth map fusion process of guided filtering, the initial depth map is used as the input image p, the edge depth map is used as the guide image I, and guided filtering is performed. The initial fused image q output by the guided filtering satisfies the local linear model, and the initial fused image for: ; Where i is the pixel index, represents the local window centered at pixel k, and Indicates in the image window The linear coefficient within is solved by minimizing the loss function. The formula for minimizing the loss function is expressed as: ; in, represents the regularization parameter, which is used to suppress overfitting caused by noise.
[0046] Linear coefficient and The closed-form solutions are expressed as follows: ; ; in, and Respectively represent the guide image I in the window The mean and variance within , Indicates that the input image p is in the window The final output is the mean value within (Fusion depth map) is obtained by averaging the results of all overlapping windows, which is expressed by the following formula: ; in, Display window Calculated from all pixels within the range The mean of Display window Calculated from all pixels within the range The mean of .
[0047] When the filter window When it is located at the edge of the guide image I, the linear coefficient Significantly increase, then enhance the initial fusion image The gradient changes in the edge direction, thereby highlighting the edge details; in the filter window When located in the flat area of the guidance image I, the linear coefficient Approaches 0, at this time the initial fusion image In flat areas, local mean filtering is used to achieve smoothing. Through this mechanism, the global structure can be preserved while local edges can be enhanced.
[0048] It can be understood that edge regions refer to locations in an image where pixel values change dramatically, typically corresponding to object outlines, texture details, or structural boundaries. Flat regions refer to areas in an image where pixel values change slowly and with minimal differences, typically corresponding to uniform areas such as object surfaces and backgrounds.
[0049] S30: Inputting the original image into the student network, so that the student network performs self-supervised training under the guidance of the soft label using a teacher-student network paradigm distillation method, and outputs a predicted depth map; S40: After splicing the original image with adjacent frame images, input the spliced images into a posture network to obtain a posture transformation matrix, and perform image reconstruction based on the posture transformation matrix and the predicted depth map to obtain a reconstructed image.
[0050] In some embodiments, step S30 may include: After splicing the original image with adjacent frame images, a spliced image is obtained; Inputting the stitched image into the posture network to obtain a posture transformation matrix; Based on the camera intrinsic parameter matrix, the predicted depth map predicted by the student network for the original image, and the pose transformation matrix, the pixels on the original image are mapped to corresponding pixels in the coordinate system of the adjacent frame image through spatial geometric transformation; The pixel values on the adjacent frame images are sampled to the viewing angle of the original image to obtain a reconstructed image.
[0051] Specifically, the original image is used as the target frame I t , the adjacent frame image adjacent to the original image is taken as the adjacent frame I s After splicing, the spliced image is obtained and sent to the pose network to estimate the adjacent frame I s Relative to target frame I t The camera pose relationship is obtained to obtain the predicted pose transformation matrix, and then the target frame I is transformed into the target frame I by combining the known camera intrinsic parameter matrix K, the predicted depth map and the pose transformation matrix. t Pixel p on the (original image) t , mapped to the adjacent frame I s The corresponding pixel point p in the coordinate system of (adjacent frame images) s , the mapping transformation formula is expressed as follows: ; Among them, K is the camera intrinsic parameter matrix, which includes camera parameters such as focal length and optical center coordinates; is the camera pose transformation matrix from the target frame to the source frame (adjacent frame images), which reflects the translation and rotation relationship between adjacent frames; is the target frame I t At pixel p t The depth estimate of p t Indicates the pixel coordinates on the target frame; p s The pixel coordinates in the source frame are obtained after the mapping transformation. The “~” in the formula refers to the pixel p in the target frame. t After transformation, the pixel position p of the source frame is obtained s It is an approximate relationship, not an absolute equation.
[0052] Furthermore, by sampling the pixel values on the adjacent frames to the perspective of the target frame (original image), the reconstructed image is finally obtained, which is expressed by the following formula: ; Among them, <·> represents the bilinear interpolation sampling operation, which is used to deal with the problem that the mapped pixel coordinates are continuous values rather than discrete integers, thereby ensuring the smoothness and accuracy of the reconstructed image. Finally, the target frame I is reconstructed by calculating s→t (reconstructed image) and target frame I t The photometric reprojection error between the two (original images) measures the difference between the two and forms a self-supervisory signal.
[0053] S50: Calculating the overall loss function of the student network. When the overall loss function is within a preset range, the student network converges to obtain a lightweight student network model. Among them, the overall loss function is the sum of the pairwise loss function, the Laplace distillation loss function, the pixel-by-pixel distillation loss function, and the basic loss function. The basic loss function is the sum of the photometric reprojection loss function and the smoothing loss function between the reconstructed image and the original image. The pixel-by-pixel distillation loss function is the pixel-by-pixel distillation loss function between the fused depth map and the predicted depth map. The pairwise loss function and the Laplace distillation loss function are determined based on the intermediate feature map of the teacher network and the intermediate feature map of the student network.
[0054] Specifically, if Figure 4 As shown, Figure 4 The student network in the embodiment of the present application adopts MobileViT, which has an efficient model architecture and optimization strategy.
[0055] First, MobileViT employs depthwise separable convolutions, which significantly reduce the number of parameters and computational complexity while maintaining the ability to extract local features. Second, the transformer utilizes a lightweight self-attention mechanism to reduce the computationally complex fully connected operations found in traditional transformers, significantly improving computational efficiency. Through these designs, MobileViT achieves performance while significantly reducing the model's computational complexity, making it ideal for resource-constrained devices. Furthermore, the student network leverages the advantages of both, combining the local perception capabilities of convolutional neural networks with the global modeling strengths of transformers. Convolutions excel at capturing information in local regions, particularly in extracting low-level visual features, while transformers excel at processing global information and long-range dependencies, effectively modeling higher-level feature representations. This combination of the two allows the student network to maintain strong feature representation capabilities while remaining lightweight.
[0056] Furthermore, the present embodiment uses a self-supervised learning mechanism to optimize the student network. Specifically, on a public sequence dataset, it fully utilizes the geometric constraints and camera pose information between adjacent video frames, and completes inter-frame image reconstruction through depth prediction and image reprojection techniques, thereby obtaining a self-supervisory signal to guide the training process of the student network.
[0057] Based on the structure of the teacher network and the student network, in order to further improve the effect of knowledge distillation and enhance the ability of the student network to fit the feature representation and distribution of the teacher network, the overall loss function of the student network in the embodiment of the present application is expressed by the following formula: ; in, represents the overall loss function, Represents the basic loss function, which consists of the photometric reprojection loss function and the smoothing loss. represents the pairwise loss of the affinity graph between the teacher network and the intermediate feature graph of the student network, is the Laplace distillation loss, which means that the intermediate feature maps of the teacher network and the student network are decoupled into corresponding scale losses at three scales. is the pixel-wise distillation loss, which represents the logarithmic error loss between the final depth estimates of the teacher network and the student network.
[0058] In some implementations, the basic loss function can be expressed as follows: ; Where C represents the number of scales predicted by the decoder of the teacher network or the decoder of the student network, and L ph represents the photometric reprojection loss, L smRepresents the smoothing loss. λ is the weight coefficient of the smoothing loss. μ is not a simple weight coefficient, but a binary mask matrix generated by Iverson (Iverson bracket), which is used to ignore pixels in occluded areas or beyond the field of view. The purpose of this design is to ensure that the photometric reprojection loss only acts on valid areas. Iverson is a mathematical symbol that can be expressed as [P], where [P] is a Boolean logical expression. The final result value depends on the truth of the condition P. Its calculation formula is expressed as: ; Specifically, μ is a binary mask matrix. When a pixel is not occluded during the reprojection process and is within the field of view, the value of μ at the corresponding position is 1; conversely, when a pixel is occluded or exceeds the field of view during the reprojection process, the value of μ at the corresponding position is 0, thereby ignoring the reconstruction error of these pixels and only calculating the photometric loss in the visible area within the field of view. The calculation formula of μ is expressed as: ; in, Indicates the target frame I t (Original image) and the image after the adjacent frame image is reprojected to the target frame perspective The photometric reprojection error calculated between (reconstructed images); Indicates the target frame I t (original image) and source frame I s The photometric error obtained by directly comparing pixels (of adjacent frames) at their original viewing angles; Indicates that a minimization operation is performed on all source frames (adjacent frame images), that is, the frame with the smallest reprojection error is selected from multiple source frames as the optimal source frame. Smaller than the photometric error without reprojection When μ=1, it means that the pixel is valid and is included in the photometric reprojection loss; when the reprojection error Greater than or equal to the error without reprojection When μ=0, it means that the pixel is invalid and is not included in the loss calculation.
[0059] The photometric reprojection loss formula is: ; in, represents the photometric reprojection loss function, pe represents the photometric reprojection error, and SSIM represents the structural similarity index; Represents the first-order norm of the image and calculates pixel-level differences; represents the weight coefficient, represents the original image (target frame), represents the reconstructed image; The smooth loss function is calculated using the smooth loss formula, which is: ; in, represents the smooth loss function, Represents the depth map The mean-normalized depth value of represents the predicted depth map obtained by inputting the original image into the student network prediction, Represents the predicted depth map The mean of and Represents the gradient calculation in the x and y directions respectively, and represents the image gradient, as a weighting factor; Based on the photometric reprojection loss function and the smoothing loss function, a basic loss function is calculated, which is expressed by the following formula: ; in, represents the base loss function, C represents the number of scales predicted by the decoder of the teacher network or the decoder of the student network, represents the weight coefficient of smoothing loss, Represents a depth prediction map A binary matrix with consistent scale. When a pixel point is blocked or out of view during the reprojection process, the corresponding position The value is 0, otherwise the value is 1.
[0060] The photometric reprojection loss function and the smoothness loss function aim to improve the photometric consistency of the image and the smoothness of the depth estimation.
[0061] In some embodiments, calculating the pairwise loss function may include: Obtaining a teacher intermediate feature map of the original image inferred by the teacher network; Obtaining a student intermediate feature map of the original image inferred by the student network; Based on the teacher's intermediate feature map and the student's intermediate feature map, an affinity map is constructed; wherein the nodes in the affinity map represent different spatial positions, and the connection relationship between two nodes represents similarity; the size of the teacher's intermediate feature map or the student's intermediate feature map is , the affinity map includes nodes, and the number of edges connected to each node is ; Indicates the connection scope; Indicates granularity; Based on the affinity graph, the formula for calculating the pairwise loss function is: ; in, , ; in, and Represents the width and height of the teacher's intermediate feature map or the student's intermediate feature map, respectively, Represents all nodes, represents the similarity between node i and node j in the affinity graph corresponding to the student network, represents the similarity between node i and node j in the affinity graph corresponding to the teacher network, and They represent the aggregated features obtained by average pooling. The aggregation operation means averaging the features of all channels in the node and compressing them into a feature vector of fixed length, thereby extracting the global representation information of the node.
[0062] The pairwise loss function minimizes the difference in affinity graphs between the teacher network and the student network, constraining the student network to more accurately learn the similarity structure of spatial positions from the perspective of pairwise relationships, and is used to guide the student network and the teacher network to maintain consistent feature data distribution between the intermediate feature map pairs.
[0063] In some embodiments, the Laplace distillation loss function L is calculated. Laplacion The calculation formula can be expressed as: ; Among them, C i is the total number of units contained in the features obtained by the i-th decoder layer, N represents the total number of decoder layers. In some embodiments, N=4 can be set to indicate that the decoder has 4 layers.
[0064] x is the spatial position index on the teacher's intermediate feature map or the student's intermediate feature map. For the i-th decoder layer, the Laplacian residual can be expressed as follows: ; ; ; ; Where x represents the spatial position index of the decoder feature of the student network or the teacher network, is the absolute difference of the residuals of the first level features of the Laplacian pyramid, is the absolute difference of the residuals of the second level features of the Laplace pyramid, It is the squared difference between the features of the teacher network and the student network after downsampling to 1 / 4 scale. Represents the Laplace residual, where i represents the decoder layer index, j represents the scale layer index of the Laplace pyramid, and F i,j represents the feature map corresponding to the j-th Gaussian pyramid of the i-th decoder. Up(·) represents an upsampling operation, which is used to adjust the feature map of the j+1th layer to the same spatial resolution as the j-th layer, thereby calculating the Laplacian residual between layers. In this embodiment of the present application, bilinear interpolation is used for upsampling, and the number of layers in the Laplacian pyramid is set to 3.
[0065] The Laplace distillation loss function compares the Laplace residuals of the teacher network and the student network on the decoder features of each layer, and constrains the student network from a fine-grained level to more accurately capture the multi-scale depth feature differences, thereby improving the robustness of depth estimation.
[0066] The Laplace distillation loss function decouples global context information and local detail information in a multi-scale manner to capture the structure and detail information at different levels of the image. By calculating the error between these corresponding decoupled features, it helps students to understand the overall depth layout of the network learning scene, thereby improving the accuracy of depth estimation.
[0067] In some implementations, the pixel-by-pixel distillation loss function is calculated and can be expressed as follows: ; in, represents the pixel-by-pixel distillation loss function, C represents the number of scales predicted by the decoder, and N i represents the total number of pixels at the corresponding scale i, represents the depth prediction value of the teacher network at the i-th scale and the j-th pixel position, represents the depth prediction value of the student network at the jth pixel position at the i-th scale.
[0068] The pixel-by-pixel distillation loss function calculates the difference in depth between the teacher network and the student network’s final output, enabling the student network to accurately imitate the teacher network’s output at each pixel.
[0069] Therefore, the final calculation formula of the overall loss function is expressed as: ; Specifically, the photometric reprojection loss function Based on the principle of photometric consistency, the photometric error between the target frame (original image) and the reprojected frame (reconstructed image) is measured by comparing the pixel differences after the source view is reprojected to the target view, thereby constraining the depth estimation result to be consistent with the photometric relationship under the target view; smoothing loss function By performing edge-aware weighting on the depth map gradient, the depth estimation result is encouraged to remain consistent in the texture smooth area, while changing flexibly in the edge area, thereby reducing the pseudo edges and noise of the depth estimation and ensuring the smoothness of the depth estimation result; Pairwise loss function By calculating the affinity graph difference between the teacher network and the student network at the intermediate feature layer, the student network is guided to learn the feature space structure of the teacher network, thereby realizing knowledge transfer at the feature level and enabling the student network to better fit the intermediate representation of the teacher network; Laplace pyramid loss function Based on the Laplacian pyramid, the feature map is decoupled into global context information and local detail information, and the differences between the teacher network and the student network at different scales are calculated respectively, which is beneficial for the student network to enhance the ability to capture edges and details while maintaining the consistency of the depth layout; pixel-by-pixel distillation loss function By comparing the final output of the teacher network with the final output of the student network pixel by pixel, minimizing the pixel-by-pixel depth difference, and constraining the output of the student network from the pixel level, making it as close as possible to the final predicted distribution of the teacher network, thus achieving end-to-end knowledge transfer.
[0070] In summary, this overall loss function integrates pixel-level, feature-level, and multi-scale depth estimation constraints, effectively guiding the student network to learn the teacher network's knowledge representation at different levels. The joint design of multiple loss terms proposed in this application not only comprehensively improves the robustness and generalization capabilities of the model, but also alleviates problems such as blurred boundaries and noise interference in depth estimation, thereby improving the overall performance of monocular depth estimation.
[0071] In the posture network, the posture network encoder receives two frames of spliced images as input, where the spliced image refers to the current frame image I t and adjacent frame images I s . The pose network estimates the motion of the camera between two adjacent frames, and its final output is the rotation matrix R and translation vector t, which describes the rotation and translation relationship required for the camera from one image to another. Among them, the rotation matrix R is expressed in the form of a 3×3 matrix, and the translation vector t is expressed in the form of a 3×1 column vector. During the training process, the pose network iteratively adjusts the rotation matrix R and the translation vector t by optimizing the aforementioned loss function to ensure that the camera motion transformation can accurately reproject the scene points from one perspective to another, so that the projection positions of the same scene points in different frame images are aligned as much as possible. This method provides geometric constraints for self-supervised monocular depth estimation, thereby achieving accurate modeling of camera motion without real pose labels.
[0072] By reconstructing the image I s→t and the current frame It The photometric reprojection error and smoothing loss are calculated. Combined with the teacher network's distillation guidance, the pixel-by-pixel distillation loss, pairwise loss, and Laplace distillation loss are calculated. These loss terms are combined into a final overall loss function. The student network optimizes this overall loss function, gradually adjusting and optimizing the model's weight parameters to improve its performance on the training data. This training process continues until the overall loss function value remains within a preset range, such as remaining stable, and converges to an optimal state. Ultimately, a small, lightweight student network model is obtained after distillation from the teacher network.
[0073] In some embodiments, after step S50, the training method of the lightweight student network model may further include: Verify the monocular depth estimation performance of the lightweight student network model; If the monocular depth estimation performance is verified to be successful, the lightweight student network model is deployed to the UAV platform; If the monocular depth estimation performance verification fails, adjusting the training parameters of the student network until the monocular depth estimation performance verification passes.
[0074] Specifically, the quantitative indicators of monocular depth estimation performance include absolute relative error (Abs Rel), square relative error (Sq Rel), root mean square error (RMSE), logarithmic root mean square error (RMSE log), and accuracy.
[0075] Among them, the absolute relative error (Abs Rel) can be expressed by the formula: ; The absolute relative error reflects the average relative error between the predicted depth and the true depth. The smaller the value, the more accurate the depth estimation.
[0076] The squared relative error (Sq Rel) can be expressed as: ; The squared relative error focuses more on measuring the impact of larger errors. When the errors of certain pixels are large, the value of this indicator will increase significantly.
[0077] The root mean square error (RMSE) can be expressed as: ; The root mean square error measures the standard deviation of the error between the predicted depth and the true depth and is used to assess the magnitude of the overall error.
[0078] The logarithmic root mean square error (RMSE log) can be expressed as: ; Logarithmic root mean square error is suitable for scenes with a wide range of depth values or drastic dynamic changes, effectively reducing the impact of distant targets on error calculation.
[0079] Accuracy can be expressed as: ; This function measures the percentage of pixels where the relative error between the predicted depth and the true depth is less than a set threshold. Three thresholds are typically used to measure depth estimation accuracy: threshold = 1.25, threshold = 1.25², and threshold = 1.25³. This effectively evaluates the performance of the depth estimation model at different error tolerances and provides a comprehensive assessment of the model's depth estimation accuracy.
[0080] In the above formula, d i and They represent the predicted depth map and the true depth map respectively, and M represents the number of valid pixels in the depth map.
[0081] In summary, these quantitative indicators jointly evaluate the performance of the model from multiple dimensions, including absolute error, relative error, overall error distribution, dynamic range adaptability, and accuracy under different error thresholds, providing a comprehensive and detailed measurement system for the depth estimation task, thereby ensuring that the proposed method has sufficient verification basis in terms of robustness and accuracy.
[0082] When the above indicators are verified to be successful, the lightweight student network model can be deployed on the UAV platform; when the above indicators are not verified to be successful, the parameters of the student network training are adjusted until the indicators are verified to be successful, and the lightweight student network model is deployed on the UAV platform.
[0083] The training method of the lightweight student network model of the embodiment of the present application, first, adopts the self-supervised learning method to train the student network model, without relying on labeled data, thereby effectively reducing the cost and time of manual labeling, and can better adapt to the actual application environment. Secondly, a knowledge distillation method based on the teacher-student network paradigm is adopted to transfer the knowledge of a teacher network with a large number of parameters and strong performance to a student network model with a smaller number of parameters, ensuring that the student network model reduces the number of parameters, calculation amount and storage requirements while maintaining high performance, and also greatly reduces the dependence on hardware computing resources, thereby achieving lightweight model. In addition, an edge extraction algorithm based on deep learning is introduced into the teacher network. By obtaining the edge detection map corresponding to the original image, the teacher network receives the original image and its corresponding edge detection image as input, and predicts the edge depth map and the initial depth map. Subsequently, the edge depth map and the initial depth map are fused using the guided filter fusion technology to fully combine the overall depth layout and edge detail information, thereby improving the overall depth estimation accuracy. Subsequently, the fused depth map is used as a guidance signal for the student network to further enhance the depth estimation capability of the student network. Ultimately, efficient and real-time depth perception was achieved on a UAV platform with limited computing resources, which has significant engineering application value and application prospects.
[0084] The present application also provides a monocular depth estimation method for a drone, which may include: The image acquired by the monocular camera of the drone at the current moment is input into the trained lightweight student network model to obtain the depth estimation value of the drone; the lightweight student network model is trained by the lightweight student network model training method.
[0085] This application proposes a monocular depth estimation method for drones. While ensuring the accuracy of monocular depth estimation, it effectively reduces the computational complexity of the algorithm model, improves the inference speed, and significantly reduces the dependence on hardware computing resources. The present invention can successfully deploy the algorithm on edge devices with limited computing resources, such as embedded platforms such as drones, thereby enhancing the application potential of monocular depth estimation algorithms in real-time perception and environmental understanding. Compared with the existing technology, the present invention further enhances the monocular depth estimation capability of drones, provides a solid technical foundation for achieving efficient and real-time visual navigation, and has significant engineering application value.
[0086] The embodiments of this application were experimentally verified using the public KITTI dataset, including quantitative metrics and time-consuming parameters. The algorithm model of this embodiment has only 1.3M parameters and a computational load of only 0.9 GFLOPs. Compared with the classic Monodepth2 method (which has 14.3M parameters and 8.0 GFLOPs), this embodiment reduces the number of parameters by 90.91% and the computational load by 88.75%. At a threshold of 1.25³, the accuracy of this embodiment can reach 98.3%, with an absolute relative error of only 0.116. To further verify the performance of the algorithm of this embodiment, time-consuming tests were conducted on different hardware platforms. Specifically, on an RTX 3070 Ti computing platform, the inference speed was 652.8 FPS, with a processing time of 1.53 milliseconds. On the resource-constrained drone embedded computing platform Jetson AGX Orin, the inference speed was 183.1 FPS, with a processing time of 5.46 milliseconds, fully meeting real-time requirements.
[0087] In order to further verify the application of the lightweight student network model proposed in the present invention deployed on a drone in the actual post-disaster search and rescue scene, the embodiment of the present application performs depth prediction on the post-disaster search and rescue scene image taken by the drone in the real scene to obtain the depth estimation image in the real scene. Figure 5 The following is a picture of a real post-disaster search and rescue scene taken by a drone's monocular camera. Figure 6 This is the monocular depth estimation image obtained before using the lightweight student network model proposed in this invention. Figure 7 This is the monocular depth estimation image obtained by using the lightweight student network model of the drone proposed by the present invention. Figure 6 and Figure 7 The comparison shows that the present invention has a significant improvement in the overall depth estimation effect. Specifically, for example, Figure 7 The present invention can more accurately capture the gradient changes of the depth edges of the rescuer's head area in the yellow protective suit and the search and rescue dog area, effectively retain and enhance the edge details of the area, and make the depth estimation results clearer in the outline and boundary areas of the target object, and richer in details, thereby comprehensively improving the monocular depth estimation accuracy of the student network in the embodiment of the present application.
[0088] The embodiments of this application effectively improve the accuracy of monocular depth estimation for drones in real-world application scenarios such as post-disaster search and rescue. In particular, they can accurately capture the edges and depth changes of injured and trapped people in complex post-disaster scenarios, providing a more accurate overall depth layout structure, the overall depth distribution of the scene, and edge information for detecting the location, outline, and spatial relationship of rescue personnel with the surrounding environment. This not only helps drones quickly locate survivors in complex environments, but also provides reliable data support for path planning and obstacle avoidance decisions, thereby improving the efficiency and safety of search and rescue missions. This fully verifies the practical value and engineering significance of this invention in actual application scenarios.
[0089] Those skilled in the art will appreciate that all or part of the features / steps of the aforementioned method embodiments may be implemented through methods, data processing systems, or computer programs. These features may be implemented without hardware, entirely through software, or through a combination of hardware and software. The aforementioned computer program may be stored in one or more computer-readable storage media, the storage medium storing the computer program. When the computer program is executed (e.g., by a processor), it executes the steps of the aforementioned embodiments of the quaternion attitude control method and robot motion prediction method based on a hybrid signal storage and computing system for robot motion prediction.
[0090] The aforementioned storage media that can store program codes include: static hard disks, solid-state hard disks, random access memories (SRAM), electrically erasable programmable read-only memories (EEPROM), erasable programmable read-only memories (EPROM), programmable read-only memories (PROM), read-only memories (ROM), optical storage devices, magnetic storage devices, flash memories, magnetic disks or optical disks and / or combinations of the above devices, that is, they can be implemented by any type of volatile or non-volatile storage devices or combinations thereof.
[0091] The present application also provides a processing device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors are used to execute the one or more computer programs stored in the memory, so that the one or more processors execute a lightweight student network model training method and a monocular depth estimation method for a drone as described in any one of the first embodiments.
[0092] The processing device of the present invention can be a chip. When the processing device is a chip, each module in the chip can be fully or partially implemented by software, hardware and a combination thereof. The above modules can be embedded in or independent of the processor in the computing device in hardware form, or can be stored in the memory in the computing device in software form, so that the processor can call and execute the operations corresponding to the above modules. The invention can be applied to fields such as artificial intelligence and metaverse that require high precision, high reliability, low power consumption, and low area cost computing power, artificial intelligence training or reasoning chip products, autonomous driving chips, VR chips, robot built-in chips, and parallel computing deep learning applications. It can be applied to fields such as artificial intelligence and metaverse that require high precision, high reliability, low power consumption, and low area cost computing power, such as artificial intelligence training or reasoning chips, autonomous driving chips, VR chips, robot built-in chips and other fields.
[0093] The present application also provides a drone, which may include a processing device. The drone may be a drone, such as a multi-rotor drone.
[0094] Since the implementation of the drone is described in detail in the first embodiment, it will not be repeated here.
[0095] The foregoing is merely a preferred embodiment of the present invention. Those skilled in the art will appreciate that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. Furthermore, under the guidance of the present invention, these features and embodiments may be modified to suit specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are intended to be within the scope of the present invention.
Claims
1. A training method for a lightweight student network model, characterized in that: include: Extracting an edge detection image of the original image, inputting the edge detection image and the original image into a pre-trained teacher network for inference, and obtaining an edge depth map and an initial depth map; Using the edge depth map as a guide map, fusing the initial depth map and the edge depth map to obtain a fused depth map, and using the fused depth map as a soft label to guide student network training; Inputting the original image into the student network, so that the student network performs self-supervised training under the guidance of the soft label using a teacher-student network paradigm distillation method, and outputs a predicted depth map; After splicing the original image with adjacent frame images, the spliced images are input into a posture network to obtain a posture transformation matrix, and image reconstruction is performed based on the posture transformation matrix and the predicted depth map to obtain a reconstructed image; Calculating the overall loss function of the student network; when the overall loss function is within a preset range, the student network converges to obtain a lightweight student network model; Among them, the overall loss function is the sum of the pairwise loss function, the Laplace distillation loss function, the pixel-by-pixel distillation loss function, and the basic loss function. The basic loss function is the sum of the photometric reprojection loss function and the smoothing loss function between the reconstructed image and the original image. The pixel-by-pixel distillation loss function is the pixel-by-pixel distillation loss function between the fused depth map and the predicted depth map. The pairwise loss function and the Laplace distillation loss function are determined based on the teacher intermediate feature map of the teacher network and the student intermediate feature map of the student network.
2. The training method of the lightweight student network model according to claim 1 is characterized in that: The edge depth map is used as a guide map, and the initial depth map and the edge depth map are fused to obtain a fused depth map, which is expressed by the following formula: ; in, represents the fused depth map, GF represents the guided filtering operation, Representing the initial depth map, including the overall depth information and spatial structure information of the original image; represents the edge depth map, including edge detail information, and serves as a guide map; r represents the radius of the guide filter window, which is used to control the local filtering range; Represents the regularization parameter, which is used to prevent overfitting and ensure the smoothness of the filter.
3. The training method of the lightweight student network model according to claim 2 is characterized in that: The guided filtering operation includes: Using the initial depth map as an input image and the edge depth map as a guide image, performing guided filtering, and outputting an initial fused image that satisfies a local linear model; When the filter window is located in the edge area of the guide image, the gradient change of the initial fused image in the edge direction is enhanced to highlight the edge details; When the filtering window is located in a flat area of the guide image, local mean filtering is performed on the corresponding flat area.
4. The training method of the lightweight student network model according to claim 1 is characterized in that: After splicing the original image with the adjacent frame images, the splicing is input into a posture network to obtain a posture transformation matrix, and image reconstruction is performed based on the posture transformation matrix and the predicted depth map to obtain a reconstructed image, including: After splicing the original image with adjacent frame images, a spliced image is obtained; Inputting the stitched image into the posture network to obtain a posture transformation matrix; Based on the camera intrinsic parameter matrix, the predicted depth map and the pose transformation matrix, mapping the pixel points on the original image to the corresponding pixel points in the coordinate system of the adjacent frame image through spatial geometric transformation; The pixel values on the adjacent frame images are sampled to the viewing angle of the original image to obtain a reconstructed image.
5. The training method of the lightweight student network model according to claim 1 is characterized in that: The calculating of the overall loss function of the student network comprises: Calculating the basic loss function, where the basic loss function is a photometric reprojection loss function and a smoothing loss function between the reconstructed image and the original image; Calculating a pixel-by-pixel distillation loss function between the fused depth map and the predicted depth map; Calculating a pairwise loss function and the Laplace distillation loss function between the teacher intermediate feature map of the teacher network and the student intermediate feature map of the student network; Based on the photometric reprojection loss function, the smoothing loss function, the pixel-by-pixel distillation loss function, the pairwise loss function, and the Laplace distillation loss function, the overall loss function is calculated using the following formula; ; in, represents the overall loss function, represents the basic loss function, which is the sum of the photometric reprojection loss function and the smoothing loss function. represents the affinity pairwise loss function between the teacher intermediate feature map of the teacher network and the student intermediate feature map of the student network, represents the Laplace distillation loss function, represents the pixel-wise distillation loss function.
6. The training method of the lightweight student network model according to claim 5 is characterized in that: The calculating a pairwise loss function between the teacher intermediate feature map of the teacher network and the student intermediate feature map of the student network includes: Based on the teacher's intermediate feature map and the student's intermediate feature map, an affinity map is constructed; wherein the nodes in the affinity map represent different spatial positions, and the connection relationship between two nodes represents similarity; the size of the teacher's intermediate feature map or the student's intermediate feature map is , the affinity map includes nodes, and the number of edges connected to each node is ; Indicates the connection scope; Indicates granularity; Based on the affinity graph, the formula for calculating the pairwise loss function is: ; in, , ; in, and Represents the width and height of the teacher's intermediate feature map or the student's intermediate feature map, respectively, Represents all nodes, represents the similarity between node i and node j in the affinity graph corresponding to the student network, represents the similarity between node i and node j in the affinity graph corresponding to the teacher network, and They represent the aggregated features obtained by average pooling. The aggregation operation means averaging the features of all channels in the node and compressing them into a feature vector of fixed length, thereby extracting the global representation information of the node. The Laplace distillation loss function between the teacher intermediate feature map of the teacher network and the student intermediate feature map of the student network is calculated and expressed as follows: ; Among them, C i is the total number of units contained in the features obtained by the decoder of the teacher network or the i-th decoder layer of the student network, N represents the total number of decoder layers of the teacher network or the student network, x is the spatial position index on the teacher's intermediate feature map, Represents the absolute difference of the residual of the first layer of the Laplacian pyramid feature, Represents the absolute difference of the residual of the second layer of the Laplacian pyramid feature, It is the squared difference between the features of the teacher network and the student network after downsampling to 1 / 4 scale.
7. The training method of the lightweight student network model according to claim 5 is characterized in that: The pixel-by-pixel distillation loss function between the fused depth map and the predicted depth map is calculated using the following formula: ; in, represents the pixel-wise distillation loss function, C represents the number of scales predicted by the decoder of the teacher network or the decoder of the student network, and N i represents the number of pixels at the i-th scale, represents the depth prediction value of the teacher network at the i-th scale and the j-th pixel point, Represents the depth prediction value of the student network at the jth pixel at the i-th scale.
8. The training method of the lightweight student network model according to claim 5 is characterized in that: The calculating of the basic loss function includes: The photometric reprojection loss function is calculated using the photometric reprojection loss formula, which is: ; in, represents the photometric reprojection loss function, pe represents the photometric reprojection error, and SSIM represents the structural similarity index; Represents the first-order norm of the image and calculates pixel-level differences; represents the weight coefficient, represents the original image, represents the reconstructed image; The smooth loss function is calculated using the smooth loss formula, which is: ; in, represents the smooth loss function, Represents the original image The predicted depth map The depth value is normalized by the mean, Indicates that the original image Input the predicted depth map predicted by the student network, Represents the predicted depth map The mean of and Represents the gradient calculation in the x and y directions respectively, and represents the image gradient, as a weighting factor; Based on the photometric reprojection loss function and the smoothing loss function, a basic loss function is calculated, which is expressed by the following formula: ; in, represents the base loss function, C represents the number of scales predicted by the decoder of the teacher network or the decoder of the student network, represents the weight coefficient of smoothing loss, Represents a predicted depth map A binary matrix with consistent scale. When a pixel point is blocked or out of view during the reprojection process, the corresponding position The value is 0, otherwise the value is 1.
9. The training method of the lightweight student network model according to claim 1 is characterized in that: After obtaining the lightweight student network model, the training method of the lightweight student network model further includes: Verify the monocular depth estimation performance of the lightweight student network model; If the monocular depth estimation performance is verified to be successful, the lightweight student network model is deployed to the UAV platform; If the monocular depth estimation performance verification fails, adjusting the training parameters of the student network until the monocular depth estimation performance verification passes.
10. A monocular depth estimation method for an unmanned aerial vehicle, characterized in that: include: The image captured by the drone's monocular camera at the current moment is input into the trained lightweight student network model to obtain the depth estimation value of the drone; The lightweight student network model is obtained by training using the lightweight student network model training method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Method and apparatus for monocular depth estimation
CN115393409A
Cited By
Track smooth control method, device and equipment for mechanical arm and storage medium
CN121893288A
Trajectory smoothing control method, device and equipment of mechanical arm and storage medium
CN121893288B
A low-altitude unmanned aerial vehicle cross-field re-identification method fusing observation space prior
CN122530883B