Image depth estimation method, system and equipment based on knowledge mutual distillation and medium

By building a training framework based on knowledge distillation, integrating the advantages of single-frame, multi-frame and stereoscopic depth estimation models, and using the scene's inherent geometric information for depth map recovery and alignment, the problem that depth estimation methods in the prior art cannot integrate image representation, temporal information and spatial information in complex scenarios, and realizing the independence and accuracy of the model.

CN120259397AActive Publication Date: 2025-07-04HUZHOU COLLEGE
View PDF 15 Cites 0 Cited by

Patent Information

Application Number
CN202510393503.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-04
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing single-frame, multi-frame and stereo-method depth estimation methods are difficult to integrate the advantages of image representation, temporal information and spatial information in complex scenarios, resulting in the inability to accurately predict the depth of dynamic objects, lighting change scenes and occlusion areas.

Method used

A training framework based on knowledge mutual distillation is constructed, and the deep knowledge mutual distillation between single-frame, multi-frame and stereo depth estimation models are integrated, and the depth map recovery and alignment are used to construct a knowledge mutual distillation model for training.

Benefits of technology

The advantages of single-frame, multi-frame and stereoscopic depth estimation models in image representation, temporal information and spatial information are realized, which ensures the independence and accuracy of the model and solves the instability of depth estimation in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259397A_ABST
    Figure CN120259397A_ABST
Patent Text Reader

Abstract

The invention discloses an image depth estimation method, system and equipment based on knowledge mutual distillation and a medium, and relates to the technical field of computer vision, the method comprises the following steps: collecting a single-frame image, a multi-frame image and a three-dimensional image, and respectively obtaining corresponding depth maps; the pixels of the single-frame image are converted into a three-dimensional space, the scale between the inherent geometric plane and the space points corresponding to the pixels is calculated, and a depth map with the recovered scale is obtained according to the scale and the depth maps of the single-frame image and the multi-frame image; aligning a single-frame depth map, a multi-frame depth map and a depth map estimated by the stereoscopic model; reconstructing pixels by using depth values aligned with corresponding scales of a single frame, multiple frames and the three-dimensional depth estimation model, and constructing a knowledge mutual distillation model according to the original pixels and the reconstructed pixels; training a depth estimation model based on a knowledge mutual distillation model, and performing depth estimation on the to-be-detected image according to the trained model; according to the method, the advantages of a single-frame model, a multi-frame model and a three-dimensional depth estimation model are fused, and the independence of a single model is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly relates to an image depth estimation method, system, device and medium based on knowledge mutual distillation. Background Art

[0002] Depth estimation is of great significance in the fields of computer vision and robotics, mainly used to obtain the distance information from objects or surfaces in the scene to the camera. Through depth estimation, the system can understand the three-dimensional structure of the scene, so as to achieve more accurate object recognition, path planning, environmental perception and other tasks. Therefore, the development of efficient and accurate depth estimation technology is crucial for improving the intelligence level of artificial intelligence systems.

[0003] Existing depth estimation methods include single-frame, multi-frame and stereo types: The single-frame method estimates depth by analyzing a single image, and has a strong perception ability for image representation information; The multi-frame method uses multiple images from different time points as input, and can strongly perceive time information, so as to analyze dynamic scenes; The stereo method relies on the disparity between two images taken simultaneously from two different viewpoints, and can provide highly accurate depth information, which is suitable for depth estimation of static scenes.

[0004] However, in the actual engineering scene, it usually contains both dynamic objects and static objects, and requires a strong ability to understand representation information. Although the above three types of methods have their own advantages when used in their respective corresponding scenes, in the actual scene usage process, their advantages are difficult to integrate, resulting in the inability to accurately predict the depth of dynamic objects, scenes with light changes, and occluded areas, thus limiting their usability. Summary of the Invention

[0005] Aiming at the deficiency that the advantages of the single-frame method, multi-frame method and stereo method in image representation, time information and space information are difficult to integrate and cannot meet complex scenes in the prior art, the present invention proposes an image depth estimation method, system, device and medium based on knowledge mutual distillation. By constructing a training framework for realizing depth knowledge mutual distillation among the single-frame depth estimation model, multi-frame depth estimation model and stereo depth estimation model, the advantages of the single-frame depth estimation model, multi-frame depth estimation model and stereo depth estimation model are integrated, so as to solve the problems existing in the prior art.

[0006] An image depth estimation method based on knowledge mutual distillation includes the following steps:

[0007] Collect single-frame images, multi-frame images and stereo images; Based on the single-frame depth estimation model, multi-frame depth estimation model and stereo depth estimation model, obtain the depth maps corresponding to the single-frame images, multi-frame images and stereo images respectively;

[0008] Convert the initial pixels of a single-frame image into three-dimensional space, calculate the scale between the scene's inherent geometric plane and the spatial point corresponding to the pixel to obtain the scene depth scale, and obtain the depth maps with restored scales according to the scene depth scale and the depth maps corresponding to the single-frame image and the multi-frame image respectively; align the depth maps corresponding to the single-frame image, the multi-frame image, and the stereo image based on the depth maps with restored scales; reconstruct the initial pixels using the depth maps of the single-frame image, the multi-frame image, and the stereo image after scale alignment respectively; construct a knowledge mutual distillation model between the single-frame depth estimation model, the multi-frame depth estimation model, and the stereo depth estimation model according to the initial pixels and the reconstructed pixels; train the single-frame depth estimation model, the multi-frame depth estimation model, and the stereo depth estimation model based on the knowledge mutual distillation model;

[0009] Perform depth estimation on the image to be measured according to the trained single-frame depth estimation model, multi-frame depth estimation model, and stereo depth estimation model.

[0010] Further, the obtaining of the depth maps corresponding to the single-frame image, the multi-frame image, and the stereo image respectively based on the single-frame depth estimation model, the multi-frame depth estimation model, and the stereo depth estimation model specifically includes:

[0011] Construct a single-frame depth estimation model and input the single-frame image I a to obtain the depth map

[0012] Construct a multi-frame depth estimation model and input the multi-frame image {I a , I a+1 , I a+2 , …, I a+n} to obtain the depth map

[0013] Construct a stereo depth estimation model and input the stereo image pair to obtain the depth map

[0014] Further, the conversion of the single-frame image pixels into three-dimensional space, the calculation of the scale between the scene's inherent geometric plane and the spatial point corresponding to the pixel to obtain the scene depth scale is expressed as:

[0015] Convert the pixel p a in the image I i (u, v) into the three-dimensional space coordinates P i (X, Y, Z), and the specific conversion process is expressed as:

[0016]

[0017] where M1 is the internal parameter matrix of the camera;

[0018] Select pixel p i and its eight adjacent points, divide them into four groups, and transform them into three-dimensional space;

[0019] According to the transformed three-dimensional space coordinates and pixel p i , construct four planes; average the normal vectors of the four planes to obtain the normal vector n(p i ) of pixel p i . The specific process is expressed as:

[0020]

[0021] where n j represents the normal vector of the j-th group of planes;

[0022] Determine whether the pixel is a valid pixel: If pixel p i is consistent with the inherent geometric information, then the pixel is a valid pixel. The determination process is expressed as:

[0023]

[0024] where is the normal vector of the inherent geometric plane, and τ is the angle threshold between the pixel normal vector and the inherent geometric plane normal vector;

[0025] Calculate the scale between the inherent geometric plane and the spatial point corresponding to the pixel, that is, the scene depth scale f s . The specific expression is:

[0026]

[0027] where n represents the number of valid pixels, h i represents the inherent geometric size corresponding to the pixel, and Y i represents the three-dimensional coordinate size obtained by solving the pixel.

[0028] Furthermore, according to the scene depth scale and the depth maps corresponding to single-frame images and multi-frame images, depth maps with restored scales are obtained respectively; the depth maps with restored scales are expressed as:

[0029]

[0030] where represents the depth map with restored scale, and D represents the relative depth map directly output by the model.

[0031] Furthermore, construct a knowledge mutual distillation model between the single-frame depth estimation model, the multi-frame depth estimation model, and the stereo depth estimation model according to the initial pixel and the reconstructed pixel, which specifically includes the following steps:

[0032] The depth values of the depth maps after scale alignment of single-frame images, multi-frame images, and stereo images are respectively used and to reconstruct pixel p i to obtain the reconstructed pixel and The specific process is expressed as:

[0033]

[0034] where Φ(·) is the reconstruction function, is the pixel reconstructed based on the depth value , k ∈ {1, 2, 3} represents the corresponding depth prediction method, is the pixel in the reference image, and P ab is the pose between the target image and the reference image;

[0035] A quantitative evaluation of the depth is constructed, expressed as:

[0036]

[0037] The teacher pixel and the student pixel are generated, expressed as:

[0038]

[0039] where T is the teacher pixel, and Stu1 and Stu2 are the student pixels, and are respectively the quantitative evaluations of the depth and .

[0040] The present invention further includes an image depth estimation system based on knowledge mutual distillation, including:

[0041] An acquisition module for acquiring single-frame images, multi-frame images, and stereo images; based on a single-frame depth estimation model, a multi-frame depth estimation model, and a stereo depth estimation model, respectively obtaining the depth maps corresponding to the single-frame images, multi-frame images, and stereo images;

[0042] The construction module of the knowledge mutual distillation model is used to convert the initial pixels of a single-frame image into three-dimensional space, calculate the scale between the inherent geometric plane of the scene and the spatial points corresponding to the pixels, obtain the scene depth scale, and respectively obtain the depth maps with restored scales according to the scene depth scale and the depth maps corresponding to the single-frame image and the multi-frame image; align the depth maps corresponding to the single-frame image, the multi-frame image, and the stereo image according to the depth maps with restored scales; respectively reconstruct the initial pixels by using the depth maps of the single-frame image, the multi-frame image, and the stereo image after scale alignment; construct a knowledge mutual distillation model between the single-frame depth estimation model, the multi-frame depth estimation model, and the stereo depth estimation model according to the initial pixels and the reconstructed pixels; train the single-frame depth estimation model, the multi-frame depth estimation model, and the stereo depth estimation model based on the knowledge mutual distillation model;

[0043] The depth estimation module is used to perform depth estimation on the image to be measured according to the trained single-frame depth estimation model, multi-frame depth estimation model, and stereo depth estimation model.

[0044] The present invention also includes an image depth estimation computer device based on knowledge mutual distillation, including: a memory, a processor, and a computer program stored in the memory. When the processor executes the computer program, the steps of the image depth estimation method based on knowledge mutual distillation are implemented.

[0045] The present invention also includes a readable storage medium. The readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the steps of the image depth estimation method based on knowledge mutual distillation are used to be executed.

[0046] The present invention provides an image depth estimation method, system, device, and medium based on knowledge mutual distillation, having the following beneficial effects:

[0047] The present invention constructs a training framework for realizing depth knowledge mutual distillation between single-frame, multi-frame, and stereo depth estimation models, and realizes the synchronous training of single-frame, multi-frame, and stereo depth estimation models; based on the inherent geometric information of the scene, the scales of the output depth maps of the single-frame and multi-frame depth estimation models are restored, avoiding the problem of inconsistent scales of the depth maps output by single-frame, multi-frame, and stereo models; a knowledge mutual distillation model between the single-frame depth estimation model, the multi-frame depth estimation model, and the stereo depth estimation model is constructed for training, avoiding the unidirectional flow of knowledge in the traditional knowledge distillation process, so that the single-frame, multi-frame, and stereo depth estimation models can be both student models and teacher models; only the dominant knowledge flows between the models during training, and the mutual interference of the models is avoided, thereby realizing the advantage fusion of the single-frame, multi-frame, and stereo depth estimation models in terms of image representation, temporal information, and spatial information respectively, and ensuring the independence of a single model. Brief Description of the Drawings

[0048] Figure 1 It is a framework diagram of neural network training based on knowledge mutual distillation in an embodiment of the present invention;

[0049] Figure 2 It is a depth map scale recovery and alignment diagram based on the inherent geometric information of the scene in an embodiment of the present invention;

[0050] Figure 3 It is a structural diagram of the depth knowledge mutual distillation model in an embodiment of the present invention. Detailed Embodiments

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0052] The present invention proposes an image depth estimation method based on knowledge mutual distillation, which can integrate the advantages of single-frame, multi-frame, and stereo depth estimation models in image representation, temporal information, and spatial information respectively, and obtain independent single-frame, multi-frame, and stereo depth estimation models. As Figure 1 shown, the model training process specifically includes three independent depth estimation models, namely a single-frame depth estimation model, a multi-frame depth estimation model, and a stereo depth estimation model; after the three models independently complete depth estimation, scale recovery and alignment will be performed based on the inherent geometric information of the scene; finally, a knowledge mutual distillation loss function is constructed to supervise the training of the single-frame, multi-frame, and stereo models. The method specifically includes the following steps:

[0053] S1. Construct a depth knowledge mutual distillation neural network training framework, including the following specific steps:

[0054] a) Construct a single-frame depth estimation model, input a single-frame image I a , and obtain a depth map

[0055] b) Construct a multi-frame depth estimation model, input multi-frame images {I a , I a+1 , I a+2 , …, I a+n}, and obtain a depth map

[0056] c) Construct a stereo depth estimation model, input a stereo image pair and obtain a depth map

[0057] d) Construct a depth knowledge mutual distillation module to realize the exploration and mutual transfer of the advantageous depth knowledge of the single-frame, multi-frame, and stereo models.

[0058] S2. Single-frame, multi-frame, and stereo model estimation depth map alignment based on the inherent geometric information of the scene (such as the height of the camera from the ground, the width of the lane lines, the width of the lane, etc.), as Figure 2 shown; including the following specific steps:

[0059] a) For the pixel p a in the image I i (u, v), convert it into three-dimensional space coordinates P i (X, Y, Z). The specific conversion process can be expressed as:

[0060]

[0061] In the formula, M1 is the internal parameter matrix of the camera, which is composed of the inherent parameters of the camera.

[0062] b) Select eight adjacent points of the pixel p i and divide them into four groups. Convert them into three-dimensional space according to Equation (1). Combine the pixel p i , and four planes can be formed. The normal vector n(p i ) describing the pixel p i can be obtained by averaging the normal vectors of the four planes. The specific process can be expressed as:

[0063]

[0064] In the formula, n j represents the normal vector of the jth group of planes.

[0065] c) If the pixel p i is consistent with the inherent geometric information, then this pixel is a valid pixel. The process of determining a valid pixel can be expressed as:

[0066]

[0067] In the formula, is the normal vector of the inherent geometric plane, and τ is the angle threshold between the pixel normal vector and the inherent geometric plane normal vector.

[0068] d) The scene depth scale f s is the scale between the inherent geometric plane and the spatial point corresponding to the pixel. The specific process can be expressed as:

[0069]

[0070] In the formula, n represents the number of valid pixels, h i represents the inherent geometric size corresponding to this pixel, and Y i represents the three-dimensional coordinate size obtained by solving this pixel.

[0071] e) Restore the scale information of the depth estimated by the single-frame and multi-frame models. The specific process can be expressed as follows:

[0072]

[0073] In the formula, represents the depth map for restoring the scale, and D represents the relative depth map directly output by the single-frame or multi-frame model.

[0074] f) Align the depth maps estimated by the single-frame, multi-frame, and stereo models.

[0075] S3. Construct a knowledge mutual distillation loss function to supervise the training of the single-frame, multi-frame, and stereo models, as Figure 3 shown; it includes the following specific steps:

[0076] a) Use the depth values after restoring the scale by the single-frame and multi-frame methods and as well as the depth value output by the stereo model to reconstruct the pixel p i , and obtain the reconstructed pixels and The specific process can be expressed as follows:

[0077]

[0078] In the formula, Φ(·) is the reconstruction function, is the pixel reconstructed based on the depth value , k ∈ {1, 2, 3} represents the corresponding depth prediction method, is the pixel in the reference image, and P ab is the pose between the target image and the reference image;

[0079] b) Construct a quantitative evaluation of the depth. The specific process can be expressed as follows:

[0080]

[0081] c) Identify the teacher pixels and student pixels. The specific process can be expressed as follows:

[0082]

[0083] In the formula, T is the teacher pixel, and Stu1 and Stu2 are the student pixels. and are the quantitative evaluations of the depths and respectively.

[0084] The present invention provides an image depth estimation method based on knowledge mutual distillation. By leveraging the advantages of single-frame, multi-frame, and stereo depth estimation models in image representation, temporal information, and spatial information respectively, a pixel-based knowledge distillation mechanism is established to achieve knowledge mutual distillation among the three methods. Ultimately, a single model can simultaneously possess the advantages in image representation, temporal information, and spatial information, solving the problem that existing methods perform unstably in engineering practical scenarios that contain both dynamic and static objects. The present invention realizes the synchronous training of single-frame, multi-frame, and stereo depth estimation models by constructing a model training framework based on knowledge mutual distillation; constructs a depth scale recovery method based on scene inherent information to achieve scale alignment of the depth maps output by single-frame, multi-frame, and stereo depth estimation models; constructs a pixel-based depth quantitative evaluation criterion to achieve knowledge mutual distillation among single-frame, multi-frame, and stereo depth estimation models; based on knowledge mutual distillation, the single-frame, multi-frame, and stereo depth estimation models increase mutual constraints during the training process, realizing knowledge selection and mutual learning to achieve complementary advantages. Compared with the prior art, the present invention solves the problem that it is difficult to integrate the advantages of single-frame, multi-frame, and stereo depth estimation models in image representation, temporal information, and spatial information respectively.

[0085] Based on the same inventive concept, the present invention also proposes an image depth estimation system based on knowledge mutual distillation, including:

[0086] An acquisition module, configured to acquire single-frame images, multi-frame images, and stereo images; and respectively obtain depth maps corresponding to the single-frame images, multi-frame images, and stereo images based on a single-frame depth estimation model, a multi-frame depth estimation model, and a stereo depth estimation model.

[0087] A construction module of the knowledge mutual distillation model, configured to convert the initial pixels of the single-frame image into three-dimensional space, calculate the scale between the scene inherent geometric plane and the spatial point corresponding to the pixel to obtain the scene depth scale, and respectively obtain the depth maps with restored scales according to the scene depth scale and the depth maps corresponding to the single-frame image and the multi-frame image; align the depth maps corresponding to the single-frame image, multi-frame image, and stereo image according to the depth maps with restored scales; respectively reconstruct the initial pixels by using the depth maps of the single-frame image, multi-frame image, and stereo image after scale alignment; construct a knowledge mutual distillation model among the single-frame depth estimation model, multi-frame depth estimation model, and stereo depth estimation model according to the initial pixels and the reconstructed pixels; and train the single-frame depth estimation model, multi-frame depth estimation model, and stereo depth estimation model based on the knowledge mutual distillation model.

[0088] A depth estimation module, configured to perform depth estimation on the image to be measured according to the trained single-frame depth estimation model, multi-frame depth estimation model, and stereo depth estimation model.

[0089] The present invention also provides a computer device for image depth estimation based on knowledge mutual distillation, including: a memory, a processor, and a computer program stored in the memory. When the processor executes the computer program, the steps of the method for image depth estimation based on knowledge mutual distillation are implemented.

[0090] The present invention also provides a readable storage medium storing a computer program, where the computer program includes program instructions. When the program instructions are executed by a processor, they are used to execute the steps of the method for image depth estimation based on knowledge mutual distillation.

[0091] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.

Claims

1. An image depth estimation method based on knowledge mutual distillation, characterized in that Including the following steps: Collect single-frame images, multi-frame images, and stereo images; based on a single-frame depth estimation model, a multi-frame depth estimation model, and a stereo depth estimation model, obtain the depth maps corresponding to the single-frame images, multi-frame images, and stereo images respectively; Convert the initial pixels of the single-frame image into three-dimensional space, calculate the scale between the scene-intrinsic geometric plane and the spatial point corresponding to the pixel to obtain the scene depth scale, and obtain the depth maps with restored scales according to the scene depth scale and the depth maps corresponding to the single-frame image and the multi-frame image respectively; Align the depth maps corresponding to the single-frame image, multi-frame image, and stereo image according to the depth maps with restored scales; Respectively use the depth maps with aligned scales of the single-frame image, multi-frame image, and stereo image to reconstruct the initial pixels; According to the initial pixels and the reconstructed pixels, construct a knowledge mutual distillation model between the single-frame depth estimation model, the multi-frame depth estimation model, and the stereo depth estimation model; train the single-frame depth estimation model, the multi-frame depth estimation model, and the stereo depth estimation model based on the knowledge mutual distillation model; Perform depth estimation on the image to be measured according to the trained single-frame depth estimation model, multi-frame depth estimation model, and stereo depth estimation model.

2. The method for image depth estimation based on knowledge mutual distillation according to claim 1, wherein The step of respectively obtaining the depth maps corresponding to the single-frame image, multi-frame image, and stereo image based on the single-frame depth estimation model, multi-frame depth estimation model, and stereo depth estimation model specifically includes: Build a single-frame depth estimation model and input a single-frame image I a , and obtain a depth map Construct a multi-frame depth estimation model, and input multi-frame images {I a , I a+1 , I a+2 , …, I a+n}, and obtain a depth map Build a stereo depth estimation model and input a stereo image pair Obtain a depth map 3. The method for image depth estimation based on knowledge mutual distillation according to claim 2, characterized in that, The step of converting the single-frame image pixels into three-dimensional space, calculating the scale between the scene-intrinsic geometric plane and the spatial point corresponding to the pixel to obtain the scene depth scale, which is expressed as: Convert the pixel p a in the image I i (u, v) into three-dimensional space coordinates P i (X, Y, Z). The specific conversion process is expressed as: where M1 is the internal parameter matrix of the camera; Select pixel p i and its eight adjacent points, divide them into four groups, and transform them into three-dimensional space; According to the converted three-dimensional space coordinates and pixel p i , four planes are formed; the average of the normal vectors of the four planes is calculated to obtain the normal vector n(p i ) of pixel p i ). The specific process is expressed as: where n j represents the normal vector of the j-th group of planes; Determine whether the pixel is a valid pixel: If pixel p i is consistent with the inherent geometric information, then the pixel is a valid pixel, and the determination process is expressed as: wherein, is the normal vector of the intrinsic geometric plane, and τ is the angle threshold between the pixel normal vector and the normal vector of the intrinsic geometric plane; Calculate the scale between the intrinsic geometric plane and the spatial point corresponding to the pixel, that is, the scene depth scale f s , which is specifically expressed as: Among them, n represents the number of effective pixels, and h i represents the inherent geometric size corresponding to this pixel, and Y i represents the three-dimensional coordinate size obtained by solving this pixel.

4. A method for image depth estimation based on knowledge mutual distillation according to claim 3, characterized in that, The step of respectively obtaining the depth maps with restored scales according to the scene depth scale and the depth maps corresponding to the single-frame image and the multi-frame image; the depth maps with restored scales are expressed as: Among them, represents the depth map for restoring the scale, and D represents the relative depth map directly output by the model.

5. A method for image depth estimation based on knowledge mutual distillation according to claim 4, characterized in that, The step of constructing a knowledge mutual distillation model between the single-frame depth estimation model, multi-frame depth estimation model, and stereo depth estimation model according to the initial pixels and the reconstructed pixels specifically includes the following steps: The depth values of the depth maps after scale alignment of single-frame images, multi-frame images, and stereo images are respectively used and to reconstruct pixel p i , and the reconstructed pixel is obtained The specific process is as follows: where, Φ(·) is the reconstruction function, is the pixel reconstructed according to the depth value k ∈ {1, 2, 3} represents the corresponding depth prediction method, is the pixel in the reference image, P ab is the pose between the target image and the reference image; Construct a quantitative evaluation of depth, which is expressed as: Generate teacher pixels and student pixels, which are expressed as: Among them, T is the teacher pixel, and Stu1 and Stu2 are student pixels. and are respectively the quantitative evaluations of the depth and respectively.

6. An image depth estimation system based on knowledge mutual distillation, characterized in that, Including: An acquisition module for collecting single-frame images, multi-frame images, and stereo images; based on a single-frame depth estimation model, a multi-frame depth estimation model, and a stereo depth estimation model, respectively obtaining the depth maps corresponding to the single-frame images, multi-frame images, and stereo images; A construction module of the knowledge mutual distillation model for converting the initial pixels of the single-frame image into three-dimensional space, calculating the scale between the scene-intrinsic geometric plane and the spatial point corresponding to the pixel to obtain the scene depth scale, and obtaining the depth maps with restored scales according to the scene depth scale and the depth maps corresponding to the single-frame image and the multi-frame image respectively; Align the depth maps corresponding to the single-frame image, multi-frame image, and stereo image according to the depth maps with restored scales; Respectively use the depth maps with aligned scales of the single-frame image, multi-frame image, and stereo image to reconstruct the initial pixels; Construct a knowledge mutual distillation model among a single-frame depth estimation model, a multi-frame depth estimation model, and a stereo depth estimation model based on the initial pixels and the reconstructed pixels; train the single-frame depth estimation model, the multi-frame depth estimation model, and the stereo depth estimation model based on the knowledge mutual distillation model; A depth estimation module for performing depth estimation on the image to be measured according to the trained single-frame depth estimation model, multi-frame depth estimation model, and stereo depth estimation model.

7. A computer device for image depth estimation based on knowledge mutual distillation, characterized in that including: A memory, a processor, and a computer program stored in the memory, and when the processor executes the computer program, the steps of the method for image depth estimation based on knowledge mutual distillation according to any one of claims 1-5 are implemented.

8. A readable storage medium, characterized in that, The readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by the processor, they are used to execute the steps of the method for image depth estimation based on knowledge mutual distillation according to any one of claims 1-5.

Citation Information

Patent Citations

  • Entity relation extraction method

    CN108733792A

  • Knowledge distillation method and system based on multi-student discussion

    CN114049513A

  • Industrial defect detection method and system under single view angle

    CN115953347A

  • Visual positioning method, device and equipment based on double knowledge distillation and memory

    CN116778140A

  • Monocular image depth estimation method and system

    CN117437274A