A monocular image depth estimation method for wearable helmets

By calculating the plane coefficients in the decoder of the deep convolutional neural network and combining the weight allocation of geometric and color information, the problem of poor depth estimation effect in the mine is solved, and more accurate depth estimation and safety situation analysis are achieved.

CN115423857BActive Publication Date: 2025-07-01CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211242648.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-11
Publication Date
2025-07-01
Estimated Expiration
2042-10-11

AI Technical Summary

Technical Problem

Due to low light, low texture and complex structure underground, the existing monocular image depth estimation method has poor effect, making it difficult to effectively identify and analyze dangerous situations, affecting the safety of underground operations.

Method used

A monocular image depth estimation method for wearable helmets is adopted to calculate the plane coefficients in the decoder of the deep convolutional neural network, indirectly predict the depth map, combine the normal vector and the origin-to-plane distance as geometric information, and allocate weights according to the color and geometric difference information, avoiding error detection and oversegment.

Benefits of technology

The effect of downward image depth estimation in mines is improved, the starting point of training constraints is enhanced, and the effect of later depth estimation is significantly improved, resulting in more accurate Manhattan normal axis and better plane constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115423857B_ABST
    Figure CN115423857B_ABST
Patent Text Reader

Abstract

The present invention discloses a monocular image depth estimation method for a wearable helmet, which relates to the technical field of image processing and includes the following steps: using a mine image sequence as training data, establishing a training model of a depth convolutional neural network model for monocular depth estimation, and calculating plane coefficients that can predict the depth map of the underground image from the plane coefficient decoder of the convolutional neural network; predicting an initial underground image depth map based on the plane coefficients, obtaining a predicted normal vector according to the Manhattan structure normal detection, so as to be similar to the aligned normal vector for constraint; through coplanar normal depth constraint estimation, extracting the depth map obtained from the initial predicted depth and the plane difference, and using the cosine similarity constraint of the two style matrices. The present invention indirectly re-predicts the depth map based on the plane coefficients that can predict the depth map of the underground image, breaks the traditional method of generating the initial depth map, and has a high starting point for training constraints, effectively improving the later depth estimation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a monocular image depth estimation method for a wearable helmet. Background Art

[0002] Coal is the foundation of the development of China's national economy. However, the environment underground in coal mines is very complex, with various toxic and harmful gases such as gas, hydrogen sulfide, sulfur dioxide, etc. In addition, a series of potential safety hazards such as roof falls, rib spalls, and sudden water inrusions often occur underground. These problems will bring about a decline in coal production capacity at worst, and casualties at worst. Therefore, it is beneficial to study the use of wearable helmets equipped with cameras and sensors to collect images of the surrounding environment in real time, perform depth prediction and estimation, thereby identifying and analyzing dangerous situations, and triggering early warning responses, which is conducive to underground workers to eliminate potential hazards in time, reduce the occurrence of dangerous accidents, and ensure the safety of underground operations. At the same time, it is also an important part of the current intelligent mine construction strategy. However, the depth of the mine ranges from dozens of meters to hundreds of meters, with many low-light and low-texture areas, and the structure is complex, which causes great trouble to the monocular depth estimation of underground images and the effect is very poor. Therefore, combining the helmets worn by underground workers, researching an integrated and intelligent wearable helmet for monocular image depth estimation methods and systems is a crucial link to ensure the safety of coal mine workers.

[0003] According to different mathematical models, monocular depth estimation methods can be divided into methods based on traditional machine learning and methods based on deep learning. Monocular depth estimation methods based on traditional machine learning generally use Markov random fields (MRFs) or conditional random fields (CRFs) to model depth relationships, and solve for depth by minimizing the energy function under the maximum a posteriori probability framework. According to whether the model contains parameters, it can be further divided into parametric learning methods and non-parametric learning methods. The former assumes that the model contains unknown parameters, and the training process is to solve for the unknown parameters; the latter uses existing data sets for similarity retrieval to infer depth and does not require learning to obtain parameters. Monocular depth estimation methods based on deep learning are generally classified hierarchically from bottom to top according to different classification criteria. The first level is single-task methods that only predict depth and multi-task methods that simultaneously predict depth and semantic information, etc. The depth and semantic information of pictures are closely related, so some work studies the joint prediction method of multi-tasks; the second level is absolute depth prediction methods and relative depth relationship prediction methods. Absolute depth refers to the actual distance from an object in the scene to the camera, while relative depth focuses on the relative distance relationship of objects in the picture. Given any picture, human vision is better at judging the relative distance relationship of objects in the scene; the third level includes supervised regression methods, supervised classification methods, and unsupervised methods.

[0004] However, since there are no rich texture regions underground in the mine, the photometric loss becomes very weak, making it impossible to train a good depth model. Although the structural rules underground in the mine are now used to select the self-supervised signals of the Manhattan world structural rules, there are many planes underground in the mine, which will predict a lot of insignificant plane normal vectors, resulting in a large error in the estimated main direction axis of the Manhattan world. The existing methods have a certain improvement effect on the results, but the effect of monocular depth estimation is also unsatisfactory, affecting the self-supervised training. Summary of the Invention

[0005] In view of the above problems, the present invention provides a method for monocular image depth estimation for a wearable helmet. By obtaining plane coefficients in the convolutional neural network decoder for monocular depth estimation, the depth map is indirectly predicted again. Mathematically, compared with directly predicting the depth of the underground mine image, the effect of the depth map can be effectively improved, making the starting point of the training constraint high and the later depth estimation effect good. Moreover, the present invention uses the normal vector and the distance from the origin to the plane as the geometric information for calculating the image, and assigns weights to the color difference information and the geometric difference information according to their importance for the underground mine image segmentation, avoiding the false detection and over-segmentation of the image caused by detecting the plane region by using the feature that the region with uniform color is a plane in the complex structure underground in the mine, and obtaining a better plane constraint effect.

[0006] The technical solution adopted by the present invention to solve the above technical problems is as follows:

[0007] A method for monocular image depth estimation for a wearable helmet, comprising the following steps:

[0008] Step a: Using the underground mine image sequence as training data, establishing a training model of a deep convolutional neural network model for monocular depth estimation, using a residual convolutional neural network as an encoder to extract the features of the underground mine image sequence, then calculating the plane coefficients that can predict the underground mine image depth map through a plane coefficient decoder, and then restoring the initial underground mine image depth map through the plane coefficients;

[0009] Step b: Estimating the dominant normal vector of the initial underground mine image depth map obtained in step a, and using its Manhattan structure normal constraint as a loss function constraint for estimation;

[0010] Step c: Using the initial underground mine image depth map and the dominant normal vector obtained in step b, obtaining the difference between the plane and the origin and the normal vector difference, and combining them with the color difference to form a coplanar difference, thereby estimating a new plane depth map, and using the style feature as a loss function for depth constraint estimation.

[0011] Further, the method for restoring the initial underground mine image depth map through the plane coefficients in step a includes:

[0012] Monocular depth estimation requires learning a dense mapping:

[0013] f θ : I(u, v) → D(u, v)

[0014] where I is the input image with scale H×W, D is the corresponding depth map of the same resolution, (u, v) are the pixel coordinates in the image space, and θ are the parameters of the mapping f.

[0015] Assume that the back-projected 3D point P corresponds to a planar part of the 3D scene. The associated plane equation in point-normal form is n·p + d = 0, where n = (a, b, c) T is the plane normal vector, and -d is the distance from the plane to the origin; using the pinhole camera model and given the camera focal lengths (f x , f y ) and the principal point (u0, v0), each pixel point p = (u, v) in the image T is mapped to the 3D point P = (X, Y, Z) through the following formula T ,

[0016]

[0017] Substituting the above 3D point P into the point-normal equation gives:

[0018]

[0019] For the image region depicting the planar 3D surface, the inverse depth is an affine function of the pixel position, where the coefficients encode the camera intrinsics and the 3D plane; by introducing and normalizing and , we get:

[0020] Z = [(αu + βv + γ)ρ] -1

[0021] Taking C = (α, β, γ, ρ) T as the plane coefficients and substituting into the above formula gives Z = h(C, u, v); and then predicting the initial downhole image depth map D i ; the mapping f θ can accurately represent

[0022]

[0023] Applying the above formula at each pixel, g θ : I(u, v) → C(u, v) is to map the input image into plane coefficients for representation, and h: (C(u, v), u, v) → D i(u, v) is remapped from the planar coefficient representation into the predicted initial downhole image depth map.

[0024] Furthermore, the specific steps of step b include:

[0025] Step b1, vanishing point detection;

[0026] Step b2, based on the vanishing point v extracted from the two-dimensional image in step b1, estimate and extract the dominant direction according to the structural lines in the image. The extraction formula is:

[0027] η ∝ K -1 v

[0028] where η ∈ R 3 is the unit vector of the dominant direction, and K is the internal parameter matrix of the camera. The double-line search method is used to extract the dominant direction from the image, and the main direction is extracted once before training;

[0029] Step b3, estimation of the plane normal vector: According to the formula Xp = D(p)K -1 p, obtain the three-dimensional coordinates Xp corresponding to each pixel p in the depth map, and then use a differentiable point-to-normal layer to estimate the plane normal vector;

[0030] where D(p) is the initial depth predicted by the depth network, and K represents the internal parameter matrix of the camera;

[0031] Step b4, adopt an adaptive method. By calculating the total number of planes in the mine, and then selecting N as the boundary value. When N > 10, select 50% of the larger planes as the planes for estimating the normal vector. When N ≤ 10, select 70% of the larger planes as the planes for estimating the normal vector; Use the Manhattan normal detection to detect and classify the plane normals belonging to the main plane, and then use the cosine similarity S to compare the normal vector n p of the estimated plane and each possible main direction η k and select the one with the best similarity as the Manhattan main direction classification result of this point, that is

[0032] n p ∈ (n1, n2,,, n 70%d )

[0033] where d is the total number of planes and n1 > n2 > n3....n 70%d ;

[0034]

[0035] where is the aligned normal, and the cosine similarity is defined as s(n p , η k ) = (n p ·ηk ) / (||n p || · ||η k ||); Let the maximum similarity of each pixel be s p max , so the Manhattan mask is defined as:

[0036]

[0037] where 1 and 0 represent the Manhattan and non-Manhattan regions respectively. Therefore, this method uses the above alignment normal as the monitoring signal, applies the Manhattan structure normal detection within the Manhattan region, and obtains the normal vector n of the estimated plane p The loss function L that is as close as possible to the alignment normal norm .

[0038] Furthermore, the loss function L norm is specifically described as:

[0039]

[0040] where N norm is the number of pixels located in the Manhattan region, indicates whether the pixel p is located in the plane region, indicates whether the pixel is located within the Manhattan plane.

[0041] Furthermore, the specific method for vanishing point detection in step b1 includes:

[0042] Step b11, line detection of the input image;

[0043] Step b12, calculate the intersection points of the above lines as candidates for the vanishing points, and then use an optimization method to obtain the optimal three vanishing points; when calculating the vanishing points, use the Harris pixel corner detection method to detect the sequence images, extract the image coordinates of the four intersection points of two sets of mutually orthogonal parallel lines in each image, and then calculate the image coordinates of the two vanishing points according to the coordinates of the four intersection points and the definition of the vanishing points.

[0044] Furthermore, the line detection method used in step b11 is: first perform boundary detection using the non-differential edge detection operator Canny, and then connect the line segments.

[0045] Furthermore, the optimization method is the least squares method or the voting method.

[0046] Furthermore, step c specifically includes:

[0047] Step c1, planar region detection: Let the three-dimensional coordinates of pixel p be Xp. Assuming that this three-dimensional point lies in a plane with a normal vector based on the aligned normal vector calculated in step b5, the distance from the plane to the origin is calculated as:

[0048]

[0049] Assume that q is an adjacent pixel of p, and the normal dissimilarity between them is defined as the Euclidean distance between two vectors:

[0050]

[0051] Pass through and respectively represent the maximum and minimum dissimilarities between all adjacent pixels. Additionally, a [·] operator is defined as

[0052] [D n (p,q)]=(D n (p,q)-D n min ) / (D n max -D n min )

[0053] Then the dissimilarity of the distance from the plane to the origin is defined as:

[0054] D n (p,q)=|d p -d q |

[0055] Then the geometric information difference combines the normalization of the two dissimilarities as:

[0056] D g (p,q)=[D n (p,q)]+[D d (p,q)]

[0057] The color information difference is calculated as:

[0058] D c (p,q)=||I p -I q ||

[0059] where I P 、I q are the RGB image pixel color values;

[0060] According to the fact that more geometric information is relied on for the information underground in the mine, weights are assigned to the color information difference and the geometric information difference for combination:

[0061] D(p,q) = 0.4 * D c (p,q) + 0.6 * D g (p,q)

[0062] Based on the difference, apply graph-based segmentation and filter out small regions to obtain planar regions.

[0063] Step c2, after detecting the planar region, call the coplanar constraint to flatten the three-dimensional points in the planar region, perform plane fitting on the three-dimensional points in the planar region, and obtain the plane parameter θ = -n / d ∈ R 3 , where the formula for solving the plane parameter is:

[0064] X T θ = 1

[0065] where X ∈ R 3×N represents the three-dimensional points in the planar region;

[0066] Then calculate the inverse depth ρ of pixel p through plane fitting p as:

[0067]

[0068] where K represents the camera intrinsic matrix;

[0069] Then convert the inverse depth ρ p to depth Use the depth obtained from plane fitting Extract the style feature as an additional signal to constrain the estimated depth.

[0070] Furthermore, the extraction of the style feature as an additional signal to constrain the estimated depth is specifically: send the predicted initial depth D p and the plane depth into the encoder of the convolutional neural network respectively to extract their style features to obtain the Gram matrices gram1 and gram2 of the two, expand the two style matrices into one-dimensional vectors, and finally calculate the cosine similarity of the two. The specific constraint is:

[0071]

[0072] where N plane is the number of pixels in the planar region M P inside, n garm1 , n gram2 are the one-dimensional vectors of the two extracted style matrices, and s(,) is the cosine similarity calculation.

[0073] The technical solution of the present invention can produce the following technical effects:

[0074] 1. Instead of directly predicting the depth map for training constraints in the deep convolutional neural network, the present invention calculates the plane coefficients in the decoder of the neural network that can predict the depth map of the underground image, and then indirectly predicts the depth map through these plane coefficients. Mathematically, the effect of the indirectly predicted depth map is far better than directly predicting the depth of the underground mine image, breaking the traditional method of generating the initial depth map, resulting in a high starting point for training constraints and good subsequent depth estimation effects;

[0075] 2. Currently, the use of the structural rules in the underground mine is to calculate the cosine similarity between the possible normal vectors of all planes and the predicted main normal vector to obtain the most likely main normal axis. This not only has a large amount of calculation, but also the calculated main direction axis may not conform to the actual situation. The present invention improves the method of selecting the main direction axis in the Manhattan structure constraint. An adaptive method is adopted. By calculating the total number of planes in the underground mine, and then taking N as the boundary value. When N is greater than 10, the larger planes (taking 50% of them) are selected as the planes for estimating the normal vector. When N is less than or equal to 10, the larger planes (taking 70% of them) are selected as the planes for estimating the normal vector. This reduces the interference caused by smaller planes leading to inaccuracies. In this way, the method of calculating the similarity using the normal vectors of larger planes not only reduces the amount of calculation but also conforms to the Manhattan law of the underground mine structure. Experiments prove that the obtained Manhattan normal axis is more accurate and the constraint is better;

[0076] 3. The present invention uses the normal vector and the distance from the origin to the plane as the geometric information for calculating the image, and assigns weights to the color difference information and geometric difference information according to their importance for the underground mine image segmentation, avoiding the false detection and over-segmentation of the image caused by detecting the plane area by using the feature that the area with uniform color is a plane in the complex structure of the underground mine, and obtaining a better plane constraint effect;

[0077] 4. The present invention innovates a new loss function. In the obtained segmented planes, the least squares method is used to calculate the plane constraint depth D p plane , without directly comparing the norm 1 with the depth D predicted by the plane coefficients p directly, because directly handling the norm 1 lacks consideration of depth feature information. Using the relevant theory in the style transfer paper, D p , D p plane are used to extract the style feature matrices gram1 and gram2 through the encoder of the convolutional neural network in the invention. The two matrices are respectively pulled into one-dimensional vectors n garm1 , n gram2 , and then the cosine similarity between them is calculated to obtain the cosine similarity, making the picture constraint more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 Overall method framework diagram of the embodiments of the present invention;

[0079] Figure 2 Framework diagram of Manhattan structure depth constraint estimation described in the embodiments of the present invention;

[0080] Figure 3 Framework diagram of color-geometry information plane segmentation estimation in the co-planar normal depth constraint estimation step described in the embodiments of the present invention;

[0081] Figure 4 Flow chart of style matrix (gram) feature loss calculation in the co-planar normal depth constraint estimation step described in the embodiments of the present invention;

[0082] Figure 5 Schematic diagram of comparison of depth maps generated by the present invention and other methods on the NYUv2 dataset. Detailed implementation manners

[0083] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.

[0084] As Figure 1 shown, the overall process of a monocular image depth estimation method for a wearable helmet according to the present invention is as follows: First, use a mine image sequence as training data to establish a training model of a deep convolutional neural network model for monocular depth estimation, and calculate plane coefficients that can predict the depth map of the underground image from the plane coefficient decoder of the convolutional neural network; then predict an initial underground image depth map based on the plane coefficients, obtain a predicted normal vector according to the Manhattan structure normal detection, so as to perform a similarity constraint with the aligned normal; through co-planar normal depth constraint estimation, extract the depth map obtained from the initial predicted depth and the plane difference, and use the style matrices of the two for cosine similarity constraint.

[0085] Monocular depth estimation needs to learn a dense mapping f θ : I(u, v) → D(u, v), where I is the input image with a scale of H×W, D is the corresponding depth map with the same resolution, (u, v) are pixel coordinates in the image space, and θ are the parameters of the mapping f. In the initial depth prediction step based on plane coefficients, assume that the back-projected 3D point P corresponds to the planar part of the 3D scene, and the associated plane equation in point-normal form is n·p + d = 0, where n = (a, b, c) T is the plane normal vector and d is the distance from the plane to the origin; use the pinhole camera model and given focal lengths (f x , fy ) and the principal point (u0, v0), each pixel p = (u, v) T is mapped to a 3D point P = (X, Y, Z) T , obtaining:

[0086]

[0087] Substituting the above 3D point P into the point normal equation gives:

[0088]

[0089] For the image region depicting a planar 3D surface, the inverse depth is an affine function of the pixel position, where the coefficients encode the camera interior and the 3D plane; by introducing and making and normalized, we get:

[0090] Z = [(αu + βv + γ)ρ] -1

[0091] Taking C = (α, β, γ, ρ) T as the plane coefficients and substituting into the above equation gives Z = h(C, u, v), and then predicting the initial downhole image depth map. The mapping f θ can be accurately expressed as Applying the above formula at each pixel, where g θ : I(u, v) → C(u, v), maps the input image to the plane coefficient representation; h: (C(u, v), u, v) → D i (u, v) remaps from the plane coefficient representation to the predicted initial depth image

[0092] Such as Figure 2As shown in the figure, first, find the structural regular lines in the mine and perform vanishing point detection. First, perform Canny operator edge detection, and then connect the line segments to complete the straight line detection of the input image. Then, calculate the intersection points of these straight lines as the candidate points for the vanishing points. Then, use optimization methods such as least squares and voting to obtain the optimal three vanishing points. The specific algorithm steps of the voting method are as follows: (1) All line segments vote for all candidate points. The voting method is to calculate the angle theta between the straight line from the midpoint of the line segment to the specified point and the original line segment. The voting value is |l|*e^(theta / (2*u^2)), where l is the length of the line segment and u is a robustness parameter, specifically set to 0.1; (2) Perform hierarchical clustering on the candidate points after voting. The clustering end condition is that the minimum distance is greater than 50 pixel points, and the minimum distance value used as the clustering end condition can be set; (3) Calculate the weighted center of gravity of the votes for each cluster as the new candidate point; the number of votes is the sum of the votes of all candidate points in the cluster; (4) Select the cluster with the highest number of votes as the first output point, and eliminate the line segments whose included angle with the connection line of all midpoint candidate points is less than 10 degrees; (5) Repeat the above steps until there are no remaining line segments, or three optimal vanishing points have been found, or there are no candidate points, then the algorithm ends. When calculating the vanishing points, use the Harris pixel corner detection method to detect the sequence images, extract the image coordinates of the four intersection points of two groups of mutually orthogonal parallel straight lines in each image, and then calculate the image coordinates of the two vanishing points according to the coordinates of the four intersection points and the definition of the vanishing points.

[0093] Next, based on the obtained vanishing point V, calculate the principal direction in the camera coordinate system. In this embodiment, the double-line search method is used to extract the dominant direction from the image. The principal direction is extracted only once before training, and both the extracted principal direction and its opposite direction are regarded as the possible directions of the normal of the main planes in the scene, such as wearable helmets, roadway walls, and belts.

[0094] η∝K -1 V

[0095] where η belongs to R 3 is the unit vector representing the principal direction, and K is the internal parameter matrix of the camera.

[0096] Then, according to the formula Xp = D(p)K -1 p, obtain the three-dimensional coordinates Xp corresponding to each pixel p in the depth map, and then use a differentiable point-to-normal layer to estimate the plane normal vector, where D(p) is the initial depth predicted by the depth network. Specifically, given the normal n p of pixel p is calculated from a set of three-dimensional points in a small neighborhood centered at point X p ; the small neighborhood here is set to an 8×8 block domain.

[0097] Next, Manhattan region detection is carried out. In this embodiment, an adaptive method is adopted. By calculating the total number of planes in the mine, and then selecting N as the boundary value. When N is greater than 10, 50% of the larger plane is selected as the plane for estimating the normal vector. When N is less than or equal to 10, 70% of the larger plane is selected as the plane for estimating the normal vector. Estimate the plane normal prediction n, and use Manhattan normal vector detection to detect and classify the plane normals belonging to the main plane, and then use the cosine similarity S(,) to compare the normal vector n of the estimated plane p and each possible main direction η k between the differences, and select the one with the best similarity as the Manhattan main direction classification result of this point, that is

[0098] n p ∈(n1,n2,,,n 70%d )

[0099] where d is the total number of planes and n1>n2>n3....n 70%d ;

[0100]

[0101] where is the aligned normal, and the cosine similarity is defined as s(n p ,η k )=(n p ·η k ) / (||n p ||·||η k ||); Let the maximum similarity of each pixel be s p max , so the Manhattan mask is defined as:

[0102]

[0103] where 1 and 0 represent the Manhattan and non-Manhattan regions respectively. Therefore, this method uses the above aligned normal as the monitoring signal, applies Manhattan structure normal detection within the Manhattan region, and obtains the normal vector n of the estimated plane p as close as possible to the loss function L of the aligned normal norm ; The loss function L norm is specifically described as:

[0104]

[0105] where, N norm is the number of pixels located in the Manhattan region, indicates whether the pixel p is located in the plane region, indicates whether the pixel is located within the Manhattan plane.

[0106] In the co-planar normal depth constraint estimation step, as Figure 3 , Figure 4 shown, in the plane region detection process carried out first, by integrating color and geometric information and relying on the principle of style transfer, the cosine similarity between the style matrices of the initial prediction and the depth map obtained from the plane is calculated for constraint. This difference constraint method takes into account color, plane normal vector, distance from the origin to the plane, and depth style features, specifically including:

[0107] Assume that the three-dimensional coordinates Xp of pixel p lie in a plane with an aligned normal, then the distance from the plane to the origin is calculated as:

[0108]

[0109] Assume that q is an adjacent pixel of p, and the normal dissimilarity between them is defined as the Euclidean distance between the two vectors:

[0110]

[0111] Through and respectively represent the maximum and minimum dissimilarities between all adjacent pixels. Additionally, a [·] operator is defined,

[0112] [D n (p,q)]=(D n (p,q)-D n min ) / (D n max -D n min )

[0113] Then the dissimilarity of the distance from the plane to the origin is defined as:

[0114] D n (p,q)=|d p -d q |

[0115] Then the geometric information dissimilarity combines the normalization of the two dissimilarities as:

[0116] D g (p,q)=[D n (p,q)]+[D d (p,q)]

[0117] The color information dissimilarity is calculated as:

[0118] D c (p,q)=||I p -I q ||

[0119] Among them, I P , I q is the RGB image pixel color value;

[0120] Since there is a lot of geometric information based on the information underground in the mine, weight values are assigned to the color information difference and the geometric information difference for combination:

[0121] D(p,q) = 0.4 * D c (p,q) + 0.6 * D g (p,q)

[0122] Based on the difference, graph-based segmentation is applied and small regions are filtered out to obtain a planar region. Compared with only using color information, the method of this embodiment fully considers the geometric information in the underground structure of the mine, avoiding false planar regions that cannot be distinguished by color and over-segmentation caused by different colors.

[0123] After detecting the planar region, the coplanarity constraint is called to flatten the three-dimensional points in the planar region, and plane fitting is performed on the three-dimensional points in the planar region. The plane parameter θ = -n / d ∈ R is obtained by solving the least squares problem 3 , where the formula for solving the plane parameter is:

[0124] X T θ = 1

[0125] where X ∈ R 3×N represents the three-dimensional points in the planar region;

[0126] Then the inverse depth ρ of pixel p is calculated through plane fitting p as:

[0127]

[0128] where K represents the camera intrinsic matrix;

[0129] Then the inverse depth ρ p is converted to depth using the depth obtained from plane fitting to extract style features as an additional signal to constrain the estimated depth;

[0130] In style transfer, the Gramian matrix can well reflect the combination between different features. Therefore, in this embodiment, the predicted initial depth D p and the plane depth are respectively fed into the encoder in the convolutional neural network to extract their style features to obtain the Gram matrices gram1 and gram2 of the two. The two style matrices are expanded into one-dimensional vectors, and finally the cosine similarity between the two is calculated. The specific constraint is:

[0131]

[0132] where N plane is the number of pixels in the planar region M P and n garm1 , n gram2 is the one-dimensional vector of the two style matrices extracted, and s(,) is the cosine similarity calculation.

[0133] Combined with Figure 5 the comparison graph shown in, Table 1 shows the depth estimation results of the present invention compared with MovingIndoor(2019), Monodepth2(2019), P 2 Net(2020) and Structdepth(2021) on the NYUv2 dataset; where RMS represents the root mean square, AbsRel represents the absolute relative difference, and σ represents the accuracy.

[0134] Table 1 Depth Estimation Results on the NYUv2 Dataset

[0135]

[0136] As can be seen from Table 1, the accuracy of the depth map predicted by the technical solution of this embodiment on the NYUv2 dataset is greater than that of other algorithms; and from the comparison results of the predicted depth maps, the effect of the image depth map predicted by this embodiment is better than that of P 2 Net(2020) and Structdepth(2021).

[0137] The preferred specific embodiments of the present invention have been described in detail above, and they do not impose any restrictive effects on the present invention. Any person skilled in the art, without departing from the scope of the technical solution of the present invention, makes any form of equivalent substitution or modification and other changes to the technical solutions and technical contents disclosed by the present invention, which are all within the content of the technical solution of the present invention and still fall within the protection scope of the present invention.

Claims

1. A monocular image depth estimation method for a wearable helmet, characterized in that, It includes the following steps: Step a: Using the mine image sequence as training data, establish a training model for a depth convolutional neural network model for monocular depth estimation. Use a residual convolutional neural network as an encoder to extract the features of the underground image sequence, and then calculate the plane coefficients that can predict the underground image depth map through a plane coefficient decoder. Then, restore the initial underground image depth map through the plane coefficients; Step b: Perform a dominant normal vector estimation on the initial underground image depth map obtained in step a, and use its Manhattan structure normal constraint as a loss function constraint for estimation; The dominant normal vector estimation adopts an adaptive method. By calculating the total number of planes underground in the mine, and then taking N as the boundary value. When N is greater than 10, 50% of the larger plane is selected as the plane for estimating the normal vector. When N is less than or equal to 10, 70% of the larger plane is selected as the plane for estimating the normal vector; Step c: Use the initial underground image depth map obtained in step a and the dominant normal vector in step b to obtain the difference between the plane and the origin, the normal vector difference, and the color difference to form a coplanar difference, thereby estimating a new plane depth map, and using the style feature as a loss function for depth constraint estimation; The extraction of the style feature as an additional signal to constrain the estimated depth is specifically as follows: the predicted initial depth D p and the planar depth are respectively fed into the encoder of the convolutional neural network to extract their style features to obtain the Gram matrices gram1 and gram2 of the two. The two style matrices are unfolded into one-dimensional vectors, and finally the cosine similarity between the two is calculated. The specific constraint is as follows: where N plane is the number of pixels in the planar region M P , n garm1 , n gram2 is the one-dimensional vector of the two style matrices extracted, and s(,) is the cosine similarity calculation.

2. The monocular image depth estimation method for a wearable helmet according to claim 1, wherein The method for restoring the initial underground image depth map through plane coefficients in step a includes: Monocular depth estimation needs to learn a dense mapping: f θ : I(u, v) → D(u, v); Where I is the input image with a scale of H×W, D is the corresponding depth map with the same resolution, (u, v) are the pixel coordinates in the image space, and θ is the parameter of the mapping f; Assume that the back-projected 3D point P corresponds to a planar part of the 3D scene, and the associated plane equation in point-normal form is np + d = 0, where n = (a, b, c) T is the plane normal vector and d is the distance from the plane to the origin; using the pinhole camera model and given the camera focal lengths (f x , f y ) and the principal point (u0, v0), each pixel point p = (u, v) in the image T is mapped to the 3D point P = (X, Y, Z) by the following formula T , Substitute the above three-dimensional point P into the point normal equation to get: For an image region depicting a planar three-dimensional surface, the inverse depth is an affine function of the pixel position, where the coefficients encode the camera intrinsics and the three-dimensional plane; by introducing and normalizing and we obtain: Z = [(αu + βv + γ)ρ] -1 ; Let \(C = (\alpha, \beta, \gamma, \rho)\) T be the plane coefficients and substitute them into the above formula to obtain \(Z = h(C, u, v)\); then predict the initial downhole image depth map \(D\) i ; the mapping \(f\) θ can be accurately expressed as Applying the above formula at each pixel, g θ : I(u, v) → C(u, v) is to map the input image into plane coefficients for representation, and h: (C(u, v), u, v) → D i (u, v) remaps from the plane coefficient representation into the predicted initial downhole image depth map.

3. A monocular image depth estimation method for a wearable helmet according to claim 2, characterized in that, Step b specifically includes: Step b1: Vanishing point detection; Step b2: Based on the vanishing point v extracted from the two-dimensional image in step b1, estimate and extract the dominant direction according to the structural lines in the image. The extraction formula is: η∝K -1 v; where η ∈ R 3 is the unit vector in the dominant direction, and K is the internal parameter matrix of the camera; the double-line search method is used to extract the dominant direction from the image, and the dominant direction is extracted once before training; Step b3, estimation of the plane normal vector: According to the formula Xp = D(p)K -1 p, the three-dimensional coordinate Xp corresponding to each pixel p in the depth map is obtained, and then a differentiable point-to-normal layer is used to estimate the plane normal vector; Where D(p) is the initial depth predicted by the depth network, and K represents the camera intrinsic matrix; Step b4, use Manhattan normal detection to detect the plane normal of the plane classified as the main plane, and then use the cosine similarity S to compare the normal vector n of the estimated plane p and each possible main direction η k to find the difference, and select the one with the best similarity as the Manhattan main direction classification result of this point, that is n p ∈(n1,n2,,,n 70%d ); where d is the total number of planes and n1 > n2 > n3....n 70%d ; Among them is the alignment normal, and the cosine similarity is defined as s(n p , η k ) = (n p · η k ) / (||n p || · ||η k ||); Let the maximum similarity of each pixel be s p max , so the Manhattan mask is defined as: Where 1 and 0 represent Manhattan and non-Manhattan regions respectively. Therefore, this method uses the above-aligned normal as the monitoring signal, applies the Manhattan structure normal detection within the Manhattan region, and obtains the normal vector n of the estimated plane p The loss function L that is as close as possible to the aligned normal norm .

4. A monocular image depth estimation method for a wearable helmet according to claim 3, characterized in that, The loss function L norm is specifically described as: where N norm is the number of pixels located in the Manhattan area, indicates whether pixel p is located in the planar area, indicates whether the pixel is located within the Manhattan plane.

5. A monocular image depth estimation method for a wearable helmet according to claim 3, characterized in that, The specific method for vanishing point detection in step b1 includes: Step b11: Straight line detection of the input image; Step b12: Calculate the intersection points of the above straight lines as candidates for the vanishing point, and then use an optimization method to obtain the optimal three vanishing points. When calculating the vanishing point, use the Harris pixel corner detection method to detect the sequence images, extract the image coordinates of the four intersection points of two groups of mutually orthogonal parallel straight lines in each image, and then calculate the image coordinates of the two vanishing points according to the coordinates of the four intersection points and the definition of the vanishing point.

6. The monocular image depth estimation method for a wearable helmet according to claim 5, characterized in that The straight line detection method adopted in step b11 is: First, perform boundary detection using the non-differentiable edge detection operator Canny, and then connect the line segments.

7. A monocular image depth estimation method for a wearable helmet according to claim 5, characterized in that The optimization method is the least squares method or the voting method.

8. A monocular image depth estimation method for a wearable helmet according to claim 3, characterized in that Step c specifically includes: Step c1: Plane region detection: Let the three-dimensional coordinates of pixel p be Xp. Assume that this three-dimensional point is located in a plane with a normal line being the aligned normal line calculated in step b4. Then the distance from the plane to the origin is calculated as: Assume that q is an adjacent pixel of p, and the normal line dissimilarity between them is defined as the Euclidean distance between two vectors: By and respectively represent the maximum and minimum dissimilarities between all adjacent pixels. Additionally, a [·] operator is defined. Then the dissimilarity of the distance from the plane to the origin is defined as: D d (p,q) = |d p - d q |; Then the geometric information dissimilarity combines the normalization of the two dissimilarities as: D g (p,q) = [D n (p,q)] + [D d (p,q)]; The color information dissimilarity is calculated as: D c (p,q) = ||I p -I q ||; where I P and I q are RGB colors; Since there is a lot of geometric information on which the underground mine information is based, weights are assigned to the color information difference and the geometric information difference for combination: D(p,q) = 0.4 * D c (p,q) + 0.6 * D g (p,q); Based on the difference, graph-based segmentation is applied and small regions are filtered out to obtain planar regions; Step c2, after detecting the planar region, call the coplanarity constraint to flatten the three-dimensional points in the planar region, perform planar fitting on the three-dimensional points in the planar region, and obtain the plane parameter θ = -n / d ∈ R by solving the least squares problem 3 , where the formula for solving the plane parameter is: X T θ = 1; where X ∈ R 3×N represents a three-dimensional point within a planar region; Then, the inverse depth ρ of pixel p is calculated by plane fitting p as follows: where K represents the camera intrinsic matrix; Then the inverse depth ρ p is converted to depth using the depth obtained from plane fitting Extract the style features as additional signals to constrain the estimated depth.