A 3D reconstruction method for imaging sonar based on differentiable spatial carving

By adopting a method based on microspace engraving in sonar image reconstruction, using echo probability to iteratively segment the space and combining rendering loss-guided backpropagation, the problem of the lack of correspondence between echo and occupancy grid in sonar image reconstruction is solved, and faster three-dimensional reconstruction efficiency is achieved.

CN118887342BActive Publication Date: 2025-05-16HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411013467.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2025-05-16
Estimated Expiration
2044-07-26

AI Technical Summary

Technical Problem

Due to the lack of elevation angle measurement, the existing sonar image reconstruction method has lost the correspondence between the echo and the occupied grid, and it is impossible to directly use the classical spatial engraving method. The NeRf model takes several hours to converge, which is inefficient.

Method used

The three-dimensional reconstruction method of imaging sonar based on microspace engraving is adopted. By establishing an occupation probability model and a rendering model, the echo probability rather than the echo intensity is used to iteratively segment the space, and combined with the rendering loss-guided backpropagation, the lack of corresponding relationships is solved.

Benefits of technology

It realizes faster three-dimensional reconstruction convergence speed, can quickly rebuild the detected objects, and improves the reconstruction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118887342B_ABST
    Figure CN118887342B_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional reconstruction method of imaging sonar based on differentiable space carving, which belongs to the technical field of underwater sonar image reconstruction. A space carving algorithm (DSC) is proposed. According to the sonar image in a known posture, the detected object is reconstructed, the occupancy probability model is iteratively refined, and each pixel is rendered by evaluating the probability of at least one echo appearing along the corresponding arc. When rendering the echo intensity image, the method of the invention can eliminate the need for the echo intensity network, so that the redundancy of the scene model is reduced. The invention adopts the above method, reconstructs the construction of a three-dimensional model according to the information collected by the underwater sonar through the occupancy probability model and the rendering model, and combines the multi-resolution hash position coding for dimensional conversion, thereby improving the model convergence speed and reconstruction quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of sonar image reconstruction, and in particular to an imaging sonar three-dimensional reconstruction method based on microscopic space carving. Background Art

[0002] Multibeam sonar (MBS) is an advanced sonar system used to inspect underwater scenes. Due to the low attenuation of sound waves in water, sonar systems receive less interference than visual sensors and are more suitable for use in turbid underwater environments. However, using sonar detection involves the problem of using sonar images with known postures for three-dimensional reconstruction. The classic space carving method for 3D reconstruction relies on inferring whether each voxel grid is part of an object. However, sonar images only provide the distance, azimuth, and intensity of the echo, lacking elevation. The lack of correspondence between the echo and the occupancy grid makes the space carving method not directly applicable, so existing methods usually use the NeRF model. The NeRF model is an ideal model for three-dimensional reconstruction from 2D visual images through iterative back propagation, and has been reconstructed using sonar images. However, due to the lack of elevation measurement of acoustic echoes, rendering a pixel in a sonar image requires occupancy evaluation on a spatial sector, while when rendering visual pixels, the evaluation is only performed along a single fragment. Therefore, the number of required samples increases dramatically, and existing methods usually take hours to converge. Summary of the invention

[0003] The purpose of the present invention is to provide an imaging sonar three-dimensional reconstruction method based on differentiable space carving, which can quickly reduce the detected object according to the sonar image, and solve the lack of correspondence between occupancy and echo by presenting the echo probability instead of the iterative segmentation space of echo intensity, and combining the back propagation guided by rendering loss, so as to reconstruct faster.

[0004] To achieve the above object, the present invention provides an imaging sonar 3D reconstruction method based on microscopic spatial carving, comprising the following steps:

[0005] Step 1: Create the sonar framework F s , according to the performance of the sonar equipment, the distance r∈(r min ,r max ), azimuth θ∈(θ min ,θ max ) and elevation angle The formula for obtaining the echo position coordinates is as follows:

[0006]

[0007] In the above formula, p (s) is the echo coordinate;

[0008] Step 2: Collect sonar images at different positions under known postures, process the collected sonar images, obtain the echo probability map EP from the original echo intensity map EI, and calculate the arc corresponding to the pixel q = (r, θ)

[0009] Step 3: Establish differentiable space carving, including occupancy probability model and rendering model. Sampling is performed to form a set P j and the set L mn , use the occupancy probability model to output the occupancy probability of the current pixel, input the occupancy probability into the rendering model, perform single-pixel rendering, and obtain the predicted echo probability map;

[0010] Step 4: Iterative training, establish a loss function group, given a set of echo probability maps of known postures, randomly select several echo probability maps to form a subset, and sample a set from the subset pixels, repeat step 3 to get the corresponding The predicted echo probability map is used to generate an echo probability map for each pixel. The occupancy probability model is iteratively updated based on the loss function group through the predicted echo probability map and the actual echo probability map until the loss function group meets the standard.

[0011] Step 5: Use the trained differentiable spatial carving to reconstruct the detected object.

[0012] Preferably, in step 2, the process of obtaining the echo probability map EP from the echo intensity map EI is as follows:

[0013] For the original echo intensity map, remove the object in it to obtain the noise background map, use the trained SCU neural network to reduce the noise of the original echo intensity map and the noise background map respectively, and obtain the reduced noise echo intensity map and the reduced noise background map, and use the reduced noise EI map to subtract the reduced noise background map to obtain the reduced noise and background removed echo intensity map;

[0014] In the echo intensity map with noise reduction and background removal, calculate the arc corresponding to pixel q = (r, θ) The formula is as follows:

[0015]

[0016] Pixel q and arc The probability of at least one echo being present is transformed by the mapping function g, and the transformation formula is as follows:

[0017]

[0018] Where I(q) and Pr(q) are the intensity and pixel value of pixel q respectively, ∈ and μ are preset values. The echo intensity map with noise reduction and background removal is transformed into the echo probability map by formula (2).

[0019] Preferably, in step three, the specific process is as follows:

[0020] When the value of pixel q=(r,θ) in the echo probability map is greater than zero, the arc At least one echo appears on the arc Sample m points to form a set P j , where 1≤j≤m;

[0021] Draw m line segments from the sonar origin and the set P j The sampling points in are connected, and n points are sampled on each line segment to form a set L mn ; where the set P j With the set L mn Composed of sampling point set u;

[0022] In the sonar frame F s In the set P j The corresponding coordinates are Convert it to world frame F w In the formula, the following is true:

[0023]

[0024] In the above formula, For the corresponding P j In world frame coordinates, is the transformation formula from sonar frame to world frame coordinate system;

[0025] Establish a position encoding network E, through multi-resolution hash coding, input u to get the formula as follows:

[0026] e=enc(u,ξ)

[0027] Then, e is input into the occupancy probability model M(e;κ), where κ is the training parameter set of the occupancy network M. M(e;κ) outputs the occupancy probability, which is then input into the rendering model.

[0028] In the rendered model, click P j The echo probability is recorded as in accordance with The occupancy probability and and O, along the line segment The projection rate estimation of is parameterized as follows:

[0029]

[0030] Among them, p o are the coordinates of point O, The occupancy probability of is given by the occupancy network M(e;κ), The occupancy probability is set to but The formula is as follows:

[0031]

[0032] Among them, T j Indicates along The transmittance of the sonar is estimated to reach The probability of a point, ρ is the volume density function, then Pr o (l jk ) and the function ρ are related as follows:

[0033]

[0034] Among them, ρ jk =ρ(t jk ), δ jk =t jk+1 -t jk , then {L jk}, the occupancy probability of 1≤k≤n is Pr o (l jk ), by occupying the network M(e;κ), the perspective rate T j The formula is as follows:

[0035]

[0036] The probability of obtaining at least one echo estimated for pixel q = (r, θ) is as follows:

[0037]

[0038] in, is the probability value of a pixel with at least one echo.

[0039] Preferably, in step 4, the process of iterative training is as follows:

[0040] Establish a loss function group, including the rendering loss L of the echo probability prob , sparsity loss L sparsity and the parallax loss L of the surface normal norm , the formula is as follows:

[0041]

[0042] Among them, λ1 and λ2 are the weights of the corresponding losses, and the rendering loss L of the echo probability at the sampling pixel isprob as follows:

[0043]

[0044] in, Represents the set of sampled pixels, the sparsity loss L sparsity The formula is as follows:

[0045]

[0046] in, Represents the set of sampling points, and the formula is as follows:

[0047] Where 1≤j≤m,1≤k≤n

[0048] The parallax loss of the surface normal is obtained as follows:

[0049]

[0050] S={v∈B|Pr o (v)>τ}

[0051]

[0052] In the above formula, the threshold τ is selected based on the tracking error, S represents the set of sampling points with occupancy probability greater than the threshold τ, and v w It represents random sampling of points within a small radius of v, where n(v) is the unit normal vector at the position of point v.

[0053] Therefore, the present invention adopts the above-mentioned imaging sonar 3D reconstruction method based on microscopic space carving, which has the following advantages:

[0054] (1) In the present invention, an occupancy probability grid and multi-resolution hash coding are used to construct a differentiable occupancy model, thereby ensuring its convergence speed. Compared with the existing Nerf method, the present invention can converge faster.

[0055] (2) In the present invention, a space carving algorithm is proposed for efficient 3D reconstruction using sonar images in a constant pose.

[0056] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 A flowchart of a three-dimensional reconstruction method of imaging sonar based on microscopic spatial carving according to the present invention;

[0058] Figure 2A reconstruction diagram of an imaging sonar three-dimensional reconstruction method based on microscopic space carving according to the present invention;

[0059] Figure 3 A schematic diagram of arc sampling in an imaging sonar three-dimensional reconstruction method based on differentiable space carving according to the present invention;

[0060] Figure 4 This is a comparison diagram of reconstruction using different methods in an imaging sonar three-dimensional reconstruction method based on microscopic spatial carving according to the present invention. DETAILED DESCRIPTION

[0061] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations. The specific model specifications need to be selected and determined according to the actual specifications of the device, and the specific selection calculation method adopts the existing technology in the field, so it will not be described in detail.

[0062] Example

[0063] like Figure 1-Figure 4 As shown, the present invention provides an imaging sonar 3D reconstruction method based on microscopic space carving, comprising the following steps:

[0064] Step 1: If Figure 2 , establish the sonar framework F for sonar s , according to the performance of the sonar equipment, the distance r∈(r min ,r max ), azimuth θ∈(θ min ,θ max ) and elevation angle The formula for obtaining the echo position coordinates is as follows:

[0065]

[0066] Among them, p (s) is the echo coordinate;

[0067] Step 2: Collect sonar images at different positions under known postures, process the collected sonar images, obtain the echo probability map EP from the original echo intensity map EI, and obtain the arc corresponding to the pixel q = (r, θ)

[0068] The process of obtaining the echo probability map EP from the echo intensity map EI is as follows:

[0069] The trained neural network SCU network is used to reduce the noise of the echo intensity map. The full name of the SCU network is Swin-Conv-Unet. The SCU network is used to obtain a filtered background image from an image without an object. The denoised image is subtracted to obtain an EI image with denoised and background removed. The denoised and background removed echo intensity map is converted into an echo probability map. First, the arc corresponding to the pixel q = (r, θ) is calculated. The formula is as follows:

[0070]

[0071] The conversion formula is as follows:

[0072]

[0073] Where I(q) and Pr(q) are the intensity and pixel value of pixel q, ∈ and μ are preset values, and g is a mapping function. The echo intensity map with noise reduction and background removal is transformed into an echo probability map using formula (2).

[0074] Step 3: Establish differentiable space carving, including occupancy probability model and rendering model. Sampling, forming a set P j , use the occupancy probability model to output the occupancy probability of the current pixel The occupancy probability is input into the rendering model, and single-pixel rendering is performed to obtain the echo probability. The specific process is as follows:

[0075] When the value of pixel q=(r,θ) in the echo probability map is greater than zero, the arc At least one echo appears on the surface. The specific sampling results are as follows: Figure 3 , along the arc Sample m points to form a set P j , where 1≤j≤m;

[0076] Draw m line segments from the sonar origin and the set P j The sampling points in are connected, and n points are sampled on each line segment to form a set L mn , where the set P j With the set L mn Composed of sampling point set u;

[0077] In the sonar frame F s In the set P j The corresponding coordinates are Convert it to world frame F w In the formula, the following is true:

[0078]

[0079] in, For the corresponding P j In world frame coordinates, is the transformation formula from sonar frame to world frame coordinate system;

[0080] Establish a position encoding network E, through multi-resolution hash coding, input u to get the formula as follows:

[0081] e=enc(u,ξ)

[0082] Then e is input into the occupancy probability model M(e;κ), where κ is the training parameter set of the occupancy network M. M(e;κ) outputs the occupancy probability, which is then input into the rendering model.

[0083] In the rendered model, click P j The echo probability is recorded as in accordance with The occupancy probability and and O, along the line segment The projection rate estimation of is parameterized as follows:

[0084]

[0085] Among them, p o are the coordinates of point O, The occupancy probability of is given by the occupancy network M(e;κ), The occupancy probability is set to but The formula is as follows:

[0086]

[0087] Among them, T j Indicates along The transmittance of the sonar is estimated to reach The probability of a point, ρ is the volume density function, then ρr o (l jk ) and the function ρ are related as follows:

[0088]

[0089] Among them, ρ jk =ρ(t jk ), δ jk =t jk+1 -t jk , then {L jk}, the occupancy probability of 1≤k≤n is Pr o (l jk ), by occupying the network M(e;κ), the perspective rate T j The formula is as follows:

[0090]

[0091] The probability of obtaining at least one echo estimated for pixel q = (r, θ) is as follows:

[0092]

[0093] in, is the probability value of a pixel with at least one echo.

[0094] Step 4: Iterative training, establish a loss function group, given a set of echo probability maps with known postures, randomly select several echo probability maps to form a subset, and sample a set from the subset pixels, repeat step 3 to obtain the predicted echo probability map. Through the predicted echo probability map and the original echo probability map, the occupancy probability model is iteratively updated based on the loss function group until the loss function group meets the standard. The specific process is as follows:

[0095] Establish a loss function group, including the rendering loss L of the echo probability prob , sparsity loss L sparsity and the parallax loss L of the surface normal norm , the formula is as follows:

[0096]

[0097] Among them, λ1 and λ2 are the weights of the corresponding losses, and the rendering loss L of the echo probability at the sampling pixel is prob as follows:

[0098]

[0099] in, Represents the set of sampled pixels, corresponding to the sparsity loss L sparsity The formula is as follows:

[0100]

[0101] in, Represents the set of sampling points, and the formula is as follows:

[0102] Where 1≤j≤m,1≤k≤n

[0103] The parallax loss of the surface normal is obtained as follows:

[0104]

[0105] S={v∈B|Pr o (v)>τ}

[0106]

[0107] In the above formula, the threshold τ is selected based on the tracking error, S represents the set of sampling points with occupancy probability greater than the threshold τ, and v w It represents random sampling of points within a small radius of v, where n(v) is the unit normal vector at the position of point v.

[0108] The specific experimental process is as follows: configure the experimental environment, conduct simulation experiments in the holooocean simulation environment, place several objects of different sizes in the simulated underwater environment and use the simulated MBS to obtain sonar images, set the additive noise to w sa ~R(0.15), R is Rayleigh distribution, and multiplicative noise is set to w sm ~N(0,0.2), and simulate the multipath effect, the detection distance is [0,8m], the horizontal aperture is [-14.4,14.4], and the vertical aperture is [];

[0109] like Figure 4 , showing the images reconstructed by the real model and the present invention and Neusis respectively. In order to compare the set differences between them, the Hausdorff distance is used for comparison, and the formula is as follows:

[0110]

[0111] Among them and is the set of points to be compared, ||ab|| represents the Euclidean distance between point a and point b,

[0112] From the definition of H(U,V), it can be seen that it measures the maximum mismatch between two point sets. First, the reconstructed model is aligned with the GT using the iterative closest point algorithm, and then several points are randomly selected from the reconstructed model and the GT model to obtain the Hausdorff distance, as shown in Equation (15). This two-step process is run multiple times to obtain the average value and root mean square (RMS) of the Hausdorff distance. The results are shown in Table 1, as follows

[0113] Table 1

[0114]

[0115] In addition to the Hausdorff distance (HD) indicating the reconstruction quality, the chamfer distance (CD) and the land moving distance (EMD) are also shown in Table 1. The time required for the invention and the comparison algorithm to complete the reconstruction, as well as the size of each object in the first column (the length, width and height of the smallest cuboid bounding box, in meters). Please note that one epoch in the table means a complete pass of the dataset through the algorithm, which means one iteration, that is, the number of training rounds, in which each image is used once to update the neural network.

[0116] from Figure 4 As can be seen from Table 1, the method of the present invention has better reconstruction details. For example, in the reconstruction of a ship and an airplane, the method of the present invention can reconstruct the mast at the stern and the propeller of the airplane, and the RMS value of the Hausdorff distance calculated by the present invention is smaller than that of the existing method.

[0117] Therefore, the present invention adopts an imaging sonar three-dimensional reconstruction method based on differentiable space carving, which can quickly reduce the detected object according to the sonar image, and solves the lack of correspondence between occupancy and echo by presenting the echo probability instead of the iterative segmentation space of echo intensity, and combines the back propagation guided by rendering loss, so as to reconstruct faster.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.

Claims

1. A 3D reconstruction method of imaging sonar based on microscopic spatial carving, characterized in that: The following steps are involved: Step 1: Create the sonar framework F s , according to the performance of the sonar equipment, the distance f∈(r min ,r max ), azimuth θ∈(θ min ,θ max ) and elevation angle The formula for obtaining the echo position coordinates is as follows: In the above formula, p (s) is the echo coordinate; Step 2: Collect sonar images at different positions under known postures, process the collected sonar images, obtain the echo probability map EP from the original echo intensity map EI, and calculate the arc corresponding to the pixel q = (r, θ) Step 3: Establish differentiable spatial carving, including the occupancy probability model and rendering model constructed using occupancy probability grid and multi-resolution hash coding. Sampling is performed to form a set P j and the set L mn , use the occupancy probability model to output the occupancy probability of the current pixel, input the occupancy probability into the rendering model, perform single-pixel rendering, and obtain the predicted echo probability map; Step 4: Iterative training, establish a loss function group, given a set of echo probability maps with known postures, randomly select several echo probability maps to form a subset, and sample a set from the subset pixels, repeat step 3 to get the corresponding The predicted echo probability map is used to predict the echo probability map of each pixel. The occupancy probability model is iteratively updated based on the loss function group through the predicted echo probability map and the actual echo probability map until the loss function group meets the standard. Step 5: Use the trained differentiable spatial carving to reconstruct the detected object.

2. The imaging sonar 3D reconstruction method based on microscopic spatial carving according to claim 1 is characterized in that: In step 2, the process of obtaining the echo probability map EP from the echo intensity map EI is as follows: For the original echo intensity map, remove the object in it to obtain the noise background map, use the trained SCU neural network to reduce the noise of the original echo intensity map and the noise background map respectively, and obtain the reduced noise echo intensity map and the reduced noise background map, and use the reduced noise EI map to subtract the reduced noise background map to obtain the reduced noise and background removed echo intensity map; In the echo intensity map with noise reduction and background removal, calculate the arc corresponding to pixel q = (r, θ) The formula is as follows: Pixel q and arc The probability of at least one echo being present is transformed by the mapping function g, and the transformation formula is as follows: Where I(q) and Pr(q) are the intensity and pixel value of pixel q respectively, ∈ and μ are preset values. The echo intensity map with noise reduction and background removal is transformed into the echo probability map by formula (2).

3. The imaging sonar 3D reconstruction method based on microscopic spatial carving according to claim 2 is characterized in that: In step three, the specific process is as follows: When the value of pixel q=(r,θ) in the echo probability map is greater than zero, the arc At least one echo appears on the arc Sample m points to form a set P j , where 1≤j≤m; Draw m line segments from the sonar origin and the set P j The sampling points in are connected, and n points are sampled on each line segment to form a set L mn ; where the set P j With the set L mn Composed of sampling point set u; In the sonar frame F s In the set P j The corresponding coordinates are Convert it to world frame F w In the formula, the following is true: In the above formula, For the corresponding P j In world frame coordinates, is the transformation formula from sonar frame to world frame coordinate system; Establish a position encoding network E, through multi-resolution hash coding, input u to get the formula as follows: e=enc(u,ξ) Then, e is input into the occupancy probability model M(e;κ), where κ is the training parameter set of the occupancy network M. M(e;κ) outputs the occupancy probability, which is then input into the rendering model. In the rendered model, click P j The echo probability is recorded as in accordance with The occupancy probability and and O, along the line segment The projection rate estimation of is parameterized as follows: Among them, p o are the coordinates of point O, The occupancy probability of is given by the occupancy network M(e;κ), The occupancy probability is set to but The formula is as follows: Among them, T j Indicates along The transmittance of the sonar is estimated to reach The probability of a point, ρ is the volume density function, then Pr o (l jk ) and the function ρ are related as follows: Among them, ρ jk =ρ(t jk ), δ jk =t jk+1 -t jk , then {L jk }, the occupancy probability of 1≤k≤n is Pr o (l jk ), By occupying the network M(e;κ), the perspective rate T j The formula is as follows: The probability of obtaining at least one echo estimated for pixel q = (r, θ) is as follows: in, is the probability value of a pixel with at least one echo.

4. The imaging sonar 3D reconstruction method based on microscopic spatial carving according to claim 3 is characterized in that: In step 4, the process of iterative training is as follows: Establish a loss function group, including the rendering loss L of the echo probability prob , sparsity loss L sparsity and the parallax loss L of the surface normal norm , the formula is as follows: Among them, λ1 and λ2 are the weights of the corresponding losses, and the rendering loss L of the echo probability at the sampling pixel is prob as follows: in, Represents the set of sampled pixels, the sparsity loss L sparsity The formula is as follows: in, Represents the set of sampling points, and the formula is as follows: Where 1≤j≤m,1≤k≤n The parallax loss of the surface normal is obtained as follows: S={v∈B|Pr o (v)>τ} In the above formula, the threshold τ is selected based on the tracking error, S represents the set of sampling points with occupancy probability greater than the threshold τ, and v w It represents random sampling of points within a small radius of v, where n(v) is the unit normal vector at the position of point v.

Citation Information

Patent Citations

  • Real-time 3d reconstruction with a depth camera

    CN106105192A

  • Arc aperture radar imaging method based on antenna phase pattern compensation and radar

    CN112034460A