Image processing method based on underwater robot formation operation

By dynamically adjusting the robot position and camera parameters in the underwater robot formation operation, collecting multi-view image data, and using multi-view stereo vision and deep learning technology for image repair, the problem of the formation of dead corner areas in the underwater environment is solved, and the integrity and quality of the three-dimensional reconstruction of the target object is improved.

CN120219937APending Publication Date: 2025-06-27GUANGZHOU MARITIME INST
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510371879.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the underwater robot formation operation, due to the complexity and uncertainty of the underwater environment, the spatial distribution of image acquisition is uneven, forming dead corners or dead corner areas, affecting the completeness and accuracy of target reconstruction.

Method used

By dynamically adjusting the robot's position and camera parameters, multi-view and multi-spectral imaging data are collected and image preprocessed. The three-dimensional point cloud model is generated and optimized using multi-view stereoscopic vision technology, combining image stitching and fusion technology and deep learning-based image repair methods to achieve seamless repair and content generation of blind spots and occlusion parts.

Benefits of technology

It effectively solves the blind spots and occlusion problems in the three-dimensional reconstruction of underwater target objects, and improves the integrity and visual quality of reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219937A_ABST
    Figure CN120219937A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method based on underwater robot formation operation, and the method comprises the steps: determining the position and observation range of an observation target object according to the task demands of the underwater robot formation operation, and deploying underwater robots around the target object to form an annular observation network; the method comprises the following steps: dynamically adjusting the pose of a robot and camera parameters in an image acquisition process, acquiring a plurality of visual angles and a plurality of spectral images so as to obtain omnibearing target object image data, and preprocessing the image, including denoising, enhancing and correcting; the definition, the contrast ratio and the integrity of an image are analyzed, the collected image data are subjected to quality evaluation, the motion trail and the sampling strategy of the robot are adjusted in a self-adaptive mode, a dead angle area and a shielding part are subjected to emphasized observation, and supplementary image data are obtained for the dead angle area and the shielding part.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to an image processing method based on underwater robot formation operation. Background Art

[0002] In the underwater robot formation operation, multiple robots form a circular observation network around the target object and collect image data from an all-round perspective in order to improve the integrity of target reconstruction. However, in actual application scenarios, due to the complexity and uncertainty of the underwater environment, the position and posture of the robot formation are difficult to accurately control, resulting in uneven spatial distribution of the collected image data. In particular, in some areas of the target object, due to the limitations of the observation perspective, it is easy to form blind spots or dead angles for image acquisition, thereby affecting the integrity and accuracy of target reconstruction. In addition, the lighting conditions in the underwater environment vary greatly, including factors such as light intensity, spectral distribution, scattering and absorption, which will have a significant impact on image quality. In the dead angle area, due to the poor lighting conditions, the image quality is often lower, the noise level is high, the contrast is low, and the details are seriously lost. These factors further aggravate the difficulty of image restoration in the dead angle area. Therefore, how to minimize the appearance of the dead angle area by optimizing the formation configuration and observation strategy in the underwater robot formation operation, and at the same time, studying effective image restoration algorithms for the existing dead angle area images, has become a key technical problem that needs to be solved to improve the integrity of underwater target reconstruction. Summary of the invention

[0003] The present invention provides an image processing method based on underwater robot formation operation, which mainly includes:

[0004] According to the mission requirements of the underwater robot formation operation, the location and observation range of the observation target are determined, and the underwater robots are deployed around the target to form a ring observation network;

[0005] Dynamically adjust the robot's position and camera parameters during image acquisition, acquire all-round target image data by acquiring multiple viewing angles and multiple spectral images, and pre-process the images, including denoising, enhancement, and correction;

[0006] By analyzing the clarity, contrast and integrity of the image, the quality of the collected image data is evaluated, the robot's motion trajectory and sampling strategy are adaptively adjusted, blind spots and occluded parts are observed, and supplementary image data is obtained for blind spots and occluded parts;

[0007] Extract significant features from multi-view images, establish point correspondences between image pairs through feature matching, estimate camera poses and 3D point coordinates, then generate and optimize a 3D point cloud model through multi-view stereo vision, integrate multi-view images, and improve the surface information of dead zones and occluded parts through the integration of multi-view images;

[0008] For the dead zones and occluded parts in the 3D model, use image stitching and fusion techniques. By analyzing the overlapping relationship of adjacent view images, estimate the surface structure and texture information of the dead zones, and achieve seamless stitching and repair of the images to obtain a visually continuous and complete target reconstruction result;

[0009] During the image repair process, introduce a deep learning-based method. By training a convolutional neural network, learn the local features and global structures of a large number of target object images, perform semantic-level image repair and content generation on the dead zones and occluded parts, and comprehensively utilize the prior knowledge and context information of the target object to achieve image repair;

[0010] According to the repaired target image and 3D model, optimize the surface mesh of the 3D model. By simplifying and smoothing the mesh geometry structure, removing noise and artifacts, and adjusting the vertex positions, establish the mapping relationship between the image and the 3D model, calculate the texture coordinates of each mesh surface in the 3D model, map the image data to the corresponding mesh surfaces, and blend the multi-view images to obtain a complete 3D reconstruction result of the target object.

[0011] The technical solution provided by the embodiments of the present invention may include the following beneficial effects:

[0012] The present invention discloses an image processing method based on underwater robot formation operation. Aiming at problems such as dead zones and occluded parts during the observation of underwater target objects, the present invention conducts collaborative operations through robot formations, dynamically adjusts the robot poses and camera parameters, collects multi-view and multi-spectral imaging data, and performs image preprocessing. By analyzing the image quality, adaptively adjusts the robot motion trajectory and sampling strategy, and conducts key observations and supplementary data collection on the dead zones and occluded parts. Utilize multi-view stereo vision technology to generate and optimize a 3D point cloud model, combine image stitching and fusion technology and a deep learning-based image repair method to achieve seamless repair and content generation of the dead zones and occluded parts. Finally, optimize the repaired target image and 3D model, establish the mapping relationship between the image and the model, and obtain a complete 3D reconstruction result of the target object. The present invention effectively solves the problems of dead zones and occlusions in the 3D reconstruction of underwater target objects, and improves the integrity and visual quality of the reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1Flow chart of an image processing method based on underwater robot formation operation of the present invention. Detailed implementation manners

[0014] To further understand the content of the present invention, the present invention will be described in detail in combination with the accompanying drawings and embodiments. The following further elaborates on the present application in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for ease of description, only parts related to the invention are shown in the drawings.

[0015] As Figure 1 , an image processing method based on underwater robot formation operation in this embodiment may specifically include:

[0016] Step S101, according to the task requirements of underwater robot formation operation, determine the position and observation range of the observation target, and deploy the underwater robots around the target to form a circular observation network.

[0017] Based on the time delay of the acoustic signal emitted by the target to each base station, use phase difference measurement to obtain the distance values from the target to each base station, and calculate the three-dimensional coordinate values of the target by using the trilateration method in combination with the known coordinates of the base stations; for the three-dimensional coordinate values of the target, use a multi-beam sonar to scan and obtain the reflected echo signal, obtain the three-dimensional point cloud data through sonar data processing, and calculate the circular observation radius value by using the minimum circumscribed sphere; divide the grid units according to the circular observation radius value, measure the transmission time delay, signal attenuation degree, and noise interference degree of the underwater acoustic channel for the grid units, and select the grid units with a score higher than the preset threshold score value as valid observation points through weighted scoring; for the valid observation points, generate a Bezier curve navigation path in combination with the kinematic constraints of the underwater robot, and generate smooth navigation trajectory points by limiting the maximum steering angle between path points; calculate the priority score value according to the distance from the current position of the underwater robot to the navigation trajectory point, and assign a navigation trajectory point identification number to the underwater robot according to the priority score value to generate a multi-robot collaborative observation navigation planning scheme.

[0018] Specifically, the acoustic signals of the target are collected by multiple underwater acoustic positioning base stations. The time delay from the target's transmitted signal to each base station is calculated through phase difference measurement. Based on the sound wave propagation speed, the distances from the target to each base station are obtained. The coordinates of the target in the three-dimensional space coordinate system are calculated using the trilateration method in combination with the known coordinates of the base stations. The sliding mean filter is used to smooth the target's movement trajectory. For the target, a multi-beam sonar is used to scan at different angles to obtain the target's reflected echo signals. The three-dimensional point cloud data of the target is obtained through sonar data processing. The radius value of the circumscribed sphere is calculated using the minimum circumscribed sphere, and the annular observation radius is obtained by multiplying it by a preset observation coefficient. The observation point positions are divided at an interval angle of 30 degrees within the annular observation radius. The annular grid cell division parameters are obtained from the observation point position data. The underwater acoustic channel transmission time delay, signal attenuation degree, and noise interference degree are measured for each grid cell. The channel quality of each grid cell is quantitatively evaluated using weighted scoring, and the grid cells with a scoring value higher than the preset threshold of 750 points are selected as valid observation points. For the valid observation point position data, a Bezier curve navigation path is generated in combination with the kinematic constraints of the underwater robot. Sharp turns are avoided by limiting the maximum steering angle between path points. An identification number of the underwater robot is assigned to the observation points, and the priority is set according to the navigation distance from the robot's current position to the specified observation point. The navigation trajectories of each underwater robot to the observation points are planned. When the underwater acoustic positioning base stations collect the acoustic signals of the target, 40 kHz ultrasonic waves are used as the positioning signals. The cross-correlation processing is used for phase difference measurement to obtain the time delay value. The sound wave propagation speed is taken as 1500 m / s. The base station coordinates are based on the GPS data of the deployment position to establish a rectangular coordinate system. After the initial position of the target is measured using the trilateration method, 32 points are taken to form a sliding time window to smooth the target's movement trajectory data. When the multi-beam sonar scans, the working frequency is set to 200 kHz, the beam angle is 1 degree, and the sonar collects data every 45 degrees around the target during scanning. A total of 8 angles of echo data are obtained. The sonar data processing uses time-domain integration to obtain the reflection intensity value. The boundary points of the target are intercepted through the intensity threshold to form point cloud data. The density of the point cloud data is 500 points per cubic meter. The radius value of the minimum circumscribed sphere is calculated using the sphere center iteration. The observation coefficient is taken as 3 times the radius of the circumscribed sphere. 12 observation point positions are divided on the annular observation path at an angle interval of 30 degrees. The annular grid cell division uses polar coordinate grids, with a radial interval of 10 m and an angle interval of 30 degrees. Detection signals are sent to each grid cell, and the round-trip time delay, received signal amplitude, and background noise intensity are recorded. The time delay weight is 0.4, the amplitude weight is 0.4, and the noise weight is 0.2. The grid cell score is calculated by weighted calculation. The full score of the score is 1000 points, and the grid cells with a score higher than 750 points are selected as valid observation positions.The Bezier curve navigation path is generated based on the third-order curve. Four adjacent waypoints are taken as control points. The curve sampling interval is 1 meter. The maximum turning angle between adjacent sampling points is limited to within 30 degrees. The robot kinematic constraints include a maximum speed of 2 meters per second and a maximum acceleration of 0.5 meters per square second. The robot number is assigned in descending order according to the distance from the observation point, and the priority value is inversely proportional to the distance.

[0019] Step S102, dynamically adjust the robot's posture and camera parameters during the image acquisition process, acquire omnidirectional target object image data by acquiring multiple viewing angles and multiple spectral images, and pre-process the image, including denoising, enhancement and correction.

[0020] A spherical coordinate system with the target object as the center is established according to the three-dimensional spatial coordinates of the target object, and sampling point positions are divided on the spherical surface according to preset latitude and longitude intervals, and the sampling point positions are driven to the sampling point positions by a servo motor; the reflectivity value of the surface of the target object is measured for the sampling point positions, and exposure parameters are obtained from a pre-established reflectivity exposure correspondence table, and the exposure parameters are used to collect visible light images and near-infrared images; Haar wavelet decomposition is performed on the visible light image and the near-infrared image to obtain decomposition coefficients, and the high-frequency components of the decomposition coefficients are removed by using a preset soft threshold coefficient to reconstruct a corrected image group; an image quality score is calculated for the corrected image group, and if the image quality score is lower than a preset threshold, the corresponding image is eliminated to obtain a group of images to be spliced; scale-invariant feature points are extracted from the group of images to be spliced, and feature point matching pairs whose errors exceed the preset threshold are removed by random consistency sampling, and a least squares optimization function is constructed based on the feature point matching pairs to solve the splicing transformation matrix.

[0021] Specifically, according to the three-dimensional spatial coordinates of the target object, the viewpoint position is generated in the spherical coordinate system established with the center of the target object as the center. Sampling points are divided on the spherical surface with a radius of 5 meters at intervals of 30 degrees in longitude and 20 degrees in latitude. The servo motor speed is set to 30 degrees per second to drive the robot pan-tilt to rotate. The pitch angle and yaw angle are adjusted to the specified viewpoint position through encoder feedback. When the position error is less than 0.1 degree, the adjustment stops. For the said viewpoint position, the surface reflectivity of the target object is measured, and the exposure parameters are queried from the pre-established reflectivity-exposure time correspondence table. When the target distance is 5 meters, a 50-mm focal length is selected, and the aperture is automatically adjusted to F4 with the change of the focal length. Visible light and near-infrared dual-spectrum images are synchronously collected to generate the original image group. The visible light image and near-infrared image in the original image group are respectively decomposed by 4-layer Haar wavelet. The soft threshold coefficient is set to 0.5 to remove the high-frequency noise components. The image entropy value is calculated for the reconstructed image. When the entropy value is greater than 5.0, it is determined that the denoising is effective. The contrast-limited histogram equalization is used to enhance the image, and the limiting coefficient is set to 0.02. The geometric distortion of the image is corrected based on the camera internal parameter data to generate the corrected image group. For the said corrected image group, the image sharpness, signal-to-noise ratio, and contrast are calculated to obtain the quality score. Images with a score lower than 80 are excluded. Scale-invariant features are used to extract the image feature points. The random sample consensus is used to remove the matching pairs with a feature point matching error greater than 2 pixels. The least-squares optimization function is constructed based on the coordinates of the matching feature points. The upper limit of the iteration number is set to 500 times, and the convergence threshold is set to 0.001 to solve the image stitching transformation matrix to complete the multi-view image stitching. During the sampling process in the spherical coordinate system, the total number of viewpoints is calculated as 360 degrees in longitude divided by the interval of 30 degrees to get 12 sampling points, and 180 degrees in latitude divided by the interval of 20 degrees to get 9 sampling points, for a total of 108 viewpoint positions. The robot pan-tilt uses a digital servo with the model number MG995, and the rotation resolution is 0.1 degree. The rotation speed is controlled by the pulse-width modulation signal at 30 degrees per second. The position closed-loop control uses a proportional coefficient of 0.5. When the position error is less than 0.1 degree, it is determined that the target position has been reached. The surface reflectivity of the target object is measured by a radiometer. The measurement wavelength range is 400 - 1000 nm, and the resolution is 10 nm. When the average reflectivity of the target object is measured to be 0.6, the exposure time is set to 1 / 100 second by looking up the table. The camera uses a CMOS sensor with a pixel size of 5.5 microns and a resolution of 2048×1536, and is equipped with a 50-mm fixed-focus lens. When collecting the dual-spectrum images, the sensitivity of the visible light camera is set to ISO400, and the near-infrared camera operates in the 800 - 1000 nm band. Haar wavelet is used for image denoising. The low-frequency coefficients and three layers of high-frequency coefficients are selected, and the soft threshold value is 0.5 times the standard deviation of each layer of coefficients. The entropy value of the reconstructed image is increased from the original 4.2 to 5.5. The contrast limit of the image enhancement is 0.02 times the original contrast to avoid over-enhancement.The image quality score is calculated by weighted averaging, with sharpness accounting for 0.4, signal-to-noise ratio for 0.3, and contrast for 0.3, on a 100-point scale. Images with an actual score of over 85 account for 70%. Feature point matching calculates descriptors using 8×8 pixel blocks. The random sampling quantity is set at 10% of the total number of feature points. When iteratively optimizing, the gradient step size for each calculation is 0.01, and iteration stops when the change amount for 10 consecutive iterations is less than 0.001.

[0022] In step S103, the quality of the acquired image data is evaluated by analyzing the sharpness, contrast, and integrity of the image, and the motion trajectory and sampling strategy of the robot are adaptively adjusted to focus on observing dead zones and occluded parts, and supplementary image data is obtained for the dead zones and occluded parts.

[0023] The Laplacian operator is used to calculate the sharpness value of the image edge, the contrast value of the image is calculated by the mean variance of pixel block grayscale, and the integrity value of the image contour is calculated based on the continuity of edge points. Weights are assigned to the sharpness value, contrast value, and contour integrity value respectively to obtain the image quality score. The three-dimensional space grid is obtained by mapping according to the camera projection matrix based on the image quality score. The grid to be resampled is marked for the area where the image quality score is lower than the threshold. For the grid to be resampled, a three-dimensional grid map is constructed, and the dead zone is obtained by counting the sampling coverage times. The occluded area is judged based on the intersection position of the ray and the depth map. For the dead zone and the occluded area, the sampling probability distribution is calculated according to the grid density, and candidate sampling points are generated around the high-probability area according to the probability weight. The ordered sampling point sequence is obtained by sorting the distances from the sampling points to the robot. For the ordered sampling point sequence, the unvisited sampling point with the closest distance is selected as the target position, and adjacent sampling points are connected by a cubic spline curve, and the trajectory satisfying the motion constraint conditions is generated according to the curve parameters.

[0024] Specifically, calculate the image edge sharpness value according to the 5×5 Laplacian operator, calculate the contrast value using the gray-scale mean variance of 16×16 pixel blocks, calculate the contour integrity value through the continuity of edge points, assign weights of 0.4, 0.3, and 0.3 to the three indicators respectively to calculate the image quality score, and map the area with a quality score lower than 75 points to the three-dimensional space grid through the camera projection matrix and mark it as the grid to be resampled. For the grid to be resampled, construct a three-dimensional grid map with a side length of 0.1 meters based on the target point cloud data, count the sampling coverage times of each grid, mark the grid with a sampling times less than 3 times as the dead angle area, emit a ray from the current viewpoint position to the target point cloud, judge the intersection position of the ray and the depth map, and mark the grid where the depth value at the intersection is less than the target point depth value as the occlusion area. For the dead angle area and the occlusion area, calculate the sampling probability distribution based on the grid density, set the probability weight of the dead angle grid to 0.7, set the probability weight of the occlusion grid to 0.3, generate candidate sampling points within the range of 0.5 meters to 5 meters around the high-probability area, the number of sampling points is 2 times the total number of dead angle and occlusion grids, and sort based on the distance from the sampling point to the current position of the robot. For the sorted sampling points, select the sampling point with the closest distance and not visited as the target position, use a cubic spline curve to connect adjacent sampling points to generate a smooth trajectory, select the adjacent three sampling points as the control points, adjust the curve parameters to control the movement speed not exceeding 0.5 meters per second, the turning radius not less than 0.8 meters, and the attitude angle change rate not exceeding 30 degrees per second to obtain supplementary image data. In image quality assessment, use the 5×5 Laplacian operator to extract the edges of the image. When the image edge response value is greater than the threshold of 50, it is determined as a clear edge point, and the proportion of the number of edge points is counted as the sharpness value. Divide the 1024×768 image into 16×16 pixel blocks, a total of 48×48 blocks, calculate the difference between the gray-scale mean of each block and the overall mean to obtain the local contrast value, detect the straight line segments of the target contour through the Hough transform, determine that the distance between the endpoints of adjacent line segments is less than 5 pixels as continuous, and the proportion of the length of the continuous line segments to the total length of the contour is used as the integrity value. The quality score uses a 100-point system. In actual image processing, the sharpness score is 82 points, the contrast score is 76 points, and the integrity score is 71 points. When mapping the grid, the camera focal length is 50 mm, the image sensor size is 1 / 2.3 inch, and the image coordinates are mapped to the world coordinates through the projection matrix. The three-dimensional grid map covers a space range of 3 meters × 3 meters × 2 meters, with a total of 6000 grid cells. Sampling statistics show that the dead angle area accounts for 15%, and the occlusion area accounts for 25%. The resolution of the depth map is the same as that of the image, and the ray is advanced step by step using the Bresenham algorithm with a step size of 1 cm. In the sampling probability distribution, the value of the dead angle area is 0.4 - 0.7, the value of the occlusion area is 0.2 - 0.3, and 400 candidate sampling points are generated through roulette.When planning the trajectory, the distance between control points is 0.5 meters, the curve parameters are normalized to the 0-1 interval, the speed curve adopts a cosine shape to achieve smooth acceleration and deceleration, the turning area adopts an arc transition with a radius of 0.8 meters, the attitude angular velocity is planned by a trapezoidal curve, and the acceleration is limited to 5 degrees per second squared.

[0025] In step S104, significant features are extracted from the multi-view images, the point correspondence between the image pairs is established through feature matching, the camera pose and the three-dimensional point coordinates are estimated, and then the three-dimensional point cloud model is generated and optimized through multi-view stereo vision. The multi-view images are integrated, and the surface information of the dead angle area and the occluded part is improved through the integration of the multi-view images.

[0026] Feature points are extracted from the multi-view images according to the Harris corner response function. The gradient direction histogram descriptor is calculated for the feature points, and the nearest neighbor ratio threshold is used to match the histogram descriptor to obtain the feature matching pairs. For the feature matching pairs, the essential matrix is estimated by the five-point algorithm, and the camera rotation matrix and translation vector are obtained by solving the essential matrix through singular value decomposition. The three-dimensional coordinates of the feature points are calculated by triangulation according to the camera rotation matrix and translation vector, and the three-dimensional coordinate points are obtained by screening through three indicators: reprojection error, parallax angle, and tracking length. For the three-dimensional coordinate points, the space is divided by a hierarchical voxel grid, the initial pose is determined by centroid alignment, and the point clouds under adjacent views are registered based on the iterative closest point algorithm to obtain the registered point cloud. For the registered point cloud, a local neighborhood search range is established, the local surface is fitted by moving least squares based on the normal vector consistency constraint, and the Poisson reconstruction is used to fill the missing area surface to obtain a continuous surface.

[0027] Specifically, significant feature points are extracted from multi-view images according to the Harris corner response function. The corner response threshold is set to 0.05, and the distribution of feature point response values is calculated. The 20% feature points with the smallest local response values are removed, and more than 300 stable feature points are retained in each image. In the 32×32 pixel area around the feature points, the gradient direction histogram descriptor is calculated according to 8×8 sub-regions. The nearest neighbor ratio threshold of 0.6 is used to match the feature descriptors, and the feature pairs with a matching distance ratio greater than 0.8 are removed. For the feature matching pairs, the five-point algorithm is used to estimate the essential matrix. The inlier determination threshold is set to 2 pixels. The feature matching pairs are screened based on the reprojection error. The camera rotation matrix and translation vector are solved by singular value decomposition. The three-dimensional coordinates of the feature points are calculated by triangulation. Three indicators, namely reprojection error, parallax angle, and tracking length, are calculated for the three-dimensional points. The minimum parallax angle is set to 5 degrees, and the three-dimensional points with a reprojection error less than 1 pixel and a tracking length greater than 3 frames are retained. For the three-dimensional coordinate points, the space is divided by a hierarchical voxel grid. The large voxel size of 0.05 meters is used for sparse point cloud registration, and the small voxel size of 0.01 meters is used to retain geometric details. The point clouds under adjacent views are registered based on the iterative closest point algorithm. The centroid alignment is used to determine the initial pose. The upper limit of the iteration number is set to 50 times, and the convergence threshold is set to 0.001 meters. For the registered point cloud, a local neighborhood radius of 0.02 meters is established for each point, and the local surface is fitted by moving least squares based on the normal vector consistency constraint, interpolated to the grid resolution of 0.01 meters, and a continuous surface is generated. The Poisson reconstruction is used to fill the missing area surface in the dead corner area and the occlusion part. The Poisson equation is solved using an octree with 8 layers, and the weight coefficient is set to 1.5. In an image with a resolution of 1920×1080, the window size of the Harris corner detector is set to 3×3, and the Gaussian smoothing kernel size is set to 5×5. On average, 3500 corners are calculated for each image. After sorting by the response value, the first 300 stable points are retained. These feature points are evenly distributed in the image, and the distance between adjacent feature points is more than 20 pixels. When calculating the feature descriptor, 64 sub-blocks are divided in the 32×32 pixel area, and the gradient direction histogram of 8 bins in each sub-block is statistically calculated to obtain a 512-dimensional feature vector. On average, 250 effective corresponding points are obtained by matching between two images. When estimating the essential matrix by the five-point algorithm, the least squares optimization is used for solving, and the inlier ratio reaches 85%, and the average reprojection error is 0.8 pixels. In the three-dimensional point reconstruction, the overlap degree between adjacent views is kept above 60%, the average parallax angle is 12 degrees, and the number of three-dimensional points is about 5000. The point cloud registration uses a two-level voxel grid of large and small sizes. On average, 15 points are included in the large voxel for rough registration, and on average, 3 points are included in the small voxel for fine registration. The registration process converges after 35 iterations, and the average value of the final point pair distance is 0.8 mm.When performing surface reconstruction, there are on average 20 points in the local neighborhood. If the deviation of the normal vector angle is less than 30 degrees, it is determined to be the same surface. After interpolation, the distance between grid vertices is 1 centimeter, and the curvature change is continuous in the surface smoothness evaluation. In Poisson reconstruction, the number of nodes in each layer of the octree is about 4 times that of the upper layer, and the side length of the grid at the leaf nodes is about 2 millimeters. The area of the filled dead corners and occluded areas accounts for 15% of the total surface area.

[0028] In step S105, for the dead corner areas and occluded parts in the three-dimensional model, using image stitching and fusion technology, by analyzing the overlapping relationship of adjacent perspective images, estimate the surface structure and texture information of the dead corner areas, achieve seamless stitching and repair of the images, and obtain a visually continuous and complete target reconstruction result.

[0029] Project the vertex coordinates of the three-dimensional point cloud model onto the image plane to obtain the proportion of the overlapping area between perspective images. If the proportion of the overlapping area is greater than the preset threshold, adjacent perspective image pairs are obtained; for the adjacent perspective image pairs, calculate the texture features of the overlapping area using the multi-scale gradient histogram, calculate the texture feature similarity score through the cross-correlation coefficient. If the similarity score is greater than the preset threshold, the image pairs to be matched are obtained; for the image pairs to be matched, perform block matching using a multi-layer image pyramid, obtain the disparity field through the region growing algorithm, and smooth the disparity field based on total variation regularization to obtain a smoothed disparity field; according to the smoothed disparity field, generate a dense deformation field using thin plate spline interpolation, and fuse the dense deformation field and the image pairs to be matched through a Laplacian pyramid to obtain a stitched image; for the dead corner areas and occluded areas in the stitched image, construct a repair mask using edge continuity constraints, and propagate the texture features of the area around the mask through Poisson image editing to obtain the repaired stitched image.

[0030] Specifically, according to the vertex coordinates of the three-dimensional point cloud model projected onto the image planes of each perspective, calculate the proportion of the overlapping area between images. Set the overlap threshold to 50%, and filter out adjacent perspective image pairs. In the overlapping area, use the multi-scale gradient histogram to extract texture features. Calculate the 8-direction gradient histogram at three scales of 8×8, 16×16, and 32×32 respectively. After normalization, calculate the texture similarity score through the cross-correlation coefficient. Image pairs with a score greater than 0.7 enter the subsequent matching. For the said image pairs, use a three-layer image pyramid for block matching, and refine the search layer by layer from 64×64 to 16×16 pixels. The block matching cost combines the pixel brightness difference and the gradient direction difference, and the weights are set to 0.6 and 0.4 respectively. Use region growing at each layer of the pyramid to determine the disparity field, and construct an optimization objective function based on total variation regularization to smooth the disparity field. For the smoothed disparity field, use thin plate spline interpolation to generate a dense deformation field. Set the interpolation grid spacing to 8 pixels, add a bending energy constraint term, and take the weight coefficient as 0.1. Construct an adaptive hybrid weight based on the image gradient at the boundary of the overlapping area, and achieve image stitching through bottom-up fusion of a four-layer Laplacian pyramid. For the stitched image, construct a repair mask in the dead corner and occlusion areas, adaptively expand the mask boundary by 8 pixels based on edge continuity and texture consistency, extract structure and texture feature blocks within a range of 80×80 pixels around the mask, with the feature block size of 16×16 pixels, and use Poisson image editing to propagate the texture, and use gradient constraints at the boundary to ensure structural continuity. In a 2048×1536 resolution image, the average coverage area of the overlapping area of adjacent perspective image pairs reaches 65%. The multi-scale gradient histogram obtains a 64-dimensional feature vector at the smallest 8×8 scale, a 256-dimensional feature vector at the medium 16×16 scale, and a 1024-dimensional feature vector at the largest 32×32 scale. The feature vectors of the three scales are concatenated and then normalized. The average cross-correlation coefficient of adjacent images reaches 0.85, and 85% of the image pairs are retained after screening. The pyramid matching starts from the top 64×64 pixel block, and the matching window searches within a range of 128×128. The pixel brightness difference is calculated using zero-mean normalized correlation, and the gradient direction difference is calculated using the 8-direction histogram similarity. Region growing expands from the highest correlation point, and an average of 500 blocks are expanded per layer. The thin plate spline interpolation uses a 4×4 grid to divide the control points, and the bending energy constraint minimizes the sum of the squares of the second derivatives of the surface. The adaptive hybrid weight is set according to the image gradient amplitude, and the mixing bandwidth increases to 48 pixels at the places with larger gradients. The scale ratio of each layer of the Laplacian pyramid is 2, and the resolution of the bottom layer is the same as the original image. The average area of the repair mask region accounts for 12% of the image. The feature block matching uses normalized cross-correlation, finds the three most similar candidate blocks within the search range, and uses the multi-grid method to solve the Poisson equation, with the convergence threshold set to 0.001 and the number of iterations limited within 100 times.

[0031] Step S106, during the image restoration process, introduce a deep learning-based method. By training a convolutional neural network, learn the local features and global structures of a large number of target object images, perform semantic-level image restoration and content generation on the dead corner areas and occluded parts, and comprehensively utilize the prior knowledge and context information of the target object to achieve image restoration.

[0032] Obtain a binary mask with annotated dead corner areas and occluded parts, and construct a five-layer encoder-decoder network according to the binary mask. The encoder extracts features through a 3×3 convolutional kernel, and the decoder uses transposed convolution for upsampling to restore the resolution; add residual blocks at the skip connections of the encoder-decoder network. The residual blocks calculate the feature weights through a channel attention mechanism to obtain feature channels with a significance greater than a preset threshold; for the feature channels, use a contrast loss function to calculate the cosine distance between the restored area and the real image patch. If the cosine distance is less than a preset distance threshold, apply a contrast loss to the feature pair, and calculate the overall loss by combining the reconstruction loss and the perceptual loss; according to the overall loss, evaluate the authenticity of the local area through a two-layer discriminator. The discriminator uses a downsampling convolutional block and a rectified linear unit with a leakage coefficient to obtain the pixel-level error and the feature-level adversarial loss; for the pixel-level error and the feature-level adversarial loss, construct a multi-head attention layer to calculate the context features at different scales. The multi-head attention layer calculates the weight matrix through a 1×1 convolution to obtain the attention feature map; generate a spatial attention mask according to the attention feature map. The spatial attention mask guides the feature propagation based on the edge significance to obtain the final restored image.

[0033] Specifically, based on the binary masks of the annotation dead zones and occluded parts, a five-layer encoder-decoder network is constructed. The encoder uses 3×3 convolutional kernels to extract features with a stride of 2, and the number of channels is 32, 64, 128, 256, and 256 in sequence. The decoder uses transposed convolution for upsampling to restore the resolution, and residual blocks are added at the skip connections. Each residual block contains two 3×3 convolutional layers. The feature weights are calculated through the channel attention mechanism, and the feature channels with a significance greater than 0.3 are retained. For the network structure, structural constraints are introduced through the contrast loss function. The cosine distance between the inpainting region and the real image patch is calculated in the feature space, and the distance threshold is set to 0.6. The contrast loss is applied to the feature pairs with a distance less than the threshold. The reconstruction loss is calculated in the pixel space, and the perceptual loss is calculated by extracting the features of the third and fourth convolutional layers. The weights are set to 0.4, 0.3, and 0.3 respectively in combination with the structural similarity evaluation. For the inpainting results, a two-layer discriminator is used to evaluate the authenticity of local regions. Each layer contains 4 downsampling convolutional blocks with a pooling kernel size of 3×3, and the activation function uses the rectified linear unit with a leakage coefficient of 0.2. The loss function combines the pixel-level mean square error and the feature-level adversarial loss, and the weight ratio is set to 1:0.1. For the inpainted image, a multi-head attention layer is constructed to capture long-range dependencies. The size of the attention map is 1 / 4 of the feature map. Four attention heads are used to focus on different-scale contexts respectively. The weight matrix is calculated through 1×1 convolution, and the temperature parameter is set to 0.07. The spatial attention mask generated based on edge saliency is combined to guide feature propagation. Based on an input image of 512×512 pixels, the average proportion of dead zones and occluded regions is 15%. The network input is preprocessed by zero-mean normalization. The size of the 32-channel feature map in the first layer of the encoder is 256×256, and finally a 16×16 feature map is obtained after 5 layers of downsampling. 256 weight coefficients are calculated through channel attention, and 87 channels with weights greater than 0.3 are retained. The residual block uses two 3×3 convolutional layers with 256 channels, and a batch normalization layer is added in the middle. In the calculation of the contrast loss, the average cosine distance is 0.52, and the feature pairs with a distance lower than the threshold of 0.6 account for 73% of the total. The reconstruction loss uses the L1 norm to calculate the pixel difference. The perceptual loss extracts features from the third and fourth convolutional layers of the VGG16 network. The structural similarity is calculated within an 11×11 window. The input of the first layer of the discriminator is a 32×32 image patch, and a 2×2 output is obtained through 4 downsampling layers. The input of the second layer of the discriminator is a 64×64 image patch. In the adversarial training, the weight of the adversarial term in the generator loss is set to 0.1, and the accuracy of the discriminator is stable between 52% and 56%. In the multi-head attention layer, the four attention heads focus on the context information in the ranges of 8×8, 16×16, 32×32, and 64×64 respectively. The attention map is reduced to 64 channels through 1×1 convolution. The spatial attention mask is generated based on the edge saliency detected by the Sobel operator, and the saliency threshold is set to 0.15. The peak signal-to-noise ratio of the inpainted image reaches 32 dB, and the structural similarity index reaches 0.91.

[0034] In step S107, according to the repaired target image and the three-dimensional model, optimize the surface mesh of the three-dimensional model. By simplifying and smoothing the mesh geometric structure, remove noise and artifacts, and adjust the vertex positions to establish a mapping relationship between the image and the three-dimensional model. Calculate the texture coordinates of each mesh surface in the three-dimensional model, map the image data to the corresponding mesh surfaces, and blend the multi-view images to obtain a complete three-dimensional reconstruction result of the target object.

[0035] Obtain the three-dimensional mesh vertex coordinates and connection relationships, and calculate the mean curvature value and normal vector of each vertex according to the three-dimensional mesh vertex coordinates; for the region where the mean curvature value is less than the preset curvature threshold, use local quadratic surface least squares fitting to obtain the smooth mesh vertex coordinates; according to the smooth mesh vertex coordinates, calculate the geometric error energy and texture error energy of each edge, and use the edge collapse algorithm based on quadratic error metric to obtain the simplified mesh; for the simplified mesh, construct a parameterized mapping based on the discrete coordinate field, impose area preservation constraints and angle preservation constraints at the boundary fixed points with the maximum curvature value, and obtain the parameterized mesh through alternating iterative optimization; project the parameterized mesh with multi-view images, calculate the projection overlap degree and viewing angle of the mesh patches; construct a blending weight map according to the projection overlap degree and viewing angle, and solve the blending weight map using the Poisson equation to obtain a seamless texture image.

[0036] Specifically, according to the three-dimensional grid vertex coordinates and connection relationships, calculate the mean curvature and normal vector of each vertex. Set the curvature threshold to 0.1. Keep the vertices with curvature greater than the threshold unchanged, and use local quadratic surface least squares fitting for mesh smoothing in the area with curvature less than the threshold. Set the smoothing iteration times to 5 times, and limit the vertex displacement within the range of 0.1 times the average edge length in each iteration. Calculate the Hausdorff distance before and after smoothing. When the distance exceeds 0.05 times the diagonal length of the model bounding box, roll back the smoothing operation. For the smoothed mesh, use the edge collapse algorithm based on quadratic error metric for simplification. Calculate the geometric and texture error energies of each edge. The geometric error is calculated by the sum of the squares of the distances from the vertex to the adjacent faces, and the texture error is calculated by the deviation of the parameter space coordinates. Set the energy threshold to 0.01, and stop collapsing when the error caused by simplification exceeds the threshold. Finally, retain 30% of the original number of patches. For the simplified mesh, construct a parameterized mesh based on the discrete coordinate field. Set boundary fixed points at the 8 vertices with the largest curvature values, apply area-preserving and angle-preserving constraints, and optimize the parameterized results through alternating iterations. Calculate the triangle distortion rate during the iteration process. Stop when the maximum distortion rate is less than 0.1 or the number of iterations exceeds 100 times. For the parameterized mesh, project multi-view images onto the mesh patches, calculate the overlap degree and viewing angle of the projection area for each patch, construct a mixed weight map based on the overlap degree and viewing angle quality. Assign a weight of 1.0 when the overlap degree is greater than 0.8, use a cosine function for smooth transition when the overlap degree is between 0.5 and 0.8, and linearly decay the weight to 0 when the viewing angle is greater than 60 degrees. Use Poisson fusion to eliminate texture seams. In a three-dimensional mesh model with 50,000 vertices, the mean curvature is calculated by the angle between the normal vectors of the vertex and the adjacent faces. 90% of the vertex curvature values are distributed between 0.05 and 0.15. After setting the threshold to 0.1, 15% of the feature vertices are retained. During the smoothing process, use a 5×5 vertex neighborhood to fit a quadratic surface, and control the root mean square value of the fitting error within 0.02 mm. The average displacement per iteration is 0.08 times the edge length. The maximum Hausdorff distance appears at sharp corners, accounting for 0.03 of the diagonal length of the bounding box. When simplifying the mesh, the initial priority queue of edge collapse contains 75,000 edges. The geometric error of each edge is represented by a 4×4 matrix, and the texture error is calculated by projection transformation in the parameter space. The simplification process terminates early when the energy accumulates to 0.008, leaving 15,000 patches. For parameterization, use 8 fixed points to divide the boundary, select vertices with curvature values exceeding 0.12 as fixed points, set the area-preserving weight to 0.6, and the angle-preserving weight to 0.4. During the iterative optimization, the triangle distortion rate of 95% drops below 0.08, and it converges after an average of 65 iterations.When performing texture mapping, the image resolution is 2048×2048. On average, each patch is covered by 3.5 viewpoints. The overlapping area accounts for 72% of the total area. The cosine value of the viewing angle is distributed between 0.5 and 1.0. Poisson fusion uses multi-scale decomposition and sets 4 decomposition levels. Iteration stops when the residual is less than 1%.

[0037] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. An image processing method based on underwater robot formation operation, characterized in that: The method comprises: According to the mission requirements of the underwater robot formation operation, the location and observation range of the observation target are determined, and the underwater robots are deployed around the target to form a ring observation network; Dynamically adjust the robot's position and camera parameters during image acquisition, acquire all-round target image data by acquiring multiple viewing angles and multiple spectral images, and pre-process the images, including denoising, enhancement, and correction; By analyzing the clarity, contrast and integrity of the image, the quality of the collected image data is evaluated, the robot's motion trajectory and sampling strategy are adaptively adjusted, blind spots and occluded parts are observed, and supplementary image data is obtained for blind spots and occluded parts; Extract salient features from multi-view images, establish point correspondences between image pairs through feature matching, estimate camera poses and 3D point coordinates, generate and optimize 3D point cloud models through multi-view stereo vision, integrate multi-view images, and improve surface information of blind spots and occluded parts through the integration of multi-view images; For the blind spots and occluded parts in the 3D model, the image stitching and fusion technology is used to analyze the overlapping relationship of adjacent view images, estimate the surface structure and texture information of the blind spots, achieve seamless stitching and restoration of images, and obtain visually continuous and complete target reconstruction results. In the process of image restoration, a deep learning-based method is introduced to train a convolutional neural network to learn the local features and global structures of a large number of target images, perform semantic-level image restoration and content generation on blind spots and occluded parts, and comprehensively utilize the prior knowledge and contextual information of the target to achieve image restoration. According to the repaired target image and 3D model, the surface mesh of the 3D model is optimized. By simplifying and smoothing the mesh geometry, removing noise and artifacts, and adjusting the vertex positions, a mapping relationship between the image and the 3D model is established. The texture coordinates of each mesh surface in the 3D model are calculated, the image data is mapped to the corresponding mesh surface, and the multi-view images are mixed to obtain a complete 3D reconstruction result of the target object.

2. The method according to claim 1, characterized in that The method of determining the position and observation range of the observation target object according to the task requirements of the underwater robot formation operation, and deploying the underwater robots around the target object to form a ring observation network includes: According to the time delay of the target object transmitting the acoustic signal to each base station, the phase difference measurement is used to obtain the distance value from the target object to each base station, and the three-dimensional coordinate value of the target object is calculated by the trilateral positioning method combined with the known coordinates of the base station; For the three-dimensional coordinate value of the target object, a multi-beam sonar scan is used to obtain a reflected echo signal, three-dimensional point cloud data is obtained through sonar data processing, and a minimum circumscribed sphere calculation is used to obtain a circular observation radius value; Divide the grid units according to the circular observation radius value, perform underwater acoustic channel transmission delay measurement, signal attenuation measurement and noise interference measurement on the grid units, and select the grid units with a score value higher than a preset threshold as valid observation points through weighted scoring; For the effective observation points, a Bezier curve navigation path is generated in combination with the kinematic constraints of the underwater robot, and a smooth navigation trajectory point is generated by limiting the maximum steering angle between the path points; A priority score is calculated based on the distance from the current position of the underwater robot to the navigation track point, and a navigation track point identification number is assigned to the underwater robot according to the priority score to generate a multi-robot collaborative observation navigation planning scheme.

3. The method according to claim 1, characterized in that The robot's posture and camera parameters are dynamically adjusted during the image acquisition process, and a full range of target image data is acquired by acquiring multiple viewing angles and multiple spectral images, and the image is preprocessed, including denoising, enhancement and correction, including: A spherical coordinate system with the target object as the center is established according to the three-dimensional spatial coordinates of the target object, sampling point positions are divided on the spherical surface according to preset longitude and latitude intervals, and the sampling point positions are driven to the sampling point positions by a servo motor; Measuring the reflectivity value of the surface of the target object at the sampling point position, obtaining exposure parameters from a pre-established reflectivity exposure correspondence table, and using the exposure parameters to collect a visible light image and a near infrared image; Performing Haar wavelet decomposition on the visible light image and the near infrared image to obtain decomposition coefficients, and using a preset soft threshold coefficient to remove high-frequency components from the decomposition coefficients and then reconstructing them to obtain a corrected image group; Calculating an image quality score for the corrected image group, and if the image quality score is lower than a preset threshold, eliminating the corresponding image to obtain an image group to be stitched; Scale-invariant feature points are extracted from the image group to be stitched, feature point matching pairs whose errors exceed a preset threshold are removed through random consistency sampling, and a least squares optimization function is constructed based on the feature point matching pairs to solve a stitching transformation matrix.

4. The method according to claim 1, characterized in that: The method analyzes the clarity, contrast and integrity of the image, evaluates the quality of the collected image data, adaptively adjusts the robot's motion trajectory and sampling strategy, focuses on observing the blind spot area and the occluded part, and obtains supplementary image data for the blind spot area and the occluded part, including: The Laplace operator is used to calculate the image edge clarity value, the image contrast value is calculated by the grayscale mean variance of the pixel block, and the image contour integrity value is calculated according to the continuity of the edge points. The clarity value, contrast value and contour integrity value are weighted to obtain the image quality score. Acquire a three-dimensional space grid through camera projection matrix mapping according to the image quality score, and mark the area where the image quality score is lower than a threshold to obtain a grid to be resampled; For the grid to be resampled, a three-dimensional grid map is constructed, the blind spot area is obtained by counting the sampling coverage times, and the occlusion area is obtained according to the intersection position of the ray and the depth map; For the blind spot area and the blocked area, the sampling probability distribution is calculated according to the grid density, and candidate sampling points are generated around the high probability area according to the probability weight, and an ordered sampling point sequence is obtained by sorting the distance from the sampling point to the robot; For the ordered sampling point sequence, the nearest unvisited sampling point is selected as the target position, a cubic spline curve is used to connect adjacent sampling points, and a trajectory that meets motion constraints is generated according to curve parameter control.

5. The method according to claim 1, characterized in that The method extracts significant features from multi-view images, establishes point correspondences between image pairs through feature matching, estimates camera posture and three-dimensional point coordinates, generates and optimizes a three-dimensional point cloud model through multi-view stereo vision, integrates multi-view images, and improves surface information of blind spots and occluded parts through the integration of multi-view images, including: Extracting feature points from the multi-view image according to the Harris corner response function, calculating the gradient direction histogram descriptor for the feature points, and matching the histogram descriptors using the nearest neighbor ratio threshold to obtain feature matching pairs; The five-point algorithm is used to estimate the essential matrix for the feature matching pair, and the camera rotation matrix and translation vector are obtained by solving the essential matrix through singular value decomposition; The three-dimensional coordinates of the feature points are calculated by triangulation according to the camera rotation matrix and the translation vector, and the three-dimensional coordinate points are obtained by screening through three indicators: reprojection error, parallax angle and tracking length; A hierarchical voxel grid is used to divide the space for the three-dimensional coordinate points, an initial pose is determined by centroid alignment, and point clouds under adjacent viewpoints are registered based on an iterative closest point algorithm to obtain a registered point cloud; A local neighborhood search range is established for the registered point cloud, a moving least squares method is used to fit the local surface based on the normal vector consistency constraint, and Poisson reconstruction is used to fill the missing area surface to obtain a continuous surface.

6. The method according to claim 1, characterized in that The method uses image stitching and fusion technology to analyze the overlapping relationship of adjacent view images, estimate the surface structure and texture information of the blind spot area, achieve seamless stitching and restoration of images, and obtain visually continuous and complete target reconstruction results, including: According to the projection of the vertex coordinates of the three-dimensional point cloud model onto the image plane, the overlapping area ratio between the perspective images is obtained. If the overlapping area ratio is greater than a preset threshold, an adjacent perspective image pair is obtained; For the adjacent view image pair, a multi-scale gradient histogram is used to calculate the texture features of the overlapping area, and a texture feature similarity score is calculated by a mutual correlation coefficient. If the similarity score is greater than a preset threshold, a pair of images to be matched is obtained; For the pair of images to be matched, a multi-layer image pyramid is used to perform block matching, a disparity field is obtained by a region growing algorithm, and the disparity field is smoothed based on total variation regularization to obtain a smoothed disparity field; According to the smoothed disparity field, a dense deformation field is generated by thin plate spline interpolation, and the dense deformation field is fused with the to-be-matched image pair by Laplace pyramid to obtain a spliced ​​image; For the blind spot area and the occluded area in the stitched image, an inpainting mask is constructed by adopting edge continuity constraint, and the texture features of the area around the mask are propagated by Poisson image editing to obtain a stitched image after inpainting.

7. The method according to claim 1, characterized in that In the image restoration process, a deep learning-based method is introduced to learn the local features and global structures of a large number of target images by training a convolutional neural network, perform semantic-level image restoration and content generation on blind spots and occluded parts, and comprehensively utilize prior knowledge and context information of the target to achieve image restoration, including: Obtain a binary mask with marked blind spots and occluded parts, and construct a five-layer encoder-decoder network based on the binary mask. The encoder extracts features through a 3×3 convolution kernel, and the decoder uses deconvolution upsampling to restore the resolution. Adding a residual block at the jump connection of the encoder-decoder network, wherein the residual block calculates feature weights through a channel attention mechanism to obtain feature channels whose significance is greater than a preset threshold; For the feature channel, a contrast loss function is used to calculate the cosine distance between the repaired area and the real image block. If the cosine distance is less than a preset distance threshold, a contrast loss is applied to the feature pair, and the overall loss is calculated by combining the reconstruction loss and the perceptual loss. According to the overall loss, the authenticity of the local area is evaluated by a two-layer discriminator, which uses a downsampled convolution block and a rectified linear unit with leakage coefficients to obtain pixel-level error and feature-level adversarial loss; In view of the pixel-level error and feature-level adversarial loss, a multi-head attention layer is constructed to calculate context features of different scales. The multi-head attention layer calculates a weight matrix through 1×1 convolution to obtain an attention feature map. A spatial attention mask is generated according to the attention feature map, and the spatial attention mask guides feature propagation based on edge saliency to obtain a final repaired image.

8. The method according to claim 1, characterized in that The method optimizes the surface mesh of the three-dimensional model according to the repaired target image and the three-dimensional model, simplifies and smoothes the mesh geometry, removes noise and artifacts, and adjusts the vertex positions to establish a mapping relationship between the image and the three-dimensional model, calculates the texture coordinates of each mesh surface in the three-dimensional model, maps the image data to the corresponding mesh surface, mixes the multi-view images, and obtains a complete three-dimensional reconstruction result of the target object, including: Acquire the coordinates and connection relationship of the 3D mesh vertices, and calculate the mean curvature value and normal vector of each vertex according to the 3D mesh vertex coordinates; For the area where the mean curvature value is less than a preset curvature threshold, local quadratic surface least squares fitting is used to obtain smooth mesh vertex coordinates; According to the vertex coordinates of the smoothed mesh, a simplified mesh is obtained by calculating the geometric error energy and the texture error energy of each edge and adopting an edge collapse algorithm based on a quadratic error metric; For the simplified grid, a parameterized mapping is constructed based on a discrete coordinate field, an area preservation constraint and an angle preservation constraint are imposed at a boundary fixed point with a maximum curvature value, and a parameterized grid is obtained through alternating iterative optimization; The parameterized grid is projected using multi-view images to calculate the projection overlap and observation angle of the grid facets; A mixed weight map is constructed according to the projection overlap and the observation angle, and the mixed weight map is solved by using Poisson's equation to obtain a seamless texture image.

Citation Information

Cited By

  • Robot inspection planning system for ocean platform jacket marine organism monitoring

    CN121095752A

  • Robot patrol planning system for offshore platform jacket marine biology monitoring

    CN121095752B