Image processing method for sampling at magnetic anomaly by underwater magnetic measuring robot
By oversampling and refining network processing of underwater magnetic survey robot images, high-resolution and new perspective images are generated, solving the problems of insufficient image resolution and generalization ability in marine exploration, and realizing efficient and accurate marine engineering operations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU DIANZI UNIV
- Filing Date
- 2023-02-22
- Publication Date
- 2026-04-10
AI Technical Summary
Existing underwater robots suffer from limited image resolution and insufficient generalization ability in ocean exploration, resulting in unvisualized exploration results and difficulties in real-time interaction, which affects the efficiency of marine engineering operations.
By combining an oversampling algorithm and a thinning network with a generalizable radial neural network, high-resolution and novel perspective images are generated through multi-view constraints and image thinning processing, and 3D models of marine targets are reconstructed.
It improves the utilization rate of image information and real-time interactive capabilities, provides efficient visualization tools, and enhances the accuracy and stability of marine engineering operations.
Smart Images

Figure CN116137053B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of ocean exploration and underwater robots, and particularly relates to a method for processing images of sampling at a magnetic anomaly by an underwater magnetic survey robot. BACKGROUND
[0002] With the development of ocean economy and the continuous progress of fluid mechanics, electromagnetism, new material science, automatic control theory and other disciplines, exploring the ocean has become an inevitable act for human beings to understand the ocean, utilize ocean resources and protect the ocean environment. However, due to the complexity of the ocean environment, the efficiency of exploring seabed mineral resources is low, and the laying and maintenance of seabed cables are difficult, which is not conducive to the development of the ocean industry. At present, underwater robots are used to undertake ocean engineering tasks, which to some extent solves the problem, but there are still deficiencies. The detection results are not visualized, and the single influencing factor makes the completion of the task in the complex seabed environment not high.
[0003] At present, computer vision is used as an auxiliary tool in the task of underwater robots exploring ocean resources. For example, in recent years, the image generation of NeRF, which is very popular in the field of artificial intelligence, can improve the detection efficiency and accuracy in ocean exploration tasks. However, as the detection depth increases, the visibility of the ocean environment decreases sharply. NeRF has two major defects, which make it difficult for this auxiliary tool to achieve its original auxiliary capabilities. First, the resolution of the generated image of NeRF depends on the resolution of the input image, which makes the image obtained under poor underwater light conditions not capable of assisting underwater robots in their work. Second, NeRF does not have generalization ability, and a set of images for each object needs to be retrained, which makes it impossible to interact in real time when applied to ocean operations. Therefore, improvements and enhancements are needed. SUMMARY
[0004] The purpose of the present application is to address the shortcomings of the prior art by providing a method for processing images of sampling at a magnetic anomaly by an underwater magnetic survey robot. This method uses the super-sampling algorithm in computer vision to reconstruct the 3D model of the image of the target object in the ocean using multiple image resources, provides multiple new perspective images, improves the utilization rate of image information and real-time interaction capability, and provides a powerful visualization tool for underwater magnetic survey tasks.
[0005] To solve the above technical problems, the technical solution of the present application is as follows:
[0006] A method for processing images of sampling at a magnetic anomaly by an underwater magnetic survey robot, comprising the following steps:
[0007] S1, collect images of the magnetic anomaly under water and construct a basic magnetic target image database;
[0008] S2, based on the super sampling strategy, the resolution of the image in the database is optimized, a group of images taken in the first step is used as input, a plurality of rays are established on each pixel in the image, thereby applying multi-view constraints at the sub-pixel level, and the image is improved from low resolution to high resolution by combining multi-view images;
[0009] S3, a refinement network is used to further optimize the generated high-resolution image, the high-resolution image generated after step S2 is refined patch by patch through the refinement strategy of the refinement network, and the image is further optimized by supplementing image details to generate a high-resolution image;
[0010] S4, a generalizable radiance neural network is used to produce a new view image by inputting the high-resolution image optimized by steps S2 and S3.
[0011] The present application has the following characteristics and beneficial effects:
[0012] The technology can effectively improve the available information and image information utilization rate of the photographed image in the computer vision ocean work. By shooting a plurality of rays on each image pixel of the input image, multi-view constraints are applied at the sub-pixel level, thereby realizing the improvement from low resolution to high resolution. The details of the super-sampled image are processed through the refinement strategy of the refinement network, the available information in the image is increased, and a data basis is provided for subsequent establishment of high-resolution 3D reconstruction and generation of new view images. By epipolar geometry, feature blocks are extracted along the epilolar line, linearly projected into 1D feature vectors, and a transformer is used to process data. The high-resolution image sampling blocks of a series of scenes are modeled in terms of new scene light color, and new view high-resolution images are generated.
[0013] By using the super sampling algorithm in computer vision, the low-resolution image obtained by sampling the seabed is optimized to a high-resolution image, which improves the ratio of available information in the image to a certain extent. Combined with the generalizable neural radiance field in the field of artificial intelligence, the algorithm model can be trained in advance, so that a plurality of image resources are used to reconstruct the image 3D model of the marine target object, provide a plurality of new view images, improve the utilization rate of image information and real-time interaction ability, and provide a powerful visualization tool for underwater magnetic measurement tasks.
[0014] Compared with the existing underwater magnetic measurement robot for detecting magnetic targets, the present application can more intuitively judge the high-resolution image of the target object, and can to a certain extent avoid the interference of unknown items in the ocean to hinder the progress of the engineering operation, thereby efficiently, high quality and stably completing the marine engineering operation. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description only constitute some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.
[0016] Figure 1 The method flow chart of the method for sampling and image processing of the underwater magnetic measuring robot for magnetic anomaly in the present application;
[0017] Figure 2 The schematic diagram of the oversampling strategy in the embodiment of the present application.
[0018] Figure 3 The schematic diagram of the encoding synthesis patch refinement network structure in the embodiment of the present application.
[0019] Figure 4 The effect comparison schematic diagram of the embodiment of the present application.
[0020] Figure 5 The structure schematic diagram of the model in the embodiment of the present application. DETAILED DESCRIPTION
[0021] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0022] In the description of the present application, it should be understood that the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second" and the like are only for the purpose of description and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features limited by "first", "second" and the like can explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0023] In the description of the present application, it should be noted that unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connecting" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0024] The embodiment provides a sampling image processing method for a magnetic anomaly of an underwater magnetic measuring robot, which considers that the resolution of an image shot by a lightweight camera under strong light loss in water is limited, so that the available information in the image is reduced, an oversampling strategy and a refinement network are used to apply multi-view constraints at a sub-pixel level, from low resolution to high resolution, so that more available information can be extracted from the image. The collected image data is input, the feature blocks are extracted by epipolar geometry, and the transformer is used to process the data to generate a new view image, which can improve the utilization rate of available information in the shot image. It can assist the efficient completion of the marine magnetic survey task, improve the accuracy of magnetic target detection, and the specific method is as follows:
[0025] Step 1: At the data acquisition level, the camera in the structure of the underwater magnetic measuring robot is used to shoot 100 images (each image has a different view angle and is uniformly distributed around the detected magnetic anomaly), as shown in Figure 1 , a basic magnetic target image database is established.
[0026] Step 2: Based on the oversampling strategy, the resolution of the images in the database is optimized. A group of images shot in step 1 are used as input, a plurality of rays are established on each pixel in the image, so that multi-view constraints are applied at a sub-pixel level, and the image is improved from low resolution to high resolution in combination with multi-view images.
[0027] As shown in Figure 2 , the oversampling strategy divides one pixel (solid line) into a plurality of sub-pixels (dashed line), and draws a ray for each sub-pixel. Therefore, compared with ordinary NeRF, more 3D points in the scene can be corresponded and constrained.
[0028] The color of the pixel p after oversampling can be represented by the following formula:
[0029] C(p)=C(r(p))=Comp(r′(p))(r′∈R(p))
[0030] where R(p) represents the set of all possible directions of the rays of pixel p in the training image, and Comp is the rasterization process of the composition of the radiance of all the incident rays contained in R(p). Ideally, the training ray directions are sampled from R(p), but this data is too large, leading to a sharp increase in the difficulty of network fitting, so in practical applications, a pixel is uniformly divided into an sxs grid pixel set S(p), and the training ray direction is selected from R'(p).
[0031] R'(p) = {r(j) | j ∈ S(p)} ∈ R(p)
[0032] Therefore, by rendering the sub-pixels, an sHxsW image can be obtained (input HxW image).
[0033] Third step: further optimize the generated high-resolution image using the refinement network. Through the patch-by-patch refinement strategy of the refinement network, the generated high-resolution image is refined patch by patch using the high-resolution image taken under ideal conditions, so as to further optimize the image and generate a high-resolution image. The refinement process is as shown in Figure 3
[0034] The refinement module encodes the generated image from the super-sampling and reference patch to synthesize a patch The encoded features of the super-sampling are connected with the encoded features of the patch , and are decoded to generate a refined patch.
[0035] After the first two steps, we can obviously see that the resolution of the picture has been significantly improved, achieving the reduction of the noise of the picture and increasing the proportion of usable information of the picture to a certain extent.
[0036] Fourth step: using the generalizable radiance neural network to generate a new view image with the image cluster generated in the second and third steps as input. The model (as shown in Figure 5 ) is composed of three modules, each stage having a different transformation neural network. Each transformation neural network follows the ViT architecture, which uses residual connections to interleave layer normalization (LN), self-attention (SA), and multi-layer perceptron (MLP). Each layer consists of LN→SA→LN→MLP.
[0037] 1) Visual feature transformation neural network, the input of this module is a set of patch linear embeddings and position encoding vectors
[0038] First, the features of the zero layer (input) are defined as:
[0039]
[0040] This module is repeated for each depth sample, so it operates on K sequences of views.
[0041]
[0042] 2) Epipolar aggregator transformation neural network, this module aggregates information along each epipolar line, generating a feature for each reference view. The input of this module is the set f1 = {f1 m |1≤m≤M} concatenated with the positional encoding.
[0043] The reference set f1 m is concatenated with the features corresponding to the view f1 k,m . Each view repeats the transformation neural network, so it operates along M sequences of epipolar samples. First, we compute
[0044]
[0045] where r 0 is a special token representing the target ray. Then, we apply a learned weighted sum along the M coregistered line samples, as follows:
[0046]
[0047]
[0048] resulting in a feature vector for each view k, where W1 are learnable weights, is the output corresponding to the target ray token.
[0049] 3) Reference view aggregator transformation neural network, the last transformation neural network aggregates the features on the reference views and predicts the color of the target ray.
[0050] Its input is the set of per-reference view features concatenated with the camera relative positional encoding. Formally, we compute
[0051]
[0052] Similarly to the previous module, we compute the mixing weights
[0053]
[0054] which are used in combination with the weights from the previous module to estimate the color of the target ray by mixing the colors along each coregistered line sample at each reference view,
[0055]
[0056] wherein is the pixel color of the mth sample along the epipolar line k. By using two sets of attention weights and β k to achieve, allowing the mixing of color from all epipolar line samples and all reference views.
[0057] The rendering network formed by the above three modules solves the generalization problem of new view image generation.
[0058] The embodiments of the present application are described in detail above with reference to the drawings, but the present application is not limited to the described embodiments. For those skilled in the art, various changes, modifications, replacements and variations of the embodiments including components are made without departing from the principles and spirits of the present application, and still fall within the protection scope of the present application.
Claims
1. A method for processing images sampled at magnetic anomalies using an underwater magnetic surveying robot, characterized in that, Includes the following steps: S1. Acquire images of underwater magnetic anomalies and construct a basic magnetic target image database; S2. Optimize the resolution of images in the database based on an oversampling strategy; The specific method of the oversampling strategy is as follows: Will In an image, a pixel is divided into multiple sub-pixels, and a ray is drawn for each sub-pixel, where H is the image height and W is the image width; The color of pixel p after oversampling can be represented by the following formula: ; in Let represent the set of all possible ray directions for pixel p in the training image. yes The grating process that comprises the radiation of all incident rays contained therein; the training ray direction from Medium sampling; dividing a pixel evenly into a... Grid pixel set The training ray direction is from Selected from, among which ; By rendering sub-pixels, a... image; S3. By refining the network patch by patch using the high-resolution image generated after optimization in step S2, further optimization of the image is achieved by supplementing image details, and a high-resolution image is generated. The specific method for step S3 is as follows: S3-1, Setting Reference Patch ; S3-2, Refining the network from oversampling and reference patches Encoded synthetic patch in the generated image ; S3-3, The oversampled encoded features, after max pooling, are combined with the patch... The encoded features are connected, and the encoding features are connected. Decode and generate detailed patches; S4. Use a generalizable radial neural network combined with the high-resolution image optimized in steps S2 and S3 as input to produce a new perspective image. The generalizable radial neural network includes a visual feature transformation neural network, an Epipolar aggregator transformation neural network, and a reference view aggregator transformation neural network. The specific method for step S4 is as follows: S4-1. Using a visual feature transformation neural network, the input is a set of patch linear embeddings and position encoding vectors consisting of view k and the m-th sampling depth index. , First, the feature connection of the zero-layer input is defined as: ; It repeats for each depth sample, therefore it operates on K view sequences. ; S4-2. The information is aggregated along each epipolar line by the transforming neural network using an Epipolar aggregator, thereby generating features for each reference view. The input is a set. Connected with the location code, Reference set Center and View For the corresponding features, the neural network is repeatedly transformed for each view, thus operating along the M epipolar sample sequence, first calculating... ; in This is a special label representing the target ray. Then, a learned weighted sum is applied along M epipolar samples, as shown below: ; ; We obtain the feature vector for each view k, where W1 is the learnable weight. This is the output corresponding to the target ray marker; S4-3, The input to the reference view aggregator transform neural network is a per-reference view feature connected to the camera's relative position encoding. The set, in form, is computed ; Calculate the mixed weights ; It is used in conjunction with weights from the Epipolar aggregator transform neural network to estimate the color of the target ray by mixing colors along each epipolar sample at each reference view. ; in It is the pixel color of the m-th sample along the epipolar line k, obtained by using two sets of attention weights. and This allows for the mixing of colors from all epipolar samples and all reference views.
2. The underwater magnetic surveying robot's image processing method for sampling locations of magnetic anomalies according to claim 1, characterized in that, In step S1, when acquiring images, the camera in the underwater magnetic surveying robot structure takes 100 images around the detected magnetic anomaly. Each image has a different perspective and is evenly distributed at the magnetic anomaly.
Citation Information
Patent Citations
Magnetic resonance image super-resolution reconstruction method and device
CN114494014A
Method for displaying object on three-dimensional model
US20200286290A1