Polar region ship safe navigation method and device

By enhancing and fusing the visible light and infrared images of polar ships, combining feature extraction and closed-loop detection, high-precision two-dimensional maps are generated, which solves the problem of polar ships being unable to accurately identify sea ice and ensures safe navigation.

CN120635404APending Publication Date: 2025-09-12WUHAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510643925.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In existing technologies, due to the geographical location and environmental influences of the polar regions, it is impossible to accurately identify and extract sea ice in images taken during polar ship navigation, resulting in untimely map updates and safety risks.

Method used

By acquiring ship and drone video data, enhancing and fusing visible light and infrared images, extracting feature points, performing closed-loop detection and global optimization, generating a two-dimensional map, and using a semantic segmentation model to extract navigable waters and control ship navigation.

Benefits of technology

It achieves accurate positioning and detection of sea ice targets under polar low-light conditions, improves the real-time perception accuracy of sea ice distribution, and ensures safe navigation of ships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635404A_ABST
    Figure CN120635404A_ABST
Patent Text Reader

Abstract

The invention relates to a polar region ship safe navigation method and device, and the method comprises the steps: obtaining the current driving data of a ship during driving in a polar region and the video data of an unmanned plane, carrying out the enhancement processing and fusion of a visible light image and an infrared image of each frame of image in the video data, and obtaining a fusion image corresponding to each frame of image; the advantages of infrared and visible light images are combined, thermal radiation and texture information are fused into one image, and detection and positioning of a sea ice target under the polar region low-illumination condition are facilitated; performing feature extraction on the fused image, performing map point search on the plurality of key frames to obtain a plurality of map point positions, and performing closed-loop detection, global optimization and two-dimensional map splicing on the plurality of key frames and the plurality of map point positions to generate a high-precision polar region two-dimensional orthographic spliced image; and sea ice extraction is performed on the two-dimensional map to obtain a navigable water area, so that the accuracy of details and the precision of sea ice extraction are improved, and ship navigation is controlled according to the navigable water area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of polar navigation safety technology, and in particular to a method and device for polar ship navigation safety. Background Art

[0002] As Arctic sea ice melts at an accelerated pace, the commercial value of Arctic shipping routes has increased significantly. One of the biggest challenges facing ships is how to avoid becoming "ice-bound." "Ice-bound" refers to the phenomenon in which a ship becomes trapped by ice during navigation, preventing it from continuing. This phenomenon usually occurs in polar regions or cold waters. Due to the thickness, density, and distribution of the ice, a ship may become surrounded by ice and lose its maneuverability. Ice-bound not only prevents a ship from sailing, but can also cause damage to the hull, increase fuel consumption, and even threaten the safety of the crew. In extreme cases, a ship may need to wait in the ice for rescue, which is not only time-consuming and labor-intensive, but can also have a negative impact on the environment.

[0003] Drones and advanced image processing algorithms offer solutions for polar navigation. Drones can extend a vessel's sensory range, capturing high-resolution, real-time aerial images, helping ships more accurately identify ice conditions. These capabilities enable ships to plan routes more effectively, avoid hazardous areas, and reduce the risk of ice entrapment. However, the polar region's geographic location and environment significantly impact image quality, making it difficult to accurately extract sea ice from maps, hindering map updates and posing safety risks to polar vessels.

[0004] Therefore, there is an urgent need to propose a method and device for safe navigation of polar ships to solve the technical problem that due to the geographical location and environment of the polar regions, the existing technology cannot accurately identify and extract sea ice in the picture, and thus cannot update the map, resulting in safety hazards for polar ships during navigation. Summary of the Invention

[0005] In view of this, it is necessary to provide a method and device for safe navigation of polar ships to solve the technical problem in the existing technology that due to the geographical location and environment of the polar regions, the sea ice in the pictures cannot be accurately identified and extracted, and the maps cannot be updated, resulting in safety hazards for polar ships during navigation.

[0006] In order to solve the above problems, in a first aspect, the present invention provides a method for safe navigation of polar ships, comprising: Obtain current navigation data of ships and video data from drones when sailing in polar regions; Performing enhancement processing and fusing the visible light image and the infrared image of each frame of the video data to obtain a fused image corresponding to each frame of the image; Performing feature extraction on the fused image to obtain a plurality of key frames, and performing map point search based on the plurality of key frames to obtain a plurality of map point positions; Performing closed-loop detection and global optimization on the multiple key frames to obtain multiple optimized key frames, and performing two-dimensional map stitching on map point positions corresponding to the multiple optimized key frames to obtain a two-dimensional map; Sea ice is extracted from the two-dimensional map to obtain navigable waters; and the ship is controlled to navigate in the polar region based on the navigable waters.

[0007] In a possible implementation, the enhancing and fusing the visible light image and the infrared image of each frame of the video data to obtain a fused image corresponding to each frame of the image includes: Performing illumination enhancement processing on the visible light image of each frame of the video data to obtain a first partial image; the first partial image is an image with prominent appearance detail information and obvious contrast; Brightness enhancement is performed on the infrared image of each frame of the video data to obtain a second partial image; the second partial image is an image with enhanced contrast and uniform overall brightness distribution; The first partial image and the second partial image of each frame of image are fused and reconstructed to obtain a corresponding fused image.

[0008] In a possible implementation, performing illumination enhancement processing on the visible light image of each frame of the video data to obtain the first partial image includes: Converting the values ​​of the red, green, and blue channels of the visible light image of each frame of the video data to obtain a brightness component; Calculating the visible light image according to an image enhancement algorithm to obtain reflection components of different scales; Calculating the probability of association between the brightness component and each reflection component to obtain an association weight for each reflection component; Obtaining an enhanced brightness image according to the brightness component, each reflection component and the associated weight; The color domain of the enhanced brightness image is reconstructed and merged to obtain a first partial image.

[0009] In a possible implementation, the step of performing brightness enhancement on the infrared image of each frame of the video data to obtain the second partial image includes: performing illumination enhancement processing on the infrared image of each frame of the video data based on a logarithmic scaling function to obtain an illumination enhanced image; Performing non-complex exponential function processing on the infrared image to obtain a non-complex exponential image; combining the illumination-enhanced image and the non-complex-exponential image to obtain a combined image; enhancing the overall brightness of the combined image to obtain a brightness-enhanced image; Normalization is performed on the brightness enhanced image to obtain a second partial image.

[0010] In a possible implementation, extracting features from the fused image to obtain a plurality of key frames, and searching for map points based on the plurality of key frames to obtain a plurality of map point locations, includes: Initializing a map based on the images of the first two frames in the fused image to obtain three-dimensional map points; Extract features from each frame of the fused image one by one according to a feature extraction algorithm to obtain a plurality of two-dimensional feature points; A plurality of key frames are obtained according to the three-dimensional map points and the plurality of two-dimensional feature points, and map point searches are performed on the plurality of key frames to obtain a plurality of map point positions.

[0011] In one possible implementation, the multiple two-dimensional feature points include two-dimensional feature points of a current frame image; obtaining multiple key frames based on the three-dimensional map points and the multiple two-dimensional feature points, and performing map point search on the multiple key frames to obtain multiple map point positions includes: Set up the initial map; Obtaining a current frame pose according to the three-dimensional map points and the two-dimensional feature points of the current frame image; When the current frame pose meets a preset condition, determining the current frame image as a key frame, and adding the key frame to the initial map to obtain an updated target map; Searching according to adjacent key frames in the target map to obtain matching points; According to the matching point, the current map point position is obtained; When the processing of the plurality of two-dimensional feature points is completed, a plurality of key frames and a map point position corresponding to each key frame are obtained.

[0012] In a possible implementation, performing closed-loop detection and global optimization on the multiple key frames to obtain multiple optimized key frames includes: Determining multiple candidate frames based on a positional relationship between the current frame and the multiple key frames; Performing closed-loop detection on similar transformations between the multiple candidate frames and the current frame using a random sampling consensus algorithm to obtain closed-loop key frames; Correcting the multiple key frames according to the closed-loop key frames to obtain multiple corrected key frames; Obtaining a correspondence between a global position and a camera position based on a difference between a GPS constraint and a timestamp in the video data; Performing coordinate system conversion on the multiple corrected key frames according to the corresponding relationship and similarity transformation to obtain multiple coordinate system key frames; After optimizing the difference, the residuals of the multiple coordinate system key frames are minimized and roughly estimated and optimized according to the map point projection error minimization principle to obtain an optimization result; the optimization result includes the optimal multiple optimized key frames.

[0013] In a possible implementation, performing two-dimensional map stitching on the map point positions corresponding to the multiple optimized keyframes to obtain a two-dimensional map includes: Determining the edges of a rectangle based on the relative posture of the map point positions; When the edge of the rectangle exceeds the current area, the stitching area is expanded and the weighted image and homography transformation are calculated; Correcting the weighted image according to the homography transformation and the preset corrected color image to obtain a corrected image, and establishing a weight pyramid; An optimal weight is determined according to the weight pyramid, and the corrected image is fused onto the rectangle of the stitching area according to the optimal weight to obtain a two-dimensional map.

[0014] In a possible implementation, extracting sea ice from the two-dimensional map to obtain navigable waters includes: Inputting the two-dimensional map into a preset semantic segmentation model to extract sea ice and obtain extracted features; the preset semantic segmentation model includes a lightweight deep convolutional neural network, a pyramid pooling module, and a coordinate attention mechanism module; Navigable waters are obtained according to the two-dimensional map and the extracted features.

[0015] In a second aspect, the present invention further provides a polar ship safety navigation device, comprising: A data acquisition module is used to obtain the current driving data of the ship and the video data of the drone when it is traveling in the polar region; An image processing module is used to enhance and fuse the visible light image and infrared image of each frame of the video data to obtain a fused image corresponding to each frame of the image; a feature extraction module, configured to extract features from the fused image to obtain a plurality of key frames, and perform map point searches based on the plurality of key frames to obtain a plurality of map point positions; A coordinate conversion module is used to perform closed-loop detection and global optimization on the multiple key frames to obtain multiple optimized key frames, and to perform two-dimensional map stitching on the map point positions corresponding to the multiple optimized key frames to obtain a two-dimensional map; The image stitching module is used to extract sea ice from the two-dimensional map to obtain navigable waters; and control the ship to navigate in the polar region based on the navigable waters.

[0016] The beneficial effects of the present invention are as follows: the current driving data of a ship and the video data of a drone when it is traveling in a polar region are obtained, the visible light image and the infrared image of each frame of the video data are enhanced and fused, and a fused image corresponding to each frame of the image is obtained; the advantages of infrared and visible light images are combined to fuse thermal radiation and texture information into one image, which is beneficial to the detection and positioning of sea ice targets under low illumination conditions in polar regions; feature extraction is performed on the fused image to obtain multiple key frames, and map point searches are performed on the multiple key frames to obtain multiple map point positions, which can realize real-time updating of map information and multiple Closed-loop detection and global optimization are performed on key frames and multiple map point positions to obtain multiple optimized key frames, and two-dimensional map stitching is performed on the map point positions corresponding to the multiple optimized key frames to obtain a two-dimensional map. Therefore, the key frames can be optimized through closed-loop detection and global optimization during the stitching process, and the stitching results can be adjusted to obtain a high-precision polar two-dimensional orthophoto stitching image; then sea ice can be extracted from the two-dimensional map to obtain navigable waters, which improves the accuracy of details and the accuracy of sea ice extraction, so as to perceive the distribution of sea ice in real time and control the navigation of ships according to the navigable waters. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A schematic flow chart of an embodiment of the method for safe navigation of polar ships provided by the present invention; Figure 2 A schematic structural diagram of an embodiment of a ship safety navigation system provided by the present invention; Figure 3 For the present invention Figure 1 A schematic flow chart of an embodiment of step S102; Figure 4 For the present invention Figure 1 A schematic flow chart of an embodiment of step S103; Figure 5 For the present invention Figure 1 A schematic flow chart of an embodiment of step S104; Figure 6 A schematic diagram of an embodiment of the sea ice extraction technology provided by the present invention; Figure 7 A schematic diagram of the structure of an embodiment of the preset semantic segmentation model provided by the present invention; Figure 8 This is a schematic structural diagram of an embodiment of the polar ship safety navigation device provided by the present invention. DETAILED DESCRIPTION

[0018] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, and are not used to limit the scope of the present invention.

[0019] like Figure 1 As shown, a specific embodiment of the present invention discloses a method for safe navigation of polar ships, comprising: S101. Acquire current driving data of a ship and video data of a drone when the ship is traveling in a polar region.

[0020] The ground vessel safe navigation method provided in the embodiment of the present application can be applied to the ground vessel safe navigation system, wherein the ground vessel safe navigation can be based on a software system running on a terminal device, and the terminal device can be a server, a tablet computer, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a mobile phone and other terminal devices. The embodiment of the present application does not impose any restrictions on the specific type of the terminal device.

[0021] When a ship is sailing in polar regions, it can obtain current driving data and also capture video data using drones. Current driving data includes but is not limited to the current driving position, basic ship data, and environmental data.

[0022] S102 : performing enhancement processing and fusing the visible light image and the infrared image of each frame of the video data to obtain a fused image corresponding to each frame of the image.

[0023] Among them, the visible light image and infrared image of each frame in the video data can be enhanced, and then the enhanced images can be fused to obtain the fused image corresponding to each frame. This can combine the advantages of infrared and visible light images, and fuse thermal radiation and texture information into one image, which is beneficial to the detection and positioning of sea ice targets under polar low-light conditions.

[0024] S103 , extracting features from the fused image to obtain multiple key frames, and searching for map points based on the multiple key frames to obtain multiple map point positions.

[0025] To reduce the computational burden of the tracking thread, a fast feature extraction and description algorithm (ORB) detector and descriptor are used to extract features from the fused image, resulting in multiple keyframes. Feature extraction is then incorporated into the preprocessing phase to complete video frame preparation and tracking. A bag-of-words model is introduced to search for matching points across multiple keyframes and identify new map points. Bundle Adjustment (BA) is then used to refine the keyframe poses and map point positions of multiple keyframes within the local map, improving the overall tracking and mapping accuracy of the system.

[0026] S104 , performing closed-loop detection and global optimization on the multiple key frames to obtain multiple optimized key frames, and performing two-dimensional map stitching on the map point positions corresponding to the multiple optimized key frames to obtain a two-dimensional map.

[0027] Among them, by detecting and calculating multiple key frames, and then performing closed-loop detection and global optimization, multiple optimized key frames are obtained, and then the coordinates of the multiple optimized key frames are transformed into the WGS-84 coordinate system through similarity transformation. At the same time, the similarity transformation is optimized, and then multiple map points are sampled. A suitable ground plane is fitted using RANSAC. An incremental image stitching method can be introduced to stitch the two-dimensional map to obtain a two-dimensional map.

[0028] S105: extracting sea ice from the two-dimensional map to obtain navigable waters; and controlling the ship to navigate in the polar region based on the navigable waters.

[0029] Among them, a semantic segmentation model can be set, and the semantic segmentation model can be a DeepLabV3+ model. By using the DeepLabV3+ model to extract sea ice from a two-dimensional map, navigable waters can be obtained, and then the navigation of ships in the polar region can be controlled based on the navigable waters.

[0030] In a specific embodiment of the present invention, the polar ship safety navigation system can be composed of three parts: a monitoring end, a service end, and a client end. Figure 2As shown in the figure. On the monitoring side, the drone completes the tasks of video acquisition and parameter transmission. The server consists of three parts: a drone visual perception enhancement module, a real-time video stream image stitching module, and a polar multi-navigable waters segmentation module. The drone visual perception enhancement module uses the integrated image restoration network (AirNet) to enhance visible light images using the AMSR algorithm, processes infrared images using the B algorithm, and fuses the two images using the GTF algorithm to improve video clarity in real time. The real-time video stream image stitching module uses a real-time positioning and map construction method to generate high-precision two-dimensional orthophoto stitched images of the polar regions through video frame extraction, feature extraction, feature matching, transformation matrix calculation, and image fusion. The polar multi-navigable waters segmentation module uses an improved DeepLabV3+ model to segment the stitched images, accurately extracting the outlines of sea ice, ships, and other features. On the client side, the system can add the processed images as a layer to the electronic nautical chart and designate the segmented sea ice areas as danger zones, allowing ship operators to intuitively perceive the sea ice distribution and choose safe navigation routes.

[0031] Compared with the existing technology, the present embodiment obtains the current driving data of the ship and the video data of the drone when it is sailing in the polar region, enhances and fuses the visible light image and infrared image of each frame in the video data to obtain a fused image corresponding to each frame; combines the advantages of infrared and visible light images, fuses thermal radiation and texture information into one image, which is conducive to the detection and positioning of sea ice targets under low illumination conditions in the polar region; extracts features from the fused image to obtain multiple key frames, and performs map point search on the multiple key frames to obtain multiple map point locations, which can realize real-time updating of map information. Closed-loop detection and global optimization are performed on multiple key frames and multiple map point positions to obtain multiple optimized key frames, and two-dimensional map stitching is performed on the map point positions corresponding to the multiple optimized key frames to obtain a two-dimensional map. Therefore, the key frames can be optimized through closed-loop detection and global optimization during the stitching process, and the stitching results can be adjusted to obtain a high-precision polar two-dimensional orthophoto stitched image. Then, sea ice can be extracted from the two-dimensional map to obtain navigable waters, which improves the accuracy of details and the precision of sea ice extraction, so as to perceive the distribution of sea ice in real time and control the navigation of ships according to the navigable waters.

[0032] In some embodiments of the present invention, Figure 3 As shown, step S102 includes: S301 , performing illumination enhancement processing on the visible light image of each frame of the video data to obtain a first portion of images; the first portion of images is an image with prominent appearance detail information and obvious contrast.

[0033] Among them, the adaptive multi-scale Retinex algorithm is used to perform illumination enhancement processing on the low-illuminance visible light image of each frame in the video data to obtain an image with prominent appearance detail information and obvious contrast, that is, the first part of the image is obtained.

[0034] S302, performing brightness enhancement on the infrared image of each frame in the video data to obtain a second partial image; the second partial image is an image with enhanced contrast and uniform overall brightness distribution.

[0035] Among them, the infrared image of each frame in the video data is processed by a brightness enhancement algorithm to obtain an image with enhanced contrast and uniform overall brightness distribution, which is called the second part of the image.

[0036] S303 : Fusing and reconstructing the first partial image and the second partial image of each frame image to obtain a corresponding fused image.

[0037] The first and second parts of each frame image are used as input, and the two parts of the image are fused and reconstructed through the gradient transfer fusion algorithm to obtain the final fused image of each frame image.

[0038] In some embodiments of the present invention, step S301 includes: The values ​​of the red, green and blue channels of the visible light image of each frame in the video data are converted to obtain the brightness component.

[0039] Among them, an adaptive multi-scale Retinex algorithm is set up, which can include the AMSR algorithm, which can well balance the problem of image color reproduction and over-enhancement. The AMSR algorithm attempts to combine the output images of different SSR algorithms together so that the weights associated with each SSR scale are adapted to the calculation based on the content of the input image. First, the visible light image of each frame in the input video data is mathematically converted to the values ​​of the red, green and blue channels of the visible light image to obtain the brightness component. , as shown in formula (1): (1) Where, 、 and Respectively expressed in The red, green, and blue pixel values ​​at position.

[0040] The visible light image is calculated according to the image enhancement algorithm to obtain reflection components of different scales.

[0041] Among them, the visible light image can be processed by the image enhancement algorithm (SSR algorithm), and the SSR output corresponding to the s-th scale can be calculated ,Then the linear stretching method is used to normalize each SSR output to the full display range, as shown in formula (2): (2) Where, is a probability parameter, which means that the output value of SSR has a 99% probability of being less than or equal to the output value. Same as above. Thus, reflection components of different scales can be obtained, that is, the output image of each SSR .

[0042] The association probability between the brightness component and each reflection component is calculated to obtain the association weight of each reflection component.

[0043] Among them, for the pixel value , the probability value associated with each SSR scale is calculated as shown in formula (3): (3) Where, Indicates SSR1 corresponding to small-scale Retinex; represents SSR2 corresponding to the medium-scale Retinex; Indicates SSR3 corresponding to large-scale Retinex.

[0044] Therefore, the brightness component can be calculated according to formula (3) to obtain the probability. In order to obtain a clear and natural image, it is necessary to combine the input image with the output image of each SSR to obtain the AMSR output image. In order to calculate the weight value related to the input image , the initial stage will set p 0 = 1. According to the above probability value, the association weight of the input image and each SSR output image is Calculated by formula (4): (4) According to the brightness component, each reflection component and the associated weight, an enhanced brightness image is obtained.

[0045] Among them, the output image of the SSR algorithm is enhanced by the image enhancement algorithm (AMSR algorithm). , brightness component and associated weights After combining, the AMSR enhanced image is generated, as shown in formula (5): (5) The color domain of the enhanced brightness image is reconstructed and merged to obtain the first partial image.

[0046] Among them, after obtaining the enhanced brightness image, the corresponding hue shift and color desaturation are prevented, the color gamut of R, G, and B is reconstructed, and finally the color components Rr, Gr, and Br are merged to obtain a color image.

[0047] In some embodiments of the present invention, step S302 includes: The infrared image of each frame in the video data is subjected to illumination enhancement processing based on a logarithmic scaling function to obtain an illumination enhanced image.

[0048] Among them, the IB algorithm is used to perform illumination enhancement processing on the infrared image to obtain the illumination enhanced image , the theoretical algorithm used is shown in formula (6): (6) Where, Represented as the original input image; is the image processed by the logarithmic scaling function; Represented as a product operation.

[0049] The infrared image is processed by a non-complex exponential function to obtain a non-complex exponential image.

[0050] The input original image is processed by a non-complex exponential function to enhance the local contrast and reduce the high-intensity pixels of the original image to obtain the output image of the non-complex exponential function. .

[0051] The illumination enhanced image and the non-complex exponential image are combined to obtain a combined image.

[0052] The improved LIP model is used to enhance the illumination image With the output image Combine to get the combined image ,Although Have and feature information.

[0053] The overall brightness of the combined image is enhanced to obtain a brightness enhanced image.

[0054] Among them, the image The overall brightness is low, so the image needs to be The overall brightness is improved, and the improved cumulative distribution function of hyperbolic secant distribution (CDF-HSD) is used to enhance the overall brightness to obtain the output image of the CDF-HSD function .

[0055] The brightness enhanced image is normalized to obtain the second image.

[0056] Among them, the image Normalization is performed to ensure that the image can be linearly scaled within a suitable range, and finally the second part of the image is obtained. .

[0057] Furthermore, step S303 includes: assuming that the visible light image is , the infrared image is , the fused image is , the scale is m × n According to the gradient transfer fusion model, it can be expressed as shown in formula (7): (7) Where, is the data fidelity term constraining the image With infrared images have similar pixel intensities; is the regularization term, that is, to fuse the images With visible light images have similar gradients. More specifically, have similar edges at corresponding locations, It is a regularization parameter that controls the trade-off between the two terms and is used to adjust the smoothness of the image. By setting different parameters, the fused image with the required information enhancement can be obtained. The smaller the value of , the more thermal infrared image information the fused image retains. The larger the value of , the more detailed information the fused image can obtain, that is, the more visible light image information is retained. In other words, when = 0, the fusion result is the input infrared image; when = +∞, the fusion result is a visible light image. =4, the fused image has a good visual effect in most cases. Therefore, the embodiment of the present invention sets =4, the image obtained using objective function fusion tends to be visually morphological, and has the characteristics of highlighting the target information in the image and retaining the edge detail information of the image.

[0058] For the norm in formula (7), in order to obtain the target information in the highlighted image, as much thermal radiation information of the infrared image as possible is retained, so most of the data fidelity terms should be zero. However, in order to achieve the gradient transfer of the visible light image, a small part of the fidelity term may be large. This results in a fused image. F With infrared images IThe difference between is the Laplace function or impulse function, not the Gaussian function, that is, p = 1. In addition, since natural images are usually piecewise smooth, gradients tend to be sparse, and large gradients correspond to edges, the norm is used to minimize the gradient difference, that is, q =1, and let y = x-v , and optimize as shown in formula (8): (8) Where: Indicates that the pixel is i Image gradient at ; and are respectively expressed as the linear operators corresponding to the horizontal and vertical first-order differences, that is, and , and Represents image pixels i The nearest neighbors to the right and below. If the pixel i is in the last row or column, then and Both i , that is, the gradient is 0. This formula (8) is a convex function with the characteristics of a global optimal solution. Its first term is regarded as a data fidelity term, and the second term is regarded as a regularization term. It can achieve an appropriate balance between retaining thermal radiation information and detailed appearance texture information. This formula is a standard The minimization problem can be solved. For the regularization parameter , it only needs to be done by utilizing Minimization technology optimizes the above formula to solve , and then determine the final fused image .

[0059] The embodiment of the present invention is affected by factors such as polar fog and light sources, and the visual perception ability of the imaging system used to obtain video images is weak. The images collected in low-light environments have problems such as low contrast, obvious noise, and invisible details in dark areas, which will seriously affect the effective application of images in high-level visual processing tasks. In order to improve the effective application of images in high-level visual processing tasks, the embodiment of the present invention proposes a method for fusion of infrared and visible light images based on illumination enhancement. The algorithm enhances infrared images and low-light visible light images, and then fuses and reconstructs the enhanced images. In polar environments, single-modal images are often difficult to accurately reflect the environment. The combination of the two can capture information more comprehensively. Infrared images can detect heat sources in dark environments, while visible light images provide the necessary color and detail information. The complementarity of this information enables drones to provide richer and more accurate environmental information, improving navigation decision-making capabilities.

[0060] In some embodiments of the present invention, Figure 4 As shown, step S103 includes: S401 , performing map initialization based on the images of the first two frames in the fused image to obtain three-dimensional map points.

[0061] Feature extraction is incorporated into the preprocessing phase. To reduce the computational burden of the tracking thread, an Oriented Fast and Rotated Brief (ORB) detector and descriptor are used. Map initialization is performed using the first two frames to obtain 3D map points.

[0062] S402 : performing feature extraction on each frame of the fused image one by one according to a feature extraction algorithm to obtain a plurality of two-dimensional feature points.

[0063] The feature extraction algorithm (ORB (Oriented FAST and Rotated BRIEF)) is used to extract features from each frame of the fused image. This algorithm detects and extracts significant corners or texture points (such as the edge of an ice cube) on the current frame, obtaining two-dimensional feature points for each frame. This results in multiple two-dimensional feature points.

[0064] S403: Obtain multiple key frames based on the three-dimensional map point and multiple two-dimensional feature points, and perform map point search on the multiple key frames to obtain multiple map point positions.

[0065] After initialization is completed, the pose of the current frame is solved using several pairs of three-dimensional map points and multiple two-dimensional feature points to obtain multiple key frames, and map point searches are performed on the multiple key frames to obtain multiple map point positions.

[0066] In some embodiments of the present invention, the plurality of two-dimensional feature points include two-dimensional feature points of the current frame image; step S403 includes: Set up the initial map; According to the 3D map points and the 2D feature points of the current frame image, the current frame pose is obtained; When the current frame pose meets the preset conditions, the current frame image is determined to be a key frame, and the key frame is added to the initial map to obtain the updated target map; Search the adjacent key frames in the target map to obtain matching points; According to the matching points, get the current map point position; When the processing of multiple two-dimensional feature points is completed, multiple key frames and the map point position corresponding to each key frame are obtained.

[0067] In a specific embodiment of the present invention, an initial map can be set based on actual conditions. When each frame in the fused image is processed one by one, the current frame is considered the current frame. After initialization, the pose of the current frame is solved using several pairs of 3D map points and 2D feature points. A strategy similar to ORB-SLAM is used to solve the matching relationship between 3D map points and 2D feature points. Based on a given 3D map point, the 2D feature point extracted from the current frame is determined to match the existing map point, i.e., the current keypoint. The camera pose is estimated using a residual function. The distance between the current frame and the previous keyframe is measured using a weighted combination of distance and rotation, compatible with monocular and binocular cameras (Large-Scale Direct-SLAM, LSD-SLAM). The weight of the translation parameter is determined based on the average inversion depth. A preset condition is that if the distance exceeds a threshold, the current frame becomes a new keyframe. This threshold ensures sufficient overlap between keyframes. After the keyframe is added, it is quickly incorporated into the initial map, resulting in an updated target map. This ensures sufficient matching relationships between subsequent frames and ensures successful tracking. Because not all feature points are observed during tracking, corresponding matching points must be found on the epipolar lines of adjacent keyframes. New map points are then drawn based on these matching points. Matching points are the two-dimensional observation data used to calculate (generate) new map points, while new map points are the three-dimensional counterparts of these matching points. This step utilizes the bag-of-words model to accelerate the search. After adding new map points, a data association step is performed to add more observations, including merging points and removing bad points. Bundle Adjustment (BA) is used to refine the poses of multiple keyframes and the positions of multiple map points in the local map.

[0068] In some embodiments of the present invention, Figure 5 As shown, step S104 includes: S501: Determine multiple candidate frames based on the positional relationship between the current frame and multiple key frames.

[0069] To achieve better consistency in large-scale scenes, loop closure detection and global optimization are still required despite the use of GPS constraints. The keyframe closest to the current frame's position is detected. All keyframes that share some visual words with the current frame are evaluated for similarity. A keyframe is considered a candidate only if its adjacent keyframes are also considered candidates, enhancing robustness.

[0070] S502: Perform closed-loop detection on similar transformations between multiple candidate frames and the current frame using a random sampling consensus algorithm to obtain closed-loop key frames.

[0071] The Random Sample Consensus (RANSAC) algorithm is used to calculate the similarity transformation between all candidate frames and the current frame. The frame with the best inlier fit is the closed-loop keyframe.

[0072] S503: Correct the multiple key frames according to the closed-loop key frame to obtain multiple corrected key frames.

[0073] Among them, the closed-loop key frame is to correct the accumulated errors caused by the long-term operation of the SLAM system and improve the global Figure 1 The key to consistency and accuracy. Eliminating cumulative error: Over long periods of time, the SLAM system's pose estimate will inevitably drift, causing trajectory and map deformation. By detecting loop closure (i.e., the current frame observes the same scene as a historical closed-loop keyframe), the system establishes a strong constraint: the estimated pose of the current frame should be very close to the pose of the closed-loop keyframe after transformation. Triggering global optimization: After identifying a closed loop, the system uses this closed-loop constraint to perform global pose graph optimization (Pose Graph Optimization), or global BA. This adjusts the poses and 3D positions of all keyframes along the entire trajectory to enforce the closed-loop constraint, significantly reducing cumulative error and improving the consistency and accuracy of the entire map and trajectory. Map fusion and correction: Within closed-loop areas, duplicate map points may exist. After closed-loop optimization, the system can fuse these duplicate map points to further improve map quality and density. Therefore, multiple keyframes can be corrected using closed-loop keyframes, resulting in multiple corrected keyframes.

[0074] S504 : Obtain the correspondence between the global position and the camera position according to the difference between the GPS constraint and the timestamp in the video data.

[0075] Among them, GPS constraints can be set and the difference of timestamps in video data can be determined. The difference between GPS constraints and video timestamps can be used to establish the correspondence between the global position and the camera position.

[0076] S505 : performing coordinate system transformation on the multiple corrected key frames according to the corresponding relationship and similarity transformation to obtain multiple coordinate system key frames.

[0077] Among them, multiple correction key frames are converted into the WGS-84 coordinate system through correspondence and similarity transformation to obtain multiple coordinate system key frames.

[0078] S506. After optimizing the difference, the residuals of multiple coordinate system key frames are minimized and roughly estimated and optimized according to the principle of minimizing the map point projection error to obtain an optimization result; the optimization result includes multiple optimal optimized key frames.

[0079] The timestamp difference and similarity transformation are optimized. Global optimization (i.e., relative pose optimization) of multiple map point positions is performed based on the principle of minimizing map point projection error and closed-loop keyframes. The principle of minimizing map point projection error minimizes the reprojection error by weighting the translation parameters, thereby minimizing the residual error for all keyframes and performing a rough estimate optimization. This narrows the range of time difference variation and yields an initial transformation. Generally, the timestamp difference is less than 60 seconds, so the time difference can be divided into several equal intervals within this range. A similarity transformation is calculated for each interval, and the optimal multiple optimized keyframes are selected. Once the optimization result is accurate, GPS constraints are implemented to mitigate tracking drift. The trajectory and map estimated by pure visual SLAM typically have only relative scale (uncertain size) and no geographic coordinates. GPS provides absolute position information in meters and geographic coordinates (latitude and longitude). GPS constraints can "scale" the SLAM map to the real-world scale and "anchor" it to the geographic coordinate system. Visual SLAM drift can be significant, especially in scenarios lacking texture, where feature matching is difficult, or when running for long periods of time. GPS can provide a globally consistent position reference. Even though its accuracy is not as high as the local accuracy of visual SLAM, it can effectively limit the accumulation of long-term drift, thereby obtaining more accurate multiple optimized keyframes.

[0080] In some embodiments of the present invention, step S104 includes: Determine the rectangle edges based on the relative poses of the map point positions; When the edge of the rectangle exceeds the current area, the stitching area is expanded and the weighted image and homography transformation are calculated; Correcting the weighted image according to the homography transformation and the preset correction color image to obtain a corrected image and establish a weight pyramid; According to the weight pyramid, the optimal weight is determined, and the corrected image is fused to the rectangle of the stitching area according to the optimal weight to obtain a two-dimensional image.

[0081] In a specific embodiment of the present invention, through the previous steps, the key frames are accurately aligned together and a stitching plane is fitted. This section describes the incremental stitching of key frames. Because the traditional graph cutting algorithm is very time-consuming, the orthophoto stitching task in the system is divided into different rectangular patches (a sub-area of ​​the image) for stitching. During the fusion process, a weighted multi-band algorithm is used to generate stitching seams. For each patch, a Laplacian pyramid and a weighted pyramid are established. Using the Laplacian pyramid, the exposure difference and the registration error of the image can be optimized, so that the brightness difference between the images becomes smoother, and the details of the image can also be retained by the underlying pyramid. By calculating the differential Gaussian pyramid, the Laplacian pyramid is quickly established. The steps of stitching a frame of image can be summarized as the following steps: (1) Determine the edge of the rectangular patch according to the relative position of the map point. When the edge exceeds the current area, the splicing area expands.

[0082] (2) Considering the row height, viewing angle and pixel position, the weighted image is calculated to ensure the orthorectification of the stitching as much as possible; then the homography transformation is calculated and applied to the corrected color image and the weighted image.

[0083] (3) Construct a Laplacian pyramid and a related weighted pyramid on the rectified image. Instead of using a weighted sum method, the optimal weight is found to fuse the image into a globally stitched rectangular patch.

[0084] When all frame image processing is completed, a two-dimensional image is obtained.

[0085] In embodiments of the present invention, traditional nautical charts may not be able to accurately depict the complex polar regions' icebergs, floating ice, and potential obstacles. Drones can rapidly acquire large numbers of high-resolution images and stitch them together to form orthophotos, providing clearer and more accurate geographic information. This helps ships identify their surroundings and promptly detect potential hazards during navigation. Drones can cover large areas and acquire real-time data in a relatively short period of time, ensuring that ships can safely navigate based on the latest imagery. However, stitching high-resolution images requires extensive computation and processing time, especially when dealing with large amounts of data, which increases hardware requirements. To address this, embodiments of the present invention employ a real-time stitching algorithm for drone video streams based on a simultaneous positioning and mapping method. By incorporating SLAM technology, this algorithm dynamically corrects the position and attitude (collectively referred to as pose) of real-time transmitted images, minimizing misalignment between adjacent images. This algorithm is based on image stitching, but its steps are more complex than traditional image stitching methods, including video frame extraction, feature extraction, feature matching, transformation matrix calculation, and image fusion.

[0086] In some embodiments of the present invention, step S105 includes: The two-dimensional map is input into a preset semantic segmentation model to extract sea ice and obtain extracted features; the preset semantic segmentation model includes a lightweight deep convolutional neural network, a pyramid pooling module, and a coordinate attention mechanism module; According to the two-dimensional map and the extracted features, navigable waters are obtained.

[0087] Among them, the preset semantic segmentation model can include a lightweight deep convolutional neural network, a pyramid pooling module, and a coordinate attention mechanism module; in order to solve the problems of the traditional semantic segmentation model in large-scale sea ice extraction, an improved DeepLabV3+ model (i.e., the preset semantic segmentation model) is adopted. Based on Sentinel-1 wide-band dual-polarization SAR imagery, Mobilenetv2 (i.e., a lightweight deep convolutional neural network) is used on the basis of the traditional DeepLabV3+ to replace Xception as the backbone network, densely connected void space pyramid pooling DenseASPP module (i.e., pyramid pooling module) and coordinate attention CA mechanism module (i.e., coordinate attention mechanism module). The technical route for sea ice extraction is as follows: Figure 6 Obtain the HH original image and HV original image, and then perform preprocessing, which includes thermal noise removal, radiometric calibration, filtering and dB conversion to obtain HH db 、HV db and HH db / HV db , a false color image is obtained by RGB color synthesis, and a sea ice dataset is obtained by labeling and cropping. The sea ice dataset is input into the improved DeepLabV3+ model, and the DeepLabV3+ model outputs the sea ice extraction results through the DenseASPP module and the MobileNetV2 module.

[0088] MobileNetV2 is a lightweight deep convolutional neural network architecture designed to provide high-performance computer vision solutions with low model size and computational cost. As the second generation of the MobileNet family, MobileNetV2 inherits the lightweight nature of its predecessor while improving performance and computational efficiency. The core design of MobileNetV2 is depthwise separable convolution, which helps reduce the number of parameters and computational complexity, accelerating sea ice feature extraction. Depthwise separable convolution consists of depthwise convolution and pointwise convolution. First, depthwise convolution uses a 3 × 3 kernel for each channel to extract features. Then, pointwise convolution uses a 1 × 1 kernel to increase the dimension of the features and generate the final output feature map. MobileNetV2 introduces an inverted residual module, which consists of an expansion layer, a depthwise convolution layer, and a projection layer. The expansion layer increases the number of channels to extract more sea ice information. Next, the depthwise convolution layer performs a depthwise separable convolution operation, and finally, a projection layer reduces the number of channels. This structure helps maximize the richness of sea ice features while reducing computational complexity. The coordination of the above modules makes MobileNetV2 lightweight and can better improve the efficiency of sea ice extraction tasks.

[0089] The core concept of the DenseASPP module is to integrate the ASPP module within a densely connected network structure. Dense connectivity involves concatenating feature maps from different layers, allowing each layer to directly access information from all previous layers. This dense connectivity allows extracted sea ice features to be transferred between different layers, maximizing their utilization. The DenseASPP module consists of the following key components: ① Multi-scale ASPP branch: Similar to the traditional ASPP module, the DenseASPP module also contains multiple parallel dilated convolution branches, each using a different dilation rate to capture sea ice information at different scales. Compared to ASPP, DenseASPP has a larger dilation rate, enabling it to extract sea ice features at a larger scale. ② Dense connectivity mechanism: Each branch in the DenseASPP module is densely connected to every other branch, allowing each branch to directly receive sea ice feature maps from every other branch. ③ 1×1 convolution: After connecting different branches, a 1×1 convolution is typically used to reduce the number of channels in the feature map, thereby reducing the computational complexity of sea ice feature extraction. Therefore, DenseASPP can make better use of the extracted sea ice features, better restore the sea ice edge information in the sea ice extraction task, and enhance the extraction effect of the sea ice edge.

[0090] In sea ice extraction tasks, small targets are often distributed along the edge of the ice. To further improve the model's ability to extract small pieces of sea ice, an attention mechanism can be introduced. The coordinate attention mechanism is a novel, lightweight attention mechanism that simultaneously considers inter-channel relationships and long-range positional information, effectively improving model accuracy while incurring minimal computational overhead. Traditional attention mechanisms, such as the classic SENet network with a channel-based attention mechanism, only consider relationships between channels to measure the importance of each channel, often ignoring the positional information of sea ice features. Although the later Convolutional Block Attention Module (CBAM) attempts to extract positional attention information through convolution after reducing the number of channels, convolution can only extract local relationships and lacks the ability to extract long-range relationships. Furthermore, the extraction of channel and spatial attention information is performed separately, resulting in limited improvement compared to SENet. To address this issue, a novel attention method that integrates spatial and channel information is used: the coordinate attention mechanism.

[0091] Combining it with the DeepLabV3+ model further improves the accuracy of fine sea ice extraction while ensuring real-time performance. The coordinate attention mechanism decomposes the global pooling into two dimensions, using a pooling kernel of (H, 1) or (1, W) to encode each channel along the vertical and horizontal coordinate directions, respectively. Based on the two features generated above, the two feature maps are further merged and then transformed using a 1×1 convolution as shown in formula (9):

[0092] Where, is a 1×1 convolution transformation function; the square brackets indicate the merging operation along the spatial dimension; is the nonlinear activation function h-Swish; is the intermediate feature. Map the intermediate feature Decomposed into two separate tensors and , is the module size reduction rate. and Transformed into a tensor with the same number of channels and activated by sigmoid, we get and , the final attention module output is shown in formula (10):

[0093] Where, 、 Respectively cOutput feature map and input feature map at the channel; 、 Respectively c The attention weights generated in the vertical and horizontal directions at the channel.

[0094] The embodiment of the present invention uses the MobileNetV2 feature extraction network to replace the original Xception network to achieve lightweight model and improve the speed of sea ice extraction. Figure 7 As shown in the figure, the MobileNetV2 feature extraction network adopts a depthwise separable convolution structure. This structure reduces the amount of computation and parameters by dividing the standard convolution operation into two steps: depthwise convolution and pointwise convolution. This reduces the model's runtime and memory usage while ensuring the accuracy of sea ice extraction. Secondly, the densely connected DenseASPP module is used to replace the ASPP module. The traditional ASPP module was proposed to capture contextual information of different receptive field sizes. It usually uses multiple parallel dilated convolution branches, and the features extracted by these branches are independent of each other. However, DenseASPP can share feature information between branches through skip connections, linking originally independent features to generate a denser feature pyramid, maximizing the utilization of sea ice extraction features. Another benefit of DenseASPP is that it can generate a larger sea ice feature receptive field without increasing the convolution kernel dilation rate, thus avoiding the kernel collapse problem that occurs when the convolution kernel dilation rate is increased in ASPP. Finally, a coordinate attention mechanism is introduced after the deep feature layer extracted by the DenseASPP module and the shallow feature layer extracted by MobileNetV2. This mechanism can adaptively weight pixels at different positions, enabling the network to better focus on important sea ice pixel positions, and simultaneously consider the relationship between each channel to measure the importance of sea ice features between different channels, thereby improving the accuracy of boundary details and the precision of sea ice extraction, so as to further improve the performance of sea ice extraction, and then obtain sea ice features through decoding.

[0095] Furthermore, the steps of combining the two-dimensional image and extracting features to obtain navigable waters are as follows: Among them, the two-dimensional image and extracted features are added as layers to the electronic nautical chart, and the segmented sea ice area is set as a dangerous area, and the navigable area is set as navigable waters. Ship drivers can intuitively perceive the sea ice distribution and choose the sailing route in the navigable waters.

[0096] The distribution and state of sea ice in this embodiment of the present invention directly impacts the safety of ship navigation. Image segmentation can clearly identify sea ice areas and their shape, size, and state (such as thickness and degree of melt), helping ships better understand their surroundings. Furthermore, the dynamic characteristics of sea ice, especially in polar waters, can change over a relatively short period of time. Image segmentation enables real-time or near-real-time monitoring, providing up-to-date sea ice information and helping ships make accurate navigation decisions. To address the issues of inaccurate detail extraction and slow extraction speed in traditional semantic segmentation models, a sea ice extraction method based on an improved DeepLabV3+ was constructed. First, the Xception backbone network was replaced with MobileNetV2, significantly reducing the number of model parameters while maintaining sea ice extraction accuracy and saving time. Second, the ASPP algorithm was improved to DenseASPP, further expanding the receptive field and obtaining denser features when extracting multi-scale features of sea ice. Finally, a coordinate attention mechanism was introduced to simultaneously focus on channel and spatial features, enhancing the extraction of detailed information about the sea ice edge.

[0097] In order to better implement the polar ship safe navigation method in the embodiment of the present invention, based on the polar ship safe navigation method, the embodiment of the present invention also provides a polar ship safe navigation device, such as Figure 8 As shown, the polar ship safety navigation device 800 includes: The data acquisition module 801 is used to acquire the current driving data of the ship and the video data of the drone when the ship is traveling in the polar region; The image processing module 802 is used to enhance and fuse the visible light image and infrared image of each frame in the video data to obtain a fused image corresponding to each frame; The feature extraction module 803 is used to extract features from the fused image to obtain multiple key frames, and to search for map points on the multiple key frames to obtain multiple map point positions; The coordinate conversion module 804 is used to perform closed-loop detection and global optimization on multiple key frames to obtain multiple optimized key frames, and to perform two-dimensional map stitching on the map point positions corresponding to the multiple optimized key frames to obtain a two-dimensional map; The image stitching module 805 is used to extract sea ice from the two-dimensional map to obtain navigable waters; and control the navigation of ships in the polar region based on the navigable waters.

[0098] The polar ship safe navigation device 800 provided in the above embodiment can implement the technical solution described in the above polar ship safe navigation method embodiment. The specific implementation principles of the above modules or units can be found in the corresponding contents of the above polar ship safe navigation method embodiment, which will not be repeated here.

[0099] The above is a detailed introduction to the polar ship safe navigation method and device provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only intended to help understand the method and core concept of the present invention. At the same time, for those skilled in the art, according to the concept of the present invention, there may be changes in the specific implementation method and scope of application. In summary, the contents of this specification should not be understood as limiting the present invention.

Claims

1. A method for safe navigation of polar ships, characterized in that: include: Obtain current navigation data of ships and video data from drones when sailing in polar regions; Performing enhancement processing and fusing the visible light image and the infrared image of each frame of the video data to obtain a fused image corresponding to each frame of the image; Performing feature extraction on the fused image to obtain a plurality of key frames, and performing map point search based on the plurality of key frames to obtain a plurality of map point positions; Performing closed-loop detection and global optimization on the multiple key frames to obtain multiple optimized key frames, and performing two-dimensional map stitching on map point positions corresponding to the multiple optimized key frames to obtain a two-dimensional map; Sea ice is extracted from the two-dimensional map to obtain navigable waters; and the ship is controlled to navigate in the polar region based on the navigable waters.

2. The method for safe navigation of polar ships according to claim 1, characterized in that: The enhancing and fusing of the visible light image and the infrared image of each frame of the video data to obtain a fused image corresponding to each frame of the image includes: Performing illumination enhancement processing on the visible light image of each frame of the video data to obtain a first partial image; the first partial image is an image with prominent appearance detail information and obvious contrast; Brightness enhancement is performed on the infrared image of each frame of the video data to obtain a second partial image; the second partial image is an image with enhanced contrast and uniform overall brightness distribution; The first partial image and the second partial image of each frame of image are fused and reconstructed to obtain a corresponding fused image.

3. The method for safe navigation of polar ships according to claim 2, characterized in that: The performing illumination enhancement processing on the visible light image of each frame of the video data to obtain a first partial image includes: Converting the values ​​of the red, green, and blue channels of the visible light image of each frame of the video data to obtain a brightness component; Calculating the visible light image according to an image enhancement algorithm to obtain reflection components of different scales; Calculating the probability of association between the brightness component and each reflection component to obtain an association weight for each reflection component; Obtaining an enhanced brightness image according to the brightness component, each reflection component and the associated weight; The color domain of the enhanced brightness image is reconstructed and merged to obtain a first partial image.

4. The method for safe navigation of polar ships according to claim 2, characterized in that: performing brightness enhancement on the infrared image of each frame of image in the video data; Get the second part of the image, including: performing illumination enhancement processing on the infrared image of each frame of the video data based on a logarithmic scaling function to obtain an illumination enhanced image; Performing non-complex exponential function processing on the infrared image to obtain a non-complex exponential image; combining the illumination-enhanced image and the non-complex-exponential image to obtain a combined image; enhancing the overall brightness of the combined image to obtain a brightness-enhanced image; Normalization is performed on the brightness enhanced image to obtain a second partial image.

5. The method for safe navigation of polar ships according to claim 2, characterized in that: The extracting features of the fused image to obtain a plurality of key frames, and searching for map points based on the plurality of key frames to obtain a plurality of map point positions, includes: Initializing a map based on the images of the first two frames in the fused image to obtain three-dimensional map points; Extract features from each frame of the fused image one by one according to a feature extraction algorithm to obtain a plurality of two-dimensional feature points; A plurality of key frames are obtained according to the three-dimensional map points and the plurality of two-dimensional feature points, and map point searches are performed on the plurality of key frames to obtain a plurality of map point positions.

6. The method for safe navigation of polar ships according to claim 5, characterized in that: The multiple two-dimensional feature points include the two-dimensional feature points of the current frame image; the multiple key frames are obtained based on the three-dimensional map points and the multiple two-dimensional feature points, and map point search is performed on the multiple key frames to obtain multiple map point positions, including: Set up the initial map; Obtaining a current frame pose according to the three-dimensional map points and the two-dimensional feature points of the current frame image; When the current frame pose meets a preset condition, determining the current frame image as a key frame, and adding the key frame to the initial map to obtain an updated target map; Searching according to adjacent key frames in the target map to obtain matching points; According to the matching point, the current map point position is obtained; When the processing of the plurality of two-dimensional feature points is completed, a plurality of key frames and a map point position corresponding to each key frame are obtained.

7. The method for safe navigation of polar ships according to claim 1, characterized in that: The performing closed-loop detection and global optimization on the multiple key frames to obtain multiple optimized key frames includes: Determining multiple candidate frames based on a positional relationship between the current frame and the multiple key frames; Performing closed-loop detection on similar transformations between the multiple candidate frames and the current frame using a random sampling consensus algorithm to obtain closed-loop key frames; Correcting the multiple key frames according to the closed-loop key frames to obtain multiple corrected key frames; Obtaining a correspondence between a global position and a camera position based on a difference between a GPS constraint and a timestamp in the video data; Performing coordinate system conversion on the multiple corrected key frames according to the corresponding relationship and similarity transformation to obtain multiple coordinate system key frames; After optimizing the difference, the residuals of the multiple coordinate system key frames are minimized and roughly estimated and optimized according to the map point projection error minimization principle to obtain an optimization result; the optimization result includes the optimal multiple optimized key frames.

8. The method for safe navigation of polar ships according to claim 7, characterized in that: The step of performing two-dimensional map stitching on the map point positions corresponding to the plurality of optimized key frames to obtain a two-dimensional map includes: Determining the edges of a rectangle based on the relative posture of the map point positions; When the edge of the rectangle exceeds the current area, the stitching area is expanded and the weighted image and homography transformation are calculated; Correcting the weighted image according to the homography transformation and the preset corrected color image to obtain a corrected image, and establishing a weight pyramid; An optimal weight is determined according to the weight pyramid, and the corrected image is fused onto the rectangle of the stitching area according to the optimal weight to obtain a two-dimensional map.

9. The method for safe navigation of polar ships according to claim 1, characterized in that: The step of extracting sea ice from the two-dimensional map to obtain navigable waters includes: Inputting the two-dimensional map into a preset semantic segmentation model to extract sea ice and obtain extracted features; the preset semantic segmentation model includes a lightweight deep convolutional neural network, a pyramid pooling module, and a coordinate attention mechanism module; Navigable waters are obtained according to the two-dimensional map and the extracted features.

10. A polar ship safety navigation device, characterized in that: include: A data acquisition module is used to obtain the current driving data of the ship and the video data of the drone when it is traveling in the polar region; An image processing module is used to enhance and fuse the visible light image and infrared image of each frame of the video data to obtain a fused image corresponding to each frame of the image; a feature extraction module, configured to extract features from the fused image to obtain a plurality of key frames, and perform map point searches based on the plurality of key frames to obtain a plurality of map point positions; A coordinate conversion module is used to perform closed-loop detection and global optimization on the multiple key frames to obtain multiple optimized key frames, and to perform two-dimensional map stitching on the map point positions corresponding to the multiple optimized key frames to obtain a two-dimensional map; The image stitching module is used to extract sea ice from the two-dimensional map to obtain navigable waters; and control the ship to navigate in the polar region based on the navigable waters.