Ship panoramic vision system and method combining distortion correction and image stitching optimization

By employing deep learning and multi-source sensor fusion technologies, adaptive distortion correction and high-precision image stitching of the ship panoramic vision system have been achieved. This has solved problems such as error cascading accumulation, failure to match weak texture features on the sea surface, and fusion of overlapping areas, providing seamless panoramic display and intelligent collision warning capabilities.

CN122434795APending Publication Date: 2026-07-21SHANGHAI MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI MARITIME UNIVERSITY
Filing Date
2026-06-23
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing ship panoramic vision systems, the independent optimization of distortion correction and image stitching modules leads to cascading error accumulation, failure in matching weak texture features of the sea surface, unstable dynamic parallax of ship motion, ghosting seams generated by overlapping area fusion, disconnection of boundaries in invalid area filling, and a lack of engineering triggering strategies based on nautical parameters and multi-source data fusion for collision warning.

Method used

By employing deep learning, computer vision, and multi-source sensor fusion technologies, adaptive joint correction and high-precision image stitching of fisheye images are achieved through adaptive joint correction, mixed feature extraction in the frequency and spatial domains, semantically guided boundary filling, and interactive panoramic intelligent early warning, forming a 360-degree all-round visual monitoring capability.

Benefits of technology

It improves image stitching accuracy, enhances the robustness and coherence of panoramic vision, and achieves seamless panoramic display and quantifiable collision risk assessment and linkage response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434795A_ABST
    Figure CN122434795A_ABST
Patent Text Reader

Abstract

The application provides a ship panoramic vision system and method for distortion correction and image stitching joint optimization, and relates to the technical field of computer vision and artificial intelligence. The method comprises the following steps: synchronously collecting fisheye images and multi-source sensing data; constructing a learnable distortion model containing distortion center offset, eliminating error accumulation by using feedback consistency loss and stitching network end-to-end joint training; extracting frequency-space mixed features, combining inertial data and optical flow to perform motion pre-compensation and multi-dimensional constraint fine stitching; realizing seamless transition of the boundary based on semantic analysis and partial differential equation collaborative filling of invalid areas; and fusing multi-source heterogeneous data and performing hierarchical intelligent early warning based on navigation parameters. The application breaks the error cascade amplification caused by traditional isolated processing, overcomes the problems of weak texture and dynamic parallax on the sea surface, and improves the robustness of ship panoramic perception.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and artificial intelligence technology, and involves deep learning, image processing, and panoramic image stitching technology. Specifically, it is a ship panoramic vision system and method that jointly optimizes distortion correction and image stitching. Background Technology

[0002] Ships need to monitor their surroundings comprehensively during navigation to ensure safety. Existing ship monitoring systems typically use multiple independent cameras to monitor different directions, requiring crew members to simultaneously observe multiple displays to obtain complete information about the surrounding situation. This results in problems such as blind spots, fragmented views, and the inability to form a unified panoramic perspective.

[0003] The unique environment of ships presents distinct technical challenges to visual surveillance. Ships at sea are subject to continuous roll, pitch, and sway movements due to wind and waves, leading to unstable image acquisition and significant parallax variations between adjacent frames. The complex and variable lighting conditions at sea, including direct sunlight, sea surface reflections, backlighting, and low-light nighttime conditions, make it difficult for traditional image processing algorithms to adapt to these wide-ranging lighting variations. The vast field of view around a ship necessitates 360-degree surveillance to identify potential collision risks. While traditional fisheye lenses can achieve wide-angle shooting, they produce severe radial and tangential distortion, with more pronounced distortion at the edges, affecting the accuracy of target recognition and distance assessment.

[0004] Existing panoramic image stitching technologies have the following main drawbacks.

[0005] First, the disconnect between distortion correction and stitching registration optimization leads to a cascading accumulation of errors. Existing methods often employ the Brown-Conrady model for independent offline calibration, assuming the distortion center coincides with the image center, failing to account for actual optical center offset. This results in residual barrel or pincushion distortion in edge regions. Furthermore, the parameter optimization of the correction module is independent of the registration requirements of the stitching network. Correction accuracy is not fed back to the stitching module, and stitching geometric errors do not correct the correction parameters in reverse. With each module executing in series and optimizing independently, the system as a whole struggles to achieve global optimization.

[0006] Second, feature extraction and matching success rates are low in low-texture sea surface scenes. Traditional algorithms such as SIFT, SURF, and ORB rely on salient features such as corners and edges, which makes feature extraction difficult and matching success rates low in environments with highly similar sea surface textures. Existing deep learning stitching networks mostly use general backbone networks (such as ResNet-50) for spatial domain feature extraction, failing to identify and enhance key geometric structures such as the horizon and ship outlines based on the characteristics of sea surface images, and also failing to introduce frequency domain information to compensate for the lack of spatial domain texture, resulting in insufficient robustness of stitching in low-contrast sea environments.

[0007] Third, there is a lack of prediction and compensation mechanisms for dynamic parallax caused by ship motion. Ships experience continuous roll, pitch, and sway motion due to wind and waves, leading to dynamic parallax changes and exposure differences between adjacent frames and between adjacent camera fields of view. Existing methods assume a fixed relative camera pose, do not incorporate motion prediction modules based on IMU or optical flow, do not construct dynamic region-of-interest masks to adaptively adjust the feature matching range, and do not utilize temporal consistency constraints between adjacent frames. This results in decreased registration accuracy when the ship is violently swaying, leading to jitter and discontinuity in the video sequence.

[0008] Fourth, the overlapping area fusion strategy is crude and prone to ghosting and seams. Existing stitching methods mostly use a simple weighted fusion strategy. When there are exposure differences, parallax, or local misalignment between stitched images, ghosting, misalignment, and seams are likely to appear in the overlapping areas, affecting the visual continuity of the panorama.

[0009] Fifth, the content filling and stitching process are disconnected and lack semantic guidance. Due to the limited field of view of fisheye lenses and installation geometric constraints, there are invalid black areas in the stitching result. Existing methods typically use cropping, solid color filling, or independent partial differential equation (PDE) image inpainting techniques for processing, without integrating the filling process with the output of the stitching network, and without utilizing the semantic information of panoramic images (such as the filling area usually being the sky or sea). This results in brightness discontinuities or texture mismatches between the filled area and the effective image boundary.

[0010] Sixth, the interactive display and obstacle warning mechanisms lack engineered trigger strategies and multi-source data fusion. Existing systems mostly use fixed bird's-eye view display modes, unable to adapt to crew needs by switching perspectives, zooming, and adaptively adjusting motion. Some systems propose a tiered warning concept, but fail to define specific trigger thresholds based on nautical parameters such as distance, relative speed, time to collision (TCPA), and closest encounter distance (DCPA), nor specify the exact method of overlaying warnings (e.g., color, flashing frequency, graphic style), and further fail to clarify the data fusion method with shipboard AIS, radar, and other navigation systems, making it difficult to achieve intelligent collision risk assessment and coordinated response.

[0011] Therefore, it is necessary to provide a ship panoramic vision system and method that combines distortion correction and image stitching optimization to solve the above-mentioned technical problems and meet the higher requirements for ship navigation safety supervision. Summary of the Invention

[0012] This invention addresses the fundamental problem in existing ship panoramic vision systems: the independent optimization of distortion correction and image stitching modules leads to cascading error accumulation and poor overall performance. Furthermore, this results in a series of issues in specific marine scenarios, including failure to match weak surface texture features, unstable dynamic parallax due to ship motion, ghosting seams during overlapping area fusion, disconnected boundaries in invalid area filling lacking semantic guidance, and a lack of engineering-based triggering strategies and multi-source data fusion mechanisms for collision warning based on nautical parameters. To address these issues, this invention provides a ship panoramic vision system and method that jointly optimizes distortion correction and image stitching. Employing deep learning, computer vision, image processing, and multi-source sensor fusion technologies, it achieves comprehensive visual monitoring across multiple aspects, including adaptive joint correction of fisheye images, high-precision image stitching facing the sea surface, semantically guided boundary collaborative filling, and interactive panoramic intelligent early warning, thus forming a 360-degree all-around visual monitoring capability.

[0013] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: A method for joint optimization of distortion correction and image stitching for ship panoramic vision includes the following steps: Step S1, Multi-source data acquisition: Raw fisheye images are acquired through fisheye cameras deployed around the ship, ship attitude angular velocity data are acquired simultaneously through the inertial measurement unit, ship target information is acquired through the automatic identification system, and ranging data is acquired through the shipborne radar. The data from each sensor are synchronized through a time protocol. Step S2, Adaptive Joint Correction: Establish a distortion model containing the coordinates of the distortion center to perform coordinate mapping correction on the original fisheye image. Use the distortion correction parameters as learnable variables and perform end-to-end joint training with the image stitching network. By jointly optimizing the feedback consistency loss in the objective function, the registration error of the image stitching network in the overlapping area is directly backpropagated and the distortion correction parameters are dynamically adjusted to output a distortion-free image. Step S3, Image stitching: Extract frequency domain and spatial domain hybrid features from the distortion-free image, estimate the optical motion field based on the distortion-free images of adjacent frames, use the attitude angular velocity data and the optical motion field to adaptively fuse and construct the motion field for inter-frame motion pre-compensation, complete the image coarse alignment, and use joint constraints including geometric consistency, boundary consistency and temporal consistency loss to perform fine reconstruction network fusion, and output the initial panoramic image. Step S4, joint boundary filling: generate a filling mask for the invalid region of the initial panoramic image and perform semantic parsing. Based on the semantic category of the region, a partial differential equation filling strategy is used to solve the invalid region at the pixel level. At the same time, minimize the joint loss of stitching and filling, so that the filling result of the partial differential equation has a reverse influence on the fusion weight of the stitching boundary, and output a seamless panoramic image. Step S5, Panoramic Display and Intelligent Early Warning: Display the seamless panoramic image, perform multi-source data state fusion on the visual detection position, the ship target information and the ranging data through Kalman filtering, establish a graded early warning triggering mechanism based on the nearest encounter time, the nearest encounter distance and the relative speed, and execute graded early warning annotation and underlying linkage actions on the seamless panoramic image.

[0014] On the other hand, the present invention also provides a ship panoramic vision system jointly optimized by distortion correction and image stitching, including a computing and perception unit configured to perform the above-described method, specifically including: The multi-source image acquisition module is used to acquire raw fisheye images and various sensor data through deployed fisheye cameras, inertial measurement units, automatic identification systems for ships, and radar equipment, and to achieve time synchronization. An adaptive joint correction module, whose image input end is connected to the image output end of the multi-source image acquisition module, is used to establish a distortion model containing the coordinates of the distortion center to correct the original fisheye image. The distortion correction parameters are jointly trained with the image stitching network, and the parameters are dynamically adjusted using feedback consistency loss. The distortion-free image is output through its output end. The image stitching network module has its image input end connected to the output end of the adaptive joint correction module, and its data input end connected to the data output end of the inertial measurement unit of the multi-source image acquisition module. It is used to extract the mixed features of frequency domain and spatial domain, use the received attitude angular velocity data and optical flow adaptive fusion to perform inter-frame motion pre-compensation and coarse alignment, and perform fine reconstruction under multi-dimensional joint constraints. The initial panoramic image is output through its output end. The boundary joint filling module, whose input end is connected to the output end of the image stitching network module, is used to perform pixel-level iterative filling of invalid regions based on semantic parsing results and partial differential equation filling strategy. It also achieves boundary optimization by inversely influencing the fusion weights through stitching and filling collaborative loss, and outputs a seamless panoramic image through its output end. The panoramic display and intelligent early warning module has its image input end connected to the output end of the boundary joint filling module, and its data input end connected to the data output end of the ship automatic identification system and radar equipment of the multi-source image acquisition module. It is used to display the seamless panoramic image, perform state fusion on the received multi-source data based on Kalman filtering, and calculate and execute hierarchical early warning visual annotation and underlying physical linkage according to nautical parameters.

[0015] Compared with the prior art, the beneficial effects of the present invention are: First, by establishing an end-to-end joint optimization of learnable distortion parameters and stitching network, the registration error is directly corrected by using feedback consistency loss, eliminating the cascading accumulation of errors between modules. In areas with severe ship sway and edge distortion, the corrected image is more suitable for stitching requirements, and the stitching accuracy is improved.

[0016] Second, by using frequency-spatial domain hybrid enhancement and sea surface attention mechanism, low-frequency structural information is used to compensate for the lack of spatial texture, guiding the network to focus on salient structures such as the sea horizon and ship outline, thereby improving the feature matching success rate in low-contrast sea surface scenes and enhancing the robustness of panoramic stitching.

[0017] Third, based on the adaptive fusion of IMU and optical flow, motion compensation and dynamic ROI masking effectively predict and compensate for inter-frame motion, maintaining registration accuracy even when the ship is violently rocking, and reducing video sequence splicing jitter by combining temporal consistency loss.

[0018] Fourth, the fine reconstruction employs joint constraints of geometric consistency, boundary consistency, and temporal consistency to suppress ghosting and misalignment in overlapping areas, resulting in natural transitions in fused areas and enhanced panoramic visual coherence.

[0019] Fifth, through semantically guided anisotropic PDE filling and splicing-filling collaborative loss coupling optimization, the filling content is semantically coordinated with the surrounding environment and the boundaries are continuous, avoiding brightness abrupt changes and texture mismatch, thus achieving seamless panoramic display.

[0020] Sixth, it clearly defines the three-level warning trigger thresholds based on TCPA / DCPA and relative speed, as well as the corresponding HMI parameters, and integrates AIS, radar, and visual data to conduct a comprehensive collision risk assessment, thereby achieving quantifiable collision risk classification warnings and coordinated responses.

[0021] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0022] Figure 1 This is a technical roadmap for the ship panoramic vision system according to the present invention; Figure 2 This is a schematic diagram comparing the image before and after correction; Figure 3 This is a schematic diagram of the stitching effect without using the specific processing mechanism for ocean surface textures; Figure 4 A schematic diagram illustrating the stitching effect using a specific processing mechanism for ocean surface textures; Figure 5 To fill in the previous panoramic view; Figure 6 This is the panoramic view after filling in the image. Detailed Implementation

[0023] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings, so as to more clearly understand the purpose, features and advantages of this invention. It should be understood that the embodiments shown in the drawings are not intended to limit the scope of this invention, but are only for illustrating the essential spirit of the technical solutions of this invention. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this invention.

[0024] Unless the context requires otherwise, throughout the specification and claims, the word “comprising” and its variations, such as “including” and “having”, shall be understood to have an open, inclusive meaning, that is, to be interpreted as “including, but not limited to”.

[0025] Throughout this specification, references to "an embodiment" or "an embodiment" indicate that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Therefore, the appearance of "in an embodiment" or "an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment. Furthermore, a particular feature, structure, or characteristic may be combined in any manner in one or more embodiments.

[0026] The singular forms “a” and “the” used in this specification and the appended claims include plural references unless otherwise expressly stated herein. It should be noted that the term “or” is generally used to mean “and / or” unless otherwise expressly stated herein.

[0027] In the following description, in order to clearly demonstrate the structure and working method of the present invention, a number of directional terms will be used. However, terms such as "front", "back", "left", "right", "outside", "inside", "outward", "inward", "up", and "down" should be understood as convenient terms and not as limiting terms.

[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention.

[0029] Example 1 This embodiment describes the overall architecture, module composition, and inter-module connections and data flow of the ship panoramic vision system jointly optimized by distortion correction and image stitching according to the present invention. Details are as follows: like Figure 1As shown, this system consists of a multi-source image acquisition module, an adaptive joint correction module, an unsupervised deep image stitching network, a semantically guided boundary filling module, and an interactive display and intelligent early warning module. The above modules are trained end-to-end through a multi-task joint learning framework.

[0030] The output of the multi-source image acquisition module is connected to the input of the adaptive joint correction module. The output of the adaptive joint correction module is connected to the input of the unsupervised depth image stitching network facing the sea surface. The output of the unsupervised depth image stitching network facing the sea surface is connected to the input of the semantically guided boundary joint filling module. The output of the semantically guided boundary joint filling module is connected to the input of the interactive panoramic display and intelligent early warning module.

[0031] The IMU data output of the multi-source image acquisition module is also connected to the input of the motion-adaptive coarse image alignment module in the unsupervised depth image stitching network facing the sea surface, providing roll and pitch angular velocity data. The AIS and radar data outputs of the multi-source image acquisition module are also connected to the input of the interactive panoramic display and intelligent early warning module, providing target information and ranging data.

[0032] The multi-source image acquisition module acquires raw fisheye images through fisheye cameras deployed in the forward, aft, left, and right directions of the ship. Simultaneously, it acquires ship attitude and angular velocity data via IMU sensors, gathers surrounding ship target information via an AIS receiver, and acquires target distance and azimuth data via shipborne radar. Data from each sensor is synchronized at the microsecond level via an industrial switch supporting a precise time protocol. The specific hardware deployment and data acquisition method of the multi-source image acquisition module are described in detail in Example 2.

[0033] The adaptive joint correction module receives the raw fisheye image from the multi-source image acquisition module and corrects the image using an improved distortion model that incorporates the distortion center coordinates (x0, y0). The corrected distortion-free image is generated by establishing a coordinate mapping lookup table and bilinear interpolation resampling. The core feature of this module is the integration of distortion correction parameters... As a learnable variable, it is jointly trained end-to-end with the subsequent stitching network. The backpropagation and dynamic adjustment of the stitching network registration error on the correction parameters are achieved by jointly optimizing the feedback consistency loss in the objective function. The correction model, correction process, and specific process of joint optimization training of the adaptive joint correction module are described in detail in Example 3.

[0034] The unsupervised depth image stitching network facing the sea surface receives the corrected multi-channel images and generates an initial panoramic image. This stitching network consists of three sub-modules: (1) The sea surface texture-aware enhanced feature extraction network, which introduces a frequency-domain to spatial-domain hybrid feature extraction mechanism and a sea surface scene attention module, uses the frequency-domain enhancement branch to compensate for the lack of spatial-domain texture and guides the network to focus on the sea horizon and ship contours; (2) The motion-adaptive rough image alignment module, which performs inter-frame motion pre-compensation based on the motion field adaptively fused by IMU and optical flow, constructs a dynamic ROI mask to adjust the feature matching range, and estimates the homography matrix after compensation to complete the rough alignment of the images; (3) The geometric consistency-aware fine reconstruction module, which adopts an encoder-decoder structure and eliminates stitching seams and ghosting under the joint constraints of geometric consistency, boundary consistency, and temporal consistency losses, and outputs a fine panoramic image. The specific implementation method of this stitching network is described in detail in Embodiment 4.

[0035] The semantic-guided boundary joint filling module receives the panoramic image output by the stitching network and fills the black invalid areas caused by fisheye projection in it. The module first identifies the semantic categories (such as sky, sea water, etc.) of the filling areas through a semantic segmentation network, adopts an anisotropic PDE filling strategy for different semantic areas, introduces boundary consistency constraints and semantic guidance terms in the energy functional, iteratively solves by the finite difference method, and establishes a stitching-filling collaborative loss to couple the filling process with the output end of the stitching network, and finally generates a complete panoramic image with seamless boundary transition. The specific implementation method of the semantic-guided boundary joint filling module is described in detail in Embodiment 5.

[0036] The interactive panoramic display and intelligent warning module receives the filled panoramic image and presents it on the display interface in a spherical projection manner. This module supports functions such as zooming, rotating, view switching, and motion-adaptive view adjustment based on the hull attitude of the IMU, and performs real-time data fusion with the ship AIS, radar, and GPS systems. Based on nautical parameters such as TCPA, DCPA, and relative speed, a three-level warning trigger mechanism is defined, and the positions of obstacles are superimposed and marked on the panoramic image with color coding, flashing frequency, and graphic styles, and corresponding linkage actions are executed. The specific implementation method of the interactive panoramic display and intelligent warning module is described in detail in Embodiment 6.

[0037] The above five modules are end-to-end co-trained through a multi-task joint learning framework, incorporating the three major tasks of distortion correction, image stitching, and ship target detection into a unified training system. The joint training adopts a gradient normalization strategy to dynamically balance the weights of each task, and ensures that the spatial understanding of the stitching network and the detection network for the overlapping areas is consistent through the inter-task consistency loss. The specific implementation method of the multi-task joint learning framework is described in detail in Embodiment 7.

[0038] The system's workflow includes: a multi-source image acquisition module acquires raw fisheye images and sensor data and aligns them synchronously; an adaptive joint correction module corrects distortion in the fisheye images and dynamically adjusts correction parameters through a joint optimization mechanism; an unsupervised depth image stitching network facing the sea surface generates an initial panoramic image through three stages: feature enhancement, motion compensation coarse alignment, and consistency constraint fine reconstruction; a semantically guided boundary joint filling module performs semantically guided PDE filling on invalid regions and achieves seamless boundary transitions through collaborative loss; and an interactive panoramic display and intelligent early warning module displays the panoramic image interactively and executes hierarchical early warnings based on the multi-source data fusion results.

[0039] Example 2 This embodiment provides a ship panoramic vision method that combines distortion correction and image stitching optimization. This method corresponds to the workflow of the panoramic vision system described in Embodiment 1 and is executed according to the following steps.

[0040] Step S1: Multi-source data acquisition One fisheye camera is installed at the bow, stern, port, and starboard sides of the ship. Each camera has a horizontal field of view of 180 degrees and a vertical field of view of 180 degrees, achieving 360-degree coverage without blind spots through proper arrangement. The fisheye cameras are 1920×1920 pixel resolution fisheye cameras with a frame rate of 30fps. The original fisheye image suffers from severe radial distortion, which becomes more pronounced closer to the edges. Let the resolution of the original image be W×H, and the image center be O(W / 2,H / 2). For any pixel P(x,y) in the image, what is its distance from the image center? The degree of radial distortion is proportional to the distance r.

[0041] Simultaneously, an inertial measurement unit (IMU) is deployed at the ship's center of gravity to collect three-axis angular velocity and acceleration data, with an output frequency of 100Hz. The IMU data is used to obtain the roll rate. and pitch angular velocity This provides information on changes in ship attitude for subsequent motion compensation. The AIS receiver and shipborne radar output target information via the NMEA 0183 / 2000 protocol, including the target ship's MMSI, name, speed, heading, position, and distance and azimuth obtained by radar ranging.

[0042] Unlike existing solutions that only acquire images, this method achieves microsecond-level time synchronization of data from various sensors through an industrial switch supporting PTP (Precise Time Protocol), providing multi-source data support for subsequent motion compensation, timing prediction, and obstacle collision warning. After acquisition, the raw fisheye image is output to step S2, the IMU data is output to step S3, and the AIS and radar data are output to step S5.

[0043] Step S2: Adaptive Joint Correction This step uses an improved fisheye distortion model to correct the original fisheye image acquired in step S1. This step is performed in the adaptive joint correction module described in Example 1. See below for the results before and after correction. Figure 2 The details are as follows: Step S2.1: Establish an improved fisheye distortion model Based on the traditional Brown-Conrady model, distortion center coordinates are introduced to establish a more accurate distortion model. Let the pixel coordinates in the original distorted image be ( , The corrected, distortion-free image coordinates are (x, y), and the distortion center coordinates are (...). , The radial distortion coefficient is , , The tangential distortion coefficient is , Radial distance r² = (x - )²+(y- If )², then the improved distortion model can be expressed as: By introducing the coordinates of the distortion center ( , This model can more accurately describe the distortion characteristics of actual fisheye lenses and solves the problem of inaccurate edge region correction caused by the traditional model assuming that the distortion center coincides with the image center.

[0044] Step 2.2: Establish a coordinate mapping lookup table To improve online inference efficiency, the coordinate mapping relationship between the corrected image and the original image is pre-calculated, and a lookup table (LUT) is established. For each integer pixel coordinate (x, y) in the corrected image, its corresponding floating-point coordinate (x', y') in the original distorted image is calculated according to the inverse distortion model.

[0045] Step 2.3: Pixel resampling using bilinear interpolation Since the calculated corresponding position (x', y') is usually not an integer coordinate, bilinear interpolation is needed to calculate the pixel value. Let the four neighboring integer coordinates of (x', y') be (x1, y1), (x2, y1), (x1, y2), and (x2, y2), where... , , , The interpolation weights are... , The interpolation formula is then: Step 2.4: End-to-end joint optimization training The core innovation of this embodiment lies in establishing a joint optimization mechanism for distortion correction and stitching registration. Distortion correction parameters... As a learnable variable, its initial value is set to... =0.1, k2=0.01, k3=0.001, p1=p2=0, x0=W / 2, y0=H / 2. The calibration module and the splicing network are jointly trained end-to-end to establish a two-way feedback mechanism.

[0046] The joint optimization objective function is defined as: Among them, splicing loss Measuring reprojection error and structural similarity of the stitched images: Correction loss Measuring the straightness preservation and distortion removal of the corrected image: Feedback consistency loss Defined as the norm of the registration error of the stitching network with respect to the gradient of the correction parameters: in, The registration error of the stitching network in the overlapping region is defined as follows: =|| - ||1 represents the L1 reprojection error between the two images in the overlapping region. Representing an image Pixel values ​​in the overlapping region, Indicates the image After applying the homography transformation H, the overlapping region is considered. The stitching network employs a fully differentiable architecture, and all operations (including feature extraction, homography estimation, deformation, and fusion) support gradient backpropagation. For steps involving non-differentiable operations, such as RANSAC, a straight-through estimator or a differentiable RANSAC is used to ensure the integrity of the gradient propagation path.

[0047] In the formula, Indicates the image The deformation result after applying homography transformation H, where ||·||1 is the L1 norm and SSIM is the structural similarity index. λ1, λ2, λ1 and μ are the weighting coefficients for balancing the loss terms; in this embodiment, they are set to λ1=0.1 and λ2=0.05. =0.5, μ=0.01. The above weight coefficients are determined by grid search on the validation set. The search range of λ1 is [0.05, 0.2], and the search range of λ2 is [0.01, 0.1]. Stable convergence can be achieved within this range.

[0048] The Adam optimizer is used during training to adjust the learning rate and parameters. =1×10 -4 splicing network learning rate =1×10 - ³. In each iteration, a pair of adjacent camera images are input simultaneously. The correction module generates a corrected image, which is then fed into the stitching network to obtain the registration error. After calculating the total loss and backpropagating, the registration error gradient of the stitching network is... Backpropagation is performed to the correction parameters, and updates are made simultaneously. This mechanism enables the correction module to receive feedback signals from the stitching module in real time, dynamically adjusting the correction parameters to ensure that the corrected image meets the registration requirements of the stitching network to the greatest extent possible, thus achieving global optimization of the panoramic image processing workflow.

[0049] Specifically, the backpropagation path from the splicing loss to the correction parameters is achieved through the following gradient propagation: / =( / )·( / ) in / The values ​​are obtained by taking the analytical partial derivatives of the correction parameters with respect to the distortion model. This is the corrected image. Since the distortion model formula is a continuously differentiable polynomial function, the partial derivative can be calculated analytically with precision.

[0050] Through this end-to-end joint optimization, for example, when the stitching network detects a large registration error in the edge overlap area, the feedback gradient will drive x0, y0 to shift towards the actual optical center, while adjusting the radial coefficient to optimize the edge straightness, effectively avoiding the error accumulation problem in the traditional serial architecture.

[0051] In one specific implementation, during training, each iteration performs the following operations: inputting the corrected image into the stitching network to obtain the registration error E_stitch; calculating the gradient update amount of the correction parameters based on the registration error: , where η is the learning rate. The joint optimization objective is: , in To splice network parameters, For the estimated homography matrix, The image after correction using the correction parameters. To compensate for the loss of structural similarity in the splicing reprojection, To correct for straightness preservation and correction loss, this end-to-end joint optimization allows the correction module to receive registration error feedback from the stitching network in real time and dynamically adjust parameters, effectively avoiding the error accumulation problem in traditional serial architectures and achieving global optimization of the panoramic image processing workflow.

[0052] Step 2.5: Alternative Correction Implementation Method As another implementation, the correction module may include a lightweight parameter prediction network Φ, which takes the feature error map output by the current stitching network as input and predicts the increment Δ of the correction parameters. Update parameters ← +Δ Similarly, incremental predictions are constrained by feedback consistency loss. This approach can be considered an equivalent variant of the direct learning in step 2.4.

[0053] After the correction is completed, the corrected distortion-free image is output to step S3.

[0054] Step S3: Unsupervised depth image stitching facing the sea surface This step stitches together the corrected multi-channel images output from step S2. This step is performed in the unsupervised depth image stitching network facing the sea surface described in Example 1.

[0055] Step 3.1: Enhanced Feature Extraction for Sea Surface Texture Perception Existing general backbone networks (such as ResNet-50) lack feature discrimination power in weak textured sea surface scenes. In this embodiment, a sea surface texture perception enhancement feature extraction network is used to replace the general backbone network.

[0056] For input image pairs and First, low-frequency structure extraction is performed through a frequency domain enhancement branch. A two-dimensional discrete Fourier transform is then performed on the input image to obtain its frequency domain representation F(I)(u,v), which is then processed by an ideal low-pass filter. (u,v) filtering preserves low-frequency structural information to compensate for the lack of spatial domain texture. ,in Cutoff frequency of (u,v) The complexity of the sea surface texture is adaptively set, with a typical value ranging from 5% to 15% of the image diagonal length; in this embodiment, 10% is used. The frequency domain enhanced feature map is obtained by taking the modulus of the inverse Fourier transform. Meanwhile, the spatial attention branch is based on the intermediate layer feature map. An attention map is generated to guide the network to focus on salient areas such as the horizon and ship outline, which have a critical impact on stitching accuracy. in, This represents the attention map, where GAP and GMP represent global average pooling and global max pooling, respectively, and ⊕ indicates channel concatenation. It is a 3×3 convolutional layer, and σ is the Sigmoid activation function.

[0057] The sea surface texture perception unit analyzes the local texture complexity. The variance is used to adaptively adjust the receptive field size for feature extraction. In weakly textured regions ( It employs a 7×7 large receptive field convolution kernel to capture a wider range of contextual information; in medium texture regions ( It employs a 5×5 medium receptive field convolution kernel, preserving local details while also considering a certain range of contextual aggregation; in areas with strong texture ( A 3×3 small receptive field convolutional kernel is used to preserve detailed features. Threshold , Automatically set based on statistical values ​​of the sea surface area, among which Take the 25th percentile of the texture variance in the sea surface area. Take the 75th percentile of the texture variance of the sea surface area.

[0058] Final fused feature map , where α and β are learnable balance coefficients, initially set to 0.5 and 0.3 respectively, and automatically updated during network training. Feature dimension d=128.

[0059] To verify the effectiveness of the sea surface texture perception enhancement mechanism, Figure 3 This is a schematic diagram of the stitching effect without using the specific processing mechanism for ocean surface textures. As a baseline, it shows the failure of the general network in the scene with weak ocean surface texture. Figure 4 This is a diagram illustrating the stitching effect achieved using a specific ocean surface texture processing mechanism. (Comparison) Figure 3As can be seen, after introducing frequency-spatial domain hybrid feature extraction and the sea surface scene attention module, the success rate of stitching weak texture areas on the sea surface is significantly improved, and key geometric structures such as the sea-line and ship outlines remain clear and continuous. Using fused feature maps and Calculate the feature similarity matrix S, where the matrix element S(i,j) represents... The i-th feature and Similarity of the j-th feature: The locations of corresponding points are estimated using a soft argmax operation to establish the correspondence between images. The RANSAC algorithm is then used to remove outlier matching points and estimate the initial homography matrix. The number of iterations N in the RANSAC algorithm is determined by the following formula: In the formula, p is the confidence probability, which is taken as 0.99, and w is the proportion of interior points.

[0060] Step 3.2: Motion-Adaptive Coarse Image Alignment To address the dynamic parallax problem caused by ship swaying, this embodiment introduces a motion pre-compensation mechanism based on IMU and optical flow, as well as a dynamic ROI mask.

[0061] Let the graph at time t be The motion field between adjacent frames is obtained through an optical flow estimation network (using FlowNet2.0). Simultaneously, roll and pitch angular velocities are acquired from the ship's IMU sensors. Calculate the motion field predicted by the IMU: Where Δt is the inter-frame time interval and f is the camera focal length (in pixels).

[0062] The fusion weight γ is adaptively adjusted based on the confidence level of optical flow estimation. ,in The matching confidence score for optical flow estimation is calculated from the forward-backward consistency error output by the optical flow network. The noise variance for IMU measurements is set to 0.01. The final motion field is the fused result: Motion-compensated image Obtained through inverse deformation: In addition, construct a dynamic ROI mask. The search range for feature matching is adaptively adjusted based on the motion amplitude detected by the IMU. Motion amplitude threshold. Determine using the following formula: =β· ,in represents the standard deviation of the combined amplitude of the roll and pitch angular velocities measured by the IMU within the current time window, and β is a scaling factor of 1.5. When the local motion amplitude || (x,y)||> When the motion is smooth, the search window is expanded to twice its original size, and the mask value is set to 1; when the motion is smooth (|| (x,y)||≤ When performing a search, the default search window is maintained, and the mask value is set to 1; the mask value for non-overlapping regions is set to 0. When calculating the feature similarity matrix, only M... ROI Matching is performed on regions with a value of 1 to suppress interference from non-overlapping regions and regions with intense motion.

[0063] The homography matrix is ​​re-estimated based on the correspondence after motion compensation. = ·ΔH, where ΔH is the incremental transform based on motion compensation. Let the coordinates of a pixel in the compensated image be (x', y'), and its coordinates in the reference image be... The coordinates of the corresponding point are (x1, y1), and they satisfy the following projection transformation relationship: [x1,y1,1 ·[x',y',1 That is, there exists a non-zero scaling factor λ such that [x1,y1,1] =λ· ·[x',y',1 Based on the above projection transformation relationship, all pixels of the compensated image are transformed to... In the coordinate system, coarse alignment of the image is achieved.

[0064] Step 3.3: Fine-grained reconstruction based on geometric consistency awareness This embodiment employs an encoder-decoder structure (similar to U-Net) to perform fine fusion on the coarsely aligned image, eliminating stitching seams and ghosting. The input to the fine reconstruction module is the coarsely aligned image pair and the mask of the overlapping region, and the output is the finely fused stitched image.

[0065] The training loss of the reconstruction module introduces a multidimensional consistency constraint: The weights for each item were determined through a grid search on the validation set and set as follows: =0.1, =0.5, =0.2, =0.3.

[0066] in, For pixel-level reconstruction loss, L1 loss is used; To perceive the loss, multi-layer feature differences extracted by a pre-trained VGG-19 network are used.

[0067] For geometric consistency loss, the geometric coordinate mapping deviation of corresponding pixels in the overlapping region is constrained.

[0068] To address the boundary consistency loss, the continuity of pixel values ​​in the reconstructed image at the stitching boundary is constrained. The difference in pixel values ​​and gradient difference on both sides of the stitching boundary are calculated to provide a smooth transition boundary for the subsequent filling module.

[0069] To mitigate temporal consistency loss, temporal consistency constraints between adjacent frames are utilized to reduce splicing jitter in the video sequence: In the formula, The static region mask is obtained through inter-frame difference and dilation operations, which constrains the splicing consistency of the static region between consecutive frames.

[0070] After fusion, the content filling and stitching processes are collaboratively optimized. This is based on the output of the stitching network and the invalid region mask. Define the splicing-filling collaborative loss: in To stitch together the stitched images output by the network, The initial result for filling the PDE. For the final merged image, This is the gradient continuity weight, typically 0.5.

[0071] In some embodiments, to enhance the temporal stability of video stitching, a long-range temporal consistency loss is added to the loss function of the fine reconstruction module. (Except for adjacent frame constraints) In addition, long-range constraints are introduced: in, λ This is the attenuation factor, with a value of 0.8; K This is the number of frames to be traced back, and its value is 5. and Representing the t-th frame and the t-th frame respectively. The stitched image output from k frames processed by the stitching network. Attenuation factor. The constraint weights of more distant historical frames on the current frame vary with the frame interval. k It decreases exponentially as it increases.

[0072] This long-range constraint is related to the temporal consistency loss of adjacent frames during the loss calculation stage. The total timing consistency loss, jointly applied to the output of the fine reconstruction module, is: in, This is the balancing weight for the long-range constraint, with a value of 0.1.

[0073] By introducing the aforementioned long-range constraints, the stitching parameters not only remain consistent between adjacent frames, but also remain stable over medium to long time scales (covering 5 consecutive frames), further differentiating it from existing schemes that rely solely on adjacent frame constraints and enhancing the temporal coherence of panoramic video.

[0074] In some embodiments, a cross-camera exposure consistency compensation submodule is added during the image stitching stage. Because lighting conditions may differ in different directions on a ship (e.g., one side is in direct sunlight while the other is in shadow), there can be significant exposure differences between images from adjacent cameras. This submodule is executed after correction is complete and before stitching feature extraction.

[0075] Let the overlapping region of the two images to be stitched be O. Calculate the cumulative distribution function (CDF) of the pixel values ​​in the overlapping region of the two images. Then, use histogram matching to... CDF mapping to The CDF is used to obtain the exposure compensation function. .Will Applied to The global pixel value ensures consistent exposure between the two images in the overlapping area, reducing color inconsistencies during subsequent stitching and fusion.

[0076] After stitching, the generated initial panoramic image Output to step S4.

[0077] Step S4: Semantically Guided Boundary Joint Filling This step fills in the black invalid areas in the initial panoramic image output in step S3. Figure 5 The image shows the pre-fill panoramic view, illustrating the black invalid areas and their distribution after stitching. This step is performed in the semantically guided boundary joint filling module described in Example 1. The following details the working process of the semantically guided boundary joint filling module, which, unlike existing technologies that treat filling as an independent post-processing step, establishes a joint optimization mechanism for filling and stitching boundaries.

[0078] Step 4.1: Invalid Region Mask and Semantic Parsing First, create an invalid region mask. The locations of invalid black pixels that need to be filled are marked. A semantic segmentation network (using the DeepLabV3+ architecture) is used to perform semantic parsing on the panoramic image, identifying categories such as sky, sea, and ship hull, and generating a semantic mask. The semantic segmentation network is trained using transfer learning: a DeepLabV3+ model pre-trained on the Cityscapes dataset is used as the initial weights, and then fine-tuned on an annotated marine panoramic dataset containing 1000 panoramic images of ships, labeled in four categories: sky, seawater, ship hull, and distant land. Cross-entropy loss is used during fine-tuning, with a learning rate of 1×10⁻⁶. -4 , 50 rounds of training.

[0079] Step 4.2: Anisotropic PDE Filling Strategy An anisotropic PDE filling strategy is adopted for different semantic regions. For regions marked as seawater, an anisotropic diffusion equation is used, with a larger diffusion coefficient in the horizontal direction (set to 1.0) and a smaller one in the vertical direction (set to 0.3) to maintain the horizontal continuity of the water ripples. For regions marked as sky, isotropic fast and smooth diffusion is used, with a diffusion coefficient of 1.0 in all directions, and the color is constrained to converge towards the average hue of the sky to maintain the consistency of the color gradient.

[0080] The energy functional of the filled region is defined as follows: In the formula, Ω represents the filled region. Ω represents the boundary of the filled region. Sem(u) represents the effective pixel values ​​on the boundary, and Sem(u) represents the semantic features of the filling content. For the reference semantic features (sky or sea), η and ζ are weighting coefficients. This represents the balancing weighting coefficient between the data fidelity and regularization terms. The first term is the data fidelity and regularization term, the second term is the boundary consistency constraint, ensuring the luminosity continuity between the filled region and the effective image boundary, and the third term is the semantic guidance term, ensuring the semantic consistency between the filled content and the surrounding environment.

[0081] The parameters in the energy functional were determined through validation set testing and were set to λ=1 (indicating that the fidelity term and the regularization term have equal weights), η=0.8, and ζ=0.5.

[0082] Step 4.3: Iterative solution using the finite difference method The corresponding partial differential equations are obtained from the Euler-Lagrange equations and solved using the finite difference method. Let the discretized image grid be... The iterative update formula is: Iteratively update the pixel values ​​of the filled region until convergence. Convergence is determined by reaching the maximum number of iterations, or by the change between two consecutive iterations being less than a preset threshold. : By minimizing collaborative loss The stitching and fusion weights and filling content are iteratively optimized until the luminance and gradient differences in the boundary regions are all less than a preset threshold, achieving a seamless transition between the filled region and the effective image. The final panoramic image is shown in the appendix. Figure 6 .

[0083] Step 4.4: Stitching-Infill Co-optimization The padding process is integrated with the output of the splicing network, and a smooth gradient transition is ensured through the splicing-padding cooperative loss. The cooperative loss is defined as: In the formula, Invalid region mask. Fill the results for PDE. For the final merged image, The gradient continuity weight is set to 0.5.

[0084] The first term of the loss constraint constrains the continuity of pixel values ​​on both sides of the fill boundary, and the second term constrains the continuity of gradients on both sides of the fill boundary. By minimizing this collaborative loss, the stitching and fusion weights and the fill content are iteratively optimized, so that the fill result inversely affects the fusion weights of the stitching boundary, until the luminance difference and gradient difference in the boundary region are both less than a preset threshold, achieving a seamless transition between the fill region and the effective image.

[0085] Step 4.5: Alternative Implementation of the Fill Module As an alternative to PDE padding, a padding branch based on a Generative Adversarial Network (GAN) can be added. This branch takes a panoramic semantic mask as conditional input, with a generator responsible for generating the padding content and a discriminator simultaneously judging the realism of the padding content and the continuity of the transition with the stitching boundaries. This GAN branch also... Loss is coupled with the splicing network.

[0086] After the infill is complete, a seamless panoramic image will be generated. Output to step S5.

[0087] Step S5: Interactive panoramic display and intelligent early warning This step displays and issues a warning for the seamless panoramic image output in step S4. This step is performed in the interactive panoramic display and intelligent warning module described in Example 1.

[0088] Step 5.1: Interactive display function The processed panoramic image is presented to the user through the display interface. The panoramic image rendering uses equidistant cylindrical projection. The display interface supports the following interactive functions: The zoom function allows users to zoom in and out of the panoramic image using the mouse wheel or touchscreen gestures, with zoom ratios ranging from 0.5x to 5x.

[0089] With view rotation, users can freely switch between different directions of view around the ship by dragging. View rotation uses a spherical projection method, supporting 360-degree full-range rotation in the horizontal direction and pitch from -45 degrees to +45 degrees in the vertical direction.

[0090] The system supports switching between bird's-eye view, panoramic view, and bridge view. The bird's-eye view is a view from directly above the ship; the panoramic view unfolds in a ring around the ship; and the bridge view simulates the actual observation direction from the bridge.

[0091] The motion-adaptive viewing angle automatically adjusts the default observation angle based on the ship's attitude and speed detected by the IMU. When the speed exceeds the preset value, the view in the bow direction is automatically magnified.

[0092] Step 5.2: Three-level early warning triggering mechanism The system integrates real-time data from the ship's AIS, radar, and GPS systems to calculate the closest encounter time (TCPA) and closest encounter distance (DCPA) between the ship and the target, combined with relative speed. Define a three-level early warning system. Typical early warning threshold parameters are set as follows: D1=1000m, D2=500m, D3=200m, T1=10min, T2=5min, v1=15kn. =0.8. The above threshold can be adaptively adjusted according to ship type, speed and sea state.

[0093] The formula for calculating the collision time TCPA is: , where θ is the angle between the target's relative motion direction and the ship's direction.

[0094] Level 1 Warning (Yellow "Attention" Level): Triggering condition is that the target distance D is ≥ 1000m and D > 500m, or the target relative speed... More than 15 sections. The display method is to mark the location of obstacles with a yellow semi-transparent rectangle, and the target ID and distance are marked with text. The border flashes at a frequency of 1Hz, and the linked action is to record logs.

[0095] Level 2 Warning (Orange "Warning" Level): Triggered by 500m ≥ D > 200m, or collision time TCPA less than 10 minutes. Displayed as an orange diamond marker with the fill color flashing at 2Hz, accompanied by a short audible alert and a pop-up notification suggesting attention.

[0096] Level 3 Warning (Red "Danger" Level): Triggering conditions are D≤200m, or TCPA less than 5 minutes, or collision risk probability. >0.8. The display method is a red octagonal danger mark, with the fill color flashing rapidly at a frequency of 4Hz, accompanied by a full-screen red border pulse effect, a continuous alarm sound, and a red light strip at the edge of the screen. The linkage action is to automatically link with the ship navigation control system and push to ECDIS.

[0097] The warning display uses color coding, flashing frequency, and graphic style overlay to highlight the location of obstacles on the panoramic image.

[0098] In one specific embodiment, the collision risk probability The parameters were calculated using a weighted fusion model based on the membership degree of multiple navigation parameters. First, the nearest encounter distance (DCPA), nearest encounter time (TCPA), and current target distance were calculated separately. and relative velocity Corresponding risk membership degree: In the formula, To preset a safe encounter distance, the setting is adaptively set according to the vessel type and sea conditions, with a typical value of 0.5 nautical miles (approximately 926 m). The preset safety time threshold is typically set to 30 minutes. This represents the distance from the risk half-life, typically taken as 500m. The relative speed risk benchmark value is typically taken as 15 sections; This is the slope coefficient of the velocity membership curve, typically taken as 0.5.

[0099] Subsequently, the four membership degrees mentioned above are combined according to their weights to obtain the collision risk probability: In the formula, , , , For each weight coefficient, satisfying + + + =1, adaptively adjusted according to ship type, speed, and sea state; under typical configuration =0.35, =0.35, =0.20, =0.10. When , , or When the membership degree increases significantly, it approaches 1, thus driving... It approaches 1; conversely, when all parameters are within a safe range, Approaching 0.

[0100] Step 5.3: Multi-source data fusion The system simultaneously receives visual detection results (target category, pixel position), AIS target information (MMSI, ship name, speed, heading, position) and radar ranging data (target distance, azimuth), and performs data fusion through Kalman filtering.

[0101] Fusion state vector This includes the target position and velocity components. Observation vector It integrates AIS location, visual detection location, and radar polar coordinate measurements. The process model adopts a uniform motion model, with the state transition matrix F=[1,0,Δt,0;0,1,0,Δt;0,0,1,0;0,0,0,1]. The observation matrix H is constructed based on measurements from different sensors, mapping the state vector to the observation space of each sensor. The filter update period is 1 second.

[0102] By iteratively updating the target state estimate through Kalman filtering, the accuracy of obstacle localization and the reliability of motion prediction are improved, enabling a comprehensive collision risk assessment based on visual detection, AIS target information, and radar ranging data.

[0103] Through steps S1 to S5, this method achieves a complete processing flow from original fisheye image acquisition to final panoramic display and intelligent early warning. Specifically, steps S2 and S3 establish a bidirectional feedback mechanism through joint optimization: the registration error of the stitching network in step S3 is backpropagated to the correction parameters in step S2 via feedback consistency loss, dynamically adjusting the distortion correction result to ensure the corrected image meets the stitching registration requirements to the greatest extent possible. Steps S3 and S4 are coupled through a stitching-filling collaborative loss, allowing the filling process to inversely influence the fusion weights of the stitching boundary, achieving a seamless boundary transition. The modules in steps S1 to S5 undergo end-to-end collaborative training using the multi-task joint learning framework described in Example 3.

[0104] Example 3: Multi-task Joint Learning Framework The modules mentioned above are trained end-to-end through a multi-task joint learning framework, which integrates the three major tasks of distortion correction, image stitching and ship target detection into a unified training system. The system’s overall robustness in complex maritime scenarios is improved through knowledge sharing between tasks.

[0105] The joint loss function for multiple tasks is defined as: In the formula, , , For dynamic weights based on training epoch t, ​​a gradient normalization strategy is used to automatically adjust them: ,in The learning rate scaling factor for task i; The balancing weight coefficient for consistency loss between tasks is a fixed weight with a value ranging from 0.05 to 0.2, determined through validation set debugging. The detection loss includes classification loss and bounding box regression loss.

[0106] The consistency loss between tasks is defined as: in, To construct the feature representation of the network at position (i,j), To detect the feature representation of the network at the corresponding location, This is a mask for the overlapping region.

[0107] This consistency loss ensures that the stitching network and the detection network maintain a consistent spatial understanding of the overlapping region. During the stitching process, this constraint enables the network to simultaneously serve two objectives in feature extraction within the overlapping region: stitching alignment and ship detection. When the detection network identifies a ship target in the overlapping region, the consistency loss prompts the stitching network to maintain higher feature fidelity in that region, allowing the stitching process to adaptively enhance its focus on navigation safety-critical areas.

[0108] The total number of training rounds is set to 200, and the initial learning rate is 1×10. - ³, using cosine annealing scheduling. Through multi-task joint training, the calibration module, stitching network, and detection network mutually reinforce each other, improving the overall robustness in complex maritime scenarios.

[0109] In summary, this invention, through the collaborative work of the aforementioned modules, achieves a complete processing flow from raw fisheye image acquisition to final panoramic display. The system employs an end-to-end deep learning architecture, integrating innovative mechanisms such as distortion correction and stitching joint optimization, enhanced sea surface texture perception, temporal motion compensation, semantically guided boundary joint filling, and intelligent early warning through multi-source data fusion. This automates the process from image input to panoramic output, significantly improving processing efficiency and stitching quality, thus meeting the needs of real-time ship monitoring.

[0110] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features, and such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for combined distortion correction and image stitching optimization of ship panoramic vision, characterized in that, Includes the following steps: Step S1, Multi-source data acquisition: Raw fisheye images are acquired through fisheye cameras deployed around the ship, ship attitude angular velocity data are acquired simultaneously through the inertial measurement unit, ship target information is acquired through the automatic identification system, and ranging data is acquired through the shipborne radar. The data from each sensor are synchronized through a time protocol. Step S2, Adaptive Joint Correction: Establish a distortion model containing the coordinates of the distortion center to perform coordinate mapping correction on the original fisheye image. Use the distortion correction parameters as learnable variables and perform end-to-end joint training with the image stitching network. By jointly optimizing the feedback consistency loss in the objective function, the registration error of the image stitching network in the overlapping area is directly backpropagated and the distortion correction parameters are dynamically adjusted to output a distortion-free image. Step S3, Image stitching: Extract frequency domain and spatial domain hybrid features from the distortion-free image, estimate the optical motion field based on the distortion-free images of adjacent frames, use the attitude angular velocity data and the optical motion field to adaptively fuse and construct the motion field for inter-frame motion pre-compensation, complete the image coarse alignment, and use joint constraints including geometric consistency, boundary consistency and temporal consistency loss to perform fine reconstruction network fusion, and output the initial panoramic image. Step S4, joint boundary filling: generate a filling mask for the invalid region of the initial panoramic image and perform semantic parsing. Based on the semantic category of the region, a partial differential equation filling strategy is used to solve the invalid region at the pixel level. At the same time, minimize the joint loss of stitching and filling, so that the filling result of the partial differential equation has a reverse influence on the fusion weight of the stitching boundary, and output a seamless panoramic image. Step S5, Panoramic Display and Intelligent Early Warning: Display the seamless panoramic image, perform multi-source data state fusion on the visual detection position, the ship target information and the ranging data through Kalman filtering, establish a graded early warning triggering mechanism based on the nearest encounter time, the nearest encounter distance and the relative speed, and execute graded early warning annotation and underlying linkage actions on the seamless panoramic image.

2. The ship panoramic vision method jointly optimized by distortion correction and image stitching according to claim 1, characterized in that, The distortion model containing the coordinates of the distortion center established in step S2 is represented by the following algebraic mapping relationship: ; ; in,( , (x, y) represents the pixel coordinates in the original distorted image, and (x, y) represents the coordinates in the corrected, distortion-free image. , () represents the learnable distortion center coordinate variable based on the mechanical physical offset setting. , , The radial distortion coefficient is... , Let r be the tangential distortion coefficient, and r be the radial distance satisfying r² = (x - 1 / 2)^2. )²+(y- ).

3. The ship panoramic vision method jointly optimized by distortion correction and image stitching according to claim 1, characterized in that, The end-to-end joint training in step S2 is based on a joint optimization objective function. The update calculation is performed, and the joint optimization objective function is defined as follows: ; in, To measure the stitching loss between network reprojection error and structural similarity, The correction loss is used to measure the degree of straightness preservation versus distortion removal. The feedback consistency loss is precisely defined as the L2 norm of the gradient of the partial derivative of the L1 reprojection registration error extracted by the image stitching network in the overlapping region with respect to all distortion correction parameters. and These are the fixed weighting coefficients used to balance the gradients.

4. The ship panoramic vision method jointly optimized by distortion correction and image stitching according to claim 1, characterized in that, The process of extracting the frequency domain and spatial domain hybrid features in step S3 specifically includes: A two-dimensional discrete Fourier transform is performed on the input image. An ideal low-pass filter, set to a cutoff frequency of 5% to 15% of the image diagonal length, is used to filter out high-frequency noise while preserving low-frequency structural information. The image is then subjected to an inverse Fourier transform to generate a frequency domain enhanced feature map. In parallel, the intermediate layer feature maps are used to generate spatial attention maps by connecting activation functions through global average pooling and global max pooling operations, guiding the feature extraction network to focus on the sea-line and ship outline regions. The variance value of local image texture complexity is analyzed and extracted. The variance value is determined and the receptive field switching is performed: in weak texture regions with variance values ​​below the first threshold, a first-size receptive field convolutional kernel is used to obtain global context features; in medium texture regions with variance values ​​between the first and second thresholds, a second-size receptive field convolutional kernel is used to balance local details and context aggregation; in strong texture regions with variance values ​​greater than or equal to the second threshold, a third-size receptive field convolutional kernel is used to retain high-frequency detail features to generate a spatial domain feature map, wherein the first size is greater than the second size and the second size is greater than the third size. The spatial domain feature map is weighted and modulated using the spatial attention map, and the modulated spatial domain feature map is then merged with the frequency domain enhanced feature map to generate a fused feature map for subsequent calculation of the feature similarity matrix.

5. The ship panoramic vision method jointly optimized by distortion correction and image stitching according to claim 1, characterized in that, The process of dynamically adaptively fusing and constructing a motion field to perform inter-frame motion pre-compensation in step S3 specifically includes: Two-dimensional optical motion field vectors between adjacent frames are obtained through an independent optical flow estimation network. Simultaneously, based on the roll and pitch angular velocities measured by the inertial measurement unit, the camera focal length constant extracted from calibration and the inter-frame time interval are combined to calculate and project the inertial predicted motion field vectors. Based on the forward and backward consistency error output by the optical flow estimation network, the optical matching confidence function is calculated. Using this confidence function and the preset ratio of sensor physical measurement noise variance, the fusion weight coefficients for balancing the sensors are adaptively adjusted and output. The optical motion field two-dimensional vector and the inertial predicted motion field vector are linearly weighted and fused using the fusion weight coefficient to generate the final fused motion field, and an inverse deformation operation is performed on the original image to obtain a pre-compensated image; The standard deviation of the combined amplitude of roll and pitch angular velocities within the current time window is calculated in real time to dynamically determine the motion amplitude threshold. When the amplitude of the local fused motion field exceeds the threshold, the spatial size of the search window corresponding to the local area is expanded to twice the original preset size. Within this area, the matrix bits of the dynamic region of interest mask are effectively set. In subsequent homography matrix estimation, feature matching calculation is only performed on the regions where the mask is effectively set.

6. The ship panoramic vision method jointly optimized by distortion correction and image stitching according to claim 1, characterized in that, The training loss relied upon in step S3, which uses joint constraints to perform fine reconstruction of the coarsely aligned image network fusion, is... Defined as the sum of the following five items: ; in, The L1 pixel-level reconstruction loss is used to measure the absolute pixel difference. A perceptual loss constructed to extract deep, multi-layer features to represent differences using a pre-trained visual model; The geometric consistency loss is a constraint that forcibly minimizes the deviation of the geometric space coordinate mapping of corresponding feature points in the overlapping region. To forcibly constrain the boundary consistency loss so that the absolute difference between pixels on both sides of the splicing boundary seam and the spatial gradient difference are in a continuous state; Constrain the temporal consistency loss of adjacent consecutive frame stitched pixels within the static region mask obtained by inter-frame difference operation in this temporal dimension; In addition, the joint constraint also includes a long-range temporal consistency loss applied to the output of the fine reconstruction network. This long-range temporal consistency loss uses a preset decay factor to perform a time-dimensional decay weighted exponential constraint calculation on the output results covering multiple backtracking historical frames and the current frame.

7. The ship panoramic vision method jointly optimized by distortion correction and image stitching according to claim 1, characterized in that, The semantically guided anisotropic partial differential equation filling strategy and end-to-end coupled optimization process in step S4 specifically include: For invalid region pixels marked as seawater by the semantic segmentation network, an anisotropic diffusion equation containing a diffusion coefficient tensor is called for numerical evolution, where the coefficient value of the tensor in the horizontal direction is greater than the coefficient value in the vertical direction; for invalid region pixels marked as sky, an isotropic fast smooth diffusion equation with absolutely equal diffusion coefficients in all directions is called, and the color gradient is constrained to approximate the average hue of the panoramic sky reference area using color Euclidean distance. The filling strategy is based on the finite difference method to iteratively solve a pre-constructed energy functional, which is composed of three core terms: a basic diffusion term that controls its smoothness and includes a data fidelity term and a regularization term; a boundary consistency constraint term that ensures that the filled pixels approximate the effective boundary pixel values ​​to maintain the continuity of boundary luminance; and a semantic guidance term that ensures that the deep semantic features of the filled content are consistent with the high-dimensional mapping of the seawater or sky semantic category. The stitching and filling collaborative loss is constructed as an objective function consisting of a weighted combination of a photometric continuous term and a gradient continuous term. The photometric continuous term measures the L2 norm difference between the fused image and the gaps between the masked filling and stitching regions, while the gradient continuous term measures the L2 norm difference between the gradient of the fused image boundary space and the gradient of the network output features. During joint backpropagation of the network, this collaborative loss is used to directly iteratively update the hyperparameter tensor that determines the stitching and fusion weights and the initial conditions of the PDE, so that the optical and gradient jumps in the boundary region are less than a minimum preset residual at convergence to achieve a seamless transition.

8. The ship panoramic vision method jointly optimized by distortion correction and image stitching according to claim 1, characterized in that, In step S5, the three-level graded early warning triggering mechanism based on multi-source data and nautical parameters is specifically executed with the following judgment and response sequence: The target distance D, collision time TCPA, and relative velocity were calculated. and collision risk probability variables; When 1000m ≥ D > 500m, or the relative velocity of the target is detected. When any condition exceeds 15, the processor's internal clock triggers a first-level warning state, driving the display module to surround the target feature area with a preset yellow semi-transparent wireframe, and superimposing the target's unique identifier and distance value. The wireframe rendering continues to flash alternately according to the first low-frequency cycle, while the event identifier is persistently written to the underlying disk log file. When either 500m≥D>200m is met, or the collision time TCPA is less than the 10-minute calculation threshold, the state machine jumps to the second-level warning state, drives the display primitive to render an orange diamond-shaped marker array, forces the interface to redraw at a frequency higher than the first low-frequency cycle to form a fast-flickering visual feedback, and simultaneously displays the dynamic countdown TCPA value in a text control on its side, and calls the operating system's audio driver module to output periodic short audio pulse interruptions; When the distance D ≤ 200m, or the TCPA is less than 5 minutes, or the calculated collision risk probability is met. When any core condition >0.8 is met, the system forcibly switches to the highest level three early warning and interception state, driving the front-end GPU unit to render a red octagonal danger mask with the highest priority, flashing it at the highest frequency cycle, and injecting a global full-screen pulsed red warning light filter into the video memory. At the same time, a continuous locking system alarm sound is triggered, and the underlying communication interface immediately sends the automatic avoidance control word to the ship navigation underlying electronic control system through the bus protocol and forcibly synchronizes the display of the status frame to the electronic chart information and information system terminal.

9. The ship panoramic vision method jointly optimized by distortion correction and image stitching according to claim 1, characterized in that, In steps S1 to S5, the underlying network parameter updates are uniformly performed by an end-to-end collaborative training scheduling method using a multi-task joint learning framework that includes a multi-task joint loss function; the multi-task joint loss function is defined as follows: ; in To correct the error terms generated by the parameter layer, The metrics generated for splicing the network, For the regression and classification network loss of ship target detection bounding boxes that are independent of the splicing network but share the backbone layer features, The equilibrium coefficient is fixed as a constant. , and In each training round t, instead of a constant tensor, a gradient normalization feedback strategy is used to dynamically reallocate computing resources and step size in the backend. Its dynamic value is strictly proportional to the reciprocal of the gradient of the objective function of the corresponding independent task to suppress the gradient annihilation effect caused by the difference in magnitude. As a loss mechanism for inter-task consistency, it removes overlapping data using a mask matrix, forcing the deep feature tensors of the image stitching network to be consistent. Intermediate feature tensor of the object detection network The Euclidean distance within this overlapping subset approaches zero when backpropagation converges, thereby driving the stitching network to be strongly correlated during feature extraction and adaptively enhance key safety navigation zones containing the ship's geometric profile.

10. A ship panoramic vision system jointly optimized by distortion correction and image stitching, characterized in that, Includes a computing and sensing unit configured to perform the method as described in any one of claims 1 to 9, specifically including: The multi-source image acquisition module is used to acquire raw fisheye images and various sensor data through deployed fisheye cameras, inertial measurement units, automatic identification systems for ships, and radar equipment, and to achieve time synchronization. An adaptive joint correction module, whose image input end is connected to the image output end of the multi-source image acquisition module, is used to establish a distortion model containing the coordinates of the distortion center to correct the original fisheye image. The distortion correction parameters are jointly trained with the image stitching network, and the parameters are dynamically adjusted using feedback consistency loss. The distortion-free image is output through its output end. The image stitching network module has its image input end connected to the output end of the adaptive joint correction module, and its data input end connected to the data output end of the inertial measurement unit of the multi-source image acquisition module. It is used to extract the mixed features of frequency domain and spatial domain, use the received attitude angular velocity data and optical flow adaptive fusion to perform inter-frame motion pre-compensation and coarse alignment, and perform fine reconstruction under multi-dimensional joint constraints. The initial panoramic image is output through its output end. The boundary joint filling module, whose input end is connected to the output end of the image stitching network module, is used to perform pixel-level iterative filling of invalid regions based on semantic parsing results and partial differential equation filling strategy. It also achieves boundary optimization by inversely influencing the fusion weights through stitching and filling collaborative loss, and outputs a seamless panoramic image through its output end. The panoramic display and intelligent early warning module has its image input end connected to the output end of the boundary joint filling module, and its data input end connected to the data output end of the ship automatic identification system and radar equipment of the multi-source image acquisition module. It is used to display the seamless panoramic image, perform state fusion on the received multi-source data based on Kalman filtering, and calculate and execute hierarchical early warning visual annotation and underlying physical linkage according to nautical parameters.