Systems and methods for underwater imagery enhancement

The DUCT algorithm and AI-based enhancement module address temporal inconsistencies in underwater imagery by ensuring color consistency across frames, enhancing clarity and accuracy for improved human and automated applications.

WO2025224087A1PCT designated stage Publication Date: 2025-10-30FNV IP BV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
PCT/EP2025/060898
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-26
Filing Date
2025-04-22
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing underwater image enhancement technologies suffer from temporal inconsistencies and flickering in sequences of enhanced images, which impair human viewing experience and compromise the effectiveness of automated processing applications, particularly in autonomous underwater vehicles and marine environment monitoring.

Method used

The system employs a Dynamic Underwater Color Transfer (DUCT) algorithm and AI-based image enhancement module to correct underwater image distortions, ensuring color consistency across frames using a blending process that adjusts color characteristics dynamically, leveraging Generative Adversarial Networks (GANs) and other neural networks trained on diverse underwater imagery datasets.

Benefits of technology

The system provides high-quality, temporally consistent underwater imagery, enhancing visual clarity, color accuracy, and operational efficiency for applications like navigation, marine life identification, and environmental monitoring, improving human perception and automated decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025060898_30102025_PF_FP_ABST
    Figure EP2025060898_30102025_PF_FP_ABST
Patent Text Reader

Abstract

A computer implemented method for enhancing underwater imagery is disclosed. A sequence of raw image frames including a current raw image frame is received. The raw image frames are processed with an Artificial Intelligence (AI) image enhancement model to provide enhanced image frames including a current enhanced image frame. Image quality scores are determined for the enhanced image frames. Based on the image quality scores, one of the enhanced image frames is set as an enhanced image template. Color characteristics of the current enhanced image frame are adjusted using a color transfer function for applying color characteristics of the enhanced image template to those of the current enhanced image frame. Unlocking insights from Geo-Data, the present invention further relates to improvements in sustainability and environmental developments: together we create a safe and liveable world.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR UNDERWATER IMAGERY ENHANCEMENTFIELD OF THE INVENTION

[0001] The present disclosure generally relates to the field of underwater image and video (imagery) processing. More particularly, the invention relates to methods and systems for enhancing underwater imagery and improved consistency from frame to frame of the enhanced imagery. Unlocking insights from Geo-Data, the present invention further relates to improvements in sustainability and environmental developments: together we create a safe and liveable world.BACKGROUND OF THE INVENTION

[0002] Underwater image enhancement has become an increasingly important area of research in computer vision, driven by the expanding interest in exploring the underwater world. Remotely operated vehicles (ROVs) and autonomous underwater vehicles (AUVs) are extensively utilized in underwater research, where clear visual conditions are essential for the visual systems onboard these vehicles. However, achieving clarity in underwater environments is challenging due to factors such as color degradation, where different light wavelengths are absorbed unevenly (with red wavelengths diminishing most at greater depths), water turbidity causing haziness from scattered light by numerous small particles, and poor illumination at deeper levels due to light absorption. These conditions can make the maneuvering of ROVs and AUVs under water quite risky. Image and video enhancement methods for underwater conditions aim to address these problems by improving the quality of the footage, leading to enhanced visibility.

[0003] Deep learning and artificial intelligence (Al) systems have been proposed for enhancing visual data from underwater imaging. Al-based techniques have emerged as powerful tools to address the challenges posed by the underwater environment, including low light conditions, varying water clarity, and color distortion due to light absorption and scattering. These technologies may employ complex neural network architectures to learn fromlarge datasets of underwater images, enabling them to automatically apply corrections that enhance visual clarity, restore natural colors, and improve overall image quality without manual intervention.

[0004] Some deep learning algorithms for underwater image enhancement may analyze input imagery to identify and mitigate issues such as haze, blur, and unnatural color casts. By adjusting parameters like contrast, brightness, and hue based on learned models, these systems can produce images that more closely resemble how the scene would appear in optimal lighting conditions above water. The training process may involve exposing the neural network to a wide range of underwater images and their enhanced counterparts, allowing the model to understand and replicate the transformations needed to correct various underwater imaging artifacts.

[0005] Despite the significant advancements offered by Al-based underwater image enhancement, there remain limitations. For example, temporal inconsistencies or flickering may be introduced in sequences of enhanced images, particularly in video footage. This phenomenon may occur when consecutive frames are processed independently, leading to variations in color correction and enhancement across frames. Temporal inconsistencies and flickering in sequences of enhanced images can impair a human viewing experience. Inconsistencies across frames of enhanced images can also compromise the effectiveness of visual recognition applications, which rely on consistent imagery for accurate object and feature detection. Furthermore, such inconsistencies can disrupt the operational algorithms of autonomous underwater vehicles (AUVs), where precise visual cues may be used for navigation and task execution. Additionally, in applications like automated monitoring of marine environments and habitats, fluctuating visual characteristics can lead to unreliable data collection and analysis, undermining efforts in conservation and scientific research.

[0006] There is thus a need for systems and methods for improved temporal consistency in enhanced underwater imagery for both human viewing and automated processing applications.BRIEF SUMMARY OF THE INVENTION

[0007] The above mentioned and other features and advantages of the disclosure will be best understood from the following description referring to the attached drawings. In thedrawings, like reference numerals denote identical parts or parts performing an identical or comparable function or operation.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to describe the manner in which the above-recited and other advantages and features of the disclosure can be obtained, a more particular description of the principles briefly described above will be rendered by reference to specific embodiments thereof which are illustrated in the appended drawings. Understanding that these drawings depict only exemplary embodiments of the disclosure and are therefore not to be considered to be limiting of its scope, the principles herein are described and explained with additional specificity and detail through the use of the accompanying drawings in which:

[0009] FIG. l is a schematic diagram of a system for enhanced consistency in underwater imagery, according to embodiments of the present disclosure;

[0010] FIG. 2 is a schematic diagram of the system of FIG. 1 further detailing data flow within underwater imagery enhancement software, according to embodiments of the present disclosure;

[0011] FIG. 3 is a flowchart of a method for enhanced consistency in underater imagery, according to embodiments of the present disclosure.DESCRIPTION OF ILLUSTRATIVE EMBODIMENTS

[0012] Embodiments contemplated by the present disclosure will now be described in more detail with reference to the accompanying drawings. The disclosed subject matter should not be construed as limited to only the embodiments set forth herein. Rather, the illustrated embodiments are provided by way of example to covey the scope of the subject matter to those skilled in the art.

[0013] FIG. l is a schematic diagram of a system for enhanced consistency in underwater imagery 10 including an underwater camera 12, a processing unit 14, a user interface 32, a data transmission system 34, a vehicle control application 28, a display device 26, data storage 30 and a visual recognition application 36. The processing unit 14 includes a processor 16, volatile memory 18 and underwater imagery enhancement software 20. The underwater imagery enhancement software includes a Dynamic Underwater Color Transfer (DUCT) algorithm 22and an Al based image enhancement module 24. Whilst the processing unit 14 is shown to be functionally separate from the underwater camera 12, they may be included in a common housing, e.g. implemented in an edge computing manner. Alternatively, the processing unit 14 may be separately housed or cloud based.

[0014] The system for enhanced consistency in underwater imagery 10 is designed to improve the quality and consistency of underwater imagery, utilizing advanced processing techniques and Al to enhance visual content captured by the underwater camera 12.The underwater camera 12 serves as an image capture device, designed to operate in subaquatic environments. The underwater camera 12 captures raw image frames that are often subject to underwater-specific distortions like color casting, reduced contrast, and blurriness due to light absorption and scattering by water particles. The raw images captured by the underwater camera 12 are subsequently processed and enhanced for clarity, color accuracy, and overall visual quality by the processing unit 14. In one example, the underwater camera 12 is mounted on a Remotely Operated Vehicle (ROV). An ROV may be tethered to a control unit on the surface and employ the underwater camera 12 for real-time visual feedback. An ROV may be utilized for underwater exploration, infrastructure inspection, and scientific research. The underwater camera 12 may be equipped on an Autonomous Underwater Vehicle (AUV) that operate autonomously, following pre-programmed missions. The integration of underwater cameras on an AUV supports navigation, obstacle avoidance, and data collection. The underwater camera 12 may be attached to dive equipment or handheld by a diver for targeted image capture such as coral reef monitoring or underwater cinematography. The underwater camera 12 may be connected to fixed installations requiring continuous observation of a specific site, such as an artificial reef. The underwater camera 12 may be incorporated in unmanned surface vehicles (US Vs) for applications involving both surface and underwater observation. USVs can launch camera systems (including the underwater camera 12) for underwater investigations while maintaining surface mobility, which may be useful in search and rescue operations or water quality assessments.

[0015] Regardless of the deployment scenario — whether mounted on ROVs, equipped on AUVs, handheld by divers, or utilized in fixed or wearable configurations — the provision of high-quality images is important. Underwater environments introduce unique distortions and image artifacts, including color casting, reduced visibility due to particulate matter, and light attenuation, which can significantly degrade the quality of captured imagery. The processing unit 14, equipped with underwater imagery enhancement software 20, is specifically designed to address these challenges, utilizing advanced algorithms and Al techniques tailored to correctsuch underwater-specific distortions, ensuring the output of clear, color-accurate, and visually compelling images.

[0016] The output of the enhanced frames from the processing unit 14, achieved through the system's Dynamic Underwater Color Transfer (DUCT) algorithm and Al-based enhancement processes, can be applied in a variety of use cases. In operations involving RO Vs and AUVs, enhanced real-time video feeds can significantly improve operator or autonomous decision-making processes. Enhanced frames serve as a superior input for visual recognition algorithms, facilitating automated identification of marine life, anomaly detection in underwater structures, and machine learning models trained on underwater data. In every scenario, the system's ability to provide enhanced underwater imagery is not just about visual improvement but about unlocking new possibilities and enhancing the efficacy of underwater operations and observations. The enhanced frames boost human perceptibility when displayed across a variety of output devices, such as high-definition monitors, virtual reality (VR) headsets, mobile devices, and immersive projection systems. The enhanced clarity and color accuracy of the images facilitate educational outreach, heritage preservation, and improve operational efficiency in public safety efforts (e.g. rescue missions) by providing greater visual precision.

[0017] The processing unit 14 includes the processor 16, volatile memory 18, and underwater imagery enhancement software 20. The processing unit 14 is configured to process, analyze, and enhance raw images captured by the underwater camera 12 using the DUCT algorithm 22 and the Al-based image enhancement module 24.

[0018] The processor 16 executes the software algorithms described herein and manages data flow within the system 10. The processor 16 can encompass various types of processors, such as Central Processing Units (CPUs) for general-purpose computing, Graphics Processing Units (GPUs) for intensive parallel processing tasks, and specialized processors like Tensor Processing Units (TPUs) and Field-Programmable Gate Arrays (FPGAs) designed to accelerate machine learning and image processing workflows. Examples of such processors where the underwater imagery enhancement software 20 has been deployed and tested for performance include low-power, compact processors like those found in Raspberry Pi 4 and Nvidia Jetson series, which are ideal for real-time applications and embedded systems, as well as high-end GPUs like the Nvidia RTX4090, known for their high computational throughput and efficiency in handling intensive image processing tasks. This varied deployment highlights that the underwater imagery enhancement software 20 described herein is particularly processing efficient and facilitates a wide range of hardware platforms.Volatile memory 18 temporarily stores data and instructions for the processor 16 during execution of the underwater imagery software. Types of volatile memory utilized within the system can include Random Access Memory (RAM), which provides the high-speed read / write capabilities essential for real-time data processing. During the execution of the DUCT algorithm 22, volatile memory 18 is tasked with storing a variety of data. Raw image frames received from the underwater camera 12 may be stored, awaiting processing by the Al-based image enhancement module 24. Enhanced image frames before and after application of color characteristics adjustments may be stored, enabling comparison and blending with an enhanced image template. Image quality scores may be stored for each enhanced frame, used in determining a selection of an enhanced image template and guiding dynamic adjustment of blending steps. Volatile memory 18 may store color statistics (mean and standard deviation across color channels) for both a current enhanced image frame and an enhanced image template for executing the color transfer function. Dynamically calculated blending parameters may be stored such as the number of blending steps (N) and the current blending factor ((|>) for managing gradual application of color characteristics across consecutive frames.

[0019] The underwater imagery enhancement software 20 comprises specialized algorithms (computer program instructions) designed to correct and enhance underwater images. The underwater image enhancement software 20 includes the DUCT algorithm 22, which adjusts the color characteristics of underwater imagery to ensure consistency across a series of images, addressing issues of color inconsistency and enhancing visual appeal. The Al-based image enhancement module 24 utilizes Al techniques, potentially based on a variety of neural network architectures, including but not limited to Generative Adversarial Networks (GANs), which have been found to be particularly effective for the present image enhancement tasks when trained specifically for underwater imagery to correct distortions and enhance image quality.The Al-based image enhancement module 24 employs artificial intelligence techniques to systematically correct common underwater image distortions. The neural network architecture, potentially including GANs, identifies and rectifies issues such as color distortion, loss of contrast, and blurriness inherent in underwater imagery. The Al-based image enhancement module 24 is trained on a dataset that encompasses a diverse array of underwater images, each paired with target images that represent the desired output quality. These target images are typically generated through expert-led enhancement processes or through advanced simulation techniques that mimic optimal underwater visibility conditions. This dual-image approach — comprising raw, distorted underwater scenes alongside their enhanced counterparts — enablesthe module to accurately learn the specific characteristics of underwater image degradation, such as color fading due to the water’s selective absorption of light wavelengths, blurriness from suspended particulate matter, and the loss of contrast caused by scattering. The training process is designed to familiarize the module with the gamut of underwater conditions, from the murkiness of coastal waters to the clearer but deeper oceanic realms. By effectively modeling the transformation from degraded to idealized visibility, the module is adept at autonomously correcting these distortions, enhancing details, and adjusting image properties to yield outputs that closely resemble the clarity, color balance, and detail definition found in the target images. This precise calibration against a well-curated dataset ensures that the AI- based image enhancement module 24 significantly improves the quality of underwater imagery. The underwater image enhancement module 24 analyzes and processes raw image frames captured by underwater camera 12 to improve image quality.

[0020] Following the initial enhancement by the Al based image enhancement module 24, the DUCT algorithm 22 further refines color characteristics of the enhanced image frames to ensure visual consistency and fidelity across a series of images. This algorithm addresses the challenge of color inconsistency across image frames by employing a color transfer function that aligns the color palette of each enhanced frame with that of a selected enhanced image template, chosen for its high image quality score. The color transfer function operates by matching mean color values and standard deviations for each color channel within the selected template to those of the enhanced image frames. In scenarios where the template is updated due to higher quality frames being processed, the algorithm adeptly applies the color characteristics in a blended mode to maintain continuity; otherwise, the algorithm defaults to a direct mode of color transfer for efficiency. In the blending mode, the algorithm incrementally applies the color characteristics of the template to subsequent frames. This blending is governed by a dynamically determined number of blending steps and a blending factor, (|), which progressively increases to control the degree of color characteristic application. The process ensures a gradual and controlled transition of colors.

[0021] The user interface 32 allows users to interact with the system, configure settings, and view processed images. It provides accessibility to system functionalities and facilitates user control over the image enhancement process.

[0022] Data transmission system 34 enables the transfer of data between the system and external devices or networks. The data transmission system 34 can be used to upload raw images for processing and download enhanced images for various applications. The datatransmission system 34 may alternatively allow remote upload of locally produced enhanced images.

[0023] The vehicle control application 28 may be part of autonomous or remotely operated underwater vehicles, and use the enhanced imagery provided by the processing unit 14 for navigation, mission planning, and / or environment interaction. The vehicle control application 28 benefits from clearer, more consistent images for improved decision-making and operational efficiency.

[0024] The display device 26 acts as the interface through which users may view the enhanced underwater images or videos produced by the system 10. The display device 26 may include monitors, immersive VR headsets, high-resolution projectors, interactive touch screens, augmented reality (AR) glasses, portable tablets and smartphones.

[0025] Data storage 30 (non-volatile) may store raw and enhanced images, along with other relevant data. Data storage 30 ensures that data is securely kept for future reference, analysis, or sharing.

[0026] The visual recognition application 36 may leverage the enhanced imagery for visual recognition tasks, such as identifying underwater features, marine life, or man-made objects. The visual recognition application 36 may utilize advanced algorithms to interpret images, benefitting from the increased clarity and consistency provided by the system 10.

[0027] The data transmission system 34, the vehicle control application 28, the display device 26, the visual recognition application and data storage 30 are examples of potential uses of the enhanced imagery provided by the system 10. It should be understood that only one or any combination of such devices may be included in the system 10.

[0028] FIG. 2 illustrates a schematic diagram of the system for enhanced consistency in underwater imagery 10 focusing on data flows. The underwater camera 12 is shown to provide a raw image frame 40 for processing by the Al based image enhancement module 24. The Al based image enhancement module 24 provides an enhanced image frame 42 to the DUCT algorithm 22. The DUCT algorithm 22 includes various software modules including a frame quality evaluation module 46 (providing an image quality score 48), a blending steps determination module 58 (providing a number of steps N 60), a best template management module 50 (providing template data 52), a blending module 62 (providing a blended color adjusted frame 64), a color transfer module 54 (providing a color adjusted image frame 56) and a scene change detection module 66 (providing a similarity score 68).The underwater camera 12 captures the raw image frames from underwater environments and outputs each raw image frame 40. These raw frames often contain distortions specific tounderwater settings, such as color degradation and blurriness due to light refraction. The raw image frame 40 may be the initial data captured by the underwater camera 12, representing the unprocessed visual information directly from the underwater scene. Alternatively, the raw image frame 40 may be the result of some initial pre-processing such as noise reduction, white balance correction, contrast adjustment and image cropping and alignment.A raw image frame 40 captured by the underwater camera 12 is composed of pixels, each representing the smallest unit of an image. These pixels carry information on light intensity and color, structured into color channels that typically correspond to the primary colors of light: red, green, and blue (RGB). The color channels for image processing and enhancement in the present disclosure are not limited to the RGB (Red, Green, Blue) model; alternatively, the YUV color space (or another color space), which separates luminance (Y) from chrominance (U and V components), can also be utilized for manipulating and enhancing underwater imagery. The aggregation of these pixels and their respective color channel values forms the complete image, capturing the visual essence of the underwater scene. The pixels in a raw image frame 40 embody the light intensity captured by the camera's sensor. Each pixel is represented by a combination of values across the RGB color channels, dictating the color and brightness of that specific point in the image. Raw image frames are composed of multiple color channels, usually RGB, each channel storing the intensity information for its respective color across the image. The intensity values in these channels may range from 0 to 255 in an 8-bit image, where 0 represents the absence of that color (black) and 255 represents the full intensity of the color. The resolution of a raw image frame, typically measured in pixels (e.g., 1920x1080), indicates the total number of pixels in the image. Underwater cameras vary widely in their capabilities, and the resolution of the raw image frames they generate can range from standard definition to ultra-high definition, depending on the camera's specifications and the intended use of the imagery.

[0029] The underwater camera 12 captures raw image frames at a specific frame rate, which denotes the number of frames recorded or displayed per second. This frame rate can vary widely depending on the camera's design and the requirements of the particular underwater operation, ranging from lower frame rates (e.g., 24 fps for cinematic look) to higher frame rates (e.g., 60 fps or more for smooth motion capture). The Al-based image enhancement module 24 and the DUCT algorithm 22 are configured to process the raw image frames sequentially, accommodating real-time data flow directly from the camera 12 as well as batch processing from stored data. In real-time processing, such as for navigational aid for RO Vs or live monitoring, the system 10 processes incoming frames on-the-fly. The Al module 24 andthe DUCT algorithm 22 provide enhanced images without significant delay, maintaining the frame rate as closely as possible to the camera's output. Alternative embodiments perform batch processing from memory whereby raw image frames stored in data storage 30 (for example) are processed at a predetermined frame or clock rate. This batch processing mode allows the system 10 to access and enhance images sequentially at a rate optimized for processing power and quality outcomes, independent of the original capture frame rate. In both processing modes, the system 10 applies consistent quality improvements through the Al based image enhancement module 24 and color corrections through the DUCT algorithm 22 frame by frame.

[0030] The Al-based image enhancement module 24 utilizes advanced machine learning algorithms, specifically trained on underwater imagery, to correct distortions observed in raw image frame 40. This module enhances overall image quality by improving clarity, adjusting colors to more natural tones, and increasing contrast to mitigate the effects of underwater scattering and absorption. The Al based image enhancement module outputs an enhanced image frame 42, which represents the visually improved version of the raw image frame 40. This frame has corrected distortions and enhanced features, ready for further processing by the DUCT algorithm 22.The Al based image enhancement module 24 may be provided by a variety of architectures as described in the following, or combinations thereof. In one example, the image enhancement module 24 may be trained using GANs that can generate realistic underwater images from inair image and depth pairings. This approach uses a generative network to produce a large dataset of synthetic underwater images, which can then be used to train an image enhancement convolutional neural network as the image enhancement module 24. This method is beneficial for underwater image enhancement, where collecting large datasets of labelled images is challenging. In another example, the Al based image enhancement module 24 may include a deep Convolutional Neural Network (CNN) that have been trained to correct color, enhance contrast, and improve overall image clarity by being trained on large datasets of underwater images. The Al based image enhancement module 24 may include, in yet another example, encoder-decoder architectures for image enhancement. The encoder-decoder architectures can learn to enhance underwater imagery by encoding the input image into a compact representation and then decoding it back to an enhanced version.

[0031] In a preferred embodiment, the Al-based image enhancement model 24 employs a GAN architecture that generates realistic underwater images. This method leverages unlabeled real underwater images to learn a representation of water column properties, using in-airimages and depth maps to generate corresponding synthetic underwater images. The training dataset comprises a diverse range of underwater images to effectively model and counter the unique challenges presented by underwater conditions, such as color degradation, haziness, and diminished illumination at greater depths. By learning from real underwater environments, the Al-based image enhancement module 24 is able to enhance image clarity and visual appeal.

[0032] Generally, the Al based image enhancement module 24, which may include GANs, CNNs, and / or encoder-decoder networks, are trained on a diverse underwater imagery dataset. This dataset includes a wide range of conditions such as different levels of turbidity, varying lighting conditions, and diverse underwater landscapes. The training process involves adjusting the model’s parameters to minimize the difference between the model's output and the target image quality. This is achieved through various loss functions, which can include pixel-wise loss for direct similarity, perceptual loss for feature similarity, and adversarial loss for realism (in the case of GANs).

[0033] In embodiments of the present disclosure, the Al-based image enhancement module 24 leverages a tailored Generative Adversarial Network (GAN) architecture optimized for underwater imagery enhancement. This module includes three scalable model configurations to address diverse requirements across computational resources and image enhancement objectives. The models range from small to large, each designed for specific application scenarios.

[0034] The small model configuration is engineered for environments with limited computational resources, such as onboard systems in remotely operated vehicles (RO Vs) or autonomous underwater vehicles (AUVs). It features a lightweight generator architecture that employs depthwise and pointwise convolutions, significantly reducing the model's computational footprint while still providing quality image enhancement. This configuration is ideal for real-time applications where speed is critical, employing streamlined encoder-decoder blocks and instance normalization to enhance efficiency and image output quality.

[0035] The medium model configuration offers a balanced solution, tailored for systems with moderate processing capabilities. This setup is based on a GAN architecture but includes advancements such as gated attention mechanisms within skip connections and instance normalization. These enhancements allow for more focused feature refinement and efficient normalization, making it suitable for applications that demand a compromise between image quality and processing speed, such as onboard analysis on research vessels.

[0036] The large model configuration is designed for maximum image quality enhancement, targeting platforms with substantial computational resources, like research labsor cloud-based processing centers. It utilizes a deep network architecture capable of processing larger-sized training images and incorporates bottleneck residual layers and a multi-scale discriminator network. This configuration is aimed at applications where the highest quality image output is paramount and computational time is not a limiting factor, such as detailed scientific analysis or documentary filmmaking.

[0037] These configurations can be deployed independently, chosen based on the specific hardware capabilities and the real-time requirements of the deployment environment, e.g. based on the processing power of the processor 16. For instance, underwater vehicles with limited onboard computing power might employ the small configuration for real-time navigation and basic imaging tasks. In contrast, the large configuration might be reserved for post-mission analysis where higher image quality can significantly enhance the value of the collected data. Furthermore, systems that support dynamic switching between configurations can adjust in real-time to changing computational resources or task requirements, optimizing the balance between enhancement quality and processing speed.

[0038] The training regime for the Al-based image enhancement module 24 involves a paired image training approach, wherein each raw underwater image is associated with a corresponding reference image for objective assessment. This approach combines multiple loss functions to guide the enhancement model toward generating images that closely match the reference images, particularly in aspects such as dehazing and color restoration.

[0039] The reference images may be provided by professional enhancement by which images taken underwater are professionally edited by experts using photo editing software or algorithms to correct color, enhance clarity, and adjust brightness and contrast to approximate the appearance of the scene as if viewed under natural, non-water conditions. The reference images may additionally or alternatively be provided by controlled conditions photography by taking photographs in controlled underwater environments with optimal lighting and visibility conditions can also provide high-quality reference images. These images serve as a benchmark for enhancing more degraded images taken in less ideal conditions. The reference images may additionally or alternatively be provided by synthetic image generation by which reference images are synthetically generated using computer graphics techniques. This method may include creating realistic 3D models of underwater scenes and rendering these models using computer graphics software that simulates the optical properties of water. For training purposes, these reference images are paired with their corresponding raw underwater images. The paired dataset supports training the Al based image enhancement module 24 as it allowsthe model to learn the transformation needed to correct the distortions and enhance the overall quality of underwater images, mimicking the quality seen in the reference images.

[0040] The Al based image enhancement module 24 according to embodiments of the present disclosure may be trained based on at least one loss as described in the following. In some embodiments, the Al based image enhancement module 24 is trained based on any combination of some or all of the following losses. The loss function may employ an adversarial loss component so that the model learns a transformation function mapping raw underwater images to their visually enhanced versions. This process involves a generative network striving to produce images indistinguishable from real, enhanced images, while a discriminative network attempts to differentiate between the real and generated images. The adversarial loss optimizes this competition, refining the generative model's ability to replicate the quality of enhanced underwater imagery accurately. The loss function may incorporate an LI loss (mean absolute error), which targets the minimization of artifacts and blurriness by calculating the mean absolute difference between the reference images and the images enhanced by the model. This ensures a closer match in visual content between the generated images and their corresponding high-quality targets. The loss function may include a further loss that employs a deep learning-based image quality assessment metric that considers both the structural and textural aspects of images, providing a detailed and perceptually relevant evaluation of image quality. The loss function may incorporate an SSIM (Structural Similarity Index Measure) loss component to quantify the perceptual difference between the enhanced images and their corresponding reference images. SSIM evaluates the visual impact of three characteristics of an image: luminance, contrast, and structure, comparing local patterns of pixel intensities that have been normalized for luminance and contrast. This approach allows for a more detailed and human-perceptible measure of image similarity. The loss function can include a perceptual loss component that utilizes intermediate representations extracted from a pretrained Visual Geometry Group Network (VGG-19) to emphasize the significance of high- level features in the image enhancement process. This method compares the enhanced images with their reference counterparts by measuring differences in the feature maps at various layers of the VGG-19 network. This approach of combined loss functions ensures that the enhanced images not only match the reference images in terms of pixel accuracy but also in textures, patterns, and structural elements that contribute to the overall perceptual quality of the images.

[0041] The loss function may employ a linear combination of the aforementioned loss components to establish a comprehensive training objective. The balance among thesecomponents is maintained through weight scaling hyperparameters, which are empirically determined.Enhancing underwater imagery using the Al based image enhancement module 24 introduces improvements in image clarity, contrast, and color accuracy. The Al-based image enhancement module 24 excels in correcting distortions specific to underwater conditions, thereby elevating the quality of each processed frame. However, while individual frame enhancement yields sharper and more vivid images, flickering across frames may remain. This flickering effect, which has been found to result, at least in part, from frame-to-frame color inconsistencies, can detract from video quality. Significant scene changes may introduce additional complexity to the enhancement process. As the underwater environment or lighting conditions shift, the color characteristics that were suitable for one scene may not perfectly suit the next, causing abrupt visual transitions.

[0042] The DUCT algorithm 22 ensures color consistency across consecutive frames, mitigating flickering and facilitating smoother transitions between scenes. By dynamically adjusting color characteristics, the DUCT algorithm 22 maintains visual continuity across the video sequence. This sophisticated method enhances individual frames processed by the AI- based image enhancement module 24 and the overall cohesiveness of the video through the DUCT algorithm 22.

[0043] The frame quality evaluation module 46 of the DUCT algorithm 22 analyzes the enhanced image frame 42 to assess its quality based on predefined criteria, generating an image quality score 48. This score reflects the frame's visual fidelity and suitability to serve as a reference in the color consistency process. The frame quality evaluation module 46 quantitatively assesses the quality of enhanced underwater images or video frames. In one embodiment, the frame quality evaluation module 46 employs a deep learning-based approach to evaluate aspects of image quality. In such an embodiment, the frame quality evaluation module 46 utilizes a pretrained deep learning model that has been exposed to a vast array of image-text pairings. This pretrained model captures a comprehensive understanding of various quality dimensions across a wide range of images, enabling it to assess both the tangible (e.g., sharpness, exposure) and intangible (e.g., aesthetics, mood) aspects of image quality. The evaluation process involves the generation of textual prompts that describe desired image qualities or attributes. These prompts can be specifically crafted to assess different aspects of image quality, such as brightness, contrast, or even more abstract qualities like the "feel" of the image. The frame quality evaluation module 46 uses these prompts to guide the assessment process, allowing for a flexible evaluation framework that can adapt to various qualitydimensions. For each enhanced image or frame, the frame quality evaluation module 46 performs inference by presenting it alongside a set of predefined prompts to the pretrained model. The model then assesses the compatibility of the image with the qualities described in the prompts, generating scores that reflect the image's adherence to these quality attributes. The output from the model is a set of quantitative scores corresponding to each assessed quality dimension. These scores provide a comprehensive evaluation of the image's quality, encompassing both its physical characteristics and its ability to convey desired abstract qualities. The individual quality dimension scores are aggregated or selectively combined to form a comprehensive quality score for each image or frame, which corresponds to the image quality score 48.In another embodiment, the frame quality evaluation module 46 includes a combination of deep learning and traditional image quality assessment techniques. A convolutional neural network (CNN) may extract pertinent features from underwater images, focusing on aspects like texture, color distribution, and edge sharpness. Pretrained networks such as VGG or ResNet, may be adapted to underwater imagery through transfer learning, to capture a range of features reflective of image quality. Algorithmic processes may capture image quality metrics such as the Structural Similarity Index (SSIM), Peak Signal-to-Noise Ratio (PSNR), and Colorfulness metric. These metrics evaluate various dimensions of image quality, including brightness, contrast, and color fidelity. A machine learning model, e.g. a neural network, may be provided and trained on a dataset of underwater images annotated with quality scores. This model learns to predict image quality scores based on the extracted features and traditional metrics, offering insights into the perceptual quality of the image. A final image quality score is a composite measure that reflects objective quality aspects and subjective perceptual qualities and output as the image quality score 48.

[0044] The image quality score 48 is a numerical value representing the assessed quality of the enhanced image frame 42 and is used to identify the most suitable frames for template selection.The best template management module 50 manages the selection and storage of enhanced image frames as templates based on their quality scores. The best template management module 50 manages multiple template types, including current, next, previous, and raw image templates. The current enhanced image template serves as the primary reference for color transfer and blending processes. It represents the highest quality frame encountered up to the current moment in the video sequence, determined by its image quality score. When a new enhanced image frame receives a higher score 48 from the frame quality evaluation module 46,the current enhanced image template is updated to this new frame. This ensures that color transfer operations use the most visually appealing frame as a reference. Upon detecting a frame with a higher quality score 48 than the current best, the algorithm may encounter two scenarios. A first scenario occurs when blending (discussed further below) is in progress. If the system is already blending between templates, the newly identified superior frame is saved as the next image template. This template will be utilized in future blending steps once the current blending process concludes. A second scenario occurs when no current blending is currently taking place. In the second scenario, the newly observed frame is set as the new image template. This initiates a new blending process to gradually transition from the previous template to this new template, ensuring a smooth change in visual appearance across frames.

[0045] The best template management module 50 also manages a raw image template, which represents the raw counterpart of the current template and is used for scene change detection (described further below) and resetting the best score under significant environmental shifts. When a new enhanced image template is detected, the raw template is updated to the raw version of enhanced image frame that has been set to the new image enhanced image template.

[0046] The template data 52 includes information on the current enhanced image template, which is a reference for color transfer, a newly observed image template that becomes the next template for color transfer, and a temporary next template stored (queued) during blending processes. Additionally, the template data includes a raw image template, which is a raw image counterpart of the current enhanced image template used in scene change detection.

[0047] The blending steps determination module 58 calculates the number of steps (N) 60 to blend the color characteristics between the current enhanced image template and a newly observed template, ensuring a smooth transition in visual appearance across frames. That is, the number of blending steps is determined based on a difference in color characteristics between the current enhanced image template and a newly observed image template so that a greater difference results in a greater number of steps. Referring to equation 1 below, this calculation is based on comparing mean pixel values (Ma and Mb) for each color channel of the current and new templates (Ta and Tb, respectively), across all pixels (X, Y) and color channels (C). A smoothness factor (q), which may be empirically set, influences the sensitivity of blending to changes in color characteristics, with a lower value resulting in more blending steps. Nmax is a maximum allowable number of blending steps, and may also be empirically determined, to balance transition smoothness with processing efficiency. The outcome, N, isused to adjust a blending factor ((|>) in subsequent processing, dictating the proportion of color characteristics from each template applied during each blending step. (equation 1)

[0048] The number of steps N 60 is output by the frame quality evaluation module 46 as the number of steps required for blending color characteristics between frames.Color transfer module 54 applies color adjustment to enhanced image frames, resulting in a color-adjusted image frame 56. The color transfer module 54 operates on the principle of adjusting the color characteristics of the enhanced image frame 42 to align with the color characteristics of a selected template (e.g., current, or next enhanced image templates). This alignment is critical for maintaining visual consistency across frames, especially in sequences where lighting conditions or environmental factors may cause abrupt color shifts. The color transfer module 54 converts the color space of the enhanced image frame and the template from RGB to a decorrelated color space. This conversion reduces correlations between the color channels, allowing for independent manipulation of color characteristics. The chosen color space is one where the correlation between the color channels is minimized (such as lap or another perceptually uniform color space), facilitating adjustments that closely mimic human visual perception. For both the enhanced image frame 42 and the template (which could be a current or next template depending on the operational context within the DUCT algorithm 22), the color transfer module 54 calculates statistical measures representing color characteristics, e.g. specifically, the mean and / or standard deviation for each color channel in the converted color space. These statistical measures capture the core color characteristics of both the enhanced image frame 42 and the enhanced image template. The color transfer module 54 adjusts the color characteristics of the enhanced image frame 42 to match those of the template by: subtracting the mean of each color channel of the enhanced image frame 42 to center its distribution, scaling the centered distribution by the ratio of the template's standard deviation to the frame's standard deviation for each color channel to adjust the spread of the color distribution of the enhanced image frame 42 to match that of the template, and adding the meanof the template's color channels to the scaled distribution to shift the color distribution enhanced image frame 42 to align with the template's central tendency. After adjusting the color characteristics, the color-corrected frame is converted back from the decorrelated color space to the standard RGB color space to prepare the color adjusted image frame 42 for display or further processing.

[0049] In other words, the color transfer module 54 adapts the color characteristics of enhanced image frames to align with those of an input enhanced image template, which includes, for each color channel, computing the mean and standard deviation for both the frame and the template and adjusting the colors of the frame by aligning mean of the frame with that of the template and / or scaling the color values in each channel of the enhanced image frame 42 to have the same standard deviation as the corresponding color channel in the template. The color transfer module 54 is applied to each enhanced image frame in relation to the chosen template.

[0050] A colour adjusted image frame 56 is output from the color transfer module 54 representing an enhanced image frame with adjusted color characteristics aligned with color characteristics of the template.

[0051] When there is a new image template that has been set based on the enhanced image frame 42 having an image quality score 48 that exceeds the best score, gradual color adjustments may be applied across N frames by the blending module 62. The blending module 62 in the Dynamic Underwater Color Transfer (DUCT) algorithm addresses the issue of abrupt color changes across video frames, which was a limitation observed in prior algorithms. The blending module 62 uses a previous template frame and a new template frame. The previous template frame corresponds to the previous best enhanced image frame prior to it being updated by the new best enhanced image frame. During the blending process color characteristics from both templates are applied to the current enhanced image frame, ensuring a gradual transition of color characteristics between frames.

[0052] The blending may be conducted according to the equation 2 below: lab = la • (1 - < >) + lb • < > (equation 2) la and lb are the frames to be blended, (|) is the blending factor that ranges from 0 to 1, and lab is the resulting blended frame. This calculation is performed on a pixel-by-pixel basis across all color channels.

[0053] The algorithm of the blending module dynamically adjusts the blending factor ((|>) based on the current blending step (step) and the total number of blending steps N, calculated as (|) = step / (N + 1). With each successive frame, the influence of the new template graduallyincreases while the influence of the previous template correspondingly decreases, leading to a smooth transition in color characteristics across frames. When the blending module reaches the final step, the algorithm directly applies the color characteristics from the new template to the current frame, marking the completion of the blending process for that set of templates. If a next template is waiting, the algorithm proceeds to blend with this new template, updating the total number of blending steps (N) for the new blending process.The blending module 62 utilizes the color transfer function of the color transfer module as a precursor to the blending operation by first applying the color transfer function to adjust the color characteristics of the current enhanced image frame to match those of the previous and new enhanced image templates. This process results in two color-adjusted frames: one where the current frame's colors are adjusted to match the previous template and another adjusted to match the new template. The final blending operation is performed pixel-by-pixel across all three color channels, combining the two color-adjusted frames based on the blending factor. The blending module 62 outputs a blended color adjusted image frame 56.

[0054] The scene change detection module 66 monitors for significant changes in the scene that may affect the relevance of the current template, using image analysis to output a similarity score 68. The score is used to determine when to update the template to better match new scene characteristics by resetting the best score. The scene change detection module 66evaluates the similarity between the raw image of the current template and the raw image of the current enhanced image frame 42. The raw images are compared to determine significant changes in the scene. The scene change detection module 66 determines a similarity score 68, which can be compared to a threshold. When the threshold is exceeded, a scene change is detected and the best score is reset to zero. In one embodiment, scene detection module 66 may leverage deep neural network features to assess similarities, focusing on both the structural integrity and textural details of the compared images. Other embodiments for comparing the raw images can be used either alone or in combination. Structural Similarity Index Measure (SSIM) is one example, which is a perception-based model that considers changes in luminance, contrast, and structure between two images. In another alternative, feature matching using descriptors could be employed including Scale-Invariant Feature Transform (SIFT) or Speeded-Up Robust Features (SURF), which extract and compare feature points between images. The number of matched features and the quality of matches can serve as indicators of scene continuity or change. A histogram comparison could also be included wherein the color or intensity histogram of an image provides a global description that can be useful for scene change detection. Histogram intersection, Chi-square distance, or Bhattacharyya distance betweenhistograms of the raw frames can quantify scene changes. Deep learning-based change detection can be deployed, especially Convolutional Neural Networks (CNNs) or Transformerbased models, which can be trained to directly classify pairs of images as "same scene" or "scene change".

[0055] Similarity score 68 provides a value indicating the degree of visual similarity between a new enhanced image frame and a current template, used to detect scene changes. FIG3 provides a flowchart of a method 300 for enhanced consistency in underwater imagery. The method 300 is performed by the processing unit 14 executing the underwater imagery enhancement software 20.

[0056] It will be understood by those skilled in the art that all the steps as illustrated in FIG. 3 do not have to be performed in the order as shown in FIG 3 Instead, some steps can be performed in parallel or earlier or later than shown in FIG. 3.

[0057] Method 300 starts at 302. The method 300 initializes variables at step 304 to prepare for processing enhanced image frames provided by the Al based image enhancement module 24. A blending flag is set to false, indicating that template blending is not initially occurring. Various template variables are initialized to none, including the current enhanced image template (Tprev), the raw image used to create the current template (TR), a newly observed image template (TEnew), and a next template if currently blending (TEnext). Initial values for the total number of blending steps (N and Nnext) and the current blending step (step) are set to 1, the best score is initialized to zero and a threshold for scene change detection (dists thresh) is set to an empirically determined value.At step 306, the algorithm receives the enhanced image frame 42 and the counterpart raw image frame 40. The processing pipeline outlined by method 300 (after initialize variables step 304) is performed on each newly received enhanced image frame 42 and generates a color adjusted version for each newly received image frame 42.

[0058] At step 308, frame quality of the enhanced image frame 42 (FE) is evaluated using the frame quality evaluation module 46 to provide a score (image quality score 48) representing a determined quality of the enhanced image frame 42. The frame quality evaluation of step 308 may evaluate image quality by considering perceptual features, such as visual clarity, contrast, and texture detail, comparing these attributes against learned models of high-quality imagery to generate the score.Step 310 compares the score to a best score that is maintained based on scores achieved by prior enhanced image frames (or a best score of zero for a first processed enhanced image frame or after a reset of the best score). If a new best score is detected, the algorithm checks if thereis a current blending process (at step 312). If there is current blending, the enhanced image frame (FR) is set as a next template (TEnext) and the template is queued (step 320) for later application. In step 322, the raw image frame (FR) corresponding to the enhanced image frame is set as the raw template (TR). Further, and in step 324, a next number of blending steps (Nnext) is determined. The number of blending steps is determined by the blending steps determination module 58 based on color characteristic differences between the new template (TEnew) and the next template (TEnext).

[0059] If it has been determined that there is no current blending in step 312, the algorithm sets the current enhanced image (FE) as the new template (TEnew) (at step 314) and sets the counterpart raw image frame (TR) as the best raw template (TR) (at step 316). Further, the next template (TEnext) is set to none and the blending flag is set to true. The blending flag is used in steps 312 and 326 to determine whether a blending process is being performed. In step 318, the number of blending steps (N) is determined in step 318 based on a difference in color characteristics between a previous template (TEprev) and the new template (TEnew), which is set in step 314. The number of blending steps N 60 is determined by the blending steps determination module 58 as described previously.In step 326, there is a second step of determining if a current blending process is being performed based on the blending flag. If there is a current blending process, the algorithm determines if the number of blending steps is complete in step 308. If the number of blending steps is complete, a final one of the color transfer processes of the blending process is applied in step 346. That is, the enhanced image frame (FE) is set to a result of the color transfer function (CT(FE, TEnew)) by which the enhanced image frame has its color characteristics adjusted according to the new template (TEnew). The color transfer function is executed by the color transfer module 54. The color transfer function adjusts the color characteristics of the enhanced image frame to match those of the new template by first aligning their color spaces, then adjusting the mean and standard deviation of the color distributions channel-wise, ensuring the enhanced image adopts the color ambiance of the template. Further, the current count of steps (step) is reset to 1 and the stored previous template (TEprev) is set as the new template (TEnew)In step 348, it is determined whether there is a queued template (see discussion of step 320 above). If not, the color adjusted enhanced image frame (from step 346) is returned. Further, the new template (TEnew) is set to none and blending is set to false.

[0060] If there is a queued template, the new template (TEnew) is set, at step 352, to the queued template (TEnext) from step (320) and the next template (TEnext) is set to none. Atstep 354, the number of blending steps (N) is set to the next number of blending steps (Nnext) from step 324 such that when the next enhanced image frame is received at step 306, the blending process begins based on that enhanced image frame and the queued template. The color adjusted enhanced frame from step 346 is returned in step 356.

[0061] Returning to the outcome of decision 338 as to whether the number of blending steps are complete, if the blending process is not yet complete, the algorithm determines a blending factor (<|>) in step 340. The blending factor is calculated as a fraction of the current blending step over the total number of blending steps plus one, determining the proportionate influence of the new template on the enhanced image frame during the blending process. In step 342, a blended color transfer is performed in step 342 using the color transfer function. That is, the color transfer function (CT(FE, TEprev)) is applied to the enhanced image frame and the previous template so that the color characteristics of the previous template are applied to the enhanced image frame and the color transfer function (CT(FE, TEnew)) is applied to the enanced image frame and the new template (set in step 314) so that the color characteristics of the new template are applied to the enhanced image frame. The color characteristics of the enhanced image frame is adjusted based on a blend ( blend(CT(FE, TEprev), CT(FE, TEnew), (|>) of the color transfers from the pervious and new templates according to the dynamically adjusted blending factor from step 340. The blending step 342 is performed by the blending module 62, which operates by linearly interpolating between the two color-adjusted enhanced image frames on a per-pixel basis across all color channels, using the blending factor (|) that varies from 0 to 1 (as discussed with reference to equation 2 above). The resulting frame is a blend of the enhanced frame adjusted to the previous template's color characteristics, and the enhanced image frame adjusted to the new template's color characteristics, ensuring a smooth transition between frames. The step count is increased by 1. The color adjusted enhanced image frame from step 342 is returned.

[0062] Returning to decision step 326, if there is no current blending and no new best score is detected in step 310, the color transfer function (CT(FE, TEprev)) is directly applied to the enhanced image frame using the previous template. The algorithm then checks for significant scene changes by comparing the current raw frame (FR) against the best raw frame (TR) using the scene change detection module 66. In one embodiment of step 330, the structural and textural similarities between the raw image frame of the current template (TR) and the raw image frame of the enhanced image frame (FR) are evaluated to produce the similarity score. Significant changes in scene composition or quality degradation, triggering a reset in in the best score if the discrepancy (similarity score) surpasses a predefined threshold (dists thresh).That is, if a major scene change is detected in step 330, indicating a divergence greater than the predefined threshold, the best score is reset to zero to accommodate for the new scene dynamics by facilitating the next received enhanced image frame being set as the new template in step 314. The enhanced image frame with the color characteristics adjusted as a result of step 328 is returned in steps 334 or 336.

[0063] Accordingly, method 300 provides a number of routes for determining and outputting a color adjusted version of the enhanced image frame. A new template may be used for applying the color transfer or a previous template may be used depending on whether the enhanced image frame has a sufficiently high image quality score. Further, the color transfer may be performed in a blended way according to a dynamic number of blending steps determined depending on a difference in mean pixel values between the next template to be used and the previous template used. Each enhanced image frame that is received is converted to a single color adjusted enhanced image frame by performance of method 300, thereby facilitating real-time, frame-by-frame, color adjustment to reduce flickering and enhancing consistency in a video output.

[0064] The invention has been described by reference to certain embodiments discussed above. It will be recognized that these embodiments are susceptible to various modifications and alternative forms well known to those of skill in the art.

[0065] Further modifications in addition to those described above may be made to the structures and techniques described herein without departing from the spirit and scope of the invention. Accordingly, although specific embodiments have been described, these are examples only and are not limiting upon the scope of the invention.

Claims

CLAIMS1. A computer implemented method for enhancing underwater imagery, comprising: receiving a sequence of raw image frames including a current raw image frame; processing the raw image frames with an Artificial Intelligence (Al) image enhancement model to provide enhanced image frames including a current enhanced image frame; determining image quality scores for the enhanced image frames; based on the image quality scores, setting one of the enhanced image frames as an enhanced image template; and adjusting color characteristics of the current enhanced image frame using a color transfer function for applying color characteristics of the enhanced image template to those of the current enhanced image frame.

2. The computer implemented method of Claim 1, wherein the enhanced image template is updated based on the image quality scores as new raw image frames are received and processed.

3. The computer implemented method of Claim 1, wherein the Artificial Intelligence image enhancement model employs a neural network-based approach trained using underwater imagery to correct underwater image distortions including color distortion, increasing contrast, and reducing blurriness.

4. The computer implemented method of claim 3, wherein the neural network-based approach is trained using underwater imagery to enhance visual features including edge definition, comer sharpness, and keypoint clarity.

5. The computer implemented method of claim 1, wherein the Artificial Intelligence (Al) image enhancement model includes a Generative Adversarial Network (GAN) configured to: generate enhanced underwater imagery through an adversarial process that iteratively refines imagery based on feedback from a discriminative model,employ separate networks for assessing and enhancing different aspects of image quality, including one network focusing on global image adjustments and another on enhancing local details, and utilize a training dataset comprising pairs of raw and professionally enhanced underwater imagery to learn optimal enhancement transformations.

6. The computer implemented method of any preceding Claim, wherein determining image quality scores for the enhanced image frames includes employing a neural network to evaluate characteristics of image quality including at least one of: brightness, contrast, colorfulness, and sharpness.

7. The computer implemented method of any preceding Claim, wherein adjusting color characteristics using the color transfer function comprises: matching, or more closely matching, mean color values of each of a plurality of color channels of the enhanced image template to those of the current enhanced image frame.

8. The computer implemented method of any preceding Claim, wherein adjusting color characteristics using the color transfer function comprises: matching, or more closely matching, standard deviation of color values around mean color values for each channel of a plurality of color channels of the enhanced image template to those of the current enhanced image frame.

9. The computer implemented method of any preceding Claim, comprising maintaining a best score of the image quality scores for the enhanced image frames, the one of the enhanced image frames being set as the enhanced image template based on the image quality score for the one of the enhanced image frames having an image quality score exceeding the best score.

10. The computer implemented method of any preceding Claim, setting one of the sequence of raw image frames as a raw image template, the one of the sequence of raw image frames being a raw image frame that has been processed to the one of the enhanced image frames set as the enhanced image template, the method comprising detecting a scene change in the current raw image frame relative to the raw image template and resetting the best score in response to detecting the scene change.

11. The computer implemented method of any preceding Claim, comprising a selective blending process so that the adjusting color characteristics of the current enhanced image frame using the color transfer function staggers application of color characteristics of the enhanced image template to those of the current enhanced image frame over a plurality of consecutive enhanced image frames.

12. The computer implemented method of Claim 11, wherein the raw image frames are received according to a frame rate and the most recent raw image frame is set as the current raw image frame, which is processed by the Al image enhancement model into the most recent enhanced image frame, which is set as the current enhanced image frame, the method comprising updating the enhanced image template when the most recent enhanced image frame has an image quality score the is greater than a best image quality score for the enhanced image frame by setting the most recent enhanced image frame as the enhanced image template, wherein the selective blending process is applied when the update of the enhanced image template is detected.

13. The computer implemented method of Claim 11 or 12, wherein the selective blending process includes a calculation for determining a number of blending steps based on a difference in color characteristics between a preceding enhanced image template preceding a most recent update of the enhanced image template and a current enhanced image template set as a result of the most recent update.

14. The computer implemented method of Claim 11, 12 or 13, wherein the selective blending process calculates a blending factor for each of the plurality of blending steps, with the blending factor designed to progressively increase in each step to dictate a gradually increasing application of the color characteristics from the enhanced image template to the current enhanced image frame.

15. The computer implemented method of any one of the preceding claims, wherein adjusting color characteristics of the current enhanced image frame is performed in two modes: a blended mode, where color transfer occurs across a dynamically determined number of blending steps when the enhanced image template changes, and a direct mode, where color transfer is applied directly without blending if no change in the enhanced image template is detected.

16. The computer implemented method of any preceding claim, wherein adjusting color characteristics of the current enhanced image frame using the color transfer function provides a color adjusted image frame and further comprises outputting a sequence of color adjusted image frames based on the sequence of raw image frames, wherein the method further comprises: applying a visual recognition algorithm to the sequence of color adjust image frames to identify underwater features; autonomously controlling an underwater vessel using the sequence of color adjust image frames; and generating a display based on the sequence of color adjusted image frames.

17. A system for enhancing underwater imagery, comprising: an underwater camera for capturing a sequence of raw image frames; a processing unit; the processing unit including an Artificial Intelligence (Al) image enhancement model configured to processes the sequence of raw image frames to generate enhanced image frames including a current enhanced image frame; the processing unit configured to: determine image quality scores for the enhanced image frames; set one of the enhanced image frames as an enhanced image template based on the image quality scores; adjust color characteristics of the current enhanced image frame by applying a color transfer function for applying color characteristics of the enhanced image template to the current enhanced image frame.

18. A non-transitory computer-readable medium having instructions stored thereon that, when executed by a processor of a system for enhancing underwater imagery, cause the system to perform operations according to the method of any one of Claims 1 to 16.

Citation Information

Cited By

  • Underwater image enhancement model training method, system and equipment

    CN121095091A

  • Image analysis method fusing physical prior

    CN121120626A

  • Tunnel panoramic video enhancement and modeling method and system based on multi-modal large model

    CN121458603A