An underwater image distortion correction method based on a multi-medium refraction model

The underwater image distortion correction method based on the multi-medium refraction model solves the problem of light refraction distortion caused by changes in the medium in the underwater image processing system, achieving high-precision object recognition and improved security inspection efficiency, and adapting to complex underwater environments.

CN120997099BActive Publication Date: 2025-12-30SICHUAN ENERGY INTERNET RES INST TSINGHUA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511525604.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2025-12-30
Estimated Expiration
2045-10-24

AI Technical Summary

Technical Problem

Underwater image processing systems suffer from light refraction distortion due to changes in the medium, affecting the position, shape, and edge information of objects, resulting in inaccurate recognition results. Furthermore, they lack adaptive modeling capabilities, making it difficult to guarantee unified analysis across devices and scenarios.

Method used

An underwater image distortion correction method based on a multi-medium refraction model is adopted. By acquiring original image data, extracting optical interference features, analyzing medium changes using a convolutional neural network, and combining adaptive correction algorithms and dynamic threshold settings, a corrected image is generated. Object contour extraction and brightness equalization are performed, standardized process parameters are updated, noise suppression and iterative optimization are integrated, and a general processing template is generated.

Benefits of technology

It significantly improves the clarity of underwater images and the accuracy of object recognition, optimizes security inspection efficiency, ensures the accuracy and stability of recognition results, and adapts to complex underwater environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997099B_ABST
    Figure CN120997099B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on the distortion correction method of underwater image of multi-medium refraction model, it is related to image processing technical field, including S1, by collecting original underwater image data and extracting data optical interference features, obtain preliminary distortion parameter set, S2, according to preliminary distortion parameter set analysis medium change influence, using convolutional neural network model processes image pixel distribution and integrates pixel intensity adjustment and local contrast enhancement, determine refraction distortion degree, S3, if refraction distortion degree exceeds preset threshold, then adaptive correction algorithm is applied to image pixel distribution in conjunction with refraction deviation compensation and dynamic threshold setting, obtain corrected image version;The distortion correction method of underwater image based on the multi-medium refraction model, significantly improve underwater image definition and object recognition accuracy, optimize security efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and specifically to an underwater image distortion correction method based on a multi-medium refraction model. Background Technology

[0002] In fields such as marine engineering, port infrastructure, and underwater pipelines, underwater security inspections serve as a crucial means of ensuring structural integrity and operational safety. The accuracy and efficiency of these inspections directly impact the reliability of critical national facilities and the level of public safety. With the deepening of marine development activities, the underwater environment is becoming increasingly complex, posing greater challenges to security inspection tasks. Image processing technology is playing an increasingly critical role, particularly in identifying hidden objects, extracting target contours, and assessing minute damage.

[0003] Currently, underwater image processing systems primarily rely on underwater cameras to acquire video images and utilize techniques such as image enhancement, contour extraction, and target detection to assist in security checks. However, underwater optical interference severely restricts image quality. This is mainly manifested in light refraction distortion caused by changes in the medium (such as air-glass-water multilayer media), which leads to deviations in the position, shape, and edge information of objects in the image, resulting in inaccurate recognition results and seriously affecting the reliability of security inspection conclusions. For example, when detecting cracks in underwater pipes or deformation of facilities, refraction distortion may cause blurred boundaries and misaligned contours of target objects in the image, affecting the judgment of automatic recognition algorithms and even leading to false detections or missed detections. Furthermore, existing technologies generally lack adaptive modeling capabilities for complex underwater medium changes. Under different water quality, current velocity, and depth conditions, image processing performance fluctuates significantly, making it difficult to guarantee stability. Moreover, due to the lack of a standardized mechanism in current underwater image processing workflows, data acquired by different devices, after undergoing their own independent processing algorithms, exhibits significantly different results, making unified analysis across devices and scenarios difficult. This not only increases the burden of manual review but also limits the versatility and reusability of image processing templates. Especially in high-risk scenarios, such as closed waters of ports or deep-water well inspections, if image distortion is not effectively corrected, it is very easy to miss small but critical safety hazards. Summary of the Invention

[0004] The purpose of this invention is to provide an underwater image distortion correction method based on a multi-medium refraction model, thereby solving the problems existing in the prior art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: an underwater image distortion correction method based on a multi-medium refraction model, comprising: S1, acquiring raw underwater image data and extracting optical interference features from the data to obtain a preliminary distortion parameter set; S2, analyzing the influence of medium changes based on the preliminary distortion parameter set, using a convolutional neural network model to process the image pixel distribution and integrating pixel intensity adjustment and local contrast enhancement to determine the degree of refraction distortion; S3, if the degree of refraction distortion exceeds a preset threshold, applying an adaptive correction algorithm to the image pixel distribution combined with refraction deviation compensation and dynamic threshold setting to obtain a corrected image version; S4, obtaining object contour data from the corrected image version and integrating region segmentation processing and brightness equalization to determine whether the object offset is acceptable. Within the specified range, the offset calibration result is obtained; S5. The offset calibration result is used to update the standardized process parameters. The image edges are smoothed using a Gaussian filtering algorithm, and kernel size selection and standard deviation parameters are applied to determine the final recognition accuracy index; S6. If the final recognition accuracy index meets the detection capability requirements, the standardized process parameters are integrated to generate a general processing template and a noise suppression mechanism and iterative optimization loop are embedded to obtain the optimized image output; S7. The hidden object features are analyzed based on the optimized image output, and edge preservation filtering and bilateral filtering expansion are combined to determine the degree of improvement in security inspection efficiency and obtain a complete underwater inspection report; S8. The adaptive correction algorithm parameters are fed back through the complete underwater inspection report, and gradient calculation integration and smoothing intensity control are integrated to update the medium change model and obtain the basis for the next cycle of image processing.

[0006] Preferably, step S1 includes acquiring raw underwater image data through sensors to obtain an initial image set; using image preprocessing techniques to denoise the initial image set to generate a denoised image set; extracting optical interference features from the denoised image set to obtain an interference feature set; if the feature values ​​in the interference feature set exceed a preset threshold, classifying the feature values ​​to determine the interference category; calculating distortion parameters based on the interference category to generate an initial distortion parameter set; using principal component analysis to reduce the dimensionality of the initial distortion parameter set to obtain an optimized distortion parameter set; and using a support vector machine algorithm to verify the parameters of the optimized distortion parameter set to obtain the final distortion parameter set.

[0007] Preferably, step S2 includes obtaining refractive index change data from the medium change characteristics, analyzing the interference of refractive index change on the image pixel distribution using an optical flow estimation algorithm to obtain a pixel offset set; if the offset value in the pixel offset set exceeds a preset threshold, performing cluster analysis on the offset value to determine the offset category; optimizing the pixel intensity adjustment parameters using a gradient descent algorithm based on the offset category to obtain an adjusted pixel intensity set; processing the adjusted pixel intensity set using a local contrast enhancement algorithm to generate an enhanced image set; extracting local texture features from the enhanced image set, comparing it with the original image using a feature matching method to obtain a texture distortion distribution; if the distortion value in the texture distortion distribution exceeds a preset threshold, performing weighted fusion on the distortion value to generate a refractive distortion degree parameter; and determining the final refractive distortion degree by performing correlation analysis between the refractive distortion degree parameter and the medium change characteristics.

[0008] Preferably, step S3 includes obtaining pixel distribution data from the original image, analyzing the impact of refraction distortion on the pixel distribution using an optical flow estimation algorithm to obtain a pixel offset set; if the offset value in the pixel offset set exceeds a preset threshold, then processing the pixel offset set using an adaptive correction algorithm, combined with refraction deviation compensation, to generate a corrected pixel set; adjusting the local feature contrast of the corrected pixel set using a dynamic threshold setting method to obtain an enhanced image set; extracting texture features from the enhanced image set, comparing it with the original image using a feature matching method, and determining the final image version.

[0009] Preferably, step S4 includes obtaining object contour data from the corrected image version, extracting the boundary pixel set of the target object using an edge detection algorithm to obtain an object contour dataset; partitioning the corrected image version using a region segmentation algorithm based on the object contour dataset to generate a region segmentation dataset containing the target object; extracting brightness distribution features from the region segmentation dataset, adjusting the brightness of the region segmentation dataset using a histogram equalization method to obtain a brightness equalization image set; if the object offset value in the brightness equalization image set exceeds a preset threshold, comparing it with the object contour dataset using a texture matching algorithm to determine whether the object offset is within an acceptable range, and obtaining an offset calibration result.

[0010] Preferably, step S5 includes processing the offset calibration data using a parameter adjustment algorithm based on the offset calibration results to generate an updated parameter dataset; obtaining the kernel size and standard deviation from the updated parameter dataset, and smoothing the image edges using a Gaussian filtering algorithm to obtain a set of smoothed edge images; if the matching degree between the texture features of the smoothed edge image set and the preset texture template is lower than a preset threshold, then performing a secondary calibration using a texture matching algorithm to determine the calibrated image dataset; processing the calibrated image dataset using a region segmentation algorithm to extract the target object region, calculating a comprehensive index of brightness distribution and edge sharpness, and obtaining the final recognition accuracy index.

[0011] Preferably, step S6 includes obtaining standardized process parameters from a preset parameter database, generating an initial general processing template through a parameter integration algorithm, and determining the preliminary structure of the initial general processing template; if the detection capability of the initial general processing template is lower than a preset threshold, then using a noise suppression algorithm to process the parameter dataset in the initial general processing template to obtain a noise-suppressed template parameter set; based on the noise-suppressed template parameter set, using an iterative optimization loop algorithm to adjust the parameters in the template parameter set to obtain an optimized general processing template; and using a region segmentation algorithm to process the optimized general processing template, extracting feature regions from the target image, and generating an optimized image output.

[0012] Preferably, step S7 includes acquiring a target underwater image from a preset image database, processing the target underwater image using a bilateral filtering algorithm to preserve edge details, and obtaining a first optimized image; for the first optimized image, using an image segmentation algorithm to separate the foreground and background regions, extracting the hidden object region, and obtaining a segmentation feature set; based on the segmentation feature set, using a convolutional neural network algorithm to analyze the hidden object features, determining whether a target object exists, and obtaining an object recognition result; comparing the object recognition result with a preset efficiency threshold, and using a report generation algorithm to integrate the object recognition result and efficiency data to obtain an underwater inspection report.

[0013] Preferably, step S8 includes acquiring a target underwater image from a preset underwater image database, processing the target underwater image using an adaptive correction algorithm, and generating a first corrected image by adjusting brightness and contrast; for the first corrected image, calculating the pixel gradient of the image using a gradient calculation method, extracting edge features, and generating an edge feature set.

[0014] Preferably, step S8 further includes adjusting the smoothness of features based on the edge feature set using a smoothing intensity control algorithm to optimize the edge features and obtain an optimized feature set; by comparing the optimized feature set with a preset medium change database, if the matching degree of the optimized feature set exceeds a preset threshold, then using a convolutional neural network algorithm to update the medium change model to obtain the basis for the next cycle of image processing.

[0015] As can be seen from the above technical solution, the present invention has the following beneficial effects:

[0016] This underwater image distortion correction method based on a multi-medium refraction model addresses the operational problems of image refraction distortion and insufficient object recognition accuracy caused by medium changes in underwater security inspections. This problem stems from underwater optical interference leading to image distortion, affecting object contour extraction and hidden feature analysis, and reducing security inspection efficiency. This invention acquires raw image data, extracts optical interference features, generates a preliminary distortion parameter set, and uses a convolutional neural network to analyze medium changes, adjust pixel distribution and local contrast, and accurately determine the degree of refraction distortion. For distortion exceeding a threshold, an adaptive correction algorithm is applied, combining refraction deviation compensation and a dynamic threshold to generate a corrected image. Further, object contours are extracted through region segmentation and brightness equalization, offset is calibrated, standardized process parameters are updated, and Gaussian filtering is used to smooth edges to ensure recognition accuracy. Finally, noise suppression and iterative optimization are integrated to generate a general processing template, which, combined with edge preservation and bilateral filtering, improves security inspection efficiency and provides feedback to update the medium model, providing a basis for the next cycle. This invention significantly improves underwater image clarity and object recognition accuracy, optimizing security inspection efficiency. Attached Figure Description

[0017] Figure 1 This is a flowchart of the underwater image distortion correction method based on a multi-medium refraction model according to the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] like Figure 1As shown, this invention provides a technical solution: an underwater image distortion correction method based on a multi-medium refraction model, comprising: S1, acquiring raw underwater image data and extracting optical interference features from the data to obtain a preliminary distortion parameter set; S2, analyzing the influence of medium changes based on the preliminary distortion parameter set, using a convolutional neural network model to process the image pixel distribution and integrating pixel intensity adjustment and local contrast enhancement to determine the degree of refraction distortion; S3, if the degree of refraction distortion exceeds a preset threshold, applying an adaptive correction algorithm to the image pixel distribution combined with refraction deviation compensation and dynamic threshold setting to obtain a corrected image version; S4, obtaining object contour data from the corrected image version and integrating region segmentation processing and brightness equalization to determine whether the object offset is within an acceptable range. S5. The offset calibration result is used to update the standardized process parameters. The image edges are smoothed by Gaussian filtering algorithm, and kernel size selection and standard deviation parameters are applied to determine the final recognition accuracy index. S6. If the final recognition accuracy index meets the detection capability requirements, the standardized process parameters are integrated to generate a general processing template and embed a noise suppression mechanism and iterative optimization loop to obtain the optimized image output. S7. The hidden object features are analyzed based on the optimized image output and combined with edge preservation filtering and bilateral filtering expansion to determine the degree of improvement in security inspection efficiency and obtain a complete underwater inspection report. S8. The adaptive correction algorithm parameters are fed back through the complete underwater inspection report and the gradient calculation integration and smoothing intensity control are integrated to update the medium change model and obtain the basis for the next cycle of image processing.

[0020] This method establishes an underwater image correction framework based on a multi-medium refraction model. First, the acquired raw underwater images typically contain complex refraction and distortion caused by multi-medium light propagation (e.g., air-glass-water). By extracting optical interference features from the image, distortion parameters affecting image quality can be initially obtained. Second, a convolutional neural network (CNN) is used to model and learn the image pixel distribution, combined with pixel intensity adjustment and local contrast enhancement strategies to more accurately reflect the distribution pattern of refraction distortion. Whether to enter the adaptive correction stage depends on whether the determined degree of refraction distortion exceeds a preset threshold. During the correction stage, the accuracy of restoring the image's geometric structure is effectively improved by introducing a refraction deviation compensation mechanism and a dynamic threshold setting based on image characteristics. Subsequently, object contour extraction is performed on the corrected image, and region segmentation algorithms and brightness equalization processing are combined to analyze the target object's offset. The obtained offset calibration results are further used to update standardized image processing parameters, followed by Gaussian filtering for edge smoothing. The final recognition accuracy of the image is calculated under kernel size and standard deviation control. If the recognition accuracy meets the requirements, a general template is generated by combining the current process parameters, noise suppression mechanism, and iterative optimization process to output an optimized image. Finally, by analyzing the features of hidden objects in the optimized image, and using techniques such as edge-preserving filtering and bilateral filtering to expand the analysis depth, the security inspection efficiency is evaluated and a complete inspection report is output. The analysis results in the report are used to update the correction algorithm and media change model in the image processing flow, providing a basis for the next cycle.

[0021] This method effectively addresses the nonlinear distortion problem in underwater images caused by complex refraction paths. By combining a deep learning model with image enhancement strategies, the accuracy and correction effect of refraction distortion estimation are improved. The adaptive correction algorithm features dynamic threshold adjustment and refraction deviation compensation, significantly enhancing the accuracy of image structure restoration. Gaussian filtering optimizes the edge transition process, ensuring image clarity and recognition stability. Edge preservation and bilateral filtering techniques demonstrate strong detail preservation capabilities in analyzing hidden object features, resulting in underwater inspection reports that outperform traditional methods in both recognition rate and analysis efficiency. Furthermore, a feedback mechanism enables iterative updates of algorithm parameters and the model, improving overall adaptability and continuous optimization capabilities, significantly enhancing the practicality and robustness of underwater image processing in complex environments.

[0022] S1 includes: acquiring raw underwater image data through sensors to obtain an initial image set; using image preprocessing techniques to denoise the initial image set to generate a denoised image set; extracting optical interference features from the denoised image set to obtain an interference feature set; if the feature values ​​in the interference feature set exceed a preset threshold, classifying the feature values ​​to determine the interference category; calculating distortion parameters based on the interference category to generate an initial distortion parameter set; using principal component analysis to reduce the dimensionality of the initial distortion parameter set to obtain an optimized distortion parameter set; and using a support vector machine algorithm to verify the parameters of the optimized distortion parameter set to obtain the final distortion parameter set.

[0023] In one possible implementation, a dedicated underwater image acquisition sensor is first deployed on an underwater platform or mobile device (such as an underwater robot). The sensor is preferably a high-resolution, low-light CMOS image sensor with a waterproof housing and red-blue filter lenses, capable of image acquisition under low-light and turbid water conditions. The acquisition operation is initiated upon control command, with the sensor continuously acquiring image data at a frequency of 1 to 5 frames per second for a duration set to 60 seconds, obtaining at least 60 still image frames, which are then combined to form an initial image set. All images are encoded in a standard image format (such as JPEG or PNG) and stored in a local storage module.

[0024] Subsequently, preprocessing is performed on each frame of the initial image set, primarily to remove random noise caused by underwater suspended particles, sensor noise, or light flicker. The preferred image preprocessing steps include: First, applying median filtering to each frame with a 3×3 pixel window. The gray values ​​of the nine neighboring pixels centered on each pixel are sorted, and the center pixel's gray value is replaced with the median, thus suppressing salt-and-pepper noise. Second, histogram equalization is performed on the entire image to enhance contrast and highlight edge areas, making subsequent feature recognition more stable. Third, the image color shift value is calculated based on the RGB channels. If the shift exceeds 10 gray levels, color recorrection based on the channel mean is performed. After preprocessing, a set of images without significant noise or lighting anomalies is obtained, which is the denoised image set.

[0025] Based on the denoised image set, optical interference feature extraction is performed frame-by-frame. This process begins by identifying high-brightness areas through local gray-level difference calculation: the entire image is scanned in a 5×5 window, and the difference between the maximum and average brightness within each window is calculated. If this difference exceeds 20 gray levels, it is marked as a high-reflectivity area. Next, gradient analysis is performed on the image, using the Sobel operator to extract edge intensity maps in the horizontal and vertical directions. The mean edge gradient is calculated for each image region; if the average gradient of a region is below 5, it is identified as a blurred region. Then, the local mean differences of the three RGB channels are analyzed. If the difference between any two channels exceeds 15, it is determined to be a color-distorted region. These three types of regions are recorded as different types of optical interference features, labeled with their image location, interference type, and numerical characteristics. Finally, the information of all interference regions forms an interference feature set, where each record includes the image number, region location, interference type, corresponding feature value, and a description of the extraction method.

[0026] Each feature value in the interference feature set is evaluated individually. These feature values ​​include brightness deviation, mean edge gradient, and color channel difference, corresponding to three typical interference types: high reflectivity, blurring, and color distortion, respectively. The following preset thresholds are used: a brightness deviation threshold of 20 (if the brightness value in a certain image area is more than 20 gray levels higher than the average of adjacent areas), it is marked as high reflectivity interference; a mean edge gradient threshold of 5 (if this value is below 5, the area is considered to have blurring); and a color channel difference threshold of 15 (if the difference between the mean values ​​of any two RGB channels is greater than 15, it is determined to be color distortion interference). These thresholds are derived from a large number of manually labeled samples to ensure applicability to various underwater imaging environments.

[0027] If any feature value exceeds its corresponding threshold, the classification rule is immediately invoked to categorize the interference type corresponding to that region as reflective, blurry, or color distorted. The classification rule is implemented using a static mapping table, which directly maps the interference type based on the feature index exceeding the threshold. Multiple interference regions may exist in each image. After processing each region individually, the classification results of all regions are merged to calculate the number of each type of interference and its average feature value in each image.

[0028] Based on this, corresponding distortion parameters are generated for each type of interference. Specifically, for reflective interference, the distortion parameter is set as the ratio of the mean brightness deviation to the average brightness of the entire image; for blurring interference, it is set as the difference between the mean edge gradient of the blurred area in the image and the average gradient of the entire image; for color distortion, it is set as the standard deviation of the RGB three-channel differences. These parameters are obtained using image pixel-level statistical calculations, with one to two quantitative indicators output for each type as a distortion description. The distortion parameters in multiple images are arranged sequentially to form the initial set of distortion parameters.

[0029] The initial set of distortion parameters is usually high-dimensional, which is not conducive to subsequent algorithm processing. Therefore, principal component analysis (PCA) is introduced for dimensionality reduction. The PCA process is as follows: First, calculate the covariance matrix of all parameter terms in the set; second, find the eigenvalues ​​and corresponding eigenvectors of the covariance matrix; third, sort the eigenvalues ​​from high to low and select the top principal components, such that the cumulative variance contribution rate of these principal components is greater than 85%. Usually, the first three principal components are selected as representatives. The multidimensional distortion parameters of each original image are then projected into a 3D vector, resulting in the optimized set of distortion parameters. This set significantly reduces the amount of data while retaining the main distortion features.

[0030] To verify the accuracy and discriminative ability of the optimized distortion parameter set, a supervised learning method, Support Vector Machine (SVM), is introduced for parameter validation. The specific steps are as follows: 200 labeled "distorted images" and "non-distorted images" are selected from historical data. The optimized distortion parameters corresponding to these images are used as training set sample features, and the image labels are used as supervisory label inputs. An SVM model is trained using a linear kernel function. After training, the model is used to predict and classify each parameter vector in the current optimized distortion parameter set. If the consistency rate between the prediction result and the manually labeled result exceeds 95%, the parameter set is considered validated and confirmed as the final distortion parameter set. Otherwise, feedback is returned to the initial feature extraction or threshold setting stage for correction. The final distortion parameter set serves as the core input for image correction, accuracy adjustment, and template generation.

[0031] S2 includes obtaining refractive index change data from medium change characteristics, using an optical flow estimation algorithm to analyze the interference of refractive index change on image pixel distribution, and obtaining a pixel offset set; if the offset value in the pixel offset set exceeds a preset threshold, cluster analysis is performed on the offset value to determine the offset category; based on the offset category, a gradient descent algorithm is used to optimize the pixel intensity adjustment parameters to obtain an adjusted pixel intensity set; the adjusted pixel intensity set is processed by a local contrast enhancement algorithm to generate an enhanced image set; local texture features are extracted from the enhanced image set, and a feature matching method is used to compare it with the original image to obtain a texture distortion distribution; if the distortion value in the texture distortion distribution exceeds a preset threshold, the distortion value is weighted and fused to generate a refractive distortion degree parameter; the refractive distortion degree parameter is correlated with the medium change characteristics to determine the final refractive distortion degree.

[0032] In one possible implementation, an environmental monitoring module is first integrated into the underwater image acquisition device. This module records the temperature, salinity, and depth information in real time during image acquisition, and records these as characteristics of medium changes. By consulting existing tables of optical medium physical parameters, the temperature, salinity, and depth are converted into underwater refractive index values. This refractive index change data is used to establish a correspondence for each image frame. For example, when the water temperature is 25 degrees Celsius, the salinity is 35 units, and the depth is 10 meters, the refractive index corresponding to the image is calculated to be 1.3408 by looking up the refractive index table.

[0033] After refractive index extraction, the refractive index change is used as one of the boundary constraints of the optical flow estimation algorithm. The Lucas-Kanade pyramid-based optical flow algorithm is employed to estimate pixel motion between images. This process includes the following steps: First, a four-layer image pyramid is constructed, and the image size is reduced sequentially. Then, starting from the smallest layer, each pixel is selected in the image, and the gray-level difference between two consecutive frames is calculated within a 3×3 window around it. Based on the assumptions of gray-level invariance and local spatial consistency, the horizontal and vertical displacement vectors of the pixel are derived. After preliminary estimation at all image layers, the image is upsampled layer by layer to the original resolution for refinement compensation, finally obtaining the complete displacement data of each pixel. The combination of all pixel displacements constitutes the pixel offset set.

[0034] Next, a threshold of 2 pixels was set for pixel offset. That is, if the total offset distance of a pixel in two consecutive frames is greater than 2 pixels, the pixel is considered to have significant interference caused by medium refraction. This threshold was derived by analyzing a large-scale underwater image dataset and can effectively distinguish pixel changes caused by natural motion and physical refraction. The entire pixel offset set was traversed, and all pixels with offset values ​​greater than 2 were marked and their offset vectors were recorded.

[0035] These significant offset vectors are fed into a K-means clustering algorithm for cluster analysis, with a preset number of three clusters corresponding to three common offset patterns: "centralized offset," "dispersed marginal offset," and "mixed offset." The clustering process uses the horizontal and vertical components of each offset vector as two-dimensional feature inputs, employing Euclidean distance as a similarity metric. Iteratively, the distance between each sample point and the three initial cluster centers is calculated, and samples are assigned to the nearest cluster center. The cluster center positions are then updated, and this process is iterated until all centers are stable. Finally, the cluster category of each pixel offset point is output, forming an offset category label map. The spatial distribution ratio and center offset intensity of each type of offset are statistically analyzed to provide a classification basis for subsequent optimization processing.

[0036] Based on the offset category label map generated by the previous clustering step, the entire image is divided into several regions with different offset types. Specifically, they are divided into three categories: the first category is concentrated offset regions, which usually appear in the image center or at the edge of the target; the second category is scattered edge offset regions, which are mostly located at the image boundaries; and the third category is mixed offset regions, which are irregularly distributed. Each type of region is processed separately, with different pixel intensity optimization targets set.

[0037] For each image region identified as having an aberration, a gradient descent optimization model is constructed using the grayscale values ​​of all pixels within the region as the variables to be optimized. The optimization objective is to minimize the mean square error between the grayscale values ​​of pixels within the region and the average grayscale values ​​of the surrounding regions, thereby improving local grayscale consistency and eliminating brightness distortion caused by refraction. The gradient descent process is as follows: First, the pixel intensity adjustment parameter is initialized to the grayscale value of the corresponding pixel in the original image, with a learning rate of 0.01 and a loss function of the squared error between each pixel within the region and the average value of its neighborhood. Then, for each pixel, the following operations are performed iteratively: the derivative of the pixel intensity with respect to the loss function is calculated, which is the product of the current grayscale value minus twice the neighborhood average value, multiplied by the learning rate, and used to correct the current pixel value. After each iteration, the pixel grayscale value is updated, and it is determined whether the change in the loss function is less than 0.001 or whether the number of iterations has reached 50. If either condition is met, the iteration terminates. This optimization process is executed in parallel by region, and the final output is the optimized set of pixel intensities, where the grayscale value of each pixel has been specifically adjusted according to the offset distribution characteristics.

[0038] After pixel intensity optimization, local contrast enhancement is performed on the entire image. The method is as follows: For each pixel, a 5×5 neighborhood is extracted, and the difference between the maximum and minimum grayscale values ​​within this region is calculated as the local contrast value of the current pixel. If this difference is less than a set minimum contrast threshold (e.g., 10 grayscale levels), the region is considered to have low contrast, and enhancement is performed. The enhancement method involves renormalizing the current pixel's grayscale value to the dynamic range of the local window. Specifically, the calculation is: subtract the local minimum value from the current grayscale value, multiply by 255, and divide by the difference between the maximum and minimum values ​​to obtain the new enhanced pixel value. This process is performed pixel-by-pixel, ensuring that each pixel receives personalized contrast enhancement based on its local grayscale environment, thereby improving the visibility of image details and the overall visual hierarchy.

[0039] Finally, the images after the above gradient descent optimization and local contrast enhancement processes constitute an enhanced image set, which is used for subsequent processing steps such as texture feature extraction and distortion analysis.

[0040] Local texture features are extracted frame-by-frame from the generated set of enhanced images. This process employs a local binary mode (LVM) approach for texture description. For each enhanced image frame, it is divided into 4×4 image blocks. Within each block, the grayscale information of the 3×3 neighborhood of each pixel is extracted. By comparing the grayscale value of the center pixel with that of its eight surrounding pixels, values ​​greater than or equal to the center pixel are recorded as 1, and values ​​less than are recorded as 0, forming an 8-bit binary code. This code is then converted to a decimal integer and used as the local texture feature value for that pixel. The average local texture feature values ​​of all pixels within each image block are calculated to form a block-level texture feature map of the enhanced image.

[0041] Then, the texture feature maps of the enhanced images are matched one-to-one with the corresponding original images. The original images are extracted using the same local binary mode method to ensure consistency in feature extraction. During the matching process, the texture feature values ​​of each corresponding image block are subtracted, and the absolute difference between the two is calculated. If this difference is greater than a set texture distortion judgment threshold, the block is considered to have texture structural distortion. This texture distortion threshold is set to 20, and experiments have verified that this value can stably distinguish between natural texture variations and structural perturbations caused by refraction. The distortion values ​​of all image blocks are recorded as a texture distortion distribution map, where each block corresponds to a specific distortion value, and is marked whether it exceeds the threshold.

[0042] After obtaining the complete texture distortion distribution, a weighted fusion operation is performed on all distortion values ​​to generate a uniform refractive distortion parameter. The weighted fusion uses a multi-factor weighted average method with the following weights: texture intensity difference accounts for 50%, meaning the larger the texture difference, the higher the weight; image patch area accounts for 30%, meaning the larger the patch's coverage area, the more significant its distortion impact; spatial continuity accounts for 20%, meaning if multiple adjacent patches have continuous distortion, the overall weight of that area is increased. The calculation method is as follows: first, each factor is normalized, converting its value to between 0 and 1, then multiplied by its corresponding weight and summed to obtain the fusion score for each image patch. Finally, the average score of all image patches is taken as the refractive distortion parameter for the entire image. The value range of this parameter is 0 to 100, with higher values ​​indicating more severe refractive interference in the image.

[0043] Finally, the calculated refractive distortion parameters are correlated with the medium change characteristics during image acquisition. The goal of this correlation analysis is to determine whether the refractive distortion in the image is primarily caused by medium change. Specifically, based on the water temperature, salinity, and depth information recorded during image acquisition, the corresponding refractive index variation range is obtained from a table and compared with the refractive distortion of images with the same refractive index in historical data to calculate the correlation coefficient. If the refractive distortion parameters of the current image show a high linear correlation with the refractive index change in the statistical data, and the correlation coefficient is greater than 0.85, then the current refractive distortion is confirmed to primarily originate from medium change, and the image is marked as an "effective refractive distortion image." Its final refractive distortion level is then output to guide subsequent adaptive image correction processing.

[0044] S3 includes obtaining pixel distribution data from the original image, analyzing the impact of refraction distortion on pixel distribution using an optical flow estimation algorithm to obtain a pixel offset set; if the offset value in the pixel offset set exceeds a preset threshold, an adaptive correction algorithm is used to process the pixel offset set, combined with refraction deviation compensation, to generate a corrected pixel set; based on the corrected pixel set, a dynamic threshold setting method is used to adjust the local feature contrast of the corrected pixel set to obtain an enhanced image set; texture features are extracted from the enhanced image set, and a feature matching method is used to compare it with the original image to determine the final image version.

[0045] In one possible implementation, the grayscale information of all pixels is first extracted from the original image to construct a complete pixel distribution data matrix. Each pixel is identified by its row and column coordinates, with a grayscale value ranging from 0 to 255, and stored in 8-bit unsigned integer format. Image frames are read and paired with the previous frame to form pixel pairs. A pyramid dense optical flow estimation algorithm is used to analyze the impact of refraction distortion on the pixel distribution. The specific steps are as follows: First, a four-layer image pyramid is constructed for both the original and reference images, with the bottom layer being the original image and the top layer being an image scaled down by a factor of four. Second, starting from the top layer of the pyramid, a 3x3 neighborhood is extracted centered on each pixel. The grayscale difference and gradient direction between the current image and the reference image are calculated, and the horizontal and vertical displacements of the pixel are iteratively solved according to optical flow constraints. Third, the optical flow results are upsampled layer by layer to the original resolution to obtain a two-dimensional offset vector for each pixel within the complete image range. The offset value is the Euclidean distance, calculated as the square root of the square of the horizontal component plus the square of the vertical component. The offset values ​​of all pixels are combined to form a pixel offset set, stored as a two-dimensional array.

[0046] To determine the presence of significant refraction effects, a shift threshold of 2 pixels was set. If a pixel's shift distance exceeds 2 pixels in consecutive image frames, it is considered an abnormal shift. This threshold was determined through annotation and analysis of images acquired at different underwater depths and with different refractive indices. When the shift exceeds 2 pixels, the image typically exhibits obvious geometric misalignment, meeting the criteria for refractive distortion. All pixels exceeding this threshold were marked, and an abnormal shift region mask was created for subsequent correction processing.

[0047] After entering the correction process, an adaptive correction algorithm is used to process abnormal offset regions. Each abnormal pixel undergoes a reverse mapping operation based on its offset vector, that is, the pixel is projected from the current image back to its theoretical position in the reference image in the reverse direction. Simultaneously, based on environmental parameters recorded during underwater image acquisition, such as water temperature, salinity, and depth, the corresponding refractive index value for the image is obtained from a table, with common values ​​between 1.3330 and 1.3400. Combining the distance from the image sensor to the imaging surface and the light incident angle, the theoretical refraction path deviation is calculated, converted into pixel coordinate compensation, and superimposed on the reverse mapping result to form the corrected pixel position. All corrected pixels are then recombined into a corrected pixel set.

[0048] To enhance the visual appeal of this image set during post-processing, a local feature contrast enhancement operation is performed. Specifically, the entire image is divided into 8x8 blocks, and the maximum, minimum, and average grayscale values ​​are calculated for each block. The dynamic threshold is set as 0.7 times the difference between the maximum and minimum values; for example, if the grayscale range of a block is 40, the dynamic threshold is 28. Then, a linear stretching operation is performed on all pixels within the block, standardizing the grayscale values ​​according to the dynamic threshold interval, thereby improving the grayscale contrast of that block. This process ensures that each image region undergoes adaptive enhancement based on its actual grayscale distribution characteristics, resulting in an enhanced image set.

[0049] Subsequently, texture consistency verification was performed on the enhanced image set. The gray-level co-occurrence matrix method was used for texture feature extraction, calculating a texture feature vector for each image patch, including four statistical indicators: energy, contrast, entropy, and correlation. After extraction, the feature vectors at the same location in the enhanced image and the original image were compared, and the difference was calculated using Euclidean distance. A texture difference threshold of 20 was set; if the difference between the feature vectors of any two corresponding image patches exceeded 20, the texture consistency of that patch was considered insufficient after correction and needed to be re-evaluated; if the difference was within 20, the correction was considered effective. This threshold was determined through extensive comparative experiments and can accommodate a certain range of visual variations while maintaining the realism of the image texture structure.

[0050] Finally, a weighted average of the texture differences of all image patches is calculated, with the weights set equally according to the area size of each patch. The texture consistency score of the entire image is then output, and the image version with the highest texture score is taken as the final image version.

[0051] S4 includes obtaining object contour data from the corrected image version, extracting the boundary pixel set of the target object using an edge detection algorithm to obtain an object contour dataset; based on the object contour dataset, performing region segmentation processing on the corrected image version using a region segmentation algorithm to generate a region segmentation dataset containing the target object; extracting brightness distribution features from the region segmentation dataset, and adjusting the brightness of the region segmentation dataset using a histogram equalization method to obtain a brightness equalization image set; if the object offset value in the brightness equalization image set exceeds a preset threshold, then using a texture matching algorithm to compare with the object contour dataset to determine whether the object offset is within an acceptable range, and obtaining the offset calibration result.

[0052] In one possible implementation, the complete image pixel matrix is ​​first read from the corrected image version and fed into the edge detection module as input data. The Canny edge detection algorithm is used to extract the boundaries of potential target objects in the image, specifically including the following steps: First, Gaussian blur is applied to the input image to remove noise, with a convolution kernel size of 5×5 and a standard deviation of 1.4. These parameters have been experimentally determined to effectively remove background interference while preserving edge structure. Second, the gradient magnitude and direction of each pixel are calculated and used to detect edge intensity and direction, respectively. Third, non-maximum suppression is applied, comparing the neighborhood of each pixel along its gradient direction and retaining only local maxima as edge candidate points. Fourth, a dual-threshold algorithm is used to connect the edge candidate points, with a low threshold of 50 and a high threshold of 150, ensuring that weak edges are preserved when connected to strong edges, forming a complete edge image. The set of points with a pixel value of 255 in this edge image is the boundary pixel set. All coordinate points are extracted and recorded as a two-dimensional coordinate array, forming an object contour dataset.

[0053] Next, guided by this contour dataset, the region segmentation module is entered. A region-growing-based image segmentation algorithm is preferred, with the following steps: Each contour point is used as a seed point, and its grayscale value is obtained as a growth reference value. The grayscale of adjacent pixels in the image is checked sequentially along an 8-neighborhood direction. If the difference between the grayscale of an adjacent pixel and the reference value is within 10, the pixel is included in the current region. The pixel grayscale difference threshold of 10 was determined by analyzing the grayscale variation range of the target region in 100 representative underwater images, effectively separating the target region from the background. The region growing process continues to expand until the current region can no longer expand outwards, repeating this process until all seed points have been grown. Finally, the image is divided into several regions, and the region containing the target contour is marked as the target region. The pixel coordinate information of each region is summarized to generate a region segmentation dataset, with the data format being a region number matrix corresponding to each pixel.

[0054] Subsequently, brightness distribution features are extracted region by region from the region segmentation dataset. The brightness feature of each region is composed of the gray values ​​of all pixels within that region, and a gray-level histogram for each region is calculated. To enhance the image's brightness hierarchy, histogram equalization is performed on each region. The processing flow is as follows: First, the cumulative distribution function of the region's gray-level histogram is calculated; second, the original gray values ​​are mapped to a range of 0 to 255 based on the cumulative probability, achieving a balanced redistribution of gray values; third, the gray values ​​of all pixels in that region are updated, and the final output is a set of brightness-equalized images. No new parameters are introduced in this step, and the gray-level range and mapping formula used are standard image processing specifications.

[0055] Next, the system detects whether objects in the brightness-equalized image set have shifted position. The method is as follows: based on the object contour data extracted during the edge detection stage, the center point of the smallest bounding rectangle surrounding the object is calculated to obtain the object's center coordinates in the original image; this calculation is then repeated for the same object region in the equalized image to obtain new center coordinates. The Euclidean distance between the two sets of center coordinates is calculated as the object offset value. If the offset value is greater than 5 pixels, the object is considered to have undergone significant displacement during processing, requiring offset calibration. This threshold is based on statistics from 1000 manually annotated images; when the offset exceeds 5 pixels, the user can visually perceive a positional inconsistency, serving as a reliable basis for offset identification.

[0056] If the offset exceeds the threshold, a texture matching algorithm is invoked to calibrate the object's position. This algorithm uses the object's outline region in the original image as a template and performs region-by-region texture matching within a 32×32 pixel sliding window in the brightness-equalized image. During each matching process, the similarity between the current window and the template region in the texture dimension is compared. Four texture feature values—contrast, energy, entropy, and correlation—are extracted using the gray-level co-occurrence matrix method, and the Euclidean distance to the corresponding index of the template is calculated. The matching position with the smallest distance is considered the calibrated object position. If the smallest matching distance is below the threshold of 20, the current offset is considered within an acceptable range, and the offset calibration result is output. This texture error threshold of 20 is also experimentally set to ensure that the matching position does not produce significant visual errors.

[0057] S5 includes processing the offset calibration data using a parameter adjustment algorithm based on the offset calibration results to generate an updated parameter dataset; obtaining the kernel size and standard deviation from the updated parameter dataset, and smoothing the image edges using a Gaussian filtering algorithm to obtain a set of smoothed edge images; if the matching degree between the texture features of the smoothed edge image set and the preset texture template is lower than a preset threshold, a texture matching algorithm is used for secondary calibration to determine the calibrated image dataset; processing the calibrated image dataset using a region segmentation algorithm, extracting the target object region, calculating the comprehensive index of brightness distribution and edge sharpness, and obtaining the final recognition accuracy index.

[0058] In one possible implementation, the remaining offset, the edge intensity of the corresponding pixel, and the local brightness difference are read pixel by pixel from the offset calibration results. First, the average, maximum, and ninth percentile values ​​of the remaining offset are statistically analyzed pixel by pixel across the entire image. Then, the average and standard deviation of the edge intensity, as well as the brightness range of that region, are statistically analyzed within each local region centered on the edge. Subsequently, the parameter adjustment algorithm generates an updated parameter dataset. This algorithm does not rely on empirical verbal descriptions but uses a fixed decision mapping table: when the average remaining offset within the image does not exceed 2, the kernel size is set to 3 and the standard deviation is set to 0.8; when the average is greater than 2 but not exceeding 4, the kernel size is set to 5 and the standard deviation is set to... 1.2; When the mean is greater than 4 but not exceeding 6, the kernel size is set to 7 and the standard deviation is set to 1.8; when the mean is greater than 6, the kernel size is set to 9 and the standard deviation is set to 2.3. The above threshold range is obtained by statistically analyzing the error distribution of 1000 labeled samples to ensure that the filtering intensity increases synchronously as the offset increases to stabilize the edges. Subsequently, the kernel size and standard deviation are extracted from each image according to the updated parameter dataset, and Gaussian filtering is performed to smooth the edges. Specifically, a square neighborhood with the same kernel size is established for each pixel. The weight of each neighboring pixel is read from the pre-generated weight table according to the distance and standard deviation. The gray level of the neighboring pixel is multiplied by the corresponding weight, and the result is summed and normalized using the sum of all weights. To obtain a new center pixel value, the image boundary is processed by copying boundary pixels to avoid size changes. After traversal, a set of smoothed edge images is obtained. Then, texture features are calculated for each smoothed image and compared with a preset texture template. Texture features are calculated within each 32x32 pixel non-overlapping block. For each block, four statistical indicators—contrast, entropy, energy, and correlation—are extracted. The preset texture template is a set of four indicators extracted using the same method from representative underwater target images acquired under standard conditions. The matching degree is calculated as the sum of the absolute differences between the four indicators and the corresponding indicators of the template, divided by four to obtain the average difference. If this average difference exceeds a threshold of 20, the texture of that block is determined to deviate from the template and requires secondary calibration. 20 is determined by statistical analysis of the subjectively perceptible threshold values ​​of 500 texture samples, achieving an optimal balance between false positives and false negatives. For blocks identified as deviating, a texture matching algorithm is performed for secondary calibration. The algorithm establishes a search window within an 8-pixel range horizontally and vertically with the center of the block as the origin, recalculates four texture indices at each position and compares them with the template, and takes the position with the smallest average difference as the calibration position. The displacement of this position relative to the original position is used for pixel coordinate fine-tuning. At the same time, the average gray level of the block is compared with the average gray level of the template block and linear stretching is performed to make the two consistent. To avoid abrupt changes at the block boundary, after completing the displacement and brightness calibration, the block is subjected to another calibration with a kernel size of 3 and a standard deviation of 0.8. Gaussian micro-smoothing; After all blocks requiring secondary calibration are processed, the calibrated image dataset is obtained and enters the region segmentation algorithm to extract the target object region. The region segmentation algorithm uses the edge intensity map as a guide to scan the entire image pixel by pixel. First, pixels with edge intensities in the first two fractions of the entire image are marked as candidate boundaries. Then, connected pixels within the candidate boundaries are used as the starting set to expand the region inwards and outwards. The expansion criterion is that the difference between the average gray level of the adjacent pixel and the current region does not exceed 12 and the difference between the average edge intensity of the adjacent pixel and the current region does not exceed 15. If both conditions are met, the pixel is included in the current region; otherwise, the expansion stops until all pixels in the entire image are reached. All pixels are labeled to a specific region or as background. Isolated regions smaller than 50 pixels are merged into their largest surrounding neighboring region to eliminate fragmentation. The final output is a region segmentation result containing only the target object and its background. After obtaining the target object region, a comprehensive index of brightness distribution and edge sharpness is calculated to form the final recognition accuracy index. The calculation process is as follows: first, calculate the brightness standard deviation of all pixels in the target region and divide the standard deviation by 255 to obtain the brightness non-uniformity ratio. Then, subtract this ratio from 1 and multiply by 100 to obtain the brightness uniformity score. Subsequently, calculate the average gradient intensity at the boundary of the target region, divide this average value by 255 and multiply by 1. The edge sharpness score is obtained from 0.00. Finally, the brightness uniformity score is added to the edge sharpness score with a weight of 0.4 and a weight of 0.6 to obtain a single numerical value for the final recognition accuracy index, ranging from 0 to 100. The higher the value, the better the brightness consistency and contour sharpness required for recognition. All parameters mentioned in this section are clearly given and explained one by one: the kernel size is 3, 5, 7, or 9, read from the updated parameter dataset to determine the filter neighborhood size; the standard deviation is 0.8, 1.2, 1.8, or 2.3 to control the weight decay rate; both are determined by a fixed mapping based on the interval of the average remaining offset; the texture block size is 3. A 2x32 matrix is ​​used to stabilize the statistics and ensure spatial positioning accuracy. The search range is 8 pixels horizontally and vertically to limit spatial displacement during secondary calibration and prevent mismatches. A texture matching threshold of 20 is used to distinguish between consistent and inconsistent textures with the template. A grayscale difference threshold of 12 and an edge intensity difference threshold of 15 are used to ensure consistency in lighting and structure within the region. A fragment incorporation threshold of 50 pixels is used to remove non-target micro-regions. The weights of the comprehensive index are brightness 0.4 and edge 0.6 to highlight the dominant role of edges in recognition. All of the above values ​​were determined and fixed in the configuration through statistical analysis and cross-validation of no less than 1000 sample images.

[0059] S6 includes obtaining standardized process parameters from a preset parameter database, generating an initial general processing template through a parameter integration algorithm, and determining the preliminary structure of the initial general processing template; if the detection capability of the initial general processing template is lower than a preset threshold, then a noise suppression algorithm is used to process the parameter dataset in the initial general processing template to obtain a noise-suppressed template parameter set; based on the noise-suppressed template parameter set, an iterative optimization loop algorithm is used to adjust the parameters in the template parameter set to obtain an optimized general processing template; the optimized general processing template is processed through a region segmentation algorithm to extract feature regions from the target image and generate an optimized image output.

[0060] In one possible implementation, a standardized set of process parameters closest to the current medium conditions is first read from a preset parameter database. The medium conditions are described by three numerical values: water temperature, salinity, and depth. The database search uses the sum of the absolute differences of these three values ​​as a similarity score for sorting, and selects the three sets of parameters with the lowest scores as a reference set. Each set of parameters includes nine fields: kernel size, standard deviation, contrast enhancement factor, brightness equalization intensity, edge gradient threshold, texture matching threshold, region growth grayscale difference threshold, region growth edge difference threshold, and final recognition accuracy threshold. The value range of each field in the database is set as follows: kernel size is an odd number from 3 to 11, and standard deviation is from 0.8 to 2.5. The following parameters are set: contrast enhancement factor (0.6-0.8), brightness equalization intensity (1-3), edge gradient threshold (30-80), texture matching threshold (10-30), region growing grayscale difference threshold (8-16), region growing edge difference threshold (10-20), and final recognition accuracy threshold (70-90). Then, a parameter integration algorithm is executed to perform weighted fusion of the three sets of reference parameters with fixed weights of 0.5, 0.3, and 0.2. These weights are determined through historical statistics of at least 1000 sample images to consistently achieve higher average recognition accuracy. After fusion, an initial general processing template is obtained, and its preliminary structure is determined. The structural order is: denoising and enhancement, refraction distortion estimation, ... The process involves pixel-level correction, local contrast enhancement, texture consistency determination, region segmentation, and index calculation, with the nine fused parameters bound step-by-step for subsequent processing. Next, the detection capability of the initial general processing template is evaluated on verification images from the same medium, with at least 30 verification images. The detection capability is defined as the arithmetic mean of the final recognition accuracy index output from the preceding process on these images. If this mean is less than a preset threshold of 80, the noise suppression algorithm is initiated; otherwise, the template is directly retained for subsequent output. The noise suppression algorithm performs two steps on the template parameter dataset: the first step is outlier removal, where parameters distributed along the verification set dimension are below the second percentile. Values ​​above the 98th percentile are truncated to their corresponding quantiles to remove occasional extreme fluctuations. The second step is sequence smoothing, which uses median smoothing with a window size of 3 along the image index dimension to eliminate short-term parameter jitter, resulting in a noise-suppressed template parameter set. Based on this, an iterative optimization loop algorithm is executed to perform a finite neighborhood search on the template parameters and improve the detection capability round by round. The iterative process includes three steps: candidate generation, template execution, and index evaluation. In the candidate generation stage, the kernel size is searched within one feasible odd number above and below the current value, the standard deviation is searched within 0.4 above and below the current value with a step size of 0.2, and the contrast enhancement factor is searched within 0.1 above and below the current value with a step size of 0.The search process involves several steps: searching with a step size of 1 for brightness equalization intensity (between 1 and 3), searching with a step size of 5 for edge gradient threshold (within 10 above and below the current value), searching with a step size of 5 for texture matching threshold (between 10 and 30), searching with a step size of 2 for region growing grayscale difference threshold (between 8 and 16), and searching with a step size of 2 for region growing edge difference threshold (between 10 and 20). The number of candidate groups is limited to no more than 50 to control computational load. During template execution, the template process is fully run on the same validation set for each candidate parameter group, and the final recognition accuracy of each image is recorded. In the performance evaluation phase, the average recognition accuracy of each candidate parameter group is calculated, and the highest accuracy is selected as the new template parameter. If the average accuracy of this round is lower than the previous round... If the improvement is less than 1 and remains less than 1 for three consecutive iterations, the process stops early, or is forcibly stopped when the number of iterations reaches 20, outputting the optimized general processing template. The methods for determining all parameters are clearly defined in the workflow. Kernel size determines the pixel size of the filtering neighborhood and is positively correlated with image structural complexity; standard deviation controls the decay rate of the filtering weights with distance and is positively correlated with residual noise intensity; contrast enhancement factor limits the local contrast stretching to prevent over-enhancement; brightness equalization intensity determines the intensity of histogram redistribution to balance the dynamic range of dark and bright areas; edge gradient threshold filters stable edge pixels before region segmentation to reduce erroneous expansion; and texture matching threshold distinguishes pixels from the template. Consistency and inconsistency directly trigger secondary texture correction. A region growth grayscale difference threshold is used to limit brightness consistency within the same region, while a region growth edge difference threshold ensures structural consistency within the same region. Finally, a recognition accuracy threshold determines whether the current template meets deployment standards. All these thresholds are determined through statistical analysis of at least 1000 training samples and at least 100 validation samples, combined with the human-perceptible limit, and are stored in a database to ensure consistency. After template optimization, the optimized general processing template is used to perform a region segmentation algorithm on the target image to generate an optimized image output. The specific process of region segmentation involves first generating a candidate boundary map based on the edge gradient threshold, and then using the candidate boundaries as a guide to perform region segmentation. The process involves region growing, with growth criteria including the difference between the average grayscale of adjacent pixels and the current region not exceeding a region growth grayscale difference threshold, and the difference between the average edge intensity of adjacent pixels and the current region not exceeding a region growth edge difference threshold. Expansion stops if either condition is not met. After expansion, fragments smaller than 50 pixels are merged into the region with the largest adjacent area to obtain a regular target region outline. Finally, the ratio of the standard deviation to the mean of brightness within the target region is calculated to obtain a brightness uniformity score, and the average gradient intensity on the target boundary is calculated to obtain an edge sharpness score. These scores are linearly combined with a brightness weight of 0.4 and an edge weight of 0.6 to form a single numerical final recognition accuracy index, which is output along with the optimized image for subsequent compliance determination or ranking decisions.

[0061] S7 includes acquiring target underwater images from a preset image database, processing the target underwater images using a bilateral filtering algorithm to preserve edge details, and obtaining a first optimized image; for the first optimized image, using an image segmentation algorithm to separate the foreground and background regions, extracting hidden object regions, and obtaining a segmentation feature set; based on the segmentation feature set, using a convolutional neural network algorithm to analyze the hidden object features, determining whether a target object exists, and obtaining an object recognition result; by comparing the object recognition result with a preset efficiency threshold, using a report generation algorithm to integrate the object recognition result and efficiency data, and obtaining an underwater inspection report.

[0062] In one possible implementation, the target underwater image is first read from a preset image database according to the task identifier and unified to a fixed size and fixed grayscale range. Then, bilateral filtering is performed to suppress noise while preserving edge details. The spatial neighborhood diameter of the bilateral filter is set to 9, the spatial range parameter is set to 5, and the intensity range parameter is set to 20. The boundary processing method is to copy the boundary pixels and calculate a weighted average pixel by pixel for the entire image to output a first optimized image. Next, the first optimized image is segmented to separate the foreground and background and extract the hidden object region. The specific process is to first calculate the grayscale histogram of the entire image and then use fixed threshold segmentation and adaptive segmentation. The combined strategy involves a global threshold of 128 for initial screening of foreground pixels, and an adaptive threshold within each 32x32 pixel local block obtained by subtracting 10 from the sum of the median and range of gray levels within the block. Pixels must simultaneously meet both threshold criteria to be classified as foreground pixels. Morphological denoising is then performed on the foreground mask, with one iteration of opening operations to remove isolated noise points and one iteration of closing operations to fill small holes. Four-neighbor connected component analysis is then performed, with connected components less than 200 pixels directly discarded. The remaining connected components are cropped using the minimum bounding rectangle, and their position, area, aspect ratio, average gray level, and edge density are recorded to form a segmentation feature set. Based on... The segmented feature set is fed into a convolutional neural network algorithm for hidden object feature analysis and target object determination. The input for the inference stage is a foreground candidate block cropped to 256x256 pixels after being scaled proportionally along its longer side and normalized to zero mean within the range using the channel mean and channel standard deviation. The network structure is fixed, consisting of five convolutional layers and three downsampling layers stacked alternately, followed by two fully connected layers with modified linear units as the activation function. Overfitting is suppressed by a dropout layer with a dropout ratio of 0.5. The two-class probability output of the last layer is used as the determination criterion, with a determination threshold of 0.8. A target object is determined to exist when the target class probability of any candidate block is not lower than 0.8. The system outputs the presence markers and corresponding candidate block location information and probability scores from the object recognition results. Then, it compares the object recognition results with a preset efficiency threshold and generates an underwater inspection report. The efficiency data consists of two parts: one is the processing time (in seconds), representing the total time from image reading to obtaining the object recognition result; the other is the recognition quality score, expressed as a percentage based on the target probability of the candidate blocks, with the highest value taken when multiple blocks exist. These two scores are added together with fixed weights to obtain a single efficiency score. The time weight is 0.4, derived by setting the maximum acceptable time to 5 seconds and linearly mapping the actual time to a percentage. The quality weight is 0.The value of 6 is the mass fraction obtained by multiplying the probability by 100. The preset efficiency threshold is 80. When the efficiency score is not lower than 80, the report conclusion is "meets the standard"; otherwise, it is "not met the standard," and the report lists the time taken, probability, reason for not meeting the standard, and suggested parameter range. To ensure that each step is reproducible and the parameters are interpretable, all parameters involved in this step and their determination methods are as follows: Spatial neighborhood diameter 9 and spatial range parameter 5 are used to limit the spatial influence range of bilateral filtering to avoid cross-boundary blurring; intensity range parameter 20 is used to control the weight attenuation across grayscale boundaries to retain true edges; global threshold 128 is used to quickly distinguish foreground and background under balanced lighting conditions; adaptive threshold block size 32x32 is used to stably segment underwater under uneven lighting; intra-block correction 10 is used to suppress false detections of weak noise; morphological iterations are 1 each to clean up isolated noise and small holes without excessively changing the shape; minimum area of ​​connected components 200 is used to filter out accidental noise spots without mistakenly deleting small targets. This value is determined by a certain number of... The algorithm uses statistical analysis of 1000 samples to determine and fix the configuration. Five statistical measures of the segmentation feature set are used to provide discriminative prior structural information to the convolutional neural network without introducing additional modules. The input size of 256x256 is used to balance memory usage and accuracy. Modified linear unit activation is used to improve convergence speed without introducing letter symbols. A dropout ratio of 0.5 is used to disable the regularization during inference while maintaining training consistency. The target probability threshold of 0.8 is determined by the inflection point of the receiver operating characteristic curve on the validation set. An efficiency threshold of 80 ensures that the deployment requirements are met in terms of both processing time and recognition quality. The upper limit of the linear mapping for the time subset (5 seconds) comes from the real-time processing requirements of the device, and the quality subset directly uses a probability percentage to maintain interpretability. The report generation algorithm has fixed output fields including image identifier, processing time, highest target probability, candidate location, efficiency score, threshold comparison conclusion, and suggested parameter range, and is archived as an underwater inspection report in a single-page structured text format.

[0063] S8 includes acquiring target underwater images from a preset underwater image database, processing the target underwater images using an adaptive correction algorithm, and generating a first corrected image by adjusting brightness and contrast; for the first corrected image, calculating the pixel gradient of the image using a gradient calculation method, extracting edge features, and generating an edge feature set; based on the edge feature set, adjusting the feature smoothness using a smoothing intensity control algorithm, optimizing the edge features, and obtaining an optimized feature set; comparing the optimized feature set with a preset medium change database, if the matching degree of the optimized feature set exceeds a preset threshold, updating the medium change model using a convolutional neural network algorithm to obtain the basis for the next cycle of image processing.

[0064] In one possible implementation, the target underwater image is first read from a preset underwater image database according to the task identifier, and the image resolution is standardized to 512 by 512 pixels and the grayscale range to 0 to 255 to ensure consistency in the processing flow. Then, the adaptive correction algorithm stage begins. This algorithm uses the global grayscale histogram of the image as input, first calculating the mean, standard deviation, and grayscale distribution skewness, and then determining the brightness and contrast adjustment parameters: when the mean is below 100, the brightness increase is set to 255 minus 10% of the mean; when the mean is above 180, the brightness decrease is set to the mean minus 180; when the standard deviation is below 40, the contrast amplification factor is set to 1.3; when the standard deviation is above 80, the contrast compression factor is set to 0.8; otherwise, the parameters remain unchanged. The adjustment parameters are applied pixel by pixel, and the first corrected image is output.

[0065] For the first corrected image, a gradient calculation method is used to extract edge features. Specifically, the Sobel operator is used to calculate gradient values ​​in both the horizontal and vertical directions. The gradient strength is the square root of the sum of the squares in both directions, and the direction is calculated using the arctangent. After traversing the entire image, points with gradient strengths greater than 20 are identified as valid edge points, and their coordinates and directions are recorded. The edge feature set is then output. The threshold of 20 was determined through gradient statistical analysis of no less than 500 image samples, which effectively distinguishes background noise from true edges.

[0066] Subsequently, a smoothing intensity control algorithm is applied to optimize the edge feature set. The algorithm first calculates the gradient variance of each edge point within its local 5x5 neighborhood. If the variance is greater than 50, it indicates the presence of edge spikes in the region, requiring smoothing. If the variance is less than 10, it indicates over-sharpening of the edge, requiring a reduction in smoothing intensity. Otherwise, the default smoothing intensity of 1.0 is maintained. Based on this judgment, the smoothing intensity is dynamically adjusted between 0.5 and 1.5, and a weighted average of the gradient values ​​of edge points is performed within the neighborhood to generate an optimized feature set.

[0067] Finally, the optimized feature set is compared with a pre-defined medium change database. The database stores statistical templates of edge features under different water temperatures, salinities, and depths, including features such as average gradient intensity, orientation distribution histogram, and edge density. The matching degree is calculated using the cosine similarity method. If the similarity between the optimized feature set and the database template exceeds a threshold of 0.85, the medium change feature corresponding to the image is considered known, and a convolutional neural network algorithm is invoked to update the medium change model. The input to the convolutional neural network is the feature vector of the optimized feature set, and the output is a new set of medium change parameters, including refractive index correction coefficients, optical attenuation factors, and local scattering weights. After the update, these parameters are stored in the model as the basis for the next image processing cycle. The threshold of 0.85 is determined by plotting subject operating characteristic curves on a validation set of 1000 sample images to ensure a balance between sensitivity and specificity.

[0068] The upper and lower limits of brightness adjustment are determined by the mean threshold of 100 and 180, respectively; the contrast adjustment is determined by the standard deviation threshold of 40 and 80; the edge detection threshold is 20; the gradient variance threshold for smoothing intensity control is 10 and 50; and the final database matching similarity threshold is 0.85. All of the above parameters were determined and fixed through large-scale sample statistical experiments.

[0069] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for underwater image rectification based on multi-medium refraction model, characterized in that, The method comprises the following steps: S1, obtaining a preliminary distortion parameter set by collecting original underwater image data and extracting optical interference features in the data; S2, analyzing the influence of medium changes based on the preliminary distortion parameter set, processing image pixel distribution using a convolutional neural network model, and integrating pixel intensity adjustment and local contrast enhancement to determine the degree of refraction distortion; S3, if the degree of refraction distortion exceeds a preset threshold, applying an adaptive correction algorithm to image pixel distribution combined with refraction deviation compensation and dynamic threshold setting to obtain a corrected image version; S4, obtaining object contour data from the corrected image version and integrating region segmentation processing and brightness equalization to determine whether the object offset is within an acceptable range to obtain an offset calibration result; S5, updating standardized process parameters using the offset calibration result, smoothing image edges through a Gaussian filter algorithm, and applying kernel size selection and standard deviation parameters to determine the final recognition accuracy index; S6, if the final recognition accuracy index meets the detection capability requirements, integrating the standardized process parameters to generate a general processing template and embedding a noise suppression mechanism and an iterative optimization loop to obtain an optimized image output; S7, analyzing hidden object features based on the optimized image output and combining edge-preserving filtering and bilateral filtering extension to determine the degree of security efficiency improvement to obtain a complete underwater inspection report; S8, updating the medium change model by feeding back the adaptive correction algorithm parameters and integrating gradient calculation and smoothing intensity control based on the complete underwater inspection report to obtain the basis for the next cycle of image processing.

2. The method of claim 1, wherein the method is based on a multi-medium refraction model. The S1 comprises: Collecting underwater original image data through a sensor to obtain an initial image set; Using image preprocessing techniques to denoise the initial image set to generate a denoised image set; Extracting optical interference features from the denoised image set to obtain an interference feature set; If the feature values in the interference feature set exceed a preset threshold, classifying the feature values to determine the interference category; According to the interference category, calculate the distortion parameters to generate an initial distortion parameter set; Using principal component analysis algorithm to reduce the dimension of the initial distortion parameter set to obtain an optimized distortion parameter set; Using support vector machine algorithm to verify the parameters for the optimized distortion parameter set to obtain the final distortion parameter set.

3. The method of claim 1, wherein the method is based on a multi-medium refraction model. The S2 comprises: Obtaining refractive index change data from the medium change features, using optical flow estimation algorithm to analyze the interference of refractive index change on image pixel distribution to obtain a pixel offset set; If the offset value in the pixel offset set exceeds a preset threshold, clustering the offset value to determine the offset category; According to the offset category, using gradient descent algorithm to optimize the pixel intensity adjustment parameter to obtain the adjusted pixel intensity set; Processing the adjusted pixel intensity set through local contrast enhancement algorithm to generate an enhanced image set; Extracting local texture features from the enhanced image set and comparing them with the original image using feature matching method to obtain texture distortion distribution; If the distortion value in the texture distortion distribution exceeds a preset threshold, weighting and fusing the distortion value to generate a refraction distortion degree parameter; Correlating the refraction distortion degree parameter with the medium change features to determine the final refraction distortion degree.

4. The method of claim 1, wherein the method is based on a multi-medium refraction model. The S3 comprises: Pixel distribution data is acquired from the original image, a light flow estimation algorithm is used to analyze the influence of refraction distortion on the pixel distribution, and a pixel offset set is obtained; If the offset value in the pixel offset set exceeds a preset threshold, an adaptive correction algorithm is used to process the pixel offset set, combined with refraction deviation compensation, to generate a corrected pixel set; According to the corrected pixel set, a dynamic threshold setting method is used to adjust the local feature contrast of the corrected pixel set, and an enhanced image set is obtained; Texture features are extracted from the enhanced image set, and a feature matching method is used to compare with the original image to determine the final image version.

5. The method of claim 1, wherein: The S4 includes: Object contour data is acquired from the corrected image version, an edge detection algorithm is used to extract the boundary pixel set of the target object, and an object contour data set is obtained; According to the object contour data set, a region segmentation algorithm is used to process the corrected image version, and a region segmentation data set containing the target object is generated; Luminance distribution features are extracted from the region segmentation data set, and a histogram equalization method is used to adjust the luminance of the region segmentation data set, and a luminance equalization image set is obtained; If the object offset value in the luminance equalization image set exceeds the preset threshold, a texture matching algorithm is used to compare with the object contour data set to determine whether the object offset is within an acceptable range, and an offset calibration result is obtained.

6. The method of claim 1, wherein: The S5 includes: According to the offset calibration result, a parameter adjustment algorithm is used to process the offset calibration data to generate an updated parameter data set; The core size and standard deviation value are obtained from the updated parameter data set, a Gaussian filter algorithm is used to smooth the image edge, and a smooth edge image set is obtained; If the texture feature of the smooth edge image set is lower than the preset texture template threshold, a texture matching algorithm is used for secondary calibration to determine the calibrated image data set; The target object region is extracted by processing the calibrated image data set through the region segmentation algorithm, and the comprehensive index of luminance distribution and edge definition is calculated to obtain the final recognition accuracy index.

7. The method of claim 1, wherein the method is based on a multi-medium refraction model. The S6 includes: Standardized process parameters are obtained from the preset parameter database, and an initial general processing template is generated by a parameter integration algorithm to determine the preliminary structure of the initial general processing template; If the detection capability of the initial general processing template is lower than the preset threshold, a noise suppression algorithm is used to process the parameter data set in the initial general processing template to obtain a noise-suppressed template parameter set; According to the noise-suppressed template parameter set, an iterative optimization loop algorithm is used to adjust the parameters in the template parameter set to obtain an optimized general processing template; The feature region is extracted from the target image by processing the optimized general processing template through the region segmentation algorithm to generate an optimized image output.

8. The method of claim 1, wherein: The S7 includes: A target underwater image in the preset image database is acquired, and a bilateral filter algorithm is used to process the target underwater image to retain edge details to obtain a first optimized image; For the first optimized image, an image segmentation algorithm is used to separate the foreground and background regions, and the hidden object region is extracted to obtain a segmentation feature set; According to the segmentation feature set, a convolutional neural network algorithm is used to analyze the hidden object features, to determine whether there is a target object, and to obtain an object recognition result; By comparing the object recognition result with a preset efficiency threshold, a report generation algorithm is used to integrate the object recognition result and the efficiency data, to obtain an underwater inspection report.

9. The method of claim 1, wherein: The S8 comprises: A target underwater image in a preset underwater image database is acquired, an adaptive correction algorithm is used to process the target underwater image, and a first corrected image is generated by adjusting the brightness and contrast; For the first corrected image, a gradient calculation method is used to calculate the image pixel gradient, to extract the edge features, and to generate an edge feature set.

10. The method of claim 9, wherein the method is based on a multi-medium refraction model. The S8 further comprises: According to the edge feature set, a smoothing strength control algorithm is used to adjust the feature smoothing degree, to optimize the edge features, and to obtain an optimized feature set; By comparing the optimized feature set with a preset medium change database, if the matching degree of the optimized feature set exceeds a preset threshold, a convolutional neural network algorithm is used to update the medium change model, to obtain the basis for the next cycle of image processing.

Citation Information

Patent Citations

  • Underwater three-dimensional measurement data correction method based on structured light three-dimensional measurement

    CN111006610A

  • Underwater image enhancement method and device based on Retinex algorithm

    CN111275644A