Underwater target recognition method and system based on 4K visible light polarization imaging

By using 4K visible light polarization imaging technology, combined with pixel-level registration of polarization parameter matrix and RGB image and multi-band attenuation compensation, the problem of recognition accuracy of traditional underwater target recognition methods in complex water environments has been solved, and high-precision, real-time underwater target recognition has been achieved.

CN120564023BActive Publication Date: 2025-10-21SICHUAN NATIONAL INNOVATION VISION UHD VIDEO TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511080042.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-10-21
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

Traditional underwater target recognition methods suffer from problems such as low imaging contrast, blurred details, and color distortion in complex water environments. Single-mode sensors are insufficient to meet the requirements for high-precision recognition, and existing multi-band attenuation compensation models do not take into account dynamic environmental changes, resulting in reduced recognition accuracy.

Method used

Employing 4K visible light polarization imaging technology, RGB intensity images and polarization parameter matrices are acquired synchronously, pixel-level registration is performed, and then encoded into four-dimensional light field tensor data. Combining cross-modal attention mechanism and multi-band attenuation weighting technology, high-resolution spectral intensity information and polarization physical features are extracted, light attenuation is dynamically compensated, a joint enhanced feature map is generated, and input into a lightweight recognition network.

Benefits of technology

It achieves high-resolution target recognition in complex underwater environments, improves the contrast between the target and the background, accurately quantifies the material reflection characteristics, suppresses scattering noise interference, and realizes efficient and accurate identification of target location, category and material information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564023B_ABST
    Figure CN120564023B_ABST
Patent Text Reader

Abstract

The application provides an underwater target recognition method and system based on 4K visible light polarization imaging, and relates to the technical field of underwater target recognition.The application uses pixel-level registration RGB and polarization four-dimensional light field tensor encoding technology to synchronously capture high-resolution spectral intensity information and polarization physical characteristics of a target, overcomes the limitation of information loss of single modal data in low-light or high-turbidity scenes, designs polar coordinate symmetry of a ring-shaped direction-sensitive convolution kernel, combines polarization degree matrix decomposition technology of the Fresnel reflection law, accurately quantifies the optical reflection characteristic difference of a material, dynamically learns the correlation weight of spatial features and polarization features based on a cross-modal attention mechanism, strengthens the joint response of target edge contours and internal textures, adopts a lightweight recognition network architecture, integrates a polarization-sensitive convolution kernel and a dynamic channel reparameterization module, realizes parallel and efficient inference of target position, category and material information, and realizes accurate recognition of underwater targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of underwater target recognition, and in particular to an underwater target recognition method and system based on 4K visible light polarization imaging. Background Art

[0002] Underwater target recognition is a core technology in fields such as underwater robot navigation and ecological monitoring. However, traditional methods face significant challenges in practical application due to the complex optical properties of water. Water's strong scattering and absorption of light, particularly the rapid attenuation of blue-green wavelengths over long distances, leads to low contrast, blurred details, and color distortion in underwater imaging. Furthermore, random scattering interference from suspended particles and microorganisms further exacerbates image degradation, making it difficult for traditional algorithms based on visible light imaging to accurately extract target outlines and texture features in turbid waters. Existing technologies often rely on single-modality sensors to acquire data, such as spectral information from RGB cameras or acoustic signals from sonar equipment. While RGB images can provide color and texture information, the signal-to-noise ratio drops sharply in low-light or high-turbidity scenarios, making it difficult to effectively distinguish targets from background. While sonar technology offers long-range detection capabilities, its limited spatial resolution makes it difficult to meet the requirements for high-precision recognition at close range. Existing methods generally neglect the use of high-resolution polarization data. 4K visible light imaging technology can capture the subtle textures and geometric structures of underwater scenes at ultra-high resolution. Combined with polarization imaging, it analyzes the directional characteristics of light wave vibrations, revealing the reflective properties and microscopic geometric features of the target surface material, thereby providing an additional dimension of physical characteristics for target recognition in complex environments. Traditional multi-band attenuation compensation models typically use fixed parameters to calculate light propagation attenuation, failing to account for the nonlinear effects of dynamic changes in water turbidity, light intensity, and target distance on light attenuation characteristics, thereby reducing the accuracy of underwater target recognition.

[0003] Therefore, it is necessary to provide an underwater target recognition method and system based on 4K visible light polarization imaging to solve the above technical problems. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention provides an underwater target recognition method and system based on 4K visible light polarization imaging, which achieves the beneficial effect of underwater target recognition that integrates multimodal polarization imaging and dynamic environmental parameter adaptation.

[0005] The present invention provides an underwater target recognition method based on 4K visible light polarization imaging, comprising:

[0006] S1: Synchronously collect the RGB intensity image and polarization parameter matrix of the underwater scene, perform pixel-level registration on the RGB intensity image and polarization parameter matrix, and then encode them into four-dimensional light field tensor data;

[0007] S2: Decompose the four-dimensional light field tensor data to obtain RGB channel data and polarization parameter channel data, and extract the spatial feature map of the RGB channel data and the polarization feature map of the polarization parameter channel data;

[0008] S3: The spatial feature map and the polarization feature map are interactively processed through the cross-modal attention mechanism to generate a joint enhanced feature map;

[0009] S4: Based on the real-time acquired water turbidity, target distance, light intensity parameters and scene depth map, a multi-band attenuation weight map is generated in combination with the pre-established water type and attenuation parameter mapping table;

[0010] S5: Perform multi-scale fusion and normalization on the multi-band attenuation weight map to generate a spatial attention mask. Based on the spatial attention mask, adjust the joint enhanced feature map by pixel-by-pixel multiplication to obtain a fused feature map.

[0011] S6: Input the fused feature map into the lightweight recognition network and output the target's location bounding box, category label and material type information.

[0012] Preferably, in step S1, the pixel-level registration processing step includes:

[0013] The polarization parameter matrix is ​​bilinearly interpolated through a pre-calibrated spatial transformation matrix to align the polarization parameter matrix with the pixel coordinates of the RGB intensity image one by one.

[0014] Preferably, in step S2, the step of extracting the polarization characteristic map includes:

[0015] Extract the polarization angle matrix from the polarization parameter channel data, perform circular direction-sensitive convolution calculation on the polarization angle matrix, and generate a surface normal change rate feature map;

[0016] Extract the polarization degree matrix from the polarization parameter channel data, decompose the polarization degree matrix into specular reflection component and diffuse reflection component based on the Fresnel reflection law, and generate a material reflection characteristic feature map;

[0017] The surface normal change rate feature map and the material reflection characteristic feature map are spliced ​​along the channel dimension to output a polarization feature map.

[0018] Preferably, the weight distribution of the convolution kernel calculated by the annular direction-sensitive convolution satisfies radial symmetry in a polar coordinate system.

[0019] Preferably, the radial weight gradient of the convolution kernel calculated by the annular direction-sensitive convolution satisfies a positive correlation with the refractive index of the material.

[0020] Preferably, in step S3, the step of generating the joint enhanced feature map includes:

[0021] Compress the spatial feature map and polarization feature map in channel dimension to generate compressed feature vectors;

[0022] Calculate the similarity matrix of the compressed feature vector and generate the spatial and polarization joint attention weights through normalization;

[0023] The spatial and polarization joint attention weights are used to perform channel weighting and spatial recalibration on the spatial feature map and the polarization feature map to generate a joint enhanced feature map.

[0024] Preferably, in step S4, obtaining the multi-band attenuation weight map includes the following steps:

[0025] Calculate the underwater light propagation path length of each pixel based on the target distance and the pixel depth value of the scene depth map;

[0026] Based on the turbidity of the water body, the absorption coefficient and scattering coefficient of the red, green and blue bands of the corresponding water body type are obtained from the pre-established water body type and attenuation parameter mapping table;

[0027] Add the absorption coefficient and scattering coefficient of the red, green and blue bands band by band to obtain the total attenuation coefficient of each band;

[0028] Calculate the ratio of the light intensity parameter to the preset light intensity standard reference value to obtain the light intensity correction factor;

[0029] Based on the total attenuation coefficient, path length matrix and light intensity correction factor, the weight values ​​of the red, green and blue bands are calculated pixel by pixel and spliced ​​into a three-channel matrix to obtain a multi-band attenuation weight map.

[0030] Preferably, in step S4, during the generation of the multi-band attenuation weight map, when the target distance is less than the preset distance threshold and the turbidity is lower than the preset turbidity threshold, the green band weight calculation is turned off, and the equivalent green band attenuation coefficient is reconstructed by interpolating the red and blue band weights.

[0031] Preferably, step S5 further includes adjusting the spatial attention mask, including:

[0032] Perform a morphological dilation operation on the spatial attention mask to generate a dilated mask;

[0033] Calculate the difference area between the expanded mask and the original spatial attention mask to generate a boundary difference matrix;

[0034] The boundary difference matrix is ​​superimposed on the original spatial attention mask to generate the adjusted spatial attention mask.

[0035] The present invention also provides an underwater target recognition system based on 4K visible light polarization imaging, which is applied to an underwater target recognition method based on 4K visible light polarization imaging, comprising:

[0036] The data acquisition and registration module is used to synchronously collect the RGB intensity image and polarization parameter matrix of the underwater scene, perform pixel-level registration processing on the RGB intensity image and polarization parameter matrix, and then encode them into four-dimensional light field tensor data;

[0037] A feature map extraction module is used to decompose the four-dimensional light field tensor data to obtain RGB channel data and polarization parameter channel data, and to extract the spatial feature map of the RGB channel data and the polarization feature map of the polarization parameter channel data;

[0038] The cross-modal feature fusion module is used to interactively process the spatial feature map and the polarization feature map through the cross-modal attention mechanism to generate a joint enhanced feature map;

[0039] A multi-band attenuation weight generation module is used to generate a multi-band attenuation weight map based on the real-time acquired water turbidity, target distance, light intensity parameters and scene depth map, combined with a pre-established water type and attenuation parameter mapping table;

[0040] The fusion feature map generation module is used to perform multi-scale fusion and normalization on the multi-band attenuation weight map to generate a spatial attention mask. Based on the spatial attention mask, the joint enhancement feature map is adjusted by pixel-by-pixel multiplication to obtain a fused feature map.

[0041] The target recognition and output module is used to input the fused feature map into the lightweight recognition network and output the target's location bounding box, category label and material type information.

[0042] Compared with related technologies, the underwater target recognition method and system based on 4K visible light polarization imaging provided by the present invention has the following beneficial effects:

[0043] The present invention utilizes pixel-level registered RGB and polarization four-dimensional light field tensor encoding technology to synchronously capture the high-resolution spectral intensity information and polarization physical properties of the target, including polarization angle and polarization degree, overcoming the limitation of single modal data in information loss in low-light or high-turbidity scenes and improving the contrast between the target and the background; through the polar coordinate symmetry design of the circular direction-sensitive convolution kernel, the target surface normal change rate characteristics are extracted, and combined with the polarization degree matrix decomposition technology based on the Fresnel reflection law, the specular reflection and diffuse reflection components are decoupled, and the differences in the optical reflection characteristics of the material are accurately quantified, thereby realizing the characterization of the physical properties of different materials; based on the cross-modal attention mechanism, the correlation weights of spatial features and polarization features are dynamically learned, and the target edge is adaptively enhanced through channel compression and spatial recalibration operations. The joint response of contours and internal textures can stably extract key target areas, especially under the obstruction of suspended matter in the water or the interference of dynamic light and shadow. Through the multi-band attenuation weight generation technology driven by real-time water parameters, combined with the pre-calibrated water optical property mapping table and the dynamic light intensity correction factor, the red, green and blue band attenuation compensation weights are calculated pixel by pixel, and the multi-scale fusion strategy is used to separate the global attenuation trend and local detail noise, effectively suppressing the scattering noise interference of turbid waters while retaining the faint features of distant targets. A lightweight recognition network architecture is adopted, integrating polarization-sensitive convolution kernels and dynamic channel reparameterization modules, and optimizing network weights through end-to-end training to achieve parallel and efficient inference of target position, category and material information, ultimately achieving the purpose of underwater target recognition based on 4K visible light polarization imaging. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is a flow chart of an underwater target recognition method based on 4K visible light polarization imaging according to the present invention;

[0045] Figure 2 This is a module structure diagram of an underwater target recognition system based on 4K visible light polarization imaging of the present invention. DETAILED DESCRIPTION

[0046] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all of the structures. Furthermore, the embodiments of the present invention and the features of the embodiments may be combined with one another unless there is a conflict.

[0047] It should also be noted that, for ease of description, only portions relevant to the present invention are shown in the accompanying drawings, rather than all of the contents. Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the various operations (or steps) as being processed sequentially, many of the operations can be performed in parallel, concurrently, or simultaneously. In addition, the order of the various operations can be rearranged. The process can be terminated when its operations are completed, but may also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0048] Example 1

[0049] A method for underwater target recognition based on 4K visible light polarization imaging, in the specific implementation process, such as Figure 1 , which shows a flow chart of an underwater target recognition method based on 4K visible light polarization imaging, including:

[0050] Step S1: synchronously collect the RGB intensity image and polarization parameter matrix of the underwater scene, perform pixel-level registration processing on the RGB intensity image and the polarization parameter matrix, and then encode them into four-dimensional light field tensor data.

[0051] Specifically, in step S1, the pixel-level registration process includes:

[0052] The polarization parameter matrix is ​​bilinearly interpolated through a pre-calibrated spatial transformation matrix to align the polarization parameter matrix with the pixel coordinates of the RGB intensity image one by one.

[0053] During the specific implementation process, the RGB intensity image of the underwater scene and the polarization parameter matrix containing polarization angle and polarization degree information are first synchronously collected through an imaging system equipped with a high-resolution 4K polarization camera. The RGB camera and the polarization camera use a hardware synchronization trigger mechanism to ensure time consistency; then the two modal data are pixel-level aligned. Specifically, in the laboratory pre-calibration stage, the spatial transformation matrix between the polarization camera and the RGB camera is calculated through multiple sets of calibration plate images. In actual application, bilinear interpolation operation is performed on the polarization parameter matrix based on the spatial transformation matrix, and the coordinates of each pixel point in the polarization parameter matrix are mapped to RG The corresponding position of the B image is determined by the registration method, eliminating the pixel offset caused by lens parallax or installation error, and achieving sub-pixel alignment; the aligned RGB intensity image and polarization parameter matrix are integrated into four-dimensional light field tensor data along the channel dimension, whose dimensions include RGB three-channel intensity values, polarization angle matrix and polarization degree matrix, forming structured input data that integrates spectral intensity and polarization physical properties. Through precise coordinate mapping and interpolation algorithms, the consistency of multimodal data spatial alignment is guaranteed, providing a high-precision data foundation for subsequent feature extraction and fusion. At the same time, the four-dimensional tensor encoding is used to effectively retain the optical properties and geometric detail information of the target.

[0054] Step S2: Decompose the four-dimensional light field tensor data to obtain RGB channel data and polarization parameter channel data, and extract the spatial feature map of the RGB channel data and the polarization feature map of the polarization parameter channel data.

[0055] Specifically, in step S2, the step of extracting the polarization characteristic map includes:

[0056] Extract the polarization angle matrix from the polarization parameter channel data, perform circular direction-sensitive convolution calculation on the polarization angle matrix, and generate a surface normal change rate feature map;

[0057] Extract the polarization degree matrix from the polarization parameter channel data, decompose the polarization degree matrix into specular reflection component and diffuse reflection component based on the Fresnel reflection law, and generate a material reflection characteristic feature map;

[0058] The surface normal change rate feature map and the material reflection characteristic feature map are spliced ​​along the channel dimension to output the polarization feature map.

[0059] Specifically, the weight distribution of the convolution kernel calculated by the annular direction-sensitive convolution satisfies the radial symmetry in the polar coordinate system.

[0060] Specifically, the radial weight gradient of the convolution kernel calculated by the annular direction-sensitive convolution satisfies a positive correlation with the refractive index of the material.

[0061] In the specific implementation process, the four-dimensional light field tensor data is first decomposed into RGB three-channel intensity data and polarization parameter channel data along the channel dimension. The RGB channel data is extracted through a convolutional neural network to extract a spatial feature map containing target texture and color information, and the polarization parameter channel data is further processed into a polarization feature map; for the polarization angle matrix in the polarization parameter channel data, a circular direction-sensitive convolution kernel is used for feature extraction. The weight distribution of the convolution kernel follows the radial symmetry design under the polar coordinate system, and the continuous change characteristics of the target surface normal direction are captured by the convolution weights uniformly distributed along the circumferential direction, generating a surface normal change rate feature map reflecting the surface geometric structure; at the same time, the polarization degree matrix is ​​extracted from the polarization parameter channel data, and the polarization degree value of each pixel point is decomposed into a mirror angle matrix based on the Fresnel reflection law. The reflection component and the diffuse reflection component quantify the differences in the reflection properties of different materials and generate a material reflection characteristic feature map that characterizes the optical properties of the material; the radial weight gradient of the annular direction-sensitive convolution kernel is positively correlated with the refractive index of the material, that is, the radial weight gradient corresponding to the high refractive index material is larger, thereby enhancing the characteristic response of the edges of highly reflective materials such as metal and glass; finally, the surface normal change rate feature map and the material reflection characteristic feature map are spliced ​​along the channel dimension to form a polarization feature map that integrates the geometric structure and material properties, providing highly discriminative polarization physical feature input for subsequent cross-modal attention fusion. Through the physical-driven design of the annular convolution kernel and the polarization decomposition mechanism of the Fresnel law, the problem of insufficient representation of the surface characteristics of underwater targets by traditional methods is effectively solved, laying the foundation for high-precision recognition in complex environments.

[0062] Step S3: The spatial feature map and the polarization feature map are interactively processed through the cross-modal attention mechanism to generate a joint enhanced feature map.

[0063] Specifically, in step S3, the step of generating the joint enhanced feature map includes:

[0064] Compress the spatial feature map and polarization feature map in channel dimension to generate compressed feature vectors;

[0065] Calculate the similarity matrix of the compressed feature vector and generate the spatial and polarization joint attention weights through normalization;

[0066] The spatial and polarization joint attention weights are used to perform channel weighting and spatial recalibration on the spatial feature map and the polarization feature map to generate a joint enhanced feature map.

[0067] In the specific implementation process, the extracted spatial feature map and polarization feature map are first compressed in channel dimension respectively, and the high-dimensional channel information of each feature map is compressed into a low-dimensional feature vector through global average pooling operation to reduce redundant information and retain key modal characteristics; then the similarity matrix between the compressed spatial feature vector and the polarization feature vector is calculated, and the correlation strength of the two modal features at different spatial positions is quantified by element-by-element dot product, and the similarity matrix is ​​converted into a spatial and polarization joint attention weight in the form of probability distribution using a normalization function; based on the generated attention weight, the original spatial feature map and polarization feature map are channel-weighted to enhance the cross-modal features. The key channel responses related to the target are detected, and the spatial dimensions of the feature map are recalibrated at the same time. The response intensity of the target edge, texture and material reflection characteristic areas is enhanced through weight distribution; finally, the spatial feature map and polarization feature map after channel weighting and spatial recalibration are added element by element to generate a joint enhanced feature map that integrates spectral characteristics and polarization physical characteristics. The complementarity of cross-modal features is dynamically captured through the attention mechanism, and the feature conflict problem caused by fixed weight fusion in traditional methods is solved. Especially in scenes with turbid water or target occlusion, it can adaptively enhance effective features and suppress noise interference, providing highly robust fusion feature input for subsequent attenuation compensation and target recognition.

[0068] Step S4: Based on the real-time acquired water turbidity, target distance, light intensity parameters and scene depth map, a multi-band attenuation weight map is generated in combination with a pre-established water type and attenuation parameter mapping table.

[0069] Specifically, in step S4, obtaining the multi-band attenuation weight map includes the following steps:

[0070] Calculate the underwater light propagation path length of each pixel based on the target distance and the pixel depth value of the scene depth map;

[0071] Based on the turbidity of the water body, the absorption coefficient and scattering coefficient of the red, green and blue bands of the corresponding water body type are obtained from the pre-established water body type and attenuation parameter mapping table;

[0072] Add the absorption coefficient and scattering coefficient of the red, green and blue bands band by band to obtain the total attenuation coefficient of each band;

[0073] Calculate the ratio of the light intensity parameter to the preset light intensity standard reference value to obtain the light intensity correction factor;

[0074] Based on the total attenuation coefficient, path length matrix and light intensity correction factor, the weight values ​​of the red, green and blue bands are calculated pixel by pixel and spliced ​​into a three-channel matrix to obtain a multi-band attenuation weight map.

[0075] Specifically, in step S4, during the generation of the multi-band attenuation weight map, when the target distance is less than the preset distance threshold and the turbidity is lower than the preset turbidity threshold, the green band weight calculation is turned off, and the equivalent green band attenuation coefficient is reconstructed by interpolating the red and blue band weights.

[0076] During the specific implementation process, first, based on the real-time target distance and pixel depth value of the scene depth map, the underwater light propagation path length is calculated pixel by pixel to reflect the actual transmission distance of the light from the target to the sensor; combined with the real-time water turbidity parameters, the absorption coefficient and scattering coefficient of the red, green and blue bands in the corresponding water environment are dynamically queried from the pre-established water type and attenuation parameter mapping table, where the pre-established water type and attenuation parameter mapping table is generated by measuring the transmittance data of water samples with different turbidity, salinity and suspended matter concentration in the laboratory, and fitting the attenuation parameters of each band according to the Beer-Lambert law; the absorption coefficient and scattering coefficient of the red, green and blue bands are added band by band to obtain the total attenuation coefficient of each band, and the light intensity is calculated according to the ratio of the real-time light intensity parameter and the preset standard light intensity reference value. Correction factor is used to compensate for the impact of ambient lighting changes on the incident light energy; based on the total attenuation coefficient, light propagation path length matrix and light intensity correction factor, the attenuation weight values ​​of the red, green and blue bands are calculated pixel by pixel, and the weight values ​​of the three bands are spliced ​​into a multi-band attenuation weight map; in specific scenarios, when the target distance is less than the preset distance threshold and the water turbidity is lower than the preset turbidity threshold, the independent calculation process of the green band weight is turned off, and bilinear interpolation is performed through the weight values ​​of the red and blue bands to reconstruct the equivalent green band attenuation coefficient, so as to reduce computing resource consumption and maintain the integrity of multi-band weights, and ultimately achieve efficient attenuation compensation in complex water environments. At the same time, the light intensity correction factor and path length matrix are used to accurately quantify the attenuation differences in different regions, providing a high-precision weight basis for subsequent multi-scale fusion and feature enhancement.

[0077] Step S5: Perform multi-scale fusion and normalization on the multi-band attenuation weight map to generate a spatial attention mask. Based on the spatial attention mask, the joint enhanced feature map is adjusted by pixel-by-pixel multiplication to obtain a fused feature map.

[0078] Specifically, step S5 also includes the step of adjusting the spatial attention mask including:

[0079] Perform a morphological dilation operation on the spatial attention mask to generate a dilated mask;

[0080] Calculate the difference area between the expanded mask and the original spatial attention mask to generate a boundary difference matrix;

[0081] The boundary difference matrix is ​​superimposed on the original spatial attention mask to generate the adjusted spatial attention mask.

[0082] In the specific implementation process, the multi-band attenuation weight map is first subjected to multi-scale fusion processing. For example, the Gaussian pyramid decomposition method is used to decompose the weight map into low-frequency global attenuation components and high-frequency local detail components of different scales. The fusion weights of the components of each scale are dynamically adjusted according to the turbidity parameters of the water body. When the turbidity is high, the weight ratio of the low-frequency component is increased to suppress the interference of high-frequency noise. When the turbidity is low, the contribution of the high-frequency component is increased to retain the target details. The fused multi-scale weight map is then normalized based on the L2 norm to generate a spatial attention mask. The spatial attention mask is further subjected to a morphological dilation operation. For example, a rectangular structure element is used to expand the edge area of ​​the spatial attention mask to generate an expansion mask, and the calculation is performed. The pixel-level difference area between the dilated mask and the spatial attention mask is calculated to extract the boundary difference matrix; the boundary difference matrix is ​​fused into the spatial attention mask in a weighted superposition manner to enhance the attenuation weight distribution of the target edge area and generate an adjusted spatial attention mask; finally, the adjusted spatial attention mask is multiplied pixel by pixel with the joint enhancement feature map, and the background noise is dynamically suppressed and the target feature response is enhanced through the attenuation weight to obtain the optimized fused feature map. Finally, through multi-scale fusion and morphological boundary enhancement strategy, the problem of over-smoothing or blurred edges of traditional attenuation compensation methods in complex water bodies is effectively solved, the feature discrimination of the target area is significantly improved, and the foundation for the high-precision output of the subsequent lightweight recognition network is laid.

[0083] Step S6: Input the fused feature map into the lightweight recognition network and output the target's location bounding box, category label, and material type information.

[0084] During the specific implementation process, the polarization-sensitive convolutional layer in the lightweight recognition network is first used to extract multi-scale features from the input features, dynamically perceive the correlation between the polarization characteristics of the target area and the spectral intensity, and automatically merge redundant channels and strengthen the high-contribution channels based on the polarization sensitivity score of the feature channel; then the target detection and material classification tasks are synchronously processed through a parallel branch structure, where the target detection branch uses the anchor box regression mechanism to predict the target position bounding box and its confidence, and the material classification branch focuses on the characteristic area of ​​the material reflectance characteristics through the attention pooling layer, and outputs the material type label in combination with the pre-trained material decision tree model; during the network training process, the reparameterization technology is used to convert the sparse network after dynamic pruning into an equivalent full-precision dense network to ensure the computational efficiency of the inference stage; the final output includes the detection results of the target position bounding box coordinates, category probability distribution and material type label, and the non-maximum suppression algorithm is used to remove overlapping detection boxes to optimize the recognition accuracy. Finally, through the lightweight network structure and multi-task collaborative optimization mechanism, millisecond-level high-precision inference is achieved to adapt to the real-time target recognition needs of underwater resource-constrained scenarios.

[0085] The working principle of the underwater target recognition method based on 4K visible light polarization imaging provided by the present invention is as follows:

[0086] The present invention synchronously collects high-resolution RGB intensity images and polarization parameter matrices, uses a pre-calibrated spatial transformation matrix to achieve pixel-level registration and encodes them into four-dimensional light field tensor data, ensuring the spatial consistency of spectral information and polarization physical properties; decomposes the four-dimensional light field tensor data of the light field, extracts the spatial feature map of the RGB channel and the polarization feature map of the polarization channel respectively, wherein the polarization feature map extracts the target surface normal change rate feature through the circular direction sensitive convolution kernel, and combines the material reflection characteristic feature decomposed by the Fresnel reflection law to accurately characterize the target geometric structure and optical properties; dynamically fuses the spatial features and polarization features through the cross-modal attention mechanism to generate a joint enhanced feature map, which strengthens the target The method dynamically generates a multi-band attenuation weight map based on real-time water parameters, obtains the absorption and scattering coefficients of different bands through a pre-calibrated mapping table, calculates the attenuation weight pixel by pixel based on the light propagation path length and the light intensity correction factor, adaptively closes the green band calculation in specific scenarios, and reconstructs the equivalent weights through red and blue band interpolation to optimize computing power; multi-scale fusion and morphological edge enhancement technology are used to generate a spatial attention mask, adjust and optimize the fusion feature map, and finally integrates the polarization-sensitive convolution kernel and the dynamic channel pruning strategy through a lightweight recognition network to output the target position, category and material information in parallel, realizing high-precision and high-efficiency target recognition in complex underwater environments.

[0087] Example 2

[0088] An underwater target recognition system based on 4K visible light polarization imaging, in the specific implementation process, such as Figure 2 As shown in FIG, a module structure diagram of an underwater target recognition system based on 4K visible light polarization imaging is shown, including:

[0089] The data acquisition and registration module 100 is used to synchronously acquire the RGB intensity image and polarization parameter matrix of the underwater scene, perform pixel-level registration processing on the RGB intensity image and the polarization parameter matrix, and then encode them into four-dimensional light field tensor data;

[0090] The feature map extraction module 200 is used to decompose the four-dimensional light field tensor data to obtain RGB channel data and polarization parameter channel data, and extract the spatial feature map of the RGB channel data and the polarization feature map of the polarization parameter channel data;

[0091] A cross-modal feature fusion module 300 is used to interactively process the spatial feature map and the polarization feature map through a cross-modal attention mechanism to generate a joint enhanced feature map;

[0092] An attenuation weight map generation module 400 is configured to generate a multi-band attenuation weight map based on the real-time acquired water turbidity, target distance, light intensity parameters, and scene depth map, combined with a pre-established water type and attenuation parameter mapping table;

[0093] A fused feature map generation module 500 is used to perform multi-scale fusion and normalization processing on the multi-band attenuation weight map to generate a spatial attention mask, and based on the spatial attention mask, adjust the joint enhanced feature map by pixel-by-pixel multiplication to obtain a fused feature map;

[0094] The target recognition and output module 600 is used to input the fused feature map into the lightweight recognition network and output the target's position bounding box, category label and material type information.

[0095] The working principle of the underwater target recognition system based on 4K visible light polarization imaging provided by the present invention is as follows:

[0096] The data acquisition and registration module 100 synchronously acquires high-resolution RGB intensity images and polarization parameter matrices, realizes pixel-level alignment through pre-calibrated spatial transformation matrix and bilinear interpolation, and encodes them into four-dimensional light field tensor data; the feature map extraction module 200 decomposes the light field tensor data, extracts the RGB spatial feature map using a convolutional neural network, and decomposes the polarization angle and polarization degree data through a circular direction-sensitive convolution kernel and the Fresnel reflection law to generate a polarization feature map that integrates the surface normal change rate and the material reflection characteristics; the cross-modal feature fusion module 300 dynamically learns the associated weights of spatial features and polarization features based on the attention mechanism, and generates a joint enhanced feature map through channel compression and spatial recalibration; the attenuation weight map generation module 400 combines the actual The fusion feature map generation module 500 uses multi-scale Gaussian pyramid decomposition and morphological boundary enhancement strategies to generate a spatial attention mask that is adaptive to the water environment, suppresses background noise and enhances target features through pixel-by-pixel multiplication. The target recognition and output module 600 is equipped with a lightweight network architecture, integrates polarization-sensitive convolution kernels and dynamic channel pruning technology, processes target positioning, classification and material recognition tasks in parallel, and outputs high-precision bounding boxes, category labels and material type information to achieve real-time and efficient recognition of complex underwater scenes.

[0097] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0098] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, magnetic disk storage, or magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0099] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

Claims

1. A method for underwater target recognition based on 4K visible light polarization imaging, characterized in that: The underwater target recognition method comprises the following steps: S1: Synchronously collect the RGB intensity image and polarization parameter matrix of the underwater scene, perform pixel-level registration on the RGB intensity image and polarization parameter matrix, and then encode them into four-dimensional light field tensor data; S2: Decompose the four-dimensional light field tensor data to obtain RGB channel data and polarization parameter channel data, and extract the spatial feature map of the RGB channel data and the polarization feature map of the polarization parameter channel data; S3: The spatial feature map and the polarization feature map are interactively processed through the cross-modal attention mechanism to generate a joint enhanced feature map; S4: Based on the real-time acquired water turbidity, target distance, light intensity parameters and scene depth map, a multi-band attenuation weight map is generated in combination with the pre-established water type and attenuation parameter mapping table; S5: Perform multi-scale fusion and normalization on the multi-band attenuation weight map to generate a spatial attention mask. Based on the spatial attention mask, adjust the joint enhanced feature map by pixel-by-pixel multiplication to obtain a fused feature map. S6: Input the fused feature map into the lightweight recognition network and output the target's location bounding box, category label and material type information; In step S4, obtaining the multi-band attenuation weight map includes the following steps: Calculate the underwater light propagation path length of each pixel based on the target distance and the pixel depth value of the scene depth map; Based on the turbidity of the water body, the absorption coefficient and scattering coefficient of the red, green and blue bands of the corresponding water body type are obtained from the pre-established water body type and attenuation parameter mapping table; Add the absorption coefficient and scattering coefficient of the red, green and blue bands band by band to obtain the total attenuation coefficient of each band; Calculate the ratio of the light intensity parameter to the preset light intensity standard reference value to obtain the light intensity correction factor; Based on the total attenuation coefficient, path length matrix and light intensity correction factor, the weight values ​​of the red, green and blue bands are calculated pixel by pixel and spliced ​​into a three-channel matrix to obtain a multi-band attenuation weight map.

2. The underwater target recognition method based on 4K visible light polarization imaging according to claim 1, characterized in that: In step S1, the pixel-level registration process includes: The polarization parameter matrix is ​​bilinearly interpolated through a pre-calibrated spatial transformation matrix to align the polarization parameter matrix with the pixel coordinates of the RGB intensity image one by one.

3. The underwater target recognition method based on 4K visible light polarization imaging according to claim 2, characterized in that: In step S2, the step of extracting the polarization characteristic map includes: Extract the polarization angle matrix from the polarization parameter channel data, perform circular direction-sensitive convolution calculation on the polarization angle matrix, and generate a surface normal change rate feature map; Extract the polarization degree matrix from the polarization parameter channel data, decompose the polarization degree matrix into specular reflection component and diffuse reflection component based on the Fresnel reflection law, and generate a material reflection characteristic feature map; The surface normal change rate feature map and the material reflection characteristic feature map are spliced ​​along the channel dimension to output a polarization feature map.

4. The underwater target recognition method based on 4K visible light polarization imaging according to claim 3, characterized in that: The weight distribution of the convolution kernel calculated by the annular direction-sensitive convolution satisfies radial symmetry in a polar coordinate system.

5. The underwater target recognition method based on 4K visible light polarization imaging according to claim 4, characterized in that: The radial weight gradient of the convolution kernel calculated by the annular direction-sensitive convolution satisfies a positive correlation with the refractive index of the material.

6. The underwater target recognition method based on 4K visible light polarization imaging according to claim 5, characterized in that: In step S3, the step of generating the joint enhanced feature map includes: Compress the spatial feature map and polarization feature map in channel dimension to generate compressed feature vectors; Calculate the similarity matrix of the compressed feature vector and generate the spatial and polarization joint attention weights through normalization; The spatial and polarization joint attention weights are used to perform channel weighting and spatial recalibration on the spatial feature map and the polarization feature map to generate a joint enhanced feature map.

7. The underwater target recognition method based on 4K visible light polarization imaging according to claim 6, characterized in that: In step S4, during the generation of the multi-band attenuation weight map, when the target distance is less than the preset distance threshold and the turbidity is lower than the preset turbidity threshold, the green band weight calculation is turned off, and the equivalent green band attenuation coefficient is reconstructed by interpolating the red and blue band weights.

8. The underwater target recognition method based on 4K visible light polarization imaging according to claim 7, characterized in that: Step S5 also includes the following steps for adjusting the spatial attention mask: Perform a morphological dilation operation on the spatial attention mask to generate a dilated mask; Calculate the difference area between the dilation mask and the spatial attention mask to generate a boundary difference matrix; The boundary difference matrix is ​​superimposed on the spatial attention mask to generate the adjusted spatial attention mask.

9. An underwater target recognition system based on 4K visible light polarization imaging, characterized in that: The underwater target recognition method based on 4K visible light polarization imaging according to any one of claims 1 to 8 is applied, wherein the underwater target recognition system comprises: The data acquisition and registration module is used to synchronously collect the RGB intensity image and polarization parameter matrix of the underwater scene, perform pixel-level registration processing on the RGB intensity image and polarization parameter matrix, and then encode them into four-dimensional light field tensor data; A feature map extraction module is used to decompose the four-dimensional light field tensor data to obtain RGB channel data and polarization parameter channel data, and to extract the spatial feature map of the RGB channel data and the polarization feature map of the polarization parameter channel data; The cross-modal feature fusion module is used to interactively process the spatial feature map and the polarization feature map through the cross-modal attention mechanism to generate a joint enhanced feature map; The attenuation weight map generation module is used to generate a multi-band attenuation weight map based on the real-time acquired water turbidity, target distance, light intensity parameters and scene depth map, combined with the pre-established water type and attenuation parameter mapping table; The fusion feature map generation module is used to perform multi-scale fusion and normalization on the multi-band attenuation weight map to generate a spatial attention mask. Based on the spatial attention mask, the joint enhancement feature map is adjusted by pixel-by-pixel multiplication to obtain a fused feature map. The target recognition and output module is used to input the fused feature map into the lightweight recognition network and output the target's location bounding box, category label and material type information.

Citation Information

Patent Citations

  • Underwater target detection and identification method based on projection polarization distance characteristics

    CN117310834A

  • Deep learning underwater polarization image target detection algorithm combined with multi-dimensional information

    CN118196608A