Underwater target detection method and system based on deep learning
Through deep learning, the depth information and reflection patterns of underwater image frames are generated, and semantic segmentation and coordinated mapping are performed, which solves the distortion problem of target detection in underwater environments and improves the robustness and stability of detection.
Patent Information
- Application Number
- CN202510666623.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In underwater environments, existing technologies have difficulty in effectively detecting underwater targets with distortion-resistant properties, especially due to the distortion of visual features caused by light attenuation, water scattering and noise interference, resulting in low detection efficiency.
Through a deep learning-based method, depth information of image frames is generated, the boundary contours and characteristic depths of underwater objects are extracted, the reflection pattern is constructed in combination with surface roughness, semantic segmentation and coordinated mapping are performed, confident semantic labels are generated, artifact noise is removed, and the mask boundaries are dynamically adjusted to adapt to changes in light and water flow.
In complex underwater environments, anti-distortion detection of underwater targets is achieved, which improves the robustness and stability of detection, ensures the continuity and integrity of target contours, and reduces the impact of environmental interference.
Smart Images

Figure CN120656049A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of underwater target detection. More specifically, the present application relates to a method and system for underwater target detection based on deep learning. Background Art
[0002] Underwater target detection, as the core support for marine engineering, environmental monitoring, military security and other fields, focuses on the accurate identification and positioning of various targets (such as submarines, fish, underwater facilities, plankton, etc.) in complex water environments. This technology faces multiple challenges such as the complex optical properties of water (such as light absorption and scattering leading to image blur), multipath effects of sonar signals, variable target morphology and background noise interference. It requires the integration of physical modeling, sensor technology and intelligent algorithms to achieve efficient detection. Its core technical paths include: First, obtaining target echo signals or image data through sonar (such as active sonar, passive sonar), optical cameras, lidar and other sensors. With the development of marine resource development, deep-sea exploration and intelligent navigation technology, underwater target detection technology is evolving towards high precision, low power consumption and multi-scenario adaptability, showing broad application prospects in marine ecological protection, underwater archaeology, oil and gas exploration and national defense security.
[0003] Deep learning-based underwater target detection addresses the challenges of target detection in complex underwater environments (such as image blur caused by light attenuation, feature distortion caused by water scattering, and sonar signal noise interference). By building an end-to-end intelligent model, it achieves efficient recognition and positioning of underwater objects (such as submarines, fish, and underwater facilities). This has become a key breakthrough in current marine technology. The core of this method is to leverage the multi-layer nonlinear mapping capabilities of deep neural networks to automatically learn robust feature representations for underwater targets, avoiding the limitations of traditional methods that rely on handcrafted features. In existing deep learning-based underwater target detection processes, target detection is often performed based on image visual features (such as edges and textures). However, in underwater environments, the visual features of targets are significantly affected by light attenuation and scattering. In addition, differences in surface roughness and reflectivity (specular / diffuse reflection, high / low reflectivity) among similar targets (such as rocks of different materials and different organisms) can cause distortion in the system's visual features, reducing the system's detection efficiency. Therefore, how to perform distortion-resistant detection of underwater targets in the presence of underwater interference has become a challenge facing the industry. Summary of the Invention
[0004] The present application provides a deep learning-based underwater target detection method and system, which can perform distortion-resistant detection of underwater targets under underwater environmental interference.
[0005] In a first aspect, the present application provides a method for underwater target detection based on deep learning, comprising the following steps: Acquire an underwater image sequence of the water area to be detected; Generating depth information of different image frames based on the color attenuation characteristics of the underwater image sequence, and then extracting the boundary contours and corresponding characteristic depths of different underwater objects in each image frame by combining all the depth information through a deep learning algorithm; Based on all feature depths and the surface roughness of various underwater objects, the reflection pattern of underwater light in different water propagation characteristics is constructed. Then, based on the reflection pattern and the structural morphology of the boundary contours of different underwater objects, all boundary contours are semantically divided to obtain the mask areas corresponding to different semantic levels for each image frame; For each image frame, the texture directional features of the mask areas corresponding to different semantic levels are determined. The main reflection directions of underwater light in different mask areas are determined by the brightness gradient distribution of each mask area. Then, a coordinated mapping is performed based on all the texture directional features and all the main reflection directions to obtain the semantic coordination degree of each image frame in each mask area. Confident semantic labels of different underwater objects in the underwater image sequence are generated based on all semantic coordination degrees.
[0006] In some embodiments, generating depth information of different image frames based on the color attenuation characteristics of the underwater image sequence specifically includes: Acquiring color attenuation characteristics of the underwater image sequence; determining the relative optical path length of each pixel in different image frames of the underwater image sequence by using the color attenuation feature; The depth information of different image frames is determined according to the relative optical path lengths of all pixels in each image frame.
[0007] In some embodiments, extracting the boundary contours and corresponding feature depths of different underwater objects in each image frame by combining all depth information with a deep learning algorithm specifically includes: Determine the depth gradient change area in each image frame using all depth information; According to the depth gradient change area in all image frames and combined with the deep learning algorithm, the boundary contours of different underwater objects in each image frame are located, and then the characteristic depth corresponding to each boundary contour is extracted.
[0008] In some embodiments, constructing the reflection pattern of underwater light in different water propagation characteristics based on all characteristic depths and surface roughness of various underwater objects specifically includes: Obtain the surface roughness of various underwater objects; The underwater light reflection response model of water bodies under different propagation conditions is established by taking into account all characteristic depths and the surface roughness of various underwater objects; The reflection pattern of underwater light in different water body propagation characteristics is output according to the underwater light reflection response model.
[0009] In some embodiments, semantically dividing all boundary contours according to the reflection pattern and the structural form of the boundary contours of different underwater objects to obtain mask areas corresponding to different semantic levels of each image frame specifically includes: Obtain the structural morphology corresponding to each boundary contour; Establishing a semantic matching relationship between the reflection pattern and the structural form by combining the structural forms corresponding to each boundary contour with the reflection pattern; All boundary contours are divided into mask areas corresponding to different semantic levels of each image frame according to the semantic matching relationship.
[0010] In some embodiments, determining the main reflection direction of underwater light in different mask areas based on the brightness gradient distribution of each mask area specifically includes: Obtain the brightness gradient distribution of each mask area; Determine the directionality index value of each mask area based on all brightness gradient distributions; The main reflection directions of underwater light in different mask areas are determined based on all directional index values.
[0011] In some embodiments, performing coordinated mapping based on all texture direction features and all main reflection directions to obtain the semantic coordination degree of each image frame in each mask area specifically includes: Determine the directional consistency between the texture directional features and the main reflection direction of each mask area; The semantic coordination degree of each image frame in each mask area is determined through all directional consistencies.
[0012] In a second aspect, the present application provides an underwater target detection system based on deep learning, comprising: An acquisition module, used for acquiring underwater image sequences of the water area to be detected; a processing module for generating depth information of different image frames based on the color attenuation characteristics of the underwater image sequence, and then extracting the boundary contours and corresponding characteristic depths of different underwater objects in each image frame by combining all the depth information through a deep learning algorithm; The processing module is further configured to construct a reflection pattern of underwater light in different water propagation characteristics based on all characteristic depths and the surface roughness of various underwater objects, and then semantically divide all boundary contours according to the reflection pattern and the structural morphology of the boundary contours of different underwater objects to obtain mask areas corresponding to different semantic levels for each image frame; The processing module is further configured to determine, for each image frame, texture directional features of mask areas corresponding to different semantic levels, determine the main reflection directions of underwater light in different mask areas based on the brightness gradient distribution of each mask area, and then perform coordinated mapping based on all texture directional features and all main reflection directions to obtain the semantic coordination degree of each image frame in each mask area; An execution module is configured to generate confident semantic labels for different underwater objects in the underwater image sequence based on all semantic coordination degrees.
[0013] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned deep learning-based underwater target detection method when executing the computer program.
[0014] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the steps of the above-mentioned deep learning-based underwater target detection method.
[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: The deep learning-based underwater target detection method and system provided in the present application obtain an underwater image sequence of a water area to be detected; generate depth information of different image frames based on the color attenuation characteristics of the underwater image sequence, and then extract the boundary contours and corresponding feature depths of different underwater objects in each image frame by combining all the depth information with a deep learning algorithm; construct a reflection pattern of underwater light in different water body propagation characteristics based on all the feature depths and the surface roughness of various underwater objects, and then semantically divide all the boundary contours according to the reflection pattern and the structural form of the boundary contours of different underwater objects to obtain mask areas corresponding to different semantic levels for each image frame; for each image frame, determine the texture direction features of the mask areas corresponding to different semantic levels, determine the main reflection direction of the underwater light in different mask areas through the brightness gradient distribution of each mask area, and then coordinate mapping is performed based on all the texture direction features and all the main reflection directions to obtain the semantic coordination degree of each image frame in each mask area; and generate confidence semantic labels for different underwater objects in the underwater image sequence based on all the semantic coordination degrees.
[0016] It can be seen that in the present application, first, after obtaining an underwater image sequence of the water area to be detected, the boundary contours and corresponding feature depths of different underwater objects in each image frame are extracted based on the color attenuation features of the underwater image sequence, and then the boundary contours of different semantic levels are semantically divided, and the depth information generated by color attenuation is combined with the light reflection pattern of the object surface roughness modeling to accurately outline the regional range (mask area) of various underwater targets, so that the system can adaptively extract the semantic differences between the target object and the background, suspended particles and noise spots in complex environments such as turbid water, light spot interference or water flow disturbance, and remove irrelevant pixels from the mask, so that subsequent feature extraction and semantic coordination are carried out in the high-confidence mask area, effectively reducing the artifact noise caused by sudden changes in ambient illumination, scattering and refraction; that is: by dynamically updating the mask boundary, the system can respond to water flow and light changes in real time, ensure that the extracted target contour is continuous and complete, suppress distortion from the source, and significantly improve The robustness and stability of underwater target detection are enhanced. Subsequently, within each mask region, the system first analyzes the primary reflection direction of underwater light based on the brightness gradient distribution. This is then combined with the texture directional features within the region and their consistency (semantic coordination) is quantified through multidimensional mapping. This method integrates physical optical priors with the boundary predictions of the semantic segmentation model, enabling cross-frame and cross-feature verification of single-frame segmentation results. (In underwater environments, light scattering and diffraction often lead to model prediction drift or edge jitter, making it difficult to accurately locate targets using pixel-level segmentation alone.) The semantic coordination measures the degree of consistency between texture and illumination variations, eliminating anomalous regions that do not conform to the object's true structure and dynamically adjusting the confidence label for each mask region. Finally, when outputting section-level semantic labels, the system prioritizes regions with high coordination, ensuring that the detection results are robust to underwater optical interference. In summary, this solution enables robust detection of underwater targets despite interference from the underwater environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart of a method for underwater target detection based on deep learning according to some embodiments of the present application; Figure 2 is a schematic diagram of a process for determining a depth gradient change area according to some embodiments of the present application; Figure 3 is a schematic diagram of a process for determining a main reflection direction according to some embodiments of the present application; Figure 4 is a schematic structural diagram of an underwater target detection system based on deep learning according to some embodiments of the present application; Figure 5This is a diagram of the internal structure of a computer device for implementing a deep learning-based underwater target detection method according to some embodiments of the present application. DETAILED DESCRIPTION
[0018] In order to better understand the technical solution in this embodiment, the technical solution in this embodiment will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0019] refer to Figure 1 , which is a flow chart of a method for underwater target detection based on deep learning according to some embodiments of the present application. The method 100 for underwater target detection based on deep learning mainly includes the following steps: In step 101, a sequence of underwater images of the water area to be detected is acquired.
[0020] In specific implementation, an underwater image sequence of the water area to be inspected can be obtained by continuous shooting using an industrial-grade underwater camera, a remote-controlled underwater robot (ROV) or an imaging system carried by an autonomous underwater vehicle. Alternatively, an existing underwater image data set can be used to obtain the image, or underwater environment simulation can be performed on a three-dimensional simulation platform to synthesize multiple continuous image frames. It should be noted that the underwater image sequence refers to a set of multiple frames of underwater images with a time sequence that are continuously acquired using an image acquisition device within a certain time range. The underwater image sequence can be used to reflect the dynamic changes of the underwater scene in the water area to be inspected over time.
[0021] In step 102, depth information of different image frames is generated based on the color attenuation characteristics of the underwater image sequence, and then the boundary contours and corresponding characteristic depths of different underwater objects in each image frame are extracted by combining all the depth information through a deep learning algorithm.
[0022] In some embodiments, generating depth information of different image frames based on the color attenuation characteristics of the underwater image sequence may be achieved by using the following steps: Acquiring color attenuation characteristics of the underwater image sequence; determining the relative optical path length of each pixel in different image frames of the underwater image sequence by using the color attenuation feature; The depth information of different image frames is determined according to the relative optical path lengths of all pixels in each image frame.
[0023] It should be noted that the color attenuation feature of the underwater image sequence refers to the non-uniform variation in brightness distribution between color channels in the underwater image due to the selective absorption and scattering of light of different wavelengths by water in the underwater image sequence. The color attenuation feature can be used to reflect the spectral attenuation behavior during water propagation, thereby assisting in inferring the relative distance of an object. Preferably, a physical model describing the color attenuation behavior of water in the propagation path of light of different wavelengths can be constructed by obtaining the RGB values of each pixel in each image frame and combining the optical properties of the water body (such as absorption index and scattering coefficient) of the water body to be detected as model parameters based on the Beer-Lambert law in the prior art. Then, using this physical model, the corresponding light intensity attenuation value is calculated for the RGB channels of each pixel. Subsequently, based on the calculated light intensity attenuation values, the light intensity attenuation rate differences between the RGB channels are derived at all pixels. Finally, a vector composed of these light intensity attenuation rate differences is used as the color attenuation feature of the underwater image sequence. In other embodiments, other methods can also be used for implementation, which are not limited here.
[0024] It should be noted that the color changes of underwater images often reflect the absorption and scattering of light in the water body; therefore, the present application obtains the color attenuation feature by determining the difference in light intensity attenuation between the RGB channels in the underwater image; for example: as the depth of the water body increases, red light (long wavelength light) will be absorbed more strongly, while blue light (short wavelength light) will propagate farther; that is: by comparing the changes in the RGB channel values in the image, the degree of light attenuation can be calculated as the color attenuation feature in the present application.
[0025] In a specific implementation, determining the relative optical path length of each pixel in different image frames of the underwater image sequence using the color attenuation feature can be achieved in the following manner: first, selecting an image frame, obtaining all light intensity attenuation rate differences between the RGB channels of the image frame reflected in the color attenuation feature, and then constructing a functional relationship model between the light intensity attenuation rate and the optical path length between the channels by combining all the light intensity attenuation rate differences with the absorption coefficient differences of light of different wavelengths in the water area to be detected. Subsequently, for each pixel in the image frame, using its color attenuation feature (light intensity attenuation rate difference) as input, substituting it into the functional relationship model between the light intensity attenuation rate and the optical path length, and inversely calculating the relative optical path length of the corresponding pixel, thereby obtaining the relative optical path length of each pixel in different image frames of the underwater image sequence. Preferably, the relative optical path length of the corresponding pixel can be calculated by combining the attenuation rate differences corresponding to the RGB values of the corresponding pixel with the absorption coefficient of the water area where the corresponding pixel is located and applying the Beer-Lambert law.
[0026] In a specific implementation, the depth information of different image frames can be determined based on the relative optical path lengths of all pixels in each image frame. The following method is used: first, a historical image sequence of the area to be inspected is obtained, and the relative optical path lengths of each pixel in each image frame and its corresponding true depth value are recorded. Then, the relative optical path lengths and corresponding depth values of all pixels are organized into training samples to form a training data set, and a deep neural network in the prior art is used for model training. A neural network model with an input layer, a hidden layer, and an output layer is designed, wherein the input layer receives the relative optical path length of each pixel, and the output layer generates the depth value corresponding to the pixel. Subsequently, the mean square error (MSE) is used as a loss function to train the neural network to optimize the model parameters. During the training process, the network weights are continuously adjusted so that the model can accurately predict the depth value based on the relative optical path length. After the training is completed, the obtained neural network model is the optical path-to-depth conversion model of the area to be inspected. Finally, the relative optical path lengths of all pixels in each image frame are input into the optical path-to-depth conversion model, and the corresponding depth value is output. Then, the set of depth values of all pixels in each image frame is used as the depth information of the corresponding image frame.
[0027] Specifically, the true depth values in the training samples are obtained through synchronous acquisition by auxiliary equipment: while shooting underwater images, a high-precision depth sensor (such as a sonar device) is used to scan the same scene to obtain the actual depth data of each position; by calibrating the spatial position relationship between the camera and the sensor (i.e., internal and external parameter calibration), the depth values measured by the sensor are accurately mapped to the corresponding pixel points in the image, forming a training sample pair of 'optical path length-true depth', thereby ensuring the reliability of the model training data. In some embodiments, extracting the boundary contours and corresponding characteristic depths of different underwater objects in each image frame by combining all depth information with a deep learning algorithm can be achieved by the following steps: Determine the depth gradient change area in each image frame using all depth information; According to the depth gradient change area in all image frames and combined with the deep learning algorithm, the boundary contours of different underwater objects in each image frame are located, and then the characteristic depth corresponding to each boundary contour is extracted.
[0028] For specific implementation, refer to Figure 2As shown, this figure is a schematic diagram of a process for determining a depth gradient change region shown in some embodiments of the present application. The depth gradient change region in each image frame is determined by all depth information, that is, the position of all pixels in each image frame that shows a significant depth change (that is, the position where the gradient amplitude is greater than a set threshold) is taken as the depth gradient change region; for example: first, an image frame is selected and the depth information corresponding to the image frame is obtained; then, based on the depth information, the depth gradient value of each pixel in the image frame in the horizontal and vertical directions is calculated respectively; then, based on the horizontal and vertical depth gradients of each pixel, the depth gradient amplitude corresponding to the pixel is calculated; then, a gradient change threshold is set, and all pixel regions with depth gradient amplitudes greater than or equal to the gradient change threshold are screened out as the depth gradient change region in the image frame (otherwise, no processing is performed); the above steps are repeated to determine the depth gradient change regions in the remaining image frames in turn. Preferably, the gradient change threshold can be set adaptively: after calculating the depth gradient amplitude for all pixels in a single frame image, its distribution characteristics (such as the mean and standard deviation) are counted, and the threshold is set to a value that is a certain multiple of the standard deviation above the mean (for example, 1.5 times).
[0029] In specific implementation, the boundary contours of different underwater objects in each image frame are located according to the depth gradient change areas in all image frames combined with the deep learning algorithm, and then the feature depth corresponding to each boundary contour is extracted. This can be achieved by the following steps: first, a training sample set is constructed through all depth gradient change areas, and then, based on the semantic segmentation neural network in the existing technology (such as U-Net or DeepLab), taking the depth gradient change area as the input area and the depth information of the corresponding image frame as the auxiliary input channel, a multi-channel input model (semantic segmentation model) is constructed, wherein the input of the multi-channel input model is the depth value and gradient distribution of each pixel point, and the output is the classification result of whether each pixel point belongs to the boundary contour; subsequently, the semantic segmentation model is trained, and then the depth gradient change area in all image frames is input into the trained semantic segmentation model, and the semantic segmentation model outputs whether each pixel point is an object boundary; finally, based on the model output result (whether the pixel point is a boundary), the boundary contours of different underwater objects in each image frame are extracted. Preferably, for each boundary contour, the depth values of all pixels in each contour area can be obtained, and then the average value of the depth values of all pixels can be used as the characteristic depth of the corresponding underwater object. In other embodiments, other methods can also be used for implementation, which is not limited here.
[0030] It should be noted that the characteristic depth described in this application refers to the statistic (such as the average value) of the depth values of all pixels within the boundary contour of the object, which is used to characterize the relative position of the object underwater; this application extracts the position of the object boundary from all depth gradient change areas through manual labeling or automated methods, and marks whether each pixel belongs to the boundary contour, and extracts the corresponding object depth value to construct a training sample set; that is, each sample in the training sample set includes a depth gradient change area and an object boundary label; in addition, preferably, the semantic segmentation model can be trained using a supervised learning method, and the model parameters are optimized through a cross-entropy loss function to improve its classification accuracy in the boundary contour recognition task; finally, after the model training is completed, this application inputs the depth gradient change areas of all image frames into the trained semantic segmentation model, wherein the semantic segmentation model outputs a classification result of whether each pixel belongs to the object boundary; based on the output results of all pixels, the boundary points are connected using a connected domain analysis method or a contour extraction algorithm (such as Canny edge detection), and finally, the boundary contours of different underwater objects are obtained by aggregating the pixels of the boundary contours in each image frame.
[0031] In step 103, the reflection pattern of underwater light in different water propagation characteristics is constructed based on all feature depths and the surface roughness of various underwater objects, and then all boundary contours are semantically divided according to the reflection pattern and the structural form of the boundary contours of different underwater objects to obtain the mask areas corresponding to different semantic levels of each image frame.
[0032] In some embodiments, constructing the reflection pattern of underwater light in different water propagation characteristics based on all characteristic depths and the surface roughness of various underwater objects can be achieved by the following steps: Obtain the surface roughness of various underwater objects; The underwater light reflection response model of water bodies under different propagation conditions is established by taking into account all characteristic depths and the surface roughness of various underwater objects; The reflection pattern of underwater light in different water body propagation characteristics is output according to the underwater light reflection response model.
[0033] It should be noted that the present application obtains the surface roughness of various underwater objects by combining structured light irradiation with three-dimensional reconstruction. As a preferred embodiment, the method may specifically include the following steps: projecting a preset phase-coded fringe pattern onto the surface of the target underwater object through a structured light projection device, and synchronously collecting a sequence of deformed fringe images formed by the structured light on the surface of the object through an underwater camera; then, based on the calibration parameters of the camera and the projection device, combined with phase unwrapping and three-dimensional reconstruction algorithms, the collected images are processed to reconstruct three-dimensional point cloud data of the object surface; then, statistical analysis (such as arithmetic mean roughness) of the surface height information in a local area of the three-dimensional point cloud is performed, and the surface roughness parameters of various underwater objects are obtained as the surface roughness of the corresponding underwater objects; in addition, preferably, the method for obtaining the surface roughness of underwater objects also includes: combining a priori database with machine learning to establish a roughness parameter library of common objects (such as fish and corals), and using a convolutional neural network (CNN) to obtain the surface roughness parameter library of common objects (such as fish and corals). The roughness is inferred from the image texture features by a CNN (Convolutional Neural Network) as the surface roughness of various underwater objects in this application, or indirectly judged by the complexity of the reflection pattern, that is, the higher the proportion of diffuse reflection (such as >70%), the higher the roughness, and the higher the proportion of specular reflection (such as <30%), the lower the roughness; if the difference is still not significant, the dynamic analysis of the reflection pattern changes of multiple frames of images (such as if the reflection direction fluctuation within 3 frames is >10° is classified as "rough") can be used to obtain the relative roughness as the surface roughness of various underwater objects in this application.
[0034] Preferably, when performing statistical analysis on the surface height information in a local area of the three-dimensional point cloud, a spherical neighborhood with a preset radius (for example, 0.5 mm to 2 mm) is selected as the local area with each target point as the center, and the height values of all points in the area relative to the fitting plane are calculated. The arithmetic mean of the absolute values of these height values is calculated to obtain the arithmetic mean roughness parameter, which is used as the surface roughness of the corresponding underwater object.
[0035] In specific implementation, the underwater light reflection response model of the water body under different propagation conditions is established through all characteristic depths and the surface roughness of various underwater objects. This can be achieved in the following way, namely: taking the characteristic depths and corresponding surface roughness of various underwater objects as input variables, and combining the optical propagation characteristic parameters of the water body (such as the absorption coefficient, scattering coefficient and refractive index of the water body in the water area to be detected, etc.) to establish a multi-parameter reflection response function model, and then use the established model as the underwater light reflection response model in this application; as a preferred embodiment, a forward modeling method based on the principle of light radiation transmission (such as based on the Henyey-Greenstein scattering model) can be used to construct a reflection function with the characteristic depth and the surface roughness as variables, and then simulate the different incident light under The reflection intensity and scattering angle distribution formed after propagating in the water to the surface of the object, thereby describing the model of light reflection behavior under different water propagation conditions (underwater light reflection response model), can also be determined by other methods in other embodiments, which are not limited here; wherein, the input variables of the multi-parameter reflection response function model include the characteristic depth and surface roughness parameters of the target object, and the absorption coefficient, scattering coefficient and anisotropic scattering parameter of the water body; the output is the reflection intensity and scattering angle distribution of different incident light after propagating in the water body to the surface of the object; the model is constructed based on the radiation transfer theory, considering the initial light intensity of the incident light on the water surface, and the product of the reflectivity of the object surface and the light intensity as boundary conditions, and simulating the propagation and reflection process of light in the water body through numerical methods.
[0036] It should be noted that the underwater light reflection response model described in this application is based on the radiation transfer equation as its theoretical basis, integrates the characteristic depth, surface roughness and optical parameters of the target object and the water body, and realizes the model of underwater light propagation and reflection behavior through numerical simulation (such as Monte Carlo or DOM algorithm); the underwater light reflection response model is constructed based on the principle of physical propagation, and the input includes characteristic depth, surface roughness, and water body optical parameters (absorption coefficient, scattering coefficient), and the output is reflection intensity and angular distribution; the light reflection response model simulates the light propagation path through Monte Carlo, combines the micro-surface model (such as Trowbridge-Reitz) to calculate the rough surface reflection, and calibrates the parameters through the synchronous data collection of sonar and structured light to verify the consistency of the model with the measured reflection pattern.
[0037] In specific implementation, the reflection pattern of underwater light in different water propagation characteristics is output according to the underwater light reflection response model, which can be achieved in the following manner, namely: first, after obtaining the optical characteristic parameters (including absorption coefficient, scattering coefficient, phase function and refractive index, etc.) under different water propagation conditions, the characteristic depths and surface roughness of various underwater objects are combined as input variables and input into the underwater light reflection response model; secondly, the underwater light reflection response model is numerically solved. Preferably, it can be simulated based on the radiation transfer equation and combined with the Monte Carlo photon tracing algorithm to reproduce the complete transmission process of light from incident to reflection in the water medium. propagation path; wherein, in the simulation process, the Henyey-Greenstein scattering model can be introduced to describe the anisotropic scattering characteristics of water particles, and the microsurface model can be used to characterize the modulation effect of the surface roughness of the object on the reflection distribution; finally, the simulation data output by the model are subjected to angular distribution fitting, spectral intensity analysis and statistical clustering processing to obtain an underwater light reflection pattern corresponding to specific water conditions; it should be noted that the reflection pattern includes a reflection energy distribution diagram under the illumination direction and a reflection intensity curve under each incident angle. In other embodiments, other methods can also be used to achieve this, which is not limited here.
[0038] In some embodiments, semantically dividing all boundary contours according to the reflection pattern and the structural morphology of the boundary contours of different underwater objects to obtain mask areas corresponding to different semantic levels of each image frame can be achieved by the following steps: Obtain the structural morphology corresponding to each boundary contour; Establishing a semantic matching relationship between the reflection pattern and the structural form by combining the structural forms corresponding to each boundary contour with the reflection pattern; All boundary contours are divided into mask areas corresponding to different semantic levels of each image frame according to the semantic matching relationship.
[0039] In a specific implementation, obtaining the structural morphology corresponding to each boundary contour can be achieved in the following manner, namely: using the collective statistical features describing the shape of the boundary contour as the structural morphology of the corresponding boundary contour; for example: first, extracting the boundary contour curve of each boundary contour from the underwater image through an image processing method (such as Canny, Sobel, or a gradient-direction-based boundary tracking method); then, using the B-spline curve fitting method for the boundary contour curve of each boundary contour, the discrete points of the contour are converted into a continuous function representation, namely: a boundary contour function; then, determining the structural morphology parameters of each boundary contour based on the boundary contour function, thereby using the structural morphology parameters as the structural morphology of the corresponding boundary contour; preferably, determining the structural morphology parameters of each boundary contour based on the boundary contour function specifically includes: calculating the local curvature through the curve derivative to obtain the curvature distribution; evaluating the boundary concavity and convexity based on the curvature change characteristics; evaluating the boundary uniformity in combination with local length fluctuation or frequency domain analysis; and obtaining the main direction and direction distribution characteristics of the boundary through tangential angle statistics. In other embodiments, other determinations can also be made, which are not limited here.
[0040] Preferably, the boundary contour curve of each boundary contour is converted into a continuous function representation using a B-spline curve fitting method. Specifically, the method includes: fitting the discrete points of the boundary contour curve with a continuous function to obtain the boundary contour function, performing a second-order difference calculation on the function to obtain the first-order derivative and the second-order derivative, and using the first-order derivative and the second-order derivative to calculate the local curvature value of each point to form the curvature distribution characteristics. The boundary concavity is further evaluated by analyzing the curvature change, the boundary uniformity is evaluated by frequency domain analysis or local length fluctuation, and the main direction and directional distribution characteristics of the boundary are determined by tangential angle statistics, ultimately obtaining the structural morphological parameters of the boundary contour.
[0041] In specific implementation, the semantic matching relationship between the reflection pattern and the structural morphology corresponding to each boundary contour is established in combination with the reflection pattern, that is, the association mapping rule between the structural morphology and the reflection pattern is used as the semantic matching relationship between the reflection pattern and the structural morphology; it can be implemented in the following way, for example: first, the structural morphology corresponding to each boundary contour (such as curvature distribution feature vector, concavity index, boundary uniformity coefficient, main direction angle distribution, etc.) and the reflection pattern (such as reflection angle-intensity distribution vector, spectral reflection energy vector) are used as input variables, and a multi-parameter semantic mapping model is established by combining the joint feature space construction and the machine learning algorithm, and then the multi-parameter semantic mapping model is used as the semantic matching relationship between the reflection pattern and the structural morphology; as a preferred embodiment, a forward modeling method based on supervised learning (such as support vector regression, random forest regression or multi-layer perceptron neural network) can be used to construct a mapping function with structural morphology parameters and reflection pattern as input, and the model parameters are optimized through cross-validation and grid search and then trained to obtain the model as the semantic matching relationship between the reflection pattern and the structural morphology in this application.
[0042] It should be noted that the semantic matching relationship is a mapping model constructed through the feature coupling of structural morphology and reflection pattern, which can be used to quantify the degree of semantic adaptation of structural morphology to reflection pattern; in addition, the semantic matching relationship between reflection pattern and structural morphology is constructed through supervised learning, and the semantic labels of training samples are manually labeled in combination with physical simulation, and divided into multiple semantic levels (such as high, medium and low) according to the degree of coupling between reflection pattern and structural morphology; when training the model, the cross entropy between the true semantic level of the sample and the model prediction result can be used as the loss function, and the loss function can be minimized by adjusting the model parameters (such as the kernel function parameters of the support vector machine and the weights of the neural network) to improve the accuracy of semantic matching.
[0043] In a specific implementation, all boundary contours are divided into mask areas corresponding to different semantic levels of each image frame according to the semantic matching relationship, that is: all boundary contours are classified according to the semantic matching relationship of the reflection pattern-structural morphology, and then the mask areas corresponding to different semantic levels of each image frame are obtained; for example: first, semantic matching calculation is performed on each boundary contour through the semantic matching relationship, and then the semantic coupling index of each boundary contour relative to the reflection pattern in the joint feature space is obtained; then, a semantic level division criterion is constructed, and then a plurality of semantic level labels (for example: high semantic level, medium semantic level, low semantic level) are obtained according to the semantic level division criterion; finally, the semantic level label of each boundary contour is bound to its contour position information in the image space, and a semantic level labeling map (mask areas of different semantic levels) is generated in the corresponding image coordinate system, wherein contour filling, boundary mask mapping or image partition reconstruction can be used to assign unique mask values to different semantic levels respectively, and generate mask areas of different semantic levels covering the original image, thereby dividing all boundary contours into different semantic levels of each image frame. Corresponding mask area; preferably, constructing a semantic level division criterion, the semantic coupling indicators of all boundary contours can be clustered to obtain multiple clusters, wherein each cluster corresponds to a semantic level; for example: first, the semantic matching score of each boundary contour is used as an input feature, and then the K-Means algorithm is used to divide all boundary contours into several clusters (for example: high semantic level, middle semantic level, low semantic level) to obtain a semantic level division criterion. In other embodiments, other methods can also be used to implement it, which is not limited here; it should be noted that the The mask area is generated separately for each frame of the underwater image sequence and corresponds one-to-one to each frame in the image sequence. It should be noted that when using the K-Means algorithm to divide the semantic hierarchy, the optimal number of clusters can be determined by the elbow rule: the sum of squares of the intra-cluster errors corresponding to different numbers of clusters (such as 2, 3, and 4 categories) is calculated, and the number of clusters at the inflection point of the curve is selected as the number of semantic levels. In addition, the initial centroid can be selected using the K-Means++ algorithm, that is, a centroid is randomly selected first, and each subsequent centroid selects the point farthest from the selected centroid, thereby avoiding clustering falling into local optimality.
[0044] It should be noted that the semantic hierarchy described in this application is a comprehensive semantic representation of the optical properties (reflection patterns) and geometric properties (structural morphology) of underwater objects, which is used to distinguish targets, backgrounds, and interference; the specific granularity is divided into: the first-level granularity includes "target objects" (such as fish, corals), "background environment" (such as uniform water bodies, sandy bottoms), and "interference noise" (such as suspended particles); the second-level granularity is subdivided on the basis of the first-level, for example, "target objects" are divided into biological and non-biological categories, and "background environment" is divided into static background and dynamic background.
[0045] In addition, it should be noted that when different boundary contours overlap in the image space, the mask value can be assigned according to the priority of the semantic level label; if the semantic levels are different, the mask area of the higher semantic level (such as the high level determined by manual annotation or clustering) covers the lower semantic level area; if the semantic levels are the same, the contour area is judged based on its size, and the contour with a larger area is given priority to retain its mask value, thereby ensuring that each pixel in the image belongs to the mask area of only one semantic level.
[0046] In step 104, for each image frame, the texture directional features of the mask areas corresponding to different semantic levels are determined, and the main reflection directions of underwater light in different mask areas are determined through the brightness gradient distribution of each mask area. Then, coordinated mapping is performed based on all texture directional features and all main reflection directions to obtain the semantic coordination degree of each image frame in each mask area.
[0047] It should be noted that the texture direction feature described in the present application refers to the main direction of the texture structure in the mask area. As a preferred embodiment, the texture direction features of the mask areas corresponding to different semantic levels can be determined by calculating the gradient direction of each pixel in each mask area through image gradient operation (such as Sobel operator, Prewitt operator, etc.); then, a gradient direction histogram of the corresponding mask area is constructed through the gradient directions of all pixels, and the gradient direction of each pixel is divided into several angle intervals, and the number of pixels in each interval is counted. Subsequently, the gradient direction with the highest frequency is extracted from the gradient direction histogram as the main texture direction of the mask area. Finally, the main texture direction is used as the texture direction feature of the corresponding mask area, thereby obtaining the texture direction features of the mask areas corresponding to different semantic levels; as a preferred embodiment, in the present application, if the mean pixel gradient amplitude in the mask area is lower than 1 / 5 of the global gradient mean, it is determined to be a texture-free semantic area (such as a uniform water body); for the texture-free semantic area, the texture direction feature is set to 0° or a null value by default, indicating that there is no significant texture direction; in addition.
[0048] It should be noted that the present application can perform gradient calculation by using the Sobel operator, the convolution kernel size is 3×3, and the boundary is processed by mirror padding to ensure the integrity of the gradient calculation of the edge pixels of the mask area; in addition, the gradient direction of each pixel is divided into several angle intervals, specifically including: dividing the gradient direction into 36 equal-width areas, and weighting the counting histogram according to the gradient amplitude of each pixel.
[0049] In some embodiments, reference Figure 3 As shown in FIG, this figure is a schematic diagram of a process for determining the main reflection direction according to some embodiments of the present application. The main reflection direction of underwater light in different mask areas can be determined by the brightness gradient distribution of each mask area by the following steps: First, in step 1041, the brightness gradient distribution of each mask area is obtained; Then, in 1042 , a directionality index value of each mask region is determined based on all brightness gradient distributions; Finally, in 1043 , the main reflection directions of underwater light in different mask areas are determined based on all the directional index values.
[0050] It should be noted that the brightness gradient distribution refers to the gradient information of the image brightness (i.e., grayscale value) changing with spatial position in each mask area, and the gradient information can be used to reflect the direction and degree of brightness change in the corresponding mask area. As a preferred embodiment, after obtaining the grayscale image of each mask area, the Sobel operator or the Prewitt operator can be applied to the grayscale image in each mask area to perform convolution operation to obtain the horizontal gradient and vertical gradient of each pixel point, and then the gradient amplitude and gradient direction of each pixel point are determined by all the horizontal gradients and vertical gradients. Subsequently, in each mask area, a gradient amplitude distribution map is constructed according to the gradient amplitude of each pixel point, and the gradient direction is statistically constructed into a gradient direction histogram according to a preset angle interval as the brightness gradient distribution of each mask area.
[0051] It should be noted that the brightness gradient distribution refers to the distribution information composed of the gradient amplitude and gradient direction of all pixels within the mask area, including the "amplitude distribution map" (indicating intensity) and the "direction histogram" (indicating direction distribution).
[0052] In a specific implementation, the directional index value of each mask area is determined according to all brightness gradient distributions, that is, the proportion of pixels in the dominant direction in the gradient direction histogram in the mask area is used as the directional index value of the corresponding mask area; for example, first, a mask area is selected, and the gradient direction histogram corresponding to the brightness gradient distribution of the mask area is obtained. Then, the direction interval with the largest number of pixels is obtained through the gradient direction histogram, and then the proportion of the number of pixels in the direction interval with the largest number of pixels to the total number of pixels in the histogram is determined. Finally, the proportion is used as the directional index value of the corresponding mask area; as a preferred embodiment, the main reflection direction of underwater light in different mask areas is determined according to all directional index values, that is, for each mask area, the corresponding dominant direction in the gradient direction histogram of the area is used as the main reflection direction of the corresponding mask area, thereby obtaining the main reflection direction of underwater light in different mask areas, that is, directly mapping the main peak direction in the brightness gradient distribution to the light reflection direction without additional refraction or phase correction, assuming that the slight refraction in the water body can be ignored.
[0053] It should be noted that the main reflection direction refers to the most important reflection angle of underwater light in the area, which is inferred based on the concentration direction of the brightness gradient.
[0054] In some embodiments, the following steps may be used to perform coordination mapping based on all texture direction features and all main reflection directions to obtain the semantic coordination degree of each image frame in each mask area: Determine the directional consistency between the texture directional features and the main reflection direction of each mask area; The semantic coordination degree of each image frame in each mask area is determined through all directional consistencies.
[0055] In specific implementation, the directional consistency between the texture directional features and the main reflection direction of each mask area is determined, that is: the relative difference between the texture directional features and the main reflection direction is used as the directional consistency between the texture directional features and the main reflection direction; for example: the absolute value of the angle between the texture directional features corresponding to each mask area and the corresponding main reflection direction is used as the directional consistency between the texture directional features and the main reflection direction of the corresponding mask area; as a preferred embodiment, the semantic coordination degree of each image frame in each mask area is determined by all directional consistencies, which can be achieved in the following way, that is: the matching degree between the texture features in the mask area and the reflected light is used as the semantic coordination degree of the corresponding mask area; for example: first, the maximum and minimum values of all directional consistencies are obtained, and then the directional consistency corresponding to a mask area is selected, and after subtracting the directional consistency from the maximum value, the difference between the maximum value and the minimum value is compared as the semantic coordination degree of the mask area, and the above steps are repeated to determine the semantic coordination degree of the remaining mask areas, thereby obtaining the semantic coordination degree of each image frame in each mask area.
[0056] It should be noted that the directional consistency refers to the measurement of the difference between the angles of two directions, and the directional consistency can be used to characterize the degree of similarity between the texture directional features and the main reflection direction of each mask area; in addition, the semantic coordination refers to the quantitative value of the degree of alignment between the texture direction and the reflection direction of the mask area at the semantic level; in addition, the coordinated mapping of texture direction and semantics described in this application is achieved through multi-dimensional constraints: as a preferred embodiment, the feature depth (relative position of the object) and the reflection mode (ratio of diffuse reflection / specular reflection) can be combined to distinguish cross-semantically similar textures; for example, although corals and rocks have similar texture directions, the feature depth of corals is shallow and the surface roughness is high (dominated by diffuse reflection), while the feature depth of rocks is deep and may be mixed reflection; secondly, in the calculation of the semantic coordination, a weight factor can be introduced to suppress the ambiguity of a single feature, thereby ensuring the semantic distinction ability in weakly correlated scenarios; wherein, the weight factor can be automatically obtained from the training data through a supervised learning algorithm.
[0057] In step 105 , confident semantic labels of different underwater objects in the underwater image sequence are generated based on all semantic coordination degrees.
[0058] In some embodiments, generating confident semantic labels for different underwater objects in the underwater image sequence based on all semantic coordination degrees may be achieved by using the following steps: Obtaining initial semantic labels of the underwater image sequence in different image frames; The confidence adjustment is performed on the initial semantic label in each image frame through all semantic coordination degrees to obtain the confidence semantic labels of different underwater objects in the underwater image sequence.
[0059] It should be noted that the initial semantic label refers to the preliminary annotation of each area or pixel in the underwater image sequence based on the image content (such as color, texture, shape and other features), wherein the image content (such as color distribution, texture pattern, edge shape, etc.) is not directly used for rule-based classification, but is automatically learned and extracted by the deep network model during the training process to improve the semantic understanding ability of underwater targets; in addition, each initial semantic label represents the semantic category to which the corresponding pixel belongs, such as "water grass", "fish school", and "bottom", etc.; preferably, the initial semantic labels of the underwater image sequence in different image frames can be obtained by the deep learning algorithm (U-Net network algorithm) in the prior art, wherein the U-Net algorithm is a deep learning model for semantic segmentation, and the U-Net network model adopts a classic encoder-decoder structure, including 4 layers of downsampling modules and 4 layers of upsampling modules; each layer adopts a 3×3 convolution kernel, followed by ReLU activation and maximum pooling; the upsampling part adopts a deconvolution operation, and The features of the coding layer at the same level are spliced to the decoding path through jump connections, thereby retaining the spatial edge information of the input image; the final output layer is a tensor of the same size as the input image, and each pixel thereof contains a confidence probability vector of the corresponding category (for example, the output dimension is [512, 512, 5], representing 5 categories of semantic labels). In the specific implementation, first, the U-Net network is trained using a historical dataset of underwater image sequences containing pixel-level annotations to obtain a pre-trained semantic segmentation model; then, the underwater image sequence is input into the pre-trained semantic segmentation model frame by frame, wherein the pre-trained semantic segmentation model not only outputs the semantic category label of each pixel, but also gives the confidence probability corresponding to each semantic category through the softmax layer at the end of the network in the U-Net network; finally, the semantic category corresponding to the maximum confidence probability of each pixel can be taken as its initial semantic label to obtain the initial semantic label of the underwater image sequence in different image frames. In other embodiments, other methods can also be used for determination, which is not limited here.
[0060] It should be noted that the historical dataset includes several annotated underwater image sequences, each image is equipped with a semantic label mask corresponding to the pixel, and the semantic categories include "fish school", "water grass", "sand", "bottom", etc.; the labels are manually annotated by human experts using tools such as LabelMe; the cross-entropy loss function is used in the training process, the optimizer is Adam, the initial learning rate is set to 1e-4, and random flipping, rotation, brightness perturbation and other methods are used for data enhancement to improve the generalization ability of the model; in addition, after each frame of the image is processed by the U-Net network, its terminal softmax layer outputs the confidence probability of the corresponding semantic category for each pixel; the semantic category corresponding to the maximum confidence value of the pixel in each semantic category can be used as the initial semantic label; at the same time, the maximum confidence value is used as the initial confidence of the label.
[0061] It should be noted that the Softmax probability refers to the output of the model for each category in a multi-classification neural network, which is converted into a normalized probability value. For example, the original scores of the pixel point output by the semantic segmentation model in the three categories of "water plants", "fish school" and "bottom" are [2.0, 1.0, 0.1] respectively, and the corresponding Softmax probability is [0.66, 0.24, 0.10], where 0.66 is the confidence probability of the semantic category "water plants", 0.24 is the confidence probability of the semantic category "fish school", and 0.10 is the confidence probability of the semantic category "bottom".
[0062] In a specific implementation, confidence adjustment is performed on the initial semantic label in each image frame using all semantic coordination degrees to obtain the confidence semantic labels of different underwater objects in the underwater image sequence. This can be achieved in the following manner: first, for each image frame, a mask region is selected, and the softmax probability of all pixels in the mask region is obtained, thereby determining the average confidence probability of each semantic category in the mask region; then, the average confidence probability is multiplied by the semantic coordination degree of the mask region to obtain the weighted confidence probability of each semantic category in the mask region; then, for each pixel in the mask region, the confidence probability of each semantic category in the pixel is linearly fused with the weighted confidence probability of the corresponding semantic category in the mask region, and the obtained result is used as the fused confidence probability of each pixel in different semantic categories. The maximum fused confidence probability of all pixels in the mask region among all semantic categories is obtained, and the corresponding semantic category is used as the confidence semantic label of each pixel, thereby obtaining the confidence semantic labels of different underwater objects in the underwater image sequence.
[0063] It should be noted that the fused confidence probability refers to the probability value obtained by weighting each pixel point in the process of generating image semantic labels, combined with its Softmax confidence in a specific semantic category and the semantic coordination degree of its mask area, which is used to reflect the confidence strength that the pixel belongs to the target semantic category; preferably, the detection threshold can be set to a fixed value (for example, 0.6), or it can be adaptively set according to the frequency of occurrence of the target category in the training set; therefore, in specific implementation, dynamic detection of underwater targets through all confidence semantic labels can be achieved in the following way, namely: first, set the detection threshold; for each image frame, extract the water The method comprises the following steps: first, detecting the pixel points whose fusion confidence probability of the semantic category of the underwater target (such as "fish school") is greater than or equal to the detection threshold, and then binarizing all the pixel points to generate a binary mask map of the underwater target. Then, a connected domain analysis is performed on the binary mask map to identify the connected area of the underwater target and determine the centroid coordinates of the connected area. Then, the centroid coordinates of the connected area of the underwater target in all adjacent image frames in the underwater image sequence are determined to determine the speed information of the underwater target. Finally, the appearance frame, disappearance frame, centroid coordinate trajectory and speed information of the underwater target in the underwater image sequence are recorded to achieve dynamic detection of the underwater target.
[0064] As a preferred embodiment, generating the binary mask map specifically includes: for each pixel point in the image, if its fusion confidence probability is greater than or equal to the set detection threshold, assigning the pixel a value of 1 in the mask, otherwise assigning it a value of 0, thereby forming a binary mask map of the underwater target; in addition, when performing connected domain analysis on the binary mask map, an 8-neighborhood judgment method can be used to improve the integrity of target extraction; each connected domain is regarded as an independent target instance, and its bounding box, pixel area and centroid coordinates are recorded; wherein the centroid coordinates are obtained by averaging the coordinates of all pixels in the connected domain, indicating the center position of the target area; in addition, the centroid coordinates of the pixels identified as the same instance (underwater target) in consecutive image frames can be connected in chronological order. , forming a centroid coordinate trajectory, and then by calculating the Euclidean distance of the centroids between adjacent frames and dividing it by the time difference between frames, the instantaneous speed (speed information) of the target in the period can be obtained; further, the speed sequence can be processed by sliding average to remove occasional jitter and improve stability; finally, the present application outputs the corresponding appearance frame number, disappearance frame number, centroid trajectory (recorded in the form of frame number and coordinates), speed sequence, direction change and other information of the detected underwater target, and uniformly stores them as a structured data file (such as JSON or CSV format); at the same time, the target trajectory overlay map or trajectory visualization video can be generated in combination with the original image to intuitively display the dynamic process and support multi-target parallel output, thereby realizing dynamic detection of underwater targets.
[0065] In addition, in another aspect of the present application, in some embodiments, the present application provides an underwater target detection system based on deep learning, referring to Figure 4 , which is a schematic structural diagram of an underwater target detection system based on deep learning according to some embodiments of the present application. The underwater target detection system based on deep learning 200 includes: an acquisition module 201, a processing module 202, and an execution module 203, which are described as follows: Acquisition module 201, in this application, acquisition module 201 is mainly used to acquire underwater image sequences of the water area to be detected; Processing module 202, in this application, is mainly used to generate depth information of different image frames based on the color attenuation characteristics of the underwater image sequence, and then extract the boundary contours and corresponding characteristic depths of different underwater objects in each image frame by combining all the depth information through a deep learning algorithm; In addition, the processing module 202 of the present application is further configured to construct a reflection pattern of underwater light in different water propagation characteristics based on all characteristic depths and the surface roughness of various underwater objects, and then semantically divide all boundary contours according to the reflection pattern and the structural morphology of the boundary contours of different underwater objects to obtain mask areas corresponding to different semantic levels for each image frame; In addition, the processing module 202 in the present application is further configured to determine, for each image frame, texture directional features of mask areas corresponding to different semantic levels, determine the main reflection directions of underwater light in different mask areas through the brightness gradient distribution of each mask area, and then perform coordinated mapping based on all texture directional features and all main reflection directions to obtain the semantic coordination degree of each image frame in each mask area; The execution module 203 in this application is mainly used to generate confident semantic labels of different underwater objects in the underwater image sequence based on all semantic coordination degrees.
[0066] In addition, the present application also provides a computer device, which includes a memory and a processor, the memory storing a code, and the processor being configured to obtain the code and execute the above-mentioned deep learning-based underwater target detection method.
[0067] In some embodiments, reference Figure 5 , which is an internal structure diagram of a computer device for implementing a method for underwater target detection based on deep learning according to some embodiments of the present application. The method for underwater target detection based on deep learning in the above embodiments can be Figure 5 The computer device 300 shown in FIG. 1 is implemented as shown in FIG. 1 , and the computer device 300 includes at least one processor 301 , a communication bus 302 , a memory 303 , and at least one communication interface 304 .
[0068] The processor 301 may be a general-purpose central processing unit (CPU), or an application-specific integrated circuit (ASIC) or one or more for controlling the execution of the deep learning-based underwater target detection method in the present application.
[0069] The communication bus 302 is used to transmit information between the above components.
[0070] Memory 303 may be, but is not limited to, a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, an optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer. Memory 303 may exist independently and be connected to processor 301 via communication bus 302. Memory 303 may also be integrated with processor 301.
[0071] Memory 303 is used to store program code for executing the solution of the present application, and is controlled by processor 301 for execution. Processor 301 is used to execute the program code stored in memory 303. The program code may include one or more software modules. The deep learning-based underwater target detection method in the above embodiment can be implemented by processor 301 and one or more software modules in the program code in memory 303.
[0072] The communication interface 304 uses any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area network (WLAN), etc.
[0073] In a specific implementation, as an example, a computer device may include multiple processors, each of which may be a single-CPU processor or a multi-CPU processor. A processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0074] The aforementioned computer device may be a general-purpose computer device or a dedicated computer device. In a specific implementation, the computer device may be a desktop computer, a portable computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of this application do not limit the type of computer device.
[0075] In addition, the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned deep learning-based underwater target detection method.
[0076] In summary, the deep learning-based underwater target detection method and system disclosed in the embodiments of the present application obtain an underwater image sequence of the water area to be detected; generate depth information of different image frames based on the color attenuation characteristics of the underwater image sequence, and then extract the boundary contours and corresponding feature depths of different underwater objects in each image frame through a deep learning algorithm combined with all depth information; construct a reflection pattern of underwater light in different water body propagation characteristics based on all feature depths and the surface roughness of various underwater objects, and then semantically divide all boundary contours according to the reflection pattern and the structural form of the boundary contours of different underwater objects to obtain mask areas corresponding to different semantic levels for each image frame; for each image frame, determine the texture direction features of the mask areas corresponding to different semantic levels, determine the main reflection direction of the underwater light in different mask areas through the brightness gradient distribution of each mask area, and then coordinate mapping is performed based on all texture direction features and all main reflection directions to obtain the semantic coordination degree of each image frame in each mask area; generate confident semantic labels for different underwater objects in the underwater image sequence based on all semantic coordination degrees; and perform anti-distortion detection of underwater targets under the interference of the underwater environment.
[0077] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0078] Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if such changes and modifications fall within the scope of the claims of the present application and their equivalents, the present application is intended to include such changes and modifications.
Claims
1. A method for underwater target detection based on deep learning, characterized in that: The steps include: Acquire an underwater image sequence of the water area to be detected; Generating depth information of different image frames based on the color attenuation characteristics of the underwater image sequence, and then extracting the boundary contours and corresponding characteristic depths of different underwater objects in each image frame by combining all the depth information through a deep learning algorithm; Based on all feature depths and the surface roughness of various underwater objects, the reflection pattern of underwater light in different water propagation characteristics is constructed. Then, based on the reflection pattern and the structural morphology of the boundary contours of different underwater objects, all boundary contours are semantically divided to obtain the mask areas corresponding to different semantic levels for each image frame; For each image frame, the texture directional features of the mask areas corresponding to different semantic levels are determined. The main reflection directions of underwater light in different mask areas are determined by the brightness gradient distribution of each mask area. Then, a coordinated mapping is performed based on all the texture directional features and all the main reflection directions to obtain the semantic coordination degree of each image frame in each mask area. Confident semantic labels of different underwater objects in the underwater image sequence are generated based on all semantic coordination degrees.
2. The method according to claim 1, wherein Generating depth information of different image frames based on the color attenuation characteristics of the underwater image sequence specifically includes: Acquiring color attenuation characteristics of the underwater image sequence; determining the relative optical path length of each pixel in different image frames of the underwater image sequence by using the color attenuation feature; The depth information of different image frames is determined according to the relative optical path lengths of all pixels in each image frame.
3. The method according to claim 1, wherein The deep learning algorithm combines all depth information to extract the boundary contours and corresponding feature depths of different underwater objects in each image frame. Specifically, Determine the depth gradient change area in each image frame using all depth information; According to the depth gradient change area in all image frames and combined with the deep learning algorithm, the boundary contours of different underwater objects in each image frame are located, and then the characteristic depth corresponding to each boundary contour is extracted.
4. The method according to claim 1, wherein The reflection pattern of underwater light in different water propagation characteristics is constructed based on all characteristic depths and the surface roughness of various underwater objects. Specifically, it includes: Obtain the surface roughness of various underwater objects; The underwater light reflection response model of water bodies under different propagation conditions is established by taking into account all characteristic depths and the surface roughness of various underwater objects; The reflection pattern of underwater light in different water body propagation characteristics is output according to the underwater light reflection response model.
5. The method according to claim 1, wherein According to the reflection pattern and the structural form of the boundary contours of different underwater objects, all boundary contours are semantically divided to obtain the mask areas corresponding to different semantic levels of each image frame, specifically including: Obtain the structural morphology corresponding to each boundary contour; Establishing a semantic matching relationship between the reflection pattern and the structural form by combining the structural forms corresponding to each boundary contour with the reflection pattern; All boundary contours are divided into mask areas corresponding to different semantic levels of each image frame according to the semantic matching relationship.
6. The method according to claim 1, wherein The main reflection directions of underwater light in different mask areas are determined by the brightness gradient distribution of each mask area. Specifically, Obtain the brightness gradient distribution of each mask area; Determine the directionality index value of each mask area based on all brightness gradient distributions; The main reflection directions of underwater light in different mask areas are determined based on all directional index values.
7. The method according to claim 1, wherein According to all texture direction features and all main reflection directions, the semantic coordination degree of each image frame in each mask area is obtained by performing coordination mapping, specifically including: Determine the directional consistency between the texture directional features and the main reflection direction of each mask area; The semantic coordination degree of each image frame in each mask area is determined through all directional consistencies.
8. An underwater target detection system based on deep learning, characterized in that: include: An acquisition module, used for acquiring underwater image sequences of the water area to be detected; a processing module for generating depth information of different image frames based on the color attenuation characteristics of the underwater image sequence, and then extracting the boundary contours and corresponding characteristic depths of different underwater objects in each image frame by combining all the depth information through a deep learning algorithm; The processing module is further configured to construct a reflection pattern of underwater light in different water propagation characteristics based on all characteristic depths and the surface roughness of various underwater objects, and then semantically divide all boundary contours according to the reflection pattern and the structural morphology of the boundary contours of different underwater objects to obtain mask areas corresponding to different semantic levels for each image frame; The processing module is further configured to determine, for each image frame, texture directional features of mask areas corresponding to different semantic levels, determine the main reflection directions of underwater light in different mask areas based on the brightness gradient distribution of each mask area, and then perform coordinated mapping based on all texture directional features and all main reflection directions to obtain the semantic coordination degree of each image frame in each mask area; An execution module is configured to generate confident semantic labels for different underwater objects in the underwater image sequence based on all semantic coordination degrees.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the underwater target detection method based on deep learning according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the underwater target detection method based on deep learning are implemented as described in any one of claims 1 to 7.
Citation Information
Cited By
Image analysis method fusing physical prior
CN121120626A
Image analysis method fusing physical prior
CN121120626B
Intelligent benthonic animal identification method and system and storage medium
CN121582573A