A High-Precision Target Detection Model and Method Based on Underwater Image Enhancement

By combining technologies such as ResNet, 2D histogram equalization and Faster R-CNN, a high-precision underwater object detection model is proposed, which solves the problem of insufficient accuracy and adaptability of underwater object detection in the prior art, and achieves higher image quality and object detection accuracy.

CN117975251BActive Publication Date: 2025-06-10OCEAN UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410156535.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-04
Publication Date
2025-06-10
Estimated Expiration
2044-02-04

AI Technical Summary

Technical Problem

The existing underwater target detection technology has limited accuracy and adaptability when dealing with complex underwater environments, and it is difficult for a single method to distinguish the characteristic importance of underwater fish images, resulting in a high rate of error detection.

Method used

A high-precision object detection model based on underwater image enhancement is proposed, combining the ResNet network module, 2D histogram equalization module and Faster R-CNN object detection module to improve image quality and object detection accuracy through deep learning and image enhancement technology.

Benefits of technology

Through the combination of image enhancement and object detection model, the contrast and clarity of underwater images are significantly improved, the adaptability to different underwater environments is enhanced, the error detection rate is reduced, and the accuracy and efficiency of object detection are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117975251B_ABST
    Figure CN117975251B_ABST
Patent Text Reader

Abstract

The present invention provides a high-precision target detection model and method based on underwater image enhancement, belonging to the technical field of underwater target image enhancement and detection. A method for underwater target detection combining deep learning and 2D histogram equalization technology is mainly proposed. This method first introduces the ResNet model for deep enhancement of color conversion and image details. Then, the 2D histogram equalization technology is applied to further improve the contrast and clarity of the image. This process not only optimizes color and detail processing but also improves the accuracy and efficiency of target detection. Finally, Faster R-CNN is used to detect the target. By combining multiple factors such as ResNet, 2D histogram equalization, and Faster R-CNN, the quality of target detection in underwater images is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of underwater target image enhancement and detection, and particularly relates to a high-precision target detection model and method based on underwater image enhancement. Background Art

[0002] In the underwater environment, image enhancement technology is crucial for ecological and economic research, especially in the exploration, protection, and management of marine resources. The underwater ecosystem, including coral reefs, fish, and other marine organisms, constitutes a complex and diverse environment, which is of great significance for maintaining ecological balance, protecting biodiversity, and promoting marine scientific research.

[0003] However, the quality of underwater images is affected by various factors, such as light transmission problems, background interference, target occlusion, and shadows. These factors greatly limit the clarity and accuracy of underwater images. Traditional underwater target detection technologies, such as methods based on feature extraction and model matching, although effective in some applications, have limited accuracy and adaptability when dealing with complex underwater environments. These methods usually rely on manual feature extraction and are difficult to adapt to the variability of the underwater environment, resulting in limited performance in practical applications.

[0004] To address these problems, deep learning methods such as Residual Network (ResNet) and Faster Region Convolutional Neural Network (Faster R-CNN) have been introduced. ResNet solves the gradient vanishing problem in deep neural networks through its skip connection mechanism, allowing for deeper learning without loss of performance, and thus showing excellent capabilities in enhancing images. In addition, as an effective target detection framework, Faster R-CNN demonstrates powerful capabilities in dealing with target detection in complex underwater environments through automated feature extraction and learning from large-scale datasets. This combination not only improves the performance of underwater image enhancement technology, especially when dealing with problems such as color distortion and uneven illumination, but also enhances the adaptability to complex underwater environments, thereby demonstrating the great potential of deep learning in improving the quality of underwater images and the accuracy of target detection.

[0005] However, due to the complex underwater environment, such as the influence of underwater plants, sediments, etc. on the background, a single method cannot distinguish the importance of features of underwater fish images, resulting in a relatively high false detection rate. Moreover, the collected fish features have different poses and deformations, which makes it difficult for traditional deep learning networks that are sensitive to target pose changes to accurately detect targets with changes. Summary of the Invention

[0006] In view of the above problems, the present invention proposes a new noise-robust depth detection framework and training method to effectively process small targets and noise data in underwater scenes.

[0007] In the first aspect of the present invention, a high-precision target detection model based on underwater image enhancement is proposed, including a ResNet network module, a 2D histogram equalization module, and a Faster R-CNN target detection module connected in sequence;

[0008] The ResNet network module is trained using the dataset Dataset to learn and extract more complex image features; sixteen residual blocks are stacked together to form a deep network structure, and a convolutional layer is set as the last layer of the model, with the output size and number of channels being the same as those of the input image, ensuring that the output is the enhanced image, and completing the depth enhancement of color conversion and image details;

[0009] The 2D histogram equalization module applies 2D histogram equalization technology to further improve the contrast and clarity of the image frames enhanced by the ResNet network module;

[0010] The Faster R-CNN target detection module uses the images jointly enhanced by the ResNet network module and the 2D histogram equalization module to construct a dataset and trains it for classifying and detecting the images to be detected.

[0011] Preferably, the production process of the dataset Dataset is as follows:

[0012] Based on the underwater short video data data of 2s, each segment of data is cut into image frame data d with a sequence length of 50 at uniform time intervals 1 ,d 2 ,…,d 50 ; the image frames in the image frame data d i ,i∈[1,50] are preprocessed, including adjusting the image size to 320×320 pixels and labeling the corresponding tags for the objects to be detected in the image, and the results are used as a set of data to integrate and obtain the dataset Dataset.

[0013] Preferably, the specific process of setting and processing data by the ResNet network module is as follows:

[0014] Setting of improved ResNet residual blocks; each residual block contains three convolutional layers, where the first and third layers use 1x1 filters, mainly responsible for adjusting the number of channels, i.e., dimensionality reduction and dimensionality increase; the second layer uses a 3x3 filter, focusing on feature extraction;

[0015] The data input is x l , and the convolutional operation is F(x l ,W l ), where W lis the weight of the convolutional layer; in addition, in order to learn the complex relationship between the input and the residual, in the residual connection, a small convolutional network Conv is introduced to obtain the output of the residual block:

[0016] x l+1 = F(x l , W l ) + Conv(x l )

[0017] Stacking 16 improved residual blocks together forms a deep network structure;

[0018] The last layer of the model is set as a convolutional layer, and the size and number of channels of its output are the same as those of the input image, so as to ensure that the output is the enhanced image, and the output formula is:

[0019] Y = E(X, W)

[0020] where X is the input image and W represents all the weights in the network; enhancing all the image frame data in each group of sequence data; obtaining the image frame d enhanced by ResNet iRes , defined as:

[0021] d iRes = Res(d i )

[0022] where Res(d i ) is the original image frame data processed by ResNet.

[0023] Preferably, the specific processing process of the 2D histogram equalization module is as follows:

[0024] S1. For the input image frame d iRes , calculate the histograms of the R, G, and B channels respectively, and the formula is:

[0025]

[0026] where H(v) represents the number of pixels with pixel value v, and I(i, j) is the pixel value of the image at the position (i, j);

[0027] S2. Use the Gaussian smoothing method to smooth the histogram, reduce image noise and improve the statistical characteristics of the histogram, and the formula is:

[0028]

[0029] where G(x, y) represents the Gaussian function, x and y are the coordinates of the pixel positions, and σ is the standard deviation, which controls the width of the Gaussian kernel;

[0030] S3. Calculate the cumulative distribution function (CDF) C(v) for the histogram of each channel, with the formula:

[0031]

[0032] S4. Use the CDF to map the original pixel values to new intensity values I′(i,j), with the formula:

[0033]

[0034] where MN is the total number of pixels in the image;

[0035] S5. Apply the mapped pixel values to the image, merge the pixels of each channel, complete the contrast enhancement, and obtain the enhanced image frame d i2D 。

[0036] Preferably, the specific structure of the Faster R-CNN object detection module is:

[0037] Use the lightweight ResNet-34 as the feature extraction network, and then generate the Region Proposal Network (RPN) based on the output of the feature extraction network;

[0038] The RPN locates potential target regions by using anchor boxes at multiple scales and aspect ratios;

[0039] The output of the RPN includes two parts, one for predicting the probability of object presence and the other for bounding box regression; the goal of bounding box regression is to adjust the position and size of the anchor boxes to better cover the real object, and the adjustment of the bounding box is represented by the following formula:

[0040]

[0041] where x r , y r , w r , h r are the center coordinates and dimensions of the real bounding box, and x a , y a , w a , h a are the center coordinates and dimensions of the corresponding anchor box;

[0042] Pool the proposed regions of different sizes into a unified size, and then classify the candidate detection boxes to output the detection results.

[0043] The second aspect of the present invention provides a high-precision object detection method based on underwater image enhancement, including the following processes:

[0044] Obtain the target image to be detected;

[0045] Input the target image to be detected into the high-precision target detection model described in any one of claims 1 to 5;

[0046] Output the target class probabilities of each RPN, P = [p 1 , p 2 , …, p k , where k is the total number of classes; and the target detection result is the class with the highest probability value, that is:

[0047]

[0048] where p i is the probability of the i-th class.

[0049] The third aspect of the present invention provides a high-precision target detection device based on underwater image enhancement. The device includes at least one processor and at least one memory, and the processor and the memory are coupled; a computer execution program of the high-precision target detection model described in the first aspect is stored in the memory; when the processor executes the computer execution program stored in the memory, the processor executes a high-precision target detection method based on underwater image enhancement.

[0050] The fourth aspect of the present invention provides a computer-readable storage medium, in which a computer program or instruction of the high-precision target detection model described in the first aspect is stored. When the program or instruction is executed by a processor, the processor executes a high-precision target detection method based on underwater image enhancement.

[0051] Compared with the prior art, the present invention has the following beneficial effects:

[0052] Color accuracy: Use the ResNet model to adjust and correct the color distortion of underwater images.

[0053] Contrast enhancement: Apply 2D histogram equalization to improve the contrast of the image and make the details clearer.

[0054] Adaptability: Combine the Faster R-CNN model to improve the adaptability to different underwater environments.

[0055] Generally speaking, the method of this research greatly improves the quality of target detection in underwater images by combining multiple factors such as ResNet, 2D histogram equalization, and Faster R-CNN. Description of the Drawings

[0056] To more clearly illustrate the technical solutions of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the following description is only one embodiment of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0057] Figure 1 It is a logic block diagram of the underwater target detection method of the present invention.

[0058] Figure 2 It is a schematic structural diagram of the ResNet model established by the present invention.

[0059] Figure 3 It is a schematic diagram of the 2D histogram equalization process of the present invention.

[0060] Figure 4 It is a schematic structural diagram of the Faster R-CNN model of the present invention.

[0061] Figure 5 It is a simple structural block diagram of the target detection device in Embodiment 2. Detailed implementation manners

[0062] The present invention will be further described below in conjunction with specific embodiments.

[0063] To improve the performance of underwater image enhancement and target detection, the present invention proposes an underwater target detection method combining deep learning and 2D histogram equalization technology. This method first introduces the ResNet model to perform deep enhancement of color conversion and image details. Then, the 2D histogram equalization technology is applied to further enhance the contrast and clarity of the image. This process not only optimizes color and detail processing but also improves the accuracy and efficiency of target detection. Finally, Faster R-CNN is used to detect the target.

[0064] The overall idea of the present invention is as Figure 1 shown. The overall model structure includes a ResNet network module, a 2D histogram equalization module, and a Faster R-CNN target detection module connected in sequence;

[0065] The ResNet network module is trained using the dataset Dataset to learn and extract more complex image features; 16 residual blocks are stacked together to form a deep network structure, and a convolutional layer is set at the last layer of the model. The output size and number of channels are the same as those of the input image, ensuring that the output is the enhanced image, thus completing the deep enhancement of color conversion and image details;

[0066] The 2D histogram equalization module applies the 2D histogram equalization technique to further enhance the contrast and clarity of the image frames enhanced by the ResNet network module.

[0067] The Faster R-CNN object detection module uses the images jointly enhanced by the ResNet network module and the 2D histogram equalization module to construct a dataset and trains it for classifying and detecting the images to be detected.

[0068] In this embodiment, a group of underwater images is taken as an example to further illustrate the method of the present invention.

[0069] 1. Collection and preprocessing of the Dataset

[0070] Collect multiple underwater video data, where the videos cover different underwater environments and conditions, and the short video data data with a duration of 2s.

[0071] At uniform time intervals, each segment of data is sliced into image frame data d with a sequence length of 50 1 , d 2 , …, d 50 ;

[0072] Preprocess the image frames in the image frame data d i , i ∈ [1, 50], including resizing the image to 320×320 pixels and labeling the corresponding tags for the objects to be detected in the image, and using the results as a set of data Dataset j ;

[0073] Repeat the above process and integrate the data processing results to obtain the dataset Dataset.

[0074] 2. ResNet network module

[0075] The ResNet model structure is as Figure 2 shown:

[0076] Improved ResNet residual block settings; each residual block contains three convolutional layers, where the first and third layers use 1x1 filters, mainly responsible for adjusting the number of channels, that is, dimensionality reduction and dimensionality increase; the second layer uses a 3x3 filter, focusing on feature extraction;

[0077] The data input is x l , and the convolution operation is F(x l , W l ), where W l is the weight of the convolutional layer; in addition, in order to learn the complex relationship between the input and the residual, in the residual connection, a small convolutional network Conv is introduced to obtain the output of the residual block:

[0078] x l+1 = F(x l , W l ) + Conv(x l )

[0079] 1) Stack 16 improved residual blocks together to form a deep network structure; this structure helps to learn and extract more complex image features.

[0080] The last layer of the model is set as a convolutional layer, and the output size and number of channels are the same as those of the input image, thus ensuring that the output is the enhanced image. The output formula is:

[0081] Y = E(X, W)

[0082] where X is the input image and W represents all the weights in the network; enhance all the image frame data in each group of sequence data; obtain the image frame d enhanced by ResNet iRes , defined as:

[0083] d iRes = Res(d i )

[0084] where Res(d i ) is the original image frame data processed by ResNet.

[0085] 3. Regarding the 2D histogram equalization module

[0086] The 2D histogram equalization process is as Figure 3 shown:

[0087] S1. For the input image frame d iRes , calculate the histograms of the R, G, and B channels respectively. The formula is:

[0088]

[0089] where H(v) represents the number of pixels with pixel value v, and I(i, j) is the pixel value of the image at the position (i, j);

[0090] S2. Use the Gaussian smoothing method to smooth the histogram, reduce image noise and improve the statistical characteristics of the histogram. The formula is:

[0091]

[0092] where G(x, y) represents the Gaussian function, x and y are the coordinates of the pixel positions, and σ is the standard deviation, which controls the width of the Gaussian kernel;

[0093] S3. Calculate the cumulative distribution function (CDF) C(v) for the histogram of each channel, with the formula:

[0094]

[0095] S4. Use the CDF to map the original pixel values to new intensity values I′(i,j), with the formula:

[0096]

[0097] where MN is the total number of pixels in the image;

[0098] S5. Apply the mapped pixel values to the image, merge the pixels of each channel, complete the contrast enhancement, and obtain the enhanced image frame d i2D .

[0099] 4. Faster R-CNN Object Detection Module

[0100] The specific structure of the Faster R-CNN object detection module is as Figure 4 shown:

[0101] Use the lightweight ResNet-34 as the feature extraction network, and then generate the Region Proposal Network (RPN) based on the output of the feature extraction network;

[0102] The RPN locates potential target regions by using anchor boxes at multiple scales and aspect ratios; in this embodiment, the anchor box sizes are [32, 64, 128] and the aspect ratios are [0.5, 1, 2];

[0103] The output of the RPN includes two parts, one for predicting the probability of object presence and the other for bounding box regression; the goal of bounding box regression is to adjust the position and size of the anchor box to better cover the real object, and the adjustment of the bounding box is represented by the following formula:

[0104]

[0105] where x r , y r , w r , h r are the center coordinates and size of the real bounding box, and x a , y a , w a , h a are the center coordinates and size of the corresponding anchor box;

[0106] Pool the proposed regions of different sizes into a unified size (7×7), and then classify the candidate detection boxes to output the detection results.

[0107] For the newly input image sequence data, it is jointly enhanced through the ResNet network module and the 2D histogram equalization module to obtain the enhanced data d i2D . Take d i2D as the input of the Faster R-CNN module, and the Faster R-CNN model outputs the target class probabilities for each PRN, P = [p 1 , p 2 , …, p k , where k is the total number of classes. And the object detection result is the class with the highest probability value, that is:

[0108]

[0109] where p i is the probability of the i-th class.

[0110] Therefore, based on this object classification and detection method, object classification and detection are performed on the input underwater sequence image (50 images) data, and the detection effect is shown in Table 1.

[0111] Table 1 Performance comparison of different models in object detection tasks

[0112]

[0113]

[0114] Among them, HOG is the histogram of oriented gradients, SVM is the support vector machine, mAP is the mean average precision, and GFLOPs is the number of floating-point operations in billions per second.

[0115] Example 2:

[0116] As Figure 5As shown, the present invention also provides a target detection device applicable to a noisy underwater scenario. The device includes at least one processor and at least one memory, and also includes a communication interface and an internal bus. A computer executable program of the high-precision target detection model as described in Embodiment 1 is stored in the memory. When the processor executes the computer executable program stored in the memory, the processor can execute a high-precision target detection method based on underwater image enhancement. The internal bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus. The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.

[0117] The device can be provided as a terminal, a server, or other forms of devices.

[0118] Figure 5 It is a block diagram of an exemplary device. The device may include one or more of the following components: a processing component, a memory, a power supply component, a multimedia component, an audio component, an input / output (I / O) interface, a sensor component, and a communication component. The processing component generally controls the overall operation of the electronic device, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component may include one or more processors to execute instructions to complete all or part of the steps of the above method. In addition, the processing component may include one or more modules to facilitate the interaction between the processing component and other components. For example, the processing component may include a multimedia module to facilitate the interaction between the multimedia component and the processing component.

[0119] The memory is configured to store various types of data to support the operation of the electronic device. Examples of such data include instructions for any application or method operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disc.

[0120] The power supply component provides power for various components of the electronic device. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device. The multimedia component includes a screen that provides an output interface between the electronic device and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component includes a front camera and / or a rear camera. When the electronic device is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0121] The audio component is configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) that is configured to receive external audio signals when the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory or sent via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals. The I / O interface provides an interface between the processing component and the peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0122] The sensor component includes one or more sensors for providing a status assessment of various aspects of the electronic device. For example, the sensor component can detect the on / off state of the electronic device, the relative positioning of components, such as the display and the keypad of the electronic device. The sensor component can also detect a change in the position of the electronic device or a component of the electronic device, the presence or absence of user contact with the electronic device, the orientation or acceleration / deceleration of the electronic device, and the temperature change of the electronic device. The sensor component may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0123] The communication component is configured to facilitate communication between the electronic device and other devices in a wired or wireless manner. The electronic device can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0124] In an exemplary embodiment, the electronic device can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0125] Embodiment 3:

[0126] The present invention also provides a computer-readable storage medium storing a computer program or instruction of the high-precision target detection model as described in Embodiment 1, and when the program or instruction is executed by a processor, it can cause the processor to execute a high-precision target detection method based on underwater image enhancement.

[0127] Specifically, a system, device, or equipment equipped with a readable storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer or processor of the system, device, or equipment reads and executes the instructions stored in the readable storage medium. In this case, the program code read from the readable medium itself can implement the functions of any one of the above embodiments, so the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of the present invention.

[0128] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-20ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tape, etc. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0129] It should be understood that the above-mentioned processor may be a central processing unit (CPU for short), or other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed and completed by a hardware processor, or executed and completed by a combination of hardware and software modules in the processor.

[0130] It should be understood that the storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium may be located in an application specific integrated circuit (ASIC for short). Of course, the processor and the storage medium may also exist as discrete components in a terminal or a server.

[0131] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0132] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0133] The foregoing is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

[0134] Although the specific embodiments of the present invention have been described above, they are not limitations on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.

Claims

1. A high-precision target detection model based on underwater image enhancement, characterized by: It includes a ResNet network module, a 2D histogram equalization module, and a Faster R-CNN target detection module connected in sequence; The ResNet network module uses the dataset Dataset for training, learning and extracting more complex image features; 16 residual blocks are stacked together to form a deep network structure, and the last layer of the model is set as a convolutional layer, whose output size and number of channels are the same as the input image, ensuring that the output is the enhanced image, completing the deep enhancement of color conversion and image details; The 2D histogram equalization module applies 2D histogram equalization technology to the image frame enhanced by the ResNet network module to further improve the contrast and clarity of the image; The Faster R-CNN target detection module uses the ResNet network module and the 2D histogram equalization module to jointly construct a data set of enhanced images, and trains the data set for classification detection of the images to be detected.

2. A high-precision target detection model based on underwater image enhancement as claimed in claim 1, characterized in that: The process of making the dataset is as follows: Based on the 2s underwater short video data, each segment of data is divided into image frame data d1, d2, ..., d with a sequence length of 50 at a uniform time interval. 50 ; Image frame data d i , the image frames in i∈[1,50] are preprocessed, including adjusting the image size to 320×320 pixels and annotating the corresponding labels of the objects to be detected in the image, and the results are used as a set of data to integrate the data set Dataset.

3. A high-precision target detection model based on underwater image enhancement as claimed in claim 1, characterized in that: The specific process of the ResNet network module setting and processing data is as follows: Improved ResNet residual block settings; each residual block contains three convolutional layers, where the first and third layers use 1x1 filters, which are mainly responsible for adjusting the number of channels, i.e., dimensionality reduction and dimensionality increase; the second layer uses 3x3 filters, focusing on feature extraction; The data input is x l , the convolution operation is F(x l ,W l ), where W l is the weight of the convolutional layer; in addition, in order to learn the complex relationship between the input and the residual, a small convolutional network Conv is introduced in the residual connection to obtain the output of the residual block: x l+1 =F(x l ,W l )+Conv(x l ) 16 improved residual blocks are stacked together to form a deep network structure; The last layer of the model is set as a convolutional layer, whose output size and number of channels are the same as the input image, thus ensuring that the output is the enhanced image. The output formula is: Y=E(X,W) Among them, X is the input image, W represents all the weights in the network; all the image frame data in each set of sequence data are enhanced; the image frame d enhanced by ResNet is obtained iRes , defined as: d iRes =Res(d i ) Among them, Res(d i ) is the original image frame data processed by ResNet.

4. A high-precision target detection model based on underwater image enhancement as claimed in claim 1, characterized in that: The specific processing process of the 2D histogram equalization module is: S1. For the input image frame d iRes , calculate the histograms of the three channels R, G, and B respectively, the formula is: Where H(v) represents the number of pixels with value v, and I(i,j) is the pixel value at position (i,j) of the image; S2. Use the Gaussian smoothing method to smooth the histogram, reduce image noise and improve the statistical characteristics of the histogram. The formula is: Where G(x,y) represents the Gaussian function, x and y are the coordinates of the pixel position, and σ is the standard deviation, which controls the width of the Gaussian kernel; S3. Calculate the cumulative distribution function (CDF) C(v) of the histogram of each channel, the formula is: S4. Use CDF to map the original pixel value to the new intensity value I′(i,j), the formula is: Where MN is the total number of pixels in the image; S5. Apply the mapped pixel values ​​to the image, merge the pixels of each channel, complete the contrast enhancement, and obtain the enhanced image frame d i2D .

5. The high-precision target detection model based on underwater image enhancement as claimed in claim 1, characterized in that: The specific structure of the Faster R-CNN target detection module is: Use lightweight ResNet-34 as the feature extraction network, and then generate the candidate detection box generation network RPN based on the output of the feature extraction network; RPN locates potential target regions by using anchor boxes at multiple scales and aspect ratios; The output of RPN consists of two parts, one for probability prediction of object existence and the other for bounding box regression; the goal of bounding box regression is to adjust the position and size of the anchor box so that it better covers the real object. The adjustment of the bounding box is expressed by the following formula: where x r ,y r ,w r ,h r are the center coordinates and size of the ground-truth bounding box, x a ,y a ,w a ,h a are the center coordinates and size of the corresponding anchor box; The proposal regions of different sizes are pooled to a uniform size, and then the candidate detection boxes are classified and the detection results are output.

6. A high-precision target detection method based on underwater image enhancement, characterized in that: The process includes: Acquire the target image to be detected; Inputting the target image to be detected into the high-precision target detection model according to any one of claims 1 to 5; Output the target category probability of each RPN, P = [p1, p2, ..., p k ], where k is the total number of categories; and the target detection result is the category with the highest probability value, that is: where p i is the probability of the ith class.

7. A high-precision target detection device based on underwater image enhancement, characterized in that: The device includes at least one processor and at least one memory, the processor and the memory are coupled; the memory stores a computer execution program of the high-precision target detection model as described in any one of claims 1 to 5; when the processor executes the computer execution program stored in the memory, the processor executes a high-precision target detection method based on underwater image enhancement.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program or instruction of the high-precision target detection model as described in any one of claims 1 to 5. When the program or instruction is executed by the processor, the processor executes a high-precision target detection method based on underwater image enhancement.

Citation Information

Patent Citations

  • An aerial photograph image tower identification card fault diagnosis method based on depth learning

    CN109376768A

  • Underwater target detection method based on improved Faster R-CNN

    CN116778311A