Water surface image target detection method under motion blur condition

By combining generative adversarial networks and target detectors, the game relationship between the generator and the discriminator is optimized, which solves the problem of motion blur in surface image target detection, achieves high-precision target detection, and ensures the safety of ship intelligent navigation.

CN120655894APending Publication Date: 2025-09-16CHINA SHIP DEV & DESIGN CENT

Patent Information

Application Number
CN202510734713.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing image motion deblurring methods fail to effectively link image quality improvement with target detection tasks in surface image detection, resulting in insufficient detection accuracy for ship intelligent navigation tasks.

Method used

A generative adversarial network is used, combined with the physical model of motion blur, to generate motion blurred images paired with clear images. Through the game relationship between the generator and the discriminator, the generator is optimized to eliminate motion blur, and a highly targeted target detector network structure is designed to improve target detection capabilities.

Benefits of technology

It effectively eliminates the interference of motion blur on surface target detection, improves target positioning and classification accuracy, enhances the ship's intelligent perception capability, and ensures navigation safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655894A_ABST
    Figure CN120655894A_ABST
Patent Text Reader

Abstract

The invention discloses a water surface image target detection method under a motion blurring condition, and relates to the technical field of detection perception, and the method comprises the steps: carrying out the blurring processing of a clear image through combining with a motion blurring physical model, so as to generate a motion blurring image matched with the clear image, and carrying out the target marking of the clear image and the corresponding motion blurring image; training a target detector by using the clear image, after the training process is converged, freezing network parameters of the target detector, training a generator and a discriminator, optimizing the generator in combination with detection feature loss, and generating a restored image by the generator by taking the motion blurred image as input, the discriminator discriminates the quality of the clear image and the quality of the restored image based on related features of target detection, and transmits a discrimination result to the generator; and after the training is completed, removing an identification branch of the discriminator. The motion blur is eliminated with target detection as guidance, the target detection capability is improved, and the safety of intelligent navigation of the ship is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of detection and perception technology, to water surface image enhancement, and in particular to a method for detecting water surface image targets under motion blur conditions. Background Art

[0002] When ships are sailing on the water, they often experience swaying motions such as heave, roll and pitch caused by environmental factors such as wind, waves and currents. This causes motion blur in the images collected by the visual sensors of smart ships, degrades the image quality and weakens the target features, which in turn has an adverse impact on the detection of surface obstacles based on visible light.

[0003] Most existing image motion deblurring methods optimize image quality by building large convolutional neural networks. This not only significantly increases the computational complexity of the network, but also fails to effectively link it with target detection tasks, making it difficult to adapt to intelligent navigation tasks for surface ships. Summary of the Invention

[0004] This invention aims to address the problem of motion blur interference in surface image target detection. Existing methods for deblurring surface images only aim to improve the quality of blurred images, but do not provide image enhancement for target detection, making them unsuitable for intelligent ship navigation. This invention proposes a surface image target detection method under motion blur conditions based on a generative adversarial network. This method eliminates motion blur using target detection as a guide, improving target detection capabilities and ensuring the safety of intelligent ship navigation.

[0005] In a first aspect, the present invention provides a method for detecting targets in a water surface image under motion blur conditions, comprising:

[0006] The clear image is blurred by combining the physical model of motion blur to generate a motion blurred image paired with the clear image, and the clear image and the corresponding motion blurred image are labeled.

[0007] The target detector is trained using clear images. After the training process converges, the network parameters of the target detector are frozen. Then, the generator and discriminator are trained through the game relationship between them. At the same time, the generator is optimized by combining the detection feature loss. The generator uses the motion blurred image as input to generate the restored image. The discriminator distinguishes the quality of the clear image and the restored image based on the relevant features of target detection, and passes the identification result to the generator, guiding the generator to learn to remove motion blur for the target detection task.

[0008] After the network training is completed, the identification branch of the discriminator is removed, and the target detection in the water surface image under motion blur conditions is completed only by the generator and detector.

[0009] In some instances, IB =K*I S +Q blurs the clear image to generate a motion blurred image paired with the clear image, where * represents the convolution operation, I S , I B are clear images and motion blurred images respectively, K is the blur kernel containing motion blur trajectory information, and Q is noise.

[0010] In some instances, Construct the game relationship between the generator and the discriminator, where N is the number of images used for network training, I S , I B 、G(I B ) are clear images, motion blurred images and restored images respectively, D is the discriminator, G is the generator, λ GP is the gradient penalty term coefficient, x is the calculation input of the gradient penalty term, E x Indicates the “[||▽ x D(x)||-1]".

[0011] In some instances, Construct detection feature loss, where and are the feature maps of the clear image and the restored image in the target detector backbone network, W, H, and C are the width, height, and number of channels of the feature map, respectively, and i, j, and k are the three-dimensional coordinates in the feature map.

[0012] In some instances, the front end of the generator performs preliminary extraction of motion blurred image features through three convolutional layers, and the back end reconstructs the resolution of the motion blurred image through two deconvolution layers and one convolution layer. Nine receptive field modules are set in the middle part of the generator. The receptive field module combines the advantages of residual network, Inception network and void convolution, and at the same time combines the 1×1 convolution set in the receptive field module to adjust the number of feature channels and set a global jump connection structure.

[0013] In some instances, the target detector uses the YOLOv3 network and optimizes the head network structure of YOLOv3. Convolution kernels with large aspect ratios (1×3, 3×5) and 1:1 aspect ratio are designed to fit the contours of ship targets under different perspectives.

[0014] In some instances, three convolutional branches are derived from the backbone network of the object detector to build the discriminator.

[0015] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:

[0016] In order to eliminate the influence of motion blur on surface target detection during the intelligent navigation of ships, the present invention proposes a surface image target detection method under motion blur conditions based on a generative adversarial network. By constructing a multi-scale discriminator on the basis of the target detector, the deblurring identification process can share the relevant features of target detection, thereby effectively linking the image de-motion blurring task with the target detection task. In addition, a detection feature loss is designed to guide the generator to minimize the detection feature error between the restored image and the clear image, thereby realizing target detection-oriented de-motion blurring. Finally, the network structure of the target detector is designed specifically according to the appearance characteristics of the ship target, so as to more effectively detect the ship target. Compared with the traditional surface target detection method, the present invention can eliminate the interference of motion blur on surface image detection, effectively improve the target positioning and classification accuracy, thereby enhancing the intelligent perception capability of the ship and ensuring the safety of ship navigation. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 This is a schematic diagram of the overall network architecture provided by an embodiment of the present invention;

[0019] Figure 2 Schematic diagram of detection feature loss provided by an embodiment of the present invention;

[0020] Figure 3 Schematic diagram of the network structure of the receptive field module provided by an embodiment of the present invention;

[0021] Figure 4 Schematic diagram of a detector head network provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0023] In the following description, specific embodiments of the present invention will be described with reference to steps and symbols performed by one or more computers, unless otherwise specified. Therefore, these steps and operations will be mentioned several times as being performed by a computer, and computer execution as referred to herein includes operations by a computer processing unit that represents electronic signals of data in a structured form. This operation converts the data or maintains it at a location in the computer's memory system, which can be reconfigured or otherwise change the operation of the computer in a manner familiar to testers in the field. The data structure in which the data is maintained is a physical location in the memory, which has specific characteristics defined by the data format. However, the principles of the present invention are described in the above text, which does not represent a limitation, and testers in the field will understand that the various steps and operations below can also be implemented in hardware.

[0024] As used herein, the terms "module" or "unit" may be considered software objects executed on the computing system. The various components, modules, engines, and services herein may be considered implementation objects on the computing system. While the devices and methods herein are preferably implemented in software, they may also be implemented in hardware and remain within the scope of protection of the present invention.

[0025] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0026] The present invention is aimed at intelligent navigation tasks of ships and is designed in combination with a generative adversarial network. The specific implementation method includes building a network game relationship, building a generator, building a detector, building a discriminator based on the detector, creating a data set, training and testing.

[0027] In the embodiment of the present invention, Figure 1 As shown, a method for detecting targets in a water surface image under motion blur conditions is provided, comprising:

[0028] The clear image is blurred by combining the physical model of motion blur to generate a motion blurred image paired with the clear image, and the clear image and the corresponding motion blurred image are labeled.

[0029] The target detector is trained using clear images. After the training process converges, the network parameters of the target detector are frozen. Then, the generator and discriminator are trained through the game relationship between them. At the same time, the generator is optimized by combining the detection feature loss. The generator uses the motion blurred image as input to generate the restored image. The discriminator distinguishes the quality of the clear image and the restored image based on the relevant features of target detection, and passes the identification result to the generator, guiding the generator to learn to remove motion blur for the target detection task.

[0030] After the network training is completed, the identification branch of the discriminator is removed, and the target detection in the water surface image under motion blur conditions is completed only by the generator and detector.

[0031] In an embodiment of the present invention, training and testing of the present invention require corresponding clear image and blurred image data. However, in real-world scenarios, once motion blur occurs in an image, it is often difficult to obtain a clear image that completely corresponds to the blurred image in that scenario. The present invention uses a physical model of motion blur to blur the clear image, generating a motion blurred image that is paired with the clear image for network training and testing. The physical model of motion blur is shown in Equation (1).

[0032] I B =K*I S +Q (1)

[0033] Among them, * represents the convolution operation, I S , I B are the clear image and the blurred image respectively, K is the blur kernel containing the motion blur trajectory information, and Q is the noise. S After convolution with blur kernel K and adding noise Q, a blurred image I is generated. B .

[0034] In addition, clear images and blurred images need to be labeled for target processing for detector training and testing.

[0035] In an embodiment of the present invention, the training process is divided into two parts. First, the detector is trained using a clear image. When the training process converges, the network parameters of the detector are frozen to solidify its target detection capability. Afterwards, based on the game relationship constructed by formula (2), the generator and the discriminator are trained, and the generator is optimized by combining the detection feature loss of formula (3). The generator takes the blurred image as input and generates a restored image. The discriminator identifies the quality of the clear image and the restored image based on the relevant features of target detection, and passes the identification results to the generator, guiding the generator to learn to remove motion blur for the target detection task.

[0036] After network training is completed, the generator has learned how to eliminate motion blur, and the detector's target detection capability has also been solidified. The detector can identify targets in the image using the generator's restored image as input. Therefore, the identification branch of the discriminator can be removed, reducing network parameters. Target detection in water surface images under motion blur conditions can be completed using only the generator and detector.

[0037] In the embodiment of the present invention, establishing a network game relationship can be achieved by the following methods:

[0038] The network structure of the present invention is as follows Figure 1 As shown in the figure, the generator is responsible for reconstructing the blurred image and generating a high-quality restored image as much as possible, striving to make the discriminator mistake the restored image for a clear image. The discriminator continuously identifies the quality of the clear image and the generator's restored image, and continuously improves its identification ability to prevent being deceived by the generator. The discriminator gives high scores to the clear image and low scores to the restored image, and at the same time feeds the identification results back to the generator. The generator will try to continuously improve the discriminator's evaluation of the restored image, thereby improving the quality of the restored image through the game relationship between the generator and the discriminator. The expression of the game relationship between the two is:

[0039]

[0040] Where N is the number of images used for network training, I S , I B 、G(I B ) are clear images, motion blurred images, and restored images, respectively. D is the discriminator and G is the generator. To ensure the stability of the training process, a gradient penalty term is added, i.e., the second term in the formula, λ GP is the coefficient of this term, and x is the calculation input of the gradient penalty term. During the training process, the discriminator continuously maximizes formula (2), that is, it reduces the restored image score D(G(I B )) to improve the clear image score D(I S ), thereby continuously optimizing the model parameters of the discriminator to ensure that it can effectively distinguish between restored images and clear images; the generator continuously minimizes formula (2), that is, allows the discriminator to improve the score of the restored image, in this way to build a game relationship between the two.

[0041] In an embodiment of the present invention, a multi-scale discriminator is constructed based on the target detector in combination with the generative adversarial network architecture, and the detection features are used to deblur and discriminate targets of different sizes in a more refined manner.

[0042] Generative adversarial networks (GANs) primarily consist of two modules: a discriminator and a generator. Through the interplay between the generator and the discriminator, they can effectively capture abstract feature differences in image data, offering significant advantages in image processing tasks. In this paper, the generator takes a motion-blurred water surface image as input and generates a restored image as clear as possible. The discriminator, built on the backbone network of the target detector, discriminates the quality of the clear image and the generator's restored image, assigning a high score if the image is clear and a low score if the image is restored. The discriminator's identification results are fed back to the generator, which improves the quality of the restored image by continuously improving the discriminator's evaluation of the restored image. This approach allows the deblurring discriminator to share the image features of the target detection backbone network, enabling it to identify deblurring based on features related to target detection, thereby effectively associating it with the target detection task. Furthermore, the multi-scale discriminator, consisting of multiple network branches, each with a receptive field of varying size, helps the network adapt to motion blur discrimination at different scales.

[0043] In an embodiment of the present invention, a detection feature loss is constructed based on a backbone network of an object detector to guide the generator to learn to de-motion blur of water surface images for object detection.

[0044] To enhance feature consistency between the clear and restored images, many deblurring methods employ pre-trained convolutional neural networks (CNNs) to design perceptual losses. While these deep CNNs can improve image restoration quality thanks to their superior feature extraction capabilities, the features they extract are often specific to image processing tasks and have little relevance to object detection.

[0045] In order to achieve deblurring that is more conducive to target detection, the present invention proposes a detection feature loss based on the pre-trained target detector backbone network. The implementation process is as follows: Figure 2 As shown in Figure 2, the pre-trained object detector backbone network extracts detection features from the restored and clear images for object classification and regression, calculates the corresponding feature errors, and then passes these errors to the generator. The generator learns image restoration for object detection tasks by minimizing this detection feature loss.

[0046] The detection feature loss expression is:

[0047]

[0048] in, and are the feature maps of the clear image and the restored image in the pre-trained detector backbone network, W, H, and C are the width, height, and number of channels of the feature map, respectively, and i, j, and k are the three-dimensional coordinates in the feature map.

[0049] In an embodiment of the present invention, the generator is constructed in the following manner:

[0050] The front end of the generator uses three convolutional layers to perform preliminary extraction of blurred image features, and the back end uses two deconvolutional layers and one convolutional layer to reconstruct the resolution of the blurred image. The middle part of the generator is equipped with nine receptive field modules, such as Figure 3 As shown in the figure, this module combines the advantages of residual networks, Inception networks, and dilated convolutions. It can simulate the characteristics of the human visual receptive field to deeply extract image features. At the same time, the 1×1 convolution set in the module adjusts the number of feature channels to improve the network's computational efficiency. In addition, by setting up a global skip connection structure, the generator does not need to learn the complete mapping from blurred images to clear images. It only needs to learn the residual between the two to complete the restoration of blurred images. This can reduce the learning difficulty of the generator and shorten the training process.

[0051] In an embodiment of the present invention, the convolution kernel size of the detector head network is designed according to the appearance characteristics of the ship target, so as to more effectively detect the ship target.

[0052] Most current object detectors are designed for general-purpose scenarios. Customizing detectors for intelligent ship navigation is particularly important. Ships are the most common obstacles in intelligent navigation, and their appearance varies with viewing angle. When viewed from the side, a ship's length is significantly greater than its width, while from the front, its aspect ratio is close to 1.

[0053] In view of the shape characteristics of the ship target, the embodiment of the present invention designs and optimizes the convolution kernel size of the target detector head network. Multiple branches are designed in the head network. The convolution kernel sizes in the branches include convolution kernels with a length larger than the width (1×3, 3×5) and convolution kernels with an aspect ratio of 1, so as to more effectively capture the characteristics of the ship target. The head network design is as follows: Figure 4 shown.

[0054] In an embodiment of the present invention, the detector can be constructed in the following manner:

[0055] The pre-trained detector is the basis of the discriminator. The discriminator identifies the image quality based on the target detection-related features, thereby associating the de-motion blurring task with the target detection task. The pre-trained detector of the present invention uses the YOLOv3 network. This model has good real-time performance while ensuring detection accuracy and is widely used in engineering tasks. At the same time, in order to effectively detect ship targets, the head network structure of YOLOv3 is optimized in the embodiment of the present invention. By designing convolution kernels with large aspect ratios (1×3, 3×5) and 1:1 aspect ratios, the contours of ship targets under different perspectives are fitted, thereby capturing ship target features more specifically.

[0056] In an embodiment of the present invention, a discriminator can be constructed based on a detector in the following manner:

[0057] After the detector is built, three convolutional branches are derived from the detector's backbone network to construct a discriminator. This discriminator combines relevant features for object detection to determine image quality. Each discriminator branch outputs a two-dimensional matrix. The value of each element in the matrix represents the discriminator's local discrimination result for an image patch, and the entire matrix represents the discriminator's discrimination result for the entire image. Each branch in the discriminator operates at a different network depth and therefore has a different receptive field size, enabling multi-scale image discrimination.

[0058] In addition, clear images and blurred images need to be labeled for target processing for detector training and testing.

[0059] The above is a detailed introduction to a method for detecting targets in water images under motion blur conditions provided by an embodiment of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, based on the ideas of the present invention, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A method for detecting targets in water surface images under motion blur conditions, characterized in that: include: The clear image is blurred by combining the physical model of motion blur to generate a motion blurred image paired with the clear image, and the clear image and the corresponding motion blurred image are labeled. The target detector is trained using clear images. After the training process converges, the network parameters of the target detector are frozen. Then, the generator and discriminator are trained through the game relationship between them. At the same time, the generator is optimized by combining the detection feature loss. The generator uses the motion blurred image as input to generate the restored image. The discriminator distinguishes the quality of the clear image and the restored image based on the relevant features of target detection, and passes the identification result to the generator, guiding the generator to learn to remove motion blur for the target detection task. After the network training is completed, the identification branch of the discriminator is removed, and the target detection in the water surface image under motion blur conditions is completed only by the generator and detector.

2. The method according to claim 1, characterized in that byI B =K*I S +Q blurs the clear image to generate a motion blurred image paired with the clear image, where * represents the convolution operation, I S , I B are clear images and motion blurred images respectively, K is the blur kernel containing motion blur trajectory information, and Q is noise.

3. The method according to claim 2, characterized in that Depend on Construct the game relationship between the generator and the discriminator, where N is the number of images used for network training, I S , I B 、G(I B ) are clear images, motion blurred images and restored images respectively, D is the discriminator, G is the generator, λ GP is the gradient penalty term coefficient, x is the calculation input of the gradient penalty term, E x Indicates [||▽ x D(x)||-1].

4. The method according to claim 3, characterized in that Depend on Construct detection feature loss, where and are the feature maps of the clear image and the restored image in the target detector backbone network, W, H, and C are the width, height, and number of channels of the feature map, respectively, and i, j, and k are the three-dimensional coordinates in the feature map.

5. The method according to claim 4, characterized in that The front end of the generator performs preliminary extraction of motion blurred image features through three convolutional layers, and the back end reconstructs the resolution of the motion blurred image through two deconvolution layers and one convolution layer. Nine receptive field modules are set in the middle part of the generator. The receptive field module combines the advantages of residual network, Inception network and void convolution, and at the same time adjusts the number of feature channels by combining the 1×1 convolution set in the receptive field module, and sets a global skip connection structure.

6. The method according to claim 5, characterized in that The target detector uses the YOLOv3 network, and the head network structure of YOLOv3 is optimized. Convolution kernels with large aspect ratios (1×3, 3×5) and 1:1 aspect ratio are designed to fit the contours of ship targets under different perspectives.

7. The method according to claim 6, characterized in that Three convolution branches are derived from the backbone network of the target detector to build the discriminator.

Citation Information

Patent Citations

  • Water surface motion blurred image enhancement method for intelligent navigation of ship

    CN119579462A

Cited By

  • Detection model generation method, monocular 3D target detection method and monocular 3D target detection device

    CN120877273A