Training and Detection Methods for Low-Quality Image Target Detection Models, Related Devices

By constructing an object detection model for low-quality images, the problem of low-quality objects detection accuracy and efficiency caused by low-quality images is solved, and efficient object detection under harsh conditions is achieved.

CN119048735BActive Publication Date: 2025-06-10HARBIN ENGINEERING UNIVERSITY SANYA NANHAI INNOVATION & DEVELOPMENT BASE +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411076362.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2025-06-10
Estimated Expiration
2044-08-07

AI Technical Summary

Technical Problem

In severe weather and complex sea surface conditions, low-quality images lead to low accuracy and efficiency of target detection, making it difficult to achieve efficient target detection.

Method used

By acquiring low-quality sample images, the initial features are extracted using the first encoder, and inputting them into the migration convolutional network for training, the target convolutional network is obtained. Then, based on the target convolution network and the target detection head, an object detection model is constructed to perform object detection on low-quality images to be detected.

Benefits of technology

It realizes efficient object detection for low-quality images, improves the accuracy and efficiency of object detection, and can effectively identify target objects under harsh conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119048735B_ABST
    Figure CN119048735B_ABST
Patent Text Reader

Abstract

The present application provides a training and detection method for a low-quality image target detection model and related devices. The training method includes: obtaining a low-quality sample image for a target area; inputting the low-quality sample image into a first encoder to obtain an initial feature of the low-quality sample image; inputting the initial feature into a transfer convolutional network and training the transfer convolutional network based on a first target feature to obtain a target convolutional network, where the first target feature is obtained based on the initial feature; obtaining a second target feature for the low-quality sample image based on the target convolutional network; inputting the second target feature into a first detection head and training the first detection head to obtain a target detection head, where the first detection head is obtained based on the first target feature; obtaining a target detection model based on the first encoder, the target convolutional network, and the target detection head; and the target detection model is used for target detection of a low-quality image to be detected. It provides technical support for achieving efficient target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a method for training and detecting a low-quality image target detection model, and related devices. Background Art

[0002] In scenarios such as offshore operations, visual target detection technology can provide a lot of help for operation decision-making. However, in scenarios such as offshore operations, the quality of the captured images often fails to meet the standards due to problems such as bad weather and complex sea conditions, resulting in low accuracy and efficiency of target detection. Therefore, how to achieve efficient target detection has become a technical problem to be solved urgently. Summary of the Invention

[0003] This application provides a method for training and detecting a low-quality image target detection model, and related devices, so as to solve at least the above technical problems existing in the prior art.

[0004] This application provides a method for training a low-quality image target detection model, and the method includes:

[0005] Obtain a low-quality sample image for a target area;

[0006] Input the low-quality sample image into a first encoder to obtain initial features of the low-quality sample image;

[0007] Input the initial features into a transfer convolutional network, and train the transfer convolutional network based on first target features to obtain a target convolutional network; wherein, the first target features are obtained based on the initial features of the low-quality sample image; the feature clarity of the first target features is stronger than that of the initial features;

[0008] Based on the target convolutional network, obtain second target features for the low-quality sample image;

[0009] Input the second target features into a first detection head, and train the first detection head to obtain a target detection head; the first detection head is obtained based on the first target features;

[0010] Based on the first encoder, the target convolutional network, and the target detection head, obtain a target detection model; the target detection model is used for target detection of a low-quality image to be detected.

[0011] In the above solution, the step of inputting the initial features into a transfer convolutional network, and training the transfer convolutional network based on first target features to obtain a target convolutional network includes:

[0012] Input the initial features into the transfer convolutional network, and use the first target feature as the feature constraint output by the transfer convolutional network to train the transfer convolutional network to obtain the target convolutional network.

[0013] In the above solution, the first target feature is obtained based on the initial features of the low-quality sample image, including:

[0014] Input the initial features of the low-quality sample image into a pre-trained image type prediction model to obtain the low-quality type of the low-quality sample image;

[0015] Based on the low-quality type of the low-quality sample image, obtain the first target feature for the low-quality sample image.

[0016] In the above solution, the obtaining of the first target feature for the low-quality sample image based on the low-quality type of the low-quality sample image includes:

[0017] Input the low-quality type of the low-quality sample image into a pre-trained content mapping network to obtain mapping coding information for the low-quality type;

[0018] Input the mapping coding information and the initial features of the low-quality sample image into a first decoder to obtain a restored sample image; the clarity of the restored sample image is stronger than that of the low-quality sample image;

[0019] Based on the restored sample image, obtain the first target feature for the low-quality sample image.

[0020] In the above solution, the obtaining of the first target feature for the low-quality sample image based on the restored sample image includes:

[0021] Input the restored sample image and the low-quality sample image into a second encoder to obtain the first target feature;

[0022] The first detection head is obtained based on the first target feature, including:

[0023] Input the first target feature into a second detection head to be trained, and train the second detection head to be trained to obtain the first detection head.

[0024] This application provides a low-quality image target detection method, and the method includes:

[0025] Obtain a low-quality image to be detected for the area to be detected;

[0026] Input the low-quality image to be detected into the target detection model to obtain the target detection result for the region to be detected; wherein, the target detection result is obtained by the first encoder in the target detection model extracting features from the low-quality image to be detected, the target convolution network in the target detection model migrating the extracted features, and the target detection head in the target detection model performing target detection based on the migrated features; the target detection result is used to represent whether there is a target object in the region to be detected.

[0027] The present application provides a training device for a low-quality image target detection model, and the device includes:

[0028] The first acquisition unit is used to acquire a low-quality sample image for the target region;

[0029] The second acquisition unit is used to input the low-quality sample image into the first encoder to obtain the initial features of the low-quality sample image;

[0030] The third acquisition unit is used to input the initial features into the migration convolution network, and train the migration convolution network based on the first target features to obtain the target convolution network; wherein, the first target features are obtained based on the initial features of the low-quality sample image; the feature clarity of the first target features is stronger than that of the initial features;

[0031] The fourth acquisition unit is used to obtain the second target features for the low-quality sample image based on the target convolution network;

[0032] The fifth acquisition unit is used to input the second target features into the first detection head and train the first detection head to obtain the target detection head; the first detection head is obtained based on the first target features;

[0033] The sixth acquisition unit is used to obtain a target detection model based on the first encoder, the target convolution network, and the target detection head; the target detection model is used to perform target detection on the low-quality image to be detected.

[0034] The present application provides a low-quality image target detection device, and the device includes:

[0035] The seventh acquisition unit is used to acquire a low-quality image to be detected for the region to be detected;

[0036] The detection unit is configured to input the low-quality image to be detected into a target detection model to obtain a target detection result for the region to be detected. The target detection result is obtained by the first encoder in the target detection model extracting features from the low-quality image to be detected, the target convolution network in the target detection model migrating the extracted features, and the target detection head in the target detection model performing target detection based on the migrated features. The target detection result is used to indicate whether there is a target object in the region to be detected.

[0037] This application provides an electronic device, including:

[0038] At least one processor; and

[0039] A memory communicatively connected to the at least one processor. Wherein,

[0040] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of this application.

[0041] This application provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute the method of this application.

[0042] In this application, a low-quality sample image for a target region is obtained; the low-quality sample image is input into a first encoder to obtain initial features of the low-quality sample image; the initial features are input into a migration convolution network, and based on first target features, the migration convolution network is trained to obtain a target convolution network. The first target features are obtained based on the initial features of the low-quality sample image, and the feature clarity of the first target features is stronger than that of the initial features. Based on the target convolution network, second target features for the low-quality sample image are obtained; the second target features are input into a first detection head, and the first detection head is trained to obtain a target detection head. The first detection head is obtained based on the first target features; based on the first encoder, the target convolution network, and the target detection head, a target detection model is obtained. The target detection model is used to perform target detection on a low-quality image to be detected. Efficient target detection can be achieved.

[0043] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of this application, nor is it used to limit the scope of this application. Other features of this application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present application will become readily understood. In the drawings, several embodiments of the present application are shown in an exemplary rather than restrictive manner, where:

[0045] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.

[0046] Figure 1 A schematic diagram of the implementation process of the training method for the low-quality image target detection model according to the embodiment of the present application is shown;

[0047] Figure 2 A schematic diagram of the acquisition process of the target convolutional network according to the embodiment of the present application is shown;

[0048] Figure 3 A schematic diagram of the acquisition process of the first detection head according to the embodiment of the present application is shown;

[0049] Figure 4 A schematic diagram of the composition structure of the image type prediction model according to the embodiment of the present application is shown;

[0050] Figure 5 A schematic diagram of the processing process of the content mapping network according to the embodiment of the present application is shown;

[0051] Figure 6 A schematic diagram of the structural comparison between the target detection head and the first detection head according to the embodiment of the present application is shown;

[0052] Figure 7 A schematic diagram of the implementation process of the low-quality image target detection method according to the embodiment of the present application is shown;

[0053] Figure 8 A schematic diagram of the composition structure of the target detection model according to the embodiment of the present application is shown;

[0054] Figure 9 A schematic diagram of the composition structure of the training device for the low-quality image target detection model according to the embodiment of the present application is shown;

[0055] Figure 10 A schematic diagram of the composition structure of the low-quality image target detection device according to the embodiment of the present application is shown;

[0056] Figure 11 A schematic diagram of the composition structure of an electronic device according to the embodiment of the present application is shown. Detailed implementation manners

[0057] To make the objectives, features, and advantages of the present application more obvious and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0058] In the embodiments of the present application, a low-quality image refers to a distorted image with a clarity less than a preset (or standard) threshold, resulting in the inability to perform (accurate) target detection on it. The low-quality sample images and low-quality images to be detected mentioned in the embodiments of the present application are all distorted images with a clarity less than a preset (or standard) threshold.

[0059] The technical solutions in the embodiments of the present application involve a solution for training a low-quality image target detection model and a solution for performing target detection on the obtained low-quality images to be detected using the trained target detection model. The technical solutions of the present application can achieve the efficient acquisition of the target detection model and thus achieve the efficient detection of the target.

[0060] The embodiments of the present application provide a method for training a low-quality image target detection model, as Figure 1 shown, the method includes:

[0061] S101: Obtain low-quality sample images for a target area.

[0062] In this step, the target area is the area where the known target object is located. A low-quality sample image refers to an image with poor quality (low clarity). For example, in a maritime operation scenario, the low-quality sample image can be a low-quality image obtained by photographing the target area due to common high-humidity conditions such as fog and rain at sea, or due to problems such as the rapid movement of boats, harsh sea conditions, complex water surface reflections, clutter environments, and strong interference from dynamic backgrounds. That is, the low-quality sample image is a low-quality image including the target object.

[0063] S102: Input the low-quality sample image into the first encoder to obtain the initial features of the low-quality sample image.

[0064] In this step, the first encoder is an encoder for image restoration. Through image restoration, a low-quality sample image can be restored to a relatively clear image, thereby improving the accuracy of subsequent target detection model training. The low-quality sample image is input into the first encoder, and the first encoder can extract features from the low-quality sample image to obtain the initial features of the low-quality sample image. The first encoder can adopt a convolutional module, a residual module, or a Swin Transformer (shifted window hierarchical vision model) module, etc. Usually, in a way of downsampling layer by layer, the features of the image (low-quality sample image) are collected at 3 to 4 scales of convolutional features, aiming to select the deep semantic information of the image. For the specific application of the first encoder, please refer to the detailed description in the relevant parts below and will not be elaborated here.

[0065] S103: Input the initial features into the transfer convolutional network, and train the transfer convolutional network based on the first target features to obtain the target convolutional network; wherein, the first target features are obtained based on the initial features of the low-quality sample image; the feature clarity of the first target features is stronger than that of the initial features.

[0066] In this step, the first target features are pre-obtained features that can be used for target detection. The first target features are obtained based on the initial features of the low-quality sample image, and their feature clarity is stronger than that of the initial features. For the specific obtaining process of the first target features, please refer to the detailed description in the relevant parts below and will not be elaborated here.

[0067] After obtaining the initial features of the low-quality sample image in step S102, input the initial features into the transfer convolutional network (the convolutional network to be trained), and train the transfer convolutional network based on the first target features to obtain the trained transfer convolutional network, that is, the target convolutional network. The target convolutional network is used for feature transfer, that is, to transfer the initial features of the low-quality sample image into features that can be used for target detection (that is, the second target features described below).

[0068] S104: Based on the target convolutional network, obtain the second target features for the low-quality sample image.

[0069] In this step, based on the trained target convolutional network, the initial features can be directly transferred into features that can be used for object detection, namely the second target features. The feature clarity of the second target features is stronger than that of the initial features. Also, since the second target features are obtained through the target convolutional network trained based on the first target features, considering that deep learning training cannot be 100% accurate, the feature clarity of the second target features is less than or equal to that of the first target features. The reason for using the target convolutional network to transfer the initial features is that compared with the solution of using multiple models to process serially to obtain the final features for object detection, directly using the trained target convolutional network to transfer the initial features to obtain the features for object detection can greatly reduce the computational amount of data, reduce the processing time, ensure the accuracy of object detection, and achieve lightweight object detection while improving the efficiency of object detection.

[0070] S105: Input the second target features into the first detection head, train the first detection head, and obtain the target detection head; the first detection head is obtained based on the first target features.

[0071] In this step, the first detection head is a detection head pre-obtained based on the first target features and is used to perform the steps of object detection. After obtaining the features for object detection, that is, the second target features, in order to ensure that the first detection head can better adapt to the second target features, thereby achieving the training consistency between the target convolutional network and the first detection head, the second target features are input into the first detection head to train it and obtain the target detection head. The target detection head is used to finally perform the steps of object detection.

[0072] S106: Based on the first encoder, the target convolutional network, and the target detection head, obtain the target detection model; the target detection model is used to perform object detection on low-quality images to be detected.

[0073] In this step, the first encoder, together with the target convolutional network and the target detection head obtained through the foregoing steps, constitutes the target detection model. This target detection model can be applied to the actual scenario to perform object detection on low-quality images to be detected.

[0074] In the solution shown in steps S101 to S106, a low-quality sample image for a target area is obtained; the low-quality sample image is input into a first encoder to obtain initial features of the low-quality sample image; the initial features are input into a transfer convolutional network, and the transfer convolutional network is trained based on first target features to obtain a target convolutional network; wherein, the first target features are obtained based on the initial features of the low-quality sample image; the feature clarity of the first target features is stronger than that of the initial features; based on the target convolutional network, second target features for the low-quality sample image are obtained; the second target features are input into a first detection head, and the first detection head is trained to obtain a target detection head; the first detection head is obtained based on the first target features; based on the first encoder, the target convolutional network, and the target detection head, a target detection model is obtained; the target detection model is used for target detection of a low-quality image to be detected. Efficient target detection can be achieved.

[0075] In an alternative solution, the inputting the initial features into a transfer convolutional network and training the transfer convolutional network based on first target features to obtain a target convolutional network includes:

[0076] Inputting the initial features into a transfer convolutional network, using the first target features as the feature constraint of the output of the transfer convolutional network, and training the transfer convolutional network to obtain a target convolutional network.

[0077] In this application, as shown in Figure 2 , the initial features obtained after passing the low-quality sample image through the first encoder are input into a transfer convolutional network, and the first target features are used as the output feature constraint of the transfer convolutional network to train the transfer convolutional network to obtain a target convolutional network. Considering that the initial features are features for image restoration extracted by the first encoder, while the first target features are features for target detection. There is a domain difference between the two, and the initial features cannot be directly used for target detection. The initial features need to be transferred before target detection can be performed. Specifically, the initial features are input into a transfer convolutional network, and the first target features are used as the feature constraint of the output of the transfer convolutional network to train the transfer convolutional network. Among them, the loss function of the transfer convolutional network is shown in formulas (1) to (4):

[0078] L = L p (f u , k u ) + L r (f u , k u ) + L t (f u , k u ) Formula (1)

[0079] Wherein,

[0080]

[0081] Among them, in Formulas (1) to (4), L is the loss function value of the transfer convolutional network. f is the initial feature, k is the first target feature. u is the number of features. L p (f u , k u ) is the point - wise feature difference between the initial feature and the first target feature. L r (f u , k u ) is the regional feature difference between the initial feature and the first target feature. L t (f u , k u ) is the channel - wise feature difference between the initial feature and the first target feature. c, h, and w are the three dimensions of the feature. c represents the number of channels of the feature, h represents the height of the feature, and w represents the width of the feature. f u,c,w,h represents the initial feature composed of u features when the number of channels is c, the feature height is h, and the feature width is w. k u,c,w,h represents the first target feature composed of u features when the number of channels is c, the feature height is h, and the feature width is w.

[0082] When the loss function of the transfer convolutional network is minimized, that is, when the value of L in the aforementioned Formula (1) is the smallest, the training of the transfer convolutional network ends, and the trained target convolutional network is obtained. The target convolutional network can directly transfer the initial features of the low - quality (sample) images extracted into features that can be used for target detection, and can achieve the lightweight design of the model without affecting the accuracy of the model.

[0083] In an alternative solution, the first target feature is obtained based on the initial features of the low - quality sample images, including:

[0084] Input the initial features of the low - quality sample images into a pre - trained image type prediction model to obtain the low - quality type of the low - quality sample images;

[0085] Based on the low - quality type of the low - quality sample images, obtain the first target feature for the low - quality sample images.

[0086] In this application, as shown in Figure 3 , input the initial features obtained by passing the low - quality sample images through the first encoder into the image type prediction model. The image type prediction model is a pre - trained model used to predict the low - quality type of low - quality sample images. The image type prediction model is trained by learning the features of images of different low - quality types in advance. The structure of the image type prediction model is as Figure 4As shown, it includes a convolutional layer, a pooling layer, and a fully connected layer. Among them, the convolutional layer and the pooling layer are used to extract features, and the fully connected layer is used to output the low-quality type (i.e., the reason causing the image to be of low quality). Exemplarily, in an offshore operation scenario, the low-quality types output by the fully connected layer include categories such as fog, rain, motion blur, sea clutter, etc. This layer is used to analyze what type of harsh conditions the current (low-quality sample) image has experienced, so as to use it as a high-level semantic prior to guide the first decoder (see the following related description) to complete the reconstruction (restoration) of the clear image from the encoded features (initial features). The low-quality type of the same low-quality sample image can be one or multiple. Exemplarily, assume there is a low-quality sample image A. Inputting the initial features of image A into the image type prediction model can obtain the low-quality type (and its proportion) of image A. For example, the reasons causing image A to be of low quality include: wave jitter (20%), water surface reflection (80%). Based on the low-quality type of the low-quality sample image, the first target feature for the low-quality sample image can be obtained. For the specific process of obtaining the first target feature, please refer to the detailed description in the following related parts and will not be elaborated here.

[0087] By predicting the low-quality type of the low-quality sample image, it provides a data basis for subsequently restoring the low-quality sample image to a clear image in a targeted manner, and provides technical support for achieving accurate object detection.

[0088] In an optional solution, obtaining the first target feature for the low-quality sample image based on the low-quality type of the low-quality sample image includes:

[0089] Inputting the low-quality type of the low-quality sample image into a pre-trained content mapping network to obtain mapping coding information for the low-quality type;

[0090] Inputting the mapping coding information and the initial features of the low-quality sample image into the first decoder to obtain a restored sample image; the clarity of the restored sample image is stronger than that of the low-quality sample image;

[0091] Based on the restored sample image, obtain the first target feature for the low-quality sample image.

[0092] In this application, as shown in Figure 3 Output (low-quality type) of the image type prediction model is used as the input of the content mapping network, and the output (mapping coding information) of the content mapping network is used as the input of the first decoder. The composition of the content mapping network refers to Figure 5As shown in the figure, the content mapping network takes the output of the image type prediction model as input, degrades it into a cause probability (the probability of each low-quality type causing image degradation, such as 80% and 20% mentioned above), expands it into a two-dimensional tensor consistent with the corresponding feature map scale, then performs a unified convolution once, and then performs two different convolutions respectively to generate γ and β with the same number of channels and the same size as the processed features (features of low-quality types). Then multiply γ by the processed features and add β. This is equivalent to regularizing each pixel point of each channel of the low-quality type features separately, thereby fusing to obtain prior information, encoding the prior information, and then embedding it into the feature space of the first decoding layer for feature transformation calculation. The structure of the first decoder is the same as that of the first encoder, and convolutional modules, residual modules, or Swin Transformer modules can also be used. The difference is that the first decoder uses a layer-by-layer upsampling method to gradually reconstruct the features extracted by the first encoder to the spatial resolution of the low-quality sample image, and finally outputs a clear image (restored sample image) with the same spatial and channel size as the low-quality sample image. It can be understood that the first decoder is a decoder for image restoration. The pre-trained content mapping network is used to convert the output of the image type prediction model into an input that the first decoder can recognize, that is, to convert the low-quality type of the low-quality sample image into mapping coding information that the first decoder can recognize. Using the mapping coding information and the initial features of the low-quality sample image as the input of the first decoder together, the first decoder specifically reconstructs the initial features based on the mapping coding information to obtain a restored sample image (a clear image). This provides a data basis for accurate object detection. The first target feature of the low-quality sample image is obtained based on the restored sample image. For the specific process, please refer to the detailed description in the relevant parts below and will not be elaborated here.

[0093] In an optional solution, obtaining the first target feature for the low-quality sample image based on the restored sample image includes:

[0094] Inputting the restored sample image and the low-quality sample image into a second encoder to obtain the first target feature;

[0095] The first detection head is obtained based on the first target feature, including:

[0096] Inputting the first target feature into a second detection head to be trained, and training the second detection head to be trained to obtain the first detection head.

[0097] In this application, the second encoder is an encoder for object detection. Refer to Figure 3As shown, the restored sample image and the low-quality sample image are input into the second encoder together. The second encoder extracts features from them, which can further ensure the accuracy of feature extraction by the second encoder and obtain clearer and more accurate first target features. The first target features are input into the second detection head to be trained, and the first detection head is obtained after training. The trained first detection head can perform object detection. As Figure 6 shown, the object detection head is obtained by training the first detection head. Compared with the first detection head, the structure of the object detection head is the structure obtained by equivalently merging the convolutional modules of the first detection head. It can be understood that a 1x1 convolution is equivalent to a special 3x3 convolution (with many 0s in the convolutional kernel), and an identity mapping is a special 1x1 convolution (with an identity matrix as the convolutional kernel), so it can also be regarded as a special 3x3 convolution. Therefore, the 1x1 convolutions in the first detection head are equivalently converted into 3x3 convolutions, and mathematical merging operations are used to merge some feature space calculations. This can improve the capacity and effective feature space of the object detection head. When using the object detection head for object detection, forward calculation operations can be reduced, and the computational efficiency of object detection can be improved. It should be noted that Figure 6 what is shown is only a schematic diagram of the structure of the object detection head. In the embodiments of the present application, the object detection head is not limited to Figure 6 the structure shown. The present application does not make specific limitations on this.

[0098] The embodiments of the present application provide an application method of a low-quality image object detection model (low-quality image object detection method). As Figure 7 shown, the method includes:

[0099] S701: Obtain a low-quality image to be detected for the area to be detected.

[0100] In this step, the area to be detected is the area where object detection is to be performed. The area to be detected may or may not contain the target object. The low-quality image to be detected is an image with poor quality (low clarity). For example, in an offshore operation scenario, the low-quality image to be detected can be a low-quality image obtained by taking a picture of the area to be detected due to common high-humidity conditions such as fog and rain at sea, or due to problems such as the rapid movement of boats, bad sea conditions, complex water surface reflections, clutter environments, and strong interference from dynamic backgrounds.

[0101] S702: Input the low-quality image to be detected into the target detection model to obtain the target detection result for the area to be detected. Among them, the target detection result is obtained by the first encoder in the target detection model extracting features from the low-quality image to be detected, the target convolution network in the target detection model migrating the extracted features, and the target detection head in the target detection model performing target detection based on the migrated features. The target detection result is used to indicate whether there is a target object in the area to be detected.

[0102] In this step, input the low-quality image to be detected into the target detection model as shown in Figure 8 to obtain the target detection result for the area to be detected. The target detection result can indicate whether there is a target object in the area to be detected and the location information of the target object in the case where there is a target object. In a scenario such as an offshore operation, the target object can be an obstacle such as a reef. By detecting whether there is a reef in the area to be detected and the location information of the reef, the boat in the offshore operation can avoid the obstacle in time or make other decisions (such as returning etc.) for the obstacle. Specifically, the target detection result is obtained by the first encoder in the target detection model extracting features from the low-quality image to be detected, the target convolution network in the target detection model migrating the extracted features, and the target detection head in the target detection model performing target detection based on the migrated features. For the specific obtaining process of the first encoder, the target convolution network and the target detection head, please refer to the relevant descriptions above and will not be elaborated.

[0103] The solutions shown in steps S701 - S702 can be regarded as the application solutions of the low-quality image target detection model. The target detection model obtained through the solutions shown in the foregoing steps S101 - S106 can perform target detection for the area to be detected. The solution of the present application is a lightweight solution, which can not only improve the accuracy of target detection but also improve the efficiency of target detection.

[0104] An embodiment of the present application provides a training device for a low-quality image target detection model, as shown in Figure 9 The device includes:

[0105] The first acquisition unit 901 is used to acquire low-quality sample images for the target area;

[0106] The second acquisition unit 902 is used to input the low-quality sample images into the first encoder to obtain the initial features of the low-quality sample images;

[0107] A third acquisition unit 903, configured to input the initial features into a transfer convolutional network, and train the transfer convolutional network based on first target features to obtain a target convolutional network; wherein, the first target features are obtained based on the initial features of the low-quality sample images; the feature clarity of the first target features is stronger than that of the initial features;

[0108] A fourth acquisition unit 904, configured to obtain second target features for the low-quality sample images based on the target convolutional network;

[0109] A fifth acquisition unit 905, configured to input the second target features into a first detection head, and train the first detection head to obtain a target detection head; the first detection head is obtained based on the first target features;

[0110] A sixth acquisition unit 906, configured to obtain a target detection model based on the first encoder, the target convolutional network, and the target detection head; the target detection model is used for performing target detection on low-quality images to be detected.

[0111] In an optional solution, the third acquisition unit 903 is configured to input the initial features into a transfer convolutional network, and use the first target features as the feature constraint output by the transfer convolutional network to train the transfer convolutional network to obtain a target convolutional network.

[0112] In an optional solution, the third acquisition unit 903 is configured to input the initial features of the low-quality sample images into a pre-trained image type prediction model to obtain the low-quality type of the low-quality sample images; and obtain first target features for the low-quality sample images based on the low-quality type of the low-quality sample images.

[0113] In an optional solution, the third acquisition unit 903 is configured to input the low-quality type of the low-quality sample images into a pre-trained content mapping network to obtain mapping coding information for the low-quality type; input the mapping coding information and the initial features of the low-quality sample images into a first decoder to obtain a restored sample image; the clarity of the restored sample image is stronger than that of the low-quality sample image; and obtain first target features for the low-quality sample images based on the restored sample image.

[0114] In an optional solution, the third acquisition unit 903 is configured to input the restored sample image and the low-quality sample image into a second encoder to obtain first target features;

[0115] The fifth acquisition unit 905 is configured to input the first target features into a second detection head to be trained, and train the second detection head to be trained to obtain a first detection head.

[0116] An embodiment of the present application provides a low-quality image target detection device, such as Figure 10 shown, the device includes:

[0117] A seventh acquisition unit 1001, configured to acquire a low-quality image to be detected for a region to be detected;

[0118] A detection unit 1002, configured to input the low-quality image to be detected into a target detection model to obtain a target detection result for the region to be detected; wherein, the target detection result is obtained by the first encoder in the target detection model extracting features from the low-quality image to be detected, the target convolutional network in the target detection model migrating the extracted features, and the target detection head in the target detection model performing target detection based on the migrated features; the target detection result is used to characterize whether there is a target object in the region to be detected.

[0119] It should be noted that for the training device of the low-quality image target detection model and the low-quality image target detection device in the embodiments of the present application, since the principles of solving problems by the training device of the low-quality image target detection model and the low-quality image target detection device are similar to those of the foregoing low-quality image target detection model training method and low-quality image target detection method, the implementation processes, implementation principles, and beneficial effects of the training device of the low-quality image target detection model and the low-quality image target detection device can refer to the descriptions of the implementation processes, implementation principles, and beneficial effects of the foregoing methods, and repeated parts will not be elaborated.

[0120] According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.

[0121] Figure 11 FIG. shows a schematic block diagram of an exemplary electronic device 1100 that can be used to implement the embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present application described herein and / or claimed.

[0122] Such as Figure 11As shown, the electronic device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the electronic device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0123] Multiple components in the electronic device 1100 are connected to the I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a magnetic disk, an optical disc, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows the electronic device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0124] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1101 executes the various methods and processes described above, such as the training and detection methods of a low-quality image target detection model. For example, in some embodiments, the training and detection methods of a low-quality image target detection model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the training and detection methods of the low-quality image target detection model described above can be executed. Alternatively, in other embodiments, the computing unit 1101 can be configured to execute the training and detection methods of the low-quality image target detection model in any other appropriate manner (e.g., by means of firmware).

[0125] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0126] The program code for implementing the methods of this application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0127] In the context of this application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0128] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0129] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0130] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server combined with a blockchain.

[0131] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved, and no limitation is made herein.

[0132] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of this application, "a plurality" means two or more unless otherwise specifically defined.

[0133] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights.

Claims

1. A method for training a low-quality image object detection model, characterized in that: The method comprises: Obtain low-quality sample images of the target area; Inputting the low-quality sample image into a first encoder to obtain initial features of the low-quality sample image; Inputting the initial features into a migration convolutional network, and training the migration convolutional network based on a first target feature to obtain a target convolutional network; wherein the first target feature is obtained based on the initial features of the low-quality sample image; and the feature clarity of the first target feature is stronger than that of the initial feature; Based on the target convolutional network, a second target feature for the low-quality sample image is obtained; Inputting the second target feature into the first detection head, training the first detection head, and obtaining a target detection head; the first detection head is obtained based on the first target feature; Based on the first encoder, the target convolutional network and the target detection head, a target detection model is obtained; the target detection model is used to perform target detection on the low-quality image to be detected; The first target feature is obtained based on the initial feature of the low-quality sample image, including: Input the initial features of the low-quality sample image into a pre-trained image type prediction model to obtain the low-quality type of the low-quality sample image; Inputting the low-quality type of the low-quality sample image into a pre-trained content mapping network to obtain mapping encoding information for the low-quality type; Inputting the mapping coding information and the initial features of the low-quality sample image into a first decoder to obtain a restored sample image; the definition of the restored sample image is better than the definition of the low-quality sample image; Based on the restored sample image, a first target feature for the low-quality sample image is obtained.

2. The method according to claim 1, characterized in that The step of inputting the initial features into a migration convolutional network and training the migration convolutional network based on the first target features to obtain a target convolutional network comprises: The initial features are input into a migration convolutional network, the first target features are used as feature constraints output by the migration convolutional network, and the migration convolutional network is trained to obtain a target convolutional network.

3. The method according to claim 1, characterized in that The obtaining, based on the restored sample image, a first target feature for the low-quality sample image comprises: Inputting the restored sample image and the low-quality sample image into a second encoder to obtain a first target feature; The first detection head is obtained based on the first target feature, including: The first target feature is input to a second detection head to be trained, and the second detection head to be trained is trained to obtain a first detection head.

4. A low-quality image object detection method, characterized in that: The method comprises: Obtain a low-quality image of the area to be detected; The low-quality image to be detected is input into the target detection model to obtain a target detection result for the area to be detected; wherein the target detection result is obtained by extracting features of the low-quality image to be detected by the first encoder in the target detection model, migrating the extracted features by the target convolution network in the target detection model, and performing target detection based on the migrated features by the target detection head in the target detection model; the target detection result is used to characterize whether there is a target object in the area to be detected; wherein the target convolution network is obtained by training the migration convolution network based on the first target feature; the first target feature is obtained based on the initial feature of the low-quality sample image; The first target feature is obtained based on the initial feature of the low-quality sample image, including: Input the initial features of the low-quality sample image into a pre-trained image type prediction model to obtain the low-quality type of the low-quality sample image; Inputting the low-quality type of the low-quality sample image into a pre-trained content mapping network to obtain mapping encoding information for the low-quality type; Inputting the mapping coding information and the initial features of the low-quality sample image into a first decoder to obtain a restored sample image; the definition of the restored sample image is better than the definition of the low-quality sample image; Based on the restored sample image, a first target feature for the low-quality sample image is obtained.

5. A training device for a low-quality image object detection model, characterized in that: The device comprises: A first acquisition unit, used to acquire a low-quality sample image of a target area; a second acquisition unit, configured to input the low-quality sample image into a first encoder to obtain initial features of the low-quality sample image; A third acquisition unit is used to input the initial feature into the migration convolution network, and train the migration convolution network based on the first target feature to obtain a target convolution network; wherein the first target feature is obtained based on the initial feature of the low-quality sample image; and the feature clarity of the first target feature is stronger than that of the initial feature; a fourth acquisition unit, configured to obtain a second target feature for the low-quality sample image based on the target convolutional network; A fifth acquisition unit is used to input the second target feature into a first detection head, train the first detection head, and obtain a target detection head; the first detection head is obtained based on the first target feature; A sixth acquisition unit, configured to obtain a target detection model based on the first encoder, the target convolutional network, and the target detection head; the target detection model is used to perform target detection on the low-quality image to be detected; The third acquisition unit is used to input the initial features of the low-quality sample image into a pre-trained image type prediction model to obtain the low-quality type of the low-quality sample image; input the low-quality type of the low-quality sample image into a pre-trained content mapping network to obtain mapping coding information for the low-quality type; input the mapping coding information and the initial features of the low-quality sample image into the first decoder to obtain a restored sample image; the clarity of the restored sample image is stronger than that of the low-quality sample image; based on the restored sample image, a first target feature for the low-quality sample image is obtained.

6. A low-quality image object detection device, characterized in that: The device comprises: A seventh acquisition unit, used for acquiring a low-quality image to be detected for the area to be detected; A detection unit, used for inputting the low-quality image to be detected into a target detection model to obtain a target detection result for the area to be detected; wherein the target detection result is obtained by extracting features of the low-quality image to be detected by a first encoder in the target detection model, migrating the extracted features by a target convolutional network in the target detection model, and performing target detection based on the migrated features by a target detection head in the target detection model; the target detection result is used to characterize whether there is a target object in the area to be detected; wherein the target convolutional network is obtained by training a migration convolutional network based on a first target feature; and the first target feature is obtained based on an initial feature of the low-quality sample image; The detection unit is used to input the initial features of the low-quality sample image into a pre-trained image type prediction model to obtain the low-quality type of the low-quality sample image; input the low-quality type of the low-quality sample image into a pre-trained content mapping network to obtain mapping coding information for the low-quality type; input the mapping coding information and the initial features of the low-quality sample image into a first decoder to obtain a restored sample image; the clarity of the restored sample image is stronger than that of the low-quality sample image; based on the restored sample image, a first target feature for the low-quality sample image is obtained.

7. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 3 or 4.

8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to make a computer execute the method according to any one of claims 1-3 or 4.

Citation Information

Patent Citations

  • Cross-band infrared target detection method and device, computer equipment and storage medium

    CN115861630A