Image generation method, model training method, detection method, device and system

By extracting and migrating the manhole cover missing manhole features from the images with manhole cover manhole covers and generating sample images, the problem of insufficient manhole covers missing manhole covers in the prior art is solved, and the detection accuracy of the target detection model is improved.

CN114882206BActive Publication Date: 2025-07-22SHANGHAI SENSETIME LINGANG INTELLIGENT TECH CO LTD
View PDF -1 Cites 0 Cited by

Patent Information

Application Number
CN202210707305.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-21
Publication Date
2025-07-22
Estimated Expiration
2042-06-21

AI Technical Summary

Technical Problem

The existing technology is difficult to obtain the missing manhole images of manhole covers in different scenarios, resulting in the incomplete training data of the target detection model and low detection accuracy.

Method used

By intercepting the first image area from the image with the manhole cover manhole, the features of the manhole cover missing manhole are migrated to the first image area, a second image area including the manhole cover missing manhole is generated, and replaced it into the target image, forming a sample image for training the object detection model.

Benefits of technology

A large number of manhole sample images covering manhole covers different scenarios were generated, which improved the sample data diversity and detection accuracy of the target detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114882206B_ABST
    Figure CN114882206B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an image generation method, a model training method, a detection method, a device, and a system. The method can extract features of manholes without manhole covers from a small number of acquired images of manholes without manhole covers, then intercept a first image region including the manhole with a manhole cover from an image of a manhole with a manhole cover, and then transfer the extracted features to the manhole with a manhole cover in the first image region to obtain a second image region including a manhole without a manhole cover. Then, a target image region where a manhole may exist can be determined from a target image, and the target image region can be replaced with the second image region, so that images of manholes without manhole covers in different scenarios can be obtained as sample images. In this way, a large number of sample images of manholes without manhole covers in different scenarios can be obtained for training a target detection model, making the sample data more diverse and the prediction results of the trained target detection model more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to an image generation method, a model training method, a detection method, an apparatus, and a system. Background Art

[0002] Since there are often a large number of manholes on the ground, and the manhole covers of these manholes often have the problem of being missing. Once the manhole cover is missing, it will pose a great safety hazard to pedestrians and vehicles, causing an adverse social impact. Therefore, it is necessary to timely detect the manholes with missing manhole covers on the ground and take corresponding protective measures. Currently, images of roads can be collected, and a pre-trained object detection model is used to detect whether the images include manholes with missing manhole covers. However, since the images of manholes with missing manhole covers are few, it is difficult to cover different scenarios, resulting in the training data of the object detection model being incomplete, and finally the accuracy of the trained object detection model is relatively low. Summary of the Invention

[0003] The present disclosure provides an image generation method, a model training method, a detection method, an apparatus, and a system.

[0004] According to a first aspect of an embodiment of the present disclosure, a method for generating a sample image is provided, where the sample image is used to train an object detection model for detecting manholes with missing manhole covers in an image, and the method includes:

[0005] Obtain a target image;

[0006] Determine a target image region from the target image, and replace the target image region with a second image region including a manhole with a missing manhole cover to obtain a sample image including a manhole with a missing manhole cover;

[0007] Wherein, the second image region is obtained based on the following method: intercept a first image region including the manhole with a manhole cover from an image of a manhole with a manhole cover, and transfer the features of the manhole with a missing manhole cover to the manhole with a manhole cover in the first image region to obtain the second image region, where the features of the manhole with a missing manhole cover are obtained by extracting features from an image of the manhole with a missing manhole cover.

[0008] In some embodiments, the transferring the features of the manhole with a missing manhole cover to the manhole with a manhole cover in the first image region to obtain the second image region includes:

[0009] Input the first image region into a pre-trained style transfer model. Through the style transfer model, transfer the features of the manhole without a manhole cover to the manhole with a manhole cover in the first image region, obtaining a second image region including a manhole without a manhole cover, where the style transfer model is trained by using a first image including a manhole with a manhole cover and a second image including a manhole without a manhole cover.

[0010] In some embodiments, the style transfer model includes a generative adversarial network. The style transfer model is trained based on a first image including a manhole with a manhole cover and a second image including a manhole without a manhole cover, including:

[0011] Use the generator of the generative adversarial network to generate an image including a manhole without a manhole cover based on the first image;

[0012] Determine a first loss based on the discrimination result of the discriminator of the generative adversarial network for the image generated by the generator and the second image;

[0013] Determine a second loss based on the similarity between a first image patch in the image generated by the generator and a second image patch in the first image, and the similarity between the first image patch in the image generated by the generator and a third image patch in the first image; where the first image patch and the second image patch are located at the same pixel position, and the first image patch and the third image patch are located at different pixel positions;

[0014] Determine a target loss based on the first loss and the second loss, and use the target loss to train the generative adversarial network.

[0015] In some embodiments, determining a target image region from a target image includes:

[0016] If the target image is the image of the manhole with a manhole cover, use the first image region as the target image region; or

[0017] If the target image is the image of the manhole with a manhole cover or an image other than the image of the manhole with a manhole cover, perform semantic segmentation processing on the target image, and select the target image region from the target image based on the result of the semantic segmentation.

[0018] In some embodiments, using the second image region including a manhole without a manhole cover to replace the target image region to obtain the sample image includes:

[0019] Determine a mask image based on the second image region including a manhole without a manhole cover;

[0020] Extract a foreground image from the second image region using the mask image, and extract a background image from the target image region using the mask image;

[0021] Fuse the foreground image and the background image, and replace the target image region with the fused image to obtain the sample image.

[0022] In some embodiments, determining a mask image based on the second image region includes:

[0023] Perform a scaling process on the second image region so that the size of the second image region is the same as the size of the first image region;

[0024] Perform gray-scale processing on the scaled second image region, then perform denoising processing, and then perform binarization processing on the denoised gray-scale image;

[0025] Perform denoising processing on the binarized image obtained by the binarization processing to obtain the mask image.

[0026] According to a second aspect of the embodiments of the present disclosure, there is provided a method for training a target detection model, the method including:

[0027] Generate a sample image using the sample image generation method mentioned in the first aspect above;

[0028] Train a preset initial model using the sample image to obtain the target detection model.

[0029] According to a third aspect of the embodiments of the present disclosure, there is provided a method for detecting a manhole with a missing manhole cover on a road, the method including:

[0030] Obtain an image of the road;

[0031] Input the image into a pre-trained target detection model, and detect the manhole with a missing manhole cover in the image through the target detection model, where the target detection model is trained by a sample image, and the sample image includes an image generated by the sample image generation method mentioned in the first aspect above.

[0032] According to a fourth aspect of the embodiments of the present disclosure, there is provided a road detection system, the road detection system including an image acquisition device and a server, the image acquisition device is located on the side of the road or the image acquisition device is mounted in a road inspection device,

[0033] The image acquisition device is used to acquire an image of the road and send it to the server;

[0034] The server is configured to input the image into a pre-trained object detection model, and detect manholes without manhole covers in the image through the object detection model. The object detection model is trained by using sample images, and the sample images include images generated by the sample image generation method mentioned in the first aspect above.

[0035] According to a fifth aspect of the embodiments of the present disclosure, there is provided a sample image generation device. The sample image is used to train an object detection model for detecting manholes without manhole covers in an image. The device includes:

[0036] An acquisition module, configured to acquire a target image;

[0037] A replacement module, configured to determine a target image area from the target image, and replace the target image area with a second image area including a manhole without a manhole cover to obtain a sample image including a manhole without a manhole cover. The second image area is obtained based on the following method: intercept a first image area including the manhole with a manhole cover from an image of a manhole with a manhole cover, and transfer the features of the manhole without a manhole cover to the manhole with a manhole cover in the first image area to obtain the second image area. The features of the manhole without a manhole cover are obtained by performing feature extraction on an image of the manhole without a manhole cover.

[0038] According to a sixth aspect of the embodiments of the present disclosure, there is provided an electronic device, which includes a processor, a memory, and computer instructions stored in the memory and executable by the processor. When the processor executes the computer instructions, the methods mentioned in the first aspect, the second aspect, and the third aspect above can be implemented.

[0039] According to a seventh aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer instructions are stored. When the computer instructions are executed, the methods mentioned in the first aspect, the second aspect, and the third aspect above are implemented.

[0040] In the embodiments of the present disclosure, considering that although it is difficult to collect images of manholes without manhole covers from actual life scenarios, a large number of images of manholes with manhole covers in different scenarios can be obtained. Therefore, the features of manholes without manhole covers can be extracted from the small number of obtained images of manholes without manhole covers. Then, a first image region including the manhole with a manhole cover can be intercepted from the image of the manhole with a manhole cover, and the extracted features can be transferred to the manhole with a manhole cover in the first image region to obtain a second image region including a manhole without a manhole cover. Then, a target image including a road scene can be obtained, a target image region where a manhole may exist can be determined from the target image, and the target image region can be replaced with the second image region, so that images of manholes without manhole covers in different scenarios can be obtained as sample images. In this way, a large number of sample images of manholes without manhole covers in different scenarios can be obtained for training the target detection model, making the sample data more diverse and the prediction results of the trained target detection model more accurate.

[0041] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0043] Figure 1 It is a schematic diagram of a method for generating sample images according to an embodiment of the present disclosure.

[0044] Figure 2 It is a flowchart of a method for generating sample images according to an embodiment of the present disclosure.

[0045] Figure 3 It is a schematic diagram of the method for generating sample images according to an embodiment of the present disclosure.

[0046] Figure 4 It is a schematic diagram of the training of a style transfer model according to an embodiment of the present disclosure.

[0047] Figure 5 It is a schematic diagram of replacing a target image region with a second image region according to an embodiment of the present disclosure.

[0048] Figure 6 It is a schematic diagram of a road detection system according to an embodiment of the present disclosure.

[0049] Figure 7 It is a schematic diagram of a road detection system according to an embodiment of the present disclosure warning a vehicle.

[0050] Figure 8Schematic diagram of a road detection system according to an embodiment of the present disclosure for prompting vehicles and pedestrians.

[0051] Figure 9 Schematic diagram of the logical structure of a sample image generation device according to an embodiment of the present disclosure.

[0052] Figure 10 Schematic diagram of the logical structure of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners

[0053] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0054] The terms used in the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. The singular forms "a", "the" and "said" used in the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality.

[0055] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0056] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present disclosure and to make the above-mentioned objects, features, and advantages of the embodiments of the present disclosure more obvious and understandable, the technical solutions in the embodiments of the present disclosure will be further described in detail below with reference to the drawings.

[0057] With the modernization of cities, various pipelines need to be laid underground in cities, such as gas or natural gas pipelines, sewers, power grids, water supply pipelines, etc. The entrances of these pipelines on the ground are generally called manholes, and manholes are usually covered by manhole covers. Since there are often a large number of manholes distributed on the ground, and the problem of missing manhole covers often occurs in these manholes. Once the manhole cover is missing, it will pose a great safety hazard to pedestrians and vehicles, causing adverse social impacts. Therefore, it is necessary to detect in time the manholes with missing manhole covers on the ground and take corresponding protective measures.

[0058] Currently, when automatically detecting manholes with missing manhole covers, it is usually to collect images of manholes in the road, and then detect the images to determine whether the images include manholes with missing manhole covers. Detecting the images usually includes two methods. One method is detection based on traditional image processing technology. This technology collects images of manholes, extracts morphological, color, geometric and other characteristics of manholes with missing manhole covers in the images, and identifies manholes with missing manhole covers based on these characteristics. This method requires a large amount of calculation and a large number of hyperparameter adjustments, and has poor generalization, making it difficult to adapt to different scenarios. Another method is the object detection technology based on the neural network model of deep learning. This method obtains a large number of images of manholes with missing manhole covers, uses these images to train the neural network model, and obtains an object detection model that can detect manholes with missing manhole covers from the images. This method has stronger generalization, can adapt to various scenarios, has fewer hyperparameters, and is also easier to deploy on lightweight computing devices.

[0059] However, the second method requires using a large number of images of manholes with missing manhole covers in different scenarios to train the neural network model, and the obtained object detection model can achieve better effects and detection accuracy. However, the scenarios of missing manhole covers of manholes are often few, and the images of manholes with missing manhole covers (especially the images of manholes with missing manhole covers in different scenarios) are usually difficult to obtain, resulting in insufficient number of training samples and difficulty in covering different scenarios, and lack of diversity, resulting in less than ideal detection effects of the finally trained object detection model.

[0060] Based on this, embodiments of the present disclosure provide a method for generating a sample image. Considering that although it is difficult to collect images of manholes without manhole covers from actual life scenarios, a large number of images of manholes with manhole covers in different scenarios can be obtained. Therefore, the features of manholes without manhole covers can be extracted from the small number of images of manholes without manhole covers that have been obtained. Then, a first image region including the manhole with a manhole cover can be intercepted from the image of the manhole with a manhole cover, and the extracted features can be migrated to the manhole with a manhole cover in the first image region to obtain a second image region including a manhole without a manhole cover. Then, a target image including a road scene can be obtained, a target image region where a manhole may exist can be determined from the target image, and the second image region can be used to replace the target image region, so that images of manholes without manhole covers in different scenarios can be obtained as sample images.

[0061] In embodiments of the present disclosure, by migrating the extracted features of the manhole without a manhole cover to the manhole with a manhole cover in the first image region, a second image region including a manhole without a manhole cover is obtained, and then the second image region is fused into the target image region that may contain a manhole in the target image. In this way, a large number of sample images of manholes without manhole covers in different scenarios can be obtained for training a target detection model, making the sample data more diverse and the prediction results of the trained target detection model more accurate.

[0062] The method for generating a sample image according to embodiments of the present disclosure can be executed by various electronic devices, such as mobile phones, computers, cloud servers, etc. For example, in some scenarios, the method can be implemented through an App installed on the electronic device. For example, assuming the target image is an image of the manhole with a manhole cover, the user can open the App and import the image of the manhole with a manhole cover, and then the sample image can be automatically generated. Or the user can import both the image of the manhole with a manhole cover and the image of the manhole without a manhole cover at the same time, and then the sample image can be automatically generated. Or, assuming the target image is not an image of the manhole with a manhole cover, the user can also import the image of the manhole with a manhole cover and the target image at the same time, and then the sample image can be automatically generated.

[0063] As Figure 1 shown, it is a schematic diagram of the method for generating a sample image according to an embodiment of the present disclosure. As Figure 2 shown, it is a flowchart of the method for generating a sample image according to embodiments of the present disclosure. The method may include the following steps:

[0064] S202. Obtain a target image;

[0065] In step S202, a target image can be obtained. The target image can be various images including scenes where manholes may exist. For example, it can be an image including a road scene. The target image may or may not include a manhole. In some scenarios, a user interaction interface can be provided, and then the target image imported by the user through the "import control" in the interaction interface can be obtained.

[0066] S204. Determine a target image area from the target image, and replace the target image area with a second image area including a manhole without a manhole cover to obtain a sample image including a manhole without a manhole cover. The second image area is obtained based on the following method: intercept a first image area including the manhole with a manhole cover from an image of a manhole with a manhole cover, and transfer the features of the manhole without a manhole cover to the manhole with a manhole cover in the first image area to obtain the second image area, where the features of the manhole without a manhole cover are obtained by extracting features from an image of a manhole without a manhole cover.

[0067] In step S204, a second image area including a manhole without a manhole cover can be generated in advance. For example, an image of a manhole with a manhole cover can be obtained. For example, a user interaction interface can be provided, and then the image of a manhole with a manhole cover imported by the user through the "import control" in the interaction interface can be obtained. The image of a manhole with a manhole cover can be an image obtained by collecting images of manholes in a road. For example, cameras are usually installed on roads, so images of manholes in the road can be collected through these cameras. Or, in order to obtain images of manholes with manhole covers in different scenarios, a road inspection device equipped with a camera can be used to inspect the road, and images of manholes with manhole covers in different sections can be collected through the camera during the inspection.

[0068] In addition, some images of manholes without a manhole cover can also be obtained. For example, they can be obtained from the network, or images of manholes without a manhole cover can be collected from road images. Since it is difficult to obtain images of manholes without a manhole cover and the quantity is often small, features of the manholes without a manhole cover can be extracted from the images of manholes without a manhole cover. After obtaining the image of a manhole with a manhole cover, a first image area including the manhole with a manhole cover can be intercepted from the image of the manhole with a manhole cover, and then the features of the manhole without a manhole cover extracted from the image of the manhole without a manhole cover can be transferred to the manhole with a manhole cover in the first image area to obtain a second image area including a manhole without a manhole cover.

[0069] Among them, the first image region may only include manholes with manhole covers, or most of the region are manholes with manhole covers. By intercepting the first image region from the image of manholes with manhole covers, other scenes in the image can be excluded, so that there is less interference in the first image region. Furthermore, when the features of manholes without manhole covers are transferred to the first image region through style transfer, a second image region including manholes without manhole covers with more natural and realistic effects can be obtained.

[0070] In order for the finally generated sample images to be more consistent with the actual scene, the target image region can be determined from the target image first. Among them, the target image region may be a region that may include manholes. For example, it may be a road region. Then, the second image region including manholes without manhole covers can be fused into the target image region of the target image to obtain a sample image corresponding to the real scene.

[0071] By pasting the second image region including manholes without manhole covers onto the target image containing various scenes, a large number of sample images including manholes without manhole covers covering various scenes can be generated, making the sample images more abundant.

[0072] In some scenes, the target image may be the image of the manhole with a manhole cover, and the target image region may be the first image region. As Figure 3 shown, that is, replacing the manhole with a manhole cover in the image of the manhole with a manhole cover with a manhole without a manhole cover, and pasting the manhole without a manhole cover obtained by style transfer back to the original position, a more natural sample image can be obtained.

[0073] In some scenes, in order to make the obtained sample images more diverse, the target image region may also be other regions in the image of the manhole with a manhole cover except the first image region. For example, it may be other road regions in the image, or regions that may include manhole covers.

[0074] In some scenes, other images including road scenes outside the image of the manhole with a manhole cover can also be obtained as the target image, and then the target image region where manholes may exist can be selected from the target image.

[0075] After determining the target image region, the second image region can be used to replace the target image region to obtain a sample image including manholes without manhole covers. In this way, a large number of sample images corresponding to different scenes can be obtained, greatly enriching the sample images.

[0076] Among them, in some scenarios, feature extraction of the image of the manhole without a manhole cover can be completed in advance. For example, the feature extraction of the obtained image of the manhole without a manhole cover can be performed in advance to obtain the features of the manhole without a manhole cover. Subsequently, for each frame of the image of the manhole with a manhole cover obtained and the first image area is intercepted, the features extracted in advance can be transferred to the manhole with a manhole cover in the first image area.

[0077] In some scenarios, feature extraction of the image of the manhole without a manhole cover can also be completed in real time. For example, the user can also input a set of images at the same time. The images include the image of the manhole with a manhole cover and the image of the manhole without a manhole cover. Then, the features of the manhole without a manhole cover can be extracted from the image of the manhole without a manhole cover, and the extracted features can be transferred to the manhole with a manhole cover in the first image area extracted from the image of the manhole with a manhole cover.

[0078] Image processing technology can be used to extract the features of the obtained image of the manhole without a manhole cover to obtain the features of the manhole without a manhole cover. For example, image processing technology can extract the morphological, color, geometric and other features of the manhole without a manhole cover in the image. Or a neural network can be pre-trained in advance, and the neural network is used to extract the features of the image of the manhole without a manhole cover to obtain the features of the manhole without a manhole cover. For example, the neural network can learn and extract the features of the manhole without a manhole cover. Similarly, transferring the extracted features to the manhole with a manhole cover in the first image area can be achieved through image processing technology or through a pre-trained neural network. For example, the neural network model can be a generative adversarial network, an autoencoder, etc. Among them, feature extraction and feature transfer can be implemented by the same neural network model or by different neural network models.

[0079] In some embodiments, transferring the features of the manhole without a manhole cover to the manhole with a manhole cover in the first image area to obtain the second image area including the manhole without a manhole cover can be achieved through a pre-trained style transfer model. For example, the first image of the manhole with a manhole cover and the second image including the manhole without a manhole cover can be obtained, and then the style transfer model can be trained using the above two images. For example, taking the style transfer model as a generative adversarial network, the generator of the generative adversarial network can be used to generate the image of the manhole without a manhole cover based on the first image, and then the discriminator of the generative adversarial network can be used to identify the generated image and the second image, and the generative adversarial network can be trained based on the discrimination result. Among them, the first image can be the image area of the manhole with a manhole cover intercepted from the image of the manhole with a manhole cover, and the second image can be the image area of the manhole without a manhole cover intercepted from the image of the manhole without a manhole cover. After the style transfer model is trained, the first image area can be input into the style transfer model (for example, into the generator of the generative adversarial network), and the second image area can be output through the style transfer model.

[0080] In some embodiments, as Figure 4 shown, the style transfer model may be a generative adversarial network. When training the generative adversarial network using a first image of a manhole with a manhole cover and a second image of a manhole without a manhole cover, the generator of the generative adversarial network may be used to generate an image including a manhole without a manhole cover based on the first image, and then a first loss may be determined based on the discrimination result of the discriminator for the image generated by the generator and the second image, and a second loss may be determined based on the similarity between a first image patch in the image generated by the generator and a second image patch in the first image, and the similarity between the first image patch in the image generated by the generator and a third image patch in the first image; wherein, the first image patch and the second image patch are located at the same pixel position, and the first image patch and the third image patch are located at different pixel positions. The third image patch may be an image patch in the first image that is located at a different pixel position from the first image patch, or may be multiple image patches. For example, in some scenarios, the third image patch may be 256 image patches in the first image that are located at different pixel positions from the first image patch.

[0081] When a general generative adversarial network is trained, the loss is only determined based on the discrimination result of the discriminator for the image generated by the generator and the real image, and the generative adversarial network is trained based on this loss. The effect of the generative adversarial network trained in this way is not yet ideal, and the effect of the sample image generated using the trained generative adversarial network still needs to be improved. To improve the accuracy of the trained generative adversarial network, in the embodiments of the present disclosure, when training the generative adversarial network, the idea of contrastive learning is introduced, that is, considering that the similarity between a certain region in the image generated by the generator and the corresponding region in the input image must be higher than the similarity between the region and other regions in the input image except the corresponding region, so that the image generated by the generator is more accurate. Based on this idea, when training the generative adversarial network, in addition to determining a first loss based on the discrimination result of the discriminator for the image generated by the generator and the second image, a second loss may further be determined based on the similarity between a first image patch in the image generated by the generator and a second image patch in the first image that is located at the same pixel position as the first image patch, and the similarity between the first image patch in the image generated by the generator and a third image patch in the first image that is located at a different pixel position from the first image patch, and then a target loss may be determined based on the first loss and the second loss, and the generative adversarial network may be trained using the target loss. The effect of the sample image generated by the generative adversarial network trained in this way will be greatly improved when performing feature transfer.

[0082] After training the generative adversarial network, the first image region may be input into the generator to obtain a second image region.

[0083] In some embodiments, the target image may be an image of a manhole with a manhole cover. The first image region in the image of the manhole with a manhole cover can be directly used as the target image region. That is, the manhole with a missing manhole cover obtained by style transfer is used to replace the original manhole with a manhole cover in the image. In this way, a more natural image can be obtained.

[0084] In some embodiments, in order to obtain more sample images of different scenarios, the target image region may also be other regions in the image of the manhole with a manhole cover except the first image region. For example, the target image region may also be other regions in the image of the manhole with a manhole cover where manholes may exist. To accurately identify these regions so as to fuse the manhole with a missing manhole cover into the appropriate position in the image, semantic segmentation processing can be performed on the image of the manhole with a manhole cover, and the target image region is selected from the manhole with a manhole cover based on the result of semantic segmentation. For example, through semantic segmentation, the road surface region, sky region, building region, etc. in the image can be determined, and then a region can be selected from the road surface region as the target image region.

[0085] Of course, in some embodiments, in order to obtain more abundant sample images to comprehensively cover different scenarios. The target image may also be other images except the image of the manhole with a manhole cover. For example, it may also be an image collected without a manhole, or other images including manholes. Then semantic segmentation can be performed on the target image, and a region where a manhole may exist is selected from the target image based on the result of semantic segmentation as the target image region.

[0086] In some embodiments, when using the second image region to replace the target image region to obtain a sample image, the second image region can be directly used to replace the target image region. However, the image obtained in this way will not be very natural. For example, there will be obvious fusion boundaries, resulting in a poor effect of the sample image.

[0087] In some embodiments, such as Figure 5As shown, in order to make the fusion of the second image region and the target image more natural and obtain a sample image with better and more realistic effects, a mask image can be determined based on the second image region first. Then, the foreground image can be extracted from the second image region using the mask image, and the background image can be extracted from the target image region using the mask image. The extracted foreground image and background image are fused, and then the image obtained by fusion is used to replace the target image region in the target image, resulting in a sample image including a manhole with a missing manhole cover. Among them, the mask image is a binary image. By performing an AND operation on the mask image and the second image region, the foreground image can be obtained. By inverting the mask image and then performing an AND operation on it and the target image region, the background image can be obtained. Then the two are fused to obtain a fused image, and the target image region is replaced with the fused image. Through this fusion method, the fusion boundary can be eliminated, and the features of the manhole with a missing manhole cover can be seamlessly fused into the real scene, obtaining a more natural sample image.

[0088] The accurate design of the mask image is undoubtedly the key to affecting the final fusion result. In some implementations, in order to obtain a more accurate mask image, the second image region can be scaled first so that the size of the second image region is the same as the size of the target image region. Since the size of the second image region output by the style transfer model is usually fixed and may be different from the size of the determined target image region. Therefore, the second image region can be scaled first so that the sizes of the two image regions are the same. Then, the scaled second image region can be grayscale processed. Grayscale processing is to convert a color image into a grayscale image to facilitate binarization of the image. After grayscale processing the image, the image can be denoised first to remove some noise in the image. Then, the denoised image can be binarized. For example, the OTSU algorithm can be used to binarize the image. Since there may still be some noise in the binarized image, the binarized image can be denoised again to obtain the mask image. For example, morphological opening operation can be used to further reduce noise in the binarized image, and then morphological closing operation can be used to connect the regions in the binarized image to obtain the final mask image. Through the above series of image processing, an accurate mask image can be obtained, and using the mask image to fuse the second image region and the target image can also obtain a more natural fusion effect.

[0089] Furthermore, the embodiments of the present disclosure also provide a method for training a target detection model, and the method includes the following steps:

[0090] Generate a sample image using the sample image generation method described in the above embodiments;

[0091] Train a preset initial model using sample images to obtain an object detection model.

[0092] Among them, the specific implementation details of generating the sample images can be referred to the description in the above embodiments, and will not be elaborated here.

[0093] Furthermore, the embodiments of the present disclosure also provide a method for detecting manholes with missing manhole covers on roads. This method can be used to detect manholes with missing manhole covers on roads, and the method includes the following steps:

[0094] Obtain an image of the road;

[0095] Input the image into a pre-trained object detection model, and detect the manholes with missing manhole covers in the image through the object detection model. Among them, the object detection model is trained by sample images, and the sample images include images generated by the sample image generation method introduced in the above embodiments.

[0096] For example, an image of the road can be collected by a camera set on the road, or a camera can also be mounted on a road inspection device. During the process of the inspection vehicle inspecting the road, an image of the road is collected by the camera, and then the pre-trained object detection model can be used to detect the collected image to timely detect the manholes with missing manhole covers on the road. Among them, in some scenarios, a detection device can be set in the road inspection device, and the above road detection method is executed by using the detection device. In some scenarios, after the camera collects an image of the road, it can also be sent to the cloud server, and the above road detection method is executed by the cloud server.

[0097] Among them, the object detection model is trained by sample images, and the sample images include images generated by the sample image generation method introduced in the above embodiments.

[0098] Furthermore, the embodiments of the present disclosure also provide a road detection system, as Figure 6 shown. This road detection system includes an image acquisition device and a server. The image acquisition device is located on the road side or in a road inspection device. For example, the image acquisition device can be fixedly set on the road, or the image acquisition device can also be mounted on a road inspection device. The road inspection device can be a movable intelligent device, such as a robot, an inspection vehicle, or an autonomous driving vehicle, etc. During the process of the road inspection device inspecting the road, images of different sections can be collected. After the image acquisition device collects an image of the road, it can be sent to the server. The server is used to input the received image into a pre-trained object detection model, and detect the manholes with missing manhole covers in the image through the object detection model.

[0099] Among them, the object detection model is trained through sample images, and the sample images include images generated by the sample image generation method introduced in the above embodiments.

[0100] In some embodiments, as Figure 7 shown, the server is further configured to, when detecting that a manhole in the image lacks a manhole cover, obtain the position information of the image acquisition device when it acquires the image, and send an alarm message to the vehicles on the road, where the alarm message includes the position information. If the image acquisition device is fixedly installed on the road, the position information of the image acquisition device can be pre-recorded in the server, and the server can determine the position information of the image acquisition device when it acquires the image based on the identification information of the image acquisition device carried in the image, so as to locate the position where the manhole lacking a manhole cover is located. If the image acquisition device is mounted on a road inspection device, the road inspection device can be mounted with a positioning device, such as GPS, and can send the positioning information simultaneously when sending the image to the server, so that the server can determine the position of the image acquisition device when the image is acquired based on the positioning information, and further locate the position where the manhole lacking a manhole cover is located. When the server detects that there is a manhole lacking a manhole cover on the road, it can push an alarm message to the vehicles on the road, reminding the vehicles that there is a problem with the manhole cover lacking at a certain position in a certain section of the road, so that the vehicles can pay attention to safety. At the same time, it can also prompt the road maintenance personnel to handle the abnormal situation in time.

[0101] In some scenarios, as Figure 8 shown, when the server detects that there is a problem with the manhole cover lacking at a certain position in a certain section of the road, it can also control the road inspection device to stay near the manhole lacking the manhole cover, and then use the road inspection device to prompt the surrounding pedestrians or vehicles through voice or visual information until the maintenance personnel come to repair it.

[0102] To further explain the sample image generation method provided by the embodiments of the present disclosure, the following is explained in conjunction with a specific embodiment.

[0103] In order to obtain more abundant images of manholes lacking manhole covers as sample images for training an object detection model to detect manholes lacking manhole covers on the road, this embodiment provides a sample image generation method. Specifically as follows:

[0104] 1. Training of the style transfer model

[0105] The original images of manholes without covers can be obtained from various channels, and a large number of original images including manholes with covers can be obtained from actual usage scenarios (such as cameras installed on roads or cameras equipped in road inspection devices). Use bounding boxes to label the area where the manhole without a cover is located in the original image of the manhole without a cover, and use bounding boxes to label the area where the manhole with a cover is located in the original image of the manhole with a cover.

[0106] Crop all the original images of manholes without covers according to the bounding boxes, extract partial areas of the manholes without covers to obtain local images of the manholes without covers, and use these local images to form dataset A. Similarly, crop the images of manholes with covers according to the bounding boxes, extract the parts of the manholes with covers to obtain local images of the manholes with covers, and use these images to form dataset B. Use datasets A and B to train a generative adversarial network to obtain a style transfer model. The specific training method can refer to the description in the above embodiments and will not be elaborated here.

[0107] 2. Obtain local images of manholes without covers using the style transfer model

[0108] Local images of manholes with covers can be obtained from dataset B and input into the style transfer model to generate corresponding local images of manholes without covers for the scene.

[0109] 3. Image fusion

[0110] (1) Obtain the local image b of the manhole without a cover generated in step 2, and the original image a including the manhole with a cover before cropping corresponding to b.

[0111] (2) Select a roi area in a reasonable area in the original image a. Among them, this reasonable area can be the area where the original manhole with a cover was located in the original image a, or other areas where manholes may exist. For example, the original image a can be first subjected to semantic segmentation, and the roi area can be determined from the original image a based on the results of semantic segmentation.

[0112] (3) Scale b to the same size as the roi area, perform gray processing and then perform regional noise reduction to generate b1;

[0113] (4) Use the OTSU algorithm to binarize b1, and use morphological opening operation to further reduce noise in the feature area, and then use morphological closing operation to connect the feature areas to generate a mask image b_mask;

[0114] (5) Perform an AND operation on the obtained mask image and b to obtain the feature region b_front of the manhole without a manhole cover. At the same time, perform an AND operation on the roi region in the original image a and the inverse of the mask image to obtain the roi background region roi_bg. Merge b_front and roi_bg and replace the roi region in the original image a to obtain the sample image.

[0115] 4. Training of the object detection model

[0116] Use the sample image generated in step 3 and the roi region corresponding to the sample image as labels; train the preset initial model to obtain the object detection model.

[0117] 5. Road detection

[0118] Use the camera set on the road or the camera carried by the road inspection device to collect the image of the road, input the image into the object detection model, and detect whether there is a manhole without a manhole cover in the image through the object detection model.

[0119] It is not difficult to understand that the solutions described in the above embodiments can be combined without conflict, and they are not listed one by one in the embodiments of the present disclosure.

[0120] Correspondingly, the embodiments of the present disclosure also provide a sample image generation device, the sample image is used to train an object detection model, and the object detection model is used to detect the manhole without a manhole cover in the image. As Figure 9 shown, the device includes:

[0121] An acquisition module 91, configured to acquire a target image;

[0122] A replacement module 92, configured to determine a target image region from the target image, and replace the target image region with a second image region including a manhole without a manhole cover to obtain a sample image including a manhole without a manhole cover; wherein, the second image region is obtained based on the following method: intercept a first image region including the manhole with a manhole cover from the image of the manhole with a manhole cover, and transfer the feature of the manhole without a manhole cover to the manhole with a manhole cover in the first image region to obtain the second image region, wherein the feature of the manhole without a manhole cover is obtained by performing feature extraction on the image of the manhole without a manhole cover.

[0123] In some embodiments, the process of transferring the feature of the manhole without a manhole cover to the manhole with a manhole cover in the first image region to obtain the second image region is as follows:

[0124] Input the first image region into a pre-trained style transfer model, and transfer the features of the manhole without a manhole cover to the manhole with a manhole cover in the first image region through the style transfer model, so as to obtain a second image region including the manhole without a manhole cover, wherein the style transfer model is trained by using a first image including a manhole with a manhole cover and a second image including a manhole without a manhole cover.

[0125] In some embodiments, the style transfer model includes a generative adversarial network, and the style transfer model is trained based on a first image including a manhole with a manhole cover and a second image including a manhole without a manhole cover. The specific training process is as follows:

[0126] Use the generator of the generative adversarial network to generate an image including a manhole without a manhole cover based on the first image;

[0127] Determine a first loss based on the discrimination result of the discriminator of the generative adversarial network for the image generated by the generator and the second image;

[0128] Determine a second loss based on the similarity between a first image patch in the image generated by the generator and a second image patch in the first image, and the similarity between the first image patch in the image generated by the generator and a third image patch in the first image; wherein the first image patch and the second image patch are located at the same pixel position, and the first image patch and the third image patch are located at different pixel positions;

[0129] Determine a target loss based on the first loss and the second loss, and use the target loss to train the generative adversarial network.

[0130] In some embodiments, when the replacement module is used to determine a target image region from a target image, it is specifically used for:

[0131] When the target image is the image of the manhole with a manhole cover, use the first image region as the target image region; or

[0132] When the target image is the image of the manhole with a manhole cover or an image other than the image of the manhole with a manhole cover, perform semantic segmentation on the target image, and select the target image region from the target image based on the result of the semantic segmentation.

[0133] In some embodiments, when the replacement module is used to replace the target image region with a second image region including a manhole without a manhole cover to obtain the sample image, it is specifically used for:

[0134] Determine a mask image based on the second image region including a manhole without a manhole cover;

[0135] Extract a foreground image from the second image region using the mask image, and extract a background image from the target image region using the mask image;

[0136] Fuse the foreground image and the background image, and replace the target image region with the fused image to obtain the sample image.

[0137] In some embodiments, when the replacement module is used to determine a mask image based on the second image region, it is specifically used for:

[0138] Perform a scaling process on the second image region so that the size of the second image region is the same as the size of the first image region;

[0139] Perform gray-scale processing on the scaled second image region and then perform denoising processing, and then perform binarization processing on the denoised gray-scale image;

[0140] Perform denoising processing on the binary image obtained by the binarization processing to obtain the mask image.

[0141] Wherein, the specific steps for the above device to execute the sample image generation method can refer to the description in the above method embodiments, and will not be elaborated here.

[0142] Furthermore, an embodiment of the present disclosure further provides an electronic device, as Figure 10 shown, the electronic device includes a processor 110, a memory 120, and computer instructions stored in the memory 120 and executable by the processor 110. When the processor 110 executes the computer instructions, the method described in any one of the above embodiments is implemented.

[0143] An embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in any one of the foregoing embodiments is implemented.

[0144] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0145] From the description of the above embodiments, those skilled in the art can clearly understand that the embodiments of the present disclosure can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions of the embodiments of the present disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present disclosure.

[0146] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.

[0147] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description of the method embodiments. The apparatus embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. When implementing the solutions of the embodiments of the present disclosure, the functions of the modules can be implemented in one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solutions of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0148] The above are only the specific implementation manners of the embodiments of the present disclosure. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principles of the embodiments of the present disclosure, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of the present disclosure.

Claims

1. A method for generating a sample image, characterized in that, The sample image is used to train an object detection model, and the object detection model is used to detect manholes without manhole covers in images. The method includes: Obtain a target image; Determine a target image area from the target image, and replace the target image area with a second image area including a manhole without a manhole cover to obtain a sample image including a manhole without a manhole cover; Wherein, the second image area is obtained based on the following method: intercept a first image area including the manhole with a manhole cover from an image of a manhole with a manhole cover, and transfer the features of the manhole without a manhole cover to the manhole with a manhole cover in the first image area to obtain the second image area. Among them, the features of the manhole without a manhole cover are obtained by extracting features from an image of a manhole without a manhole cover; Wherein, the step of transferring the features of the manhole without a manhole cover to the manhole with a manhole cover in the first image area to obtain the second image area includes: inputting the first image area into a pre-trained style transfer model, and transferring the features of the manhole without a manhole cover to the manhole with a manhole cover in the first image area through the style transfer model to obtain a second image area including a manhole without a manhole cover. Among them, the style transfer model includes a generative adversarial network, and the style transfer model is trained based on the following method: Use the generator of the generative adversarial network to generate an image including a manhole without a manhole cover based on the first image of the manhole with a manhole cover; Determine a first loss based on the discrimination result of the discriminator of the generative adversarial network on the image generated by the generator and the second image including a manhole without a manhole cover; Determine a second loss based on the similarity between a first image patch in the image generated by the generator and a second image patch in the first image, and the similarity between the first image patch in the image generated by the generator and a third image patch in the first image; wherein, the first image patch and the second image patch are located at the same pixel position, and the first image patch and the third image patch are located at different pixel positions; Determine a target loss based on the first loss and the second loss, and use the target loss to train the generative adversarial network.

2. The method according to claim 1, characterized in that, Determining a target image area from a target image includes: If the target image is the image of the manhole with a manhole cover, use the first image area as the target image area; or If the target image is the image of the manhole with a manhole cover or other images except the image of the manhole with a manhole cover, perform semantic segmentation processing on the target image, and select the target image area from the target image based on the result of semantic segmentation.

3. The method according to claim 1 or 2, characterized in that, Replacing the target image area with a second image area including a manhole without a manhole cover to obtain the sample image includes: Determine a mask image based on the second image area including a manhole without a manhole cover; Extract a foreground image from the second image area using the mask image, and extract a background image from the target image area using the mask image; Fuse the foreground image and the background image, and use the fused image to replace the target image area to obtain the sample image.

4. The method according to claim 3, wherein Determining a mask image based on the second image region includes: Performing a scaling process on the second image region so that the size of the second image region is the same as the size of the first image region; Performing gray processing on the scaled second image region and then performing denoising processing, and then performing binarization processing on the denoised gray image; Performing denoising processing on the binary image obtained by the binarization processing to obtain the mask image.

5. A method for training an object detection model, characterized in that, The method includes: Generating a sample image by using the sample image generation method according to any one of claims 1-4; Training a preset initial model by using the sample image to obtain the target detection model.

6. A method for detecting manholes with missing manhole covers on roads, characterized in that, The method includes: Obtaining an image of a road; Inputting the image into a pre-trained target detection model, and detecting manholes with missing manhole covers in the image through the target detection model, wherein the target detection model is trained by a sample image, and the sample image includes an image generated by the sample image generation method according to any one of claims 1-4.

7. A road detection system, characterized in that, The road detection system includes an image acquisition device and a server, the image acquisition device is located on the side of the road or the image acquisition device is carried on a road inspection device, The image acquisition device is used to acquire an image of the road and send it to the server; The server is used to input the image into a pre-trained target detection model, and detect manholes with missing manhole covers in the image through the target detection model, wherein the target detection model is trained by a sample image, and the sample image includes an image generated by the sample image generation method according to any one of claims 1-4.

8. The road detection system according to claim 7, characterized in that, The server is further used to, when it is detected that the image includes a manhole with a missing manhole cover, obtain the position information when the image acquisition device acquires the image, and push an alarm message to vehicles on the road, wherein the alarm message includes the position information.

9. The road detection system according to claim 8, wherein The image acquisition device is carried on a road inspection device, and the server is further used to, when it is detected that the image includes a manhole with a missing manhole cover, control the road inspection device to move to a position near the manhole with the missing manhole cover to send a prompt message through the road inspection device.

10. A sample image generation device, characterized in that, The sample image is used to train a target detection model, and the target detection model is used to detect manholes with missing manhole covers in an image. The device includes: An acquisition module, configured to acquire a target image; A replacement module, configured to determine a target image region from the target image, and replace the target image region with a second image region including a manhole with a missing manhole cover to obtain a sample image including a manhole with a missing manhole cover; wherein the second image region is obtained in the following manner: intercepting a first image region including the manhole with a manhole cover from an image of a manhole with a manhole cover, and migrating the features of the manhole with a missing manhole cover to the manhole with a manhole cover in the first image region to obtain the second image region, wherein the features of the manhole with a missing manhole cover are obtained by performing feature extraction on an image of the manhole with a missing manhole cover. Among them, migrating the feature of the manhole without a manhole cover to the manhole with a manhole cover in the first image area to obtain the second image area includes: inputting the first image area into a pre-trained style transfer model, and through the style transfer model, migrating the feature of the manhole without a manhole cover to the manhole with a manhole cover in the first image area to obtain a second image area including a manhole without a manhole cover. Among them, the style transfer model includes a generative adversarial network, and the style transfer model is trained based on the following method: Using the generator of the generative adversarial network to generate an image including a manhole without a manhole cover based on the first image of the manhole with a manhole cover; Determining a first loss based on the discrimination result of the discriminator of the generative adversarial network on the image generated by the generator and the second image including a manhole without a manhole cover; Determining a second loss based on the similarity between the first image block in the image generated by the generator and the second image block in the first image, and the similarity between the first image block in the image generated by the generator and the third image block in the first image; wherein, the first image block and the second image block are located at the same pixel position, and the first image block and the third image block are located at different pixel positions; Determining a target loss based on the first loss and the second loss, and using the target loss to train the generative adversarial network.

11. An electronic device, characterized in that, The electronic device includes a processor, a memory, and computer instructions stored in the memory and executable by the processor. When the processor executes the computer instructions, the method described in any one of claims 1-6 is implemented.

12. A computer-readable storage medium, characterized in that, Computer instructions are stored in the computer-readable storage medium, and when the computer instructions are executed, the method described in any one of claims 1-6 is implemented.