Image generation method, model training method, detection method, device and system

By migrating the manhole cover missing manhole characteristics in the images with manhole cover manhole cover to generate sample images, the problem of insufficient data on manhole cover missing manhole cover detection model is solved and the detection accuracy is improved.

CN114882207BActive Publication Date: 2025-08-22SHANGHAI SENSETIME LINGANG INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210707324.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-21
Publication Date
2025-08-22
Estimated Expiration
2042-06-21

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect the missing manhole cover manhole cover through the target detection model. This is mainly because there are fewer images of missing manhole cover manhole cover, resulting in insufficient training data and low model accuracy.

Method used

By migrating the characteristics of manhole cover missing manholes from images with manhole cover manholes, sample images are generated, and using style migration models such as generating adversarial networks, the characteristics of manhole cover missing manholes are migrated to images with manholes with manholes, and diverse sample images are generated for training the target detection model.

Benefits of technology

A large number of manhole sample images of missing manhole covers in different scenarios were generated, which improved the diversity of training data and improved the detection accuracy of the target detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114882207B_ABST
    Figure CN114882207B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide an image generation method, a model training method, a detection method, an apparatus, and a system. Features of manholes with missing manhole covers can be extracted from a small number of images of manholes with missing manhole covers that have been obtained, and then these features can be transferred to the manholes with manhole covers in the images of manholes with manhole covers, so that images of manholes with missing manhole covers in different scenarios can be obtained as sample images. In this way, a large number of sample images of manholes with missing manhole covers in different scenarios can be obtained for training target detection models, so that the sample data can be more diversified and the prediction results of the trained target detection model can be more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to an image generation method, a model training method, a detection method, a device, and a system. Background Art

[0002] Since there are often numerous manholes on the ground, these often lack covers. Once a manhole cover is missing, it poses a significant safety hazard to pedestrians and vehicles, resulting in adverse social impacts. Therefore, it is necessary to promptly detect manholes with missing covers and take appropriate protective measures. Currently, it is possible to collect images of roads and use pre-trained object detection models to detect whether the images contain manholes with missing covers. However, due to the limited number of images of manholes with missing covers, it is difficult to encompass different scenarios. This results in incomplete training data for the object detection model, and the resulting trained object detection model is relatively low in accuracy. Summary of the Invention

[0003] The present disclosure provides a sample image generation method, a model training method, a detection method, a device and a system.

[0004] According to a first aspect of an embodiment of the present disclosure, a method for generating a sample image is provided. The sample image is used to train a target detection model. The target detection model is used to detect manholes with missing manhole covers in an image. The method includes:

[0005] Get images of covered manholes;

[0006] The features of the manhole with missing manhole cover are transferred to the manhole with manhole cover in the image of the manhole with manhole cover, so as to obtain a sample image including the manhole with missing manhole cover, wherein the features of the manhole with missing manhole cover are obtained by extracting features from the image of the manhole with missing manhole cover.

[0007] In some embodiments, migrating features of a manhole with a missing manhole cover to the manhole with a manhole cover in the image of the manhole with a manhole cover to obtain a sample image including the manhole with a missing manhole cover includes:

[0008] Inputting the image of the manhole with a manhole cover into a pre-trained style transfer model, and migrating the features of the manhole with a missing manhole cover to the manhole with a manhole cover in the image of the manhole with a manhole cover through the style transfer model, thereby obtaining a sample image including the manhole with a missing manhole cover; wherein the style transfer model is obtained by training a first image including the manhole with a manhole cover and a second image including the manhole with a missing manhole cover; or

[0009] A first image region including the manhole with a manhole cover is captured from the image of the manhole with a manhole cover, and the first image region is input into the style transfer model. The features of the manhole with missing manhole cover are transferred to the manhole with a manhole cover in the first image region through the style transfer model to obtain a second image region including the manhole with missing manhole cover. The first image region is replaced with the second image region to obtain the sample image.

[0010] In some embodiments, the style transfer model includes a generative adversarial network, and the style transfer model is trained based on a first image of a manhole with a manhole cover and a second image of a manhole without a manhole cover, including:

[0011] generating, using a generator of the generative adversarial network, an image of a manhole with a missing manhole cover based on the first image;

[0012] Determine a first loss based on a discrimination result of the discriminator of the generative adversarial network on the image generated by the generator and the second image;

[0013] determining a second loss based on a similarity between a first image block in an image generated by the generator and a second image block in the first image, and a similarity between a first image block in the image generated by the generator and a third image block in the first image, wherein the first image block and the second image block are located at the same pixel position, and the first image block and the third image block are located at different pixel positions;

[0014] A target loss is determined based on the first loss and the second loss, and the generative adversarial network is trained using the target loss.

[0015] In some embodiments, replacing the first image area with the second image area to obtain the sample image includes:

[0016] determining a mask image based on the second image area;

[0017] extracting a foreground image from the second image region using the mask image, and extracting a background image from the first image region using the mask image;

[0018] The foreground image and the background image are fused, and the first image region is replaced with the fused image to obtain the sample image.

[0019] In some embodiments, determining a mask image based on the second image area includes:

[0020] performing scaling processing on the second image area so that a size of the second image area is consistent with a size of the first image area;

[0021] Performing grayscale processing on the scaled second image area and then performing denoising processing, and then performing binarization processing on the denoised grayscale image;

[0022] The binary image obtained by the binarization process is subjected to denoising processing to obtain the mask image.

[0023] According to a second aspect of an embodiment of the present disclosure, a method for training an object detection model is provided, the method comprising:

[0024] Generate a sample image using the sample image generation method mentioned in the first aspect above;

[0025] The preset initial model is trained using the sample images to obtain the target detection model.

[0026] According to a third aspect of an embodiment of the present disclosure, a method for detecting a manhole with a missing manhole cover in a road is provided, the method comprising:

[0027] Acquire an image of the road;

[0028] The image is input into a pre-trained target detection model, and the target detection model is used to detect manholes with missing manhole covers in the image, wherein the target detection model is trained through sample images, and the sample images include images generated by the sample image generation method mentioned in the first aspect above.

[0029] According to a fourth aspect of an embodiment of the present disclosure, a road detection system is provided, comprising an image acquisition device and a server, wherein the image acquisition device is located on a road side or in a road inspection device.

[0030] The image acquisition device is used to acquire images of the road and send them to the server;

[0031] The server is used to input the image into a pre-trained target detection model, and detect manholes with missing manhole covers in the image through the target detection model, wherein the target detection model is trained through sample images, and the sample images include images generated by the sample image generation method mentioned in the first aspect above.

[0032] In some embodiments, the server is also used to obtain the location information of the image acquisition device when it captures the image when a manhole with a missing manhole cover is detected in the image, and push alarm information to vehicles on the road, wherein the alarm information includes the location information.

[0033] In some embodiments, the image acquisition device is installed in a road inspection device, and the server is also used to control the road inspection device to move to a position near the manhole with missing manhole cover when it is detected that the image includes a manhole with missing manhole cover, so as to issue a prompt message through the road inspection device.

[0034] According to a fifth aspect of an embodiment of the present disclosure, a device for generating a sample image is provided, wherein the sample image is used to train a target detection model, wherein the target detection model is used to detect a manhole with a missing manhole cover in an image, the device comprising:

[0035] An acquisition module, used for acquiring images of manholes with manhole covers;

[0036] A migration module is used to migrate the features of the manhole with missing manhole cover to the manhole with manhole cover in the image of the manhole with manhole cover, so as to obtain a sample image including the manhole with missing manhole cover, wherein the features of the manhole with missing manhole cover are obtained by extracting features from the image of the manhole with missing manhole cover.

[0037] According to the sixth aspect of an embodiment of the present disclosure, an electronic device is provided, comprising a processor, a memory, and computer instructions stored in the memory for execution by the processor. When the processor executes the computer instructions, the methods mentioned in the first, second and third aspects above can be implemented.

[0038] According to a seventh aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer instructions are stored. When the computer instructions are executed, the methods mentioned in the first, second and third aspects above are implemented.

[0039] In the disclosed embodiment, although it is difficult to collect images of manholes with missing manhole covers from real-life scenarios, a large number of images of manholes with manhole covers in different scenarios can be obtained. Therefore, the features of manholes with missing manhole covers can be extracted from the small number of images of manholes with missing manhole covers that have been obtained, and then these features can be transferred to the manholes with manhole covers in the images of manholes with manhole covers, so that images of manholes with missing manhole covers in different scenarios can be obtained as sample images. In this way, a large number of sample images of manholes with missing manhole covers in different scenarios can be obtained for training the target detection model, so that the sample data can be more diversified and the prediction results of the trained target detection model can be more accurate.

[0040] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0042] Figure 1 It is a schematic diagram of a sample image generation method according to an embodiment of the present disclosure.

[0043] Figure 2 This is a flowchart of a sample image generation method according to an embodiment of the present disclosure.

[0044] Figure 3 is a schematic diagram of a sample image generation method according to an embodiment of the present disclosure.

[0045] Figure 4 This is a schematic diagram of replacing a first image area with a second image area according to an embodiment of the present disclosure.

[0046] Figure 5 2 is a schematic diagram of training a style transfer model according to an embodiment of the present disclosure.

[0047] Figure 6 A schematic diagram of a road detection system according to an embodiment of the present disclosure.

[0048] Figure 7 A schematic diagram of a road detection system alerting a vehicle according to an embodiment of the present disclosure.

[0049] Figure 8 A schematic diagram of a road detection system providing prompts to vehicles and pedestrians according to an embodiment of the present disclosure.

[0050] Figure 9 It is a logical structure diagram of a sample image generating device according to an embodiment of the present disclosure.

[0051] Figure 10 It is a schematic diagram of the logical structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0052] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0053] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. The singular forms "a", "the" and "the" used in this disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items. In addition, the term "at least one" herein means any combination of at least two of any one or more of a plurality of.

[0054] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining."

[0055] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present disclosure and to make the above-mentioned purposes, features and advantages of the embodiments of the present disclosure more obvious and easy to understand, the technical solutions in the embodiments of the present disclosure are further described in detail below with reference to the accompanying drawings.

[0056] With the modernization of cities, various pipelines need to be laid underground, such as gas or natural gas pipelines, sewers, power pipelines, and water pipes. The above-ground entrances to these pipelines are generally called manholes, and manholes are generally covered by manhole covers. Since there are often a large number of manholes on the ground, these manholes often have missing covers. Once a manhole cover is missing, it poses a serious safety hazard to pedestrians and vehicles, causing adverse social impacts. Therefore, it is necessary to promptly detect manholes with missing covers on the ground and take appropriate protective measures.

[0057] Currently, automated detection of manholes with missing covers typically involves capturing images of manholes along roads and then performing detection to determine whether any of these manholes are found. This image detection typically involves two methods. One method relies on traditional image processing techniques. This technique captures images of manholes and extracts morphological, color, and geometric features from these images. Based on these features, the manholes are identified. This method requires extensive computation and hyperparameter tuning, resulting in poor generalization and difficulty adapting to diverse scenarios. Another method involves object detection based on deep learning neural network models. This method captures a large number of images of manholes with missing covers and uses these images to train a neural network model, resulting in an object detection model capable of detecting manholes with missing covers from images. This method offers greater generalization and adaptability to a variety of scenarios. It also requires fewer hyperparameters and is easier to deploy on lightweight computing devices.

[0058] However, the second approach requires training the neural network model with a large number of images of manholes with missing covers in a variety of scenarios. Only then can the trained object detection model achieve good results and detection accuracy. However, scenes with missing manhole covers are often rare, and images of manholes with missing covers (especially images of manholes with missing covers in different scenarios) are often difficult to obtain. This results in an insufficient number of training samples, a difficulty in covering different scenarios, and a lack of diversity, resulting in suboptimal detection results for the trained object detection model.

[0059] Based on this, the embodiment of the present disclosure provides a sample image generation method. Considering that although images of manholes with missing manhole covers are difficult to collect from real-life scenarios, a large number of images of manholes with manhole covers in different scenarios can be obtained, the features of manholes with missing manhole covers can be extracted from the small number of images of manholes with missing manhole covers that have been obtained, and then these features can be transferred to the manholes with manhole covers in the images of manholes with manhole covers, so that images of manholes with missing manhole covers in different scenarios can be obtained as sample images. In this way, a large number of sample images of manholes with missing manhole covers in different scenarios can be obtained for training target detection models, so that the sample data can be more diversified and the prediction results of the trained target detection model can be more accurate.

[0060] The sample image generation method of the disclosed embodiments can be executed by various electronic devices, such as mobile phones, computers, and cloud servers. For example, in some scenarios, the method can be implemented using an app installed on the electronic device. For example, a user can open the app, import an image of a manhole with a manhole cover, and automatically generate a sample image. Alternatively, the user can import both an image of a manhole with a manhole cover and an image of a manhole without a manhole cover, and then automatically generate a sample image.

[0061] like Figure 1 FIG. 1 is a schematic diagram of a sample image generation method according to an embodiment of the present disclosure. Figure 2 FIG. 1 is a flow chart of a method for generating a sample image according to an embodiment of the present disclosure. The method may include the following steps:

[0062] S202, acquiring an image of a manhole with a manhole cover;

[0063] In step S202, images of covered manholes can be obtained. For example, a user interaction interface can be provided, and then images of covered manholes imported by the user through the "import control" in the interaction interface can be obtained. The images of covered manholes can be images obtained by capturing images of manholes on the road. For example, cameras are usually installed on the road, so images of manholes on the road can be captured by these cameras. Alternatively, in order to obtain images of covered manholes in different scenarios, a road inspection device equipped with a camera can be used to inspect the road, and during the inspection process, images of covered manholes in different road sections can be captured by the camera.

[0064] S204. Migrate the features of the manhole with missing manhole cover to the manhole with manhole cover in the image of the manhole with manhole cover, to obtain a sample image including the manhole with missing manhole cover, wherein the features of the manhole with missing manhole cover are obtained by extracting features from the image of the manhole with missing manhole cover.

[0065] In step S204, some images of manholes with missing manhole covers can be obtained, for example, they can be obtained from the Internet, or road images can be collected to obtain images of manholes with missing manhole covers. Since images of manholes with missing manhole covers are difficult to obtain, the number of them is often small. Therefore, feature extraction can be performed on the images of manholes with missing manhole covers to obtain the features of manholes with missing manhole covers. After obtaining the images of manholes with missing manhole covers, the features of manholes with missing manhole covers extracted from the images of manholes with missing manhole covers can be transferred to the manholes with missing manhole covers in the images of manholes with missing manhole covers to obtain sample images of manholes with missing manhole covers. Through this style transfer method, a large number of sample images of manholes with missing manhole covers in different scenarios can be obtained.

[0066] In some scenarios, feature extraction can be performed in advance on images of manholes with missing covers. For example, feature extraction can be performed on an image of a manhole with missing covers to obtain the features of the manhole. Each subsequent frame of an image of a manhole with a covered manhole can be transferred to the manhole with a covered manhole in that image.

[0067] In some scenarios, feature extraction of images of manholes with missing manhole covers can also be completed in real time. For example, the user can also input a set of images at the same time, which includes images of manholes with manhole covers and images of manholes with missing manhole covers. Then, the features of the manholes with missing manhole covers can be extracted from the images of manholes with missing manhole covers, and the extracted features can be transferred to the manholes with manhole covers in the images of manholes with covered manholes.

[0068] Image processing technology can be used to extract features from the obtained image of the manhole with missing manhole cover to obtain the features of the manhole with missing manhole cover. For example, image processing technology can be used to extract the morphological, color, geometric and other features of the manhole with missing manhole cover in the image. Alternatively, a neural network can be pre-trained to extract features from the image of the manhole with missing manhole cover to obtain the features of the manhole with missing manhole cover. For example, the features of the manhole with missing manhole cover can be learned and extracted through a neural network. Similarly, the extracted features can be transferred to the manhole with manhole cover in the image of the manhole with manhole cover by image processing technology or by a pre-trained neural network. For example, the neural network model can generate adversarial networks, autoencoders, etc. Among them, feature extraction and feature transfer can be implemented by the same neural network model or by different neural network models.

[0069] In some embodiments, migrating features of a manhole with a missing manhole cover to a manhole with a manhole cover in an image of a manhole with a manhole cover, and obtaining a sample image including a manhole with a missing manhole cover, can be achieved by using a pre-trained style transfer model. For example, a first image of a manhole with a manhole cover and a second image including a manhole with a missing manhole cover can be obtained, and then the style transfer model can be trained using the two images. For example, taking a generative adversarial network as an example of a style transfer model, the generator of the generative adversarial network can be used to generate an image of a manhole with a missing manhole cover based on the first image, and then the discriminator of the generative adversarial network can be used to discriminate between the generated image and the second image, and the generative adversarial network can be trained based on the discrimination result.

[0070] After training the style transfer model, the style transfer model can be used to transfer features of the manhole with missing manhole cover to the covered manhole in the image of the covered manhole. For example, in some embodiments, the image of the covered manhole can be directly input into the pre-trained style transfer model, and the style transfer model can be used to transfer features of the manhole with missing manhole cover to the covered manhole in the image of the covered manhole, thereby obtaining a sample image that includes the manhole with missing manhole cover.

[0071] Of course, the above method is to perform full image migration on the image of the manhole with a manhole cover, that is, input the image of the manhole with a manhole cover into the style model to obtain a sample image, which is convenient and fast. However, this method requires the neural network to first identify the manhole area from the image and then complete the feature migration. Since there may be other objects in the image, in the case of a complex scene in the image, due to the presence of more interference, the use of the full image migration method may result in the final generated sample image effect being less than ideal. Therefore, in some embodiments, such as Figure 3 As shown, a first image region including a manhole with a manhole cover can be first captured from an image of a manhole with a manhole cover, and the first image region can be input into a style transfer model. The features of the manhole with a missing manhole cover can be transferred to the manhole with a manhole cover in the first image region through the style transfer model to obtain a second image region including the manhole with a missing manhole cover. Then, the second image region can be used to replace the first image region to obtain a sample image.

[0072] This method first captures a first image region from an image that includes covered manholes. Since this first image region typically contains only covered manholes, or is mostly composed of covered manholes, it presents less interference. A style transfer model is then used to transfer the features of the manholes with missing covers to the covered manholes in the first image region, resulting in a second image region that includes the manholes with missing covers. This second image region is then used to replace the first image region to create a sample image. This method results in a more refined sample image.

[0073] In some embodiments, when replacing the first image area with the second image area to obtain a sample image, the first image area can be directly replaced with the second image area. However, the image obtained in this way will not be very natural. For example, there will be a more obvious fusion boundary, which makes the sample image effect poor.

[0074] In some embodiments, as Figure 4As shown, in order to make the second image area and the image of the covered manhole more naturally merged, and to obtain a better and more realistic sample image, a mask image can be first determined based on the second image area. The mask image is then used to extract the foreground image from the second image area, and the background image is extracted from the first image area of ​​the mask image. The extracted foreground image and background image are fused, and the fused image is then used to replace the first image area in the image of the covered manhole, thereby obtaining a sample image including the manhole with the missing manhole cover. The mask image is a binary image. By performing an AND operation on the mask image and the second image area, the foreground image can be obtained. By inverting the mask image and then performing an AND operation on the mask image with the first image area, the background image can be obtained. The foreground image and background image are then fused, and the fused image is used to replace the first image area. This fusion method eliminates the fusion boundary, seamlessly integrating the features of the manhole with the missing manhole cover into the real scene, and obtaining a more natural sample image.

[0075] The accurate design of the mask image is undoubtedly crucial to the final fusion result. In some implementations, to obtain a more accurate mask image, the second image region can be scaled to match the size of the first image region. Since the second image region output by the style transfer model is typically fixed in size and may differ from the size of the first image region, the second image region can be scaled to match the size of the first image region. The scaled second image region can then be grayscaled. Grayscale processing converts a color image into a grayscale image to facilitate binarization. After grayscale processing, the image can be denoised to remove some noise. The denoised image can then be binarized, for example, using the OTSU algorithm. Since the binarized image may still contain some noise, the binarized image can be denoised again to obtain the mask image. For example, a morphological opening operation can be used to further reduce noise in the binary image, and then a morphological closing operation can be used to connect the regions in the binary image to obtain the final mask image. Through this series of image processing, a precise mask image can be obtained. Using the mask image to fuse the second image region with the image of the covered manhole can also produce a more natural fusion effect.

[0076] In some embodiments, as Figure 5As shown, the style transfer model can be a generative adversarial network. When the generative adversarial network is trained using a first image of a manhole with a manhole cover and a second image of a manhole with a missing manhole cover, the generator of the generative adversarial network can be used to generate an image of the manhole with a missing manhole cover based on the first image. Then, a first loss can be determined based on the discrimination result of the image generated by the generator and the second image by the discriminator. The second loss can be determined based on the similarity between a first image block in the image generated by the generator and a second image block in the first image, and the similarity between a first image block in the image generated by the generator and a third image block in the first image; wherein the first image block and the second image block are located at the same pixel position, and the first image block and the third image block are located at different pixel positions. The third image block can be an image block in the first image that is located at a different pixel position from the first image block, or it can be multiple image blocks. For example, in some scenarios, the third image block can be 256 image blocks in the first image that are located at different pixel positions from the first image block.

[0077] When training a general generative adversarial network, the loss is determined based solely on the discriminator’s discrimination results between the image generated by the generator and the real image, and the generative adversarial network is trained based on this loss. The effect of the generative adversarial network trained in this way is not ideal, and the effect of the sample images generated by the trained generative adversarial network needs to be improved. In order to improve the accuracy of the trained generative adversarial network, the embodiment of the present disclosure introduces the idea of ​​contrastive learning when training the generative adversarial network, that is, considering that the similarity between a certain area in the image generated by the generator and the corresponding area of ​​the area in the input image must be higher than the similarity between the area and other areas of the input image except the corresponding area, so that the image generated by the generator is more accurate. Based on this idea, when training a generative adversarial network, in addition to determining the first loss based on the discriminator's discrimination results on the image generated by the generator and the second image, the second loss can be further determined based on the similarity between the first image block in the image generated by the generator and the second image block in the first image at the same pixel position as the first image block, and the similarity between the first image block in the image generated by the generator and the third image block in the first image at a different pixel position from the first image block. Then, the target loss can be determined based on the first loss and the second loss, and the target loss can be used to train the generative adversarial network. The generative adversarial network trained in this way can greatly improve the effect of generating sample images when performing feature transfer.

[0078] After the generative adversarial network is trained, the image of the manhole with a manhole cover can be input into the generator of the generative adversarial network to obtain a sample image, or the first image area can be input into the generator to obtain the second image area.

[0079] Furthermore, the present disclosure also provides a method for training a target detection model, the method comprising the following steps:

[0080] Generate a sample image using the sample image generation method described in the above embodiment;

[0081] The preset initial model is trained using sample images to obtain the target detection model.

[0082] The specific implementation details of generating the sample image may refer to the description in the above embodiment and will not be repeated here.

[0083] Furthermore, the embodiments of the present disclosure also provide a method for detecting manholes with missing manhole covers in roads. The method can be used to detect manholes with missing manhole covers in roads. The method includes the following steps:

[0084] Acquire an image of the road;

[0085] The image is input into a pre-trained target detection model, and the target detection model is used to detect manholes with missing manhole covers in the image, wherein the target detection model is trained through sample images, and the sample images include images generated by the sample image generation method introduced in the above embodiment.

[0086] For example, a camera installed on the road can be used to capture images of the road, or a camera can be installed in a road inspection device. During the inspection of the road by the inspection vehicle, the camera can be used to capture images of the road. The captured images can then be detected using a pre-trained target detection model to promptly detect manholes with missing covers on the road. In some scenarios, a detection device can be installed in the road inspection device, and the above-mentioned road detection method can be executed using the detection device. In some scenarios, after the camera captures the image of the road, it can also be sent to a cloud server, and the above-mentioned road detection method can be executed by the cloud server.

[0087] The target detection model is obtained by training sample images, and the sample images include images generated by the sample image generation method introduced in the above embodiment.

[0088] Furthermore, the present disclosure also provides a road detection system, such as Figure 6As shown, the road detection system includes an image acquisition device and a server. The image acquisition device is located on the side of the road or in a road inspection device. For example, the image acquisition device can be fixed on the road, or the image acquisition device can also be mounted in a road inspection device. The road inspection device can be a movable intelligent device, for example, a robot, a patrol car, or an autonomous driving vehicle, etc. During the inspection of the road by the road inspection device, images of different road sections can be collected. After collecting the image of the road, the image acquisition device can send it to the server. The server is used to input the received image into a pre-trained target detection model, and detect manholes with missing manhole covers in the image through the target detection model.

[0089] The target detection model is obtained by training sample images, and the sample images include images generated by the sample image generation method introduced in the above embodiment.

[0090] In some embodiments, as Figure 7 As shown, the server is further configured to, upon detecting an image containing a manhole with a missing manhole cover, obtain the location information of the image acquisition device at the time the image was captured and issue an alert to vehicles on the road, wherein the alert includes the location information. If the image acquisition device is fixed on the road, the location information of the image acquisition device can be pre-recorded in the server. The server can then determine the location information of the image acquisition device at the time the image was captured based on the device's identification information carried in the image, thereby locating the manhole with a missing manhole cover. If the image acquisition device is mounted on a road inspection device, the device can also be equipped with a positioning device, such as a GPS. When transmitting the image to the server, the server can simultaneously transmit positioning information, allowing the server to determine the location of the image acquisition device at the time the image was captured and, therefore, locate the manhole with a missing manhole cover. Upon detecting a manhole with a missing manhole cover on the road, the server can issue an alert to vehicles on the road, alerting them to a missing manhole cover at a certain location on a certain road section, so that they can exercise caution. Furthermore, the server can also alert road maintenance personnel to promptly address the abnormal situation.

[0091] In some scenarios, such as Figure 8 As shown, when the server detects that a manhole cover is missing at a certain location on a certain road section, it can also control the road inspection device to stay near the manhole where the cover is missing, and prompt the surrounding pedestrians or vehicles through voice or visual information until the maintenance personnel come to repair it.

[0092] In order to further explain the sample image generation method provided by the embodiment of the present disclosure, it is explained below in conjunction with a specific embodiment.

[0093] In order to obtain more images of manholes with missing covers as sample images for training the target detection model to detect manholes with missing covers on roads, this implementation provides a sample image generation method. The details are as follows:

[0094] 1. Training of style transfer model

[0095] A limited number of original images of manholes with missing manhole covers can be obtained from various channels, and a large number of original images of manholes with manhole covers can be obtained from actual usage scenarios (such as cameras installed in roads or cameras mounted in road inspection devices). The areas where the manholes with missing manhole covers are located in the original images of the manholes with missing manhole covers are marked with bounding boxes, and the areas where the manholes with manhole covers are located in the original images of the manholes with manhole covers are marked with bounding boxes.

[0096] All original images of manholes with missing covers are cropped according to the bounding box, and the partial area of ​​the manholes with missing covers is extracted to obtain partial images of the manholes with missing covers. These partial images are used to form data set A. Similarly, the images of manholes with covered covers are cropped according to the bounding box, and the part of the manholes with covered covers is extracted to obtain partial images of the manholes with covered covers. These images are used to form data set B. The generative adversarial network is trained using the A and B data sets to obtain a style transfer model. The specific training method can be referred to the description in the above embodiment and will not be repeated here.

[0097] 2. Use the style transfer model to obtain a partial image of the manhole with missing manhole cover

[0098] A partial image of a manhole with a manhole cover can be obtained from the B dataset and input into the style transfer model to generate a partial image of a manhole with a missing manhole cover in the corresponding scene.

[0099] 3. Image Fusion

[0100] (1) Obtain the partial image b of the manhole with missing manhole cover generated in step 2, and the original image a including the manhole with manhole cover before cropping corresponding to b.

[0101] (2) Selecting a ROI region from a reasonable area in the original image a, where the reasonable area can be the area where the manhole with a manhole cover is located in the original image a, or other areas where manholes may exist. For example, semantic segmentation can be performed on the original image a first, and the ROI region can be determined from the original image a based on the results of the semantic segmentation.

[0102] (3) Scale b to the same size as the roi area, perform grayscale processing and regional noise reduction to generate b1;

[0103] (4) Using the OTSU algorithm, b1 is binarized and the feature region is further denoised using the morphological opening operation. The feature region is then connected using the morphological closing operation to generate a mask image b_mask;

[0104] (5) Perform an AND operation on the obtained mask image and b to obtain the feature area b_front of the manhole with missing manhole cover. At the same time, perform an inverse AND operation on the ROI area in the original image a and the mask image to obtain the ROI background area roi_bg. B_front and roi_bg are fused and replaced with the ROI area in the original image a to obtain the sample image.

[0105] 4. Training of target detection model

[0106] Use the sample image generated in step 3 and the ROI area corresponding to the sample image as labels; train the preset initial model to obtain the target detection model.

[0107] 5. Road detection

[0108] The road image is captured using a camera installed on the road or a camera mounted in a road inspection device. The image is input into a target detection model, which then detects whether there are manholes with missing covers in the image.

[0109] It is not difficult to understand that the solutions described in the above embodiments can be combined when there is no conflict, and they are not listed one by one in the embodiments of the present disclosure.

[0110] Accordingly, the embodiment of the present disclosure further provides a sample image generating device, wherein the sample image is used to train a target detection model, and the target detection model is used to detect a manhole with a missing manhole cover in an image, such as Figure 9 As shown, the device includes:

[0111] An acquisition module 91 is used to acquire an image of a manhole with a manhole cover;

[0112] The migration module 92 is used to migrate the features of the manhole with missing manhole cover to the manhole with manhole cover in the image of the manhole with manhole cover, and obtain a sample image including the manhole with missing manhole cover, wherein the features of the manhole with missing manhole cover are obtained by extracting features from the image of the manhole with missing manhole cover.

[0113] In some embodiments, the migration module is used to migrate the features of the manhole with missing manhole cover to the manhole with manhole cover in the image of the manhole with manhole cover, and when obtaining the sample image including the manhole with missing manhole cover, specifically to:

[0114] Inputting the image of the manhole with a manhole cover into a pre-trained style transfer model, and migrating the features of the manhole with a missing manhole cover to the manhole with a manhole cover in the image of the manhole with a manhole cover through the style transfer model, thereby obtaining a sample image including the manhole with a missing manhole cover; wherein the style transfer model is obtained by training a first image including the manhole with a manhole cover and a second image including the manhole with a missing manhole cover; or

[0115] A first image region including the manhole with a manhole cover is captured from the image of the manhole with a manhole cover, and the first image region is input into the style transfer model. The features of the manhole with missing manhole cover are transferred to the manhole with a manhole cover in the first image region through the style transfer model to obtain a second image region including the manhole with missing manhole cover. The first image region is replaced with the second image region to obtain the sample image.

[0116] In some embodiments, the style transfer model includes a generative adversarial network. When the style transfer model is trained based on a first image of a manhole with a manhole cover and a second image of a manhole without a manhole cover, the specific training process is as follows:

[0117] generating, using a generator of the generative adversarial network, an image of a manhole with a missing manhole cover based on the first image;

[0118] Determine a first loss based on a discrimination result of the discriminator of the generative adversarial network on the image generated by the generator and the second image;

[0119] determining a second loss based on a similarity between a first image block in an image generated by the generator and a second image block in the first image, and a similarity between a first image block in the image generated by the generator and a third image block in the first image, wherein the first image block and the second image block are located at the same pixel position, and the first image block and the third image block are located at different pixel positions;

[0120] A target loss is determined based on the first loss and the second loss, and the generative adversarial network is trained using the target loss.

[0121] In some embodiments, the migration model is used to replace the first image region with the second image region to obtain the sample image, specifically to:

[0122] determining a mask image based on the second image area;

[0123] extracting a foreground image from the second image region using the mask image, and extracting a background image from the first image region using the mask image;

[0124] The foreground image and the background image are fused, and the first image region is replaced with the fused image to obtain the sample image.

[0125] In some embodiments, when the migration module is used to determine the mask image based on the second image area, it is specifically used to:

[0126] performing scaling processing on the second image area so that a size of the second image area is consistent with a size of the first image area;

[0127] Performing grayscale processing on the scaled second image area and then performing denoising processing, and then performing binarization processing on the denoised grayscale image;

[0128] The binary image obtained by the binarization process is subjected to denoising processing to obtain the mask image.

[0129] The specific steps of the sample image generation method executed by the above apparatus may refer to the description in the above method embodiment, which will not be repeated here.

[0130] Furthermore, the present disclosure also provides an electronic device, such as Figure 10 As shown, the electronic device includes a processor 110, a memory 120, and computer instructions stored in the memory 120 for execution by the processor 110. When the processor 110 executes the computer instructions, the method described in any one of the above embodiments is implemented.

[0131] An embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which implements the method described in any of the aforementioned embodiments when the program is executed by a processor.

[0132] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0133] Through the description of the above implementation methods, it can be seen that those skilled in the art can clearly understand that the embodiments of the present disclosure can be implemented by means of software plus the necessary general hardware platform. Based on this understanding, the technical solution of the embodiments of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments of the present disclosure.

[0134] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.

[0135] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the functions of each module can be implemented in the same one or more software and / or hardware when implementing the embodiment of the present disclosure. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the embodiment. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0136] The above is only a specific implementation of the embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the embodiment of the present disclosure. These improvements and modifications should also be regarded as the scope of protection of the embodiment of the present disclosure.

Claims

1. A sample image generation method, characterized in that: The sample image is used to train a target detection model, and the target detection model is used to detect manholes with missing manhole covers in the image. The method includes: Get images of covered manholes; Migrating the features of the manhole with missing manhole cover to the manhole with manhole cover in the image of the manhole with manhole cover, to obtain a sample image including the manhole with missing manhole cover, wherein the features of the manhole with missing manhole cover are obtained by extracting features from the image of the manhole with missing manhole cover; The step of migrating the features of the manhole with missing manhole cover to the manhole with manhole cover in the image of the manhole with manhole cover to obtain a sample image including the manhole with missing manhole cover includes: intercepting a first image region including the manhole with manhole cover from the image of the manhole with manhole cover, inputting the first image region into a style transfer model, migrating the features of the manhole with missing manhole cover to the manhole with manhole cover in the first image region through the style transfer model to obtain a second image region including the manhole with missing manhole cover, and replacing the first image region with the second image region to obtain the sample image; The replacing the first image area with the second image area to obtain the sample image includes: determining a mask image based on the second image area; extracting a foreground image from the second image region using the mask image, and extracting a background image from the first image region using the mask image; The foreground image and the background image are fused, and the first image region is replaced with the fused image to obtain the sample image.

2. The method according to claim 1, characterized in that Determining a mask image based on the second image area includes: performing scaling processing on the second image area so that a size of the second image area is consistent with a size of the first image area; Performing grayscale processing on the scaled second image area and then performing denoising processing, and then performing binarization processing on the denoised grayscale image; The binary image obtained by the binarization process is subjected to denoising processing to obtain the mask image.

3. A target detection model training method, characterized in that: The method comprises: Generating a sample image using the sample image generation method according to any one of claims 1 to 2; The preset initial model is trained using the sample images to obtain the target detection model.

4. A method for detecting missing manhole covers in roads, characterized in that: The method comprises: Acquire an image of the road; The image is input into a pre-trained target detection model, and the target detection model is used to detect manholes with missing manhole covers in the image, wherein the target detection model is trained through sample images, and the sample images include images generated by the sample image generation method according to any one of claims 1-2.

5. A road detection system, characterized in that: The road detection system includes an image acquisition device and a server. The image acquisition device is located on the road side or in a road inspection device. The image acquisition device is used to acquire images of the road and send them to the server; The server is used to input the image into a pre-trained target detection model, and detect manholes with missing manhole covers in the image through the target detection model, wherein the target detection model is trained through sample images, and the sample images include images generated by the sample image generation method according to any one of claims 1-2.

6. The road detection system according to claim 5, characterized in that: The server is also used to obtain the location information of the image acquisition device when it captures the image when a manhole with a missing manhole cover is detected in the image, and push alarm information to vehicles on the road, wherein the alarm information includes the location information.

7. The road detection system according to claim 5, characterized in that: The image acquisition device is installed in the road inspection device, and the server is also used to control the road inspection device to move to a position near the manhole with missing manhole cover when it is detected that the image includes a manhole with missing manhole cover, so as to issue a prompt message through the road inspection device.

8. A sample image generating device, characterized in that: The sample image is used to train a target detection model, and the target detection model is used to detect manholes with missing manhole covers in the image. The device includes: An acquisition module, used for acquiring images of manholes with manhole covers; a migration module, configured to migrate features of the manhole with missing manhole cover to the manhole with a manhole cover in the image of the manhole with a manhole cover, thereby obtaining a sample image including the manhole with missing manhole cover, wherein the features of the manhole with missing manhole cover are obtained by extracting features from the image of the manhole with missing manhole cover; The step of migrating the features of the manhole with missing manhole cover to the manhole with manhole cover in the image of the manhole with manhole cover to obtain a sample image including the manhole with missing manhole cover includes: intercepting a first image region including the manhole with manhole cover from the image of the manhole with manhole cover, inputting the first image region into a style transfer model, migrating the features of the manhole with missing manhole cover to the manhole with manhole cover in the first image region through the style transfer model to obtain a second image region including the manhole with missing manhole cover, and replacing the first image region with the second image region to obtain the sample image; The replacing the first image area with the second image area to obtain the sample image includes: determining a mask image based on the second image area; extracting a foreground image from the second image region using the mask image, and extracting a background image from the first image region using the mask image; The foreground image and the background image are fused, and the first image region is replaced with the fused image to obtain the sample image.

9. An electronic device, characterized in that: The electronic device includes a processor, a memory, and computer instructions stored in the memory and executable by the processor. When the processor executes the computer instructions, the method according to any one of claims 1 to 4 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which, when executed, implement the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Well cover missing detection system and method based on deep learning

    CN108474866A

  • Defect sample generation method and device, electronic equipment and storage medium

    CN111681162A