A detection method, device and computer device for water surface floating objects
By combining the water body segmentation network and semantic segmentation large model and performing the fusion of significance detection, the problems of low detection accuracy and high application threshold in the prior art are solved, and higher detection accuracy and lower usage threshold are achieved.
Patent Information
- Application Number
- CN202411630425.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-11-15
AI Technical Summary
The prior art has low accuracy and high application threshold requirements in river floating objects detection, especially due to different water styles and diverse types of floating objects in different places, the model training results are not ideal, and the pixel limits on effective targets in the image are high, resulting in poor practical application effects.
By inputting the initial water surface image to a fully trained water body segmentation network for water body segmentation processing, a water body image carrying the segmentation results is obtained; then the image of the area where the water surface floating objects are located in the water body image is input into a preset semantic segmentation model for detection and processing, and combining the significance detection results for fusion processing are obtained to obtain the detection results of water surface floating objects.
This method effectively improves the accuracy of detection of floating objects on the water surface, reduces the pixel requirements for effective image targets, reduces the demand for training data sets, and lowers the threshold for use.
Smart Images

Figure CN119151964B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly to a method, device, and computer equipment for detecting floating objects on the water surface. Background Art
[0002] With the rapid development of the economy and the continuous advancement of urbanization and industrialization, human activities and natural factors have caused serious pollution to the environment. A large number of floating objects appear on the river surface, seriously damaging the river channel landscape and water ecological environment, and directly threatening human survival and development. Monitoring and treatment of river channel floating objects have become key issues in river channel supervision. To solve this problem, existing technologies include a monitoring scheme based on a fixed-position camera and an automatic inspection scheme based on an unmanned aerial vehicle (UAV). Among them, the UAV automatic inspection scheme can not only achieve complete monitoring of the entire river channel but also achieve real-time positioning of the detected floating objects. However, most of the recognition algorithms are supervised image classification, object detection, and semantic segmentation algorithms based on deep learning. Such algorithms require a large amount of training data sets. For river channel floating object data, due to the different water area styles, various types, and different sizes of floating objects in each region, the amount of effective data collected is very limited, and the model training results are often not ideal. Moreover, such algorithms have high pixel limitations for effective targets in images. The size of river channel floating objects in UAV images is often smaller than the pixel requirements of effective targets, and the actual application effect is poor.
[0003] Currently, there is no effective solution to the problems of low accuracy and high application threshold requirements in the detection of floating objects on the water surface in the existing technology. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, device, and computer equipment for detecting floating objects on the water surface in view of the above technical problems.
[0005] In a first aspect, this application provides a method for detecting floating objects on the water surface. The method includes:
[0006] Input the obtained initial water surface image into a trained water body segmentation network for water body segmentation processing to obtain a water body image carrying the segmentation result;
[0007] Input the image of the area where the floating object on the water surface is located in the water body image into a preset large semantic segmentation model for detection processing to obtain a first segmentation probability result corresponding to the area where the floating object on the water surface is located;
[0008] Perform saliency detection on the image of the area where the floating object on the water surface is located in the water body image to obtain a second segmentation probability result corresponding to the floating object;
[0009] Fuse the first segmentation probability result and the second segmentation probability result to obtain the detection result corresponding to the floating objects on the water surface.
[0010] In one embodiment, obtaining the initial water surface image includes:
[0011] Obtain the initial image;
[0012] According to the preset cropping size, crop the initial image to obtain at least one initial water surface image.
[0013] In one embodiment, according to the preset cropping size, the process of cropping the initial image includes:
[0014] Obtain the preset overlapping region cropping rule, where the overlapping region cropping rule characterizes the size of the overlapping region between two adjacent initial water surface images;
[0015] According to the preset cropping size and the overlapping region cropping rule, crop the initial image to obtain the initial water surface image, where the size of the overlapping region is greater than 0.
[0016] In one embodiment, each initial water surface image corresponds to a cropping index; input the obtained initial water surface image into a trained water body segmentation network for water body segmentation processing to obtain a water body image carrying the segmentation result, including:
[0017] Input the initial water surface image into the water body segmentation network for water body segmentation processing to obtain the water body segmentation result, where the water body segmentation result is in the form of a probability matrix;
[0018] Restore the water body segmentation result to the initial water surface image to obtain a water body sub-image carrying the water body segmentation result;
[0019] Based on the cropping index, splice the water body sub-images, where the probability matrix of the area where the water body is located in each water body sub-image is determined according to the water body segmentation result;
[0020] When the area where the water body segmentation result is located is in the overlapping region of at least two water body sub-images, fuse the water body segmentation results in the water body sub-images and calculate to obtain a water body image carrying the segmentation result.
[0021] In one embodiment, input the image of the area where the floating objects on the water surface are located in the water body image into a preset semantic segmentation large model for detection processing, including:
[0022] The water body image is cropped based on a preset cropping size and overlapping area cropping rule, and the cropped water body image and the corresponding water body detection instruction template are input into a well-trained object detection large model to obtain object bounding boxes corresponding to all the floating objects on the water surface;
[0023] Based on the object bounding boxes, the images of the areas where the floating objects on the water surface are located are extracted from the cropped water body image, and the images of the areas where the floating objects on the water surface are located are input into a semantic segmentation large model for detection processing.
[0024] In one embodiment, after the water body image is cropped based on a preset cropping size, and the cropped water body image and the corresponding water body detection instruction template are input into a well-trained object detection large model to obtain object bounding boxes corresponding to all the floating objects on the water surface, the method further includes:
[0025] The images of the areas where the floating objects on the water surface are located are extracted from the water body image, and the images of the areas where the floating objects on the water surface are located are classified to obtain the classification corresponding to the floating objects on the water surface;
[0026] If the category of the floating object on the water surface belongs to a preset abnormal detection category, the images of the areas where the floating objects on the water surface are located corresponding to the floating objects on the water surface are input into a semantic segmentation large model for detection processing.
[0027] In one embodiment, after obtaining the detection results corresponding to the floating objects on the water surface, it further includes:
[0028] The detection results are restored to the images of the areas where the floating objects on the water surface are located in the water body image, and the cropped water body images carrying the detection results are stitched based on the cropping index, cropping size and overlapping area cropping rule;
[0029] When the area where the detection results are located is in at least two cropped water body images, the probability matrices of the cropped water body images for the detection results are fused to calculate a water body image carrying the final probability matrix.
[0030] In one embodiment, performing saliency detection on the images of the areas where the floating objects on the water surface are located in the water body image to obtain a second segmentation probability result corresponding to the floating objects on the water surface, includes:
[0031] Performing saliency detection on the images of the areas where the floating objects on the water surface are located in the water body image to obtain the eigenvalue corresponding to each pixel point in the images of the areas where the floating objects on the water surface are located;
[0032] Calculating the similarity between the eigenvalue corresponding to each pixel point and the eigenvalue corresponding to the adjacent pixel point, and determining the second segmentation probability result based on the target similarity less than the preset similarity threshold in the similarity.
[0033] In one embodiment, obtaining a water body segmentation network includes:
[0034] Obtaining a preset water body image training set, where the water body image training set carries water body feature labels;
[0035] Inputting the water body image training set into a preset initial water body segmentation network for training to obtain a training water body segmentation prediction result, calculating a loss function result based on the training water body segmentation prediction result and the water body feature labels, and backpropagating the gradient of the loss function result to the initial water body segmentation network for iterative training to generate a trained complete water body segmentation network.
[0036] In a second aspect, the present application also provides a water surface floating object detection device. The device includes:
[0037] An acquisition module, configured to input an obtained initial water surface image into a trained complete water body segmentation network for water body segmentation processing to obtain a water body image carrying a segmentation result;
[0038] A calculation module, configured to input an image of the area where the water surface floating object is located in the water body image into a preset semantic segmentation large model for detection processing to obtain a first segmentation probability result corresponding to the area where the water surface floating object is located; performing saliency detection on the image of the area where the water surface floating object is located in the water body image to obtain a second segmentation probability result corresponding to the water surface floating object;
[0039] A generation module, configured to perform fusion processing on the first segmentation probability result and the second segmentation probability result to obtain a detection result corresponding to the water surface floating object.
[0040] For the above-mentioned method, device and computer equipment for detecting water surface floating objects, first input the obtained initial water surface image into a trained complete water body segmentation network for water body segmentation processing to obtain a water body image carrying a segmentation result; input the image of the area where the water surface floating object is located in the water body image into a semantic segmentation large model for detection processing to obtain a first segmentation probability result corresponding to the area where the water surface floating object is located; perform saliency detection on the image of the area where the water surface floating object is located in the water body image to obtain a second segmentation probability result corresponding to the water surface floating object; perform fusion processing on the first segmentation probability result and the second segmentation probability result to obtain a detection result corresponding to the water surface floating object. Through the present application, it is possible to combine a supervised water body segmentation network with a semantic segmentation large model that does not require additional training, greatly reducing the required training set size, lowering the usage threshold, and further, the present application obtains the detection result for the water surface floating object through the calculated first segmentation probability result and the second segmentation probability result, effectively improving the accuracy of processing water surface floating objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 Schematic flowchart of the detection method for floating objects on the water surface in one embodiment;
[0042] Figure 2 Schematic flowchart of the detection method for floating objects on the water surface in a preferred embodiment;
[0043] Figure 3 Schematic flowchart of the abnormal segmentation processing in one embodiment;
[0044] Figure 4 Original image captured by a drone in one embodiment;
[0045] Figure 5 Water body image obtained by water body segmentation processing in one embodiment;
[0046] Figure 6 Detection result obtained after abnormal segmentation processing in one embodiment;
[0047] Figure 7 Mask image of the area where floating objects on the water surface are located in one embodiment;
[0048] Figure 8 Structural block diagram of the detection device for floating objects on the water surface in one embodiment;
[0049] Figure 9 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners
[0050] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0051] In one embodiment, as Figure 1 shown, a detection method for floating objects on the water surface is provided, including the following steps:
[0052] Step S110: Input the acquired initial water surface image into a trained water body segmentation network for water body segmentation processing to obtain a water body image carrying the segmentation result.
[0053] Specifically, the above initial water surface image is an image taken of a water surface such as a river channel or a lake surface, and there can be one or more such initial water surface images. This application does not impose excessive restrictions on the method of obtaining the initial water surface image. For example, when only one initial water surface image is obtained, this initial water surface image is input into the above-mentioned water body segmentation network for water body segmentation processing to obtain a segmentation result for the water body area, and this segmentation result is restored to the initial water surface image to obtain the above-mentioned water body image. Similarly, a drone can also be used to take pictures of the river channel to obtain an initial image. Moreover, in practical applications, considering that the memory of the high-resolution images taken will occupy more space, the initial images taken can be cropped to obtain multiple smaller-sized above-mentioned initial water surface images. Further, the above-mentioned initial water surface images are input into a well-trained water body segmentation network for water body segmentation processing. Among them, the water body segmentation network includes but is not limited to the ViT-Adapter network. The water body segmentation network preferably adopts a fully supervised network. Based on the above-mentioned water body segmentation network, semantic segmentation of the water body area can be completed. Each initial water surface image corresponds to a water body segmentation result. The above-mentioned water body segmentation result is in the form of a probability matrix, and different pixel points correspond to their respective probability values. For example, the water body area generally corresponds to 1, and the non-water body area corresponds to 0. And it can be understood that in practical applications, for the area where the water body area and the non-water body area intersect, the probability result is generally a decimal between 0 and 1. In some preferred embodiments, when the initial water surface image is composed of multiple smaller-sized images cropped from the initial image taken by the drone, the segmentation result is first restored to the corresponding initial water surface image, and then all the initial water surface images are stitched together according to the cropping rules to obtain the above-mentioned water body image. It can be understood that at this time, the segmentation result includes the water body segmentation results corresponding to multiple initial water surface images.
[0054] Step S120: Input the image of the area where the floating objects on the water surface in the water body image into a preset large semantic segmentation model for detection processing to obtain a first segmentation probability result corresponding to the area where the floating objects on the water surface are located.
[0055] Specifically, a local image of the area where the floating objects on the water surface are located in the water body image is obtained. That is, in this embodiment, in order to reduce the calculation amount and improve the calculation efficiency, the area where the floating objects on the water surface are located in the water body image is first detected and segmented, and a local image of the area where the floating objects on the water surface are located is segmented from the water body image. The segmentation method includes, but is not limited to, manually calibrating the above local image, or completing the recognition and segmentation of the local image through a preset algorithm. The above semantic segmentation large model is a large model that does not require additional training, such as SAM (Segment Anything Model). The above semantic segmentation large model is used to detect and segment the image of the area where the floating objects on the water surface are located, and output a first segmentation probability result corresponding to the area where the floating objects on the water surface are located. Among them, in some preferred embodiments, the above first segmentation probability result is in the form of a probability matrix, and different pixel points correspond to corresponding probability values. This probability value represents the classification result of the corresponding position of the pixel point, such as the probability that the point is a floating object on the water surface, and / or the probability that the point is water body, etc.
[0056] Step S130: Perform saliency detection on the image of the area where the floating objects on the water surface are located in the water body image to obtain a second segmentation probability result corresponding to the floating objects on the water surface.
[0057] Specifically, perform saliency detection on the image of the area where the floating objects on the water surface are located, focusing on the area where the floating objects on the water surface are located in the image. In some embodiments, the above saliency detection can be completed using the HC algorithm (Histogram-based Contrast) or the RC algorithm (Region-based Contrast), etc. By performing saliency detection on the image of the area where the floating objects on the water surface are located, the area where the floating objects on the water surface are located is distinguished, and a corresponding second segmentation probability result is obtained. Similarly, in some preferred embodiments, the second segmentation probability result is in the form of a probability matrix, and different pixel points correspond to corresponding probability values. This probability value represents the classification result of the corresponding position of the pixel point.
[0058] In summary, in some embodiments, it is necessary to detect the image of the area where the same water surface floating object is located twice based on two different detection methods, and correspondingly obtain the first segmentation probability result and the second segmentation probability result. Although the above two detection probability results are both detection results of the same target, the probability values corresponding to the same pixel points usually will not be exactly the same. In some embodiments, the number of detections and algorithms for the image of the area where the water surface floating object is located are not limited. The image of the area where the water surface floating object is located can be detected multiple times through different detection algorithms. That is, the above semantic segmentation large model can include multiple large models for detecting water surface floating objects, such as the PerSAM model, the CaR (CLIP as RNN) model, etc. Correspondingly, the above first segmentation probability result also includes multiple probability results, and each probability result corresponds to each of the above semantic segmentation large models. Similarly, the above saliency detection algorithm can include multiple algorithms for detecting the features of water surface floating objects, such as the HC algorithm, the RC algorithm, etc. Correspondingly, the above second segmentation probability result also includes multiple probability results, and each probability result corresponds to each of the above saliency detection algorithms.
[0059] Step S140, fuse the first segmentation probability result and the second segmentation probability result to obtain a detection result corresponding to the water surface floating object.
[0060] Specifically, after obtaining all the first segmentation probability results and the second segmentation probability results for the same water surface floating object, fuse the first segmentation probability result and the second segmentation probability result. Among them, in some preferred embodiments, the above fusion process usually refers to multiplying the probability values corresponding to the same pixel points in the first segmentation probability result and the second segmentation probability result. After fusing the first segmentation probability result and the second segmentation probability result, a unique detection result can be obtained for a water surface floating object. This detection result represents the probability that the position of each pixel point in the image is a water surface floating object. In some preferred embodiments, after obtaining the detection result in the form of a probability matrix, the probability values in this detection result can be filtered according to a preset probability threshold. For example, if the probability value detected as a water surface floating object is greater than 0.8, the position corresponding to this pixel point is determined as the water surface floating object area, so that the final detection result for the water surface floating object can be obtained after filtering.
[0061] Through steps S110 to S140, the initial water surface image is input into a trained water body segmentation network for water body segmentation processing to obtain the corresponding water body image. Then, the image of the area where the floating objects on the water surface in the water body image is input into a large semantic segmentation model that does not require additional training for detection processing to obtain the first segmentation probability result, and the image of the area where the floating objects on the water surface is subjected to saliency detection to obtain the second segmentation probability result. That is, in this application, a supervised semantic segmentation algorithm is combined with a zero-shot detection algorithm for floating objects on the water surface, which not only ensures the detection accuracy but also reduces the computational cost, and has lower requirements for the effective pixels of the floating objects on the water surface. Further, in this application, the first segmentation probability result and the second segmentation probability result are combined to jointly obtain the detection result for the floating objects on the water surface, effectively improving the accuracy of the detection of the floating objects on the water surface.
[0062] In some embodiments, obtaining the initial water surface image includes:
[0063] Obtaining an initial image;
[0064] According to a preset cropping size, the initial image is cropped to obtain at least one initial water surface image.
[0065] Specifically, first, an initial image is obtained. The initial image is an image to be detected captured by a high-resolution camera. In some preferred embodiments, high-resolution drone images can be used. After obtaining the initial image, according to the preset cropping size, the initial image is cropped to obtain a plurality of smaller initial water surface images. Among them, in this embodiment, according to the cropping method, the high-resolution image is cropped in a continuous or overlapping manner, and the cropping size can be set by those skilled in the art according to actual needs. Through this embodiment, the obtained initial image can be cropped, avoiding the problem of high memory occupancy of the original high-resolution initial image, thereby greatly improving the computational efficiency in subsequent processing.
[0066] In some embodiments, according to a preset cropping size, cropping the initial image includes:
[0067] Obtaining a preset overlapping area cropping rule, where the overlapping area cropping rule characterizes the size of the overlapping area between two adjacent initial water surface images;
[0068] According to the preset cropping size and the overlapping area cropping rule, the initial image is cropped to obtain the initial water surface image, where the size of the overlapping area is greater than 0.
[0069] Specifically, in this embodiment, the initial image is cropped according to the overlapping method. First, a preset overlapping area cropping rule is obtained. The overlapping area cropping rule indicates the size of the overlapping area between two adjacent images obtained by cropping when the initial image is cropped. The size of the overlapping area cannot be equal to 0. In some preferred embodiments, the size of the overlapping area is preferably set to 50% of the size of the initial water surface image obtained by cropping. According to the preset cropping size and the above overlapping area cropping rule, the cropping process of the initial image is completed to obtain the initial water surface image. Among them, the size of the initial water surface image is preferably set to 512×512. This embodiment provides an overlapping cropping method for the initial image, which is beneficial to improving the accuracy of identifying water surface floating objects in the subsequent image stitching.
[0070] In some of these embodiments, each initial water surface image corresponds to a cropping index; the obtained initial water surface image is input into a trained water body segmentation network for water body segmentation processing to obtain a water body image carrying the segmentation result, including:
[0071] The initial water surface image is input into the water body segmentation network for water body segmentation processing to obtain a water body segmentation result, where the water body segmentation result is in the form of a probability matrix;
[0072] The water body segmentation result is restored to the initial water surface image to obtain a water body sub-image carrying the water body segmentation result;
[0073] Based on the cropping index, each water body sub-image is stitched, where the probability matrix of the area where the water body included in each water body sub-image is located is determined according to the water body segmentation result;
[0074] When the area where the water body segmentation result is located is in the overlapping area of at least two water body sub-images, the water body segmentation results in the water body sub-images are fused, and a water body image carrying the segmentation result is calculated.
[0075] Specifically, each initial water surface image is input into the water body segmentation network for water body segmentation processing to obtain a water body segmentation result corresponding to each initial water surface image. The water body segmentation result is in the form of a probability matrix of the area where the water body is located. Then the water body segmentation result is restored to the initial water surface image to obtain a water body sub-image. The water body sub-image corresponds to the initial water surface image. Therefore, the cropping index of the water body sub-image is the same as that of the initial water surface image, and the water body sub-image carries the corresponding water body segmentation result. It can be understood that the water body segmentation result is the probability matrix of the area where the water body is located in the water body sub-image.
[0076] It can be understood that the above water body segmentation result is in the form of a probability matrix, and each pixel point has a corresponding probability value. Restoring the water body segmentation result to the original water surface image means multiplying the probability matrix by the corresponding pixel points of the original water surface image pixels.
[0077] Then, based on the cropping index, each water body sub-image is stitched. Since when cropping the original image, there are overlapping regions of a certain size between adjacent original water surface images or adjacent water body sub-images, and the same water body region may be located in multiple water body sub-images at the same time, and there may be some differences in the segmentation results obtained by segmenting different original water surface images. Therefore, when the region where the water body segmentation result is located is in the overlapping region of at least two water body sub-images, the water body segmentation results for the same water body region in each water body sub-image are fused. This fusion is generally to multiply the corresponding water body segmentation results, that is, the probability matrices point by point. In summary, after stitching and fusing each water body sub-image, a water body image carrying the segmentation result is obtained. This water body image includes multiple water body sub-images, and this segmentation result includes the water body segmentation results carried in multiple water body sub-images. In this embodiment, the obtained original image is cropped into multiple original water surface images with overlapping regions, which can effectively avoid the problem of high memory occupancy of the original high-resolution original image. And when stitching, fusing the water body segmentation results for the same water body region in adjacent water body sub-images can effectively improve the accuracy of water body segmentation, and can also play the effect of preventing jagged edges at the graphic boundary and improving the picture quality of the image.
[0078] In some of these embodiments, the image of the region where the floating objects on the water surface in the water body image is input into a preset semantic segmentation large model for detection processing, including:
[0079] The water body image is cropped based on a preset cropping size and overlapping region cropping rule, and the cropped water body image and the corresponding water body detection instruction template are input into a trained object detection large model to obtain target boxes corresponding to all the floating objects on the water surface;
[0080] Based on the target boxes, the image of the region where the floating objects on the water surface is extracted from the cropped water body image, and the image of the region where the floating objects on the water surface is input into the semantic segmentation large model for detection processing.
[0081] Specifically, in this embodiment, the above water body image is obtained by splicing the initial water surface image carrying the water body segmentation result. That is, the difference between the water body image and the initial image is that it carries the segmentation result for the water body area. In this embodiment, when obtaining the target boxes corresponding one by one to the floating objects on the water surface, an unsupervised object detection large model can be used to complete the task. This object detection large model preferably adopts the GroundDINO (Grounding DINO) model. The water body image and the corresponding water body detection instruction template are input into the above object detection large model to obtain the target boxes containing the floating objects on the water surface. Among them, the water body detection instruction template is generally a prompt template containing prompt words. This template is generally used to guide the model to generate a piece of text or instruction of a specific type, theme or format. By adding keywords, phrases or specific format requirements in the template, the above object detection large model can be guided to generate results that meet specific requirements. For example, in this application, the keywords / prompt words can be set as the floating objects on the water surface to be extracted, including but not limited to words such as "floats", "floating", "anomly", etc. Further, the obtained target boxes for the floating objects on the water surface can be summarized according to a preset threshold. In some preferred embodiments, this threshold can be set to 0.1. In summary, based on the water body detection instruction module and the well-trained object detection large model, the target boxes corresponding one by one to the above floating objects on the water surface can be obtained.
[0082] After obtaining the target boxes corresponding to the required floating objects on the water surface, the local image of the area where the floating objects on the water surface are located can be extracted from the water body image, and the local image of the area where the floating objects on the water surface are located is input into the above semantic segmentation large model for precise segmentation processing of the floating objects on the water surface. In practical applications, a water body image may include multiple floating objects on the water surface, or there may be a situation where a large area of water body image only includes one or more small-volume floating objects on the water surface. Through this embodiment, the local area image of the water body image where the floating objects on the water surface are located can be extracted, so that in subsequent semantic segmentation, it is not necessary to detect and segment the entire image, and only the extracted local small image needs to be detected, thereby greatly improving the efficiency of image processing and reducing the waste of computing resources.
[0083] In some of these embodiments, the above method further includes:
[0084] Extract the image of the area where the floating objects on the water surface are located from the water body image, and perform classification processing on the image of the area where the floating objects on the water surface are located to obtain the classification corresponding to the floating objects on the water surface;
[0085] If the category of the floating objects on the water surface belongs to the preset anomaly detection category, input the image of the area where the floating objects on the water surface are located corresponding to the floating objects on the water surface into the semantic segmentation large model for detection processing.
[0086] Specifically, after extracting the local image of the area where the water surface floating objects are located, the local image of the area where the water surface floating objects are located is first classified to obtain the classification result corresponding to the water surface floating objects. In practical applications, it is generally necessary to perform subsequent detection work on some water surface floating objects to facilitate the cleaning work of relevant personnel, such as garbage floating on the water surface, etc.; however, for other types of water surface floating objects, such as aquatic plants floating on the water surface, ships traveling on the water surface, etc., there is no need to perform subsequent cleaning work or subsequent detection tasks. Therefore, in this embodiment, the image of the area where the water surface floating objects are located obtained can be classified. If it is detected that the category of the water surface floating objects belongs to the preset abnormal detection category, such as water surface garbage, then the image of the area where the water surface floating objects are located is input into the semantic segmentation large model for subsequent detection processing. If it is detected that the category of the water surface floating objects does not belong to the preset abnormal detection category, then no subsequent processing is required for it. Through this embodiment, the area where the water surface floating objects belonging to the abnormal detection category can be extracted targeted and the subsequent required detection steps can be performed, without the need to perform complete detection steps on all detected water surface floating objects, thereby greatly improving the efficiency of image processing while ensuring the detection effect.
[0087] In some of these embodiments, the above method further includes:
[0088] Restore the detection result to the image of the area where the water surface floating objects are located in the water body image, and perform splicing processing on the cropped water body image carrying the detection result based on the cropping index, cropping size, and overlapping area cropping rule;
[0089] When the area where the detection result is located is in at least two cropped water body images, fuse the probability matrices of the cropped water body images for the detection result, and calculate to obtain the water body image carrying the final probability matrix.
[0090] Specifically, after obtaining the detection results for all water surface floating objects, it is necessary to splice the segmented and cropped images, so as to better and more intuitively obtain the positional relationship of the water surface floating objects on the water body and the relative positional relationship between all water surface floating objects.
[0091] When performing stitching, it is necessary to correspond to the detection process of floating objects on the water surface in the above text. First, restore the finally obtained detection result of floating objects on the water surface to the corresponding position in the image of the area where the floating objects on the water surface are located. This detection result carries a probability matrix for floating objects on the water surface. Then, perform stitching processing on the cropped water body image with the restored detection result according to the cropping index, cropping size, and overlapping area cropping rule. It can be understood that the cropping index of the cropped water body image is the same as that of the corresponding water body sub-image at the corresponding position, and is also the same as the cropping index of the initial water surface image at the corresponding position. Among them, the cropping index is set according to the position of each initial water surface image in the initial image when cropping the initial image to obtain multiple initial water surface images. Each cropping index includes the position information of the corresponding initial water surface image.
[0092] When stitching each cropped water body image based on the cropping index, since the initial image was cropped based on the preset overlapping area cropping rule in the above text, when stitching, if the area where the detection result is located is in the overlapping area of at least two cropped water body images, it is necessary to fuse the detection results carried in each cropped water body image. In practical applications, the detection results of the same floating object on the water surface carried in different cropped water body images are usually not exactly the same. Therefore, when stitching, it is necessary to fuse the detection results of the same floating object on the water surface in the cropped water body images to obtain a more accurate final detection result for the floating object on the water surface. And in practical applications, it can also play the role of preventing jagged edges at the graphic boundary and improving the picture quality of the image.
[0093] In some embodiments, perform saliency detection on the image of the area where the floating object on the water surface is located in the water body image to obtain a second segmentation probability result corresponding to the floating object on the water surface, including:
[0094] Perform saliency detection on the image of the area where the floating object on the water surface is located in the water body image to obtain the eigenvalue corresponding to each pixel point in the image of the area where the floating object on the water surface is located;
[0095] Calculate the similarity between the eigenvalue corresponding to each pixel point and the eigenvalue corresponding to the adjacent pixel point, and determine the second segmentation probability result based on the target similarity that is less than the preset similarity threshold among the similarities.
[0096] Specifically, perform saliency detection on the image of the area where the floating objects on the water surface are located. In practical applications, the methods of saliency detection include, but are not limited to, using saliency segmentation algorithms, such as the HC algorithm and the RC algorithm, or using saliency segmentation models, such as the ITTI model, the SGAN model (Saliency Generative Adversarial Network), etc. to complete. In some preferred embodiments, a pre-trained convolutional neural network can also be used to extract image features to obtain the eigenvalue or eigenvector corresponding to each pixel point. After detecting the eigenvalues corresponding to each pixel point, calculate the similarity between the eigenvalue corresponding to each pixel point and the eigenvalue corresponding to the adjacent pixel points. If it is detected that the similarity between the eigenvalue corresponding to a certain pixel point and the eigenvalue corresponding to its adjacent pixel points is too low, it indicates that this point is probably the area where the floating object is located. Similarly, if it is detected that the similarity between the eigenvalue corresponding to a certain pixel point and the eigenvalue corresponding to its adjacent pixel points is relatively high, it indicates that this point and its surrounding adjacent pixel points are probably not in the area where the floating object is located. Among them, the calculation of the similarity between pixel points can be realized by calculating the cosine similarity distance between feature points. In summary, traverse all pixel points, filter to obtain the saliency map of the image of the area where the floating object is located, and obtain the second segmentation probability result. Among them, the second segmentation probability result preferably exists in the form of a probability matrix. In the probability matrix of the second segmentation probability result, different probability values represent the probability that the corresponding pixel point belongs to the floating object. Through this embodiment, the saliency map of the image of the area where the floating object on the water surface is located can be calculated. By fusing the calculated second segmentation probability result with the first segmentation probability result calculated above, a more accurate detection and segmentation result can be obtained.
[0097] In some of these embodiments, obtaining a water body segmentation network includes:
[0098] Obtain a preset water body image training set, where the water body image training set carries water body feature labels;
[0099] Input the water body image training set into a preset initial water body segmentation network for training to obtain a training water body segmentation prediction result. Calculate the loss function result according to the training water body segmentation prediction result and the water body feature labels, and reverse transmit the gradient of the loss function result to the initial water body segmentation network for iterative training to generate a trained complete water body segmentation network.
[0100] Specifically, the above water body segmentation network can be any neural network based on deep learning, and a suitable network can be selected according to the device performance in actual use. In some preferred embodiments, the ViT-Adapter network can be used.
[0101] This application also provides a preferred embodiment of a method for detecting floating objects on the water surface. Figure 2In a preferred embodiment, it is a detection method for floating objects on the water surface.
[0102] Step S210: Obtain an initial image, and perform cropping processing on the initial image according to a preset cropping size and the overlapping area cropping rule to obtain multiple initial water surface images. Among them, the above initial image is preferably an image captured by a high-resolution drone, and the initial image is cropped according to a preset cropping size and the overlapping area cropping rule. The cropping size is preferably set to a target length and width of 512, and the overlapping area cropping rule can be set to preferably crop the initial image with a 50% overlapping area.
[0103] Step S220: Perform water body segmentation processing on all initial water surface images respectively to obtain water body segmentation results, and restore the water body segmentation results to the initial water surface images to obtain water body sub-images, and splice the respective water body sub-images to obtain a water body image carrying the segmentation results. Among them, the water body segmentation processing can be completed through a water body segmentation network, and this water body segmentation network can be based on any neural network of deep learning. In this preferred embodiment, ViT-Adapter is taken as an example for illustration. The network framework is mainly divided into two parts. The first part is the classification Backbone network BEIT, which consists of a patch embedding module and 4 Transformer Blocks; the second part is the Adapter module, which is mainly used for the information interaction work between the Adapter branch and the feature map of the ViT network. This module includes a spatial feature extraction module for extracting spatial features from the input initial water surface image, a spatial feature injection module for inputting the spatial features into the ViT, and a multi-scale feature extraction module for extracting multi-scale spatial information from the ViT. In practical applications, the input initial water surface image is first cropped into 16×16 non-overlapping blocks through a Patch Embedding module. Then, these blocks will be flattened into D-dimensional encodings. At this time, the resolution of the feature layer will become 1 / 16 of the original image, and the feature channels are 16×3. Then, these encoded feature blocks plus position encodings are sent to the Transformer Block of the ViT. In another branch, the input image is input into the spatial feature extraction module to generate three spatial feature layers with different resolutions of 1 / 8, 1 / 16, and 1 / 32. After being flattened, these feature maps are spliced together as the input of the spatial feature injection module. The ViT network used in this study includes 4 TransformerBlocks, and each Block contains L / N Transformer layers. Before each Transformer Block, first use the spatial feature injection module to inject the spatial feature layer obtained before the previous Block into the ViT network; after the TransformerBlock, use the multi-scale feature extraction module to extract multi-level spatial features from the output feature layer of the Block. After 4 times of feature interaction, a high-quality multi-scale feature layer will be obtained, and then this feature layer is separated to generate three feature layers of 1 / 8, 1 / 16, and 1 / 32. Finally, use a 2×2 deconvolution layer to upsample the 1 / 8 feature layer to a 1 / 4 feature layer. Finally, input the above-obtained 4 feature layers into the Decoder network for segmentation prediction.
[0104] Step S230: Detect whether all initial water surface images have been traversed. If so, jump to step S240; if not, jump to step S220.
[0105] Step S240: Crop the water body image again according to the above-mentioned cropping index, and input the cropped water body image and the water body detection instruction template into the trained large object detection model to obtain the object bounding boxes corresponding to the floating objects on the water surface. Based on the object bounding boxes, extract the images of the areas where the floating objects on the water surface are located. Among them, first input the cropped water body image and the corresponding water body detection instruction template into the large object detection model to obtain the object bounding boxes of the floating objects. The water body detection instruction template includes prompt words for the floating objects to be extracted, including but not limited to words such as "floats", "floating", "anomaly", etc. Multiple prompt words are input sequentially, and the obtained object bounding boxes are aggregated according to a preset threshold (such as 0.1).
[0106] Step S250: Perform anomaly segmentation processing on all the images of the areas where the floating objects on the water surface are located, and fuse the obtained first segmentation probability result and the second segmentation probability result to obtain the detection result. Among them, the anomaly segmentation processing includes: inputting the image of the area where the floating object on the water surface is located into a large semantic segmentation model (such as SAM) for segmentation processing to obtain the first segmentation probability result corresponding to the area where the floating object on the water surface is located; inputting the image of the area where the floating object on the water surface is located into a saliency segmentation module for saliency segmentation to obtain the second segmentation probability result; multiplying the first segmentation probability result and the second segmentation probability result to obtain the final detection result, and filtering the probability values in the detection result according to a preset probability threshold to obtain the mask image of the area where the floating object on the water surface is located. Figure 3 It is a schematic flowchart of the anomaly segmentation processing in an embodiment.
[0107] Step S260: Determine whether all the images of the areas where the floating objects on the water surface are located have been traversed. If so, jump to step S270; if not, jump to step S250.
[0108] Step S270: Restore the detection result to the image of the area where the corresponding floating object on the water surface is located, and perform splicing processing on the cropped water body images carrying the detection result according to the above-mentioned cropping size and overlapping area cropping rules to obtain the final detection result.
[0109] In summary, Figure 4 is the original image captured by the drone, Figure 5 is the water body image after water body segmentation processing, Figure 6 is the detection result obtained after anomaly segmentation processing, Figure 7 is the mask image of the area where the floating object on the water surface is located.
[0110] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0111] Based on the same inventive concept, an embodiment of the present application also provides a detection device for water surface floating objects for implementing the detection method of water surface floating objects involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the detection device for water surface floating objects provided below can refer to the limitations on the detection method of water surface floating objects in the above text, and will not be repeated here.
[0112] In one embodiment, as Figure 8 shown, a detection device for water surface floating objects is provided, including: an acquisition module 81, a calculation module 82, and a generation module 83, where:
[0113] The acquisition module 81 is configured to input the acquired initial water surface image into a trained water body segmentation network for water body segmentation processing to obtain a water body image carrying the water body segmentation result;
[0114] The calculation module 82 is configured to input the image of the area where the water surface floating object is located in the water body image into a preset semantic segmentation large model for detection processing to obtain a first segmentation probability result corresponding to the area where the water surface floating object is located; perform saliency detection on the image of the area where the water surface floating object is located in the water body image to obtain a second segmentation probability result corresponding to the water surface floating object;
[0115] The generation module 83 is configured to perform fusion processing on the first segmentation probability result and the second segmentation probability result to obtain a detection result corresponding to the water surface floating object.
[0116] Each module in the above detection device for water surface floating objects can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or independent of it, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0117] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structural diagram may be as shown in Figure 9 . The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to the detection of floating objects on the water surface. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for detecting floating objects on the water surface.
[0118] Those skilled in the art can understand that Figure 9 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0119] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties.
[0120] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0121] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0122] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for detecting floating objects on a water surface, characterized in that: The method comprises: Input the acquired initial water surface image into a well-trained water body segmentation network for water body segmentation processing to obtain a water body image with segmentation results; including: inputting the initial water surface image into the water body segmentation network for water body segmentation processing to obtain a water body segmentation result, wherein the water body segmentation result is in the form of a probability matrix; restoring the water body segmentation result to the initial water surface image to obtain a water body sub-image carrying the water body segmentation result; splicing each of the water body sub-images based on a clipping index, wherein a probability matrix of a water body area included in each of the water body sub-images is determined according to the water body segmentation result; each of the initial water surface images corresponds to a clipping index; When the area where the water body segmentation result is located is located in the overlapping area of at least two water body sub-images, the water body segmentation results in the water body sub-images are fused to calculate the water body image carrying the segmentation results; Inputting the image of the area where the floating object on the water surface is located in the water body image into a preset semantic segmentation model for detection processing, and obtaining a first segmentation probability result corresponding to the area where the floating object on the water surface is located; wherein the first segmentation probability result is in the form of a probability matrix; Performing saliency detection on the image of the area where the floating object on the water surface is located in the water body image to obtain a second segmentation probability result corresponding to the floating object on the water surface; wherein the second segmentation probability result is in the form of a probability matrix; The first segmentation probability result and the second segmentation probability result are fused to obtain a detection result corresponding to the water surface floating object.
2. The method according to claim 1, characterized in that Acquiring the initial water surface image, comprising: Get the initial image; The initial image is cropped according to a preset cropping size to obtain at least one initial water surface image.
3. The method according to claim 2, characterized in that The step of cropping the initial image according to a preset cropping size includes: Acquire a preset overlapping region clipping rule, wherein the overlapping region clipping rule represents the size of the overlapping region between two of the initial water surface images at adjacent positions; The initial image is cropped according to a preset cropping size and the overlapping area cropping rule to obtain the initial water surface image, wherein the size of the overlapping area is greater than 0.
4. The method according to claim 1, characterized in that The step of inputting the image of the area where the floating object is located in the water body image into a preset semantic segmentation model for detection processing includes: The water body image is cropped based on a preset cropping size and overlapping area cropping rules, and the cropped water body image and the corresponding water body detection instruction template are input into a well-trained target detection large model to obtain target frames corresponding to all the floating objects on the water surface; An image of the area where the floating objects on the water surface are located is extracted from the cropped water body image based on the target frame, and the image of the area where the floating objects on the water surface are located is input into the semantic segmentation large model for detection processing.
5. The method according to claim 4, characterized in that The method further comprises: cropping the water body image based on a preset cropping size, and inputting the cropped water body image and the corresponding water body detection instruction template into a well-trained target detection large model, and obtaining target frames corresponding to all the floating objects on the water surface. Extracting an image of the area where the floating object is located from the water body image, and classifying the image of the area where the floating object is located to obtain a classification corresponding to the floating object; If the category of the floating object on the water surface belongs to the preset abnormality detection category, the image of the area where the floating object on the water surface is located corresponding to the floating object on the water surface is input into the semantic segmentation large model for detection processing.
6. The method according to claim 4, characterized in that After obtaining the detection result corresponding to the floating object on the water surface, the method further includes: Restoring the detection result to the image of the area where the floating object on the water surface is located in the water body image, and splicing the cropped water body image carrying the detection result based on the cropping index, cropping size and overlapping area cropping rules; When the area where the detection result is located is located in at least two of the cropped water body images, the probability matrices of the cropped water body images for the detection results are fused to calculate a water body image carrying a final probability matrix.
7. The method according to claim 1, characterized in that The performing saliency detection on the image of the area where the floating object on the water surface is located in the water body image to obtain a second segmentation probability result corresponding to the floating object on the water surface includes: Performing saliency detection on the image of the area where the floating object on the water surface is located in the water body image to obtain a feature value corresponding to each pixel point in the image of the area where the floating object on the water surface is located; The similarity between the feature value corresponding to each of the pixel points and the feature value corresponding to the adjacent pixel points is calculated, and the second segmentation probability result is determined based on a target similarity among the similarities that is less than a preset similarity threshold.
8. The method according to claim 1, characterized in that Obtaining the water body segmentation network includes: Obtaining a preset water body image training set, wherein the water body image training set carries a water body feature label; The water body image training set is input into a preset initial water body segmentation network for training to obtain a training water body segmentation prediction result, a loss function result is calculated based on the training water body segmentation prediction result and the water body feature label, and the gradient of the loss function result is reversely transmitted to the initial water body segmentation network for iterative training to generate a fully trained water body segmentation network.
9. A device for detecting floating objects on a water surface, characterized in that: The device comprises: An acquisition module is used to input the acquired initial water surface image into a well-trained water body segmentation network for water body segmentation processing to obtain a water body image carrying a segmentation result; including: inputting the initial water surface image into the water body segmentation network for water body segmentation processing to obtain a water body segmentation result, wherein the water body segmentation result is in the form of a probability matrix; restoring the water body segmentation result to the initial water surface image to obtain a water body sub-image carrying the water body segmentation result; splicing each of the water body sub-images based on a clipping index, wherein a probability matrix of a water body area included in each of the water body sub-images is determined according to the water body segmentation result; each of the initial water surface images corresponds to a clipping index; When the area where the water body segmentation result is located is located in the overlapping area of at least two water body sub-images, the water body segmentation results in the water body sub-images are fused to calculate the water body image carrying the segmentation results; A calculation module, for inputting the image of the area where the floating object on the water surface is located in the water body image into a preset semantic segmentation large model for detection processing, and obtaining a first segmentation probability result corresponding to the area where the floating object on the water surface is located; performing saliency detection on the image of the area where the floating object on the water surface is located in the water body image, and obtaining a second segmentation probability result corresponding to the floating object on the water surface; wherein the first segmentation probability result is in the form of a probability matrix, and the second segmentation probability result is in the form of a probability matrix; A generation module is used to fuse the first segmentation probability result and the second segmentation probability result to obtain a detection result corresponding to the water surface floating object.
Citation Information
Patent Citations
Image scene multi-object segmentation method based on target identification and saliency detection
CN105760886A
Water surface floating object identification method based on semantic segmentation and image anomaly detection
CN116824352A