Multi-target detection method and system on urban river water based on DCBFFNet
By adopting the deep learning model DCBFFNet in urban river water object detection, combining densely connected bidirectional feature fusion module and multi-scale object detection module, the problem of insufficient robustness and generalization capabilities in the existing technology is solved, and high-precision multi-object detection is achieved.
Patent Information
- Application Number
- CN202210309109.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-03-28
AI Technical Summary
The existing technology has problems with robustness and poor generalization capabilities in the classification and identification of urban river water objects. Especially when the water surface environment is complex, the target types are diverse and the scale changes are large, it is difficult to achieve high-precision multi-object detection.
The urban river water multi-object detection method based on the deep learning model DCBFFNet is adopted. By constructing a densely connected bidirectional feature fusion module with object scale sensitivity, combined with the object detection module of multi-scale features, the multi-layer feature map anchor frame size is designed for multi-objects, and the clustering algorithm is used to optimize the anchor frame size to achieve high-precision detection of multi-scale and multi-category targets.
It improves the accuracy and robustness of urban river water object detection, shortens training time, enhances the generalization ability of the detection model, and can effectively deal with multi-category targets in complex water surface environments.
Smart Images

Figure CN114973054B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a method and system for detecting multiple targets on water in urban rivers based on a deep learning model DCBFFNet (Densely Connected Bidirectional Feature Fusion Network, referred to as DCBFFNet). Background Art
[0002] In terms of river and lake health management, timely discovery and salvage of floating objects on the river surface is very important to avoid water pollution caused by the accumulation of floating objects. Early commonly used floating object detection methods such as background difference method and inter-frame difference method are mainly based on bottom-level features and middle-level features (such as color, texture, etc.), and detection is carried out by combining feature extraction with classifiers. In multi-category target detection and recognition tasks, multiple classifiers (such as SVM) are required for classification, which consumes a lot of time in training classifiers. This type of detection method is based on manual feature construction, which is greatly affected by feature selection, target morphology and background changes, and has poor robustness and generalization ability.
[0003] With the rapid development of deep learning, deep learning-based target detection technology has been widely used in many fields due to its powerful feature expression ability, such as aerospace, robot navigation, intelligent security and industrial detection. However, due to the complex water surface environment (such as shore reflection, target reflection, water surface, etc.), the variety of surface targets and the large scale change (from small plastic bottles to large ships), its application effect in water conservancy is not good. In the detection and recognition of multi-category surface targets, surface targets are piled up with the water flow, resulting in occlusion between targets, and the target object category is blurred and difficult to judge. Some small targets such as plastic bottles and cans are small in size and occupy a small image area. Compared with large targets, they lack appearance information, and high-level features lack discriminability, making it difficult to distinguish from the background and achieve accurate positioning, which increases the difficulty of detection. In addition, from the perspective of data sets, compared with single-category target detection tasks, multi-category target detection tasks require higher-quality data sets. In addition to considering the total sample size, it is also necessary to consider the balance between target samples of different categories. Too many or too few samples of a certain target in the data set will affect the learning ability of the detection model. Therefore, it is necessary to study a new and high-performance method for detecting water objects in urban rivers to achieve the complete process from constructing data sets to target detection and recognition. Summary of the invention
[0004] Purpose of the invention: The purpose of the present invention is to overcome the deficiencies of the prior art in the classification and identification of water objects in urban rivers, and to provide a high-precision method and system for detecting and identifying multiple objects on water in urban rivers based on deep learning, so as to provide advanced technology for smart water conservancy and river chief system.
[0005] Technical solution: To achieve the purpose of the present invention, the technical solution adopted by the present invention is: a method for detecting multiple targets on water in urban rivers based on DCBFFNet, comprising the following steps:
[0006] (1) Extract images of water objects related to urban river management from the video stream, including bottle images, waterweed images, mixture images, and ship images. Use the contrast-constrained adaptive histogram equalization method to reduce the noise of the images. Then use rectangular boxes to semantically annotate the locations and category labels of various water objects to construct a target detection dataset.
[0007] (2) Construct an object scale-sensitive densely connected bidirectional feature fusion deep learning model DCBFFNet, including:
[0008] Construct a feature layer selection module for multi-object characteristics, study the scale parameters of each type of detection object in the data set and the receptive field scale parameters of different convolutional layers and the differences in the extracted features at the corresponding scales, and select the feature maps of different scales and resolutions that need to be fused in the model in a targeted manner;
[0009] Construct a dense bidirectional feature fusion module based on multi-scale features, design two transmission connection blocks TCM and TCB with different structures according to the selected feature maps of different scales and resolutions, and construct two mutually inverse transmission paths from top to bottom and from bottom to top respectively with multiple interconnected transmission connection blocks TCM and multiple interconnected transmission connection blocks TCB. Take the selected multi-layer feature maps as input, and complete the top-down and bottom-up feature transmission and feature fusion respectively based on dense connection.
[0010] Design the anchor box size of multi-layer feature maps for multiple objects. Utilize the sampling mapping relationship between the original image and different feature maps, use the clustering algorithm to count the scale classification of various types of water object rectangular annotation boxes mapped on different feature maps, and determine the anchor box size on each layer of feature map selected. The distance in the clustering algorithm is expressed as d = 1-IoU (bboxes, anchor);
[0011] Construct a multi-scale feature-based object detection module, use the output of the dense bidirectional feature fusion module as input, and perform category prediction and regression on multi-scale and multi-category objects;
[0012] (3) Iteratively train the DCBFFNet model based on the optimal anchor box sizes and the constructed object detection dataset;
[0013] (4) Implement the localization and class prediction of various objects on water based on the trained model, and visualize the detection results.
[0014] Furthermore, the feature map selection step in step (2) includes:
[0015] (a) Traverse the marked box scales of various water objects including bottle class, waterweed class, boat class, and mixture class in the entire detection dataset, and visually display them in the form of dot distribution diagrams respectively;
[0016] (b) Calculate the receptive field sizes of different feature layers, compare them with the object scales obtained in (a), and select the feature layers whose receptive fields match the object scales. The receptive field formula is as follows:
[0017]
[0018] where, l k is the receptive field of the kth layer, l k-1 is the receptive field of the (k - 1)th layer, f k is the convolutional kernel size of the kth layer, and s i is the stride of the ith layer;
[0019] (c) Visualize the multiple feature layers obtained in (b), compare the differences in the features obtained by different convolutional layers, and select five feature layers that respectively contain object detail information and high-level semantic information from them.
[0020] Preferably, only four feature maps are used as detection feature maps, and the layer with the largest scale and highest resolution among the selected feature maps is only used to provide target features, and classification and regression are not performed on this layer.
[0021] Preferably, the dense connection between multiple interconnected transmission connection blocks TCM and multiple interconnected transmission connection blocks TCB in step (2) adopts a selective connection method, connecting the TCM / TCB containing low / high-level features with the TCM / TCB containing high / low-level semantic features successively to achieve cross-layer feature interaction; the process of feature fusion between different layers is expressed as:
[0022] X p = Φ p {Φ f {T k (X k )}}
[0023] where, X k represents the feature map to be fused, and T kIndicates that the feature map is upsampled or downsampled to ensure the scale consistency of feature fusion; Φ f Indicates the feature fusion method; Φ p Represents the convolution processing of the fused features; when fusing features, the elements at corresponding positions on corresponding channels of the same scale are added.
[0024] Preferably, the feature layer selection module is composed of seven groups of sequentially connected convolution structures and classification layers and regression layers, the convolution structure is used for feature extraction, and the classification layer and regression layer realize binary classification and regression of "background-target"; after feature difference analysis, the four-layer feature maps C3, C4, C5, and C6 output by the fourth to seventh convolution structures are selected as detection feature maps; in the dense bidirectional feature fusion module, the first, second, and third transmission connection blocks TCM connected in sequence are used to fuse C6, C5, C4, C3 and the feature maps output by the third convolution structure in the feature layer selection module from top to bottom; wherein the first transmission connection block TCM takes the feature map after the convolution operation of C6 and C5 as input, and outputs it to the second transmission connection block TCM after fusion, the second transmission connection block TCM takes the feature map after the convolution operation of C4 and the output of the first transmission connection block TCM as input, and outputs it to the third transmission connection block TCM after fusion, and the third transmission connection block TCM takes the feature map after the convolution operation of C3 and C2 and the feature map output by the second transmission connection block TCM The output is used as input, and the feature map obtained after fusion is used as the first scale feature map in the target detection module; the first transmission connection block TCB takes the output of the third transmission connection block TCM and the output of the second transmission connection block TCM as input, and the feature map obtained after fusion is used as the second scale feature map in the target detection module, and the other output is sent to the second transmission connection block TCB; the second transmission connection block TCB takes the output of the first transmission connection block TCB and the element-by-element fusion features of the three transmission connection blocks TCM outputs as input, and the feature map obtained after fusion is used as the third scale feature map in the target detection module, and the other output is sent to the third transmission connection block TCB; the third transmission connection block TCB has one input as the output of the second transmission connection block TCB, and the other input takes the element-by-element fusion features of the three transmission connection blocks TCM outputs and the C6 convolution operation output as input, and the feature map obtained after fusion is used as the fourth scale feature map in the target detection module; the target detection module performs detection, recognition and regression of multi-category objects based on the feature maps of the first scale to the fourth scale.
[0025] Preferably, in the step (2), the optimized K-means clustering algorithm is used to respectively count the scale classifications of the rectangular annotation boxes of the four types of water objects mapped on different feature maps, and determine the sizes of the anchor boxes on the selected feature maps of each layer. The specific process is as follows: Under the feature maps of each layer, different K values are pre-selected, and the clustering effects under different K values are compared according to the maximum possible recall rate parameter, and the optimal K value of the layer is selected. Based on the clustering effects under the optimal K value obtained at each layer, anchor boxes with different numbers and aspect ratios are designed in different feature maps; based on the aspect ratio a of the anchor boxes obtained with the optimal K value, r , the model can calculate the specific width and height of the anchor box:
[0026]
[0027] in,
[0028]
[0029] m represents the number of feature maps; S max , S min Respectively represent the ratio of the anchor box to the image; S k Indicates the ratio of the feature layer anchor box to the image.
[0030] A multi-target detection system for urban rivers based on DCBFFNet, including:
[0031] The target detection dataset construction module is used to extract images of water objects related to urban river management from the video stream, including bottle images, waterweed images, mixture images, and ship images, perform noise reduction on the images, and then use rectangular frames to semantically annotate the locations and category labels of various water objects to construct the target detection dataset;
[0032] The deep learning model is a scale-sensitive densely connected bidirectional feature fusion deep learning model DCBFFNet. Its construction process includes: constructing a feature layer selection module for multi-object characteristics, studying the scale parameters of each type of detection object in the data set and the receptive field scale parameters of different convolutional layers and the differences in the extracted features under the corresponding scales, and selecting feature maps of different scales and resolutions that need to be fused in the model in a targeted manner; constructing a dense bidirectional feature fusion module based on multi-scale features, designing two different structures of transmission connection blocks TCM and TCB according to the selected feature maps of different scales and resolutions, and constructing multiple interconnected transmission connection blocks TCM and multiple interconnected transmission connection blocks TCB respectively. The two mutually inverse transmission paths, top-down and bottom-up, take the selected multi-layer feature maps as input, and complete the top-down and bottom-up feature transmission and feature fusion respectively based on dense connection; design the anchor box size of the multi-layer feature map for multiple objects, use the sampling mapping relationship between the original map and different feature maps, and use the clustering algorithm to count the scale classification of the rectangular annotation boxes of various water objects mapped on different feature maps, and determine the anchor box size on each layer of the selected feature map. The distance in the clustering algorithm is represented by d=1-IoU(bboxes,anchor); construct an object detection module based on multi-scale features, take the output of the dense bidirectional feature fusion module as input, and perform category prediction and regression on multi-scale and multi-category targets;
[0033] The model training module is used to iteratively train the DCBFFNet model based on the optimal anchor box size and the constructed object detection dataset;
[0034] And the water multi-target detection module is used to locate and predict the categories of various objects on the water based on the trained model, and visualize the detection results.
[0035] Preferably, in the DCBFFNet model, the feature layer selection module is composed of seven groups of sequentially connected convolution structures and classification layers and regression layers. The convolution structure is used for feature extraction, and the classification layer and regression layer realize the binary classification and regression of "background-target". After feature difference analysis, the four-layer feature maps C3, C4, C5, and C6 output by the fourth to seventh convolution structures are selected as detection feature maps. In the dense bidirectional feature fusion module, the first, second, and third transmission connection blocks TCM connected in sequence are used to fuse C6, C5, C4, C3 and the feature maps output by the third convolution structure in the feature layer selection module from top to bottom. The first transmission connection block TCM takes the feature maps after the convolution operation of C6 and C5 as input, and outputs them to the second transmission connection block TCM after fusion. The second transmission connection block TCM takes the feature map after the convolution operation of C4 and the output of the first transmission connection block TCM as input, and outputs them to the third transmission connection block TCM after fusion. The third transmission connection block TCM takes the feature map after the convolution operation of C3 and C2 and the feature map output by the second transmission connection block TCM. The output of the connection block TCM is used as input, and the feature map obtained after fusion is used as the first scale feature map in the target detection module; the first transmission connection block TCB takes the output of the third transmission connection block TCM and the output of the second transmission connection block TCM as input, and the feature map obtained after fusion is used as the second scale feature map in the target detection module, and the other output is sent to the second transmission connection block TCB; the second transmission connection block TCB takes the output of the first transmission connection block TCB and the element-by-element fusion features of the three transmission connection block TCM outputs as input, and the feature map obtained after fusion is used as the third scale feature map in the target detection module, and the other output is sent to the third transmission connection block TCB; the third transmission connection block TCB has one input as the output of the second transmission connection block TCB, and the other input takes the element-by-element fusion features of the three transmission connection block TCM outputs and the C6 convolution operation output as input, and the feature map obtained after fusion is used as the fourth scale feature map in the target detection module; the target detection module performs detection, recognition and regression of multi-category objects based on the first to fourth scale feature maps.
[0036] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the computer program is loaded into the processor, the method for detecting multiple targets on water in urban rivers based on DCBFFNet is implemented.
[0037] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the DCBFFNet-based method for detecting multiple targets on water in urban rivers.
[0038] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:
[0039] 1) In the target detection and recognition model based on anchor frames, the selection of anchor frames affects the final effect. Different from the previous design method of setting anchor frames with the same aspect ratio on multiple layers of feature maps according to the scale of the target in the original image, the present invention determines the number and aspect ratio of anchor frames on feature maps of different scales according to the distribution of the features of four types of water objects, namely bottles, water plants, boats, and mixtures, on different feature maps, thereby improving the effectiveness of target detection on each layer of feature maps, shortening the training time, and improving detection accuracy.
[0040] 2) In the present invention, a scale-sensitive densely connected bidirectional feature fusion deep learning model DCBFFNet is constructed. Combining the scale characteristics of four types of aquatic objects with the differences in the receptive field size and extracted features of different convolutional layers, multiple layers of feature maps of different scales and resolutions are selected from the model, and two transmission connection blocks with different structures are constructed. Taking multiple layers of feature maps of different scales as input, feature transmission and feature bidirectional fusion are performed based on dense connection, which improves the full fusion of feature layers of different object classes and the classification and recognition performance of aquatic objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flow chart of a method for detecting multiple targets on water in urban rivers based on a densely connected bidirectional feature fusion deep learning model DCBFFNet that is sensitive to object scale, provided by the present invention;
[0042] Figure 2 It is a schematic diagram of the process of constructing a detection data set in the present invention;
[0043] Figure 3 It is a network structure diagram of the DCBFFNet model constructed by the present invention;
[0044] Figure 4 is a network structure diagram of a transmission connection block TCM in the present invention;
[0045] Figure 5 It is a network structure diagram of the transmission connection block TCB in the present invention;
[0046] Figure 6 It is a visualization effect diagram of the output feature maps of different convolutional layers in the present invention (taking Conv4_3 layer and Conv5_3 layer as examples);
[0047] Figure 7 This is an experimental diagram of the multi-target detection method on urban rivers based on DCBFFNet in the present invention;
[0048] Figure 8 This is a PR (Precision-Recall) curve comparison chart of the DCBFFNet algorithm in the present invention and other classic algorithms. DETAILED DESCRIPTION
[0049] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0050] The present invention discloses a method for detecting multiple targets on water in urban rivers based on a deep learning model DCBFFNet. Figure 1 , which shows the algorithm flow of an embodiment of the present invention, Figure 3 It is a network structure diagram of the DCBFFNet model in the present invention, and the method flow specifically includes the following steps:
[0051] Step 1: Collect images and build a multi-target detection dataset
[0052] Images of water objects involved in urban river management are extracted from the video stream, and the images taken are subjected to noise reduction and annotation to construct a target detection data set. In this embodiment, in order to maintain the diversity of water objects, the VideoCaputer method of the OpenCV library is used to capture image data of river monitoring videos at different times, different weather conditions and different scenes. The capture time interval t is determined by the water flow speed in the video monitoring area. In this example, t=3, that is, an image is captured from the video image every 3 seconds. In addition, from the perspective of the sample size requirements and sample balance requirements contained in the construction of the training data set, combined with the water objects involved in urban river management, 560 bottle images, 2000 waterweed images, 2168 boat images, and 2550 mixed images are selected from the captured images according to the spatial position and morphological characteristics of the images of different objects. And the four types of objects are annotated with rectangular boxes for position and semantics, and a target detection data set containing four types of objects is constructed.
[0053] In addition, based on some problems existing in the four types of object images (such as uneven illumination, etc.), the image is enhanced using the limited contrast adaptive histogram equalization method. The limited contrast adaptive histogram equalization algorithm divides the image into multiple sub-regions, performs histogram equalization on each sub-region, and suppresses noise amplification and local contrast enhancement by limiting the height of the local histogram. The specific process is as follows:
[0054] 1) Divide the sub-region into M*M sliding windows to obtain its local mapping function m(i):
[0055]
[0056] And the slope S of the local mapping function m(i):
[0057]
[0058] Among them, C DF (i) represents the cumulative distribution function of the local histogram of the sliding window; histogram H ist (i) represents the cumulative distribution function C DF The derivative of (i);
[0059] 2) By setting the threshold T, the histogram is truncated and divided. The truncated part will be evenly distributed in the entire grayscale stage, which can increase the histogram height without changing the total area. The increased height is expressed as L, and the final histogram can be expressed as:
[0060]
[0061] in,
[0062]
[0063] S max To limit the maximum mapping function slope, we can change S max and its corresponding maximum histogram height H max To obtain different enhancement effects.
[0064] Step 2: Construct an object scale-sensitive densely connected bidirectional feature fusion deep learning model DCBFFNet
[0065] Based on the acquired scale features of four types of water objects and the differences in features extracted by different convolutional layers, the present invention constructs a densely connected bidirectional feature fusion deep learning model DCBFFNet that is sensitive to object scale. The model mainly includes three innovative modules: a feature layer selection module, a dense bidirectional feature fusion module and a target detection module. The feature layer selection module consists of 7 groups of sequentially connected convolution structures, classification layers and regression layers. The convolution structure is used for feature extraction, and the classification layer and regression layer realize binary classification and regression of the target. At the same time, the module provides the multi-scale feature map required by the model to realize the detection of multi-scale objects. The dense bidirectional feature fusion module constructs two transmission connection blocks with different structures to realize the transmission and bidirectional fusion of multi-scale features. At the same time, it connects the feature layer selection module and the target detection module, so that the model realizes two-step cascade regression. The feature layer selection module performs binary classification and regression of "background-target", and the target detection module performs detection, recognition and regression of multi-category objects based on the binary classification results.
[0066] Among them, when selecting feature maps of different scales and resolutions that need to be fused in the feature layer selection module, an object-oriented feature layer selection method is adopted. This method studies the scale parameters of each type of detection object in the data set and the receptive field scale parameters of different convolutional layers and the differences in the extracted features at the corresponding scales; then, based on the effectiveness and diversity of feature selection, feature maps of different scales and resolutions that need to be fused in the model are selected in a targeted manner. The specific feature map selection steps mainly include:
[0067] (a) Traversing the scale of the marker boxes of the four types of water objects, namely bottles, water plants, boats, and mixtures, in the entire detection dataset, and visually presenting them in the form of dot distribution graphs;
[0068] (b) Calculate the receptive field sizes of different feature layers and compare them with the object scale obtained in (a). Select the feature layer whose receptive field matches the object scale. The receptive field formula is as follows:
[0069]
[0070] Among them, l k-1 is the receptive field of the k-1th layer, f k is the size of the convolution kernel of the kth layer, s i is the step size of the i-th layer;
[0071] (c) Visualization of multiple feature layers obtained in (b) (schematic diagram as shown in Figure 6 As shown in the figure, the differences in features obtained by different convolutional layers are compared, and five feature layers containing object detail information and high-level semantic information are selected, denoted as {C2, C3, C4, C5, C6}; however, in order to reduce the inference time, only four layers of feature maps are used as detection feature maps, and the layer C2 with the largest scale and highest resolution in the selected feature map is only used to provide target features, and classification and regression are not performed at this layer.
[0072] The feature maps selected in this example are the feature maps output by the five convolutional layers conv3_3, conv4_3, conv5_3, fc7, and conv6_2, and the scale of the feature maps decreases layer by layer by 2x, which are recorded as {C2, C3, C4, C5, C6} respectively.
[0073] In addition, based on the multi-layer feature layers obtained by the above method, this example designs multiple interconnected transmission connection blocks TCM and TCB to form a dense bidirectional fusion module based on multi-scale features, realizing multi-scale feature transmission and feature bidirectional fusion. According to the selected feature maps of different scales and resolutions, two different structures of transmission connection blocks TCM and TCB are constructed respectively. Multiple interconnected transmission connection blocks TCM and multiple interconnected transmission connection modules TCB form a top-down path and a bottom-up path respectively. The model performs feature bidirectional fusion through these two inverse paths. At the same time, features are transmitted between TCM and TCB in a dense connection manner. The dense connection method is selective connection. The low-level feature maps output by TCM are transmitted layer by layer to the high-level features of TCB and fused to realize cross-layer feature interaction. Feature fusion is performed directly between different layers, which reduces the information loss caused by the deepening of the convolution layer. The result of feature fusion can be obtained by the following formula:
[0074]
[0075] Among them, X k represents the feature map to be fused, T k Indicates that the feature map is upsampled or downsampled to ensure the scale consistency of feature fusion; Indicates the fusion method, that is, the channel splicing method or the corresponding element addition method. Indicates the convolution processing of the fused features to avoid the overlapping effect caused by upsampling. The present invention adopts the method of adding the elements at corresponding positions on the corresponding channels of the feature maps of the same scale to perform feature fusion. The specific fusion process can be specifically expressed as:
[0076]
[0077] Represents the feature vector of the fused feature map at (i, j), Represents the feature vector where the feature to be fused is located at (i, j). Represents the feature vector at (i, j) after scaling the feature of layer l to the feature scale of layer k.
[0078] In the dense bidirectional feature fusion module, since the number of channels among the selected feature maps {C2, C3, C4, C5, C6} is different, it is necessary to reduce the channel number through 3*3 convolution before transmission and fusion to keep the same number of channels (channel=256) and output the feature maps {P2, P3, P4, P5, P6}.
[0079] In addition, the dense bidirectional feature fusion module forms a top-down transmission path through three sequentially connected transmission connection blocks TCM, in which the first transmission connection block TCM takes feature maps P5 and P6 as two inputs respectively. After 3*3 convolution enhancement, one of them uses 2*2 deconvolution to double upsample the smaller feature map P6 to make it reach the same scale as P5, and then fuses the feature maps of the two inputs element by element. The fused features are output to the second transmission connection block TCM on the one hand, and on the other hand, they are passed to the second transmission connection block TCB after eliminating the overlapping effect through 3*3 convolution; the second transmission connection block TCM has the same structure as the first transmission connection block TCM, and the feature maps are transmitted in the same way. The feature map P4 is fused with the feature map output by the first transmission connection block TCM, and the fused feature maps are respectively passed to the third transmission connection block TCM and the first transmission connection block TCB; the third transmission connection block TCM takes the feature maps P2, P3 and the output of the second transmission connection block TCM as three inputs, one of which uses 3*3 convolution to downsample the feature map P2 to P3 scale, and the other uses 2*2 deconvolution to upsample the feature map output by the second transmission connection block TCM to P3 scale. Finally, the three feature maps are fused element by element, and after using 3*3 convolution to eliminate the overlapping effect, they are passed as input to the first transmission connection block TCB on the one hand, and on the other hand, they are used as the first-scale feature map K3 in the target detection module. In addition, three sequentially connected transmission connection blocks TCB form a bottom-up transmission path, and each TCM is selectively connected with the TCB, that is, each transmission connection block TCM is connected and fused layer by layer with the transmission connection block TCB at a higher level than itself. The specific connection and fusion method is as follows: the first transmission connection block TCB takes the output of the third transmission connection block TCM and the output of the second transmission connection block TCM as two inputs, one of which uses 3*3 convolution to downsample the output of the third transmission connection block TCM to the scale of the first transmission connection block TCB, and then fuses it element by element with the other input. The feature obtained after fusion Figure 1 On the one hand, it is transmitted to the second transmission connection block TCB, and on the other hand, it is used as the second scale feature map K4 in the target detection module; the second transmission connection block TCB has the same structure as the first transmission connection block TCB, but the difference is that one input of the second transmission connection block TCB is the output of the first transmission connection block TCB, and the other input is the fused feature of the output of the second and third transmission connection blocks TCM, that is, the second transmission connection block TCB fuses the output of the first transmission connection block TCB and the output of the second and third transmission connection blocks TCM, and the fused feature Figure 1On the one hand, it is transmitted to the third transmission connection block TCB, and on the other hand, it is used as the third scale feature map K5 in the target detection module; the third transmission connection block TCB has one input as the output of the second transmission connection block TCB, and the other input as the element-by-element fusion feature of the output of the second and third transmission connection blocks TCM and the C6 convolution operation output as input, and the feature map obtained after fusion is used as the fourth scale feature map K6 in the target detection module.
[0080] In addition, the DCBFFNet model uses anchor boxes as the basis for classification and regression. Whether the anchor box design matches the scale of the four types of water objects affects the final detection effect of the model. Therefore, this example constructs an object-oriented multi-layer feature map anchor box size design module to design anchor boxes that match the scales of the four types of water objects. This is completely different from the difficulty of single-category target detection, because there may be contradictions between the sizes of anchor boxes for different targets, so balancing the sizes and numbers of anchor boxes for different targets is the key. This method uses the sampling mapping relationship between the original image and different feature maps, and uses the optimized K-means clustering algorithm to respectively count the scale classification of the rectangular annotation boxes of the four types of water objects mapped on different feature maps, and determine the size of the anchor boxes on the selected feature maps of each layer, where d=1-IoU(bboxes,anchor) is used to represent the distance, instead of the Euclidean distance in the K-means algorithm. Here, IoU(bboxes,anchor) represents the intersection-over-union ratio of the target annotation box bboxes and the anchor box anchor. The specific process is as follows: Under each layer of feature map, pre-select different K values (K∈[3,9]) and maximize the recall rate parameter, compare the clustering effects under different K values, select the optimal K value of the layer, and based on the optimal K value obtained at each layer, different numbers and aspect ratios of anchor boxes can be obtained in different feature maps. The aspect ratio of the anchor box obtained based on the optimal K value is a r , the model can calculate the specific width and height of the anchor box:
[0081]
[0082] in,
[0083]
[0084] m represents the number of feature graphs, and in the present invention, m=4; S max , S min Respectively represent the ratio of the anchor box to the image (S max =0.9, S min =0.2); S k Indicates the ratio of the feature layer anchor box to the image.
[0085] In this example, according to this method, the optimal K values of each feature layer are obtained, which are {5, 3, 3, 3} respectively. The aspect ratios of the anchor boxes set on the C3, C4, C5, and C6 layers are {2:5, 2:3, 1, 3:2, 5:2}, {1:2, 1:1, 2:1}, {2:3, 1:1, 3:2}, and {1:3, 1:1, 3:1} respectively. That is, the number of anchor boxes set for each position unit on the four layers of feature maps is {5, 3, 3, 3} respectively, and the anchor box scales are set to {24 2 ,48 2 ,96 2 ,192 2}.
[0086] Step 3: Train the DCBFFNet model
[0087] This step trains the DCBFFNet model. The training mode uses densely sampled anchor boxes on feature maps of different scales as training samples. According to the number of anchor boxes and aspect ratios obtained in step 3, densely sample anchor boxes of different numbers and aspect ratios on the feature maps of different scales obtained in step 2. The DCBFFNet model is iteratively trained with an initial learning rate of 0.0001 and a fixed step-size decay learning rate. The DCBFFNet model uses the focal loss function as the classification loss function and Smooth_L1 as the regression loss function, and optimizes the model parameters through iterative training.
[0088] The total loss function includes the classification and regression loss functions of the feature layer selection module and the classification and regression loss functions of the target detection module, which can be expressed as:
[0089]
[0090] Among them, i represents the i-th anchor box in a mini-batch; N FLSM , N ODM Respectively represent the number of positive samples in the feature layer selection module and the target detection module; L FL represents the classification loss function, L smooth-L1 represents the regression loss function.
[0091] Step 4: Test the DCBFFNet model
[0092] Based on the trained model, the positioning and category prediction of four types of objects on the water are realized, and the prediction box and category label are used to complete the visualization of the detection results. Since the dense bidirectional feature fusion module connects the feature layer selection module and the target detection module to realize two-step cascade regression, two classification predictions are required in the prediction and reasoning stage, namely the binary classification of "background-target" and the multi-classification of multiple types of objects.
[0093] Based on the same inventive concept, an urban river water multi-target detection system based on DCBFFNet is disclosed in an embodiment of the present invention, including: a target detection data set construction module, which is used to extract images of water objects related to urban river management from video streams, including bottle images, waterweed images, mixture images, and ship images, perform noise reduction processing on the taken images, and then use rectangular frames to semantically annotate the positions and category labels of various water objects to construct a target detection data set; a deep learning model, which is an object scale-sensitive dense connection bidirectional feature fusion deep learning model DCBFFNet, and its construction process includes: constructing a feature layer selection module for multi-object characteristics, respectively studying the scale parameters of each type of detection object in the data set and the differences in the receptive field scale parameters of different convolutional layers and the extracted features at the corresponding scales, and selectively selecting feature maps of different scales and resolutions that need to be fused in the model; constructing a dense bidirectional feature fusion module based on multi-scale features, designing two transmission connection blocks TCM and TCB with different structures according to the selected feature maps of different scales and resolutions, and multiple The interconnected transmission connection blocks TCM and multiple interconnected transmission connection blocks TCB respectively construct two mutually inverse transmission paths from top to bottom and from bottom to top, take the selected multi-layer feature map as input, and complete the top-down and bottom-up feature transmission and feature fusion respectively based on dense connection; design the anchor box size of the multi-layer feature map for multiple objects, use the sampling mapping relationship between the original map and different feature maps, and use the clustering algorithm to respectively count the scale classification of the rectangular annotation boxes of various water objects mapped on different feature maps, and determine the anchor box size on each layer of the selected feature map, where the distance in the clustering algorithm is represented by d=1-IoU(bboxes,anchor); construct a target detection module based on multi-scale features, take the output of the dense bidirectional feature fusion module as input, and perform category prediction and regression for multi-scale and multi-category targets; a model training module is used to iteratively train the DCBFFNet model based on the optimal anchor box size and the constructed target detection dataset; and a water multi-target detection module is used to realize the positioning and category prediction of various water objects based on the trained model, and visualize the detection results. The specific implementation of each module of the above system can be found in the above method embodiment and will not be described in detail.
[0094] Based on the same inventive concept, an embodiment of the present invention discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, the DCBFFNet-based urban river water multi-target detection method is implemented.
[0095] Based on the same inventive concept, an embodiment of the present invention discloses a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the DCBFFNet-based method for detecting multiple targets on water in urban rivers is implemented.
[0096] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A multi-target detection method for urban river water based on DCBFFNet. It is characterized in that The following steps are involved: (1) Extract images of water objects related to urban river management from the video stream, including bottle images, waterweed images, mixture images, and boat images. Use the restricted contrast adaptive histogram equalization method to reduce the noise of the images. Then use rectangular annotation boxes to semantically annotate the locations and category labels of various water objects to construct a target detection dataset. (2) Construct an object scale-sensitive densely connected bidirectional feature fusion deep learning model DCBFFNet, including: Construct a feature layer selection module for multi-object characteristics, study the scale parameters of each type of detection object in the data set and the receptive field scale parameters of different convolutional layers and the differences in the extracted features at the corresponding scales, and select the feature maps of different scales and resolutions that need to be fused in the model in a targeted manner; Construct a dense bidirectional feature fusion module based on multi-scale features, design two transmission connection blocks TCM and TCB with different structures according to the selected feature maps of different scales and resolutions, and construct two mutually inverse transmission paths from top to bottom and from bottom to top respectively with multiple interconnected transmission connection blocks TCM and multiple interconnected transmission connection blocks TCB. Take the selected multi-layer feature maps as input, and complete the top-down and bottom-up feature transmission and feature fusion respectively based on dense connection. Design the anchor frame size of multi-layer feature maps for multiple objects. Utilize the sampling mapping relationship between the original image and different feature maps, use the clustering algorithm to count the scale classification of the rectangular annotation boxes of various water objects mapped on different feature maps, and determine the anchor frame size on each layer of the selected feature map. The distance in the clustering algorithm is represented by d = 1-IoU (bboxes, anchor), where IoU (bboxes, anchor) represents the intersection-over-union ratio of the rectangular annotation box bboxes and the anchor box anchor. Construct a multi-scale feature-based object detection module, use the output of the dense bidirectional feature fusion module as input, and perform category prediction and regression on multi-scale and multi-category objects; (3) Iteratively train the DCBFFNet model based on the optimal anchor box size and the constructed object detection dataset; (4) Based on the trained model, the positioning and category prediction of various objects on the water are realized, and the detection results are visualized; The feature layer selection module is composed of seven groups of sequentially connected convolution structures, classification layers and regression layers. The convolution structure is used for feature extraction, and the classification layer and regression layer realize the binary classification and regression of "background-target". After feature difference analysis, the four-layer feature maps C3, C4, C5 and C6 output by the fourth to seventh convolution structures are selected as detection feature maps. In the dense bidirectional feature fusion module, the first, second and third transmission connection blocks TCM connected in sequence are used to fuse C6, C5, C4, C3 and the feature maps output by the third convolution structure in the feature layer selection module from top to bottom. The first transmission connection block TCM takes the feature maps after the convolution operation of C6 and C5 as input, and outputs them to the second transmission connection block TCM after fusion. The second transmission connection block TCM takes the feature map after the convolution operation of C4 and the output of the first transmission connection block TCM as input, and outputs them to the third transmission connection block TCM after fusion. The third transmission connection block TCM takes the feature map after the convolution operation of C3 and C2 and the output of the second transmission connection block TCM as The first transmission connection block TCB takes the output of the third transmission connection block TCM and the output of the second transmission connection block TCM as input, and the feature map obtained after fusion is used as the second scale feature map in the target detection module, and the other output is sent to the second transmission connection block TCB. The second transmission connection block TCB takes the output of the first transmission connection block TCB and the element-by-element fusion features of the three transmission connection blocks TCM outputs as input, and the feature map obtained after fusion is used as the third scale feature map in the target detection module, and the other output is sent to the third transmission connection block TCB. The third transmission connection block TCB has one input as the output of the second transmission connection block TCB, and the other input takes the element-by-element fusion features of the three transmission connection blocks TCM outputs and the C6 convolution operation output as input, and the feature map obtained after fusion is used as the fourth scale feature map in the target detection module; the target detection module performs detection, recognition and regression of multi-category objects based on the feature maps of the first scale to the fourth scale.
2. According to claim 1, a method for detecting multiple targets on water in urban rivers based on DCBFFNet, It is characterized in that The feature map selection step in step (2) includes: (a) Traversing the scales of the rectangular annotation boxes of various types of water objects in the entire detection data set, including bottles, water plants, boats, and mixtures, and visually presenting them in the form of dot distribution maps; (b) Calculate the receptive field size of different feature layers and compare it with the scale of the rectangular annotation box of the object obtained in (a). Select the feature layer whose receptive field matches the scale of the rectangular annotation box of the object. The receptive field formula is as follows: Among them, l k is the receptive field of the kth layer, l k-1 is the receptive field of the k-1th layer, f k is the size of the convolution kernel of the kth layer, s i is the step size of the i-th layer; (c) Visualize the multiple feature layers obtained in (b), compare the differences in features obtained by different convolutional layers, and select five feature layers that contain object detail information and high-level semantic information respectively.
3. According to claim 2, a method for detecting multiple targets on water in urban rivers based on DCBFFNet, It is characterized in that Only four layers of feature maps are used as detection feature maps. The layer with the largest scale and highest resolution in the selected feature map is only used to provide target features, and classification and regression are not performed on this layer.
4. According to claim 1, a method for detecting multiple targets on water in urban rivers based on DCBFFNet, It is characterized in that In the step (2), the dense connection between the multiple interconnected transmission connection blocks TCM and the multiple interconnected transmission connection blocks TCB adopts a selective connection method, and the TCM / TCB containing low / high-level features is successively connected with the TCM / TCB containing high / low-level semantic features to realize cross-layer interaction of features; The process of feature fusion between different layers is expressed as: X p =Φ p {F f {T x (X k )}} Among them, X p represents the result of feature fusion, X k represents the feature map to be fused, T k Indicates that the feature map is upsampled or downsampled to ensure the scale consistency of feature fusion; Φ f Indicates the feature fusion method; Φ p Represents the convolution processing of the fused features; when fusing features, the elements at corresponding positions on corresponding channels of the same scale are added.
5. According to claim 1, a method for detecting multiple targets on water in urban rivers based on DCBFFNet, It is characterized in that In the step (2), the optimized K-means clustering algorithm is used to respectively count the scale classifications of the rectangular annotation boxes of the four types of water objects mapped on different feature maps, and determine the size of the anchor boxes on the selected feature maps of each layer. The specific process is as follows: Under the feature maps of each layer, different K values are pre-selected, and the clustering effects under different K values are compared according to the maximum possible recall rate parameter, and the optimal K value of the layer is selected. Based on the clustering effects under the optimal K value obtained at each layer, anchor boxes with different numbers and aspect ratios are designed in different feature maps; based on the aspect ratio a of the anchor boxes obtained with the optimal K value, r , the model can calculate the specific width and height of the anchor box: in, m represents the number of feature maps; S max , S min Respectively represent the ratio of the anchor box to the image; S k Indicates the ratio of the feature layer anchor box to the image.
6. A multi-target detection system for urban rivers based on DCBFFNet. It is characterized in that include: The module for constructing the target detection dataset is used to extract images of water objects related to urban river management from the video stream, including bottle images, waterweed images, mixture images, and ship images, perform noise reduction on the images, and then use rectangular annotation boxes to semantically annotate the locations and category labels of various water objects to construct the target detection dataset. The deep learning model is a scale-sensitive densely connected bidirectional feature fusion deep learning model DCBFFNet. Its construction process includes: constructing a feature layer selection module for multi-object characteristics, studying the scale parameters of each type of detection object in the data set and the receptive field scale parameters of different convolutional layers and the differences in the extracted features at the corresponding scales, and selecting feature maps of different scales and resolutions that need to be fused in the model in a targeted manner; constructing a dense bidirectional feature fusion module based on multi-scale features, designing two different structures of transmission connection blocks TCM and TCB according to the selected feature maps of different scales and resolutions, and multiple interconnected transmission connection blocks TCM and multiple interconnected transmission connection blocks TCB respectively construct two mutually inverse transmission paths from top to bottom and from bottom to top to select multiple layers. The feature map is used as input, and the top-down and bottom-up feature transmission and feature fusion are completed based on dense connection. The anchor box size of the multi-layer feature map for multiple objects is designed. The sampling mapping relationship between the original map and different feature maps is used, and the clustering algorithm is used to count the scale classification of the rectangular annotation boxes of various water objects mapped on different feature maps, and the anchor box size on each layer of the selected feature map is determined. The distance in the clustering algorithm is represented by d=1-IoU(bboxes,anchor), where IoU(bboxes,anchor) represents the intersection-over-union ratio of the rectangular annotation box bboxes and the anchor box anchor. A target detection module based on multi-scale features is constructed, and the output of the dense bidirectional feature fusion module is used as input to perform category prediction and regression on multi-scale and multi-category targets. Model training module, used to iteratively train the DCBFFNet model based on the optimal anchor box size and the constructed object detection dataset; And the water multi-target detection module is used to realize the positioning and category prediction of various objects on the water based on the trained model, and visualize the detection results; In the DCBFFNet model, the feature layer selection module consists of seven groups of sequentially connected convolutional structures, classification layers, and regression layers. The convolutional structure is used for feature extraction, and the classification layer and regression layer realize the binary classification and regression of "background-target". After feature difference analysis, the four-layer feature maps C3, C4, C5, and C6 output by the fourth to seventh convolutional structures are selected as detection feature maps. In the dense bidirectional feature fusion module, the first, second, and third transmission connection blocks TCM connected in sequence are used to fuse C6, C5, C4, C3 and the feature maps output by the third convolutional structure in the feature layer selection module from top to bottom. The first transmission connection block TCM takes the feature maps after the convolution operation of C6 and C5 as input, and outputs them to the second transmission connection block TCM after fusion. The second transmission connection block TCM takes the feature map after the convolution operation of C4 and the output of the first transmission connection block TCM as input, and outputs them to the third transmission connection block TCM after fusion. The third transmission connection block TCM takes the feature map after the convolution operation of C3 and C2 and the feature map output by the second transmission connection block TCM. The TCM output is used as input, and the feature map obtained after fusion is used as the first scale feature map in the target detection module; the first transmission connection block TCB takes the third transmission connection block TCM output and the second transmission connection block TCM output as input, and the feature map obtained after fusion is used as the second scale feature map in the target detection module, and the other output is sent to the second transmission connection block TCB; the second transmission connection block TCB takes the first transmission connection block TCB output and the three transmission connection block TCM output element-by-element fusion features as input, and the feature map obtained after fusion is used as the third scale feature map in the target detection module, and the other output is sent to the third transmission connection block TCB; the third transmission connection block TCB has one input as the second transmission connection block TCB output, and the other input takes the three transmission connection block TCM outputs and the C6 convolution operation output element-by-element fusion features as input, and the feature map obtained after fusion is used as the fourth scale feature map in the target detection module; the target detection module performs detection, recognition and regression of multi-category objects based on the first to fourth scale feature maps.
7. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, It is characterized in that When the computer program is loaded into the processor, the DCBFFNet-based method for detecting multiple targets on water in urban rivers is implemented according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by the processor, the method for detecting multiple targets on water in urban rivers based on DCBFFNet according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Multi-scale underwater fish school detection method based on attention module
CN114170497A