Optimization Implementation Method and System for Object Detection Based on Multi-Layer Fusion Edge Enhancement Neck Network

By designing a multi-layer fusion edge enhancement module in the neck network, the high resolution and edge information of shallow features are effectively utilized, and the problem of poor performance of the existing technology in small-scale and transparent object detection is solved, and the target detection performance is significantly improved.

CN115035394BActive Publication Date: 2025-06-13SUZHOU GOLDENPOINT INTERNET OF THINGS TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210798462.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-08
Publication Date
2025-06-13
Estimated Expiration
2042-07-08

AI Technical Summary

Technical Problem

Existing neck networks perform poorly when dealing with small-scale and transparent targets, failing to effectively utilize high-resolution and edge information features in multi-layer features.

Method used

A multi-layer fusion edge enhancement neck network is designed to fuse the high resolution and edge information of shallow features into the high-level features through reverse and forward fusion paths, and use the edge enhancement module of residual learning to improve the edge perception ability of image targets.

Benefits of technology

The target detection system's ability to identify small-size and transparent targets has been significantly improved, compared with the cutting-edge PANet network, the AP50 improvement on the garbage detection dataset, achieving a 2.69% improvement on the key categories of small-size and transparent targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035394B_ABST
    Figure CN115035394B_ABST
Patent Text Reader

Abstract

An optimized implementation method and system for object detection based on a multi-layer fusion edge enhancement neck network. Five layers of input features are extracted from an image by the backbone network. The second to fifth layer features are used for optimized feature fusion using a reverse fusion path and a forward fusion path to obtain three layers of output features, and then through an object classification and regression network, an optimized detection result is obtained. The present invention utilizes the difference in the ability of different layers of the neural network to extract high-level and low-level features, that is, the advantage of the shallow network in extracting low-level features such as image contours, edges, and textures, and the characteristic of high resolution of the shallow features, to fuse valuable shallow information into high-level features for subsequent detection and classification tasks. By utilizing high-resolution and edge information, the neck network can better handle the detection problems of small-size and transparent objects, improving the accuracy without being restricted by data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of image processing, specifically an optimized implementation method and system for small-scale and transparent target detection based on a multi-layer fusion edge-enhanced neck network. Background Art

[0002] The neck network (Bottleneck layer) is located between the backbone network and the detection head and is used to fuse and extract the features of the backbone network. The existing Feature Pyramid Network (FPN) continuously upsamples the deep features and adds them to the shallow features, so that the final data features contain multi-scale information and improve the subsequent detection performance. The Path Aggregation Network (PANet) adds a fusion path from shallow to deep on the basis of FPN. By continuously downsampling the shallow features and fusing them with the deep features, the performance is further improved. However, the current neck network does not consider well the differences in semantic information contained in different layers of the network, especially the lack of targeted extraction of high-resolution features and edge information features that are crucial for judging small targets and transparent targets. Therefore, the performance is not good when facing difficult targets.

[0003] Some existing improved neck networks only perform one-way splicing of network features from shallow to deep. For multi-layer features, especially the extraction degree of valuable information in shallow network features is insufficient; the particularity of each layer of features is ignored during the fusion process, and the advantages of large resolution and rich edge information of shallow network features cannot be explored targeted. There are limitations in the case of small-scale and transparent targets. Summary of the Invention

[0004] In view of the above deficiencies of the existing technology, the present invention proposes an optimized implementation method and system for target detection based on a multi-layer fusion edge-enhanced neck network. By using the difference in the ability of different layers of the neural network to extract high-level and low-level features, that is, the advantages of the shallow network in extracting low-level features such as image contours, edges, and textures, and the characteristics of high-resolution shallow features, the valuable information in the shallow layer is fused into the high-level features to perform subsequent detection and classification tasks. By utilizing high-resolution and edge information, the neck network can better handle the detection problems of small-size and transparent targets, improving the accuracy rate without being restricted by data.

[0005] The present invention is realized through the following technical solutions:

[0006] The present invention relates to an optimized implementation method for target detection based on a multi-layer fusion edge-enhanced neck network. Five layers of input features are extracted from the image by the backbone network, and the second to fifth layer features are used for optimized feature fusion using a reverse fusion path and a forward fusion path to obtain three layers of output features, and then through the target classification and regression network, an optimized detection result is obtained.

[0007] The second to fifth layer features mentioned above refer to the last four layers of features extracted by a commonly used backbone network for target detection feature extraction, where: the shallow features have a large resolution and a small number of channels. For each deeper layer, the number of channels of the features will double, while the resolution will be halved. Eventually, four features with different resolutions and numbers of channels will be obtained.

[0008] The optimized feature fusion mentioned above means: starting from the fifth to the second layer features, layer by layer, upsampling is performed to compress the number of channels of the deep layer features and expand the resolution, and then they are concatenated with the shallow network features until the second layer features are reached, realizing a reverse fusion path from deep to shallow; after fusing to the shallowest layer features, downsampling is performed to reduce the resolution, so as to concatenate with the third to fifth layer features in the reverse fusion path, realizing a forward fusion path from shallow to deep. Specifically, it includes: after respectively performing upsampling processing on the third to fifth layer features, they are respectively ① fused with the second layer features, enhanced by edge enhancement using the residual learning method, and then downsampled, and ② directly enhanced by edge enhancement using the residual learning method and fused with the downsampled features obtained from path ① to obtain intermediate features; the intermediate features are successively downsampled once, fused with the results after upsampling the fourth layer features, downsampled twice, and fused with the results after upsampling the fifth layer features to obtain the corresponding three-layer output features.

[0009] The upsampling mentioned above specifically includes:

[0010] ① For features with an input dimension of W×H×C, first, the number of channels is halved through a CSP layer, and the output feature dimension is: W×H×C / 2.

[0011] ② For the features with a dimension of W×H×C / 2 output by the CSP layer, further feature extraction is performed through a convolutional layer with a convolution kernel of 1 and a stride of 1, and the output feature 1 dimension is: W×H×C / 2.

[0012] ③ For the features with a dimension of W×H×C / 2 output by the convolutional layer, the features are upsampled by a factor of two through the nearest neighbor interpolation method, and the output feature dimension is: 2W×2H×C / 2.

[0013] The downsampling mentioned above specifically includes:

[0014] ① For features with an input dimension of W×H×C, first, the features are downsampled by a factor of two through the nearest neighbor interpolation method, and the output feature dimension is: W / 2×H / 2×C.

[0015] ② For the features with a dimension of W / 2×H / 2×C obtained by downsampling, the CSP layer is used to further extract features, and the output feature dimension is: W / 2×H / 2×C.

[0016] The edge enhancement using the residual learning method includes two branches, specifically: ① performing an identity mapping on the input features, and ② first performing a Gaussian blur operation on the input features through a Gaussian filter to obtain features with weakened irrelevant edge information, and then convolving the features through a Sobel operator to obtain the gradient map of the features. Based on the feature gradient map, calculate its phase map and amplitude map, splice them in the channel dimension, and finally input a 3×3 convolution for channel dimension compression to one dimension, which is used as edge enhancement information and superimposed on the input features of the identity mapping in branch ① to achieve edge enhancement of the features.

[0017] Technical effects

[0018] The present invention extracts multi-layer features in the backbone network, designs an innovative neck network, fuses the second to fifth layer features in a path aggregation manner, and applies the designed edge enhancement module to the features of the second and third layers. Through the Gaussian blur algorithm, the edge perception ability of the algorithm for image targets is enhanced, and the performance of object detection is improved. Compared with the prior art, the present invention effectively uses the large-resolution features and edge features of the front network features for object detection inference, significantly improves the recognition ability of the object detection system for transparent objects and small-scale targets in the image, and the overall AP50 of the present invention is improved by 0.8% compared with the advanced PANet network on the garbage detection dataset used, and a 2.69% AP50 improvement is achieved on the key categories of common small-scale and transparent targets. Description of the drawings

[0019] Figure 1 It is a schematic diagram of the system structure of the present invention;

[0020] Figure 2 It is a schematic diagram of the upsampling module;

[0021] Figure 3 It is a schematic diagram of the downsampling module;

[0022] Figure 4 It is a schematic diagram of the edge enhancement module;

[0023] Figure 5 It is an effect diagram of the embodiment of the present invention. Detailed implementation manners

[0024] Such as Figure 1As shown in the figure, this embodiment relates to an object detection system based on a multi-layer fusion edge enhancement neck network, including: a backbone neural network, a multi-layer fusion edge enhancement neck network, and a detection network, where: the backbone neural network extracts five layers of features with different scales and semantics in the input image data and outputs the second to fifth layer features to the multi-layer fusion edge enhancement neck network respectively. The multi-layer fusion edge enhancement neck network fully fuses the high-level semantic information with the low-level high-resolution and edge features and obtains three layers of output features. The detection network performs object box prediction based on the three layers of output features.

[0025] The backbone neural network mentioned above uses, but is not limited to, the DarkNet53 backbone network proposed by Redmin J, Farhadi A in Yolov3: An incremental improvement[J]. arXiv preprint arXiv:1804.02767, 2018 to extract image features. Specifically, the dark2, dark3, dark4, and dark5 layer features of DarkNet53 are used as the input to the neck network. Among them, dark2 is the shallowest layer. For each deeper layer, the number of channels of the features doubles, and the length and width become half of the original.

[0026] The different scales and semantics mentioned above refer to the outputs of the second to fifth convolutional layers of the backbone neural network. The deep networks (the fourth and fifth layers) contain high-level semantic information, which helps in object classification. The shallow networks (the second and third layers) have a larger resolution, which helps in the localization of small-scale objects; at the same time, due to the smaller receptive field of the shallow networks, they are better at focusing on the details and low-level edge and contour information in the image, which also helps in the localization of transparent objects.

[0027] In this embodiment, the neck network connects the output corresponding to the third to fifth layer features of the backbone network to the subsequent detection network. Since the useful information of the second layer features has been fused into the high-level features through the neck network, considering that the second layer features are relatively shallow and do not contain complete and effective information for subsequent detection, and will bring additional computational complexity, the neck network discards the second layer features when outputting.

[0028] The described multi-layer fusion edge enhancement neck network includes: three upsampling modules, two edge enhancement modules, and three downsampling modules, where: The first upsampling module processes the features of the fifth layer of the backbone neural network through a CSP layer, a convolutional layer, and an upsampling layer to obtain features with the number of channels compressed by half and the resolution doubled, and outputs them to the second upsampling module. The second upsampling module processes the information obtained by concatenating the output of the first upsampling module and the features of the fourth layer of the backbone neural network in the channel dimension through a CSP layer, a convolutional layer, and an upsampling layer to obtain features and outputs them to the third upsampling module; The third upsampling module processes the information obtained by concatenating the output of the second upsampling and the features of the third layer of the backbone neural network in the channel dimension through a CSP layer, a convolutional layer, and an upsampling layer to obtain features and outputs them to the first edge enhancement module; The first edge enhancement module performs feature edge enhancement processing on the result information obtained by concatenating the output of the third upsampling module and the output features of the second layer of the backbone neural network in the channel dimension to obtain the first edge enhancement feature; The second edge enhancement module performs feature edge enhancement processing on the output information of the convolutional layer in the third upsampling module to obtain the second edge enhancement feature; The first downsampling module performs downsampling layer and CSP layer processing on the first edge enhancement feature to obtain the output result of the first downsampling module; The second downsampling module performs CSP layer and convolutional layer processing on the information obtained by concatenating the second edge enhancement feature and the output of the first downsampling module in the channel dimension to obtain the output result of the second downsampling module; The third downsampling module performs CSP layer and convolutional layer processing on the features obtained by concatenating the output of the second downsampling module and the output of the convolutional layer of the first upsampling module in the channel dimension, and the third downsampling feature obtained is concatenated with the output of the convolutional layer of the first upsampling module in the channel dimension, and finally passes through the CSP layer processing to obtain the final feature; For the described final feature, the CSP layer outputs of the second and third downsampling modules will be used as the multi-scale outputs of the neck network.

[0029] The described detection network includes: an object classification and regression network and a non-maximum suppression module, where: The object classification and regression network respectively performs three-layer convolutional layer processing on the features obtained from the multi-scale outputs of the three neck networks to obtain the class and position results of all candidate boxes. The non-maximum suppression module performs overlapping box removal processing based on the candidate box information to obtain the final position and class results of object detection.

[0030] In the multi - level fusion neck network of this embodiment, first, the features of the dark5 layer are input. Through the up - sampling module, the number of channels of the features is halved, and the length and width are doubled at the same time, so as to splice with the features output by the dark4 layer in the channel dimension. And so on, this structure can continuously fuse the dark4 features into the dark3 features, and then fuse the dark3 features into the dark2 features. Thus, the feature fusion from deep to shallow is completed. Then, the fused dark2 features pass through the down - sampling module, the number of channels is doubled, and the length and width are halved, so as to splice with the penultimate layer output features with the same dimension in the previous dark3 up - sampling module. In this way, the shallow features can be continuously fused with the deep features again. Finally, the neck network will output the dark3, dark4 and dark5 features.

[0031] As Figure 2 shown, each up - sampling module includes: a CSP structure, a convolutional layer and an up - sampling layer, where: the CSP structure compresses the channels of the features with an input dimension of W×H×C, and the output feature dimension is W×H×C / 2. The convolutional layer does not change the feature dimension and further extracts features. The up - sampling layer uses nearest - neighbor interpolation to expand the feature resolution, and the output feature dimension is 2W×2H×C / 2.

[0032] As Figure 3 shown, each down - sampling module includes: a CSP structure and a down - sampling layer, where: for the features with an input dimension of W×H×C, the CSP layer does not change the feature dimension for feature extraction, and the down - sampling layer uses nearest - neighbor interpolation to reduce the feature resolution to obtain features with a dimension of W / 2×H / 2×C. Finally, it is spliced with the high - level features in the channel dimension, and the final output feature dimension is W / 2×H / 2×2C.

[0033] The CSP structure is based on the CSP module method of the residual network proposed by Wang C Y, Liao H Y M, Wu Y H, et al. in CSPNet: A new backbone that can enhance learning capability of CNN[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition workshops. 2020: 390 - 391.

[0034] As Figure 4As shown, each edge enhancement module includes: an identity mapping branch and a residual edge branch, where: the identity mapping branch performs an identity mapping on the input features, and the residual edge branch excludes noise edges through Gaussian blur, then extracts the phase map and amplitude map of the feature gradient through a Sobel operator, and obtains edge information through a 3×3 convolution and superimposes it on the original features to achieve edge enhancement of the features.

[0035] This embodiment relates to an object detection method for the above system, including network training in the offline stage and object detection in the online stage.

[0036] In this embodiment, the dataset used in the offline stage is implemented by adopting but not limited to the MS-COCO dataset. At the same time, data augmentation operations such as Mosaic, Mixup, and HSV are used to improve the generalization ability and performance of the network.

[0037] The network training mentioned above is implemented by adopting but not limited to the method proposed by Ge Z, Liu S, Wang F, et al. in "Yolox: Exceeding yolo series in 2021" [J]. arXiv preprint arXiv:2107.08430, 2021. This network training uses classification loss, confidence loss, and localization loss. Among them, the classification part is supervised by cross-entropy loss, the confidence is supervised by Smooth L1 loss, and the localization is supervised by IoU loss. The final total loss is the sum of the three loss values.

[0038] The working process of the system: First, use the backbone network to extract four layers of features from the input image, input them into the multi-layer fusion edge neck network, and output three layers of network features to the subsequent detection network. The detection network obtains the category, confidence, and location of the target box, and multiplies the category probability and the confidence as the final category probability. The category probability and the location will obtain the final prediction box through the non-maximum suppression algorithm.

[0039] The non-maximum suppression algorithm mentioned above adopts the method proposed by Neubeck A, Gool L V. in "Efficient Non-Maximum Suppression" [C] / / International Conference on Pattern Recognition. IEEE Computer Society, 2006: 850-855.

[0040] The object detection in the online stage specifically includes the following steps:

[0041] Step 1: Image preprocessing: Input an image of any size, scale the long side of the image to 640, scale the short side while maintaining the aspect ratio, place the scaled image in the upper left corner, and fill the remaining part with gray of RGB color value 114 until it reaches the size of 640×640.

[0042] Step 2: Feature extraction: Use DarkNet53 to extract the features of the dark2, dark3, dark4, and dark5 layers respectively. Among them, the feature dimension (W, H, C) of dark2 is (160, 160, 128), the feature dimension of dark3 is (80, 80, 256), the feature dimension of dark4 is (40, 40, 512), and the feature dimension of dark5 is (20, 20, 1024).

[0043] Step 3: Feature fusion of the neck network: Input the features of the dark2, dark3, dark4, and dark5 layers, and through fusion, output three layers of features with feature dimensions of (80, 80, 256), (40, 40, 512), and (20, 20, 1024).

[0044] Step 4: Regression and classification of target boxes. For the three layers of features, perform regression and classification of target boxes through the detection network respectively. The obtained position will be converted according to the ratio of the feature to the original image, and the prediction box results obtained from the three features will be spliced together.

[0045] Step 5: Output detection boxes. The prediction boxes will retain the box with the highest confidence in the area through the non-maximum suppression algorithm and delete the remaining overlapping redundant boxes. The IoU threshold for overlapping deletion set in this embodiment is 0.65, and the confidence threshold for each box is set to 0.3. Finally, the algorithm will output the coordinates of the upper left corner and the lower right corner of each prediction box, the category, and its confidence.

[0046] As Figure 5 shown, through specific actual experiments, under the specific environment settings of the constructed garbage target detection dataset, set the backbone neural network to DarkNet53, with its corresponding parameters width = 0.67, height = 0.75, the input image size is 640×640, the aspect ratio of the image remains unchanged, the extra area is filled with (114, 114, 114), the number of training rounds is 100, the batch size is 32, and the learning rate is set to 1.5625×10 -4, the SGD optimizer of the Pytorch framework is used to adjust the learning rate. Refer to the data augmentation method proposed by Ge Z, Liu S, Wang F, et al. in "Yolox: Exceeding yolo series in 2021" [J]. arXiv preprint arXiv:2107.08430, 2021. to improve the performance. The IoU threshold of the NMS algorithm is set to 0.65, and the object confidence threshold is 0.01. Running the above method with the above parameter configurations, the overall performance and the performance of each category in the garbage detection scenario are as follows:

[0047]

[0048]

[0049] In this embodiment, a comparative experiment between the multi-layer fusion edge-enhanced neck network and the advanced neck network:

[0050]

[0051] Comparison results between the present invention and the advanced PANet on key categories containing a large number of transparent objects and small-scale objects.

[0052]

[0053] Compared with the prior art, the multi-layer edge-enhanced neck network designed by the present invention can generally achieve further improvement on the basis of the advanced neck network structure under the garbage object detection dataset used in the present invention. Taking AP50 as an example, it is improved by 5.4% compared with the design without a neck network, 4.1% compared with FPN, and 0.8% compared with PANet. The improvement in object detection performance mainly benefits from the attention of the neck network of the present invention to large-resolution and edge features. Therefore, in the key categories of common small-scale and transparent objects, the algorithm of the present invention can achieve an average further improvement of 3.74% compared with the most advanced PANet. Therefore, the present invention can effectively improve the detection performance of the object detection algorithm on difficult objects such as transparent objects and small objects.

[0054] The above specific implementation can be locally adjusted by those skilled in the art in different ways without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation. All implementation solutions within its scope are subject to the present invention.

Claims

1. An optimized implementation method for object detection based on a multi-layer fusion edge-enhanced neck network, characterized in that, five layers of input features are extracted from the image by the backbone network, and the second to fifth layer features are subjected to optimized feature fusion using a reverse fusion path and a forward fusion path to obtain three layers of output features, and then through an object classification and regression network to obtain an optimized detection result; The second to fifth layer features refer to the last four layers of features extracted by a commonly used object detection feature extraction backbone network, where: the shallow features have a large resolution and a small number of channels, and the number of channels of each deeper layer of features will double, while the resolution will be halved; finally, four features with different resolutions and channel numbers will be obtained; The optimized feature fusion means: from the fifth to the second layer features, layer by layer, the channel number of the deep layer features is compressed by upsampling and the resolution is enlarged and spliced with the shallow network features until the second layer features are reached, realizing a reverse fusion path from deep to shallow; after fusing to the shallowest layer features, the resolution is reduced by downsampling, so as to be spliced with the third to fifth layer features in the reverse fusion path, realizing a forward fusion path from shallow to deep, specifically including: after respectively performing upsampling processing on the third to fifth layer features, respectively ① fusing with the second layer features and performing edge enhancement using the residual learning method and then performing downsampling processing, and ② directly performing edge enhancement using the residual learning method and fusing with the downsampled features obtained in path ① to obtain intermediate features; the intermediate features are successively subjected to one downsampling, fusing with the upsampled result of the fourth layer features, two downsamplings, and fusing with the upsampled result of the fifth layer features to obtain the corresponding three layers of output features; The edge enhancement using the residual learning method includes two branches, specifically: ① performing an identity mapping on the input features, and ② first performing a Gaussian blur operation on the input features through a Gaussian filter to obtain features that weaken irrelevant edge information, and then performing convolution on the features through a Sobel operator to obtain the gradient map of the features; based on the feature gradient map, calculate its phase map and amplitude map, splice them in the channel dimension, and finally input a 3×3 convolution for channel dimension compression to one dimension, and superimpose it on the input features of the identity mapping in branch ① as edge enhancement information to realize edge enhancement of the features.

2. The optimized implementation method for object detection based on a multi-layer fusion edge-enhanced neck network according to claim 1, characterized in that, the upsampling specifically includes: ① For features with an input dimension of W×H×C, first halve the number of channels through a CSP layer, and the output feature dimension is: W×H×C / 2; ② For features with a dimension of W×H×C / 2 output by the CSP layer, further feature extraction is performed through a convolutional layer with a convolution kernel of 1 and a stride of 1, and the output feature 1 dimension is: W×H×C / 2; ③ For features with a dimension of W×H×C / 2 output by the convolutional layer, the features are upsampled by a factor of two through the nearest neighbor interpolation method, and the output feature dimension is: 2W×2H×C / 2.

3. The method for optimizing and implementing object detection based on a multi-layer fusion edge-enhanced neck network according to claim 1, characterized in that the downsampling specifically includes: ① For features with an input dimension of W×H×C, first downsample the features by a factor of 2 using the nearest neighbor interpolation method, and the output feature dimension is: W / 2×H / 2×C; ② For the features with a dimension of W / 2×H / 2×C obtained by downsampling, use the CSP layer to further extract features, and the output feature dimension is: W / 2×H / 2×C.

4. An object detection system based on a multi-layer fusion edge-enhanced neck network for implementing the method according to any one of claims 1 to 3, characterized in that it includes: a backbone neural network, a multi-layer fusion edge-enhanced neck network, and a detection network, where: the backbone neural network extracts five-layer features with different scales and semantics in the input image data and outputs the second to fifth layer features to the multi-layer fusion edge-enhanced neck network respectively, the multi-layer fusion edge-enhanced neck network fully fuses the high-level semantic information with the low-level high-resolution and edge features and obtains three-layer output features, and the detection network performs object bounding box prediction based on the three-layer output features.

5. The object detection system based on a multi-layer fusion edge-enhanced neck network according to claim 4, characterized in that The described multi-layer fusion edge enhancement neck network includes: three upsampling modules, two edge enhancement modules, and three downsampling modules. Among them: The first upsampling module processes the features of the fifth layer of the backbone neural network through a CSP layer, a convolutional layer, and an upsampling layer, obtaining features with the number of channels compressed by half and the resolution doubled, and outputs them to the second upsampling module. The second upsampling module processes the information obtained by concatenating the output of the first upsampling module and the features of the fourth layer of the backbone neural network in the channel dimension through a CSP layer, a convolutional layer, and an upsampling layer, obtaining features and outputting them to the third upsampling module; The third upsampling module processes the information obtained by concatenating the output of the second upsampling and the features of the third layer of the backbone neural network in the channel dimension through a CSP layer, a convolutional layer, and an upsampling layer, obtaining features and outputting them to the first edge enhancement module; The first edge enhancement module performs feature edge enhancement processing on the result information obtained by concatenating the output of the third upsampling module and the output features of the second layer of the backbone neural network in the channel dimension, obtaining the first edge enhancement feature; The second edge enhancement module performs feature edge enhancement processing on the output information of the convolutional layer in the third upsampling module, obtaining the second edge enhancement feature; The first downsampling module performs downsampling layer and CSP layer processing on the first edge enhancement feature, obtaining the output result of the first downsampling module; The second downsampling module performs CSP layer and convolutional layer processing on the information obtained by concatenating the second edge enhancement feature and the output of the first downsampling module in the channel dimension, obtaining the output result of the second downsampling module; The third downsampling module performs CSP layer and convolutional layer processing on the features obtained by concatenating the output of the second downsampling module and the output of the convolutional layer of the first upsampling module in the channel dimension, and concatenates the obtained third downsampling features with the output of the convolutional layer of the first upsampling module in the channel dimension, and finally obtains the final features after CSP layer processing; For the described final features, the CSP layer outputs of the second and third downsampling modules will be used as the multi-scale outputs of the neck network.

6. The object detection system based on the multi-layer fusion edge enhancement neck network according to claim 4, characterized in that, the described detection network includes: an object classification and regression network and a non-maximum suppression module. Among them: The object classification and regression network respectively performs three-layer convolutional layer processing on the features obtained from the multi-scale outputs of the three neck networks, obtaining the class and position results of all candidate boxes. The non-maximum suppression module performs overlapping box removal processing based on the candidate box information, obtaining the final position and class results of object detection.

7. An object detection method based on the system according to any one of claims 4 to 6, characterized in that, it includes network training in the offline stage and object detection in the online stage. Among them, object detection in the online stage includes the following steps: The first step: Image preprocessing: Input an image of any size, scale the long side of the image to 640, scale the short side while maintaining the aspect ratio, place the scaled image in the upper left corner, and fill the remaining part with gray with an RGB color value of 114, filling it to a size of 640×640. Step 2: Feature extraction: Use DarkNet53 to extract the features of the dark2, dark3, dark4, and dark5 layers respectively; among them, the feature dimensions (W, H, C) of dark2 are (160, 160, 128), the feature dimensions of dark3 are (80, 80, 256), the feature dimensions of dark4 are (40, 40, 512), and the feature dimensions of dark5 are (20, 20, 1024); Step 3: Feature fusion of the neck network: Input the features of the dark2, dark3, dark4, and dark5 layers, and through fusion, output three layers of features with feature dimensions of (80, 80, 256), (40, 40, 512), and (20, 20, 1024); Step 4: Regression and classification of the target boxes; for the three layers of features, perform regression and classification of the target boxes through the detection network respectively; the positions obtained by regression will be converted according to the ratio of the features to the original image, and the prediction box results obtained from the three features will be concatenated together; Step 5: Output the detection boxes; the prediction boxes will use the non-maximum suppression algorithm to retain the box with the highest confidence in the region and delete the remaining overlapping redundant boxes; the IoU threshold for overlapping deletion set in this embodiment is 0.65, and at the same time, the confidence threshold for each box is set to 0.3; finally, the algorithm will output the coordinates of the upper left and lower right corners of each prediction box, the category, and its confidence.

Citation Information

Patent Citations

  • Construction site safety helmet identification method and system

    CN113989726A