Multi-scale adaptive target detection method, system and equipment

Multi-scale target features are extracted through multi-branch deep convolutional neural network and bidirectional fusion strategy, combined with regional suggestion network and non-maximum suppression algorithm, the problems of accuracy and efficiency in traditional methods in multi-scale target detection are solved, and higher detection accuracy and accuracy are achieved.

CN119992054AInactive Publication Date: 2025-05-13CHINA SHIPBUILDING RES INST (SEVENTH RES INST OF CHINA STATE SHIPBUILDING CORP)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510079878.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-18
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing object detection methods have challenges in dealing with multi-scale targets. Traditional convolutional neural networks are difficult to take into account the target characteristics of different scales at the same time, resulting in a decrease in detection accuracy, especially in complex scenarios, occlusion and overlap between targets increase detection difficulty.

Method used

Multi-branch deep convolutional neural network combined with a two-way fusion strategy is used to extract target features of different scales, and target candidate regions are generated through the regional suggestion network, overlapping candidate boxes are removed using a non-maximum suppression algorithm, and bounding box coordinates are finally restored through the size proportional relationship.

Benefits of technology

The ability to extract features of targets at different scales is improved, the detection accuracy of small-scale targets and the positioning accuracy of large-scale targets is enhanced, the interference of overlapping candidate boxes is reduced, and the accuracy of detection results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992054A_ABST
    Figure CN119992054A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale adaptive target detection method, system and device, and the method comprises the steps: receiving an input image, carrying out the normalization processing, adjusting the size, recording a proportional relation, extracting target features, carrying out the bidirectional fusion, obtaining a feature image, generating a target candidate region, and traversing the target candidate region. And finally, taking the reserved candidate frame as a target detection result, and restoring the target detection result to an original coordinate system to obtain target position and category information. The system comprises an image preprocessing module, a feature extraction network module, a feature fusion module, a target detection module and a positioning module. The device comprises a processor and a memory. Therefore, on one hand, through the multi-branch deep convolutional neural network and the bidirectional fusion strategy, the feature extraction capability of different-scale targets can be improved, the detection precision of small-scale targets and the positioning accuracy of large-scale targets can be enhanced, and on the other hand, the interference of overlapped candidate frames can be reduced by using a non-maximum suppression algorithm. And the accuracy of the detection result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a multi-scale adaptive target detection method, system and device. Background Art

[0002] As an important task in the field of computer vision, object detection has a wide range of applications in many fields, such as intelligent security, autonomous driving, industrial inspection, etc. Its main goal is to accurately identify the location (usually represented by a bounding box) and category of the target object in an image or video.

[0003] In the development history of target detection, early methods such as sliding window-based methods slide windows at different scales and positions and use manually designed features (such as HOG, SIFT, etc.) to classify the image area within the window to determine whether it contains the target object. However, this method has low computational efficiency and limited detection effect on complex scenes and multi-scale targets.

[0004] With the rise of deep learning technology, target detection has made significant breakthroughs. Target detection methods based on deep learning are mainly divided into two categories: one-stage detection methods and two-stage detection methods. Among them, one-stage detection methods (such as YOLO series, SSD, etc.) directly predict the category and position of the target on the feature map, which has the advantages of fast detection speed, but is relatively weak in detecting small targets and positioning accuracy. The two-stage detection method (such as Faster R-CNN, etc.) first generates candidate regions that may contain targets through the region proposal network (RPN), and then classifies and accurately locates these candidate regions. The detection accuracy is relatively high, but the detection speed is slow.

[0005] Despite this, existing target detection methods still face challenges when dealing with multi-scale targets. Since targets may appear at different scales in images, from large-scale distant objects to small-scale nearby objects, traditional convolutional neural networks often find it difficult to simultaneously take into account target features of different scales during feature extraction. For small-scale targets, a large amount of information may be lost in the feature map after multiple downsampling, resulting in a decrease in detection accuracy. For large-scale targets, the network may not be able to effectively learn their global features, affecting the accuracy of positioning. In addition, in complex scenes, there may be occlusions and overlaps between targets, which also increases the difficulty of target detection. Therefore, in order to improve the performance of target detection in multi-scale targets and complex scenes, there is an urgent need for a new target detection method that can adaptively process targets of different scales and improve detection accuracy and efficiency. Summary of the invention

[0006] The present invention aims to solve the technical problems in the above-mentioned technology at least to some extent.

[0007] To this end, the first aspect of the present invention discloses a multi-scale adaptive target detection method, comprising the following steps:

[0008] S1. Receive an input image, and perform normalization processing on the input image, mapping the pixel values ​​of the input image to a preset range;

[0009] S2. resizing the normalized input image and recording the size ratio between the input image and the resized input image;

[0010] S3. extracting target features of the resized input image using a multi-branch deep convolutional neural network;

[0011] S4. bidirectionally fusing the target features to obtain a feature image;

[0012] S5. Generate a target candidate region of the feature image using a region proposal network, and predict the existence probability of the target at each position and the position and size information of the target bounding box;

[0013] S6. For any two candidate frames in the target candidate region, use a non-maximum suppression algorithm to compare the intersection-and-union ratio and confidence between the two candidate frames. When the intersection-and-union ratio of the two candidate frames is greater than a preset intersection-and-union ratio threshold, remove the one with the smaller confidence among the two candidate frames, repeat this process, and the remaining candidate frames after traversing the target candidate region are taken as the target detection results;

[0014] S7. According to the size ratio relationship, the bounding box coordinates in the target detection result are restored to the original image coordinate system to obtain the final target position and category information.

[0015] In addition, the multi-scale adaptive target detection method disclosed in the present invention may also have the following additional technical features:

[0016] In one embodiment of the present invention, in step S1, the preset range is [0, 1] or [-1, 1].

[0017] In one embodiment of the present invention, in step S3, the multi-branch deep convolutional neural network includes:

[0018] A small-scale branch deep convolutional neural network for capturing large-scale target features in the resized input image using a large-stride convolutional layer;

[0019] A mesoscale branch deep convolutional neural network, for capturing mesoscale target features in the resized input image using a medium-stride convolutional layer;

[0020] A large-scale branched deep convolutional neural network is used to capture small-scale target features in the resized input image using a small-step convolutional layer.

[0021] In one embodiment of the present invention, in step S3, a residual connection structure is set in each branch deep convolutional neural network of the multi-branch deep convolutional neural network.

[0022] In one embodiment of the present invention, in step S4, the target features are bidirectionally fused, including:

[0023] Top-down fusion, which is used to start from the large-scale branch, gradually map the large-scale target features to the same size as the small-scale features through upsampling operations and perform weighted sum fusion;

[0024] Bottom-up fusion is used to start from the small-scale branch, gradually map the small-scale target features to the same size as the large-scale features through upsampling operations, and perform weighted sum fusion.

[0025] In one embodiment of the present invention, the large-scale target feature fusion weight in the top-down fusion is the ratio of the current training iteration number to the total training iteration number, and the small-scale target feature fusion weight in the bottom-up fusion is the ratio of the current training iteration number to the total training iteration number.

[0026] A second aspect of the present invention discloses a multi-scale adaptive target detection system, comprising:

[0027] An image preprocessing module, used for receiving an input image, normalizing the input image, adjusting the size of the input image after normalization, and recording the size ratio relationship;

[0028] A feature extraction network module, used to construct a multi-branch deep convolutional neural network and extract target features of the resized input image;

[0029] A feature fusion module, used for bidirectionally fusing the target features to obtain a feature image;

[0030] A target detection module is used to construct a region proposal network to generate a target candidate region of the feature image, remove candidate frames with excessive overlap in the target candidate region using a non-maximum suppression algorithm, and retain the candidate frame with the highest confidence as the final target detection result;

[0031] A positioning module is used to restore the bounding box coordinates in the target detection result to the original image coordinate system according to the size ratio relationship to obtain the final target position and category information.

[0032] A third aspect of the present invention discloses a multi-scale adaptive target detection device, comprising:

[0033] A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the multi-scale adaptive target detection method is implemented when the processor executes the computer program.

[0034] According to the multi-scale adaptive target detection method, system and device disclosed in the present invention, on the one hand, the ability to extract features of targets of different scales can be improved through a multi-branch deep convolutional neural network and a bidirectional fusion strategy, thereby enhancing the detection accuracy of small-scale targets and the positioning accuracy of large-scale targets; on the other hand, the non-maximum suppression algorithm can be used to reduce the interference of overlapping candidate boxes, thereby improving the accuracy of the detection results.

[0035] Additional contents and advantages of the present invention will be given in the following description or may be understood through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The technical solutions and beneficial effects of the present invention will become apparent and easily understood from the following contents in conjunction with the accompanying drawings, wherein:

[0037] Figure 1 is a flow chart of the multi-scale adaptive target detection method of the present invention;

[0038] Figure 2 It is a structural block diagram of the multi-scale adaptive target detection system of the present invention. DETAILED DESCRIPTION

[0039] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0040] The multi-scale adaptive target detection method, system and device disclosed in the present invention will be described below with reference to the accompanying drawings.

[0041] In the embodiments of the present invention, the multi-scale adaptive target detection method, system and device can be used in an intelligent security monitoring system, which needs to accurately detect people and objects of different distances and sizes. Small-scale pedestrians close to the camera (such as people walking near the bottom of the picture), medium-scale objects at a moderate distance from the camera (such as vehicles parked in the middle of the picture) and large-scale buildings or scene elements far from the camera (such as high-rise buildings in the distance) may appear in the monitoring picture at the same time, and there may be partial occlusion between these targets; Figure 1 and Figure 2 As shown, a multi-scale adaptive target detection method includes the following steps:

[0042] S1. Receive the input image I(x,y), where x and y represent the coordinate positions of the image pixels, and perform normalization on the input image I(x,y). Suppose the normalization function is N(·), then the normalized image I n (x, y) = N[I(x, y)], maps the pixel values ​​of the input image to a preset range, the preset range is [0, 1] or [-1, 1];

[0043] Specifically, the image preprocessing module receives the input image I(x, y) from the surveillance camera with a resolution of 1280×720 (W×H), and first performs normalization to map the pixel values ​​to the range of [0,1];

[0044] S2. Resize the normalized input image. Suppose the transformation function for resizing is T(·). The resized image I a (x,y)=T[I n (x,y)] to adapt it to the network input requirements and record the size ratio relationship between the input image and the resized input image, specifically, and Among them, X, Y are the width and height of the original image, X a ,Y a The width and height of the adjusted image are used for subsequent restoration of the target position;

[0045] The resized image size is 640×360 (W×H), and the size ratio relationship between the input image and the resized image is recorded. and

[0046] S3. Extract target features of the resized input image using a multi-branch deep convolutional neural network;

[0047] Multi-branch deep convolutional neural network, including:

[0048] Small-scale branch deep convolutional neural network L(·) for convolutional layers with large stride l s (Assuming the step size s1 = 4) captures the large-scale target features in the resized input image. Specifically, its convolution operation can be expressed as Among them, ω l (i, j) is the weight of the low-scale branch convolution kernel at position (i, j), k = 3 is the half-size of the convolution kernel, which is used to quickly capture large-scale target features in the image. Its receptive field is large and can cover a large range of image areas. Its output feature is F l (x,y);

[0049] Medium-scale branch deep convolutional neural network M(·) for convolutional layers with medium stride length m s (Assuming the step size s2 = 2) captures the mid-scale target features in the resized input image. Specifically, its convolution operation can be expressed as Among them, ω m (i, j) is the weight of the medium-scale branch convolution kernel at position (i, j), k = 3 is the half-size of the convolution kernel, which is used to quickly capture the medium-scale target features in the image. Its perception field is moderate and can cover a moderate range of image areas. Its output feature is F m (x,y);

[0050] A large-scale branched deep convolutional neural network H(·) is used to adopt a small-step convolutional layer h s (Assuming step size s2 = 1) captures the small-scale target features in the resized input image. Specifically, its convolution operation can be expressed as Among them, ω h (i, j) is the weight of the high-scale branch convolution kernel at position (i, j), k = 3 is the half-size of the convolution kernel, which is used to quickly capture small-scale target features in the image. Its perception field is small and can cover a smaller range of image areas. Its output feature is F h (x,y);

[0051] For each branch of the multi-branch deep convolutional neural network, a residual connection structure is set respectively. Let the residual block function be R(·), then the feature representation after residual connection is F ref (x,y)=R[F(x,y)+F skip (x,y)], where F(x,y) is the feature output of the convolution operation of the current layer, and F skip (x, y) is the feature of the skip connection to solve the gradient vanishing problem in deep network training and ensure that the network can effectively learn feature representations of different scales;

[0052] Assume that in a certain training iteration, the current training iteration number i = 100 and the total training iteration number I = 500;

[0053] S4. bidirectionally fuse the target features to obtain a feature image;

[0054] The target features are bidirectionally fused, including:

[0055] Top-down fusion, assuming that the high-scale feature is F h (x, y), the upsampling function is U(·), which is used to start from the large-scale branch, and gradually map the large-scale target features to the same size as the small-scale features through upsampling operations (assuming the upsampling multiple is 2) and perform weighted sum fusion. The top-down fused features

[0056] Bottom-up fusion, assuming low-scale features as F l (x, y), the downsampling function is D(·), which is used to start from the small-scale branch, gradually map the small-scale target features to the same size as the large-scale features through upsampling operations and perform weighted sum fusion. The bottom-up fused features

[0057] Thus, the fused feature image F is obtained fusion (x,y);

[0058] Large-scale target feature fusion weights in top-down fusion is the ratio of the current training iteration number i to the total training iteration number I, and the fusion weight of small-scale target features in bottom-up fusion is the ratio of the current training iteration number i to the total training iteration number I;

[0059] in, is the fusion weight of large-scale target features in top-down fusion, Fusion weights of small-scale target features in bottom-up fusion;

[0060] In addition, these weights can also be optimized through training to achieve adaptive adjustment of different scale features during fusion, so that the network can automatically adjust the contribution of each scale feature according to the actual scale distribution of the target in the input image;

[0061] S5. Use the region proposal network RPN to generate target candidate regions of the feature image and predict the existence probability of the target at each position and the position and size information of the target bounding box;

[0062] The region proposal network RPN is constructed by fusion (x, y) by sliding the window (assuming the window size is 3×3), assuming the window position is (x ω ,y ω ), predict the probability of the target existing at each position P(x ω ,y ω ) and the location and size information of the target bounding box (x min ,y min ,x max ,y max ), the prediction process can be achieved through a series of convolutional layers and fully connected layers, specifically, P(x ω ,y ω ),(x min ,y min ,x max ,ymax ) = RPN(F fusion (x, y)], where F fusion (x, y) is the fused feature map;

[0063] S6. For any two candidate bounding boxes in the target candidate region, using the non-maximum suppression algorithm NMS, compare the intersection over union IoU and the confidence p between the two candidate bounding boxes. Let the candidate bounding box set be where p i is the confidence of the i-th candidate bounding box. When the intersection over union IoU(B i and B j ) of the two candidate bounding boxes is greater than the preset intersection over union threshold T i , B j ), remove the one with the smaller confidence p IoU and p i and B j . Repeat this process, and the remaining candidate bounding boxes after traversing the target candidate region are used as the target detection results; i and p i

[0064] For targets such as vehicles and pedestrians in the picture, generate a series of candidate regions and related information. Assume that 5 candidate bounding boxes are generated Assume that the confidences of these 5 candidate bounding boxes are p1 = 0.6, p2 = 0.7, p3 = 0.5, p4 = 0.8, p5 = 0.4 respectively;

[0065] Let two candidate bounding boxes and Intersection over union where Area(·) represents the area of the region. Assume that the preset intersection over union threshold T IoU = 0.5. When IoU(B i , B j ) > T IoU , remove the one with the smaller confidence. For example, calculate IoU(B i , B j ) = 0.6. Since p1 = 0.6 < p2 = 0.7, remove B1. Repeat this process, and the remaining candidate bounding boxes after traversing the target candidate region are used as the target detection results;

[0066] S7. According to the size ratio relationship, restore the bounding box coordinates in the target detection results to the original image coordinate system to obtain the final target position and category information. Specifically, let the restoration coordinate function be C(·), and obtain the final target position and category information where (x, y, x e , y e) are the bounding box coordinates in the adjusted image coordinate system, and c is the target category;

[0067] Assuming that the coordinates in the adjusted image coordinate system are (100, 50, 200, 150), the final target position and category information The calculation method is, Where c is the target category (assuming that a car is detected, the category is marked as "car"), so as to accurately locate the position of vehicles and pedestrians in the original monitoring image and determine their categories;

[0068] A multi-scale adaptive target detection system 100, comprising:

[0069] The image preprocessing module 101 is used to receive an input image, normalize the input image, adjust the size of the normalized input image, and record the size ratio relationship;

[0070] A feature extraction network module 102 is used to construct a multi-branch deep convolutional neural network and extract target features of the resized input image;

[0071] The feature fusion module 103 is used to perform bidirectional fusion on target features to obtain a feature image;

[0072] The target detection module 104 is used to construct a region proposal network to generate a target candidate region of the feature image, use a non-maximum suppression algorithm to remove candidate boxes with too high overlap in the target candidate region, and retain the candidate box with the highest confidence as the final target detection result;

[0073] The positioning module 105 is used to restore the bounding box coordinates in the target detection result to the original image coordinate system according to the size ratio relationship to obtain the final target position and category information.

[0074] A multi-scale adaptive target detection device, comprising:

[0075] A processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a multi-scale adaptive target detection method is implemented.

[0076] In summary, the multi-scale adaptive target detection method, system and device disclosed in the present invention can, on the one hand, improve the ability to extract features of targets of different scales, enhance the detection accuracy of small-scale targets and the positioning accuracy of large-scale targets through a multi-branch deep convolutional neural network and a bidirectional fusion strategy; on the other hand, it can use a non-maximum suppression algorithm to reduce the interference of overlapping candidate boxes, thereby improving the accuracy of the detection results.

[0077] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A multi-scale adaptive target detection method, characterized in that: The following steps are involved: S1. Receive an input image, and perform normalization processing on the input image, mapping the pixel values ​​of the input image to a preset range; S2. resizing the normalized input image and recording the size ratio between the input image and the resized input image; S3. extracting target features of the resized input image using a multi-branch deep convolutional neural network; S4. bidirectionally fusing the target features to obtain a feature image; S5. Generate a target candidate region of the feature image using a region proposal network, and predict the existence probability of the target at each position and the position and size information of the target bounding box; S6. For any two candidate frames in the target candidate area, use a non-maximum suppression algorithm to compare the intersection-and-union ratio and confidence between the two candidate frames. When the intersection-and-union ratio of the two candidate frames is greater than a preset intersection-and-union ratio threshold, remove the one with the smaller confidence of the two candidate frames, repeat this process, and the remaining candidate frames after traversing the target candidate area are taken as the target detection results; S7. According to the size ratio relationship, the bounding box coordinates in the target detection result are restored to the original image coordinate system to obtain the final target position and category information.

2. The multi-scale adaptive target detection method according to claim 1, characterized in that: In step S1, the preset range is [0, 1] or [-1, 1].

3. The multi-scale adaptive target detection method according to claim 1, characterized in that: In step S3, the multi-branch deep convolutional neural network includes: A small-scale branch deep convolutional neural network for capturing large-scale target features in the resized input image using a large-stride convolutional layer; A mesoscale branch deep convolutional neural network, for capturing mesoscale target features in the resized input image using a medium-stride convolutional layer; A large-scale branched deep convolutional neural network is used to capture small-scale target features in the resized input image using a small-step convolutional layer.

4. The multi-scale adaptive target detection method according to claim 1, characterized in that: In step S3, a residual connection structure is set in each branch of the multi-branch deep convolutional neural network.

5. The multi-scale adaptive target detection method according to claim 1, characterized in that: In step S4, the target features are bidirectionally fused, including: Top-down fusion, which is used to start from the large-scale branch, gradually map the large-scale target features to the same size as the small-scale features through upsampling operations and perform weighted sum fusion; Bottom-up fusion is used to start from the small-scale branch, gradually map the small-scale target features to the same size as the large-scale features through upsampling operations, and perform weighted sum fusion.

6. The multi-scale adaptive target detection method according to claim 5, characterized in that: The fusion weight of the large-scale target features in the top-down fusion is the ratio of the current training iteration number to the total training iteration number, and the fusion weight of the small-scale target features in the bottom-up fusion is the ratio of the current training iteration number to the total training iteration number.

7. A multi-scale adaptive target detection system, characterized in that: include: An image preprocessing module, used for receiving an input image, normalizing the input image, adjusting the size of the input image after normalization, and recording the size ratio relationship; A feature extraction network module, used to construct a multi-branch deep convolutional neural network and extract target features of the resized input image; A feature fusion module, used for bidirectionally fusing the target features to obtain a feature image; A target detection module is used to construct a region proposal network to generate a target candidate region of the feature image, remove candidate frames with excessive overlap in the target candidate region using a non-maximum suppression algorithm, and retain the candidate frame with the highest confidence as the final target detection result; A positioning module is used to restore the bounding box coordinates in the target detection result to the original image coordinate system according to the size ratio relationship to obtain the final target position and category information.

8. A multi-scale adaptive target detection device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the multi-scale adaptive target detection method according to any one of claims 1 to 6 when executing the computer program.

Citation Information

Patent Citations

  • Remote sensing image multi-scale target detection and identification method based on lightweight network

    CN114170526A

  • Multi-scale flame image detection method, device, equipment and storage medium

    CN114639060A

  • Remote sensing image directional target detection method based on multi-feature aggregation and interaction

    CN114926747A

  • Remote sensing image target detection method

    CN116258953A