Shielding target detection method based on feature pyramid network

By introducing an occlusion object detection method based on feature pyramid network in AI vision algorithms, and using the weighted feature pyramid network for feature fusion, the problem of low recognition accuracy under occlusion conditions is solved, and higher object detection accuracy and better object recognition capabilities are achieved.

CN120047782APending Publication Date: 2025-05-27UNIT 32002 OF THE CHINESE PEOPLES LIBERATION ARMY

Patent Information

Application Number
CN202411978320.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing AI vision algorithm has low recognition accuracy under occlusion conditions and has less research work, so it has not formed a unified improvement framework.

Method used

A method of occlusion object detection based on feature pyramid network is proposed, multi-scale shallow features are extracted through backbone network, and feature fusion is used to eliminate semantic information and position information imbalance, and multi-scale deep features are obtained.

Benefits of technology

It improves the accuracy of object detection under occlusion conditions, reduces interference caused by occlusion, and enhances the ability to identify targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047782A_ABST
    Figure CN120047782A_ABST
Patent Text Reader

Abstract

The invention discloses an occluded target detection method based on a feature pyramid network, and belongs to the technical field of image processing. The method comprises the following steps: acquiring an image acquired by image acquisition equipment, and performing data preprocessing on the image; performing feature extraction on the preprocessed image through a backbone network to obtain multi-scale shallow layer features; fusing the multi-scale shallow-layer features by using a weight feature pyramid network to obtain multi-scale deep-layer features for eliminating imbalance of semantic information and position information; determining the target category and position information of each prediction frame according to the multi-scale deep features; and removing redundant detection results, and outputting a corresponding optimal detection result for each target. The method is wide in coverage and high in expansibility, and can be conveniently integrated into an existing target detection algorithm framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to an occluded target detection method based on a feature pyramid network. Background Art

[0002] With the development of deep learning, target detection technology has been widely applied in fields such as security, autonomous driving, and military. However, deep learning belongs to the category of statistical learning and highly depends on high-quality and large-scale data. In real scenarios, due to the existence of occlusion, target detection remains a challenging task. The human visual system enables humans to continue and infer through the contours present in the scene even when part of the information of an object is occluded or lost, so as to judge the attributes of the object. However, it is still difficult for a computer vision system based on deep learning to effectively detect occluded objects. The neural network model obtained through training cannot fully learn complex occlusion patterns, has weak generalization ability, and poor processing ability for scenes not present in the training set.

[0003] In view of the fact that targets will form multi-scale phenomena under occlusion conditions and the extracted features will be interfered, researchers have proposed various feature pyramid algorithms to improve the accuracy of target detection. Generally speaking, related research can be divided into three development stages:

[0004] (1) Using an image feature pyramid: The input image is scaled or enlarged to different sizes and features are extracted separately to form a multi-scale feature pyramid for detection;

[0005] (2) Using multi-level feature maps: Utilize the characteristics of multi-scale multi-level feature maps extracted by a deep convolutional neural network, and extract features of different scales from different layers of the network for detection;

[0006] (3) Using a feature pyramid: The feature pyramid adds three lines of bottom-up, top-down, and lateral connections on the basis of multi-level feature maps, and performs detection after fusing and reorganizing the multi-level feature maps.

[0007] First of all, each level of the image feature pyramid has strong semantic information, but the features of each level are carried out independently, consuming huge computing resources and seriously slowing down the detection process.

[0008] Secondly, the multi-level feature maps utilize the characteristics of multi-scale feature maps extracted by a deep convolutional neural network, and extract features of different scales from different levels of the network for detection, which can utilize both low-level features and high-level features. This method does not increase the computational amount additionally, but this method does not fully consider the internal connection between feature maps of different levels, resulting in the top-level features losing position information and the low-level features losing detail information.

[0009] Finally, in order to effectively utilize the feature maps of different scales generated by the backbone network in the deep convolutional neural network, the Feature Pyramid Network (FPN) uses three lines, namely bottom-up, top-down, and lateral connections, to fuse features at different levels. The bottom-up process is the forward process of feature extraction by the backbone network. The top-down process is to upsample the feature layers extracted by the backbone network. During the upsampling process, the feature maps of the same size generated bottom-up are fused through lateral connections. However, existing fusion methods treat all features Figure 1 equally and do not consider that the contributions of feature maps of different scales to the output features are usually not equal. SUMMARY OF THE INVENTION

[0010] In the actual occlusion environment, the recognition accuracy of existing AI vision algorithms will decrease significantly, and there is little research work on this problem, and no unified improvement framework has been formed. In response to this, starting from improving the recognition accuracy of AI vision algorithms in the case of occlusion, the present invention proposes an occlusion target detection scheme based on the Feature Pyramid Network. This scheme has a wide coverage and strong scalability and can be easily integrated into the existing object detection algorithm framework.

[0011] The first aspect of the present invention discloses an occlusion target detection method based on the Feature Pyramid Network, and the method includes:

[0012] Step S1: Obtain an image collected by an image acquisition device and perform data preprocessing on the image;

[0013] Step S2: Extract features from the preprocessed image through a backbone network to obtain multi-scale shallow features;

[0014] Step S3: Use a weighted feature pyramid network to fuse the multi-scale shallow features to obtain multi-scale deep features that eliminate the imbalance of semantic information and position information;

[0015] Step S4: Determine the target category and position information of each prediction box according to the multi-scale deep features;

[0016] Step S5: Remove redundant detection results and output the corresponding optimal detection result for each target.

[0017] According to the method of the first aspect of the present invention, in step S1, an optical instrument device is used as the acquisition device to collect an image, and the data preprocessing includes cropping and magnifying the image; wherein, the size of the preprocessed image is 3×H 0 ×W 0 , H 0 ×W 0 represents the image size, which are the height and width respectively, and the number of channels of the image is 3.

[0018] According to the method of the first aspect of the present invention, in step S2, the preprocessed image is input into a backbone network based on a convolutional neural network structure to extract multi-scale shallow features; wherein, multi-scale shallow features are obtained for each convolutional block using Resnet residuals.

[0019] According to the method of the first aspect of the present invention, in step S3, a weighted feature pyramid network is used to fuse the multi-scale shallow features, and the multi-scale shallow features are used as input features. The weight network in the weighted feature pyramid network learns the importance of each input feature, and multi-scale deep features are obtained through weighted fusion. The calculation process is as follows:

[0020]

[0021] w i = w[i - 1], i ∈ [1, 2]

[0022] M i-1 = w 1 · f 1×1 (L i-1 ) + w 2 · (F upsample (M i ))

[0023] where F Maxpooling represents the max pooling operation, F contact represents the concatenation operation, F upsample represents the upsampling operation, w represents the weights obtained by the weight network, and f represents the convolution operation;

[0024] The convolutional layer in the weight network processes the upsampled feature layer to deal with the aliasing effect, and a 1×1 convolutional layer is used to replace the 3×3 convolutional layer for eliminating the aliasing effect, thereby retaining the activation function. The calculation process of the weighted feature pyramid network is as follows:

[0025]

[0026] where P represents the deep features after retaining the activation function.

[0027] According to the method of the first aspect of the present invention, in step S4, the deep features after retaining the activation function are input into the detection network. The network layers in the detection network determine the confidence of the target in the prediction box according to the learned parameters, and output the target with a confidence exceeding the threshold as prediction information.

[0028] The second aspect of the present invention discloses an occlusion target detection system based on a feature pyramid network. The system includes a processing unit configured to perform:

[0029] Obtain an image collected by an image acquisition device and perform data preprocessing on the image;

[0030] Extract features from the preprocessed image through a backbone network to obtain multi-scale shallow features;

[0031] Use a weighted feature pyramid network to fuse the multi-scale shallow features to obtain multi-scale deep features that eliminate the imbalance between semantic information and location information;

[0032] Determine the target category and location information of each prediction box according to the multi-scale deep features;

[0033] Remove redundant detection results and output the corresponding optimal detection result for each target.

[0034] According to the system of the second aspect of the present invention, the processing unit is specifically configured to: use an optical instrument device as the acquisition device to collect images, and the data preprocessing includes cropping and enlarging the images; wherein, the size of the preprocessed image is 3×H 0 ×W 0 H 0 ×W 0 represents the image size, which are the height and width respectively, and the number of channels of the image is 3.

[0035] According to the system of the second aspect of the present invention, the processing unit is specifically configured to: input the preprocessed image into a backbone network based on a convolutional neural network structure to extract multi-scale shallow features; wherein, use Resnet residuals to obtain multi-scale shallow features for each convolutional block

[0036] According to the system of the second aspect of the present invention, the processing unit is specifically configured to: use a weighted feature pyramid network to fuse the multi-scale shallow features, and use the multi-scale shallow features as input features. The weight network in the weighted feature pyramid network learns the importance of each input feature, and obtains multi-scale deep features through weighted fusion The calculation process is:

[0037]

[0038] w i = w[i - 1], i ∈ [1, 2]

[0039] M i-1 = w 1 ·f1×1 (L i-1 ) + w 2 ·(F upsample (M i ))

[0040] wherein, F Maxpooling represents a max pooling operation, F contact represents a concatenation operation, F upsample represents an upsampling operation, w represents the weight obtained by the weight network, and f represents a convolution operation;

[0041] The convolutional layer in the weight network is used to process the upsampled feature layer to cope with the aliasing effect, and a 1×1 convolutional layer is used to replace the 3×3 convolutional layer for eliminating the aliasing effect, so as to retain the activation function. The calculation process of the weight feature pyramid network is as follows:

[0042]

[0043] wherein, P represents the deep feature after retaining the activation function.

[0044] According to the system of the second aspect of the present invention, the processing unit is specifically configured to: input the deep feature after retaining the activation function into the detection network, and the network layer in the detection network determines the confidence of the target in the prediction box according to the learned parameters, and outputs the target with the confidence exceeding the threshold as prediction information.

[0045] The third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements a method for detecting occluded targets based on a feature pyramid network according to the first aspect of the present disclosure.

[0046] The fourth aspect of the present invention discloses a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements a method for detecting occluded targets based on a feature pyramid network according to the first aspect of the present disclosure.

[0047] In summary, the present invention designs an algorithm for improving the recognition accuracy of AI vision algorithms under occlusion conditions. The key point of this algorithm lies in the designed weighted feature pyramid network. According to the actual occlusion scenario, the present invention adaptively optimizes the feature fusion process using the weighted feature pyramid network, and the recombined features can reduce the interference caused by occlusion and are more conducive to target recognition. The proposed weighted feature pyramid network in the present invention redesigned the feature fusion process of the feature pyramid network. Different from the existing feature pyramid networks, the weighted network can eliminate the imbalance between semantic information and position information of multi-scale shallow feature information, so that the finally obtained multi-scale deep feature information can better improve the target detection accuracy under occlusion conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 It is a schematic diagram of the occlusion target detection process;

[0050] Figure 2 It is a schematic diagram of the occlusion target detection network;

[0051] Figure 3 It is a schematic diagram of the weighted network. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0053] The first aspect of the present invention discloses an occlusion target detection method based on a feature pyramid network, and the method includes:

[0054] Step S1: Obtain the image collected by the image acquisition device and perform data preprocessing on the image;

[0055] Step S2: Extract features from the preprocessed image through the backbone network to obtain multi-scale shallow features;

[0056] Step S3: Use the weighted feature pyramid network to fuse multi-scale shallow features to obtain multi-scale deep features that eliminate the imbalance between semantic information and location information;

[0057] Step S4: Determine the target category and location information of each prediction box according to the multi-scale deep features;

[0058] Step S5: Remove redundant detection results and output the corresponding optimal detection result for each target.

[0059] According to the method of the first aspect of the present invention, in step S1, an optical instrument device is used as the acquisition device to acquire an image, and the data preprocessing includes cropping and enlarging the image; wherein, the size of the preprocessed image is 3×H 0 ×W 0 , H 0 ×W 0 represents the image size, which are the height and width respectively, and the number of channels of the image is 3.

[0060] According to the method of the first aspect of the present invention, in step S2, the preprocessed image is input into a backbone network based on a convolutional neural network structure to extract multi-scale shallow features; wherein, multi-scale shallow features are obtained for each convolutional block using Resnet residuals

[0061] According to the method of the first aspect of the present invention, in step S3, the weighted feature pyramid network is used to fuse the multi-scale shallow features, and the multi-scale shallow features are used as input features. The weight network in the weighted feature pyramid network learns the importance of each input feature, and multi-scale deep features are obtained through weighted fusion The calculation process is as follows:

[0062]

[0063] w i = w[i - 1], i ∈ [1, 2]

[0064] M i-1 = w 1 ·f 1×1 (L i-1 ) + w 2 ·(F upsample (M i ))

[0065] where, F Maxpooling represents the max pooling operation, F contact represents the concatenation operation, F upsample represents the upsampling operation, w represents the weight obtained by the weight network, and f represents the convolution operation;

[0066] The convolutional layer in the weight network processes the upsampled feature layer to address the aliasing effect, and uses a 1×1 convolutional layer to replace the 3×3 convolutional layer for eliminating the aliasing effect, thereby retaining the activation function. The calculation process of the weighted feature pyramid network is as follows:

[0067]

[0068] Among them, P represents the deep feature after retaining the activation function.

[0069] According to the method of the first aspect of the present invention, in step S4, the deep feature after retaining the activation function is input into the detection network. The network layer in the detection network determines the confidence of the target in the prediction box according to the learned parameters, and outputs the target with a confidence exceeding the threshold as the prediction information.

[0070] The second aspect of the present invention discloses an occlusion target detection system based on a feature pyramid network. The system includes a processing unit, and the processing unit is configured to execute:

[0071] Obtain the image collected by the image acquisition device and perform data preprocessing on the image;

[0072] Extract features from the preprocessed image through the backbone network to obtain multi-scale shallow features;

[0073] Use the weighted feature pyramid network to fuse the multi-scale shallow features to obtain multi-scale deep features that eliminate the imbalance between semantic information and location information;

[0074] Determine the target category and location information of each prediction box according to the multi-scale deep features;

[0075] Remove redundant detection results and output the corresponding optimal detection result for each target.

[0076] According to the system of the second aspect of the present invention, the processing unit is specifically configured to: use an optical instrument device as the acquisition device to collect images, and the data preprocessing includes cropping and magnifying the images; among them, the size of the preprocessed image is 3×H 0 ×W 0 , H 0 ×W 0 represents the image size, which are the height and width respectively, and the number of channels of the image is 3.

[0077] According to the system of the second aspect of the present invention, the processing unit is specifically configured to: input the preprocessed image into the backbone network based on the convolutional neural network structure to extract multi-scale shallow features; among them, use Resnet residuals to obtain multi-scale shallow features for each convolutional block

[0078] For the system according to the second aspect of the present invention, the processing unit is specifically configured to: fuse multi-scale shallow features by using a weighted feature pyramid network, and use the multi-scale shallow features as input features. The weight network in the weighted feature pyramid network learns the importance of each input feature, and obtains multi-scale deep features through weighted fusion The calculation process is as follows:

[0079]

[0080] w i = w[i - 1], i ∈ [1, 2]

[0081] M i-1 = w 1 · f 1×1 (L i-1 ) + w 2 · (F upsample (M i ))

[0082] where F Maxpooling represents the max pooling operation, F contact represents the concatenation operation, F upsample represents the upsampling operation, w represents the weights obtained by the weight network, and f represents the convolution operation;

[0083] The convolutional layer in the weight network is used to process the upsampled feature layer to cope with the aliasing effect, and a 1×1 convolutional layer is used to replace the 3×3 convolutional layer for eliminating the aliasing effect, so as to retain the activation function. The calculation process of the weighted feature pyramid network is as follows:

[0084]

[0085] where P represents the deep features after retaining the activation function.

[0086] For the system according to the second aspect of the present invention, the processing unit is specifically configured to: input the deep features after retaining the activation function into the detection network. The network layer in the detection network determines the confidence of the target in the prediction box according to the learned parameters, and outputs the target with a confidence exceeding the threshold as prediction information.

[0087] First Embodiment

[0088] The above system includes five modules: input, feature extraction, feature fusion, detection, and output.

[0089] Such as Figure 1As shown in the figure, the input module pre - processes the photos taken by image acquisition devices such as cameras and then sends them to the feature extraction module; the feature extraction module extracts multi - scale shallow features from the input pictures through the backbone network; the feature fusion module uses the designed Weighted Feature Pyramid Networks (FPN) to fuse the shallow features, obtains multi - scale deep features that eliminate the imbalance problem of semantic information and location information, and sends them to the detection module for object prediction; the detection module determines the object category and location information of each prediction box according to the input multi - scale deep features; the output module aims to retain one detection result for each object and remove other redundant detection results, and outputs the final detection results after annotation.

[0090] The second embodiment

[0091] Taking the improvement of the recognition accuracy of the AI vision algorithm under occlusion as the starting point, a system for verifying and improving the performance of the AI vision algorithm under occlusion is proposed, as Figure 2 shown.

[0092] Input module: The input module uses optical instrument equipment to collect image data, and pre - processes the collected images, such as cropping and magnifying, and then uses them as the input of the object detection algorithm.

[0093] Feature extraction module: Input the image data into the feature extraction module based on the convolutional neural network structure for feature extraction, and obtain the multi - scale shallow feature information of the input image through the backbone network.

[0094] Feature fusion module: Input the multi - scale shallow feature information into the feature fusion module, and use the Weighted Feature Pyramid Network to obtain the fused and re - organized multi - scale deep feature information.

[0095] Detection network module: Input the multi - scale deep feature information into the detection network module. The network layers in the detection network determine the confidence of the object in the prediction box according to the learned parameters, and output the object with a high confidence as the prediction information.

[0096] Output module: Post - process the prediction information to remove redundant results, so that each object only retains one prediction result. The prediction result after post - processing is used as the final detection result for annotation and then output.

[0097] The third embodiment

[0098] Step S1: Obtain image data. After pre - processing, the input size is 3×H 0 ×W 0 , where H 0 ×W 0is the size of the input image, and 3 is the number of channels of the input image.

[0099] Step S2: Input the image data into the backbone network based on the convolutional neural network structure for feature extraction. Each convolutional block of Resnet can obtain multi-scale shallow feature information.

[0100] Step S3: Use the multi-scale shallow feature information as the input. The weighted feature pyramid network uses the weight network to learn the importance of each input feature, and then performs weighted fusion to obtain multi-scale deep feature information. As shown in, the calculation process is as follows: Such as Figure 3 shown, the calculation process is:

[0101]

[0102] w i = w[i - 1], i ∈ [1, 2]

[0103] M i-1 = w 1 · f 1×1 (L i-1 ) + w 2 · (F upsample (M i ))

[0104] In the formula, F Maxpooling represents the max pooling operation, and F contact represents the concatenation operation.

[0105] Finally, considering that the convolutional layer in the weight network has completed the processing of the upsampled feature layer to cope with the aliasing effect, a 1×1 convolutional layer is used to replace the 3×3 convolutional layer in the original FPN that eliminates the aliasing effect. This approach reduces the number of parameters while retaining the activation function, ensuring the non-linear expression ability of the network. The entire calculation process of the weighted feature pyramid is:

[0106]

[0107] Step S4: Input the multi-scale deep feature information P into the detection network module. The network layer in the detection network determines the confidence of the target in the prediction box according to the learned parameters, and outputs the target with a high confidence as the prediction information.

[0108] Step S5: Send the prediction information into the output module. The output module uses a post-processing algorithm to remove redundant results, so that each target only retains one prediction result, and labels it as the final detection result for output.

[0109] A third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements a method for occluded object detection based on a feature pyramid network described in the first aspect of the present disclosure.

[0110] A fourth aspect of the present invention discloses a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements a method for occluded object detection based on a feature pyramid network described in the first aspect of the present disclosure.

[0111] In summary, the present invention designs an algorithm for improving the recognition accuracy of the AI vision algorithm under occlusion conditions. The key point of this algorithm lies in the designed weighted feature pyramid network. According to the actual occlusion scenario, the present invention uses the weighted feature pyramid network to adaptively optimize the feature fusion process. The recombined features can reduce the interference caused by occlusion and are more conducive to object recognition. The proposed weighted feature pyramid network in the present invention redesigned the feature fusion process of the feature pyramid network. Different from the existing feature pyramid network, the weighted network can eliminate the imbalance between the semantic information and position information of the multi-scale shallow feature information, so that the finally obtained multi-scale deep feature information can better improve the object detection accuracy under occlusion conditions.

[0112] Please note that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification. The above embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for detecting occluded objects based on a feature pyramid network, characterized in that: The method comprises: Step S1, obtaining an image captured by an image acquisition device, and performing data preprocessing on the image; Step S2, extracting features from the preprocessed image through a backbone network to obtain multi-scale shallow features; Step S3, using a weighted feature pyramid network to fuse multi-scale shallow features to obtain multi-scale deep features that eliminate the imbalance between semantic information and position information; Step S4: determining the target category and location information of each prediction box according to the multi-scale deep features; Step S5: remove redundant detection results and output the corresponding optimal detection result for each target.

2. The method for detecting occluded objects based on a feature pyramid network according to claim 1, characterized in that: In step S1, an optical instrument is used as an acquisition device to acquire an image, and the data preprocessing includes cropping and enlarging the image; wherein the size of the preprocessed image is 3×H0×W0, H0×W0 represents the image size, which are height and width respectively, and the number of channels of the image is 3.

3. The method for detecting occluded objects based on a feature pyramid network according to claim 2, characterized in that: In step S2, the preprocessed image is input into a backbone network based on a convolutional neural network structure to extract multi-scale shallow features; wherein, each convolution block obtains multi-scale shallow features using the Resnet residual 4. The method for detecting occluded objects based on a feature pyramid network according to claim 3, characterized in that: In step S3, the multi-scale shallow features are fused using the weighted feature pyramid network. As input features, the weight network in the weighted feature pyramid network learns the importance of each input feature and obtains multi-scale deep features through weighted fusion. The calculation process is: w i =w[i-1],i∈[1,2] M i-1 =w1·f 1×1 (L i-1 )+w2·(F upsample (M i )) Among them, F Maxpooling represents the maximum pooling operation, F contact Represents the splicing operation F upsample represents the upsampling operation, w represents the weight obtained by the weight network, and f represents the convolution operation; The convolution layer in the weight network completes the processing of the upsampled feature layer to deal with the aliasing effect, and the 1×1 convolution layer is used to replace the 3×3 convolution layer that eliminates the aliasing effect, thereby retaining the activation function. The calculation process of the weighted feature pyramid network is: Among them, P represents the deep features after retaining the activation function.

5. The method for detecting occluded objects based on a feature pyramid network according to claim 4, characterized in that: In step S4, the deep features after retaining the activation function are input into the detection network. The network layer in the detection network determines the confidence of the target in the prediction box according to the learned parameters, and outputs the target whose confidence exceeds the threshold as the prediction information.

6. An occluded target detection system based on a feature pyramid network, characterized in that: The system comprises a processing unit configured to perform: Acquire images acquired by an image acquisition device and perform data preprocessing on the images; The preprocessed image is subjected to feature extraction through the backbone network to obtain multi-scale shallow features; The weighted feature pyramid network is used to fuse multi-scale shallow features to obtain multi-scale deep features that eliminate the imbalance between semantic information and position information. Determine the target category and location information of each prediction box based on multi-scale deep features; Remove redundant detection results and output the corresponding optimal detection result for each target.

7. The occluded target detection system based on feature pyramid network according to claim 6, characterized in that: The processing unit is specifically configured as follows: An optical instrument is used as an acquisition device to acquire images, and the data preprocessing includes cropping and enlarging the image; wherein the size of the preprocessed image is 3×H0×W0, H0×W0 represents the image size, which is height and width respectively, and the number of channels of the image is 3; The preprocessed image is input into the backbone network based on the convolutional neural network structure to extract multi-scale shallow features; wherein, each convolution block obtains multi-scale shallow features using the Resnet residual 8. The occluded target detection system based on feature pyramid network according to claim 7, characterized in that: The processing unit is specifically configured as follows: The weighted feature pyramid network is used to fuse multi-scale shallow features. As input features, the weight network in the weighted feature pyramid network learns the importance of each input feature and obtains multi-scale deep features through weighted fusion. The calculation process is: w i =w[i-1],i∈[1,2] M i-1 =w1·f 1×1 (L i-1 )+w2·(F upsample (M i )) Among them, F Maxpooling represents the maximum pooling operation, F contact Represents the splicing operation F upsample represents the upsampling operation, w represents the weight obtained by the weight network, and f represents the convolution operation; The convolution layer in the weight network completes the processing of the upsampled feature layer to deal with the aliasing effect, and the 1×1 convolution layer is used to replace the 3×3 convolution layer that eliminates the aliasing effect, thereby retaining the activation function. The calculation process of the weighted feature pyramid network is: Among them, P represents the deep features after retaining the activation function; The deep features after retaining the activation function are input into the detection network. The network layer in the detection network determines the confidence of the target in the prediction box according to the learned parameters, and outputs the target whose confidence exceeds the threshold as the prediction information.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the occluded target detection method based on a feature pyramid network as described in any one of claims 1-5.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method for detecting occluded objects based on a feature pyramid network according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Target detection method and device based on improved feature pyramid network structure

    CN111985503A

  • Method and device for detecting forbidden articles in complex environment based on multi-scale feature fusion

    CN117765378A

  • Multi-view SAR target identification method based on multi-scale feature fusion of feature pyramid

    CN119107526A

Cited By

  • Shielding target high-precision identification method based on data enhancement

    CN121884059A