Target detection method and system based on adaptive feature extraction and multi-scale enhancement

By adopting adaptive feature extraction and multi-scale enhancement methods in object detection, the limitations of traditional technology in object detection are solved, and higher detection accuracy and object detection capabilities in complex scenarios are achieved.

CN120014354AActive Publication Date: 2025-05-16NANJING COMM INST OF TECH

Patent Information

Application Number
CN202510103122.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Traditional object detection technology has limitations when dealing with targets at different scales, including difficulty in adapting to target scale changes, inefficient fusion of features at different scales, and lack of flexibility in balancing the difficulty of target detection at different scales, resulting in low detection accuracy and robustness.

Method used

The target detection method based on adaptive feature extraction and multi-scale enhancement is adopted, and the features of different scales are extracted through the expanded convolution block of parallel structures are dynamically adjusted, and the feature contribution degree is further adjusted through the multi-scale enhanced feature pyramid network, and the crossover and ratio threshold is adjusted using the scale adaptive crossover and ratio loss function.

Benefits of technology

It improves the accuracy and effectiveness of target detection on targets of different scales, enhances the detection capabilities of targets in complex scenarios, and improves the overall detection performance and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014354A_ABST
    Figure CN120014354A_ABST
Patent Text Reader

Abstract

The invention relates to a target detection method and system based on adaptive feature extraction and multi-scale enhancement. The method comprises the following steps: carrying out standardization preprocessing on an original image to obtain a to-be-detected original image; inputting a to-be-detected original image into the expansion convolution block of the parallel structure, and capturing feature information of the to-be-detected original image under different scales; dynamically adjusting the contribution degree of each channel feature through a target normalization method to obtain an optimal feature vector; the method comprises the steps of obtaining an optimal feature vector, scaling the optimal feature vector to obtain a compressed target attention weight, calculating a gating signal for the target attention weight by adopting a channel attention gating method, and generating a fusion feature map; deep features are extracted through a multi-scale enhanced feature pyramid network; and outputting a target detection result through the classifier and the regression device. The accuracy and effectiveness of target detection on targets of different scales can be improved, and the target detection capability in a complex scene is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and target detection, and in particular to a target detection method and system based on adaptive feature extraction and multi-scale enhancement. Background Art

[0002] Object detection is a core task in the field of computer vision. Its purpose is to identify and locate the target object of interest in an image or video. Object detection not only needs to identify the object category in the image, but also needs to accurately mark their location. It can be applied in many fields such as smart security, smart robots, and smart transportation. In traditional object detection, the input image is usually extracted through a convolutional neural network (CNN) to obtain key information in the image; then the position of the target object is marked using a bounding box, which is a rectangular box that can accurately circle the scope of the target object; finally, the category of the target object is determined, and the output of the convolutional neural network is used to determine which category the object in the image belongs to.

[0003] However, in the field of target detection, traditional technologies have many limitations when dealing with targets of different scales. These include: it is difficult to effectively adapt to changes in target scale during feature extraction, resulting in insufficient detection capabilities for targets of extreme scales; the feature fusion method of different scales is not efficient enough and cannot fully integrate the advantages of features at different levels; the loss function lacks flexibility in balancing the difficulty of detecting targets of different scales, affecting detection accuracy and robustness.

[0004] Therefore, traditional target detection methods often have low accuracy and effectiveness in detecting targets of different scales, and have the problem of insufficient target detection capabilities in complex scenarios. Summary of the invention

[0005] Based on this, in order to solve the above technical problems, a target detection method and system based on adaptive feature extraction and multi-scale enhancement are provided, which can improve the accuracy and effectiveness of target detection on targets of different scales and enhance the detection capability of targets in complex scenes.

[0006] A target detection method based on adaptive feature extraction and multi-scale enhancement, the method comprising:

[0007] Obtaining an input original image, and performing standardization preprocessing on the original image to obtain an original image to be detected;

[0008] Inputting the original image to be detected into a dilated convolution block of a parallel structure, capturing feature information of the original image to be detected at different scales through convolution layers with different dilation rates in the dilated convolution block, extracting image features, and obtaining feature maps of each channel;

[0009] Determine a target normalization method based on the image features, dynamically adjust the contribution of each channel feature through the target normalization method to obtain an optimal feature vector; and scale the optimal feature vector to obtain a compressed target attention weight, calculate a gating signal for the target attention weight using a channel attention gating method, and generate a fused feature map according to the gating signal and the channel feature map;

[0010] The fused feature map is input into a multi-scale enhanced feature pyramid network, and the multi-scale enhanced feature pyramid network scales the upper and lower feature maps to match the scale of the middle layer based on the fused feature map, thereby realizing multi-scale feature fusion and extracting deep features;

[0011] The deep features are input into the classifier and regressor, and the IoU threshold is dynamically adjusted using a scale-adaptive IoU loss function to output the target detection result.

[0012] In one embodiment, obtaining an input original image and performing standardization preprocessing on the original image to obtain the original image to be detected includes:

[0013] Obtain an input original image, and obtain a pixel matrix of the original image;

[0014] According to a predetermined data normalization rule, the pixel matrix is ​​normalized to obtain an original image to be detected.

[0015] In one embodiment, the original image to be detected is input into a dilated convolution block of a parallel structure, and feature information of the original image to be detected at different scales is captured through convolution layers with different dilation rates in the dilated convolution block, and image features are extracted to obtain feature maps of each channel, including:

[0016] Inputting the original image to be detected into the dilated convolution block of the parallel structure in the form of a tensor;

[0017] Parallel feature extraction is performed through convolutional layers with different dilation rates in the dilated convolution block to obtain feature maps at different scales and receptive fields;

[0018] The feature maps are fused in the channel dimension by using feature splicing to obtain fused channel feature maps.

[0019] In one embodiment, a target normalization method is determined based on the image features, and the contribution of each channel feature is dynamically adjusted by the target normalization method to obtain an optimal feature vector, including:

[0020] Using various normalization methods to perform normalization processing on each of the channel feature maps in spatial dimensions to obtain a normalized processing result;

[0021] Using an activation function to select an optimal normalization method based on the normalization processing result as a target normalization method;

[0022] The target normalization method is used to dynamically adjust the contribution of each channel feature to obtain the optimal feature vector.

[0023] In one embodiment, scaling the optimal feature vector to obtain a compressed target attention weight includes:

[0024] Scaling the optimal feature vector by a set of learnable weight parameters and scaling it by a parameter-free normalization technique to obtain a compressed feature vector;

[0025] A specific scalar is introduced into the compressed feature vector to normalize the scale, and scale normalization is performed to obtain the compressed target attention weight.

[0026] In one embodiment, a channel attention gating method is used to calculate a gating signal for the target attention weight, and a fusion feature map is generated according to the gating signal and the channel feature map, including:

[0027] Using a channel attention gating method, applying an activation function to the target attention weight to calculate a gating signal;

[0028] Using the target attention weight to perform weighted processing on the channel feature map to obtain a weighted feature map;

[0029] Multiplying the gated signal by the weighted feature map element by element to obtain a multiplied result;

[0030] The multiplication results are summed along the first dimension to obtain a fused feature map.

[0031] In one embodiment, the fused feature map is input into a multi-scale enhanced feature pyramid network, and the multi-scale enhanced feature pyramid network scales the upper and lower feature maps to match the scale of the middle layer based on the fused feature map to achieve multi-scale feature fusion and extract deep features, including:

[0032] Inputting the fused feature map into a multi-scale enhanced feature pyramid network, and scaling the features in the fused feature map through a linear scaling layer in the multi-scale enhanced feature pyramid network;

[0033] Performing pixel-by-pixel addition operation on the scaled features and the original features and inputting the resultant features into a fusion layer in the multi-scale enhanced feature pyramid network;

[0034] The fusion layer performs convolution operation on the input features, extracts and implements multi-scale feature fusion, and obtains deep features of the fusion feature map.

[0035] In one embodiment, the method further comprises:

[0036] Determine a real frame and a predicted frame of the target in the original image to be detected, and calculate an intersection-over-union ratio between the real frame and the predicted frame;

[0037] Calculate the Euclidean distance between the center point of the real box and the predicted box, and calculate the diagonal length of the minimum closure area covering the real box and the predicted box;

[0038] A target size function is defined based on the diagonal length, and a scale-adaptive intersection-over-union loss function is obtained according to the intersection-over-union ratio, the Euclidean distance, the diagonal length, and the target size function.

[0039] In one embodiment, the deep features are input into a classifier and a regressor, a scale-adaptive IoU loss function is used to dynamically adjust the IoU threshold, and a target detection result is output, including:

[0040] The deep features are input into a classifier and a regressor, and the classifier predicts the target category according to the deep features to obtain the probability distribution of the target belonging to different categories;

[0041] The target position and size are predicted by the regressor to obtain position and size prediction results;

[0042] The probability distribution, position and size prediction results are output as target detection results.

[0043] A target detection system based on adaptive feature extraction and multi-scale enhancement, the system comprising:

[0044] The image processing module is used to obtain the input original image and perform standardization preprocessing on the original image to obtain the original image to be detected;

[0045] An adaptive feature extraction module, used for inputting the original image to be detected into a dilated convolution block of a parallel structure, capturing feature information of the original image to be detected at different scales through convolution layers with different dilation rates in the dilated convolution block, extracting image features, and obtaining feature maps of each channel;

[0046] A feature processing and fusion module, for determining a target normalization method based on the image features, dynamically adjusting the contribution of each channel feature through the target normalization method to obtain an optimal feature vector; scaling the optimal feature vector to obtain a compressed target attention weight, using a channel attention gating method to calculate a gating signal for the target attention weight, and generating a fused feature map based on the gating signal and the channel feature map;

[0047] A multi-scale feature enhancement module, used for inputting the fused feature map into a multi-scale enhanced feature pyramid network, and scaling the upper and lower feature maps to match the scale of the middle layer based on the fused feature map through the multi-scale enhanced feature pyramid network, so as to achieve multi-scale feature fusion and extract deep features;

[0048] The target detection module is used to input the deep features into the classifier and regressor, dynamically adjust the intersection-over-union ratio threshold using a scale-adaptive intersection-over-union ratio loss function, and output the target detection result.

[0049] The above-mentioned target detection method and system based on adaptive feature extraction and multi-scale enhancement use a parallel structured dilated convolution sequence to extract features from the image through a step-by-step expansion rate, effectively expanding the network receptive field, and can better capture contextual information in the image and improve the accuracy of large target detection; the adaptive normalization method extracts channel attention weights, which can dynamically adjust the importance of feature channels, enhance the sensitivity to key features, and help improve the overall detection effect; the multi-scale enhanced feature pyramid network is used to fully integrate feature maps of different scales, and can effectively detect targets of different scales; due to the use of a scale-adaptive intersection-over-union loss function, the loss contribution of targets of different sizes can be reasonably balanced, thereby improving the overall performance and adaptability of the target detection method. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 This is a diagram of an application environment of a target detection method based on adaptive feature extraction and multi-scale enhancement in one embodiment;

[0051] Figure 2 1 is a flow chart of a target detection method based on adaptive feature extraction and multi-scale enhancement in one embodiment;

[0052] Figure 3 A schematic diagram of a multi-scale enhanced feature pyramid network structure and a bidirectional fusion module in one embodiment;

[0053] Figure 4 A schematic diagram of a medium-scale adaptive intersection-over-union loss function in an embodiment;

[0054] Figure 5It is a structural block diagram of an object detection system based on adaptive feature extraction and multi-scale enhancement in one embodiment;

[0055] Figure 6 is a schematic diagram of the structure of an adaptive feature extraction module in one embodiment;

[0056] Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0058] The object detection method based on adaptive feature extraction and multi-scale enhancement provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Figure 1 As shown, the application environment includes a computer device 110. The computer device 110 can obtain an input original image and perform standardization preprocessing on the original image to obtain an original image to be detected; the computer device 110 can input the original image to be detected into a dilated convolution block of a parallel structure, capture feature information of the original image to be detected at different scales through convolution layers with different dilation rates in the dilated convolution block, extract image features, and obtain feature maps of each channel; the computer device 110 can determine a target normalization method based on image features, dynamically adjust the contribution of each channel feature through the target normalization method, and obtain an optimal feature vector; and scale the optimal feature vector to obtain The compressed target attention weight is used to calculate the gating signal for the target attention weight using the channel attention gating method, and a fused feature map is generated based on the gating signal and the channel feature map; the computer device 110 can input the fused feature map into the multi-scale enhanced feature pyramid network, and through the multi-scale enhanced feature pyramid network, based on the fused feature map, the upper and lower feature maps are scaled to match the scale of the middle layer, so as to achieve multi-scale feature fusion and extract deep features; the computer device 110 can input the deep features into the classifier and regressor, and dynamically adjust the intersection-and-union ratio threshold using the scale-adaptive intersection-and-union ratio loss function, and output the target detection result. Among them, the computer device 110 can be, but is not limited to, various personal computers, laptops, smart phones, robots and other devices.

[0059] In one embodiment, Figure 2 As shown, a target detection method based on adaptive feature extraction and multi-scale enhancement is provided, comprising the following steps:

[0060] Step 202: obtain an input original image, and perform standardization preprocessing on the original image to obtain an original image to be detected.

[0061] The computer device can obtain the input original image and then preprocess the original image. Specifically, in one embodiment, a target detection method based on adaptive feature extraction and multi-scale enhancement can also include a process of image preprocessing, which specifically includes: obtaining the input original image and obtaining the pixel matrix of the original image; normalizing the pixel matrix according to a predetermined data normalization rule to obtain the original image to be detected.

[0062] In this embodiment, the input raw image data preprocessing method may be to standardize the raw image data to ensure that the data input into the adaptive feature extraction has a consistent format and scale. The standardization operation for the raw image is based on a general data preprocessing method, including but not limited to pixel value normalization, size adjustment, etc.

[0063] For example, I represents the original image, C represents the number of color channels, H and W represent the height and width of the image respectively, then Here, for the convenience of calculation and expression, part of the data of a specific color channel in I can be shown in the following table:

[0064]

[0065]

[0066] In this embodiment, the computer device can read the original image data, obtain the pixel matrix and related data of the image, such as the image width W and height H, the number of channels C, etc.; and perform normalization on the pixel matrix according to a predetermined data normalization rule. For example, if the pixel value range of image I is [min val ,max val ], then by the formula x c =(I-min val ) / (max val -min val ) maps pixel values ​​to the interval [0,1]; according to the input requirements of the target detection network, the image is resized, and operations such as cropping and scaling can be used to unify the image to a specific width and height such as 256×256×3.

[0067] Step 204, input the original image to be detected into the dilated convolution block of the parallel structure, capture the feature information of the original image to be detected at different scales through the convolution layers with different dilation rates in the dilated convolution block, extract the image features, and obtain the feature maps of each channel.

[0068] In one embodiment, a target detection method based on adaptive feature extraction and multi-scale enhancement is provided, which may also include an adaptive feature extraction process, and the specific process includes: inputting the original image to be detected in the form of a tensor into a dilated convolution block of a parallel structure; performing parallel feature extraction through convolution layers with different dilation rates in the dilated convolution block to obtain feature maps at different scales and receptive fields; and fusing each feature map in the channel dimension by feature splicing to obtain fused feature maps of each channel.

[0069] Among them, the parallel structure of the dilated convolution block can extract features from the preprocessed image data through a step-by-step dilated convolution sequence. For example, the input image is Where H is the height, W is the width, and C is the number of channels; the dilated convolution is k×k, and the default K=3; the receptive field size actually covered by the convolution kernel with a dilation rate of d is RF=(k-1)·d+1; the receptive field size of multiple layers is RF n =RF n-1 +(k-1)·d; the minimum receptive field that can cover the target can be described as

[0070] The expansion rate of each branch is D = {d0 = 1, d1 = 2, ..., d n =n} is a sequence that is expanded step by step, and RF is calculated n Whether the minimum receptive field is met Requirement; select the minimum number of parallel branches N so that the total receptive field of all branches

[0071] The input image data is sent to the dilated convolution branch in sequence for feature extraction and splicing in the channel dimension. The feature map is Where N is the number of branches. Convolutional layers with different expansion rates can simultaneously capture the feature information of images at different scales, avoiding the problem of feature information loss or weakening caused by multiple convolution operations in traditional serial structures.

[0072] In this embodiment, the process of performing adaptive feature extraction may specifically include: Where H is the height, W is the width, and C is the number of channels. The format is adjusted to a tensor suitable for network input, such as The image data X is sequentially input into the dilated convolution block of the parallel structure in the form of tensors for adaptive feature extraction.

[0073] In the dilated convolution sequence of the parallel structure, the convolution operation is performed according to the set dilation rate and convolution kernel size. Taking the dilated convolution layer with a dilation rate of 2 as an example, for the input feature map, its convolution operation will skip one pixel in space for sampling, thereby expanding the receptive field; after the parallel feature extraction of the dilated convolution layer of N branches, N feature maps with different scales and receptive fields are obtained; then, these N feature maps are fused in the channel dimension by feature splicing; assuming that the number of channels of each feature map is C, the number of channels of the fused feature map will become N×C, and the size of the fused feature map will be The rich feature information extracted under different expansion rates is integrated to better represent various features in the image, especially when processing images containing objects of different scales, with stronger adaptability and feature expression capabilities. Then, the fused feature map U can be sent to the adaptive normalization method for further processing.

[0074] Specifically, in this embodiment, the input image is For the convenience of calculation, H = 5, W = 5, C = 3 are used for calculation here; the minimum receptive field of the covered target f(x) = min(RF total )=H+W-1=5+5-1=9; According to the formula RF total =k0+(k0-1)×(d1-1)+(d1+(k1-1)×(d2-1))+...+(k n-1 +{k n-1 -1)), the number of branches of the convolution layer is calculated, the receptive field size of the first branch is 3+(3-1)×(2-1)=5; the receptive field size of the second branch is 5+(3+(3-1)×(2-1))=9; the number of branches of the final dilated convolution N=2 and the dilation rate D={d0=1,d1=2} are obtained. In this embodiment, for the convolution layer with a dilation rate of d0=1, according to the convolution formula Output each pixel value of the feature map, where w mn is the convolution kernel weight, b c is the bias term, b c is the pixel value at the corresponding position of the input feature map. Assume that the convolution kernel weights are randomly initialized in the range of [-0.1, 0.1] and the bias term is initialized to 0; after the convolution layer, the output feature map size is C1×222×222, and the number of channels output by the convolution layer is C1=64; for the convolution layer with an expansion rate of d1=2, the calculation formula is However, since the expansion rate is 2, one pixel will be skipped for sampling in space, and the output feature map size is C2×220×220. The calculation method is the same for convolutional layers with different expansion rates.

[0075] Next, the N branch feature maps can be fused in the channel dimension by feature concatenation. The size of the fused feature map is C = [1, 2, ..., C], where C is the number of channels. The fused feature map will be sent to the channel attention gating for further processing.

[0076] Step 206, determine the target normalization method based on the image features, dynamically adjust the contribution of each channel feature through the target normalization method to obtain the optimal feature vector; and scale the optimal feature vector to obtain the compressed target attention weight, use the channel attention gating method to calculate the gating signal for the target attention weight, and generate a fusion feature map based on the gating signal and the channel feature map.

[0077] First, the computer device can perform normalization operations based on image features, use the normalization method in the standard library, calculate the feature vectors separately, and take the optimal value as the gating weight.

[0078] The adaptive normalization method is applied to learn the importance of channel feature maps of different branches, and dynamically adjust the contribution of each channel feature, thereby enhancing the sensitivity to key features. In this embodiment, a normalized standard library can be uniformly adopted, and the standard library includes but is not limited to global average pooling F GAP =GAP(U)=max 1≤i≤H,1≤j≤W U(i,j), global maximum pooling L1 norm L2 norm And other general methods.

[0079] Among them, global average pooling: For feature maps of different channels, global average pooling operation is performed in the spatial dimension to convert the two-dimensional feature map into a numerical value. The calculation formula is: Assume that some pixel values ​​of the first channel feature map are as follows Here we only show the calculation example of the 4 pixels in the upper left corner of the image. The actual calculation covers the entire H×W range. The calculation can be obtained This operation is performed on all channels in turn to obtain a feature vector with a dimension of [C, 1, 1].

[0080] Global maximum pooling: For feature maps of different channels, perform global maximum pooling operations to convert the two-dimensional feature map into a numerical value. The calculation formula is Taking the first channel as an example, we can calculate After performing this operation on all channels, the corresponding global maximum pooling result can be obtained, and finally a feature vector with a dimension of [C, 1, 1] is formed.

[0081] Calculate the L1 norm: The calculation formula is Taking the first channel as an example, the calculation process is: Finally, the L1 norm value of the channel is calculated; in the same way, the L1 norms of all channels are calculated in turn, and finally a feature vector with a dimension of [C, 1, 1] is formed.

[0082] Calculate the L2 norm: The calculation formula is where x c (i, j) represents the pixel value at position (i, j) in the cth channel; also taking the first channel as an example, calculate The L2 norm of all channels is calculated in sequence, and finally a feature vector with dimension [C, 1, 1] is formed.

[0083] Specifically, in one embodiment, a target detection method based on adaptive feature extraction and multi-scale enhancement is provided, which may also include a process for obtaining an optimal feature vector, and the specific process includes: using various normalization methods to perform normalization processing on each channel feature map in the spatial dimension to obtain a normalized processing result; using an activation function to select the best normalization method as the target normalization method based on the normalization processing result; using the target normalization method to dynamically adjust the contribution of each channel feature to obtain the optimal feature vector.

[0084] The computer device can generate the feature map of each channel To normalize the spatial dimension, global average pooling F can be used. GAP =GAP(U)=max 1≤i≤H,1≤j≤W U(i,j), global maximum pooling L1 norm L2 norm And other general methods.

[0085] Adaptively select the best normalization method dynamically according to the input data, that is, automatically learn and select the normalization method that best suits the current task and data during the training process, described as U out =ω1F GAP (U)+ω2F GMP (U)+ω3F L1 (U)+ω4F L2 (U), where the weight of the normalized standard library ω i satisfy And ω i ≥0; Use the activation function Softmax to adaptively select the best normalization method. The output range of Softmax is between [0,1] to obtain the best feature representation The feature vector of .

[0086] Apply the Softmax activation function and take the normalized value as input. Through Softmax(x)=[0.25,0.4,1.0,0.5477], the output weight is [0.1816,0.2123,0.3826,0.2345]; take the maximum weight as the optimal normalization method and get the feature vector

[0087] In one embodiment, a target detection method based on adaptive feature extraction and multi-scale enhancement is provided, which may also include a process of scaling features. The specific process includes: scaling the optimal feature vector through a set of learnable weight parameters, and scaling it using a parameter-free normalization technique to obtain a compressed feature vector; introducing a specific scalar into the compressed feature vector to normalize the scale, performing scale normalization, and obtaining a compressed target attention weight.

[0088] The obtained feature vector is passed through a set of learnable weight parameters τ=[τ1,...,τ i ,...,τ c ] is scaled, and its expression is z c =τ c ·s c , assuming that the learnable weight parameter τ is a vector whose dimension matches the dimension of the feature vector, and its value is randomly initialized, for example, τ = [0.5, 0.6, ..., 0.4], and the optimal feature vector is assumed to be s c =[1.2,1.3,...,1.1], then the scaled vector z c The elements of are calculated as z1 = 0.5 × 1.2 = 0.6, z2 = 0.6 × 1.3 = 0.78, and the scaled complete eigenvector z is calculated in sequence. c ; Then, the parameter-free standardization technique is used for feature compression, and the expression is Among them, ∈=10 -8 , with the scaled eigenvector z just calculated c For example, after normalization, the complete compressed feature vector is calculated in sequence In the process of performing scale normalization, a specific scalar is introduced Right c The scale is normalized to avoid the situation when the number of channels C is large. The scale of is too small to ensure the stability and effectiveness of the feature scale; finally, the optimal channel attention weight G = (g1,...,g c ).

[0089] In one embodiment, a target detection method based on adaptive feature extraction and multi-scale enhancement is provided, which may also include a feature fusion process, and the specific process includes: using a channel attention gating method to use an activation function on the target attention weight to calculate a gating signal; using the target attention weight to weight the channel feature map to obtain a weighted feature map; multiplying the gating signal and the weighted feature map element by element to obtain a multiplied result; summing the multiplied result along the first dimension to obtain a fused feature map.

[0090] Weights for N branches The activation function tanh is used to calculate the gate signal, the gate signal is multiplied element by element with the stacked feature map U, and the results are summed along the first dimension to obtain the final fused feature map, which can be expressed as Shape

[0091] Specifically, after normalization, the gated adaptive module The tanh activation function is used to obtain the gate weight information. The output range of tanh is between [-1,1], which helps to maintain numerical stability during training and allows features to be flexibly adjusted in the positive and negative directions to adapt to different feature distributions. Its expression is in, is the feature map after channel attention weighting, is the normalized scale vector, and ψ c They are gate weight and bias respectively, and the scale of each channel in the feature map is dynamically adjusted by learning these parameters. Assuming the gate weight Random Initialization The bias is also randomly initialized ψ c =0.1; the normalized eigenvector As an example, the first element of Calculate the result G of each channel after being processed by the gated adaptive module in turn C , forming a new feature vector, the dimension is still [N,C,1,1].

[0092] Residual connection and fusion, through the residual connection method, the processed features are fused with the original features. Taking the first channel as an example, assuming that the feature vector example element of the first channel of the original feature map is x c =[0.5,0.6,...,1.1], then the new feature vector element of the first channel after fusion is calculated as This operation is performed on each channel in turn to obtain a complete new fused feature map, completing the processing flow of the entire module on the input feature map. The fused feature map can be passed to the subsequent network module for further processing, such as participating in the next layer of convolution, pooling or classification operations.

[0093] In step 208, the fused feature map is input into a multi-scale enhanced feature pyramid network. The multi-scale enhanced feature pyramid network scales the upper and lower feature maps to match the scale of the middle layer based on the fused feature map, thereby achieving multi-scale feature fusion and extracting deep features.

[0094] like Figure 3 As shown, the multi-scale enhanced feature pyramid network uses a bidirectional fusion module to fuse the features extracted by the backbone network. The bidirectional fusion module includes a linear scaling layer and a fusion layer. Further, the linear scaling layer uses upsampling and downsampling to scale the features of different layers.

[0095] The weighted fusion feature map is sent to the multi-scale enhanced feature pyramid network. Through the bidirectional fusion module, the feature maps of different scales are combined to construct a feature pyramid that can cover multi-scale targets. Among them, the bidirectional fusion module includes a linear scaling layer and a fusion layer. The linear scaling layer includes but is not limited to linear interpolation, linear filtering and linear transformation; the fusion layer uses a 3×3 convolution kernel and a convolution operation with a step size of 1 to extract the deep features of the image.

[0096] Specifically, in one embodiment, a target detection method based on adaptive feature extraction and multi-scale enhancement is provided, which may also include a multi-scale feature fusion process, and the specific process includes: inputting the fused feature map into the multi-scale enhanced feature pyramid network, and scaling the features in the fused feature map through the linear scaling layer in the multi-scale enhanced feature pyramid network; performing pixel-by-pixel addition operation on the scaled features and the original features and inputting them into the fusion layer in the multi-scale enhanced feature pyramid network; performing convolution operation on the input features through the fusion layer, extracting and realizing multi-scale feature fusion, and obtaining deep features of the fused feature map.

[0097] The linear scaling layer starts from the lowest resolution feature layer and uses the upsampling operation F up =UpSampling(F high ,α), where α is the scaling factor, F high The upsampling operation starts from the feature layer F2 and gradually increases its resolution to the same as that of the adjacent feature layer F3. The upsampling method used is assumed to be bilinear interpolation, and the scaling rate α = 1.5; taking a channel in the feature layer as an example, the upsampled feature map is F up, the pixel value calculation at position (i, j) is based on the weighted average of the pixels around the corresponding position in the original feature map F1; for example, for F up (10,10)=w1×F3(4,4)+w2×F3(4,5)+w3×F3(5,4)+w4×F3(5,5), where w1, w2, w3, and w4 are weights determined according to the bilinear interpolation algorithm, and their sum is 1.

[0098] Then the low-level features are downsampled using a convolutional layer with a step size of 2. down = DownSampling(F low ,β), scaling rate β=0.75,F low is the low-level feature map, and the feature layer F with gradually decreasing resolution is obtained down Then the scaled features are combined with the intermediate layer features F orig Perform pixel-by-pixel addition operation F add =F orig +F up +F down , to enhance the expressiveness of features. Specifically, the downsampling operation process uses a convolution layer with a stride of 2 to downsample the low-level feature F4. The step size is used to control the resolution to be halved. After downsampling, a feature layer with a gradually reduced resolution is obtained. The goal is to make its resolution the same as that of the adjacent feature layer F3. For convolution downsampling with a stride of 2, taking one channel as an example, when the convolution kernel slides on the feature map, a sampling calculation is performed every 2 pixels. Assuming that the convolution kernel size is 3×3, the calculation formula for each pixel value of the output feature map is , where w mn is the convolution kernel weight, and F4(2i+m,2j+n) is the pixel value at the corresponding position of the input feature map; for example, for the output feature map

[0099] The above upsampled feature layer F up Perform pixel-by-pixel addition operation with the intermediate layer F3; similarly, perform pixel-by-pixel addition operation on the downsampled feature layer F down The pixel-by-pixel addition operation is performed on the intermediate layer F3 to obtain the fused feature layer F fusion , F fused =Conv(F add ,K 3×3 ).

[0100] The fusion layer uses a convolution layer with a convolution kernel size of 3×3, 64 input channels, and 128 output channels to further extract and fuse features through convolution to achieve deeper feature extraction. in is the convolution kernel weight corresponding to the output channel C, is the input feature map F fusion The pixel value at the corresponding position.

[0101] Multiple superposition of bidirectional fusion, that is, repeating the steps of upsampling fusion from low resolution to high resolution and downsampling fusion from high resolution to low resolution, pixel-by-pixel addition operation, fusion layer convolution operation, etc.; each round of fusion will make the information interaction between feature maps of different scales more sufficient, low-resolution feature maps obtain more detailed information, and high-resolution feature maps incorporate more semantic information. After multiple superposition of bidirectional fusion, a complete multi-scale feature pyramid can be formed. The pyramid contains feature maps of multiple scales, and each scale feature map integrates multiple rounds of fusion information from low resolution to high resolution and from high resolution to low resolution. It can fully express the features of targets of different scales in the image and provide rich and effective feature input for subsequent target detection classifiers and regressors.

[0102] Step 210, input the deep features into the classifier and regressor, use the scale-adaptive IoU loss function to dynamically adjust the IoU threshold, and output the target detection result.

[0103] In one embodiment, a target detection method based on adaptive feature extraction and multi-scale enhancement is provided, which may also include a process of obtaining a scale-adaptive intersection-over-union loss function, the specific process including: determining a real box and a predicted box of a target in an original image to be detected, and calculating the intersection-over-union ratio between the real box and the predicted box; calculating the Euclidean distance between the center points of the real box and the predicted box, and calculating the diagonal length of the minimum closure area covering the real box and the predicted box; defining a target size function based on the diagonal length, and obtaining a scale-adaptive intersection-over-union loss function according to the intersection-over-union ratio, Euclidean distance, diagonal length, and target size function.

[0104] In the scale-adaptive intersection-over-union loss function, scale weights are assigned according to the size of the target. In the loss calculation process, the scale weight α is used to balance the loss contribution of targets of different sizes. Small targets are given larger weights to optimize the detection accuracy of small targets, while maintaining the accuracy of large target detection, so that more attention can be paid to the detection effects of targets of different scales.

[0105] The schematic diagram of the scale-adaptive intersection-over-union loss function is as follows Figure 4 As shown, the real bounding box of the target is B g =[x g1 ,y g1 ,x g2 ,y g2 ], the predicted box is B p =[x p1 ,yp1 ,x p2 ,y p2 ], the intersection of the predicted box and the real box In this embodiment, it is assumed that in the target detection task scenario, the coordinate information of the real bounding box and the predicted box is as follows: the real bounding box B g : The coordinates of the upper left corner are (20,30), and the coordinates of the lower right corner are (60,80); prediction box B p :The coordinates of the upper left corner are (25,35) and the coordinates of the lower right corner are (55,75); then, calculate the true bounding box B g and prediction box B p The coordinates of the intersection part are determined by taking the maximum value of the two boxes on the horizontal and vertical coordinates as the upper left corner coordinate and the minimum value as the lower right corner coordinate; the upper left corner coordinate of the intersection box is (max(20,25),max(30,35))=(25,35), and the lower right corner coordinate is (min(60,55),min(80,75))=(55,75); the predicted box area (B p )=(55-25)×(75-35)=1200;real frame area erea(B g )=(60-20)×(80-30)=2000;intersection area(B p ∩B g )=(55-25)×(75-35)=1200;The area of ​​the union is area(B p ∪B g )=area(B p )+area(B g )-area(B p ∩B g )=1200+2000-1200=2000; then the intersection and union ratio Next, the Euclidean distance d between the predicted box and the center point of the real box is calculated = (B p ,B g ), the calculation formula is Where (x p ,y p ) is the coordinate of the center point of the prediction box, (x g ,y g ) is the coordinate of the center point of the real box. The coordinate of the center point of the predicted box Real frame center coordinates Substituting into the distance formula we get

[0106] Calculate the diagonal length c(B) of the minimum closure area covering the predicted box and the true box p ,B g), specifically including: calculating by obtaining the maximum difference between the upper left corner and lower right corner coordinates of the two boxes. For example, if the coordinates of the upper left corner of the predicted box are (x p1 ,y p1 ), the coordinate of the lower right corner is (x p2 ,y p2 ), the coordinate of the upper left corner of the real box is (x g1 ,y g1 ), the coordinate of the lower right corner is (x g2 ,y g2 ), then c(B p ,B g )=max(|x p2 -x g1 |,|x g2 -x p1 |)+max(|y p2 -y g1 |,|y g2 -y p1 |)=max(|55-20|,|60-25|)+max(|75-30|,|80-35|)=35+45=80.

[0107] Next, define the target size function s(B) to determine the size of the target based on the area of ​​the target box, etc. Where (x p1 ,y p1 ) and (x p2 ,y p2 ) is the diagonal coordinate of the prediction box; Where (x g1 ,y g1 ) and (x g2 ,y g2 ) is the diagonal coordinate of the real box. For the predicted box area For the real frame area

[0108] From this, we can calculate The scale-adaptive intersection-over-union loss function established is: The parameter μ ranges from -∞<α<1, and μ=-3 in this embodiment. Its function is to adjust the zoom level of small targets. The parameter ν is the rate at which larger objects are restored to the standard intersection-over-union ratio, and its range is 1<ν<∞. In this embodiment, ν=16. By adjusting the values ​​of μ and v, the network can adaptively adjust the detection accuracy of targets of different scales.

[0109] Through the above complete data flow and specific numerical calculations, the calculation process of each part of adaptive feature extraction, multi-scale enhanced feature pyramid network, and scale-adaptive intersection-over-union loss function and the overall application method are demonstrated. In the actual object detection network training process, such loss calculations are performed based on a large amount of sample frame (real frame and predicted frame) data, and the network parameters are adjusted according to the loss value through optimization algorithms (such as gradient descent, etc.), so that the network can adaptively adjust the detection accuracy of small and large targets to improve the overall object detection performance.

[0110] In one embodiment, a target detection method based on adaptive feature extraction and multi-scale enhancement is provided, which may also include a process of outputting target detection results. The specific process includes: inputting deep features into a classifier and a regressor, predicting target categories based on the deep features through the classifier, and obtaining a probability distribution of targets belonging to different categories; predicting target position and size through the regressor to obtain position and size prediction results; and outputting the probability distribution, position and size prediction results as target detection results.

[0111] That is, the fused features are processed by the classifier and regressor to output the final target detection results, including information such as the target category, location, and size.

[0112] The classifier uses structures such as fully connected layers and softmax functions to predict the category of the feature map and output the probability distribution of each target belonging to different categories. For example, for an object detection task containing N categories, the classifier outputs an N-dimensional probability vector, indicating the possibility of the target belonging to each category.

[0113] The regressor uses a convolutional layer or a fully connected layer to predict the target position and size of the feature map. For example, it predicts the center coordinates (x, y) of the target detection box and parameters such as width w and height h.

[0114] According to the output results of the classifier and regressor, the final target detection results are determined, including the target category, location, size and other information. For example, the category with the highest classification probability is selected as the target category, the location and size of the target in the image are determined according to the coordinates and size predicted by the regressor, and this information is organized into a target detection result list for output.

[0115] It should be understood that, although the various steps in the above-mentioned flow chart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above-mentioned flow chart may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0116] In one embodiment, Figure 5 As shown, a target detection system based on adaptive feature extraction and multi-scale enhancement is provided, including: an image processing module 510, an adaptive feature extraction module 520, a multi-scale feature enhancement module 530 and a target detection module 540, wherein:

[0117] The image processing module 510 is used to obtain an input original image and perform standardization preprocessing on the original image to obtain an original image to be detected;

[0118] The adaptive feature extraction module 520 is used to input the original image to be detected into the dilated convolution block of the parallel structure, capture the feature information of the original image to be detected at different scales through the convolution layers with different dilation rates in the dilated convolution block, extract the image features, and obtain the feature maps of each channel;

[0119] The adaptive feature extraction module 520 is further used to determine a target normalization method based on image features, dynamically adjust the contribution of each channel feature through the target normalization method, and obtain an optimal feature vector; and scale the optimal feature vector to obtain a compressed target attention weight, and use a channel attention gating method to calculate a gating signal for the target attention weight, and generate a fusion feature map according to the gating signal and the channel feature map;

[0120] The multi-scale feature enhancement module 530 is used to input the fused feature map into the multi-scale enhanced feature pyramid network, and scale the upper and lower feature maps to match the scale of the middle layer based on the fused feature map through the multi-scale enhanced feature pyramid network to achieve multi-scale feature fusion and extract deep features;

[0121] The target detection module 540 is used to input the deep features into the classifier and regressor, dynamically adjust the IoU threshold using the scale-adaptive IoU loss function, and output the target detection result.

[0122] In one embodiment, the image processing module 510 is further used to obtain an input original image and a pixel matrix of the original image; and perform normalization processing on the pixel matrix according to a predetermined data normalization rule to obtain the original image to be detected.

[0123] In one embodiment, the adaptive feature extraction module 520 is also used to input the original image to be detected in the form of a tensor into a dilated convolution block of a parallel structure; perform parallel feature extraction through convolution layers with different dilation rates in the dilated convolution block to obtain feature maps at different scales and receptive fields; and fuse the feature maps in the channel dimension using feature splicing to obtain fused feature maps of each channel.

[0124] In one embodiment, the adaptive feature extraction module 520 is also used to use various normalization methods to perform normalization processing on each channel feature map in the spatial dimension to obtain a normalized processing result; use an activation function to select the best normalization method as the target normalization method based on the normalization processing result; use the target normalization method to dynamically adjust the contribution of each channel feature to obtain the optimal feature vector.

[0125] In one embodiment, the adaptive feature extraction module 520 is also used to scale the optimal feature vector through a set of learnable weight parameters, and scale it using a parameter-free normalization technique to obtain a compressed feature vector; introduce a specific scalar into the compressed feature vector to normalize the scale, perform scale normalization, and obtain a compressed target attention weight.

[0126] In one embodiment, the adaptive feature extraction module 520 is also used to use an activation function on the target attention weight using a channel attention gating method to calculate a gating signal; use the target attention weight to weight the channel feature map to obtain a weighted feature map; multiply the gating signal and the weighted feature map element by element to obtain the multiplied result; sum the multiplied results along the first dimension to obtain a fused feature map.

[0127] In one embodiment, the structure of the adaptive feature extraction module is as follows: Figure 6As shown in the figure, it mainly includes a parallel structured dilated convolution module, adaptive normalization, activation function, and channel attention gating module, wherein: the parallel structured dilated convolution module extracts features from the preprocessed image data through a step-by-step dilated convolution sequence and concatenates them in the channel dimension; then the adaptive normalization method is applied to learn the importance of channel feature maps of different branches, and dynamically adjust the contribution of each channel feature, thereby enhancing the sensitivity of the model to key features; the channel attention gating module adopts a normalized standard library, and adaptively selects the best normalization method dynamically according to the input data; the activation function softmax is used to adaptively select the best normalization method to obtain the best feature vector.

[0128] In one embodiment, the multi-scale feature enhancement module 530 is also used to input the fused feature map into the multi-scale enhanced feature pyramid network, scale the features in the fused feature map through the linear scaling layer in the multi-scale enhanced feature pyramid network; perform pixel-by-pixel addition operation on the scaled features and the original features and input them into the fusion layer in the multi-scale enhanced feature pyramid network; perform convolution operation on the input features through the fusion layer to extract and realize multi-scale feature fusion, and obtain the deep features of the fused feature map.

[0129] In one embodiment, a target detection system based on adaptive feature extraction and multi-scale enhancement is provided, which may also include a scale-adaptive intersection-over-union loss function module, which is used to determine the real box and the predicted box of the target in the original image to be detected, and calculate the intersection-over-union ratio between the real box and the predicted box; calculate the Euclidean distance between the center points of the real box and the predicted box, and calculate the diagonal length of the minimum closure area covering the real box and the predicted box; define the target size function based on the diagonal length, and obtain the scale-adaptive intersection-over-union loss function according to the intersection-over-union ratio, Euclidean distance, diagonal length, and target size function.

[0130] In one embodiment, the target detection module 540 is also used to input deep features into a classifier and a regressor, and use the classifier to predict the target category based on the deep features to obtain a probability distribution of the target belonging to different categories; use the regressor to predict the target position and size to obtain position and size prediction results; and output the probability distribution, position and size prediction results as target detection results.

[0131] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 7As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a target detection method based on adaptive feature extraction and multi-scale enhancement is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0132] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0133] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of a target detection method based on adaptive feature extraction and multi-scale enhancement are implemented.

[0134] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of a target detection method based on adaptive feature extraction and multi-scale enhancement are implemented.

[0135] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0136] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0137] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A target detection method based on adaptive feature extraction and multi-scale enhancement, characterized in that: The method comprises: Obtaining an input original image, and performing standardization preprocessing on the original image to obtain an original image to be detected; Inputting the original image to be detected into a dilated convolution block of a parallel structure, capturing feature information of the original image to be detected at different scales through convolution layers with different dilation rates in the dilated convolution block, extracting image features, and obtaining feature maps of each channel; Determine a target normalization method based on the image features, dynamically adjust the contribution of each channel feature through the target normalization method to obtain an optimal feature vector; and scale the optimal feature vector to obtain a compressed target attention weight, calculate a gating signal for the target attention weight using a channel attention gating method, and generate a fused feature map according to the gating signal and the channel feature map; The fused feature map is input into a multi-scale enhanced feature pyramid network, and the multi-scale enhanced feature pyramid network scales the upper and lower feature maps to match the scale of the middle layer based on the fused feature map, thereby realizing multi-scale feature fusion and extracting deep features; The deep features are input into the classifier and regressor, and the IoU threshold is dynamically adjusted using a scale-adaptive IoU loss function to output the target detection result.

2. The target detection method based on adaptive feature extraction and multi-scale enhancement according to claim 1, characterized in that: Obtaining an input original image and performing standardization preprocessing on the original image to obtain an original image to be detected, including: Obtain an input original image, and obtain a pixel matrix of the original image; According to a predetermined data normalization rule, the pixel matrix is ​​normalized to obtain an original image to be detected.

3. The target detection method based on adaptive feature extraction and multi-scale enhancement according to claim 1, characterized in that: The original image to be detected is input into a dilated convolution block of a parallel structure, and feature information of the original image to be detected at different scales is captured through convolution layers with different dilation rates in the dilated convolution block, and image features are extracted to obtain feature maps of each channel, including: Inputting the original image to be detected into the dilated convolution block of the parallel structure in the form of a tensor; Parallel feature extraction is performed through convolutional layers with different dilation rates in the dilated convolution block to obtain feature maps at different scales and receptive fields; The feature maps are fused in the channel dimension by using feature splicing to obtain fused channel feature maps.

4. The target detection method based on adaptive feature extraction and multi-scale enhancement according to claim 1, characterized in that: Determining a target normalization method based on the image features, dynamically adjusting the contribution of each channel feature through the target normalization method, and obtaining an optimal feature vector, including: Using various normalization methods to perform normalization processing on each of the channel feature maps in spatial dimensions to obtain a normalized processing result; Using an activation function to select an optimal normalization method based on the normalization processing result as a target normalization method; The target normalization method is used to dynamically adjust the contribution of each channel feature to obtain the optimal feature vector.

5. The target detection method based on adaptive feature extraction and multi-scale enhancement according to claim 4, characterized in that: The optimal feature vector is scaled to obtain a compressed target attention weight, including: Scaling the optimal feature vector by a set of learnable weight parameters and scaling it by a parameter-free normalization technique to obtain a compressed feature vector; A specific scalar is introduced into the compressed feature vector to normalize the scale, and scale normalization is performed to obtain the compressed target attention weight.

6. The target detection method based on adaptive feature extraction and multi-scale enhancement according to claim 5, characterized in that: A channel attention gating method is used to calculate a gating signal for the target attention weight, and a fusion feature map is generated according to the gating signal and the channel feature map, including: Using a channel attention gating method, applying an activation function to the target attention weight to calculate a gating signal; Using the target attention weight to perform weighted processing on the channel feature map to obtain a weighted feature map; Multiplying the gated signal by the weighted feature map element by element to obtain a multiplied result; The multiplication results are summed along the first dimension to obtain a fused feature map.

7. The target detection method based on adaptive feature extraction and multi-scale enhancement according to claim 1, characterized in that: The fused feature map is input into a multi-scale enhanced feature pyramid network. The multi-scale enhanced feature pyramid network scales the upper and lower feature maps to match the scale of the middle layer based on the fused feature map, realizes multi-scale feature fusion, and extracts deep features, including: Inputting the fused feature map into a multi-scale enhanced feature pyramid network, and scaling the features in the fused feature map through a linear scaling layer in the multi-scale enhanced feature pyramid network; Performing pixel-by-pixel addition operation on the scaled features and the original features and inputting the resultant features into a fusion layer in the multi-scale enhanced feature pyramid network; The fusion layer performs convolution operation on the input features, extracts and implements multi-scale feature fusion, and obtains deep features of the fusion feature map.

8. The target detection method based on adaptive feature extraction and multi-scale enhancement according to claim 1, characterized in that: The method further comprises: Determine a real frame and a predicted frame of the target in the original image to be detected, and calculate an intersection-over-union ratio between the real frame and the predicted frame; Calculate the Euclidean distance between the center point of the real box and the predicted box, and calculate the diagonal length of the minimum closure area covering the real box and the predicted box; A target size function is defined based on the diagonal length, and a scale-adaptive intersection-over-union loss function is obtained according to the intersection-over-union ratio, the Euclidean distance, the diagonal length, and the target size function.

9. The target detection method based on adaptive feature extraction and multi-scale enhancement according to claim 1, characterized in that: The deep features are input into the classifier and regressor, and the scale-adaptive IoU loss function is used to dynamically adjust the IoU threshold to output the target detection results, including: The deep features are input into a classifier and a regressor, and the classifier predicts the target category according to the deep features to obtain the probability distribution of the target belonging to different categories; The target position and size are predicted by the regressor to obtain position and size prediction results; The probability distribution, position and size prediction results are output as target detection results.

10. An object detection system based on adaptive feature extraction and multi-scale enhancement, characterized in that: The system comprises: The image processing module is used to obtain the input original image and perform standardization preprocessing on the original image to obtain the original image to be detected; An adaptive feature extraction module, used for inputting the original image to be detected into a dilated convolution block of a parallel structure, capturing feature information of the original image to be detected at different scales through convolution layers with different dilation rates in the dilated convolution block, extracting image features, and obtaining feature maps of each channel; The adaptive feature extraction module is further used to determine a target normalization method based on the image features, dynamically adjust the contribution of each channel feature through the target normalization method, and obtain an optimal feature vector; and scale the optimal feature vector to obtain a compressed target attention weight, and use a channel attention gating method to calculate a gating signal for the target attention weight, and generate a fusion feature map according to the gating signal and the channel feature map; A multi-scale feature enhancement module, used for inputting the fused feature map into a multi-scale enhanced feature pyramid network, and scaling the upper and lower feature maps to match the scale of the middle layer based on the fused feature map through the multi-scale enhanced feature pyramid network, so as to achieve multi-scale feature fusion and extract deep features; The target detection module is used to input the deep features into the classifier and regressor, dynamically adjust the intersection-over-union ratio threshold using a scale-adaptive intersection-over-union ratio loss function, and output the target detection result.

Citation Information

Patent Citations

  • Video object detection and segmentation method based on space-time double-branch network

    CN110097568A

  • Multi-scale single-stage target detection method based on RetinaNet

    CN115861772A

  • Multi-scale nixie tube detection method based on improved YOLO adaptive attention-feature enhancement network

    CN117095155A

  • Target positioning and counting method based on point domain feature learning

    CN117765072A

  • Infrared weak and small target detection method based on feature fusion

    CN118279707A

Cited By

  • Cross-scale space target detection method and device

    CN120526129A