Visual detection method for tiny target

Through adaptive region segmentation and deep learning super-resolution reconstruction technology, high-resolution target images are generated, and combined with feature pyramid network and attention mechanism, the problem of low micro-object detection accuracy is solved, achieving more stable and efficient micro-object detection.

CN120182701APending Publication Date: 2025-06-20CHONGQING NORMAL UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510268567.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In visual detection, small targets lack feature information in low-resolution images, resulting in low detection accuracy, and data enhancement strategies may destroy the target structure and affect detection performance.

Method used

Adaptive area segmentation algorithm is used to extract small target areas, use deep learning super-resolution reconstruction model to generate high-resolution target images, and enhance target features through feature pyramid networks and attention mechanisms. Finally, target detection models are used to predict target positions and categories.

Benefits of technology

It effectively improves the detection accuracy and stability of micro-targets, especially in the micro-target recognition scenarios under complex backgrounds, and improves the application value in security monitoring and medical image analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182701A_ABST
    Figure CN120182701A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of visual detection, and discloses a visual detection method for a tiny target, which comprises the following steps of: extracting a tiny target region from an original image, and determining a target boundary range according to target edge features and background contrast by adopting a self-adaptive region segmentation algorithm; introducing an attention mechanism at the output end of the feature pyramid network, calculating the attention weight of the tiny target area, and enhancing the significance of the target features in the feature map; and based on the enhanced feature map, adopting a target detection model, and predicting the position and category information of the tiny target through regression and classification branches. The method effectively improves the detection precision and stability of the tiny target, is especially suitable for a tiny target recognition scene under a complex background, and has an important application value in the fields of security monitoring, medical image analysis and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of visual detection, and particularly relates to a visual detection method for tiny targets. Background Art

[0002] During the visual detection process for tiny targets, there is a core technical problem: how to effectively improve the visibility and detection accuracy of tiny targets in the image acquisition stage and the processing stage. Specifically, in the image acquisition stage, limited by the device resolution and environmental factors, tiny targets occupy very few pixels in the original image, resulting in insufficient feature information and making it difficult to be accurately recognized by subsequent algorithms. Entering the image processing stage, directly performing target detection on low-resolution images is likely to cause missed detection or false detection of tiny targets. To solve this problem, upsampling or super-resolution reconstruction techniques are usually adopted, attempting to amplify the visual representation of tiny targets by increasing the pixel density. However, a simple upsampling operation may lead to image detail distortion or introduce artifacts, instead reducing the detection performance.

[0003] Furthermore, in the feature extraction stage, the feature pyramid network attempts to balance high-level semantic information and low-level detail information by fusing feature information at different levels. However, due to the low pixel ratio of tiny targets, their low-level features are often submerged by background noise, and the high-level features are insufficient in resolution, resulting in information loss. This feature fusion process is difficult to balance the detail retention and semantic enhancement of tiny targets. In addition, during the training process, data augmentation strategies for tiny targets, such as random scaling, cropping, rotation, and color transformation, although they can increase sample diversity, may also damage the original structure and context information of the targets, affecting the model's ability to localize and classify tiny targets.

[0004] The above problems are interrelated, ultimately leading to unstable detection results for tiny targets, especially in complex background or target-dense scenarios, where the detection performance significantly deteriorates. How to effectively enhance the feature representation of tiny targets without sacrificing image quality, while avoiding the destruction of target information by data augmentation, has become an urgent technical problem to be solved. Summary of the Invention

[0005] To solve the problems existing in the prior art, the present invention provides a visual detection method for tiny targets, which effectively improves the detection accuracy and stability of tiny targets, and is particularly suitable for the scenario of identifying tiny targets in complex backgrounds, having important application values in fields such as security monitoring and medical image analysis.

[0006] To achieve the above object, the present invention provides the following solution:

[0007] A visual detection method for tiny targets, the method comprising:

[0008] Extract the tiny target region from the original image. Using an adaptive region segmentation algorithm, determine the target boundary range according to the target edge features and background contrast;

[0009] According to the segmented tiny target region, utilize a super-resolution reconstruction model based on deep learning. Through multi-scale feature extraction and residual connection, reconstruct the high-resolution target image;

[0010] Input the reconstructed high-resolution target image into the Feature Pyramid Network. Adopt an adaptive feature fusion mechanism to dynamically adjust the weight ratio of low-level features and high-level features according to the target pixel proportion, and determine the final target image feature distribution;

[0011] At the output end of the Feature Pyramid Network, introduce an attention mechanism to calculate the attention weight of the tiny target region and enhance the saliency of the target features in the feature map;

[0012] Based on the enhanced feature map, adopt a target detection model. Through the regression and classification branches, predict the location and category information of the tiny target.

[0013] Preferably, extracting the tiny target region from the original image and using an adaptive region segmentation algorithm to determine the target boundary range according to the target edge features and background contrast includes:

[0014] Obtain the original image and convert it into a grayscale image, and perform preprocessing using Gaussian filtering;

[0015] Calculate the image edge features according to a preset gradient threshold to generate an edge feature map;

[0016] Calculate the background difference value through local contrast analysis to determine the potential target region;

[0017] Combine the edge feature map and the background difference value, and adopt an adaptive threshold segmentation algorithm to extract the candidate target region;

[0018] Perform morphological processing on the candidate target region;

[0019] Judge the target boundary based on region connectivity analysis to obtain the final target region;

[0020] According to the final target region, obtain an image result with target boundary markings.

[0021] Preferably, according to the segmented tiny target region, using a super-resolution reconstruction model based on deep learning to reconstruct the high-resolution target image through multi-scale feature extraction and residual connection includes:

[0022] Obtain the segmented tiny target region as input data;

[0023] Preprocess the input data using a super-resolution reconstruction model based on deep learning to extract multi-scale features;

[0024] Fuse the multi-scale features through residual connections to generate intermediate feature maps;

[0025] Perform super-resolution reconstruction based on the intermediate feature maps to obtain a high-resolution target image;

[0026] If there is distortion in the high-resolution target image, use a distortion correction algorithm for correction;

[0027] If there are artifacts in the high-resolution target image, use an artifact removal algorithm for removal;

[0028] Output the final high-resolution target image to complete the reconstruction process.

[0029] Preferably, the super-resolution reconstruction model based on deep learning includes three parts: feature extraction, non-linear mapping, and image reconstruction;

[0030] Among them, the feature extraction part is a 3×3 convolutional layer that transforms the input three-channel image into a multi-channel feature map for subsequent operations;

[0031] The non-linear mapping part is a combination and superposition of a receptive field fusion unit and a channel information fusion unit. At the same time, ResNet is introduced to implement residual learning, and a receptive field and channel information fusion block RCFB and a receptive field and channel information fusion group RCFG are constructed;

[0032] The image reconstruction part uses scaled convolution.

[0033] Preferably, input the reconstructed high-resolution target image into a feature pyramid network, adopt an adaptive feature fusion mechanism, and dynamically adjust the weight ratio of low-level features and high-level features according to the target pixel ratio to determine the final target image feature distribution, including:

[0034] Process the high-resolution target image using a feature pyramid network to extract multi-level feature maps;

[0035] Based on a preset pixel ratio threshold, determine whether the target pixel ratio exceeds the threshold;

[0036] If the target pixel ratio exceeds the threshold, increase the weight ratio of low-level features;

[0037] If the target pixel ratio does not exceed the threshold, increase the weight ratio of high-level features;

[0038] According to the dynamically adjusted weight ratio, use an adaptive feature fusion mechanism to fuse the multi-level feature maps;

[0039] Obtain the enhanced feature value of the target area through the fused feature map;

[0040] Determine the final target image feature distribution according to the enhanced feature value.

[0041] Preferably, at the output end of the feature pyramid network, introduce an attention mechanism to calculate the attention weight of the tiny target area and enhance the saliency of the target feature in the feature map, including:

[0042] For the feature map X with dimensions C×H×W, where C is the number of channels, and H and W represent the height and width of the input image, use pooling kernels of (H, 1) and (1, W) to perform one-dimensional feature encoding along the horizontal and vertical directions of the channels, and obtain feature maps Z in the height and width directions h and Z w , with sizes of C×H×1 and C×1×W;

[0043] Concatenate the obtained feature maps Z h and Z w Then use a 1×1 convolution and a non-linear activation function to obtain an intermediate feature map f∈R C / r×1×(H+W) , where r is the downsampling ratio, and then decompose f along the spatial dimension into f h ∈R C / r×H×1 and f w ∈R C / r×1×W Two separate tensors, and obtain f h ∈R C×H×1 and f w ∈R C×1×W through two 1×1 convolution transformations, and obtain two-direction attention weights g h and g w using the activation function Sigmoid(x), and finally multiply the attention weights g h and g w by the input feature as the output feature map.

[0044] Preferably, based on the enhanced feature map, adopt an object detection model to predict the position and category information of the tiny target through regression and classification branches, including:

[0045] Use the object detection model to process the enhanced feature map, predict the bounding box coordinate values of the tiny target through the regression branch, and predict the category probability values of the tiny target through the classification branch;

[0046] Determine the position information and category information of the tiny target according to the predicted bounding box coordinate values and category probability values.

[0047] Compared with the prior art, the beneficial effects of the present invention are:

[0048] The present invention discloses a visual detection method and system for tiny targets. For the tiny target regions extracted from the original image, the present invention uses an adaptive region segmentation algorithm to determine the target boundary, and then uses a super-resolution reconstruction model of deep learning to generate high-resolution target images. Next, the saliency of the target features is enhanced through a feature pyramid network and an attention mechanism, and a target detection model is used to predict the position and category information. If the detection confidence is insufficient, the present invention will perform secondary reconstruction and detection. During the training process, a local data augmentation strategy based on the target structure is adopted, and the parameters of each model are optimized iteratively. This method effectively improves the detection accuracy and stability of tiny targets, and is especially suitable for the recognition scenario of tiny targets under complex backgrounds, and has important application value in the fields of security monitoring, medical image analysis, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0050] Figure 1 Schematic flow chart of a visual detection method for tiny targets according to an embodiment of the present invention;

[0051] Figure 2 Schematic diagram of the receptive field fusion unit according to an embodiment of the present invention;

[0052] Figure 3 Schematic diagram of the channel information fusion unit according to an embodiment of the present invention;

[0053] Figure 4 Schematic diagram of the comparison between the deconvolution method and the scaled convolution method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0056] Embodiment 1

[0057] As Figure 1As shown in the figure, the present invention provides a visual detection method for tiny targets, and the method includes:

[0058] Extract the tiny target area from the original image, and use an adaptive region segmentation algorithm to determine the target boundary range according to the target edge feature and the background contrast.

[0059] According to the segmented tiny target area, use a super-resolution reconstruction model based on deep learning to reconstruct a high-resolution target image through multi-scale feature extraction and residual connection.

[0060] Input the reconstructed high-resolution target image into a feature pyramid network, and adopt an adaptive feature fusion mechanism to dynamically adjust the weight ratio of low-level features and high-level features according to the target pixel ratio, and determine the final target image feature distribution.

[0061] At the output end of the feature pyramid network, introduce an attention mechanism to calculate the attention weight of the tiny target area and enhance the saliency of the target feature in the feature map.

[0062] Based on the enhanced feature map, use a target detection model to predict the position and category information of the tiny target through regression and classification branches.

[0063] In this embodiment, extracting the tiny target area from the original image and using an adaptive region segmentation algorithm to determine the target boundary range according to the target edge feature and the background contrast includes:

[0064] Obtain the original image and convert it into a grayscale image, and perform preprocessing using Gaussian filtering. Calculate the image edge feature according to a preset gradient threshold to generate an edge feature map. Calculate the background difference value through local contrast analysis to determine the potential target area. Combine the edge feature map and the background difference value, and use an adaptive threshold segmentation algorithm to extract the candidate target area. Perform morphological processing on the candidate area to remove noise interference. Judge the target boundary based on region connectivity analysis to obtain the final target area. Output the image result with the target boundary marked.

[0065] In this embodiment, according to the segmented tiny target area, using a super-resolution reconstruction model based on deep learning to reconstruct a high-resolution target image through multi-scale feature extraction and residual connection includes:

[0066] Obtain the segmented tiny target regions as input data. Use a super-resolution reconstruction model of deep learning to preprocess the input data and extract multi-scale features. Fuse the multi-scale features through residual connections to generate an intermediate feature map. Perform super-resolution reconstruction based on the intermediate feature map to obtain a high-resolution target image. If there is distortion in the high-resolution target image, use a distortion correction algorithm for correction. If there are artifacts in the high-resolution target image, use an artifact removal algorithm for removal. Output the final high-resolution target image to complete the reconstruction process.

[0067] Specifically, the super-resolution reconstruction model of deep learning includes three parts: feature extraction, non-linear mapping, and image reconstruction;

[0068] Among them, the feature extraction part is a 3×3 convolutional layer that converts the input three-channel image into a multi-channel feature map for subsequent operations;

[0069] The non-linear mapping part is a combined stack of receptive field fusion units and channel information fusion units. At the same time, ResNet is introduced to implement residual learning, constructing a receptive field and channel information fusion block RCFB and a receptive field and channel information fusion group RCFG;

[0070] The image reconstruction part uses scaled convolution.

[0071] Among them, Figure 2 is a schematic diagram of the receptive field fusion unit. The input feature map is respectively convolved with three convolutional kernels of 3×3, 5×5, and 7×7 to obtain three three-dimensional feature maps with different receptive fields, named Feature1, Feature2, and Feature3 respectively. Subsequently, the pixel values of these three feature maps are added and input into the channel information fusion unit (including dimensionality reduction, dimensionality increase, and non-linear mapping) to obtain their respective channel information weights, named weight a, weight b, and weight c respectively. Then Feature1, Feature2, and Feature3 are multiplied by weight a, weight b, and weight c respectively to obtain the feature maps Feature1', Feature2', and Feature3' after channel weight recalibration, and then added and input into the next module. The final step to achieve receptive field fusion is Figure 2The right half of it. After weight guidance, the feature maps Feature1', Feature2' and Feature3' each retain their own receptive fields. Subsequently, the three different receptive fields are fused by adding pixel values, enabling the information between feature maps obtained through convolution operations of different scales to complement each other, thus achieving a more comprehensive feature extraction and feature mapping effect. This method enables the network to adaptively select different-sized convolutional kernels, that is, different receptive fields. Small convolutional kernels and small receptive fields are used in regions with concentrated information distribution density, while large convolutional kernels and large receptive fields are used in regions with dispersed information distribution density, thereby enhancing the network's adaptability and robustness.

[0072] As Figure 3 shown, the image super-resolution reconstruction algorithm based on channel information fusion mainly includes the following four steps: (1) The feature map U output by multiple convolutional layers undergoes average pooling operation, and the spatial information of each channel is compressed into a feature value, generating a channel information feature map s of size 1x1xC. s carries the original channel information of the three-dimensional feature map U; (2) s undergoes dimensionality reduction through the fully connected layer fc0 to generate a feature map z of size . The dimensionality reduction mainly has two functions. First, it reduces the parameters of the fully connected layer and reduces the computational amount; second, it reduces the function complexity, reduces overfitting, and enhances the network's generalization ability. (3) The activation function is used to perform a non-linear mapping on z to generate a feature map z' of size , enabling the channel information fusion unit to have non-linear expression ability. (4) Three fully connected layers are constructed to perform dimensionality increase operations on z' to generate three feature maps a, b, and c of size 1x1xC, which carry the channel information recalibration weights of the feature maps Featurel, Feature2, and Feature3 respectively. Subsequently, a, b, c are multiplied pairwise with Feature1, Feature2, and Feature3 respectively, and Feature1', Feature2', and Feature3' are output respectively to achieve channel information fusion.

[0073] The super-resolution network structure that combines receptive field fusion and channel information fusion has the following main advantages: (1) The multi-scale convolutional layer realizes receptive field fusion, which improves the receptive field of the network without increasing the network depth and alleviates the problem of gradient disappearance; it also improves the adaptability of the network to different input image information and the difference in information density of different regions in the same image, making the LR→HR prediction function constructed by the network more accurate, thereby improving the quality of the reconstructed image. (2) Introducing the squeeze-and-excitation mechanism to achieve channel information fusion can, on the one hand, make full use of the channel information of the image, alleviate the pressure of spatial feature extraction, speed up network training, and improve network efficiency; on the other hand, it can strengthen the extraction of the channel features of the original image and improve the reconstruction quality. (3) The feature map s is actually obtained by adding three three-dimensional feature maps output by the multi-scale convolutional layer and then performing average pooling. Therefore, it carries the channel information fused from different branches. Finally, the channel information feature maps a, b, and c used for recalibration are all uniformly guided by the global information carried in the feature map s. This structure of global fusion and branch excitation not only plays the self-adaptive selection role of each branch but also retains the expression of each branch for global information, and can maintain the effective transmission of information in complex network mappings.

[0074] ResNet adopts a head-to-tail connection method, enabling the intermediate structure of the network to only learn the difference between the output and the input, which can greatly reduce the learning obstacles of the network and alleviate the problem of vanishing gradients. In the proposed algorithm RCFSR, the ResNet residual structure is also introduced. In the non-linear mapping module and the residual setting strategy in RCFSR, the first layer is a Receptive Field and Channel Information Fusion Block (RCFB). This module consists of a receptive field fusion unit, a channel information fusion unit, and a common convolution. At the same time, the ResNet structure is introduced, and the input and output are directly connected. The second layer is a Receptive Field and Channel Information Fusion Group (RCFG), which is composed of multiple RCFBs and a common convolution stacked together. At the same time, the ResNet structure is introduced. The third layer is the RCFSR network. The input image first undergoes a common convolution for shallow feature extraction to expand the number of image channels, and then enters the non-linear mapping module composed of multiple RCFGs and a common convolution for feature mapping. Finally, it undergoes upsampling processing based on scaled convolution to generate the super-resolution reconstructed image. The settings of RCFB and RCFG can ensure that each unit in the network only needs to learn the difference between the output and the input of this unit, which can better improve the propagation efficiency of information in the deep network. The number of RCFB and RCFG represents the network depth. If the network is too shallow, it will lead to insufficient feature learning. If the network is too deep, it will bring problems such as vanishing gradients and an overly large model. The "residual learning" structure alleviates the problem of vanishing gradients to a certain extent, but the expansion of the network depth is still limited. Therefore, the number of RCFB and RCFG is an important factor affecting the network quality and efficiency.

[0075] The RCFSR network proposed by the present invention uses a scaled convolution strategy to replace the commonly used transposed convolution method for upsampling the image. Without increasing the algorithm design cost, it solves the checkerboard effect problem and reduces the graininess of the reconstructed image. Figure 4 As a comparison diagram of the transposed convolution method and the scaled convolution method, it actually replaces the transposed convolution module in the upsampling layer with an "interpolation + common convolution" structure. Scaled convolution is a simple and effective upsampling strategy, which is an effective combination of the traditional interpolation method and the convolutional neural network. The structure mainly includes two parts: (1) Interpolation processing: The expression of the linear interpolation method is as follows:

[0076]

[0077] Among them, (x0, y0) and (x1, y1) are known pixel points, (x, y) is the pixel point to be solved, and x ∈ [x0, x1] x ∈ [x0, y1], y ∈ [y0, y1]. It can be seen from the expression that the linear interpolation method actually calculates the weights according to the distances between x0, x, and x1, and then calculates the y value by weighted calculation of the distances between y0, y, and y1. Bilinear interpolation is the superposition of the results of linear interpolation in the x and y directions respectively. Suppose the size of the original image is h × w, and the scaling factor is "4", then the size of the target image is 4h × 4w, and the pixel coordinates corresponding to a certain pixel point (i, j) on the target image in the original image are Among them, and are often not integers, so it is necessary to and round up and down respectively to obtain the coordinates of the 4 pixel points closest to on the original image, and calculate the pixel value to be interpolated through the pixel values of these four pixel points. The interpolation method has a small amount of calculation, simple operation, and is easy to implement. However, a single interpolation process is too shallow for feature extraction. Therefore, convolution processing needs to be added after interpolation.

[0078] (2) Ordinary convolution operation: Use a 3x3 convolution kernel to repair the feature map after interpolation processing to make up for the deficiency of shallow feature extraction in the interpolation operation. Transposed convolution is the inverse process of ordinary convolution, which is a one-to-many uncertainty problem, while ordinary convolution calculation is a many-to-one deterministic problem. Therefore, there is no problem of "uneven overlap" and it can effectively solve the checkerboard effect problem. The two steps of the scaling convolution method are interpolation and ordinary convolution. Interpolation is used to enlarge the image, and ordinary convolution is used to repair the image and extract deep features. Although it lacks the end-to-end integrity of the transposed convolution method, the scaling convolution method connects two extremely simple structures, with a small overall calculation amount and low complexity. While not compromising the quality of the reconstructed image, it effectively solves the checkerboard effect problem. Therefore, overall, the performance of the scaling convolution is better than that of the transposed convolution.

[0079] In this embodiment, the reconstructed high-resolution target image is input into the feature pyramid network, and an adaptive feature fusion mechanism is adopted to dynamically adjust the weight ratio of low-level features and high-level features according to the target pixel ratio to determine the final target image feature distribution, including:

[0080] The feature pyramid network is used to process the high-resolution target image to extract multi-level feature maps. Based on a preset pixel occupancy threshold, it is determined whether the target pixel occupancy exceeds the threshold. If the target pixel occupancy exceeds the threshold, the weight ratio of the low-level features is increased. If the target pixel occupancy does not exceed the threshold, the weight ratio of the high-level features is increased. According to the dynamically adjusted weight ratio, an adaptive feature fusion mechanism is used to fuse the multi-level feature maps. Through the fused feature maps, the enhanced feature values of the target area are obtained. According to the enhanced feature values, the final target image feature distribution is determined.

[0081] In this embodiment, at the output end of the feature pyramid network, an attention mechanism is introduced to calculate the attention weights of the tiny target areas and enhance the saliency of the target features in the feature maps, including:

[0082] For the feature map X with dimensions C×H×W, where C is the number of channels, and H and W represent the height and width of the input image, pooling kernels of (H, 1) and (1, W) are used to perform one-dimensional feature encoding on the channels along the horizontal and vertical directions to obtain feature maps Z h and Z w , with sizes C×H×1 and C×1×W;

[0083] The obtained feature maps Z h and Z w are concatenated, and then a 1×1 convolution and a non-linear activation function are used to obtain an intermediate feature map f ∈ R C / r×1×(H+W) , where r is the downsampling ratio, and then f is decomposed along the spatial dimension into f h ∈ R C / r×H×1 and f w ∈ R C / r×1×W two separate tensors, and f h ∈ R C×H×1 and f w ∈ R C×1×W are obtained through two 1×1 convolution transformations. The attention weights g h and g w in the two directions are obtained using the activation function Sigmoid(x), and finally the attention weights g h and g w are multiplied by the input features as the output feature map.

[0084] In this embodiment, based on the enhanced feature map, a target detection model is used to predict the location and category information of the tiny targets through the regression and classification branches, including:

[0085] The target detection model is used to process the enhanced feature map. The bounding box coordinate values of the tiny targets are predicted through the regression branch, and the category probability values of the tiny targets are predicted through the classification branch;

[0086] Determine the position information and category information of the tiny target according to the predicted bounding box coordinate values and category probability values.

[0087] In this embodiment, according to the detection result, calculate the confidence difference between the tiny target and the background. If the confidence is lower than the preset threshold, return to the super-resolution reconstruction step for secondary reconstruction and detection.

[0088] Use the object detection model to process the enhanced feature map. Predict the bounding box coordinate values of the tiny target through the regression branch, and predict the category probability values of the tiny target through the classification branch. Calculate the confidence difference value between the tiny target and the background according to the predicted bounding box coordinate values and category probability values. If the confidence difference value is lower than the preset threshold, obtain the original image of the tiny target area for super-resolution reconstruction. Perform secondary reconstruction on the original image of the tiny target area through the super-resolution reconstruction algorithm to obtain a high-resolution image. Use the attention mechanism to process the high-resolution image and calculate the attention weight value of the tiny target area. According to the attention weight value, use the adaptive feature fusion mechanism to fuse the multi-level feature maps to obtain an enhanced feature map. Use the object detection model to process the enhanced feature map and re-predict the bounding box coordinate values and category probability values of the tiny target.

[0089] In this embodiment, during the training process, adopt a local data augmentation strategy based on the target structure, scale and rotate the area of the tiny target, and retain the original structure and context information of the target.

[0090] Adopt a local data augmentation strategy based on the target structure to perform scaling and rotation operations on the tiny target area, and retain the original structure and context information of the target. According to the scaled and rotated tiny target area, obtain the original image data of the target area and input it into the super-resolution reconstruction algorithm for processing. Reconstruct the original image of the target area through the super-resolution reconstruction algorithm to obtain a high-resolution image, and calculate the bounding box coordinate values of the tiny target in the image. Use the attention mechanism to process the high-resolution image, calculate the attention weight value of the tiny target area, and determine the saliency of the target area. According to the attention weight value, use the adaptive feature fusion mechanism to fuse the multi-level feature maps to obtain an enhanced feature map. Use the object detection model to process the enhanced feature map and re-predict the bounding box coordinate values and category probability values of the tiny target. If the confidence difference value between the re-predicted category probability value and the background is lower than the preset threshold, return to the super-resolution reconstruction step for secondary reconstruction and detection.

[0091] In this embodiment, through iterative training and optimization, adjust the parameters of the super-resolution reconstruction model, feature pyramid network and object detection model to improve the detection accuracy and stability of tiny targets.

[0092] The parameters of the super-resolution reconstruction model are adjusted by using an iterative training method to optimize the reconstruction effect. According to the optimized super-resolution reconstruction model, the input image is reconstructed to obtain a high-resolution image. Multilevel feature extraction is performed on the high-resolution image through a feature pyramid network to generate a feature map. The extracted feature map is processed by using an object detection model to predict the bounding box coordinate values and class probability values of the tiny objects. If the difference value between the predicted class probability value and the confidence level of the background is lower than a preset threshold, the super-resolution reconstruction step is returned for secondary reconstruction. According to the high-resolution image after secondary reconstruction, features are re-extracted and object detection is performed to update the bounding box coordinate values and class probability values of the tiny objects. Through multiple iterative trainings and optimizations, the detection accuracy and stability of the tiny objects are gradually improved.

[0093] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A visual detection method for small targets, characterized in that: The method comprises: Extract the tiny target area from the original image, and use the adaptive region segmentation algorithm to determine the target boundary range according to the target edge features and background contrast; According to the segmented tiny target area, a super-resolution reconstruction model based on deep learning is used to reconstruct a high-resolution target image through multi-scale feature extraction and residual connection. The reconstructed high-resolution target image is input into the feature pyramid network, and an adaptive feature fusion mechanism is used to dynamically adjust the weight ratio of low-level features and high-level features according to the target pixel ratio to determine the final target image feature distribution. At the output end of the feature pyramid network, an attention mechanism is introduced to calculate the attention weight of the small target area and enhance the significance of the target feature in the feature map. Based on the enhanced feature map, the target detection model is adopted to predict the location and category information of tiny targets through regression and classification branches.

2. The method according to claim 1, characterized in that: Extract the tiny target area from the original image, use the adaptive region segmentation algorithm, and determine the target boundary range based on the target edge features and background contrast: Get the original image and convert it into a grayscale image, and use Gaussian filtering for preprocessing; Calculate the image edge features according to the preset gradient threshold and generate an edge feature map; Calculate the background difference value through local contrast analysis to determine the potential target area; Combining the edge feature map and background difference value, an adaptive threshold segmentation algorithm is used to extract the candidate target area; Perform morphological processing on the candidate target area; Determine the target boundary based on regional connectivity analysis and obtain the final target area; According to the final target area, an image result with target boundary marking is obtained.

3. The method according to claim 1, characterized in that According to the segmented tiny target area, a super-resolution reconstruction model based on deep learning is used to reconstruct a high-resolution target image through multi-scale feature extraction and residual connection, including: Get the segmented tiny target area as input data; Use a deep learning super-resolution reconstruction model to preprocess the input data and extract multi-scale features; Multi-scale features are fused through residual connections to generate intermediate feature maps; Perform super-resolution reconstruction based on the intermediate feature map to obtain a high-resolution target image; If there is distortion in the high-resolution target image, a distortion correction algorithm is used to correct it; If there are artifacts in the high-resolution target image, an artifact removal algorithm is used to remove them; Output the final high-resolution target image to complete the reconstruction process.

4. The method according to claim 3, characterized in that The deep learning super-resolution reconstruction model includes three parts: feature extraction, nonlinear mapping and image reconstruction; Among them, the feature extraction part is a 3×3 convolution layer, which converts the input three-channel image into a multi-channel feature map for subsequent operations; The nonlinear mapping part is a combination of the receptive field fusion unit and the channel information fusion unit. At the same time, ResNet is introduced to realize residual learning, and the receptive field and channel information fusion block RCFB and the receptive field and channel information fusion group RCFG are constructed; The image reconstruction part uses scaled convolution.

5. The method according to claim 1, characterized in that The reconstructed high-resolution target image is input into the feature pyramid network, and an adaptive feature fusion mechanism is used to dynamically adjust the weight ratio of low-level features and high-level features according to the target pixel ratio to determine the final target image feature distribution, including: The feature pyramid network is used to process the high-resolution target image and extract multi-level feature maps; Based on a preset pixel ratio threshold, determine whether the target pixel ratio exceeds the threshold; If the target pixel ratio exceeds the threshold, the weight ratio of the low-level features is increased; If the target pixel ratio does not exceed the threshold, the weight ratio of the high-level features is increased; According to the dynamically adjusted weight ratio, an adaptive feature fusion mechanism is used to fuse multi-level feature maps; Obtain enhanced feature values ​​of the target area through the fused feature map; According to the enhanced feature values, the final target image feature distribution is determined.

6. The method according to claim 1, characterized in that At the output end of the feature pyramid network, an attention mechanism is introduced to calculate the attention weight of the small target area and enhance the significance of the target feature in the feature map, including: For a feature map X with dimensions C×H×W, where C is the number of channels, H and W represent the height and width of the input image, a pooling kernel of (H, 1) and (1, W) is used to encode the channels in one dimension along the horizontal and vertical directions to obtain a feature map Z in both the height and width directions. h and Z w , the sizes are C×H×1 and C×1×W; The feature map Z h and Z w Concatenate and then use 1×1 convolution and nonlinear activation function to obtain an intermediate feature map f∈R C / r×1×(H+W) , where r is the downsampling ratio, and then f is decomposed into f along the spatial dimension h ∈R C / r×H×1 and f w ∈R C / r×1×W Two separate tensors, transformed by two 1×1 convolutions to get f h ∈R C×H×1 and f w ∈R C×1×W , use the activation function Sigmoid(x) to obtain the attention weights g in both directions h and g w , and finally the attention weight g h and g w Multiply the input features as the output feature map.

7. The method according to claim 1, characterized in that Based on the enhanced feature map, the target detection model is used to predict the location and category information of small targets through regression and classification branches, including: The enhanced feature map is processed using the target detection model. The bounding box coordinates of the tiny target are predicted through the regression branch, and the category probability values ​​of the tiny target are predicted through the classification branch. According to the predicted bounding box coordinate values ​​and category probability values, the location information and category information of the tiny target are determined.

Citation Information

Cited By

  • Super-resolution method for tiny target recognition, electronic equipment and storage medium

    CN120672583A

  • Real-time target detection method and system for intelligent image processing

    CN120765917A

  • A real-time target detection method and system for intelligent image processing

    CN120765917B