Image processing method based on non-uniform sampling and scale adaptive supervision
Through the image processing method of non-uniform sampling and scale-adaptive supervision, the problem of poor accuracy in small target image processing is solved, and efficient, small target feature extraction and accurate detection are achieved in oil field scenarios.
Patent Information
- Application Number
- CN202511287103.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-10
AI Technical Summary
When processing small target images, especially in oilfield scenarios, existing technologies suffer from blurred target features, missing texture information, and difficulty in effectively extracting key features of small targets. In addition, existing methods lack adaptability in multi-scale target processing, resulting in poor processing accuracy.
An image processing method with non-uniform sampling and scale-adaptive supervision is adopted to restore high-frequency details of the image through a super-resolution reconstruction network. The SAS DETR model, which combines an adaptive instance-level grid generator and Transformer, dynamically adjusts the sampling density and feature extraction, introduces a saliency weight map and a scale-sensitive intersection-over-union metric, and optimizes the bounding box regression loss function.
It effectively preserves detailed features and significantly improves the processing accuracy of small targets without affecting the processing performance of large and medium targets, reduces computational redundancy, and improves the detection accuracy of small targets on oilfield datasets.
Smart Images

Figure CN120765459A_ABST
Abstract
Description
Technical Field
[0001] The invention discloses an image processing method based on non-uniform sampling and scale adaptive supervision, and belongs to the technical field of image processing. Background Art
[0002] In the field of computer vision, small target processing has always been a very challenging research direction, especially in practical applications in industrial scenarios such as oil fields, where this problem is particularly prominent. Traditional target processing algorithms have made significant progress in general scenarios, but they still face many technical bottlenecks when dealing with small target processing tasks. Oil field operation sites are usually characterized by complex environments, long monitoring distances, and variable lighting conditions. This makes the pixel ratio of the target in the image generally low, resulting in blurred target features and missing texture information. At the same time, the complex industrial background will generate a large amount of interference information, further increasing the difficulty of distinguishing small targets from the background. Most existing processing methods use a uniform downsampling strategy to process the input image. Although this processing method can reduce computational complexity, it will inevitably cause the loss of key details of small targets and lack the ability to adaptively process multi-scale targets.
[0003] In recent years, deep learning-based object processing algorithms have made significant progress. Two-stage processors such as FasterR-CNN generate candidate boxes through a region proposal network, single-stage processors such as the YOLO series directly predict target locations, and Transformer-based processors such as DETR use self-attention mechanisms to model global relationships. However, these methods all have significant limitations when processing small targets: CNN-based methods are limited by the size of the receptive field and have difficulty capturing the global contextual information of small targets; while Transformer-based methods, while capable of global modeling, are insufficient in feature extraction for small targets and have high computational complexity. Furthermore, existing methods have poor adaptability to the scale changes of small targets when processing oilfield scenarios, making it difficult to effectively extract the key features of small targets. There is an urgent need for a feature enhancement and efficient processing method for small targets. Summary of the Invention
[0004] The purpose of the present invention is to provide an image processing method based on non-uniform sampling and scale adaptive supervision to solve the problem of poor accuracy in small object image processing in the prior art.
[0005] An image processing method based on non-uniform sampling and scale adaptive supervision, comprising: S1. Obtain the image to be processed and send it to the super-resolution reconstruction network CAMixerSR to restore the high-frequency details of the image and generate a super-resolution image; S2, sequentially input the super-resolution image into the adaptive instance-level grid generator, the warping sampling network and the first sampler; S3, the first sampler performs non-uniform warping sampling on the super-resolution image based on the saliency weight map to obtain a non-uniform image; S4, inputting the non-uniform image into the backbone layer to obtain a non-uniform feature map and sending it to the second sampler; The output of the adaptive instance-level grid generator is divided into a plurality of rectangular block sets, which are sent to the inverse warping sampling network to extract features, and the inverse warping sampling network realizes multi-scale feature inverse warping through inverse transformation of bilinear mapping for each block set, and aligns the features to the original image coordinate system; S5, the second sampler obtains the restored feature map and inputs it into the Transformer-based SAS DETR model to obtain the image processing result.
[0006] The CAMixerSR includes a content-aware mixer and a residual connection module. The content-aware mixer uses convolution operation for detail simple regions in image regions and uses self-attention mechanism processing for detail rich regions in image regions according to the complexity of the image regions. The residual connection module fuses multi-scale features to output a super-resolution image.
[0007] The adaptive instance-level grid generator models the target bounding box distribution of the super-resolution image through a Gaussian kernel function, adaptively calculates saliency coefficients, and dynamically generates a saliency weight map. Adaptive calculation of saliency coefficients includes: ; In the formula, is the saliency coefficient of the i-th bounding box, is a scale control factor, is the width of the i-th bounding box, is the height of the i-th bounding box.
[0008] S3 includes constructing a separable one-dimensional sampling grid along the x-axis and y-axis of the super-resolution image, adjusting the grid density through the saliency weight map, oversampling high saliency regions and undersampling low saliency regions, and performing pixel value calculation of non-integer coordinates through bilinear interpolation.
[0009] The Transformer-based SAS DETR model includes inputting the recovered feature map into the backbone layer to obtain multi-level features, and then inputting the multi-level features into the top-down score adjustment layer Top-down Score Modulations, encoder Encoder, query refinement layer Query Refinement, and decoder Decoder in sequence. The encoder is provided with multiple encoder layers connected in series.
[0010] The input of the encoder layer is divided into three branches, one is directly input to the position update layer Update CorrespondingPositions, one is input to the global score predictor GSP, and one is input to the global threshold predictor GTP; The result of GSP is input into the foreground confidence layer to obtain the foreground confidence. The foreground confidence and the filtering threshold of each layer of feature map calculated by GTP are input into the dynamic token filtering layer Dynamic TokenFiltering. The dynamic token filtering layer is connected to the layer filtering layer. The output of the layer filtering layer includes background features and saliency features. The background features are filtered through the self-attention mechanism Self Attention, and the obtained saliency features are input into the position update layer. The position update layer outputs the processing results of the encoder layer.
[0011] Scale-Adaptive Salience Supervision is added to the dynamic token filtering layer and the layer filtering layer.
[0012] The global threshold predictor processes multi-scale features. The global threshold predictor includes four parallel paths. The first path adds the input features to itself and then inputs the features into the convolution layer CONV to obtain the output features of the first path. The structures of the other three paths are the same. The output features of the previous path are upsampled, then multiplied with the input features of the current path, and then added to the input features of the current path. The features are then input into the convolution layer to obtain the output features of the current path. The output features of each channel are fused with the output features of the previous channel, and finally the total output features are output by the fourth channel. Then, they are sequentially input into the normalization module LN, the activation function layer GELU, the linear layer Linear, and the global average pooling layer GAP to obtain the filtering threshold.
[0013] The scale-sensitive intersection-over-union (SIoU) metric is introduced as the loss function for bounding box regression, and the Hungarian algorithm is used to construct the loss function of the SAS DETR model.
[0014] The loss functions of the SAS DETR model include: ; Where, is the loss function, Matching losses for Hungary, is the denoising loss, is the encoder-assisted loss, is the focal loss for saliency supervision, 、 、 、 They are 、 、 、 The corresponding weight.
[0015] Compared with the existing technology, the present invention has the following beneficial effects: the present invention effectively retains detailed features, and the processing accuracy of small targets on oil field datasets is significantly improved; the processing effect of small targets is significantly improved without affecting the processing performance of large and medium targets. The dynamic token filtering strategy reduces computational redundancy and takes into account both efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is the overall flow chart of the present invention; Figure 2 This is a schematic diagram of the SAS DETR model structure based on Transformer; Figure 3 Schematic diagram of the global threshold predictor structure. DETAILED DESCRIPTION
[0017] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0018] An image processing method based on non-uniform sampling and scale-adaptive supervision, such as Figure 1 ,include: S1. Obtain the image to be processed and send it to the super-resolution reconstruction network CAMixerSR to restore the high-frequency details of the image and generate a super-resolution image; S2, inputting the super-resolution image into the adaptive instance-level grid generator, the warped sampling network and the first sampler in sequence; S3, the first sampler performs non-uniform warping sampling on the super-resolution image based on the saliency weight map to obtain a non-uniform image; S4, inputting the non-uniform image into a backbone layer to obtain a non-uniform feature map and sending the non-uniform feature map to a second sampler; The output of the adaptive instance-level grid generator is divided into a plurality of rectangular patch sets, which are input into an inverse warping sampling network to extract features, and the inverse warping sampling network realizes multi-scale feature inverse warping through inverse transformation of bilinear mapping for each patch set to align the features to the original image coordinate system. S5, the second sampler obtains the restored feature map and inputs the feature map into a Transformer-based SAS DETR model to obtain an image processing result.
[0019] The CAMixerSR includes a content-aware mixer and a residual connection module, the content-aware mixer uses convolution operation for detail simple regions in image regions and uses self-attention mechanism processing for detail rich regions in image regions according to image region complexity, and the residual connection module fuses multi-scale features to output a super-resolution image.
[0020] The adaptive instance-level grid generator models the target bounding box distribution of the super-resolution image through a Gaussian kernel function, adaptively calculates saliency coefficients, and dynamically generates a saliency weight map. Adaptive calculation of saliency coefficients includes: ; In the formula, is the saliency coefficient of the i-th bounding box, is a scale control factor, is the width of the i-th bounding box, is the height of the i-th bounding box.
[0021] S3 includes constructing separable one-dimensional sampling grids along the x-axis and y-axis of the super-resolution image, adjusting the grid density through the saliency weight map, oversampling high saliency regions and undersampling low saliency regions, and performing pixel value calculation of non-integer coordinates through bilinear interpolation.
[0022] The Transformer-based SAS DETR model is as follows: Figure 2 , including inputting the recovered feature map into the backbone layer to obtain multi-level features, and inputting the multi-level features into the top-down score adjustment layer Top-down ScoreModulations, encoder Encoder, query improvement layer Query Refinement, and decoder Decoder in sequence, wherein the encoder is provided with multiple encoder layers connected in series.
[0023] The input of the encoder layer is divided into three branches, one is directly input to the position update layer Update CorrespondingPositions, one is input to the global score predictor GSP, and one is input to the global threshold predictor GTP; The result of GSP is input into the foreground confidence layer to obtain the foreground confidence. The foreground confidence and the filtering threshold of each layer of feature map calculated by GTP are input into the dynamic token filtering layer Dynamic TokenFiltering. The dynamic token filtering layer is connected to the layer filtering layer. The output of the layer filtering layer includes background features and saliency features. The background features are filtered through the self-attention mechanism Self Attention, and the obtained saliency features are input into the position update layer. The position update layer outputs the processing results of the encoder layer.
[0024] Scale-Adaptive Salience Supervision is added to the dynamic token filtering layer and the layer filtering layer.
[0025] The global threshold predictor processes multi-scale features such as Figure 3 ,The global threshold predictor includes four parallel paths. The first path adds the input features and itself to the ,features, and then inputs them into the convolutional layer CONV to obtain the ,output features of the first path; The structures of the other three paths are the same. The output features of the previous path are upsampled, then multiplied with the input features of the current path, and then added to the input features of the current path. The features are then input into the convolution layer to obtain the output features of the current path. The output features of each channel are fused with the output features of the previous channel, and finally the total output features are output by the fourth channel. Then, they are sequentially input into the normalization module LN, the activation function layer GELU, the linear layer Linear, and the global average pooling layer GAP to obtain the filtering threshold.
[0026] The scale-sensitive intersection-over-union (SIoU) metric is introduced as the loss function for bounding box regression, and the Hungarian algorithm is used to construct the loss function of the SAS DETR model.
[0027] The loss functions of the SAS DETR model include: ; Where, is the loss function, Matching losses for Hungary, is the denoising loss, is the encoder-assisted loss, is the focal loss for saliency supervision, 、 、 、 They are 、 、 、 The corresponding weight.
[0028] Tables 1 through 3 provide detailed comparisons of our proposed method with current mainstream object processing methods on multiple datasets. We used ResNet50 as the backbone network and conducted comprehensive evaluations on MS COCO2017, AI-TOD-V2, and self-built oilfield datasets to verify the effectiveness and generalization capabilities of our method. The evaluation results of the object detection framework's improvement on the object detection model are shown in Table 1.
[0029] Table 1 Evaluation results of the improvement effect of the target detection framework on the target detection model ; In Table 1, * indicates that the present invention was used for training. The first vertical column is for various methods, and the first horizontal column is for various accuracy indicators. From the results of the MS COCO2017 dataset shown in Table 1, compared with the original model, after training using the target processing framework proposed by the present invention, the three classic model indicators have been greatly improved, among which FCOS has the largest improvement, AP has increased by 4.9%, and small target indicator AP has increased by 1.3%. s The improvement is 6.2%. This shows that the target processing framework has universal enhancement capabilities for different architectural models and pays more attention to small targets.
[0030] The evaluation results of each method on the MS COCO2017 dataset are shown in Table 2.
[0031] Table 2 Evaluation results of various methods on the MS COCO2017 dataset ; From the results of various methods on the MS COCO2017 dataset shown in Table 2, SAS DETR achieved an AP of 50.4%, which is significantly better than other advanced models. Compared with the baseline model Salience DETR, SAS DETR achieved a 1.2% improvement in AP indicators. Specifically, under the same training conditions, the SAS DETR proposed in the present invention achieved +4.4% AP compared to Deformable-DETR, while the number of training iterations was greatly reduced; under the same number of training iterations, it achieved +1.4% AP compared to DINO. It is worth noting that compared with other models, SAS DETR has an AP of s The improvement in these indicators is generally greater than the improvement in AP, which further proves the superiority of the model for small target processing tasks.
[0032] The evaluation results of each method on the AI-TOD-V2 dataset are shown in Table 3.
[0033] Table 3 Evaluation results of various methods on the AI-TOD-V2 dataset ; In the AI-TOD-V2 dataset evaluation shown in Table 3, SAS DETR outperformed other processors in most evaluation indicators. Specifically, the present invention achieved an AP of 25.2% on the AI-TOD-V2 test set, which is 1.8% higher than Salience DETR. In addition, thanks to the scale-adaptive saliency supervision, small objects can more effectively utilize the feature information within the processing box. The present invention observed a significant improvement in small object processing performance, including AP vt and AP t The improvements were 2.4% and 1.9% respectively, which further verified the advantages of the proposed method in small target processing tasks.
[0034] The evaluation results of each method on the self-built oil field dataset are shown in Table 4.
[0035] Table 4 Evaluation results of various methods on self-built oilfield dataset ; In the evaluation of the self-built oilfield dataset shown in Table 4, SAS DETR, combined with a non-uniform sampling framework, improved mAP by 6.78%. The AP for the "smoking" category increased by 7.7%, and the AP for the "mobile phone" category increased by 8.7%. This demonstrates the adaptability and practicality of our method in real-world industrial scenarios.
[0036] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image processing method based on non-uniform sampling and scale adaptive supervision, characterized in that: include: S1. Obtain the image to be processed and send it to the super-resolution reconstruction network CAMixerSR to restore the high-frequency details of the image and generate a super-resolution image; S2, inputting the super-resolution image into the adaptive instance-level grid generator, the warped sampling network and the first sampler in sequence; S3, the first sampler performs non-uniform distortion sampling on the super-resolution image based on the saliency weight map to obtain a non-uniform image; S4, input the non-uniform image into the backbone layer to obtain a non-uniform feature map and send it to the second sampler; The output of the adaptive instance-level grid generator is divided into several rectangular block sets and fed into the anti-distortion sampling network to extract features. The anti-distortion sampling network performs multi-scale feature anti-distortion on each block set through the inverse transformation of the bilinear mapping and aligns the features to the original image coordinate system. S5. The second sampler obtains the restored feature map and inputs it into the Transformer-based SAS DETR model to obtain the image processing result.
2. The image processing method based on non-uniform sampling and scale adaptive supervision according to claim 1, characterized in that: CAMixerSR includes a content-aware mixer and a residual connection module. The content-aware mixer uses convolution operations on areas with simple details in the image region according to the complexity of the image region, and uses a self-attention mechanism to process areas with rich details in the image region. The residual connection module fuses multi-scale features and outputs a super-resolution image.
3. The image processing method based on non-uniform sampling and scale adaptive supervision according to claim 2, characterized in that: The adaptive instance-level grid generator models the target bounding box distribution of the super-resolution image through a Gaussian kernel function, adaptively calculates the saliency coefficient, and dynamically generates a saliency weight map; Adaptive calculation of significance coefficients includes: ; Where, For the The significance coefficient of the bounding box, is the scale control factor, For the The width of the bounding box, For the The height of the bounding box.
4. The image processing method based on non-uniform sampling and scale adaptive supervision according to claim 3, characterized in that: S3 includes constructing separable one-dimensional sampling grids along the x-axis and y-axis of the super-resolution image, adjusting the grid density through the saliency weight map, oversampling high-saliency areas, undersampling low-saliency areas, and calculating pixel values of non-integer coordinates through bilinear interpolation.
5. The image processing method based on non-uniform sampling and scale adaptive supervision according to claim 4, characterized in that: The Transformer-based SAS DETR model includes inputting the recovered feature map into the backbone layer to obtain multi-level features, and then inputting the multi-level features into the top-down score adjustment layer Top-down Score Modulations, encoder Encoder, query refinement layer Query Refinement, and decoder Decoder in sequence. The encoder is provided with multiple encoder layers connected in series.
6. The image processing method based on non-uniform sampling and scale adaptive supervision according to claim 5, characterized in that: The input of the encoder layer is divided into three branches, one is directly input to the position update layer UpdateCorresponding Positions, one is input to the global score predictor GSP, and one is input to the global threshold predictor GTP; The result of GSP is input into the foreground confidence layer to obtain the foreground confidence. The foreground confidence and the filtering threshold of each layer of feature map calculated by GTP are input into the dynamic token filtering layer. The dynamic token filtering layer is connected to the layer filtering layer. The output of the layer filtering layer includes background features and saliency features. The background features are filtered through the self-attention mechanism, and the obtained saliency features are input into the position update layer. The position update layer outputs the processing results of the encoder layer.
7. The image processing method based on non-uniform sampling and scale adaptive supervision according to claim 6, characterized in that: Scale-Adaptive Salience Supervision is added to the dynamic token filtering layer and the layer filtering layer.
8. The image processing method based on non-uniform sampling and scale adaptive supervision according to claim 7, characterized in that: The global threshold predictor processes multi-scale features. The global threshold predictor includes four parallel paths. The first path adds the input features to itself and then inputs the features into the convolution layer CONV to obtain the output features of the first path. The structures of the other three paths are the same. The output features of the previous path are upsampled, then multiplied with the input features of the current path, and then added to the input features of the current path. The features are then input into the convolution layer to obtain the output features of the current path. The output features of each channel are fused with the output features of the previous channel, and finally the total output features are output by the fourth channel. Then, they are sequentially input into the normalization module LN, the activation function layer GELU, the linear layer Linear, and the global average pooling layer GAP to obtain the filtering threshold.
9. The image processing method based on non-uniform sampling and scale adaptive supervision according to claim 8, characterized in that: The scale-sensitive intersection-over-union (SIoU) metric is introduced as the loss function for bounding box regression, and the Hungarian algorithm is used to construct the loss function of the SAS DETR model.
10. The image processing method based on non-uniform sampling and scale adaptive supervision according to claim 9, characterized in that: The loss functions of the SAS DETR model include: ; Where, is the loss function, Matching losses for Hungary, is the denoising loss, is the encoder-assisted loss, is the focal loss for saliency supervision, 、 、 、 They are 、 、 、 The corresponding weight.
Citation Information
Patent Citations
Small target detection method based on scale adaptive saliency filtering
CN118887379A
Unmanned aerial vehicle aerial image target detection method based on learnable non-uniform sampling
CN119762995A
Complex scene small target detection system and method based on mask attention and context feature optimization
CN120219695A
Self-supervised dynamic scene monocular depth estimation method based on efficient fine tuning of parameters
CN120612354A
Sample distribution-informed denoising & rendering
US20230065183A1