Scene understanding method and system for multi-level image feature extraction

By combining DFFormer and MAE-enhanced ViT, fine-grained local features and global context information of the image are extracted and fused, the limitations of the prior art in understanding complex scenes are solved, and a more accurate and robust feature extraction effect is achieved.

CN120070914APending Publication Date: 2025-05-30CHINESE PEOPLE'S PUBLIC SECURITY UNIVERSITY +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510226420.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has limitations in extracting fine-grained local features and global context information in complex scenarios, and is poorly robust to noise and occlusion.

Method used

DFFormer is used to extract fine-grained local scene features, and global scene features are extracted through MAE-enhanced ViT, combining dynamic filtering and self-supervised training to achieve efficient extraction and fusion of features.

Benefits of technology

It realizes efficient extraction and fusion of image features, enhances the ability to understand complex scenes, and improves the robustness and generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070914A_ABST
    Figure CN120070914A_ABST
Patent Text Reader

Abstract

The invention discloses a scene understanding method and system for multi-level image feature extraction, and the method comprises the steps: obtaining input image data, and carrying out the data preprocessing, and obtaining the preprocessed image data; extracting fine-grained local scene features of the preprocessed image data based on DFFormer to obtain local scene features; extracting global features based on MAE enhancement to obtain global scene features; fusing the local scene features and the global scene features to obtain fused features; and performing classified output according to the fusion features. The limitation of an existing scene understanding method in the aspect of extracting fine-grained local features and global context information is solved. By combining the dynamic filtering capability of the DFFormer and the global feature representation capability of the ViT enhanced by the MAE, efficient extraction and fusion of image features are realized, and more accurate input features are provided for a scene understanding task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a scene understanding method and system for multi-level image feature extraction. Background Art

[0002] Image feature extraction is the basis of scene understanding tasks. Traditional feature extraction methods rely on manually designed features (such as SIFT, HOG, etc.), and often fall short when dealing with complex visual scenes. In recent years, with the development of deep learning, methods based on convolutional neural networks (CNNs) and self-attention mechanisms (such as the Transformer architecture) have rapidly emerged, significantly improving the effect of feature extraction. However, these technologies have their own advantages and limitations in feature extraction, and how to combine their advantages is the focus of current research.

[0003] The limitations of traditional feature extraction methods are specifically as follows:

[0004] Before the rise of deep learning, feature extraction mainly relied on manually designed image descriptors.

[0005] For example: SIFT (Scale-Invariant Feature Transform): Extracts key points and their local descriptions in an image, but is sensitive to illumination changes and affine transformations.

[0006] HOG (Histogram of Oriented Gradients): Describes image features by extracting the distribution of gradient directions, but performs poorly when dealing with complex scenes and non-rigid deformations.

[0007] LBP (Local Binary Pattern): Suitable for texture analysis, but unable to capture global information.

[0008] These methods work well for simple scenes, but are insufficient in global understanding and semantic representation of complex scenes, and are less robust to noise and occlusion. Summary of the Invention

[0009] In view of the above problems, the present invention is proposed to provide a scene understanding method and system for multi-level image feature extraction that overcomes or at least partially solves the above problems.

[0010] According to one aspect of the present invention, a scene understanding method for multi-level image feature extraction is provided. The scene understanding method includes:

[0011] Obtain input image data and perform data preprocessing to obtain preprocessed image data;

[0012] Extract fine-grained local scene features of the preprocessed image data based on DFFormer to obtain local scene features;

[0013] Based on MAE-enhanced global feature extraction, global scene features are obtained;

[0014] The local scene features and the global scene features are fused to obtain fused features;

[0015] Classification output is performed according to the fused features.

[0016] Optionally, the obtaining of the input image data and the performing of data preprocessing to obtain preprocessed image data specifically includes:

[0017] During the data preprocessing stage, various random enhancement operations are performed on the input image data, including:

[0018] Random cropping and scaling transformation: randomly crop and scale the spatial size and position of the image to enhance the model's adaptability to images of different resolutions;

[0019] Gaussian blur: reduce the model's dependence on overly sharp images through random blurring, simulating the blurred image situation in the real scene;

[0020] Elastic transformation: perform elastic deformation on the image to enhance the model's robustness to shape changes;

[0021] Random erasing: occlude part of the image area to simulate the scene where part of the target is occluded;

[0022] Random perspective transformation: perform random perspective projection on the image to enhance the model's adaptability to angle changes.

[0023] Optionally, the extracting of the fine-grained local scene features of the preprocessed image data based on DFFormer to obtain local scene features specifically includes:

[0024] Adaptively generate filter weights according to the features of the input image to adapt to the scene detail features of various images;

[0025] The overall architecture of DFFormer is based on the MetaFormer framework and has the characteristic of hierarchical design;

[0026] Multi-stage downsampling, between each stage, use convolution operations to gradually downsample the image scene features;

[0027] Feature maps of various resolutions retain the complete information from fine-grained to high-level semantics after layer-by-layer processing.

[0028] Optionally, the overall architecture of DFFormer being based on the MetaFormer framework specifically includes:

[0029] The first stage is low-level feature extraction: use large convolution kernels and strides for downsampling to extract low-level scene features;

[0030] The middle stage is mid-level feature extraction: introducing multi-level MetaFormer blocks to gradually extract the mid-level scene features of the image;

[0031] High-level feature extraction: using dynamic filtering combined with deeper MetaFormer blocks to capture high-level scene semantic features;

[0032] Each MetaFormer block completes feature modeling through two core components:

[0033] The Token Mixer is dynamic filtering or global filtering: dynamically modeling local or global scene information;

[0034] The MLP module: used to enhance the non-linear expression ability and further fuse scene features.

[0035] Optionally, the global feature extraction based on MAE enhancement to obtain global scene features specifically includes:

[0036] MAE self-supervised training: In the pre-training stage, randomly mask some regions of the input image and reconstruct the masked regions through the decoder;

[0037] The encoder only extracts scene features from the unmasked regions, learning context relationships and global dependency features;

[0038] Global feature extraction: MAE-enhanced ViT extracts long-range global scene semantic information of the image, especially in complex scenes, capturing the object relationships and global layout in the image.

[0039] Optionally, the fusion of the local scene features and the global scene features to obtain the fused features specifically includes:

[0040] Convert the feature map generated by DFFormer into attention. First, encode the features using a linear transformation operation, and then, through an activation function such as Softmax to ensure that these scores can represent a probability distribution, thus forming attention weights. Subsequently, integrate the attention mechanism guided by DFFormer features into the Vision Transformer (ViT) architecture, and choose to introduce the attention module after a specific layer of ViT. The specific implementation method is to multiply the generated attention map with the feature map output by the corresponding layer of ViT channel by channel;

[0041] The fused features contain both global scene semantic information and retain fine-grained local scene features.

[0042] Optionally, the classification output according to the fused features specifically includes:

[0043] The fused features are output through a lightweight classification head and used as the input for the scene detection task.

[0044] Multi-task support: Output global scene features for classification, and output multi-scale feature maps for detection or segmentation tasks.

[0045] The present invention also provides a scene understanding system for multi-level image feature extraction, which applies the above-mentioned scene understanding method for multi-level image feature extraction. The scene understanding system includes:

[0046] An image data preprocessing module, configured to obtain input image data and perform data preprocessing to obtain preprocessed image data.

[0047] A local feature extraction module based on DFFormer, configured to extract fine-grained local scene features of the preprocessed image data based on DFFormer to obtain local scene features.

[0048] A global feature extraction module enhanced by MAE, configured to obtain global scene features based on global feature extraction enhanced by MAE.

[0049] A feature fusion module, configured to fuse the local scene features and the global scene features to obtain fused features.

[0050] A feature output module, configured to perform classification output according to the fused features.

[0051] A scene understanding method and system for multi-level image feature extraction provided by the present invention. The scene understanding method includes: obtaining input image data and performing data preprocessing to obtain preprocessed image data; extracting fine-grained local scene features of the preprocessed image data based on DFFormer to obtain local scene features; obtaining global scene features based on global feature extraction enhanced by MAE; fusing the local scene features and the global scene features to obtain fused features; and performing classification output according to the fused features. It solves the limitations of existing scene understanding methods in extracting fine-grained local features and global context information. By combining the dynamic filtering ability of DFFormer with the global feature representation ability of MAE-enhanced ViT, it realizes the efficient extraction and fusion of image features and provides more accurate input features for scene understanding tasks.

[0052] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the description. And in order to make the above and other objects, features and advantages of the present invention more obvious and understandable, the following specifically describes the specific embodiments of the present invention. Brief Description of the Drawings

[0053] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0054] Figure 1 Overall framework diagram of the algorithm;

[0055] Figure 2 Flowchart of self-supervised task enhanced ViT;

[0056] Figure 3 It is a schematic logical flow diagram of a scene understanding method for multi-level image feature extraction provided by an embodiment of the present invention. Detailed implementation manners

[0057] The following will describe the exemplary embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0058] The terms "including" and "having" and any variations thereof in the description of the embodiments and claims of the present invention and the accompanying drawings are intended to cover non-exclusive inclusion. For example, including a series of steps or units.

[0059] The following will further describe the technical solutions of the present invention in detail in conjunction with the accompanying drawings and embodiments.

[0060] The present invention proposes an image feature extraction method and system based on the DFFormer framework and MAE self-supervised task enhancement, aiming to solve the limitations of existing scene understanding methods in extracting fine-grained local features and global context information.

[0061] By combining the dynamic filtering ability of DFFormer with the global feature representation ability of MAE-enhanced ViT, the efficient extraction and fusion of image features are realized, providing more accurate input features for scene understanding tasks.

[0062] Embodiment 1

[0063] As Figure 1 and Figure 2 shown, a scene understanding method for multi-level image feature extraction includes:

[0064] 1. Image data preprocessing module, including:

[0065] To enhance the robustness of the scene understanding model and improve its generalization ability, the present invention performs various random enhancement operations on the input image data during the data preprocessing stage, including but not limited to:

[0066] Random cropping and scaling transformation: Randomly crop and scale the spatial size and position of the image to enhance the model's adaptability to images of different resolutions.

[0067] Gaussian blur: Reduce the model's dependence on overly sharp images through random blurring, simulating the situation of blurred images in real-world scenes.

[0068] Elastic transformation: Elastically deform the image to enhance the model's robustness to shape changes.

[0069] Random erasing: Occlude part of the image area to simulate the scene where some objects are occluded.

[0070] Random perspective transformation: Perform random perspective projection on the image to enhance the model's adaptability to angular changes.

[0071] This not only improves the data diversity during the training process of the model but also helps the model better adapt to complex changes in the actual scene.

[0072] 2. The local feature extraction module based on DFFormer includes:

[0073] DFFormer (Dynamic FilterFormer) is the core framework for feature extraction in the present invention, responsible for extracting fine-grained local scene features of the image, including:

[0074] (1) Dynamic filtering module

[0075] Dynamic weight generation:

[0076] The convolution kernels of traditional convolutional neural networks are fixed, while the dynamic filtering module can adaptively generate filter weights according to the features of the input image, thus better adapting to the scene detail features of different images.

[0077] The dynamic filter shows extremely strong flexibility in extracting image edges, textures, and other local scene features.

[0078] (2) MetaFormer framework

[0079] The overall architecture of DFFormer is based on the MetaFormer framework and has the characteristics of a hierarchical design. Specifically, it includes:

[0080] The first stage (low-level feature extraction): Use a large convolution kernel (such as 7×7) and a relatively large stride (such as 4) for downsampling to extract low-level scene features;

[0081] Intermediate stage (mid-level feature extraction): Introduce multi-level MetaFormer blocks to gradually extract mid-level scene features of the image;

[0082] High-level feature extraction: Use a dynamic filtering module combined with deeper MetaFormer blocks to capture high-level scene semantic features.

[0083] Each MetaFormer block completes feature modeling through two core components:

[0084] Token Mixer (dynamic filtering or global filtering): Responsible for dynamically modeling local or global scene information.

[0085] MLP module: Used to enhance the non-linear expression ability and further fuse scene features.

[0086] (3) Multi-stage downsampling

[0087] Between each stage, use convolutional operations to achieve gradual downsampling of the image scene features. Feature maps of different resolutions retain the complete information from fine-grained to high-level semantics after layer-by-layer processing.

[0088] 3. MAE-enhanced global feature extraction module, including:

[0089] To make up for the deficiency of DFFormer in capturing global context information, the present invention combines a ViT (Vision Transformer) module based on MAE (Masked Autoencoder):

[0090] MAE self-supervised training:

[0091] In the pre-training stage, randomly mask some regions of the input image and reconstruct the masked regions through the decoder.

[0092] The encoder only extracts scene features from the unmasked regions, learning context relationships and global dependency features.

[0093] Global feature extraction:

[0094] MAE-enhanced ViT can extract long-range global scene semantic information of the image. Especially in complex scenes, it can effectively capture the object relationships and global layout in the image.

[0095] 4. Feature fusion module, including:

[0096] The feature fusion module of the present invention aims to organically combine the local scene features extracted by DFFormer with the MAE-enhanced global scene features:

[0097] Attention Weight Calculation:

[0098] Convert the feature map generated by DFFormer into attention. First, encode the features using a linear transformation operation, and then, through an activation function such as Softmax, ensure that these scores can represent a probability distribution, thus forming attention weights. Pixel-wise feature multiplication:

[0099] Multiply the generated attention map and the feature map output by the corresponding layer of ViT channel-wise, enabling ViT to utilize this enhanced information for subsequent processing and decision-making. The fused features contain both global scene semantic information and fine-grained local scene features.

[0100] 5. Feature Output Module, including:

[0101] Final Features: The fused features are output through a lightweight classification head (MLP Head) and used as the input for the scene detection task.

[0102] Multi-task Support: Through the modular design of features, it can not only output global scene features for classification but also output multi-scale feature maps for detection or segmentation tasks.

[0103] Embodiment 2

[0104] As Figure 3 shown, a scene understanding method for multi-level image feature extraction includes:

[0105] 1. Data Preprocessing Module

[0106] The present invention preprocesses the input image through a series of data augmentation techniques to make the input data more diverse, thereby improving the generalization ability of the scene understanding model. Specifically, it includes the following operations:

[0107] (1) Random Shearing and Scale Transformation

[0108] During each training, randomly crop a part of the image area and resize it to a fixed size (e.g., 224×224).

[0109] The cropping ratio range is 0.8, 1.0. The cropping position is random.

[0110] Purpose: Enhance the robustness of the model to targets of different scales and positions.

[0111] (2) Gaussian Blur

[0112] Apply Gaussian blur with a random degree (the range of σ is 0.1, 2.0.) to the input image.

[0113] Simulate the blurring situation in reality caused by inaccurate camera focusing or movement, and enhance the model's adaptability to blurred images.

[0114] (3) Elastic transformation

[0115] Apply random non - linear deformations to the pixel grid in the image.

[0116] This operation enhances the model's perception ability of non - rigid deformation targets by controlling the deformation amplitude and elastic parameters.

[0117] (4) Random erasing

[0118] Randomly select a rectangular area in the image and set the pixels in this area to zero or fill them randomly.

[0119] This operation simulates the scenario where some targets are occluded, such as the situation where a target is occluded by an obstacle.

[0120] (5) Random perspective transformation

[0121] Apply a random perspective projection to the input image to change the perspective of the image.

[0122] Purpose: Simulate the perspective effect brought by the change of camera angle in the real scene and improve the model's adaptability to different shooting angles.

[0123] (6) Normalization and standardization

[0124] Perform normalization and standardization processing on the enhanced image:

[0125]

[0127] where μ and σ are the mean and standard deviation of the image pixels respectively.

[0128] Ensure that the numerical range of the input image is adapted to the feature extraction module of the model.

[0129] 2. Local feature extraction module based on DFFormer

[0130] The core of this module is the DFFormer framework, which is designed based on a hierarchical feature extraction architecture to gradually extract fine - grained local scene features of the image.

[0131] (1) Overall structure of DFFormer

[0132] Four - stage feature extraction:

[0133] Each stage consists of a down - sampling layer and multiple MetaFormer blocks.

[0134] The input resolution gradually decreases at each stage, while the feature dimension gradually increases.

[0135] For example: the input resolution starts from 224×224 and gradually decreases to 112×112, 56×56, 28×28, and 14×14.

[0136] Number of feature channels:

[0137] The number of channels at each stage gradually increases, such as 64, 128, 320, 512.

[0138] (2) Dynamic filtering module

[0139] Dynamic weight generation:

[0140] Use a fully connected layer to generate the weights of the dynamic filter, and the weights vary with the input features.

[0141] Each dynamic filter can capture the feature patterns at different positions in the input image.

[0142] Efficient feature fusion:

[0143] The dynamic filter weights the input features through element-wise operations, strengthening the expression of local edge and texture information.

[0144] (3) MetaFormer block

[0145] Each MetaFormer block contains the following components:

[0146] Token Mixer: The dynamic filtering module is used for local feature modeling.

[0147] MLP module: Consists of two fully connected layers, used to enhance the non-linear expression ability of features.

[0148] Normalization layer: Use LayerNorm to normalize features, ensuring the stability of the feature distribution.

[0149] (4) Downsampling layer

[0150] Downsampling is achieved through convolution operations. The convolution kernel size is 3×3, the stride is 2, and the number of channels gradually increases.

[0151] The downsampled feature map is passed to the next stage.

[0152] 3. MAE-enhanced global feature extraction module

[0153] Based on ViT (Vision Transformer), use MAE (Masked Autoencoder) self-supervised task pre-training to enhance the global scene feature extraction ability.

[0154] (1) MAE Self-Supervised Task

[0155] Masking Strategy:

[0156] 75% of the pixels of the input image are randomly masked, leaving only 25% visible.

[0157] The pixel values of the masked part are replaced with a fixed padding value.

[0158] Reconstruction Task:

[0159] The encoder only extracts features from the visible part.

[0160] The decoder reconstructs the masked area, aiming to minimize the reconstruction error (such as Mean Squared Error - MSE).

[0161] (2) ViT Architecture

[0162] Patch Processing:

[0163] The image is divided into patches of a fixed size (e.g., 16×16), and each patch is flattened into a vector.

[0164] Self-Attention Mechanism:

[0165] ViT captures the dependencies between any patches through the self-attention mechanism to form a global scene feature representation.

[0166] (3) Feature Output

[0167] The global scene features output by ViT include long-range context information in the image and are especially good at capturing the global layout and scene relationships of objects.

[0168] 4. Feature Fusion Module

[0169] In the present invention, the feature maps generated by DFFormer are converted into attention: First, the features are encoded using a linear transformation or other form of operation; then, these scores are possibly passed through an activation function such as Softmax to ensure that they represent a probability distribution, thus forming the attention weights A df , when integrating the attention mechanism derived from DFFormer features into the Vision Transformer (ViT) architecture, we choose to introduce the attention module after a specific layer of ViT. The specific implementation is to multiply the generated attention map with the feature map output by the corresponding layer of ViT channel-wise, enabling ViT to utilize this enhanced information for subsequent processing and decision-making.

[0170] (1) Weight Calculation

[0171] The relevant formula is as follows:

[0172] A df = Softmax(Linear(Z fft ))

[0173] (2) Feature multiplication

[0174] Multiply the attention weights of the DFFormer features with the ViT features pixel by pixel to form the fused features:

[0175] Zfused = Adf·Z vit

[0176] (3) Feature normalization

[0177] The fused features are normalized by LayerNorm to ensure the stability of their numerical distribution.

[0178] 5. Final feature output module

[0179] Classification head:

[0180] Use an MLP head (multi-layer perceptron) to classify the fused features.

[0181] The output class probabilities are used for the image scene classification task.

[0182] Multi-task support:

[0183] The fused features can also be used for object detection (through the region proposal network) and image segmentation (by generating a segmentation map through the decoder).

[0184] Experimental results and performance evaluation

[0185] (1) Datasets and metrics

[0186] The experiments were conducted on datasets such as ImageNet, COCO, and Pascal VOC.

[0187] The main evaluation metrics include classification accuracy (Top-1 Accuracy), object detection mAP (mean Average Precision), and segmentation mIoU (mean Intersection over Union).

[0188] (2) Performance improvement

[0189] In the ImageNet classification task, the Top-1 accuracy was increased to 86.7%.

[0190] In the COCO object detection task, the mAP was increased by 3.5 percentage points.

[0191] In the Pascal VOC segmentation task, the mIoU reached 82.3%.

[0192] (3) Computational efficiency

[0193] The present invention is trained and inferred on an NVIDIA A100 GPU. The dynamic filtering design of the DFFormer module significantly reduces the computational complexity, and the inference speed is increased by 17% compared to traditional ViT.

[0194] The present invention provides an efficient, flexible, and robust method for image scene understanding, providing a solid foundation for various computer vision tasks. The proposed method for image scene feature extraction based on DFFormer and MAE enhancement effectively fuses local details and global context features of images through dynamic filtering and self-supervised enhancement, with broad application value and high practicality.

[0195] Beneficial effects: Complementary nature of global and local scene features: DFFormer is good at extracting fine-grained local scene features of images; the MAE-enhanced ViT focuses on global context modeling.

[0196] Through the feature fusion of the two, the complementarity of global and local information is effectively achieved.

[0197] Dynamic feature modeling: The dynamic filter can generate convolutional weights in real time according to the input data, more flexibly adapting to diverse image scene contents.

[0198] Self-supervised enhancement: The MAE self-supervised task enables the encoder to learn high-quality image scene representations without relying on large-scale labeled data.

[0199] Adapting to multi-task requirements: The output features are suitable for both global tasks such as classification and tasks that require fine-grained features such as segmentation and detection.

[0200] Strong robustness: A variety of enhancement operations in the image preprocessing stage significantly improve the generalization ability of the method in complex scenes.

[0201] The above specific implementation manners further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A scene understanding method for multi-level image feature extraction, characterized in that: The scene understanding method comprises: Acquire input image data, and perform data preprocessing to obtain preprocessed image data; Extract fine-grained local scene features of preprocessed image data based on DFFormer to obtain local scene features; Based on MAE-enhanced global feature extraction, global scene features are obtained; Fusing the local scene feature with the global scene feature to obtain a fused feature; Classification output is performed according to the fusion features.

2. A scene understanding method for multi-level image feature extraction according to claim 1, characterized in that: The step of obtaining input image data and performing data preprocessing to obtain preprocessed image data specifically includes: In the data preprocessing stage, a variety of random enhancement operations are performed on the input image data, including: Random cropping and scaling: Randomly crop and scale the spatial size and position of the image to enhance the model's adaptability to images of different resolutions; Gaussian blur: random blurring reduces the model's reliance on overly sharp images and simulates image blur in real scenes. Elastic transformation: elastically deform the image to enhance the robustness of the model to shape changes; Random erasing: Block parts of the image to simulate a scene where part of the target is blocked; Random perspective transformation: Perform random perspective projection on the image to enhance the model’s adaptability to angle changes.

3. The scene understanding method of multi-level image feature extraction according to claim 1, characterized in that: The method of extracting fine-grained local scene features of preprocessed image data based on DFFormer to obtain local scene features specifically includes: Generate filter weights adaptively based on the features of the input image to adapt to the scene detail features of various images; The overall architecture of DFFormer is based on the MetaFormer framework and features a layered design. Multi-stage downsampling, between each stage, convolution operations are used to achieve gradual downsampling of image scene features; Feature maps of multiple resolutions retain complete information from fine-grained to high-level semantics after layer-by-layer processing.

4. The scene understanding method of multi-level image feature extraction according to claim 1, characterized in that: The overall architecture of DFFormer is based on the MetaFormer framework and specifically includes: The first stage is low-level feature extraction: downsampling is performed using large convolution kernels and strides to extract low-level scene features; The middle stage is mid-level feature extraction: a multi-level MetaFormer block is introduced to gradually extract mid-level scene features of the image; High-level feature extraction: Dynamic filtering combined with deeper MetaFormer blocks is used to capture high-level scene semantic features; Each MetaFormer block accomplishes feature modeling through two core components: Token Mixer is a dynamic filter or a global filter: it dynamically models local or global scene information; MLP module: used to improve nonlinear expression capabilities and further integrate scene features.

5. The scene understanding method of multi-level image feature extraction according to claim 1, characterized in that: The global feature extraction based on MAE enhancement to obtain the global scene feature specifically includes: MAE self-supervised training: In the pre-training stage, parts of the input image are randomly masked, and the masked areas are reconstructed through the decoder; The encoder only extracts scene features from unmasked areas and learns contextual relationships and global dependency features. Global Feature Extraction: MAE-enhanced ViT extracts long-range global scene semantic information of images, especially in complex scenes, capturing object relationships and global layout in images.

6. The scene understanding method of multi-level image feature extraction according to claim 1, characterized in that: The local scene features and the global scene features are fused, The fusion features obtained specifically include: Convert the feature map generated by DFFormer to attention. First, use linear transformation operations to encode the features. Then, use activation functions such as Softmax to ensure that these scores can represent probability distributions and form attention weights. Then, integrate the attention mechanism guided by DFFormer features into the Vision Transformer (ViT) architecture. Choose to introduce the attention module after a specific layer of ViT. The specific implementation method is to multiply the generated attention map with the feature map output by the corresponding layer of ViT by channel. The fused features contain both global scene semantic information and fine-grained local scene features; The fused features contain both global scene semantic information and fine-grained local scene features.

7. The scene understanding method of multi-level image feature extraction according to claim 1, characterized in that: The classification output according to the fusion feature specifically includes: The fused features are output through a lightweight classification head as the input of the scene detection task; Multi-task support: output global scene features for classification, and output multi-scale feature maps for detection or segmentation tasks.

8. A scene understanding system for multi-level image feature extraction, using a scene understanding method for multi-level image feature extraction as described in any one of claims 1 to 7, characterized in that: The scene understanding system comprises: An image data preprocessing module is used to obtain input image data and perform data preprocessing to obtain preprocessed image data; A local feature extraction module based on DFFormer is used to extract fine-grained local scene features of preprocessed image data based on DFFormer to obtain local scene features; MAE enhanced global feature extraction module, used for extracting global features based on MAE enhancement to obtain global scene features; A feature fusion module, used to fuse the local scene feature with the global scene feature to obtain a fused feature; The feature output module is used to perform classification output according to the fusion feature.

Citation Information

Cited By

  • Image feature extraction method and device, and electronic equipment

    CN121010774A

  • Agricultural disease hash retrieval method based on DV stabilization and adaptive feature enhancement

    CN121278128A

  • Dv stabilized and adaptive feature enhanced hash retrieval method for agricultural diseases

    CN121278128B