Image processing method for simulating human eye visual perception mode
By simulating the image processing method of the human eye's visual perception mode, a visual transformer model is constructed. Combined with the dynamic position encoder and the selective fusion attention module, the problem of high computational complexity in the transformer model is solved, efficient multi-level feature interaction and fusion is achieved, and the accuracy and efficiency of image understanding tasks are improved.
Patent Information
- Application Number
- CN202510845426.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-03
AI Technical Summary
The existing transformer model has high computational complexity in high-resolution image processing. The self-attention mechanism leads to excessively high global computational complexity, making it difficult to achieve effective interaction and fusion of multi-level features, limiting the improvement of model efficiency and recognition accuracy.
An image processing method that simulates the visual perception mode of the human eye is adopted. By constructing a visual transformer model, combining a dynamic position encoder, a selective fusion attention module and a deep convolutional feedforward neural network, the rapid scanning, focusing on the area of interest and selective attention of the human eye are simulated to achieve efficient interaction and fusion of multi-level features.
It effectively solves the problem of high global computational complexity in the self-attention mechanism, significantly enhances feature extraction capabilities, improves the robustness and generalization capabilities of the model, achieves a balance between efficiency and recognition accuracy, and improves the accuracy of image understanding tasks.
Smart Images

Figure CN120747604A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and image processing, and in particular to an image processing method that simulates human visual perception patterns. Background Art
[0002] In recent years, transformer architectures have achieved remarkable success in transforming computer vision tasks dominated by convolutional neural networks. Many vision works based on transformer architectures have achieved more advanced performance on a range of downstream tasks, including image recognition and dense prediction tasks (e.g., object detection, instance segmentation, and semantic segmentation). The Vision Transformer is one of the earliest attempts. It directly captures the long-range visual dependencies of image features by applying the self-attention mechanism to images, demonstrating powerful global context modeling capabilities. However, as the image resolution increases, the computational complexity of the self-attention mechanism grows quadratically, which to some extent limits its application in high-resolution image feature processing.
[0003] The self-attention process is the primary cause of this application bottleneck. This mechanism blindly compares the similarity of individual image tokens within a segmented image with all other image tokens, resulting in significant computational redundancy. Recent research has focused on addressing the computational complexity of the self-attention mechanism, and a series of effective strategies have been proposed. Some approaches consider combining the self-attention module with downsampling techniques, replacing the standard self-attention module to reduce the computational complexity and memory burden of the Transformer model. Others, inspired by convolutional neural networks, propose adopting local attention to alleviate the computational complexity of the Transformer model. However, these approaches rely heavily on image features within a single window, limiting the semantic information reliance of the global feature modeling of the original self-attention mechanism. Recent approaches have indirectly achieved effective interaction between coarse-grained and fine-grained perception by alternating local and global attention. However, stacking layers of different attention modules fails to effectively promote information exchange between coarse-grained and fine-grained features at different layers. To address this issue, some methods attempt to enhance the information aggregation capabilities of each layer of the Transformer model by mimicking biological vision mechanisms. However, the main focus of these methods is still on the fusion of local and global features, while neglecting the effective integration of multi-level features in spatial and channel dimensions. Existing technologies still have obvious shortcomings in this regard, mainly due to the inherent limitations of existing biological visual perception modeling, which makes it difficult to achieve effective interaction and fusion of multi-level features, resulting in limited feature expression capabilities, thereby restricting the improvement of model efficiency and recognition accuracy. Therefore, how to design a multi-level feature interaction and fusion mechanism that is closer to human visual perception patterns and achieve an effective balance between efficiency and accuracy in the transformer model remains a technical challenge that needs to be solved urgently. Summary of the Invention
[0004] The purpose of the invention is to overcome the defects of the above-mentioned prior art and provide an image processing method that simulates the visual perception mode of the human eye. By combining complex neural network theory with modular design, it simulates the multi-level perception mechanism of human vision, constructs an efficient and robust image feature extraction model, effectively balances efficiency and recognition accuracy, and demonstrates excellent performance in a variety of image understanding tasks, and has stronger robustness and generalization capabilities.
[0005] In order to solve the above method problems, the present invention adopts the following technical solutions to achieve the goal:
[0006] An embodiment of the present invention provides an image processing method that simulates a human visual perception mode, including:
[0007] Obtain the target image to be processed;
[0008] Preprocessing the target image to obtain a standardized input image;
[0009] Constructing a visual transformer model, inputting the standardized input image into the visual transformer model for feature extraction, and obtaining a feature extraction result;
[0010] Constructing a task model corresponding to the target image understanding task, inputting the feature extraction result into the task model to obtain an initial prediction result;
[0011] Comparing the difference between the initial prediction result and the actual result to construct a loss function;
[0012] Performing gradient updates on the task model using a back-propagation algorithm based on the loss function to complete the overall training process of the task model;
[0013] The test target image to be predicted is input into the task model that has been trained to obtain an image that simulates the visual perception mode of the human eye.
[0014] Preferably, preprocessing the target image to obtain a standardized input image comprises:
[0015] Adjust the target image to the preset image size using bilinear interpolation method;
[0016] The same data augmentation method is applied to the target image adjusted to a preset image size to obtain a data-enhanced target image;
[0017] The target image after data augmentation is subtracted from the mean of all target images and then divided by the standard deviation to obtain a standardized input image.
[0018] Preferably, the visual transformer model is set up with four stages, each stage contains multiple unified visual transformer modules, and a downsampling module is set between two consecutive stages to perform feature aggregation and reduce the resolution, wherein each visual transformer module includes a dynamic position encoder, a selective fusion attention module and a deep convolutional feedforward neural network.
[0019] Preferably, the dynamic position encoder is used to dynamically integrate the position information of the two-dimensional image into all image blocks;
[0020] The selective fusion attention module includes the neighborhood attention module, the dynamic aggregation attention module and the selective attention module, which are used to simulate the rapid scanning, attention to the region of interest and selective attention in the human visual perception mode respectively;
[0021] The neighborhood attention module and the dynamic aggregation attention module respectively extract the local features and aggregation features of the image through a dual-path design, and then perform feature screening and focusing through the selection attention module;
[0022] Deep convolutional feedforward neural networks are used to provide nonlinear transformations and enhance local feature extraction, and image understanding tasks are performed based on the extracted features.
[0023] As an example, a visual transformer model is constructed, including:
[0024] Construct a dynamic positional encoder, which involves sliding the convolution kernel across the spatial dimensions of the input feature map and computing the weighted sum of a 3×3 neighborhood of pixels within the local region.
[0025] Constructing a selection fusion attention module includes building a dual-path structure combining a neighborhood attention module and a dynamic aggregation attention module, and further constructing a selection attention module.
[0026] The neighborhood attention module performs fixed window attention calculation centered on the query, and the dynamic aggregation attention module quickly focuses on the region of interest based on the dynamic aggregation attention mechanism;
[0027] The attention module is selected to perform weighted fusion of the outputs of the two paths in the spatial or channel dimension;
[0028] Build a deep convolutional feedforward neural network, add additional deep convolution operations, and generate the next layer of features through residual connections.
[0029] Preferably, feature extraction is performed, including:
[0030] The normalized input image is fed into a convolutional prior module to obtain the initial feature extraction result;
[0031] Inputting the initial feature extraction result into a downsampling module to obtain downsampled output image features;
[0032] The downsampled output image features are input into the selective fusion attention module, which dynamically integrates the position information of the input two-dimensional image features into all image tokens through the dynamic position encoder DPE;
[0033] By selecting the fusion attention module SFA, the three visual perception modes of the human eye are simulated: rapid scanning of input two-dimensional image features, focusing on the region of interest, and selective attention. Local features and focus features are extracted from image features respectively. These two features are used to quickly filter visual information in the spatial and channel dimensions, perform feature focusing, and obtain focused output features that remove information redundancy.
[0034] The focused output features are transformed nonlinearly through the deep convolutional feedforward neural network DWFFN and the local features are enhanced and extracted.
[0035] Preferably, the convolution prior module includes three 3×3 convolution modules, three batch normalization modules, and three nonlinear activation functions;
[0036] The downsampling module reduces the feature image resolution and increases the number of channels through a 3×3 overlapping convolution module and a layer normalization module.
[0037] Preferably, a task model corresponding to the target image understanding task is constructed, including:
[0038] Build a classification task model, using the visual transformer model as the basic architecture for feature extraction; connect a fully connected layer or convolutional layer as the classification module, and use the overall classification category of the target image as the output parameter of the classification task model;
[0039] Build an object detection task model, using the visual transformer model as a feature extractor; build an object detection module based on the visual transformer, and use the bounding box coordinates of each detection box and its corresponding class label as the output parameters of the object detection task model through regression;
[0040] Build a semantic segmentation task model, using the visual transformer model as the semantic segmentation backbone network; connect to a fully convolutional network (FCN) or DeepLabV3 architecture, and use pixel-level classification and semantic labels for each pixel as output parameters of the semantic segmentation task model;
[0041] Build an instance segmentation task model and use the visual transformer model as the feature extraction part of instance segmentation to obtain high-order visual features of the image. Use Mask R-CNN or a segmentation head based on the visual transformer to use each pixel-level mask and each pixel-level segmentation as the output parameters of the segmentation task model.
[0042] Construct an industrial defect detection task model, use the visual transformer model as the feature extraction part of the industrial defect detection task, and build a detection module based on the visual transformer to locate and classify defects on the surface of industrial products.
[0043] Preferably, the feature extraction result is input into the task model to obtain an initial prediction result, including:
[0044] The feature extraction results of the visual transformer model are input into the classification module in the image classification task model to obtain the initial category prediction results of the target image;
[0045] Input the feature extraction results of the visual transformer model into the object detection module in the object detection task model to obtain the initial object detection prediction results of the target image;
[0046] Input the feature extraction results of the visual transformer model into the semantic segmentation module in the semantic segmentation task model to obtain the initial semantic segmentation prediction results of the target image;
[0047] Input the feature extraction results of the visual transformer model into the instance segmentation module in the instance segmentation task model to obtain the initial instance segmentation prediction results of the target image;
[0048] The feature extraction results of the visual transformer model are input into the defect detection module in the industrial defect detection task model to obtain the initial defect detection prediction results of the target image.
[0049] Preferably, the task model is gradient updated using a back propagation algorithm according to the loss function to complete the overall training process of the task model, including:
[0050] Calculate the loss function value and use the backpropagation algorithm to calculate the gradient of the loss function relative to the model parameters;
[0051] Use the optimization algorithm to update the task model parameters according to the gradient and adjust the model to minimize the loss function value;
[0052] The training process of the task model is completed through multiple rounds of iterations.
[0053] Compared with the prior art, the present invention has the following beneficial effects:
[0054] The present invention provides an image processing method that simulates the visual perception pattern of the human eye, breaks through the inherent limitations of existing methods in simulating biological visual perception modeling, and effectively solves the technical bottleneck in balancing model efficiency and recognition accuracy. The present invention introduces a dynamic aggregation attention mechanism to focus on key areas of the image more efficiently and accurately, effectively solving the problem of high global computational complexity in the self-attention mechanism; at the same time, by introducing a selective attention module based on attention features, it is used to quickly screen out more effective visual information from two dimensions, namely spatial and channel, to achieve more accurate feature focusing, and effectively solve the information redundancy problem in the feature extraction process. In addition, by introducing a specially designed dual-path selective fusion attention module to simulate the focusing perception pattern of human vision, the feature extraction capability of the visual transformer model is significantly enhanced, while avoiding the accuracy loss that may be caused by removing the self-attention mechanism, thereby effectively solving the balance problem between model efficiency and recognition accuracy.
[0055] The present invention comprehensively and completely explores the three functions of the human eye's visual perception mode, and completely models them through a simple and effective method, realizing a universal image processing method that conforms to human eye vision. The present invention has stronger robustness and generalization ability, and achieves excellent performance in various image understanding tasks, significantly improving the accuracy of each task, and maintaining or improving model efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute an improper limitation of the present invention. In the drawings:
[0057] Figure 1 This is a flow chart of an image processing method simulating a human visual perception mode according to an embodiment of the present invention;
[0058] Figure 2 A schematic diagram of the overall architecture of a converter model according to an embodiment of the present invention;
[0059] Figure 3 Schematic diagram of the principles of various functional modules in a converter according to an embodiment of the present invention;
[0060] Figure 4 The figure is a schematic diagram of the application process of an embodiment of the present invention. DETAILED DESCRIPTION
[0061] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The exemplary embodiments and descriptions of the present invention are used to explain the present invention but should not be construed as limiting the present invention.
[0062] Technologies, methods, and apparatuses known to those skilled in the art are not discussed in detail herein, but should be considered as part of the specification.
[0063] As a specific embodiment of the present invention, Figure 1 As shown, the present invention provides an image processing method that simulates the visual perception mode of the human eye, which specifically includes the following steps:
[0064] Step S101: Acquire a target image to be processed.
[0065] Target images are images required for training and testing for a specific image understanding task, covering a variety of representative scene types. Specifically, target images include but are not limited to natural scenes and industrial scenes, such as roads, zoos, and beaches, as well as industrial environments such as metal surfaces with defects and object picking in factories. The diversity and complexity of target images enable them to effectively support the training and inference of deep learning models in various application areas. When describing target images below, they refer to a scene image in an image understanding task.
[0066] Step S102: pre-process the target image to obtain a standardized input image.
[0067] Specific preprocessing methods include:
[0068] Adjust the target image to the preset image size using bilinear interpolation method;
[0069] The same data augmentation method is applied to the target image adjusted to the preset image size to obtain the data-enhanced target image to ensure fairness in performance comparison;
[0070] The target image after data augmentation is subtracted from the mean of all target images and then divided by the standard deviation to obtain a standardized input image.
[0071] Step S103: construct a visual transformer model, input the standardized input image into the visual transformer model for feature extraction, and obtain a feature extraction result.
[0072] The present invention provides a novel visual transformer model based on the human visual perception mode. The visual transformer model is constructed in the Pytorch deep learning framework using a multi-level visual transformer architecture similar to the traditional one.
[0073] Specifically, the traditional multi-stage visual transformer architecture typically consists of an initial layer and four stages, similar in structure to a convolutional neural network. For an image classification task, the target image is first resized to 224×224 and subjected to data augmentation and normalization to obtain a standardized input image. This standardized input image is then divided into several image tokens, called visual tags, by the initial layer and passed as input features to the subsequent multi-stage feature extraction stage. Each stage contains multiple unified transformer modules, each composed of a multi-head self-attention module and a feedforward neural network. The feature extraction results of each stage are output to the next stage by stacking the transformer modules. A downsampling module is placed between stages to reduce the resolution of the feature map by merging visual tags while increasing the number of channels. The output features of the final transformer module are passed through a fully connected classification layer to obtain a probability distribution for predicting the target image's category.
[0074] Compared with the traditional multi-level transformer architecture, the present invention replaces the typical self-attention module with a novel selective fusion attention module based on a dual-path design. This module significantly enhances the feature extraction capability of the visual transformer model by simulating the different visual perception levels of the human eye's understanding of the scene, effectively solving the problems of excessive global computational complexity and redundant feature information in the self-attention module, while avoiding the loss of accuracy caused by replacing the self-attention module, thereby effectively solving the balance problem between model efficiency and recognition accuracy.
[0075] Specifically, each stage of the visual transformer model contains a different number of visual transformer modules, and three variants with different model sizes are proposed by scaling the network width (i.e., the number of channels) and depth (i.e., the number of blocks used in each stage). For example, for the smallest model variant DFViT-T, the number of visual transformer modules in each stage is 2, 2, 2, 2, and the number of channels is set to 32, 64, 128, and 256, respectively. For the smaller model variant DFViT-XS, the number of visual transformer modules in each stage is 2, 2, 9, 2, and the number of channels is set to 64, 128, 256, and 512, respectively. For the larger model variant DFViT-S, the number of visual transformer modules in each stage is 3, 4, 18, 4, and the number of channels is set to 64, 128, 256, and 512, respectively.
[0076] In the visual transformer modules, each visual transformer module includes a dynamic position encoder, a selective fusion attention module and a deep convolutional feedforward neural network.
[0077] Formally, given the input features of the lth visual transformer module, the input-output relationship of the entire visual transformer module can be expressed as follows:
[0078]
[0079] Among them, x l is the input feature of the lth visual transformer module, x l ′ is the intermediate output feature of the lth selected fusion attention module, x l+1 is the final output feature of the lth visual transformer module, LN is layer normalization, DPE is dynamic position encoder, SFA is selection fusion attention module, and DWFFN is deep convolutional feedforward neural network.
[0080] The construction of the visual transformer model specifically includes the following steps:
[0081] 31) Build a dynamic position encoder
[0082] In one embodiment, Figure 3 As shown in the figure, dynamic position encoding is achieved through 3×3 depthwise convolution and residual connections. The core of 3×3 depthwise convolution for position encoding is that it calculates the weighted sum of pixels in a local area (3×3 neighborhood) by sliding the convolution kernel across the spatial dimensions (H×W) of the input feature map, capturing the relative position relationship between pixels in that area. This local operation naturally introduces spatial position information because the convolution calculation depends on the spatial arrangement of pixels. The weight of the convolution kernel encodes the neighboring relationship of each position during the sliding process, thereby achieving implicit position awareness.
[0083] 32) Constructing a selection fusion attention module
[0084] Compared to the traditional standard self-attention module, this paper innovatively designs a visual transformer module based on the human visual perception model. During visual perception, the eye rapidly moves within a fixed range to capture the entire scene in its field of view and quickly locate regions of interest, thereby selectively focusing on these areas. During viewpoint shifts, fixation scanning and focusing on regions of interest typically occur simultaneously, and attention is often redistributed following this movement, a process known as attention switching. To implement this entire process, we adopt a dual-path design that combines neighborhood attention (NA) and dynamic aggregation attention (DFA) to design a tailored selective fusion attention block. Specifically, the selective fusion attention block comprises two parallel attention paths: the neighborhood attention path and the dynamic aggregation attention path. The neighborhood attention path performs fixed-window attention computation centered on the query. Meanwhile, the dynamic aggregation attention path uses the dynamic aggregation attention mechanism to rapidly focus on regions of interest. Furthermore, to achieve attention switching, we introduce a selective attention module that selectively fuses the output information of the two attention paths in spatial or channel dimensions, thereby comprehensively integrating multi-level visual information.
[0085] Formally, given the input features of the lth selective fusion attention module, the inputs of both paths are features after layer normalization. Then, through the selective attention module, the outputs of these two paths are weightedly fused in the spatial or channel dimension, and finally the features of the next layer are generated through residual connection. The specific operation is as follows:
[0086]
[0087] in, is the output feature of the layer normalization, NA is the neighborhood attention module, is the output feature of the neighborhood attention module, DFA is the dynamic aggregation attention module, is the output feature of the neighborhood attention module, and SAM is the selection attention module.
[0088] 32a) Constructing a dynamic aggregation attention module
[0089] Dynamic aggregated attention enables coarse-grained global-information perception of each query, enabling efficient focus on key regions. To simulate the rapid focus of attention across different regions achieved through perspective shifts during eye movements, a probability distribution of each token relative to global information is first calculated. In this way, the weights of the tokens are associated with global information, forming a preliminary attention guide for the global context. These probabilities are then aggregated and grouped, and self-attention is calculated within the groups. This grouping mechanism not only significantly reduces computational complexity but also introduces a sparse inter-token attention mechanism, allowing the model to efficiently capture local correlations between different tokens. Specifically, the key vector of the input features, obtained by linear transformation, is used as the basis for calculating the probability distribution and aggregation operations. To simplify computation, only the key vector is average pooled based on experience, and the result is used as a single global information vector.
[0090] Formally, given an input feature X, we first perform a linear transformation to obtain the key K, then perform average pooling to obtain a single global information G, and finally calculate the similarity score between the two. The similarity score is implemented through matrix multiplication, which is as follows:
[0091]
[0092] Among them, W K is a trainable weight matrix, A represents the attention map, with dimensions (b,n,1,h×w), where b is the batch size, n is the number of feature channels, and h×w is the spatial dimension. In this way, a single global information interacts with the features of each spatial position of the key vector to derive the corresponding attention weight.
[0093] Next, to stabilize numerical computation and transform the attention map into a probability distribution, we perform log-softmax normalization on the attention map:
[0094]
[0095] This operation maps the similarity score A to a log-probability space P, which represents the importance of each spatial position relative to the category embedding.
[0096] We then sort the normalized attention distribution to select the most spatially relevant tokens. We generate a sorted index by sorting the attention weights from low to high, with the goal of putting more important tokens at the top, making it easier to group them later. The specific operation is as follows:
[0097] S=argsort(P,dim=-1) (5)
[0098] Among them, argsort(*) is the sorting operator, S is the sorted index, and the dimension is (b,n,h×w), which represents the sorting order of each input token at different spatial positions.
[0099] After sorting, we group the tokens into k groups of m tokens each (assuming the number of tokens is divisible), as follows:
[0100] G1={x1,…,x m},G2={x m+1 ,…,x 2m},…,G k ={x ((k-1))m+1 ,…,x km} (6)
[0101] Among them, G k Group the k-th image token.
[0102] In order to capture the rich contextual dependencies between tokens within a group, we i Apply a multi-head self-attention mechanism. This mechanism can process multiple different subspace representations in parallel, thereby enhancing the model's expressive power. Finally, restore the tokens within each group to their original positions. The specific operation is as follows:
[0103]
[0104] Among them, X i represents the i-th input feature, Z represents the set of tokens restored to the global sequence, MHSA is the multi-head self-attention mechanism, Concat(·) is the concatenation operation, and InverseSort is the inverse sorting operation used to remap the grouped tokens to the global original token order.
[0105] 32b) Constructing a neighborhood attention module
[0106] Neighborhood attention expands the receptive field of each query point to its nearest neighboring pixels. By calculating attention weights within a fixed window around the query point, it focuses on information in local regions. Neighborhood attention effectively enhances the expressiveness of local features, enabling the model to better understand and process detailed information in an image. Specifically, the neighborhood attention mechanism defines a fixed-size window around each query location, limiting the scope of attention calculation, thereby reducing computational complexity and focusing on important local regions. In this way, the model is able to capture local contextual information and improve its perception of image details. For example, in image recognition tasks, neighborhood attention helps the model more accurately identify subtle features such as edges and textures, thereby improving overall recognition accuracy. Furthermore, the introduction of the neighborhood attention mechanism enhances the model's sensitivity to local variations, making it more adaptable to images with complex backgrounds or diverse objects. Combined with dynamic aggregated attention paths, neighborhood attention provides the model with rich multi-layered information, further improving its performance in various vision tasks.
[0107] 32c) Constructing the selective attention module
[0108] The selective attention module selectively applies channel-selective attention (SCA) or spatial-selective attention (SSA) in each layer to dynamically fuse the outputs of the two attention paths, thereby achieving comprehensive integration of multi-level visual information. Channel-selective attention can dynamically adjust the weights of each channel, effectively fusing features from the two paths in the channel dimension, and enhancing the model's ability to focus on important channels. Spatial-selective attention can dynamically adjust the weights of each spatial position, effectively fusing features from the two paths in the spatial dimension, and enhancing the model's ability to focus on important spatial regions.
[0109] For channel-selective attention, a 1×1 convolution is used to reduce the channel dimension, followed by batch normalization and activation function processing. Next, another 1×1 convolution layer is used to double the number of channels. A softmax function is then used to generate channel weights for the output features of the two paths. Finally, the normalized weights are applied to the channel dimension of the output features of the two paths, and the fused output features are obtained through weighted summation.
[0110] For spatial selective attention, a series of 1×1 convolutional layers are used to gradually reduce the number of channels in the input features. Each convolutional layer is followed by batch normalization and an activation function to enhance the nonlinear representation of the features. In particular, the final 1×1 convolutional layer reduces the number of channels to 2. A softmax function is then used to generate attention weights based on the spatial features. Finally, the normalized weights are applied to the spatial dimensions of the output features of the two paths, and the fused output features are obtained through a weighted summation.
[0111] By introducing this module, the transformer model can flexibly and selectively fuse attention information in the channel or spatial dimension in each layer according to the different input features. It not only enhances the expression of local and global features, but also improves the model's adaptive learning and generalization capabilities in complex visual tasks through a flexible weight adjustment mechanism.
[0112] 33) Building a deep convolutional feedforward neural network
[0113] Finally, the feedforward neural network in the traditional multi-stage converter architecture consists of two fully connected layers (FC) connected by a nonlinear activation function. In contrast, the deep convolutional feedforward neural network used in the present invention further improves the modeling effect of the feedforward neural network on local features by adding additional deep convolution operations. The operation of the deep convolutional feedforward network can be expressed as:
[0114]
[0115] Among them, X represents the output features of the selected fusion attention module, σ represents the nonlinear activation function, DWConv represents the deep convolution operation, FC represents the fully connected layer, and DWFFN represents the deep convolutional feedforward neural network.
[0116] The normalized input image is input to the visual transformer model for feature extraction, which specifically includes the following steps:
[0117] First, the normalized input image is fed into a convolutional prior module, which includes three 3×3 convolutional modules, three batch normalization modules, and three nonlinear activation functions to obtain the initial feature extraction results. Compared with the initial layer used in the traditional multi-level visual transformer architecture, although the number of parameters and computational complexity is slightly increased, it can stabilize the network optimization process and improve network performance; then, the initial feature extraction results are input into the downsampling module, which uses a 3×3 overlapping convolution module and a layer normalization module to reduce the feature image resolution and increase the number of channels to obtain the downsampled output image features; then, the downsampled output image features are input into the novel selective fusion attention module based on dual-path design. The downsampled output image features pass through the dynamic position encoder (DPE) in succession to dynamically integrate the position information of the input two-dimensional image features into all image tokens, and then pass through the selective fusion attention module. The attention module (SFA) simulates the three visual perception modes of the human eye: rapid scanning of input two-dimensional image features, focusing on areas of interest, and selective attention. These modes are implemented through neighborhood attention (NA), dynamic aggregation attention (DFA), and selective attention module (SAM). The module first extracts local features and focused features from image features, and then quickly filters more effective visual information in both spatial and channel dimensions to achieve more accurate feature focusing, significantly enhance image extraction capabilities, and obtain focused output features that remove information redundancy. Finally, the focused output features are nonlinearly transformed through a deep convolutional feedforward neural network (DWFFN) and the local features are further enhanced and extracted. The overall network structure of the visual transformer model is as follows: Figure 2 shown.
[0118] Step S104: construct a task model corresponding to the target image understanding task, input the feature extraction results into the task model, and obtain an initial prediction result.
[0119] In this implementation, we used the target image understanding task as an example. These tasks included, but were not limited to, image classification, object detection, instance segmentation, semantic segmentation, and industrial defect detection. The underlying architecture of all task models employed a visual transformer model, supplemented by various task modules, including classification, object detection, instance segmentation, semantic segmentation, and defect detection.
[0120] The specific method for constructing the task model is as follows:
[0121] 41a) Build a classification task model
[0122] First, the visual transformer model is used as the infrastructure (Backbone) for feature extraction; then, a fully connected layer or convolutional layer is connected as the classification module, and the overall classification category of the target image is used as the output parameter of the classification task model to complete the construction of the classification task model.
[0123] 42a) Build a target detection task model
[0124] First, a visual transformer model is used as a feature extractor. Then, a visual transformer-based object detection module is constructed, such as a variant of the YOLO or Faster R-CNN detection model. The bounding box coordinates of each detection box and its corresponding category label are used as the output parameters of the object detection task model through regression.
[0125] Furthermore, in this embodiment, the object detection module of the RetinaNet model is adopted, and the construction of the object detection task model is completed by replacing the basic architecture of the RetinaNet model with the visual transformer model.
[0126] 43a) Build a semantic segmentation task model
[0127] First, the visual transformer model is used as the backbone network for semantic segmentation; then, a fully convolutional network (FCN) or DeepLabV3 architecture is connected to it. Based on the features extracted by the visual transformer, pixel-level classification and the semantic label of each pixel are used as the output parameters of the semantic segmentation task model.
[0128] Furthermore, in this embodiment, the semantic segmentation module of the Semantic FPN model and the UperNet model is adopted, and the basic architecture of the Semantic FPN model and the UperNet model is replaced with the visual transformer model to complete the construction of the semantic segmentation task model.
[0129] 44a) Build an instance segmentation task model
[0130] First, a visual transformer model is used as the feature extraction part of instance segmentation to obtain high-order visual features of the image; then, through Mask R-CNN or a segmentation head based on the visual transformer, the pixel-level mask of each instance and the pixel-level segmentation of each detected instance are used as the output parameters of the instance segmentation task model.
[0131] Furthermore, in this embodiment, the instance segmentation module of the Mask R-CNN model is adopted, and the construction of the instance segmentation task model is completed by replacing the basic architecture of the Mask R-CNN model with the visual transformer model.
[0132] 45a) Building an industrial defect detection task model
[0133] First, a visual transformer model is used as the feature extraction part of the industrial defect detection task to ensure high expressiveness in extracting detailed features of industrial products. Subsequently, a detection module is constructed based on the visual transformer, which usually combines convolutional neural networks (CNN) and region extraction technology to locate and classify defects on the surface of industrial products.
[0134] Furthermore, in this embodiment, the defect detection module of the RetinaNet model is adopted, and the basic architecture of the RetinaNet model is replaced with a visual transformer model to complete the construction of the industrial defect detection task model.
[0135] The core of each task model is feature extraction based on the visual transformer model (as the backbone), followed by task-specific predictions through different task modules (such as classification and object detection). Each module uses different network layers and output methods based on the characteristics of its task to achieve an accurate understanding of the target image.
[0136] The feature extraction results are input into the task model to obtain the initial prediction results. The specific steps are as follows:
[0137] 41b) Image classification task
[0138] The feature extraction results of the visual transformer model are further input into the classification module in the image classification task model, and the classification category of the target image is output to obtain the initial category prediction result of the target image.
[0139] 42b) Object detection task
[0140] The feature extraction results of the visual transformer model are further input into the target detection module in the target detection task model, and the bounding box coordinates of each detection box in the target image and its corresponding category label are output to obtain the initial target detection prediction results of the target image.
[0141] 43b) Semantic segmentation task
[0142] The feature extraction results of the visual transformer model are further input into the semantic segmentation module in the semantic segmentation task model, which outputs the pixel-level classification of each object in the target image and the semantic label of each pixel to obtain the initial semantic segmentation prediction result of the target image.
[0143] 44b) Instance segmentation task
[0144] The feature extraction results of the visual transformer model are further input into the instance segmentation module in the instance segmentation task model, which outputs the pixel-level mask of each instance in the target image and the pixel-level segmentation of each detected instance to obtain the initial instance segmentation prediction result of the target image.
[0145] 45b) Industrial defect detection tasks
[0146] The feature extraction results of the visual transformer model are further input into the defect detection module in the industrial defect detection task model, and the location bounding box and defect category of each defect in the target image are output to obtain the initial defect detection prediction results of the target image.
[0147] Step S105 , comparing the difference between the initial prediction result and the actual result, and constructing a loss function.
[0148] The specific method is as follows:
[0149] The initial predictions of the target image obtained from several image understanding tasks are compared with the actual results. This comparison usually uses different evaluation criteria depending on the task type. Image classification tasks calculate the difference between the predicted category and the actual category, and commonly used comparison methods include cross-entropy loss. Object detection tasks use box regression loss (such as Smooth L1 loss or Intersection over Union loss) and classification loss to compare, respectively measuring the regression error of the bounding box and the prediction error of the target category. Semantic segmentation tasks use the cross-entropy loss function for pixel-level classification to calculate the difference between each pixel category. Instance segmentation tasks compare pixel-level masks with the actual segmentation results. Common loss functions include Dice loss, cross-entropy loss, and Intersection over Union loss. Industrial defect detection tasks use a combination of loss functions such as bounding box regression loss, classification loss, and localization error based on the prediction of defect location and category.
[0150] Step S106: Based on the loss function, the back propagation algorithm is used to perform gradient update on the task model to complete the overall training process of the task model.
[0151] The specific method is as follows:
[0152] First, the loss function is calculated and its gradient with respect to the model parameters is computed using the backpropagation algorithm. Next, an optimization algorithm (such as SGD or AdamW) is used to update the task model parameters based on the calculated gradients, adjusting the model to reduce the loss function. Finally, the task model is trained through multiple epochs, each consisting of multiple batches. During training, the task model parameters are gradually optimized to improve the model's performance in the target image understanding task.
[0153] Step S107: input the test target image to be predicted into the trained task model to obtain an image that simulates the visual perception mode of the human eye.
[0154] Furthermore, the entire training process and the inference process of the image to be predicted are performed on the GPU server to accelerate calculations and improve processing efficiency.
[0155] Typically, the target image understanding task requires first pre-training the model on the target image in the image classification task, and then applying the weights of the pre-trained model to other image understanding tasks to fine-tune the visual transformer model to evaluate the generalization ability of the visual transformer model.
[0156] The trained visual transformer model is then performance-verified and deployed in a specific environment. Deployment environments primarily include cloud, local, and edge environments, each suited to different application scenarios and requirements. Cloud environments leverage cloud computing platforms (such as AWS SageMaker, Google Cloud AIPlatform, and Microsoft Azure), supporting large-scale distributed deployment, high-concurrency request processing, and elastic scaling, making them suitable for scenarios requiring centralized management and high-performance computing. Local environments are primarily deployed on enterprise-owned servers or local hardware, enabling model inference using containerization tools (such as Docker) and service frameworks (such as FastAPI, Flask, or TorchServe). These environments are suitable for scenarios with high requirements for data privacy and security. Edge environments target resource-constrained embedded systems or mobile devices, optimizing inference performance through lightweight frameworks (such as TensorFlow Lite, PyTorch Mobile, or ONNX Runtime) to meet low-latency, offline prediction, or real-time computing requirements. These deployment environments provide flexible options for diverse model applications.
[0157] The following is a further detailed description of an image processing method simulating the human eye's visual perception mode according to the present invention through the accompanying drawings and specific embodiments, as follows:
[0158] like Figure 4 The figure shows the overall flow chart of this embodiment, which primarily includes a computer, an image acquisition device, a server, and the constructed transformer model. During actual model application, the computer controls the image acquisition device to acquire target images, while simultaneously completing model construction and testing. After being processed by the computer, the captured images are transmitted along with the model and related data to the server for training. After model training, the server performs data inference based on the trained model and returns the inference results to the computer for storage and visualization.
[0159] To further validate the effectiveness of the converter model presented in this paper, extensive experiments were conducted on various benchmark datasets representing common scenarios, as well as on real-world industrial datasets. The experimental results are shown in Tables 1, 2, 3, 4, 5, and 6. All performance comparisons were conducted using consistent training and testing parameters. The proposed model achieved optimal performance across a variety of benchmark datasets and real-world industrial datasets.
[0160] Table 1 Comparison of classification performance of ImageNet-1K image classification dataset
[0161]
[0162]
[0163] Table 2 Comparison of object detection performance on the COCO dataset
[0164]
[0165]
[0166] Table 3 Comparison of instance segmentation performance on the COCO dataset
[0167]
[0168]
[0169] Table 4 Comparison of semantic segmentation performance on the ADE20K dataset
[0170]
[0171] Table 5 Comparison of instance segmentation performance on industrial real-world robotic arm datasets
[0172]
[0173]
[0174] Table 6 Comparison of defect detection performance on the industrial relay defect detection dataset
[0175]
[0176] As shown in Table 1, the performance comparison of the present invention with other methods on the image classification task is demonstrated; as shown in Table 2, the performance comparison of the present invention with other methods on the target detection task is demonstrated; as shown in Table 3, the performance comparison of the present invention with other methods on the instance segmentation task is demonstrated, and Table 5 shows the instance segmentation performance comparison of the present invention with other methods in the industrial robot arm picking scenario, in order to verify the transfer learning ability of the converter model; as shown in Table 4, the performance comparison of the present invention with other methods on the semantic segmentation task is demonstrated; finally, as shown in Table 6, the performance comparison of the present invention with other methods on the industrial relay defect detection task is demonstrated.
[0177] It can be seen from the above experiments and embodiments that the present invention provides an image processing method that simulates the visual perception pattern of the human eye, breaks through the limitations of the existing technology in simulating biological visual perception modeling, and solves the technical bottleneck in balancing model efficiency and recognition accuracy. By introducing a dynamic aggregation attention mechanism, the complexity of global calculations in the self-attention mechanism is significantly reduced, while the focusing efficiency and accuracy of key areas of the image are improved. In addition, the present invention also adopts a selective attention module based on attention features to quickly screen and focus on effective visual information in the spatial and channel dimensions, solving the problem of information redundancy in feature extraction. Combined with a specially designed dual-path selection fusion attention module, the present invention effectively simulates the focusing perception pattern of human vision, enhances the feature extraction capability of the visual transformer model, and avoids the loss of accuracy that may result from removing the self-attention mechanism.
[0178] In summary, this paper comprehensively models the three core functions of human visual perception, proposing a simple and effective image processing method with enhanced robustness and generalization capabilities. This method excels in a variety of image understanding tasks, significantly improving task accuracy while effectively maintaining or improving model efficiency. Therefore, this paper is not only theoretically innovative but also has broad practical application prospects.
[0179] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. An image processing method simulating the visual perception mode of the human eye, characterized in that: include: Obtain the target image to be processed; Preprocessing the target image to obtain a standardized input image; Constructing a visual transformer model, inputting the standardized input image into the visual transformer model for feature extraction, and obtaining a feature extraction result; Constructing a task model corresponding to the target image understanding task, inputting the feature extraction result into the task model to obtain an initial prediction result; Comparing the difference between the initial prediction result and the actual result to construct a loss function; Performing gradient updates on the task model using a back-propagation algorithm based on the loss function to complete the overall training process of the task model; The test target image to be predicted is input into the task model that has been trained to obtain an image that simulates the visual perception mode of the human eye.
2. The image processing method simulating the human eye visual perception mode according to claim 1, characterized in that: Preprocessing the target image to obtain a standardized input image includes: Adjust the target image to the preset image size using bilinear interpolation method; The same data augmentation method is applied to the target image adjusted to a preset image size to obtain a data-enhanced target image; The target image after data augmentation is subtracted from the mean of all target images and then divided by the standard deviation to obtain a standardized input image.
3. The image processing method simulating the human eye visual perception mode according to claim 1, characterized in that: The visual transformer model is set up in four stages, each stage contains multiple unified visual transformer modules. A downsampling module is set between two consecutive stages to aggregate features and reduce the resolution. Each visual transformer module includes a dynamic position encoder, a selective fusion attention module and a deep convolutional feedforward neural network.
4. The image processing method simulating the human eye visual perception mode according to claim 3, characterized in that: The dynamic position encoder is used to dynamically integrate the position information of the two-dimensional image into all image blocks; The selective fusion attention module includes the neighborhood attention module, the dynamic aggregation attention module and the selective attention module, which are used to simulate the rapid scanning, attention to the region of interest and selective attention in the human visual perception mode respectively; The neighborhood attention module and the dynamic aggregation attention module respectively extract the local features and aggregation features of the image through a dual-path design, and then perform feature screening and focusing through the selection attention module; Deep convolutional feedforward neural networks are used to provide nonlinear transformations and enhance local feature extraction, and image understanding tasks are performed based on the extracted features.
5. The image processing method simulating the human eye visual perception mode according to claim 1, characterized in that: Build a visual transformer model, including: Construct a dynamic positional encoder, which involves sliding the convolution kernel across the spatial dimensions of the input feature map and computing the weighted sum of a 3×3 neighborhood of pixels within the local region. Constructing a selective fusion attention module, including building a dual-path structure combining a dynamic aggregation attention module and a neighborhood attention module, and further constructing a selective attention module; The neighborhood attention module performs fixed window attention calculation centered on the query, and the dynamic aggregation attention module quickly focuses on the region of interest based on the dynamic aggregation attention mechanism; The attention module is selected to perform weighted fusion of the outputs of the two paths in the spatial or channel dimension; Build a deep convolutional feedforward neural network, add additional deep convolution operations, and generate the next layer of features through residual connections.
6. The image processing method simulating the human eye visual perception mode according to claim 1, characterized in that: Perform feature extraction, including: The normalized input image is fed into a convolutional prior module to obtain the initial feature extraction result; Inputting the initial feature extraction result into a downsampling module to obtain downsampled output image features; The downsampled output image features are input into the selective fusion attention module, which dynamically integrates the position information of the input two-dimensional image features into all image tokens through the dynamic position encoder DPE; By selecting the fusion attention module SFA, the three visual perception modes of the human eye are simulated: rapid scanning of input two-dimensional image features, focusing on the region of interest, and selective attention. Local features and focus features are extracted from image features respectively. These two features are used to quickly filter visual information in the spatial and channel dimensions, perform feature focusing, and obtain focused output features that remove information redundancy. The focused output features are transformed nonlinearly through the deep convolutional feedforward neural network DWFFN and the local features are enhanced and extracted.
7. The image processing method simulating the human eye visual perception mode according to claim 1, characterized in that: The convolution prior module includes three 3×3 convolution modules, three batch normalization modules, and three nonlinear activation functions; The downsampling module reduces the feature image resolution and increases the number of channels through a 3×3 overlapping convolution module and a layer normalization module.
8. The image processing method simulating the human eye visual perception mode according to claim 1, characterized in that: Build a task model corresponding to the target image understanding task, including: Build a classification task model, using the visual transformer model as the basic architecture for feature extraction; connect a fully connected layer or convolutional layer as the classification module, and use the overall classification category of the target image as the output parameter of the classification task model; Build an object detection task model, using the visual transformer model as a feature extractor; build an object detection module based on the visual transformer, and use the bounding box coordinates of each detection box and its corresponding class label as the output parameters of the object detection task model through regression; Build a semantic segmentation task model, using the visual transformer model as the semantic segmentation backbone network; connect to a fully convolutional network (FCN) or DeepLabV3 architecture, and use pixel-level classification and semantic labels for each pixel as output parameters of the semantic segmentation task model; Build an instance segmentation task model and use the visual transformer model as the feature extraction part of instance segmentation to obtain high-order visual features of the image. Use Mask R-CNN or a segmentation head based on the visual transformer to use each pixel-level mask and each pixel-level segmentation as the output parameters of the segmentation task model. Construct an industrial defect detection task model, use the visual transformer model as the feature extraction part of the industrial defect detection task, and build a detection module based on the visual transformer to locate and classify defects on the surface of industrial products.
9. The image processing method simulating the human eye visual perception mode according to claim 1, characterized in that: Inputting the feature extraction results into the task model to obtain an initial prediction result includes: The feature extraction results of the visual transformer model are input into the classification module in the image classification task model to obtain the initial category prediction results of the target image; Input the feature extraction results of the visual transformer model into the object detection module in the object detection task model to obtain the initial object detection prediction results of the target image; Input the feature extraction results of the visual transformer model into the semantic segmentation module in the semantic segmentation task model to obtain the initial semantic segmentation prediction results of the target image; Input the feature extraction results of the visual transformer model into the instance segmentation module in the instance segmentation task model to obtain the initial instance segmentation prediction results of the target image; The feature extraction results of the visual transformer model are input into the defect detection module in the industrial defect detection task model to obtain the initial defect detection prediction results of the target image.
10. The image processing method simulating the human eye visual perception mode according to claim 1, characterized in that: Performing a gradient update on the task model using a back propagation algorithm according to the loss function, including: Calculate the loss function value and use the backpropagation algorithm to calculate the gradient of the loss function relative to the model parameters; Use the optimization algorithm to update the task model parameters according to the gradient and adjust the model to minimize the loss function value; The training process of the task model is completed through multiple rounds of iterations.