Multi-resolution fusion image recognition method based on human vision mode

Through a multi-resolution fusion image recognition method based on human visual mode, combined with a convolutional neural network and a support vector machine model, the problems of insufficient fusion of multi-level information and difficulty in optimizing feature weights in image recognition technology are solved, and high accuracy and robust image recognition effect is achieved.

CN120125902APending Publication Date: 2025-06-10FIRST INSTITUTE OF OCEANOGRAPHY MNR +1

Patent Information

Application Number
CN202510233135.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing image recognition technology is difficult to fully capture multi-level information of images at different resolutions, resulting in insufficient fusion of detailed features and global information, especially under the influence of complex backgrounds or noise, and it is difficult to optimize feature weights.

Method used

A multi-resolution fusion image recognition method based on human visual mode is adopted, and image recognition and classification are combined with support vector machine models through image acquisition and preprocessing, multi-resolution image generation, convolutional neural network feature extraction, weighted fusion and human visual attention mechanism optimization.

Benefits of technology

It significantly improves the accuracy and robustness of image recognition, fully retains the details and global information of the image, optimizes the weight and attention of features, and improves the generalization ability and classification effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125902A_ABST
    Figure CN120125902A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, in particular to a multi-resolution fusion image recognition method based on a human vision mode, and the method comprises the steps: obtaining a target image through image collection equipment, carrying out the preprocessing of the image, carrying out the processing of the preprocessed image according to different resolutions, and generating images of a plurality of resolutions; the method comprises the following steps of: extracting features of images with different resolutions through a convolutional neural network, performing weighted fusion on the extracted multi-resolution image features, and then introducing a human visual attention mechanism to optimize the fused features, thereby improving the precision of image recognition. And finally, inputting the optimized multi-resolution features into a support vector machine model, performing image recognition and classification, further removing noise and enhancing classification confidence through a post-processing step, and outputting a final image recognition result. According to the method, the image recognition precision in a complex scene can be effectively improved, and the method has relatively high robustness and is suitable for various image classification tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular, to a multi-resolution fusion image recognition method based on the human visual pattern. Background Art

[0002] With the rapid development of computer vision technology, image recognition has been widely applied in many fields such as autonomous driving, medical image analysis, security monitoring, and industrial inspection. In these applications, the core of the image recognition task is to extract effective features from the input image and perform accurate classification.

[0003] Although multi-resolution image processing has become one of the effective methods to improve the accuracy of image recognition, there are still some key problems in the existing image recognition technologies. First, most traditional image processing methods rely on single-resolution image input, ignoring the multi-level information of images at different resolutions. This way cannot fully capture the detailed features and global information of images. Especially when the important regions or details in the image are relatively blurred, the classification accuracy may be significantly reduced. Second, feature fusion is a difficult point in multi-resolution image recognition. How to effectively fuse features at different scales in multi-resolution images so that the information at each scale can be complementary and maximize the classification accuracy is still an unsolved technical problem. In addition, although convolutional neural networks perform well in feature extraction, how to optimize the weights of features and the attention to important features under the influence of complex backgrounds or noises is still a key challenge. Summary of the Invention

[0004] Based on the above purposes, the present invention provides a multi-resolution fusion image recognition method based on the human visual pattern.

[0005] A multi-resolution fusion image recognition method based on the human visual pattern includes the following steps: S1, Image acquisition and preprocessing: Obtain the target image through an image acquisition device, and perform preprocessing on the image. The preprocessing includes denoising and enhancing the contrast; S2, Multi-resolution image generation: Process the preprocessed target image at different resolutions to generate images at multiple resolutions; S3, Multi-resolution image feature extraction: Respectively perform feature extraction on the images at multiple resolutions through a convolutional neural network. The extracted features include edge features, texture features, shape features, and color features; S4, Feature fusion: Perform weighted fusion on the multi-resolution image features extracted in S3 to generate a multi-resolution feature set; S5, Introduction of attention mechanism: On the basis of feature fusion, introduce the human visual attention mechanism to optimize the multi-resolution feature set; S6, Image Recognition and Classification: Input the multi-resolution feature set into the support vector machine model for image recognition and classification; S7, Post-processing and Result Output: Post-process the recognition result and output the final image recognition result.

[0006] Optionally, S1 includes: S11, Image Acquisition: Obtain the target image through an image acquisition device. The target image is the original image to be recognized. The image acquisition device includes an infrared camera and an RGB camera, and the specific selection depends on the requirements of the image recognition task; S12, Image Denoising: Denoise the obtained target image to reduce random noise in the image; S13, Contrast Enhancement: Enhance the contrast of the denoised target image.

[0007] Optionally, S2 includes: S21, Select Resolution Level: Determine multiple resolution levels according to the requirements of the image recognition task; S22, Image Scaling: Scale the preprocessed target image according to the selected resolution level; S23, Resolution Image Generation: Generate multiple images with different resolutions based on the scaling process; S24, Image Quality Assessment: Assess the quality of each generated image with a different resolution.

[0008] Optionally, S3 includes: S31, Select Convolutional Neural Network Structure: To extract effective features from images with different resolutions, select a convolutional neural network based on the ResNet-50 structure; S32, Image Input: Input images with different resolutions into the convolutional neural network for processing to extract their respective features.

[0009] Optionally, S4 includes: S41, Feature Weighted Fusion: Use a weighted fusion algorithm to fuse features with different resolutions to generate a multi-resolution feature set; S42, Dimension Adjustment: After feature fusion, the generated multi-resolution feature set may have a high dimension. Therefore, use the principal component analysis method to reduce the dimension of the multi-resolution feature set for subsequent classification and recognition tasks. After dimension reduction, generate the final multi-resolution feature set.

[0010] Optionally, S5 includes:; S51, Selective attention mechanism: To introduce the attention mechanism, a combination of channel attention mechanism and spatial attention mechanism is adopted to adaptively select the most critical feature parts from two dimensions (i.e., channel and spatial), thereby optimizing the feature representation; S52, Fusion and weighting of attention mechanism: After introducing the channel and spatial attention mechanisms, these two parts of attention information are combined to optimize the representation of the feature map.

[0011] Optionally, S6 includes: S61, Feature processing and model training: Construct and train a support vector machine model, and process the multi-resolution feature set; S62, Support vector machine model prediction: Through the trained support vector machine model, predict the input feature vector to identify the image category.

[0012] Optionally, median filtering and Gaussian filtering are used for post-processing in S7 to denoise the recognition result and improve the classification accuracy.

[0013] Optionally, after denoising the recognition result, the final image recognition result is output by enhancing the classification confidence.

[0014] Advantages of the present invention: In the present invention, by combining multi-resolution image processing and convolutional neural network feature extraction, the accuracy of image recognition is significantly improved. First, in the image preprocessing stage, through steps such as denoising, enhancing contrast, and equalization, the quality of the image is effectively improved, noise and unnecessary interference are removed, image details are enhanced, and the efficiency and accuracy of subsequent feature extraction are ensured. Through the generation of multi-resolution images, the feature information of the image at different scales can be captured, which enables the full retention of the detail information and global information of the image during fusion. Second, using a convolutional neural network for feature extraction not only effectively extracts multi-dimensional features such as edges, textures, shapes, and colors, but also adaptively learns the non-linear relationships in the image using deep learning algorithms, greatly enhancing the representation ability of the image and improving the recognition accuracy.

[0015] In the aspect of feature fusion, the present invention optimizes multi - resolution image features through a weighted fusion method, generating a multi - resolution feature set. These feature sets avoid redundancy while retaining information. A dynamic weighting mechanism is adopted to weight according to the importance of each resolution image, ensuring that key features receive higher weights and improving the model's ability to focus on important details. In addition, the introduction of the human visual attention mechanism enables the model to automatically focus on important regions and features in the image, suppressing irrelevant or noisy information and further enhancing the accuracy of classification. By combining channel attention and spatial attention mechanisms, the present invention simulates the selective attention of the human visual system to images, enabling the model to classify images more accurately, especially in the recognition of complex backgrounds and detailed regions, improving the classification effect and generalization ability.

[0016] In the present invention, by introducing a support vector machine model, the present invention not only improves the processing speed but also optimizes the credibility of the classification results. Through normalization processing and optimized feature input, the support vector machine model can better process high - dimensional feature data, improve classification efficiency, and reduce the risk of overfitting. The post - processing steps include denoising and enhancing classification confidence, further optimizing the quality of the classification results. During the classification post - processing, methods such as median filtering and Gaussian filtering are used to remove noise, ensuring the accuracy and stability of the classification results. And the confidence enhancement improves the reliability of classification decisions through smoothing and correction, and the finally output classification results have higher credibility. Brief Description of the Drawings

[0017] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention; Figure 2 It is a schematic flowchart of the S2 process according to an embodiment of the present invention. Detailed Embodiments

[0019] The present invention will be described in detail below in conjunction with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well - known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the drawing part is only for more specific description of the embodiments and is not intended to specifically limit the present invention.

[0020] Such as Figure 1 - Figure 2As shown in the figure, a multi-resolution fusion image recognition method based on the human visual pattern includes the following steps: S1, Image acquisition and preprocessing: Obtain the target image through an image acquisition device, and perform preprocessing on the image. The preprocessing includes denoising and enhancing the contrast; S2, Multi-resolution image generation: Process the preprocessed target image at different resolutions to generate images of multiple resolutions; S3, Multi-resolution image feature extraction: Respectively perform feature extraction on the images of multiple resolutions through a convolutional neural network. The extracted features include edge features, texture features, shape features, and color features; S4, Feature fusion: Perform weighted fusion on the multi-resolution image features extracted in S3 to generate a multi-resolution feature set; S5, Introduction of attention mechanism: On the basis of feature fusion, introduce the human visual attention mechanism to optimize the multi-resolution feature set; S6, Image recognition and classification: Input the multi-resolution feature set into a support vector machine model for image recognition and classification; S7, Post-processing and result output: Perform post-processing on the recognition result and output the final image recognition result.

[0021] S1 includes: S11, Image acquisition: Obtain the target image through an image acquisition device. The target image is the original image to be recognized. The image acquisition device includes an infrared camera and an RGB camera, and the specific selection depends on the requirements of the image recognition task; S12, Image denoising: Perform denoising processing on the obtained target image to reduce the random noise in the image. The specific steps are as follows: Gaussian filtering: Use a two-dimensional Gaussian filter to perform convolution processing on the target image to remove the high-frequency noise in the target image. The convolution kernel formula of the Gaussian filter is: ; Among them, is the standard deviation of the Gaussian distribution, which controls the smoothing intensity of the filtering. x and y are the horizontal and vertical distances between the current pixel and the neighboring pixels; Median filtering: Eliminate salt-and-pepper noise by replacing each pixel value with the median of the pixels in its neighborhood, which helps to retain the edge information of the image and remove outliers at the same time; Bilateral filtering: Combine the pixel spatial information and the pixel value difference, and use bilateral filtering to smooth the target image. The bilateral filtering formula is: ; Among them, is the neighboring pixel value, is the Gaussian function in the spatial domain, which controls the spatial range, It is a Gaussian function based on pixel value differences to control similarity; By denoising through the above method, the details and edge information of the image are retained while the noise is removed; S13, Enhance contrast: Enhance the contrast of the denoised target image to improve the visibility of important features in the image, specifically including: S131, Histogram equalization: Improve the contrast of the image by stretching the range of the gray-scale distribution of the image. The histogram equalization method adjusts the gray-scale distribution of the image to make the histogram distribution of the image more uniform and reduce the compression of the brightness interval. The steps of this method are: (1) Calculate the gray-scale histogram of the image; (2) Calculate the cumulative distribution function (CDF), and the formula is: ; Among them, is the number of pixels with gray value j, is the cumulative distribution function; (3) Map each pixel value, and the new pixel value is: ; Among them, is the minimum value of the cumulative distribution function, M and N are the width and height of the image, and L is the number of gray levels; S132, Logarithmic transformation: Enhance the low-gray regions in the image by performing a logarithmic transformation on the gray values of the image, making the dark details of the image clearer, expressed as: ; Among them, is the gray value of the original image, is the gray value of the enhanced image, and c is a constant used to adjust the image brightness; Through the image acquisition and preprocessing process, combined with the denoising and contrast enhancement processing methods, the image quality is optimized, making the image clearer, with higher contrast and richer details, providing higher-quality input for subsequent image feature extraction and image recognition, and improving the accuracy and robustness of the entire image recognition process.

[0022] S2 includes: S21, Select the resolution level: According to the requirements of the image recognition task, determine multiple resolution levels. Common resolution levels include low resolution, medium resolution, and high resolution. The specific resolution levels can be set according to the detail requirements of the image and the processing ability. For example, the low resolution is 128×128 pixels, the medium resolution is 256×256 pixels, and the high resolution is 512×512 pixels. The setting of the resolution level is flexibly adjusted according to the application scenario and the complexity of the image features; S22, Image scaling: Scale the preprocessed target image according to the selected resolution level; S23, Resolution image generation: Based on the scaling process, generate multiple images with different resolutions. Each resolution image retains the main features of the original image but presents the image content with different precisions to meet the requirements of different recognition tasks. For example, the low-resolution image is used to capture the overall shape and general structure, the medium-resolution image is used to capture detailed information, and the high-resolution image is used for precise detail recognition; S24, Image quality assessment: Perform quality assessment on each generated image with different resolutions. The quality assessment methods include calculating the clarity and contrast of the image to ensure that each resolution image meets the requirements in terms of detailed information and global structure.

[0023] S3 includes: S31, Select the convolutional neural network structure: In order to extract effective features from images with different resolutions, select a convolutional neural network based on the ResNet-50 structure. ResNet-50 is a deep residual network with 50 layers. By introducing residual connections, it effectively alleviates the problem of gradient disappearance in deep neural networks and enhances the feature extraction ability. This network structure is suitable for processing multi-scale images and can extract multi-level and multi-dimensional image features from images with low resolution to high resolution; The ResNet-50 network structure includes the following layers: Input layer: The input layer accepts the preprocessed multi-resolution images. The size of the input image depends on the specific resolution, such as 128×128, 256×256, or 512×512. All images will be normalized, and the pixel values are mapped to the interval [0,1]; Convolutional layer: The first convolutional layer of this network applies 64 3×3 convolutional kernels to perform convolution operations on the input image. The size of the convolutional kernel is 3×3, the stride is 1, and the padding is 1 to ensure that the image size remains unchanged after convolution, which is expressed as: ; where I is the input image, K is the convolutional kernel, b is the bias term, and i, j are the indices of the output image; Residual Block: The core structure of ResNet-50 is residual learning. By using skip connections, the input is directly added to the output to solve the vanishing gradient problem in deep networks. The calculation formula for each residual block is: ; where, represents the convolutional operation of the residual block, x is the input of the skip connection, which is directly added to the convolutional result to form the final output; Pooling Layer: Between every two convolutional layers, a max pooling layer is inserted to reduce the spatial size of the feature map while retaining the most important feature information. The max pooling operation is performed through the following formula: ; where, is the pixel value within the pooling region. The size of the pooling window is 2×2 and the stride is 2; Fully Connected Layer: Finally, the feature map is processed through the fully connected layer to output a multi-dimensional feature vector; S32, Image Input: Images of different resolutions are input into the convolutional neural network for processing to extract their respective features. The specific operations are as follows: Edge Feature Extraction: Through the low-level convolutional layer, the convolutional neural network can extract the edge features of the image. The patterns learned by the convolutional kernel will respond to the structural information such as edges and contours in the image; Texture Feature Extraction: Through the middle-level convolutional layer, the convolutional neural network will identify the texture features in the image, such as details and patterns; Shape Feature Extraction: Through the high-level convolutional layer, the convolutional neural network will learn the approximate shape and contour information in the image; Color Feature Extraction: By inputting the image into multiple convolutional kernels, the color information in the image is extracted. Especially in the RGB space, color features are very useful for image classification.

[0024] S4 includes: S41, Feature Weighted Fusion: A weighted fusion algorithm is used to fuse the features of different resolutions to generate a multi-resolution feature set. The specific operations are: S412, Calculate the Feature Importance of Each Resolution Image: Using the content information of the image (such as edge sharpness, texture complexity, etc.), calculate the importance of each resolution image in the feature fusion process. The importance index can be calculated through information-based metrics (such as information gain, entropy, etc.); S413, Dynamic Weighting Mechanism: For the image features of each resolution, different weights are assigned through the dynamic weighting mechanism. The assignment of this weight is based on the detail information and importance of each resolution image and is calculated using an adaptive weighting algorithm. The specific weighting process is: ; Among them, is the feature of the i-th resolution image, is the weight assigned to the i-th image feature, and satisfies the weight normalization condition: ; The calculation of the weight can be achieved through the following steps: Calculate the feature importance of each resolution image (such as information gain, feature sensitivity, etc.) Normalize these importance values to the interval [0, 1] as the weighting coefficient of each resolution image feature; S414, Feature fusion operation: After calculating the weights of each resolution image feature, all image features are weighted and superimposed according to the weights. After weighting the image features of each resolution (such as edge features, texture features, etc.), they are merged into a unified multi-resolution feature set according to the weight ratio. The feature set after weighted fusion can fully fuse the low-level features and high-level features in images of different resolutions, retain the global and local information of the image, and improve the accuracy of image recognition; S42, Dimension adjustment: After feature fusion is completed, the generated multi-resolution feature set may have a high dimension. Therefore, the principal component analysis method is used to reduce the dimension of the multi-resolution feature set for subsequent classification and recognition tasks. After dimension reduction, the final multi-resolution feature set is generated.

[0025] S5 includes: After feature fusion, the generated multi-resolution feature set contains rich information in the image. However, some regions of these features may be more critical for the image recognition task, while other regions contribute less to the recognition. For this reason, the human visual attention mechanism is introduced, aiming to simulate the focusing behavior of the human visual system on important regions, and further optimize the image features by enhancing important features and suppressing unimportant features, improving the accuracy and efficiency of image recognition; S51, Selective attention mechanism: To introduce the attention mechanism, a combination of channel attention mechanism and spatial attention mechanism is adopted to adaptively select the most critical feature parts from two dimensions (i.e., channel and space), thereby optimizing the feature representation. The specific steps include: S511, Channel attention mechanism: In a convolutional neural network, the output feature map of each layer contains multiple channels (i.e., different feature dimensions). The channel attention mechanism enhances the attention to key channel features by adaptively adjusting the importance of each channel. In this step, calculate the attention weight of each channel and use the weighted average method to weight the channels. The specific algorithm is as follows: (1) Calculate the average feature response of each channel, expressed as: ; Among them, represents the feature map value at position , H and W are the height and width of the feature map, is the average response value of the channel; (2) Calculate the channel weight, expressed as: ; Among them, is the attention weight of channel c, is the weight parameter obtained through training, is the sigmoid function; S512, Spatial Attention Mechanism: The spatial attention mechanism further optimizes the spatial distribution of the feature map by evaluating the importance of each position in the image space. The spatial attention mechanism dynamically adjusts the spatial weight of the feature map based on the information intensity of each spatial position in the input feature map. The specific algorithm is as follows: (1) Calculate the feature response of each spatial position, expressed as: ; Among them, represents the spatial response value at position , C is the number of channels of the feature map; (2) Calculate the spatial attention weight, expressed as: ; Among them, is the attention weight of spatial position , is the weight parameter obtained through training, is the sigmoid activation function; S52, Fusion and Weighting of Attention Mechanisms: After introducing the channel and spatial attention mechanisms, combine these two parts of attention information to optimize the representation of the feature map. The specific operation is as follows: Weight the channels of the feature map: Apply the channel attention weight to the feature map of each channel, expressed as: ; Among them, is the value of the weighted feature map, is the channel attention weight; Weight the space of the feature map: Apply the spatial attention weight to the feature of each spatial position, expressed as: ; Among them, is the final feature map after channel and spatial attention weighting; The multi-resolution feature map optimized by introducing the attention mechanism can adaptively enhance the features of key regions in the image while suppressing irrelevant or noisy regions, thereby improving the accuracy and robustness of image recognition. Finally, the multi-resolution feature map optimized by the attention mechanism will be used in subsequent image recognition and classification steps.

[0026] S6 includes: S61, Feature processing and model training: Construct and train a support vector machine model, and process the multi-resolution feature set. The specific steps are as follows: S611, Feature flattening and vectorization: Flatten the multi-resolution feature set extracted and fused from images of different resolutions and convert it into a one-dimensional feature vector. Let each feature set be , , , , where represents the nth feature dimension, and the final feature vector X is expressed as: ; S612, Data standardization: Since the support vector machine is sensitive to the scale of input features, the feature vector is standardized to have zero mean and unit variance. The standardization process is as follows: ; where X is the original feature vector, is the mean of all training sample features, q is the standard deviation, is the standardized feature vector; S613, Support vector machine model training: Use the standardized feature vector as input to train the support vector machine model. The goal of the support vector machine model is to find an optimal hyperplane by maximizing the inter-class margin, which can effectively distinguish images of different classes. Specifically, the support vector machine model trains the model by solving the following optimization problem: ; where w is the normal vector of the classification hyperplane, b is the bias term, is the feature vector of the ith sample, is the corresponding label (class); This optimization problem aims to find an optimal classification hyperplane to maximize the margin between classes; S62, Support vector machine model prediction: Use the trained support vector machine model to predict the input feature vector and identify the image class. The specific steps are as follows: Input test data: Input the feature vector of the image to be classified into the trained support vector machine model for classification prediction; Calculating the decision function: For the input feature vector , the support vector machine model calculates the decision function , which is used to determine the class to which the sample belongs. The calculation formula of the decision function is: ; where w and b are parameters obtained from the training process, is the feature vector of the image to be classified; Classification result: According to the value of the decision function, the image is assigned to the class with the maximum decision value. If >0, the image is classified as the positive class. If <0, the image is classified as the negative class. For multi-class classification problems, the "one-vs-one" or "one-vs-all" strategy can be used.

[0027] In S7, median filtering and Gaussian filtering are used for post-processing to denoise the recognition result and improve the classification accuracy; After the image classification is completed, misclassification or noise usually appears in the recognition result, especially in complex scenarios or when the edges are blurred. Therefore, the recognition result is first denoised to improve the classification accuracy. The specific operations are as follows: Applying median filtering: Perform spatial filtering on the classification result. Remove the noise points in the image recognition result through median filtering. Median filtering selects the median of the pixels in its neighborhood for each pixel point to replace it, thereby removing the noise. The specific algorithm is expressed as: ; where , , , are the values of the current pixel and its neighboring pixels, is the output value after processing; Applying Gaussian filtering: After median filtering, use Gaussian filtering to further smooth the classification result and remove high-frequency noise.

[0028] After denoising the recognition result, the final image recognition result is output by enhancing the classification confidence; Confidence weighting: When using the support vector machine model, each classification result usually has a corresponding confidence value. For each classification result, calculate its confidence value and weight the final result according to the classification confidence. The specific calculation is: ; where refers to the final output after the confidence of the image classification result is enhanced, is the confidence value of the i-th classification result, is the weight of the result, and N is the total number of classification results.

[0029] The present invention encompasses any alternatives, modifications, equivalent methods, and solutions made to the essence and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention without the description of these details. Additionally, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion to the essence of the present invention.

[0030] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A multi-resolution fusion image recognition method based on human visual patterns, characterized in that: The following steps are involved: S1, image acquisition and preprocessing: acquiring the target image through the image acquisition device and preprocessing the image, including denoising and contrast enhancement; S2, multi-resolution image generation: the pre-processed target image is processed at different resolutions to generate images with multiple resolutions; S3, multi-resolution image feature extraction: Convolutional neural network is used to extract features from images of multiple resolutions. The extracted features include edge features, texture features, shape features, and color features. S4, feature fusion: weighted fusion of the multi-resolution image features extracted in S3 to generate a multi-resolution feature set; S5, introduction of attention mechanism: based on feature fusion, the human visual attention mechanism is introduced to optimize the multi-resolution feature set; S6, image recognition and classification: input the multi-resolution feature set into the support vector machine model for image recognition and classification; S7, post-processing and result output: post-process the recognition result and output the final image recognition result.

2. The multi-resolution fusion image recognition method based on human visual mode according to claim 1 is characterized in that: The S1 includes: S11, image acquisition: acquiring a target image through an image acquisition device, wherein the target image is an original image to be recognized, and the image acquisition device includes an infrared camera and an RGB camera; S12, image denoising: performing denoising processing on the acquired target image to reduce random noise in the image; S13, contrast enhancement: performing contrast enhancement on the denoised target image.

3. The multi-resolution fusion image recognition method based on human visual mode according to claim 1 is characterized in that: The S2 includes: S21, selecting a resolution level: determining multiple resolution levels according to the requirements of the image recognition task; S22, image scaling: scaling the preprocessed target image according to the selected resolution level; S23, generating a resolution image: generating a plurality of images with different resolutions based on the scaling process; S24, image quality assessment: perform quality assessment on each generated image with different resolutions.

4. The multi-resolution fusion image recognition method based on human visual mode according to claim 1, characterized in that: The S3 includes: S31, select the convolutional neural network structure: In order to extract effective features from images of different resolutions, a convolutional neural network based on the ResNet-50 structure is selected; S32, image input: Input images of different resolutions into the convolutional neural network for processing and extract their respective features.

5. The multi-resolution fusion image recognition method based on human visual mode according to claim 4 is characterized in that: The S4 includes: S41, feature weighted fusion: using weighted fusion algorithm to fuse features of different resolutions to generate a multi-resolution feature set; S42, dimensionality adjustment: principal component analysis is used to reduce the dimensionality of the multi-resolution feature set, and after dimensionality reduction, the final multi-resolution feature set is generated.

6. The multi-resolution fusion image recognition method based on human visual mode according to claim 5, characterized in that: The S5 includes: S51, select the attention mechanism: use the combination of channel attention mechanism and spatial attention mechanism to optimize feature representation; S52, fusion and weighting of attention mechanism: After introducing the channel and spatial attention mechanisms, these two parts of attention information are combined to optimize the representation of feature maps.

7. The multi-resolution fusion image recognition method based on human visual mode according to claim 6 is characterized in that: The S6 includes: S61, Feature processing and model training: Build and train a support vector machine model and process multi-resolution feature sets; S62, support vector machine model prediction: predict the input feature vector through the trained support vector machine model to identify the image category.

8. The multi-resolution fusion image recognition method based on human visual mode according to claim 1, characterized in that: In S7, median filtering and Gaussian filtering are used for post-processing to denoise the recognition results and improve the accuracy of classification.

9. The multi-resolution fusion image recognition method based on human visual mode according to claim 8, characterized in that: After the recognition result is denoised, the final image recognition result is output by enhancing the classification confidence.

Citation Information

Patent Citations

  • Bill identification method and device based on machine vision

    CN119068504A

  • Food and nutrient estimation, dietary assessment, evaluation, prediction and management

    WO2022133190A1

Cited By

  • Biological diversity intelligent monitoring method and system based on cloud platform

    CN120726307A

  • Cloud-based intelligent biodiversity monitoring methods and systems

    CN120726307B