Image recognition method and device, equipment and medium

By dynamically adjusting the size of the convolution kernel according to the image area complexity in the feature extraction layer of the deep learning model, the problems of low recognition accuracy and slow processing speed when processing high-resolution images are solved, and higher recognition accuracy and faster processing speed are achieved.

CN119942182APending Publication Date: 2025-05-06CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411919708.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional image recognition technology uses low recognition accuracy and slow processing speed due to computing resource limitations when processing high-resolution images.

Method used

A deep learning model was designed, and the feature extraction layer dynamically adjusts the size of the convolution kernel according to the complexity of different regions in the image, enhancing the ability to capture details in high-resolution images while reducing computational costs.

Benefits of technology

It improves the recognition accuracy and processing speed of high-resolution images, enhances the model's ability to capture detailed features, and reduces the overall calculation amount.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942182A_ABST
    Figure CN119942182A_ABST
Patent Text Reader

Abstract

The invention provides an image recognition method and device, equipment and a medium, and belongs to the technical field of computer vision, and the method comprises the steps: obtaining a to-be-recognized image, the resolution of the to-be-recognized image being higher than the standard resolution; inputting the to-be-recognized image into a pre-trained deep learning model to obtain a classification recognition result of the to-be-recognized image; wherein the deep learning model comprises an input layer used for receiving the image to be recognized; the plurality of feature extraction layers are used for executing a multi-layer convolution operation and a multi-layer pooling operation on the to-be-recognized image so as to extract a plurality of image features, the size of a convolution kernel used when the feature extraction layer performs convolution operation on each region in the to-be-recognized image and the complexity of each region accord with a positive correlation relationship; and the classification layer is used for carrying out image classification on the to-be-identified image according to the plurality of image features to obtain a classification identification result of the to-be-identified image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to an image recognition method, device, equipment and medium. Background Art

[0002] With the development of computer vision technology, image recognition is increasingly used, from face recognition to autonomous driving, from medical image analysis to industrial inspection, image recognition technology is profoundly affecting people's lives. Among them, high-resolution images usually contain rich details, so the recognition of high-resolution images is of great significance for application scenarios such as medical image analysis and industrial inspection, which require detailed distinction of object features.

[0003] However, traditional image recognition technology is usually designed to process standard resolution images. When processing high-resolution images, it is often difficult to directly and effectively process high-resolution images due to computing resource limitations, resulting in problems such as low recognition accuracy and slow processing speed. Summary of the invention

[0004] In view of the above problems, the embodiments of the present application provide an image recognition method, apparatus, device and medium to overcome the above problems or at least partially solve the above problems.

[0005] In a first aspect of an embodiment of the present application, an image recognition method is provided, the method comprising:

[0006] Acquire an image to be identified, wherein the resolution of the image to be identified is higher than the standard resolution;

[0007] Inputting the image to be identified into a pre-trained deep learning model to obtain a classification and recognition result of the image to be identified;

[0008] Wherein, the deep learning model includes:

[0009] An input layer, used for receiving the image to be recognized;

[0010] A plurality of feature extraction layers, used to perform multi-layer convolution operations and multi-layer pooling operations on the image to be identified, so as to extract a plurality of image features, wherein the size of the convolution kernel used by the feature extraction layer when performing the convolution operation on each region in the image to be identified is positively correlated with the complexity of each region;

[0011] The classification layer is used to classify the image to be identified according to the multiple image features to obtain a classification recognition result of the image to be identified.

[0012] According to a second aspect of the embodiments of the present application, an image recognition device is provided, the device comprising:

[0013] A data acquisition module, used for acquiring an image to be identified, wherein the resolution of the image to be identified is higher than the standard resolution;

[0014] A model reasoning module is used to input the image to be identified into a pre-trained deep learning model to obtain a classification and recognition result of the image to be identified;

[0015] Wherein, the deep learning model includes:

[0016] An input layer, used for receiving the image to be recognized;

[0017] A plurality of feature extraction layers, used to perform multi-layer convolution operations and multi-layer pooling operations on the image to be identified, so as to extract a plurality of image features, wherein the size of the convolution kernel used by the feature extraction layer when performing the convolution operation on each region in the image to be identified is positively correlated with the complexity of each region;

[0018] The classification layer is used to classify the image to be identified according to the multiple image features to obtain a classification recognition result of the image to be identified.

[0019] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the image recognition method described in the first aspect are implemented.

[0020] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the image recognition method described in the first aspect are implemented.

[0021] According to a fifth aspect of the embodiments of the present application, a computer program product is provided, including a computer program / instruction, which, when executed by a processor, implements the steps of the image recognition method as described in the first aspect.

[0022] The embodiments of the present application include the following advantages: when the deep learning model processes the image to be recognized (i.e., the high-resolution image), the feature extraction layer in the model dynamically adjusts the convolution kernel size according to the complexity of different regions in the image, that is, a larger convolution kernel is used for more complex regions to extract more contextual information, and a smaller convolution kernel is used for simpler regions to reduce the computational cost while ensuring sufficient information capture capability. In this way, the deep learning model's ability to capture detailed features in high-resolution images can be enhanced, while reducing the overall computational complexity of the model, thereby improving the recognition accuracy and processing speed of high-resolution images. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0024] Figure 1 is an implementation flow chart of an image recognition method in an embodiment of the present application;

[0025] Figure 2 is a schematic diagram of an implementation process of an image recognition method in an embodiment of the present application;

[0026] Figure 3 is a structural schematic diagram of an image recognition device in an embodiment of the present application;

[0027] Figure 4 It is a schematic diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0028] To facilitate understanding of the technical solutions provided by the present application, the main technical concepts involved in the embodiments of the present application are briefly described below.

[0029] Computer Vision (CV) large models: refers to large deep learning models used in the field of computer vision. Such models have more parameters and deeper network structures, and can handle complex visual tasks such as image classification, object detection, and semantic segmentation. CV large models are trained with large-scale data sets and can achieve high accuracy in various visual tasks. With the improvement of hardware computing power and the optimization of algorithms, CV large models can show better performance in high-resolution image processing.

[0030] Convolutional Neural Networks (CNN): A feedforward neural network with a deep structure that includes convolution calculations. It is used to process data with a grid structure, such as images and videos. It extracts local features of data through convolution and pooling operations, and has the characteristics of local connection, weight sharing, and translation invariance.

[0031] Standard resolution: usually refers to a resolution of 1920*1080 (corresponding to 2 million pixels), and in a few cases refers to a resolution of 2560*1440 (corresponding to 4 million pixels); images with a resolution higher than the standard resolution are considered high-resolution images.

[0032] In the related technologies, when processing high-resolution images, traditional image recognition technology faces many challenges. This is because the amount of data in high-resolution images is relatively large, and traditional image recognition methods are often difficult to directly and effectively process these images due to limited computing resources, resulting in low recognition accuracy and slow processing speed. In addition, since high-resolution images contain rich detail information, this also puts higher requirements on feature extraction algorithms. How to effectively use this detail information has become a key issue.

[0033] Take the traditional CNN as an example. As one of the core technologies of image recognition, it performs well in processing standard resolution images, but it has some inherent limitations when processing high-resolution images. For example, in order to reduce computational complexity and speed up training, CNN usually uses downsampling operations to reduce the spatial dimension of the image. However, this operation often leads to the loss of detailed information in the image, especially those subtle features (i.e., detail features) that are crucial for target recognition. In high-resolution images, these detail features are usually the key to distinguishing different objects. The loss of these detail features will greatly affect the recognition accuracy of high-resolution images. Therefore, traditional CNN often suffers from a decrease in recognition accuracy when processing such images. In addition, the amount of data in high-resolution images is huge. Inputting images of original size into CNN for processing will greatly increase the computational burden, resulting in a slowdown in recognition speed.

[0034] In response to the problems existing in the above-mentioned related technologies, the embodiments of the present application propose an image recognition method, device, equipment and medium, and design a feature extraction layer to dynamically adjust the convolution kernel size according to the complexity of different areas in the image, so as to enhance the deep learning model's ability to capture detailed features in high-resolution images, while reducing the overall computational complexity of the model, thereby improving the recognition accuracy and processing speed of high-resolution images.

[0035] In the following, in combination with the accompanying drawings, an image recognition method, apparatus, device and medium provided by the embodiments of the present application are described in detail through some embodiments and their application scenarios.

[0036] First, refer to Figure 1 FIG. 1 is a flowchart of an implementation of an image recognition method provided in an embodiment of the present application, and the method includes the following steps:

[0037] Step S11: Acquire an image to be identified, wherein the resolution of the image to be identified is higher than the standard resolution.

[0038] The standard resolution is 1920*1080 or 2560*1440.

[0039] In specific implementation, an interactive interface such as a web page may be pre-established to receive a high-resolution image (ie, an image to be recognized) input by a user.

[0040] Step S12: inputting the image to be identified into a pre-trained deep learning model to obtain a classification and recognition result of the image to be identified, wherein the deep learning model includes:

[0041] An input layer, used for receiving the image to be recognized;

[0042] A plurality of feature extraction layers, used to perform multi-layer convolution operations and multi-layer pooling operations on the image to be identified, so as to extract a plurality of image features, wherein the size of the convolution kernel used by the feature extraction layer when performing the convolution operation on each region in the image to be identified is positively correlated with the complexity of each region;

[0043] The classification layer is used to classify the image to be identified according to the multiple image features to obtain a classification recognition result of the image to be identified.

[0044] In this embodiment, the present application provides a deep learning model suitable for processing high-resolution images (which may be a CV large model or a convolutional neural network, etc.). The deep learning model includes an input layer, multiple high-resolution feature extraction layers (i.e., multiple feature extraction layers), and a classification layer, wherein: the input layer is used to receive a high-resolution image; the high-resolution feature extraction layer may include a convolution layer and a batch normalization layer (and an activation function layer), and the convolution layer uses a variable convolution kernel size to adapt to feature extraction requirements of different scales; the classification layer is used to classify the image according to the extracted features to obtain a classification recognition result.

[0045] Specifically, after dividing the input image into multiple regions of corresponding sizes according to its resolution, the feature extraction layer dynamically adjusts the convolution kernel size according to the complexity of the region, thereby using a larger convolution kernel to perform convolution operations on regions with higher complexity to extract more contextual information; and using a smaller convolution kernel to perform convolution operations on regions with lower complexity to reduce computational costs while ensuring sufficient information capture capabilities.

[0046] By adopting the technical solution of the embodiment of the present application, when the deep learning model processes the image to be recognized (i.e., the high-resolution image), the feature extraction layer in the model will dynamically adjust the convolution kernel size according to the complexity of different regions in the image, that is, use a larger convolution kernel for more complex regions to extract more contextual information, and use a smaller convolution kernel for simpler regions to reduce the computational cost while ensuring sufficient information capture capability. In this way, the deep learning model's ability to capture detailed features in high-resolution images can be enhanced, while reducing the overall computational complexity of the model, thereby improving the recognition accuracy and processing speed of high-resolution images.

[0047] As a possible implementation, the deep learning model further includes:

[0048] The feature fusion layer is used to perform feature fusion processing on the multiple image features based on the attention mechanism to obtain multiple fusion features, and input the multiple fusion features into the classification layer so that the classification layer determines the classification recognition result of the image to be recognized according to the multiple fusion features.

[0049] In this embodiment, the feature fusion layer uses an attention mechanism to perform feature fusion processing on multiple image features, which can help the model identify important features and suppress features that are not related to image recognition, thereby further improving the recognition accuracy and speed of the model when processing high-resolution images.

[0050] Optionally, feature fusion processing is performed on the multiple image features based on an attention mechanism to obtain multiple fusion features, including:

[0051] Determining attention weights for the multiple image features respectively;

[0052] According to the attention weights corresponding to the multiple image features, the multiple image features are subjected to feature fusion processing to obtain multiple fusion features;

[0053] The attention weight corresponding to a single image feature is determined by the following formula:

[0054]

[0055] Among them, A i represents the attention weight corresponding to a single image feature, N represents the number of the multiple image features, and the energy value e i =Wx i +b, W and b represent the weight vector and bias learned by the deep learning model during the model training process, x u represents the i-th image feature.

[0056] In this embodiment, for the N image features x1, x2, ..., x extracted by multiple feature extraction layers, N The feature fusion layer first transforms each image feature x i are mapped to a new space to produce an energy value e i , this energy value is calculated by a learnable weight vector W and a bias b through a linear transformation (an activation function (such as softmax) can also be added).

[0057] The weight vector W and bias b are parameters learned by the model through the training process, which can reflect the importance of each image feature and the energy value e i After the exponential function conversion, the energy values ​​of all image features are normalized to obtain each image feature x. i The corresponding attention weight A i , which ensures that the sum of the attention weights of all features is 1.

[0058] The feature fusion layer then performs feature fusion processing on each part of the multiple image features in a weighted sum manner based on the attention weights corresponding to each of the multiple image features, so as to obtain multiple fused features.

[0059] It is understandable that the attention mechanism can not only help the model identify important features and suppress irrelevant features to improve the recognition accuracy and speed of the model when processing high-resolution images, but also enhance the interpretability of the model, that is, it can intuitively show which parts of the features contribute the most to the decision-making process based on the attention weights.

[0060] Optionally, the feature fusion layer is further used to perform the following steps:

[0061] Determining importance scores for the multiple fusion features respectively;

[0062] Select the first k fusion features with the highest importance scores and input them into the classification layer, where k is a positive integer;

[0063] The importance score of a single fusion feature is determined by the following formula:

[0064]

[0065] Among them, I f Represents the importance score of a single fusion feature, w i represents the weight coefficient of the image feature fi, and n represents the number of image features associated with a single fusion feature.

[0066] In this embodiment, in order to further streamline the model, the feature fusion layer performs feature selection on the multiple fused features obtained, and only retains those features that are most critical to the classification task. This not only reduces the amount of calculation, but also improves the generalization ability and efficiency of the model.

[0067] Specifically, the feature fusion layer calculates the importance score I of each fused feature f To perform feature selection, the importance score is to associate each image feature f with the relevant fusion feature. i (i.e., the weight coefficient w of each image feature used to obtain the fusion feature in the feature fusion stage) i The weight coefficient w is obtained by accumulating the product of its own value. i It reflects the importance of each image feature for the classification task, which can be learned during the model training process.

[0068] Importance score of fusion features I f The larger the fusion feature is, the more critical it is to the decision-making of the model. Therefore, all fusion features can be sorted in descending order according to the importance score, and the first k features are selected and input into the classification layer for processing. Among them, one of the k features can be a hyperparameter, which can be adjusted according to the actual application scenario and the required model complexity.

[0069] It is understandable that by introducing feature selection, not only can the most representative and discriminative features be effectively screened, but the computational burden of the model can also be reduced, thereby improving the speed and accuracy of image recognition. In addition, feature selection also helps prevent overfitting, allowing the model to perform better when faced with new data.

[0070] As a possible implementation, the deep learning model is trained by the following steps:

[0071] Step S21: performing data enhancement processing on the training data set, wherein the data enhancement processing includes at least one of random cropping, horizontal flipping, rotation, scaling, and brightness and contrast adjustment.

[0072] In specific implementation, the image samples in the training data set can be denoised to eliminate the noise interference that may exist in the image samples, and the diversity of image samples can be increased through data enhancement processing (such as brightness adjustment, contrast enhancement, rotation, translation, etc.), thereby improving the robustness and generalization ability of the model and ensuring that the model performs well when facing new data.

[0073] Step S22: The training data set after data enhancement is standardized by batch normalization technology.

[0074] In the specific implementation, batch normalization technology is used to speed up the training process. Specifically, the input of each batch of data is standardized to ensure that the input of each layer has the same distribution, thereby improving the stability of the model and alleviating the problem of gradient disappearance or gradient explosion in deep networks.

[0075] Exemplarily, the batch normalization algorithm is defined as follows:

[0076]

[0077] in, represents the data after normalization, x i Represents the data in the training data set, μ and σ 2 Respectively represent the mean and variance of each data in the training data set, and ∈ represents a set constant.

[0078] Step S23: Use the standardized training data set to train the deep learning model, and dynamically adjust the learning rate of the deep learning model through an adaptive moment estimation optimization algorithm during the training process.

[0079] In specific implementation, an adaptive moment estimation optimization algorithm is used to perform model optimization training, such as dynamically adjusting the learning rate by maintaining the first-order moment estimation and second-order moment estimation of the gradient to accelerate convergence and improve model training efficiency. It can be understood that the adaptive moment estimation optimization algorithm can adaptively adjust the learning rate according to historical gradient information to ensure that the training process is efficient.

[0080] Exemplarily, the adaptive moment estimation optimization algorithm is defined as follows:

[0081]

[0082] Among them, y o,c Represents the unique hot encoding of the true label corresponding to the image sample, p o,c Represents the probability distribution predicted by the model, L represents the moment estimate of the gradient, and C represents the preset number of image classifications.

[0083] As a possible implementation manner, the complexity of each region is determined according to the edge strength of each region.

[0084] In this embodiment, considering that areas with high edge strength often contain more details and structural information, which means that these areas are more likely to contain features that are important to the recognition task, the edge strength is used as an indicator to measure the complexity of the region. Based on this indicator, the feature extraction layer can intelligently assign the most suitable convolution kernel size to each image area, thereby reducing computational efficiency while ensuring the quality of feature extraction. As a result, the performance of the model in processing high-resolution images can be improved, especially in image recognition application scenarios such as medical image analysis or industrial detection that require detailed distinction of object features. By dynamically adjusting the convolution kernel size according to the regional edge strength, the recognition accuracy and speed can be significantly improved.

[0085] As a possible implementation, the classification layer receives input data from other layers through a residual structure.

[0086] In this implementation, it is considered that when the network becomes very deep, the training process may encounter the problem of gradient disappearance or gradient explosion, which may make the model difficult to converge or the training unstable.

[0087] To solve the above problems, this application introduces a residual structure (i.e., residual block), which allows the network to bypass one or more layers by directly adding the input to the output of the layer. This design can be seen as adding a "shortcut" or "skip connection" to the network, thereby helping the gradient to flow back to the early layers of the network more easily, thereby improving the training effect of the deep network.

[0088] It is understandable that in the design of the residual block, the input signal is processed by one or more convolutional layers, and after being processed by these layers, the original input signal is directly added to the output signal processed by the convolutional layer. This direct addition method forms a so-called jump connection, which allows information and gradients to be transmitted more smoothly during training. If the dimensions of the input and output do not match, the dimensions can be adjusted by setting an additional convolutional layer to ensure that the dimensions of the two are consistent so that the addition operation can be performed.

[0089] By adopting the residual structure, the model can still be effectively trained even when the network becomes very deep, thereby avoiding the gradient vanishing problem common in traditional deep networks, thereby improving the training efficiency of the model and the final recognition performance.

[0090] Optionally, each residual block in the residual structure contains two convolutional layers, each followed by a batch normalization layer and a rectified linear unit (ReLU) activation function layer. It can be understood that the addition of the batch normalization layer helps to speed up the training process and contributes to the stable convergence of the model, while the ReLU activation function layer introduces nonlinearity to the network, allowing the network to learn more complex feature mappings.

[0091] For example, refer to Figure 2 The schematic diagram of the implementation process of the image recognition method shown in FIG. 1 includes the following steps:

[0092] A1: Build a deep learning model suitable for high-resolution images, the model including an input layer, multiple high-resolution feature extraction layers, a feature fusion layer, and a classification layer.

[0093] Among them, the deep learning model is built based on a convolutional neural network, and the input layer is used to receive a high-resolution image; the high-resolution feature extraction layer includes a convolution layer, a ReLU activation function layer and a batch normalization layer, and the convolution layer can be connected to the ReLU activation function layer and the batch normalization layer in a parallel or serial manner; the feature fusion layer uses an attention mechanism to enhance the weights of important features; the classification layer is used to perform image classification according to the image features after feature fusion processing.

[0094] It should be noted that the design of the attention mechanism is to enable the model to focus on the most relevant parts of the input features, thereby enhancing the performance of the model. Specifically, the feature fusion layer in the deep learning model tells the model the importance of each feature by assigning different attention weights to each feature, so that the model can automatically learn which features are more important for the final recognition task.

[0095] A2: It uses a feature extraction algorithm based on an improved convolutional neural network, including image preprocessing, image feature extraction through multi-layer convolution and pooling operations, nonlinear transformation, and fully connected layer classification and recognition.

[0096] In this step, the input image is preprocessed by denoising, enhancing, etc., and then the preprocessed image is input into the model, where multiple high-resolution feature extraction layers extract image features through multi-layer convolution and pooling operations. It can be understood that for the image input to the model, multi-layer convolution and pooling operations will work together to capture local features in the image and gradually reduce its spatial dimension; among them, the convolution operation can capture different feature patterns of the image, while pooling helps to extract the most significant features and reduce the amount of calculation.

[0097] The high-resolution feature extraction layer then uses a nonlinear transformation function to perform a nonlinear transformation on the extracted image features to enhance the expressiveness of the features. Optionally, the nonlinear transformation function f(x) is defined as: f(x) = max(0, x), that is, a nonlinear transformation is performed using the ReLU function. The ReLU function is an activation function that can help the network learn more complex feature representations; in addition, the ReLU function is simple to calculate and can alleviate the gradient vanishing problem, which is especially important for deep networks.

[0098] Finally, the transformed features are processed by the feature fusion layer and input into the fully connected layer (i.e., classification layer) for classification and recognition. Specifically, the fully connected layer converts the feature map into category probabilities, thereby achieving image classification. In the aforementioned process, each layer of the model is connected to the previous layer through a weight matrix, and these weights are continuously adjusted through the training process to better fit the training data.

[0099] A3: Use efficient model training strategies, including data augmentation, batch normalization, residual connections, and adaptive moment estimation optimization algorithms.

[0100] It is understandable that by introducing residual connections in the deep network, that is, directly adding the input and the output of the layer to form a skip connection, the flow of gradients in the deep network can be promoted, making it easier for the model to learn the identity mapping, and even if the newly added layers of the model fail to learn useful information, the original performance of the model can still be maintained.

[0101] A4: Apply the trained model to actual image recognition tasks to quickly and accurately recognize input high-resolution images.

[0102] Based on the above examples, the present application provides an image recognition method, which improves the accuracy and speed of image recognition by constructing a specific deep learning model, adopting an improved feature extraction algorithm, and introducing an efficient training strategy and an attention mechanism in the feature fusion layer.

[0103] Specifically, by using the attention mechanism in the feature fusion layer, the feature fusion layer can assign different attention weights to different features according to their importance, so that the model can focus more on those features that are critical to image recognition. Among them, the attention weight is calculated through a learnable weight vector and bias, which enables the model to self-regulate during the training process and focus on the important features that best represent the image category. This can avoid the problem of important detail information loss caused by operations such as global average pooling in traditional CNN, thereby improving the recognition accuracy of the model.

[0104] By using variable convolution kernel sizes to adapt to feature extraction requirements at different scales, the problem of extracting detail information in high-resolution images is solved. Specifically, for each region in the image, the model dynamically adjusts the convolution kernel size according to the complexity of the region: regions with high complexity use larger convolution kernels to capture more details, while regions with low complexity use smaller convolution kernels to reduce unnecessary computational burden. Among them, the regional complexity can be determined based on the local edge strength, so that the model can intelligently identify which areas need more attention and perform targeted feature extraction. In this way, the model's ability to capture detailed features can be enhanced, while also reducing the overall computational complexity of the model, making image recognition both fast and accurate, and making the model suitable for processing high-resolution images that contain rich details.

[0105] For the method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present application are not limited by the order of the actions described, because according to the embodiments of the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.

[0106] Second, Figure 3 : is a schematic diagram of the structure of an image recognition device according to an embodiment of the present application, the device comprising:

[0107] The data acquisition module 310 is used to acquire an image to be identified, wherein the resolution of the image to be identified is higher than the standard resolution;

[0108] The model inference module 320 is used to input the image to be identified into a pre-trained deep learning model to obtain a classification and recognition result of the image to be identified;

[0109] Wherein, the deep learning model includes:

[0110] An input layer, used for receiving the image to be recognized;

[0111] A plurality of feature extraction layers, used to perform multi-layer convolution operations and multi-layer pooling operations on the image to be identified, so as to extract a plurality of image features, wherein the size of the convolution kernel used by the feature extraction layer when performing the convolution operation on each region in the image to be identified is positively correlated with the complexity of each region;

[0112] The classification layer is used to classify the image to be identified according to the multiple image features to obtain a classification recognition result of the image to be identified.

[0113] By adopting the technical solution of the embodiment of the present application, when the deep learning model processes the image to be recognized (i.e., the high-resolution image), the feature extraction layer in the model will dynamically adjust the convolution kernel size according to the complexity of different regions in the image, that is, use a larger convolution kernel for more complex regions to extract more contextual information, and use a smaller convolution kernel for simpler regions to reduce the computational cost while ensuring sufficient information capture capability. In this way, the deep learning model's ability to capture detailed features in high-resolution images can be enhanced, while reducing the overall computational complexity of the model, thereby improving the recognition accuracy and processing speed of high-resolution images.

[0114] Optionally, the deep learning model also includes:

[0115] The feature fusion layer is used to perform feature fusion processing on the multiple image features based on the attention mechanism to obtain multiple fusion features, and input the multiple fusion features into the classification layer so that the classification layer determines the classification recognition result of the image to be recognized according to the multiple fusion features.

[0116] Optionally, the feature fusion layer is further used to perform the following steps:

[0117] Determining importance scores for the multiple fusion features respectively;

[0118] Select the first k fusion features with the highest importance scores and input them into the classification layer, where k is a positive integer;

[0119] The importance score of a single fusion feature is determined by the following formula:

[0120]

[0121] Among them, I f Represents the importance score of a single fusion feature, w i Represents the image feature f i The weight coefficient n represents the number of image features associated with a single fusion feature.

[0122] Optionally, feature fusion processing is performed on the multiple image features based on an attention mechanism to obtain multiple fusion features, including:

[0123] Determining attention weights for the multiple image features respectively;

[0124] According to the attention weights corresponding to the multiple image features, the multiple image features are subjected to feature fusion processing to obtain multiple fusion features;

[0125] The attention weight corresponding to a single image feature is determined by the following formula:

[0126]

[0127] Among them, A i represents the attention weight corresponding to a single image feature, N represents the number of the multiple image features, and the energy value e i =Wx i +b, W and b represent the weight vector and bias learned by the deep learning model during the model training process, x i represents the i-th image feature.

[0128] Optionally, the device further comprises a model training module, configured to perform the following steps:

[0129] Performing data enhancement processing on the training data set, wherein the data enhancement processing includes at least one of random cropping, horizontal flipping, rotation, scaling, and brightness and contrast adjustment;

[0130] The training data set after data enhancement is standardized through batch normalization technology;

[0131] The deep learning model is trained using the standardized training data set, and the learning rate of the deep learning model is dynamically adjusted during the training process using an adaptive moment estimation optimization algorithm.

[0132] Optionally, the complexity of each region is determined according to the edge strength of each region.

[0133] Optionally, the classification layer receives input data from other layers through a residual structure.

[0134] It should be noted that the device embodiment is similar to the method embodiment, so the description is relatively simple, and the relevant parts can be referred to the method embodiment.

[0135] The present application also provides an electronic device, referring to Figure 4 , Figure 4 Schematic diagram of an electronic device proposed in an embodiment of the present application. Figure 4 As shown, the electronic device 100 includes: a memory 110 and a processor 120. The memory 110 and the processor 120 are connected via a bus communication. A computer program is stored in the memory 110. The computer program can be run on the processor 120 to implement the steps in the image recognition method disclosed in the embodiment of the present application.

[0136] The embodiment of the present application also provides a computer-readable storage medium on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the image recognition method disclosed in the embodiment of the present application is implemented.

[0137] The embodiment of the present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the image recognition method disclosed in the embodiment of the present application.

[0138] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0139] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0140] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, systems, devices, storage media, and program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0141] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0142] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1The steps for the functions specified in one or more boxes.

[0143] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.

[0144] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0145] The above is a detailed introduction to an image recognition method, device, equipment and medium provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. An image recognition method, characterized in that: The method comprises: Acquire an image to be identified, wherein the resolution of the image to be identified is higher than the standard resolution; Inputting the image to be identified into a pre-trained deep learning model to obtain a classification and recognition result of the image to be identified; Wherein, the deep learning model includes: An input layer, used for receiving the image to be recognized; A plurality of feature extraction layers, used to perform multi-layer convolution operations and multi-layer pooling operations on the image to be identified, so as to extract a plurality of image features, wherein the size of the convolution kernel used by the feature extraction layer when performing the convolution operation on each region in the image to be identified is positively correlated with the complexity of each region; The classification layer is used to classify the image to be identified according to the multiple image features to obtain a classification recognition result of the image to be identified.

2. The method according to claim 1, characterized in that The deep learning model also includes: The feature fusion layer is used to perform feature fusion processing on the multiple image features based on the attention mechanism to obtain multiple fusion features, and input the multiple fusion features into the classification layer so that the classification layer determines the classification recognition result of the image to be recognized according to the multiple fusion features.

3. The method according to claim 2, characterized in that The feature fusion layer is also used to perform the following steps: Determining importance scores for the multiple fusion features respectively; Select the first k fusion features with the highest importance scores and input them into the classification layer, where k is a positive integer; The importance score of a single fusion feature is determined by the following formula: Among them, I f Represents the importance score of a single fusion feature, w i Represents the image feature f i The weight coefficient n represents the number of image features associated with a single fusion feature.

4. The method according to claim 2, characterized in that: The multiple image features are subjected to feature fusion processing based on the attention mechanism to obtain multiple fusion features, including: Determining attention weights for the multiple image features respectively; According to the attention weights corresponding to the multiple image features, the multiple image features are subjected to feature fusion processing to obtain multiple fusion features; The attention weight corresponding to a single image feature is determined by the following formula: Among them, A i represents the attention weight corresponding to a single image feature, N represents the number of the multiple image features, and the energy value e i =Wx i +b, W and b represent the weight vector and bias learned by the deep learning model during model training, x i represents the i-th image feature.

5. The method according to claim 1, characterized in that: The deep learning model is trained by the following steps: Performing data enhancement processing on the training data set, wherein the data enhancement processing includes at least one of random cropping, horizontal flipping, rotation, scaling, and brightness and contrast adjustment; The training data set after data enhancement is standardized through batch normalization technology; The deep learning model is trained using the standardized training data set, and the learning rate of the deep learning model is dynamically adjusted during the training process using an adaptive moment estimation optimization algorithm.

6. The method according to any one of claims 1 to 5, characterized in that: The complexity of each region is determined according to the edge strength of each region.

7. The method according to any one of claims 1 to 5, characterized in that: The classification layer receives input data from other layers through a residual structure.

8. An image recognition device, characterized in that: The device comprises: A data acquisition module, used for acquiring an image to be identified, wherein the resolution of the image to be identified is higher than the standard resolution; A model reasoning module is used to input the image to be identified into a pre-trained deep learning model to obtain a classification and recognition result of the image to be identified; Wherein, the deep learning model includes: An input layer, used for receiving the image to be recognized; A plurality of feature extraction layers, used to perform multi-layer convolution operations and multi-layer pooling operations on the image to be identified, so as to extract a plurality of image features, wherein the size of the convolution kernel used by the feature extraction layer when performing the convolution operation on each region in the image to be identified is positively correlated with the complexity of each region; The classification layer is used to classify the image to be identified according to the multiple image features to obtain a classification recognition result of the image to be identified.

9. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the image recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the image recognition method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Feature extraction method and device, electronic equipment, medium and computer program product

    CN121937831A