Image recognition method and device based on convolutional neural network model and terminal equipment

By converting spatial domain convolution operations in convolutional neural network models into frequency domain multiplication operations, the problem of high computational cost for global feature extraction in large-size images is solved, enabling efficient image recognition on devices with lower computing power.

CN115690488BActive Publication Date: 2026-04-17SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
Filing Date
2022-10-13
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing convolutional neural network models are computationally intensive when extracting global features of images, making them difficult to apply effectively on devices with low computing power, especially for processing large images.

Method used

Fast Fourier Transform is used to convert spatial domain convolution operations into frequency domain multiplication operations. A convolutional neural network model is then used to perform global frequency domain convolution on the image, reducing the computational cost of global feature extraction.

Benefits of technology

It effectively reduces the computational complexity of convolutional neural network models in extracting global features from large images, improves recognition efficiency, and facilitates deployment on devices with lower computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115690488B_ABST
    Figure CN115690488B_ABST
Patent Text Reader

Abstract

The application is suitable for the field of image recognition technology, and provides an image recognition method and device based on a convolutional neural network model and a terminal device, wherein the convolutional neural network model performs frequency domain global convolution on a to-be-recognized image based on fast Fourier transform, and the image recognition method comprises the following steps: inputting the to-be-recognized image into the trained convolutional neural network model, sequentially performing feature extraction and recognition on the to-be-recognized image through the convolutional neural network model, and obtaining a recognition result. The application can reduce the calculation amount of the convolutional neural network model for extracting global features and improve the model efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image recognition technology, and in particular relates to image recognition methods, devices, terminal equipment, and computer-readable storage media based on convolutional neural network models. Background Technology

[0002] Feature extraction and matching are important tasks in many computer vision applications, widely used in image retrieval, object detection, and other image recognition tasks. When extracting features from an image, image features include global features and local features. Global features refer to the overall attributes of the image, while local features are features extracted from local regions of the image.

[0003] In existing technologies, convolutional neural networks are widely used to extract global features of images because convolutional operations have good hardware support. However, since convolutional neural networks cannot capture global information at once, multiple convolutional layers need to be stacked to increase the receptive field, which increases the number of model parameters and computational cost. Summary of the Invention

[0004] This application provides an image recognition method, apparatus, and terminal device based on a convolutional neural network model, which can reduce the computational load when the convolutional neural network model performs global feature extraction, thereby improving model efficiency.

[0005] In a first aspect, embodiments of this application provide an image recognition method based on a convolutional neural network model. The convolutional neural network model performs frequency domain global convolution on the image to be recognized based on a fast Fourier transform. The image recognition method includes:

[0006] The image to be identified is input into the trained convolutional neural network model, and the convolutional neural network model sequentially extracts and identifies features from the image to obtain the recognition result.

[0007] Secondly, embodiments of this application provide a method for training a convolutional neural network model, including:

[0008] Obtain the constructed convolutional neural network model, and input the sample image into the convolutional neural network for training until the convolutional neural network meets the preset requirements, thus obtaining the convolutional neural network model;

[0009] The aforementioned convolutional neural network performs frequency domain global convolution on the sample images based on the Fast Fourier Transform.

[0010] Thirdly, embodiments of this application provide an image recognition device, including:

[0011] The input module and the trained convolutional neural network model, which is based on the Fast Fourier Transform to perform frequency domain global convolution on the image to be recognized;

[0012] The above-mentioned input module is used to input the image to be recognized into the above-mentioned convolutional neural network model;

[0013] The aforementioned convolutional neural network model is used to sequentially extract and recognize features from the image to be recognized, thereby obtaining the recognition result.

[0014] Fourthly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the image recognition method based on a convolutional neural network model described in the first aspect or the convolutional neural network model training method described in the second aspect.

[0015] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the image recognition method based on a convolutional neural network model described in the first aspect or the convolutional neural network model training method described in the second aspect.

[0016] Sixthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the image recognition method based on a convolutional neural network model as described in any one of the first aspects or the convolutional neural network model training method described in the second aspect.

[0017] The beneficial effects of the embodiments in this application compared with the prior art are:

[0018] In this embodiment, the image to be recognized is input into the trained convolutional neural network model. The convolutional neural network model sequentially extracts features and recognizes the image to be recognized to obtain the recognition result. Since the global convolution in the frequency domain of the image to be recognized is based on the Fast Fourier Transform, converting the convolution operation in the spatial domain into a multiplication operation in the frequency domain, the computational load when the convolutional neural network extracts global features can be reduced, improving the recognition efficiency of the convolutional neural network model. It also facilitates deployment on devices with lower computing power. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0020] Figure 1This is a schematic flowchart of an image recognition method based on a convolutional neural network model provided in an embodiment of this application;

[0021] Figure 2 This is a schematic diagram of the structure of the convolutional neural network model provided in the embodiments of this application;

[0022] Figure 3 This is a schematic diagram of the structure of the second convolutional module provided in an embodiment of this application;

[0023] Figure 4 This is a flowchart illustrating the convolutional neural network model training method provided in the embodiments of this application;

[0024] Figure 5 This is a schematic diagram of the structure of the image recognition device provided in the embodiments of this application;

[0025] Figure 6 This is a schematic diagram of the structure of the convolutional neural network model training device provided in the embodiments of this application;

[0026] Figure 7 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. Detailed Implementation

[0027] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0028] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0029] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0030] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0031] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.

[0032] Example 1:

[0033] Figure 1 A flowchart illustrating an image recognition method based on a convolutional neural network model according to an embodiment of the present invention is shown below in detail:

[0034] The image to be identified is input into the trained convolutional neural network model, and the convolutional neural network model sequentially extracts and identifies features from the image to obtain the recognition result.

[0035] The aforementioned convolutional neural network model performs frequency domain global convolution processing on the image to be recognized based on the Fast Fourier Transform.

[0036] Specifically, when extracting global features from an image, convolutional neural networks (CNNs) typically require stacking multiple convolutional layers to increase the receptive field, thereby extracting global features through a larger receptive field. However, this also increases the number of parameters and computational cost of the CNN model, making its computational complexity excessive. Furthermore, as the size of the image to be recognized increases, especially at large sizes (e.g., 112*112), the computational complexity of the CNN model rapidly exceeds that of a 7*7 convolution, making it impractical for real-world applications. Therefore, when the CNN model extracts features from the input image to be recognized, a Fast Fourier Transform (FFT) is performed on the image to be recognized. This transforms the image into a frequency domain image, converting spatial domain convolution operations into frequency domain multiplication operations, thus reducing the computational cost in the global feature extraction process.

[0037] In this embodiment, the image to be recognized is input into the trained convolutional neural network model. The convolutional neural network model sequentially extracts features and recognizes the image to be recognized to obtain the recognition result. Since the global convolution in the frequency domain of the image to be recognized is based on the Fast Fourier Transform, converting the convolution operation in the spatial domain into a multiplication operation in the frequency domain, the computational load in the process of extracting global features of large-size images by the convolutional neural network model can be effectively reduced when extracting global features of the image to be recognized, thereby improving the recognition efficiency of the convolutional neural network model and facilitating deployment on devices with lower computing power.

[0038] In some embodiments, the above-described image recognition method based on a convolutional neural network model further includes:

[0039] Obtain the image to be recognized.

[0040] Optionally, the image to be identified may be an image captured by a camera device or an image frame in a video stream captured by a camera device.

[0041] Optionally, since different image recognition tasks may require different images to be recognized, and the camera equipment used and the rules for acquiring the images to be detected may also differ, the corresponding images to be recognized are obtained according to the acquisition methods and rules for each application field. For example, for a face recognition task, it is necessary to acquire face images as images to be recognized and to identify the facial features in the face images.

[0042] In this embodiment of the application, based on the images required for image recognition tasks in various application fields, corresponding acquisition methods and acquisition rules are adopted to obtain images to be recognized that meet the requirements of image recognition tasks, so as to perform image recognition tasks.

[0043] In some embodiments, the convolutional neural network model includes a feature extraction module and a recognition module. The steps described above involve sequentially extracting features and recognizing the image to be recognized using the convolutional neural network model to obtain a recognition result, including:

[0044] A1. The above feature extraction module is used to extract features from the image to be identified.

[0045] A2. Based on the above recognition module, the extracted features are recognized to obtain the recognition results.

[0046] Optionally, since image recognition includes different tasks such as image classification and object detection, different image recognition tasks employ different recognition methods for the same feature. Therefore, the feature extraction module extracts features from the input image to be recognized and uses the extracted features as input to the recognition module. Based on the image recognition task, corresponding recognition is performed to obtain the recognition result. The recognition module may include one or more recognition units, each performing different recognition tasks. For example, it may include a pedestrian detection unit and an object detection unit. The pedestrian detection unit performs pedestrian detection based on the extracted features, or the extracted feature map is input into both the pedestrian detection unit and the object detection unit for pedestrian and object detection tasks.

[0047] In this embodiment, the feature extraction module extracts features from the image to be recognized, and the recognition module obtains the extracted features for corresponding recognition to obtain the corresponding recognition results, thereby improving the recognition efficiency of each image recognition task.

[0048] In some embodiments, the feature extraction module includes a first convolution module and a second convolution module, and step A1 includes:

[0049] A11. Based on the first convolutional module described above, local features are extracted from the image to be identified to obtain a local feature map;

[0050] A12. Based on the second convolution module mentioned above, the fast Fourier transform is used to perform frequency domain global convolution on the local feature map to obtain the global feature map.

[0051] Optionally, the above convolutional neural network model can be constructed based on existing convolutional neural networks. Shallow convolutional layers can be used as the first convolutional module, and deep convolutional layers or self-attention layers can be replaced with second convolutional modules. The first convolutional module performs convolution processing on the image to be recognized using ordinary convolution, extracting local features of the image and outputting a local feature map. This local feature map is then used as input to the second convolutional module. The second module performs a Fast Fourier Transform on the local feature map to obtain a local feature map in the frequency domain, and extracts global features based on the local feature map in the frequency domain to obtain a global feature map. For example, as... Figure 2 The convolutional neural network model shown has three convolutional layers, the first three being ordinary convolutional modules, and three deeper convolutional layers, which are the second convolutional modules. The second convolutional modules are connected to the recognition module. The image to be recognized is used as the input of the first convolutional module for local feature extraction, and the output features are input to the second convolutional module for global feature extraction. The global feature map output by the second convolutional module is used as the input of the recognition module for recognition, thereby outputting the corresponding recognition result.

[0052] It should be noted that the first and second convolutional modules in the above convolutional neural network model can also adopt an alternating structure, that is, the first convolutional module is connected to the second convolutional module, and the output of the second convolutional module is connected to another first convolutional module. Figure 2 The structure shown is a stacked structure, which means that the feature map output by the feature extraction module is the global feature map extracted by the second convolution module (that is, the recognition module recognizes based on the global features extracted by the second convolution module). It does not limit the specific structure of the first convolution module (ordinary convolution layer) and the second convolution module provided in the embodiments of this application in the convolutional neural network model.

[0053] In this embodiment, the local features of the image to be recognized are extracted by the first convolutional module of the convolutional neural network model, and the obtained local feature map is used as the input of the second convolutional module to perform fast Fourier transform processing on the local feature map. Then, global features are extracted based on the obtained local feature map in the frequency domain. Therefore, the amount of computation in the global feature extraction process is reduced. Moreover, the obtained global features contain both local and global features, which improves the recognition accuracy of the convolutional neural network model.

[0054] In some embodiments, the second convolutional module includes a first branch and a second branch, and step A12 includes:

[0055] The aforementioned local feature map is split along the channel direction to obtain a first local feature map and a second local feature map, and the first local feature map and the second local feature map are respectively input into the first branch and the second branch.

[0056] Optionally, when the second convolutional module performs global feature extraction on the input local feature map, the local feature map is first split along the channel direction (e.g., evenly split into two parts along the channel direction) to obtain the first local feature map and the second local feature map. The first local feature map is then input into the first branch, and the second local feature map is input into the second branch, so as to extract global features from the first local feature map and the second local feature map respectively.

[0057] The first branch and the second branch mentioned above respectively use Fast Fourier Transform to perform frequency domain global convolution on the input local feature map to obtain the first global feature map and the second global feature map.

[0058] Optionally, the first branch performs a Fast Fourier Transform on the first local feature map to obtain a first local feature map in the frequency domain, and then performs a global convolution on the first local feature map in the frequency domain to obtain a first global feature map; the second branch performs a Fast Fourier Transform on the second local feature map to obtain a second local feature map in the frequency domain, and then performs a global convolution on the second local feature map in the frequency domain to obtain a second global feature map.

[0059] The first global feature map and the second global feature map are concatenated along the channel direction to obtain the global feature map.

[0060] Optionally, since the first local feature map and the second local feature map are obtained by splitting the local feature map along the channel direction, after extracting the features of the first local feature map and the second local feature map from the first branch and the second branch respectively, the obtained first global feature map and the second global feature map are spliced ​​along the channel direction to obtain a complete global feature map, so as to perform subsequent processing based on the complete global feature map of the input image to be recognized.

[0061] In this embodiment, since the local feature map is split into two parts along the channel direction and input to the first branch and the second branch, the number of channels in the local feature map is halved. Furthermore, in the global feature extraction process, the local feature map is subjected to frequency domain global convolution based on the fast Fourier transform. Therefore, the global feature extraction is performed based on the fast Fourier transform of the local feature map with halved channel number, which reduces the computational complexity of global feature extraction and thus reduces the requirements for device computing power, making it easier to deploy applications on devices with lower computing power.

[0062] In some embodiments, when the first branch and the second branch perform frequency domain global convolution based on the input local feature map, the process includes:

[0063] The first branch performs Fast Fourier Transform on the first local feature map and the corresponding weight matrix based on the column dimension to obtain the first local feature map and weight matrix in the frequency domain.

[0064] The first local feature map in the frequency domain and the weight matrix are multiplied point by point to obtain the first frequency domain feature map;

[0065] The first feature map in the frequency domain is processed by inverse fast Fourier transform to obtain the first global feature map.

[0066] Optionally, during the process of extracting global features from the first local feature map in the first branch, the first local feature map and its weight matrix are processed by Fast Fourier Transform along the row dimension to convert them into the first local feature map and weight matrix in the frequency domain. This allows the first local feature map and weight matrix in the frequency domain to be multiplied point by point when the weight matrix is ​​used to perform global convolution on the first local feature map. In other words, the convolution operation in the spatial domain is converted into a multiplication operation in the frequency domain, thereby obtaining the first frequency domain feature map (global feature). Then, the first frequency domain feature map is subjected to inverse Fast Fourier Transform to obtain the first global feature map in the spatial domain. This effectively reduces the computational load when using a large convolution kernel to extract global features.

[0067] The second branch performs Fast Fourier Transform on the aforementioned second local feature map and the corresponding weight matrix based on the row dimension to obtain the second local feature map and weight matrix in the frequency domain.

[0068] The second local feature map in the frequency domain and the weight matrix are multiplied point by point to obtain the second frequency domain feature map;

[0069] The second feature map in the frequency domain is processed by inverse fast Fourier transform to obtain the second global feature map.

[0070] Optionally, during the process of extracting global features from the second local feature map in the second branch, the second local feature map and its weight matrix are processed by Fast Fourier Transform along the row dimension to convert them into the second local feature map and weight matrix in the frequency domain. This allows the second local feature map and weight matrix in the frequency domain to be multiplied point by point when the weight matrix is ​​used to perform global convolution on the second local feature map. In other words, the convolution operation in the spatial domain is converted into a multiplication operation in the frequency domain, thereby obtaining the second frequency domain feature map (global features). Then, the second frequency domain feature map is subjected to inverse Fast Fourier Transform to obtain the second global feature map in the spatial domain. This effectively reduces the computational cost when using a large convolution kernel to extract global features.

[0071] The computational complexity of the second convolutional module in extracting global features can be expressed as:

[0072]

[0073] That is, the computational complexity of the first branch and the second branch performing global convolution on the local feature map based on the fast Fourier transform is O(CHW(log2[H]+log2[W]).

[0074] Optionally, in order to reduce the computational complexity of the convolutional neural network model, when performing a fast Fourier transform on the first local feature map and its weight matrix along the row dimension, a one-dimensional fast Fourier transform is performed on the first local feature map and its weight matrix to obtain a one-dimensional form (such as an array) of the first local feature map and its weight matrix. For example, the first local feature map is represented as (C, H, W), and a one-dimensional fast Fourier transform is performed on the first local feature map along the column dimension to obtain a numerical form of the first local feature map with H elements.

[0075] In this embodiment, since the local feature map and its weight matrix in the spatial domain are converted into the frequency domain form based on the Fast Fourier Transform for global feature extraction, the convolution operation in the spatial domain is converted into a simple multiplication operation, reducing the computational load of the convolutional neural network model. At the same time, since the local feature map is processed by Fast Fourier Transform from different dimensions and global features are extracted, global features of different dimensions are obtained. Then, the obtained global features are concatenated to obtain global features including the width and height directions, which reduces the computational load when extracting global features from the local feature map. This makes it easier to deploy and apply the convolutional neural network model on devices with low computing power.

[0076] In some embodiments, the second convolutional module further includes a location embedding module, and the step of splitting the local feature map along the channel direction further includes:

[0077] The location embedding module extracts features from the local feature map to obtain a location feature map. The location feature map is then added to the local feature map to obtain a local feature map containing location features.

[0078] Optionally, before performing global feature extraction on the local feature map, the local feature map is convolved by the location embedding module to extract the location features in the local feature map, generate a location feature map, and add the location feature map to the local feature map according to the pixel position to obtain a local feature map with embedded location features.

[0079] Optionally, the aforementioned location embedding module is a two-dimensional convolutional module that performs convolution processing on the input image to be recognized to generate a two-dimensional location feature map. That is, the size of the location feature map is made consistent with the resolution of the input local feature map, so that the location feature map can be directly added to the local feature map to obtain a local feature map containing location features. For example, the location embedding module adopts a two-layer lightweight convolutional network structure, namely a simple structure of "convolution + normalization processing + activation function + convolution". Since it is necessary to generate a two-dimensional location feature map, the convolutional layer can use a 3×3 depthwise separable convolution to perform convolution processing on the local feature map to generate the location feature map.

[0080] In this embodiment, since the positional features of an image can enhance the ability to describe and distinguish the content of an image, before extracting global features based on the local feature map, the positional features of the local feature map are extracted based on the weight matrix of the local feature map, and the positional features are embedded into the local feature map so that the local feature map contains positional features. Therefore, the recognition accuracy can be improved when performing image recognition based on the global feature map containing positional features.

[0081] In some embodiments, the network structure of the second convolutional module is as follows: Figure 3As shown, the system may include a location embedding module, a first branch, and a second branch. The location embedding module embeds the location features of the local feature map into the local feature map. The local feature map with embedded location features is then split into a first local feature map and a second local feature map, which are input into the first and second branches respectively. The first and second branches perform a one-dimensional Fast Fourier Transform (FFT) on the input local feature map and its corresponding weight matrix based on the row and column dimensions, respectively, to obtain a one-dimensional local feature map and weight matrix in the frequency domain. The local feature map in the frequency domain is then multiplied point-by-point with its corresponding weight matrix to obtain a first frequency domain feature map and a second frequency domain feature map. These are then transformed to the spatial domain using an inverse FFT to obtain a first global feature map and a second global feature map. Finally, the first and second global feature maps are concatenated to obtain the global feature map of the image to be recognized. Solid arrows indicate that the data stream (feature data) is in real number form, while dashed arrows indicate that the data stream is in imaginary number format, i.e., the feature data in the frequency domain.

[0082] Corresponding to the image recognition method based on the convolutional neural network model mentioned above, Figure 4 A flowchart illustrating a convolutional neural network model training method provided in an embodiment of this application is shown below:

[0083] Obtain the constructed convolutional neural network model, and input the sample images into the convolutional neural network for training until the convolutional neural network meets the preset requirements, thus obtaining the convolutional neural network model.

[0084] The aforementioned convolutional neural network performs global convolution in the frequency domain on the sample images based on the Fast Fourier Transform.

[0085] Optionally, before training the convolutional neural network (CNN), the CNN can be pre-built according to user needs. This means setting the network structure of the CNN based on the user's image recognition task requirements (e.g., it can be built based on existing ResNet or VGGNet networks) to achieve the corresponding image recognition task. For example, if a user needs to train a CNN model for object detection, and in order to achieve better detection results for objects of different sizes in the image, a CNN can be built based on the SSD (Single Shot MultiBox Detector) network structure. This allows for the detection of objects at different feature scales.

[0086] Specifically, a pre-constructed convolutional neural network (CNN) is acquired, and corresponding sample images are used as input to train the CNN until it meets preset requirements (e.g., the recognition accuracy of the CNN reaches a preset threshold, such as 0.99). Training then stops, resulting in a trained CNN model. During CNN training, the sample images and weight matrices undergo Fast Fourier Transform (FFT) processing to convert them into frequency domain form. Global features are then extracted from the sample images using the weight matrix in the frequency domain. This involves converting the spatial domain input image to the frequency domain using FFT, followed by multiplication operations on the converted frequency domain sample images. This rapid global feature extraction effectively reduces the computational load, especially when the image resolution is high.

[0087] Optionally, the aforementioned sample images are labeled sample images corresponding to the image recognition task performed by the user, so that they can be directly used for training without further annotation. When using sample images for training, a portion of the sample images can be used as the training set, and another portion as the validation and test sets to adjust the convolutional neural network and obtain a good convolutional neural network model. For example, when training a convolutional neural network model for pedestrian re-identification, the Market1501 dataset can be used as the training set. Market1501 contains 32,217 images of 1,501 pedestrians captured by 6 cameras. Each pedestrian is captured by at least 2 cameras, and each camera may contain multiple images, which are then divided into training and test sets.

[0088] Optionally, when using sample images as input to train a convolutional neural network, the sample images input for a single training run can be one or more, such as 100 images. When inputting sample images in batches, the number of input sample images is represented as the batch size. Correspondingly, the feature shape of the sample images extracted by the convolutional neural network can be represented in a four-dimensional format (B, H, W, C), where B represents the batch size, H represents the height, W represents the width, and C represents the channel.

[0089] In this embodiment, a convolutional neural network is pre-constructed according to user needs. Labeled sample images are used as input to the constructed convolutional neural network for training until the network meets preset requirements, resulting in a trained convolutional neural network model. Since the input image is converted from the spatial domain to the frequency domain using the Fast Fourier Transform, converting spatial domain convolution operations into frequency domain multiplication operations, the computational load in the global feature extraction process is effectively reduced. This lowers the computational complexity of the convolutional neural network model for large-scale input images, making it easier to deploy and run on devices with lower computing power, while also improving the computational speed of the convolutional neural network model.

[0090] Corresponding to the above-mentioned image recognition methods or convolutional neural network model training methods, the following section introduces image recognition methods based on convolutional neural network models based on some application scenarios.

[0091] (I) Remote Sensing Detection

[0092] High-resolution remote sensing images are characterized by their inclusion of complex information structures and natural scenes. A single remote sensing image often contains a large amount of information on various land features and geomorphological elements, such as buildings, sites, vegetation, and farmland. Target detection in remote sensing images has always been a hot research topic. Existing target detection models for remote sensing images mostly have deep structures and complex connection channels. However, remote sensing image data is more abundant and covers a larger area than natural images. Using ordinary convolution with large convolution kernels to extract global features from remote sensing images results in excessive computational complexity and low detection efficiency, limiting the deployment and use of these models in scenarios with limited computing resources. The image recognition method based on a convolutional neural network model provided in this application addresses the problem of high computational cost in extracting global features from large-size images using convolution. By converting spatial domain convolution operations into frequency domain multiplication operations, global feature extraction of the image to be detected is performed, effectively reducing the computational cost of global feature extraction and thus improving the detection efficiency of the convolutional neural network model.

[0093] First, the acquired image to be detected is input into the convolutional neural network model (object detection model). The first convolutional module extracts local features from the image to obtain feature information such as edges, corners, and lines, resulting in a local feature map. This local feature map is then used as the input to the second convolutional module. In the second convolution module, the positional features of the image to be detected are embedded into the local feature map through the position embedding module, so that it contains more positional information. Then, the local feature map is split along the channel direction to obtain the first local feature map and the second local feature map, which are input into the first branch and the second branch. The first local feature map and the second local feature map are processed by one-dimensional fast Fourier transform along the column dimension and the row dimension to obtain the first local feature map and the second local feature map in the frequency domain. At the same time, the corresponding weight matrices of the first local feature map and the second local feature map are processed by one-dimensional fast Fourier transform, so that the weight matrix in the frequency domain is multiplied point by point with the corresponding first local feature map and the second local feature map in the frequency domain to obtain the first frequency domain feature map and the second frequency domain feature map. After performing inverse fast Fourier transform, they are concatenated to obtain the complete global feature map. Finally, the detection (recognition) module detects the global feature map and outputs the corresponding detection results. Because the local feature maps are split and calculated in the process of global feature extraction, the amount of computation is reduced. Furthermore, one-dimensional fast Fourier transforms are performed on each feature map to convert the convolution operation in the spatial domain into a multiplication operation determined by the frequency domain. This greatly reduces the amount of computation required for global feature extraction and effectively improves the efficiency of target detection in remote sensing images.

[0094] (II) Facial Recognition

[0095] Facial recognition technology is currently widely used in fields such as smart access control and security monitoring. Since facial recognition requires extracting facial features from images for identification, high-resolution images are typically acquired for facial recognition. For example, facial recognition in smart access control systems requires accurate identification to improve security, which necessitates high-resolution cameras. However, smart access control systems have limited computing resources, and extracting global features from high-resolution images is inefficient, which is not conducive to practical applications. The image recognition method based on a convolutional neural network model provided in this application can be deployed in smart access control systems, effectively reducing the computational load of extracting global features from high-resolution images and improving facial recognition efficiency.

[0096] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0097] Example 2:

[0098] Corresponding to the image recognition method based on the convolutional neural network model described in the above embodiments, Figure 5 The diagram shows a structural block diagram of an image recognition device based on a convolutional neural network model provided in an embodiment of this application. For ease of explanation, only the parts related to the embodiments of this application are shown.

[0099] Reference Figure 5 The device includes: an input module 51 and a convolutional neural network model 52. The convolutional neural network model performs frequency-domain global convolution on the image to be recognized based on the Fast Fourier Transform.

[0100] Input module 51 is used to input the image to be recognized into the above convolutional neural network model;

[0101] The convolutional neural network model 52 is used to extract and recognize features from the above-mentioned image to be recognized in sequence, and obtain the recognition result.

[0102] In this embodiment, the image to be recognized is input into the trained convolutional neural network model. The convolutional neural network model sequentially extracts features and recognizes the image to be recognized to obtain the recognition result. Since the global convolution in the frequency domain of the image to be recognized is based on the Fast Fourier Transform, converting the convolution operation in the spatial domain into a multiplication operation in the frequency domain, the computational load in the process of extracting global features of large-size images by the convolutional neural network model can be effectively reduced when extracting global features of the image to be recognized, thereby improving the recognition efficiency of the convolutional neural network model and facilitating deployment on devices with lower computing power.

[0103] In some embodiments, the image recognition device further includes:

[0104] The image acquisition module is used to acquire the image to be recognized.

[0105] In some embodiments, the above convolutional neural network model 52 includes:

[0106] The feature extraction unit is used to extract features from the image to be identified.

[0107] The recognition unit is used to identify the extracted features and obtain the recognition result.

[0108] In some embodiments, the feature extraction unit includes:

[0109] The first convolutional unit is used to extract local features from the image to be identified to obtain a local feature map.

[0110] The second convolutional unit is used to perform frequency domain global convolution on the aforementioned local feature map using Fast Fourier Transform to obtain the global feature map.

[0111] In some embodiments, the second convolutional unit includes:

[0112] The splitting unit is used to split the above local feature map along the channel direction to obtain a first local feature map and a second local feature map.

[0113] The first branch unit is used to perform frequency domain global convolution on the first local feature map using fast Fourier transform to obtain the first global feature map;

[0114] The second branch unit is used to perform frequency domain global convolution on the second local feature map using fast Fourier transform to obtain the second global feature map;

[0115] The splicing unit is used to splice the first global feature map and the second global feature map along the channel direction to obtain a global feature map.

[0116] In some embodiments, the first branch unit includes:

[0117] The first transformation unit is used to perform fast Fourier transform processing on the first local feature map and the corresponding weight matrix based on the column dimension to obtain the first local feature map and weight matrix in the frequency domain.

[0118] The first convolutional unit is used to multiply the first local feature map in the frequency domain and the weight matrix point by point to obtain the first frequency domain feature map.

[0119] The first inverse transform unit is used to perform a fast inverse Fourier transform on the first feature map in the frequency domain to obtain the first global feature map.

[0120] The aforementioned second branch unit includes:

[0121] The second transformation unit is used to perform fast Fourier transform processing on the above-mentioned second local feature map and the corresponding weight matrix based on the row dimension to obtain the second local feature map and weight matrix in the frequency domain.

[0122] The second convolutional unit is used to multiply the second local feature map in the frequency domain and the weight matrix point by point to obtain the second frequency domain feature map.

[0123] The second inverse transform unit is used to perform a fast inverse Fourier transform on the second feature map in the frequency domain to obtain the second global feature map.

[0124] In some embodiments, the second convolutional unit further includes:

[0125] The location embedding unit is used to extract features from the local feature map through the location embedding module to obtain a location feature map, and add the location feature map to the local feature map to obtain a local feature map containing location features.

[0126] Corresponding to the training method of the convolutional neural network model described in the above embodiments, Figure 6 This paper shows a structural block diagram of a convolutional neural network model training device provided in an embodiment of this application. (Refer to...) Figure 6 The device includes:

[0127] Training module 61 is used to obtain the constructed convolutional neural network model and input the sample image into the convolutional neural network for training until the convolutional neural network meets the preset requirements to obtain the convolutional neural network model. The convolutional neural network performs frequency domain global convolution on the sample image based on fast Fourier transform.

[0128] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0129] Example 3:

[0130] Figure 7 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Figure 7 As shown, the terminal device 7 of this embodiment includes: at least one processor 70 ( Figure 7 The diagram shows only one processor, a memory 71, and a computer program 72 stored in the memory 71 and executable on the at least one processor 70, which, when executed, performs the steps in any of the above-described method embodiments.

[0131] For example, the computer program 72 can be divided into an input module 51 and a convolutional neural network model 52. The convolutional neural network model performs frequency domain global convolution on the image to be recognized based on the fast Fourier transform. The specific functions of each module are as follows:

[0132] Input module 51 is used to input the image to be recognized into the above convolutional neural network model;

[0133] The convolutional neural network model 52 is used to extract and recognize features from the above-mentioned image to be recognized in sequence, and obtain the recognition result.

[0134] Alternatively, the above computer program 72 can be divided into a training module 61, the specific functions of which are as follows:

[0135] Training module 61 is used to obtain the constructed convolutional neural network model and input the sample image into the convolutional neural network for training until the convolutional neural network meets the preset requirements to obtain the convolutional neural network model. The convolutional neural network performs frequency domain global convolution on the sample image based on fast Fourier transform.

[0136] The terminal device 7 can be a desktop computer, laptop, handheld computer, or cloud server, etc. This terminal device may include, but is not limited to, a processor 70 and a memory 71. Those skilled in the art will understand that... Figure 7 This is merely an example of terminal device 7 and does not constitute a limitation on terminal device 7. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0137] The processor 70 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0138] In some embodiments, the memory 71 may be an internal storage unit of the terminal device 7, such as a hard disk or memory of the terminal device 7. In other embodiments, the memory 71 may be an external storage device of the terminal device 7, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the terminal device 7. Furthermore, the memory 71 may include both internal and external storage units of the terminal device 7. The memory 71 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 71 can also be used to temporarily store data that has been output or will be output.

[0139] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0140] This application also provides a network device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the steps in any of the above method embodiments.

[0141] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0142] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.

[0143] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0144] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0145] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0146] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0147] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0148] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for image recognition based on a convolutional neural network model, characterized in that, The convolutional neural network model is based on the Fast Fourier Transform (FFT) to perform frequency domain global convolution on the image to be recognized. The image recognition method includes: The image to be identified is input into the trained convolutional neural network model, and the convolutional neural network model sequentially extracts and identifies features from the image to be identified to obtain the recognition result. The convolutional neural network model includes a feature extraction module and a recognition module. The process of sequentially extracting and recognizing features from the image to be recognized using the convolutional neural network model to obtain a recognition result includes: The feature extraction module performs feature extraction on the image to be identified. The extracted features are identified based on the identification module to obtain the identification result; The feature extraction module includes a first convolution module and a second convolution module. The second convolution module performs a frequency domain global convolution on the image to be identified based on a Fast Fourier Transform. The feature extraction module extracts features from the image to be identified, including: Based on the first convolution module, local feature extraction is performed on the image to be identified to obtain a local feature map; Based on the second convolution module, a fast Fourier transform is used to perform frequency domain global convolution on the local feature map to obtain a global feature map; The second convolutional module includes a first branch and a second branch. The step of performing a frequency-domain global convolution on the local feature map using a fast Fourier transform based on the second convolutional module to obtain a global feature map includes: The local feature map is split along the channel direction to obtain a first local feature map and a second local feature map, and the first local feature map and the second local feature map are respectively input into the first branch and the second branch; The first branch and the second branch respectively use Fast Fourier Transform to perform frequency domain global convolution on the input local feature map to obtain the first global feature map and the second global feature map; The first global feature map and the second global feature map are concatenated along the channel direction to obtain a global feature map; The second convolutional module further includes a location embedding module, which, before splitting the local feature map along the channel direction, also includes: The location embedding module extracts features from the local feature map to obtain a location feature map, and then adds the location feature map to the local feature map to obtain a local feature map containing location features. The location embedding module is a two-dimensional convolution module.

2. The image recognition method of claim 1, wherein, The first branch and the second branch respectively use Fast Fourier Transform to perform frequency domain global convolution on the input local feature maps to obtain a first global feature map and a second global feature map, including: The first branch performs Fast Fourier Transform on the first local feature map and the corresponding weight matrix based on the column dimension to obtain the first local feature map and weight matrix in the frequency domain. The first local feature map in the frequency domain and the weight matrix are multiplied point by point to obtain the first frequency domain feature map; The first feature map in the frequency domain is processed by inverse fast Fourier transform to obtain the first global feature map; The second branch performs Fast Fourier Transform on the second local feature map and the corresponding weight matrix based on the row dimension to obtain the second local feature map and weight matrix in the frequency domain. The second local feature map in the frequency domain and the weight matrix are multiplied point by point to obtain the second frequency domain feature map; The second feature map in the frequency domain is processed by inverse fast Fourier transform to obtain the second global feature map.

3. An image recognition apparatus characterized by comprising: include: The input module and the trained convolutional neural network model, which performs frequency domain global convolution on the image to be recognized based on the fast Fourier transform; The input module is used to input the image to be recognized into the convolutional neural network model; The convolutional neural network model is used to sequentially extract and recognize features from the image to be recognized, and obtain the recognition result. The convolutional neural network model includes a feature extraction module and a recognition module. The feature extraction and recognition of the image to be recognized are performed sequentially to obtain the recognition result, including: The feature extraction module performs feature extraction on the image to be identified. The extracted features are identified based on the identification module to obtain the identification result; The feature extraction module includes a first convolution module and a second convolution module. The second convolution module performs a frequency domain global convolution on the image to be identified based on a Fast Fourier Transform. The feature extraction module extracts features from the image to be identified, including: Based on the first convolution module, local feature extraction is performed on the image to be identified to obtain a local feature map; Based on the second convolution module, a fast Fourier transform is used to perform frequency domain global convolution on the local feature map to obtain a global feature map; The second convolutional module includes a first branch and a second branch. The step of performing a frequency-domain global convolution on the local feature map using a fast Fourier transform based on the second convolutional module to obtain a global feature map includes: The local feature map is split along the channel direction to obtain a first local feature map and a second local feature map, and the first local feature map and the second local feature map are respectively input into the first branch and the second branch; The first branch and the second branch respectively use Fast Fourier Transform to perform frequency domain global convolution on the input local feature map to obtain the first global feature map and the second global feature map; The first global feature map and the second global feature map are concatenated along the channel direction to obtain a global feature map; The second convolutional module further includes a location embedding module, which, before splitting the local feature map along the channel direction, also includes: The location embedding module extracts features from the local feature map to obtain a location feature map, and then adds the location feature map to the local feature map to obtain a local feature map containing location features. The location embedding module is a two-dimensional convolution module.

4. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the image recognition method as described in claim 1 or 2.

5. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 4. When the computer program is executed by the processor, it implements the image recognition method as described in claim 1 or 2.

Citation Information

Patent Citations

  • Underwater fish target detection method and device and storage medium

    CN113869330A