Image recognition method and device based on convolutional neural network model and terminal equipment
By dynamically adjusting the weight matrix in a convolutional neural network for global feature extraction, and combining local and global feature processing, the problem of low accuracy in recognizing images of different sizes is solved, achieving a more efficient image recognition effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
- Filing Date
- 2022-10-13
- Publication Date
- 2026-05-22
AI Technical Summary
Existing convolutional neural networks cannot effectively adapt to input images of different sizes when performing feature extraction, resulting in a decrease in recognition accuracy.
Global feature extraction is performed based on the weight matrix of the image to be identified. The weight matrix is dynamically adjusted to adapt to images of different sizes. Combining local and global feature extraction, a multi-layer convolutional neural network model is used for feature recognition.
This improved the accuracy of convolutional neural networks in recognizing images of different sizes, enhancing the model's adaptability and recognition performance.
Smart Images

Figure CN115578590B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image recognition technology, and in particular relates to image recognition methods, devices, terminal equipment, and computer-readable storage media based on convolutional neural network models. Background Technology
[0002] Convolutional Neural Networks (CNNs) are a special type of artificial neural network that have become one of the most commonly used techniques in speech analysis and image recognition, and are therefore a hot research topic in artificial neural networks. A CNN is typically a multi-layered neural network, with each layer consisting of multiple feature maps, and each feature map composed of multiple independent neurons. Neurons within the same feature map share weights (i.e., the convolutional kernel). CNNs reduce the connections between layers by sharing weights, thereby mitigating the risk of overfitting.
[0003] Existing convolutional neural networks typically use a small kernel matrix (i.e., weight matrix) and a large feature map for convolution operations when performing feature extraction in order to save parameters. However, they share a single static weight matrix and cannot adapt well to input images of different sizes. Summary of the Invention
[0004] This application provides an image recognition method, apparatus, and terminal device based on a convolutional neural network model, which can improve the image recognition performance of the convolutional neural network model for images of different sizes.
[0005] In a first aspect, embodiments of this application provide an image recognition method based on a convolutional neural network model. The convolutional neural network model performs global feature extraction on the image to be recognized based on a weight matrix of the image to be recognized. The image recognition method includes:
[0006] The image to be identified is input into the trained convolutional neural network model, and the convolutional neural network model sequentially extracts and identifies features from the image to obtain the identification result.
[0007] Secondly, embodiments of this application provide a method for training a convolutional neural network model, including:
[0008] Obtain the constructed convolutional neural network model, and input the sample image into the convolutional neural network for training until the convolutional neural network meets the preset requirements, thus obtaining the convolutional neural network model;
[0009] The convolutional neural network described above extracts global features from the sample images based on the weight matrix of the sample images.
[0010] Thirdly, embodiments of this application provide an image recognition device, including:
[0011] The input module and the trained convolutional neural network model, which performs global feature extraction on the image to be identified based on the weight matrix of the image to be identified;
[0012] The input module is used to input the image to be recognized into the convolutional neural network model;
[0013] The convolutional neural network model is used to sequentially extract and recognize features from the image to be recognized, thereby obtaining the recognition result.
[0014] Fourthly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the image recognition method based on a convolutional neural network model described in the first aspect or the convolutional neural network model training method described in the second aspect.
[0015] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the image recognition method based on a convolutional neural network model described in the first aspect or the convolutional neural network model training method described in the second aspect.
[0016] In a sixth aspect, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the image recognition method based on a convolutional neural network model as described in any one of the first aspects or the convolutional neural network model training method described in the second aspect.
[0017] The beneficial effects of the embodiments in this application compared with the prior art are:
[0018] In this embodiment, the image to be recognized is input into the trained convolutional neural network model. The convolutional neural network model sequentially extracts features and recognizes the image to be recognized to obtain a recognition result. Since the convolutional neural network model performs global feature extraction on the image to be recognized based on the weight matrix of the image to be recognized, when performing image recognition on images of different sizes, the weight matrix can be dynamically adjusted according to each image to extract global features, enabling the convolutional neural network model to have good recognition performance for input images of different sizes. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0020] Figure 1 This is a schematic flowchart of an image recognition method based on a convolutional neural network model provided in an embodiment of this application;
[0021] Figure 2 This is a schematic diagram of the structure of the convolutional neural network model provided in the embodiments of this application;
[0022] Figure 3 This is a schematic diagram of the structure of the convolution module for dynamically extracting the weight matrix provided in an embodiment of this application;
[0023] Figure 4 This is a schematic diagram of the structure of the second convolution module provided in the embodiments of this application;
[0024] Figure 5 This is a flowchart illustrating the convolutional neural network model training method provided in the embodiments of this application;
[0025] Figure 6 This is a schematic diagram of the structure of the image recognition device based on the convolutional neural network model provided in the embodiments of this application;
[0026] Figure 7 This is a schematic diagram of the structure of the convolutional neural network model training device provided in the embodiments of this application;
[0027] Figure 8 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. Detailed Implementation
[0028] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0029] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0030] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0031] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0032] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.
[0033] Example 1:
[0034] Figure 1 A flowchart illustrating an image recognition method based on a convolutional neural network model according to an embodiment of the present invention is shown below in detail:
[0035] The image to be identified is input into a trained convolutional neural network model. The convolutional neural network model then sequentially extracts and identifies features from the image to obtain the identification result.
[0036] The convolutional neural network model described above performs global feature extraction on the image to be recognized based on the weight matrix of the image to be recognized, in order to perform image recognition. That is, it extracts global features from the image to be recognized based on the weight matrix of the currently input image to be recognized.
[0037] The weight matrix mentioned above refers to the convolution kernel of a convolutional neural network. In image processing, given a small region of an input image, the pixels are weighted and averaged to become each corresponding pixel in the output image. The weights are defined by a function, which is the convolution kernel. For different input images, the weight matrix (convolution kernel) of the convolutional neural network remains constant.
[0038] Specifically, in image recognition tasks, the images to be recognized vary in size. Using the same fixed weight matrix cannot dynamically adjust the weight matrix based on the input image. Furthermore, traditional convolutional operations only have a local receptive field and cannot effectively extract global features, affecting the accuracy of image recognition. Therefore, during the feature extraction process of the image to be recognized input into the aforementioned convolutional neural network model, global features of the image to be recognized are extracted based on its weight matrix. Recognition is then performed based on the obtained features to obtain the image recognition result. This allows the convolutional neural network model to simultaneously possess a global receptive field and dynamic weights, thereby improving the recognition accuracy of the convolutional neural network model. For example, for image A to be recognized, global features of image A are extracted based on its weight matrix 'a', and recognition is performed based on the obtained global features to obtain the recognition result of image A. When recognizing image B, global features of image B are extracted and recognized based on its weight matrix 'b'.
[0039] In this embodiment, the image to be recognized is input into a trained convolutional neural network (CNN) model. The CNN model sequentially extracts features from the image and performs recognition to obtain the recognition result. Because the CNN model extracts global features from the image based on its weight matrix during feature extraction, and because it possesses both a global receptive field and dynamic weights, it can effectively extract features from images of different sizes, thereby improving the recognition accuracy of the CNN model.
[0040] In some embodiments, the above-described image recognition method based on a convolutional neural network model further includes:
[0041] Obtain the image to be recognized.
[0042] Optionally, the image to be identified may be an image captured by a camera device or an image frame in a video stream captured by a camera device.
[0043] Optionally, since different image recognition tasks may require different images to be recognized, and the camera equipment used and the rules for acquiring the images to be detected may also differ, the corresponding images to be recognized should be obtained according to the acquisition methods and rules for each application field. For example, for a pedestrian re-identification task, images acquired by multiple installed cameras need to be obtained as images to be recognized in order to perform the pedestrian re-identification task.
[0044] In this embodiment of the application, based on the images required for image recognition tasks in various application fields, corresponding acquisition methods and acquisition rules are adopted to obtain images to be recognized that meet the requirements of image recognition tasks, so as to perform image recognition tasks.
[0045] In some embodiments, the convolutional neural network model includes a feature extraction module and a recognition module. The steps described above involve sequentially extracting and recognizing features from the image to be recognized using the convolutional neural network model to obtain a recognition result, including:
[0046] A1. The above feature extraction module is used to extract features from the image to be identified.
[0047] A2. Based on the above recognition module, the extracted features are recognized to obtain the recognition results.
[0048] Optionally, since image recognition includes different tasks such as image classification and object detection (e.g., face recognition, pedestrian detection), different image recognition tasks employ different recognition methods for the same feature. Therefore, the feature extraction module extracts features from the input image to be recognized and uses the extracted features as input to the recognition module. Based on the image recognition task, corresponding recognition is performed to obtain the recognition result. The recognition module may include one or more recognition units, each performing different recognition tasks. For example, it may include a face recognition unit and an object detection unit. The face recognition unit performs face recognition based on the extracted features, or the extracted feature distribution is input to both the face recognition unit and the object detection unit for both face recognition and object detection tasks.
[0049] In this embodiment, the feature extraction module extracts features from the image to be recognized, and the recognition module obtains the extracted features for corresponding recognition to obtain the corresponding recognition results, thereby improving the recognition efficiency of each image recognition task.
[0050] In some embodiments, the feature extraction module includes a first convolution module and a second convolution module, and step A1 includes:
[0051] A11. Based on the first convolutional module described above, local features are extracted from the image to be identified to obtain a local feature map;
[0052] A12. Based on the second convolution module, global feature extraction is performed on the local feature map according to the weight matrix of the local feature map to obtain the global feature map.
[0053] Optionally, the above convolutional neural network model can be constructed based on existing convolutional neural networks. Shallow convolutional layers are used as the first convolutional module, and deep convolutional layers or self-attention layers are replaced with second convolutional modules. The first convolutional module performs convolution processing on the image to be recognized using ordinary convolution, extracting local features of the image and outputting a local feature map. This local feature map is then used as input to the second convolutional module. The second module obtains the weight matrix of the local feature map through dynamic convolution and extracts global features from the local feature map based on this weight matrix, resulting in a global feature map. For example, as... Figure 2 The convolutional neural network model shown has two layers: the first two layers are ordinary convolutional first convolutional modules, and the four deeper convolutional layers are second convolutional modules. The second convolutional modules are connected to the recognition module. The image to be recognized is used as the input of the first convolutional module for local feature extraction, and the output features are input to the second convolutional module for global feature extraction. The global feature map output by the second convolutional module is used as the input of the recognition module for recognition, thereby outputting the corresponding recognition result.
[0054] It should be noted that the first and second convolutional modules in the above convolutional neural network model can also adopt an alternating structure, that is, the first convolutional module is connected to the second convolutional module, and the output of the second convolutional module is connected to another first convolutional module. Figure 2 The structure shown is a stacked structure, which means that the feature map output by the feature extraction module is the global feature map extracted by the second convolution module (that is, the recognition module recognizes based on the global features extracted by the second convolution module). It does not limit the specific structure of the first convolution module (ordinary convolution layer) and the second convolution module provided in the embodiments of this application in the convolutional neural network model.
[0055] In this embodiment, since the local features of the image to be identified are extracted by the first convolutional module of the convolutional neural network model, and the obtained local feature map is used as the input of the second convolutional module, global feature extraction is performed on the local feature map according to the weight matrix of the local feature map. Therefore, the obtained global features contain both local features and global features, which improves the recognition accuracy of the convolutional neural network model.
[0056] In some embodiments, the second convolutional module includes a first branch and a second branch, and step A12 includes:
[0057] The aforementioned local feature map is split along the channel direction to obtain a first local feature map and a second local feature map, and the first local feature map and the second local feature map are respectively input into the first branch and the second branch.
[0058] Optionally, when the second convolutional module performs global feature extraction on the input local feature map, the local feature map is first split along the channel direction (e.g., evenly split into two parts along the channel direction) to obtain the first local feature map and the second local feature map. The first local feature map is then input into the first branch, and the second local feature map is input into the second branch, so as to extract global features from the first local feature map and the second local feature map respectively.
[0059] The first branch and the second branch described above respectively extract features from the corresponding local feature maps based on the weight matrix of the input local feature maps, to obtain the first global feature map and the second global feature map.
[0060] Optionally, the first branch generates a weight matrix of the first local feature map based on the first local feature map, and extracts global features from the first local feature map based on the weight matrix to obtain a first global feature map; the second branch generates a weight matrix of the second local feature map based on the second local feature map, and extracts global features from the second local feature map based on the weight matrix to obtain a second global feature map.
[0061] The first global feature map and the second global feature map are concatenated along the channel direction to obtain the global feature map.
[0062] Optionally, since the first local feature map and the second local feature map are obtained by splitting the local feature map along the channel direction, after extracting the features of the first local feature map and the second local feature map from the first branch and the second branch respectively, the obtained first global feature map and the second global feature map are spliced along the channel direction to obtain a complete global feature map, so as to perform subsequent processing based on the complete global feature map of the input image to be recognized.
[0063] In this embodiment, since the local feature map is split into two parts along the channel direction and input to the first branch and the second branch, the number of channels of the local feature map is halved. Therefore, global feature extraction is performed based on the local feature map with halved channel number, which reduces the computational complexity of global feature extraction, thereby reducing the requirements for device computing power and making it easier to deploy applications on devices with low computing power.
[0064] In some embodiments, when the first branch and the second branch extract global features based on the input local feature map, the process includes:
[0065] Based on the first local feature map, a column-dimensional weight matrix is generated to obtain the first weight matrix. Then, the first local feature map is convolved along the column dimension based on the first weight matrix to obtain the first global feature map.
[0066] Specifically, when the first branch performs feature extraction on the first local feature map, it performs convolution processing along the column dimension (i.e., the H dimension) based on the first local feature map to dynamically generate the weight matrix of the first local feature map, thereby obtaining the first weight matrix. Then, it performs global feature extraction on the first local feature map based on the first weight matrix to obtain the first global feature map.
[0067] Optionally, to reduce the computational complexity of the convolutional neural network model, when dynamically generating the first weight matrix along the column dimension based on the first local feature map, a one-dimensional convolution is performed on the first local feature map along the column dimension. That is, a convolution operation is performed on each row of pixels in the height direction. For example, if the first local feature map is represented as (B, C, H, W), a lightweight one-dimensional convolution module is used to perform convolution processing along the column dimension on the first local feature map to generate the first weight matrix (B, C, H, 1) of the first local feature map. Since the first weight matrix is obtained by performing convolution operations on each row of pixels in the column dimension, the first weight matrix has a global receptive field. Therefore, after obtaining the first weight matrix of the first local feature map, a cyclic convolution is performed on the first local feature map along the column dimension based on the first weight matrix to extract the global features of the first local feature map and obtain the first global feature map. For example, assuming the height of the first local feature map is 4 and the batch size is 1, since the first weight matrix is obtained by performing a convolution operation along the column dimension on the first feature map, the first local feature map is represented as x = (x0, x1, x2, x3), and the first weight matrix (equivalent to the convolution kernel) is represented as w = (w0, w1, w2, w3). This is equivalent to treating each row of pixels of the first local feature map as an element, so that the first local feature map and the first weight matrix have the same resolution (e.g., size (C, H, 1)).
[0068] Optionally, when performing one-dimensional convolution processing on the first local feature map to obtain the first weight matrix, depthwise separable convolution is used to reduce the number of parameters and computational cost. For example, ... Figure 3As shown, the one-dimensional convolutional module employs a lightweight network structure with two convolutional layers. It performs average pooling along the column dimension on the first local feature map (B, C, H, W) to reduce its dimensionality, resulting in a first local feature map of shape (B, C, H, 1). Then, it performs depthwise separable convolution with a 1×3 kernel on the first local feature map, followed by HardSwish activation and batch normalization. This is followed by another 1×3 depthwise separable convolution to output the first weight matrix. The HardSwish activation function is an improvement on the Swish non-linear activation function. While the Swish activation function can improve the accuracy of convolutional neural networks to some extent, its computational cost is high, making it unsuitable for use on embedded mobile devices. The HardSwish activation function can be implemented as a segmentation function to reduce memory accesses, improving the accuracy of the convolutional neural network while facilitating deployment on embedded mobile devices.
[0069] Based on the local feature map of the second branch, a row-dimensional weight matrix is generated to obtain the second weight matrix. Then, the second local feature map is convolved along the row dimension based on the second weight matrix to obtain the second global feature map.
[0070] Optionally, when performing feature extraction on the second local feature map in the second branch, a weight matrix of the second local feature map is dynamically generated based on the second local feature map to obtain a second weight matrix. Then, a convolution process is performed based on the second weight matrix and the second local feature map to obtain a second global feature map.
[0071] Optionally, when generating the second weight matrix based on the second local feature map, a one-dimensional convolution is performed on the second local feature map along the row dimension (i.e., the horizontal dimension). This means performing a convolution operation on each column of pixels in the width direction. For example, if the second local feature map is represented as (B, C, H, W), a lightweight one-dimensional convolution module is used to perform convolution along the W dimension to generate the second weight matrix (B, C, I, W) of the second local feature map. After obtaining the second weight matrix of the second local feature map, a circular convolution is performed on the second local feature map along the row dimension based on the second weight matrix to extract the global features of the second local feature map, thus obtaining the second global feature map.
[0072] Optionally, when performing one-dimensional convolution processing on the aforementioned second local feature map to obtain the second weight matrix, depthwise separable convolution can be used, similar to the first branch, to reduce the number of parameters and computational cost. For example, the one-dimensional convolution module adopts the lightweight network structure shown in the figure, where average pooling is performed along the row dimension to obtain the second weight matrix of shape (B, C, 1, W), and the 1×3 depthwise separable convolution is replaced with a 3×1 depthwise separable convolution.
[0073] In this embodiment, since the local feature maps are cyclically convolved with one-dimensional weight matrices from different dimensions to obtain global features of different dimensions, and the obtained global features are concatenated to obtain global features including width and height directions, the computational load when extracting global features from local feature maps is reduced. At the same time, the weight matrices of the first and second local feature maps are dynamically extracted respectively, and global features are extracted from the corresponding local feature maps based on their weight matrices. This makes the convolutional neural network model adaptable to images to be recognized at different scales, thereby improving the image recognition performance of the convolutional neural network model.
[0074] In some embodiments, the second convolutional module further includes a location embedding module, and the step of splitting the local feature map along the channel direction further includes:
[0075] The location embedding module extracts features from the local feature map based on the weight matrix of the local feature map to obtain a location feature map. The location feature map is then added to the local feature map to obtain a local feature map containing location features.
[0076] Optionally, before performing global feature extraction on the local feature map based on the weight matrix of the local feature map, the local feature map is convolved by the location embedding module to dynamically generate the weight matrix of the local feature map. Local features are then extracted from the local feature map based on the weight matrix to generate a location feature map. The location feature map is then added to the local feature map to obtain a local feature map embedded with location features.
[0077] Optionally, the aforementioned location embedding module is a two-dimensional convolutional module that performs convolution processing on the input sample image to generate a two-dimensional location feature map. That is, the size of the location feature map is made consistent with the resolution of the input local feature map, so that the location feature map can be directly added to the local feature map to obtain a local feature map containing location features. For example, the location embedding module adopts a two-layer lightweight convolutional network structure, namely a simple structure of "convolution + normalization processing + activation function + convolution". Since it is necessary to generate a two-dimensional location feature map, the convolutional layer can use a 3×3 depthwise separable convolution to perform convolution processing on the local feature map to generate the location feature map.
[0078] In this embodiment, since the positional features of an image can enhance the ability to describe and distinguish the content of an image, before extracting global features based on the local feature map, the positional features of the local feature map are extracted based on the weight matrix of the local feature map, and the positional features are embedded into the local feature map so that the local feature map contains positional features. Therefore, the recognition accuracy can be improved when performing image recognition based on the global feature map containing positional features.
[0079] In some embodiments, the network structure of the second convolutional module is as follows: Figure 4 As shown, the system may include a location embedding module, a first branch, and a second branch. The location embedding module embeds the location features of the local feature map into the local feature map. The local feature map with embedded location features is then split into a first local feature map and a second local feature map, which are input into the first branch and the second branch respectively. The first branch and the second branch extract global features based on the row and column dimensions, respectively, according to the weight matrices of the first and second local feature maps, resulting in a first global feature map and a second global feature map. The first and second global feature maps are then concatenated to obtain the global feature map. At this point, the first branch performs a circular convolution operation along the column dimension on the first local feature map according to the first weight matrix to extract the global features of the first local feature map. This circular convolution operation can be represented as:
[0080]
[0081] w H =f(w,H)
[0082] x P =x + f(pe,H)
[0083] Among them, y i This represents the i-th element of the first global feature output, and H refers to the row number (i.e., height) of the first local feature map. It refers to the k-th weight (convolution kernel) in the first weight matrix, where w H This means that if the length of the first weight matrix is not H, then an interpolation operation is performed on the first weight matrix to make its length H. H This refers to the periodic extension of the first local feature map at point H. This refers to taking the (k+i)th element from the input x (the first local feature map), x P It refers to the first local feature map with added positional features, where p is pe, representing the positional feature corresponding to the first local feature map.
[0084] Corresponding to the image recognition method based on the convolutional neural network model mentioned above, Figure 5 A flowchart illustrating a convolutional neural network model training method provided in an embodiment of this application is shown below:
[0085] Obtain the constructed convolutional neural network model, and input the sample images into the convolutional neural network for training until the convolutional neural network meets the preset requirements, thus obtaining the convolutional neural network model.
[0086] Optionally, a convolutional neural network (CNN) can be constructed according to user needs (e.g., for face recognition tasks, the CNN needs to be configured with a classification head to classify and recognize face images). This involves setting the network structure of the CNN based on user requirements (e.g., it can be built based on existing VGGNet or AlexNet network structures) and initializing the constructed CNN. For example, to alleviate the gradient vanishing and gradient exploding problems that exist when the CNN has many layers, a CNN can be constructed based on the ResNet18 residual network structure, which includes 17 convolutional layers and 1 fully connected layer.
[0087] Specifically, a pre-constructed convolutional neural network is obtained, and the corresponding sample images are used as input to train the convolutional neural network until it meets preset requirements (such as the recognition accuracy of the convolutional neural network reaching a preset threshold, such as 0.988). At this point, training of the convolutional neural network stops, resulting in a trained convolutional neural network model. During training, the convolutional neural network extracts global features from the input image based on the weight matrix of the input image, which is the input sample image.
[0088] Optionally, before the convolutional neural network extracts global features from the input image based on its weight matrix, it generates a weight matrix for the input image through convolution, and then extracts global features from the input image based on this weight matrix. It should be noted that since the convolutional neural network uses dynamic weights, not a fixed weight matrix, during the training process, when updating the parameters of the convolutional neural network according to the loss function, updating the weight matrix involves updating the parameters of the convolutional module that generates the weight matrix of the input image.
[0089] Optionally, the aforementioned sample images are labeled sample images corresponding to the image recognition task performed by the user, so that they can be directly used for training without further annotation. For example, when training a convolutional neural network model for face recognition, the CelebA face attribute dataset can be used as the training set. CelebA contains 202,599 face images for 10,177 identities, each image with feature labels, which can be used as input to the convolutional neural network for training. When using sample images for training, a portion of the sample images can be used as the training set, and another portion as the validation and test sets to adjust the convolutional neural network and obtain a good convolutional neural network model.
[0090] Optionally, when using sample images as input to train a convolutional neural network, the sample images input for a single training run can be one or more, such as 100 images. When inputting sample images in batches, the number of input sample images is represented as the batch size. Correspondingly, the feature shape of the sample images extracted by the convolutional neural network can be represented in a four-dimensional format (B, H, W, C), where B represents the batch size, H represents the height, W represents the width, and C represents the channel.
[0091] In this embodiment, a convolutional neural network is constructed according to user requirements. Labeled sample images are used as input to the convolutional neural network for training until the convolutional neural network meets preset requirements, resulting in a trained convolutional neural network model. Since the convolutional neural network includes a dynamic convolution module for global feature extraction of the input image based on the weight matrix of the input image, feature extraction is performed on each sample image according to the weight matrix corresponding to each sample image during the training process. Each sample image does not need to use the same static weight matrix, which enables the convolutional neural network to adapt well to sample images of different scales during training, reduces the difficulty of multi-scale training of the convolutional neural network, and enables the convolutional neural network model to have good recognition performance for multi-scale images.
[0092] Corresponding to the above-mentioned image recognition methods or convolutional neural network model training methods, the following section introduces image recognition methods based on convolutional neural network models based on some application scenarios.
[0093] (I) Pedestrian Detection
[0094] Pedestrian detection has always been a hot topic and a challenge in computer vision research. The problem it aims to solve is to identify all pedestrians in an image or video frame, including their position and size, typically represented by rectangular boxes. Pedestrian detection can be combined with techniques such as pedestrian tracking and pedestrian re-identification for applications in autonomous driving, intelligent transportation, and other fields, thus possessing significant application value. Since the images to be detected vary in size, the image recognition method based on a convolutional neural network model provided in this application specifically addresses this issue. It dynamically acquires the weight matrix of each image to be identified and extracts global features from the images, thereby improving the accuracy of pedestrian detection.
[0095] First, the acquired image to be identified is input into a convolutional neural network model (pedestrian detection model). The first convolutional module extracts local features from the image to obtain edge, corner, and line features, resulting in a local feature map. This local feature map is then used as input to the second convolutional module. Since pedestrian detection detects pedestrians in an image, there are strong spatial relationships between pedestrian features (such as body and head). Therefore, a location embedding module performs convolution processing on the local feature map to obtain a location feature map. This location feature map is then added to the local feature map to include location information. The local feature map is then split along the channel direction to obtain a first and a second local feature map, which are input into the first and second branches. Convolution processing is performed on the first and second local feature maps along the column and row dimensions to obtain a weight matrix for the image to be identified. Convolution processing is then performed on the corresponding local feature maps according to the weight matrix to obtain a first and a second global feature map corresponding to the first and second local feature maps. These are then concatenated along the channel direction to obtain a complete global feature map. The recognition module obtains the global feature map output by the second convolutional module, detects existing pedestrian features based on the global feature map, and outputs the corresponding detection results.
[0096] (II) Target Detection
[0097] In object detection tasks, the sizes of the image to be detected and the target are often not fixed. For example, detection tasks related to autonomous driving may need to detect large trucks and animals at the same time, and medical lesion detection tasks may need to detect lesions of different sizes at the same time. When the scale of the objects to be detected differs greatly, it is usually done by acquiring images of different sizes to be identified to detect targets of different sizes. However, the model usually has difficulty adapting well to the detection of targets with large size differences. In this case, based on the convolutional neural network model provided in this application, the weight matrix can be dynamically adjusted according to the input image to be identified, and the image to be identified can be detected according to the weight matrix of each image to be identified, thereby achieving better detection results.
[0098] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0099] Example 2:
[0100] Corresponding to the image recognition method based on the convolutional neural network model described in the above embodiments, Figure 6 A structural block diagram of an image recognition device provided in an embodiment of this application is shown. For ease of explanation, only the parts related to the embodiments of this application are shown.
[0101] Reference Figure 6 The device includes: an input module 61 and a convolutional neural network model 62, wherein the convolutional neural network model performs global feature extraction on the image to be identified based on the weight matrix of the image to be identified.
[0102] Input module 61 is used to input the image to be recognized into the above convolutional neural network model;
[0103] The convolutional neural network model 62 is used to extract and recognize features from the above-mentioned image to be recognized in sequence, and obtain the recognition result.
[0104] In this embodiment, the image to be recognized is input into a trained convolutional neural network (CNN) model. The CNN model sequentially extracts features from the image and performs recognition to obtain the recognition result. Because the CNN model extracts global features from the image based on its weight matrix during feature extraction, and because it possesses both a global receptive field and dynamic weights, it can effectively extract features from images of different sizes, thereby improving the recognition accuracy of the CNN model.
[0105] In some embodiments, the image recognition device further includes:
[0106] The image acquisition module is used to acquire the image to be recognized.
[0107] In some embodiments, the convolutional neural network model 62 includes:
[0108] The feature extraction unit is used to extract features from the image to be identified.
[0109] The recognition unit is used to identify the extracted features and obtain the recognition result.
[0110] In some embodiments, the feature extraction unit includes:
[0111] The first convolutional unit is used to extract local features from the image to be identified to obtain a local feature map.
[0112] The second convolutional unit is used to extract global features from the local feature map based on the weight matrix of the local feature map, so as to obtain a global feature map.
[0113] In some embodiments, the second convolutional unit includes:
[0114] The splitting unit is used to split the above local feature map along the channel direction to obtain a first local feature map and a second local feature map.
[0115] The first branch unit is used to extract global features from the first local feature map based on the weight matrix of the first local feature map to obtain the first global feature map;
[0116] The second branch unit is used to extract global features from the second local feature map based on the weight matrix of the second local feature map, so as to obtain the second global feature map.
[0117] The splicing unit is used to splice the first global feature map and the second global feature map along the channel direction to obtain a global feature map.
[0118] In some embodiments, the first branch unit includes:
[0119] The first weighting unit is used to perform convolution processing on the first local feature image to generate the first weight matrix;
[0120] The global feature extraction unit is used to perform global feature extraction on the first local feature map according to the first weight matrix mentioned above, so as to obtain the first global feature map;
[0121] The aforementioned second branch unit includes:
[0122] The second weighting unit is used to perform convolution processing on the second local feature image to generate the second weighting matrix;
[0123] The global feature extraction unit is used to extract global features from the second local feature map based on the second weight matrix mentioned above, so as to obtain the second global feature map.
[0124] In some embodiments, the second convolutional unit further includes:
[0125] The location embedding unit is used to extract features from the local feature map based on the weight matrix of the local feature map to obtain a location feature map, and to add the location feature map to the local feature map to obtain a local feature map containing location features.
[0126] Corresponding to the convolutional neural network model training method described in the above embodiments, Figure 7 This paper shows a structural block diagram of a convolutional neural network model training device provided in an embodiment of this application. (Refer to...) Figure 7 The device includes:
[0127] Training module 71 is used to obtain the constructed convolutional neural network model and input the sample image into the convolutional neural network for training until the convolutional neural network meets the preset requirements to obtain the convolutional neural network model. The convolutional neural network performs global feature extraction on the sample image based on the weight matrix of the sample image.
[0128] In this embodiment, a convolutional neural network is constructed according to user requirements. Labeled sample images are used as input to the convolutional neural network for training until the convolutional neural network meets preset requirements, resulting in a trained convolutional neural network model. Since the convolutional neural network includes a dynamic convolution module for global feature extraction of the input image based on the weight matrix of the input image, feature extraction is performed on each sample image according to the weight matrix corresponding to each sample image during the training process. Each sample image does not need to use the same static weight matrix, which enables the convolutional neural network to adapt well to sample images of different scales during training, reduces the difficulty of multi-scale training of the convolutional neural network, and enables the convolutional neural network model to have good recognition performance for multi-scale images.
[0129] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0130] Example 3:
[0131] Figure 8 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Figure 8 As shown, the terminal device 8 of this embodiment includes: at least one processor 80 ( Figure 8 The diagram shows only one processor, a memory 81, and a computer program 82 stored in the memory 81 and executable on the at least one processor 80, which, when executing the computer program 82, performs the steps in any of the above method embodiments.
[0132] For example, the computer program 82 can be divided into one or more modules / units, which are stored in the memory 81 and executed by the processor 80 to complete this application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 82 in the terminal device 81. For example, the computer program 82 can be divided into an input module 61 and a convolutional neural network model 62, wherein the convolutional neural network model performs global feature extraction on the image to be recognized based on the weight matrix of the image to be recognized. The specific functions of each module are as follows:
[0133] Input module 61 is used to input the image to be recognized into the above convolutional neural network model;
[0134] The convolutional neural network model 62 is used to extract and recognize features from the above-mentioned image to be recognized in sequence, and obtain the recognition result.
[0135] Alternatively, the computer program 82 can be divided into a training module 71, the specific functions of which are as follows:
[0136] Training module 71 is used to obtain the constructed convolutional neural network model and input the sample image into the convolutional neural network for training until the convolutional neural network meets the preset requirements to obtain the convolutional neural network model. The convolutional neural network performs global feature extraction on the sample image based on the weight matrix of the sample image.
[0137] The terminal device 8 can be a desktop computer, laptop, handheld computer, or cloud server, etc. This terminal device may include, but is not limited to, a processor 80 and a memory 81. Those skilled in the art will understand that... Figure 8 This is merely an example of terminal device 8 and does not constitute a limitation on terminal device 8. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.
[0138] The processor 80 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0139] In some embodiments, the memory 81 may be an internal storage unit of the terminal device 8, such as a hard disk or memory of the terminal device 8. In other embodiments, the memory 81 may be an external storage device of the terminal device 8, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the terminal device 8. Furthermore, the memory 81 may include both internal and external storage units of the terminal device 8. The memory 81 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 81 can also be used to temporarily store data that has been output or will be output.
[0140] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0141] This application also provides a network device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the steps in any of the above method embodiments.
[0142] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0143] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.
[0144] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0145] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0146] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0147] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0148] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0149] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. An image recognition method based on a convolutional neural network model, characterized in that, The convolutional neural network model extracts global features from the image to be identified based on the weight matrix of the image to be identified; The image recognition method includes: The image to be identified is input into the trained convolutional neural network model, which then sequentially extracts and identifies features from the image to obtain the identification result. The convolutional neural network model includes a feature extraction module and a recognition module. The process of sequentially extracting and recognizing features from the image to be recognized using the convolutional neural network model to obtain a recognition result includes: The feature extraction module performs feature extraction on the image to be identified. The extracted features are identified based on the identification module to obtain the identification result; The feature extraction module includes a first convolution module and a second convolution module. The second convolution module performs global feature extraction on the image to be identified based on the weight matrix of the image to be identified. The feature extraction of the image to be identified by the feature extraction module includes: Based on the first convolution module, local feature extraction is performed on the image to be identified to obtain a local feature map; Based on the second convolution module, global features are extracted from the local feature map according to the weight matrix of the local feature map to obtain a global feature map; The second convolutional module includes a first branch and a second branch. The step of extracting global features from the local feature map based on the weight matrix of the local feature map using the second convolutional module to obtain a global feature map includes: The local feature map is split along the channel direction to obtain a first local feature map and a second local feature map, and the first local feature map and the second local feature map are respectively input into the first branch and the second branch; The first branch and the second branch respectively extract features from the corresponding local feature maps based on the weight matrix of the input local feature maps to obtain a first global feature map and a second global feature map. The first branch generates a weight matrix of the first local feature map based on the first local feature map, and the second branch generates a weight matrix of the second local feature map based on the second local feature map. The first global feature map and the second global feature map are concatenated along the channel direction to obtain a global feature map; The first branch and the second branch respectively extract features from the corresponding local feature maps based on the weight matrix of the input local feature maps, obtaining a first global feature map and a second global feature map, including: A column-dimensional weight matrix is generated based on the first local feature map to obtain a first weight matrix. The first local feature map is then convolved along the column dimension based on the first weight matrix to obtain a first global feature map. Generate a row-dimensional weight matrix based on the input image of the second branch to obtain a second weight matrix, and perform convolution processing on the second local feature map along the row dimension based on the second weight matrix to obtain a second global feature map.
2. The image recognition method as described in claim 1, characterized in that, The second convolutional module further includes a location embedding module, which, before splitting the local feature map along the channel direction, also includes: The location embedding module extracts features from the local feature map based on the weight matrix of the local feature map to obtain a location feature map, and then adds the location feature map to the local feature map to obtain a local feature map containing location features.
3. A method for training a convolutional neural network model, characterized in that, include: Obtain the constructed convolutional neural network model, and input the sample image into the convolutional neural network for training until the convolutional neural network meets the preset requirements to obtain the convolutional neural network model. The convolutional neural network model is the convolutional neural network model used in the image recognition method as described in claim 1 or 2. The convolutional neural network extracts global features from the sample image based on the weight matrix of the sample image.
4. An image recognition device, characterized in that, include: The input module and the trained convolutional neural network model, which performs global feature extraction on the image to be identified based on the weight matrix of the image to be identified; The input module is used to input the image to be recognized into the convolutional neural network model; The convolutional neural network model is used to sequentially extract and recognize features from the image to be recognized, and obtain the recognition result. The convolutional neural network model includes a feature extraction module and a recognition module. The convolutional neural network model sequentially extracts and recognizes features from the image to be recognized, obtaining recognition results, including: The feature extraction module performs feature extraction on the image to be identified. The extracted features are identified based on the identification module to obtain the identification result; The feature extraction module includes a first convolution module and a second convolution module. The second convolution module performs global feature extraction on the image to be identified based on the weight matrix of the image to be identified. The feature extraction of the image to be identified by the feature extraction module includes: Based on the first convolution module, local feature extraction is performed on the image to be identified to obtain a local feature map; Based on the second convolution module, global features are extracted from the local feature map according to the weight matrix of the local feature map to obtain a global feature map; The second convolutional module includes a first branch and a second branch. The step of extracting global features from the local feature map based on the weight matrix of the local feature map using the second convolutional module to obtain a global feature map includes: The local feature map is split along the channel direction to obtain a first local feature map and a second local feature map, and the first local feature map and the second local feature map are respectively input into the first branch and the second branch; The first branch and the second branch respectively extract features from the corresponding local feature maps based on the weight matrix of the input local feature maps to obtain a first global feature map and a second global feature map. The first branch generates a weight matrix of the first local feature map based on the first local feature map, and the second branch generates a weight matrix of the second local feature map based on the second local feature map. The first global feature map and the second global feature map are concatenated along the channel direction to obtain a global feature map; The first branch and the second branch respectively extract features from the corresponding local feature maps based on the weight matrix of the input local feature maps, obtaining a first global feature map and a second global feature map, including: A column-dimensional weight matrix is generated based on the first local feature map to obtain a first weight matrix. The first local feature map is then convolved along the column dimension based on the first weight matrix to obtain a first global feature map. Generate a row-dimensional weight matrix based on the input image of the second branch to obtain a second weight matrix, and perform convolution processing on the second local feature map along the row dimension based on the second weight matrix to obtain a second global feature map.
5. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the image recognition method as described in claim 1 or 2, or the convolutional neural network model training method as described in claim 3.
6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the image recognition method as described in claim 1 or 2, or the convolutional neural network model training method as described in claim 3.