Image recognition method, device and electronic device

By generating heat maps through the attention mechanism model, the problem of high object position labeling cost in neural network model training is solved, and efficient image recognition is achieved.

CN115424122BActive Publication Date: 2025-09-23GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210975639.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-15
Publication Date
2025-09-23
Estimated Expiration
2042-08-15

AI Technical Summary

Technical Problem

Existing neural network models require the locations of objects in training images to be labeled during training, resulting in high training costs and low recognition efficiency.

Method used

The attention mechanism model is used to obtain the classification vector of image feature information, generate a heat map to represent the object position, omit the object position labeling process, and directly label the object through the heat map.

Benefits of technology

It reduces model training costs, improves recognition efficiency, and can perform multi-label reasoning simultaneously, shortening the overall recognition time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115424122B_ABST
    Figure CN115424122B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose an image recognition method, device, and electronic device. The method includes: obtaining feature information of an image to be recognized; obtaining a classification vector corresponding to the feature information based on an attention mechanism model, the attention mechanism model is used to obtain a corresponding classification vector based on the feature information; based on the classification vector, obtaining a heat map corresponding to the object included in the image to be recognized; obtaining the position information of the object corresponding to the heat map in the image to be recognized based on the heat map, the position information is used to mark the object in the image to be recognized. Thus, through the above-mentioned method, the model used to recognize the image no longer needs to be trained using training images marked with the position of the object, so that the process of marking the position of the object corresponding to the training image can be omitted, thereby reducing the training cost of the model. Moreover, in the method provided in the present application, the position of the object can be directly determined by the heat map, thereby improving the recognition efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more specifically, to an image recognition method, device, and electronic device. Background Art

[0002] With the development of technology, neural network models can be used to classify and identify the content in images. For example, neural network models can be used to identify objects such as people and animals in an image and even mark the location of the identified objects. However, neural network models that can be used for classification and recognition must first be trained using training data. However, the training cost and recognition efficiency of these neural network models need to be optimized. Summary of the Invention

[0003] In view of the above problems, the present application proposes an image recognition method, device and electronic device to improve the above problems.

[0004] In a first aspect, the present application provides an image recognition method, which includes: obtaining feature information of an image to be recognized; obtaining a classification vector corresponding to the feature information based on an attention mechanism model, wherein the attention mechanism model is used to obtain a corresponding classification vector based on the feature information; based on the classification vector, obtaining a heat map corresponding to an object included in the image to be recognized; obtaining position information of the object corresponding to the heat map in the image to be recognized based on the heat map, wherein the position information is used to mark the object in the image to be recognized.

[0005] In the second aspect, the present application provides an image recognition device, which includes: a feature acquisition unit for acquiring feature information of an image to be identified; a classification unit for acquiring a classification vector corresponding to the feature information based on an attention mechanism model, and the attention mechanism model is used to obtain a corresponding classification vector based on the feature information; a heat map acquisition unit for obtaining a heat map corresponding to an object included in the image to be identified based on the classification vector; and a position acquisition unit for acquiring position information of the object corresponding to the heat map in the image to be identified based on the heat map, and the position information is used to mark the object in the image to be identified.

[0006] In a third aspect, the present application provides an electronic device, which includes at least a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the above method.

[0007] In a fourth aspect, the present application provides a computer-readable storage medium, in which program code is stored, wherein the above method is executed when the program code is executed by a processor.

[0008] The present application provides an image recognition method, apparatus, and electronic device. After obtaining feature information of an image to be recognized, a classification vector corresponding to the feature information can be obtained based on an attention mechanism. Furthermore, based on the classification vector, a heat map corresponding to the objects included in the image to be recognized is obtained. The location information of the objects corresponding to the heat map in the image to be recognized is then obtained based on the heat map, so that the objects in the image to be recognized can be annotated based on the location information. Thus, through the above-described method, after obtaining the feature information of the image to be recognized, a corresponding classification vector can be directly obtained based on an attention mechanism model, and further, a heat map corresponding to the objects included in the image to be recognized can be obtained based on the classification vector. Furthermore, if the heat map can represent the location of the object in the image to be recognized, the object can be annotated in the image to be recognized using the heat map. This eliminates the need for the image recognition model to be trained using training images annotated with object locations, thereby omitting the process of annotating the object locations corresponding to the training images and reducing the model training cost. Furthermore, in the method provided by the present application, the location of the object can be determined directly using the heat map, thereby improving recognition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0010] Figure 1 A schematic diagram showing an application scenario of the image recognition method proposed in an embodiment of the present application;

[0011] Figure 2 A schematic diagram showing another application scenario of the image recognition method proposed in an embodiment of the present application;

[0012] Figure 3 A flowchart of an image recognition method proposed in one embodiment of the present application is shown;

[0013] Figure 4 A schematic diagram of a heat map in an embodiment of the present application is shown;

[0014] Figure 5 A flowchart of an image recognition method proposed in another embodiment of the present application is shown;

[0015] Figure 6 A flowchart of an image recognition method proposed in another embodiment of the present application is shown;

[0016] Figure 7A schematic diagram showing the position information of an object in an embodiment of the present application;

[0017] Figure 8 A flowchart of obtaining the preset threshold corresponding to each category in an embodiment of the present application is shown;

[0018] Figure 9 A schematic diagram illustrating obtaining heat areas of a reference image corresponding to multiple reference thresholds in an embodiment of the present application is shown;

[0019] Figure 10 A schematic diagram of obtaining a first pixel mean and a second pixel mean in an embodiment of the present application is shown;

[0020] Figure 11 The following is a structural block diagram of an image recognition device proposed in an embodiment of the present application;

[0021] Figure 12 A structural block diagram of another image recognition device proposed in an embodiment of the present application is shown;

[0022] Figure 13 A structural block diagram of another electronic device for executing the image recognition method according to an embodiment of the present application is shown;

[0023] Figure 14 It is a storage unit in an embodiment of the present application for storing or carrying program codes for implementing the image recognition method in accordance with an embodiment of the present application. DETAILED DESCRIPTION

[0024] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0025] With the development of neural network model technology, neural network models can now be used to classify and identify image content in more and more cases. For example, neural network models can be used to identify objects such as people and animals in an image and even mark the location of the identified objects. However, neural network models that can be used for image classification and recognition generally require training to achieve the required recognition function.

[0026] However, the inventors found in their research that the training process of the relevant neural network model still has the problem of training cost that needs to be optimized. For example, in the training process of the relevant neural network model, training images are needed to train the neural network model, and the training images need to be labeled before they can achieve the corresponding training effect. In the labeling process of the training images, not only the classification of the objects included in the training images needs to be labeled, but also the positions of the objects included in the training images need to be labeled (for example, the objects are marked with a frame in the image), which makes the labeling cost of the training images relatively high. Moreover, when there are many classifications of objects included in the training images, it is also inconvenient to label the positions of the objects. In addition, the inventors also found that in order to realize the classification and positioning of objects in the image, the relevant neural network model needs to include more functional modules, resulting in the efficiency of the recognition process to be improved.

[0027] Therefore, after discovering the above problems during research, the inventors proposed an image recognition method, device, and electronic device in this application that can improve the above problems. In the image recognition method provided in the embodiment of the present application, after obtaining the feature information of the image to be recognized, a classification vector corresponding to the feature information can be obtained based on the attention mechanism, and then, based on the classification vector, a heat map corresponding to the object included in the image to be recognized is obtained. Then, based on the heat map, the position information of the object corresponding to the heat map in the image to be recognized is obtained, so that the object can be marked in the image to be recognized based on the position information.

[0028] Thus, after obtaining the feature information of the image to be identified, the corresponding classification vector can be directly obtained based on the attention mechanism model, so as to further obtain the heat map corresponding to the object included in the image to be identified through the classification vector. Moreover, in the case where the heat map can represent the position of the object in the image to be identified, the object can be marked in the image to be identified through the heat map, so that the model used to identify the image no longer needs to be trained using training images marked with the position of the object, so that the process of marking the position of the object corresponding to the training image can be omitted, thereby reducing the training cost of the model. Moreover, in the method provided in the present application, the position of the object can be determined directly through the heat map, thereby improving the recognition efficiency.

[0029] Before further describing the embodiments of the present application in detail, an application environment involved in the embodiments of the present application is introduced.

[0030] The following first introduces the application scenarios involved in the embodiments of this application.

[0031] In the embodiment of the present application, the image recognition method provided can be executed by an electronic device. In this manner, all steps in the image recognition method provided by the embodiment of the present application can be executed by the electronic device. For example, Figure 1 As shown, when all steps in the image recognition method provided in the embodiment of the present application can be executed by an electronic device, all steps can be executed by the processor of the electronic device 100.

[0032] Furthermore, the image recognition method provided in the embodiments of the present application can also be executed by a server. Accordingly, in this server-side execution mode, the server can begin executing the steps of the image recognition method provided in the embodiments of the present application in response to a trigger instruction. The trigger instruction can be sent by the electronic device used by the user, or can be triggered locally by the server in response to some automated event.

[0033] In addition, the image recognition method provided in the embodiment of the present application can also be executed by the electronic device and the server in a collaborative manner. In this manner, some steps of the image recognition method provided in the embodiment of the present application are executed by the electronic device, while other steps are executed by the server. For example, Figure 2 As shown, the electronic device 100 can perform an image recognition method including: sending an image to be recognized to a server, which then executes the server 200 to obtain feature information of the image to be recognized; obtaining a classification vector corresponding to the feature information based on an attention mechanism model; obtaining a heat map corresponding to objects included in the image to be recognized based on the classification vector; and obtaining location information of the objects corresponding to the heat map in the image to be recognized based on the heat map. The server 200 then returns the location information to the electronic device 100 so that the electronic device 100 can annotate the objects in the image to be recognized.

[0034] It should be noted that in this method of collaborative execution by the electronic device and the server, the steps respectively executed by the electronic device and the server are not limited to the methods introduced in the above examples. In actual applications, the steps respectively executed by the electronic device and the server can be dynamically adjusted according to actual conditions.

[0035] It should be noted that the electronic device 100 is Figure 1 and Figure 2In addition to the smartphone shown in the figure, it can also be a tablet computer, a smart watch, an intelligent voice assistant and other devices. Server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud computing, cloud storage, network services, cloud communications, middleware services, CDN (Content Delivery Network), and artificial intelligence platforms. In the case where the image recognition method provided in the embodiment of the present application is executed by a server cluster or distributed system composed of multiple physical servers, different steps in the image recognition method can be executed by different physical servers respectively, or can be executed in a distributed manner by a server built based on a distributed system.

[0036] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0037] See also Figure 3 , an embodiment of the present application provides an image recognition method, the method comprising:

[0038] S110: Acquire feature information of the image to be identified.

[0039] In the embodiment of the present application, the image to be identified can be understood as an image to be subjected to content recognition, wherein the identified content can be the category of an object included in the image to be identified and the position of the object in the image to be identified.

[0040] In the embodiment of the present application, there are multiple ways to obtain the image to be identified.

[0041] As one approach, user input can be performed. In this approach, the user can operate the electronic device to obtain the image to be recognized. Alternatively, the electronic device can capture an image in response to a user-triggered capture operation, and then use the captured image as the image to be recognized. Alternatively, the electronic device can select an image from the electronic device's photo album as the image to be recognized in response to a user-triggered selection operation.

[0042] As another approach, the electronic device can select the image to be evaluated. For example, still taking the image acquisition scenario as an example, in the image acquisition scenario, the image stored in the electronic device can be used as the image to be recognized.

[0043] The feature information of the image to be identified is information used to characterize the features of objects included in the image to be identified. Optionally, the feature information of the image to be identified can be obtained using a pre-trained network model. The obtained feature information can include a feature map and label embedding corresponding to the image to be identified.

[0044] S120: Obtain a classification vector corresponding to the feature information based on the attention mechanism model, where the attention mechanism model is used to obtain the corresponding classification vector based on the feature information.

[0045] In an embodiment of the present application, feature information can be input into a pre-trained attention mechanism model, and a classification vector output by the attention mechanism model can be obtained. It should be noted that in an embodiment of the present application, an initial attention mechanism model can be obtained first, and then the initial attention mechanism model can be trained through training images, thereby obtaining an attention mechanism model that can be used to output a corresponding classification vector based on the feature information. In particular, because in an embodiment of the present application, the position information of the object in the image to be identified is obtained through a subsequent heat map, and the training image used to train the initial attention mechanism model can only be labeled for object classification, and the positions of the objects included in the training image do not need to be labeled.

[0046] During the initial training of the attention mechanism model, if the classification vector output by the attention mechanism model of the current training stage is obtained, the attention mechanism model of the current training stage can be adjusted based on the output classification vector and the loss function until the attention mechanism model whose output classification vector meets the classification requirements is trained. Among them, the loss function used in the training process can be a multivariate binary classification loss function based on sigmoid. For example, the BCE (Binary Cross Entropy) multi-label loss function can be used. The BCE (Binary Cross Entropy) multi-label loss function can be defined as:

[0047]

[0048] Where P is the classification score predicted by the model. In multi-label image classification tasks, since the feature map of an image (e.g., a training image) responds differently to multiple feature points of an object (the object included in the training image) in different channels, and these responses are weighted and combined to determine the multi-label classification performance of the trained model, constraints on the loss function in multi-label tasks can also provide weak supervision on the hot regions where the object is located.

[0049] The classification vectors obtained by the attention mechanism model represent the classification and corresponding location information of the objects included in the image to be identified. Moreover, the attention mechanism model operates based on a multi-head attention mechanism, so that the output classification vectors can correspond to multiple categories of objects. In turn, the classification vectors obtained by the attention mechanism model can directly represent the classification and corresponding location information of multiple objects included in the image to be identified.

[0050] S130: Based on the classification vector, a heat map corresponding to the object included in the image to be identified is obtained.

[0051] The classification vector output by the attention mechanism model can represent the classification and location information of the objects included in the image to be identified, and the classification vector can then be used to obtain a heat map corresponding to the objects included in the image to be identified. As one method, obtaining a heat map corresponding to the objects included in the image to be identified based on the classification vector includes: obtaining a heat map corresponding to a specified aspect ratio based on the classification vector, adjusting the size of the heat map with the specified aspect ratio, and using the adjusted image as the heat map corresponding to the object, wherein the size of the adjusted image is the same as the size of the image to be identified. Optionally, the classification vector can be resized to obtain a heat map corresponding to the specified aspect ratio, and then bilinear interpolation can be performed on the heat map with the specified aspect ratio to enlarge the size of the heat map with the specified aspect ratio to the same size as the image to be identified, thereby obtaining a heat map corresponding to the object. For example, if the output classification vector is bs*80*196, resizing the classification vector results in bs*80*14*14, which can be understood as a heat map with the specified aspect ratio.

[0052] S140: Obtaining position information of an object corresponding to the heat map in the image to be identified according to the heat map, where the position information is used to mark the object in the image to be identified.

[0053] It should be noted that the heat map can represent the content it includes through the pixel values ​​of the pixels it includes, and thus the location of the corresponding object can be determined through the pixel values ​​of the pixels included in the heat map. Figure 4 As shown, in the heat map, the object corresponding to the heat map (for example, Figure 4The pixel values ​​at the position (such as the person shown in the figure) will be different from the pixel values ​​at other positions. Therefore, after obtaining the heat map, the position information of the position of the object can be calculated based on the pixel values ​​of the pixels included in the heat map, and then the object corresponding to the heat map can be marked in the image to be identified based on the position information. Optionally, the center position of the corresponding object can be calculated by the pixel values ​​of the pixels included in the heat map, and then the width and height of the object can be calculated by the pixel values ​​of the pixels in the heat map, and then the position information of the object can be determined based on the center position, width and height of the object.

[0054] It should be noted that the heat map obtained in the embodiments of the present application is actually in color. In the heat map, the location of the corresponding object can be identified by color. For example, the pixels at the location of the object corresponding to the heat map can be red, and the pixels in the area outside the corresponding object in the heat map can be other colors.

[0055] This embodiment provides an image recognition method. After obtaining feature information of an image to be recognized, a classification vector corresponding to the feature information can be obtained based on an attention mechanism. Furthermore, based on the classification vector, a heat map corresponding to the object included in the image to be recognized is obtained. The position information of the object corresponding to the heat map in the image to be recognized is then obtained based on the heat map, so that the object in the image to be recognized can be annotated based on the position information. Thus, through the above-described method, after obtaining the feature information of the image to be recognized, a corresponding classification vector can be directly obtained based on an attention mechanism model, and a heat map corresponding to the object included in the image to be recognized can be further obtained based on the classification vector. Furthermore, if the heat map can represent the position of the object in the image to be recognized, the object can be annotated in the image to be recognized using the heat map, thereby eliminating the need for the model used for image recognition to be trained using training images annotated with the object's position. This allows the object's position to be annotated in the training image to be omitted, thereby reducing the model's training cost. Furthermore, in the method provided in this application, the object's position can be determined directly using the heat map, thereby improving recognition efficiency. Furthermore, the classification vector can also represent the category of the object included in the image to be identified, so that the solution provided in the embodiment of the present application can simultaneously perform multi-label reasoning (i.e., determine the category of the object included in the image) and the location of the object, thereby shortening the overall recognition time.

[0056] See also Figure 5 , an embodiment of the present application provides an image recognition method, the method comprising:

[0057] S210: Acquire feature information of the image to be identified.

[0058] S220: Process the label vector based on the self-attention mechanism module to obtain a first processing result.

[0059] The self-attention mechanism module may perform a self-attention operation on the label vector to obtain a first processing result. In the process of performing the self-attention operation, the query, key, and value may all be the label vector.

[0060] For example, if the network used to obtain the corresponding feature information based on the image to be recognized is swin, the query can be the label embedding (bs*80*2048) of 80 categories, the key is the feature map with location information (bs*196*2048 for swin), and the value is the feature map without location information.

[0061] S230: Process the first processing result and the feature map based on the cross-attention mechanism module to obtain a second processing result.

[0062] The cross-attention mechanism module may perform a cross-attention operation on the first processing result and the feature map to obtain a second processing result. In the process of performing the cross-attention operation, the key and value may be the input feature map.

[0063] It should be noted that in the embodiment of the present application, the attention mechanism model may include a transformer model. Optionally, the self-attention mechanism module and the cross-attention mechanism module may be in the decoder of the transformer model. In this case, the label vector can be input into the self-attention mechanism module of the decoder of the transformer model, and then the first processing result and the feature map can be input into the cross-attention mechanism module of the decoder of the transformer model to obtain the second processing result output by the cross-attention mechanism module.

[0064] S240: Process the second processing result through a fully connected layer to obtain a classification vector.

[0065] S250: Based on the classification vector, a heat map corresponding to the object included in the image to be identified is obtained.

[0066] S260: Obtaining position information of an object corresponding to the heat map in the image to be identified according to the heat map, where the position information is used to mark the object in the image to be identified.

[0067] The present embodiment provides an image recognition method, which, through the above-mentioned method, enables to obtain the feature information of the image to be recognized directly based on the attention mechanism model to obtain the corresponding classification vector, so as to further obtain the heat map corresponding to the object included in the image to be recognized through the classification vector. Moreover, in the case where the heat map can represent the position of the object in the image to be recognized, the object can be annotated in the image to be recognized through the heat map, so that the model used for image recognition no longer needs to be trained using training images with the object position annotated, so that the process of annotating the object position corresponding to the training image can be omitted, thereby reducing the training cost of the model. Moreover, in the method provided in the present application, the position of the object can be determined directly through the heat map, thereby improving the recognition efficiency. Moreover, in this embodiment, the attention mechanism model adopted includes a self-attention mechanism module and a cross-attention mechanism module, so that the self-attention mechanism module and the cross-attention mechanism module can be run successively, so that the classification vector finally obtained can more accurately represent the category and position of the object in the object to be recognized.

[0068] See also Figure 6 , an embodiment of the present application provides an image recognition method, the method comprising:

[0069] S310: Acquire feature information of the image to be identified.

[0070] S320: Obtain a classification vector corresponding to the feature information based on the attention mechanism model, where the attention mechanism model is used to obtain the corresponding classification vector based on the feature information.

[0071] S330: Based on the classification vector, obtain a heat map corresponding to the object included in the image to be identified.

[0072] S340: Obtain target pixels included in the heat map, where the target pixels include pixels whose corresponding pixel values ​​are greater than a preset threshold.

[0073] It should be noted that in some cases, not all pixels in the heat map can be used to calculate the position information of the object corresponding to the heat map. In this case, the pixels included in the heat map can be filtered by a preset threshold. Among them, obtaining the target pixel based on the heat map can also be understood as obtaining the mask map corresponding to the heat map. In this case, the pixels in the mask map can be understood as the target pixels corresponding to the heat map.

[0074] After obtaining the classification vector output by the attention mechanism model, it can be input into a fully connected layer for multiple classification, thereby obtaining the classification of the objects included in the image to be recognized. In addition, if the objects included in the image to be recognized have multiple classifications, a heat map will be generated for each classification.

[0075] In this way, the target pixels included in the heat map are obtained, and the target pixels include pixels whose corresponding pixel values ​​are greater than a preset threshold, including: obtaining the classification of the object corresponding to the heat map, obtaining the preset threshold corresponding to the classification, and different classifications have different preset thresholds. From the pixels included in the heat map, pixels whose corresponding pixel values ​​are greater than the preset threshold corresponding to the classification are obtained as target pixels.

[0076] S350: Obtaining position information of an object corresponding to the heat map based on the position and pixel value of the target pixel. The position information is used to mark the object in the image to be identified.

[0077] In an embodiment of the present application, the position information corresponding to the object is used to characterize the area where the object is located in the image. Furthermore, the position information of the object can be used to determine which areas in the image are areas where the object corresponding to the heat map is located. As a way, the position of the target pixel includes the horizontal coordinate and the vertical coordinate of the coordinate area where the target pixel is located, and the position information of the object includes the horizontal coordinate of the center position of the object, the vertical coordinate of the center position, the width and the height. Correspondingly, in this way, the position information of the object corresponding to the heat map is obtained based on the position and pixel value of the target pixel, including: obtaining the horizontal coordinate of the center position of the object corresponding to the heat map based on the pixel value and the horizontal coordinate of the target pixel, obtaining the vertical coordinate of the center position of the object corresponding to the heat map based on the pixel value and the vertical coordinate of the target pixel, obtaining the width of the object corresponding to the heat map based on the pixel value of the target pixel, the horizontal coordinate of the target pixel and the horizontal coordinate of the center position, and obtaining the height of the object corresponding to the heat map based on the pixel value of the target pixel, the vertical coordinate of the target pixel and the vertical coordinate of the center position.

[0078] After obtaining the horizontal coordinate, vertical coordinate, width and height of the center position of the object, the area occupied by the object in the image to be identified can be determined, and then the object can be marked in the image to be identified. For example, if the image to be identified is Figure 7 As shown, the identified objects include Figure 7 , then after determining the horizontal coordinate, vertical coordinate, width (W) and height (H) of the center position of the puppy based on the heat map corresponding to the puppy, the location of the puppy can be marked in the image to be identified. Figure 7 Mark the corresponding box around the puppy.

[0079] As a method, before obtaining the location information of the corresponding object through the heat map, the heat map can be normalized. The formula used for normalization is as follows:

[0080]

[0081] CAM(x,y) on the left side of the equal sign represents the normalized pixel value, while CAM(x,y) on the right side represents the pixel value before normalization. minCAM(x,y) represents the minimum pixel value in the heat map before normalization, and maxCAM(x,y) represents the maximum pixel value in the heat map before normalization. After obtaining the normalized heat map, the mask image can be filtered using a preset threshold.

[0082] It should be noted that the preset thresholds corresponding to normalized and unnormalized heat maps may be different, so that for the same heat map, different preset thresholds may correspond to the heat map that has not been normalized and the heat map that has been normalized. For example, in the process of determining the target pixel through the heat map that has not been normalized, the preset threshold used is also the unnormalized preset threshold. In the process of determining the target pixel through the heat map that has been normalized, the preset threshold used is also the normalized preset threshold.

[0083] After obtaining the mask image, the horizontal coordinate and vertical coordinate of the center position of the object corresponding to the heat map can be calculated according to the following formula. The formula is as follows:

[0084]

[0085]

[0086] Among them, M(x,y) represents the pixel value of the pixel (i.e., the target pixel) in the mask image corresponding to the heat map, x represents the horizontal coordinate corresponding to the pixel value, and y represents the vertical coordinate corresponding to the pixel value. c The horizontal coordinate of the calculated center position, y c The vertical coordinate representing the calculated center position.

[0087] As a way, such as Figure 8 As shown, in the case where each classified object has a separate preset threshold, the corresponding preset threshold can be obtained in advance for each classified object. Then, before obtaining the target pixel included in the heat map, the target pixel includes the corresponding pixel value greater than the preset threshold and further includes:

[0088] S341: Acquire multiple reference images corresponding to multiple categories.

[0089] In the case where there are multiple categories of objects, the reference image corresponding to each category can be understood as an image including objects of the corresponding category. The multiple reference images corresponding to a category can be understood as all of the multiple reference images including objects of the corresponding category, and the styles of the objects included in the multiple reference images corresponding to the same category and the positions of the objects in the reference images can be different. Exemplarily, in the case where multiple categories include categories such as people, animals, houses, and vehicles, multiple reference images corresponding to people can be obtained, multiple reference images corresponding to animals can be obtained, multiple reference images corresponding to houses can be obtained, and multiple reference images corresponding to vehicles can be obtained. In the multiple reference images corresponding to people can only include people, the multiple reference images corresponding to animals can only include animals, the multiple reference images corresponding to houses can only include houses, and the multiple reference images corresponding to vehicles can only include vehicles.

[0090] S342: Obtain heat areas corresponding to multiple reference images corresponding to the current classification and corresponding to multiple reference thresholds, the heat areas represent position information of objects in the corresponding reference images, and the current classification is the classification currently performing preset threshold calculation among the multiple classifications.

[0091] After obtaining multiple reference images corresponding to the current classification, the heat areas of the multiple reference images corresponding to multiple reference thresholds can be obtained based on the method provided in the embodiment of the present application. Among them, the heat area of ​​the reference image corresponding to the reference threshold can be understood as after obtaining the heat map corresponding to the reference image, the corresponding target pixel is obtained from the corresponding heat map based on the reference threshold corresponding to the reference image, and then the position information of the object included in the reference image is calculated based on the target pixel. The area identified by the position information is the heat area of ​​the reference image corresponding to the reference threshold. Then, in the case where a single reference image corresponds to multiple reference thresholds, a heat area can be calculated for each reference threshold.

[0092] For example, if the multiple reference images corresponding to the current classification include reference image A, reference image B, and reference image C. When the multiple reference thresholds include reference threshold T1, reference threshold T2, and reference threshold T3. For reference image A, the heat areas for reference threshold T1, reference threshold T2, and reference threshold T3 will be calculated respectively. For reference image B, the heat areas for reference threshold T1, reference threshold T2, and reference threshold T3 will be calculated respectively. For reference image C, the heat areas for reference threshold T1, reference threshold T2, and reference threshold T3 will be calculated respectively. Figure 9As shown, taking reference image A as an example, the heat area calculated for reference threshold T1 is heat area Q1, the heat area calculated for reference threshold T2 is heat area Q2, and the heat area calculated for reference threshold T3 is heat area Q3.

[0093] It should be noted that the values ​​of the multiple reference thresholds may be incremented in sequence, and the differences between adjacent reference thresholds may be the same. For example, when the reference thresholds are also normalized, the smallest reference threshold among the multiple reference thresholds may be 0.4, the largest reference threshold may be 0.75, and the other reference thresholds may be evenly distributed between 0.4 and 0.75 with an interval of 0.05.

[0094] S343: Obtain the first pixel mean of the heat area of ​​each of the multiple reference images corresponding to multiple reference thresholds.

[0095] After obtaining multiple reference image heat regions corresponding to multiple reference thresholds, the first pixel mean can be calculated for each heat region. For a single heat region, the pixel values ​​of all pixels in the heat region can be summed and then divided by the number of pixels to obtain the first pixel mean corresponding to the heat region.

[0096] S344: Filter out images that meet specified conditions from multiple reference images to obtain remaining reference images, wherein the specified conditions include reference images whose first pixel mean value of a heat area corresponding to a first reference threshold is greater than the first reference mean value, or reference images whose first pixel mean value of a heat area corresponding to a second reference threshold is greater than the second reference mean value, the first reference threshold value is the smallest reference threshold value among the multiple reference threshold values, and the second reference threshold value is the largest reference threshold value among the multiple reference threshold values.

[0097] For example, if the smallest reference threshold value among the multiple reference threshold values ​​is 0.4, the largest reference threshold value is 0.75, and the other reference threshold values ​​are evenly distributed between 0.4 and 0.75 with intervals of 0.05, the first reference threshold value may be 0.4, and the second reference threshold value may be 0.75. The first reference mean and the second reference mean may be the same, for example, both the first reference mean and the second reference mean may be 0.6.

[0098] S345: Obtaining a preset threshold corresponding to the current classification according to the remaining reference images.

[0099] Optionally, a preset threshold corresponding to the current classification is obtained based on the remaining reference images, including: obtaining the heat areas of the remaining reference images corresponding to multiple reference thresholds, obtaining the second pixel means of the heat areas corresponding to the multiple reference thresholds, obtaining multiple second pixel means, and using the reference threshold corresponding to the second pixel mean that is the same as the third reference mean among the multiple second pixel means as the preset threshold corresponding to the current classification.

[0100] After obtaining the remaining reference images, the second pixel mean can be calculated for each reference threshold. Unlike the first pixel mean, which is the mean calculated for the pixels in each heat area, the second pixel mean is the mean of the pixels in all heat areas calculated for the same reference threshold across multiple images. For example, Figure 10 As shown, when the remaining reference images obtained include reference image A and reference image B, after calculating the first pixel mean corresponding to each heat area, the first pixel mean corresponding to the same reference threshold can be further averaged to obtain the second pixel mean corresponding to the reference threshold. For example, the reference threshold corresponding to the first pixel mean Z1 and the first pixel mean Z4 is T1, then the first pixel mean Z1 and the first pixel mean Z4 are averaged to obtain the second pixel mean Z7 of the heat area corresponding to the reference threshold T1. Similarly, the reference threshold corresponding to the first pixel mean Z2 and the first pixel mean Z5 is T2, then the first pixel mean Z2 and the first pixel mean Z5 are averaged to obtain the second pixel mean Z8 of the heat area corresponding to the reference threshold T2. The reference threshold corresponding to the first pixel mean Z3 and the first pixel mean Z6 is T3, then the first pixel mean Z3 and the first pixel mean Z6 are averaged to obtain the second pixel mean Z9 of the heat area corresponding to the reference threshold T3.

[0101] After obtaining multiple second pixel means, each second pixel mean can be compared with the third reference mean to obtain the second pixel mean that is the same as the third reference mean, and the reference threshold corresponding to the second pixel mean can be used as the preset threshold corresponding to the current classification. Figure 10 Taking the illustrated case as an example, if the third reference mean is Z8, then the reference threshold T2 corresponding to the second pixel mean Z8 can be used as the preset threshold for the current classification. Optionally, the third reference threshold can be 0.6.

[0102] In the process of obtaining the preset threshold corresponding to the current classification based on the remaining reference images, in addition to directly obtaining the preset threshold corresponding to the current classification based on the remaining reference images, a portion of the reference images can also be selected from the remaining reference images to obtain the preset threshold corresponding to the current classification. In this case, based on the reference images of the portion, the heat areas of the reference images of the portion corresponding to multiple reference thresholds are obtained, and then the second pixel mean is calculated according to the aforementioned method. The number of partial reference images selected from the remaining reference images can be determined according to the number of remaining reference images. For example, when the number of remaining reference images is greater than 50, the number of partial reference images selected from the reference images of the sound source is 50. For another example, the selected partial reference images can be half of the remaining reference images.

[0103] It should be noted that in some cases, the calculated position information for some small objects may not be accurate enough. In this case, the position information of small objects can be adjusted. Among them, small objects can be understood as objects that occupy less than a specified area in the image, which can be 10% or 20%.

[0104] As a method, the position information of the object corresponding to the heat map is obtained based on the position and pixel value of the target pixel, including: obtaining the initial position information of the object corresponding to the heat map based on the position and pixel value of the target pixel, obtaining the pixel mean of the heat area determined based on the position information, and if the pixel mean is less than a fourth reference mean, using the initial position information as the position information of the object corresponding to the heat map. If the pixel mean is not less than the fourth reference mean, reducing a preset threshold based on a preset ratio to obtain a reduced threshold, obtaining an updated target pixel based on the reduced threshold, and obtaining the position information of the object corresponding to the heat map based on the position and pixel value of the updated target pixel.

[0105] It should be noted that after obtaining the initial position information, the area size corresponding to the initial position information can be calculated. If the calculated area size corresponding to the initial position information accounts for less than the specified proportion in the image to be identified, the pixel mean of the heat area determined based on the initial position information will be obtained. Otherwise, the initial position information will be directly used as the position information of the object corresponding to the heat map.

[0106] The present embodiment provides an image recognition method, which, through the above-mentioned method, enables to obtain the corresponding classification vector directly based on the attention mechanism model after obtaining the feature information of the image to be recognized, so as to further obtain the heat map corresponding to the object included in the image to be recognized through the classification vector. Moreover, in the case where the heat map can represent the position of the object in the image to be recognized, the object can be marked in the image to be recognized through the heat map, so that the model used to recognize the image no longer needs to be trained using the training image with the object position marked, so that the process of marking the object position corresponding to the training image can be omitted, thereby reducing the training cost of the model. Moreover, in the present embodiment, a corresponding preset threshold for determining the target pixel can be separately configured for each classified object, so that the object marking process can be better targeted to improve the accuracy of object position marking.

[0107] See also Figure 11 , an embodiment of the present application provides an image recognition device 400, the device 400 including:

[0108] The feature acquisition unit 410 is used to acquire feature information of the image to be identified.

[0109] The classification unit 420 is used to obtain a classification vector corresponding to the feature information based on the attention mechanism model, and the attention mechanism model is used to obtain the corresponding classification vector according to the feature information.

[0110] Among them, the classification unit 420 is specifically used to input feature information into the attention mechanism model to obtain the classification vector output by the attention mechanism model, wherein the loss function used by the attention mechanism model during the training process is a multi-label loss function.

[0111] The heat map acquisition unit 430 is used to obtain a heat map corresponding to the object included in the image to be identified based on the classification vector.

[0112] The position acquisition unit 440 is used to obtain the position information of the object corresponding to the heat map in the image to be identified based on the heat map, and the position information is used to mark the object in the image to be identified.

[0113] As one approach, the attention mechanism model includes a self-attention mechanism module and a cross-attention mechanism module, and the feature information includes a feature map and a label vector corresponding to the image to be identified. Accordingly, the heat map acquisition unit 430 is specifically configured to process the label vector based on the self-attention mechanism module to obtain a first processing result; process the first processing result and the feature map based on the cross-attention mechanism module to obtain a second processing result; and process the second processing result through a fully connected layer to obtain a classification vector.

[0114] As a method, the heat map acquisition unit 430 is specifically used to obtain a heat map of a specified aspect ratio according to the classification vector; adjust the size of the heat map of the specified aspect ratio, and use the adjusted image as the heat map corresponding to the object, wherein the size of the adjusted image is the same as the size of the image to be identified.

[0115] As a method, the position acquisition unit 440 is specifically used to obtain the target pixels included in the heat map, and the target pixels include pixels whose corresponding pixel values ​​are greater than a preset threshold; based on the position and pixel value of the target pixels, the position information of the object corresponding to the heat map is obtained.

[0116] Optionally, the position of the target pixel includes the horizontal coordinate and the vertical coordinate of the coordinate area where the target pixel is located, and the position information of the object includes the horizontal coordinate of the center position of the object, the vertical coordinate of the center position, the width, and the height. Correspondingly, the position acquisition unit 440 is specifically used to obtain the horizontal coordinate of the center position of the object corresponding to the heat map based on the pixel value and the horizontal coordinate of the target pixel; obtain the vertical coordinate of the center position of the object corresponding to the heat map based on the pixel value and the vertical coordinate of the target pixel; obtain the width of the object corresponding to the heat map based on the pixel value of the target pixel, the horizontal coordinate of the target pixel, and the horizontal coordinate of the center position; obtain the height of the object corresponding to the heat map based on the pixel value of the target pixel, the vertical coordinate of the target pixel, and the vertical coordinate of the center position.

[0117] Optionally, the position acquisition unit 440 is specifically used to obtain the classification of the object corresponding to the heat map; obtain the preset threshold corresponding to the classification, and different classifications have different preset thresholds; from the pixels included in the heat map, obtain the pixel whose corresponding pixel value is greater than the preset threshold corresponding to the classification as the target pixel.

[0118] As a way, such as Figure 12 As shown, the apparatus 400 further includes:

[0119] The preset threshold determination unit 450 is used to obtain multiple reference images corresponding to multiple categories; obtain heat areas corresponding to multiple reference thresholds of the multiple reference images corresponding to the current category, the heat areas represent the position information of the objects in the corresponding reference images, and the current category is the category currently performing the preset threshold calculation among the multiple categories; obtain the first pixel mean of the heat areas corresponding to the multiple reference images; filter out the images that meet the specified conditions in the multiple reference images to obtain the remaining reference images, wherein the specified conditions include the reference images whose first pixel mean of the heat area corresponding to the first reference threshold is greater than the first reference mean, or the reference images whose first pixel mean of the heat area corresponding to the second reference threshold is greater than the second reference mean, the first reference threshold is the smallest reference threshold among the multiple reference thresholds, and the second reference threshold is the largest reference threshold among the multiple reference thresholds; obtain the preset threshold corresponding to the current category based on the remaining reference images.

[0120] Optionally, the preset threshold determination unit 450 is specifically used to obtain the heat areas of the remaining reference images corresponding to multiple reference thresholds; obtain the second pixel means of the heat areas corresponding to the multiple reference thresholds to obtain multiple second pixel means; and use the reference threshold corresponding to the second pixel mean that is the same as the third reference mean in the multiple second pixel means as the preset threshold corresponding to the current classification.

[0121] Optionally, the preset threshold determination unit 450 is specifically configured to obtain initial position information of an object corresponding to the heat map based on the position and pixel value of the target pixel, obtain a pixel mean of the heat region determined based on the position information; if the pixel mean is less than a fourth reference mean, use the initial position information as the position information of the object corresponding to the heat map. If the pixel mean is not less than the fourth reference mean, reduce the preset threshold based on a preset ratio to obtain a reduced threshold; obtain an updated target pixel based on the reduced threshold; and obtain the position information of the object corresponding to the heat map based on the updated position and pixel value of the target pixel.

[0122] This embodiment provides an image recognition device that eliminates the need for training a model for image recognition using training images labeled with object locations. This eliminates the need to label the object locations corresponding to the training images, reducing model training costs and improving recognition efficiency.

[0123] It should be noted that the device embodiment in this application corresponds to the aforementioned method embodiment. The specific principles in the device embodiment can be found in the contents of the aforementioned method embodiment and will not be repeated here.

[0124] The following will be combined Figure 13 An electronic device provided by this application is described.

[0125] See also Figure 13 Based on the above-mentioned image recognition method and apparatus, the embodiments of the present application also provide another electronic device 100 that can execute the above-mentioned image recognition method. The electronic device 100 includes one or more (only one is shown in the figure) processors 102, a memory 104, and a network module 106 that are coupled to each other. The memory 104 stores a program that can execute the content of the above-mentioned embodiments, and the processor 102 can execute the program stored in the memory 104.

[0126] The processor 102 may include one or more processing cores. The processor 102 utilizes various interfaces and circuits to connect various components within the electronic device 100. It executes instructions, programs, code sets, or instruction sets stored in the memory 104, and accesses data stored in the memory 104 to perform various functions and process data within the electronic device 100. Optionally, the processor 102 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 102 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 102 and may be implemented separately via a communication chip.

[0127] The memory 104 may include a random access memory (RAM) or a read-only memory (ROM). The memory 104 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 104 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created by the terminal 100 during use (such as a phone book, audio and video data, chat history data), etc.

[0128] The network module 106 is used to receive and transmit electromagnetic waves, realize the mutual conversion between electromagnetic waves and electrical signals, and thus communicate with a communication network or other devices, such as communicating with an audio playback device. The network module 106 may include various existing circuit components for performing these functions, such as an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, a memory, etc. The network module 106 can communicate with various networks such as the Internet, an enterprise intranet, a wireless network, or communicate with other devices via a wireless network. The above-mentioned wireless network may include a cellular telephone network, a wireless local area network, or a metropolitan area network. For example, the network module 106 can exchange information with a base station.

[0129] Please refer to Figure 14 , which shows a block diagram of a computer-readable storage medium provided in an embodiment of the present application. The computer-readable medium 800 stores program code, which can be called by a processor to execute the method described in the above method embodiment.

[0130] The computer-readable storage medium 800 can be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer-readable storage medium 800 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 800 has storage space for program code 810 for executing any of the method steps described above. These program codes can be read from or written to one or more computer program products. The program code 810 can be compressed, for example, in a suitable form.

[0131] In summary, the present application provides an image recognition method, device, and electronic device. After obtaining the feature information of the image to be recognized, the classification vector corresponding to the feature information can be obtained based on the attention mechanism, and then, based on the classification vector, a heat map corresponding to the object included in the image to be recognized is obtained. Then, the position information of the object corresponding to the heat map in the image to be recognized is obtained based on the heat map, so that the object is labeled in the image to be recognized based on the position information. Therefore, through the above method, after obtaining the feature information of the image to be recognized, the corresponding classification vector can be directly obtained based on the attention mechanism model, so as to further obtain the heat map corresponding to the object included in the image to be recognized through the classification vector. Moreover, in the case where the heat map can represent the position of the object in the image to be recognized, the object can be labeled in the image to be recognized through the heat map, so that the model used to recognize the image no longer needs to be trained using training images with the object positions labeled, so that the process of labeling the object positions corresponding to the training images can be omitted, thereby reducing the training cost of the model. Moreover, in the method provided in this application, the position of the object can be directly determined through the heat map. Compared with target detection technologies based on RNN (Recurrent Neural Network), which need to first estimate the area where the object may exist and then perform classification and frame regression, this method does not need to detect the possible area, thereby improving recognition efficiency.

[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An image recognition method, characterized in that: The method comprises: Acquire feature information of the image to be identified, the feature information including a feature map and a label vector corresponding to the image to be identified; Processing the label vector by a self-attention mechanism module based on an attention mechanism model to obtain a first processing result, wherein the attention mechanism model is used to obtain a corresponding classification vector according to feature information; Processing the first processing result and the feature map based on a cross attention mechanism module of the attention mechanism model to obtain a second processing result; Processing the second processing result through a fully connected layer to obtain a classification vector; Based on the classification vector, obtaining a heat map corresponding to the object included in the image to be identified; The position information of the object corresponding to the heat map in the image to be identified is obtained according to the heat map, and the position information is used to mark the object in the image to be identified.

2. The method according to claim 1, characterized in that The step of obtaining a heat map corresponding to an object included in the image to be identified based on the classification vector includes: Obtaining a heat map corresponding to a specified aspect ratio according to the classification vector; The size of the heat map with the specified aspect ratio is adjusted, and the adjusted image is used as the heat map corresponding to the object, wherein the size of the adjusted image is the same as the size of the image to be identified.

3. The method according to claim 1, characterized in that The obtaining, according to the heat map, position information of an object corresponding to the heat map in the image to be identified includes: Obtaining target pixels included in the heat map, where the target pixels include pixels whose corresponding pixel values ​​are greater than a preset threshold; The position information of the object corresponding to the heat map is obtained based on the position and pixel value of the target pixel.

4. The method according to claim 3, characterized in that The position of the target pixel includes the abscissa and ordinate of the coordinate area where the target pixel is located, the position information of the object includes the abscissa, ordinate, width, and height of the center position of the object, and the position information of the object corresponding to the heat map obtained based on the position and pixel value of the target pixel includes: Obtaining the abscissa of the center position of the object corresponding to the heat map based on the pixel value of the target pixel and the abscissa; Obtaining the vertical coordinate of the center position of the object corresponding to the heat map based on the pixel value of the target pixel and the vertical coordinate; Obtaining a width of the object corresponding to the heat map based on the pixel value of the target pixel, the abscissa of the target pixel, and the abscissa of the center position; Based on the pixel value of the target pixel, the ordinate of the target pixel, and the ordinate of the center position, the height of the object corresponding to the heat map is obtained.

5. The method according to claim 3, characterized in that The step of obtaining target pixels included in the heat map, wherein the target pixels include pixels whose corresponding pixel values ​​are greater than a preset threshold, includes: Obtaining a classification of an object corresponding to the heat map; Obtaining a preset threshold corresponding to the classification, where different classifications correspond to different preset thresholds; From the pixels included in the heat map, pixels whose corresponding pixel values ​​are greater than a preset threshold corresponding to the classification are obtained as target pixels.

6. The method according to claim 5, characterized in that The step of obtaining a target pixel included in the heat map, wherein the target pixel includes a corresponding pixel value greater than a preset threshold, further includes: Acquire multiple reference images corresponding to multiple categories; Obtaining heat regions corresponding to multiple reference thresholds in each of the multiple reference images corresponding to a current classification, wherein the heat regions represent position information of objects in the corresponding reference images, and the current classification is the classification currently undergoing preset threshold calculation among the multiple classifications; Obtaining first pixel averages of heat areas corresponding to multiple reference thresholds in each of the multiple reference images; Screening out images that meet specified conditions from the multiple reference images to obtain remaining reference images, wherein the specified conditions include reference images whose first pixel mean value of a heat area corresponding to a first reference threshold is greater than the first reference mean value, or reference images whose first pixel mean value of a heat area corresponding to a second reference threshold is greater than the second reference mean value, the first reference threshold being the smallest reference threshold value among the multiple reference threshold values, and the second reference threshold being the largest reference threshold value among the multiple reference threshold values; A preset threshold corresponding to the current classification is obtained according to the remaining reference images.

7. The method according to claim 6, characterized in that The obtaining, according to the remaining reference images, a preset threshold corresponding to the current classification includes: Acquire heat regions of the remaining reference images corresponding to the multiple reference thresholds; Obtaining second pixel means of the heat area corresponding to multiple reference thresholds to obtain multiple second pixel means; The reference threshold corresponding to the second pixel mean value that is the same as the third reference mean value among the multiple second pixel mean values ​​is used as the preset threshold value corresponding to the current classification.

8. The method according to claim 3, characterized in that The obtaining of the position information of the object corresponding to the heat map based on the position and pixel value of the target pixel includes: Based on the position and pixel value of the target pixel, the initial position information of the object corresponding to the heat map is obtained. Obtaining a pixel mean value of a heat area determined based on the initial position information; If the pixel mean is less than the fourth reference mean, the initial position information is used as the position information of the object corresponding to the heat map.

9. The method according to claim 8, characterized in that The method further comprises: If the pixel mean is not less than a fourth reference mean, reducing the preset threshold based on a preset ratio to obtain a reduced threshold; Obtaining an updated target pixel based on the lowered threshold; The position information of the object corresponding to the heat map is obtained based on the position and pixel value of the updated target pixel.

10. The method according to claim 1, characterized in that The loss function used by the attention mechanism model during training is a multi-label loss function.

11. An image recognition device, characterized in that: The device comprises: A feature acquisition unit, configured to acquire feature information of an image to be identified, wherein the feature information includes a feature map and a label vector corresponding to the image to be identified; a heat map acquisition unit, configured to process the label vector based on a self-attention mechanism module of an attention mechanism model to obtain a first processing result, wherein the attention mechanism model is configured to obtain a corresponding classification vector based on feature information; process the first processing result and the feature map based on a cross-attention mechanism module of the attention mechanism model to obtain a second processing result; and process the second processing result through a fully connected layer to obtain a classification vector; The heat map acquisition unit is further configured to obtain a heat map corresponding to the object included in the image to be identified based on the classification vector; A position acquisition unit is used to obtain position information of an object corresponding to the heat map in the image to be identified based on the heat map, and the position information is used to mark the object in the image to be identified.

12. An electronic device, characterized in that: The method comprises a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program code, wherein when the program code is executed by a processor, the method according to any one of claims 1 to 10 is executed.

Citation Information

Patent Citations

  • Defect coarse positioning method and device based on weak supervision mode

    CN112070733A