Image recognition method and device, equipment, medium and product

By introducing the loss of the category activation map segmentation task as a training constraint in the image recognition model, the problem of low accuracy of the fine-grained image classification method is solved, and more accurate prediction of the foreground area and image classification is achieved.

CN120070931APending Publication Date: 2025-05-30HANGZHOU NETEASE ZHIQI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411907063.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The accuracy of the fine-grained image classification method in the prior art is poor, mainly due to the existence of only global label signals, resulting in inaccurate prediction and incomplete positioning of target objects.

Method used

By introducing the loss of the category activation map segmentation task as a training constraint, the initial image recognition model is optimized so that it can predict the foreground area images more accurately, thereby improving the accuracy of fine-grained image classification.

Benefits of technology

By generating a complete and accurate category activation map, this method alleviates the problems of inaccurate prediction and incomplete positioning in fine-grained image classification, improves the accuracy of image classification, and reduces the dependence on additional annotation information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070931A_ABST
    Figure CN120070931A_ABST
Patent Text Reader

Abstract

The invention discloses an image recognition method and device, equipment, a medium and a product, and the method comprises the steps: inputting a to-be-recognized image into a target image recognition model, and obtaining a global image classification result of the to-be-recognized image; determining a category activation mapping graph of the to-be-recognized image based on the target image recognition model; determining a foreground region image of the to-be-recognized image based on the category activation mapping graph of the to-be-recognized image; inputting the foreground region image of the to-be-recognized image into the target image recognition model to obtain a foreground region classification result of the to-be-recognized image; determining a target classification result of the to-be-recognized image based on the global image classification result and the foreground region classification result of the to-be-recognized image; wherein the target image recognition model is obtained by training the initial image recognition model at least with the loss of the category activation map segmentation task as a constraint. According to the technical scheme provided by the embodiment of the invention, the accuracy of the fine-grained image classification method can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and in particular relates to an image recognition method, device, equipment, medium and product. Background Art

[0002] In recent years, with the continuous expansion of the Internet industry and traffic, how to accurately and quickly identify and filter harmful information from a large amount of online data has also become a new task and challenge. Among them, the most direct challenge is how to obtain a richer and more fine-grained label recognition ability to cope with the detection and screening of increasingly diverse harmful images and their variants online.

[0003] In related technologies, for the fine-grained image classification task, the idea of weakly supervised learning is generally used to complete the localization and secondary recognition of key regions of the target category in the image. However, when the weakly supervised learning method generates a class activation mapping from the feature map of the model, it can often only recognize local features with high class discrimination, such as the root texture of the animal's head or the plant's leaf, which ultimately leads to inaccurate localization of the complete object.

[0004] Therefore, the accuracy of the fine-grained image classification method in related technologies is poor. Summary of the Invention

[0005] Embodiments of this application provide an implementation different from related technologies to solve the technical problem of poor accuracy of the fine-grained image classification method in related technologies.

[0006] In a first aspect, this application provides an image recognition method, including:

[0007] Input the image to be recognized into a target image recognition model to obtain the global image classification result of the image to be recognized;

[0008] Determine the class activation mapping of the image to be recognized based on the target image recognition model;

[0009] Determine the foreground region image of the image to be recognized based on the class activation mapping of the image to be recognized;

[0010] Input the foreground region image of the image to be recognized into the target image recognition model to obtain the foreground region classification result of the image to be recognized;

[0011] Determine the target classification result of the image to be recognized based on the global image classification result and the foreground region classification result of the image to be recognized;

[0012] Among them, the target image recognition model is obtained by training an initial image recognition model with at least the loss of a class activation map segmentation task as a constraint. The class activation map segmentation task is a task of predicting a class activation map for a sample image using the image segmentation label of the sample image as a segmentation supervision signal, and the image segmentation label includes the class labels corresponding to each pixel in the sample image.

[0013] In a second aspect, the present application provides an image recognition device, including:

[0014] A first input unit configured to input an image to be recognized into a target image recognition model to obtain a global image classification result of the image to be recognized;

[0015] A first determination unit configured to determine a class activation map of the image to be recognized based on the target image recognition model;

[0016] A second determination unit configured to determine a foreground region image of the image to be recognized based on the class activation map of the image to be recognized;

[0017] A second input unit configured to input the foreground region image of the image to be recognized into the target image recognition model to obtain a foreground region classification result of the image to be recognized;

[0018] A third determination unit configured to determine a target classification result of the image to be recognized based on the global image classification result and the foreground region classification result of the image to be recognized;

[0019] Among them, the target image recognition model is obtained by training an initial image recognition model with at least the loss of a class activation map segmentation task as a constraint. The class activation map segmentation task is a task of predicting a class activation map for a sample image using the image segmentation label of the sample image as a segmentation supervision signal, and the image segmentation label includes the class labels corresponding to each pixel in the sample image.

[0020] In a third aspect, the present application provides an electronic device, including:

[0021] A processor; and

[0022] A memory configured to store executable instructions of the processor;

[0023] Among them, the processor is configured to execute the method in the first aspect, or any one of the possible implementation manners of the first aspect by executing the executable instructions.

[0024] Fourthly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the first aspect or any method in each possible implementation manner of the first aspect.

[0025] Fifthly, an embodiment of the present application provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the first aspect or any method in each possible implementation manner of the first aspect.

[0026] In the present application, the image to be recognized is input into the target image recognition model to obtain the global image classification result of the image to be recognized; the class activation map of the image to be recognized is determined based on the target image recognition model; the foreground region image of the image to be recognized is determined based on the class activation map of the image to be recognized; the foreground region image of the image to be recognized is input into the target image recognition model to obtain the foreground region classification result of the image to be recognized; the target classification result of the image to be recognized is determined based on the global image classification result and the foreground region classification result of the image to be recognized; wherein, the target image recognition model is trained with at least the loss of the class activation map segmentation task as a constraint. The class activation map segmentation task is a task of predicting the class activation map of the sample image with the image segmentation label of the sample image as the segmentation supervision signal. The image segmentation label includes the class labels corresponding to each pixel in the sample image. In the training of the initial image recognition model, the loss of the class activation map segmentation task can be introduced as a constraint condition. The loss of the class activation map segmentation task is used to measure the difference between the class activation map generated by the model and the real image segmentation. Therefore, the performance of the model in the class activation map segmentation task can be optimized, so as to guide the model to learn to generate a complete and accurate class activation map to predict the foreground region image during training, thereby alleviating the problem of inaccurate prediction and incomplete localization of the local activation of the target object caused by only the global label signal in the fine-grained image classification in the related art, and thus achieving the technical effect of improving the accuracy of the fine-grained image classification method. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings required for the description of the embodiments or the related art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts. In the drawings:

[0028] Figure 1 It is a schematic flowchart of an image recognition method provided by an embodiment of the present application;

[0029] Figure 2 Schematic diagram of the principle of the image recognition method provided by an embodiment of the present application;

[0030] Figure 3 Schematic diagram of the principle of determining the class activation mapping provided by an embodiment of the present application;

[0031] Figure 4 Schematic diagram of the segmentation label generation platform provided by an embodiment of the present application;

[0032] Figure 5 Schematic diagram of the principle of training the initial image recognition model provided by an embodiment of the present application;

[0033] Figure 6 Schematic diagram of the structure of the image recognition device provided by an embodiment of the present application;

[0034] Figure 7 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application.

[0035] Figure 8 Schematic diagram of the storage medium provided by an embodiment of the present application. Detailed implementation manners

[0036] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application.

[0037] Terms such as "first" and "second" in the description, claims and drawings of the embodiments of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the present solution can be implemented in an order other than the order illustrated or described in the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0038] First, some terms in the embodiments of the present application will be explained below to facilitate understanding by those skilled in the art.

[0039] Fine-grained image classification: Fine-grained image classification has been a hot research topic in the field of computer vision in recent years. Fine-grained means that the task requires more detailed sub-classification of images that belong to the same basic category in terms of subject nature (cars, plants, birds, etc.). For example, identifying the specific brand type of car, the category described by plant leaves, the specific type of animal, etc. However, due to the subtle inter-class differences and large intra-class differences between subcategories, fine-grained image classification is more difficult than coarse-grained (general) image classification tasks. Generally speaking, the difficulty of such tasks requires the algorithm side to design solutions to guide the model to focus on key areas in the image, such as the shape of the bird's beak, the outline and veins of the leaves, etc., because these areas are crucial for feature discrimination of classification.

[0040] Weakly supervised learning: refers to a deep learning method. Specifically, weakly supervised methods refer to a learning method that trains machine learning models with less labeled information or more coarse-grained labeled data. For example: in image classification tasks, although there are classification labels in the training set data, the granularity of the labels does not match the granularity of the features pointed to by the labels. The most classic example is: we only know that a picture contains a certain object, but we don’t know which specific area is the object. In this case, the model needs to infer more detailed features from the coarse-grained labels.

[0041] Class Activation Mapping (CAM): Class activation mapping, also known as class heat map, saliency map, etc., is a visualization technique mainly used in convolutional neural networks in deep learning to identify and explain which areas in images or time series data are most critical to the model's decision-making process.

[0042] Multimodal Large Language Models (MLLM): Multimodal large language models are a cutting-edge technology solution in the field of artificial intelligence. These models aim to achieve the alignment and understanding of cross-modal information through deep learning techniques and, compared to single-modal models, are capable of completing more complex and diverse cross-modal tasks. Specifically, MLLM models undergo pre-training learning through various cross-modal tasks involving different modalities (such as images, videos, text, audio, etc.). Eventually, these models are equipped with the ability to process various types of data and solve different cross-modal tasks, such as Visual Question Answering (VQA), Image-text Retrieval, Image Caption, etc. Taking the field of computer vision as an example: Multimodal large language models such as Grounding-DINO and SAM have successfully achieved the injection of text prompt information, improving the performance of visual tasks such as open-world object detection and general semantic segmentation.

[0043] Grounding-DINO: Grounding-DINO is an open-world object detection model that deeply integrates natural language and image understanding, focusing on open-vocabulary object detection and language-guided object localization. It can accurately locate and detect corresponding objects in an image according to any text description and output the predicted coordinate results.

[0044] Swin-T: As a lightweight version in the Swin Transformer series, Swin-T is a backbone network based on the Transformer architecture. The Transformer architecture initially achieved remarkable results in the field of natural language processing and has subsequently been widely applied to computer vision tasks. By introducing innovative designs such as hierarchical structures, local self-attention mechanisms, and moving windows, Swin-T has successfully applied the Transformer to image feature extraction and has become one of the important backbone networks in the field of computer vision.

[0045] SAM: Namely, the Segment Anything Model, the SAM model is a prompt-based general segmentation large model that can segment any image without any annotation. Its input includes an image and a prompt, which can be key points, coordinate regions, text, or masks, used to indicate the target to be segmented. The output is a segmentation mask, representing the probability that each pixel in the image belongs to the foreground or background.

[0046] ViT-H: ViT-H is a variant of the Vision Transformer (ViT) model, representing the high-resolution variant in Vision Transformer. ViT-H processes image data by dividing the image data into a series of image patches and inputting these image patches into the Transformer model as a sequence.

[0047] Non-Maximum Suppression (NMS): NMS is an algorithm used to remove duplicate detection boxes. In object detection tasks, the algorithm may predict multiple overlapping or similar bounding boxes, which may be generated for the same object. The purpose of NMS is to select the most suitable bounding box from these overlapping bounding boxes and remove those redundant bounding boxes to reduce the results of duplicate detections.

[0048] With the growth of the Internet industry and traffic, accurately and quickly identifying harmful information from massive data has become a new challenge. The main problem faced by the fine-grained image classification task is to obtain rich and fine-grained label recognition capabilities to deal with diverse harmful images and their variants. However, the fine-grained image recognition methods in related technologies usually can only identify high-distinguishability local features, resulting in inaccurate object localization and thus affecting the accuracy of fine-grained image classification.

[0049] To solve this technical problem, this application provides an image recognition method, device, equipment, medium and product, which is used to solve the technical problem of poor accuracy of the fine-grained image classification method in related technologies.

[0050] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problem. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the accompanying drawings.

[0051] Please refer to Figures 1 to 2 , Figure 1 which is a schematic flowchart of an image recognition method provided by an exemplary embodiment of this application, Figure 2 and is a schematic diagram of the principle of the image recognition method provided by the embodiment of this application. This method can be applied to a computing device, and this method at least includes the following steps 11 to step 15:

[0052] Step 11: Input the image to be recognized into the target image recognition model to obtain the global image classification result of the image to be recognized. Among them, the target image recognition model is obtained by training the initial image recognition model with at least the loss of the class activation mapping graph segmentation task as a constraint. The class activation mapping graph segmentation task is a task of predicting the class activation mapping graph of the sample image with the image segmentation label of the sample image as the segmentation supervision signal. The image segmentation label includes the class label corresponding to each pixel in the sample image.

[0053] The image to be recognized may contain objects or scenes. The target image recognition model is a trained image recognition model. The global image classification result of the image to be recognized is at least one predicted class and its probability of the image to be recognized. The predicted class is the class of the possible objects or scenes in the image to be recognized.

[0054] Among them, when training the initial image recognition model, the loss of the class activation mapping graph segmentation task is introduced as a constraint condition. The loss of the class activation mapping graph segmentation task is used to measure the difference between the class activation mapping graph generated by the model and the real image segmentation. Therefore, when using the loss of the class activation mapping graph segmentation task as a constraint to train the initial image recognition model, the performance of the initial image generation model in the class activation mapping graph segmentation task can be optimized, so as to guide the model to learn to generate a complete and accurate class activation mapping graph to predict the foreground region image during training, thereby alleviating the problem of inaccurate prediction and incomplete localization of the target object's local activation caused by only the global label signal in the fine-grained image classification in the related technology, and thus improving the accuracy of the fine-grained image classification method.

[0055] In this application, the total number of recognizable classes of the initial image recognition model (or target image recognition model) can be expressed as n, and the class set can be defined as C, C = {c 1 , c 2 ,..., c n}. When training the initial image recognition model with sample images, only the global classification label L x of the sample image is provided, L x = c, c ∈ C. The image segmentation label can be expressed as L y , w and h are the length and width of the input image of the target image recognition model.

[0056] It should be noted that for the target image recognition model or the initial image recognition model shown in the drawings of this application, only the model parts related to this application are shown in the drawings, which does not represent the true model structure of the target image recognition model or the initial image recognition model.

[0057] Step 12: Determine the class activation mapping graph of the image to be recognized based on the target image recognition model.

[0058] The target image recognition model can automatically generate a class activation map for each predicted category in the image to be recognized based on the relationship between the output of the model and the input image to be recognized without relying on additional annotation information. The class activation map of the image to be recognized shows the activation degree of each region in the image to be recognized for a specific category, and the darker the color, the higher the activation degree.

[0059] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the principle for determining the class activation map provided by the embodiments of the present application.

[0060] In some embodiments, determining the class activation map of the image to be recognized based on the target image recognition model may include the following steps 121 to 122:

[0061] Step 121, taking the output of the last convolutional layer of the target image recognition model as the feature map of the image to be recognized, and the feature map of the image to be recognized includes multiple channel parts;

[0062] The target image recognition model may include multiple feature layers. The target image recognition model processes the image to be recognized and outputs a feature map F ∈ R C*W*H , where C represents the number of channels of the feature map, W and H respectively represent the width and height of the feature map, and the k-th channel part of the feature map can be expressed as F k .

[0063] In Figure 3 the corresponding embodiment, the feature map includes n channels and the number of categories is 1.

[0064] Step 122, multiplying each channel part in the multiple channel parts by the channel weight corresponding to the channel part to obtain the class activation map of the image to be recognized.

[0065] First, it is necessary to obtain the channel weights of each channel part.

[0066] Specifically, connect the feature map F to a global average pooling operation (Global Average Pooling, GAP), multiply the result of the global average pooling operation on the feature map F by different channel weights W i for each channel part, and send it to a fully connected layer (Fully Connected Layer, FC Layer) to output a classification result. Thus, the prediction score result S for obtaining the c-th category can be obtained from the feature map output by the last convolutional layer in the classification task c(Before the activation function score mapping), and at the same time obtain the weight contribution value of each channel part in the feature map to the final prediction That is, the channel weights of each channel part. The prediction score result S of the c-th category can be represented by the following formula 1 c Calculation process:

[0067] Formula 1:

[0068] Among them, G(F k ) is the result of the global average pooling operation of the k-th channel part of the feature map. For the category c, For the prediction score result of this category, the importance of the k-th channel of the feature map (i.e., the channel weight of the k-th channel part corresponding to the category c), is a learnable parameter, and the channel weights will be continuously optimized and adjusted during the training of the initial image recognition model.

[0069] Furthermore, use the learned channel weights of each channel part to complete the construction process of the class activation mapping diagram for each prediction category.

[0070] Specifically, for each category, multiply each channel part of the feature map by the channel weight corresponding to this channel part, and then sum all the weighted channel parts to obtain the class activation mapping diagram for each prediction category. The calculation process of the class activation mapping diagram M of the prediction category c can be represented by the following formula 2 c Calculation process:

[0071] Formula 2:

[0072] Among them, F k (x, y) is the activation value of the (x, y) coordinate in the k-th channel part of the feature map.

[0073] Step 13, determine the foreground region image of the image to be recognized based on the class activation mapping diagram of the image to be recognized;

[0074] In an image, the foreground region is the part where the main body or the main object of interest in the image is located. These objects can be people, animals, objects, buildings, etc.

[0075] Analyze the class activation mapping diagram of the image to be recognized. According to the level of activation, the foreground region in the image to be recognized can be determined, and this panoramic region is cropped from the original image to be recognized to obtain the foreground region image of the image to be recognized.

[0076] In some embodiments, determining the foreground region image of the image to be recognized based on the class activation mapping diagram of the image to be recognized may include:

[0077] For the class activation mapping diagrams of each predicted class of the image to be recognized, mark the elements in the class activation mapping diagram that are higher than the preset class activation threshold Th seg as positive samples, and mark the rest as negative samples to obtain an initial foreground region mask;

[0078] Remove the isolated noisy positive samples in the initial foreground region mask that are not connected to other positive samples to obtain a continuous foreground region mask;

[0079] Determine the initial foreground region image of the image to be recognized according to the minimum bounding rectangle corresponding to the continuous foreground region mask;

[0080] Scale the initial foreground region image of the image to be recognized to the same input size as the image to be recognized to obtain the foreground region image of the image to be recognized.

[0081] In some embodiments, positive samples can be marked as 1 and negative samples can be marked as 0. After obtaining the minimum bounding rectangle corresponding to the continuous foreground region mask, cut out the minimum bounding rectangle from the image to be recognized to obtain the initial foreground region image of the image to be recognized.

[0082] Step 14, input the foreground region image of the image to be recognized into the target image recognition model to obtain the foreground region classification result of the image to be recognized;

[0083] Input the foreground region image of the image to be recognized into the target image recognition model for secondary prediction to obtain the foreground region classification result of the image to be recognized.

[0084] Step 15, determine the target classification result of the image to be recognized based on the global image classification result and the foreground region classification result of the image to be recognized.

[0085] After obtaining the global image classification result and the foreground region classification result of the image to be recognized, compare and analyze the global image classification result and the foreground region classification result, and comprehensively judge the target classification result of the image to be recognized according to the comparison result.

[0086] For example, if the global classification result and the foreground classification result are the same or highly similar, the global classification result can be directly used as the final target classification result; if there are differences between the two, the reasons for the differences need to be further analyzed and a final judgment should be made according to the actual situation.

[0087] In some embodiments, determining the target classification result of the image to be recognized based on the global image classification result and the foreground region classification result of the image to be recognized includes:

[0088] Perform a weighted sum processing on the global image classification result and the foreground region classification result of the image to be recognized to obtain the target classification result of the image to be recognized.

[0089] Among them, the weights of the global image classification result and the foreground region classification result can be reasonably set according to the application scenario.

[0090] The image recognition method provided by the embodiments of the present application optimizes the performance of the initial image recognition model in generating the class activation mapping diagram by introducing the loss of the class activation mapping diagram segmentation task as a training constraint, enabling the target image recognition model to more accurately predict the foreground region image, thereby improving the accuracy of fine-grained image classification and enhancing the generalization ability of the target image recognition model. In addition, the image recognition method provided by the embodiments of the present application does not rely on additional annotation information, but uses the ability of the target image recognition model itself to generate the class activation mapping diagram and determines the foreground region image accordingly, greatly reducing the cost and time of manual annotation and improving work efficiency.

[0091] In some embodiments, the steps of training the initial image recognition model with the loss of the class activation mapping diagram segmentation task as a constraint include the following steps 110 to 140:

[0092] Step 110, determine the class activation mapping diagram of the sample image according to the initial image recognition model;

[0093] In some embodiments, determining the class activation mapping diagram of the sample image according to the initial image recognition model may include the following steps 1101 to 1102:

[0094] Step 1101, use the output of the last convolutional layer of the initial image recognition model as the feature map of the sample image, and the feature map of the sample image includes multiple channel parts;

[0095] Step 1101, multiply each channel part in the multiple channel parts by the channel weight to be trained corresponding to the channel part to obtain the class activation mapping diagram of the sample image.

[0096] The specific implementation steps of steps 1101 to 1102 are the same as the principles of steps 121 to 122 in the above embodiments. For details, please refer to the description of the above embodiments and will not be elaborated here.

[0097] Step 120, determine the image segmentation label of the sample image based on the preset multimodal large model;

[0098] In this embodiment, utilize the excellent generalization ability and zero-shot learning ability of the preset multimodal large model to determine the image segmentation label of the sample image and implement a high-quality pseudo-label generation task.

[0099] In some embodiments, the present application provides a split label generation platform. Refer to Figure 4 , Figure 4 which is a schematic diagram of the split label generation platform provided by the embodiments of the present application. The split label generation platform includes a user input side and an annotation platform side, and this split label generation platform can be used to determine the image segmentation labels of sample images based on a preset multimodal large model.

[0100] Specifically, the sample image A i (image input) and the class label A t (text input) of the sample image, a preset classification score confidence threshold, a preset candidate box confidence threshold, a preset redundant box post-processing threshold, a preset segmentation threshold, etc. are input into this split label generation platform, and the split label generation platform automatically outputs the image segmentation labels of the sample image. The image segmentation labels of the sample image include the class labels corresponding to each pixel in the sample image.

[0101] The construction of this split label generation platform makes the generation of image segmentation labels more efficient and accurate, and realizes end-to-end one-key segmentation annotation data acquisition without adding additional manual annotation work. The split label generation platform can be seen in the description of the following embodiments.

[0102] In some embodiments, determining the image segmentation labels of sample images based on a preset multimodal large model may include the following steps 01 to step 03:

[0103] Step 01, determining the candidate box hint information corresponding to the sample image based on the first preset multimodal large model, where the candidate box hint information includes the position information of at least one target candidate box;

[0104] The candidate box hint information corresponding to the sample image includes hints of the image regions of possible objects of interest (target objects, such as objects, faces, etc.). The candidate box hint information is given in the form of candidate boxes and is used to mark the possible target positions in the sample image. The candidate box hint information includes one or more target candidate boxes, and each target candidate box corresponds to a possible position.

[0105] The set H of target candidate boxes can be represented by the following formula 3:

[0106] Formula 3, H = {H bbox1 , S bbox2 ,..., H bboxn};

[0107] The first preset multimodal large model can be any type of multimodal large model that can process and fuse multiple modalities of data (such as images, texts, etc.).

[0108] In some embodiments, the first preset multimodal large model may be the Grounding-DINO model. More specifically, it may be the Grounding-DINO pre-trained model with Swin-T as the backbone network.

[0109] In some embodiments, determining the candidate box prompt information corresponding to the sample image based on the first preset multimodal large model may include the following steps 011 to step 012:

[0110] Step 011: Input the sample image into the first preset multimodal large model to obtain the initial candidate box information corresponding to the sample image. Among them, the initial candidate box information includes the prediction results of at least one initial candidate box. The prediction results of the initial candidate box include the position information of the initial candidate box, the predicted category of the initial candidate box, the classification confidence of the predicted category of the initial candidate box, and the foreground confidence of the initial candidate box.

[0111] Specifically, taking the Grounding-DINO pre-trained model with Swin-T as the backbone network as the first preset multimodal large model as an example, input the sample image into the Grounding-DINO pre-trained model for object detection to obtain the prediction result set S of the sample image. The prediction result set S of the sample image can be represented by the following formula 4:

[0112] Formula 4: S = {S bbox1 , S bbox2 ,..., S bboxn};

[0113] Each element in S represents the prediction results of the detected initial candidate boxes.

[0114] Among them, the prediction result S bbox of the initial candidate box can be represented by the following formula 5:

[0115] Formula 5: S bbox = {Class, Score cls , Score box , x, y, w, h};

[0116] Among them, Class is the predicted category of the initial candidate box; Score cls is the classification confidence of the predicted category of the initial candidate box; Score box is the foreground confidence of the initial candidate box; x, y, w, h respectively refer to the position coordinates of the initial candidate box.

[0117] Step 012: Based on the preset classification score confidence threshold Threshold cls and the preset candidate box confidence threshold Thresholdbox Perform low-quality box filtering and redundant box filtering on at least one initial candidate box based on a preset redundancy box post-processing threshold IoU to obtain at least one target candidate box.

[0118] Since there are low-quality prediction boxes and redundant prediction boxes in the original detection of the first preset multimodal large model, certain post-processing screening operations are required, including low-quality box filtering and redundant box filtering.

[0119] In some embodiments, performing low-quality box filtering and redundant box filtering on at least one initial candidate box based on a preset classification score confidence threshold, a preset candidate box confidence threshold, and a preset redundancy box post-processing threshold may include the following steps 021 to step 022:

[0120] Step 021, for each initial candidate box, if the classification confidence Score cls of the initial candidate box is less than the preset classification score confidence threshold Threshold cls or the foreground confidence Score box of the initial candidate box is less than the preset candidate box confidence threshold Threshold box , filter the initial candidate box as a low-quality candidate box for low-quality filtering to obtain at least one initial candidate box after low-quality box filtering;

[0121] Step 022, perform redundant box filtering on at least one initial candidate box after low-quality box filtering according to the preset redundancy box post-processing threshold.

[0122] Since there may still be cases where the original detections of the first preset multimodal large model are close to each other and the predicted target is the same object, it is also necessary to perform redundant box filtering on at least one initial candidate box after low-quality box filtering, so as to retain the best candidate box for each target object and remove those redundant candidate boxes with a high degree of overlap at the same time.

[0123] In some embodiments, any type of candidate box redundant filtering method or candidate box deduplication method can be used to perform redundant box filtering on at least one initial candidate box after low-quality box filtering according to the preset redundancy box post-processing threshold IoU, and the present application does not limit this.

[0124] In a specific embodiment provided by the embodiments of the present application, perform redundant box filtering on at least one initial candidate box after low-quality box filtering according to the non-maximum suppression method and the preset redundancy box post-processing threshold IoU.

[0125] Step 02: Input the candidate box prompt information and the sample image into the second preset multimodal large model to obtain the pixel-level prediction result of the sample image. Among them, the pixel-level prediction result includes the prediction probability that each pixel in the sample image belongs to the foreground object.

[0126] Use all elements in the target candidate box set H as the candidate box prompt information box prompt , and combine it with the sample image A i , and jointly input them into the second preset multimodal large model for forward inference to obtain the pixel-level prediction result M of the sample image. M is a matrix of size w*h (w and h represent the width and height of the sample image respectively), where each element represents the prediction probability of the second preset multimodal large model for whether the pixel belongs to the foreground object, and the value of each element is in the range of [0,1].

[0127] Among them, the second preset multimodal large model can be any type of multimodal large model, and the specific type of the second preset multimodal large model is not limited in this application.

[0128] In a specific embodiment provided by this application, the second preset multimodal large model can be the SAM pre-trained model with ViT-H as the backbone network.

[0129] Step 03: Determine the image segmentation label of the sample image based on the pixel-level prediction result of the sample image.

[0130] In some embodiments, determining the image segmentation label of the sample image based on the pixel-level prediction result of the sample image includes the following steps 031 to 032:

[0131] Step 031: For each element in the pixel-level prediction result of the sample image, perform binarization processing on the prediction probability corresponding to the element according to the preset segmentation threshold.

[0132] Specifically, traverse the prediction probability corresponding to each element in the pixel-level prediction result M of the sample image, compare the prediction probability with the preset segmentation threshold Threshold seg , and set the prediction probability corresponding to the element greater than Threshold seg to 1, otherwise set it to 0 to complete the binarization processing operation.

[0133] Step 032: Perform label formatting operation and scaling processing on the pixel-level prediction result after binarization processing to obtain the image segmentation label of the sample image. The image segmentation label of the sample image meets the input format for training the initial image recognition model.

[0134] Although the pixel-level prediction results after binarization can clearly show the distinction between the foreground and the background, they may not meet the input format requirements for training the initial image recognition model. Therefore, it is necessary to perform label formatting operations on the pixel-level prediction results after binarization, which may include adjusting the size, resolution, or format of the labels, etc., to ensure that they match the input format of the initial image recognition model.

[0135] Meanwhile, if the sample image has been scaled or cropped during the preprocessing stage, then the generated image segmentation labels also need to be scaled accordingly to maintain consistency with the sample image.

[0136] Step 130, calculate the loss of the class activation map segmentation task based on the class activation map of the sample image and the image segmentation label of the sample image;

[0137] Use the image segmentation label of the sample image as the segmentation supervision signal, and use the class activation map of the sample image as the segmentation prediction result of the initial image recognition model to optimize the class activation map segmentation task of the initial image recognition model.

[0138] In some embodiments, the cross-entropy loss function can be used to calculate the loss of the class activation map segmentation task based on the class activation map of the sample image and the image segmentation label of the sample image. Specifically, the loss L of the class activation map segmentation task can be calculated by the following formula 6 seg :

[0139] Formula 6:

[0140] where N represents the total number of samples (i.e., the total number of pixels in the class activation map), C represents the total number of predicted classes, pred i,j represents the prediction result of whether the i-th sample belongs to the predicted class j, and y i,j is the true label of whether the i-th sample in the image segmentation label belongs to the j-th predicted class.

[0141] Step 140, adjust the parameters of the initial image recognition model based on the loss of the class activation map segmentation task.

[0142] Specifically, after calculating the loss of the class activation map segmentation task, the gradient of the loss of the class activation map segmentation task with respect to the parameters of the initial image recognition model can be calculated using the backpropagation algorithm, and the parameters of the initial image recognition model can be updated using an optimization algorithm according to the gradient. This process is iterated until the loss converges or reaches the preset number of iterations, ultimately improving the performance of the initial image recognition model in the class activation map segmentation task.

[0143] In some embodiments, when adjusting the parameters of the initial image recognition model based on the loss of the category activation map segmentation task, the training channel weights corresponding to each channel part of the output feature map of the last convolutional layer of the initial image recognition model are adjusted.

[0144] In some embodiments, the initial image recognition model is also trained with the loss of the global image classification task as a constraint. The global image classification task is a task of performing global image classification prediction on a sample image with the sample label of the sample image as the classification label supervision signal.

[0145] In some embodiments, the steps of training the initial image recognition model with the loss of the global image classification task as a constraint include:

[0146] Input the sample image into the initial image recognition model to obtain the global image classification result of the sample image;

[0147] Calculate the loss of the global image classification task according to the global image classification result of the sample image and the class label of the sample image;

[0148] Adjust the parameters of the initial image recognition model based on the loss of the global image classification task.

[0149] In some embodiments, the loss L of the global image classification task can be calculated by the following formula 7 img :

[0150] Formula 7:

[0151] where C represents the total number of categories, pred j represents the probability that the sample image belongs to category j, and y j represents whether the class label of the sample image belongs to category j. If so, it is assigned 1, otherwise 0.

[0152] In some embodiments, the initial image recognition model is also trained with the loss of the foreground region classification task as a constraint. The foreground region classification task is a task of performing foreground region classification prediction on a sample image with the class label of the sample image as the classification label supervision signal.

[0153] In some embodiments, the steps of training the initial image recognition model with the loss of the foreground region classification task as a constraint include:

[0154] Determine the category activation map of the sample image according to the initial image recognition model;

[0155] Generate the foreground region classification result of the sample image according to the category activation map of the sample image;

[0156] Input the foreground region image of the sample image into the initial image recognition model to obtain the foreground region classification result of the sample image;

[0157] Calculate the loss of the foreground region classification task according to the foreground region classification result of the sample image and the sample label of the sample image.

[0158] In some embodiments, the loss L of the foreground region classification task can be calculated by the following formula 8 region :

[0159] Formula 8:

[0160] where C represents the total number of categories, pred j represents the probability that the foreground region image of the sample image belongs to category j, and y j represents whether the category label of the sample image belongs to category j. If so, it is assigned 1, otherwise 0.

[0161] Please refer to Figure 5 , Figure 5 , which is a schematic diagram of the principle for training the initial image recognition model provided by the embodiments of the present application.

[0162] In some embodiments, the initial image recognition model is also trained with the loss of the global image classification task and the loss of the foreground region classification task as constraints. The global image classification task is a task of performing global image classification prediction on the sample image with the sample label of the sample image as the classification label supervision signal, and the foreground region classification task is a task of performing foreground region classification prediction on the sample image with the sample label of the sample image as the classification label supervision signal.

[0163] Specifically, in addition to being trained with the loss of the class activation map segmentation task as a constraint, the initial image recognition model is also trained with the loss of the global image classification task and the loss of the foreground region classification task as constraints. The global image classification task, the foreground region classification task, and the class activation map segmentation task are jointly used as the optimization tasks of the initial image recognition model, and the initial image recognition model is trained using a multi-task joint optimization training method to complete the recognition optimization of the fine-grained classification task.

[0164] In some embodiments, training the initial image recognition model according to the loss of the class activation map segmentation task, the loss of the global image classification task, and the loss of the foreground region classification task includes:

[0165] Calculate the multi-task joint loss according to the loss of the global image classification task, the loss of the foreground region image generation task, and the loss of the class activation map segmentation task;

[0166] Adjust the parameters of the initial image recognition model based on the multi-task joint loss to obtain the target image recognition model.

[0167] In some embodiments, calculate the multi-task joint loss according to the loss of the global image classification task, the loss of the foreground region image generation task, and the loss of the class activation map segmentation task, including:

[0168] Perform weighted summation on the loss of the global image classification task, the loss of the foreground region image generation task, and the loss of the class activation map segmentation task to obtain the multi-task joint loss.

[0169] In some embodiments, the multi-task joint loss Loss can be calculated by the following formula 9:

[0170] Formula 9: Loss = α * L img + β * L region + γ * L seg ;

[0171] Where α is the weight value of the global image classification task, β is the weight value of the foreground region image generation task, γ is the loss of the class activation map segmentation task, the values of α, β, and γ can be set according to the model training effect, and the sum of α, β, and γ is 1.

[0172] In some embodiments, the initial values of α, β, and γ can be all set to 1 / 3.

[0173] The present application provides a method of inputting an image to be recognized into a target image recognition model to obtain a global image classification result of the image to be recognized; determining a class activation map of the image to be recognized based on the target image recognition model; determining a foreground region image of the image to be recognized based on the class activation map of the image to be recognized; inputting the foreground region image of the image to be recognized into the target image recognition model to obtain a foreground region classification result of the image to be recognized; and determining a target classification result of the image to be recognized based on the global image classification result and the foreground region classification result of the image to be recognized. The target image recognition model is obtained by training an initial image recognition model with at least the loss of a class activation map segmentation task as a constraint. The class activation map segmentation task is a task of predicting a class activation map for a sample image with an image segmentation label of the sample image as a segmentation supervision signal. The image segmentation label includes class labels corresponding to each pixel in the sample image. In the training of the initial image recognition model, the loss of the class activation map segmentation task can be introduced as a constraint condition. The loss of the class activation map segmentation task is used to measure the difference between the class activation map generated by the model and the real image segmentation. Therefore, the performance of the model in the class activation map segmentation task can be optimized, so as to guide the model to learn to generate a complete and accurate class activation map to predict the foreground region image during training, thereby alleviating the problem of inaccurate prediction and incomplete localization of local activation of the target object caused by only the global label signal in fine-grained image classification in the related art, and thus achieving the technical effect of improving the accuracy of the fine-grained image classification method.

[0174] Figure 6 FIG. 4 is a schematic structural diagram of an image recognition device provided by an exemplary embodiment of the present application;

[0175] Wherein, the device includes:

[0176] A first input unit 61, configured to input an image to be recognized into a target image recognition model to obtain a global image classification result of the image to be recognized;

[0177] A first determination unit 62, configured to determine a class activation map of the image to be recognized based on the target image recognition model;

[0178] A second determination unit 63, configured to determine a foreground region image of the image to be recognized based on the class activation map of the image to be recognized;

[0179] A second input unit 64, configured to input the foreground region image of the image to be recognized into the target image recognition model to obtain a foreground region classification result of the image to be recognized;

[0180] A third determination unit 65, configured to determine a target classification result of the image to be recognized based on the global image classification result and the foreground region classification result of the image to be recognized;

[0181] Among them, the target image recognition model is obtained by training the initial image recognition model with at least the loss of the class activation map segmentation task as a constraint. The class activation map segmentation task is a task of predicting the class activation map of a sample image with the image segmentation label of the sample image as the segmentation supervision signal. The image segmentation label includes the class labels corresponding to each pixel in the sample image.

[0182] In some embodiments, when the third determination unit 65 is used to determine the target classification result of the image to be recognized based on the global image classification result and the foreground region classification result of the image to be recognized, it is specifically used for:

[0183] Performing a weighted sum processing on the global image classification result and the foreground region classification result of the image to be recognized to obtain the target classification result of the image to be recognized.

[0184] In some embodiments, the device further includes a training unit. When the training unit is used to train the initial image recognition model with the loss of the class activation map segmentation task as a constraint, it is specifically used for:

[0185] Determining the class activation map of the sample image according to the initial image recognition model;

[0186] Determining the image segmentation label of the sample image based on a preset multimodal large model;

[0187] Calculating the loss of the class activation map segmentation task according to the class activation map of the sample image and the image segmentation label of the sample image;

[0188] Adjusting the parameters of the initial image recognition model based on the loss of the class activation map segmentation task.

[0189] In some embodiments, when the training unit is used to determine the class activation map of the sample image according to the initial image recognition model, it is specifically used for:

[0190] Taking the output of the last convolutional layer of the initial image recognition model as the feature map of the sample image. The feature map of the sample image includes multiple channel parts;

[0191] Multiplying each channel part in the multiple channel parts by the trainable channel weight corresponding to the channel part to obtain the class activation map of the sample image.

[0192] In some embodiments, when the training unit is used to determine the image segmentation label of the sample image based on a preset multimodal large model, it is specifically used for:

[0193] Determine candidate box prompt information corresponding to the sample image based on a first preset multimodal large model, where the candidate box prompt information includes the position information of at least one target candidate box;

[0194] Input the candidate box prompt information and the sample image into a second preset multimodal large model to obtain a pixel-level prediction result of the sample image, where the pixel-level prediction result includes the prediction probability that each pixel in the sample image belongs to the foreground target;

[0195] Determine the image segmentation label of the sample image based on the pixel-level prediction result of the sample image.

[0196] In some embodiments, when the training unit is used to determine the image segmentation label of the sample image based on the pixel-level prediction result of the sample image, it is specifically used for:

[0197] For each element in the pixel-level prediction result of the sample image, perform binarization processing on the prediction probability corresponding to the element according to a preset segmentation threshold;

[0198] Perform label formatting operation and scaling processing on the binarized pixel-level prediction result to obtain the image segmentation label of the sample image, and the image segmentation label of the sample image meets the input format for model training of the initial image recognition model.

[0199] In some embodiments, when the training unit is used to determine the candidate box prompt information corresponding to the sample image based on the first preset multimodal large model, it is specifically used for:

[0200] Input the sample image into the first preset multimodal large model to obtain initial candidate box information corresponding to the sample image, where the initial candidate box information includes the prediction results of at least one initial candidate box, and the prediction results of the initial candidate box include the position information of the initial candidate box, the predicted category of the initial candidate box, the classification confidence of the predicted category of the initial candidate box, and the foreground confidence of the initial candidate box;

[0201] Perform low-quality box filtering processing and redundant box filtering processing on at least one initial candidate box based on a preset classification score confidence threshold, a preset candidate box confidence threshold, and a preset redundant box post-processing threshold to obtain at least one target candidate box.

[0202] In some embodiments, when the training unit is used to perform low-quality box filtering processing and redundant box filtering processing on at least one initial candidate box based on a preset classification score confidence threshold, a preset candidate box confidence threshold, and a preset redundant box post-processing threshold, it is specifically used for:

[0203] For each initial candidate box, if the classification confidence of the initial candidate box is less than the preset classification score confidence threshold or the foreground confidence of the initial candidate box is less than the preset candidate box confidence threshold, the initial candidate box is taken as a low-quality candidate box for low-quality filtering, and at least one initial candidate box after low-quality box filtering processing is obtained;

[0204] Perform redundant box filtering processing on at least one initial candidate box after low-quality box filtering processing according to the preset redundant box post-processing threshold.

[0205] In some embodiments, the initial image recognition model is also trained with the loss of the global image classification task as a constraint. The global image classification task is a task of performing global image classification prediction on a sample image with the sample label of the sample image as a classification label supervision signal.

[0206] In some embodiments, the initial image recognition model is also trained with the loss of the foreground region classification task as a constraint. The foreground region classification task is a task of performing foreground region classification prediction on a sample image with the sample label of the sample image as a classification label supervision signal.

[0207] In some embodiments, the initial image recognition model is also trained with the loss of the global image classification task and the loss of the foreground region classification task as constraints. The global image classification task is a task of performing global image classification prediction on a sample image with the sample label of the sample image as a classification label supervision signal. The foreground region classification task is a task of performing foreground region classification prediction on a sample image with the sample label of the sample image as a classification label supervision signal.

[0208] In some embodiments, the training unit is further configured to train the initial image recognition model according to the loss of the class activation map segmentation task, the loss of the global image classification task, and the loss of the foreground region classification task. Specifically, it is configured to:

[0209] Calculate a multi-task joint loss according to the loss of the global image classification task, the loss of the foreground region image generation task, and the loss of the foreground region classification task;

[0210] Adjust the parameters of the initial image recognition model based on the multi-task joint loss to obtain the target image recognition model.

[0211] In some embodiments, when the training unit is used to calculate the multi-task joint loss according to the loss of the global image classification task, the loss of the foreground region image generation task, and the loss of the foreground region classification task, it is specifically configured to:

[0212] Perform weighted summation on the loss of the global image classification task, the loss of the foreground region image generation task, and the loss of the class activation map segmentation task to obtain the multi-task joint loss.

[0213] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, the device can execute the above method embodiments, and the foregoing and other operations and / or functions of each module in the device respectively correspond to the corresponding processes in each method in the above method embodiments. For the sake of brevity, they will not be elaborated here.

[0214] The device of the embodiments of the present application has been described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in mature storage media in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage media is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0215] Figure 7 is a schematic block diagram of an electronic device provided by an embodiment of the present application. The electronic device may include:

[0216] A memory 701 and a processor 702. The memory 701 is used to store a computer program and transmit the program code to the processor 702. In other words, the processor 702 can call and run the computer program from the memory 701 to implement the method in the embodiments of the present application.

[0217] For example, the processor 702 can be used to execute the above method embodiments according to the instructions in the computer program.

[0218] In some embodiments of the present application, the processor 702 may include, but is not limited to:

[0219] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.

[0220] In some embodiments of the present application, the memory 701 includes, but is not limited to:

[0221] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), Programmable ROM (PROM), Erasable PROM (EPROM), Electrically Erasable PROM (EEPROM), or flash memory. The volatile memory can be a Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate Synchronous DRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synch Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0222] In some embodiments of the present application, the computer program can be divided into one or more modules, and the one or more modules are stored in the memory 701 and executed by the processor 702 to complete the method provided by the present application. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0223] As Figure 7 shown, the electronic device may further include:

[0224] A transceiver 703, which can be connected to the processor 702 or the memory 701.

[0225] Among them, the processor 702 can control the transceiver 703 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 703 can include a transmitter and a receiver. The transceiver 703 may further include an antenna, and the number of antennas can be one or more.

[0226] It should be understood that the various components in the electronic device are connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.

[0227] This application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the methods in the above method embodiments. Or rather, the embodiments of this application also provide a computer program product containing computer instructions. When the computer instructions are executed by a computer, the computer executes the methods in the above method embodiments.

[0228] Figure 8 Schematic diagram of the storage medium provided by the embodiments of this application. As Figure 8 shown, it depicts a program product 800 for implementing the above method according to an exemplary embodiment of this application. It can adopt a portable compact disc read-only memory (CDROM) and include program code, and can run on a computer device, such as a mobile phone. However, the program product of this application is not limited to this. In this application, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component.

[0229] When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0230] Those of ordinary skill in the art will realize that the modules and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0231] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or modules can be in electrical, mechanical, or other forms.

[0232] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0233] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.

[0234] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. An image recognition method, characterized in that: include: Inputting the image to be identified into the target image recognition model to obtain a global image classification result of the image to be identified; Determining a class activation map of the image to be recognized based on the target image recognition model; Determining a foreground area image of the image to be identified based on the class activation map of the image to be identified; Inputting the foreground area image of the image to be identified into the target image recognition model to obtain a foreground area classification result of the image to be identified; Determining a target classification result of the image to be identified based on a global image classification result of the image to be identified and a foreground area classification result; Among them, the target image recognition model is obtained by training the initial image recognition model with at least the loss of the class activation map segmentation task as a constraint. The class activation map segmentation task is a task of predicting the class activation map of the sample image using the image segmentation label of the sample image as the segmentation supervision signal. The image segmentation label includes the class label corresponding to each pixel in the sample image.

2. The method according to claim 1, characterized in that The determining the target classification result of the image to be identified based on the global image classification result of the image to be identified and the foreground area classification result includes: A weighted summation process is performed on the global image classification result of the image to be identified and the foreground area classification result to obtain a target classification result of the image to be identified.

3. The method according to claim 1, characterized in that The steps of training the initial image recognition model using the loss of the class activation map segmentation task as a constraint include: Determining a class activation map of the sample image according to an initial image recognition model; Determining an image segmentation label of the sample image based on a preset multimodal large model; Calculating the loss of the class activation map segmentation task based on the class activation map of the sample image and the image segmentation label of the sample image; Parameters of the initial image recognition model are adjusted based on the loss of the class activation map segmentation task.

4. The method according to claim 3, characterized in that The step of determining the category activation map of the sample image according to the initial image recognition model comprises: Using the output of the last convolutional layer of the initial image recognition model as a feature map of the sample image, wherein the feature map of the sample image includes multiple channel parts; Each channel part of the multiple channel parts is multiplied by the channel weight to be trained corresponding to the channel part to obtain a category activation map of the sample image.

5. The method according to claim 3, characterized in that: The determining the image segmentation label of the sample image based on the preset multimodal large model includes: Determining candidate box prompt information corresponding to the sample image based on the first preset multimodal large model, wherein the candidate box prompt information includes position information of at least one target candidate box; Inputting the candidate box prompt information and the sample image into a second preset multimodal large model to obtain a pixel-level prediction result of the sample image, wherein the pixel-level prediction result includes a prediction probability that each pixel in the sample image belongs to a foreground object; An image segmentation label of the sample image is determined based on a pixel-level prediction result of the sample image.

6. The method according to claim 5, characterized in that The determining the image segmentation label of the sample image based on the pixel-level prediction result of the sample image comprises: For each element in the pixel-level prediction result of the sample image, binarization processing is performed on the prediction probability corresponding to the element according to a preset segmentation threshold; The pixel-level prediction result after binarization is subjected to label formatting and scaling to obtain an image segmentation label of the sample image, wherein the image segmentation label of the sample image satisfies an input format for model training of the initial image recognition model.

7. An image recognition device, characterized in that: include: A first input unit, used to input the image to be identified into the target image recognition model to obtain a global image classification result of the image to be identified; A first determining unit, configured to determine a class activation map of the image to be recognized based on the target image recognition model; A second determining unit, configured to determine a foreground area image of the image to be identified based on the class activation map of the image to be identified; A second input unit, used for inputting the foreground area image of the image to be identified into the target image recognition model to obtain a foreground area classification result of the image to be identified; A third determining unit, configured to determine a target classification result of the image to be identified based on a global image classification result of the image to be identified and a foreground area classification result; Among them, the target image recognition model is obtained by training the initial image recognition model with at least the loss of the class activation map segmentation task as a constraint. The class activation map segmentation task is a task of predicting the class activation map of the sample image using the image segmentation label of the sample image as the segmentation supervision signal. The image segmentation label includes the class label corresponding to each pixel in the sample image.

8. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 6 by executing the executable instructions.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.