Image recognition method and device, equipment, storage medium and computer program product
By aligning image regions with their descriptive text through region-wise training, the model's ability to recognize fine details is improved, addressing the limitations of existing image recognition models.
Patent Information
- Application Number
- CN202510361403.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-15
AI Technical Summary
When performing image recognition, existing image recognition models usually align the entire image with the corresponding text description, resulting in the inability to recognize fine objects in the image and poor image recognition effect.
An image recognition model with regional fine-grained alignment training is used to align the image area with area description, and image recognition is performed through preset image recognition models to improve the recognition accuracy of image fine targets.
By carefully aligning the image area and area description, the accuracy of the fine object recognition of the image recognition model is improved and the image recognition effect is improved.
Smart Images

Figure CN120318803A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to an image recognition method, device, equipment, storage medium and computer program product. Background Art
[0002] Currently, fine-grained image recognition is becoming increasingly important in image recognition applications. For example, in wildlife monitoring, different species of birds or animals are recognized. However, when performing image recognition, related image recognition models usually align the entire image with the corresponding text description, resulting in the defect that the model cannot recognize fine targets in the image and the image recognition effect is poor. Summary of the Invention
[0003] The main purpose of the present application is to provide an image recognition method, device, equipment, storage medium and computer program product, aiming to solve the technical problem that when performing image recognition, related image recognition methods usually align the entire image with the corresponding text description, resulting in the defect that the model cannot recognize fine targets in the image and the image recognition effect is poor.
[0004] To achieve the above object, the present application provides an image recognition method, and the image recognition method includes:
[0005] Input the image to be recognized into a preset image recognition model, where the preset image recognition model is a model obtained through region fine-grained alignment training, and the region fine-grained alignment training is training for aligning an image region with a region description corresponding to the image region;
[0006] Perform image recognition on the image to be recognized through the preset image recognition model to obtain an image recognition result of the image to be recognized.
[0007] Optionally, before inputting the image to be recognized into the preset image recognition model, it further includes:
[0008] Obtain an image-text pair, where the image-text pair includes an image sample and a short text description corresponding to the image sample;
[0009] Perform global fine-grained alignment training on the initial image recognition model based on the short text description and the image sample to obtain a global fine-grained loss value of the initial image recognition model;
[0010] Obtain an image region in the image sample and a region description corresponding to the image region, where the image region is used to represent the positions of each object in the image sample;
[0011] Perform region fine-grained alignment training on the initial image recognition model based on the image region and the region description to obtain the region fine-grained loss value of the initial image recognition model;
[0012] Adjust the initial image recognition model according to the global fine-grained loss value and the region fine-grained loss value to obtain a preset image recognition model.
[0013] Optionally, the performing region fine-grained alignment training on the initial image recognition model based on the image region and the region description to obtain the region fine-grained loss value of the initial image recognition model includes:
[0014] Perform region feature aggregation on the image region to obtain the region feature of the image region;
[0015] Perform text encoding on the region description to obtain the semantic feature vector of the region description;
[0016] Perform region fine-grained alignment training on the initial image recognition model based on the region feature of the image region and the semantic feature vector of the region description to obtain the region fine-grained loss value of the initial image recognition model.
[0017] Optionally, the performing region fine-grained alignment training on the initial image recognition model based on the region feature of the image region and the semantic feature vector of the region description to obtain the region fine-grained loss value of the initial image recognition model includes:
[0018] Calculate the region feature similarity between the region feature of the image region and the semantic feature vector of the region description;
[0019] Perform region fine-grained alignment training on the initial image recognition model based on the region feature similarity to obtain the region fine-grained loss value of the initial image recognition model.
[0020] Optionally, before adjusting the initial image recognition model according to the global fine-grained loss value and the region fine-grained loss value to obtain a preset image recognition model, it further includes:
[0021] Select a target image region from the image regions, where the target image region corresponds to multiple region descriptions, and the multiple region descriptions include the correct description of the target image region and the semantic similar descriptions corresponding to the correct description;
[0022] Perform high-difficulty fine-grained alignment training on the initial image recognition model based on the target image region, the correct description, and the semantic similar descriptions to obtain the high-difficulty fine-grained loss value of the initial image recognition model;
[0023] Correspondingly, adjusting the initial image recognition model according to the global fine-grained loss value and the regional fine-grained loss value to obtain a preset image recognition model includes:
[0024] Adjusting the initial image recognition model according to the global fine-grained loss value, the regional fine-grained loss value, and the high-difficulty fine-grained loss value to obtain a preset image recognition model.
[0025] Optionally, the high-difficulty fine-grained alignment training of the initial image recognition model based on the target image region, the correct description, and the semantically similar description to obtain the high-difficulty fine-grained loss value of the initial image recognition model includes:
[0026] Performing regional feature aggregation on the target image region to obtain the regional feature of the target image region;
[0027] Performing text encoding on the correct description and the semantically similar description to obtain the semantic feature vector of the correct description and the semantic feature vector of the semantically similar description;
[0028] Performing regional fine-grained alignment training on the initial image recognition model based on the regional feature of the target image region, the semantic feature vector of the correct description, and the semantic feature vector of the semantically similar description to obtain the regional fine-grained loss value of the initial image recognition model.
[0029] Optionally, the performing regional fine-grained alignment training on the initial image recognition model based on the regional feature of the target image region, the semantic feature vector of the correct description, and the semantic feature vector of the semantically similar description to obtain the regional fine-grained loss value of the initial image recognition model includes:
[0030] Calculating the positive sample similarity between the regional feature of the target image region and the semantic feature vector of the correct description;
[0031] Calculating the negative sample similarity between the regional feature of the target image region and the semantic feature vector of the correct description;
[0032] Performing regional fine-grained alignment training on the initial image recognition model based on the positive sample similarity and the negative sample similarity to obtain the regional fine-grained loss value of the initial image recognition model.
[0033] Optionally, the adjusting the initial image recognition model according to the global fine-grained loss value, the regional fine-grained loss value, and the high-difficulty fine-grained loss value to obtain a preset image recognition model includes:
[0034] Obtain the regional loss weight value corresponding to the regional fine-grained loss value, and obtain the high-difficulty loss weight value corresponding to the high-difficulty fine-grained loss value;
[0035] Calculate the total loss value of the initial image recognition model according to the global fine-grained loss value, the regional fine-grained loss value, the regional loss weight value, the high-difficulty fine-grained loss value, and the high-difficulty loss weight value;
[0036] Adjust the initial image recognition model according to the total loss value to obtain a preset image recognition model.
[0037] Optionally, before obtaining the global fine-grained loss value of the initial image recognition model by performing global fine-grained alignment training on the initial image recognition model based on the short text description and the image sample, further include:
[0038] Generate a corresponding long text description for the image sample through a multimodal large model;
[0039] Correspondingly, obtaining the global fine-grained loss value of the initial image recognition model by performing global fine-grained alignment training on the initial image recognition model based on the short text description and the image sample includes:
[0040] Perform global fine-grained alignment training on the initial image recognition model based on the long text description, the short text description, and the image sample to obtain the global fine-grained loss value of the initial image recognition model.
[0041] Optionally, obtaining the global fine-grained loss value of the initial image recognition model by performing global fine-grained alignment training on the initial image recognition model based on the long text description, the short text description, and the image sample includes:
[0042] Perform image encoding on the image sample to obtain the semantic feature vector of the image sample;
[0043] Perform text encoding on the long text description and the short text description to obtain the semantic feature vector of the long text description and the semantic feature vector of the short text description;
[0044] Perform global fine-grained alignment training on the initial image recognition model based on the semantic feature vector of the image sample, the semantic feature vector of the long text description, and the semantic feature vector of the short text description to obtain the global fine-grained loss value of the initial image recognition model.
[0045] Optionally, global fine-grained alignment training is performed on the initial image recognition model using the semantic feature vector of the image sample, the semantic feature vector of the long text description, and the semantic feature vector of the short text description to obtain the global fine-grained loss value of the initial image recognition model, including:
[0046] Calculate the global similarity based on the semantic feature vector of the image sample, the semantic feature vector of the long text description, and the semantic feature vector of the short text description;
[0047] Perform global fine-grained alignment training on the initial image recognition model according to the global similarity to obtain the global fine-grained loss value of the initial image recognition model.
[0048] In addition, to achieve the above object, the present application also proposes an image recognition device, the image recognition device includes:
[0049] An image input module, configured to input an image to be recognized into a preset image recognition model, where the preset image recognition model is a model obtained through regional fine-grained alignment training, and the regional fine-grained alignment training is training for aligning an image region with a regional description corresponding to the image region;
[0050] An image recognition module, configured to perform image recognition on the image to be recognized through the preset image recognition model to obtain an image recognition result of the image to be recognized.
[0051] Optionally, the image recognition device further includes:
[0052] A model training module, configured to obtain an image-text pair, where the image-text pair includes an image sample and a short text description corresponding to the image sample; perform global fine-grained alignment training on the initial image recognition model based on the short text description and the image sample to obtain the global fine-grained loss value of the initial image recognition model; obtain an image region in the image sample and a regional description corresponding to the image region, where the image region is used to represent the position of each object in the image sample; perform regional fine-grained alignment training on the initial image recognition model based on the image region and the regional description to obtain the regional fine-grained loss value of the initial image recognition model; adjust the initial image recognition model according to the global fine-grained loss value and the regional fine-grained loss value to obtain a preset image recognition model.
[0053] Optionally, the model training module is further configured to perform regional feature aggregation on the image region to obtain the regional feature of the image region; perform text encoding on the regional description to obtain the semantic feature vector of the regional description; and perform regional fine-grained alignment training on the initial image recognition model based on the regional feature of the image region and the semantic feature vector of the regional description, so as to obtain the regional fine-grained loss value of the initial image recognition model.
[0054] Optionally, the model training module is further configured to calculate the regional feature similarity between the regional feature of the image region and the semantic feature vector of the regional description; and perform regional fine-grained alignment training on the initial image recognition model based on the regional feature similarity, so as to obtain the regional fine-grained loss value of the initial image recognition model.
[0055] Optionally, the model training module is further configured to select a target image region from the image regions, where the target image region corresponds to multiple regional descriptions, and the multiple regional descriptions include the correct description of the target image region and the semantic similar descriptions corresponding to the correct description; perform high-difficulty fine-grained alignment training on the initial image recognition model based on the target image region, the correct description, and the semantic similar descriptions, so as to obtain the high-difficulty fine-grained loss value of the initial image recognition model; and adjust the initial image recognition model according to the global fine-grained loss value, the regional fine-grained loss value, and the high-difficulty fine-grained loss value, so as to obtain a preset image recognition model.
[0056] Optionally, the model training module is further configured to perform regional feature aggregation on the target image region to obtain the regional feature of the target image region; perform text encoding on the correct description and the semantic similar descriptions to obtain the semantic feature vector of the correct description and the semantic feature vector of the semantic similar descriptions; and perform regional fine-grained alignment training on the initial image recognition model based on the regional feature of the target image region, the semantic feature vector of the correct description, and the semantic feature vector of the semantic similar descriptions, so as to obtain the regional fine-grained loss value of the initial image recognition model.
[0057] In addition, to achieve the above object, the present application further provides an image recognition device, where the image recognition device includes a memory, a processor, and an image recognition program stored on the memory and executable on the processor, and the image recognition program is configured to implement the image recognition method as described above.
[0058] In addition, to achieve the above object, the present application further provides a storage medium, where an image recognition program is stored on the storage medium, and when the image recognition program is executed by a processor, the image recognition method as described above is implemented.
[0059] In addition, to achieve the above object, the present application further provides a computer program product, which includes an image recognition program. When the image recognition program is executed by a processor, the image recognition method described above is implemented.
[0060] One or more technical solutions proposed by the present application have at least the following technical effects:
[0061] In the present application, an image to be recognized is input into a preset image recognition model. Among them, the preset image recognition model is a model obtained through region fine-grained alignment training. Region fine-grained alignment training is training for aligning an image region with a region description corresponding to the image region. The preset image recognition model is used to perform image recognition on the image to be recognized to obtain an image recognition result of the image to be recognized; since the present application trains the image recognition model by finely aligning the image region with the region description corresponding to the image region, and uses the trained image recognition model to recognize the model to be recognized, the recognition accuracy of the fine target of the image is improved, and thus the image recognition effect is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0063] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0064] Figure 1 It is a schematic flowchart of the first embodiment of the image recognition method of the present application;
[0065] Figure 2 It is a schematic flowchart of the second embodiment of the image recognition method of the present application;
[0066] Figure 3 It is a schematic flowchart of the third embodiment of the image recognition method of the present application;
[0067] Figure 4 It is a schematic diagram of model training of an embodiment of the image recognition method of the present application;
[0068] Figure 5 It is a schematic module structure diagram of the image recognition device in the embodiment of the present application;
[0069] Figure 6 It is a schematic device structure diagram of the hardware operating environment involved in the image recognition method in the embodiment of the present application.
[0070] The realization, functional features, and advantages of the present application will be further described in conjunction with embodiments with reference to the accompanying drawings. Specific Embodiments
[0071] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not used to limit the present application.
[0072] To better understand the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific embodiments.
[0073] Currently, fine-grained image recognition is becoming increasingly important in image recognition applications, such as in the fields of medical imaging, wildlife monitoring, and autonomous driving. In these fields, the ability to distinguish subtle differences between objects and scenes is crucial. For example, in medical imaging, the model must be able to distinguish various types of tumors or lesions, which may have very similar visual features but require different treatment methods. Similarly, in wildlife monitoring, the model must be able to identify different species of birds or animals, which may have only subtle differences in appearance. However, when performing image recognition, the relevant image recognition models usually align the entire image with the corresponding text description, resulting in the defect that the model cannot recognize fine targets in the image and the image recognition effect is poor.
[0074] Therefore, to overcome the above defects, the present application provides a solution, which includes: inputting the image to be recognized into a preset image recognition model, where the preset image recognition model is a model obtained through region fine-grained alignment training, and the region fine-grained alignment training is training for aligning the image region with the region description corresponding to the image region, performing image recognition on the image to be recognized through the preset image recognition model to obtain the image recognition result of the image to be recognized; since the present application trains the image recognition model by finely aligning the image region with the region description corresponding to the image region and performs recognition on the model to be recognized through the trained image recognition model, the recognition accuracy of the fine targets in the image is improved, and thus the image recognition effect is improved.
[0075] It should be noted that the execution subject of this embodiment can be an image recognition device with functions of data processing, network communication, and program running, such as a computer, or other electronic devices that can achieve the same or similar functions. This embodiment is not limited thereto.
[0076] Based on this, the embodiments of the present application provide an image recognition method, referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of the image recognition method of the present application.
[0077] In the first embodiment, the image recognition method includes:
[0078] Step S100: Input the image to be recognized into a preset image recognition model, where the preset image recognition model is a model obtained through region fine-grained alignment training, and the region fine-grained alignment training is training for aligning an image region with a region description corresponding to the image region.
[0079] It should be understood that the image to be recognized can refer to an image that needs to be recognized and can be any type of picture, such as a landscape photo, a portrait photo, a medical imaging diagram, etc. The preset image recognition model can refer to a vision-language model (such as FG-CLIP (Fine-Grained Contrastive Language-Image Pre-training)) that has undergone a specific training process (such as region fine-grained alignment training). The preset image recognition model can capture the detailed information of local regions in the image and accurately match the image regions with text descriptions.
[0080] The region fine-grained alignment training can refer to a training method that aligns the local regions in the image (such as the object within the bounding box) with the corresponding detailed text description (such as "a white seagull with black spots on one wing") to enable the model to understand the semantic information of the local regions. The image region can refer to a local region segmented or detected in the image, which can be defined by a bounding box or a segmentation mask and represents an object or a scene segment in the image. For example, the "wing region of a bird" in a wildlife photo. The region description can be a detailed text annotation of the image region, describing its attributes (color, shape), functions, or context relationships. For example, "the wing has black spots".
[0081] It can be understood that the steps of the region fine-grained alignment training include: obtaining a visual localization data set, where the visual localization data set includes the image regions in the image samples and the region descriptions corresponding to the image regions; and performing region fine-grained alignment training on the initial image recognition model based on the image regions and the region descriptions to obtain the preset image recognition model.
[0082] In a specific implementation, in order to improve the fine alignment between the image and the text, this embodiment designs a high-quality visual localization data set. This data set contains 12 million pictures and 40 million detailed descriptions of the corresponding detection boxes, ensuring that each region is accurately annotated with a context-rich description. By creating such a wide-ranging and richly annotated data set, the model can learn precise and context-rich representations, thus significantly improving the performance of the model in tasks that require fine understanding.
[0083] Step S200: Perform image recognition on the image to be recognized through the preset image recognition model to obtain the image recognition result of the image to be recognized.
[0084] For ease of understanding, the following is an example, but it does not limit the present application. As an example, assume that the image recognition task is bird recognition, and the input image is an image to be recognized of a bird (such as "Black-headed Gull"). The traditional CLIP model may confuse similar birds (such as "Black-headed Gull" and "Common Tern") only through global image-text alignment because their overall appearances are similar (white feathers, black wings). The fine-grained alignment process of the FG-CLIP model in the present application is as follows: 1. Region division: The FG-CLIP model extracts key regions in the image (such as the beak, wings, feet). 2. Region-text alignment: The feature of the beak region is aligned with the text description of "red slender beak"; the feature of the wing region is aligned with the text description of "black at the end of the wing, without spots". 3. Compare the learning results: The FG-CLIP model learns the unique attributes of "Black-headed Gull" (red beak, pure black wingtips), rather than relying only on global features. 4. After inputting the image to be recognized into the trained FG-CLIP model, the output result of the FG-CLIP model is: accurately recognized as "Black-headed Gull", rather than other similar species.
[0085] In this embodiment, the image recognition model is trained by finely aligning the image region with the corresponding region description of the image region, and the image to be recognized is recognized through the trained image recognition model, thereby improving the recognition accuracy of the fine-grained target of the image and further improving the image recognition effect.
[0086] Refer to Figure 2 , Figure 2 which is a schematic flowchart of the second embodiment of the image recognition method of the present application. Based on the first embodiment shown above Figure 1 , the second embodiment of the image recognition method of the present application is proposed.
[0087] In the second embodiment, before the step S100, it further includes:
[0088] Step S10: Obtain an image-text pair, where the image-text pair includes an image sample and a short text description corresponding to the image sample.
[0089] It should be understood that in order for the preset image recognition model to capture semantic details at both the global level and the local level, in this embodiment, the preset image recognition model is obtained by performing global fine-grained alignment training and regional fine-grained alignment training on the initial image recognition model.
[0090] It is understandable that an image-text pair can refer to paired data consisting of an image sample and a corresponding short text description. For example, an image of a bird can correspond to the short text description "A bird with blue feathers stands on a branch". The short text description can refer to a concise text summary of the image content, which can include the main objects and their basic attributes (such as color, action, position, etc.).
[0091] Step S20: Based on the short text description and the image sample, perform global fine-grained alignment training on the initial image recognition model to obtain the global fine-grained loss value of the initial image recognition model.
[0092] It should be understood that the initial image recognition model can refer to a pre-trained model that has not undergone fine-grained training (such as the original CLIP model). The initial image recognition model already has basic cross-modal alignment capabilities but lacks the ability to capture fine-grained details. Global fine-grained alignment training can refer to a training method that aligns the entire image features with the global text description through contrastive learning, enabling the model to learn the fine-grained matching between the image and the text in terms of overall semantics. For example, distinguishing the global features of a "cheetah" and a "leopard".
[0093] Step S30: Obtain the image regions in the image sample and the region descriptions corresponding to the image regions, where the image regions are used to represent the positions of various objects in the image sample.
[0094] It is understandable that an image region can refer to a local region segmented or detected in an image, which can be defined by a bounding box or a segmentation mask, representing an object or a scene segment in the image. For example, the "wing region of a bird" in a wildlife photo. The region description can be a detailed text annotation of the image region, describing its attributes (color, shape), function, or context relationship. For example, "The wing has black spots".
[0095] Step S40: Based on the image regions and the region descriptions, perform region fine-grained alignment training on the initial image recognition model to obtain the region fine-grained loss value of the initial image recognition model.
[0096] It should be understood that region fine-grained alignment training can refer to a training method that aligns local regions in the image (such as the object within the bounding box) with the corresponding detailed text description (such as "A white seagull with a wing having black spots") to enable the model to understand the semantic information of the local regions.
[0097] Step S50: Adjust the initial image recognition model according to the global fine-grained loss value and the region fine-grained loss value to obtain a preset image recognition model.
[0098] It can be understood that the global fine-grained loss value can be used to measure the matching degree of the model to the global image-text pair. The smaller the global fine-grained loss value, the stronger the global alignment ability. The regional fine-grained loss value can be used to measure the loss value of the matching degree between the local region features and the region description. The smaller the regional fine-grained loss value, the stronger the model's ability to capture details.
[0099] In a specific implementation, adjusting the initial image recognition model according to the global fine-grained loss value and the regional fine-grained loss value can be to weight and sum the global fine-grained loss value and the regional fine-grained loss value, calculate the total loss value, and minimize the total loss value through gradient descent, so that the model can learn global semantics and local details simultaneously.
[0100] In this embodiment, the preset image recognition model is obtained by performing global fine-grained alignment training and regional fine-grained alignment training on the initial image recognition model, so that the preset image recognition model can capture semantic details at both the global level and the local level, and further improve the recognition accuracy of fine targets.
[0101] Refer to Figure 3 , Figure 3 which is a schematic flowchart of the third embodiment of the image recognition method of the present application. Based on the second embodiment shown above Figure 2 , the third embodiment of the image recognition method of the present application is proposed.
[0102] In the third embodiment, before the step S20, it further includes:
[0103] Step S11: Generate a corresponding long text description for the image sample through a multimodal large model.
[0104] It should be understood that, in order to enhance the semantic alignment at the global level, in this embodiment, a corresponding long text description is first generated for the image sample through a multimodal large model, and then the initial image recognition model is subjected to global fine-grained alignment training based on the long text description, the short text description, and the image sample to obtain the global fine-grained loss value of the initial image recognition model.
[0105] It can be understood that the multimodal large model (Multimodal Large Language Models, LMMs) can refer to a model that can process images and texts simultaneously and can input an image to generate a detailed long text description. The long text description can refer to the detailed text generated by the multimodal large model, including fine-grained information such as the attributes, positions, and relationships of the objects in the image.
[0106] Correspondingly, the step S20 includes:
[0107] Step S20': Based on the long text description, the short text description, and the image sample, perform global fine-grained alignment training on the initial image recognition model to obtain the global fine-grained loss value of the initial image recognition model.
[0108] In a specific implementation, a multimodal large model (LMMs) is used to generate long descriptions for images, thus greatly enhancing semantic alignment at the global level. This process introduces 1.6 billion long text-image pairs, providing an unprecedented data scale that enables the FG-CLIP model to capture nuances at the global semantic layer, thereby enhancing its ability to perceive complex and detailed information.
[0109] In this embodiment, first, a multimodal large model is used to generate a corresponding long text description for the image sample, and then based on the long text description, the short text description, and the image sample, global fine-grained alignment training is performed on the initial image recognition model to obtain the global fine-grained loss value of the initial image recognition model, so as to enhance semantic alignment at the global level and further improve the accuracy of image recognition.
[0110] Further, to improve the training effect of the global fine-grained alignment training, the step S20' includes: performing image encoding on the image sample to obtain the semantic feature vector of the image sample; performing text encoding on the long text description and the short text description to obtain the semantic feature vector of the long text description and the semantic feature vector of the short text description; and performing global fine-grained alignment training on the initial image recognition model based on the semantic feature vector of the image sample, the semantic feature vector of the long text description, and the semantic feature vector of the short text description to obtain the global fine-grained loss value of the initial image recognition model.
[0111] It should be understood that performing global fine-grained alignment training on the initial image recognition model based on the semantic feature vector of the image sample, the semantic feature vector of the long text description, and the semantic feature vector of the short text description to obtain the global fine-grained loss value of the initial image recognition model may include: calculating the global similarity based on the semantic feature vector of the image sample and the semantic feature vector of the long text description; and performing global fine-grained alignment training on the initial image recognition model according to the global similarity to obtain the global fine-grained loss value of the initial image recognition model.
[0112] For ease of understanding, reference is made to Figure 4 for illustration, but it does not limit the present application. Figure 4 This is a schematic diagram of model training for an embodiment of the image recognition method of the present application. As an example, assume the image sample is as Figure 4As shown, the short text description corresponding to the graphical sample is "Modern computer workstation with a flat-panel display and wireless keyboard and mouse". The long text description corresponding to the image sample is generated by a multi-modal large model: "The image shows a modern computer workstation. At the core is a large flat-panel display located in the center of a dark arc-shaped desk, with an abstract art wallpaper dominated by orange and blue displayed on the screen. Below the display is a set of streamlined wireless keyboard and mouse, neatly placed on the desktop. On the left side of the display is a cylindrical container with a metal border, possibly used for placing drinks". The image sample is encoded by an Image Encoder to obtain the semantic feature vector CLS of the image sample img . The long text description and the short text description are encoded by a Text Encoder to obtain the semantic feature vector of the long text description and the semantic feature vector of the short text description Based on the semantic feature vector CLS of the image sample img , the semantic feature vector of the long text description and the semantic feature vector of the short text description The initial image recognition model is trained for global fine-grained alignment through a global fine-grained loss function to obtain the global fine-grained loss value of the initial image recognition model. The global fine-grained loss function is as follows:
[0113]
[0114] Among them, L global represents the global fine-grained loss value of the initial image recognition model, N represents the total number of image-text pairs, v i represents the semantic feature vector CLS of the i-th image sample img , t i represents the semantic feature vector of the text description of the i-th image sample. The semantic feature vector of the text description of the i-th image sample includes the semantic feature vector of the long text description and the semantic feature vector of the short text description s(·) represents the cosine similarity calculation function, s(v i , t i ) represents the global similarity between the semantic feature vector CLS of the i-th image sample img and the text feature. τ represents the temperature coefficient, and τ can be set in advance. For example, τ is set in advance to 0.07
[0115] In the third embodiment, in order to improve the training effect of the regional fine-grained alignment training, step S40 includes: performing regional feature aggregation on the image region to obtain the regional feature of the image region; performing text encoding on the regional description to obtain the semantic feature vector of the regional description; and performing regional fine-grained alignment training on the initial image recognition model based on the regional feature of the image region and the semantic feature vector of the regional description to obtain the regional fine-grained loss value of the initial image recognition model.
[0116] It should be understood that performing regional fine-grained alignment training on the initial image recognition model based on the regional feature of the image region and the semantic feature vector of the regional description to obtain the regional fine-grained loss value of the initial image recognition model may include: calculating the regional feature similarity between the regional feature of the image region and the semantic feature vector of the regional description; and performing regional fine-grained alignment training on the initial image recognition model based on the regional feature similarity to obtain the regional fine-grained loss value of the initial image recognition model.
[0117] For ease of understanding, reference Figure 4 is made for illustration, but it does not limit the present application. Figure 4 FIG. is a schematic diagram of model training for an embodiment of the image recognition method of the present application. As an example, assume that the image sample is as Figure 4 shown, and the image regions in the image sample are such as Figure 4 the purple bounding box, the red bounding box, and the gold bounding box in. The regional description corresponding to the purple bounding box is "a large flat panel display or placed in the center of a dark curved table", the regional description corresponding to the red bounding box is "a fashionable wireless keyboard and mouse device", and the regional description corresponding to the gold bounding box is "a cylindrical tan mug with a metal edge, possibly containing a beverage". Perform regional feature aggregation Roi Align on the image region to obtain the regional feature of the image region and Perform text encoding TextEncoder on the regional description to obtain the semantic feature vector of the regional description Based on the regional feature of the image region and and the semantic feature vector of the regional description Perform regional fine-grained alignment training on the initial image recognition model through the regional fine-grained loss function to obtain the regional fine-grained loss value of the initial image recognition model, where the regional fine-grained loss function is as follows:
[0118]
[0119] In the formula, L regional represents the regional fine-grained loss value of the initial image recognition model, K represents the total number of image regions, r iDenote the regional feature of the i-th image region, and the regional feature of the i-th image region includes the regional feature and l i Denote the semantic feature vector of the i-th image region s(·) represents the cosine similarity calculation function, s(r i ,l i ) represents the regional feature similarity between the regional feature of the i-th image region and the semantic feature vector of the regional description. τ represents the temperature coefficient, and τ can be preset. For example, τ is preset to 0.07.
[0120] In the third embodiment, before the step S50, it further includes:
[0121] Step S41: Select a target image region from the image regions, where the target image region corresponds to multiple regional descriptions, and the multiple regional descriptions include the correct description of the target image region and the semantically similar descriptions corresponding to the correct description.
[0122] It should be understood that in order to enable the preset image recognition model to capture semantic details at both the global level and the local level, and to distinguish the subtle differences in semantic similarity pairs, in this embodiment, a high-difficulty fine-grained alignment training is also performed on the initial image recognition model.
[0123] It can be understood that the target image region can refer to a specific region selected from the image regions for further high-difficulty fine-grained alignment training. The semantically similar description can refer to a description that is semantically similar but different from the correct description of the target image region, which is used to improve the model's recognition ability of subtle differences.
[0124] Step S42: Perform high-difficulty fine-grained alignment training on the initial image recognition model based on the target image region, the correct description, and the semantically similar description to obtain the high-difficulty fine-grained loss value of the initial image recognition model.
[0125] It should be understood that a challenging target image region is selected from the image regions, and the target image region can correspond to multiple similar regional descriptions. The target image region, its correct description, and the semantically similar description are input into the initial image recognition model. By optimizing the model parameters, the initial image recognition model can accurately distinguish descriptions that are semantically similar but different, improving the robustness and discrimination ability of the model.
[0126] In a specific implementation, to further enhance the robustness and discrimination ability of the model, a large corpus containing 10 million high-difficulty fine-grained negative samples was designed. By incorporating these challenging negative samples into the training process, the FG-CLIP model learned to distinguish the subtle differences between words with similar semantics but different attributes, thus significantly improving its performance in various downstream tasks.
[0127] Furthermore, to improve the training effect of high-difficulty fine-grained alignment training, step S42 includes: performing regional feature aggregation on the target image region to obtain the regional feature of the target image region; performing text encoding on the correct description and the semantically similar description to obtain the semantic feature vector of the correct description and the semantic feature vector of the semantically similar description; performing regional fine-grained alignment training on the initial image recognition model based on the regional feature of the target image region, the semantic feature vector of the correct description, and the semantic feature vector of the semantically similar description to obtain the regional fine-grained loss value of the initial image recognition model.
[0128] It should be understood that performing regional fine-grained alignment training on the initial image recognition model based on the regional feature of the target image region, the semantic feature vector of the correct description, and the semantic feature vector of the semantically similar description to obtain the regional fine-grained loss value of the initial image recognition model may include: calculating the positive sample similarity between the regional feature of the target image region and the semantic feature vector of the correct description; calculating the negative sample similarity between the regional feature of the target image region and the semantic feature vector of the correct description; performing regional fine-grained alignment training on the initial image recognition model based on the positive sample similarity and the negative sample similarity to obtain the regional fine-grained loss value of the initial image recognition model.
[0129] For ease of understanding, reference is made to Figure 4 for illustration, but it does not limit the present application. Figure 4 FIG. is a schematic diagram of model training for an embodiment of the image recognition method of the present application. As an example, assume that the image sample is as Figure 4 shown, the target image region is as Figure 4 the red bounding box in, the correct description of the target image region is "a fashionable wireless keyboard and mouse device", and the semantically similar descriptions corresponding to the correct description are "a wooden wireless keyboard and mouse device" and "a velvet wireless keyboard and mouse device". Perform regional feature aggregation Roi Align on the target image region to obtain the regional feature of the target image region Perform text encoding Text Encoder on the correct description and the semantically similar description to obtain the semantic feature vector of the correct description and the semantic feature vector of the semantically similar description Based on the regional feature of the target image region Semantic feature vector of correct description And the semantic feature vector of semantically similar descriptions Perform region fine-grained alignment training on the initial image recognition model through the region fine-grained loss function to obtain the region fine-grained loss value of the initial image recognition model, where the region fine-grained loss function is as follows:
[0130]
[0131] In the formula, L hard Represents the region fine-grained loss value of the initial image recognition model, K represents the total number of image regions, and r i Represents the region feature of the i-th image region, that is, the region feature of the target image region Represents the semantic feature vector of the i-th correct description Represents the semantic feature vector of the i-th semantically similar description M represents the number of semantically similar descriptions, Represents the positive sample similarity between the region feature of the i-th target image region and the semantic feature vector of the correct description, Represents the negative sample similarity between the region feature of the target image region and the semantic feature vector of the correct description, τ represents the temperature coefficient, and τ can be preset. For example, τ is preset to 0.07.
[0132] Correspondingly, the step S50 includes:
[0133] Step S50': Adjust the initial image recognition model according to the global fine-grained loss value, the region fine-grained loss value, and the high-difficulty fine-grained loss value to obtain a preset image recognition model.
[0134] In a specific implementation, for these 3 types of data, this embodiment designs different training stages. In the first stage, the FG-CLIP model only uses global fine-grained alignment training to adjust the global representations of images and texts. In the second stage, region fine-grained alignment training and high-difficulty fine-grained alignment training are introduced on this basis to further improve the model's understanding of fine details using regional text data.
[0135] This embodiment also performs high-difficulty fine-grained alignment training on the initial image recognition model based on the target image region, the correct description of the target region, and the semantically similar description, so that the preset image recognition model can not only capture semantic details at the global level, but also capture semantic details at the local level, and can also distinguish the subtle differences in semantically similar pairs, thereby further improving the accuracy of image recognition.
[0136] Further, to improve the model training effect, step S50' includes: obtaining the regional loss weight value corresponding to the regional fine-grained loss value, and obtaining the high-difficulty loss weight value corresponding to the high-difficulty fine-grained loss value; calculating the total loss value of the initial image recognition model according to the global fine-grained loss value, the regional fine-grained loss value, the regional loss weight value, the high-difficulty fine-grained loss value, and the high-difficulty loss weight value; adjusting the initial image recognition model according to the total loss value to obtain a preset image recognition model.
[0137] For ease of understanding, reference is made to Figure 4 for illustration, but it does not limit this application. Figure 4 FIG. is a schematic diagram of model training for an embodiment of the image recognition method of this application. As an example, assume the image sample is as Figure 4 shown, and the total loss value of the initial image recognition model is calculated by the following formula:
[0138] L = L global + α * L regional + β * L hard
[0139] In the formula, L represents the total loss value of the initial image recognition model, L global represents the global fine-grained loss value of the initial image recognition model, α represents the regional loss weight value corresponding to the regional fine-grained loss value, which can be preset to 0.1, L regional represents the regional fine-grained loss value of the initial image recognition model, β represents the high-difficulty loss weight value corresponding to the high-difficulty fine-grained loss value, which can be preset to 0.5, L hard represents the high-difficulty fine-grained loss value of the initial image recognition model.
[0140] Compared with related methods, the FG-CLIP model has significant improvements in various benchmark tasks. The comprehensive improvement in the model training process of this embodiment enables the model to achieve excellent performance in capturing subtle visual details and has obtained good results in tasks such as "fine-grained understanding", "detection box classification", "long-description image text retrieval", "short-description image text retrieval", and "open vocabulary object detection". In addition, when the FG-CLIP model is used as the backbone of a multi-modal large model, it also shows performance improvements in tasks involving attribute analysis, object localization, and reducing output hallucinations.
[0141] It should be noted that the above examples are only for understanding this application and do not limit the image recognition method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0142] This application also provides an image recognition device. Please refer toFigure 5 , the image recognition device includes:
[0143] An image input module 10 for inputting an image to be recognized into a preset image recognition model, where the preset image recognition model is a model obtained through region fine-grained alignment training, and the region fine-grained alignment training is training for aligning an image region with a region description corresponding to the image region;
[0144] An image recognition module 20 for performing image recognition on the image to be recognized through the preset image recognition model to obtain an image recognition result of the image to be recognized.
[0145] The image recognition device provided in this application adopts the image recognition method in the above embodiment, and can solve the technical problem that in the related image recognition method, the entire image is usually aligned with the corresponding text description during image recognition, resulting in the defect that the model cannot recognize fine targets in the image and the image recognition effect is poor. Compared with the prior art, the beneficial effects of the image recognition device provided in this application are the same as those of the image recognition method provided in the above embodiment, and other technical features in the image recognition device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.
[0146] This application provides an image recognition device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image recognition method in the first embodiment above.
[0147] Next, refer to Figure 6 , which shows a schematic structural diagram of an image recognition device suitable for implementing the embodiments of this application. The image recognition device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The shown image recognition device is only an example and should not bring any limitation to the functions and usage scopes of the embodiments of this application.
[0148] Such as Figure 6As shown in the figure, the image recognition device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the ROM (Read Only Memory) 1002 or the program loaded from the storage device 1003 into the RAM (Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the image recognition device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, an LCD (Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the image recognition device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an image recognition device with various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems can be implemented or had.
[0149] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0150] The image recognition device provided by the present application adopts the image recognition method in the above embodiment, and can solve the technical problem of the defect that the related image recognition method usually aligns the whole image with the corresponding text description when performing image recognition, so that the model cannot recognize the fine targets in the image and the image recognition effect is poor. Compared with the prior art, the beneficial effects of the image recognition device provided by the present application are the same as those of the image recognition method provided by the above embodiment, and the other technical features in the image recognition device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0151] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0152] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0153] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the image recognition method in the above embodiments.
[0154] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory), or flash memory, optical fibers, CD-ROM (CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or apparatus. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0155] The above computer-readable storage medium can be included in the image recognition device; it can also exist separately and not be assembled into the image recognition device.
[0156] The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed by the image recognition device, the image recognition device is caused to execute the above image recognition method.
[0157] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a LAN (Local Area Network) or a WAN (Wide Area Network), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0158] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0159] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0160] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above image recognition method, which can solve the technical problem that related image recognition methods usually align the entire image with the corresponding text description when performing image recognition, resulting in the defect that the model cannot recognize fine targets in the image and the image recognition effect is poor. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the image recognition method provided by the above embodiment, and will not be elaborated here.
[0161] This application also provides a computer program product, including a computer program, which implements the image recognition method as described above when executed by a processor.
[0162] The computer program product provided by this application can solve the technical problem that related image recognition methods usually align the entire image with the corresponding text description when performing image recognition, resulting in the defect that the model cannot recognize fine targets in the image and the image recognition effect is poor. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as those of the image recognition method provided by the above embodiment, and will not be elaborated here.
[0163] The above are only partial embodiments of this application, and do not limit the patent scope of this application accordingly. Any equivalent structural transformation made under the technical concept of this application by using the content of the specification and drawings of this application, or directly / indirectly applied in other related technical fields, is included in the patent protection scope of this application.
[0164] This application discloses A1. An image recognition method, the image recognition method includes:
[0165] Input the image to be recognized into a preset image recognition model, wherein the preset image recognition model is a model obtained through region fine-grained alignment training, and the region fine-grained alignment training is training for aligning an image region with the region description corresponding to the image region;
[0166] Perform image recognition on the image to be recognized through the preset image recognition model to obtain the image recognition result of the image to be recognized.
[0167] A2. The image recognition method as described in A1, before inputting the image to be recognized into the preset image recognition model, further includes:
[0168] Obtain an image-text pair, wherein the image-text pair includes an image sample and a short text description corresponding to the image sample;
[0169] Perform global fine-grained alignment training on the initial image recognition model based on the short text description and the image sample to obtain the global fine-grained loss value of the initial image recognition model;
[0170] Obtain the image regions in the image sample and the region descriptions corresponding to the image regions, where the image regions are used to represent the positions of various objects in the image sample;
[0171] Perform region fine-grained alignment training on the initial image recognition model based on the image regions and the region descriptions to obtain the region fine-grained loss value of the initial image recognition model;
[0172] Adjust the initial image recognition model according to the global fine-grained loss value and the region fine-grained loss value to obtain a preset image recognition model.
[0173] A3. The image recognition method as described in A2, where the performing region fine-grained alignment training on the initial image recognition model based on the image regions and the region descriptions to obtain the region fine-grained loss value of the initial image recognition model includes:
[0174] Perform region feature aggregation on the image regions to obtain the region features of the image regions;
[0175] Perform text encoding on the region descriptions to obtain the semantic feature vectors of the region descriptions;
[0176] Perform region fine-grained alignment training on the initial image recognition model based on the region features of the image regions and the semantic feature vectors of the region descriptions to obtain the region fine-grained loss value of the initial image recognition model.
[0177] A4. The image recognition method as described in A3, where the performing region fine-grained alignment training on the initial image recognition model based on the region features of the image regions and the semantic feature vectors of the region descriptions to obtain the region fine-grained loss value of the initial image recognition model includes:
[0178] Calculate the region feature similarity between the region features of the image regions and the semantic feature vectors of the region descriptions;
[0179] Perform region fine-grained alignment training on the initial image recognition model based on the region feature similarity to obtain the region fine-grained loss value of the initial image recognition model.
[0180] A5. The image recognition method as described in A2, before adjusting the initial image recognition model according to the global fine-grained loss value and the region fine-grained loss value to obtain a preset image recognition model, further includes:
[0181] Select a target image region from the image region, where the target image region corresponds to multiple region descriptions, and the multiple region descriptions include the correct description of the target image region and the semantically similar descriptions corresponding to the correct description;
[0182] Perform high-difficulty fine-grained alignment training on the initial image recognition model based on the target image region, the correct description, and the semantically similar descriptions to obtain the high-difficulty fine-grained loss value of the initial image recognition model;
[0183] Correspondingly, adjusting the initial image recognition model according to the global fine-grained loss value and the regional fine-grained loss value to obtain a preset image recognition model includes:
[0184] Adjust the initial image recognition model according to the global fine-grained loss value, the regional fine-grained loss value, and the high-difficulty fine-grained loss value to obtain a preset image recognition model.
[0185] A6. The image recognition method as described in A5, where the performing high-difficulty fine-grained alignment training on the initial image recognition model based on the target image region, the correct description, and the semantically similar descriptions to obtain the high-difficulty fine-grained loss value of the initial image recognition model includes:
[0186] Perform regional feature aggregation on the target image region to obtain the regional features of the target image region;
[0187] Perform text encoding on the correct description and the semantically similar descriptions to obtain the semantic feature vector of the correct description and the semantic feature vector of the semantically similar descriptions;
[0188] Perform regional fine-grained alignment training on the initial image recognition model based on the regional features of the target image region, the semantic feature vector of the correct description, and the semantic feature vector of the semantically similar descriptions to obtain the regional fine-grained loss value of the initial image recognition model.
[0189] A7. The image recognition method as described in A6, where the performing regional fine-grained alignment training on the initial image recognition model based on the regional features of the target image region, the semantic feature vector of the correct description, and the semantic feature vector of the semantically similar descriptions to obtain the regional fine-grained loss value of the initial image recognition model includes:
[0190] Calculate the positive sample similarity between the regional features of the target image region and the semantic feature vector of the correct description;
[0191] Calculate the negative sample similarity between the regional features of the target image region and the semantic feature vector of the correct description;
[0192] Perform region fine-grained alignment training on the initial image recognition model based on the positive sample similarity and the negative sample similarity to obtain the region fine-grained loss value of the initial image recognition model.
[0193] A8. The image recognition method as described in A5, wherein adjusting the initial image recognition model according to the global fine-grained loss value, the region fine-grained loss value, and the high-difficulty fine-grained loss value to obtain a preset image recognition model includes:
[0194] Obtain the region loss weight value corresponding to the region fine-grained loss value, and obtain the high-difficulty loss weight value corresponding to the high-difficulty fine-grained loss value;
[0195] Calculate the total loss value of the initial image recognition model according to the global fine-grained loss value, the region fine-grained loss value, the region loss weight value, the high-difficulty fine-grained loss value, and the high-difficulty loss weight value;
[0196] Adjust the initial image recognition model according to the total loss value to obtain a preset image recognition model.
[0197] A9. The image recognition method as described in any one of A2 to A8, before performing global fine-grained alignment training on the initial image recognition model based on the short text description and the image sample to obtain the global fine-grained loss value of the initial image recognition model, further includes:
[0198] Generate a corresponding long text description for the image sample through a multimodal large model;
[0199] Correspondingly, performing global fine-grained alignment training on the initial image recognition model based on the short text description and the image sample to obtain the global fine-grained loss value of the initial image recognition model includes:
[0200] Perform global fine-grained alignment training on the initial image recognition model based on the long text description, the short text description, and the image sample to obtain the global fine-grained loss value of the initial image recognition model.
[0201] A10. The image recognition method as described in A9, wherein performing global fine-grained alignment training on the initial image recognition model based on the long text description, the short text description, and the image sample to obtain the global fine-grained loss value of the initial image recognition model includes:
[0202] Perform image encoding on the image sample to obtain the semantic feature vector of the image sample;
[0203] Perform text encoding on the long text description and the short text description to obtain the semantic feature vector of the long text description and the semantic feature vector of the short text description;
[0204] Perform global fine-grained alignment training on the initial image recognition model based on the semantic feature vector of the image sample, the semantic feature vector of the long text description, and the semantic feature vector of the short text description to obtain the global fine-grained loss value of the initial image recognition model.
[0205] A11. The image recognition method as described in A10, where the performing global fine-grained alignment training on the initial image recognition model based on the semantic feature vector of the image sample, the semantic feature vector of the long text description, and the semantic feature vector of the short text description to obtain the global fine-grained loss value of the initial image recognition model includes:
[0206] Calculate the global similarity based on the semantic feature vector of the image sample, the semantic feature vector of the long text description, and the semantic feature vector of the short text description;
[0207] Perform global fine-grained alignment training on the initial image recognition model according to the global similarity to obtain the global fine-grained loss value of the initial image recognition model.
[0208] This application also discloses B12. An image recognition device, where the image recognition device includes:
[0209] An image input module, configured to input an image to be recognized into a preset image recognition model, where the preset image recognition model is a model obtained through region fine-grained alignment training, and the region fine-grained alignment training is training for aligning an image region with the region description corresponding to the image region;
[0210] An image recognition module, configured to perform image recognition on the image to be recognized through the preset image recognition model to obtain the image recognition result of the image to be recognized.
[0211] B13. The image recognition device as described in B12, where the image recognition device further includes:
[0212] A model training module, configured to obtain image-text pairs, where the image-text pairs include image samples and short text descriptions corresponding to the image samples; perform global fine-grained alignment training on an initial image recognition model based on the short text descriptions and the image samples to obtain a global fine-grained loss value of the initial image recognition model; obtain image regions in the image samples and region descriptions corresponding to the image regions, where the image regions are used to represent the positions of various objects in the image samples; perform region fine-grained alignment training on the initial image recognition model based on the image regions and the region descriptions to obtain a region fine-grained loss value of the initial image recognition model; and adjust the initial image recognition model according to the global fine-grained loss value and the region fine-grained loss value to obtain a preset image recognition model.
[0213] B14. The image recognition device according to B13, wherein the model training module is further configured to perform region feature aggregation on the image regions to obtain region features of the image regions; perform text encoding on the region descriptions to obtain semantic feature vectors of the region descriptions; and perform region fine-grained alignment training on the initial image recognition model based on the region features of the image regions and the semantic feature vectors of the region descriptions to obtain a region fine-grained loss value of the initial image recognition model.
[0214] B15. The image recognition device according to B14, wherein the model training module is further configured to calculate a region feature similarity between the region features of the image regions and the semantic feature vectors of the region descriptions; and perform region fine-grained alignment training on the initial image recognition model based on the region feature similarity to obtain a region fine-grained loss value of the initial image recognition model.
[0215] B16. The image recognition device according to B13, wherein the model training module is further configured to select target image regions from the image regions, where the target image regions correspond to multiple region descriptions, and the multiple region descriptions include correct descriptions of the target image regions and semantic similar descriptions corresponding to the correct descriptions; perform high-difficulty fine-grained alignment training on the initial image recognition model based on the target image regions, the correct descriptions, and the semantic similar descriptions to obtain a high-difficulty fine-grained loss value of the initial image recognition model; and adjust the initial image recognition model according to the global fine-grained loss value, the region fine-grained loss value, and the high-difficulty fine-grained loss value to obtain a preset image recognition model.
[0216] B17. The image recognition device as described in B16, wherein the model training module is further configured to perform regional feature aggregation on the target image region to obtain the regional features of the target image region; perform text encoding on the correct description and the semantically similar description to obtain the semantic feature vector of the correct description and the semantic feature vector of the semantically similar description; and perform regional fine-grained alignment training on the initial image recognition model based on the regional features of the target image region, the semantic feature vector of the correct description, and the semantic feature vector of the semantically similar description to obtain the regional fine-grained loss value of the initial image recognition model.
[0217] The present application also discloses C18. An image recognition device, comprising: a memory, a processor, and an image recognition program stored on the memory and executable on the processor, where when the image recognition program is executed by the processor, it implements the image recognition method as described above.
[0218] The present application also discloses D19. A storage medium, on which an image recognition program is stored, where when the image recognition program is executed by a processor, it implements the image recognition method as described above.
[0219] The present application also discloses E20. A computer program product, comprising an image recognition program, where when the image recognition program is executed by a processor, it implements the image recognition method as described above.
Claims
1. An image recognition method, characterized in that, The described image recognition method includes: Inputting the image to be recognized into a preset image recognition model, where the preset image recognition model is a model obtained through region fine-grained alignment training, and the region fine-grained alignment training is training for aligning an image region with the region description corresponding to the image region; Performing image recognition on the image to be recognized through the preset image recognition model to obtain the image recognition result of the image to be recognized.
2. The image recognition method according to claim 1, wherein, Before inputting the image to be recognized into the preset image recognition model, it further includes: Obtaining an image-text pair, where the image-text pair includes an image sample and a short text description corresponding to the image sample; Performing global fine-grained alignment training on the initial image recognition model based on the short text description and the image sample to obtain the global fine-grained loss value of the initial image recognition model; Obtaining the image regions in the image sample and the region descriptions corresponding to the image regions, where the image regions are used to represent the positions of various objects in the image sample; Performing region fine-grained alignment training on the initial image recognition model based on the image regions and the region descriptions to obtain the region fine-grained loss value of the initial image recognition model; Adjusting the initial image recognition model according to the global fine-grained loss value and the region fine-grained loss value to obtain a preset image recognition model.
3. The image recognition method according to claim 2, wherein The performing region fine-grained alignment training on the initial image recognition model based on the image regions and the region descriptions to obtain the region fine-grained loss value of the initial image recognition model includes: Performing region feature aggregation on the image regions to obtain the region features of the image regions; Performing text encoding on the region descriptions to obtain the semantic feature vectors of the region descriptions; Performing region fine-grained alignment training on the initial image recognition model based on the region features of the image regions and the semantic feature vectors of the region descriptions to obtain the region fine-grained loss value of the initial image recognition model.
4. The image recognition method according to claim 3, wherein The performing region fine-grained alignment training on the initial image recognition model based on the region features of the image regions and the semantic feature vectors of the region descriptions to obtain the region fine-grained loss value of the initial image recognition model includes: Calculating the region feature similarity between the region features of the image regions and the semantic feature vectors of the region descriptions; Performing region fine-grained alignment training on the initial image recognition model based on the region feature similarity to obtain the region fine-grained loss value of the initial image recognition model.
5. The image recognition method according to claim 2, wherein, Before adjusting the initial image recognition model according to the global fine-grained loss value and the region fine-grained loss value to obtain a preset image recognition model, it further includes: Selecting target image regions from the image regions, where the target image regions correspond to multiple region descriptions, and the multiple region descriptions include the correct description of the target image region and the semantic similar descriptions corresponding to the correct description; Performing high-difficulty fine-grained alignment training on the initial image recognition model based on the target image region, the correct description, and the semantically similar description to obtain the high-difficulty fine-grained loss value of the initial image recognition model; Correspondingly, the adjusting the initial image recognition model according to the global fine-grained loss value and the regional fine-grained loss value to obtain a preset image recognition model includes: Adjusting the initial image recognition model according to the global fine-grained loss value, the regional fine-grained loss value, and the high-difficulty fine-grained loss value to obtain a preset image recognition model.
6. The image recognition method according to claim 5, characterized in that, The performing high-difficulty fine-grained alignment training on the initial image recognition model based on the target image region, the correct description, and the semantically similar description to obtain the high-difficulty fine-grained loss value of the initial image recognition model includes: Performing regional feature aggregation on the target image region to obtain the regional feature of the target image region; Performing text encoding on the correct description and the semantically similar description to obtain the semantic feature vector of the correct description and the semantic feature vector of the semantically similar description; Performing regional fine-grained alignment training on the initial image recognition model based on the regional feature of the target image region, the semantic feature vector of the correct description, and the semantic feature vector of the semantically similar description to obtain the regional fine-grained loss value of the initial image recognition model.
7. An image recognition device, characterized in that, The image recognition device includes: An image input module, configured to input an image to be recognized into a preset image recognition model, where the preset image recognition model is a model obtained through regional fine-grained alignment training, and the regional fine-grained alignment training is training for aligning an image region with a regional description corresponding to the image region; An image recognition module, configured to perform image recognition on the image to be recognized through the preset image recognition model to obtain an image recognition result of the image to be recognized.
8. An image recognition device, characterized in that, The image recognition device includes: a memory, a processor, and an image recognition program stored on the memory and executable on the processor. When the image recognition program is executed by the processor, the image recognition method according to any one of claims 1 to 6 is implemented.
9. A storage medium, characterized in that, An image recognition program is stored on the storage medium. When the image recognition program is executed by a processor, the image recognition method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that, The computer program product includes an image recognition program. When the image recognition program is executed by a processor, the image recognition method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Image-text matching method and device, equipment, storage medium and computer program product
CN120997624A