Recognition method based on mixed prompt, electronic equipment and readable storage medium

By obtaining the pairs to be recognized for the text and images to be recognized, determining the auxiliary prompts for the text and images, and performing multiple rounds of fusion feature generation, the problem of insufficient target recognition accuracy in the prior art is solved, and a higher accuracy recognition result is achieved.

CN120448858AActive Publication Date: 2025-08-08HANGZHOU HUACHENG SOFTWARE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510346442.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-08-08
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

In the prior art, the image and text-based object recognition method has insufficient accuracy, especially in scenarios with low data quality.

Method used

By obtaining the pairs to be recognized for the text and images to be recognized, text assistive prompts, image assistive prompts and fusion assistive prompts are determined, text fusion features are obtained, and text and image features are deeply fused using multiple rounds of fusion technology to generate target fusion features to improve recognition accuracy.

Benefits of technology

It improves the accuracy of target recognition, reduces the dependence on data quality, and achieves more accurate picture and text matching results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448858A_ABST
    Figure CN120448858A_ABST
Patent Text Reader

Abstract

The invention discloses an identification method based on mixed prompt, electronic equipment and a readable storage medium. The method comprises the following steps: obtaining a to-be-identified pair; the pair to be recognized comprises a text to be recognized and an image to be recognized; determining a text auxiliary prompt matched with the to-be-recognized text, an image auxiliary prompt matched with the to-be-recognized image and a fusion auxiliary prompt matched with the to-be-recognized pair, and obtaining text features of the to-be-recognized text and text fusion features obtained after the text auxiliary prompt, the image auxiliary prompt and the fusion auxiliary prompt are fused; performing multiple rounds of fusion on the text fusion feature and the image feature of the to-be-recognized image by using text auxiliary prompt, image auxiliary prompt and fusion auxiliary prompt to obtain a target fusion feature; and based on the target fusion feature, determining a target recognition result matched with the to-be-recognized pair. In this way, the precision of target recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a recognition method based on hybrid prompts, an electronic device, and a readable storage medium. Background Art

[0002] With the development of object recognition technology, multimodal data-based target recognition solutions have gained wider application. Existing technologies typically extract features from both image and text data and then match them to determine the recognition result. However, the accuracy of the recognition result still depends on the quality of the data, resulting in low recognition accuracy in some scenarios. Therefore, improving the accuracy of target recognition has become an urgent issue. Summary of the Invention

[0003] The main technical problem solved by this application is to provide a recognition method based on hybrid prompts, an electronic device and a readable storage medium, which can improve the accuracy of target recognition.

[0004] In order to solve the above technical problems, the first aspect of the present application provides a recognition method based on hybrid prompts, including: obtaining a pair to be recognized; the pair to be recognized includes a text to be recognized and an image to be recognized; determining a text auxiliary prompt that matches the text to be recognized, an image auxiliary prompt that matches the image to be recognized, and a fusion auxiliary prompt that matches the pair to be recognized, and obtaining a text fusion feature obtained by fusing the text features of the text to be recognized with the text auxiliary prompts, the image auxiliary prompts and the fusion auxiliary prompts; using the text auxiliary prompts, the image auxiliary prompts and the fusion auxiliary prompts, the text fusion feature and the image feature of the image to be recognized are fused for multiple rounds to obtain a target fusion feature; based on the target fusion feature, the target recognition result of the pair to be recognized is determined.

[0005] To solve the above technical problems, the second aspect of the present application provides an electronic device, which includes: a memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method described in the first aspect above.

[0006] In order to solve the above technical problems, the third aspect of the present application provides a computer-readable storage medium on which program data is stored. When the program data is executed by a processor, the method described in the first aspect is implemented.

[0007] The above scheme obtains a pair to be recognized, including text to be recognized and an image to be recognized, thereby generating multimodal data. Based on the text to be recognized, a text-related auxiliary prompt associated with the text is determined. Based on the image to be recognized, an image-related auxiliary prompt associated with the image is determined. Based on the pair to be recognized, a fusion-related auxiliary prompt associated with the text and image is determined. A text fusion feature is obtained by fusing the text features of the text to be recognized with the text-related auxiliary prompt, image-related auxiliary prompt, and fusion-related auxiliary prompt. This allows for a preliminary interactive fusion of features between the image and text within the text fusion feature. Using the text-related auxiliary prompt, image-related auxiliary prompt, and fusion-related auxiliary prompt, the text fusion feature is fused with the image features of the image to be recognized in multiple rounds. Using the three prompts as a medium, the image features are fused with the text fusion feature in multiple rounds, generating a target fusion feature that is a deep fusion of features corresponding to the two modalities. Based on the target fusion feature, the target in the pair to be recognized is identified, and a target recognition result is obtained based on the more accurate image-text matching result provided by the target fusion feature, thereby reducing reliance on data quality and improving the accuracy of the target recognition result. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. Among them:

[0009] Figure 1 This is a flow chart of an embodiment of the recognition method based on hybrid prompts of the present application;

[0010] Figure 2 This is a flow chart of another embodiment of the recognition method based on mixed prompts of the present application;

[0011] Figure 3 This is a schematic diagram of an application scenario of an implementation method of obtaining text fusion features in this application;

[0012] Figure 4 This is a schematic diagram of an application scenario of an implementation method of obtaining target fusion features in this application;

[0013] Figure 5 This is a schematic structural diagram of an embodiment of the electronic device of the present application;

[0014] Figure 6 It is a structural diagram of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0015] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them, and different implementation methods can be adaptively combined. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0016] The terms "system" and "network" are often used interchangeably in this document. The term "and / or" is simply a description of an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " generally indicates that the related objects are in an "or" relationship. Furthermore, "multiple" in this document means two or more than two.

[0017] The hybrid prompt-based recognition method provided in this application is used for target recognition, and its corresponding execution subject is a processing unit capable of performing data processing.

[0018] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of the recognition method based on mixed prompts of the present application, which includes:

[0019] S101: Obtain a pair to be recognized; the pair to be recognized includes a text to be recognized and an image to be recognized.

[0020] Specifically, a pair to be recognized including a text to be recognized and an image to be recognized is obtained, thereby obtaining multimodal data.

[0021] It can be understood that the text to be recognized and the image to be recognized in the pair to be recognized match each other, wherein the image to be recognized is collected from the target scene and the text to be recognized matches the target scene, or the image to be recognized is obtained from a preset data source and the text to be recognized matches the preset data source.

[0022] Optionally, the text to be recognized includes at least a description text describing the image to be recognized. In some implementation scenarios, the text to be recognized also includes a preset text, which is used to indicate the recognition of the target in the image.

[0023] S102: Determine the text auxiliary prompts for the text to be recognized, the image auxiliary prompts for the image to be recognized, and the fusion auxiliary prompts for the pair to be recognized, and obtain the text fusion features obtained by fusing the text features of the text to be recognized with the text auxiliary prompts, the image auxiliary prompts, and the fusion auxiliary prompts.

[0024] Specifically, based on the text to be recognized, text auxiliary prompts associated with the text are determined; based on the image to be recognized, image auxiliary prompts associated with the image are determined; based on the pair to be recognized, fusion auxiliary prompts associated with the text and image are determined; and text fusion features are obtained by fusing the text features of the text to be recognized with the text auxiliary prompts, image auxiliary prompts and fusion auxiliary prompts, thereby enabling preliminary feature interaction fusion of the image and text in the text fusion features.

[0025] It should be noted that there are three types of prompts, namely text-assisted prompts, image-assisted prompts and fusion-assisted prompts. Among them, text-assisted prompts are strongly associated with the text features corresponding to the text to be recognized, image-assisted prompts are strongly associated with the image features corresponding to the image to be recognized, and fusion-assisted prompts are strongly associated with the fusion features corresponding to the text features and image features. The text fusion features are obtained by fusing the text features with the text-assisted prompts, image-assisted prompts and fusion-assisted prompts.

[0026] In some implementation scenarios, feature extraction is performed on the text to be recognized and the image to be recognized in the recognition pair respectively to obtain text features corresponding to the text to be recognized and image features corresponding to the image to be recognized; based on the text features, the global characters in the text to be recognized are determined and text auxiliary prompts associated with the global characters are generated, wherein the global characters include at least conjunctions and prepositions; based on the image features, the image to be recognized is divided into multiple image areas and image auxiliary prompts associated with at least part of the image areas are generated; the text features are fused with the image features to obtain image-text fusion features; based on the image-text fusion features, the image-text matching information is determined and a fusion auxiliary prompt associated with the image-text matching information is generated; the text features are fused with the text auxiliary prompts, the image auxiliary prompts and the fusion auxiliary prompts to obtain text fusion features.

[0027] In some implementation scenarios, feature extraction is performed on the text to be recognized and the image to be recognized in the recognition pair respectively to obtain text features corresponding to the text to be recognized and image features corresponding to the image to be recognized; based on the text features, the target text used to describe the target and the non-target text other than the target text in the text to be recognized are determined, and text auxiliary prompts matching the target text and the non-target text are generated; based on the image features, the image to be recognized is divided into a background area and a foreground area, and image auxiliary prompts matching the background area and the foreground area are generated; the text features are fused with the image features to obtain text-image fusion features; based on the text-image fusion features, the associated characters associated with the image information are determined and fusion auxiliary prompts corresponding to the associated characters are generated; the text features are fused with the text auxiliary prompts, the image auxiliary prompts and the fusion auxiliary prompts to obtain text fusion features.

[0028] Optionally, the current text fusion feature is updated to the text feature, and multiple rounds of construction are performed on the text auxiliary prompts, image auxiliary prompts and fusion auxiliary prompts to obtain the final round of text fusion features, text auxiliary prompts, image auxiliary prompts and fusion auxiliary prompts, thereby improving the accuracy of each type of prompt and ensuring the accuracy of the preliminary feature interaction fusion of images and texts in the text fusion feature.

[0029] S103: Using text auxiliary prompts, image auxiliary prompts and fusion auxiliary prompts, the text fusion features and the image features of the image to be recognized are fused for multiple rounds to obtain target fusion features.

[0030] Specifically, text-assisted prompts, image-assisted prompts and fusion-assisted prompts are used to perform multiple rounds of fusion of text fusion features and image features of the image to be identified. Thus, using the three prompts as a medium, the image features and text fusion features are fused multiple rounds to obtain the target fusion features after deep fusion of features corresponding to the data of the two modalities.

[0031] In some implementation scenarios, the text fusion features of the current round are fused with the text auxiliary prompts, image auxiliary prompts, and fusion auxiliary prompts to obtain the text fusion features of the next round. The image features of the current round are spliced with the text auxiliary prompts, image auxiliary prompts, and fusion auxiliary prompts and feature fusion is performed. The fused features are decomposed into the text auxiliary prompts, image auxiliary prompts, fusion auxiliary prompts, and image features of the next round. After multiple rounds of iteration, the text fusion features of the final round are fused with the image features to obtain the target fusion features. In this case, the image features are spliced with the text auxiliary prompts, image auxiliary prompts, and fusion auxiliary prompts and then cross-fused locally and globally. The dimensionality of the fusion remains unchanged and is decomposed into the text auxiliary prompts, image auxiliary prompts, fusion auxiliary prompts, and image features of the next round in the order of splicing.

[0032] In some implementation scenarios, the text-assisted prompts, image-assisted prompts, and fused-assisted prompts of the current round are fused and spliced to obtain spliced prompt features. The spliced prompt features are fused with the text features of the current round, and the fused features are adjusted to the same dimension as the text fusion features to obtain the text fusion features of the next round. The image features of the current round are fused with the spliced prompt features and decomposed into the image features, text-assisted prompts, image-assisted prompts, and fused-assisted prompts of the next round. After multiple rounds of iteration, the text fusion features of the final round are used as the target fusion features. The spliced prompt features are adjusted to the same position as the image features and spliced with the image features before feature fusion. The fused features are decomposed by channel into the text-assisted prompts, image-assisted prompts, fused-assisted prompts, and image features of the next round.

[0033] S104: Determine the target recognition result of the target pair to be recognized based on the target fusion feature.

[0034] Specifically, based on the target fusion features, the target in the identification pair is identified, so as to obtain the target recognition result by reducing the dependence on data quality based on the more accurate image-text matching results fed back by the target fusion features.

[0035] It can be understood that the target fusion feature is obtained by deeply fusing the features of the text and the image. Based on the target fusion feature, the target included in the to-be-recognized pair can be identified to obtain the target recognition result.

[0036] In some implementation scenarios, based on the target fusion features, the target category and target coordinates of the target matching included in the to-be-recognized pair are determined, and the target category and target coordinates are used as the target recognition result.

[0037] In some implementations, based on the target fusion features, a target box in the image to be recognized corresponding to the target included in the pair to be recognized is determined. A target identifier that matches the target category is set for the target, and the target box in the image to be recognized and the target identifier that matches the target are used as the target recognition result.

[0038] Optionally, the above method is implemented based on a target recognition model, which includes an image encoder, a text encoder, an early fusion module, a multimodal fusion module, and a prediction module. The image encoder is used to extract image features, the text encoder is used to extract text features, the early fusion module is used to generate text fusion features, text auxiliary prompts, image auxiliary prompts, and fusion auxiliary prompts, the multimodal fusion module is used to generate target fusion features, and the prediction module is used to output target recognition results. Furthermore, the image encoder and text encoder are pre-trained with fixed parameters, and the early fusion module, multimodal fusion module, and prediction module are supervised for training.

[0039] Specifically, the parameters of the pre-trained image encoder and text encoder remain unchanged. During the supervised training process, the parameters of the early fusion module, the multimodal fusion module and the prediction module are adjusted. Among them, the early fusion module enables the image and text to produce preliminary feature interaction fusion. The multimodal fusion module uses three types of prompts as a medium to achieve deep fusion of image features and text features. Finally, the target recognition result is output through the prediction module, realizing end-to-end output from the pair to be identified to the target recognition result, thereby improving the efficiency and accuracy of recognition.

[0040] The above scheme obtains a pair to be recognized, including text to be recognized and an image to be recognized, thereby generating multimodal data. Based on the text to be recognized, a text-related auxiliary prompt associated with the text is determined. Based on the image to be recognized, an image-related auxiliary prompt associated with the image is determined. Based on the pair to be recognized, a fusion-related auxiliary prompt associated with the text and image is determined. A text fusion feature is obtained by fusing the text features of the text to be recognized with the text-related auxiliary prompt, image-related auxiliary prompt, and fusion-related auxiliary prompt. This allows for a preliminary interactive fusion of features between the image and text within the text fusion feature. Using the text-related auxiliary prompt, image-related auxiliary prompt, and fusion-related auxiliary prompt, the text fusion feature is fused with the image features of the image to be recognized in multiple rounds. Using the three prompts as a medium, the image features are fused with the text fusion feature in multiple rounds, generating a target fusion feature that is a deep fusion of features corresponding to the two modalities. Based on the target fusion feature, the target in the pair to be recognized is identified, and a target recognition result is obtained based on the more accurate image-text matching result provided by the target fusion feature, thereby reducing reliance on data quality and improving the accuracy of the target recognition result.

[0041] See also Figure 2 , Figure 2 : is a flow chart of another embodiment of the recognition method based on mixed prompts of the present application, the method comprising:

[0042] S201: Obtain a pair to be recognized; the pair to be recognized includes a text to be recognized and an image to be recognized.

[0043] Specifically, a pair of to-be-recognized text and to-be-recognized images that match each other is obtained, thereby obtaining multimodal data.

[0044] S202: Acquire text features of the text to be recognized and image features of the image to be recognized.

[0045] Specifically, feature recognition is performed on the text to be recognized to obtain text features, and feature recognition is performed on the image to be recognized to obtain image features.

[0046] Optionally, the text to be recognized includes a description text and a preset text. The text string corresponding to the description text is first converted into a token sequence according to the ID corresponding to the token, and then the token is vectorized and output as a text embedding vector sequence, which is then spliced with the target token vector corresponding to the preset text. The spliced text embedding vector sequence is sent to the text encoder to obtain the text features corresponding to the text to be recognized. The image to be recognized is divided into multiple image blocks, which are then converted into image embedding vectors through a linear layer. The image embedding vectors are sent to the image encoder to obtain the image features corresponding to the image to be recognized. Among them, the text encoder can use the BERT (Bidirectional Encoder Representations from Transformers) model or other pre-trained models, and the image encoder can use the ViT (Visual Transformer) model or other pre-trained models.

[0047] S203: Based on the text features, a text auxiliary prompt is obtained; based on the image features, an image auxiliary prompt is obtained; based on the text features and the image features, a fusion auxiliary prompt is obtained.

[0048] Specifically, based on text features, the information included in the text is analyzed to obtain text auxiliary prompts, and based on image features, the information included in the image is analyzed to obtain image auxiliary prompts.

[0049] Furthermore, the text features are fused with the image features, and based on the fused features, the text and image information is analyzed to obtain fusion auxiliary prompts, so that the text and image information can achieve preliminary interaction.

[0050] It should be noted that, based on text features and image features, fusion auxiliary prompts are obtained, including: fusing text features and image features to obtain image-text fusion features, and based on the image-text fusion features, determining the image-text correlation between the characters in the text to be recognized and the image area in the image to be recognized; based on the image-text correlation, setting character weights for each character in the text to be recognized, and obtaining fusion auxiliary prompts based on the characters in the text to be recognized and their corresponding character weights.

[0051] Specifically, the text features and image features are fused to obtain the text-image fusion features. Based on the text-image fusion features, the text-image correlation between the characters in the text to be recognized and the image area in the image to be recognized is determined. Based on the features obtained after the text-image interaction, with the help of information obtained from the image to be recognized, words with a higher probability of change are obtained from the text to be recognized, that is, words with a higher correlation with the target in the image are obtained, and then the text-image correlation between the characters and the image area is determined.

[0052] Furthermore, based on the correlation between image and text, a character weight is set for each character in the text to be recognized, so that a higher weight is set for characters that are more correlated with the image. Based on the characters in the text to be recognized and their corresponding character weights, a fusion auxiliary prompt is obtained, thereby locating words that are more correlated with the target in the image.

[0053] S204: Fusing the text feature with the text auxiliary prompt, the image auxiliary prompt, and the fusion auxiliary prompt, and adjusting them to the same dimension as the text feature to obtain a text fusion feature.

[0054] Specifically, the text features are fused with text auxiliary prompts, image auxiliary prompts and fusion auxiliary prompts to achieve preliminary interaction between image and text features, and the dimensions of the fused features are adjusted to obtain text fusion features and ensure that the dimensions of the text fusion features are consistent with those of the text features.

[0055] Optionally, after the text features are fused with text auxiliary prompts, image auxiliary prompts and fusion auxiliary prompts and adjusted to the same dimension as the text features to obtain the text fusion features, it also includes: updating the current text fusion features to text features; returning to the steps of obtaining text auxiliary prompts based on text features, obtaining image auxiliary prompts based on image features, and obtaining fusion auxiliary prompts based on text features and image features, until the preset conditions are met and the final text fusion features, text auxiliary prompts, image auxiliary prompts and fusion auxiliary prompts are obtained.

[0056] Specifically, the text fusion features of the current round are used as the text features of the next round, so that the text fusion features, text auxiliary prompts, image auxiliary prompts and fusion auxiliary prompts are extracted and generated in multiple rounds to improve the accuracy of text fusion features and the three types of prompts.

[0057] In some implementation scenarios, text fusion features, text auxiliary hints, image auxiliary hints, and fusion auxiliary hints are generated by the early fusion module. Figure 3 , Figure 3 This is a schematic diagram of an application scenario of an implementation method of obtaining text fusion features in this application. The early fusion module includes a linear layer, a splicing layer and a probability layer, and the early fusion module also includes a routing embedding vector and a dynamic prompt as parameters to be adjusted.

[0058] Specifically, text-assisted prompts l is the length of the prompt, d t It is the text feature dimension. Text-assisted prompts are mainly responsible for global prompts shared by all sentences, not from the specific prompt features of a certain sentence. Image-assisted prompts are derived from image embedding features through a linear layer, and the output is P vImage-assisted prompts fuse feature information from images to assist prompt learning and are a supplement to information learning.

[0059] Furthermore, the embedded features of the text and image after passing through the encoder are passed through two linear layers respectively, and then spliced together in the embedding dimension. The spliced result is expressed as Among them, d v is the feature dimension of the image embedding vector. During the training phase, the routing embedding vector is expressed as After the dot product of f and r, the probability layer outputs the weight of the characters in the text to be recognized, and finally the learnable dynamic prompt Multiply to get the fusion auxiliary prompt P f Fusion-assisted prompts are dynamic prompts within learning prompts. With the help of information obtained from the image content, words with a high probability of change are learned and words with a high correlation with the target in the image are determined.

[0060] S205: Using text auxiliary prompts, image auxiliary prompts and fusion auxiliary prompts, the text fusion features and the image features of the image to be recognized are fused for multiple rounds to obtain target fusion features.

[0061] Specifically, text-assisted prompts, image-assisted prompts and fusion-assisted prompts are used to perform multiple rounds of fusion of text fusion features and image features of the image to be identified. Thus, using the three prompts as a medium, the image features and text fusion features are fused multiple rounds to obtain the target fusion features after deep fusion of features corresponding to the data of the two modalities.

[0062] It should be noted that, using text-assisted prompts, image-assisted prompts and fusion-assisted prompts, the text fusion features and the image features of the image to be identified are fused for multiple rounds to obtain target fusion features, including: splicing the text-assisted prompts, image-assisted prompts and fusion-assisted prompts of the current round to obtain splicing prompt features; fusing the text fusion features of the current round with the splicing prompt features, and adjusting them to the same dimension as the text fusion features to obtain the text fusion features of the next round; fusing the image features of the current round with the splicing prompt features, and decomposing them into the image features, text-assisted prompts, image-assisted prompts and fusion-assisted prompts of the next round; returning to the step of splicing the text-assisted prompts, image-assisted prompts and fusion-assisted prompts of the current round to obtain the splicing prompt features, until a preset number of rounds are iterated, and the text fusion features of the final round are used as the target fusion features.

[0063] Specifically, the text-assisted prompts, image-assisted prompts, and fusion-assisted prompts of the current round are spliced together to obtain spliced prompt features, thereby integrating multiple types of prompts.

[0064] Furthermore, the text fusion features of the current round are fused with the splicing prompt features and adjusted to the same dimension as the text fusion features to obtain the text fusion features of the next round, thereby ensuring that the dimensions of the text fusion features of each round are consistent and the text fusion features can be continuously fused with the splicing prompt features in each round. The image features of the current round are fused with the splicing prompt features, and the fused features are decomposed to obtain the image features, text auxiliary prompts, image auxiliary prompts and fusion auxiliary prompts of the next round, thereby fusing the image features into various types of prompts. Using various types of prompts as a medium, a preset number of rounds are iterated to achieve multiple rounds of fusion of features corresponding to the data of the two modalities of image and text until the final round. The text fusion features of the final round are used as the target fusion features to improve the accuracy of the target fusion features.

[0065] It should be noted that the image features of the current round are fused with the splicing prompt features and decomposed into the image features, text auxiliary prompts, image auxiliary prompts and fusion auxiliary prompts of the next round, including: feature fusion of the splicing prompt features and adjusting them to the same dimension as the image features to obtain fusion prompt features; feature fusion of the fusion prompt features and the image features of the current round are spliced together, and decomposed into the image features, text auxiliary prompts, image auxiliary prompts and fusion auxiliary prompts of the next round.

[0066] Specifically, the spliced features are fused and dimensionally adjusted to obtain fused prompt features of the same dimension as the image features, facilitating their fusion. The fused prompt features are then evaluated with the image features of the current round and fused, clarifying the dimensions corresponding to the fused features. This allows the fused features to be decomposed by dimension to obtain the next round of image features, text-assisted prompts, image-assisted prompts, and fused prompts, ensuring the fusion effect of image features and various prompt types, as well as the accuracy of the decomposition.

[0067] In some implementation scenarios, the target fusion feature is generated by a multimodal fusion module. Figure 4 , Figure 4 This is a schematic diagram of an application scenario of an embodiment of obtaining target fusion features in this application. The multimodal fusion module includes multiple Transformer layers and linear layers, wherein: Figure 4 In the paper, only one round of feature transformation process is adaptively given.

[0068] Specifically, obtain the text fusion feature T0 and text auxiliary prompt P output by the early fusion module s , Fusion Auxiliary Prompt P f and image-assisted prompts P vThe text fusion feature T0 is concatenated with the splicing prompt features corresponding to the three types of prompts in the sequence length dimension, and then sent to the transformer layer for fusion. After adjusting the dimension, the next round of text fusion feature T1 is obtained. The splicing prompt feature is fused through the transformer layer to obtain the fusion prompt feature. The fusion prompt feature is adjusted through the linear layer to obtain the new text auxiliary prompt P s ′ , Fusion Auxiliary Prompt P f ′ and image-assisted prompts P v ′ , the output dimension of the linear layer is the image feature dimension.

[0069] Furthermore, the new prompt P s ′ 、P f ′ and P v ′ After being concatenated with the image feature V0 in the sequence dimension, it is sent to the transformer layer to decompose the image feature V1 of the next round and the text auxiliary prompt P of the next round obtained after adjustment by the linear layer. s1 , Fusion Auxiliary Prompt P f1 and image-assisted prompts P v1 Similarly, the newly generated text fusion feature T1 and the concatenated prompt features corresponding to the three types of prompts are concatenated in the sequence length dimension and sent to the next layer, iterating a preset number of rounds.

[0070] S206: Determine the target recognition result of the target pair to be recognized based on the target fusion feature.

[0071] Specifically, based on the target fusion features, the target in the identification pair is identified, so as to obtain the target recognition result based on the more accurate image-text matching result fed back by the target fusion features.

[0072] In some implementation scenarios, based on the target fusion features, the target recognition results, classifications and regions of the matching pairs to be identified are determined, including: based on the target fusion features, determining the target category and target position that match the targets included in the pairs to be identified; based on the target position, generating a target frame in the image to be identified, and obtaining a target recognition result including the target category and target frame.

[0073] Specifically, based on the target fusion features, the category and position of the target included in the to-be-identified pair are analyzed to obtain the target category and target position matched by the target. Based on the target position, the corresponding target frame is generated in the to-be-identified image to obtain the target recognition result including the target category and target frame, making the target recognition result more comprehensive and intuitive.

[0074] Optionally, the target recognition result is output by the prediction module. After the target fusion feature and the image to be recognized are sent to the prediction module, the prediction module outputs the predicted target category classhead and the target position Bboxhead on the image to be recognized.

[0075] It can be understood that the above method is implemented based on the target recognition model, which includes image encoding, text encoder, early fusion module, multimodal fusion module and prediction module; wherein, the image encoder is used to extract image features, the text encoder is used to extract text features, the early fusion module is used to generate text fusion features, text auxiliary prompts, image auxiliary prompts and fusion auxiliary prompts, the multimodal fusion module is used to generate target fusion features, and the prediction module is used to output target recognition results; and, the image encoder and the text encoder have fixed parameters after pre-training, and the early fusion module, the multimodal fusion module and the prediction module are supervised.

[0076] Optionally, the entire model is trained end-to-end using the RefCOCO, RefCOCO+, and RefCOCOg datasets, without requiring any special data preprocessing. The input image size can be set to 640×640, and the text sequence length can be set to 40. If the length is less than the set sequence length, empty tokens are padded after the end token to maintain the same input length for each batch. The padded tokens are represented by a mask. In different specific implementation scenarios, the image size and text sequence length can be customized, and this application does not impose specific restrictions on this.

[0077] See also Figure 5 , Figure 5 This is a structural diagram of an embodiment of an electronic device of the present application. The electronic device 30 includes a memory 301 and a processor 302 coupled to each other, wherein the memory 301 stores program data (not shown in the figure), and the processor 302 calls the program data to implement the method in any of the above embodiments. For an explanation of the relevant content, please refer to the detailed description of the above method embodiments, which will not be repeated here.

[0078] See also Figure 6 , Figure 6 This is a structural diagram of an embodiment of a computer-readable storage medium of the present application. The computer-readable storage medium 40 stores program data 400. When the program data 400 is executed by the processor, the method in any of the above embodiments is implemented. For an explanation of the relevant content, please refer to the detailed description of the above method embodiments, which will not be repeated here.

[0079] It should be noted that the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of this embodiment.

[0080] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0081] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0082] The above description is merely an implementation method of the present application and does not limit the scope of protection of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the scope of protection of the present application.

Claims

1. A recognition method based on mixed prompts, characterized in that: The method comprises: Obtaining a pair to be recognized; the pair to be recognized includes a text to be recognized and an image to be recognized; Determine a text auxiliary prompt for the text to be recognized, an image auxiliary prompt for the image to be recognized, and a fusion auxiliary prompt for the image to be recognized, and obtain a text fusion feature obtained by fusing the text feature of the text to be recognized with the text auxiliary prompt, the image auxiliary prompt, and the fusion auxiliary prompt; Using the text auxiliary prompt, the image auxiliary prompt and the fusion auxiliary prompt, the text fusion feature and the image feature of the image to be recognized are fused for multiple rounds to obtain a target fusion feature; Based on the target fusion features, a target recognition result of the to-be-recognized matching pair is determined.

2. The recognition method based on hybrid prompts according to claim 1, characterized in that The determining of the text auxiliary prompt matching the to-be-recognized text, the image auxiliary prompt matching the to-be-recognized image, and the fusion auxiliary prompt matching the to-be-recognized image, and obtaining a text fusion feature obtained by fusing the text feature of the to-be-recognized text with the text auxiliary prompt, the image auxiliary prompt, and the fusion auxiliary prompt, includes: Acquiring text features of the text to be recognized and image features of the image to be recognized; Based on the text features, the text auxiliary prompt is obtained; based on the image features, the image auxiliary prompt is obtained; based on the text features and the image features, the fusion auxiliary prompt is obtained; The text feature is fused with the text auxiliary prompt, the image auxiliary prompt and the fusion auxiliary prompt, and adjusted to the same dimension as the text feature to obtain the text fusion feature.

3. The hybrid prompt-based recognition method according to claim 2, characterized in that: The obtaining of the fusion auxiliary prompt based on the text feature and the image feature includes: fusing the text feature and the image feature to obtain a text-image fusion feature, and determining a text-image relevance between characters in the to-be-recognized text and an image region in the to-be-recognized image based on the text-image fusion feature; Based on the image-text relevance, a character weight is set for each character in the text to be recognized, and based on the characters in the text to be recognized and their corresponding character weights, the fusion auxiliary prompt is obtained.

4. The hybrid prompt-based recognition method according to claim 2, characterized in that: After fusing the text feature with the text auxiliary prompt, the image auxiliary prompt, and the fusion auxiliary prompt and adjusting them to have the same dimension as the text feature to obtain the text fusion feature, the method further includes: Updating the current text fusion feature to the text feature; Return to the steps of obtaining the text auxiliary prompt based on the text feature, obtaining the image auxiliary prompt based on the image feature, and obtaining the fusion auxiliary prompt based on the text feature and the image feature, until the preset conditions are met to obtain the final text fusion feature, the text auxiliary prompt, the image auxiliary prompt and the fusion auxiliary prompt.

5. The hybrid prompt-based recognition method according to claim 1, characterized in that: The method of using the text auxiliary prompt, the image auxiliary prompt, and the fusion auxiliary prompt to perform multiple rounds of fusing the text fusion feature and the image feature of the image to be recognized to obtain a target fusion feature includes: splicing the text auxiliary prompt, the image auxiliary prompt and the fusion auxiliary prompt of the current round to obtain a splicing prompt feature; Fusing the text fusion feature of the current round with the splicing prompt feature and adjusting it to the same dimension as the text fusion feature to obtain the text fusion feature of the next round; Fusing the image features of the current round with the splicing prompt features, and decomposing them into the image features, the text auxiliary prompts, the image auxiliary prompts, and the fusion auxiliary prompts of the next round; Return to the step of splicing the text auxiliary prompt, the image auxiliary prompt and the fusion auxiliary prompt of the current round to obtain the splicing prompt feature, until a preset number of rounds are iterated, and the text fusion feature of the final round is used as the target fusion feature.

6. The hybrid prompt-based recognition method according to claim 5, characterized in that: The fusing of the image features of the current round with the splicing prompt features and decomposing the features into the image features, the text auxiliary prompts, the image auxiliary prompts and the fusion auxiliary prompts of the next round includes: Performing feature fusion on the splicing prompt feature and adjusting it to the same dimension as the image feature to obtain a fused prompt feature; The fusion prompt feature is spliced with the image feature of the current round to perform feature fusion, and is decomposed into the image feature of the next round, the text auxiliary prompt, the image auxiliary prompt and the fusion auxiliary prompt.

7. The hybrid prompt-based recognition method according to claim 5, characterized in that: Determining the target recognition result, classification and region of the target to be recognized based on the target fusion feature includes: Determining, based on the target fusion features, target categories and target positions that match the targets included in the to-be-identified pairs; A target frame is generated in the image to be recognized based on the target position, and the target recognition result including the target category and the target frame is obtained.

8. The hybrid prompt-based recognition method according to any one of claims 1 to 7, characterized in that: The method is implemented based on a target recognition model, which includes an image encoder, a text encoder, an early fusion module, a multimodal fusion module and a prediction module; wherein the image encoder is used to extract the image features, the text encoder is used to extract the text features, the early fusion module is used to generate the text fusion features, the text auxiliary prompts, the image auxiliary prompts and the fusion auxiliary prompts, the multimodal fusion module is used to generate the target fusion features, and the prediction module is used to output the target recognition result; and the image encoder and the text encoder have fixed parameters after pre-training, and the early fusion module, the multimodal fusion module and the prediction module are supervised.

9. An electronic device, characterized in that: include: A memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having program data stored thereon, characterized in that: When the program data is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Document identification method and device based on multiple modes, equipment and storage medium

    CN115131801A

  • Image processing method and device, storage medium and electronic equipment

    CN117540221A

  • Image annotation processing method and device, electronic equipment and medium

    CN119600384A