A Few-Shot Multimodal Fine-Grained Entity Classification Method and System for Healthcare

By adopting a multi-modal fine-grained entity classification task in the medical field, a multi-view information enhancement framework and a hierarchical cross-modal semantic alignment network are solved, the problem of multi-modal data annotation in the medical field is time-consuming and resource-consuming, the classification accuracy is improved, and accurate and reliable entity classification results are provided.

CN119917669BActive Publication Date: 2025-05-30NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510414327.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-30
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

In the medical field, acquiring high-quality multimodal data requires a lot of time and resources, resulting in a degradation in the performance of multimodal fine-grained entity classification under limited annotation data.

Method used

The multi-view information enhancement framework is adopted to enhance text information using context-aware semantic interpolation method from the text perspective, and to enhance image information using a hierarchical multi-grained visual semantic pyramid module from the image perspective, and to integrate text and image information through a hierarchical cross-modal semantic alignment network.

Benefits of technology

It improves the accuracy of fine-grained entity classification and can provide accurate and reliable entity classification results with few samples and multi-modal data, suitable for the medical field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917669B_ABST
    Figure CN119917669B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of electronic digital data processing, and in particular to a few-shot multi-modal fine-grained entity classification method and system for medical use. The method includes the following steps: constructing a few-shot training set, a few-shot validation set, and a few-shot test set, and expanding the text information to obtain an expanded training set; obtaining the text representation of sentences; obtaining the global visual representation, local object representation, and cross-modal semantic representation of images; obtaining the optimized single-modal representation; obtaining the multi-modal fusion representation; calculating the prediction probability and binary cross-entropy loss of each fine-grained entity category; performing iterative optimization training until a set number of times is reached, conducting performance verification on the few-shot validation set, and carrying out the final effect evaluation on the few-shot test set. The method and system provided by the present invention can improve the performance of fine-grained classification by effectively extracting hierarchical multi-modal features, and can provide a more accurate and reliable entity classification technology for the medical field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital data processing, and in particular, to a few-shot multi-modal fine-grained entity classification method and system for medical applications. Background Art

[0002] In the current era of rapid development of digital information, the amount of data shows an explosive growth trend, and a large amount of data exists in text form. According to statistics, the global Internet generates approximately 2.5 trillion bytes of data every day, and more than 80% of it is unstructured text. Facing this situation, information extraction technology has become the core means to extract structured knowledge from massive text. In the field of natural language processing, entity classification, as a basic task, has evolved from symbolic rule systems to statistical learning models. However, the massive text data generated in the Internet era presents highly fine-grained and specialized characteristics. Traditional methods cannot distinguish fine-grained entities in the medical field, which poses higher requirements for entity classification technology and gives rise to the research direction of fine-grained entity classification.

[0003] Fine-grained entity classification is a fundamental and important task of extracting knowledge from unstructured text. It aims to assign context-sensitive fine-grained semantic types to entities in the context. Fine-grained entity classification serves many downstream natural language processing tasks, such as relation extraction and entity linking, etc. Therefore, it is the basis for constructing knowledge graphs. In Internet media, news newspapers and periodicals, text often appears together with images, and images also contain rich semantic information, which provides additional help for the fine-grained entity classification task. Thus, multi-modal fine-grained entity classification utilizes the rich fine-grained visual cues contained in images to assist in fine-grained entity classification.

[0004] However, in real application scenarios, in the medical field, it takes a lot of time for experts to label a large amount of high-quality multi-modal data. In addition, training using such a large-scale dataset requires a considerable amount of time and computing resources. Therefore, this is a task that consumes both manpower and resources. When the high-quality labeled multi-modal data is insufficient, the performance will drop significantly. Therefore, how to perform accurate multi-modal fine-grained entity classification with limited labeled data is an urgent problem to be solved. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a few-shot multi-modal fine-grained entity classification method and system for medical applications. By performing information enhancement from the perspectives of text and image respectively, and using a hierarchical cross-modal semantic alignment network to fuse text information and image information from a hierarchical multi-granularity visual semantic pyramid module, the accuracy of fine-grained entity classification is improved.

[0006] The present invention is realized through the following technical solutions:

[0007] A few-shot multi-modal fine-grained entity classification method for healthcare, comprising the following steps:

[0008] S1: Construct a few-shot training set, a few-shot validation set, and a few-shot test set based on the healthcare FIGER dataset, and expand the text information of the texts in the few-shot training set through semantic interpolation to obtain an expanded training set;

[0009] S2: Organize each sentence in each sample in the expanded training set, the few-shot validation set, and the few-shot test set to obtain the representation of the sentence, and then input the representation of the sentence into the text feature extraction model to obtain the text representation of the sentence;

[0010] S3: Integrate the images in each sample in the expanded training set, the few-shot validation set, and the few-shot test set through a hierarchical multi-granularity visual semantic pyramid module to obtain the global visual representation, the local object representation, and the cross-modal semantic representation of the image;

[0011] S4: Optimize the text representation of the sentence, the global visual representation, the local object representation, and the cross-modal semantic representation of the image in a single-modal form using the multi-head self-attention mechanism to obtain the optimized single-modal representation;

[0012] S5: Input the optimized single-modal representation into the hierarchical cross-modal semantic alignment network for multi-modal feature fusion to obtain the text-crossed object representation, the text-crossed cross-modal semantic representation, and the text-crossed visual representation, and splice the text-crossed object representation, the text-crossed cross-modal semantic representation, and the text-crossed visual representation to obtain the multi-modal fusion representation;

[0013] S6: Input the multi-modal fusion representation into the classifier to calculate the prediction probability and binary cross-entropy loss of each fine-grained entity category;

[0014] S7: Perform iterative optimization training on the expanded training set through binary cross-entropy loss until the set number of times is reached. After determining the relevant parameters, perform performance verification on the few-shot validation set and conduct the final effect evaluation on the few-shot test set to quantify the performance of the model in the few-shot multi-modal fine-grained entity classification task.

[0015] Further, in step S1, the following method is used to construct a few-shot training set, a few-shot validation set, and a few-shot test set based on the healthcare FIGER dataset:

[0016] S11: Based on a mapping file from Wikipedia titles to unique identifiers of the knowledge graph, retrieve the unique identifiers of the knowledge graph corresponding to the entity mentions in the context of each sentence in the healthcare FIGER dataset;

[0017] S12: Based on a mapping file from knowledge graph entities to Wikidata entities, obtain the Wikidata identifier corresponding to the entity mention in the context of each sentence according to the unique identifier of the knowledge graph corresponding to the entity mention in the context of each sentence.

[0018] S13: Obtain images from the Wikidata webpage according to the Wikidata identifier corresponding to the entity mention in the context of the sentence, so as to obtain a corresponding image for the entity mention in the context of the sentence, and thus form a sample consisting of a sentence and a corresponding image.

[0019] S14: Divide the formed samples into a training set, a validation set and a test set according to a set ratio, and screen the samples in the training set. Delete the categories corresponding to the fine-grained entity classifications with the number of samples less than the set quantity, and randomly delete the samples with the number exceeding the set quantity in the remaining categories to form a few-shot training set. At the same time, delete the samples corresponding to the corresponding classifications in the validation set and the test set, and randomly delete the samples with the number exceeding the set quantity in the remaining categories in the validation set to form a few-shot validation set and a few-shot test set.

[0020] Further, the method for obtaining the text representation of the sentence in step S2 is as follows:

[0021] S21: Organize each sentence in the augmented training set, the few-shot validation set and the few-shot test set respectively according to formula (1) to obtain the representation of the sentence:

[0022] (1);

[0023] Where: represents the representation of the sentence, represents the classification token, represents the entity mention, represents the separator token, represents the text context of the sentence;

[0024] S22: Process the representation of the sentence through the tokenizer of the text feature extraction model to obtain the input token sequence, and then obtain the -th input token contextualized word representation according to formula (2) through the core architecture of the text feature extraction model, and use the contextualized word representation of the last hidden layer as the text representation of the whole sentence:

[0025] (2);

[0026] Where: represents the -th input token contextualized word representation, Denote the text feature extraction model, Denote the th input token, Denote the parameters of the text feature extraction model, Denote the dimension of the hidden state of the text feature extraction model, Denote that the dimension of the hidden state of the text feature extraction model is vector space.

[0027] Furthermore, in step S3, the following method is used to obtain the global visual representation, local object representation and cross-modal semantic representation of the image:

[0028] S31: Input the image in each sample of the augmented training set, few-shot validation set and few-shot test set into the hierarchical multi-granularity visual semantic pyramid module. The global visual encoder of the hierarchical multi-granularity visual semantic pyramid module adjusts the image to the set number of pixels, and takes the last hidden layer state of the global visual encoder as the global visual representation of the image;

[0029] S32: The local visual encoder of the hierarchical multi-granularity visual semantic pyramid module processes the image according to formula (3) to obtain the local object representation of the image:

[0030] (3);

[0031] Where: Denote the local object representation of the image, Denote the object detection model, Denote the serial number of the image, Denote the number of detected objects, Denote the feature dimension of the detected object, Denote vector space of dimension;

[0032] S33: The additional visual encoder of the hierarchical multi-granularity visual semantic pyramid module processes the image according to formula (4) to obtain the natural language text describing the image, adds a classification mark in front of the natural language text describing the image, and adds a separator mark at the end of the natural language text describing the image, and then inputs it into the text feature extraction model. The text feature extraction model processes it according to formula (5) to obtain the cross-modal semantic representation of the image:

[0033] (4);

[0034] (5);

[0035] Where: Denote the natural language text describing the image, Denote the image description generation model that constitutes the additional visual encoder, Denote the cross-modal semantic representation of the image, Denote the text feature extraction model, Denote the sequence length of the natural language text describing the image, Denote the dimension of the hidden state of the text feature extraction model, Denote A vector space of dimension Denote the classification token, Denote the separation token.

[0036] Optimized. In step S4, the text representation of the sentence, the global visual representation of the image, the local object representation, and the cross-modal semantic representation are optimized in a single-modal form using the multi-head self-attention mechanism according to Equation (6) to obtain the optimized single-modal representation:

[0037] (6);

[0038] Where: Denote the optimized single-modal representation, Denote the multi-head self-attention mechanism, Denote the query matrix of the multi-head self-attention mechanism, Denote the key matrix of the multi-head self-attention mechanism, Denote the value matrix of the multi-head self-attention mechanism, Denote the concatenation operation, Denote the Output of the Denote the number of attention heads, Denote the linear transformation weight matrix.

[0039] Furthermore, in step S5, the following method is used to input the optimized single-modal representations into the hierarchical cross-modal semantic alignment network for multi-modal feature fusion respectively to obtain the text-crossed object representation, the text-crossed cross-modal semantic representation, and the text-crossed visual representation, and the text-crossed object representation, the text-crossed cross-modal semantic representation, and the text-crossed visual representation are concatenated to obtain the multi-modal fusion representation:

[0040] S51: The hierarchical cross-modal semantic alignment network uses the global visual representation in the optimized unimodal representation as the query matrix, the text representation in the optimized unimodal representation as the key matrix and value matrix to obtain the text representation of image perception. Then it uses the text representation in the optimized unimodal representation as the query matrix, the text representation of image perception as the key matrix and value matrix to obtain the refined text representation of image perception. Next, it uses the text representation in the optimized unimodal representation as the query matrix, the global visual representation in the optimized unimodal representation as the key matrix and value matrix, and finally obtains the refined text-perceived image representation. Then, the refined text representation of image perception and the refined text-perceived image representation are passed through a gate control function to obtain the text-crossed visual representation;

[0041] S52: The hierarchical cross-modal semantic alignment network uses the local object representation in the optimized unimodal representation as the query matrix, the text representation in the optimized unimodal representation as the key matrix and value matrix to obtain the text representation of local object perception. Then it uses the text representation in the optimized unimodal representation as the query matrix, the local object representation in the optimized unimodal representation as the key matrix and value matrix to obtain the refined text representation of local object perception. Next, it uses the text representation in the optimized unimodal representation as the query matrix, the local object representation in the optimized unimodal representation as the key matrix and value matrix, and finally obtains the refined text-perceived local object representation. Then, the refined text representation of local object perception and the refined text-perceived local object representation are passed through a gate control function to obtain the text-crossed object representation;

[0042] S53: The hierarchical cross-modal semantic alignment network uses the cross-modal semantic representation in the optimized unimodal representation as the query matrix, the text representation in the optimized unimodal representation as the key matrix and value matrix to obtain the cross-modal perceived text representation. Then it uses the text representation in the optimized unimodal representation as the query matrix, the cross-modal semantic representation in the optimized unimodal representation as the key matrix and value matrix to obtain the refined cross-modal perceived text representation. Next, it uses the text representation in the optimized unimodal representation as the query matrix, the cross-modal semantic representation in the optimized unimodal representation as the key matrix and value matrix, and finally obtains the refined text-perceived cross-modal semantic representation. Then, the refined cross-modal perceived text representation and the refined text-perceived cross-modal semantic representation are passed through a gate control function to obtain the text-crossed cross-modal semantic representation;

[0043] S54: Concatenate the text-crossed object representation, the text-crossed cross-modal semantic representation, and the text-crossed visual representation to obtain the text-perceived multi-granularity cross-modal representation;

[0044] S55: Concatenate the text representation in the optimized unimodal representation and the text-perceived multi-granularity cross-modal representation to obtain the multi-modal fusion representation.

[0045] Optimized. In step S6, the following method is adopted to input the multi-modal fusion representation into the classifier to calculate the prediction probability and binary cross-entropy loss of each fine-grained entity category:

[0046] S61: The classifier calculates the prediction probability of each fine-grained entity category according to Equation (7):

[0047] (7);

[0048] Where: represents the predicted vector length type vector, represents the linear transformation matrix, represents the text representation in the optimized single-modal representation, represents the text-aware multi-granularity cross-modal representation, represents the multi-modal fusion representation, represents the th prediction probability of the fine-grained entity category, represents the total number of fine-grained entity categories in the augmented training set, represents the prediction probability of the last fine-grained entity category;

[0049] S62: Calculate the binary cross-entropy loss according to Equation (8):

[0050] (8);

[0051] Where: represents the binary cross-entropy loss, represents the th sample category label.

[0052] Optimized. In step S7, the following method is adopted to perform iterative optimization training on the augmented training set through the binary cross-entropy loss:

[0053] During the training process, the binary cross-entropy loss is monitored in real time. If the binary cross-entropy loss does not decrease for three consecutive iterations, the early stopping strategy is used to terminate the training. Otherwise, the training is terminated until the set number of iterations is reached.

[0054] Furthermore, during the performance verification on the few-shot validation set and the final effect evaluation on the few-shot test set, if the prediction probability of the th fine-grained entity category is greater than 0.5, the fine-grained entity is classified as the type corresponding to this fine-grained entity category. If there are multiple prediction probabilities of the th fine-grained entity category greater than 0.5, the fine-grained entity is classified as the types corresponding to these fine-grained entity categories. There are multiple type labels for the fine-grained entity category. If If the prediction probability of each fine-grained entity category is less than or equal to 0.5, then the type corresponding to the highest value of the prediction probability of the fine-grained entity category is selected as the final prediction type.

[0055] A few-shot multi-modal fine-grained entity classification system for medical use, which is used to execute a few-shot multi-modal fine-grained entity classification method described in any one of the above, and includes a medical-based FIGER dataset, a context-aware semantic difference module, a text feature extraction model, a hierarchical multi-granularity visual semantic pyramid module, a single-modal representation optimization module, a hierarchical cross-modal semantic alignment network, and a classifier;

[0056] The medical-based FIGER dataset is used to provide a dataset for fine-grained entity classification for the few-shot multi-modal fine-grained entity classification system for medical use;

[0057] The context-aware semantic difference module is used to expand text information for the text in the few-shot training set to obtain an expanded training set;

[0058] The text feature extraction model is used to obtain the text representation of the sentence for each sample in the expanded training set, the few-shot validation set, and the few-shot test set;

[0059] The hierarchical multi-granularity visual semantic pyramid module is used to integrate the images in each sample in the expanded training set, the few-shot validation set, and the few-shot test set to obtain the global visual representation, local object representation, and cross-modal semantic representation of the image. Among them, the global visual encoder in the hierarchical multi-granularity visual semantic pyramid module is used to obtain the global visual representation of the image, the local visual encoder is used to obtain the local object representation, and the additional visual encoder is used to obtain the cross-modal semantic representation;

[0060] The single-modal representation optimization module is used to optimize the text representation of the sentence, the global visual representation of the image, the local object representation, and the cross-modal semantic representation in a single-modal form using the multi-head self-attention mechanism to obtain the optimized single-modal representation;

[0061] The hierarchical cross-modal semantic alignment network is used to perform multi-modal feature fusion on the optimized single-modal representation to obtain a multi-modal fusion representation;

[0062] The classifier is used to calculate the prediction probability and binary cross-entropy loss of each fine-grained entity category, and classify the few-shot multi-modal fine-grained entities for medical use.

[0063] Advantages of the invention:

[0064] A few-shot multi-modal fine-grained entity classification method and system provided by the present invention has the following advantages:

[0065] The present invention innovatively proposes a multi - perspective information enhancement framework. From the text perspective, a context - aware semantic interpolation method is used to enhance text information. From the image perspective, a hierarchical multi - granularity visual semantic pyramid is used to enhance image information. Only a small number of multi - modal samples are used for training. By effectively extracting hierarchical multi - modal features, the performance of fine - grained classification can be further improved, providing a more accurate and reliable entity classification technology for the medical field, with broad prospects in practical applications and huge commercial value. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 is a schematic flowchart of the present invention.

[0067] Figure 2 is a schematic structural diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] A few - shot multi - modal fine - grained entity classification method for medical use, which includes the following steps, and its flowchart is as Figure 1 shown:

[0069] S1: Construct a few - shot training set, a few - shot validation set, and a few - shot test set based on the medical FIGER dataset, and expand the text information in the few - shot training set through semantic interpolation to obtain an expanded training set;

[0070] The medical - based FIGER is a dataset for fine - grained entity classification, with a two - level hierarchical structure derived from the public - domain knowledge graph Freebase. The training set consists of Wikipedia articles automatically annotated using the distant - supervision method, which encodes information from anchor - text links. The test set is manually annotated. The multi - modal fine - grained entity classification dataset (MFIGER) is constructed based on FIGER for multi - modal fine - grained entity classification.

[0071] Specifically, the following method can be used to construct a few - shot training set, a few - shot validation set, and a few - shot test set based on the medical FIGER dataset:

[0072] S11: Based on a mapping file from Wikipedia titles to unique identifiers (mids) of the knowledge graph (Freebase), retrieve the unique identifiers of the knowledge graph corresponding to the entity mentions in the context of each sentence in the medical - based FIGER dataset;

[0073] S12: Based on a mapping file from knowledge - graph entities to Wikidata entities, obtain the Wikidata identifiers corresponding to the entity mentions in the context of each sentence according to the unique identifiers of the knowledge graph corresponding to the entity mentions in the context of each sentence;

[0074] S13: Obtain an image from the Wikidata webpage based on the Wikidata identifier corresponding to the entity mention in the sentence context, so as to obtain a corresponding image for the entity mention in the sentence context, and thus form a sample consisting of a sentence and a corresponding image;

[0075] S14: Divide the formed samples into a training set, a validation set, and a test set according to a set ratio, and screen the samples in the training set. Delete the categories corresponding to fine-grained entity classifications with the number of samples less than the set quantity, and randomly delete the samples in the remaining categories that are more than the set quantity to form a few-shot training set. At the same time, delete the samples corresponding to the corresponding classifications in the validation set and the test set, and randomly delete the samples in the remaining categories in the validation set that are more than the set quantity to form a few-shot validation set and a few-shot test set.

[0076] Specifically, it can be divided into a training set, a validation set, and a test set according to a ratio of 7:1:2.

[0077] When screening the samples in the training set, it can be set that each category has 5 training samples. First, screen out the categories with the number of samples not less than 5 from the training set, and then only retain the samples of these categories in the training set, the validation set, and the test set. Then randomly delete the samples in the remaining categories in the training set and the validation set that are more than 5 to form a few-shot training set, a few-shot validation set, and a few-shot test set. Finally, the number of samples in each category in the few-shot training set and the few-shot validation set is only 5.

[0078] Considering that entity mentions belonging to the same category have similar text contexts, and these sentences may contain the same or similar words or phrases, the present invention proposes a context-aware semantic interpolation method to expand text information by replacing the contexts of entities belonging to the same category. Specifically, in the case where there are only 5 training samples in each category in the few-shot training set, by replacing the contexts of entity mentions belonging to the same category, the training data scale of the few-shot training set is expanded to 5 times to obtain an augmented training set, enhancing the text information of the few-shot training set.

[0079] Sample Generated augmented training sample The formula is as follows:

[0080] ;

[0081] Where: represents the th sample entity mention, represents the th sample context, is the a sample picture, indicating the th sample class label, indicating that these two samples have the same type set, which is the context of the jth sample, indicating the number of samples set for each class in the few-shot training set.

[0082] S2: For each sample in the augmented training set, few-shot validation set, and few-shot test set, organize the sentences respectively to obtain the representation of the sentences, and then input the representation of the sentences into the text feature extraction model to obtain the text representation of the sentences;

[0083] Specifically, the following method can be used to obtain the text representation of the sentences:

[0084] S21: For each sample in the augmented training set, few-shot validation set, and few-shot test set, organize the sentences respectively according to formula (1) to obtain the representation of the sentences:

[0085] (1);

[0086] Where: represents the representation of the sentence, represents the classification token, represents the entity mention, represents the separator token, represents the text context of the sentence;

[0087] and are two special token combinations in the text feature extraction model. [CLS] (Classification Token) is placed at the beginning of the input text, and its corresponding final hidden state is often used as the overall feature representation of the text for classification tasks. [SEP] (Separator Token) is used to separate different text segments, such as distinguishing two sentences when processing sentence pairs.

[0088] S22: Process the representation of the sentence through the tokenizer of the text feature extraction model to obtain the input token sequence, and then obtain the context-aware token representation of the th input token according to formula (2) through the core architecture of the text feature extraction model, and use the context-aware token representation of the last hidden layer as the text representation of the entire sentence:

[0089] (2);

[0090] Where: represents the context-aware token representation of the th input token, Denote the text feature extraction model, Denote the th input token, Denote the parameters of the text feature extraction model, Denote the dimension of the hidden state of the text feature extraction model, Denote that the dimension of the hidden state of the text feature extraction model is vector space.

[0091] After the above steps, the text information is enhanced, and the text representation is extracted using the text feature extraction model. Finally, the enhanced deep contextualized text representation is obtained, and more effective information is extracted from a small number of text samples, providing a basis for subsequent multi-modal fusion and fine-grained classification.

[0092] S3: Integrate the images in each sample of the augmented training set, few-shot validation set, and few-shot test set through the hierarchical multi-granularity visual semantic pyramid module to obtain the global visual representation, local object representation, and cross-modal semantic representation of the images;

[0093] Specifically, the following method can be used to obtain the global visual representation, local object representation, and cross-modal semantic representation of the images:

[0094] S31: Input the images in each sample of the augmented training set, few-shot validation set, and few-shot test set into the hierarchical multi-granularity visual semantic pyramid module. The global visual encoder of the hierarchical multi-granularity visual semantic pyramid module adjusts the images to the set pixels and uses the last hidden layer state of the global visual encoder as the global visual representation of the images;

[0095] The global visual encoder here is a pre-trained Vision Transformer model (ViT). The global visual representation from the Vision Transformer model (ViT) provides the global semantic features of the entire image, which helps to enhance the model's ability to grasp the comprehensive content and context information of the image.

[0096] The set pixels here are 224×224 pixels.

[0097] The global visual representation is as follows:

[0098] ;

[0099] Where: Denote the global visual representation, Denote the serial number of the image, Denote processing the content in the parentheses through the Vision Transformer model, Represent the parameters of the Vision Transformer model, represent a vector space of dimension represent the number of image tokens, represent the dimension of the hidden state of the Vision Transformer model.

[0100] S32: The local visual encoder of the hierarchical multi-granularity visual semantic pyramid module processes the image according to Equation (3) to obtain the local object representation of the image:

[0101] (3);

[0102] where: represent the local object representation of the image, represent the object detection model, represent the number of detected objects, represent the feature dimension of the detected objects, represent a vector space of dimension;

[0103] The local object representation from object detection can capture fine-grained semantic information embedded in the visual context. Since few-shot multi-modal fine-grained entity classification is a multi-class multi-label classification problem and the fine-grained types are organized into a certain hierarchical structure, it is necessary to extract not only fine-grained semantic information from the text context but also fine-grained visual cues from the visual context. The objects detected from the image provide strong visual cues for the set of fine-grained types predicted for entity mentions.

[0104] The local visual encoder here can use the object detector of the object detection model (VinVL).

[0105] The cross-modal semantic representation from image captioning provides a high-level understanding and semantic interpretation of the image content. Image captioning enables the neural network model to understand the image and generate natural language descriptions to depict the content, scene, or objects in the image. The image captioning generation model (OFA) can be used as an additional visual encoder. Specifically, the following method can be adopted to obtain the cross-modal semantic representation of the image.

[0106] S33: The additional visual encoder of the hierarchical multi-granularity visual semantic pyramid module processes the image according to Equation (4) to obtain the natural language text describing the image, adds a classification token in front of the natural language text describing the image, and adds a separator token at the end of the natural language text describing the image, and then inputs it into the text feature extraction model. The text feature extraction model processes it according to Equation (5) to obtain the cross-modal semantic representation of the image:

[0107] (4);

[0108] (5);

[0109] Wherein: represents the natural language text describing the image, represents an image description generation model that constitutes an additional visual encoder, represents the cross-modal semantic representation of the image, represents a text feature extraction model, represents the sequence length of the natural language text describing the image, represents the dimension of the hidden state of the text feature extraction model, represents a vector space of dimension, represents a classification token, represents a separator token.

[0110] The above hierarchical multi-granularity visual semantic pyramid module extracts the global semantic perception, local object relationships, and cross-modal semantic interpretation features of the image. Through a multi-level and multi-granularity feature integration mechanism, it injects more comprehensive and detailed information into the image representation, improving the accuracy and semantic richness of subsequent multi-modal information interaction.

[0111] S4: Optimize the text representation of the sentence, the global visual representation of the image, the local object representation, and the cross-modal semantic representation of the image in a single-modal form using the multi-head self-attention mechanism to obtain the optimized single-modal representation;

[0112] Specifically, the text representation of the sentence, the global visual representation of the image, the local object representation, and the cross-modal semantic representation of the image are optimized in a single-modal form using the multi-head self-attention mechanism according to Equation (6) to obtain the optimized single-modal representation:

[0113] (6);

[0114] Wherein: represents the optimized single-modal representation, which can be the optimized text representation, the optimized global visual representation, the optimized local object representation of the image, or the optimized cross-modal semantic representation of the image.

[0115] represents the multi-head self-attention mechanism, represents the query matrix of the multi-head self-attention mechanism, represents the key matrix of the multi-head self-attention mechanism, represents the value matrix of the multi-head self-attention mechanism, represents a concatenation operation, represents the linear transformation weight matrix, represents the output of the th attention head,

[0116] and represents the number of attention heads. Here, the output of the th attention head ;

[0117] where: represents the attention mechanism, represents the weight matrix for the th attention head to transform the query matrix, represents the weight matrix for the th attention head to transform the key matrix, represents the weight matrix for the th attention head to transform the value matrix.

[0118] After the above operations, the unimodal representation can be further optimized and strengthened to comprehensively capture the intra-modal semantic associations, laying a foundation for subsequent multi-modal fusion.

[0119] S5: Input the optimized unimodal representations into the hierarchical cross-modal semantic alignment network for multi-modal feature fusion respectively to obtain the text-crossed object representation, text-crossed cross-modal semantic representation, and text-crossed visual representation, and concatenate the text-crossed object representation, text-crossed cross-modal semantic representation, and text-crossed visual representation to obtain the multi-modal fusion representation;

[0120] Specifically, the following method can be used to input the optimized unimodal representations into the hierarchical cross-modal semantic alignment network for multi-modal feature fusion respectively to obtain the text-crossed object representation, text-crossed cross-modal semantic representation, and text-crossed visual representation, and concatenate the text-crossed object representation, text-crossed cross-modal semantic representation, and text-crossed visual representation to obtain the multi-modal fusion representation:

[0121] S51: The hierarchical cross-modal semantic alignment network uses the global visual representation in the optimized unimodal representation as the query matrix, the text representation in the optimized unimodal representation as the key matrix and value matrix to obtain the text representation of image perception. Then it uses the text representation in the optimized unimodal representation as the query matrix, the text representation of image perception as the key matrix and value matrix to obtain the refined text representation of image perception. Next, it uses the text representation in the optimized unimodal representation as the query matrix, the global visual representation in the optimized unimodal representation as the key matrix and value matrix, and finally obtains the refined text-perceived image representation. Then, the refined text representation of image perception and the refined text-perceived image representation are passed through a gate control function to obtain the text-crossed visual representation;

[0122] Specific control formulas are as follows:

[0123] ;

[0124] Where: Among them represents the gating signal, represents the Sigmoid activation function, represents the trainable parameter matrix used to transform the refined text representation of image perception when calculating the gating signal, , represents the trainable parameter matrix used to transform the refined text-perceived image representation when calculating the gating signal, , represents dimensional vector space, represents the refined text representation of image perception, represents the refined text-perceived image representation, represents the training parameter, , represents the trainable bias vector, represents the text-crossed visual representation, , represents dimensional vector space;

[0125] Through the operation of this step, the global semantic features of text and image can be fused.

[0126] S52: The hierarchical cross-modal semantic alignment network uses the local object representation in the optimized unimodal representation as the query matrix, and the text representation in the optimized unimodal representation as the key matrix and value matrix to obtain a text representation with local object perception. Then, it uses the text representation in the optimized unimodal representation as the query matrix, and the local object representation in the optimized unimodal representation as the key matrix and value matrix to obtain a refined text representation with local object perception. Next, it uses the text representation in the optimized unimodal representation as the query matrix, and the local object representation in the optimized unimodal representation as the key matrix and value matrix, and finally obtains a refined local object representation with text perception. Then, the refined text representation with local object perception and the refined local object representation with text perception are used to obtain a text cross-object representation through a gate control function;

[0127] The specific control formula is as shown in the control formula of the text cross-visual representation.

[0128] Through the operation of this step, the text and local object features can be fused.

[0129] S53: The hierarchical cross-modal semantic alignment network uses the cross-modal semantic representation in the optimized unimodal representation as the query matrix, and the text representation in the optimized unimodal representation as the key matrix and value matrix to obtain a cross-modal perception text representation. Then, it uses the text representation in the optimized unimodal representation as the query matrix, and the cross-modal semantic representation in the optimized unimodal representation as the key matrix and value matrix to obtain a refined cross-modal perception text representation. Next, it uses the text representation in the optimized unimodal representation as the query matrix, and the cross-modal semantic representation in the optimized unimodal representation as the key matrix and value matrix, and finally obtains a refined cross-modal semantic representation with text perception. Then, the refined cross-modal perception text representation and the refined cross-modal semantic representation with text perception are used to obtain a text cross-cross-modal semantic representation through a gate control function;

[0130] The specific control formula is as shown in the control formula of the text cross-visual representation.

[0131] Through the operation of this step, the text and cross-modal semantic features can be fused.

[0132] S54: Concatenate the text cross-object representation, the text cross-cross-modal semantic representation, and the text cross-visual representation to obtain a text perception multi-granularity cross-modal representation;

[0133] Specifically, it is as shown in the following formula:

[0134] ;

[0135] Where: represents the text perception multi-granularity cross-modal representation, represents Activation function, represents a linear transformation matrix, represents the text cross-object representation, represents the text cross-modal semantic representation, represents concatenating the text cross-visual representation, the text cross-object representation, and the text cross-modal semantic representation.

[0136] S55: Concatenate the text representation in the optimized unimodal representation with the text-aware multi-granularity cross-modal representation to obtain the multi-modal fusion representation.

[0137] The present invention uses a hierarchical cross-modal semantic alignment network to grasp the internal connection between the text modality and the visual modality. That is, a hierarchical cross-modal semantic alignment network composed of stacked cross-modal Transformers is used to perform hierarchical multi-modal feature fusion on the text representation in the optimized unimodal representation, the global visual representation in the optimized unimodal representation, the local object representation in the optimized unimodal representation, and the cross-modal semantic representation in the optimized unimodal representation respectively.

[0138] S6: Input the multi-modal fusion representation into the classifier to calculate the prediction probability of each fine-grained entity category and the binary cross-entropy loss;

[0139] Optimized, in step S6, the following method is used to input the multi-modal fusion representation into the classifier to calculate the prediction probability of each fine-grained entity category and the binary cross-entropy loss:

[0140] S61: The classifier calculates the prediction probability of each fine-grained entity category according to formula (7):

[0141] (7);

[0142] Where: represents the predicted vector length type vector, represents a linear transformation matrix, represents the text representation in the optimized unimodal representation, represents, represents the multi-modal fusion representation, represents the prediction probability of the th fine-grained entity category, represents the total number of fine-grained entity categories in the augmented training set,

[0143] S62: Calculate the binary cross-entropy loss according to formula (8):

[0144] (8);

[0145] Among them: represents the binary cross-entropy loss, represents the th sample class label.

[0146] S7: Iteratively optimize and train on the augmented training set through the binary cross-entropy loss until the set number of times is reached. After determining the relevant parameters, perform performance verification on the few-shot validation set and conduct the final effect evaluation on the few-shot test set to quantify the performance of the model in the few-shot multi-modal fine-grained entity classification task.

[0147] Specifically, the following method can be used to iteratively optimize and train on the augmented training set through the binary cross-entropy loss:

[0148] During the training process, continuously monitor the binary cross-entropy loss. If the binary cross-entropy loss does not decrease for three consecutive iterations, use the early stopping strategy to terminate the training; otherwise, terminate the training until the set number of iterations is reached.

[0149] Furthermore, during the process of performing performance verification on the few-shot validation set and conducting the final effect evaluation on the few-shot test set, if the prediction probability of the th fine-grained entity category is greater than 0.5, classify the fine-grained entity as the type corresponding to this fine-grained entity category. If there are multiple prediction probabilities of the th fine-grained entity category greater than 0.5, classify the fine-grained entity as the types corresponding to these fine-grained entity categories. There are multiple type labels for the fine-grained entity category. If the prediction probabilities of the th fine-grained entity category are all less than or equal to 0.5, select the type corresponding to the highest value of the prediction probability of the fine-grained entity category as the final prediction type.

[0150] Verify the effectiveness of the method of the present invention by performing the few-shot multi-modal fine-grained entity classification task on the few-shot validation set. The method of the present invention is compared with 5 few-shot fine-grained entity classification methods, namely Fine-tuning, Prompt-based MLM, PLET, ALIGNIE, FiveFine, 2 multi-modal named entity recognition methods UMT, MGCMT, and 1 multi-modal fine-grained entity classification method MOVCNet.

[0151] The performance comparison data between the present invention and other methods in the few-shot multi-modal fine-grained entity classification task are shown in Table 1:

[0152] Table 1

[0153]

[0154] Experiments show that: whether as a whole, or at the coarse-grained or fine-grained level, this method outperforms all single-text and multi-modal methods, achieving the optimal classification performance, effectively proving the rationality and effectiveness of the few-shot multi-modal fine-grained entity classification method with multi-view information enhancement proposed by the present invention.

[0155] A few-shot multi-modal fine-grained entity classification system for medical use, which is used to execute a few-shot multi-modal fine-grained entity classification method for medical use as described in any one of the above, and the schematic diagram of the system structure is as Figure 2 shown, which includes a medical-based FIGER dataset, a context-aware semantic difference module, a text feature extraction model (BERT), a hierarchical multi-granularity visual semantic pyramid module, a single-modal representation optimization module, a hierarchical cross-modal semantic alignment network, and a classifier;

[0156] The medical-based FIGER dataset is used to provide a dataset for fine-grained entity classification for the few-shot multi-modal fine-grained entity classification system for medical use;

[0157] The context-aware semantic difference module is used to expand the text information in the few-shot training set to obtain an expanded training set;

[0158] The text feature extraction model is used to obtain the text representation of the sentence in each sample in the expanded training set, the few-shot validation set, and the few-shot test set;

[0159] The hierarchical multi-granularity visual semantic pyramid module is used to integrate the images in each sample in the expanded training set, the few-shot validation set, and the few-shot test set to obtain the global visual representation, local object representation, and cross-modal semantic representation of the image, where the global visual encoder in the hierarchical multi-granularity visual semantic pyramid module is used to obtain the global visual representation of the image, the local visual encoder is used to obtain the local object representation, and the additional visual encoder is used to obtain the cross-modal semantic representation;

[0160] The single-modal representation optimization module is used to optimize the text representation of the sentence, the global visual representation, local object representation, and cross-modal semantic representation of the image in a single-modal form by applying the multi-head self-attention mechanism to obtain the optimized single-modal representation;

[0161] The hierarchical cross-modal semantic alignment network is used to perform multi-modal feature fusion on the optimized single-modal representation to obtain a multi-modal fusion representation;

[0162] The classifier is used to calculate the prediction probability and binary cross-entropy loss of each fine-grained entity category, and classify the few-shot multi-modal fine-grained entities for medical use.

[0163] In summary, the few-shot multi-modal fine-grained entity classification method and system for medical use provided by the present invention only uses a small number of multi-modal samples for training. By effectively extracting hierarchical multi-modal features, the performance of fine-grained classification can be further improved, providing a more accurate and reliable entity classification technology for the medical field, with broad prospects in practical applications and huge commercial value.

[0164] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A medical-oriented few-sample multimodal fine-grained entity classification method, characterized by: The steps include: S1: Based on the medical FIGER dataset, a few-shot training set, a few-shot validation set, and a few-shot test set are constructed. The text in the few-shot training set is expanded with text information through semantic interpolation to obtain an expanded training set. S2: Arrange the sentences in each sample in the expanded training set, the few-shot validation set, and the few-shot test set to obtain the representation of the sentence, and then input the representation of the sentence into the text feature extraction model to obtain the text representation of the sentence; S3: Integrate the images in each sample in the expanded training set, few-shot validation set, and few-shot test set through a hierarchical multi-granularity visual semantic pyramid module to obtain the global visual representation, local object representation, and cross-modal semantic representation of the image; S4: The text representation of the sentence, the global visual representation of the image, the local object representation and the cross-modal semantic representation are optimized in a unimodal form using a multi-head self-attention mechanism to obtain an optimized unimodal representation; S5: Input the optimized unimodal representations into the hierarchical cross-modal semantic alignment network for multimodal feature fusion to obtain text cross-object representation, text cross-modal semantic representation, and text cross-visual representation, and concatenate the text cross-object representation, text cross-modal semantic representation, and text cross-visual representation to obtain a multimodal fusion representation; S6: Input the multimodal fusion representation into the classifier to calculate the prediction probability and binary cross entropy loss of each fine-grained entity category; S7: Perform iterative optimization training on the expanded training set using binary cross entropy loss until the set number of times is reached. After determining the relevant parameters, perform performance verification on the few-shot validation set and conduct final effect evaluation on the few-shot test set to quantify the performance of the model in the few-shot multimodal fine-grained entity classification task.

2. According to claim 1, a medical-oriented few-sample multimodal fine-grained entity classification method is characterized by: In step S1, the following method is used to construct a few-sample training set, a few-sample validation set, and a few-sample test set based on the medical FIGER dataset: S11: Based on a mapping file from Wikipedia titles to knowledge graph unique identifiers, retrieve the knowledge graph unique identifier corresponding to the entity mention in each sentence context of the medical FIGER dataset; S12: Based on a mapping file from knowledge graph entities to Wikidata entities, according to the unique identifier of the knowledge graph corresponding to the entity mention in each sentence context, obtain the Wikidata identifier corresponding to the entity mention in the sentence context; S13: obtaining an image from a Wikidata webpage according to a Wikidata identifier corresponding to the entity mention in the sentence context, thereby obtaining a corresponding image for the entity mention in the sentence context, so that a sentence and a corresponding image constitute a sample; S14: Divide the constructed samples into a training set, a validation set, and a test set according to a set ratio, and screen the samples in the training set, delete the categories corresponding to the fine-grained entity classifications with fewer samples than the set number, and randomly delete the samples in the retained categories that exceed the set number to form a few-sample training set, and at the same time delete the samples corresponding to the corresponding categories in the validation set and the test set, and randomly delete the samples in the categories retained in the validation set that exceed the set number to form a few-sample validation set and a few-sample test set.

3. The medical-oriented few-sample multimodal fine-grained entity classification method according to claim 1, characterized in that: The method for obtaining the text representation of the sentence in step S2 is as follows: S21: Arrange the sentences in each sample in the expanded training set, the few-shot validation set, and the few-shot test set according to formula (1) to obtain the representation of the sentence: (1); in: Indicates the expression of the sentence, Indicates the classification mark, Indicates entity mention, Indicates a separator mark. Represents the textual context of the sentence; S22: The representation of the sentence is processed by the word segmenter of the text feature extraction model to obtain the input word sequence, and then the core architecture of the text feature extraction model is used to obtain the first The contextualized word representation of the input word unit is used, and the contextualized word representation of the last hidden layer is used as the text representation of the entire sentence: (2); in: Indicates The contextualized word representation of the input tokens, represents the text feature extraction model, Indicates Input word, Represents the parameters of the text feature extraction model, Represents the dimension of the hidden state of the text feature extraction model, The dimension representing the hidden state of the text feature extraction model is The vector space of .

4. The medical-oriented few-sample multimodal fine-grained entity classification method according to claim 1, characterized in that: In step S3, the global visual representation, local object representation and cross-modal semantic representation of the image are obtained by the following method: S31: inputting the image of each sample in the expanded training set, the few-shot validation set and the few-shot test set into the hierarchical multi-granularity visual semantic pyramid module, the global visual encoder of the hierarchical multi-granularity visual semantic pyramid module adjusts the image to the set pixel, and uses the last hidden layer state of the global visual encoder as the global visual representation of the image; S32: The local visual encoder of the hierarchical multi-granularity visual semantic pyramid module processes the image according to formula (3) to obtain the local object representation of the image: (3); in: represents the local object representation of the image, represents the target detection model, Indicates the sequence number of the image. Indicates the number of detected targets, represents the feature dimension of the detected target, express dimensional vector space; S33: The additional visual encoder of the hierarchical multi-granularity visual semantic pyramid module processes the image according to formula (4) to obtain the natural language text describing the image, adds a classification tag in front of the natural language text describing the image, adds a separation tag after the natural language text describing the image, and then inputs it into the text feature extraction model. The text feature extraction model processes according to formula (5) to obtain the cross-modal semantic representation of the image: (4); (5); in: represents the natural language text describing the image, represents the image description generation model that constitutes the additional visual encoder, Represents the cross-modal semantic representation of images, represents the text feature extraction model, represents the length of the sequence of natural language text describing the image, Represents the dimension of the hidden state of the text feature extraction model, express dimensional vector space, Indicates the classification mark, Indicates a delimiter mark.

5. The medical-oriented few-sample multimodal fine-grained entity classification method according to claim 1, characterized in that: In step S4, the text representation of the sentence, the global visual representation of the image, the local object representation, and the cross-modal semantic representation are optimized in a unimodal form using a multi-head self-attention mechanism according to formula (6) to obtain the optimized unimodal representation: (6); in: represents the optimized unimodal representation, represents the multi-head self-attention mechanism, represents the query matrix of the multi-head self-attention mechanism, represents the key matrix of the multi-head self-attention mechanism, represents the value matrix of the multi-head self-attention mechanism, Represents a splicing operation, Indicates The output of an attention head is represents the number of attention heads, Represents the linear transformation weight matrix.

6. The medical-oriented few-sample multimodal fine-grained entity classification method according to claim 1, characterized in that: In step S5, the optimized unimodal representations are respectively input into the hierarchical cross-modal semantic alignment network for multimodal feature fusion using the following method to obtain text cross-object representation, text cross-modal semantic representation, and text cross-visual representation. The text cross-object representation, text cross-modal semantic representation, and text cross-visual representation are concatenated to obtain a multimodal fusion representation: S51: The hierarchical cross-modal semantic alignment network uses the global visual representation in the optimized unimodal representation as the query matrix, and the text representation in the optimized unimodal representation as the key matrix and the value matrix to obtain the image-perceived text representation. The text representation in the optimized unimodal representation is used as the query matrix, and the image-perceived text representation is used as the key matrix and the value matrix to obtain the refined image-perceived text representation. The text representation in the optimized unimodal representation is used as the query matrix, and the global visual representation in the optimized unimodal representation is used as the key matrix and the value matrix to finally obtain the refined text-perceived image representation. The refined image-perceived text representation and the refined text-perceived image representation are then gated to obtain the text cross-visual representation. S52: The hierarchical cross-modal semantic alignment network uses the local object representation in the optimized unimodal representation as the query matrix, and the text representation in the optimized unimodal representation as the key matrix and the value matrix to obtain the local object-aware text representation, uses the text representation in the optimized unimodal representation as the query matrix, and the local object representation in the optimized unimodal representation as the key matrix and the value matrix to obtain the refined local object-aware text representation, and then uses the text representation in the optimized unimodal representation as the query matrix, and the local object representation in the optimized unimodal representation as the key matrix and the value matrix to finally obtain the refined text-aware local object representation, and then uses the refined local object-aware text representation and the refined text-aware local object representation through the gate control function to obtain the text cross-object representation; S53: The hierarchical cross-modal semantic alignment network uses the cross-modal semantic representation in the optimized unimodal representation as the query matrix, and the text representation in the optimized unimodal representation as the key matrix and the value matrix to obtain the cross-modal perceived text representation, uses the text representation in the optimized unimodal representation as the query matrix, and the cross-modal semantic representation in the optimized unimodal representation as the key matrix and the value matrix to obtain the refined cross-modal perceived text representation, and then uses the text representation in the optimized unimodal representation as the query matrix, and the cross-modal semantic representation in the optimized unimodal representation as the key matrix and the value matrix to finally obtain the refined text-aware cross-modal semantic representation, and then uses the refined cross-modal perceived text representation and the refined text-aware cross-modal semantic representation through the gate control function to obtain the text cross-modal semantic representation; S54: splicing text cross-object representation, text cross-modal semantic representation, and text cross-visual representation to obtain a multi-granular cross-modal representation of text perception; S55: Concatenate the text representation in the optimized unimodal representation with the text-aware multi-granularity cross-modal representation to obtain a multimodal fusion representation.

7. The medical-oriented few-sample multimodal fine-grained entity classification method according to claim 1, characterized in that: In step S6, the following method is used to input the multimodal fusion representation into the classifier to calculate the prediction probability and binary cross entropy loss of each fine-grained entity category: S61: The classifier calculates the predicted probability of each fine-grained entity category according to formula (7): (7); in: Indicates the predicted vector length type vector, represents the linear transformation matrix, represents the text representation in the optimized unimodal representation, Representing text-aware multi-granular cross-modal representations, represents multimodal fusion representation, Indicates The predicted probability of fine-grained entity categories, represents the total number of fine-grained entity categories in the expanded training set, Represents the predicted probability of the last fine-grained entity category; S62: Calculate the binary cross entropy loss according to formula (8): (8); in: represents the binary cross entropy loss, Indicates Sample category labels.

8. The medical-oriented few-sample multimodal fine-grained entity classification method according to claim 7, characterized in that: In step S7, the following method is used to perform iterative optimization training on the expanded training set using binary cross entropy loss: During the training process, the binary cross entropy loss is monitored in real time. If the binary cross entropy loss does not decrease after three consecutive iterations, the early stopping strategy is used to terminate the training. Otherwise, the training is terminated after the set number of iterations.

9. The medical-oriented few-sample multimodal fine-grained entity classification method according to claim 7, characterized in that: During the performance verification on the few-shot validation set and the final effect evaluation on the few-shot test set, if If the predicted probability of a fine-grained entity category is greater than 0.5, the fine-grained entity is classified as the type corresponding to the fine-grained entity category. If the predicted probabilities of multiple fine-grained entity categories are greater than 0.5, the fine-grained entities are classified into the types corresponding to these fine-grained entity categories. The fine-grained entity categories have multiple type labels. If the prediction probabilities of all fine-grained entity categories are less than or equal to 0.5, the type corresponding to the value of the fine-grained entity category with the highest prediction probability is selected as the final prediction type.

10. A medical-oriented few-sample multimodal fine-grained entity classification system, used to execute a medical-oriented few-sample multimodal fine-grained entity classification method as claimed in any one of claims 1 to 9, characterized in that: Including medical-based FIGER dataset, context-aware semantic difference module, text feature extraction model, hierarchical multi-granularity visual semantic pyramid module, single-modal representation optimization module, hierarchical cross-modal semantic alignment network and classifier; The medical-based FIGER dataset is used to provide a fine-grained entity classification dataset for a medical-oriented few-sample multimodal fine-grained entity classification system; The context-aware semantic difference module is used to expand text information for the text in the few-sample training set to obtain an expanded training set; The text feature extraction model is used to obtain text representation of sentences from each sentence in the expanded training set, the few-sample validation set and the few-sample test set; The hierarchical multi-granularity visual semantic pyramid module is used to integrate the images in each sample in the expanded training set, the few-shot validation set and the few-shot test set to obtain the global visual representation, local object representation and cross-modal semantic representation of the image, wherein the global visual encoder in the hierarchical multi-granularity visual semantic pyramid module is used to obtain the global visual representation of the image, the local visual encoder is used to obtain the local object representation, and the additional visual encoder is used to obtain the cross-modal semantic representation; The unimodal representation optimization module is used to optimize the text representation of the sentence, the global visual representation of the image, the local object representation and the cross-modal semantic representation in a unimodal form by applying a multi-head self-attention mechanism to obtain an optimized unimodal representation; The hierarchical cross-modal semantic alignment network is used to fuse the optimized single-modal representation with multi-modal features to obtain a multi-modal fusion representation; The classifier is used to calculate the prediction probability and binary cross entropy loss of each fine-grained entity category, and classify medical-oriented few-sample multimodal fine-grained entities.

Citation Information

Patent Citations

  • Multi-modal human body large model training method and system considering multi-granularity semantic alignment

    CN119625351A

  • Image generation method and apparatus based on artificial intelligence, electronic device, computer readable storage medium, and computer program product

    WO2025007653A1