Automatic otoendoscope image description system based on large language model

Through the automatic description system of endoscopic image based on large language models, the text description of endoscopic images is automatically generated using multimodal deep learning technology, which solves the problem of complex interpretation of endoscopic images, and realizes accurate description and low-cost deployment, which significantly improves the accuracy and efficiency of clinical diagnosis.

CN120147762AInactive Publication Date: 2025-06-13TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH

Patent Information

Application Number
CN202510621985.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The interpretation of endoscopic images of the ear is complex and can easily lead to misdiagnosis or misdiagnosis. The existing technology has not yet effectively solved the problem of automatic interpretation of endoscopic images.

Method used

The automatic description system of endoscopic image based on large language models is adopted, including data annotation module, model training module and model inference module, and text description of endoscopic images is automatically generated through multimodal deep learning technology.

Benefits of technology

The generation of accurate endoscopic image descriptions is achieved, which reduces the interpretation burden of doctors, reduces computational costs, supports local deployment, and shows good practicality and accuracy in actual clinical environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147762A_ABST
    Figure CN120147762A_ABST
Patent Text Reader

Abstract

The invention provides an ear endoscope image automatic description system based on a large language model, and the system comprises a data annotation module which carries out the annotation of an abnormal condition in an ear endoscope image, and converts the data of the ear endoscope image into a picture text pair data set; the model training module comprises a multi-modal large language model and is used for training the multi-modal understanding capability of images and texts of the model by using the picture texts; and the model reasoning module is used for reasoning the new ear endoscope image to obtain the text description of the ear endoscope image. According to the system, visual understanding and language generation capabilities of a large language model are combined, accurate otoendoscope image description can be automatically generated, and clinical work of otolaryngologists is effectively supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of otoscope detection devices, and more specifically, to an automatic otoscope image description system based on a large language model. Background Art

[0002] Otoscope examination is an important method for diagnosing ear diseases in otolaryngology. However, due to the complexity of the ear structure and individual differences, the accurate interpretation of otoscope images is also a challenge for experienced otologists. Currently, otoscope examination is mainly used to observe the external auditory canal and tympanic membrane, and to identify various diseases including infections, tympanic membrane perforation, cerumen obstruction, and other structural abnormalities. However, the interpretation of otoscope images not only requires rich clinical experience but also is affected by individual differences in ear structure, which leads to misdiagnosis or missed diagnosis from time to time, increasing the treatment cost and health risk of patients. Therefore, it is very necessary to develop an automatic otoscope image description system to assist doctors in preliminary diagnosis and reduce the work burden.

[0003] Currently, many studies have attempted to apply deep learning techniques to the otology field, such as the classification of pure tone audiograms, the diagnosis of otitis media, etc., but they mainly focus on the classification tasks of certain ear diseases and there is no research on the automatic interpretation of otoscope images. Summary of the Invention

[0004] In view of the problems in the background art, the present invention proposes an automatic otoscope description system based on a large language model, including: a data annotation module, which annotates abnormal conditions in otoscope images and converts otoscope image data into a picture-text pair dataset; a model training module, which includes a multimodal large language model and trains the multimodal understanding ability of the model for images and texts using the picture-text pair dataset; a model inference module, which infers a new picture-text pair dataset to obtain a text description of the otoscope image.

[0005] The present invention proposes a novel multimodal deep learning system aimed at automatically interpreting otoscope images. The system utilizes the visual processing and language generation of a large language model to automatically generate accurate descriptions of otoscope images.

[0006] The advantages of the system of the present invention include: 1. Generate accurate descriptions of otoscope images to provide accurate and rapid auxiliary diagnosis for otolaryngologists. 2. Have the advantage of low computational cost, with very low training overhead, without the need for high-end graphics processing units (GPUs), and can complete training and inference on ordinary consumer-grade graphics cards. 3. Have the ability to be locally deployed and can be applied in medical institutions with limited resources, especially in low-income and remote areas. 4. This system supports custom uploading of otoscope images and has good scalability and compatibility. Brief Description of the Drawings

[0007] To better understand the present invention, the present invention will be described in more detail by referring to the specific embodiments shown in the accompanying drawings. These drawings only depict typical embodiments of the present invention and should not be considered as limiting the protection scope of the present invention.

[0008] Figure 1 It is a schematic diagram of the system structure of the present invention.

[0009] Figure 2 It is a schematic diagram of the data annotation module of the present invention.

[0010] Figure 3 It is a technical schematic diagram of the model training and inference module of the present invention.

[0011] Figure 4 It is an example of the result display of the present invention. Specific embodiments

[0012] The following describes the embodiments of the present invention with reference to the accompanying drawings, so that those skilled in the art can better understand the present invention and be able to implement it. However, the listed embodiments are not intended to limit the present invention. Without conflict, the following embodiments and the technical features in the embodiments can be combined with each other, and the same components are represented by the same reference numerals.

[0013] In one embodiment, as Figure 1 shown, the system of the present invention includes a data annotation module, an image preprocessing module, a model training and inference module, and a result display module.

[0014] The data annotation module adopts an online interactive annotation system developed based on the Gradio framework. Through an online interactive method, annotators mark the abnormal conditions of the ear canal and eardrum, including the state of the eardrum (such as intact, perforated, retracted, etc.), the state of the external auditory canal (such as the presence of cerumen, inflammation), and other abnormalities (such as cholesteatoma, fungal infection), etc. Through the data annotation module, the data set changes from a picture data set to a picture-text pair data set (image-text pairs). That is, each otoscope picture will have a corresponding text description. For example, "The eardrum is xxx, with xxx on the surface, and white hyphae can be seen...".

[0015] Preferably, the annotation tool supports multi-person collaborative annotation to improve the annotation efficiency, accuracy, and annotation quality.

[0016] Preferably, the data annotation module has an expert consensus process. When annotating each sample, the annotation system displays the annotation results of other annotators for this sample in real time. For the annotation cases where there are differences between the current annotator and the other annotators, expert discussions and judgments are carried out to reach an agreement, so as to improve the stability and consistency of the annotation quality.

[0017] The annotation of otoscope images requires extremely high professionalism. There may be significant subjective differences among different annotators, and the quality of data annotation directly affects the training effect of the model. How to effectively coordinate the differences among annotators and ensure the consistency of annotated data is the main difficulty in the annotation process. The online annotation of the data annotation module of the present invention supports multi-person collaborative annotation, can display the results of other annotators in real time, and resolves differences through expert discussions to finally reach a consensus, thereby improving the consistency and quality of the annotated data.

[0018] Figure 2 An example is shown. In an online interactive annotation system, User A made a text description (i.e., annotation, the annotation details are not shown in the figure) of an otoscope image, and User C also made a text description (i.e., annotation, the annotation details are not shown in the figure) of an otoscope image. Other users (such as User B) can make their own annotations on the above two annotations (agree with one of the annotations, or add new annotations).

[0019] The annotated data can be exported in a standardized data format (such as csv or json files) for subsequent data analysis and model training.

[0020] The system of the present invention includes an image preprocessing module, which is used to perform standardized processing on the collected otoscope images to improve the efficiency and effect of model training.

[0021] 1) Perform size standardization on the otoscope images, uniformly adjust them to a resolution suitable for model input (such as 224×224 pixels) to reduce the computational complexity.

[0022] 2) Denoise the images to eliminate redundant noises, such as halos, overexposure, etc., to ensure the purity of visual information.

[0023] 3) Adopt data augmentation strategies, including random rotation, brightness adjustment, random cropping, etc., to increase the diversity of data, avoid overfitting during model training, and improve the generalization ability of the model.

[0024] 4) Remove the patient privacy-sensitive information (such as patient name, ID number, medical record number, etc.) embedded in the data to ensure the anonymity and security of the data.

[0025] Through these preprocessing means, the model of the system of the present invention can more effectively learn useful feature information from the images.

[0026] The system of the present invention further includes a model training module and a model inference module. The model training module and the inference module are based on the deep learning technology of a multi-modal large language model (for example, the large language model is Llama-2).

[0027] The model training module includes: a tokenizer, an embedding layer neural network, a vision encoder, a linear projection layer, and a large language model.

[0028] The tokenizer splits the text description of the otoscope image (image-text pair dataset) output by the data acquisition and annotation module into individual words or phrases. Then, the embedding layer neural network converts these words or phrases into numerical vectors (i.e., text features) for feature extraction. These numerical vectors are the representations used by the model to understand the meaning of the text and can capture the relationships and context information between words. After passing through the tokenizer and the embedding layer network, the image-text pair dataset is converted into numerical vectors (in the field of deep learning, such vectors are called features). The input data is text, and the output vector is the text feature (denoted as T).

[0029] The vision encoder is used to process the image data in the image-text pair dataset, that is, the vision encoder is used to extract the visual features of the otoscope image. In the present invention, the vision encoder uses a Vision Transformer (ViT) network to extract visual features. The specific steps include: 1) The otoscope image in the image-text pair dataset is segmented into several small image patches.

[0030] First, for the input image , where H, W, and C represent the height, width, and number of channels of the image respectively. ViT segments the image into fixed-size image patches. Assuming the size of the image patch is (preferably, P = 14 is adopted in an embodiment of the present invention), then each image patch contains pixel values. The entire image will be segmented into image patches. The segmented image patches retain the local information of the image, facilitating subsequent processing.

[0031] 2) Each image patch is flattened into a vector with a length of (in an embodiment of the present invention, it is ), the flattened image patch vector is mapped to a fixed feature dimension 𝐷 (taking 1408 in an embodiment of the present invention), enabling it to be used as the input of the encoder Transformer of ViT. This step is achieved through a linear projection layer: , where is the embedding vector of the 𝑖-th image patch, is the trainable weight of the linear projection layer, is the bias.

[0032] 3) The encoder Transformer itself does not have the ability to handle sequence order, so position encoding needs to be added to each image patch to retain spatial position information. Each image patch embedding vector plus the corresponding learnable position encoding : . The position encoding has the same dimension as the embedding vector (D) and can be learned through the model.

[0033] 4) After linear projection and position encoding, all image patch embedding vectors are concatenated into a sequence: , where is the total number of patches, and Z is the input of the encoder Transformer. After passing through multiple layers of the encoder Transformer network, global visual features are extracted. The extracted visual features are also a numerical vector, representing the important information contained in the image.

[0034] The role of the linear projection layer is to map the visual features extracted by the visual encoder to a feature space that matches the input of the language model, enabling the visual features to be fused with text features for multi-modal training and inference. This process is called "alignment". Specifically as follows: 1) After being processed by the visual encoder, the input otoscope image is transformed into a global visual feature representation, denoted as , where is the number of patches after the image is segmented, and D is the embedding dimension of the visual features. Each visual feature vector contains the local information of the corresponding image patch and the global context information fused through the encoder Transformer network. In order to enable the visual features to be used as the input of the language model, the visual features must be mapped to the embedding space of the language model, that is, the dimension is converted from D to the embedding dimension of the language model This process is achieved through the following linear transformation: , where is the aligned visual feature representation, which is consistent with the embedding dimension of the language model, is the weight matrix of the linear projection layer, which is a parameter learned during the training process. is the bias vector of the linear projection layer, which is also a trainable parameter. This linear transformation ensures that the visual features not only retain the original image information but also adapt to the input requirements of the language model.

[0035] 2), the aligned visual features are input into the large language model (Llama-2 is used in this invention) and are used for inference or generation of multimodal tasks after being fused with text features. The implementation details are as follows: is appended to the front end of the input sequence of the language model to form an extended input sequence: , where T is the feature sequence generated after tokenization and embedding of the text input (such as the user's instruction or question). provides visual information context to guide the language model to generate relevant natural language descriptions based on the image. To enable the large language model to distinguish visual information from text information, special start tokens are added before and after the visual feature to prompt the language model that the current processing is visual features.

[0036] The training process of the model training module includes two stages: the pre-training stage and the fine-tuning stage.

[0037] Pre-training stage: Pre-train on a large-scale general image-text pair dataset (such as Conceptual Captions, SBUCaptions, and LAION). The purpose of this step is to enable the model of this invention to first have a strong multi-modal understanding ability of images + text. This is because, since the currently collected ear endoscope dataset is relatively small, and the dataset scale required for training a large language model is very large (in the millions), so first use the publicly available large-scale image-text pair datasets on the network to train the model (the formats of these datasets and the labeled ear endoscope dataset of this invention are all pictures and corresponding texts). These datasets are pictures and texts of natural scenes and do not include datasets in the medical field such as inner ear endoscopes.

[0038] The model training module continuously iterates and learns to obtain the correlation between vision and text from the training set. By adjusting the model's parameters in each iteration process, it minimizes the difference between the predicted result and the actual result. In one embodiment, the training parameters are as follows: the batch size is set to 256, the AdamW optimizer is used, and the initial learning rate is and linear decay is used, and the number of training steps is 20,000 steps. In this stage, the model optimization goal is to minimize the gap between the text generated by the language model and the corresponding ground truth text. The visual features are used as part of the input sequence after linear projection to guide the language model to generate natural language output. Assume the training data is image-text pairs , where is the input image, is the text label sequence corresponding to the image , and the length is T.

[0039] Based on the conditional context (including visual features and previous text), the large language model generates the probability of the t-th word . The optimization goal is to maximize the generated probability or minimize the negative log-likelihood (i.e., cross-entropy loss). The definition of the cross-entropy loss is as follows: , where T is the total length of the target text sequence, is the t-th word of the target text sequence, is all the words before the t-th word in the target text sequence. The model calculates the probability of the language model generating under the current condition through the learning of . The smaller the loss value, the closer the text generated by the model is to the target text.

[0040] Fine-tuning stage: 1) Divide the data into a training set and a validation set (preferably, the ratio is 80% vs 20%). 2) Only fine-tune on the otoscope image dataset to enhance the model's understanding ability of otology-specific images. This way can not only retain the general knowledge obtained during the pre-training process but also optimize for a specific field. In one embodiment, the parameters during fine-tuning are as follows: the batch size is set to 8, the AdamW optimizer is used, and the initial learning rate is and linear decay is used, and the number of training steps is 400 steps. This stage also uses the validation set to evaluate the model performance. Specifically, the model uses the following two traditional natural language processing task quantization metrics BLEU and CIDEr, and an artificial evaluation metric key point coverage rate KPC dedicated to otoscope images.

[0041] The calculation formula of BLEU (Bilingual Evaluation Understudy) is as follows: , where is the exact matching ratio in the generated text , and is the weight, usually set to a uniform distribution. BP (brevity penalty) penalizes the case where the length deviation of the generated text is too large. The BLEU metric is used to evaluate the language similarity between the descriptive text generated by the model and the annotated text.

[0042] The calculation formula of CIDEr (Consensus-based Image Description Evaluation) is as follows: , where is the of the generated text, is the of the annotated text, and TF-IDF measures the importance of terms. The CIDEr metric is applicable to visual description tasks, quantifies the similarity between the generated text and the expert text, and emphasizes the semantic relevance of the description.

[0043] KPC (Key Points Coverage) is a manual evaluation metric applicable to otoscope images, and its calculation formula is as follows: , where is the total number of key points in the annotated text, is the total number of key points accurately described in the generated text. For example, if the annotated text is "The eardrum is intact and there are brown dry scabs in the ear canal" and the generated text is "There are brown dry scabs in the ear canal", then KPC is 50% (one key point is accurately described).

[0044] The accuracy of the generated text in terms of semantics and language structure is verified by the above traditional metrics BLEU and CIDEr, indicating that the model has the ability to generate high-quality descriptions. The key point coverage metric KPC through manual evaluation further assesses whether the generated text conveys the necessary clinical information in the medical scenario, and can directly reflect the effectiveness of the model in real medical applications. The model performs excellently in multiple metrics, not only generating descriptions consistent with experts, but also accurately capturing key clinical information, which fully proves the rationality and effectiveness of the system.

[0045] To reduce the demand for computing resources, only the parameters of the linear projection layer are iteratively updated during the fine-tuning stage, and the parameters of other parts are kept frozen, thus enabling the training and inference of the model to be completed on ordinary consumer-grade graphics cards.

[0046] The model training module of the present invention uses a large language model. To perform multi-modal understanding of ear endoscope images using a large language model, the problem of deep fusion of visual features and text features needs to be solved. The scale of the ear endoscope dataset is small and the features are highly specialized. Directly using a large-scale general model will result in insufficient adaptability of the model to specific otological scenarios.

[0047] Combined with the characteristics of the ear endoscope scenario, the present invention optimizes the difficulties of small dataset size and unique features in the ear endoscope dataset on the basis of the two-stage training process. For example, in the fine-tuning stage, the description generation task is mainly optimized. By adjusting the model parameters and optimizing the visual-text feature alignment, the semantic understanding of specific structures in ear endoscope images is enhanced.

[0048] The system of the present invention also includes a model inference module. The inference and training processes of the model are basically the same. The difference is that during model training, data is input into the model repeatedly, and the model updates the network parameters according to this data. The model inference module means that the model has been trained and can be deployed. New data (data not encountered during the model training stage) can be input into the model, and the model outputs the corresponding text description.

[0049] The model inference module of the present invention works as follows: After the model training is completed, the model inference module uses the trained model weights to automatically interpret new ear endoscope images. The inference process is as follows: 1) The input text prompt generates the corresponding text features through the tokenizer and the embedding layer. 2) The visual features of the ear endoscope image are extracted through the visual encoder. 3) The visual features are mapped to the text feature space through the linear projection layer for alignment. 4) The aligned visual features and text features are concatenated and then input into the large language model. 5) The large language model generates a description text based on the concatenated feature vectors, specifically describing the key pathological features in the ear endoscope image. For example, the model output may be "A red-purple area, accompanied by granulation tissue and purulent secretions, can be seen on the surface of the tympanic membrane."

[0050] The results of the entire inference process can be directly presented to the doctor through the result display module to assist in clinical diagnosis and decision-making.

[0051] The model inference module of the present invention employs a large language model. The inference of large speech models usually requires powerful computing power, so the system is mostly deployed in the cloud. However, the privacy protection of medical data has always been a key issue. Without relying on the cloud, achieving efficient model inference and deployment while ensuring patient privacy and meeting actual computing power requirements is a technical challenge.

[0052] By optimizing the training process and model architecture, the present invention significantly reduces the computational cost of the system, thus enabling low-cost local deployment. The system can run on in-hospital servers or ordinary computers without relying on high-performance GPUs. This design not only avoids the privacy risk of uploading sensitive data to the cloud but also ensures that the system can still operate efficiently in resource-constrained environments, providing a safe and reliable solution for medical institutions.

[0053] The system of the present invention further includes: a result display module. The result display module is a web-based interactive dialogue system that can read pictures from the database and automatically output a detailed description of the otoscope image, including the conditions of the ear canal and eardrum, such as whether there is infection, perforation, effusion, or other abnormalities. The generated description can be used as a reference for doctors in clinical diagnosis and treatment decisions. At the same time, the result display module has an input end that can receive user-customized uploaded otoscope images to increase scalability and compatibility.

[0054] The system of the present invention supports local deployment. Users can run the model on in-hospital servers or ordinary computers, avoiding uploading sensitive data to the cloud, thereby ensuring patient privacy. In another embodiment, the system of the present invention further includes a data acquisition module. The data acquisition module consists of a handheld otoscope and supporting data acquisition software. The handheld otoscope is used to obtain high-quality otoscope images, and these image data cover typical symptoms of various ear diseases and are widely representative. During the data acquisition process, the images collected by the otoscope are stored in the database in real time through the supporting data acquisition software to ensure the integrity and consistency of the data. In addition, a unified operation specification is strictly followed during the acquisition process, including the angle of the otoscope, light intensity, imaging distance, etc., to ensure the high quality and standardization of the images to meet the subsequent training requirements. The data acquisition module refers to existing otoscope data collection methods (such as Livingstone19, Zeng22).

[0055] The feasibility of the present invention has been verified through multiple groups of experiments. In terms of commonly used image description quantification metrics (such as BLEU, CIDEr) and evaluation metrics such as key point coverage based on manual evaluation by real doctors, the present invention shows better performance compared to other existing methods, especially achieving a significant improvement in clinical relevance. The model has been preliminarily applied in the otolaryngology department of a tertiary hospital, demonstrating its good practicality and accuracy in the actual clinical environment. For example, for Figure 2 the endoscope image of the inner ear, a description was automatically generated by the system of the present invention, such as Figure 4 shown.

[0056] The above-described embodiments are only relatively preferred specific embodiments of the present invention. The present specification uses the phrases "in one embodiment", "in another embodiment", "in yet another embodiment", or "in other embodiments", which may each refer to one or more of the same or different embodiments according to the present disclosure. Ordinary variations and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included in the protection scope of the present invention.

Claims

1. An automatic description system for ear endoscope images based on a large language model, characterized in that: include: A data annotation module, which annotates abnormalities in the ear endoscope image and converts the ear endoscope image data into a picture-text pair dataset; A model training module, which includes a multimodal large language model, uses the image-text pair dataset to train the model's multimodal understanding capabilities of images and texts; The model reasoning module performs reasoning on new ear endoscope images to obtain a text description of the ear endoscope images.

2. The system according to claim 1, characterized in that The data annotation module annotates abnormalities of the ear canal and the eardrum in the ear endoscope image in an online interactive manner.

3. The system according to claim 1, characterized in that It also includes an image preprocessing module, which is located between the data annotation module and the model training module. The image preprocessing module performs the following operations: 1) Standardize the size of otoendoscopic images to reduce computational complexity; 2) De-noising the image; 3) Adopt data enhancement strategies to increase data diversity; 4) Remove patient privacy-sensitive information embedded in the data to ensure the anonymity and security of the data.

4. The system according to claim 1, characterized in that The model training module includes: Word segmenter: It splits the image-text pair dataset output by the data collection and annotation module into words or phrases; Embedding layer neural network: it converts the words or phrases into text features; Visual encoder: extracting visual features of the ear endoscope image from the image data in the image-text pair dataset; Linear projection layer: which maps the visual features extracted by the visual encoder to a feature space that matches the language model input; and Large language model: The text features and the visual features are integrated to perform multimodal training and reasoning to obtain multimodal understanding capabilities of images and texts.

5. The system according to claim 4, characterized in that The visual encoder extracts visual features by the following operations: 1) Segmenting the ear endoscope image in the image-text pair dataset into a plurality of image blocks; 2) Flatten each image block, and the flattened image block is mapped to a fixed feature dimension; 3) Add position encoding to each image block to retain spatial position information; 4) After linear projection and position encoding, all image block embedding vectors are concatenated into a sequence.

6. The system according to claim 4, characterized in that The linear projection layer is mapped by the following linear transformation: , in, is the visual feature representation of the otoendoscopic image, is the number of image blocks after the image is segmented, H and W represent the height and width of the image respectively, represents the image patch size, D is the embedding dimension of the visual feature, is the aligned visual feature representation, which is consistent with the embedding dimension of the language model. is the weight matrix of the linear projection layer, which is the parameter learned during training.

7. The system according to claim 1, characterized in that The model training module performs pre-training operations: Trained on a large-scale general image-text dataset; The model parameters are adjusted during iteration to minimize the difference between the predicted results and the actual results: by learning the image-text pairs, the probability of the language model generating the target text sequence under the current conditions is calculated. The smaller the loss value, the closer the text generated by the model is to the target text.

8. The system according to claim 7, characterized in that The model training module performs fine-tuning operations, including: Split the data into training and validation sets; Only on the otoendoscopic image dataset, the parameters of the linear projection layer were fine-tuned to enhance the model's understanding of otolaryngology-specific images. The adjusted indicators included: BLEU, which was used to evaluate the language similarity between the descriptive text generated by the model and the annotated text, CIDEr, which was used to quantify the similarity between the generated text and the expert text, and the key point coverage KPC, a manual evaluation indicator for otoendoscopic images.

9. The system according to claim 4, characterized in that The model reasoning module performs the following operations: 1) The input text prompt is passed through the word segmenter and embedding layer to generate corresponding text features; 2) Extract visual features from otoendoscopic images through a visual encoder; 3) Map the visual features to the text feature space through a linear projection layer and align them with the text feature space; 4) The aligned visual features and text features are concatenated and input into the large language model; 5) The large language model generates a text based on the concatenated feature vectors to describe the key pathological features in the ear endoscopy image.

10. The system according to claim 6, characterized in that Special start markers are added before and after the visual feature to prompt the large language model that the current processing is a visual feature.

Citation Information

Patent Citations

  • Endoscope image report generation method based on review network and storage medium

    CN115954078A

  • Medical image report generation method based on visual priori and cross-modal alignment network

    CN117393098A

  • Cervical panoramic image few-sample classification method based on visual guidance and language prompt

    CN118230052A

  • Traffic event detection system based on multi-modal data

    CN119293551A

Cited By

  • Intelligent otology image recognition method based on multi-modal artificial intelligence

    CN122156757A