Image multimodal hierarchical labeling method and semi-automatic image data labeling method

By employing a multimodal hierarchical image annotation method, the problem of insufficient description of visual features of lesions in the annotation of digestive endoscopy images was solved, which improved the robustness and interpretability of the model, enhanced the accuracy of lesion segmentation and diagnosis, and reduced the hallucination rate of the model.

CN121686459BActive Publication Date: 2026-04-14SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
Filing Date
2026-02-11
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing methods for annotating digestive endoscopy images lack detailed descriptions of the visual features of lesions, which leads to the risk of hallucinations in the trained artificial intelligence models. These models are unable to perform a complete process of visual perception, feature description, logical reasoning, and final diagnosis. Furthermore, the annotated data confuses clinical diagnosis with visual manifestations, affecting the interpretability and accuracy of the models.

Method used

We employ a multimodal hierarchical image annotation method. By establishing image-level annotations, independently annotating key elements, decoupling visual information, and cleaning text, we fill in the logical deduction process, use polygons or rectangles to annotate lesion areas, and remove non-visual descriptions, thus constructing high-precision and highly logical training data.

Benefits of technology

It improves the robustness and interpretability of the model, reduces background noise interference, enhances the model's ability to understand complex endoscopic scenes, reduces the model's hallucination rate, and improves the accuracy of lesion segmentation and diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121686459B_ABST
    Figure CN121686459B_ABST
Patent Text Reader

Abstract

The present application relates to image multimodal layered labeling method and semi-automatic image data labeling method, wherein the image multimodal layered labeling method comprises the following steps: establishing picture level labeling, establishing at least one picture level bounding box for each effective picture to cover the whole effective picture, and filling in the overall description of the effective picture, if the current effective picture contains key elements, indicating that the key elements are contained; using the labeling box to independently label each key element in the effective picture; decoupling the visual information of the overall description of the effective picture and cleaning the text, removing the non-visual description, extracting the visual description and filling in, and extracting the subject noun in the visual description and filling in; and filling in the logical process from the visual description to the diagnostic conclusion. Eliminate non-visual description, keep pure visual description, help to eliminate model illusion, improve model robustness. By filling in the logical process from the visual description to the diagnostic conclusion, the model interpretability is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of technology, and in particular to a multimodal image layered annotation method and a semi-automatic image data annotation method. Background Technology

[0002] With the rapid development of deep learning technology, artificial intelligence has been widely applied in the field of gastrointestinal endoscopy-assisted diagnosis. The training and optimization of high-performance, large-scale multimodal models are highly dependent on massive amounts of high-quality medical image annotation data. Traditional gastrointestinal endoscopy image annotation methods only focus on "what it is"—determining the target category (classification task) or "where it is"—identifying the target location (detection task).

[0003] Existing annotation methods often only provide simple classification labels (such as "polyp") or rough bounding boxes, lacking detailed descriptions of the visual features of lesions (such as shape, color, and structure), and often confuse clinical diagnosis (which needs to be combined with medical history) with visual manifestations. This leads to the risk of "hallucination" in the trained artificial intelligence models, which cannot perform the complete thought process of "visual perception - feature description - logical reasoning - final diagnosis" like a doctor. Summary of the Invention

[0004] To address at least some of the problems mentioned above in the prior art, the present invention provides an image multimodal hierarchical annotation method, comprising the following steps:

[0005] Establish image-level annotations. Create at least one image-level selection box for each valid image to cover the entire valid image, and fill in the overall description of the valid image. If the current valid image contains key elements, mark it as containing key elements.

[0006] Use annotation boxes to independently annotate each key element in the valid image;

[0007] Visual information decoupling and text cleaning are performed on the overall descriptions of valid images to remove non-visual descriptions, extract and fill in the visual descriptions, and extract and fill in the main nouns from the visual descriptions; and

[0008] Complete the logical process from visual description to diagnostic conclusion.

[0009] Furthermore, it also includes indicating whether there are any signs of treatment.

[0010] Furthermore, using annotation boxes to independently annotate each key element in the valid image includes:

[0011] Indicate the categories of key elements;

[0012] Use polygonal boxes to label key elements with clearly defined boundaries, and fit them to the edges;

[0013] Use a rectangle to label key elements that have no explicit boundaries, and the rectangle label should completely encompass the key element.

[0014] Furthermore, the overall description of a valid image includes both non-visual and visual descriptions;

[0015] Extracting the main nouns from a purely visual description involves removing adjectives and retaining nouns.

[0016] Furthermore, it also includes validating the annotations of each valid image, and saving the verified annotations as annotation data.

[0017] Furthermore, it also includes pre-screening the input images, retaining valid images including:

[0018] Before starting the annotation process, the images are first evaluated for validity. An option is set to indicate whether an image is irrelevant. If an image is irrelevant, it is marked as such, the reason for its irrelevance is provided, and the image is removed.

[0019] Furthermore, the key elements include one or more of the following:

[0020] Space-occupying lesions / cancer foci, polyps, ulcers, erosions, wounds, surgical instruments, inflammation, foreign bodies, residual feces, bleeding, residual gastric contents, postoperative changes, medications, and Category I elements; among which Category I elements include one or more of mucus, bile, scars, stones, xanthelasma, and unidentified substances.

[0021] Furthermore, the fields for each valid image annotation include: whether it is an irrelevant image, the reason for the irrelevantness, whether there was endoscopic treatment, endoscopic description, description identifiable solely from the endoscopic image, descriptive subject, whether there are key elements, category of key elements, image identification, and annotation level, among which:

[0022] The endoscope description field should contain a complete description of the valid image.

[0023] Fill in the visual description based solely on the description fields that can be identified from the endoscopic image;

[0024] The description field should be filled with the subject's name;

[0025] The logical process of deriving diagnostic conclusions from visual descriptions when filling in the image identification field;

[0026] Select either image level or key element level for the label level field.

[0027] The present invention also provides a semi-automatic image data annotation method, comprising:

[0028] The model automatically detects suspected lesion areas in images and outputs polygonal masks or rectangular boxes.

[0029] The annotator confirms whether the boundaries of the polygon mask or rectangle are accurate, and deletes, adjusts or merges inaccurate polygon masks or rectangles.

[0030] The model automatically generates a preliminary description based on visual features. The annotator decouples the visual information and supplements the reasoning chain, and finally outputs structured data, including: removing non-visual descriptions, extracting and filling in visual descriptions, extracting and filling in the main nouns in the visual descriptions; and filling in the logical process of deriving the diagnostic conclusion from the visual descriptions.

[0031] Furthermore, structured data includes the following fields:

[0032] Whether the image is irrelevant, the cause is irrelevant, whether endoscopic treatment was performed, endoscopic description, description identifiable solely from the endoscopic image, subject of the description, presence of key elements, category of key elements, image identification, and annotation level, among which:

[0033] The endoscope description field should contain a complete description of the valid image.

[0034] Fill in the visual description based solely on the description fields that can be identified from the endoscopic image;

[0035] The description field should be filled with the subject's name;

[0036] The logical process of deriving diagnostic conclusions from visual descriptions when filling in the image identification field;

[0037] Select either image level or key element level for the label level field.

[0038] The present invention has at least the following beneficial effects:

[0039] The image multimodal hierarchical annotation method of the present invention retains pure visual descriptions, that is, descriptions that can be identified solely from endoscopic images, by decoupling visual information and cleaning text, while eliminating non-visual descriptions such as location and pathology. This prevents the model from learning to fabricate location based on images or making erroneous diagnoses without biopsy, which helps to eliminate model illusions and improve the robustness of the model.

[0040] Compared to the common rectangular bounding box detection, this invention forces the use of polygonal bounding boxes to label key elements with clear boundaries, improving boundary localization accuracy, significantly reducing the interference of background noise on feature extraction, and improving the IOU (Intersection over Union) index of lesion segmentation.

[0041] This invention enables the trained model to output diagnostic evidence by filling in the logical process from visual description to diagnostic conclusion, thereby enhancing interpretability.

[0042] The present invention requires that all elements be labeled, including non-lesion elements such as "medication", "surgical instruments", and "residue", so that the trained model can understand the complex scene under endoscopy, and not just identify lesions. Attached Figure Description

[0043] To further illustrate the above and other advantages and features of the various embodiments of the present invention, a more specific description of the embodiments of the invention will be presented with reference to the accompanying drawings. It is to be understood that these drawings depict only typical embodiments of the invention and are therefore not intended to limit its scope. In the drawings, identical or corresponding parts will be indicated by identical or similar reference numerals for clarity.

[0044] Figure 1 The flowchart of an image multimodal hierarchical annotation method according to an embodiment of the present invention is shown. Detailed Implementation

[0045] It should be noted that the components in the accompanying drawings may be shown exaggerated for illustrative purposes and may not be to scale.

[0046] In this invention, the various embodiments are merely intended to illustrate the solutions of the invention and should not be construed as limiting.

[0047] In this invention, unless otherwise specified, the quantifiers “a” and “one” do not exclude scenarios involving multiple elements.

[0048] It should also be noted that, in the embodiments of the present invention, only a portion of the parts or components may be shown for clarity and simplicity. However, those skilled in the art will understand that, under the teachings of the present invention, the required parts or components can be added as needed for specific scenarios.

[0049] It should also be noted that within the scope of this invention, the terms "same", "equal", and "equal to" do not mean that the two values ​​are absolutely equal, but allow for a certain reasonable error. In other words, the terms also cover "substantially the same", "substantially equal", and "substantially equal to".

[0050] It should also be noted that in the description of this invention, the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not explicitly or implicitly suggest that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0051] Furthermore, the embodiments of the present invention describe the process steps in a specific order. However, this is only for the convenience of distinguishing each step, and is not a limitation on the order of each step. In different embodiments of the present invention, the order of each step can be adjusted according to the process.

[0052] Existing image annotation methods are as follows: First, image-level classification, which categorizes the entire digestive endoscopy image into categories such as normal, polyps, and ulcers; second, object detection, which uses rectangular bounding boxes to select and locate lesion regions in the image; third, semantic segmentation, which performs pixel-level masking annotation on lesion regions in the image to achieve accurate segmentation of lesion contours; and fourth, endoscopic visual question answering, which constructs an endoscopic visual question answering (VQA) dataset based on the dataset formed by the aforementioned classification, detection, and segmentation annotation.

[0053] The inventors discovered that the existing annotation schemes described above have many shortcomings when applied to the training of large-scale multimodal models, specifically:

[0054] Rectangular bounding box annotation has inherent limitations. The area selected by this annotation method often includes a large number of normal mucous membrane and other background areas, which introduces redundant background features into the annotation data. This causes the model to learn noisy features during training, thereby affecting the training effect of the model.

[0055] There is a problem of confusion in the descriptive information of the labeled text. Existing labeled texts often directly reuse clinical doctors' diagnostic reports. However, diagnostic reports contain non-visually observable spatial location information such as "15cm from the anus" and pathological information such as "pathological examination shows adenoma" that can only be obtained through biopsy pathological examination. When such non-visual descriptions are included in the labeled data, it is easy to cause logical biases during model training.

[0056] The lack of a diagnostic thought process and the fact that existing annotations only directly provide the lesion judgment results without recording the visual judgment basis makes the model training lack interpretable feature support, resulting in a lack of interpretability of the model.

[0057] The annotation system is a single-level structure, which only annotates the lesion itself or the entire image independently. It lacks hierarchical association annotation between "overall visual overview" and "local lesion details", which cannot support the model to establish the relationship between global visual features and local visual features. This limits the model's ability to understand the overall image of digestive endoscopy and perform detailed local analysis.

[0058] Training large-scale multimodal models requires labeled data that can support the model in reasoning and analyzing visual features (i.e., understanding the visual basis for determining the lesion category, "why it is") and providing a detailed description of the visual features of the lesions (i.e., clearly defining what the lesions look like). Existing labeling methods can no longer meet the training needs of large-scale multimodal models.

[0059] To address the aforementioned issues, this invention provides an image multimodal hierarchical annotation method based on visual information decoupling. Through a strict shape selection strategy (polygon / rectangle), text information cleaning and hierarchical layering (original report text vs. pure visual description), and explicit annotation of thought processes, high-precision, logically strong, and strictly aligned image and text training data is constructed to improve the diagnostic accuracy and interpretability of endoscopic artificial intelligence models.

[0060] Figure 1 The flowchart of an image multimodal hierarchical annotation method according to an embodiment of the present invention is shown.

[0061] like Figure 1 As shown, an image multimodal hierarchical annotation method includes the following steps:

[0062] Step 1: Pre-screen the input images and retain valid images.

[0063] Before labeling begins, the images are first evaluated for validity. An "Irrelevant Image" option is set. If the image is an in vitro photograph, completely blurry (the gastrointestinal wall structure is not clearly visible), or a non-endoscopic image, it is marked as "Yes," and a reason for irrelevance is provided. The image is then removed. Reasons for irrelevance include non-endoscopic image, complete blurriness, etc.

[0064] It should be noted that slightly blurry, reflective, or duplicated images are not considered irrelevant images and should be properly labeled.

[0065] Step 2: Create image-level annotations. Create at least one image-level selection box for each valid image to cover the entire valid image, and fill in the overall description of the valid image. If the current valid image contains key elements, then mark it as containing key elements.

[0066] In one embodiment, the endoscopy description is entered at this image-level description level; that is, the overall description of the current image is extracted from the original report. If there is no original report, it must be entered manually. The overall description of a valid image includes visual information and non-visual descriptions. In one embodiment, the overall description of a valid image is entered in the "Endoscopy Description" field.

[0067] If the current valid image contains key elements, the category of the key elements is indicated (see Table 1 for category classification), forming a "whole-part" contextual relationship. In one embodiment, a "Does a key element exist?" option is set. If an image contains key elements, it is marked as "yes," and the category of the key elements is filled in; otherwise, it is marked as "no."

[0068] Step 3: Use annotation boxes to independently annotate each key element in the valid image. Specifically: indicate the category of the key element; use polygonal boxes to annotate key elements with clear boundaries, fitting the edges; use rectangular boxes to annotate key elements without clear boundaries. Step 3 is element-level annotation.

[0069] Key elements include one or more of the following:

[0070] Space-occupying lesions / cancer foci, polyps, ulcers, erosions, wounds, surgical instruments, inflammation, foreign bodies, residual feces, bleeding, residual gastric contents, postoperative changes, medication, and Class I elements;

[0071] The first category of elements includes one or more of the following: mucus, bile, scars, stones, xanthelasma, and unidentified substances.

[0072] Key elements are categorized based on the following preset list of IDs. This invention standardizes endoscopic features into 14 categories (including "others"), each with a unique ID used to train the model's classification head.

[0073] Table 1 shows the categories of key elements.

[0074]

[0075] If you select "Other" with ID -1, you must enter the specific name separated by commas. If the key element belongs to the first category, select "Other" and enter the specific name.

[0076] Use polygonal frames to fit as closely as possible to the edges of key elements, and minimize the background within the polygonal frame.

[0077] Rectangular bounding boxes are suitable for critical elements that cover a large area (e.g., occupying more than 50% of the entire image) and have no clear boundaries (such as extensive erosion, extensive inflammation, or full-image hemorrhage). The rectangular bounding box must completely encompass the critical element, with no residue outside the box.

[0078] The principle for labeling key elements is to label everything that should be labeled, including lesions, residues, surgical instruments, etc.

[0079] Step 4: Perform visual information decoupling and text cleaning on the overall description of the valid images, remove non-visual descriptions, extract and fill in the visual descriptions, and extract and fill in the main nouns in the visual descriptions.

[0080] Non-visual descriptions include invisible information such as location, size, pathology, and time, and fall under the category of noise reduction processing.

[0081] Remove parts of the overall description of the valid images that cannot be determined from the images (such as location information, pathological results, etc.), and retain only the visually visible information (such as shape, color, texture, etc.).

[0082] In one embodiment, the visual description is entered in the "Description that can be identified from endoscopic images only" field.

[0083] Extracting the main nouns from a purely visual description involves removing adjectives and retaining nouns, such as extracting "mucosa" from "congested and swollen mucosa". The main nouns can be various tissues of the digestive tract.

[0084] Step 5: Fill in the logical process of deriving the diagnostic conclusion from the visual description.

[0085] Step 6: Indicate whether there are any treatment marks on the image.

[0086] In one embodiment, an "Existence of endoscopic treatment" option is set. If there are traces of treatment, it is marked as yes; otherwise, it is marked as no.

[0087] Step 7: Validate the annotations for each valid image. Once valid, save the annotation data as annotation data. The annotation data is in JSON format.

[0088] Verify that image-level annotations correspond one-to-one with the images, that required fields are not empty, and that image-level and element-level annotations do not contradict each other. After successful verification, save the data as high-quality annotation data (JSON format).

[0089] Built-in system verification rules:

[0090] Each image must have one and only one image-level annotation record.

[0091] All required fields (description, category) for the selected area (rectangle or polygon) cannot be empty.

[0092] Image-text consistency check: Image-level annotations and element-level annotations must not contradict each other.

[0093] Since both image-level and element-level annotations require specifying the categories of key elements, contradictions may arise, necessitating verification.

[0094] Contradictory Scenario 1: The "Key Element Category" field at the image level indicates the existence of a certain type of key element, but no corresponding key element record of the same category can be found in the element-level annotations of the same image.

[0095] Contradictory Scenario 2: The "Key Element Category" field at the image level does not contain a certain type of key element, but a record of that type of key element exists in the element-level annotation of the same image.

[0096] In the annotation process, this invention adopts a unified and standardized field definition for "overall image level" and "key element level", as shown in Table 2.

[0097] Table 2 shows the field definitions for endoscopic image annotation.

[0098]

[0099] In one embodiment of the present invention, the image-level annotation fields include: whether it is an irrelevant image, irrelevant cause, whether there is endoscopic treatment, endoscopic description, description that can be identified solely from the endoscopic image, description subject, whether there are key elements, key element category, and image identification.

[0100] The annotation fields at the key element level include the key element category.

[0101] Traditional annotations only tell the model "this is cancer". This invention tells the model "because the surface is uneven and the color is reddish, it is cancer" through the "image identification" field, so that the trained model has the ability to output diagnostic evidence.

[0102] The image multimodal hierarchical annotation method based on visual information decoupling of the present invention has the following advantages:

[0103] A text annotation strategy based on visual information decoupling was adopted: the medical report text was decomposed into a three-level processing flow of "original report", "visual description" and "main nouns" to eliminate non-visual prior noise in multimodal model training.

[0104] An adaptive shape endoscopic feature annotation strategy was adopted: the annotation logic of dynamically selecting polygons or rectangles based on the clarity of lesion boundaries (Focal vs Diffuse) was protected. In particular, a fully enclosed rectangle was used for "bleeding / extensive erosion", while a polygon-fitting classification method was used for "space-occupying lesion / polyp".

[0105] Image annotation data structure containing explicit chain-of-thought: This protects the data structure that introduces an "image identification" field in the image annotation attributes, specifically for recording the reasoning logic from visual features (color, morphology) to medical conclusions.

[0106] A hierarchical annotation process that links image-level and element-level annotations was adopted: protecting the annotation method that simultaneously performs overall scene description (image-level) and independent lesion segmentation (element-level) in an endoscopic image, and establishing the semantic inclusion relationship and consistency verification between the two.

[0107] This invention also provides a semi-automatic image data annotation method applicable to scenarios involving large-scale data construction or limited manual annotation resources. The system preloads a pre-trained endoscopy recognition model (such as a segmentation model or detection model) and a large-scale endoscopy image model (visual language model), automatically generating candidate key region masks and preliminary visual description text. Annotators only need to manually review and correct the corresponding regions and text.

[0108] The specific implementation steps are as follows:

[0109] Step 1: The model automatically detects suspected lesion areas in the image and outputs either a polygon mask or a rectangular bounding box. The polygon mask annotates key elements with clearly defined boundaries. The rectangular bounding box annotates key elements without clearly defined boundaries, ensuring that the bounding box completely encompasses the key elements.

[0110] Step 2: The annotator confirms whether the boundaries of the polygon mask or rectangle are accurate, and deletes, adjusts or merges any inaccurate polygon masks or rectangles.

[0111] Step 3: The model automatically generates a preliminary description based on visual features (such as "red raised lesion"). The annotator decouples the visual information and supplements the inference chain, outputting structured data. At the same time, the model automatically records modification logs for quality tracking.

[0112] The annotation process for visual information decoupling and reasoning chain supplementation includes: removing non-visual descriptions, extracting and filling in visual descriptions, extracting and filling in the main nouns in the visual descriptions, and filling in the logical process of deriving the diagnostic conclusion from the visual descriptions.

[0113] The structured data has the same structure as the manually labeled fields shown in Table 2.

[0114] Compared to fully manual annotation, semi-automatic image data annotation methods are about 2–3 times more efficient, while retaining the oversight of human annotators to control annotation consistency and accuracy.

[0115] The two image data annotation methods of the present invention are not only applicable to digestive endoscopy (gastroscopy, colonoscopy), but can also be applied to the annotation of other endoscopic images such as bronchoscopy and hysteroscopy with slight modifications (such as changing the list of key element categories).

[0116] The image data annotation method of the present invention can be used to construct a medical student teaching system, and generate teaching test questions using data from the "image identification" field.

[0117] To verify the validity of the data generated by the annotation method of this invention, a complete Supervised Fine-Tuning (SFT) process was constructed.

[0118] SFT training data construction: Based on the data structure described in Table 2, the labeled JSON data is converted into a multimodal instruction fine-tuning format.

[0119] Logical mapping: Visual dialogue data is constructed by utilizing the two-level characteristics of image-level and key element-level of this invention.

[0120] Input: Includes endoscopic images and instructions (such as "Please describe the visual features of the selected area in the image and give a diagnosis").

[0121] Output: The visual feature response is based on “descriptions identifiable from endoscopic images alone”, the reasoning chain (CoT) is based on “image identification”, and the diagnostic result is based on the name corresponding to the “key element category ID”.

[0122] In the training data construction stage (i.e. the data annotation stage mentioned above), this invention strictly removes non-visual descriptions (such as pathological results) contained in the original report, retains only the decoupled visual descriptions, and forces the model to learn the alignment relationship between visual features and text.

[0123] Model fine-tuning settings:

[0124] Base model: Qwen 2.5-VL-7B was selected as the pre-trained base model. This model has strong general visual understanding capabilities, but it has shortcomings in fine-grained recognition and logical reasoning in the field of professional medical endoscopy.

[0125] Training method: Full fine-tuning technique is used to input the constructed endoscopy-specific instruction dataset into the model for iterative training, so that it can adapt to visual perception and medical logic under endoscopy.

[0126] The model, fine-tuned using data generated with different annotation standards, was compared with the original base model (Qwen 2.5-VL-7B) on an independent endoscopic test set. The tests focused on the accuracy of lesion detection, the quality of description generation, and whether "illusions" (i.e., fabricating non-existent features or non-visual information) were generated. The results are shown in Table 3.

[0127] Table 3 shows the evaluation results for the four models.

[0128]

[0129] The comparison of diagnostic accuracy shows that training the model with structured data generated by the annotation method of this invention significantly improves the model's ability to identify lesion types (such as space-occupying lesions, polyps, and ulcers).

[0130] The evaluation results on description quality show that the base model tends to generate a large number of long descriptions with insufficient effective information; the annotation method based on existing technical solutions can only generate a single category description; the model trained using only non-decoupled image description annotations has illusions and inaccurate outputs; while the model fine-tuned based on the method of this invention can accurately output professional terms such as "adductor polyp" and "clear boundary".

[0131] Based on the interpretability evaluation results, it is evident that the model trained using the structured data generated by the annotation method of this invention has learned the thought process of "because feature A is seen, therefore lesion B is inferred," rather than directly guessing the result. Existing technical solutions can only provide the final result and lack interpretability.

[0132] The evaluation results based on the hallucination rate show that the hallucination rate of the model trained using the structured data generated by the annotation method of this invention is significantly lower than that of other annotation schemes. Thanks to the visual information decoupling of this invention, the model no longer fabricates information such as "pathological examination shows adenocarcinoma" that cannot be seen from the image.

[0133] The experimental results above demonstrate that the data generated using the annotation method of this invention can significantly improve the professional performance of multimodal large models in the field of endoscopy. In particular, the extremely low illusion rate and high interpretability verify the effectiveness of this invention in solving the two technical problems of "poor image-text alignment" and "lack of logical reasoning chains".

[0134] While some embodiments of the present invention have been described in this application, those skilled in the art will understand that these embodiments are merely illustrative. Numerous variations, alternatives, and improvements will arise in those skilled in the art under the teachings of this invention without departing from its scope. The appended claims are intended to define the scope of the invention and thereby cover methods and structures within the scope of the claims themselves and their equivalents.

Claims

1. A multimodal hierarchical annotation method for images, characterized in that, Includes the following steps: Establish image-level annotations. Create at least one image-level selection box for each valid image to cover the entire valid image, and fill in the overall description of the valid image. If the current valid image contains key elements, mark it as containing key elements. Use annotation boxes to independently annotate each key element in the valid image; Visual information decoupling and text cleaning are performed on the overall descriptions of valid images to remove non-visual descriptions, extract and fill in the visual descriptions, and extract and fill in the main nouns from the visual descriptions; and Complete the logical process from visual description to diagnostic conclusion.

2. The image multimodal hierarchical annotation method according to claim 1, characterized in that, Also includes: Indicate whether there are any traces of treatment.

3. The image multimodal hierarchical annotation method according to claim 1, characterized in that, Using annotation boxes to independently annotate each key element in a valid image includes: Indicate the categories of key elements; Use polygonal boxes to label key elements with clearly defined boundaries, and fit them to the edges; Use a rectangle to label key elements that have no explicit boundaries, and the rectangle label should completely encompass the key element.

4. The image multimodal hierarchical annotation method according to claim 1, characterized in that, The overall description of a valid image includes both non-visual and visual descriptions; Extracting the main nouns from a purely visual description involves removing adjectives and retaining nouns.

5. The image multimodal hierarchical annotation method according to claim 1, characterized in that, It also includes validating the annotations for each valid image, and saving the annotation data after passing the verification.

6. The image multimodal hierarchical annotation method according to claim 1, characterized in that, It also includes pre-screening the input images to retain valid images, including: Before starting the annotation process, the images are first evaluated for validity. An option is set to indicate whether an image is irrelevant. If an image is irrelevant, it is marked as such, the reason for its irrelevance is provided, and the image is removed.

7. The image multimodal hierarchical annotation method according to claim 2, characterized in that, Key elements include one or more of the following: Space-occupying lesions / cancer foci, polyps, ulcers, erosions, wounds, surgical instruments, inflammation, foreign bodies, residual feces, bleeding, residual gastric contents, postoperative changes, medications, and Category I elements; among which Category I elements include one or more of mucus, bile, scars, stones, xanthelasma, and unidentified substances.

8. The image multimodal hierarchical annotation method according to claim 1, characterized in that, Each valid image annotation includes the following fields: whether it is an irrelevant image, the reason for its irrelevantness, whether endoscopic treatment was performed, endoscopic description, description identifiable solely from the endoscopic image, descriptive subject, presence of key elements, category of key elements, image identification, and annotation level. Among these: The endoscope description field should contain a complete description of the valid image. Fill in the visual description based solely on the description fields that can be identified from the endoscopic image; The description field should be filled with the subject's name; The logical process of deriving diagnostic conclusions from visual descriptions when filling in the image identification field; Select either image level or key element level for the label level field.

9. A semi-automatic image data annotation method, characterized in that, include: The model automatically detects suspected lesion areas in images and outputs polygonal masks or rectangular boxes. The annotator confirms whether the boundaries of the polygon mask or rectangle are accurate, and deletes, adjusts or merges inaccurate polygon masks or rectangles. The model automatically generates a preliminary description based on visual features. The annotator decouples the visual information and supplements the inference chain, and finally outputs structured data, including: removing non-visual descriptions, extracting and filling in visual descriptions, and extracting and filling in the main nouns in the visual descriptions. And to fill in the logical process of deriving the diagnostic conclusion from the visual description.

10. The semi-automatic image data annotation method according to claim 9, characterized in that, Structured data contains the following fields: Whether the image is irrelevant, the cause is irrelevant, whether endoscopic treatment was performed, endoscopic description, description identifiable solely from the endoscopic image, subject of the description, presence of key elements, category of key elements, image identification, and annotation level, among which: The endoscope description field should contain a complete description of the valid image. Fill in the visual description based solely on the description fields that can be identified from the endoscopic image; The description field should be filled with the subject's name; The logical process of deriving diagnostic conclusions from visual descriptions when filling in the image identification field; Select either image level or key element level for the label level field.

Citation Information

Patent Citations

  • Model training method, training device and medical image report marking method

    CN114582470A

  • Medical report image labeling method and device

    CN118397637A