Image understanding reasoning system and method based on operation scene
By introducing a combination of image encoder, perception decoder and multimodal large language model in surgical scenarios, the problem of lack of pixel-level reasoning in surgical scenarios is solved, and more refined pixel-level understanding and reasoning is achieved, improving the interpretability and accuracy of surgical AI.
Patent Information
- Application Number
- CN202510147756.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The prior art lacks a comprehensive solution for pixel-level reasoning in surgical scenarios, making it difficult to support advanced surgical robots and precise instrumental operation.
A system of image understanding and reasoning based on surgical scenarios is proposed, including an image encoder, a perception decoder and a multimodal large language model. The surgical image is encoded into image features through an image encoder, and the perception decoder encodes image features and object queries into visual symbols, and combines surgical text instructions to understand and reason, and outputs surgical text response and surgical segmentation mask response.
Pixel-level reasoning in surgical scenarios is realized, the interpretability and accuracy of surgical AI is improved, the domain gap between natural images and surgical images is bridged, and more refined pixel-level understanding and reasoning tasks are supported.
Smart Images

Figure CN120071355A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing in surgical scenarios, and particularly to an image understanding and reasoning system and method based on surgical scenarios. Background Art
[0002] With the development of artificial intelligence (AI) technology, the field of computer-aided intervention (CAI) has undergone revolutionary changes. The application of AI in surgery has not only revolutionized the surgical process but also significantly improved patient safety, surgical precision, and postoperative recovery time. However, despite significant progress, the current surgical AI capabilities still have obvious deficiencies or limitations.
[0003] Surgical visual question answering extends the capabilities of traditional medical visual question answering to the surgical field, focusing on real-time decision support during surgery. The latest advancements include contrastive pre-training, knowledge-aware multimodal representations, and answer query decoders. Although these methods improve question answering accuracy, they usually lack the ability to explain their decisions through visual localization. This limitation has driven the development of more interpretable methods, but still faces challenges in dealing with multiple instruments and answering queries about non-existent objects.
[0004] For example, the surgical stage prediction system and method disclosed in the publication number CN118213051A, the system includes: a visual model module for extracting visual information from surgical video data; a language model module for converting the visual information into a natural language description and obtaining a first surgical stage prediction result based on the natural language description; a multimodal understanding module for obtaining a second surgical stage prediction result and surgical process explanation information based on the visual information, the natural language description, and the first surgical stage prediction result. However, it lacks a comprehensive solution for pixel-level reasoning - which is crucial for advanced surgical robots and precise instrument operation. Summary of the Invention
[0005] The technical problem to be solved by the present invention is: to overcome the deficiencies of the prior art and provide an image understanding and reasoning system and method based on surgical scenarios.
[0006] The technical solution adopted by this application to solve its technical problems is:
[0007] An image understanding and reasoning system based on surgical scenarios, comprising an image encoder, a perception decoder, and a multimodal large language model;
[0008] The image encoder is configured to receive surgical images and encode the surgical images into image features;
[0009] The perception decoder is used to encode image features and learnable object queries into visual symbols and send them to the multi-modal large language model, and decode the information output by the multi-modal large language model;
[0010] The multi-modal large language model is used to receive surgical text instructions and the visual symbols of the perception decoder, combine the surgical text instructions and visual symbols for understanding and reasoning, and output surgical text responses and surgical segmentation mask responses.
[0011] The input of the perception decoder includes a set of learnable object queries. The perception decoder automatically captures all objects of interest, probes the features of the object queries from the image features by using the masked cross-attention mechanism, and establishes a relationship model between objects through the self-attention mechanism;
[0012] Object queries can be decoded into segmentation masks and object categories through a simple FFN layer; through the perception decoder, object information can be efficiently encoded into object-centered visual symbols T ov so as to provide object information in the image and object information referenced by the user for the multi-modal large language model.
[0013] It also includes a perception prior embedding module;
[0014] The perception decoder encodes image features and learnable object queries into pixel-centered visual symbols T pv and object-centered visual symbols T ov aiming to capture both local details and global semantic information of the image, thereby supporting more refined pixel-level understanding and reasoning tasks;
[0015] The perception prior embedding module is used to connect the pixel-centered visual symbol T pv with the object-centered visual symbol T ov to form a visual symbol T v and provide perception prior information for the multi-modal large language model;
[0016] The perception prior embedding module maps the visual symbol T v to the text embedding space of the multi-modal large language model through a visual mapper;
[0017] The multi-modal large language model maps the hidden state of the output segmentation symbol to the visual space through a text mapper.
[0018] Specifically, the perception prior embedding module described in this application maps the visual symbol T v to the text embedding space of the multi-modal large language model through a visual mapper. The visual mapper adopts an MLP (Multi-Layer Perceptron).
[0019] Since the visual symbol Tv It includes pixel-centered and object-centered symbols. Therefore, the visual mapper of this application includes two MLPs, which respectively process one type of visual symbol.
[0020] In addition, the multimodal large language model outputs through a text mapper <seg>The hidden state of the symbol is mapped to the visual space. The text mapper of this application uses a simple MLP.
[0021] An image understanding and reasoning method based on a surgical scenario, applied to the above-mentioned image understanding and reasoning system based on a surgical scenario, includes the following steps:
[0022] Step 1: Construct a dataset;
[0023] Step 2: Establish a surgical scenario graph conversation task (PG-SSGC) based on pixel-level localization with unified symbolic representation;
[0024] Step 3: Form a visual symbol T based on the perceptual prior embedding strategy v ;
[0025] Step 4: Optimize based on the adaptation strategy and output a surgical text response and a surgical segmentation mask.
[0026] The dataset includes 65K conversations, which are based on 10K surgical regions with segmentation masks. Each conversation contains detailed text instructions and responses, and is aligned with the visual annotations of the responses.
[0027] The following sub-steps are included in Step 2:
[0028] 2-1: Set the surgical understanding and reasoning task in a unified form;
[0029] 2-2: Receive input information and process related tasks using a unified instruction form.
[0030] In 2-1, text information including surgical instructions, questions, and descriptions is encoded as a text symbol T t ;
[0031] The dense features of the surgical scenario are denoted as pixel-centered visual symbols T pv ;
[0032] The surgical instrument and tissue features that can be decoded into a segmentation mask are encoded as object-centered visual symbols T ov ;
[0033] The surgical understanding and reasoning task is uniformly represented in the following form:
[0034]
[0035] Among them, LLM represents a multimodal large language model, out represents the output, and in represents the input.
[0036] In this formalization, the surgical scene graph conversation task based on pixel-level localization aims to generate surgical scene descriptions where specific phrases are directly associated with corresponding instrument and tissue segmentation masks. Each bracketed phrase corresponds to a unique segmentation mask in the surgical scene, creating a dense localization description that aligns text interpretation with visual regions.
[0037] The surgical scene graph conversation task based on pixel-level localization can receive visual input and text input and output text responses, segmentation symbols, segmentation masks, and labels. Therefore, the unified form in 2-1 is used in 2-2 to process tasks such as image-level surgical description generation, image-level surgical visual question answering, region-level surgical scene graph generation, surgical instrument segmentation based on text prompts, and surgical scene graph conversation based on pixel-level localization.
[0038] During this process, there are mainly two special symbols: and <seg>Before entering the large language model, the symbols are replaced by visual symbols, and in the output of the large language model <seg>The symbol is decoded into a segmentation mask.
[0039] Then a typical instruction might be " \nDescribe the image and output a text response with an interleaved segmentation mask." The output follows the format: " Kidney <seg>is the focus of the surgical procedure, and various instruments are placed in their respective positions. The one located in the upper left corner Grasping forceps <seg>Organizational operations are being actively carried out. Meanwhile, those located in the upper right region Monopolar curved scissors <seg>Currently in an idle state, waiting to be further used during the operation.”
[0040] The perceptual prior embedding strategy in Step 3 includes the following steps:
[0041] 3-1: Fuse the image features output by the image encoder with the object queries output by the perceptual decoder, where R represents the real number space, H represents the height of the input image, W represents the width of the input image, C represents the number of feature channels, and N represents the number of learnable object queries; q
[0042] Derive the mask score MS for the object queries of each pixel using the segmentation mask obtained from the object queries and the corresponding confidence scores , where its calculation formula is as follows:
[0043]
[0044] where Softmax represents the Softmax function, ⊙ represents element-wise multiplication, and dim represents the dimension;
[0045] 3-2: Calculate the weighted average of the object queries based on the mask score MS, and obtain the weighted object query corresponding to each pixel, and the visual symbol T centered on the pixel pv is obtained by adding the weighted object query to the image feature F, and its calculation formula is as follows:
[0046]
[0047] 3-3: Use the foreground object query as the object-centered visual symbol T ov , and the object-centered visual symbol T ov is concatenated with the pixel-centered visual symbol T pv to form the visual symbol T v =(T pv , T ov ), which is provided as input to the multimodal large language model as rich perceptual prior information.
[0048] The adaptation strategy in Step 4 includes:
[0049] 4-1: Use the LoRA model adaptation technique to perform instruction fine-tuning on the multimodal large language model;
[0050] 4-2: Perform full-scale fine-tuning on the visual mapper and the text mapper;
[0051] 4-3: Encode the necessary task requirements as symbols and input them into the multimodal large language model. Decode the output symbols of the multimodal large language model into text responses and segmentation mask responses according to the task definition, which are represented as follows:
[0052]
[0053] In the formula, is the text regression loss, is the cross-entropy loss, is the dice loss, is the segmentation loss, α represents the α coefficient, and β represents the β coefficient.
[0054] Compared with the prior art, the present application has the following beneficial effects:
[0055] The present application proposes an image understanding and reasoning system and method based on the surgical scenario. Through its streamlined architecture and instruction fine-tuning method, it effectively bridges the domain gap between natural images and surgical images, and at the same time realizes "one model rules all" across different granularity levels and cognitive complexities, achieving precise pixel-level reasoning, which is an important step towards a comprehensive surgical AI system.
[0056] The present application proposes a surgical scenario graph dialogue task for pixel-level localization, which maintains compatibility with existing image-level and object-level surgical understanding and reasoning tasks while achieving pixel-level reasoning. This task form naturally extends the vision-based localization dialogue generation framework to the surgical field, provides a unified method for pixel-level reasoning, makes up for the obvious defects or deficiencies in surgical AI capabilities, and provides a unified framework for comprehensive scene understanding.
[0057] The newly constructed dataset supports model training and robust evaluation across various surgical understanding and reasoning tasks. This dataset is built on the basis of existing surgical datasets and provides rich annotations for pixel-level reasoning tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 is a schematic diagram of the system architecture of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0060] This application constructs an image understanding and reasoning system based on the surgical scenario using the OMG-LLaVA (Large Vision-Language Model) architecture.
[0061] Specifically, referring to Figure 1 , an image understanding and reasoning system based on the surgical scenario includes an image encoder, a perception decoder, and a multimodal large language model;
[0062] The image encoder is used to receive surgical images and encode them into image features;
[0063] The perception decoder is used to encode the image features and learnable object queries into visual symbols and send them to the multimodal large language model, and decode the information output by the multimodal large language model;
[0064] The multimodal large language model is used to receive surgical text instructions and the visual symbols of the perception decoder, combine the surgical text instructions and visual symbols for understanding and reasoning, and output surgical text responses and surgical segmentation mask responses.
[0065] The input of the perception decoder includes a set of learnable object queries. The perception decoder automatically captures all objects of interest, probes the features of the object queries from the image features using the masked cross-attention mechanism, and establishes a relationship model between objects through the self-attention mechanism;
[0066] Object queries can be decoded into segmentation masks and object categories through a simple FFN layer; through the perception decoder, the system can efficiently encode object information into object-centered visual symbols T ov , thus providing object information in the image and object information referenced by the user for the multimodal large language model.
[0067] To maintain the complete capabilities of the system, this application does not fine-tune the perception decoder to adapt to the output of the multimodal large language model, but instead sets up a perception prior embedding module to address this challenge.
[0068] The perception decoder encodes the image features and learnable object queries into pixel-centered visual symbols T pv and object-centered visual symbols T ov , aiming to capture both the local details and global semantic information of the image, thereby supporting more refined pixel-level understanding and reasoning tasks;
[0069] The perception prior embedding module is used to connect the pixel-centered visual symbols T pv with the object-centered visual symbols T ov to form visual symbols T v , providing perception prior information for the multimodal large language model.
[0070] Specifically, the perception prior embedding module described in this application maps the visual symbol T v to the text embedding space of the multimodal large language model through a visual mapper. The visual mapper uses an MLP (Multi-Layer Perceptron).
[0071] Since the visual symbol T v contains pixel-centered and object-centered symbols, the visual mapper in this application includes two MLPs, which respectively process one type of visual symbol.
[0072] In addition, the multimodal large language model outputs the <seg>The hidden state of the symbol is mapped to the visual space. The text mapper of this application uses a simple MLP.
[0073] This architecture provides better potential for future expansion to video analysis and vision-cue-based tasks.
[0074] The understanding and reasoning method of the above image understanding and reasoning system based on the surgical scenario includes the following steps:
[0075] Step 1: Construct a dataset; To support the training and evaluation of Omini-SVLA, this application constructs a dataset for localizing everything in surgical scenario graph conversations. The dataset integrates multiple existing surgical datasets (such as EndoVis2017 and EndoVis2018) and generates new annotations through a hybrid annotation process. For example, we use the tissue bounding boxes in the surgical scenario graph generated dataset as visual cues for SAM to generate high-quality tissue segmentation masks. This integration across data sources makes the dataset more comprehensive and capable of supporting the training and evaluation of multiple tasks. The newly constructed dataset supports model training and robust evaluation across various surgical understanding and reasoning tasks.
[0076] The dataset contains 65K unique conversations based on 10K surgical regions with segmentation masks. Each conversation contains detailed text instructions and responses and is aligned with corresponding visual annotations (such as segmentation masks, bounding boxes, etc.). This text-visual alignment annotation form enables the model to learn how to generate corresponding visual outputs according to text instructions, thus achieving more natural multimodal interaction.
[0077] The uniqueness of the dataset lies in its multi-task support, rich annotation form, annotation generation based on GPT-4, pixel-level localization visual annotation, integration across data sources, and support for pixel-level reasoning tasks. These features enable the dataset to provide comprehensive support for the training and evaluation of surgical AI systems.
[0078] Furthermore, the dataset not only supports traditional image-level tasks (such as surgical description generation, visual question answering), but also object-level tasks (such as surgical scenario graph generation) and pixel-level tasks (such as surgical instrument segmentation based on text prompts, surgical scenario graph conversations based on pixel-level localization). This multi-task support enables the dataset to be used for training and evaluating the performance of the model at multiple granularity levels, while traditional datasets are usually designed for a single task.
[0079] Meanwhile, this application utilizes the powerful context learning ability of GPT-4 to generate rich dialogue annotations and detailed surgical descriptions. Different from traditional manual annotations or simple structured annotations, the annotations generated by GPT-4 are more natural and diverse, capable of simulating complex interactions in real surgical scenarios. For example, we generated multi-turn surgical dialogues through GPT-4, enabling the model to learn how to reason and answer in context.
[0080] A unique feature of the dataset construction in this application lies in its pixel-level localization visual annotations. This application provides segmentation masks for surgical instruments and also generates high-quality tissue segmentation masks through SAM (Segment Anything Model). These pixel-level annotations are paired with text instructions and responses, enabling the model to learn how to perform precise visual localization and reasoning at the pixel level.
[0081] Another unique feature of the dataset is its support for surgical scene graph dialogue tasks based on pixel-level localization. This task requires the model to associate specific phrases with corresponding instrument and tissue segmentation masks when generating surgical scene descriptions, and the pixel-level reasoning task provides rich training and evaluation resources.
[0082] Step 2: To bridge the gap between different granularity levels in surgical AI tasks and simultaneously achieve pixel-level reasoning, this application establishes a surgical scene graph dialogue task (PG-SSGC) based on pixel-level localization with a unified symbolic representation; the following sub-steps are included in Step 2:
[0083] 2-1: Set the surgical understanding and reasoning tasks in a unified form; in 2-1, encode the text information including surgical instructions, questions, and descriptions into text symbols T t ;
[0084] Denote the dense features of the surgical scene as visual symbols T centered on pixels pv ;
[0085] Encode the surgical instrument and tissue features that can be decoded into segmentation masks as visual symbols T centered on objects ov ;
[0086] The surgical understanding and reasoning tasks are uniformly represented in the following form:
[0087]
[0088] Among them, LLM represents a multimodal large language model, out represents the output, and in represents the input.
[0089] In this formalization, the surgical scene graph dialogue task based on pixel-level localization aims to generate surgical scene descriptions where specific phrases are directly associated with corresponding instrument and tissue segmentation masks. Each bracketed phrase corresponds to a unique segmentation mask in the surgical scene, creating a dense localization description that aligns text interpretation with visual regions.
[0090] 2-2: Receive input information and process related tasks using a unified instruction form. The surgical scene graph dialogue task based on pixel-level localization can receive visual input and text input and output text responses, segmentation symbols, segmentation masks, and labels. Therefore, in 2-2, the unified form in 2-1 is used to process tasks such as image-level surgical description generation, image-level surgical visual question answering, region-level surgical scene graph generation, surgical instrument segmentation based on text prompts, and surgical scene graph dialogue based on pixel-level localization.
[0091] During this process, there are mainly two special symbols: and <seg>Before entering the large language model, the symbols are replaced by visual symbols, and in the output of the large language model <seg>The symbol is decoded into a segmentation mask.
[0092] Then a typical instruction might be " \nPlease describe the image and output a text response with an interleaved segmentation mask." The output follows the format: " Kidney <seg>is the focus of the surgical procedure, and various instruments are placed in their respective positions. The one located in the upper left corner Grasping forceps <seg>Organizational operations are being actively carried out. Meanwhile, in the upper right region Monopolar curved scissors <seg>Currently in an idle state, waiting to be further used during the operation.
[0093] Step 3: Form the visual symbol T based on the perceptual prior embedding strategy v ;
[0094] The perceptual prior embedding strategy in Step 3 includes the following steps:
[0095] 3-1: Fuse the image features output by the image encoder with the object queries output by the perceptual decoder, where R represents the real number space, H represents the height of the input image, W represents the width of the input image, C represents the number of feature channels, and N represents the number of learnable object queries. q represents the number of learnable object queries.
[0096] Use the segmentation mask obtained from the object query and the corresponding confidence score to derive the mask score MS for the object query of each pixel, where its calculation formula is as follows:
[0097]
[0098] where Softmax represents the Softmax function, ⊙ represents element-wise multiplication, and dim represents the dimension.
[0099] 3-2: Calculate the weighted average of the object query and obtain the weighted object query corresponding to each pixel, and the pixel-centered visual symbol T pv is obtained by adding the weighted object query to the image feature F, and its calculation formula is as follows:
[0100]
[0101] 3-3: Use the foreground object query as the object-centered visual symbol T ov , and the object-centered visual symbol T ov is concatenated with the pixel-centered visual symbol T pv to form the visual symbol T v =(T pv , T ov ), which is provided as input to the multimodal large language model with rich perceptual prior information.
[0102] Although OMG-LLaVA demonstrates impressive capabilities in natural image understanding and reasoning, directly applying it to the surgical scenario would yield suboptimal results due to the significant domain gap between natural images and surgical images. This gap is manifested in several key aspects. Natural images typically contain everyday objects with consistent appearances, while surgical scenarios feature specialized instruments, complex tissue textures, and varying lighting conditions. Additionally, natural image descriptions rely on common-sense knowledge, whereas surgical scenario interpretation requires domain-specific expertise regarding surgical procedures, instruments, and anatomical structures. There are also substantial differences in the reasoning patterns: natural image reasoning generally focuses on spatial and temporal relationships, while surgical reasoning requires understanding tool-tissue interactions and surgical procedure logic.
[0103] Therefore, to bridge this domain gap while maintaining the powerful reasoning capabilities of OMG-LLaVA, this application proposes a strategy of instruction tuning for surgical domain adaptation. The aim is to preserve the pre-trained visual concepts and relationship understanding of OMG-LLaVA while injecting surgical domain expertise. This is achieved by maintaining the general reasoning ability and language understanding ability of the model during the adaptation process, while introducing surgical vocabulary, domain-specific concepts, and adapting to surgical visual patterns and instrument-tissue relationships. The key challenge lies in learning knowledge transfer from the natural image domain to the surgical domain. This application addresses this problem through a specialized instruction tuning process that carefully calibrates the learning to prevent catastrophic forgetting while ensuring the complementary integration of general knowledge and domain-specific knowledge.
[0104] Instruction tuning aims to enable the model to understand and execute multiple tasks through multi-task learning. Based on diverse instructions (such as generating surgical descriptions, answering surgical-related questions, performing pixel-level segmentation, etc.), the model learns how to execute different tasks according to different instructions during training. This strategy not only improves the model's performance on specific tasks but also enhances the model's generalization ability, enabling it to handle multiple tasks.
[0105] Based on this, the image understanding and reasoning method based on the surgical scenario further includes Step Four: optimizing based on the adaptation strategy and outputting a surgical text response and a surgical segmentation mask.
[0106] The adaptation strategy in Step Four includes:
[0107] 4-1: Using LoRA (Low-Rank Adaptation) technology to perform instruction tuning on the multi-modal large language model;
[0108] 4-2: Performing full-scale tuning on the visual mapper and the text mapper;
[0109] 4-3: In addition to using the text regression loss in this application It also applies the cross-entropy loss and Dice loss for segmentation loss to supervise the segmentation mask for [SEG] symbol decoding, and apply α coefficient and β coefficient to the cross-entropy loss and Dice loss for weighting. In the inference stage, encode the necessary task requirements as symbols and input them into the multi-modal large language model, and decode the output symbols of the multi-modal large language model into text responses and segmentation mask responses according to the task definition, which are expressed as follows:
[0110]
[0111] where is the instruction fine-tuning loss, is the text regression loss, is the cross-entropy loss, is the Dice loss, is the segmentation loss, α represents the α coefficient, and β represents the β coefficient.
[0112] Through this adaptation strategy, while maintaining the basic capabilities of OMG-LLaVA, robust performance is achieved in surgical scene understanding and reasoning. The model demonstrates enhanced capabilities to understand surgical scenes with domain expertise, generate accurate surgical descriptions and segmentations, and perform complex reasoning on surgical procedures, while maintaining its general vision-language understanding capabilities.
[0113] To bridge the significant domain gap between natural images and surgical images, this application designs LoRA to perform instruction fine-tuning on the multi-modal large language model, only updating some parameters, thus retaining the general knowledge learned by the model in the pre-training stage. This method effectively prevents catastrophic forgetting, enabling the model to adapt to specific tasks in the surgical domain while maintaining its general capabilities.
[0114] Based on the above adaptation strategy, the multi-modal large language model can handle multiple tasks without modifying the architecture. This strategy makes the model more flexible, capable of adapting to different task requirements without designing a specific model for each task.
[0115] At the same time, it enables the multi-modal large language model to utilize pre-existing world knowledge and acquire specific domain surgical expertise simultaneously. This method achieves "one model to rule them all", capable of effectively handling tasks at all granularity levels while maintaining high performance in pixel-level reasoning.
[0116] When building the system of this application, based on OMG-LLaVA, the pre-trained ConvNext-L is used as the image encoder, and the decoder in OMG-Seg is used as the perception decoder. The xtuner code library is adopted to build the model and data flow. The image is resized to 1024×1024. In the instruction fine-tuning stage, the initial learning rate is set to 2×10 -4 , the image encoder remains frozen, the perception decoder, visual mapper, and text mapper are fully fine-tuned, and the large language model is fine-tuned using LoRA. The maximum sequence length in the large language model is set to 2,048. All training is carried out on four NVIDIA A6000 GPUs with 48GB of video memory.
[0117] The above are only optional embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural transformation made using the content of the specification of the present invention under the concept of the present invention, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present invention.< / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg>
Claims
1. Image understanding and reasoning system based on surgical scenes, characterized by: Includes image encoder, perceptual decoder, and multimodal large language model; The image encoder is used to receive the surgical image and encode the surgical image into image features; The perceptual decoder is used to encode image features and learnable object queries into visual symbols and send them to the multimodal large language model, and decode the information output by the multimodal large language model; The multimodal large language model is used to receive surgical text instructions and visual symbols of the perception decoder, perform understanding reasoning in combination with the surgical text instructions and the visual symbols, and output surgical text responses and surgical segmentation mask responses.
2. The image understanding and reasoning system based on surgical scenes according to claim 1 is characterized in that: The input of the perceptual decoder includes a set of learnable object queries. The perceptual decoder automatically captures all objects of interest, detects the features of object queries from image features using a masked cross-attention mechanism, and builds a relationship model between objects through a self-attention mechanism.
3. The image understanding and reasoning system based on surgical scenes according to claim 2 is characterized in that: It also includes a perceptual prior embedding module; The perceptual decoder encodes images and learnable object queries into pixel-centric visual symbols T pv and object-centric visual notation T ov , aims to capture both local details and global semantic information of an image, thereby supporting more sophisticated pixel-level understanding and reasoning tasks; The perceptual prior embedding module is used to transform the pixel-centered visual symbol T pv With object-centric visual notation T ov Connect to form the visual symbol T v , providing perceptual prior information for multimodal large language models; The perceptual prior embedding module transforms the visual symbol T v Mapping to the text embedding space of a multimodal large language model; The multimodal large language model maps the hidden states of the output segmentation tokens to the visual space through the text mapper.
4. Image understanding reasoning method based on surgical scenes, characterized in that: The image understanding and reasoning system based on surgical scenes as described in any one of claims 1 to 3 comprises the following steps: Step 1: Build a dataset; Step 2: Establish a surgical scene graph dialogue task based on unified symbolic representation; Step 3: Form visual symbol T based on perceptual prior embedding strategy v ; Step 4: Optimize based on the adaptation strategy and output the surgical text response and surgical segmentation mask.
5. The image understanding reasoning method based on surgical scenes according to claim 4 is characterized in that: The dataset includes 65K dialogues based on 10K surgical regions with segmentation masks, each containing detailed text instructions and responses, aligned with visual annotations of the responses.
6. The image understanding reasoning method based on surgical scenes according to claim 4 is characterized in that: The step 2 includes the following sub-steps: 2-1: Set surgical comprehension and reasoning tasks in a unified format; 2-2: Receive input information and use a unified instruction format to process related tasks.
7. The image understanding reasoning method based on surgical scenes according to claim 6 is characterized in that: In 2-1, the text information including surgical instructions, questions and descriptions is encoded into text symbols T t ; The dense features of the surgical scene are recorded as pixel-centered visual symbols T pv ; Encode surgical instruments and tissue features as object-centric visual symbols T that can be decoded into segmentation masks ov ; The surgical understanding and reasoning tasks are uniformly expressed in the following form: Among them, LLM stands for multimodal large language model, out stands for output, and in stands for input.
8. The image understanding reasoning method based on surgical scenes according to claim 7 is characterized in that: In 2-2, the unified form in 2-1 is used to process image-level surgical description generation, image-level surgical visual question and answer, region-level surgical scene graph generation, surgical instrument segmentation based on text prompts, and surgical scene graph dialogue tasks based on pixel-level positioning.
9. The image understanding reasoning method based on surgical scenes according to claim 8 is characterized in that: The perceptual prior embedding strategy in step 3 includes the following steps: 3-1: Image features output by the image encoder With perceptual decoder Output object query Fusion, where R represents the real space, H represents the height of the input image, W represents the width of the input image, C represents the number of feature channels, and N q represents the number of learnable object queries; using the segmentation masks obtained from object queries and the corresponding confidence scores A mask score MS is derived for each pixel object query, where The calculation formula is as follows: Among them, Softmax represents the Softmax function, ⊙ represents element-by-element multiplication, and dim represents the dimension; 3-2: Calculate the weighted average of the object query Q based on the mask score MS, and obtain the weighted object query corresponding to each pixel, the visual symbol T centered on the pixel pv It is obtained by adding the weighted object query to the image feature F, which is calculated as follows: 3-3: Foreground object query as object-centered visual symbol T ov , object-centered visual notation T ov With pixel-centric visual symbols T pv Connect to form the visual symbol T v =(T pv , T ov ), which is provided as input to the multimodal large language model to perceive the prior information.
10. The image understanding reasoning method based on surgical scenes according to claim 9 is characterized in that: The step 4 includes the following sub-steps: 4-1: Use LoRA model adaptation technology to fine-tune the multimodal large language model; 4-2: Fully fine-tune the visual mapper and text mapper; 4-3: Encode the necessary task requirements as symbols and input them into the multimodal large language model. According to the task definition, decode the output symbols of the multimodal large language model into text responses and segmentation mask responses, which are expressed as follows: In the formula, is the text regression loss, is the cross entropy loss, is the dice loss, is the segmentation loss, α represents the α coefficient, and β represents the β coefficient.
Citation Information
Patent Citations
Surgical stage prediction system and method
CN118213051A
Online operation video instrument tracking system based on text promptable
CN117789921A
Visual big language model construction method for multi-source ship remote sensing image interpretation
CN117994677A
Method, model and device for enhancing visual perception capability of multi-modal large language model
CN118585954A
Large language model training method and device, reasoning method and device, equipment and storage medium
CN118673325A
Cited By
Unified anomaly detection method and system based on conditional adapter
CN120526232A