Text interactive minimally invasive surgery instrument segmentation method and system

By generating instrument text prompts and extracting features from a large visual-language model, flexible and accurate segmentation of minimally invasive surgical instruments is achieved. This solves the shortcomings of existing models in instrument category differentiation and adaptability, and improves segmentation accuracy and generalization ability.

CN119228820BActive Publication Date: 2026-03-20TONGJI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-12
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing minimally invasive surgical instrument segmentation models struggle to distinguish between different instrument categories and cannot adapt to the ever-increasing variety of surgical instruments, resulting in insufficient segmentation accuracy and safety.

Method used

By generating textual prompts for instruments and extracting visual-text feature pairs using a large visual-language model, mask probability maps are generated and weighted fusion is performed to enhance the recognition of difficult-to-segment regions and achieve accurate segmentation of surgical images.

Benefits of technology

It improves the accuracy and generalization ability of segmentation of minimally invasive surgical instruments, and can achieve considerable segmentation accuracy on standard datasets and new categories, overcoming the limitations of low accuracy and poor generalization ability of existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119228820B_ABST
    Figure CN119228820B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of computer vision, and discloses a text interactive minimally invasive surgical instrument segmentation method and system, text prompt information of an instrument is generated, and feature extraction is performed on surgical images and the text prompt information based on a vision-language large model; visual features are strengthened and refined according to text features, and mask probability graphs corresponding to various instruments are generated; a plurality of mask probability graphs corresponding to different text prompt information of the same instrument category are weighted and fused; a difficult segmentation area is obtained by analyzing the corresponding relationship between the fused mask probability graph and a preset label, the difficult segmentation area in the surgical image is input into a preset image encoder after being masked, and a corresponding decoder is added to perform image reconstruction. The method can flexibly and accurately perform mask segmentation prediction on different surgical instruments on the input surgical image, and breaks through the limitations of low segmentation precision, poor generalization ability and fixed prediction categories of existing methods.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, in particular to a text interactive minimally invasive surgical instrument segmentation method and system. BACKGROUND

[0002] Minimally invasive surgery refers to a surgical procedure that uses endoscopic equipment to access the body through a small incision. It is the preferred treatment method for disciplines such as gastroenterology, cardiothoracic surgery, general surgery, and urology. In recent years, with the deep integration of new generation information technology and the medical industry, minimally invasive surgery has begun to evolve towards intelligentization. Minimally invasive surgery robots can improve the flexibility, stability, and accuracy of physician operations and assist physicians in more difficult and precise surgeries. Considering the small and fragile operating space inside the patient's body, robotic-assisted minimally invasive surgery must use surgical instrument segmentation algorithms to obtain the precise position and pose of the surgical instrument, thereby improving the physician's perception and understanding of the surgical process and ensuring the safe and stable operation of the surgery.

[0003] With the development of deep neural networks and medical image analysis technology, researchers in related fields have proposed a variety of surgical instrument segmentation technical solutions. Recent research shows that although existing segmentation models can provide relatively accurate binary foreground mask prediction results containing instruments, they still cannot distinguish between different instrument categories. The reasons are as follows: the most advanced methods today, such as S3Net, TraSeTR, and MATIS, all focus on the visual end. They first use an image encoder to extract the visual features of the input image, and then use an instance segmentation decoder to convert the extracted visual features into a binary surgical instrument mask. Finally, the predicted mask is classified using visual semantic features. However, since these models do not consider the text semantics of instrument descriptions, they only use visual features to understand abstract category information, which is limited by the similarity of surgical instruments and the scarcity of pixel annotations. Their segmentation performance is still not ideal, and they cannot meet the precision and safety requirements of clinical surgery. In addition, with the rapid development of minimally invasive surgery, the number of instrument types used has increased dramatically. Current minimally invasive surgical instrument segmentation algorithms are not sufficient to adapt to the increasing number of surgical instrument types. Each time a new instrument category is introduced, data needs to be re-labeled and models need to be re-trained, which seriously hinders the practical application of surgical instrument segmentation technology in the field of minimally invasive surgery. SUMMARY

[0004] The present application provides a text interactive minimally invasive surgical instrument segmentation method to solve the problems of weak instrument differentiation ability and fixed and unchangeable prediction label space in the prior art.

[0005] Correspondingly, the present application also provides a text interactive minimally invasive surgical instrument segmentation system, an electronic device, and a computer readable storage medium for ensuring the implementation and application of the above method.

[0006] To solve the above technical problems, the application discloses a text interactive minimally invasive surgical instrument segmentation method, the method comprising:

[0007] generating text prompt information of the instrument; the text prompt information includes instrument category name, instrument appearance description and instrument function description;

[0008] extracting features of the surgical image and the text prompt information based on a visual-linguistic large model to obtain a visual-text feature pair; the visual-text feature pair includes visual features and text features;

[0009] According to the text features, the visual features are strengthened and refined to generate a mask probability graph corresponding to each type of instrument;

[0010] The multiple mask probability graphs corresponding to different text prompt information of the same instrument category are weighted and fused to obtain a fused mask probability graph;

[0011] The difficult segmentation area in the surgical image is input into a preset image encoder after being masked, and a corresponding decoder is added to perform image reconstruction; wherein the difficult segmentation area is obtained by analyzing the corresponding relationship between the fused mask probability graph and the preset label.

[0012] The application also discloses a text interactive minimally invasive surgical instrument segmentation system, the system comprising:

[0013] The text prompt generation module is used for generating text prompt information of the instrument; the text prompt information includes instrument category name, instrument appearance description and instrument function description;

[0014] The encoder module is used for extracting features of the surgical image and the text prompt information based on a visual-linguistic large model to obtain a visual-text feature pair; the visual-text feature pair includes visual features and text features;

[0015] The mask decoder module is used for strengthening and refining the visual features according to the text features to generate a mask probability graph corresponding to each type of instrument;

[0016] The mixed prompt prediction fusion module is used for weighting and fusing multiple mask probability graphs corresponding to different text prompt information of the same instrument category to obtain a fused mask probability graph;

[0017] The difficult segmentation area strengthening module is used for inputting the difficult segmentation area in the surgical image into a preset image encoder after being masked, and adding a corresponding decoder to perform image reconstruction; wherein the difficult segmentation area is obtained by analyzing the corresponding relationship between the fused mask probability graph and the preset label.

[0018] The application further discloses an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the method according to any one of the application when executing the program.

[0019] The application further discloses a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of the application.

[0020] In the application, the generation of the text prompt information including the instrument category name, the instrument appearance description and the instrument function description helps to alleviate the problem of poor model semantic segmentation performance caused by the similarity of surgical instruments and the scarcity of pixel labeling, and provides important text prompts for text interactive minimally invasive surgical instrument segmentation. Based on the visual-linguistic large model, the visual-linguistic feature pairs including visual features and text features are obtained by extracting features from the surgical image and the text prompt information, and the alignment of the visual features and the text features is realized. The visual features are strengthened and refined according to the text features, and the mask probability graph corresponding to each instrument category is generated, so that the model can more accurately identify the position and category of the surgical instrument, thereby improving the precision of the minimally invasive surgical instrument segmentation. The multiple mask probability graphs corresponding to different text prompt information of the same instrument category are weighted and fused to obtain a fused mask probability graph, thereby improving the adaptability of the model to diversified input texts. By analyzing the corresponding relationship between the fused mask probability graph and the preset label, the difficult segmentation area is mined, and then the difficult segmentation area in the surgical image is input into a preset image encoder after being masked, and a corresponding decoder is added to perform image reconstruction, thereby enhancing the segmentation and prediction accuracy of the model for the difficult segmentation area and the edge details of the surgical instrument.

[0021] The method in the application can flexibly and accurately perform mask segmentation prediction on different surgical instruments on the input surgical image. Moreover, the method in the application can obtain considerable segmentation accuracy on new categories that are not seen in the training stage in addition to obtaining accurate prediction results on the basis classes trained on the standard data set, thereby breaking through the limitations of the existing methods such as low segmentation accuracy, poor generalization ability and fixed prediction categories.

[0022] Additional aspects and advantages of the application will be described in the following description section, which will become apparent from the following description or will be understood by those skilled in the art through practice of the application. BRIEF DESCRIPTION OF DRAWINGS

[0023] The above and / or additional aspects and advantages of the application will become apparent and be readily understood from the following description, taken in conjunction with the accompanying drawings, in which:

[0024] Figure 1A flowchart of a text interactive minimally invasive surgery instrument segmentation method provided by an embodiment of the present application is shown in FIG. 1.

[0025] Figure 2 A text prompt generation module schematic diagram provided by an embodiment of the present application is shown in FIG. 3.

[0026] Figure 3 A network structure schematic diagram provided by an embodiment of the present application is shown in FIG. 4.

[0027] Figure 4 A structure schematic diagram of a text interactive minimally invasive surgery instrument segmentation system provided by an embodiment of the present application is shown in FIG. 5.

[0028] Figure 5 A structure schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 6. DETAILED DESCRIPTION

[0029] Embodiments of the present application are described in detail below with reference to the accompanying drawings, in which examples of embodiments are shown, and the same or similar numerals throughout the drawings denote the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as a limitation on the present application.

[0030] Those skilled in the art can understand that, unless specifically stated, the singular forms "a", "an" and "the" used herein also include the plural forms. It should be further understood that the phrase "comprising" used in the specification of the present application means that a feature, integer, step, operation, element and / or component exists, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or there can be intermediate elements. In addition, "connected" or "coupled" used herein can include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any single unit and all combinations of the associated listed items.

[0031] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as that generally understood by those skilled in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have meanings consistent with those in the context of the prior art, and unless specifically defined as such, should not be interpreted in an idealized or overly formal sense.

[0032] The scheme provided by the embodiments of the present application can be executed by any electronic device, which can be a terminal device or a server. The server can be a physical server, a server cluster composed of multiple physical servers, a distributed system, or a cloud server providing cloud computing services. The terminal can be a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be connected directly or indirectly through wired or wireless communication, which is not limited in the present application. For the technical problems existing in the prior art, the text interactive minimally invasive surgical instrument segmentation method and system provided by the present application aims to solve at least one of the technical problems of the prior art.

[0033] The technical scheme of the present application and how the technical scheme of the present application solves the above technical problems will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described again in some embodiments. The embodiments of the present application will be described below with reference to the drawings.

[0034] The embodiments of the present application provide a possible implementation manner, as shown in Figure 1 A flowchart of a text interactive minimally invasive surgical instrument segmentation method is provided, which can be executed by any electronic device, and can be executed on a server side or a terminal device.

[0035] As shown in Figure 1 The method can include the following steps:

[0036] Step 101, generating instrument text prompt information; the text prompt information includes instrument category name, instrument appearance description and instrument function description.

[0037] In the embodiments of the present application, the prompt text information containing the instrument category name, the instrument appearance description and the instrument function description is constructed, which can provide an important text prompt input set for the text interactive minimally invasive surgical instrument segmentation model constructed by the embodiments of the present application (the method described in the following steps 102-105).

[0038] Step 102, performing feature extraction on the surgical image and the text prompt information based on a vision-language large model to obtain a vision-text feature pair; the vision-text feature pair includes a vision feature and a text feature.

[0039] In the embodiments of the present application, the vision feature of the input surgical image and the text feature of the text prompt are first extracted to form a vision-text feature pair. To ensure that the extracted vision feature and the text feature have high semantic consistency, the image encoder and the text encoder used in the embodiments of the present application come from a vision-language large model (CLIP).

[0040] Step 103, according to the text features, the visual features are enhanced and refined, and the mask probability graph corresponding to each type of instrument is generated.

[0041] The visual features are enhanced and refined by using the guidance of the text features, and then a high-quality mask probability graph is generated from the enhanced and refined visual features, so that the model can more accurately identify the position and category of the surgical instrument, which is of great significance to improve the precision of minimally invasive surgical instrument segmentation.

[0042] Step 104, weighting and fusing multiple mask probability graphs corresponding to different text prompt information of the same instrument category to obtain a fused mask probability graph.

[0043] For an input text prompt set containing multiple text prompts of the same instrument, multiple mask probability graphs will be generated after the text prompt information contained therein is input into the model. In the embodiment of the present application, different mask probability graphs of the same instrument category are weighted and fused, so as to fully exploit the knowledge compressed in the visual-linguistic large model encoder, to improve the precision and generalization ability of the model in surgical instrument segmentation, and thus to improve the adaptability of the model to diversified input texts.

[0044] Step 105, inputting the difficult segmentation area in the surgical image into a preset image encoder after masking, and adding a corresponding decoder to perform image reconstruction; wherein the difficult segmentation area is obtained by analyzing the corresponding relationship between the fused mask probability graph and the preset label.

[0045] In the embodiment of the present application, the visual feature representation ability of the model for the difficult segmentation area is enhanced during the model training phase, which enhances the segmentation and prediction accuracy of the model for the difficult segmentation area and edge details of the surgical instrument.

[0046] In the embodiments of the present application, the generation of text prompt information including instrument category name, instrument appearance description and instrument function description helps to alleviate the problem of unsatisfactory model semantic segmentation performance caused by the similarity of surgical instruments and the scarcity of pixel labeling, and provides important text prompts for text interactive minimally invasive surgical instrument segmentation. Based on the visual-linguistic large model, the features of the surgical image and the text prompt information are extracted to obtain a visual-text feature pair including visual features and text features, and the alignment of visual and text features is realized. According to the text features, the visual features are strengthened and refined to generate a mask probability graph corresponding to each instrument category, so that the model can more accurately identify the position and category of the surgical instrument to improve the precision of minimally invasive surgical instrument segmentation. The multiple mask probability graphs corresponding to different text prompt information of the same instrument category are weighted and fused to obtain a fused mask probability graph, thereby improving the adaptability of the model to diversified input text. By analyzing the correspondence between the fused mask probability graph and the preset label, the difficult segmentation area is mined, and then the difficult segmentation area in the surgical image is input into the preset image encoder after being masked, and the corresponding decoder is added to perform image reconstruction, thereby enhancing the segmentation prediction accuracy of the model for the difficult segmentation area and edge details of the surgical instrument.

[0047] Based on the above, the method in the embodiments of the present application can flexibly and accurately perform mask segmentation prediction on different surgical instruments on the input surgical image. Moreover, in addition to being able to obtain accurate prediction results on the basis classes trained on the standard data set, the method in the embodiments of the present application can also obtain considerable segmentation accuracy on new classes not seen in the training stage, breaking through the limitations of existing methods such as low segmentation accuracy, poor generalization ability and fixed prediction categories.

[0048] In an optional embodiment, the text prompt information of the instrument is generated, including:

[0049] For different instrument categories, the text prompt information of the instrument is generated based on a large language model.

[0050] In the embodiments of the present application, the text prompt generation module generates the text prompt information of the instrument. Optionally, as shown in Figure 2 The text prompt generation module constructs a prompt text set containing the names of each surgical instrument category and the appearance description of the instrument. This module develops three types of prompt words with increasing complexity. First, the inherent name of each instrument category is used as a prompt word. Second, inspired by the CLIP prompt template, a prompt word template is designed, which is "the surgical instrument area represented by [category name]". Third, the large language model GPT-4 released in recent years is used to generate prompt words, and a prompt question template is designed: "Please describe the appearance of [category name] in endoscopic surgery. Change the description to a phrase with a subject, and do not use colons.

[0051] By generating more detailed appearance description text prompts for each surgical instrument, a text prompt set {T1, T2, T3} suitable for a surgical instrument segmentation framework is finally constructed, where T1 represents the inherent name of the instrument category, which can be referred to as the instrument category name, such as "bipolar forceps", "monopolar curved scissors"; T2 represents the prompt information generated according to the description word "surgical instrument region represented by [class name]", such as "surgical instrument region represented by [bipolar forceps]", "surgical instrument region represented by [monopolar curved scissors]"; T3 represents the prompt information generated according to the prompt template "Please describe the appearance of [class name] in endoscopic surgery. Change the description to a phrase with a subject, and do not use colons.", such as "Monopolar curved scissors show an elongated handle, a curved cutting edge for precise cutting, and an insulated shaft allowing control of the application of cutting and coagulation electrical energy."

[0052] In an optional embodiment, the vision-language large model comprises an image encoder and a text encoder;

[0053] Based on the vision-language large model, the surgical image and the text prompt information are feature extracted to obtain a vision-text feature pair, comprising:

[0054] Based on the image encoder, the surgical image is feature extracted to obtain a vision feature;

[0055] Based on the text encoder, the text prompt information is feature extracted to obtain a text feature.

[0056] In the embodiments of the present application, the surgical image and the text prompt information are feature extracted based on the encoder module, thereby forming a vision-text feature pair. Specifically, the image encoder (vision transformer, ViT) in the vision-language large model is fine-tuned in the training process. For example, Figure 3 For the input surgical image I (height and width are H and W respectively), the image encoder is used to obtain the output of the image encoder, and the output of the image encoder is represented as , where I is in the range of 1 to 12; N is the number of vision tokens (i.e. the sequence length of the vision feature); D is the channel dimension of the vision feature. Assuming that the image block conversion projection window length and width set by the image encoder are both P, then the sequence length of the vision feature is:

[0057]

[0058] Meanwhile, in order to aggregate multi-scale vision features, the embodiments of the present application simulate the multi-scale feature fusion mechanism widely used in convolutional neural network architecture: the output features of the 4th layer, the 8th layer and the 12th layer of the image encoder are extracted, then they are reconstructed into the structure of feature maps and these features are fused step by step through the feature pyramid network (FPN).

[0059] A text encoder (Bert) in the visual-linguistic large model is adopted, and its parameters are kept frozen during the training process. As shown in Figure 3 For the input text prompt T, the text encoder can convert it into a text feature in vector form:

[0060] F T ∈R 1×D

[0061] where D is the channel dimension of the text feature, and its value is equal to the channel dimension of the visual feature.

[0062] In an optional embodiment, the visual feature is reinforced and refined according to the text feature to generate a mask probability map corresponding to each type of instrument, including:

[0063] The visual feature and the text feature are decoded based on an attention mechanism to obtain an attention-based text prompt feature;

[0064] The text prompt feature is processed based on a convolution operation to obtain the mask probability map.

[0065] In the embodiments of the present application, a mask decoder module is designed to reinforce and refine the extracted visual feature according to the guidance of the text feature, and then generate a mask probability map corresponding to each type of surgical instrument. As shown in Figure 3 The mask decoder module includes two sub-modules: an attention-based prompt module and a convolution-based prompt module. The former can be regarded as a global decoding process based on attention operation, and the latter is a local decoding process based on convolution operation.

[0066] In an optional embodiment, the visual feature and the text feature are decoded based on an attention mechanism to obtain an attention-based text prompt feature, including:

[0067] Self-attention is calculated within the visual feature to obtain a self-attention feature;

[0068] Cross-attention of the self-attention visual feature and the text prompt feature is calculated to obtain a cross-attention feature;

[0069] The cross-attention feature is input into a feedforward network to output an attention-based text prompt feature.

[0070] In the attention-based prompt module, first, the visual feature F IInternal computational self-attention (SA) is used to enhance and activate regions in the visual features that may contain foreground elements. The self-attention operation consists of two parts: layer normalization (LN) and skip connections. The process is described in the following formula:

[0071]

[0072] Next, calculate With F T Cross-attention (CA), according to F T guidance The instrument area is located and further activated. Similarly, this process can be written as:

[0073]

[0074] Subsequently, a feedforward network consisting of fully connected layers is applied to further enhance the visual features guided by textual features:

[0075]

[0076] Treating SA, CA, and FFN as a single decoding unit, this application embodiment uses a total of three attention-based decoding units sequentially to obtain attention-based text prompt features.

[0077] In an optional embodiment, the text prompt features are processed based on convolutional operations to obtain a mask probability map, including:

[0078] Reconstruct the text prompt features;

[0079] The text features are converted into convolution parameters, and the reconstructed text prompt features are convolved using the convolution parameters to obtain the convolution output features; the convolution parameters include weights and parameters.

[0080] The Sigmoid function is used to transform the convolution output features into a mask probability map.

[0081] Specifically, the convolution-based prompt module designs a convolution-based feature enhancement scheme, which further refines the text prompt features through local convolution. Therefore, it is first necessary to analyze the features of text prompts. Reconstruct the text to restore its spatial dimension: extract the text features F T The parameters are converted into convolutional parameters, namely weights and biases, and then used to refine the text cue features. Convolution is performed. The conversion process is completed through a fully-connected (FC) layer, and the fully-connected layer converts the text feature F T into a vector with a channel dimension of D x k x k + 1, where "k x k" represents the size of the convolution kernel, and "+1" is an additional dimension considering the bias. This operation can decompose the text feature F T into a convolution weight w e R 1×D×k×k and a bias b e R 1 , which is then used for convolution, and the final output is converted into a mask probability map S using a Sigmoid function:

[0082]

[0083] The combination of the above attention-based prompt module and the convolution-based prompt module enables the model to more accurately identify the position and category of surgical instruments, which is of great significance to improving the precision of minimally invasive surgical instrument segmentation.

[0084] In an optional embodiment, a plurality of mask probability maps corresponding to different text prompt information of the same instrument category are weighted and fused to obtain a fused mask probability map, including:

[0085] For a plurality of mask probability maps corresponding to different text prompt information of the same instrument category, a plurality of weight mappings corresponding to the plurality of mask probability maps are output based on a visual-text threshold network;

[0086] Based on the weight mappings, the plurality of mask probability maps are weighted and fused to obtain a fused mask probability map;

[0087] The visual-text threshold network includes a plurality of residual blocks.

[0088] The embodiments of the present application implement the weighted fusion of a plurality of mask probability maps based on the hybrid prompt prediction fusion module as shown in Figure 3 The specific implementation manner is as follows:

[0089] First, for a given hybrid prompt set of each target instrument category (including the name of the instrument, the prompt word template, the instrument appearance description generated by the large language model, etc.), it is marked as T = {T1, T2, T3}, and the text feature converted from the i th prompt is marked as Then, each is input into a mask decoder module together with the visual feature F I , so as to obtain a mask probability map S iTo fuse the mask probability maps in the set and obtain the final segmentation mask prediction, the hybrid cue prediction fusion module designs a visual-text threshold network G to perform weighted fusion of mask probability maps corresponding to multiple text cues. The visual-text threshold network G consists of 3 residual blocks. The visual features F... I and text features Before entering G after connecting, first enter F. T ∈R 1×D Copy to To match F I The dimension. In this embodiment of the application, taking the example of three mask probability maps corresponding to different text prompts for the same device category, the output of the visual-text threshold network G is three weight mappings, corresponding to three mask probability maps. Through weighted calculation, the final fused mask probability map S can be obtained. et ∈R H×W The final mask prediction M is determined by S. et The threshold is then applied.

[0090] The hybrid prompt prediction fusion module can fully exploit the compressed knowledge in the visual-language large model encoder, enhance the model's adaptability to complex input text descriptions, and thus improve the model's segmentation accuracy and generalization ability.

[0091] In an optional embodiment, by means of... Figure 3 The difficult segmentation region enhancement module shown in the figure masks the difficult segmentation regions in the surgical image, inputs them into a preset image encoder, and adds a corresponding decoder to perform image reconstruction.

[0092] The difficult segmentation region enhancement module in this embodiment differs from the random masking strategy based on masked autoencoders (MAE). It aims to focus the image encoder on challenging difficult segmentation regions. Through a non-random difficult segmentation region mining mechanism, it significantly enhances the model's ability to represent the visual features of difficult segmentation regions during the model training phase, addressing the performance deficiencies of existing methods in instrument type identification and instrument edge segmentation. This module first analyzes the correspondence between the fused mask probability map and the ground truth labels to identify difficult segmentation regions. Then, it masks these difficult segmentation regions in the original image (i.e., the surgical image) and inputs them into the image encoder, adding a corresponding decoder to perform image reconstruction. In this embodiment, the difficult segmentation region enhancement module shares the same image encoder as the surgical instrument segmentation model.

[0093] Specifically, given a segmentation mask probability map, the difference between it and the ground truth label is calculated to obtain difficult segmentation regions. Then, based on a pre-set masking ratio, these difficult regions are selectively masked to generate a masked image. The masked image is then input into an image encoder for visual feature extraction, and an image reconstruction decoder is built to reconstruct the complete image region from the unmasked areas. This module guides the image encoder to focus on key, difficult-to-identify regions in the image, improving its ability to extract and represent visual features of difficult areas of surgical instruments, thereby enhancing the model's accuracy.

[0094] like Figure 2 As shown in the embodiment of this application, during the training process, the loss L between the fused mask probability map and the original input image is calculated. seg And the loss L between the difficult segmentation region and the corresponding region in the original input image. rec To improve training accuracy.

[0095] Based on the same principles as the methods provided in the embodiments of this application, the embodiments of this application also provide a text-interactive minimally invasive surgical instrument segmentation system, such as... Figure 4 As shown, the system includes:

[0096] The text prompt generation module 401 is used to generate text prompt information for the device; the text prompt information includes the device category name, device appearance description and device function description;

[0097] Encoder module 402 is used to extract features from surgical images and text prompts based on a large vision-language model to obtain visual-text feature pairs; the visual-text feature pairs include visual features and text features;

[0098] The mask decoder module 403 is used to enhance and refine visual features based on text features, and generate mask probability maps corresponding to various types of instruments.

[0099] The hybrid prompt prediction fusion module 404 is used to perform weighted fusion of multiple mask probability maps corresponding to different text prompt information of the same device category to obtain a fused mask probability map.

[0100] The difficult segmentation region enhancement module 405 is used to mask the difficult segmentation regions in the surgical image and input them into a preset image encoder, and add a corresponding decoder to perform image reconstruction; wherein, the difficult segmentation regions are obtained by parsing the correspondence between the fused mask probability map and the preset label.

[0101] In the embodiments of the present application, flexible and accurate mask segmentation prediction can be performed on different surgical instruments on the input surgical image. Moreover, in addition to obtaining accurate prediction results on the basic classes trained on the standard data set, the method in the embodiments of the present application can also obtain considerable segmentation accuracy on new classes not seen in the training stage, breaking through the limitations of low segmentation accuracy, poor generalization ability and fixed prediction classes of existing methods.

[0102] The text interactive minimally invasive surgical instrument segmentation system provided in the embodiments of the present application can realize the method embodiments Figures 1 to 3 The processes realized in the method embodiments are not repeated here to avoid repetition.

[0103] The text interactive minimally invasive surgical instrument segmentation system in the embodiments of the present application can execute the text interactive minimally invasive surgical instrument segmentation method provided in the embodiments of the present application, and the implementation principles are similar. The actions performed by each module and unit in the text interactive minimally invasive surgical instrument segmentation system in the embodiments of the present application are corresponding to the steps in the text interactive minimally invasive surgical instrument segmentation method in the embodiments of the present application. For the detailed function description of each module of the text interactive minimally invasive surgical instrument segmentation system, refer to the description of the corresponding text interactive minimally invasive surgical instrument segmentation method shown in the foregoing, which will not be repeated here.

[0104] Based on the same principles as the methods shown in the embodiments of the present application, the embodiments of the present application also provide an electronic device, which can include but is not limited to a processor and a memory. The memory is used to store a computer program, and the processor is used to execute the text interactive minimally invasive surgical instrument segmentation method shown in any optional embodiment of the present application by calling the computer program. Compared with the prior art, the text interactive minimally invasive surgical instrument segmentation method provided in the present application can perform flexible and accurate mask segmentation prediction on different surgical instruments on the input surgical image. Moreover, in addition to obtaining accurate prediction results on the basic classes trained on the standard data set, the method in the embodiments of the present application can also obtain considerable segmentation accuracy on new classes not seen in the training stage, breaking through the limitations of low segmentation accuracy, poor generalization ability and fixed prediction classes of existing methods.

[0105] In one optional embodiment, an electronic device is also provided, as shown in Figure 5 , the electronic device 500 shown in Figure 5 may be a server, including a processor 501 and a memory 503. The processor 501 and the memory 503 are connected, such as through a bus 502. Optionally, the electronic device 500 can also include a transceiver 504. It should be noted that in actual application, the transceiver 504 is not limited to one, and the structure of the electronic device 500 does not constitute a limitation on the embodiments of the present application.

[0106] The processor 501 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It can implement or execute the various exemplary logical blocks, modules and circuits described in connection with the disclosure. The processor 501 can also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0107] The bus 502 can include a path for transmitting information between the above-mentioned components. The bus 502 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 502 can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, Figure 5 In the figure, only one thick line is used to represent the bus, but it does not mean that there is only one bus or only one type of bus.

[0108] The memory 503 can be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, an optical disk storage (including a compact disk, a laser disk, an optical disk, a digital versatile disk, a Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but not limited to this.

[0109] The memory 503 is configured to store application program codes for implementing the solutions of the present application, and the processor 501 is configured to execute the application program codes stored in the memory 503.

[0110] The electronic device includes, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (for example, a car navigation terminal), and the like, and a stationary terminal such as a digital TV, a desktop computer, and the like. Figure 5 The electronic device shown is merely an example, and should not impose any limitation on the functions and use range of the embodiments of the present application.

[0111] The server provided in the present application can be a stand-alone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs, and basic cloud computing services such as big data and artificial intelligence platforms. The terminal can be a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, and the like, but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the present application.

[0112] The embodiments of the present application provide a computer readable storage medium, which stores a computer program. When the computer program is executed on a computer, the computer can execute the corresponding content in the foregoing method embodiments.

[0113] It should be understood that, although each step in the flowchart of the accompanying drawings is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and they can be executed in other sequences. Moreover, at least part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or sub-steps or stages of other steps.

[0114] It should be noted that the computer readable storage medium in the above application can also be a computer readable signal medium or a combination of a computer readable storage medium and a computer readable signal medium. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In this application, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained in the computer readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, a RF (radio frequency) or the like, or any suitable combination of the above.

[0115] The computer readable medium described above can be contained in the electronic device described above; or can exist separately and not be assembled into the electronic device.

[0116] The computer readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.

[0117] According to an aspect of the present application, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the text interactive minimally invasive surgery instrument segmentation method and system provided in the various optional implementation manners described above.

[0118] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0119] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0120] The modules involved in the embodiments of the present application can be implemented in software or hardware. In some cases, the name of a module does not constitute a limitation on the module itself. For example, the text prompt generation module can also be described as a "text prompt generation module for generating text prompt information of a medical instrument".

[0121] The above description is merely that of the preferred embodiments of the present application and of the principles thereof. It is to be understood that the disclosed scope of protection of the present application is not limited to the specific combinations of technical features described above, but also covers other technical solutions formed by any combination of the technical features described above or equivalent features thereof, without departing from the disclosed concept. For example, the technical solutions formed by replacing the above-described features with technical features disclosed in the present application (but not limited to) having similar functions.

Claims

1. A text-interactive minimally invasive surgical instrument segmentation method, characterized in that, The method includes: Generate text prompts for the device; the text prompts include the device category name, device appearance description, and device function description; Based on the visual-language large model, feature extraction is performed on the surgical images and the text prompts to obtain visual-text feature pairs; the visual-text feature pairs include visual features and text features. The visual features are enhanced and refined based on the text features to generate mask probability maps corresponding to various types of instruments, including: The visual features and text features are decoded based on an attention mechanism to obtain attention-based text prompt features, including: Self-attention is calculated within the visual features to obtain self-attention features; Calculate the cross-attention between the self-attention visual features and the text prompt features to obtain cross-attention features; The cross-attention features are input into the feedforward network, and the attention-based text prompt features are output: For the input surgical image I, use ∈ This represents the output of the I-th layer of the image encoder; N is the sequence length of the visual features; D is the channel dimension of the text features; The text prompt features are processed using convolution operations to obtain the mask probability map: in, For convolution weights, For bias; Weighted fusion of multiple mask probability maps corresponding to different text prompts for the same device category is performed to obtain a fused mask probability map; After masking the difficult segmentation regions in the surgical image, the images are input into a preset image encoder and a corresponding decoder is added to perform image reconstruction; wherein, the difficult segmentation regions are obtained by parsing the correspondence between the fused mask probability map and the preset labels.

2. The text-interactive minimally invasive surgical instrument segmentation method according to claim 1, characterized in that, The text prompt information for the generated device includes: For different instrument categories, text prompts for the instruments are generated based on a large language model.

3. The text-interactive minimally invasive surgical instrument segmentation method according to claim 1, characterized in that, The large-scale visual-language model includes an image encoder and a text encoder; The process of extracting features from surgical images and text prompts based on a large visual-language model to obtain visual-text feature pairs includes: The visual features are obtained by extracting features from the surgical image based on the image encoder. Based on the text encoder, feature extraction is performed on the text title information to obtain the text features.

4. The text-interactive minimally invasive surgical instrument segmentation method according to claim 1, characterized in that, The process of processing the text prompt features based on convolutional operations to obtain the mask probability map includes: The text prompt features are reconstructed; The text features are converted into convolution parameters, and the reconstructed text prompt features are convolved using the convolution parameters to obtain the convolution output features; the convolution parameters include weights and parameters. The convolution output features are transformed into the mask probability map using the Sigmoid function.

5. The text-interactive minimally invasive surgical instrument segmentation method according to claim 1, characterized in that, The weighted fusion of multiple mask probability maps corresponding to different text prompts for the same device category to obtain a fused mask probability map includes: For multiple mask probability maps corresponding to different text prompts for the same device category, multiple weight mappings are output based on a vision-text threshold network. Based on the weight mapping, multiple mask probability maps are weighted and fused to obtain the fused mask probability map; The visual-text threshold network comprises multiple layers of residual blocks.

6. A text-interactive minimally invasive surgical instrument segmentation system, characterized in that, The system includes: The text prompt generation module is used to generate text prompt information for the device; the text prompt information includes the device category name, device appearance description and device function description; The encoder module is used to extract features from the surgical images and the text prompts based on a large vision-language model to obtain vision-text feature pairs; the vision-text feature pairs include visual features and text features. A mask decoder module is used to enhance and refine the visual features based on the text features, generating mask probability maps corresponding to various types of instruments, including: The visual features and text features are decoded based on an attention mechanism to obtain attention-based text prompt features, including: Self-attention is calculated within the visual features to obtain self-attention features; Calculate the cross-attention between the self-attention visual features and the text prompt features to obtain cross-attention features; The cross-attention features are input into the feedforward network, and the attention-based text prompt features are output: For the input surgical image I, use ∈ This represents the output of the I-th layer of the image encoder; N is the sequence length of the visual features; D is the channel dimension of the text features; The text prompt features are processed using convolution operations to obtain the mask probability map: in, For convolution weights, For bias; The hybrid prompt prediction fusion module is used to perform weighted fusion of multiple mask probability maps corresponding to different text prompt information of the same device category to obtain a fused mask probability map. The difficult segmentation region enhancement module is used to mask the difficult segmentation regions in the surgical image, input them into a preset image encoder, and add a corresponding decoder to perform image reconstruction; wherein, the difficult segmentation regions are obtained by parsing the correspondence between the fused mask probability map and the preset labels.

7. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method of any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Online operation video instrument tracking system based on text promptable

    CN117789921A