Multi-modal reasoning segmentation to enhance language models for healthcare reasoning analysis

The multi-modal reasoning segmentation (MMRS) method addresses the limitations of existing language models by integrating text, images, and clinical records to enhance medical decision-making, improving accuracy and transparency in the medical domain.

WO2025109380A1PCT designated stage expired Publication Date: 2025-05-30NEC LAB EURO GMBH

Patent Information

Application Number
PCT/IB2024/051422
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-02-15
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing language models lack the capability to provide complex reasoning integrating multiple modalities (texts, images) effectively, especially in the medical domain, leading to misinterpretation of segmented images and lack of transparency in decision-making processes.

Method used

The proposed method employs multi-modal reasoning segmentation (MMRS) to enhance language models by integrating text, images, and clinical records. This involves tokenization, multi-modal representation learning, and joint training of text and image decoders to generate segmented images and textual explanations, thereby improving the accuracy and transparency of medical decision-making.

Benefits of technology

MMRS enhances the accuracy and efficiency of medical decision-making by providing reliable and comprehensive medical reasoning, reducing diagnostic errors, and improving treatment planning through the integration of multiple data modalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024051422_30052025_PF_FP_ABST
    Figure IB2024051422_30052025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method for improving language models includes tokenization processing to concatenate an obtained image and texts with designated tags. An image encoder is trained to encode image information of the image. A two token representation is generated by a multi-modal generative agent using the encoded image information and the obtained texts. A text encoder and the image encoder are trained to encode the obtained image and texts. A text decoder and an image decoder are combined for the multi-modal reasoning segmentation by training the text decoder and the image decoder to generate a segmented image and segmented textual information using the two token representation in combination with an output of the text encoder and an output of the image encoder. The method has applications including, but not limited to, use cases in medical Al / healthcare, for example, for optimizing medical diagnosis or treatment or supporting decision making.
Need to check novelty before this filing date? Find Prior Art

Description

MULTI-MODAL REASONING SEGMENTATION TO ENHANCE LANGUAGE MODELS FOR HEALTHCARE REASONING ANALYSISCROSS-REFERENCE TO PRIOR APPLICATION

[0001] Priority is claimed to U.S. Provisional Application Serial No. 63 / 601,814 filed on November 22, 2023, the entire contents of which is hereby incorporated by reference herein. FIELD

[0002] The present invention relates to artificial intelligence (Al) and machine learning (ML), and in particular to a method, system, data structure, computer program product and computer-readable medium for improving language models using multi-modal reasoning segmentation having applications to the medical domain and healthcare, and other domains, providing for transparency and traceability.SUMMARY

[0003] In an embodiment, the present invention provides a computer-implemented, machine learning method for improving language models using multi-modal representation learning and multi-modal reasoning segmentation. Tokenization processing is applied to concatenate an obtained image and obtained texts with designated tags. An image encoder is trained to encode image information of the image and align information for different modalities in a same latent space. A two token representation is generated by a multi-modal generative agent for the multimodal representation learning using the encoded image information and the obtained texts, the two token representation comprising a representation of multi-modal input. A text encoder and the image encoder are trained to encode the obtained image and the obtained texts. A text decoder and an image decoder are combined for the multi-modal reasoning segmentation by training the text decoder and the image decoder to generate a segmented image and segmented textual information using the two token representation in combination with an output of the text encoder and an output of the image encoder. The method has applications including, but not limited to, use cases in medical Al / healthcare, for example, for optimizing medical diagnosis or treatment or supporting decision making.BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Embodiments of the present invention will be described in even greater detail below based on the exemplary figures. The present invention is not limited to the exemplary embodiments. All features described and / or illustrated herein can be used alone or combined in different combinations in embodiments of the present invention. The features and advantages of various embodiments of the present invention will become apparent by reading the following detailed description with reference to the attached drawings which illustrate the following:

[0005] FIG. 1 schematically illustrates a method and system architecture for generating multi-modal reasoning segmentation outputs according to an embodiment of the present invention;

[0006] FIG. 2 schematically illustrates an overview of a model layer and expected output after a service layer according to an embodiment of the present invention;

[0007] FIG. 3 schematically illustrates an explainer in service layer (truthful multi-document text summarizer) according to an embodiment of the present invention; and

[0008] FIG. 4 is a block diagram of an exemplary processing system, which can be configured to perform any and all operations disclosed herein.DETAILED DESCRIPTION

[0009] Embodiments of the present invention provide truthful explanations for segmented images by combining those with segmented text information. Embodiments of the invention provide solutions the technical problem of segmented images being misinterpreted by specialists when making decisions. Embodiments of the present invention can be especially advantageously applied in domains with the need of high transparency and traceability.

[0010] The advancement of large language models (LLMs) has opened new avenues for their application in the medical domain, promising improved healthcare solutions. However, the complexity of medical decision making goes beyond what general-purpose LLMs can handle, as they do not have the technical capability of providing complex reasoning integrating from multiple modalities (texts, images) for the medical domain. To alleviate the complex reasoning obstacles for multi-model setups, previous work proposed multi-modal reasoning segmentation, which aims to mask the comparisons of two images, and provide the reasoning texts based on implicit instructions. However, this is unclear, especially for medical domain usages, and especially with different data formats like clinical records, and the technical problem of differentiating not only images, but also multiple medical document resources remains an open issue.

[0011] To address these technical challenges, embodiments of the present invention introduce a framework to enhance the utility of LLMs through Medical Multi-Modal Reasoning Segmentation (MMRS). MMRS goes beyond the conventional approaches that often focus solely on image segmentation masks or not domain-specific usage. Instead, MMRS integrates multiple data modalities, such as text, images, and clinical records into LLMs, extract and further differentiate the meaningful insights from diverse resources (both texts and images), with reliable and more comprehensive medical reasoning. This helps the practitioners to reduce diagnostic errors, streamline treatment planning, and speed up their decision making process.Ultimately, this contributes to the deployment of more precise, personalized, and efficient medical solutions.

[0012] Embodiments of the present invention provide an approach to exploit the different dimensions of multi-modal data to reason image segmentation. By combining image analysis and text processing, embodiments of the present invention aim to improve the accuracy and efficiency of LLMs and decision making in the medical field. To achieve this, embodiments of the present invention learn segmentations for different data modalities simultaneously. This is advantageous in that the segmentations describe the same area of interest so that, e.g., the text segmentation can be used to explain the image segmentation. This addresses the technical problem that the Al-based image segmentation might be misinterpreted by the medical staff, in other words addressing the technical problem of lack of explainability and not being able to determine why the specific segment was highlighted. The segmentation of text can be understood as a direct link to facts or ground truth information. Comparing multimodal data generative Al to traditional image segmentation approaches, embodiments of the present invention enable essentially improved accuracy and enhanced understanding due to the multimodal data aspect and the power of generative Al.

[0013] In a first aspect, the present invention provides a computer-implemented, machine learning method for improving language models using multi-modal representation learning and multi-modal reasoning segmentation. Tokenization processing is applied to concatenate an obtained image and obtained texts with designated tags. An image encoder is trained to encode image information of the image and align information for different modalities in a same latent space. A two token representation is generated by a multi-modal generative agent for the multimodal representation learning using the encoded image information and the obtained texts, the two token representation comprising a representation of multi-modal input. A text encoder and the image encoder are trained to encode the obtained image and the obtained texts. A text decoder and an image decoder are combined for the multi-modal reasoning segmentation by training the text decoder and the image decoder to generate a segmented image and segmented textual information using the two token representation in combination with an output of the text encoder and an output of the image encoder.

[0014] In a second aspect, the present invention provides the method according to the first aspect, wherein the segmented textual information is tokenized into words or subwords and include special tokens that are document identifiers associated with the obtained texts, to identify the origin of the segmented texts.

[0015] In a third aspect, the present invention provides the method according to the first aspect or the second aspect, further comprising obtaining a vector representation of the segmented textual information using an embedding layer.

[0016] In a fourth aspect, the present invention provides the method according to any of the first to third aspects, further comprising obtaining contextual information of source text associated with the obtained texts based on an encoder layer that uses the vector representation of the segmented textual information to generate a sequence of contextualized representations, the encoder layer including a self-attention mechanism that is configured to utilize a positional encoding of the source text.

[0017] In a fifth aspect, the present invention provides the method according to any of the first to fourth aspects, further comprising training a sequence to sequence decoder and a sequence to vector decoder using the sequence of contextualized representations.

[0018] In a sixth aspect, the present invention provides the method according to any of the first to fifth aspects, wherein the sequence to sequence decoder generates a summary of the segmented textual information, wherein the sequence to vector decoder generates a vector representation using as input the sequence of contextualized representations, and wherein the vector representation represents an original document or a list of documents associated with the obtained texts.

[0019] In a seventh aspect, the present invention provides the method according to any of the first to sixth aspects, further comprising classifying the vector representation of the sequence of contextualized representations with the original document or the list of documents using a dense layer.

[0020] In an eighth aspect, the present invention provides the method according to any of the first to seventh aspects, further comprising implementing a loss function for comparing a predicted sequence with the summary of the segmented textual information.

[0021] In a ninth aspect, the present invention provides the method according to any of the first to eighth aspects, further comprising implementing a loss function for comparing a result of the classification of the vector representation of the sequence of contextualized representations with the document identifiers.

[0022] In a tenth aspect, the present invention provides the method according to any of the first to ninth aspects, further comprising concatenating the classification of the vector representation of the sequence of contextualized representations and each sentence of the summary of the segmented textual information.

[0023] In an eleventh aspect, the present invention provides the method according to any of the first to tenth aspects, further comprising transforming the concatenation of the classificationof the vector representation of the sequence of contextualized representations and each sentence of the summary of the segmented textual information into human-readable text using the words or the subwords, the human-readable text including links to one or more documents associated with the source text.

[0024] In a twelfth aspect, the present invention provides the method according to any of the first to eleventh aspects, further comprising combining the human-readable text with the segmented image.

[0025] In a thirteenth aspect, the present invention provides the method according to any of the first to twelfth aspects, wherein the obtained image and the obtained texts are obtained from a dataset comprising the image, a text-based question, one or more multi-model data resources, and expected segmentation results as ground truth.

[0026] In a fourteenth aspect, the present invention provides a computer system for improving language models using multi-modal representation learning and multi-modal reasoning segmentation comprising one or more processors, which, alone or in combination, are configured to perform a machine learning method for improving language models using multimodal representation learning and multi-modal reasoning segmentation according to any of the first to thirteenth aspects.

[0027] In a fifteenth aspect, the present invention provides a tangible, non-transitory computer-readable medium for improving language models using multi-modal representation learning and multi-modal reasoning segmentation which, upon being executed by one or more hardware processors, provide for execution of a machine learning method according to any of the first to thirteenth aspects.

[0028] FIG. 1 provides an overview of a system architecture according to an embodiment of the present invention. In the following, individual elements of the system are introduced and described, in particular the data layer 100, the model layer 102, and the service layer 104. FIG. 2 schematically illustrates an overview of a model layer and expected output after a service layer according to an embodiment of the present invention. Components of both will be described in the following paragraphs. The data layer 100 describes the input to the system. This includes the consideration of public databases, customer databases (e.g., patient records in a hospital), and medical devices (e.g., to record the images of interest) 106. Public databases serve as valuable resources, providing a wide range of medical data for research and analysis. These databases may include anonymized patient records, clinical trial data, or publicly available medical literature. Customer databases, specifically patient records in a hospital setting, are another component of the data layer. These records contain valuable information about individual patients, their medical history, diagnoses, and treatments. By integrating this data, the systemcan consider personalized patient information, enabling more accurate and tailored analysis. Medical devices are advantageous to consider as they capture the images of interest for diagnosis and analysis. By incorporating this information, the system can learn to segment and extract relevant information and identify potential abnormalities or indications of specific medical conditions. As a minimum, the system uses the target image including a text-based question (e.g., “Does it show evidence for cancer?”) 108, related knowledge such as images combined with captions (e.g., training material for medical stuff), free -text documents including images (e.g., any kind of documents) 110, and the ground truth (e.g., the expected segmentation result of the target image) 112.

[0029] The model layer 102 constitutes a core architecture of the system, including various components. This includes the configuration for handling the tokenized input format, and the design of both encoders and decoders tailored for processing both text and image data within a multi-modal framework. The input includes a medical image required for segmentation purposes, followed by a question (e.g., please segment the tumor from the image and provide the treatment options) such as 108, and multiple sources of evidential support.

[0030] Given the multi-modal inputs, they undergo tokenization 114, employing designated tags, to distinguish between textual and image contents (e.g., <txt>, 200). Following the tokenization 114, the image data is directed towards the image encoders 116 to capture and encode relevant image information and compress the modules for standardization (see FIG. 2, Step A). The compression may involve techniques like using a perceiver resampler (see Moor, M.; Huang, Q.; Wu, S.; Yasunaga, M.; Zakka, C.; Dalmia, Y.; Reis, E.; Rajpurkam, P.; and Leskovec, J. 2023. Med-Flamingo: a Multimodal Medical Few-shot Learner. arXiv preprint arXiv:2307. 15189, hereinafter Moor et al., which is hereby incorporated by reference herein). Subsequently, the processed data proceeds through a multi-modal generative agent 118, capable of generating both image and text outputs. In embodiments, a perceiver resampler may include techniques or algorithms that use few-shot learning in the development of an Al model with a small amount of training data. The Flamingo system described in the above reference is specifically adapted to the medical domain. To facilitate the segmentation of image and text for the purpose of contradiction detection and generation, an additional submodule is introduced to enhance the quality of image and text segmentation. In the multi-modal segmentation phase (see FIG. 2, Step B), the image encoder 120 and text encoder 122 first encode both image and textual inputs, and along with the token representation of the output from the multi-modal generative agent 118: <seg_img> 202 and <seg_txt> 204, are then fed into additional image and text decoders 124 and 126 for further processing. The output 206 of the multi-modal framework is expected to be a segmented image indicating the area of required medical support, and theexplanation of text segmentation (i.e., evidential support derived from the input resources). This multi-modal output 206 serves the purpose of delivering enhanced treatment options for the patient.

[0031] The service layer 104 takes as input the output from the image decoder 124 and text decoder 126. As part of the service layer 104, the segmented text 128 is converted into a summary (explainer) 130. The text summarization task is to be truthful, and therefore LLMs are not appropriate for this task due to their technical problems in hallucinating facts. The aim is to generate a summary not only from the segmented text snippets across different documents but also to link those back to the original documents. Thus, an Al model is trained according to embodiments of the present invention for generating a text summary where each sentence is linked to one or more documents, in particular the source of the content of the sentence.

[0032] FIG. 3 illustrated the Al architecture for a multi-document summarizer according to an embodiment of the present invention. The approach according to embodiments of the present invention relies on a transformer model-based approach. Compared to existing technology, an advantage and improvement in computational functionality of the approach according to embodiments of the present invention is that the text segmentation mechanism according to embodiments of the present invention provides to not have to feed the model with long input. When the input sequence is too long, a self-attention mechanism may struggle to effectively capture long-range dependencies. Further, in contrast to existing technology, embodiments of the present invention provide to be able to train simultaneously a sequence to sequence and sequence to vector decoder, both taking the same input from the encoder layer.

[0033] The encoder layer 300 includes a self-attention mechanism and considers the positional encoding of the input text 302. The self-attention mechanism of the encoder layer 300 may be an attention mechanism for relating different positions of a single sequence in order to compute a representation of the sequence. Including the self-attention mechanism within the encoder layer 300 results in the input sequence paying attention to itself. The segmented text 302 (input) includes certain portions of “Text” with a black highlights. These specific highlights indicates the identified segments or sentence that correspond to an explanation. As described herein, the output of the multi-modal framework (e.g., FIG. 2.) is expected to be a segmented image indicated an area of the required medical support and the explanation of text segmentation (i.e., evidential support derived from the input resources). As a pre-processing step, the input text 302 is tokenized 304 into word or subwords. A subword of a word may include a word that is obtained by deleting the letters at some non-necessarily adjacent positions in the word. Each token is passed from the input layer 306 through an embedding layer 308 to obtain a vector representation. Those are passed to the encoder layer 300 to capture the contextual informationfrom the source text. In embodiments, the contextual information from a source text may include the details, background, or additional information that helps to understand the meaning or significance of the text. For example, in the sentence “Julia is bom in Heidelberg,” the word “Heidelberg” which provides geographical contextual information. In the process, the input sentences are separated with special tokens to help the model to understand the source of the sentences, therefore the special tokens are document identifiers.

[0034] The sequence to sequence (Seq2Seq) decoder 310 is responsible for deriving from the sequence of contextualized representations (output of the encoder layer 300) the summary of the segmented text. The sequence to vector (Seq2Vector) decoder 312 is responsible for deriving a vector representation from the sequence of contextualized representations which represents the original document or list of documents. Cross-attention may include a mechanism that mixes two different embedding sequences i.e., it pays attention to two sequences, including the order of the elements, at the same time. Cross-attention may be useful when going from an input sentence to an output sentence in the context of language translation. For example, the input sentence may represent one input sequence and the translation may represent the second input sequence (the two sequences may include a different number of words). The subsequent dense layer 314 classifies this vector. Softmax may be used as an activation function. An activation function may decide whether a neuron should be activated or not. This means that it will decide whether the neuron's input to the network is important or not in the process of prediction using simpler mathematical operations. Considered here is a multi-class classification scenario as the sequence of contextualized representations can originate from several documents. As the training is end-to-end, this mechanism ensures that the encoder layer 300 learns a representation function which is suitable for both deriving a text summary and linking the learned representation to the source documents. This is ensured through two loss functions (Lossl and Loss2), one for comparing the predicted sequence (i.e., text summary) with the target sequence (i.e., text summary), and one for comparing the classification result of the vector with the expected target classes (i.e., document identifiers). In both cases, a cross-entropy-based loss function is used. The loss functions are aggregated (e.g., summed up). During the training process, the contextual representation of the special tokens is removed or masked (masking layer 316) when passing the sequence to the sequence to vector decoder 312 as those can be considered as the ground truth.

[0035] In the last step, the output of both decoders is concatenated, in particular each sentence is concatenated with a citation reference (concatenate layer 318). Subsequently, the predicted tokens (e.g., a numerical vector of words as a machine internal representation of words) are transformed into human-readable text (e.g., using the initial word or sub wordvocabulary). The mechanism of being able to trace the origin of each sentence in the truthful multi-document generated summary back to the original documents is another computational improvement over existing technology provided by embodiments of the present invention.

[0036] The service layer 104 combines the output of the explainer 130 with the segmented image 132. In other words, the annotation in the image is justified with a text explanation - as seen in 206 of FIG. 2. This enables more efficient and more reliable decision making, e.g., in respect of assigning treatments.

[0037] Embodiments of the present invention thus provide for general improvements to computers in Al and machine learning systems to provide enhanced computer functionality for multi-modal reasoning segmentation, providing for transparency and explainability, and increased security, reliability, trustworthiness and accuracy of the Al and machine learning systems. Moreover, embodiments of the present invention can be practically applied to use cases to effect further improvements in a number of technical fields including, but not limited to, medical (e.g., digital medicine, personalized healthcare, Al-assisted drug or vaccine development, disease prediction, medical reports, health check analysis, medical image analysis, etc.), material development, public safety and smart cities (e.g., automated traffic or vehicle control, smart districts, smart buildings, smart industrial plants, smart agriculture, energy management, etc.). Embodiments of the present invention can be especially advantageously applied in domains that benefit from increased transparency and traceability.

[0038] In a first exemplary application of an embodiment of the present invention in the medical domain, a use case is in the context of Al-assisted health check analysis. In this use case, achieving a comprehensive and integrated examination of health data is of paramount importance for both patients and healthcare professionals. However, the current health check systems typically only provide a textual report, which is hindered by real-world technological complexities. The complexities arise from the diverse nature of health data, including medical images and textual reports. To address these complexities, the Medical Multi-Modal Reasoning Segmentation (MMRS) framework according to an embodiment of the present invention provides a solution for obtaining a thorough and integrated analysis of health data, including both images and texts. The data source in this use case includes medical image data (X-rays, MRIs, CT scans), measurements from sensors or medical devices, measurements from laboratory devices on a patient sample, health check reports containing textual information summarizing the patient's medical history, test results, and any observations made by healthcare providers during the health checkup. Application of the method according to an embodiment of the present invention takes both the input image and texts with documentation for evidential support, and returns (1) segmented image to mark the health check results, (2) segmented textsproviding reasonings of diagnoses, and (3) personalized recommendations for follow-up track. Thus, as an output, the method returns comprehensive health check assessment with integrated image segmentation, and text segmentation, with follow-up recommendations. Technicity includes integrating the method with health check systems, to enhance the computational functionalities of those systems, enabling early detection, personalized recommendations, and efficient healthcare delivery, ultimately contributing to improved health check monitoring, more comprehensive health assessment with transparency of diagnoses, and efficient healthcare delivery. Thus, the outputs can be used by the health check systems to provide diagnoses, prescriptions and / or treatments in an automated fashion.

[0039] In a second exemplary application of an embodiment of the present invention in the medical domain, a use case is in the context of Al-assisted medical image analysis, for example the analysis of medical images to determine if individuals have cancer or other diseases. This task is highly complex and requires extensive medical knowledge and expertise. It is crucial to understand that even with advanced Al -based solutions for image segmentation, mistakes can occur, and there is always a risk of misinterpreting the images. The danger lies in the possibility of misinterpreting “why” a particular area in an image was marked. Al models may highlight regions based on patterns and features they have learned from training data, but the underlying reasons for these markings may not always be apparent to the medical staff. This highlights the need for clear communication and collaboration between Al systems and human experts to ensure accurate and reliable diagnoses. The data source in this use case includes public databases, customer databases (e.g., patient records in a hospital), and medical devices (e.g., to record the images of interest). The input includes measurements taken from medical devices including images taken from imaging devices such as ultrasound equipment, X-ray equipment or other medical imaging devices. Those provide the target image, related knowledge (e.g., documents with images), and the ground truth for training the Al model. Application of the method according to an embodiment of the present invention analyses the different dimensions of multi-modal data to reason image and text segmentation. As an output, the method returns segmented images and segmented text documents. The segmented text documents explain the segmented image. Technicity includes integrating the method with mammography systems specifically designed for breast cancer screening and diagnosis with improved computational functionality providing for segmented images and explanations. These devices aid radiologists in detecting and analyzing breast abnormalities, such as tumors or calcifications, improving early detection and diagnosis. For example, the outputs can be used to identify abnormalities, or to flag images or parts of images in an automated manner, while providing explanations for the decisions, for example by integrating with the medical imaging system.

[0040] In an embodiment, the present invention provides a method for generating medical multi-modal reasoning segmentation outputs in a transparent and comprehensive manner, the method comprising the steps of1. Collect a dataset consisting of (a) image required to segment; (b) a text-based question as the instruction; (c) multi-modal data resources (e.g., images combined with captions, publications in medical field, a sensor network (e.g., in a smart hospital)); and (d) the expected segmentation results as ground truth for comparison (see FIG. 1, Data Layer).2. Train the model layer with the configuration of handling the tokenized input format, the design of both encoders and decoders tailored for segmenting both text and image data within a multi-modal framework. a. Apply the tokenization processing step to concatenate image and texts with the designated tags (e.g., , <txt>) (see FIG. 2, Model Layer). Although FIG. 2 depicts the use of tags such as , and <txt> other tags may be used. b. Train an image encoder to encode relevant image information and align different modalities information in the same latent space (see FIG. 2, Model Layer, Step A). In embodiments, the image encoder learns a representation of the image. The representation is numerical (e.g., a vector). When describing the latent space it refers to the space where the learned representation exists. In addition to the image modality, the current disclosure also describes a corresponding text modality. A representation for the text modality is also learned, and both learned representations share the same latent space, i.e., the process includes aligning the different modalities or aligning the learned representations. For example, if the image and corresponding text describe the same thing, the learned representations may be close in the latent space. c. Select a multi-modal generative agent to generate the output of both images and texts to serve as the representation of multi-modal input (see FIG. 2, Model Layer, Step A). Examples of a multi-modal generative agent may include Flamingo and VL-BERT. i. The output would be two token representation: <seg_txt>, <seg_img> which would be served as an input for text and image reasoning segmentation. d. Train a text encoder and an image encoder to encode relevant image and textual information (see FIG. 2, Model Layer, Step B). In embodiments, the image encoder at this step may be a different instance of the image encoder of step (b). e. Train a text decoder and an image decoder to segment the image and textual data (see FIG. 2, Model Layer, Step B). In embodiments the text decoder and the image decoder include a learned mapping function which may take a given input and try to leam / find a function to decode from the input the expected output (e.g., ground truthinformation). The decoders may be instantiated with a recurrent neural network (RNN) or they may be part of a decoder network of a transformer.3. Train an explainer to generate segmented textual data summaries across documents and justify the segmented image with the segmented textual output (see FIG. 3, Service Layer).4. The textual explanation from the explainer is combined with the segmented image from Step 2 to form the final output.

[0041] Embodiments of the present invention provide for the following improvements and technical advantages over existing technology:1) Combining a text and image decoder for multi-modal reasoning segmentation to gather relationships between textual and visual cues (see Steps 2d and 2e of the method above). This is enabled by training the text and image decoder simultaneously with the output of a multi-modal generative Al model additionally combined with image and text specific decoders.2) Providing a two step approach, consisting of multi-modal representation learning and multi-modal reasoning segmentation, which provides modularity and flexibility (see Steps 2a to 2e of the method above). This enables to first learn the contextualized representations from both images and texts, and the representations are channeled to the second step of learning the segmentation reasoning from different modalities.3) Providing a mechanism of being able to trace the origin of each sentence in the truthful multi -document generated summary back to the original documents. This is enabled by learning a sequence to sequence and a sequence to vector decoder simultaneously, both using the same encoder as input.

[0042] In contrast to existing technology, the Medical Multi-Modal Reasoning Segmentation (MMRS) framework according to embodiments of the present invention represents a significant leap forward in the realm of medical applications for large language models (LLMs). Existing approaches for reasoning segmentation often rely on singular data modalities (see Ivankay, A.; Rigotti, M.; and Frossard, P. 2023. DARE: Towards Robust Text Explanations in Biomedical and Healthcare Applications. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 11499-11533, which is hereby incorporated by reference herein), lack the robust integration of diverse sources (see Moor et al.), and do not have the technical capabilities or design to handle specific domain use cases such as in the medical domain (see Lai, X.; Tian, Z.; Chen, Y.; Li, Y.; Yuan, Y.; Liu, S.; and Jia, J. 2023. LISA: Reasoning Segmentation via Large Language Model. arXiv preprint arXiv:2308.00692, which is hereby incorporated by reference herein). For using multimodality (texts, images) for a reliable decision-making process, MMRS pioneers a more robust approach by seamlessly integrating textual data, images, and clinical records into LLMs. The proposedframework according to embodiments of the present invention not only (i) generates medical claims, but also (ii) differentiates meaningful insights from the medical multi-modal resources. This results in reliable MMRS outputs from LLMs.

[0043] Referring to FIG. 4, a processing system 400 can include one or more processors 402, memory 404, one or more input / output devices 406, one or more sensors 408, one or more user interfaces 410, and one or more actuators 412. Processing system 400 can be representative of each computing system disclosed herein.

[0044] Processors 402 can include one or more distinct processors, each having one or more cores. Each of the distinct processors can have the same or different structure. Processors 402 can include one or more central processing units (CPUs), one or more graphics processing units (GPUs), circuitry (e.g., application specific integrated circuits (ASICs)), digital signal processors (DSPs), and the like. Processors 402 can be mounted to a common substrate or to multiple different substrates.

[0045] Processors 402 are configured to perform a certain function, method, or operation (e.g., are configured to provide for performance of a function, method, or operation) at least when one of the one or more of the distinct processors is capable of performing operations embodying the function, method, or operation. Processors 402 can perform operations embodying the function, method, or operation by, for example, executing code (e.g., interpreting scripts) stored on memory 404 and / or trafficking data through one or more ASICs. Processors 402, and thus processing system 400, can be configured to perform, automatically, any and all functions, methods, and operations disclosed herein. Therefore, processing system 400 can be configured to implement any of (e.g., all of) the protocols, devices, mechanisms, systems, and methods described herein.

[0046] For example, when the present disclosure states that a method or device performs task “X” (or that task “X” is performed), such a statement should be understood to disclose that processing system 400 can be configured to perform task “X”. Processing system 400 is configured to perform a function, method, or operation at least when processors 402 are configured to do the same.

[0047] Memory 404 can include volatile memory, non-volatile memory, and any other medium capable of storing data. Each of the volatile memory, non-volatile memory, and any other type of memory can include multiple different memory devices, located at multiple distinct locations and each having a different structure. Memory 404 can include remotely hosted (e.g., cloud) storage.

[0048] Examples of memory 404 include a non-transitory computer-readable media such as RAM, ROM, flash memory, EEPROM, any kind of optical storage disk such as a DVD, a Blu-Ray® disc, magnetic storage, holographic storage, a HDD, a SSD, any medium that can be used to store program code in the form of instructions or data structures, and the like. Any and all of the methods, functions, and operations described herein can be fully embodied in the form of tangible and / or non-transitory machine-readable code (e.g., interpretable scripts) saved in memory 404.

[0049] Input-output devices 406 can include any component for trafficking data such as ports, antennas (i.e., transceivers), printed conductive paths, and the like. Input-output devices 406 can enable wired communication via USB®, DisplayPort®, HDMI®, Ethernet, and the like. Input-output devices 406 can enable electronic, optical, magnetic, and holographic, communication with suitable memory 406. Input-output devices 406 can enable wireless communication via WiFi®, Bluetooth®, cellular (e.g., LTE®, CDMA®, GSM®, WiMax®, NFC®), GPS, and the like. Input-output devices 406 can include wired and / or wireless communication pathways.

[0050] Sensors 408 can capture physical measurements of environment and report the same to processors 402. User interface 410 can include displays, physical buttons, speakers, microphones, keyboards, and the like. Actuators 412 can enable processors 402 to control mechanical forces.

[0051] Processing system 400 can be distributed. For example, some components of processing system 400 can reside in a remote hosted network service (e.g., a cloud computing environment) while other components of processing system 400 can reside in a local computing system. Processing system 400 can have a modular design where certain modules include a plurality of the feature s / functions shown in FIG. 4. For example, I / O modules can include volatile memory and one or more processors. As another example, individual processor modules can include read-only-memory and / or local caches.

[0052] While subject matter of the present disclosure has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive. Any statement made herein characterizing the invention is also to be considered illustrative or exemplary and not restrictive as the invention is defined by the claims. It will be understood that changes and modifications may be made, by those of ordinary skill in the art, within the scope of the following claims, which may include any combination of features from different embodiments described above.

[0053] The terms used in the claims should be construed to have the broadest reasonable interpretation consistent with the foregoing description. For example, the use of the article “a” or “the” in introducing an element should not be interpreted as being exclusive of a plurality of elements. Likewise, the recitation of “or” should be interpreted as being inclusive, such that therecitation of “A or B” is not exclusive of “A and B,” unless it is clear from the context or the foregoing description that only one of A and B is intended. Further, the recitation of “at least one of A, B and C” should be interpreted as one or more of a group of elements consisting of A, B and C, and should not be interpreted as requiring at least one of each of the listed elements A, B and C, regardless of whether A, B and C are related as categories or otherwise. Moreover, the recitation of “A, B and / or C” or “at least one of A, B or C” should be interpreted as including any singular entity from the listed elements, e.g., A, any subset from the listed elements, e.g., A and B, or the entire list of elements A, B and C.

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented method for improving language models using multi-modal representation learning and multi-modal reasoning segmentation, the computer-implemented method comprising: applying tokenization processing to concatenate an obtained image and obtained texts with designated tags; training an image encoder to encode image information of the image and align information for different modalities in a same latent space; generating, by a multi-modal generative agent for the multi-modal representation learning, a two token representation using the encoded image information and the obtained texts, the two token representation comprising a representation of multi-modal input; training a text encoder and the image encoder to encode the obtained image and the obtained texts; and combining a text decoder and an image decoder for the multi-modal reasoning segmentation by training, , the text decoder and the image decoder to generate a segmented image and segmented textual information using the two token representation in combination with an output of the text encoder and an output of the image encoder.

2. The computer-implemented method according to claim 1, wherein the segmented textual information is tokenized into words or subwords and include special tokens that are document identifiers associated with the obtained texts, to identify the origin of the segmented texts.

3. The computer-implemented method according to claims 1 or 2, further comprising obtaining a vector representation of the segmented textual information using an embedding layer.

4. The computer-implemented method according to claim 3, further comprising obtaining contextual information of source text associated with the obtained texts based on an encoder layer that uses the vector representation of the segmented textual information to generate a sequence of contextualized representations, the encoder layer including a self-attention mechanism that is configured to utilize a positional encoding of the source text.

5. The computer-implemented method according to claim 4, further comprising training, a sequence to sequence decoder and a sequence to vector decoder using the sequence of contextualized representations.

6. The computer-implemented method according to claim 5, wherein the sequence to sequence decoder generates a summary of the segmented textual information, wherein the sequence to vector decoder generates a vector representation using as input the sequence ofcontextualized representations, and wherein the vector representation represents an original document or a list of documents associated with the obtained texts.

7. The computer-implemented method according to claim 6, further comprising classifying the vector representation of the sequence of contextualized representations with the original document or the list of documents using a dense layer.

8. The computer-implemented method according to claim 7, further comprising implementing a loss function for comparing a predicted sequence with the summary of the segmented textual information.

9. The computer-implemented method according to claim 7, further comprising implementing a loss function for comparing a result of the classification of the vector representation of the sequence of contextualized representations with the document identifiers.

10. The computer-implemented method according to claim 7, further comprising concatenating the classification of the vector representation of the sequence of contextualized representations and each sentence of the summary of the segmented textual information.

11. The computer-implemented method according to claim 10, further comprising transforming the concatenation of the classification of the vector representation of the sequence of contextualized representations and each sentence of the summary of the segmented textual information into human-readable text using the words or the subwords, the human-readable text including links to one or more documents associated with the source text.

12. The computer-implemented method according to claim 11, further comprising combining the human-readable text with the segmented image.

13. The computer-implemented method according to any of the preceding claims, wherein the obtained image and the obtained texts are obtained from a dataset comprising the image, a text-based question, one or more multi-model data resources, and expected segmentation results as ground truth.

14. A computer system for improving language models using multi-modal representation learning and multi-modal reasoning segmentation, the computer system comprising one or more hardware processors which, alone or in combination, are configured to provide for execution of the following steps: applying tokenization processing to concatenate an obtained image and obtained texts with designated tags; training an image encoder to encode image information of the image and align information for different modalities in a same latent space;generating, by a multi-modal generative agent for the multi-modal representation learning, a two token representation using the encoded image information and the obtained texts, the two token representation comprising a representation of multi-modal input; training a text encoder and the image encoder to encode the obtained image and the obtained texts; and combining a text decoder and an image decoder for the multi-modal reasoning segmentation by training, the text decoder and the image decoder to generate a segmented image and segmented textual information using the two token representation in combination with an output of the text encoder and an output of the image encoder.

15. A tangible, non-transitory computer-readable medium having instructions thereon which, upon being executed by one or more processors, provide for improving language models using multi-modal representation learning and multi-modal reasoning segmentation by execution of the following steps: applying tokenization processing to concatenate an obtained image and obtained texts with designated tags; training an image encoder to encode image information of the image and align information for different modalities in a same latent space; generating, by a multi-modal generative agent for the multi-modal representation learning, a two token representation using the encoded image information and the obtained texts, the two token representation comprising a representation of multi-modal input; training a text encoder and the image encoder to encode the obtained image and the obtained texts; and combining a text decoder and an image decoder for the multi-modal reasoning segmentation by training, the text decoder and the image decoder to generate a segmented image and segmented textual information using the two token representation in combination with an output of the text encoder and an output of the image encoder.

Citation Information

Patent Citations

  • Data volume sculptor for deep learning acceleration

    US62636018P0

Cited By

  • Glioma postoperative radiotherapy prediction method based on iconomics of multi-modal MRI (Magnetic Resonance Imaging)

    CN120413056A

  • Inference model training method and device, server, storage medium and product

    CN120525015A

  • Medical multi-modal model training method and device, electronic equipment and storage medium

    CN120853191A

  • Medical text-oriented interpretable high-precision classification model and attribution analysis method

    CN121980363A