Transesophageal medical image analysis system and method based on vision-language large model

Through a system based on vision-language big model, intelligent analysis of transesophageal medical images is solved, and the subjectivity and high load problems of image analysis in traditional methods are achieved, achieving efficient and accurate diagnostic results.

CN120108683APending Publication Date: 2025-06-06TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411935389.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The traditional transesophageal ultrasound image analysis method relies on the experience of doctors and is highly subjective, which leads to inconsistency in diagnosis, missed diagnosis and misdiagnosis. Doctors need to invest a lot of time and energy in analyzing complex images, which can easily lead to fatigue and inaccurate judgment.

Method used

A system based on vision-language big model is adopted, and cross-modal correlation learning is carried out through natural language text generation module, text prompt processing module and multimodal big model learning module to generate accurate analysis results of transesophageal medical images.

Benefits of technology

Accurate analysis of transesophageal medical images is achieved, subtle pathological features are effectively identified, the reliability of disease discovery is improved, the risk of missed diagnosis and misdiagnosis is reduced, the work burden of doctors in the image analysis process is reduced, and the accuracy and efficiency of diagnosis is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108683A_ABST
    Figure CN120108683A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical image processing, in particular to a transesophageal medical image analysis system based on a vision-language large model, and the system comprises a natural language text generation module which is used for recognizing an input transesophageal medical image and generating text prompt data related to the image; the text prompt processing module is used for processing the text prompt data, extracting key information in the text prompt data and constructing key information base data; and the multi-modal large model learning module is used for performing cross-modal association learning on the key information base data and the transesophageal medical image and outputting an analysis result of the transesophageal medical image. The model can understand the corresponding relationship between the features in the image and the text prompts by simultaneously processing the image and the corresponding text prompts for fusion learning, so that an accurate transesophageal medical image analysis result is output, the workload of a doctor in the image analysis process is effectively reduced, and the diagnosis efficiency is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a transesophageal medical image analysis system and method based on a large vision-language model. Background Art

[0002] As an important cardiac imaging technique, transesophageal ultrasound (TEE) is widely used in the field of cardiology because of its advantages of providing high-resolution and real-time imaging. TEE can clearly display the structure and function of the heart, including valves, heart chambers and their blood flow dynamics, providing key data for the evaluation of complex cases. However, although TEE provides detailed imaging information, traditional image analysis methods still rely on the doctor's experience and skills and face many challenges.

[0003] First, the interpretation of TEE images is highly subjective, and different doctors may come to completely different judgments on the same image. This difference not only affects the uniformity of clinical decision-making, but may also lead to missed diagnosis and misdiagnosis of the disease, which in turn has an adverse effect on the patient's treatment effect. Secondly, processing complex TEE images requires doctors to invest a lot of time and energy, especially in emergency and intensive care environments, where doctors face great pressure. In these cases, the analysis and interpretation of images not only requires doctors to have solid professional knowledge, but also requires them to stay focused under high-intensity work pressure. This high-load work style can easily lead to fatigue, which in turn affects the accuracy of judgment. Summary of the invention

[0004] The purpose of the present invention is to provide a transesophageal medical image analysis system and method based on a visual-language large model, which can perform intelligent analysis on transesophageal medical images, output accurate analysis results of transesophageal medical images, reduce the workload of doctors during image analysis, and improve the accuracy and efficiency of diagnosis.

[0005] In order to solve the above technical problems, the embodiments of the present invention provide a technical solution as follows:

[0006] A transesophageal medical image analysis system based on a vision-language large model, comprising:

[0007] A natural language text generation module, used to recognize the input transesophageal medical image and generate text prompt data related to the image;

[0008] A text prompt processing module, used to process the text prompt data, extract key information from the text prompt data, and construct key information database data;

[0009] The multimodal large model learning module is used to perform cross-modal association learning on the key information library data and the transesophageal medical images, and output the analysis results of the transesophageal medical images.

[0010] Furthermore, the natural language text generation module includes a GPT model.

[0011] Furthermore, the text prompt data includes pathological descriptions and annotations of anatomical structures.

[0012] Furthermore, the text prompt processing module includes an ALBERT model, and the text prompt processing module can realize cross-layer parameter sharing between multiple Transformer layers of the ALBERT model, and extract key information of the text prompt data through the Embedding layer and the Transformer layer with cross-layer parameter sharing.

[0013] Furthermore, the multimodal large model learning module adopts the MTTR model, including: a text encoder, used to process the input key information library data and generate corresponding text features; a spatiotemporal feature extractor, used to extract visual features from transesophageal medical images; a multimodal Transformer, used to fuse text features and visual feature data, and output a predicted sequence of the processing object; an output end, used to calculate the similarity between each sequence and the key information library data based on the predicted sequence, and select the sequence with the highest similarity as the analysis result for output.

[0014] In order to solve the above technical problems, the embodiments of the present invention provide a technical solution as follows:

[0015] A transesophageal medical image analysis method based on a visual-language large model comprises the following steps:

[0016] Step S1: inputting the transesophageal medical image into a natural language text generation module to generate text prompt data related to the transesophageal medical image;

[0017] Step S2: inputting the text prompt data into a text prompt processing module, extracting key information from the text prompt data, and constructing key information database data;

[0018] Step S3: input the key information library data and the transesophageal medical images into the multimodal large model learning module for cross-modal association learning, and output the analysis results of the transesophageal medical images.

[0019] Furthermore, the step S1 includes dividing the transesophageal medical image into superior vena cava image, mid-segment horizontal image, transgastric horizontal image, and transgastric deep image according to different orientations, and generating corresponding text prompt data respectively through the GPT model.

[0020] Furthermore, the text prompt data includes pathological descriptions and annotations of anatomical structures.

[0021] Further, the step S2 includes: using the ALBERT model to pre-process the input text prompt data, including word segmentation, stop word removal, and vocabulary construction;

[0022] The Embedding layer is optimized using ALBERT’s unique factorized embedding parameterization mechanism;

[0023] Cross-layer parameter sharing between multiple Transformer layers of ALBERT;

[0024] Through the Embedding layer and the Transformer layer with cross-layer parameter sharing, key information in the text prompt data is extracted.

[0025] Further, step S3 includes adopting the MTTR model, processing the input key information library data through a text encoder, generating corresponding text features, extracting visual features from transesophageal medical images, fusing text features and visual features through a multimodal Transformer, and outputting a prediction sequence of the object, encoding and decoding the prediction sequence using a multimodal Transformer, outputting a decoded object prediction sequence, calculating the similarity between the object prediction sequence and the key information library data based on the decoded object prediction sequence, and selecting the sequence with the highest similarity as the analysis result for output.

[0026] Compared with the prior art, the transesophageal medical image analysis system and method based on the vision-language large model provided by the present invention realizes the precise segmentation of transesophageal medical images by simultaneously processing the image and the corresponding text prompts through the fusion learning of the vision and language large model. It can effectively identify subtle pathological features, such as heart valve lesions, blood flow abnormalities, etc., and can understand the correspondence between the features in the image and the text prompts, so as to output accurate transesophageal medical image analysis results, thereby improving the reliability of disease discovery, reducing the risk of missed diagnosis and misdiagnosis, effectively reducing the workload of doctors in the image analysis process, and significantly improving the accuracy and efficiency of diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0028] Figure 1Schematic diagram of the architecture of a transesophageal medical image analysis system based on a visual-language large model in an embodiment of the present invention;

[0029] Figure 2 It is a flow chart of the transesophageal medical image analysis method based on the visual-language large model in an embodiment of the present invention. DETAILED DESCRIPTION

[0030] To make the purpose, technical solutions and advantages of the present invention clearer, the following will be described in detail with reference to the accompanying drawings. However, it will be appreciated by those skilled in the art that in the various embodiments of the present invention, many technical details are provided to enable the reader to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed for protection in the claims of the present application can be implemented. The terms "comprise", "include" and any variation thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices.

[0031] In the following detailed description, many specific details are described to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that the well-known algorithms are not shown in detail to avoid blurring the main idea of ​​the present invention; and the GPT model, ALBERT model, MTTR model, Embedding layer and Transformer layer technologies involved in the following effect embodiments are prior arts that can be retrieved.

[0032] like Figure 1 As shown, one embodiment of the present invention relates to a transesophageal medical image analysis system based on a visual-language large model, comprising:

[0033] A natural language text generation module is used to identify the input transesophageal medical image and generate text prompt data related to the image; a text prompt processing module is used to process the text prompt data, extract key information from the text prompt data, and construct key information library data; a multimodal large model learning module is used to perform cross-modal association learning between the key information library data and the transesophageal medical image, and output the analysis results of the transesophageal medical image.

[0034] In one embodiment, it involves a transesophageal medical image analysis system based on a visual-language large model, wherein the natural language text generation module includes a GPT model, which recognizes the input transesophageal medical image based on the powerful text generation capability of GPT-4, and generates a series of text prompt data related to it. These text prompt data include but are not limited to pathological descriptions, such as abnormal changes in cardiac structure, changes in hemodynamics, etc., and annotations of anatomical structures, such as the relationship between atria, ventricles and their valves. In this way, it can provide doctors with important clinical background to help them better understand and analyze the image content; the GPT model used in this module can also be replaced by other artificial intelligence models, such as KIMI, Doubao, Wenxin Yiyan, etc. Doctors can make their own choices based on feedback content and convenience of use.

[0035] In one embodiment, it relates to a transesophageal medical image analysis system based on a visual-language large model, wherein the text prompt processing module adopts a BERT model, and the BERT model uses its powerful natural language understanding ability to extract and process text prompt data through a deep learning algorithm, extract key information from the text prompt data, and construct key information library data. The key information library data not only provides rich contextual information for subsequent image analysis, but also ensures the timeliness and accuracy of information through continuous updating and optimization, thereby improving the performance of the entire system. Preferably, the text prompt processing module adopts an ALBERT model, and the text prompt processing module can realize cross-layer parameter sharing between multiple Transformer layers of ALBERT, and extract key information of the text prompt data through the Embedding layer and the Transformer layer with cross-layer parameter sharing. The efficiency and performance of the model are improved through its optimized parameterization technology, ensuring that the model has high efficiency and robustness when processing text prompt data. When in use, the text prompt processing module first pre-processes the input text prompt data, including steps such as word segmentation, stop word removal, and vocabulary construction to ensure the clarity and consistency of the data. This process will be helpful for subsequent text prompt data analysis and key information extraction. Secondly, the Factorized Embedding Parameterization mechanism unique to ALBERT is used to optimize the Embedding layer. Specifically, the traditional Embedding matrix is ​​decomposed from the original size V×H (V is the vocabulary size, H is the hidden layer size), the vocabulary is first mapped to a lower dimensional space (V×E), and then this low dimensional space is mapped to the original hidden layer size (E×H). This decomposition method significantly reduces the number of parameters of the model while keeping the output dimension L×H (L is the sentence length) unchanged, especially when E is much smaller than H, the effect is significant; thirdly, the text prompt processing module realizes cross-layer parameter sharing between multiple Transformer layers of ALBERT, and multiple Transformer layers share the same parameters instead of using independent parameter sets for each layer. This strategy not only reduces the overall number of model parameters, but also promotes the flow of information between layers, helping the model learn more stable and consistent feature representations; finally, through the optimized Embedding layer and the Transformer layer with cross-layer parameter sharing, the text prompt processing module can effectively extract key information from the text prompt data and build key information library data.

[0036] In one embodiment, a transesophageal medical image analysis system based on a visual-language large model is involved, wherein the multimodal large model learning module adopts the Referring Image Segmentation (RIS) multimodal large model, in which the key information library data and transesophageal medical images are input into the RIS multimodal large model for joint learning. The model can capture the correlation between images and texts through the fusion of cross-modal features and realize intelligent processing. During the training process, the model will continuously optimize its internal parameters to improve the detection accuracy. In addition, the system can continuously learn and improve in real time according to the doctor's operation and judgment feedback. The model continuously optimizes its parameters to improve the accuracy and flexibility of image analysis, and finally realizes the intelligent recognition of complex pathological conditions, ensuring that the segmentation results outputted by it are more in line with the actual clinical needs, and realizing the accurate detection and analysis of transesophageal medical images. Preferably, the module adopts the MTTR model in RIS and applies it to the correlation between transesophageal medical images and text libraries. The model structure mainly includes a text encoder, a spatiotemporal feature extractor, a multimodal Transformer and an output end. Among them, the text encoder is used to process the input key information library data and generate corresponding text features; the spatiotemporal feature extractor is used to extract visual features from transesophageal medical images; the multimodal Transformer is used to fuse text features and visual feature data and output the predicted sequence of the processing object; the output end is used to calculate the similarity between each sequence and the key information library data based on the predicted sequence, and select the sequence with the highest similarity as the analysis result for output.

[0037] The present invention proposes a transesophageal medical image detection system based on a visual-language large model, which improves the analysis and detection capabilities of transesophageal ultrasound images by integrating advanced artificial intelligence technology; by utilizing the text prompt data generated by the natural language generation module, the key information library data processed by the text prompt processing module, and the fusion learning of the key information library data and transesophageal ultrasound images by the multimodal large model module, it can realize intelligent analysis of transesophageal medical images and output accurate analysis results of transesophageal medical images, thereby reducing the workload of doctors in the image analysis process and forming an efficient and accurate diagnostic tool to improve the speed and accuracy of doctors' disease discovery in transesophageal ultrasound detection and improve diagnostic efficiency.

[0038] like Figure 2 As shown, one embodiment of the present invention relates to a transesophageal medical image analysis method based on a visual-language macro model, comprising the following steps:

[0039] Step S1: Input the transesophageal medical image into the natural language text generation module to generate text prompt data related to the transesophageal medical image; the text prompt data includes pathological descriptions and annotations of anatomical structures. In one embodiment, the transesophageal medical images are divided into four categories according to different orientations: superior vena cava images, mid-segment horizontal images, transgastric horizontal images, and transgastric deep images. Transesophageal medical images containing the above four orientations are collected from multiple medical databases to ensure the diversity and richness of samples. The selected images need to cover different pathological conditions of different patients, including healthy hearts and various heart lesions (such as valvular heart disease, atrial fibrillation, etc.). These images are input into the natural language text generation module to generate text prompt data corresponding to the transesophageal medical images in the above four orientations. Preferably, the natural language text generation module adopts the GPT model. After the above transesophageal medical images are input into GPT, the generated text prompt data examples are as follows:

[0040] Superior vena cava image: The superior vena cava view shows the superior and posterior aspects of the heart, clearly showing the connection of the superior vena cava to the right atrium. This view is useful for observing the anatomy of the right atrium, especially when evaluating superior vena cava blood flow and its effect on right atrial function.

[0041] Mid-section horizontal image: The mid-section horizontal view provides a cross-sectional view of the middle of the heart, showing the main structures of the ventricles and atria. This view is particularly suitable for analyzing the thickness of the ventricular wall, the size of the ventricle and its contractile function, providing an important basis for the evaluation of diseases such as heart failure.

[0042] Transgastric horizontal image description: The transgastric horizontal view shows the left and back sides of the heart from a special angle, which is particularly suitable for observing the details of the left ventricle, left atrium and aortic root. This view is particularly important when evaluating left ventricular outflow tract stenosis or aortic valve disease.

[0043] Transgastric deep image description: The transgastric deep view provides a detailed view of the heart base and underlying structures, including the pericardium, inferior vena cava, and other adjacent tissues. This view is important for evaluating pericardial lesions and inferior vena cava dysfunction.

[0044] Step S2: Input the text prompt data into a text prompt processing module, extract key information from the text prompt data, and construct key information database data; in an exemplary example, the text prompt processing module adopts an ALBERT model, and the specific implementation method of the module is as follows:

[0045] First, the input text prompt data is preprocessed, including word segmentation, stop word removal, vocabulary building, etc., to ensure the clarity and consistency of the data. This process will help with subsequent text analysis and key information extraction.

[0046] Secondly, the embedding layer is optimized using ALBERT's unique factorized embedding parameterization mechanism. Specifically, the traditional embedding matrix is ​​decomposed from its original size V×H (V is the vocabulary size, H is the hidden layer size), the vocabulary is first mapped to a lower-dimensional space (V×E), and then this low-dimensional space is mapped to the original hidden layer size (E×H). This decomposition method significantly reduces the number of model parameters while keeping the output dimension L×H (L is the sentence length) unchanged, especially when E is much smaller than H, the effect is significant.

[0047] Thirdly, the text prompt processing module implements cross-layer parameter sharing between multiple Transformer layers of ALBERT. This means that multiple Transformer layers share the same parameters instead of using independent parameter sets for each layer. This strategy not only reduces the overall number of parameters of the model, but also promotes the flow of information between layers, helping the model learn more stable and consistent feature representations.

[0048] Finally, through the optimized Embedding layer and the Transformer layer with cross-layer parameter sharing, the text prompt processing module can effectively extract rich semantic features from the input text prompt data, form key information, and build key information library data. These key information libraries will provide strong support for subsequent task processing (such as classification, generation, etc.), making the analysis effect of real-time processing of transesophageal medical images more outstanding.

[0049] Step S3: Input the key information library data and the transesophageal medical image into the multimodal large model learning module for cross-modal association learning, and output the analysis results of the transesophageal medical image. The multimodal large model learning module adopts the Referring Image Segmentation (RIS) multimodal large model. In this module, the key information library data and the transesophageal medical image are input into the RIS multimodal large model for joint learning. Preferably, the MTTR model is adopted, the input key information library data is processed by a text encoder to generate corresponding text features, visual features are extracted from the transesophageal medical image, the text features and visual features are data-fused by a multimodal Transformer, and the predicted sequence of the object is output, the predicted sequence is encoded and decoded by a multimodal Transformer, and the decoded object prediction sequence is output, and the similarity between the object prediction sequence and the key information library data is calculated according to the decoded object prediction sequence, and the sequence with the highest similarity is selected as the analysis result for output. In an exemplary example, the specific method of this step is implemented as follows: the input key information library data is processed by a text encoder to generate corresponding text features and generate an embedded representation of the text, and the formula is as follows:

[0050]

[0051] Among them, T is the key information database data.

[0052] Extract visual features from transesophageal medical images, which are transesophageal dynamic ultrasound videos. By analyzing the relationship between video frames, it is possible to capture the dynamic changes and subtle structural changes of the heart, identify the heart movement state at different time points, and provide more accurate diagnostic information. The visual features of each frame are extracted using the following formula:

[0053]

[0054] Among them, V is the video frame sequence, which ensures that the extracted video features can accurately reflect the dynamically changing information.

[0055] The text features and visual features are fused through the multimodal Transformer, and the predicted sequence of the object is output. The text features and visual features are linearly projected to the same dimension and then spliced ​​into a multimodal sequence by frame. The formula is as follows:

[0056]

[0057] Use the multimodal Transformer to encode and decode these multimodal sequences and output the decoded object prediction sequence. The formula is as follows:

[0058]

[0059] Based on the decoded object prediction sequences, the similarity between these sequences and the key information database data is calculated, and the sequences with the highest similarity are selected as the analysis results for output. This process ensures that the system can accurately identify and classify pathological features in transesophageal medical images, providing doctors with efficient diagnostic support.

[0060] The transesophageal medical image analysis method based on the visual-language large model provided by the present invention realizes the accurate segmentation of transesophageal medical images by simultaneously processing the fusion learning of the image and the corresponding text prompts by the visual-language large model, and can effectively identify subtle pathological features, such as heart valve lesions, blood flow abnormalities, etc., and can understand the correspondence between the features in the image and the text prompts, so as to output accurate transesophageal medical image analysis results, thereby improving the reliability of disease discovery, reducing the risk of missed diagnosis and misdiagnosis, effectively reducing the workload of doctors in the image analysis process, and significantly improving diagnostic efficiency.

[0061] Those skilled in the art will appreciate that the above-mentioned embodiments are specific examples for implementing the present invention, and in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present invention.

Claims

1. A transesophageal medical image analysis system based on a visual-language large model, characterized in that: include: A natural language text generation module, used to recognize the input transesophageal medical image and generate text prompt data related to the image; A text prompt processing module, used to process the text prompt data, extract key information from the text prompt data, and construct key information database data; The multimodal large model learning module is used to perform cross-modal association learning on the key information library data and the transesophageal medical images, and output the analysis results of the transesophageal medical images.

2. The transesophageal medical image analysis system based on a visual-language large model according to claim 1, characterized in that: The natural language text generation module includes a GPT model.

3. The transesophageal medical image analysis system based on a visual-language large model according to claim 2, characterized in that: The text prompt data includes pathological description and annotation of anatomical structure.

4. The transesophageal medical image analysis system based on a visual-language large model according to claim 1, characterized in that: The text prompt processing module includes an ALBERT model, and the text prompt processing module can realize cross-layer parameter sharing between multiple Transformer layers of the ALBERT model, and extract key information of the text prompt data through the Embedding layer and the Transformer layer with cross-layer parameter sharing.

5. The transesophageal medical image analysis system based on a visual-language large model according to claim 1, characterized in that: The multimodal large model learning module adopts the MTTR model, including: a text encoder, used to process the input key information library data and generate corresponding text features; a spatiotemporal feature extractor, used to extract visual features from transesophageal medical images; a multimodal Transformer, used to fuse text features and visual feature data, and output a predicted sequence of the processing object; an output end, used to calculate the similarity between each sequence and the key information library data based on the predicted sequence, and select the sequence with the highest similarity as the analysis result for output.

6. A transesophageal medical image analysis method based on a visual-language large model, characterized in that: The steps include: Step S1: inputting the transesophageal medical image into a natural language text generation module to generate text prompt data related to the transesophageal medical image; Step S2: inputting the text prompt data into a text prompt processing module, extracting key information from the text prompt data, and constructing key information database data; Step S3: input the key information library data and the transesophageal medical images into the multimodal large model learning module for cross-modal association learning, and output the analysis results of the transesophageal medical images.

7. The method for transesophageal medical image analysis based on a visual-language large model according to claim 6, characterized in that: The step S1 includes dividing the transesophageal medical image into superior vena cava image, mid-segment horizontal image, transgastric horizontal image, and transgastric deep image according to different orientations, and generating corresponding text prompt data respectively through the GPT model.

8. The method for transesophageal medical image analysis based on a visual-language large model according to claim 7, characterized in that: The text prompt data includes pathological description and annotation of anatomical structure.

9. The method for transesophageal medical image analysis based on a visual-language large model according to claim 6, characterized in that: The step S2 comprises: using the ALBERT model to pre-process the input text prompt data, including word segmentation, stop word removal, and vocabulary construction; The Embedding layer is optimized using ALBERT’s unique factorized embedding parameterization mechanism; Cross-layer parameter sharing between multiple Transformer layers of ALBERT; Through the Embedding layer and the Transformer layer with cross-layer parameter sharing, key information in the text prompt data is extracted.

10. The method for transesophageal medical image analysis based on a visual-language large model according to claim 6, characterized in that: The step S3 includes adopting the MTTR model, processing the input key information library data through a text encoder, generating corresponding text features, extracting visual features from transesophageal medical images, fusing text features and visual features through a multimodal Transformer, and outputting a prediction sequence of the object, encoding and decoding the prediction sequence using a multimodal Transformer, outputting a decoded object prediction sequence, calculating the similarity between the object prediction sequence and the key information library data based on the decoded object prediction sequence, and selecting the sequence with the highest similarity as the analysis result for output.