A method for intelligently describing the ultrasound image content of liver space-occupying lesions using LLM

Through multimodal data fusion and intelligent description technology, the noise interference and multimodal data fusion problems of liver space-occupying lesions ultrasound images are solved, high-quality and fluent descriptive text is generated, the accuracy and practicality of diagnosis are improved, and real-time optimization is supported.

CN120543547BActive Publication Date: 2025-10-03THE FIRST AFFILIATED HOSPITAL OF WENZHOU MEDICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511036815.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-10-03
Estimated Expiration
2045-07-28

AI Technical Summary

Technical Problem

The existing intelligent description technology of liver space-occupying lesions ultrasound images has problems such as excessive noise interference, complex artifacts, difficulty in multimodal data fusion, poor quality of generated text, and lack of real-time interaction and optimization capabilities, which leads to highly subjective diagnostic results and high misdiagnosis and missed diagnosis rates.

Method used

It adopts multimodal data acquisition, cross-modal contrastive learning network, dynamic context-aware decoder, uncertainty calibration module and real-time interactive optimization mechanism to generate accurate and fluent descriptive text through multimodal data fusion, feature extraction, text generation and optimization.

Benefits of technology

It achieves efficient fusion of multimodal information, improves the accuracy and reliability of descriptions, reduces misdiagnosis and missed diagnoses, enhances the medical professionalism and clinical practicality of the text, and supports real-time optimization and improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543547B_ABST
    Figure CN120543547B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of medical image processing technology, and discloses a method for intelligently describing the content of ultrasound images of liver space-occupying lesions using LLM. A multimodal data acquisition module is used to obtain a liver ultrasound image sequence, a patient's historical medical record text, and a blood biochemical index vector. The features are extracted and generated into an embedding vector through a cross-modal contrast learning network, and the embedding vectors are aligned. The aligned image embedding vector is input into a dynamic context-aware decoder, and a hierarchical multi-head attention mechanism is used to generate a semantic tag sequence describing the text. The confidence is evaluated through an uncertainty calibration module, and when it is lower than the threshold, the text is optimized through a post-processing reordering mechanism. A real-time interactive optimization mechanism is also provided to update the model based on doctor feedback. The method integrates multimodal data, improves the accuracy and reliability of the description, optimizes the text quality, adapts to clinical needs, and assists in the diagnosis of liver diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a method for intelligently describing the contents of ultrasound images of liver space-occupying lesions using LLM. Background Art

[0002] Liver space-occupying lesions are a type of disease that manifests as an abnormal mass on liver imaging examinations. Their accurate diagnosis is crucial for clinical treatment decision-making. Ultrasound examination, a widely used method for screening and diagnosing liver diseases, offers advantages such as ease of use, lack of radiation, and real-time dynamic observation. However, the interpretation of ultrasound images is highly dependent on the physician's experience and expertise. Different physicians may differ in their description and diagnosis of the same image, leading to a high degree of subjectivity in the diagnostic results.

[0003] In the traditional liver ultrasound diagnostic process, doctors spend a considerable amount of time observing ultrasound images, identifying lesion characteristics such as location, morphology, size, and echogenicity, and then making a comprehensive assessment based on the patient's clinical information. This process is not only inefficient but also susceptible to factors such as physician fatigue and mood. For inexperienced physicians, accurately describing ultrasound images of liver space-occupying lesions and making a reliable diagnosis is even more difficult, potentially leading to missed diagnoses and misdiagnoses, delaying treatment.

[0004] The rapid development of artificial intelligence technology, especially the widespread application of deep learning in medical image analysis, has provided new solutions to the difficult problem of interpreting liver ultrasound images. However, the current intelligent description technology for ultrasound images of liver space-occupying lesions still faces many challenges. On the one hand, ultrasound images are subject to noise interference, numerous artifacts, and complex anatomical structures, making it difficult to accurately extract lesion features. On the other hand, existing intelligent description models often have difficulty in effectively fusing multimodal data, such as ultrasound images with historical medical records and blood biochemical indicators, resulting in a lack of comprehensiveness and accuracy in the description results.

[0005] Furthermore, existing models lack effective mechanisms for evaluating and optimizing text quality when generating descriptive text. The generated text may suffer from semantic ambiguity, logical incoherence, and inaccurate use of medical terminology, failing to meet the actual needs of clinical diagnosis. Furthermore, in actual clinical application, models must possess real-time interaction and continuous optimization capabilities, enabling timely adjustments and improvements based on physician feedback. However, most current models lack these capabilities. Therefore, it is urgent to develop an efficient, accurate, and clinically practical method for intelligently describing the content of liver space-occupying lesions in ultrasound images using LLM. Summary of the Invention

[0006] The purpose of the present invention is to provide a method for intelligently describing the content of ultrasound images of liver space-occupying lesions using LLM, so as to solve the problems raised in the above background technology.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a method for intelligently describing the content of ultrasound images of liver space-occupying lesions using LLM, the method comprising:

[0008] Acquiring a liver ultrasound image sequence and associated clinical text data through a multimodal data acquisition module, the multimodal data including real-time ultrasound image frames, patient history text, and blood biochemical index vectors; performing feature extraction on the ultrasound image frames based on a cross-modal contrastive learning network to generate an image embedding vector, and encoding the clinical text data into a text embedding vector using a pre-trained language model; constructing a joint embedding space, and aligning the image embedding vector with the text embedding vector using a cosine similarity loss function;

[0009] The aligned image embedding vector is input into a dynamic context-aware decoder, which uses a hierarchical multi-head attention mechanism to dynamically assign attention weights based on the segmentation mask of the image region, where the first layer of attention focuses on the lesion edge features and the second layer of attention integrates global anatomical structure information; a semantic tag sequence describing the text is iteratively generated through a gated recurrent unit;

[0010] The generated semantic tags are evaluated for confidence based on the uncertainty calibration module, and the detailed granularity of the generated text is dynamically adjusted according to the confidence threshold. When the confidence is lower than the preset threshold, the post-processing re-ranking mechanism is activated, and the fluency of the candidate descriptions is optimized through the n-gram language model.

[0011] Preferably, the cross-modal contrastive learning network includes:

[0012] A dual-tower neural network structure was constructed. The image encoding tower used a modified ResNet-50 architecture, removing the final fully connected layer and replacing it with a deformable convolutional module to capture the irregular morphology of lesions in ultrasound images. The text encoding tower used the RoBERTa model and added an adapter layer to adapt to medical terminology.

[0013] A momentum contrast loss function is introduced into the joint embedding space. The embedding vectors of historical batch samples are cached in a queue to calculate the contrast loss of positive and negative sample pairs. The positive sample pairs are composed of ultrasound images of the same case and their associated text, while the negative sample pairs are generated by randomly sampling data from different cases.

[0014] An adaptive temperature coefficient is used to adjust the difficulty of contrastive learning. The temperature coefficient is dynamically adjusted according to the similarity distribution of the current batch of samples, and the coefficient change is smoothed by an exponential moving average update strategy.

[0015] Preferably, the hierarchical multi-head attention mechanism of the dynamic context-aware decoder includes:

[0016] The ultrasound image is segmented into 256×256 pixel local regions, and lesion candidate boxes are generated through the region proposal network. Multi-scale feature maps are extracted for each candidate box, and the first-layer attention module adopts a channel-weighted aggregation strategy to calculate the correlation score between each channel feature map and the lesion description keywords.

[0017] The second-layer attention module constructs a global anatomical topology map, models the spatial relationship between liver vascular distribution and lesions through a graph convolutional network, and generates an anatomical context feature vector;

[0018] In the decoding stage, the residual attention connection mechanism is adopted to weightedly splice the attention outputs of different levels, and the information flow transmission ratio is controlled by the gated linear unit.

[0019] Preferably, the implementation of the uncertainty calibration module includes:

[0020] Output a confidence estimate vector at each time step of the decoder. The vector contains word-level and sentence-level uncertainty indicators. The word-level indicator is calculated by the entropy value of the token prediction probability, and the sentence-level indicator is estimated by the variance of the hidden state output by the bidirectional GRU.

[0021] An uncertainty-guided beam search strategy is constructed. During the search process with a beam width of k, the retention weight of candidate sequences is dynamically adjusted according to the confidence index. When the confidence of three consecutive tokens is detected to be lower than the threshold, a local backtracking mechanism is triggered to regenerate the previous text.

[0022] The final generated description text is subjected to entity verification based on the medical knowledge graph, and the detected contradictory entities are semantically corrected through the graph reasoning module.

[0023] Preferably, the post-processing reordering mechanism includes:

[0024] Build a multi-candidate generation pipeline to generate n sets of differential description texts in parallel by adjusting the decoding temperature parameters;

[0025] Perform semantic analysis on each text set, extract key medical entities and construct an entity relationship graph. Use a graph matching algorithm to calculate the topological similarity with the standard diagnostic report.

[0026] An integrated ranking model is used to comprehensively score the candidate texts. The model integrates three indicators: vocabulary overlap, entity accuracy, and syntactic complexity, and generates the final optimized description text through a gradient boosting decision tree.

[0027] Preferably, the configuration of the deformable convolution module includes:

[0028] In the third and fourth stages of ResNet-50, a deformable convolution layer is inserted, and each deformable convolution kernel is set with 9 offset sampling points;

[0029] The offset learning is guided by the segmentation mask of the lesion area. A dual-path supervision strategy is adopted in the training stage. The main path calculates the image classification loss, and the auxiliary path calculates the L2 regularization loss between the offset field and the true segmentation boundary.

[0030] A dynamic kernel selection mechanism is enabled during the inference phase to automatically select a deformable convolution kernel size of 3×3 or 5×5 based on the texture complexity of the input image.

[0031] Preferably, the method for constructing the anatomical topology map includes:

[0032] Extract the three-dimensional coordinate template of liver blood vessels from the standard medical atlas and map the template to the current ultrasound image coordinate system using a non-rigid registration algorithm;

[0033] An anatomical structure recognition network based on supervoxel segmentation was constructed, 3D sparse convolution was used to extract vascular branch features, and an iterative clustering algorithm was used to generate vascular centerline topology.

[0034] Three types of nodes are defined in the topological graph: lesion core area, main vascular intersection, and liver lobe boundary marker. The edge weight is determined by the Euclidean distance and blood flow direction.

[0035] Preferably, the construction of the medical knowledge graph includes:

[0036] Entity-relationship triplets were extracted from clinical guideline literature to establish a liver disease ontology library, which contains 7 major entity types and 23 relationship types.

[0037] A dual attention graph neural network is used for knowledge representation learning, the entity embedding layer adopts the TransE algorithm, and the relationship embedding layer adopts the rotation encoding strategy;

[0038] A dynamic entity linking mechanism is set up. When an entity mention is detected in the generated text, the associated pathological features, treatment plans, and prognostic indicator information in the atlas are synchronously queried and injected into the attention context window of the reranking model.

[0039] Preferably, the training method of the integrated sorting model includes:

[0040] Construct a multi-dimensional feature engineering project to extract word-level features, sentence-level features, and paragraph-level features of the candidate text. Word-level features include medical term coverage and negation density, while sentence-level features include dependency tree depth and semantic role labeling completeness.

[0041] A stratified sampling strategy was used to construct the training set, with stratification based on lesion type, ultrasound equipment model, and description length;

[0042] A hybrid loss function is designed to combine ranking loss and classification loss. The ranking loss uses the Listwise loss function, and the classification loss uses the focal loss function to deal with the problem of category imbalance.

[0043] Preferably, the method further includes a real-time interactive optimization mechanism:

[0044] During the ultrasound examination, descriptive text is displayed synchronously, and correction annotations are collected through the doctor feedback interface;

[0045] Build an online incremental learning framework that encodes the corrected labeled data into comparative learning samples and updates model parameters through an elastic weight consolidation algorithm, preserving important parameters while preventing catastrophic forgetting.

[0046] Set up a model version rollback mechanism. When a user corrects the same type of error five times in a row, it automatically triggers the loading and fine-tuning of the historical optimal model.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] The method proposed in this invention for intelligently describing the ultrasound image content of liver space-occupying lesions using LLM has significant beneficial effects in many aspects. In terms of data processing, a multimodal data acquisition module acquires liver ultrasound image sequences, patient historical medical records, and blood biochemical index vectors, achieving the fusion of multi-source information. This approach overcomes the limitations of single-modality data and enables the model to analyze lesion characteristics from a more comprehensive perspective. For example, combined with the patient's historical medical records, the model can understand the patient's past medical history and treatment status, assisting in determining the nature of the current lesion; blood biochemical indicators can provide information about liver function and metabolism, further enriching the understanding of the lesion. This fusion of multimodal data greatly improves the accuracy and reliability of the description and reduces misdiagnosis and missed diagnoses due to incomplete information.

[0049] Cross-modal contrastive learning networks play a key role in feature extraction and vector generation. The image encoding tower uses an improved ResNet-50 architecture and introduces a deformable convolution module, which can accurately capture the irregular shapes of lesions in ultrasound images. Compared with traditional convolution, deformable convolution can more flexibly adapt to the complex shape changes of lesions by setting multiple offset sampling points and guiding learning based on the lesion area segmentation mask, thereby improving the feature extraction ability of small and irregular lesions. The text encoding tower utilizes the RoBERTa model and adds an adapter layer to adapt to medical terminology, enhancing the understanding and encoding capabilities of medical text. At the same time, the application of momentum contrast loss function and adaptive temperature coefficient optimizes the alignment of image and text vectors in the joint embedding space, enabling the model to better associate multimodal information and laying a solid foundation for generating accurate descriptive text.

[0050] The hierarchical multi-head attention mechanism of the dynamic context-aware decoder is a major innovation of the present invention. The first layer of attention focuses on the edge features of the lesion, which can accurately outline the boundaries of the lesion and provide key information for describing the morphology of the lesion. The second layer of attention integrates global anatomical structure information. By constructing an anatomical topology map, a comprehensive analysis is performed from aspects such as the spatial relationship between the liver vascular distribution and the lesion, so that the generated description text not only contains the characteristics of the lesion itself, but also reflects its relationship with the surrounding anatomical structure, which is more in line with the actual needs of clinical diagnosis. This hierarchical attention mechanism can dynamically allocate attention weights, effectively integrate information at different levels, and significantly improve the quality and clinical practicality of the description text.

[0051] The uncertainty calibration module and post-processing re-ranking mechanism further ensure the accuracy and fluency of the descriptive text. The uncertainty calibration module evaluates the confidence of the generated semantic tags through vocabulary-level and sentence-level uncertainty indicators. When the confidence is lower than the preset threshold, the post-processing re-ranking mechanism is triggered. The multi-candidate generation pipeline generates multiple groups of differential descriptive texts by adjusting the decoding temperature parameters, and then optimizes the candidate texts from multiple dimensions through semantic parsing, entity relationship graph construction and integrated ranking model comprehensive scoring. The integrated ranking model integrates indicators such as vocabulary overlap, entity accuracy, and syntactic complexity, and uses a gradient boosting decision tree to generate the final optimized descriptive text, effectively solving problems such as semantic ambiguity and logical incoherence in the generated text, and improving the readability and medical professionalism of the text.

[0052] The real-time interactive optimization mechanism makes the present invention more adaptable and practical in clinical applications. The generated descriptive text is displayed synchronously during the ultrasound examination process, which is convenient for doctors to view and provide feedback and corrections in real time. The online incremental learning framework encodes the doctor's revised annotation data into comparative learning samples and updates the model parameters through the elastic weight consolidation algorithm. It can not only learn new knowledge but also prevent catastrophic forgetting, so that the model can be continuously optimized and improved. When the same type of error occurs continuously, the model version rollback mechanism automatically loads the historical optimal model and performs fine-tuning training to ensure that the model can maintain high accuracy and stability in the face of complex situations, and continue to provide reliable diagnostic support for doctors. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 This is a diagram showing the working principle of the method for intelligently describing the content of liver space-occupying lesions ultrasound images using LLM according to the present invention;

[0054] Figure 2 This is a schematic diagram of the workflow of the uncertainty calibration module;

[0055] Figure 3 Schematic diagram of the working principle of the deformable convolution module;

[0056] Figure 4 Schematic diagram of the construction and application of medical knowledge graph. DETAILED DESCRIPTION

[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0058] See also Figure 1-4 The present invention provides a method for intelligently describing the content of liver space-occupying lesions ultrasound images using LLM, and its overall implementation method includes:

[0059] Multimodal Data Acquisition: The multimodal data acquisition module collects liver ultrasound image sequences and associated clinical text data. This multimodal data includes real-time ultrasound image frames, patient medical history text, and blood biochemical marker vectors. This comprehensive data collection provides a rich information foundation for subsequent analysis and processing.

[0060] Feature Extraction and Vector Generation: Using a cross-modal contrastive learning network, features are extracted from acquired ultrasound image frames to generate image embedding vectors. Simultaneously, a pre-trained language model is used to encode clinical text data into text embedding vectors. This step transforms data from different modalities into a unified vector representation, facilitating subsequent processing and fusion.

[0061] Constructing a joint embedding space and vector alignment: A joint embedding space is constructed, and the cosine similarity loss function is used to align the image and text embedding vectors. This approach can shorten the distance between the image and text vectors of the same case in the joint space, allowing for better fusion of their information and providing a more accurate basis for subsequent description generation.

[0062] Descriptive Text Generation: The aligned image embedding vectors are fed into a dynamic, context-aware decoder. This decoder employs a hierarchical multi-head attention mechanism to dynamically assign attention weights based on the segmentation mask of the image region. The first layer of attention focuses on lesion edge features, while the second layer incorporates global anatomical structure information. A gated recurrent unit then iteratively generates a sequence of semantically labeled descriptive text.

[0063] Uncertainty Calibration and Post-Processing: The uncertainty calibration module assesses the confidence of generated semantic tags and dynamically adjusts the granularity of the generated text based on a set confidence threshold. When the confidence falls below the preset threshold, a post-processing re-ranking mechanism is activated. The n-gram language model is used to optimize the fluency of candidate descriptions to improve the quality and accuracy of the generated text.

[0064] The implementation of the present invention will be further described below with reference to Examples 1 to 5.

[0065] Embodiment 1:

[0066] In this embodiment, the specific implementation of the cross-modal contrastive learning network and the deformable convolution module is further explained.

[0067] The cross-modal contrastive learning network is implemented by building a dual-tower neural network structure. The image coding tower adopts an improved ResNet-50 architecture. In order to more accurately capture the irregular morphology of lesions in ultrasound images, the final fully connected layer is removed and replaced with a deformable convolution module. The deformable convolution module inserts the deformable convolution layer in the third and fourth stages of ResNet-50. Each deformable convolution kernel is set with 9 offset sampling points. The coordinates of the offset sampling points are set as , ,The setting of these sampling points helps the model adaptively learn the irregular shape of the lesion.

[0068] During the training phase, the offset learning is guided by the lesion region segmentation mask, and a dual-path supervision strategy is adopted. The main path calculates the image classification loss , the auxiliary path calculates the offset field and the true segmentation boundary Regularization loss , the total loss function ,in Is the balance coefficient, used to adjust the weight between the two losses. In the inference phase, the dynamic kernel selection mechanism is enabled to automatically select according to the texture complexity of the input image. or For example, when the image texture complexity index is Greater than a threshold When selecting Convolution kernel, otherwise choose This can improve computational efficiency while ensuring model performance.

[0069] The text encoding tower uses the RoBERTa model and adapts to medical terminology by adding an adapter layer. The momentum contrast loss function is introduced in the joint embedding space. The embedding vectors of historical batch samples are cached in the queue to calculate the contrast loss of positive and negative sample pairs. Let the positive sample pair be , the negative sample pair is ( ), the contrast loss function is:

[0070]

[0071] in, is the number of positive sample pairs, is the total number of sample pairs, Represents the calculation similarity function, The temperature coefficient is used to adjust the difficulty of contrastive learning. The temperature coefficient is dynamically adjusted according to the similarity distribution of the current batch of samples, and the exponential moving average update strategy is used to smooth the coefficient changes.

[0072] Example 2:

[0073] This embodiment describes in detail the hierarchical multi-head attention mechanism of the dynamic context-aware decoder and the method for constructing the anatomical topology map.

[0074] The hierarchical multi-head attention mechanism of the dynamic context-aware decoder first segments the ultrasound image into The local area of ​​the pixel is used to generate the lesion candidate frame through the region candidate network. Multi-scale feature maps are extracted for each candidate frame. The first layer of attention module adopts the channel weighted aggregation strategy to calculate the correlation score between each channel feature map and the lesion description keyword. Let the channel feature map be , the keyword vector is , the correlation score calculation formula is:

[0075]

[0076] in, is the dimension of the feature map and keyword vector, and the score can be used to highlight the channel features related to lesion description.

[0077] The second-layer attention module constructs a global anatomical topology map. The three-dimensional coordinate template of the liver blood vessels is extracted from the standard medical atlas, and the template is mapped to the current ultrasound image coordinate system through a non-rigid registration algorithm. An anatomical structure recognition network based on supervoxel segmentation is constructed, and vascular branch features are extracted using three-dimensional sparse convolution. The vascular centerline topology is generated through an iterative clustering algorithm. Three types of nodes are defined in the topology map: lesion core area, main vascular intersection, and liver lobe boundary marker. The edge weight is determined by the Euclidean distance. and blood flow direction Jointly decided, the edge weight calculation formula is:

[0078]

[0079] in, is a weight coefficient used to balance the influence of Euclidean distance and blood flow direction on edge weights. During the decoding phase, a residual attention connection mechanism is used to weightedly concatenate attention outputs from different levels. Gated linear units are used to control the information flow ratio, thereby better integrating information from different levels and generating more accurate description text.

[0080] Example 3:

[0081] The uncertainty calibration module outputs a confidence estimate vector at each time step of the decoder, which contains word-level and sentence-level uncertainty indicators. The word-level indicator is calculated by the entropy value of the token prediction probability. Let the token prediction probability distribution be , vocabulary-level uncertainty indicator The calculation formula is:

[0082]

[0083] in, is the vocabulary size. The sentence-level indicator is estimated by the variance of the hidden state output by the bidirectional GRU. Let the hidden state output by the bidirectional GRU be , statement-level uncertainty index The calculation formula is:

[0084]

[0085] in, is the dimension of the hidden state, is the mean of the hidden states.

[0086] Construct an uncertainty-guided beam search strategy with a beam width of During the search process, the retention weight of candidate sequences is dynamically adjusted based on the confidence index. When the confidence of three consecutive tags is detected to be below the threshold, a local backtracking mechanism is triggered to regenerate the previous text. The final description text is then subject to entity verification based on the medical knowledge graph.

[0087] The medical knowledge graph was constructed by extracting entity-relationship triplets from clinical guideline literature to establish a liver disease ontology library, which contains seven major entity types and 23 relationship types. A dual-attention graph neural network was used for knowledge representation learning, with the TransE algorithm employed in the entity embedding layer and a rotation encoding strategy in the relationship embedding layer. A dynamic entity linking mechanism was implemented. When an entity mention was detected in the generated text, the graph was simultaneously queried for associated pathological features, treatment options, and prognostic indicators, and the information was injected into the attention context window of the reranking model, thereby improving the medical accuracy and reliability of the descriptive text.

[0088] Example 4:

[0089] The post-processing re-ranking mechanism builds a multi-candidate generation pipeline, which is generated in parallel by adjusting the decoding temperature parameters. Perform semantic analysis on each text group, extract key medical entities and construct entity relationship graph, and calculate the topological similarity with the standard diagnosis report through graph matching algorithm. and entity relationship diagram of standard diagnostic report , the topological similarity calculation formula is:

[0090]

[0091] in, and are the edge sets of the two graphs respectively.

[0092] An integrated ranking model is used to comprehensively score candidate texts. The model integrates three indicators: vocabulary overlap, entity accuracy, and syntactic complexity. A multi-dimensional feature engineering is constructed to extract word-level features, sentence-level features, and paragraph-level features of candidate texts. Word-level features include medical term coverage and negation density, and sentence-level features include dependency tree depth and semantic role labeling completeness. A stratified sampling strategy is used to construct a training set, which is divided into layers according to lesion type, ultrasound equipment model, and description length. A hybrid loss function is designed, combining ranking loss and classification loss, where the ranking loss uses the Listwise loss function. , the classification loss uses the focal loss function To deal with the problem of class imbalance, the total loss function is:

[0093]

[0094] in, is a weight coefficient used to balance the ranking loss and classification loss. The final optimized description text is generated through the gradient boosting decision tree to improve the quality and accuracy of the description text.

[0095] Example 5:

[0096] In the actual operation of ultrasound examinations, real-time interactive optimization mechanisms play a key role. When a doctor uses ultrasound equipment to examine a patient's liver, the system simultaneously displays descriptive text based on the current ultrasound image on the user interface. This text is the initial result of a series of steps described above, including multimodal data acquisition, feature extraction, vector alignment, descriptive text generation, uncertainty calibration, and post-processing.

[0097] To make the generated descriptive text more in line with actual clinical needs, doctors can operate it through a specially designed feedback interface. The feedback interface is simple and intuitive, allowing doctors to easily mark and correct errors, missing information, or inaccurate statements in the generated text. For example, if the description of the lesion size in the generated text does not match the actual ultrasound image measurement results, the doctor can directly modify the value in the text and add a note to explain the correct measurement basis; if the text omits certain key liver anatomical structure information, the doctor can complete it and mark the importance of this information.

[0098] The system builds an online incremental learning framework, the main function of which is to convert the doctor's revised annotation data into comparative learning samples that the model can learn. Assume that the revised annotation data set is , each piece of corrected annotation data is recorded as ( , To correct the number of labeled data). The encoding process will Processing, using a specific encoding function , converting it into a contrastive learning sample ,Right now These comparative learning samples contain rich information, such as the differences before and after text correction, key medical knowledge points added by doctors, etc., which can help the model better understand its own mistakes and learn the correct description method.

[0099] When updating model parameters, the elastic weight consolidation algorithm is used. The core purpose of this algorithm is to retain important parameters in previous training when learning new correction data, to prevent the model from forgetting the useful knowledge learned previously due to excessive learning of new data, that is, to avoid catastrophic forgetting. For each parameter in the model ( , is the total number of model parameters), the algorithm calculates its importance score This score is determined based on the degree of influence of the parameter on the model performance during the previous training process. The formula for calculating the importance score is:

[0100]

[0101] in, Represents the loss function of the model in the previous training process. When updating the parameters, the parameters will be adjusted according to this importance score. The update formula is:

[0102]

[0103] in, is a parameter The old value of is the new value after update, is the learning rate, which controls the step size of parameter updates; is the loss function currently calculated based on the corrected labeled data; is the weight decay coefficient, which is used to balance the degree of retaining old knowledge and learning new knowledge; It is a reference parameter value determined in previous training, usually the parameter value of the model in a certain stable state.

[0104] In addition, the system has also set up a model version rollback mechanism. When the system detects that the doctor has corrected the same type of error five times in a row, this mechanism will be automatically triggered. For example, if the five consecutive corrections are all focused on the incorrect description of a specific liver space-occupying lesion (such as hepatic hemangioma), the system will load the historical optimal model. The historical optimal model is the best performing model version determined during the previous training process based on a series of evaluation indicators (such as the similarity between the generated text and the standard diagnostic report, clinician's evaluation feedback, etc.). After loading, the model will be fine-tuned using the currently accumulated corrected annotation data.

[0105] During fine-tuning training, a smaller learning rate is used to avoid over-adjusting model parameters, ensuring that the model can learn new correction information while retaining its original good performance, thereby further improving the accuracy of the description of such lesions and better assisting doctors in diagnosing liver space-occupying lesions.

[0106] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0107] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for intelligently describing the content of liver space-occupying lesions ultrasound images using LLM, characterized in that: include: Acquiring a liver ultrasound image sequence and associated clinical text data through a multimodal data acquisition module, the multimodal data including real-time ultrasound image frames, patient history text, and blood biochemical index vectors; performing feature extraction on the ultrasound image frames based on a cross-modal contrastive learning network to generate an image embedding vector, and encoding the clinical text data into a text embedding vector using a pre-trained language model; constructing a joint embedding space, and aligning the image embedding vector with the text embedding vector using a cosine similarity loss function; The aligned image embedding vector is input into a dynamic context-aware decoder, which uses a hierarchical multi-head attention mechanism to dynamically assign attention weights based on the segmentation mask of the image region, where the first layer of attention focuses on the lesion edge features and the second layer of attention integrates global anatomical structure information; a semantic tag sequence describing the text is iteratively generated through a gated recurrent unit; The generated semantic tags are evaluated for confidence based on the uncertainty calibration module, and the detailed granularity of the generated text is dynamically adjusted according to the confidence threshold. When the confidence is lower than the preset threshold, the post-processing re-ranking mechanism is activated, and the fluency of the candidate descriptions is optimized through the n-gram language model.

2. The method for intelligently describing the content of liver space-occupying lesions ultrasound images using LLM according to claim 1, characterized in that: The cross-modal contrastive learning network includes: A dual-tower neural network structure was constructed. The image encoding tower used a modified ResNet-50 architecture, removing the final fully connected layer and replacing it with a deformable convolutional module to capture the irregular morphology of lesions in ultrasound images. The text encoding tower used the RoBERTa model and added an adapter layer to adapt to medical terminology. A momentum contrast loss function is introduced into the joint embedding space. The embedding vectors of historical batch samples are cached in a queue to calculate the contrast loss of positive and negative sample pairs. The positive sample pairs are composed of ultrasound images of the same case and their associated text, while the negative sample pairs are generated by randomly sampling data from different cases. An adaptive temperature coefficient is used to adjust the difficulty of contrastive learning. The temperature coefficient is dynamically adjusted according to the similarity distribution of the current batch of samples, and the coefficient change is smoothed by an exponential moving average update strategy.

3. The method for intelligently describing the content of liver space-occupying lesions ultrasound images using LLM according to claim 1, characterized in that: The hierarchical multi-head attention mechanism of the dynamic context-aware decoder includes: The ultrasound image is segmented into 256×256 pixel local regions, and lesion candidate boxes are generated through the region proposal network. Multi-scale feature maps are extracted for each candidate box, and the first-layer attention module adopts a channel-weighted aggregation strategy to calculate the correlation score between each channel feature map and the lesion description keywords. The second-layer attention module constructs a global anatomical topology map, models the spatial relationship between liver vascular distribution and lesions through a graph convolutional network, and generates an anatomical context feature vector; In the decoding stage, the residual attention connection mechanism is adopted to weightedly splice the attention outputs of different levels, and the information flow transmission ratio is controlled by the gated linear unit.

4. The method for intelligently describing the content of liver space-occupying lesions ultrasound images using LLM according to claim 1, characterized in that: The implementation of the uncertainty calibration module includes: Output a confidence estimate vector at each time step of the decoder. The vector contains word-level and sentence-level uncertainty indicators. The word-level indicator is calculated by the entropy value of the token prediction probability, and the sentence-level indicator is estimated by the variance of the hidden state output by the bidirectional GRU. An uncertainty-guided beam search strategy is constructed. During the search process with a beam width of k, the retention weight of candidate sequences is dynamically adjusted according to the confidence index. When the confidence of three consecutive tokens is detected to be lower than the threshold, a local backtracking mechanism is triggered to regenerate the previous text. The final generated description text is subjected to entity verification based on the medical knowledge graph, and the detected contradictory entities are semantically corrected through the graph reasoning module.

5. The method for intelligently describing the content of liver space-occupying lesions ultrasound images using LLM according to claim 1, characterized in that: The post-processing reordering mechanism includes: Build a multi-candidate generation pipeline to generate n sets of differential description texts in parallel by adjusting the decoding temperature parameters; Perform semantic analysis on each text set, extract key medical entities and construct an entity relationship graph. Use a graph matching algorithm to calculate the topological similarity with the standard diagnostic report. An integrated ranking model is used to comprehensively score the candidate texts. The model integrates three indicators: vocabulary overlap, entity accuracy, and syntactic complexity, and generates the final optimized description text through a gradient boosting decision tree.

6. The method for intelligently describing the content of liver space-occupying lesions ultrasound images using LLM according to claim 2, characterized in that: The configuration of the deformable convolution module includes: In the third and fourth stages of ResNet-50, a deformable convolution layer is inserted, and each deformable convolution kernel is set with 9 offset sampling points; The offset learning is guided by the segmentation mask of the lesion area. A dual-path supervision strategy is adopted in the training stage. The main path calculates the image classification loss, and the auxiliary path calculates the L2 regularization loss between the offset field and the true segmentation boundary. A dynamic kernel selection mechanism is enabled during the inference phase to automatically select a deformable convolution kernel size of 3×3 or 5×5 based on the texture complexity of the input image.

7. The method for intelligently describing the content of liver space-occupying lesions ultrasound images using LLM according to claim 3, characterized in that: The method for constructing the anatomical topology map includes: Extract the three-dimensional coordinate template of liver blood vessels from the standard medical atlas and map the template to the current ultrasound image coordinate system using a non-rigid registration algorithm; An anatomical structure recognition network based on supervoxel segmentation was constructed, 3D sparse convolution was used to extract vascular branch features, and an iterative clustering algorithm was used to generate vascular centerline topology. Three types of nodes are defined in the topological graph: lesion core area, main vascular intersection, and liver lobe boundary marker. The edge weight is determined by the Euclidean distance and blood flow direction.

8. The method for intelligently describing the content of liver space-occupying lesions ultrasound images using LLM according to claim 4, characterized in that: The construction of the medical knowledge graph includes: Entity-relationship triplets were extracted from clinical guideline literature to establish a liver disease ontology library, which contains 7 major entity types and 23 relationship types. A dual attention graph neural network is used for knowledge representation learning, the entity embedding layer adopts the TransE algorithm, and the relationship embedding layer adopts the rotation encoding strategy; A dynamic entity linking mechanism is set up. When an entity mention is detected in the generated text, the associated pathological features, treatment plans, and prognostic indicator information in the atlas are synchronously queried and injected into the attention context window of the reranking model.

9. The method for intelligently describing the content of liver space-occupying lesions ultrasound images using LLM according to claim 5, characterized in that: The training method of the integrated sorting model includes: Construct a multi-dimensional feature engineering project to extract word-level features, sentence-level features, and paragraph-level features of the candidate text. Word-level features include medical term coverage and negation density, while sentence-level features include dependency tree depth and semantic role labeling completeness. A stratified sampling strategy was used to construct the training set, with stratification based on lesion type, ultrasound equipment model, and description length; A hybrid loss function is designed to combine ranking loss and classification loss. The ranking loss uses the Listwise loss function, and the classification loss uses the focal loss function to deal with the problem of category imbalance.

10. The method for intelligently describing the content of liver space-occupying lesions ultrasound images using LLM according to claim 1, characterized in that: Also includes real-time interactive optimization mechanism: During the ultrasound examination, descriptive text is displayed synchronously, and correction annotations are collected through the doctor feedback interface; Build an online incremental learning framework that encodes the corrected labeled data into comparative learning samples and updates model parameters through an elastic weight consolidation algorithm, preserving important parameters while preventing catastrophic forgetting. Set up a model version rollback mechanism. When a user corrects the same type of error five times in a row, it automatically triggers the loading and fine-tuning of the historical optimal model.

Citation Information

Patent Citations

  • Children MPP auxiliary diagnosis system based on multi-modal time series data modeling

    CN120015296A

  • Multimodal data collection and diagnosis of rare ocular diseases

    WO2025147362A1