Method for automatically generating Chinese breast ultrasound video diagnosis report based on multi-modal large model
Through the automatic generation method of breast ultrasound video diagnostic report based on Chinese multimodal large model, the problems of high data labeling cost, data bias and reporting language style differences in the prior art are solved, and detailed breast ultrasound diagnostic report is quickly and accurately generated, improving the utilization efficiency of medical resources and the quality of reports.
Patent Information
- Application Number
- CN202510126635.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-06-03
AI Technical Summary
The existing technology has problems such as high data labeling cost, data bias, unclear internal computer system of the model, and large differences in reporting language style in medical imaging processing, resulting in insufficient generalization performance of large models in the generation of Chinese medical reports and low quality of generating reports.
The automatic generation method of breast ultrasound video diagnostic report based on Chinese multimodal large model is adopted. Through the improved video data set processing scheme, image classification model and InternVL model, a breast ultrasound report generation model is constructed to generate a diagnostic report containing detailed indicators such as breast size, shape, echo characteristics, blood flow conditions and possible lesion types.
It realizes the rapid and accurate generation of breast ultrasound diagnostic reports containing detailed indicators, reduces diagnostic errors caused by human factors, improves the confidence and standardization of reports, reduces the work burden of doctors, and improves the efficiency of medical resources utilization.
Smart Images

Figure CN120089270A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for automatically generating a Chinese breast ultrasound video diagnostic report, belonging to the technical field of ultrasonic medical image processing. Background Art
[0002] Medical images are indispensable in modern medicine, and their important contributions in diagnosis, treatment planning, and disease monitoring have been widely recognized. Medical images provide a non-invasive visual display of the patient's physiological condition and provide important information for clinical diagnosis to formulate precise diagnostic and treatment strategies. However, with the continuous growth of medical image data, professional doctors are faced with the challenge of processing and interpreting a large amount of medical data, which urgently needs to be solved in the ever-changing medical field. In China, an ultrasound doctor has to read hundreds of medical images and provide diagnostic results every day, and radiologists need to process a large number and variety of medical images, which may lead to fatigue and cognitive burden, thus increasing the risk of missed diagnosis and misdiagnosis. For junior ultrasound doctors, understanding and interpreting complex medical images may be a daunting task. Such potential problems will further expand the risk of medical accidents and pose a potential threat to patient safety and medical quality.
[0003] At present, with the decline of computing power costs and the continuous action of Moore's law, natural language processing technology and machine vision technology have made great progress, and a series of large models (large language models, LLMs) with high performance in general fields have emerged. Moreover, with the help of a small amount of data and fine-tuning technology, large models can achieve excellent performance in specific fields. However, there are various obstacles in the generation of Chinese reports by large models in the medical field, and there are the following difficulties: First, the fine-tuning of large models in the medical field requires high-quality medical data sets, and the construction and annotation costs of such medical data sets are too high. The annotators need to have basic knowledge in the medical field, which restricts the development of data annotation in the medical field and also leads to insufficient generalization performance of large models in the generation of Chinese medical reports. Second, there is a data bias problem in medical data sets, and there are significant differences in the distribution frequencies of different types of data in the data sets. Directly using the data sets with data bias for the fine-tuning of medical report generation tasks will lead to a decline in the performance of large models and cause repeated reading phenomena. Third, the internal calculation mechanism of large language models is not clear enough, and there are potential risks in terms of technical ethics. Therefore, it is difficult to be trusted by doctors and patients in practical applications. Fourth, there are differences in language styles in medical reports, and the report styles written by different doctors are very different, resulting in large language models being difficult to learn their characteristics, and thus the generated reports are difficult to achieve unity in report formats.
[0004] Therefore, it is necessary to design an automatic generation algorithm for Chinese breast ultrasound video diagnostic reports based on Chinese multi-modal large models to solve the above problems. The method proposed in the present invention aims to provide a system that can automatically generate ultrasound diagnostic reports containing detailed indicators such as breast size, shape, echo characteristics, blood flow conditions, and possible lesion types through an improved video dataset processing scheme, an image classification model, and the InternVL (Scholar's Vision) large model fine-tuned based on LoRa. This method directly generates breast ultrasound diagnostic reports based on ultrasound videos, improving the examination efficiency of ultrasound physicians. At the same time, combined with the image classification model, it improves the confidence of the generated reports and ensures medical quality, thereby providing fast and accurate ultrasound diagnosis and treatment services for patients. Summary of the Invention
[0005] In order to solve the problem that the traditional report generation has certain drawbacks and is difficult to meet the actual needs of the medical industry, the present invention further proposes an automatic generation method for Chinese breast ultrasound video diagnostic reports based on multi-modal large models.
[0006] The technical solutions adopted by the present invention to solve the above problems are as follows: The steps of the present invention include:
[0007] Step 1: Preprocess the medical image dataset provided by the hospital. Sample and crop breast ultrasound images according to frames, process the report data through optical character recognition technology, and use ChatGLM4 for text cleaning to generate a concise and clear structured medical dataset;
[0008] Step 2: Apply the BioViL encoder to implement the image multi-classification task. Capture the image features of medical images through the medical image encoder to form feature vectors. Use BCE loss to measure the loss during training, and use the Jaccard function to evaluate the difference between the multi-classification prediction value and the true value. Finally, train a breast image multi-classification model to generate multiple classification categories corresponding to breast images;
[0009] Step 3: Apply the InternVL model, combined with the LoRa fine-tuning method, use the {frame, caption} frame-ultrasound report text pair dataset for training to construct a breast ultrasound report generation model. At the same time, construct enhanced prompt words, apply the image classifier obtained in Step 2 to perform multi-classification on ultrasound images to assist the large language model in better understanding various key indicators in breast report generation. Combine the enhanced prompt words with the frame-ultrasound report text pair, and use the supervised fine-tuning scheme to fine-tune the InternVL model, thereby automatically generating a complete breast ultrasound diagnosis and treatment report.
[0010] Further, the preprocessing of the image in step 1 includes sampling and cropping methods specific to breast ultrasound images to meet the requirements of model report generation. Optical character recognition technology is used to identify and clean the ultrasound impression report using ChatGLM4.
[0011] Further, the process of processing the input image to obtain the classification vector in step 2 includes the following steps:
[0012] Step 201, Image feature vector extraction: For the input image I, encode the image vector through the image encoder of the BioViL model to form an image vector of size batch size×128×4×4;
[0013] Step 202, Pooling: After passing through the average pooling layer, obtain a pooling vector with dimensions batch size×128×1×1;
[0014] Step 203, Feature vector flattening: Flatten the pooling vector to obtain a feature vector with dimensions batch size×2048;
[0015] Step 204, Dimension adjustment: Through the first fully connected layer, output a feature vector of batch size×512; Subsequently, through the second fully connected layer, obtain a final output classification result vector of batch size×num classed, where numclassed is the number of classes in multi-classification. Finally, adjust according to the classification result vector using the sigmoid operator to obtain a classification probability with a value range of [0,1]. If the classification probability is greater than 0.5, it is classified into this category;
[0016] In model training, BCE loss is adopted, and the loss function is:
[0017]
[0018] where, x i represents the output of the model, y i represents the target label, σ(x) represents the sigmoid function, and N represents the total number of samples.
[0019] Further, the process of constructing the breast ultrasound report generation model is as follows:
[0020] Step 301, Based on the image classifier in step 2, use the labeled breast ultrasound diagnosis report data to generate a fine-tuning dataset; The style of the enhanced prompt words in the fine-tuning dataset is: "The estimated classification result of the above image is: <finding>”, where <finding>It represents the classification result of the above image classifier on the preset classification categories, which is used to guide the large model to focus on the above classification metrics, so as to generate an accurate and error-free breast ultrasound report;
[0021] Step 302: Use the LoRa fine-tuning method based on the PEFT library to perform supervised fine-tuning on the InternVL4B model in the way of multi-image dialogue. After the fine-tuning is completed, the obtained breast ultrasound report generation model can generate an ultrasound impression report with a language style close to that of a sonographer and accurate key indicators according to the input breast ultrasound images;
[0022] Step 303: Processing scheme for multi-ultrasound video input: For multiple ultrasound videos {v 1 , v 2 ,..., v n}, perform average segmentation on each video respectively to obtain the key frames {f i,1 , f i,2 ,..., f i,n} corresponding to each video, and reconstruct the input as {f 1,i , f 2,i ,..., f n,i};
[0023] Among them, f i,n represents the nth sampled frame of the video with the serial number i in the input;
[0024] Comprehensively use the information in each video, use the reconstructed input in combination with the above breast ultrasound report generation model, output the classification results {p 1 , p 2 ,..., p n}, calculate the average classification result , and infer the breast information in the video based on the average classification result to construct a reinforcement prompt, so as to generate a more accurate breast ultrasound impression report based on the reconstructed input.
[0025] The beneficial effects of the present invention are:
[0026] 1. The present invention uses efficient image processing technology and natural language processing technology. First, it performs average sampling and cropping on breast ultrasound images, unifies the resolution, uses optical character recognition technology to process picture reports, and uses a large model to clean the text, effectively constructing a structured data set. Then, it obtains the image classification result through the BioViL encoder, and further constructs a proprietary reinforcement prompt to guide the fine-tuning of the InternVL model, so as to generate a breast ultrasound impression report with accurate classification and a language style unified with that of doctors;
[0027] 2. The present invention realizes the automatic generation of breast ultrasound diagnostic reports, greatly shortening the report generation time. In the traditional method, doctors need to manually write reports, while this invention can quickly process a large amount of ultrasound image data and generate reports, reducing the waiting time for patients to obtain diagnostic results. By constructing a multi-classification dataset to train an image classification model, various features of the breast can be more accurately identified, such as size, shape, echo characteristics, blood flow conditions, and lesion types, providing a more reliable basis for report generation and reducing diagnostic errors caused by human factors. The generated reports have a language style similar to that of ultrasound doctors, with accurate key indicators, making the reports highly standardized and consistent, facilitating information exchange and comparison between different medical institutions;
[0028] 3. The present invention can assist doctors in quickly completing report writing, enabling doctors to devote more time and energy to the diagnosis and treatment of difficult cases, improving the utilization efficiency of medical resources. At the same time, it provides new ideas and methods for the intelligent processing of medical images, promotes the deep integration of the medical field and artificial intelligence technology, and helps to promote the process of medical informatization. Therefore, the present invention is of great significance for promoting the technological progress and application development in related fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a flowchart of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0030] DETAILED DESCRIPTION OF THE INVENTION I: As Figure 1 shown, a method for automatically generating a Chinese breast ultrasound video diagnostic report based on a multi-modal large model, the specific steps include:
[0031] Step 1: Preprocess the medical image dataset provided by the hospital. Sample and crop breast ultrasound images according to frames, process report data through optical character recognition technology, and use ChatGLM4 for text cleaning to generate a concise and clear structured medical dataset; the preprocessing of the images includes sampling and cropping methods specific to breast ultrasound images to meet the requirements of model report generation, using optical character recognition technology to identify and using ChatGLM4 to clean the ultrasound impression report;
[0032] Step 2: Apply the BioViL encoder to implement the image multi-classification task. Capture the image features of medical images through a medical image encoder to form feature vectors. Use the BCE loss to measure the loss during training, and use the Jaccard function to evaluate the difference between the multi-classification prediction value and the true value. Finally, train a breast image multi-classification model for generating multiple classification categories corresponding to breast images; the process of processing the input image to obtain the classification vector includes the following steps:
[0033] Step 201, Image Feature Vector Extraction: For the input image I, encode the image vector through the image encoder of the BioViL model to form an image vector with a size of batch size×128×4×4;
[0034] Step 202, Pooling: After passing through the average pooling layer, obtain a pooling vector with a dimension of batch size×128×1×1;
[0035] Step 203, Feature Vector Flattening: Flatten the pooling vector to obtain a feature vector with a dimension of batch size×2048;
[0036] Step 204, Dimension Adjustment: Through the fully connected layer 1, output a feature vector of batch size×512; Subsequently, through the fully connected layer 2, obtain a classification result vector of the final output batch size×num classed, where numclassed is the number of categories in multi-classification. Finally, perform sigmoid operator adjustment according to the classification result vector to obtain a classification probability with a value range of [0,1], and it is considered that if the classification probability is greater than 0.5, it belongs to this category;
[0037] During model training, BCEloss is adopted, and the loss function is:
[0038]
[0039] where, x i represents the output of the model, y i represents the target label, σ(x) represents the sigmoid function, and N represents the total number of samples;
[0040] Step 3, Apply the InternVL model, combined with the LoRa fine-tuning method, use the {frame, caption} frame-ultrasound report text pair dataset for training to construct a breast ultrasound report generation model. At the same time, construct enhanced prompt words, use the image classifier obtained in Step 2 to perform multi-classification on ultrasound images to assist the large language model to better understand various key indicators in breast report generation, and combine the enhanced prompt words with the frame-ultrasound report text pair to fine-tune the InternVL model using the supervised fine-tuning scheme, so as to automatically generate a complete breast ultrasound diagnosis and treatment report;
[0041] The process of constructing the breast ultrasound report generation model is as follows:
[0042] Step 301, Based on the image classifier in Step 2, use the labeled breast ultrasound diagnosis report data to generate a fine-tuning dataset; The style of the enhanced prompt words in the fine-tuning dataset is: "The estimated classification result of the above image is: <finding>”, where <finding>It represents the classification result of the above image classifier on the preset classification categories, which is used to guide the large model to focus on the above classification metrics, so as to generate an accurate and error-free breast ultrasound report;
[0043] Step 302: Use the LoRa fine-tuning method based on the PEFT library to perform supervised fine-tuning on the InternVL4B model in the way of multi-image conversation. After the fine-tuning is completed, the generated breast ultrasound report model can generate an ultrasound impression report with a language style close to that of a sonographer and accurate key indicators according to the input breast ultrasound image;
[0044] Step 303: Processing scheme for multi-ultrasound video input: For multiple segments of ultrasound videos {v 1 , v 2 ,..., v n}, perform average segmentation on each video respectively to obtain the key frames {f i,1 , f i,2 ,..., f i,n} corresponding to each video, and reconstruct the input as {f 1,i , f 2,i ,..., f n,i};
[0045] Among them, f i,n represents the nth sampled frame of the video with the serial number i in the input;
[0046] Comprehensively use the information in each video, use the reconstructed input in combination with the above breast ultrasound report generation model, output the classification results {p 1 , p 2 ,..., p n}, calculate the average classification result p, infer the breast information in the video based on the average classification result p, construct a strengthened prompt, and thus generate a more accurate breast ultrasound impression report based on the reconstructed input.
[0047] Among them, the breast ultrasound report generation model described in this embodiment includes:
[0048] A basic model for image multi-classification, which uses a medical image encoder to capture image features, such as breast size, shape, echo characteristics, blood flow conditions, and possible lesion types, to provide classification category information of breast images for report generation and enhance the report generation performance of the model;
[0049] The InternVL-4B model based on LoRa fine-tuning, based on the image classification information provided by the image classifier, generates a fine-tuning data set with strengthened prompts, and performs supervised fine-tuning in the way of multi-image conversation, and finally generates an accurate and professional ultrasound impression report according to the input prominent ultrasound image.
[0050] Among them, when constructing the breast ultrasound report generation model, InternVL-4B is used as the base model, which has certain basic performance in the image-text dialogue task. To overcome the repetition phenomenon of the model, enhanced prompt words are introduced to provide external knowledge guidance, thereby guiding the model to reduce hallucinations and repetition phenomena and generate an ultrasound impression report that conforms to the actual breast images.
[0051] Among them, the InternVL model refers to the Shusheng Wanxiang model.
[0052] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the technical solution content of the present invention and is based on the technical essence of the present invention, any simple modification, equivalent replacement, and improvement made to the above embodiments still fall within the protection scope of the technical solution of the present invention.< / finding> < / finding> < / finding> < / finding>
Claims
1. A method for automatically generating Chinese breast ultrasound video diagnostic reports based on a multimodal large model, characterized in that: The specific steps include: Step 1: Preprocess the medical imaging dataset provided by the hospital, sample and crop the breast ultrasound images according to the frames, process the report data using optical character recognition technology, and use ChatGLM4 for text cleaning to generate a concise and clear structured medical dataset; Step 2: Apply the BioViL encoder to implement the image multi-classification task. The medical image encoder is used to capture the image features of the medical image to form a feature vector. The BCE loss is used to measure the loss in training. The Jaccard function is used to evaluate the difference between the multi-classification prediction value and the true value. Finally, a breast image multi-classification model is trained to generate multiple classification categories corresponding to the breast image. Step 3. Apply the InternVL model and the LoRa fine-tuning method, use the {frame, caption} frame-ultrasound report text to train the data set, build a breast ultrasound report generation model, and at the same time, build enhanced prompt words. Apply the image classifier obtained in step 2 to perform multi-classification of ultrasound images to assist the large language model to better understand the various key indicators in breast report generation. Combine the enhanced prompt words with the frame-ultrasound report text pairs, and use the supervised fine-tuning scheme to fine-tune the InternVL model, so as to automatically generate a complete breast ultrasound diagnosis and treatment report.
2. The method for automatically generating a Chinese breast ultrasound video diagnosis report based on a multimodal large model according to claim 1 is characterized in that: The image preprocessing in step 1 includes sampling and cropping methods specific to breast ultrasound images to suit the requirements of model report generation, recognition using optical character recognition technology, and cleaning of ultrasound impression reports using ChatGLM4.
3. The method for automatically generating Chinese breast ultrasound video diagnostic report based on multimodal large model according to claim 1 is characterized in that: The process of processing the input image to obtain the classification vector in step 2 includes the following steps: Step 201, image feature vector extraction: for the input image I, the image vector is encoded by the image encoder of the BioViL model to form an image vector of batch size × 128 × 4 × 4; Step 202, pooling: After the average pooling layer, a pooling vector with a dimension of batch size × 128 × 1 × 1 is obtained; Step 203, flattening the feature vector: flatten the pooled vector to obtain a feature vector with a dimension of batch size × 2048; Step 204, dimension adjustment: output a feature vector of batch size × 512 through the fully connected layer 1; then, through the fully connected layer 2, obtain a classification result vector of the final output batch size × num classes, where num classes is the number of multi-classification categories; finally, perform sigmoid operator adjustment based on the classification result vector to obtain a classification probability in the range of [0, 1], and consider that a classification probability greater than 0.5 is classified into this category; BCEloss is used in model training, and the loss function is: Among them, x i Represents the output of the model, y i represents the target label, σ(x) represents the sigmoid function, and N represents the total number of samples.
4. The method for automatically generating Chinese breast ultrasound video diagnostic report based on multimodal large model according to claim 1 is characterized in that: The process of building a breast ultrasound report generation model is as follows: Step 301, based on the image classifier in step 2, using the labeled breast ultrasound diagnosis report data, a fine-tuning data set is generated; the style of the enhanced prompt words in the fine-tuning data set is: "The estimated classification result of the above image is: <finding>",in <finding> It indicates the classification result based on the above image classifier on the preset classification category, which is used to guide the large model to focus on the above classification indicators, so as to generate an accurate breast ultrasound report;< / finding> < / finding> Step 302, using the LoRa fine-tuning method based on the PEFT library, supervised fine-tuning the InternVL4B model in a multi-image dialogue manner. After the fine-tuning is completed, the obtained breast ultrasound report generation model can generate an ultrasound impression report with a language style close to that of ultrasound physicians and accurate key indicators based on the input breast ultrasound image; Step 303, a processing scheme for multiple ultrasound video inputs: for multiple ultrasound video segments {v1, v2, ..., v n }, split the video equally and get the key frame {f i,1 ,f i,2 ,...,f i,n }, refactor the input to {f 1,i ,f 2,i ,...,f n,i }; Among them, f i,n Represents the nth sampled frame of the video with sequence number i in the input; Comprehensively use the information in each video, use the reconstructed input combined with the above breast ultrasound report generation model, and output the classification results {p1,p2,...,p n }, find the average classification result Based on the average classification results Infer breast information from videos and construct enhanced prompt words to generate more accurate breast ultrasound impression reports based on reconstructed input.