Automatic thyroid ultrasound image diagnosis report generation method

Thyroid ultrasound image feature extraction and semantic mapping are performed through the improved ResNet-50 network and Transformer architecture, combined with the ChatGLM3-6B model and P-TuningV2 fine-tuning technology, detailed ultrasound diagnostic reports are automatically generated, solving the problem of time-consuming and labor-intensive generation of diagnostic reports in the existing technology and poor diagnosis of rare diseases, achieving efficient and accurate generation of diagnostic reports, and enhancing the interpretability and credibility of the reports.

CN119920393APending Publication Date: 2025-05-02HARBIN INST OF TECH +1
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202411836807.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-02

Smart Images

  • Figure CN119920393A_ABST
    Figure CN119920393A_ABST
Patent Text Reader

Abstract

The invention discloses a thyroid ultrasound image diagnosis report automatic generation method, and relates to a diagnosis report automatic generation method. The invention aims to solve the problems that manual analysis is time-consuming and labor-consuming and depends on professional skills of doctors in a current thyroid ultrasound image diagnosis report. The method comprises the steps that 1, a ResNet-50 network is improved, feature extraction is conducted on a segmented thyroid ultrasound image, and the network comprises a specific recognition module for thyroid size and shape features and thyroid nodule features; step 2, applying a Transform architecture to perform semantic mapping and description of the image, capturing global semantic information of the image through a self-attention mechanism, and generating a compact feature vector; and step 3, based on the ChatGLM3-6B pre-training model, in combination with a P-Tuning V2 fine tuning method, training is carried out by using the marked thyroid ultrasound diagnosis report data, and an ultrasound report generation model is constructed. The invention belongs to the technical field of ultrasonic medical image processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for automatically generating a diagnosis report, and belongs to the technical field of ultrasonic medical image processing. Background Art

[0002] As the incidence of thyroid diseases increases, the demand for medical image analysis, especially ultrasound image analysis, also increases. Manual analysis of ultrasound images and writing diagnostic reports is time-consuming and laborious, and requires high professional skills of doctors. Especially in densely populated areas, radiologists face a huge workload. Therefore, it is of great significance to develop a system that can automatically generate ultrasound diagnostic reports.

[0003] At present, the field of medical report generation is experiencing a wave of technological innovation. The integration of deep learning and natural language processing technologies makes it possible to automatically generate medical reports, such as the Medical-VLBERT model and the VisualAI-Chat model. However, the development of this field is not without obstacles. First, the construction of high-quality medical datasets faces high annotation costs, and the annotation process requires professional knowledge, which limits the size of the available datasets, thereby affecting the training effect and generalization ability of the model. Secondly, for the diagnosis of rare diseases, existing models often perform poorly, because there are fewer case data for rare diseases, and it is difficult for the model to learn enough features from them. In addition, the internal mechanism of deep learning models is complex, and their decision-making process often lacks transparency, making it difficult to provide clear explanations to doctors and patients, which may affect clinical decisions and patients' trust in the results. At the same time, ensuring the consistency of the language and format of the generated medical reports is also a challenge, which is crucial to maintaining the professionalism and standardization of medical documents. Finally, the need for joint diagnosis by multiple experts has not been fully reflected in the model design, while in actual clinical scenarios, the collaborative diagnosis of multiple experts can often improve the accuracy of diagnosis and the quality of reports.

[0004] Therefore, it is necessary to design a method for automatically generating a thyroid ultrasound image diagnosis report to solve the above problems. The method proposed in the present invention aims to provide a system that can automatically generate an ultrasound diagnosis report containing detailed indicators such as thyroid size, shape, echo characteristics, blood flow conditions, and possible lesion types through an improved feature extraction network, deep semantic mapping, and ChatGLM3-6B model fine-tuned based on P-TuningV2, as well as a specific text feature extraction model and FunctionCall mechanism. This method not only improves the accuracy and efficiency of diagnosis, but also enhances the interpretability and credibility of the report by integrating the knowledge and experience of multiple expert diagnoses, thereby providing doctors and patients with more reliable and professional medical services. Summary of the invention

[0005] The present invention aims to solve the problems that manual analysis in current thyroid ultrasound image diagnostic reports is time-consuming and laborious, and relies on the professional skills of doctors, and further proposes a method for automatically generating thyroid ultrasound image diagnostic reports.

[0006] The technical solution adopted by the present invention to solve the above-mentioned problem is: the steps of the present invention include:

[0007] Step 1: Improve the ResNet-50 network to extract features from the segmented thyroid ultrasound images. The network includes specific recognition modules for thyroid size, shape features, and thyroid nodule features.

[0008] Step 2: Apply the Transformer architecture to perform semantic mapping and description of the image, capture the global semantic information of the image through the self-attention mechanism, and generate a compact feature vector;

[0009] Step 3. Based on the ChatGLM3-6B pre-trained model and combined with the P-TuningV2 fine-tuning method, the labeled thyroid ultrasound diagnostic report data is used for training to build an ultrasound report generation model. At the same time, the text feature extraction model is trained to extract key descriptions of the thyroid gland in the report to support the calculation of the TI-RADS score. The model inference results are then combined with the calculation results of the TI-RADS score to automatically generate a complete diagnostic report including thyroid size, shape, echo characteristics, and blood flow indicators.

[0010] Furthermore, the image preprocessing in step 1 includes normalization, resizing, and data augmentation strategies specific to thyroid ultrasound images to adapt to the input requirements of the improved ResNet-50 network.

[0011] Furthermore, in step 2, the feature vector output by the ResNet-50 network is processed, the self-attention mechanism of the Transformer architecture is used for feature interaction, and the deep semantic representation of the image is learned through a multi-layer Transformer encoder, and finally a set of compact feature vectors is generated for generating a diagnosis report.

[0012] Furthermore, the process of obtaining the feature vector based on the Transformer architecture described in step 2 includes feature vector extraction, flattening, dimension adjustment, position encoding, and inputting the processed feature map into the Transformer model for encoding as follows:

[0013] Step 201, feature vector extraction: Use the output of the last convolutional layer of ResNet-50 as the feature vector. The output feature map size of this layer is 7×7, and each position has 2048 channels, so the size of the feature vector is 7×7×2048;

[0014] Step 202, feature vector flattening: flatten the 7×7×2048 feature map into a one-dimensional vector to obtain a feature vector of size 7×7×2048=100352;

[0015] Step 203, dimensionality adjustment: map the 100352-dimensional feature vector to 256 dimensions through a linear layer;

[0016] Step 204, position coding: adding position coding to the 7×7×256 feature map;

[0017] Step 205, input to Transformer: input the feature map with position encoding added to the Transformer model; each 7×7 block will be used as a separate sequence element in the Transformer;

[0018] Step 206, Transformer encoder: The feature map is processed by a multi-layer Transformer encoder; each layer of the encoder includes a self-attention module and a feed-forward neural network;

[0019] Step 207, output feature vector: After being processed by the Transformer encoder, the final compact feature vector can be obtained.

[0020] Furthermore, the process of constructing the thyroid ultrasound report generation model in step 3 is:

[0021] Step 301: Based on the ChatGLM3-6B pre-trained model, the annotated thyroid ultrasound diagnosis report data is used to fine-tune the ChatGLM3-6B model using the P-TuningV2 fine-tuning method;

[0022] Step 302: After fine-tuning is completed, the obtained ultrasound report generation model can generate a structured and accurate diagnosis report based on the input ultrasound image;

[0023] Step 303: designing and training a text feature extraction model to capture the description of the thyroid gland in the ultrasound report text;

[0024] Step 304: Use the FunctionCall function in ChatGLM3-6B to call the TI-RADS calculation function during the report generation process, and integrate the calculation results into the report, thereby generating an ultrasound report containing a detailed quantitative analysis.

[0025] Furthermore, the text feature extraction model is based on the word frequency-inverse frequency method, and its mathematical expression is as follows:

[0026]

[0027] Among them, n i represents the number of times word i appears in the text, N represents the number of all words contained in the article, W represents the total number of documents contained in the corpus, and w i Represents the number of documents containing word i.

[0028] The beneficial effects of the present invention are as follows: the present invention integrates advanced image processing technology and natural language processing technology to form an efficient and accurate automatic diagnosis report generation process; the method first improves the ResNet-50 network to extract features of thyroid ultrasound images, which can effectively obtain the size, shape and nodule features of the thyroid gland, and then uses the Transformer structure to perform semantic mapping and image description, thereby obtaining deep semantic information; then, using the ChatGLM3-6B pre-training model and combining the P-TuningV2 fine-tuning technology and the proprietary text feature extraction model, a structured and content-accurate medical diagnosis report can be generated;

[0029] The present invention realizes the process of automated image analysis and report generation, and provides doctors with a comprehensive diagnostic report including quantitative analysis; this not only reduces the doctor's reliance on personal experience during the diagnosis process, but also improves the efficiency and accuracy of diagnosis, and helps to improve the overall quality of medical services; in addition, the method of the present invention also helps to achieve the optimal allocation of medical resources, by reducing the workload of doctors in routine diagnosis, so that they can devote more time and energy to the diagnosis and treatment of complex cases, thereby promoting scientific and technological progress in the medical industry and improving the level of medical services;

[0030] The present invention can significantly improve the efficiency and accuracy of generating diagnostic reports, reduce the workload of doctors, and promote the intelligent and automated process in the field of medical image analysis; therefore, the present invention is of great significance for promoting technological progress and application development in related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0032] Specific implementation method 1: Figure 1 As shown, a method for automatically generating a thyroid ultrasound image diagnosis report comprises the following specific steps:

[0033] Step 1: Improve the ResNet-50 network and extract features from the segmented thyroid ultrasound images. The network includes specific recognition modules for thyroid size, shape features, and thyroid nodule features. Preprocess the images, including normalization, resizing, and data enhancement strategies specific to thyroid ultrasound images, to meet the input requirements of the improved ResNet-50 network.

[0034] Step 2: Apply the Transformer architecture to perform semantic mapping and description of the image, capture the global semantic information of the image through the self-attention mechanism, and generate a compact feature vector; process the feature vector output by the ResNet-50 network, use the self-attention mechanism of the Transformer architecture for feature interaction, and learn the deep semantic representation of the image through a multi-layer Transformer encoder, and finally generate a set of compact feature vectors for generating a diagnosis report; the process of obtaining the feature vector based on the Transformer architecture includes feature vector extraction, flattening, dimension adjustment, position encoding, and inputting the processed feature map into the Transformer model for encoding as follows:

[0035] Step 201, feature vector extraction: Use the output of the last convolutional layer of ResNet-50 as the feature vector. The output feature map size of this layer is 7×7, and each position has 2048 channels, so the size of the feature vector is 7×7×2048;

[0036] Step 202, feature vector flattening: flatten the 7×7×2048 feature map into a one-dimensional vector to obtain a feature vector of size 7×7×2048=100352;

[0037] Step 203, dimensionality adjustment: map the 100352-dimensional feature vector to 256 dimensions through a linear layer;

[0038] Step 204, position coding: adding position coding to the 7×7×256 feature map;

[0039] Step 205, input to Transformer: input the feature map with position encoding added to the Transformer model; each 7×7 block will be used as a separate sequence element in the Transformer;

[0040] Step 206, Transformer encoder: The feature map is processed by a multi-layer Transformer encoder; each layer of the encoder includes a self-attention module and a feed-forward neural network;

[0041] Step 207, output feature vector: After being processed by the Transformer encoder, the final compact feature vector can be obtained;

[0042] Step 3: Based on the ChatGLM3-6B pre-trained model and the P-TuningV2 fine-tuning method, the labeled thyroid ultrasound diagnosis report data is used for training to build an ultrasound report generation model. At the same time, a text feature extraction model is trained to extract key descriptions of the thyroid gland in the report to support the calculation of the TI-RADS score. The model inference results are then combined with the calculation results of the TI-RADS score to automatically generate a complete diagnostic report including thyroid size, shape, echo characteristics, and blood flow indicators. The process of building a thyroid ultrasound report generation model is as follows:

[0043] Step 301: Based on the ChatGLM3-6B pre-trained model, the annotated thyroid ultrasound diagnosis report data is used to fine-tune the ChatGLM3-6B model using the P-TuningV2 fine-tuning method;

[0044] Step 302: After fine-tuning is completed, the obtained ultrasound report generation model can generate a structured and accurate diagnosis report based on the input ultrasound image;

[0045] Step 303: designing and training a text feature extraction model to capture the description of the thyroid gland in the ultrasound report text;

[0046] Step 304: Use the FunctionCall function in ChatGLM3-6B to call the TI-RADS calculation function during the report generation process, and integrate the calculation results into the report, thereby generating an ultrasound report containing a detailed quantitative analysis.

[0047] Specific implementation method 2: Figure 1 As shown, the text feature extraction model is based on the word frequency-inverse frequency method, and its mathematical expression is as follows:

[0048]

[0049] Among them, n i represents the number of times word i appears in the text, N represents the number of all words contained in the article, W represents the total number of documents contained in the corpus, and w i Represents the number of documents containing word i.

[0050] The thyroid ultrasound report generation model described in this embodiment includes:

[0051] Fine-tuning model based on P-TuningV2 method: learns specific language patterns and expertise of thyroid ultrasound reports to generate structured and accurate diagnostic reports based on input ultrasound images;

[0052] Text feature extraction model: used to capture the description of the thyroid gland in the ultrasound report text, such as "unclear boundaries", "uneven internal echoes", "with coarse calcifications" and other key features, which are used to calculate the TI-RADS score according to medical standards;

[0053] FunctionCall calls TI-RADS calculation functions: used to integrate calculation results into reports, thereby generating ultrasound reports containing detailed quantitative analysis.

[0054] Among them, the improved ResNet-50 network integrates a module based on the attention mechanism to enhance the model's learning and feature extraction capabilities for the thyroid nodule area, and adopts a deeply separable convolutional structure to reduce the amount of calculation and improve the performance of the model in real-time diagnosis. In specific operations, in order to meet the input standards of the ResNet-50 network, the image is first subjected to a series of preprocessing operations, including normalization, resizing, and data enhancement strategies specific to thyroid ultrasound images to meet the input requirements of the improved ResNet-50 network. The preprocessed image is then sent to the already constructed and trained ResNet-50 network, which, with its deep convolutional structure, can automatically learn and extract key features such as size, shape, and thyroid nodules in the image.

[0055] Among them, the Transformer structure is used to perform semantic mapping and generate image descriptions to obtain feature vectors. Specifically, before applying the Transformer structure, the feature vectors output by the ResNet-50 network are first processed as necessary. Subsequently, with the help of the Transformer's self-attention mechanism, these feature vectors interact with each other to capture the semantic information of the image at a global level. Through the processing of multiple Transformer encoder layers, the feature vectors continuously learn the high-level semantic features of the image. Finally, through a series of linear transformations and normalization steps, a set of refined feature vectors are formed.

[0056] In step 3, a text feature extraction model is trained to extract descriptive text about the thyroid gland from the report, which will be used to calculate the TI-RADS score later. Combined with the FunctionCall function of the ChatGLM3-6B model, the TI-RADS calculation function can be called when generating a report, and the calculation results can be integrated into the report, thus generating a complete and accurate thyroid ultrasound diagnosis report.

[0057] Among them, when developing the thyroid ultrasound report generation model, the ChatGLM3-6B model, which was pre-trained on large-scale text data, was used as the starting point. The model already has strong language understanding and processing capabilities. Subsequently, the model was fine-tuned with P-TuningV2 using the labeled thyroid ultrasound diagnosis report, so that it can learn and master the language characteristics and professional knowledge points of the thyroid ultrasound report. After fine-tuning, the model is able to generate a structured and accurate diagnosis report. In addition, a text feature extraction model is designed, which can extract key thyroid descriptive information features from ultrasound reports, which are crucial for subsequent TI-RADS scoring. Through the FunctionCall function of the ChatGLM3-6B model, the TI-RADS scoring function can be automatically called during the report generation process, and the scoring results can be included in the report, so as to generate an ultrasound diagnosis report containing detailed quantitative analysis.

[0058] Among them, the ChatGLM3-6B pre-trained model: provides powerful text processing capabilities for fine-tuning the model to generate diagnostic reports that meet medical standards;

[0059] Thyroid ultrasound diagnosis report dataset: After screening and processing according to the data format requirements of ChatGLM3-6B, a large amount of diagnosis report text data was obtained for subsequent model fine-tuning;

[0060] Fine-tuning model based on P-TuningV2 method: learns specific language patterns and expertise of thyroid ultrasound reports to generate structured and accurate diagnostic reports based on input ultrasound images;

[0061] Text feature extraction model: used to capture the description of the thyroid gland in the ultrasound report text, such as "unclear boundaries", "uneven internal echoes", "with coarse calcifications" and other key features, which are used to calculate the TI-RADS score according to medical standards;

[0062] FunctionCall calls TI-RADS calculation functions: used to integrate calculation results into reports, thereby generating ultrasound reports containing detailed quantitative analysis.

[0063] The above is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technician familiar with this profession can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement made to the above embodiments without departing from the content of the technical solution of the present invention, based on the technical essence of the present invention, within the spirit and principles of the present invention, still fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for automatically generating a thyroid ultrasound image diagnosis report, characterized in that: The specific steps include: Step 1: Improve the ResNet-50 network to extract features from the segmented thyroid ultrasound images. The network includes specific recognition modules for thyroid size, shape features, and thyroid nodule features. Step 2: Apply the Transformer architecture to perform semantic mapping and description of the image, capture the global semantic information of the image through the self-attention mechanism, and generate a compact feature vector; Step 3. Based on the ChatGLM3-6B pre-trained model and combined with the P-Tuning V2 fine-tuning method, the labeled thyroid ultrasound diagnostic report data is used for training to build an ultrasound report generation model. At the same time, the text feature extraction model is trained to extract key descriptions of the thyroid gland in the report to support the calculation of the TI-RADS score. The model inference results are then combined with the calculation results of the TI-RADS score to automatically generate a complete diagnostic report including thyroid size, shape, echo characteristics, and blood flow indicators.

2. The method for automatically generating a thyroid ultrasound image diagnosis report according to claim 1, characterized in that: The image preprocessing in step 1 includes normalization, resizing, and data augmentation strategies specific to thyroid ultrasound images to adapt to the input requirements of the improved ResNet-50 network.

3. The method for automatically generating a thyroid ultrasound image diagnosis report according to claim 1, characterized in that: In step 2, the feature vector output by the ResNet-50 network is processed, the self-attention mechanism of the Transformer architecture is used for feature interaction, and the deep semantic representation of the image is learned through a multi-layer Transformer encoder, and finally a set of compact feature vectors is generated for generating a diagnostic report.

4. The method for automatically generating a thyroid ultrasound image diagnosis report according to claim 1, characterized in that: The process of obtaining the feature vector based on the Transformer architecture described in step 2 includes feature vector extraction, flattening, dimension adjustment, position encoding, and inputting the processed feature map into the Transformer model for encoding as follows: Step 201, feature vector extraction: Use the output of the last convolutional layer of ResNet-50 as the feature vector. The output feature map size of this layer is 7×7, and each position has 2048 channels, so the size of the feature vector is 7×7×2048; Step 202, feature vector flattening: flatten the 7×7×2048 feature map into a one-dimensional vector to obtain a feature vector of size 7×7×2048=100352; Step 203, dimensionality adjustment: map the 100352-dimensional feature vector to 256 dimensions through a linear layer; Step 204, position coding: adding position coding to the 7×7×256 feature map; Step 205, input to Transformer: input the feature map with position encoding added to the Transformer model; each 7×7 block will be used as a separate sequence element in the Transformer; Step 206, Transformer encoder: The feature map is processed by a multi-layer Transformer encoder; each layer of the encoder includes a self-attention module and a feed-forward neural network; Step 207, output feature vector: After being processed by the Transformer encoder, the final compact feature vector can be obtained.

5. The method for automatically generating a thyroid ultrasound image diagnosis report according to claim 1, characterized in that: The process of building a thyroid ultrasound report generation model in step 3 is as follows: Step 301: Based on the ChatGLM3-6B pre-trained model, the annotated thyroid ultrasound diagnosis report data is used to fine-tune the ChatGLM3-6B model using the P-Tuning V2 fine-tuning method; Step 302: After fine-tuning is completed, the obtained ultrasound report generation model can generate a structured and accurate diagnosis report based on the input ultrasound image; Step 303: designing and training a text feature extraction model to capture the description of the thyroid gland in the ultrasound report text; Step 304: Use the Function Call function in ChatGLM3-6B to call the TI-RADS calculation function during the report generation process, and integrate the calculation results into the report, thereby generating an ultrasound report containing a detailed quantitative analysis.

6. The method for automatically generating a thyroid ultrasound image diagnosis report according to claim 1, characterized in that: The text feature extraction model is based on the word frequency-inverse frequency method, and its mathematical expression is as follows: Among them, n i represents the number of times word i appears in the text, N represents the number of all words contained in the article, W represents the total number of documents contained in the corpus, and w i Represents the number of documents containing word i.

Citation Information

Patent Citations

  • Thyroid ultrasound characteristic tumor grading system based on Rocchio algorithm

    CN113887228A

  • Thyroid nodule grading identification system and method

    CN114067161A

  • Thyroid autonomous scanning system based on multi-modal generative dialogue

    CN117017355A

  • Medical image segmentation method based on CNN and Transform fusion network

    CN117173412A

  • Image segmentation method based on self-attention and computer equipment

    CN117522896A