Radiology report auxiliary writing method and system based on intelligent follow-up visit
By constructing a multimodal model and text generation model, combined with massive historical data, the problems of low efficiency and insufficient accuracy of traditional radio follow-up reports are solved, and intelligent identification and analysis are realized to generate efficient and accurate follow-up reports.
Patent Information
- Application Number
- CN202411880427.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-06
AI Technical Summary
The writing of traditional radio follow-up examination reports relies on manual operations and is inefficient. Due to the intervention of subjective judgments, the consistency and accuracy of the report may be affected, making it difficult to meet the deep learning and adaptive needs of massive historical data.
A multimodal model based on intelligent follow-up is adopted. By constructing a multimodal training data set and preprocessing it, a multimodal model is built to generate differential text analysis, and a follow-up report text is generated by combining the text generation model and historical reports.
It realizes intelligent identification and analysis of radio follow-up data, generates accurate follow-up reports, reduces the burden on doctors, and improves report writing efficiency and medical service quality.
Smart Images

Figure CN119943249A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a radiology report auxiliary writing technology, and in particular to a radiology report auxiliary writing method and system based on intelligent follow-up. Background Art
[0002] Radiological follow-up is the use of radiological examination technology to conduct clinical diagnosis and treatment of the same patient at a certain time interval to detect the development of the disease, the effect of treatment, etc. Radiological follow-up examinations are mainly for patients with tumors, chronic diseases or postoperative patients, and are of great significance to patient management and disease control. Radiologists compare the imaging data of current and historical examinations, evaluate the progression of the disease, diagnose the disease, and complete the writing of the radiological report text of the current examination.
[0003] The traditional process of writing radiological follow-up examination reports is highly dependent on manual operations. Doctors need to carefully review and compare the imaging data of the patient's two examinations before and after, and record and describe the changes and development of the disease in the parts involved in the examination in the radiological report item by item compared with the last examination, especially the changes in the size, shape, number, etc. of the lesions. However, this manual writing of follow-up examination reports is extremely challenging for doctors' memory of the changes before and after the lesions, and it takes a long time to write and is inefficient. In addition, due to the intervention of subjective judgment, the consistency and accuracy of the report may be affected.
[0004] In addition, with the rapid growth of medical imaging data, doctors are increasingly faced with challenges in processing large amounts of imaging data, which further exacerbates the limitations of the traditional radiology follow-up examination report writing process. The existing report generation method based on natural language processing has made some progress, but has not yet fully utilized follow-up data to improve the accuracy of report generation.
[0005] For example, prior art 1: CN112712869B provides a method for dynamically acquiring and generating follow-up structured image reports from historical structured image reports. During the operation, the doctor still needs to click on the system multiple times, and because the imaging manifestations of each examination may vary due to changes in the condition or treatment, the target of continuous comparison may change, which makes it difficult to use a structured fixed template for before-and-after comparison, and it is difficult to meet the actual needs in clinical practice.
[0006] For example, the prior art: 2: CN115293128A proposes a method for integrating radiological image and text data features to realize automatic generation of radiological reports. Although this method improves the writing efficiency in the report writing task of a single radiological examination, it does not take into account the specific application scenario of radiological follow-up report writing.
[0007] For example, in the prior art: 3: CN110766730B, a post-processing method for image registration and follow-up evaluation is proposed to match the anatomical structures of the two examinations before and after. However, these systems support limited lesion types and contrast methods, and are expensive, making it difficult to meet a wide range of clinical needs.
[0008] Currently, there is an urgent need for a more efficient, accurate and reproducible solution to optimize the writing process of radiology follow-up examination reports. Summary of the invention
[0009] The present invention aims to address the problems in the prior art that the massive historical data cannot be fully utilized for deep learning, lacks adaptability and self-learning capabilities, and is still limited to fixed control models in practical applications and cannot be flexibly adjusted with changes in the environment. The present invention provides a method and system for assisting in writing radiological reports based on intelligent follow-up.
[0010] In order to solve the above technical problems, the present invention is solved by the following technical solutions:
[0011] The present invention has significant technical effects due to the adoption of the above technical solution:
[0012] A radiology report auxiliary writing method based on intelligent follow-up includes:
[0013] Construct a multimodal training dataset D and preprocess it to form the input of the multimodal model MF;
[0014] Multimodal model MF construction and training;
[0015] Obtain differential text analysis; generate differential text analysis through multimodal model MF;
[0016] The text generation model G combines the difference analysis text with the historical report to generate the follow-up report text.
[0017] As a preference:
[0018] Acquire a number of patient follow-up data in conjunction with the radiological imaging database; the patient follow-up data includes patient ID, examination type, examination time, radiological images, and radiological report text;
[0019] Integrate the obtained patient follow-up data, match the radiological images and radiological report texts of the same patient ID and the same examination type at two moments before and after, and construct a multimodal dataset D;
[0020] The dataset D is constructed as follows:
[0021]
[0022] Where i is the patient number, N is the number of patients, and I i,tis the radiographic image of the ith patient at time t; B i,t is the bounding box information of the lesion ROI area at time t, T i,t is the lesion description text of the i-th patient at time t, T compare text for differential analysis of follow-up data;
[0023] Preprocess the image data I and convert it into image vector X img ;
[0024] The lesion bounding box B i,t-1 and B i,t The boundary coordinates are formatted as normal text with the following structure:
[0025] <box>(Z min ,Y min ,X min ,Z max ,Y max ,X max ) < / box>;
[0026] Convert the bounding box B into a vector sequence X through the word segmenter box , obtain the text description T of the patient's lesion at the previous and next moments and convert it into a text vector X text ;
[0027] References <ref>The markers are connected to the bounding box vector X box and the text vector X text , construct the data pair of the lesion box position and its corresponding text description:
[0028] <ref> <box> X box < / box>X text < / ref>;
[0029] The image vector X img , the bounding box vector X box and the text vector X text Integrate to form a unified input sequence.
[0030] Preferably: the multimodal model MF comprises a text encoding module, a visual encoding module, a multimodal fusion module and a differentiated text generation module;
[0031] The text encoding module is composed of multiple layers of Transformers, which captures the semantic features of the diagnosis description text.
[0032] The visual encoding model uses ViT to extract multi-dimensional features of images;
[0033] The multimodal fusion module introduces a medical hierarchical encoder and a cross-modal alignment attention mechanism to achieve the same semantic space mapping of image and text features;
[0034] The differentiated text generation module is based on the joint modeling of one-dimensional text and three-dimensional images and uses a contrastive learning framework to further capture the temporal feature changes of the lesions;
[0035] As a preferred embodiment: the medical image hierarchical encoder is a multi-dimensional feature hierarchical position encoder for comprehensively characterizing medical images,
[0036]
[0037] in, represents the splicing operation, P d , P h , P w , P t Encoding functions representing depth, height, width, and time dimensions, respectively;
[0038] Time Code P t The periodic function of (t) is:
[0039]
[0040] Among them, d t Dimensions that encode time;
[0041] The spatial dimensions of medical images, namely width, height and depth are encoded as:
[0042] P spatial (d,h,w)=[P d (d),P h (h),P w (w)];
[0043] The final multi-level position encoding P final for:
[0044]
[0045] As a preference: image cross-modal alignment attention mechanism, through the cross-modal attention mechanism to fuse V i , T i and P final , and obtain the fused multimodal feature H i ,
[0046] H i =CrossModalAttention(V i ,T i ,P final )+λ·W proj
[0047] Among them, W proj is the projection matrix from medical image features to text feature space; λ is the balance coefficient.
[0048] As a preferred option: the contrastive learning framework of the differentiated text generation module for joint modeling of one-dimensional text and three-dimensional images is built based on Transformer.
[0049] The objective function of differential contrast learning constructed by lesion changes is:
[0050]
[0051] Among them, F text is the feature representation of the text, N is the number of samples; ΔF att The significant changes in the lesion area, M diff is the difference mask matrix;
[0052] M diff =I(|ΔF att |); where I is the indicator function
[0053] ΔF att =AΘΔF, where ΔF is the feature of lesion change; A is the attention weight;
[0054] ΔF=F t+1 -F t ; Among them, F t is the medical image feature at time t, F t+1 is the medical image feature at time t+1;
[0055] W q , W k is the learned parameter; Θ represents element-by-element multiplication;
[0056] Get the differential changes of lesions T compare ;
[0057]
[0058] Among them, T position is the change in the position of the lesion; ΔV is the change in the shape and size of the lesion, T signal is the signal intensity change of the lesion;
[0059] T position =Decoder(ΔF att ·M diff );
[0060] ΔV=Volume(M diff ),
[0061] T signal =Decoder(ΔF att ·M diff )
[0062] Where, ΔF att The significant changes in the lesion area, M diff is the difference mask matrix.
[0063] In order to solve the above technical problems, the present invention also provides a radiology report auxiliary writing system based on intelligent follow-up, which includes:
[0064] The preprocessing module constructs and preprocesses the multimodal training dataset D to form the input of the multimodal model MF;
[0065] Multimodal model MF construction and training module;
[0066] Obtain a differential text analysis module; generate differential text analysis through a multimodal model MF;
[0067] Follow-up report text generation module, text generation model G combines the difference analysis text and historical report to generate the follow-up report text
[0068] The present invention uses a multimodal model to intelligently identify key changes in follow-up data, generate accurate follow-up reports, reduce the burden on doctors, and improve report writing efficiency and medical service quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 is a flow chart of the present invention;
[0070] Figure 2 It is a schematic diagram of the structure of the multimodal model of the present invention;
[0071] Figure 3 It is a schematic diagram of embodiment 3 of the present invention. DETAILED DESCRIPTION
[0072] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0073] Example 1
[0074] The radiology report auxiliary writing method based on intelligent follow-up collects and organizes the patient follow-up data in the radiology imaging database to construct a data set, and then processes and converts it into the proposed input data of the model MF.
[0075] The radiological examination data of the same patient ID and the same examination type at two moments before and after are matched from the radiological image database. In this embodiment, a total of 50,000 patients' historical imaging data and report results are extracted. Each imaging result data contains the annotated lesion frame information and the comparative description information text of the patient's image lesion frame at different periods. The lesion frame annotation and the lesion description information text are both annotated by the doctor, and the follow-up difference text is used as a label. The total data set D used to train the multimodal model MF is constructed in the following form:
[0076]
[0077] Where i represents the patient number, the dataset D has a total of 50,000 samples, and N is 50,000. The dataset is split to obtain the description text T of the difference between the two examination results of each patient's image compare , each patient can construct a dataset pair represented by d i =[Current moment image I t , the image I at a certain moment in the past t-1 , the bounding box information B of the lesion ROI area at the current moment t and lesion text description T t , the image lesion bounding box B at a certain moment in the past t-1 and lesion text description T t-1 ], its corresponding label l i =[T compare ] composition, see the following table for specific examples:
[0078]
[0079]
[0080] The image data I is converted into an image vector X by performing data format conversion, intensity normalization, cropping, block division and standardization operations. img ;
[0081] Furthermore, the coordinates of the lesion box are normalized and marked accordingly, that is, Bbox is represented as: min ,Y min ,X min ,Z max ,Y max ,X max ) < / box>;
[0082] Used when processing image description text <ref>Mark it and connect it to the bounding box one-to-one, that is, <ref> <box> X box < / box>X text < / ref>;
[0083] The formatted bounding box text is treated as normal text and converted into a vector sequence X through a tokenizer. box ; Get the text description T of the patient's lesion in the images before and after and convert it into a text vector X text ;
[0084] Furthermore, the above vectors are integrated to form a unified input sequence.
[0085] Example 2
[0086] Based on Example 1, this example involves the construction and training of a multimodal model MF, specifically involving the following processes:
[0087] Build a multimodal model MF, including a text encoding module, a visual encoding module, a multimodal fusion module and a differentiated text generation module. The model MF realizes the multidimensional mapping of image data and text data, ensuring that images and texts can be integrated and unified in the same vector space. This model provides a basis for realizing the automatic analysis of images to generate differentiated texts;
[0088] The text encoding module adopts a decoder-only architecture, which is composed of L stacked Transformer layers, and then introduces a fully connected layer and a normalized residual connection layer to improve the stability of the model training process. text Input into the text encoder to get the text feature vector:
[0089] T i =Transformer(X text );
[0090] The visual encoding module uses ViT to extract multi-dimensional features of medical images. Input image vector X img , extract its feature vector V through the ViT visual encoder i :
[0091] V i =ViT(X img );
[0092] The multimodal encoding module introduces a medical hierarchical encoder and a cross-modal alignment attention mechanism to achieve the same semantic space mapping of image and text features;
[0093] The hierarchical position encoder decouples the spatial and temporal dimensions of the image, giving the model MF multimodal processing and reasoning capabilities. Specific implementation process: Expand the position encoding to the temporal encoding P t , Deep Coding P d , highly encoded P h , width encoding P w The four parts define the hierarchical position encoder as follows:
[0094]
[0095] in represents the splicing operation, P d ,P h ,P w ,P t , respectively represent the encoding functions of depth, height, width and time dimensions, where the time encoding P t The periodic function of (t) can be expressed as:
[0096]
[0097] where d t The dimension of temporal coding, the spatial dimension (width, height, depth) coding of medical images is uniformly expressed as:
[0098] P spatial (d,h,w)=[P d (d),P h (h),P w (w)]
[0099] The final multi-level position encoding can be expressed as:
[0100]
[0101] Fusion of V via cross-modal attention mechanism i , T i and P final , and obtain the fused multimodal feature H i :
[0102] H i =CrossModalAttention(V i ,T i ,P final )+λ·W proj
[0103] Where W proj is the projection matrix from medical image features to text feature space. λ is the balance coefficient.
[0104] The contrastive learning framework of the differentiated text generation module for the joint modeling of one-dimensional text and three-dimensional image is built based on Transformer to capture the difference in lesion features between the before and after images over time. Given the medical image feature F at time t, t and the medical image feature F at time t+1 t+1 , defining the characteristic of lesion change ΔF=F t+1 -F t , and then introduce the attention weight A to capture the significant changes in the lesion area: ΔF att =AΘΔF, where
[0105]
[0106] W k ,W q is a learnable parameter. Θ represents element-wise multiplication. Then define the difference mask matrix: M diff =I(|ΔF att |), I represents the indicator function.
[0107] By introducing the above lesion change representation, the differential contrast learning objective function is constructed as follows:
[0108]
[0109] Among them, F text is the feature representation of the text, and N is the number of samples. From this, we can get the difference in the changes of the lesions:
[0110] Determine whether the lesion has increased or subsided: T position =Decoder(ΔF att ·M diff )
[0111] Determine the size and morphological changes of the lesion: ΔV = Volume (M diff ), such as the lesion volume increases by 30%, and the shape changes from round to irregular circle
[0112] Determine the signal strength increase and change: T signal =Decoder(ΔF att ·M diff ), such as the nodule signal intensity is significantly enhanced.
[0113] Combining the above difference descriptions, we get the final text description of the difference between the current image and the historical image data:
[0114]
[0115] Through the above method, the model can effectively integrate the one-dimensional text sequence and the three-dimensional image features, and improve the modeling ability of the model MF for spatiotemporal information. After training, the model MF can accurately capture the characteristics of lesions and automatically analyze the dynamic changes, and obtain the difference analysis text, which includes the changes in the location of lesions (new additions, disappearance), the size and shape of lesions, and the signal intensity (enhanced, weakened or unchanged).
[0116] Example 3
[0117] Based on the above embodiment, this embodiment provides a specific implementation of generating a follow-up report:
[0118] Based on the difference analysis text obtained in Example 2 and the report text at time t-1 as input, a complete radiology report text at time t is automatically generated through a text generation model G;
[0119] The text generation model G is based on the Transformer decoder structure, accurately learns the writing standards and medical terminology of historical reports, and generates radiological follow-up reports that meet medical standards by integrating the image difference analysis text into the historical report text;
[0120] The input and output examples of the model in the embodiment are as follows:
[0121]
[0122]
[0123] Example 4
[0124] This embodiment provides a radiology report auxiliary writing system for intelligent follow-up, which is implemented by the radiology report auxiliary writing method based on intelligent follow-up, and specifically includes:
[0125] Deploy the multimodal model MF and the text generation model G into the hospital's PACS system to efficiently generate radiology follow-up reports;
[0126] Through deep connectivity with the hospital's PACS system, combined with the patient's unique identifier (patient ID), examination type and examination time, the patient's follow-up images and related data are automatically retrieved to provide efficient data support for subsequent processing; after the MF model is deployed, even in the absence of bounding box annotation and text description input, it can still use the multimodal pre-trained feature comparison mechanism to perform automated difference analysis on follow-up images and historical images, comprehensively capture the local changes and heterogeneous features between images, and output difference analysis text; the text generation model automatically generates radiology follow-up report text based on the difference analysis text output by the MF model and combined with the patient's historical radiology report.< / ref> < / ref>
Claims
1. A radiology report auxiliary writing method based on intelligent follow-up, the method comprising: Construct a multimodal training dataset D and preprocess it to form the input of the multimodal model MF; Multimodal model MF construction and training; Get differential text analysis; Generate differential text analysis through multimodal model MF; The text generation model G combines the difference analysis text with the historical report to generate the follow-up report text.
2. The method for assisting writing of radiological reports based on intelligent follow-up according to claim 1, characterized in that: Acquire a number of patient follow-up data in conjunction with the radiological image database; the patient follow-up data includes patient ID, examination type, examination time, radiological images, and radiological report text; Integrate the obtained patient follow-up data, match the radiological images and radiological report texts of the same patient ID and the same examination type at two moments before and after, and construct a multimodal dataset D; The dataset D is constructed as follows: Where i is the patient number, N is the number of patients, and I i,t is the radiographic image of the ith patient at time t; B i,t is the bounding box information of the lesion ROI area at time t, T i,t is the lesion description text of the i-th patient at time t, T compare text for differential analysis of follow-up data; Preprocess the image data I and convert it into image vector X img ; The lesion bounding box B i,t-1 and B i,t The boundary coordinates are formatted as normal text with the following structure: <box>(Z min ,Y min ,X min ,Z max ,Y max ,X max )< / box>; Convert the bounding box B into a vector sequence X through the word segmenter box , obtain the text description T of the patient's lesion at the previous and next moments and convert it into a text vector X text ; References <ref>The markers are connected to the bounding box vector X box and the text vector X text , construct the data pair of the lesion box position and its corresponding text description:< / ref> <ref><box>X box < / box>X text < / ref>; The image vector X img , the bounding box vector X box and the text vector X text Integrate to form a unified input sequence.
3. The method for assisting writing of radiological reports based on intelligent follow-up according to claim 1 is characterized in that: The multimodal model MF includes a text encoding module, a visual encoding module, a multimodal fusion module, and a differentiated text generation module; The text encoding module is composed of multiple layers of Transformers, which captures the semantic features of the diagnosis description text. The visual encoding model uses ViT to extract multi-dimensional features of images; The multimodal fusion module introduces a medical hierarchical encoder and a cross-modal alignment attention mechanism to achieve the same semantic space mapping of image and text features; The differentiated text generation module is based on the joint modeling of one-dimensional text and three-dimensional images and uses a contrastive learning framework to further capture the temporal feature changes of the lesions.
4. The method for assisting writing of radiological reports based on intelligent follow-up according to claim 1, characterized in that: Medical Image Hierarchical Encoder is a multi-dimensional feature hierarchical position encoder used to comprehensively characterize medical images. in, represents the splicing operation, P d , P h , P w , P t Encoding functions representing depth, height, width, and time dimensions, respectively; Time Code P t The periodic function of (t) is: Among them, d t Dimensions that encode time; The spatial dimensions of medical images, namely width, height and depth are encoded as: P spatial (d,h,w)=[P d (d),P h (h),P w (w)]; The final multi-level position encoding P final for:
5. The method for assisting writing of radiological reports based on intelligent follow-up according to claim 1 is characterized in that: Image cross-modal alignment attention mechanism, through the cross-modal attention mechanism to fuse V i 、T i and P final , and obtain the fused multimodal feature H i , H i =CrossModalAttention(V i ,T i ,P final )+λ·W proj Among them, W proj is the projection matrix from medical image features to text feature space; λ is the balance coefficient.
6. The method for assisting writing of radiological reports based on intelligent follow-up according to claim 1, characterized in that: The contrastive learning framework of the differentiated text generation module for joint modeling of one-dimensional text and three-dimensional images is built based on Transformer. The objective function of differential contrast learning constructed by lesion changes is: Among them, F text is the feature representation of the text, N is the number of samples; ΔF att The significant changes in the lesion area, M diff is the difference mask matrix; M diff =I(|ΔF att |); where I is the indicator function ΔF att =AΘΔF, where ΔF is the feature of lesion change; A is the attention weight; ΔF=F t+1 -F t ; Among them, F t is the medical image feature at time t, F t+1 is the medical image feature at time t+1; W q , W k is the learned parameter; Θ represents element-by-element multiplication; Get the differential changes of lesions T compare ; Among them, T position is the change in the position of the lesion; ΔV is the change in the shape and size of the lesion, T signal is the change in signal intensity of the lesion; T position =Decoder(ΔF att ·M diff ); ΔV=Volume(M diff ),; T signal =Decoder(ΔF att ·M diff ) Where, ΔF att The significant changes in the lesion area, M diff is the difference mask matrix.
7. The radiology report auxiliary writing system based on intelligent follow-up is characterized by: include: The preprocessing module constructs and preprocesses the multimodal training dataset D to form the input of the multimodal model MF; Multimodal model MF construction and training module; Obtain a differential text analysis module; generate differential text analysis through a multimodal model MF; Follow-up report text generation module,The text generation model G combines the difference analysis text and the,history report to generate the follow-up report text.
Citation Information
Patent Citations
Image registration and follow-up evaluation methods, storage media and computer equipment
CN110766730B
System and method for dynamically acquiring follow-up data from historical imaging structured reports
CN112712869B
Radiology report generation model training method and system based on multi-modal contrast learning
CN115293128A
Cited By
Clinical research report automatic writing system and method based on artificial intelligence
CN120473070A
Focus area determination method and device, electronic equipment and storage medium
CN121074499A