Medical Image Report Generation Method Based on Multimodal Large Model Preference Alignment Technology
Through multimodal large model preference alignment technology, combining image and report data, the expert diagnosis words and sentence structure can be analyzed to achieve accurate matching between images and text, and the problem of insufficient reporting accuracy and standardization in the existing technology is solved, the timeliness and accuracy of reports is improved, and the work burden of medical staff is reduced.
Patent Information
- Application Number
- CN202510522551.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The existing technology relies on large-scale unoptimized data training, resulting in insufficient accuracy and standardization of medical imaging reports generation, failure to fully consider the doctor's diagnostic style and professional preferences, and the generated reports are biased from the actual doctor's diagnostic methods and word habits, and the clinical status cannot be updated in time, which affects the timeliness and accuracy of the diagnosis.
Through the multimodal big model preference alignment technology, image and report data are collected, image texture characteristics and vector relationships between image textures and lesion descriptions, expert diagnostic words and sentence structures are analyzed, text generation parameters are adjusted, image and text are accurately matched, and report content is updated in real time to reflect the latest medical conditions.
It improves the accuracy and practicality of medical reports, ensures that reports meet clinical needs, reduces the burden on medical staff, promotes the standardization and objectivity of diagnostic processes, and improves the timeliness and accuracy of reports.
Smart Images

Figure CN120032790B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical imaging technology, and in particular to a method for generating medical imaging reports based on multi-modal large model preference alignment technology. Background Art
[0002] With the development of artificial intelligence technology, large language models have demonstrated powerful text processing and generation capabilities. However, in the real world, there is a large amount of image-modal data in addition to text-modal data. Therefore, multi-modal large models have also developed rapidly. Multi-modal large models can process both image and text inputs simultaneously, understand images according to text instructions, and output responses. The medical imaging field is a strong multi-modal scenario. Doctors need to assist in diagnosis based on patients' images. Among them, writing reports based on images is one of the most important links. Writing imaging reports relies on extremely high medical imaging knowledge and takes a lot of time for doctors. At the same time, due to the uneven level of doctors, the quality of the written reports is also uncontrollable. Multi-modal large models have powerful image-text understanding capabilities and can complete vertical field image-text understanding tasks with high quality after being fine-tuned with domain data. Therefore, the generation of medical imaging reports based on multi-modal large models has become a research hotspot.
[0003] Among them, the method for generating medical imaging reports based on multi-modal large model preference alignment technology mainly refers to the technology of using multi-modal data (such as different types of medical imaging data such as CT, MRI, etc.) and machine learning models to automatically generate medical imaging reports. This technology integrates different modal imaging data, uses machine learning models to analyze and interpret the data, and thus automatically generates detailed diagnostic reports. The purpose is to improve the generation speed and accuracy of medical imaging reports, reduce the workload of medical staff, and provide more standardized and objective diagnostic information, so as to help doctors make better judgment decisions.
[0004] The prior art relies on large-scale unoptimized data for training, resulting in insufficient accuracy and standardization of the generated reports. The prior art often ignores the systematic integration of expert experience and fails to fully consider the diagnostic styles and professional preferences of doctors, resulting in deviations between the generated medical imaging reports and the actual diagnostic methods and word usage habits of doctors, making it impossible for doctors to quickly and accurately grasp the key information of patients' conditions, which constitutes an obstacle to doctors' diagnostic decisions. The prior art reacts slowly when dealing with real-time changing medical data and cannot update the clinical status in a timely manner, restricting the timeliness of diagnostic information. This not only slows down the medical diagnosis process but also has a negative impact on the subsequent treatment of patients. Summary of the Invention
[0005] The purpose of the present invention is to solve the disadvantages existing in the prior art, and propose a method for generating medical imaging reports based on multi-modal large model preference alignment technology.
[0006] To achieve the above object, the present invention adopts the following technical solutions: A medical image report generation method based on multi-modal large model preference alignment technology, comprising the following steps:
[0007] S1: Based on medical image data, collect known medical report data corresponding to the images, analyze the lesion descriptions and image interpretation vocabulary in the report text, and establish image-text association data;
[0008] S2: Based on the image-text association data, introduce image and text features through the VLM model, analyze the vector relationship between the image texture features and the lesion descriptions, identify text generation errors, and obtain text mapping weights;
[0009] S3: Based on the text mapping weights, analyze the expert's diagnostic terms, sentence structures, and information arrangement order, judge the matching error between the expert's report annotations and the generation parameters, and adjust the image-text sorting rules and sentence expression priorities to obtain preference alignment optimization parameters;
[0010] S4: Based on the preference alignment optimization parameters, analyze the historical lesion change trend, screen the abnormal image areas, and adjust the report generation content according to the current data, update and optimize the text generation parameters to obtain real-time text adjustment parameters;
[0011] S5: Based on the real-time text adjustment parameters, analyze the consistency between the report content and the image data, compare the matching degree between the diagnostic text and the image features, and adjust the abnormal or mismatched content in the report text to obtain a context-consistent diagnostic result.
[0012] The present invention is improved in that the image-text association data includes feature identifiers, text labels, and association mapping tables, the text mapping weights include feature vectors, description vectors, and weight indicators, the preference alignment optimization parameters include annotation preferences, sorting rules, and expression priorities, the real-time text adjustment parameters include change trend analysis results, abnormal identification marks, and content update instructions, and the context-consistent diagnostic results include text matching degrees, consistency evaluation results, and accuracy indicators.
[0013] The present invention is improved in that the acquisition steps of the image-text association data are specifically as follows:
[0014] S111: Based on medical image data, extract the key image features of the lesion area, use the image pixel distribution, edge gradient, and gray-level co-occurrence matrix features, and combine the geometric shape parameters of the lesion area to analyze the edge sharpness, density distribution, and texture complexity of the lesion area to obtain the image feature parameters of the lesion area;
[0015] S112: Based on the image feature parameters of the lesion area, collect the known medical report data corresponding to the image, extract the words related to the lesion description in the report text, count the occurrence frequency of each type of word, and at the same time calculate its distribution in different lesion types to obtain the lesion description word distribution parameters;
[0016] S113: Based on the lesion description word distribution parameters, use the formula:
[0017] ;
[0018] Calculate the correlation degree between the image features and the words, screen the key matching items with high relevance, and establish the image-text association data, where represents the image feature and the word the correlation degree between them, represents the image feature in the dimension the parameter value on, represents the word in the dimension the distribution parameter value on, represents the common dimension number of the image feature and the word.
[0019] The improvement of the present invention is that the specific steps for obtaining the text mapping weight are as follows:
[0020] S211: Based on the image-text association data, extract the image features and text features through the VLM model, analyze the mutual mapping relationship between the texture information of the image and the keywords in the text description, calculate the feature vectors of the image features and the text description, and screen the image feature vectors with high matching degree to obtain the image-text matching vector set;
[0021] S212: Based on the image-text matching vector set, calculate the image feature mapping weight, analyze and adjust the error value in the mapping weight, and use the formula:
[0022] ;
[0023] Optimize the feature weight to obtain the optimized image feature mapping weight , where represents the initial image feature mapping weight, represents the text feature vector, represents the image feature vector, represents the total number of feature dimensions;
[0024] S213: Based on the optimized image feature mapping weights, perform the matching of image features and text features, screen the image-text combinations with high matching degrees, eliminate the samples with low matching degrees, and then judge the deviation degree of the generated text to obtain the text mapping weights.
[0025] The improvement of the present invention is that the obtaining step of the preference alignment optimization parameter is specifically as follows:
[0026] S311: Based on the text mapping weights, collect the annotation data of radiologists in the reports, analyze the diagnostic terms, sentence structures, and information arrangement orders in the expert reports, extract the word frequency distributions of common medical terms, calculate the occurrence probabilities of sentence structures, and summarize the arrangement orders of different types of information in the expert reports to obtain the language features of the expert reports.
[0027] S312: Invoke the language features of the expert reports, compare the matching errors between the expert report annotations and the generated text in terms of vocabulary selection, sentence structure, and information sorting, analyze the vocabulary matching degree, sentence similarity, and information arrangement consistency, and adopt the formula:
[0028] ;
[0029] Calculate the text generation matching error value , where represents the total number of matching items, is the weight of the th matching item, and are the probability values of the th matching item of the expert report and the generated text at the nd feature point respectively, is the number of feature points for each matching item;
[0030] S313: Based on the text generation matching error value, adjust the image-text sorting rules and the priority of sentence expressions, reset the text expression rules, optimize the information arrangement order, and adjust the sentence weights of the generated text to obtain the preference alignment optimization parameter.
[0031] The improvement of the present invention is that the obtaining step of the real-time text adjustment parameter is specifically as follows:
[0032] S411: Based on the preference alignment optimization parameter, synchronously receive the patient's medical images and clinical data, analyze the historical lesion change trend, calculate the pixel change amount, density difference value, and contour offset amount of the current lesion area in multi-temporal images, compare them with the lesion evolution reference indicators, and screen the areas where the changes exceed the benchmark value to obtain the lesion change trend offset value.
[0033] S412: Based on the offset value of the lesion change trend, calculate the distribution density of the abnormal regions in the current image, identify the regions with prominent changes in the image, extract their morphological features, texture parameters, and spatial distribution, and compare with the historical image data to determine the change rate of the abnormal regions in the current image;
[0034] S413: Invoke the change rate of the abnormal regions to analyze the abnormal regions in the current image. Combining with the gray-scale features of the image, calculate the boundary gradient change value, mean deviation, and regional contrast of the abnormal regions, using the formula:
[0035] ;
[0036] Obtain the real-time text adjustment parameter , where, represents the boundary gradient change value of the th abnormal region, represents the key weight of the th abnormal region, represents the mean deviation of the current abnormal region, represents the mean of the historical lesion images, represents the th regional contrast of the abnormal region, represents the total number of abnormal regions in the image.
[0037] The improvement of the present invention is that the step of obtaining the context-consistent diagnosis result is specifically as follows:
[0038] S511: Based on the real-time text adjustment parameter, invoke the image data and the diagnostic report text, compare the matching situation between the image region description and the text content, calculate the consistency between the text and the image features, screen out the text segments with low matching degree, and obtain the text matching abnormal segments;
[0039] S512: Based on the text matching abnormal segments, analyze the text deviation features, calculate the difference degree between the text description and the image data, using the formula:
[0040] ;
[0041] Obtain the text deviation degree , where, represents the th numerical description in the text, represents the corresponding numerical value of the image data, represents the th term description in the text, represents the standard term description in the image data, is the total number of numerical items to be evaluated, is the total number of term items, The coefficient for adjusting the influence of term differences;
[0042] S513: Invoke the text deviation degree, adjust the image report text, optimize the matching relationship between the terms and values of the abnormal description, correct the text segment, and obtain a context-consistent diagnosis result.
[0043] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0044] In the present invention, by combining medical image data with the corresponding report, a high degree of automation in data processing is achieved, ensuring that the generated medical report is more accurate in details. The relationship between the image and the text is deeply analyzed using the VLM model, achieving accurate feature vector mapping, improving the scientific nature of text generation. Through a refined preference alignment process, the diagnostic preferences and actual operation habits of experts can be accurately reflected in report generation. At the same time, the diagnostic terms, sentence structures, and information arrangement order of experts are analyzed to better match the expert's annotation and diagnostic thinking mode, making the report more in line with clinical needs, improving the practicality and professionalism of the medical report. Real-time updating of text generation parameters ensures that the medical report can immediately reflect the latest medical conditions, improving the timeliness and accuracy of the report, effectively alleviating the work pressure of medical staff, and promoting the standardization and objectification of the diagnostic process. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is the main step flow chart of the present invention;
[0046] Figure 2 It is the flow chart for obtaining the image-text associated data in the present invention;
[0047] Figure 3 It is the flow chart for obtaining the text mapping weight in the present invention;
[0048] Figure 4 It is the flow chart for obtaining the preference alignment optimization parameters in the present invention;
[0049] Figure 5 It is the flow chart for obtaining the real-time text adjustment parameters in the present invention;
[0050] Figure 6 It is the flow chart for obtaining the context-consistent diagnosis result in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0052] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality" is two or more, unless otherwise specifically defined.
[0053] Embodiment
[0054] Please refer to Figure 1 , the present invention provides a technical solution: a method for generating a medical image report based on the preference alignment technology of a multimodal large model, including the following steps:
[0055] S1: Based on the medical image data, extract the key image features of the lesion area, collect the known medical report data corresponding to the image, analyze the lesion descriptions and image interpretation vocabulary in the report text, and perform matching processing on the image features and the vocabulary to establish image-text association data;
[0056] S2: Based on the image-text association data, introduce image and text features through the VLM model, analyze the vector relationship between the image texture features and the lesion descriptions, adjust the feature mapping weights, and identify the text generation errors to optimize the image feature matching degree to obtain the text mapping weights;
[0057] S3: Based on the text mapping weights, collect the annotation data of radiologists in the report, analyze the diagnostic terms, sentence structures and information arrangement order of the experts, judge the matching errors between the report annotations of the experts and the generation parameters, and adjust the image-text sorting rules and the priority of sentence expressions to obtain the preference alignment optimization parameters;
[0058] S4: Based on the preference alignment optimization parameters, synchronously receive the medical images and clinical data of the patient, analyze the historical lesion change trend, screen the abnormal image areas, and adjust the report generation content according to the current data, update and optimize the text generation parameters to match the clinical needs to obtain the real-time text adjustment parameters;
[0059] S5: Based on the real-time text adjustment parameters, analyze the consistency between the report content and the image data, compare the matching degree between the diagnostic text and the image features, adjust the abnormal or mismatched content in the report text, and optimize the accuracy of the report content to obtain the context-consistent diagnostic results.
[0060] The image-text associated data includes feature identifiers, text tags, and association mapping tables. The text mapping weights include feature vectors, description vectors, and weight metrics. The preference alignment optimization parameters include annotation preferences, sorting rules, and expression priorities. The real-time text adjustment parameters include change trend analysis results, anomaly recognition markers, and content update instructions. The context consistency diagnosis results include text matching degrees, consistency evaluation results, and accuracy metrics.
[0061] Please refer to Figure 2 , and the steps for obtaining the image-text associated data are specifically as follows:
[0062] S111: Based on the medical image data, extract the key image features of the lesion area. Using the image pixel distribution, edge gradient, and gray-level co-occurrence matrix features, combined with the geometric shape parameters of the lesion area, analyze the edge sharpness, density distribution, and texture complexity of the lesion area to obtain the lesion area image feature parameters;
[0063] First, obtain the image information of the lesion area. The image information of the lesion area can be obtained through CT, MRI, or ultrasound images. The image data contains feature information such as pixel distribution, edge gradient, and gray-level co-occurrence matrix. Taking the CT image as an example, in the gray-scale space, the density differences of different tissues result in different gray-level distributions. Calculate the average gray value, variance, and histogram distribution of the pixels in the lesion area of the image, extract its density features. At the same time, use the gradient operator to calculate the edge changes in the lesion area to obtain the edge gradient value, and combine the Canny operator to detect the edge intensity to evaluate the sharpness of the lesion edge. In addition, the gray-level co-occurrence matrix is used to analyze the detailed texture features in the image, and the texture complexity is measured by calculating parameters such as contrast, uniformity, and entropy value. For example, assuming that the average gray value of the lesion area is 120 and the standard deviation is 30, then its normalized gray distribution can be calculated. The edge change rate obtained by the gradient operator is 5.6%, the texture contrast obtained by the gray-level co-occurrence matrix analysis is 0.68, the uniformity is 0.85, and the entropy value is 2.3. According to the features, combined with the geometric shape parameters of the lesion, such as area, perimeter, circularity, etc., the morphological features of the lesion area can be further analyzed. For example, if the area of a certain lesion area is 25 mm2 and the perimeter is 18 mm, then calculate its circularity Get 0.97, indicating that the lesion morphology is close to circular. Normalize all the calculated lesion image feature parameters for subsequent matching analysis, and finally obtain the lesion area image feature parameters.
[0064] S112: Based on the lesion area image feature parameters, collect the known medical report data corresponding to the image, extract the words related to the lesion description in the report text, count the occurrence frequency of each type of word, and at the same time calculate its distribution in different lesion types to obtain the lesion description word distribution parameters;
[0065] Call the known medical report data corresponding to the image. Medical reports usually contain description information of lesions, such as size, density, and edge characteristics, etc. First, extract the words related to lesion description in the report text and count the frequency of word occurrences. For example, in 1000 reports, "high density" appears 230 times, "low density" appears 170 times, "fuzzy edge" appears 320 times, and "clear edge" appears 410 times. Classify all words and calculate their distribution in different lesion types. For example, the proportion of the word "high density" in benign lesions is 60%, and in malignant lesions is 40%. While the proportion of "fuzzy edge" in malignant lesions is 75%, and only accounts for 25% in benign lesions. For each type of lesion description word, calculate the importance coefficient of the word according to its proportion in different lesion types. For example, when the word occurrence proportion exceeds 50%, its importance coefficient is set to 1, otherwise it is set to 0.5 to improve the classification accuracy and obtain the distribution parameters of lesion description words.
[0066] S113: Based on the distribution parameters of lesion description words, use the formula:
[0067] ;
[0068] Calculate the correlation degree between the image features and the words, screen the key matching items with high correlation, and establish the image-text association data. Among them, represents the correlation degree between the image feature and the word represents the parameter value of the image feature in the dimension ; represents the parameter value of the word in the dimension ; represents the common number of dimensions of the image feature and the word;
[0069] Based on the distribution parameters of lesion description words, calculate the correlation degree between the image features and the words. Set the matching threshold to 0.6. If there are the following data:
[0070] Image feature parameters (lesion image data)
[0071] (gray mean value), (edge gradient), (texture contrast);
[0072] Distribution parameters of lesion description words
[0073] (parameter value corresponding to "high density"), (parameter value corresponding to "clear edge"), (parameter value corresponding to "smooth texture");
[0074] Calculate the score for each dimension item by item:
[0075] Calculate the first dimension :
[0076] ;
[0077] Calculate the second dimension :
[0078] ;
[0079] Calculate the third dimension :
[0080] ;
[0081] Calculate the total
[0082] ;
[0083] Since the matching threshold is 0.6 and the calculated is much less than 0.6, the matching degree between the image feature and the vocabulary is low and does not meet the matching condition, and this data point will be discarded.
[0084] Please refer to Figure 3 , and the specific steps for obtaining the text mapping weight are as follows:
[0085] S211: Based on the image-text association data, extract the image features and text features through the VLM model, analyze the mutual mapping relationship between the texture information of the image and the keywords in the text description, calculate the feature vectors of the image features and the text description, and screen the image feature vectors with high matching degree to obtain the image-text matching vector set;
[0086] Preprocess the image data to ensure that the resolution, grayscale value, and texture features of the images are comparable under the same standard. At the same time, convert the text description into structured data for feature extraction. For example, the gray-level co-occurrence matrix in the image data can be used to calculate features such as the contrast, correlation, and entropy of the image, while the keywords in the text data can have their weights calculated through word frequency statistics or TF-IDF. Analyze the local texture information of the image through the VLM model, map the pixel block features of the image into the feature vector space, and perform similarity calculation between the vectorized text description and the image feature vector. During the calculation process, it is necessary to normalize the data in different dimensions to ensure that the image features and text features can be compared on the same scale. Use cosine similarity to calculate the matching degree between the image feature vector and the text feature vector, set a similarity threshold, and filter out the image feature vectors with high matching degrees. The threshold can be determined by data statistics methods. For example, calculate the mean and standard deviation of the similarity distribution of all feature vectors, and use a certain multiple of the mean plus the standard deviation as the threshold. During implementation, images with different resolutions can be input into the model to extract their corresponding texture features, and check whether the keywords in the text description appear in the interpretation text corresponding to the image to obtain the image-text matching vector set.
[0087] S212: Based on the image-text matching vector set, calculate the image feature mapping weight, analyze and adjust the error value in the mapping weight, using the formula:
[0088] ;
[0089] Optimize the feature weight to obtain the optimized image feature mapping weight , where, represents the initial image feature mapping weight, represents the text feature vector, represents the image feature vector, represents the total number of feature dimensions, is a positive number to avoid a zero denominator;
[0090] Perform hierarchical calculation on the matching degree between the image feature vector and the text feature vector, calculate the weight factors at different levels based on the similarity, and adjust the mapping weight. By calculating the absolute error between the image feature and the text feature, obtain the initial error value, then calculate the total error sum on all feature dimensions to determine the error proportion. In specific implementation, the total number of feature dimensions is 2, that is, feature vectors in two dimensions are considered. Assume that in two dimensions, the text feature vector and the image feature vector have values of and in the first dimension, and and In the second dimension, simultaneously set to a very small positive number 0.001 to avoid the situation where the denominator is zero in the calculation. The initial image feature mapping weight is . First, calculate the absolute error in each feature dimension, and then sum the errors for use in the formula:
[0091] The absolute error in the first dimension:
[0092] ;
[0093] The absolute error in the second dimension:
[0094] ;
[0095] Next, calculate the total error:
[0096] ;
[0097] Substitute into the formula to calculate the optimized image feature mapping weight :
[0098] ;
[0099] ;
[0100] ;
[0101] ;
[0102] The optimized image feature mapping weight is , indicating that the mapping weight has decreased after optimization, reflecting the adjustment of the original weight to make it more accurately reflect the actual feature matching situation.
[0103] S213: According to the optimized image feature mapping weight, perform image feature and text feature matching, screen out the image-text combinations with high matching degrees, and eliminate the samples with low matching degrees. Then, judge the deviation degree of text generation to obtain the text mapping weight;
[0104] Based on the optimized image feature mapping weights, perform the matching of image features and text features. First, sort according to the matching degree of image features and text features, set a matching degree threshold, screen out the image-text combinations with high matching degrees, and eliminate the samples with low matching degrees. The calculation of the matching degree can be carried out through similarity calculation methods. For example, the normalized Euclidean distance is used to calculate the relative matching degree between different image-text pairs. Suppose the optimized matching weight of a certain image-text combination is 0.75, and the threshold is 0.7, then this combination is retained; otherwise, it is eliminated. After elimination, the remaining image-text combinations are used to calculate the text generation error. The error calculation method can adopt the text reconstruction method, convert the image features into text descriptions, and calculate the edit distance between it and the original text description to measure the deviation degree of text generation. If the original text description is "The density of the lung lesion increases" and the generated text is "The density of the lung lesion enhances", and the calculated edit distance is 1, it indicates that the error of text generation is small and a higher weight can be assigned to obtain the text mapping weight.
[0105] Please refer to Figure 4 , and the specific steps for obtaining the preference alignment optimization parameters are as follows:
[0106] S311: Based on the text mapping weights, collect the annotation data of radiology experts in the reports, analyze the diagnostic terms, sentence structures, and information arrangement orders in the experts' reports, extract the word frequency distributions of common medical terms, calculate the occurrence probabilities of sentence structures, and summarize the arrangement orders of different types of information in the experts' reports to obtain the language features of the experts' reports;
[0107] Collect the annotation data of radiology experts in the reports. First, obtain the sources of the experts' annotation data, including historical imaging reports, manually annotated imaging diagnosis data, and auxiliary annotation texts generated by imaging analysis systems. To ensure the standardization of the annotation data, it is necessary to perform word segmentation on the report content, remove irrelevant characters, and establish a common diagnostic vocabulary for experts. Calculate the occurrence probabilities of different terms through word frequency statistics, and summarize the distribution of each vocabulary in different imaging types. For example, in lung imaging reports, terms such as "nodule" and "ground-glass opacity" have a higher occurrence frequency, while in brain imaging reports, terms such as "ischemia" and "hemorrhage" are more common. The distribution frequency of terms can be calculated using the relative word frequency calculation formula. After setting the benchmark number of reports, calculate the relative occurrence probabilities of each term in the experts' reports. In addition, to analyze the sentence structure of the experts' reports, it is necessary to use grammar parsing technology to identify the subject-predicate-object structure in different reports and count the occurrence proportions of different sentence patterns. For example, in imaging reports, short sentences account for a relatively high proportion, such as "A small nodule can be seen in the lower lobe of the lung", while complete descriptive sentence patterns are less common, such as "The CT scan of the patient's lungs shows a small nodule shadow in the lower lobe". To further analyze the information arrangement order, it is necessary to break down the report content into different information blocks, such as "lesion type", "lesion location", "lesion degree", etc., and count the average arrangement order of each information block. For example, in imaging reports, doctors usually describe the location of the lesion first, then the specific type of the lesion, and finally the degree of the lesion. This order can be determined by calculating the average position index of each information block in the report, and the language characteristics of the experts' reports are obtained.
[0108] S312: Invoke the language characteristics of the experts' reports, compare the matching errors between the annotation of the experts' reports and the generated text in terms of vocabulary selection, sentence structure, and information sorting, analyze the vocabulary matching degree, sentence pattern similarity, and information arrangement consistency, and use the formula:
[0109] ;
[0110] Calculate the text generation matching error value , where represents the total number of matching items, is the weight of the th matching item, and are the probability values of the th feature points of the expert report and the generated text in the th matching item respectively, is the number of feature points for each matching item;
[0111] For calculating the overlap between the words used by experts and the words in the generated text in terms of lexical matching degree, a matching threshold is set. For example, when the number of occurrences of the same term in two texts differs by no more than a specific ratio, it is considered a match. Then, the sentence pattern similarity is analyzed. The syntactic tree matching algorithm is used to calculate the similarity in grammatical structure between the expert text and the generated text, comparing the arrangement of the subject, predicate, and object. If the sentence patterns are basically the same, the similarity is high; if there are significant adjustments in the sentence pattern structure in the generated text, the similarity is low. At the same time, for the consistency of information arrangement, it is necessary to count whether the order of each information block generated in the text is consistent with the order in the expert report, calculate the index deviation of each information block in the text, and take the average value of all deviation values for quantification. To comprehensively quantify the matching error, a formula is used to calculate the text generation matching error value. If the probability value of a feature point in the expert report for a certain matching item is , the corresponding probability value of the generated text is , calculate the absolute error between the feature points:
[0112] ;
[0113] Let , then calculate the normalized error:
[0114] ;
[0115] If the weight term , the matching error is calculated as follows:
[0116] ;
[0117] This result indicates that the matching error value of text generation is small, indicating that the generated text and the expert report have a high degree of consistency in terms of vocabulary, sentence pattern structure, and information arrangement.
[0118] S313: Based on the text generation matching error value, adjust the sorting rule of the image text and the priority of sentence pattern expression, re-set the text expression rule, optimize the information arrangement order, and adjust the sentence pattern weight of the generated text to obtain the preference alignment optimization parameter;
[0119] First, based on the calculation results of the matching error, filter out the text content with low matching degree and adjust its vocabulary selection to make it closer to the common terms in the expert report. The adjustment strategies include replacing the low-matching vocabulary in the generated text with the terms with higher word frequencies in the expert report and setting vocabulary replacement rules. For example, for the common "hyperplasia" in the imaging report, if the expert text prefers to use "thickening", it will be automatically replaced in the generated text. In addition, for the text with low matching degree in sentence structure, adjust the syntactic structure, rewrite the common long sentence structure into shorter sentences that are more in line with the expert's writing habits. For example, rewrite "The lung CT examination shows scattered small nodules in both lungs with clear boundaries" into "CT shows scattered small nodules in both lungs, with clear boundaries". In terms of information arrangement optimization, according to the arrangement error value of the information blocks, re-adjust the arrangement order of the generated text. For example, if the expert report usually describes the lesion location first and then the lesion type, adjust the arrangement rule of the generated text to conform to this order to obtain the preference alignment optimization parameters.
[0120] Please refer to Figure 5 , and the specific steps for obtaining the real-time text adjustment parameters are as follows:
[0121] S411: Based on the preference alignment optimization parameters, synchronously receive the patient's medical images and clinical data, analyze the historical lesion change trend, calculate the pixel change amount, density difference value, and contour offset amount of the current lesion area in multi-temporal images, compare them with the reference indicators of lesion evolution, and filter out the areas where the changes exceed the benchmark value to obtain the lesion change trend offset value;
[0122] Synchronously receive the patient's medical images and clinical data, call the historical lesion change trend data, process the multi-temporal images, extract the lesion area features at multiple consecutive time points, calculate the pixel change amount, density difference value, and contour offset amount. Among them, the pixel change amount is calculated by statistically analyzing the pixel gray level difference between adjacent time points, the density difference value is evaluated based on the pixel aggregation situation, and the contour offset amount is calculated by using the edge detection method to calculate the change of the lesion contour. For example, for a patient's lung image, when calculating the pixel change amount, select the pixel gray level values of the lesion area and calculate the gray level change rate between the current frame and the previous frame. If the pixel gray level value of a certain area changes from 120 to 135, then its change rate is (135 - 120) / 120 = 12.5%. The density difference value can be obtained by statistically analyzing the number of high-gray-level pixels per unit area, and the contour offset amount can be determined by calculating the displacement distance of the edge points at different time points. For example, if the coordinates of a certain contour point in two frames of images are (50, 60) and (55, 65) respectively, then its offset amount is calculated as , compare the above calculation results with the reference indicators of lesion evolution, and screen the areas where the changes exceed the baseline values. For example, set the pixel change threshold to 10%, the density difference threshold to 15%, and the contour offset threshold to 5 pixel points. Then, if the pixel change amount in a certain area exceeds 10%, or the density difference value exceeds 15%, or the contour offset amount exceeds 5 pixel points, it is determined as a significantly changed lesion area, and the lesion change trend offset value is obtained.
[0123] S412: Based on the lesion change trend offset value, calculate the distribution density of abnormal areas in the current image, identify the areas with prominent changes in the image, extract their morphological features, texture parameters and spatial distribution, and compare with the historical image data to determine the change rate of abnormal areas in the current image;
[0124] Calculate the distribution density of abnormal areas in the current image, call the screened lesion change areas, count the number of lesion pixel points per unit area, and determine the spatial distribution characteristics of the lesions. For example, for a certain lung image area, select a 10×10 pixel grid unit, count the number of high gray value pixel points in each unit, and calculate its proportion. If the proportion of lesion pixels in a certain grid unit is higher than the set threshold, then this unit is marked as a high-density lesion area, and then identify the areas with prominent changes in the image, extract their morphological features, texture parameters and spatial distribution. The morphological features can be obtained by calculating the area, perimeter and shape parameters of the lesion area. For example, the area of a certain lesion area is 200 pixel points and the perimeter is 50 pixel points. The texture parameters are calculated based on the gray level co-occurrence matrix for uniformity, contrast, entropy, etc. For example, when calculating the contrast parameter of a lesion area, assume that the mean square value of the difference between pixel values in a certain area and adjacent pixels is 25, then the contrast value is 25. Finally, compare the abnormal areas in the current image with the historical image data, calculate the area growth rate and shape change rate of the lesion area at different time points, and determine the change rate of the abnormal area.
[0125] S413: Call the change rate of abnormal areas, analyze the abnormal areas in the current image, combine with the gray level features of the image, calculate the boundary gradient change value, mean deviation and regional contrast of the abnormal areas, and use the formula:
[0126] ;
[0127] Obtain the real-time text adjustment parameters , where represents the boundary gradient change value of the th abnormal area, represents the key weight of the th abnormal area, represents the mean deviation of the current abnormal area, represents the mean of the historical lesion image, represents the The regional contrast of an abnormal area represents the total number of abnormal areas in the image;
[0128] Analyze the abnormal areas in the current image, calculate the boundary gradient change value, mean deviation, and regional contrast of the abnormal areas. The boundary gradient change value is obtained by calculating the gray-scale gradient change rate of the pixels at the lesion boundary. For example, if the gray-scale gradients of adjacent pixels at the boundary of a certain lesion in two frames of images are 10 and 15 respectively, then the gradient change rate is calculated as (15 - 10) / 10 = 50%. The mean deviation calculates the difference between the gray-scale mean value of the lesion area in the current image and the gray-scale mean value of the historical lesion. For example, if the current gray-scale mean value of the lesion is 130 and the historical mean value is 120, then the mean deviation is |130 - 120| = 10. The regional contrast is obtained by calculating the ratio of the gray-scale mean square error between the lesion area and the surrounding background area. For example, if the gray-scale variance of the lesion area is 400 and the gray-scale variance of the background area is 100, then the regional contrast is . If the current image contains 3 abnormal areas, and their boundary gradient change values are respectively 、 and , and their critical weights are respectively 、 and , the mean deviation of the current abnormal area , the mean value of the historical lesion image , and the regional contrasts are respectively 、 and , substitute into the formula for calculation:
[0129] ;
[0130] ;
[0131] ;
[0132] This result indicates the real-time text adjustment parameter , meaning that the degree of change of the abnormal areas in the current image is relatively high, and it is necessary to adjust the image text description according to this parameter to match the actual image characteristics.
[0133] Please refer to Figure 6 , and the specific steps for obtaining the context-consistent diagnosis result are as follows:
[0134] S511: Based on the real-time text adjustment parameter, call the image data and the diagnostic report text, compare the matching situation between the image area description and the text content, calculate the consistency between the text and the image characteristics, screen out the text segments with low matching degree, and obtain the text matching abnormal segments;
[0135] Call the imaging data and diagnostic report text, extract the key descriptive information of the imaging region, including parameters such as lesion morphology, size, boundary clarity, density, and signal intensity, convert it into a standardized numerical vector, and extract the corresponding descriptive content for the diagnostic report text, perform text parsing and key attribute extraction. The parsed text features include disease name, lesion location, lesion size, and lesion development trend, etc. For each imaging region and text description content, calculate the similarity respectively. Among them, the calculation method for text and imaging feature matching uses the cosine similarity method to calculate the matching degree between the text segment and the imaging data vector. If the lesion size description in the text is "2.5 cm in diameter", and the actual diameter calculated from the imaging data is 3.0 cm, the similarity calculation method is as follows: , where, is the text vector, is the imaging data vector. If the calculated similarity is lower than the standard, it is determined that there is a deviation between the text description and the imaging data. Further screen the text segments with abnormal matching, extract the text regions with low matching degree, and obtain the text matching abnormal segments.
[0136] S512: Based on the text matching abnormal segments, analyze the text deviation features, calculate the difference degree between the text description and the imaging data, using the formula:
[0137] ;
[0138] Get the text deviation degree , where, represents the th numerical description in the text, represents the corresponding imaging data value, represents the th term description in the text, represents the standard term description in the imaging data, is the total number of numerical items to be evaluated, is the total number of term items, is the coefficient to adjust the influence of term differences;
[0139] First, calculate the deviation between the numerical values involved in the text description and the imaging data. For the numerical descriptions in the text, such as lesion size, lesion signal intensity, lesion range, etc., compare them with the numerical values calculated from the imaging data, and calculate the sum of squared deviations. If the lesion size described in the text is "2.5 cm", and the imaging measurement value is "3.0 cm", the deviation calculation is In addition, the difference of term description is calculated, and the medical terms in the text are extracted and compared with the standard medical term library. For example, the text describes "mass", while the imaging standard term is "nodule". The difference value of the terms between the two is calculated, and the text deviation is calculated using the formula. Assume that the text description contains 3 numerical items, and their numerical descriptions are , and , and the image data values are , and , calculate the sum of squared deviations:
[0140] ;
[0141] ;
[0142] Assume that the text terms are “mass” and “irregular border”, and the corresponding image standard terms are “nodule” and “fuzzy border”. ,set up , then the term deviation is calculated as:
[0143] ;
[0144] calculate :
[0145] ;
[0146] The result shows that the text deviation is 1.15, indicating that there is a certain deviation between the text description and the image data, and text correction is needed.
[0147] S513: calling the text deviation degree, adjusting the image report text, optimizing the matching relationship between the terms and numerical values of the abnormal description, correcting the text fragments, and obtaining the context-consistent diagnosis result;
[0148] First, correct the numerical descriptions with large deviations. For example, if the text describes "2.5cm" and the image measurement value is "3.0cm", the text needs to be adjusted to "3.0cm" to ensure that the numerical value is consistent with the image data. Then optimize the term description, for example, correct "lump" to "nodule" to improve the term matching degree. For the description errors of lesion features, calculate the morphology, signal, size and other parameters in the text, and compare them with the standard features of the image data. If the image data determines that the lesion boundary is clear, but the text description is "fuzzy boundary", it needs to be adjusted to "clear boundary" to complete the text correction and obtain a diagnosis result consistent with the context.
[0149] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A method for generating medical image reports based on the preference alignment technology of multi-modal large models, characterized in that, It includes the following steps: S1: Based on medical image data, collect known medical report data corresponding to the images, analyze the lesion descriptions and image interpretation vocabulary in the report text, and establish image-text association data; S2: Based on the image-text association data, introduce image and text features through the VLM model, analyze the vector relationship between the image texture features and the lesion descriptions, and obtain the text mapping weights; S3: Based on the text mapping weights, analyze the expert diagnosis terms, sentence structures, and information arrangement order, judge the matching error between the expert's report annotations and the generation parameters, and adjust the image-text sorting rules and sentence expression priorities to obtain the preference alignment optimization parameters; S4: Based on the preference alignment optimization parameters, analyze the historical lesion change trends, screen the abnormal regions of the images, and adjust the report generation content according to the current data, update and optimize the text generation parameters, and obtain the real-time text adjustment parameters; The specific steps for obtaining the real-time text adjustment parameters are as follows: S411: Based on the preference alignment optimization parameters, synchronously receive the patient's medical images and clinical data, analyze the historical lesion change trends, calculate the pixel change amount, density difference value, and contour offset amount of the current lesion region in multi-temporal images, and compare with the lesion evolution reference indicators to screen the regions where the changes exceed the benchmark value to obtain the lesion change trend offset value; S412: Based on the lesion change trend offset value, calculate the distribution density of the abnormal regions in the current image, identify the regions with prominent changes in the image, extract their morphological features, texture parameters, and spatial distribution, and compare with the historical image data to determine the abnormal region change rate of the current image; S413: Invoke the abnormal region change rate to analyze the abnormal regions in the current image, combine the image gray-scale features, calculate the boundary gradient change value, mean deviation, and regional contrast of the abnormal regions, and use the formula: ; Obtain real-time text adjustment parameters , where represents the boundary gradient change value of the th abnormal region, represents the key weight of the th abnormal region, represents the mean deviation of the current abnormal region, represents the mean of historical lesion images, represents the th regional contrast of the abnormal region, represents the total number of abnormal regions in the image; S5: Based on the real-time text adjustment parameters, analyze the consistency between the report content and the image data, compare the matching degree between the diagnostic text and the image features, and adjust the abnormal or mismatched content in the report text to obtain the context-consistent diagnostic results.
2. The method for generating a medical image report based on the multi-modal large model preference alignment technology according to claim 1, wherein The image-text association data includes feature identifiers, text tags, and association mapping tables. The text mapping weights include feature vectors, description vectors, and weight indicators. The preference alignment optimization parameters include annotation preferences, sorting rules, and expression priorities. The real-time text adjustment parameters include change trend analysis results, abnormal recognition marks, and content update instructions. The context-consistent diagnostic results include text matching degrees, consistency evaluation results, and accuracy indicators.
3. The method for generating a medical image report based on the multi-modal large model preference alignment technology according to claim 1, wherein The specific steps for obtaining the image-text association data are as follows: S111: Based on medical image data, extract the key image features of the lesion region, use the image pixel distribution, edge gradient, and gray-level co-occurrence matrix features, and combine the geometric shape parameters of the lesion region to analyze the edge clarity, density distribution, and texture complexity of the lesion region to obtain the lesion region image feature parameters; S112: Based on the image feature parameters of the lesion area, collect the known medical report data corresponding to the image, extract the words related to the lesion description in the report text, count the occurrence frequency of each type of word, and at the same time calculate its distribution in different lesion types to obtain the lesion description word distribution parameters; S113: Based on the lesion description word distribution parameters, use the formula: ; Calculate the correlation between image features and vocabulary, screen out the key matching items with high correlation, and establish image-text association data, where, represents the image feature and the vocabulary the correlation between them, represents the image feature in the dimension the parameter value on it, represents the vocabulary in the dimension the distribution parameter value on it, represents the common number of dimensions of the image feature and the vocabulary.
4. The method for generating a medical image report based on the multi-modal large model preference alignment technology according to claim 1, wherein The specific steps for obtaining the text mapping weight are as follows: S211: Based on the image-text association data, extract the image features and text features through the VLM model, analyze the mutual mapping relationship between the texture information of the image and the keywords in the text description, calculate the feature vectors of the image features and the text description, and screen the image feature vectors with a high degree of matching to obtain the image-text matching vector set; S212: Based on the image-text matching vector set, calculate the image feature mapping weight, analyze and adjust the error value in the mapping weight, and use the formula: ; Optimize the feature weights to obtain the optimized image feature mapping weights , where represents the initial image feature mapping weights, represents the text feature vector, represents the image feature vector, represents the total number of feature dimensions; S213: According to the optimized image feature mapping weight, perform the matching of image features and text features, screen the image-text combinations with a high degree of matching, eliminate the samples with a low degree of matching, and then judge the deviation degree of the text generation to obtain the text mapping weight.
5. The method for generating a medical image report based on the multi-modal large model preference alignment technology according to claim 1, wherein The specific steps for obtaining the preference alignment optimization parameters are as follows: S311: Based on the text mapping weight, collect the annotation data of the radiologist in the report, analyze the diagnostic terms, sentence structures, and information arrangement order in the expert report, extract the word frequency distribution of common medical terms, calculate the occurrence probability of the sentence structures, and summarize the arrangement order of different types of information in the expert report to obtain the language features of the expert report; S312: Invoke the language features of the expert report, compare the matching errors between the expert report annotation and the generated text in terms of word selection, sentence structure, and information sorting, analyze the word matching degree, sentence similarity, and information arrangement consistency, and use the formula: ; Calculate the matching error value of the generated text , where represents the total number of matching items, is the weight of the th matching item, and are the probability values of the th feature point of the expert report and the generated text for the th matching item respectively, is the number of feature points for each matching item; S313: Based on the text generation matching error value, adjust the image-text sorting rule and the priority of sentence expression, re-set the text expression rule, optimize the information arrangement order, and adjust the sentence weight of the generated text to obtain the preference alignment optimization parameters.
6. The method for generating a medical image report based on the multi-modal large model preference alignment technology according to claim 1, wherein The specific steps for obtaining the context-consistent diagnosis result are as follows: S511: Based on the real-time text adjustment parameters, call the image data and the diagnostic report text, compare the matching situation between the image area description and the text content, calculate the consistency between the text and the image features, and screen the text segments with a low degree of matching to obtain the text matching abnormal segments; S512: Based on the text matching abnormal segments, analyze the text deviation features, calculate the difference degree between the text description and the image data, and use the formula: ; Obtain the text deviation degree , where represents the th numerical description in the text, represents the corresponding image data value, represents the th term description in the text, represents the standard term description in the image data, is the total number of numerical items to be evaluated, is the total number of term items, is the coefficient for adjusting the influence of term differences; S513: Invoke the text deviation degree, adjust the image report text, optimize the matching relationship between the terms and numerical values of the abnormal description, and correct the text segment to obtain the context-consistent diagnosis result.
Citation Information
Patent Citations
PET / MRI (positron emission tomography / magnetic resonance imaging) image data governance model training method and device and data governance method
CN119601206A
Medical image report automatic quality control error correction system and method
CN119724466A