Visual language model-oriented medical image text generation method and system

By performing connected component analysis and feature extraction on medical images, structured natural language descriptions are generated, which solves the problem of strict requirements for text information quality in medical visual language models and improves the performance of the models.

CN120913735APending Publication Date: 2025-11-07SHANDONG UNIV

Patent Information

Application Number
CN202510753760.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing medical visual language models have strict requirements for the quality of text information. Low-quality or blurry text may introduce noise, leading to a decline in model performance. This problem is particularly prominent in medical image processing, and the training data lacks professionalism and consistency.

Method used

By acquiring labeled medical images, connected component analysis is performed to extract morphological features of the tumor region, such as size, shape, location, cavity features, and edge shape. These features are then converted into structured natural language descriptions, and high-quality text descriptions are generated using preset text templates.

Benefits of technology

The generated text descriptions are highly consistent and professional, ensuring that key features are accurately described, providing reliable supervision signals for model training, and improving model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913735A_ABST
    Figure CN120913735A_ABST
Patent Text Reader

Abstract

The invention provides a medical image text generation method and system oriented to a visual language model, and belongs to the technical field of medical image processing. The method comprises the following steps: acquiring a labeled image of a medical image; performing connected domain analysis on the annotated image, and detecting and counting the number of tumor areas in the annotated image; for each tumor area, extracting morphological features including size, shape, position, cavity features and edge shape; according to a preset text template, the extracted morphological features are converted into structured natural language description, and final medical image text description is output. It is ensured that the generated medical image description has high consistency and specialty, and subjectivity and difference of manual writing are avoided. It can be ensured that key morphological characteristics (such as tumor size, shape and position) are accurately and completely described, and reliable supervision signals are provided for subsequent model training.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of medical image processing, and particularly relates to a medical image text generation method and system oriented to a vision language model. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute prior art.

[0003] With the continuous progress of deep learning technology, especially under the driving of the increasing computing power and data resources, traditional models relying on a single modality (such as images or texts) have gradually been difficult to meet the needs of high-precision tasks. Among many application fields, medical image processing is particularly significant. Since medical images are usually information-intensive and complex in structure, it is often difficult to capture subtle differences and semantic features of diseases by relying only on image modalities. Therefore, the combination of multi-modal information, especially the combination of image and text information, has become an important means to improve the performance of models.

[0004] In recent years, with the rapid development of artificial intelligence technology in the field of medical image analysis, vision language models (VLMs) have become an important tool for medical image understanding. Such models can perform high-value tasks such as medical image segmentation, diagnosis report generation, and image retrieval by combining visual information and natural language descriptions. However, existing medical vision language models have strict requirements for the quality of text information. High-quality, structured text not only provides clear semantic guidance for the model, but also provides stronger supervision signals during model learning. On the contrary, low-quality or ambiguous text may introduce noise, which may weaken the performance of the model. This problem is particularly prominent in the field of medical image processing.

[0005] Currently, most of the medical text data available for training vision language models have uneven quality, lack of professionalism or consistency. For example: different doctors have different ways of expressing the same image, and there is a lack of standardized templates. It is easy to have problems of information redundancy or lack, so that part of the description contains irrelevant content, and key features (such as tumor shape, location) may not be mentioned. SUMMARY

[0006] To overcome the shortcomings of the prior art, this invention provides a method and system for generating medical image text for visual language models. This method can identify and understand tumor regions in medical image data (e.g., MRI images of breast cancer) and generate accurate natural language descriptions that can be used by a "Visual Language Model (VLM)". These descriptions not only express morphological features in the image (such as the size, shape, and location of the tumor) but also provide semantic information, enabling the model to understand the image content at a higher level.

[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0008] The first aspect of this invention provides a method for generating medical image text based on a visual language model;

[0009] A method for generating medical image text based on a visual language model includes:

[0010] Obtain an annotated image of a medical image; the annotated image is a binary image, wherein the first pixel value of the binary image represents the target tumor region in the binary image; and the second pixel value represents the background region of the binary image;

[0011] Connectivity analysis is performed on the labeled image to detect and count the number of tumor regions in the labeled image;

[0012] For each tumor region, morphological features including size, shape, location, cavity characteristics, and edge shape are extracted;

[0013] Based on the preset text template, the extracted morphological features are converted into structured natural language descriptions, and the final medical image text description is output.

[0014] As a further technical solution, connected component analysis is performed on the labeled image to detect and count the number of tumor regions in the labeled image, including:

[0015] Obtain the set of contour boundary points of the target tumor region and the background region in the labeled image;

[0016] C i ={(x,y)∈Ω| There exists a path connecting (x0,y0) and (x,y) such that I(x k ,y k )=255}

[0017] Among them, C i It is the i-th connected region, and the value of i is the number of tumor regions in the single slice to be extracted.

[0018] As a further technical solution, the process of extracting the size feature of the tumor region is:

[0019] The pixel area of the tumor region is calculated.

[0020] The size of the tumor region is classified according to a preset area threshold.

[0021] As a further technical solution, the process of extracting the shape feature of the tumor region is: obtaining the minimum circumscribed rectangle of the tumor region, and judging the shape of the tumor region according to the aspect ratio of the minimum circumscribed rectangle.

[0022] As a further technical solution, the process of extracting the position feature of the tumor region is:

[0023] The center point coordinates of the tumor region are calculated, and the position region of the labeled image is divided; according to the relationship between the center point coordinates and the divided position region, the position feature of the tumor region is determined.

[0024] As a further technical solution, the process of extracting the hollow feature of the tumor region is:

[0025] A hierarchical contour detection algorithm is used to identify the sub-contour in the target region, and when the sub-contour exists, it is determined that the target region contains a hollow.

[0026] As a further technical solution, the process of extracting the edge shape of the tumor region is:

[0027] The convex hull of the tumor region is obtained; the convex hull is the smallest convex polygon containing the contour point set of the tumor region.

[0028] The concave region between the original contour and the convex hull is detected.

[0029] When the ratio of the concave depth to the size of the target region exceeds a preset threshold, it is determined that the edge of the target region has a concave.

[0030] The second aspect of the present application provides a medical image text generation system for a visual language model.

[0031] A medical image text generation system for a visual language model, comprising:

[0032] A labeled image acquisition module configured to acquire a labeled image of a medical image.

[0033] A connected domain analysis module configured to perform connected domain analysis on the labeled image, detect and count the number of tumor regions in the labeled image.

[0034] A morphological feature extraction module configured to extract morphological features including size, shape, position, hollow feature and edge shape for each tumor region.

[0035] The text description output module is configured to convert the extracted morphological features into a structured natural language description according to a preset text template, and output a final medical image text description.

[0036] The third aspect of the present application provides a computer readable storage medium, which stores a program, and the program is executed by a processor to implement the steps of the medical image text generation method based on a visual language model according to the first aspect of the present application.

[0037] The fourth aspect of the present application provides an electronic device, which comprises a memory, a processor, and a program stored in the memory and executable on the processor, and the processor executes the program to implement the steps of the medical image text generation method based on a visual language model according to the first aspect of the present application.

[0038] The above one or more technical solutions have the following beneficial effects:

[0039] The present application generates a medical image description with high consistency and professionalism through a standardized feature extraction algorithm and a preset text template, avoids the subjectivity and difference of manual writing, and ensures that the key morphological features (such as tumor size, shape, position, etc.) are accurately and completely described, providing reliable supervision signals for subsequent model training.

[0040] The method provided by the present application has strong applicability, although breast cancer MRI is taken as an example for verification, but the method is applicable to various region annotation images, including other types of tumors (such as lung cancer, brain tumor, etc.) or organ structure delineation (such as liver, heart, etc.), not limited by image type, and has high portability and practical application value.

[0041] The advantages of the additional aspects of the present application will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0042] The drawings accompanying the specification of the present application form part of the present application and serve to provide further understanding of the present application, the illustrative embodiments of the present application and their descriptions serve to explain the present application, and do not constitute an improper limitation on the present application.

[0043] Figure 1 The method flowchart of the first embodiment.

[0044] Figure 2 The annotation image position region division schematic diagram of the first embodiment.

[0045] Figure 3A system structure diagram of a second embodiment. DETAILED DESCRIPTION

[0046] It should be noted that the following detailed description is exemplary in nature and is intended to provide further description of the application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0047] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments in accordance with the present application.

[0048] In the case of no conflict, the embodiments in the application and the features in the embodiments can be combined with each other.

[0049] Embodiment one

[0050] The embodiment discloses a medical image text generation method for a visual language model;

[0051] As shown in the figure, a medical image text generation method for a visual language model comprises: Figure 1

[0052] Step S1, acquiring a labeled image of a medical image;

[0053] Step S2, performing connected component analysis on the labeled image, detecting and counting the number of tumor regions in the labeled image;

[0054] Step S3, for each tumor region, extracting morphological features including size, shape, position, hollow feature and edge shape;

[0055] Step S4, according to a preset text template, converting the extracted morphological features into structured natural language description, and outputting the final medical image text description.

[0056] In this embodiment, the tumor labeled image of breast cancer is adopted, and the labeled image is a binary image, wherein the first pixel value of the binary image represents the target tumor region in the binary image, which is marked with pure white; the second pixel value represents the background region of the binary image, which is marked with pure black.

[0057] Step S2, performing connected component analysis on the labeled image, detecting and counting the number of tumor regions in the labeled image.

[0058] Specifically, the labeled image is represented by a two-dimensional function, as follows:

[0059] I(x,y)∈{0,255};

[0060] ​Where, I(x, y) = 255 represents white (tumor area), I(x, y) = 0 represents black (background area).

[0061] The contour boundary point set of the target tumor area and the background area in the labeled image is obtained, and is represented as follows:

[0062]

[0063] Where, Ω is the white area set, N(x, y) is the neighborhood of pixel (x, y) (generally 8-neighborhood), (x', y') represents a pixel point in the neighborhood of (x, y), I(x, y) is the gray value or label value (that is, 0 or 1 of the binary image) of the pixel point (x, y) in the image;

[0064] The number of tumor areas is extracted by connected domain, as follows:

[0065] C i = {(x, y) e Ω | there is a path connecting (x0, y0) and (x, y) such that I(x k ,y k ) = 255};

[0066] Where, C i is the i-th connected region; the value of i is the number of tumor areas (white areas) in the single slice to be extracted. The connected domain C i is a set of points (x, y) in the region Ω, which are connected to a starting point (x0, y0) and have a pixel value of 255. In other words, if there is a path from the starting point (x0, y0) to (x, y), and the pixel value I(x k ,y k ) = 255 at each point on the path, then these points belong to the same connected domain C i .

[0067] Further, for the convenience of subsequent judgment of all features except the number of tumors, the number of tumor areas can also be identified by introducing contour detection. Specifically:

[0068] The pixels along the edge of the tumor area are traversed, and the positions of the boundary points are recorded to form an ordered contour point set Γ = ({(x1, y1), (x2, y2), … (x n ,y n}), which satisfies:

[0069] And ‖(x i ,y i )-(x i+1 ,y i+1 )‖1 = 1

[0070] where ||·||1 represents the Manhattan distance (L1 norm) between the point (x i ,y i ) and the point (x i+1 ,y i+1 ) is equal to 1;

[0071] Based on the above method, when judging subsequent features, a tumor region contour Γ in a slice is taken as a unit.

[0072] Step S3, for each tumor region, morphological features including size, shape, position, cavity feature and edge shape are extracted.

[0073] Step S31, in the process of obtaining the size of the tumor region, according to the tumor region C i obtained in step S2, the size of the current tumor region can be obtained by only calculating the number of pixel points in a single tumor region. By setting a first threshold and a second threshold, the size of the tumor region is divided into three categories: “small, medium, large”, 0-first threshold belongs to small, first threshold-second threshold belongs to medium, and second threshold or above belongs to large. The setting method of the first threshold and the second threshold is: count the area of the tumor region in all slices in the data set respectively, and draw the corresponding column chart or point chart, observe the specific value distribution, so that the number ratio of images of three sizes is approximately 1:1:1, which can ensure the reasonable distribution of data described by different sizes, thereby enhancing the utilization rate of visual language model for text information of tumor size.

[0074] Further, in this step, the size feature can also be obtained by obtaining the contour of the tumor region in step S2. Specifically, Green's formula is used to integrate continuous curves and is commonly used to calculate the area of the region surrounded by a simple closed curve. In the pixel application of the present embodiment, a discrete Green's formula can be used, as follows:

[0075]

[0076] In the formula, area is the area contained by the contour; x i , y i are the horizontal and vertical coordinates of the i-th point on the contour, respectively;

[0077] Step S32, in the process of obtaining the shape of the tumor region, for the contour point set Γ=({(x1,y1),(x2,y2),…(x i ,y i}), first determine the minimum and maximum values of X and Y:

[0078] x min = MINi (x i )

[0079] x max = MIN i (x i )

[0080] y min = MAX i (y i )

[0081] y max = MAX i (y i )

[0082] According to the determined minimum and maximum values of X and Y, an initial circumscribed matrix (non-minimum circumscribed matrix) is constructed, the upper left corner coordinates of which are (x min ,y min ), the length h = x max -x min , the width w = y max -y min , and the area area = h x w.

[0083] The point set Γ is rotated around the angle θ, and the new coordinates after rotation are:

[0084]

[0085] According to the new area calculated from the new coordinates after rotation, all θ are traversed to find the minimum area, that is, the minimum circumscribed rectangle is found.

[0086] Further, through the proportional relationship of the long numerical value, it can be considered that when the proportion is between 0.8-1.2, the shape is a square, and when the proportion is less than 0.8 or greater than 1.2, the shape is a rectangle.

[0087] Step S33, the position coordinates (x min ,y min ) of the upper left corner of the minimum circumscribed rectangle and the length and width h, w are obtained. In order to determine the specific position of the tumor area, the center point coordinates of the minimum circumscribed matrix need to be known, and the specific calculation method is:

[0088]

[0089] The obtained (x center ,y center ) is the center point coordinates.

[0090] Further, in combination with the attached Figure 2The whole image is divided into regions, and in this embodiment, is divided into 49 small regions. One 3*3 region in the upper left, upper right, lower left, and lower right of the 49 small regions is taken as one large region (here, there are 4 large regions), and then the 1*1 small region in the center of the image is directly taken as one large region (here, there is 1 large region). After the 5 large regions are removed, the whole image is left with four 1*3 or 3*1 regions, which are taken as another 4 large regions. In this way, the whole image is divided into 9 large regions.

[0091] If the center point coordinate falls in any one of the 9 large regions, text information is directly generated, and the order from left to right and from top to bottom is "left top", "middle top", "right top", "left middle", "middle middle", "right middle", "left bottom", "middle bottom", and "right bottom", i.e., "left top", "middle top", "right top", "left middle", "middle middle", "right middle", "left bottom", "middle bottom", and "right bottom". (Corresponding to the numbers 1-9 in the figure)

[0092] In step S34, in the process of judging whether the tumor region contains a cavity, for the contour point set Γ = ({(x1, y1), (x2, y2), … (x i , y i )}, when there are multiple contour point sets Γ1, Γ2, … Γ n , it is possible that one contour is contained in another contour, i.e., Γ i ∈Γ j . At this time, if the coordinates of all points of Γ i are in Γ j , i.e., Γ i ∈Γ j , Γ i is a black cavity in Γ j .

[0093] In step S35, according to the determined tumor region contour Γ = {p1, p2, …, p n}, p i = (x i , y i ), whether there is a depression is judged by judging the circumscribed convex hull of the tumor region. Specifically, the convex hull refers to the smallest convex polygon containing the contour point set Γ, denoted as Conv(Γ), which satisfies the following properties:

[0094]

[0095] A concave is a point on the original contour that falls inside the convex hull, but not on the convex hull boundary. Counting the above case as one concave:

[0096] D k = (p start , p end , p farthest )

[0097] where p start , p end are two points on the convex hull, forming a segment of the convex hull boundary; p farthest is the point on the original contour that is farthest from the segment;

[0098] The depth of the concave is defined as:

[0099]

[0100] The specific formula is:

[0101]

[0102] In the process of judging whether the contour has a concave, it is judged whether the depth of each concave is greater than a proportional threshold. By normalizing the concave depth d k and the contour size (the diagonal length L of the minimum circumscribed rectangle is used in this embodiment):

[0103]

[0104] where w and h are the width and height of the minimum circumscribed rectangle, respectively.

[0105] When the ratio of the two is greater than a certain value (0.08 is taken here), it is considered that the concave is large enough, and it is determined that the contour has a concave, that is:

[0106] Step S4, according to the preset text template, converting the extracted morphological features into structured natural language description, outputting the final medical image text description.

[0107] First, determine how many tumor regions there are in the slice binary image, then traverse these tumor regions, that is, the outer contour of the white region, in each traversal loop, extract the size, shape, position, whether it contains a hollow, and the edge line shape of the single outer contour, and finally output to the preset text template. Take three labeled images as an example, the output content is shown in the following table:

[0108] Table 1 final medical image text description table

[0109]

[0110] Further, in the embodiment, verification is performed on three breast cancer data sets DUKE, BreaDM and ISPY1 respectively, specifically: using a visual language large model LViT, training a model once when the text information is empty, and then training a model using the text information generated by the tool, and the results of the two are shown in Table 2:

[0111] Table 2 Model training results under different text information

[0112] Indicator: Dice DUKE BreaDM No textual information 0.80 0.74 Textual information acquired by the invention 0.836 0.79

[0113] The evaluation index is the Dice coefficient, and the calculation formula of the Dice coefficient is as follows:

[0114]

[0115] wherein N represents the total number of pixels, C represents the number of categories, p ij , y ij respectively represent the probability of the i-th pixel being predicted as the j-th category and the corresponding true label, and the Dice coefficient is used to represent the overlap degree of the prediction result and the true label.

[0116] As can be seen from the table, on the two data sets, the effect is improved by 3.6% and 5% respectively using the text information made by the tool and not using the text information made by the tool, proving the effectiveness of the text information obtained by the application.

[0117] Embodiment two

[0118] The embodiment discloses a medical image text generation system for a visual language model;

[0119] As Figure 3 shown, a medical image text generation system for a visual language model comprises:

[0120] The labeled image acquisition module is configured to acquire a labeled image of a medical image.

[0121] The connected domain analysis module is configured to perform connected domain analysis on the labeled image, detect and count the number of tumor regions in the labeled image.

[0122] The morphological feature extraction module is configured to extract morphological features including size, shape, position, hollow feature and edge shape for each tumor region.

[0123] The text description output module is configured to convert the extracted morphological features into structured natural language description according to a preset text template, and output the final medical image text description.

[0124] Embodiment three

[0125] The embodiment aims to provide a computer-readable storage medium.

[0126] A computer-readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the steps of the medical image text generation method based on a visual language model according to the embodiment 1.

[0127] Embodiment four

[0128] The embodiment aims to provide an electronic device.

[0129] An electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor implements the steps of the medical image text generation method based on a visual language model according to the embodiment 1 when executing the program.

[0130] The steps and methods involved in the devices of the above embodiments two, three and four correspond to the embodiment one, and the specific embodiments can be referred to the relevant description part of the embodiment one. The term "computer-readable storage medium" should be understood as including a single medium or multiple media of one or more instruction sets; and should also be understood as including any medium capable of storing, encoding or carrying instruction sets for execution by a processor and causing the processor to perform any method of the present application.

[0131] Those skilled in the art should understand that each module or step of the present application described above can be realized by a general computer device, alternatively, they can be realized by program codes executable by a computing device, so that they can be stored in a storage device for execution by a computing device, or they can be respectively manufactured into each integrated circuit module, or a plurality of modules or steps among them can be manufactured into a single integrated circuit module to realize. The present application is not limited to any specific combination of hardware and software.

[0132] Although the specific embodiments of the present application are described above in combination with the accompanying drawings, it is not a limitation on the protection scope of the present application, and those skilled in the art should understand that various modifications or changes made by those skilled in the art on the basis of the technical solutions of the present application without creative labor are still within the protection scope of the present application.

Claims

1. A medical image text generation method for a visual language model, characterized in that, The method comprises the following steps: obtaining a labeled image of a medical image; the labeled image is a binary image, wherein a first pixel value of the binary image represents a target tumor region in the binary image; a second pixel value represents a background region of the binary image; performing connected component analysis on the labeled image to detect and count the number of tumor regions in the labeled image; for each tumor region, extracting morphological features including size, shape, position, cavity feature and edge shape; According to a preset text template, the extracted morphological features are converted into a structured natural language description, and a final medical image text description is output. 2.The medical image text generation method for visual language model of claim 1, wherein, The connected component analysis on the labeled image detects and counts the number of tumor regions in the labeled image, which comprises: obtaining a set of contour boundary points of the target tumor region and the background region in the labeled image; C i = {(x, y) e Ω | there exists a path connecting (x0, y0) and (x, y) such that I(x k ,y k ) = 255} where C i is the i-th connected region, and i is the number of tumor regions in the single slice to be extracted. 3.The medical image text generation method of claim 1, wherein, The process of extracting the size feature of the tumor region is: calculate the pixel area of the tumor region; According to a preset area threshold, the size of the tumor region is classified. 4.The medical image text generation method of claim 1, wherein, The process of extracting the shape feature of the tumor region is:

5. The medical image text generation method for visual language model according to claim 1, wherein, obtaining the minimum circumscribed rectangle of the tumor region, and judging the shape of the tumor region according to the aspect ratio of the minimum circumscribed rectangle. The process of extracting the position feature of the tumor region is:

6. The medical image text generation method for visual language model according to claim 1, wherein, calculate the center point coordinates of the tumor region, and divide the position region of the labeled image; according to the relationship between the center point coordinates and the divided position region, the position feature of the tumor region is determined. The process of extracting the cavity feature of the tumor region is:

7. The medical image text generation method oriented to a visual language model according to claim 1, wherein, use hierarchical contour detection algorithm to identify the sub-contour in the target region, and when there is a sub-contour, it is determined that the target region contains a cavity. The process of extracting the edge shape of the tumor region is: obtaining the convex hull of the tumor region; the convex hull is the smallest convex polygon containing the contour point set of the tumor region; detect the concave region between the original contour and the convex hull; 8.A medical image text generation system oriented to a visual language model, characterized in that: when the ratio of the concave depth to the size of the target region exceeds a preset threshold, it is determined that the edge of the target region has a concave. The method comprises the following steps: a labeled image acquisition module configured to obtain a labeled image of a medical image; a connected component analysis module configured to perform connected component analysis on the labeled image to detect and count the number of tumor regions in the labeled image; a morphological feature extraction module configured to extract morphological features including size, shape, position, cavity feature and edge shape for each tumor region; 9. A computer-readable storage medium having stored thereon a program, characterized in that, a text description output module configured to convert the extracted morphological features into a structured natural language description according to a preset text template, and output a final medical image text description.

10. An electronic device comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, The program is executed by the processor to realize the steps of the medical image text generation method based on the visual language model in any one of claims 1-7. The processor executes the program to realize the steps of the medical image text generation method based on the visual language model in any one of claims 1-7.

Citation Information

Patent Citations

  • Deformation degree evaluation and screening method and device for document image

    CN112435218A

  • Facilitating artificial intelligence integration into systems using a distributed learning platform

    CN113366580A

  • Key structure reconstruction method and device based on three-dimensional image and computer equipment

    CN115115772A

  • Multi-modal medical examination report automatic generation method and system

    CN119092032A

Cited By

  • Cell image classification method based on morphological semantic guidance

    CN121937808A