Human tissue measurement method, recognition method, program, storage medium, and system
Through computer-implemented neural network technology, the automatic centralized identification and measurement of human tissue from medical images has solved the problem of relying on manual intervention and manual measurement in the prior art, and improved the accuracy and efficiency of measurement.
Patent Information
- Application Number
- CN202411832133.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-12-12
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art relies on the expertise of the radiologist when interpreting CT scan images, and the process is manual, repetitive, and cumbersome, resulting in large measurement errors and variability.
Using a computer-implemented method, a trained neural network is used to output human tissue fragments from a medical image set, calculate bounding boxes, determine bounding boxes intersections, and merge fragments by generating bounding boxes to measure their size.
Improves the accuracy and efficiency of medical image measurement, reduces user bias, ensures anatomically correct representation, and reduces the cumbersomeness of manual measurement.
Smart Images

Figure CN120147217A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer programs and systems, and more particularly to methods, systems, and programs for measuring human tissue from a set of medical images. Background Art
[0002] Medical imaging has become increasingly important in clinical trials. In oncology, 95% of studies use medical imaging to evaluate treatment efficacy measured by progression-free survival (e.g., time to disease progression) or objective response rate (ORR).
[0003] There are many systems and programs for analyzing medical images such as CT scans. For example, 3DS MEDIDATA RAVE Imaging provides an IT infrastructure for collecting imaging data for centralized image review based on international best practice recommendations: the RECIST (Response Evaluation Criteria in Solid Tumors) criteria. The document "Eisenhauer, E.A., et.al, New response evaluation criteria in solid tumours: revised RECIST guideline (version 1.1). European journal of cancer, 45(2), 228 - 247, (2009)" is a revised guideline for the RECIST criteria, which defines the standard methods for measuring solid tumors and the definition of objective assessment of changes in tumor size adopted in adult and pediatric cancer clinical trials. 3DS MEDIDATA RAVE Imaging relies on the expertise of radiologists. The process of measuring solid tumors is manual, repetitive, cumbersome, and associated with measurement errors and variability among experts. Figure 1 An example of the annotation 1010 - 1040 of the identified lesions and 50% consistency 1050 using the RECIST process is shown. The annotations 1010 - 1040 have multiple differences due to variability among radiologists, who are responsible for manually modifying medical images to identify and measure target lesions in the human body and to identify new lesions. In practice, radiologists select one image from the image set based on their best practice, and they perform measurements of the tissue on that image.
[0004] A method for predicting 2D segmentation of tumors on CT scans has been developed. One of these methods is described in the following paper: Yan, K., et al, MULAN: Multitask Universal Lesion Analysis Network for Joint Lesion Detection, Tagging, and Segmentation, "MICCAI 2019, which relies on the maskrcnn model, which specifically processes CT scans to predict the 2D segmentation of tumors. All CT scan slices are annotated with all tumors detected on them. However, these measurements are of limited help to radiologists because he has to annotate individual slices of each lesion to comply with the RECIST criteria; in other words, radiologist intervention is still required to interpret the CT scans. In addition, for performance reasons, this method requires a dedicated model for each organ.
[0005] A method for detecting the 3D bounding boxes of relevant tumors is described in the related paper: Cai, J., Harrison, A. P., Zheng, Y., Yan, K., Huo, Y., Xiao, J.,... & Lu, L. (2020). Lesion-harvester: Iteratively mining unlabeled lesions and hard-negative examples at scale. IEEE Transactions on Medical Imaging, 40(1), 59-70, but it does not provide automatic segmentation or RECIST measurements. Again, radiologist intervention is still required to interpret the CT scans.
[0006] Therefore, the main problem with the above methods is that the accuracy of CT scan interpretation depends on the expertise of radiologists, and the process is also manual, repetitive, and cumbersome.
[0007] In this context, there is still a need for an improved method for measuring human tissue from a set of medical images representing human tissue. Summary of the Invention
[0008] Accordingly, a computer-implemented method for measuring human tissue is provided, where the human tissue is measured from a set of medical images representing the human tissue. The method includes: obtaining a trained neural network configured to output human tissue segments from the set of medical images. The method further includes: applying the trained neural network to the set of medical images. Thereby, the method identifies one or more segments of human tissue for at least two medical images in the set of medical images. The method further includes: calculating bounding boxes. Each bounding box encloses a segment. The method further includes: determining the intersection between a pair of bounding boxes. The method further includes: if the intersection of the pair of bounding boxes is non-empty, determining the intersection between the segments enclosed by the pair of bounding boxes. The method further includes: if the intersection between the segments is non-empty, merging the segments by calculating a generated bounding box enclosing the segments. The method further includes: measuring the size of the segments included in the generated bounding box.
[0009] The method further includes one or more of the following:
[0010] - Before determining the intersection between the segments, the method further includes:
[0011] o calculating a 2D bounding box for each identified segment from the at least two images in the set of medical images, each 2D bounding box enclosing the identified segment; and
[0012] o determining the intersection between at least two of the calculated 2D bounding boxes, and the determination of the intersection between the segments is only performed for segments where the intersection between the 2D bounding boxes is non-empty;
[0013] - The method further includes: iteratively determining the intersection on pairs of bounding boxes by repeating the determination and merging for each remaining pair of bounding boxes;
[0014] - The method further includes: if the intersection between the segments is empty, keeping the pair of bounding boxes enclosing each segment;
[0015] - The method further includes: obtaining the distance between the images including the segments enclosed by the pair of bounding boxes, and the determination of the intersection between the segments enclosed by the pair of bounding boxes is only performed when the distance between the images including the segments is below a predetermined threshold;
[0016] - The trained neural network is further configured to output labels identifying the human tissue segments, and the determination of the intersection on pairs of bounding boxes is performed for segments having the same label;
[0017] - Determining the intersection between the segments includes:
[0018] o Obtain a mask for each segment; and
[0019] o Calculate the intersection between the obtained masks, and perform the merging of the segments only for the segments for which the intersection between the obtained masks is non-empty.
[0020] - The trained neural network is further configured to output a 2D bounding box surrounding the human tissue segment, and the calculation of the 2D bounding box is performed by the trained neural network;
[0021] - Measuring the size of the segment further includes: selecting the segment included in the generated bounding box having the longest diameter;
[0022] - The medical image set includes a set of CT-SCAN images, a set of MRI images, PET scan images, or ultrasound images;
[0023] - The human tissue represented by the medical image set corresponds to a lesion in an organ or an aneurysm or the tissue of an organ.
[0024] There is also provided a method for identifying the evolution of human tissue. The method for identifying the evolution of human tissue includes: obtaining a current medical image set representing the human tissue of a patient. The method for identifying the evolution of human tissue further includes: using a method for obtaining current measurement results of the size of the segments. The method for identifying the evolution of human tissue further includes: obtaining historical measurement results obtained from a historical medical image set representing the human tissue of the patient. The method for identifying the evolution of human tissue further includes: calculating the difference between the current measurement results and the historical measurement results. The method for identifying the evolution of human tissue further includes: thereby identifying the evolution of the human tissue.
[0025] There is also provided a computer program including instructions that, when executed by a computer, cause the computer to perform the methods disclosed herein.
[0026] There is also provided a computer-readable storage medium having a computer program recorded thereon.
[0027] There is also provided a system including a processor coupled to a memory, the memory having a computer program recorded thereon. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Non-limiting examples will now be described with reference to the accompanying drawings, in which:
[0029] Figure 1 An example of the prior art is shown;
[0030] Figure 2 A flowchart of an example of the method is shown;
[0031] Figure 3 illustrates an example of the system; and
[0032] Figures 4 to 9 illustrates an example of the method. DETAILED DESCRIPTION
[0033] Referring Figure 2 to the flowchart of, a computer-implemented method for measuring human tissue from a set of medical images representing human tissue is presented. The method (also referred to as the "measurement method") includes obtaining S10 a trained neural network, the trained neural network being configured to output human tissue segments from images in the set of medical images. The method further includes applying S20 the trained neural network to the set of medical images. The method thereby identifies S30 one or more segments of human tissue for at least two medical images in the set of medical images. The method also calculates S40 the bounding boxes enclosing each of the identified segments. The method also determines S50 the intersection between pairs of the bounding boxes. If the intersection of a pair of bounding boxes is non-empty, the method determines S60 the intersection between the segments enclosed by the pair of bounding boxes. If the intersection between the segments is non-empty, the method merges S70 the segments by calculating a generated bounding box enclosing the segments. The method also measures S80 the size of the segments included in the generated bounding box.
[0034] This method improves the measurement of human tissue based on a set of medical images representing human tissue.
[0035] Notably, the method obtains an accurate measurement of human tissue due to the way the segments are determined. In fact, the method accurately determines an anatomically correct representation of human tissue because the segments included in the generated bounding box are ensured to belong to the same human tissue. For example, a human tissue may appear as a single segment in one image of the set of images, while another image may show multiple segments of the tissue. By construction, the generated bounding box includes segments from two images in the set of images representing the same tissue. Thus, the method eliminates user bias in determining the appropriate segments for performing the measurement of human tissue.
[0036] It should be noted that there seems to be an error in the original text in step S30 and subsequent steps in the flowchart description in ID=13. The correct step numbers in the description of the method should be S10, S20, S30, S40, S50, S60, S70, S80 as corrected in the translation.As described above, the method accurately determines an anatomically correct representation of human tissue due to the calculation of generating bounding boxes. This is all due to the specific steps performed by the method. First, the method utilizes the paradigm of a neural network to identify one or more segments of human tissue in at least two images in an image set. The method calculates the bounding boxes enclosing each segment and determines the intersection between pairs of bounding boxes. Thus, the method performs a rough detection of segments that may belong to the same human tissue segment (including those in the bounding box pairs). This rough detection is computationally efficient and particularly fast. If the intersection of the bounding box pair is non-empty, the method determines the intersection between the segments enclosed by the bounding box pair. In other words, the method performs a fine-grained determination of segments belonging to the same human tissue. This fine-grained determination enables the method to avoid false positives that may have been retained in the rough detection, e.g., bounding boxes that intersect but do not have intersecting segments. Thus, the segments merged by calculating the generating bounding boxes include segments belonging to the same human tissue segment.
[0037] A method for identifying (also referred to as "identification method") the evolution of human tissue is also provided. The identification method includes obtaining a current medical image set representing a patient's human tissue. Obtaining the current medical image set may include retrieving such an image set from a non-volatile storage device or retrieving the image set from a network.
[0038] The identification method further includes using the measurement method to obtain a current measurement of the size of the segment. In other words, the measurement method takes the current medical image set as input and outputs a measurement of the size of the segment included in the generated bounding box. The current medical image set is obtained at any time.
[0039] The identification method further includes obtaining historical measurement results obtained from a historical medical image set representing a patient's human tissue. In other words, the measurement method takes the historical medical image set as input and outputs a measurement of the size of the segment included in the generated bounding box. The historical medical image set represents the same human tissue as the current measurement. The historical medical image set was taken at a time before the time of obtaining the current medical image set (thus, the times of obtaining the historical medical image set and the current medical image set are different, e.g., several hours or days or even months).
[0040] The identification method further includes calculating the difference between the current measurement result and the historical measurement result. The difference can be a simple subtraction. Thus, the identification method identifies the evolution of human tissue. In fact, the identification indicates the evolution of human tissue according to the change in the size of the segment from which the measurement result is determined (as represented by the difference).
[0041] The method disclosed herein is computer-implemented. This means that the steps (or substantially all steps) of the method are performed by at least one computer or any system. Thus, the steps of the method are performed by a computer, possibly fully automatically or semi-automatically. In an example, the triggering of at least some steps of the method can be performed through user-computer interaction. The level of user-computer interaction required can depend on the level of automation foreseen and be balanced with the need to fulfill the user's wishes. In an example, this level can be user-defined and / or pre-defined.
[0042] A typical example of the computer implementation of a method is to use a system suitable for this purpose to perform the method. The system includes a processor coupled to a memory and may include a graphical user interface (GUI). A computer program including instructions for performing the method is recorded on the memory. The memory may also store a database. The memory is any hardware suitable for such storage and may include several physically distinct parts (e.g., one for the program and possibly one for the database).
[0043] Figure 3 An example of the system is shown, where the system is a client computer system, such as a user's workstation.
[0044] The client computer of the example includes a central processing unit (CPU) 3010 connected to an internal communication bus 3000, and a random access memory (RAM) 3070 also connected to the bus. The client computer is also provided with a graphics processing unit (GPU) 3110, which is associated with a video random access memory 3100 connected to the bus. The video RAM 3100 is also known as a frame buffer in the art. A mass storage device controller 3020 manages access to a mass storage device such as a hard disk drive 3030. Mass memory devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including for example semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks. Any of the foregoing devices can be supplemented or incorporated by a specially designed ASIC (application specific integrated circuit). A network adapter 3050 manages access to a network 3060. The client computer may also include a haptic device 3090, such as a cursor control device, a keyboard, etc. A cursor control device is used in the client computer to allow a user to selectively position a cursor at any desired location on a display 3080. In addition, the cursor control device allows the user to select various commands and input control signals. The cursor control device includes a plurality of signal generation devices for inputting control signals to the system. Generally, the cursor control device can be a mouse, and the buttons of the mouse are used to generate signals. Alternatively or additionally, the client computer system may include a touchpad and / or a touch screen.
[0045] The computer program may include instructions executable by a computer, the instructions including means for causing the above system to perform the method. The program may be recorded on any data storage medium, including the memory of the system. The program may be implemented, for example, in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations thereof. The program may be implemented as an apparatus, such as a product tangibly embodied in a machine-readable storage device for execution by a programmable processor. Method steps may be performed by a programmable processor executing an instruction program to perform the functions of the method by operating on input data and generating output. Thus, the processor may be programmable and coupled to receive data and instructions from a data storage system, at least one input device, and at least one output device, and to send data and instructions to the data storage system, at least one input device, and at least one output device. The application program may be implemented in a high-level procedural or object-oriented programming language, or if desired, in assembly or machine language. In any case, the language may be a compiled or interpreted language. The program may be a full installation program or an update program. The application of the program on the system results in any case of instructions for performing the method. The computer program may alternatively be stored and executed on a server in a cloud computing environment, the server communicating with one or more clients via a network. In this case, the processing unit executes the instructions included by the program, thereby causing the method to be performed on the cloud computing environment.
[0046] A medical image set may be formed by a data structure including (e.g., by being stored in a physical memory such as a hard disk drive) a plurality (e.g., two or more) of medical images.
[0047] In an example, the medical image set may include a set of CT-SCAN images, or a set of MRI images, or a set of PET scan images or a set of ultrasound images. A CT scan (Computed Tomography) is a specific type of medical image generated by sending X-rays through the human body. Signals are read and analyzed to reconstruct a dense volume of the body. In fact, the volume is divided into slices with their normal direction from feet to head. MRI (Magnetic Resonance Imaging) is a specific type of medical imaging technique that uses magnetic fields and computer-generated radio waves to create detailed images of organs and tissues inside the body. A PET scan (Positron Emission Tomography) is an imaging test that can help reveal the metabolic or biochemical function of your tissues and organs. A PET scan uses a radioactive drug called a tracer to show typical and atypical metabolic activity. Ultrasound imaging (ultrasonography) uses high-frequency sound waves and generates an image based on the reflection of the waves off body structures. The intensity (amplitude) of the sound signal and the time it takes for the wave to pass through the body provide the information needed to generate the image.
[0048] A medical image set represents human tissue. As is known in the art, tissue is a group of cells that are closely approximated and are organized to perform one or more specific functions. Thus, human tissue can correspond to a lesion in an organ (such as the liver, heart, or bone), or can correspond to tissue associated with an aneurysm, or can be an organ (e.g., a part or the whole of an organ). A lesion can be an abnormality in the human body without assuming pathogenicity in the human tissue. In an example, in the context of oncology, the term lesion can be used to refer to a tumor.
[0049] A medical image set can be images sorted according to a reference axis. In other words, each image can include a representation of human tissue viewed from a view (e.g., orthogonal to the reference axis) at a position on the reference axis. The reference axis can be set by convention, for example, along the standard z-axis of the human part that includes the human tissue. In this case, each view can include a view on the x-y plane (i.e., an axial or transverse view of the human tissue). The images in the medical image set can include position information of the images relative to the reference axis, i.e., the position on the reference axis along which the representation of the human tissue is observed. The images in the medical image set can be sorted according to such information. In an example, the images in the medical image set can be uniformly distributed, i.e., the images uniformly sample a part of the human body along the reference axis (with the same interval between two consecutive images in the ordered set). For example, the images can be separated by 1 mm or more along the reference axis.
[0050] The method obtains the S10 trained neural network. This means that the computerized system that executes the method can directly access the neural network (e.g., the neural network runs on the computerized system or is executed by the computerized system) or indirectly access the neural network (e.g., the computerized system can provide an input to the neural network and obtain the output of the neural network for the input). The trained neural network is configured to output human tissue segments from the images in the set. In other words, the trained neural network performs segmentation of the objects in the images. As is known in the art, image segmentation (also known as object segmentation) groups the pixels belonging to the same object by classifying each pixel value in the image into a specific category.
[0051] A neural network can be a function that includes a set of connected nodes, also known as "neurons". Each neuron receives inputs and outputs the results to other neurons connected to it. The neurons and the connections linking each of them have weight values, and the weight values are adjusted via training. In the context of the present disclosure, by a "trained neural network", it means that the weights of the neural network are explicitly saved after performing learning / training on a dataset. The dataset can include a set of medical images, and each set of medical images includes one or more annotated segments of human tissue. The dataset can include medical images from the DeepLession dataset, for example, as set forth in the following literature: Yan K, Wang X, Lu L, Summers RM, DeepLesion: automated mining of large-scale lesion annotations and universal lesion detection with deep learning, J Med Imaging (Bellingham), 2018.
[0052] The trained neural network can be a deep convolutional neural network, such as the Mask R-CNN neural network, as set forth in the following literature: Kaiming He, Georgia Gkioxari, Piotr Dollár, Ross Girshick Mask R-CNN, arXiv:1703.06870.
[0053] The method applies S20 the trained neural network to the set of medical images. In other words, the medical images in the set of medical images are provided as inputs to the trained neural network. Thus, the trained neural network identifies one or more segments of human tissue for at least two images in the set of medical images; the trained neural network outputs one or more identified segments of human tissue. In other words, the method processes the set of medical images as inputs to the trained neural network and outputs one or more segments of human tissue according to the weight values of the trained neural network. As is known from the field of machine learning itself, the processing of the input by the neural network includes applying operations to the input, and the operations are defined by data including weight values.
[0054] A segment can be any set of pixels corresponding to a region of the corresponding medical image (i.e., the medical image to which the trained neural network is applied). The pixels of the segment can include colors indicating the presence of human tissue. Thus, the segment can be referred to as a 2D segment because the trained neural network performs 2D segmentation.
[0055] The method calculates the S30 bounding box, which encloses each segment that has been recognized by the trained neural network. A "bounding box" refers to any geometric shape in any dimension (such as two or more dimensions, e.g., three dimensions), such as a two-dimensional rectangle or a three-dimensional cuboid. In the case of three dimensions, the bounding box can be referred to as a bounding volume. The bounding box encloses each segment. In other words, the bounding box includes (or encloses) the pixels of the segment that indicate the presence of human tissue. The method can consider the spatial position of the segments. For example, the method can convert the segments into 3D segments, where the 3D position of the segment can be the position relative to the x-y plane of its corresponding medical image and the z-axis defined by the reference axis. The method can determine the size of the bounding box based on the difference between the positions along the reference axis of the image pairs, or in the case of uniformly distributed image pairs, based on the difference between two consecutive image pairs.
[0056] The method determines the S410 intersection between the paired bounding boxes. The determination of the intersection can include calculating the intersection of at least one point (or line or surface) that belongs to the corresponding bounding boxes in the bounding box pair. The method can hold the position of each bounding box according to the position of the segments in the corresponding image and relative to the reference axis for determining the intersection.
[0057] The intersection of the bounding box pair can be empty or non-empty. The intersection is empty when there is no geometric intersection or crossing / overlap between two elements (e.g., points, lines, or surfaces), where each element corresponds to the corresponding bounding box in the bounding box pair. If there is a geometric intersection or crossing / overlap between two elements of each corresponding bounding box in the bounding box pair, the intersection is non-empty.
[0058] If the intersection S410 of the bounding box pair is non-empty, the method determines S420 the intersection between the segments enclosed by the bounding box pair. The determination of the intersection between the segments can include determining the intersection of at least one pair of points (or lines or surfaces) that belong to the corresponding segments enclosed by the bounding boxes in the bounding box pair.
[0059] If the intersection S420 between the segments is non-empty, the method merges S430 the segments by calculating the generated bounding box that encloses the segments. In other words, the method creates a new bounding box that encloses the segments. The method can discard the bounding box pair that encloses the segments when creating the generated bounding box.
[0060] In an example, the method can further include: if the intersection between the segments is empty, keeping the bounding box pair that encloses each segment. In other words, when the intersection is empty, the segments are not merged. Thus, the method ensures that the segments are not merged unnecessarily, for example, when the segments do not correspond to the same human tissue or when the segments are too far apart.
[0061] The method measures the size of the segments included in the S50 result bounding box. The method can measure the diameter or the cross-section of each segment.
[0062] In an example, measuring the size of the segments may further include selecting the segment included in the generated bounding box having the longest diameter (i.e., the longest distance). The method can determine the longest diameter by calculating all the distances between the lines of the pairs of points connecting the contours of the segment and keeping the line having the maximum distance as the longest diameter. The method can also determine the short diameter. The short diameter (or short distance) can be determined from the line connecting two points of the contour and orthogonal to the line corresponding to the longest diameter. The short diameter may correspond to the line having the longest distance among the lines orthogonal to the line corresponding to the longest diameter.
[0063] The measurement can be based on such a longest diameter. For example, the measurement of S50 can consist of the length of such a longest diameter. This is the case where the segment represents human tissue such as a non-lymph node lesion.
[0064] Alternatively, in the case where the segment represents human tissue such as a lymph node, the measurement result consists of the largest segment in the segment, which is orthogonal to such a longest diameter. In other words, the short diameter.
[0065] Thus, different types of measurements can be provided, and the choice of the measurement type can depend on the type of human tissue represented on the image, such as the RECIST diameter of a tumor lesion, the diameter of an aneurysm, or the measurement of the size of an organ. Thus, measuring human tissue is equivalent to obtaining at least one distance between two points on the representation of the human tissue. The distance can be the Euclidian distance.
[0066] Since the method applies a trained neural network to a medical image set, one or more segments of human tissue are obtained in a particularly fast and accurate manner.
[0067] The medical image set representing human tissue can be regarded as a sampling of the human tissue along a reference axis. Due to the spatial position of the bounding box, the calculation of the bounding box allows determining which segments belong to the same body tissue. In other words, the bounding box provides 3D sensing to one or more segments in order to allow determining which segments belong to the same body tissue based on the distribution of the image along the reference axis.
[0068] To identify which segments belong to the same body tissue, the method determines the intersection between pairs of bounding boxes. Thus, the method detects as early as possible when the segments do not intersect. The intersection between pairs of bounding boxes can be regarded as a rough approximation. The method moves to a fine-grained calculation: after determining the intersection between pairs of bounding boxes, when it is determined that such an intersection is non-empty, the method then determines the intersection between the pairs of segments that discard the erroneously associated segments.
[0069] When the method merges two segments by computationally generating a bounding box, the determination of the human tissue is refined for its measurement. The generated bounding box encloses the intersecting segments, so the method can measure the size of segments belonging to the same body tissue. This indeed improves the accuracy of the measurement.
[0070] The method may further include: calculating a 2D bounding box for each segment identified for each of at least two images from the medical image set before determining the intersection between the segments. Each 2D bounding box encloses the identified segment. The 2D bounding box can be a 2D parallelepiped (e.g., square, rectangle) with the minimum area to enclose the identified segment. The method may preserve the 2D position of the segment to locate the corresponding 2D bounding box, e.g., a position such as the center of the 2D bounding box.
[0071] The trained neural network may also be configured to output a 2D bounding box enclosing a segment of human tissue. Calculating the 2D bounding box may be performed by the trained neural network. In other words, the trained neural network may be applied to the medical image set. The trained neural network thus outputs the corresponding 2D bounding box for each corresponding segment. For example, the trained neural network may be a deep convolutional neural network, such as the Mask R-CNN neural network that outputs segments and the bounding box of each segment.
[0072] The method may further include determining the intersection between at least two of the calculated 2D bounding boxes. In other words, the method determines the intersection in the 2D coordinates of the calculated 2D bounding boxes. For example, determining the intersection between 2D bounding boxes may include determining the intersection of at least a pair of points (or lines) belonging to the corresponding 2D bounding boxes.
[0073] The determination of the intersection between segments may be performed only for segments where the intersection between the 2D bounding boxes is not empty. That is, when the intersection between the 2D bounding boxes is empty, the method does not pursue the determination of the intersection between the segments. Compared with the intersection detection between segments, this allows for a relatively fast and more accurate intersection detection than using 3D bounding boxes.
[0074] The method may further include iteratively determining the intersection S410 of the S40 bounding box pairs by repeating the determination of S410, S420, and merging S430 for each remaining pair of bounding boxes. That is, the method may repeatedly determine the intersection S410, followed by determining the intersection between the segments enclosed by the segment pairs S420, and then merging S430 of the segments (if the intersection between the segments is non-empty). When the method calculates the bounding box of each segment, the method may continue to perform the determination S410 between other pairs of bounding boxes in another iteration, after merging the segments S430, or after determining that the intersection of the pair of bounding boxes is empty, or when determining that the intersection between the segments is empty.
[0075] The method may perform the measurement S50 after the iterative determination S40 is completed, for example, when it is determined that there are no longer any intersecting bounding box pairs with non-empty intersections.
[0076] The method may further include obtaining the distance between images including the segments enclosed by the pairs of bounding boxes. The distance between the images may be obtained from the order of the images according to a reference axis. In other words, the distance may correspond to the difference in the positions of each image relative to the reference axis. The determination S420 of the intersection between the segments enclosed by the pairs of bounding boxes may be performed only when the distance between the images including the segments is below a predetermined threshold.
[0077] The trained neural network may also be configured to output a label identifying the human tissue segment. The label may be any piece of information indicating the human tissue. The determination of the intersection of the pairs of bounding boxes is performed for segments having the same label. In other words, the determination is performed for segments representing the same body tissue captured by the set of images. This ensures that the measurement of the human tissue is anatomically correct.
[0078] Determining the intersection between the segments may include obtaining a mask for each segment. The mask may be a binary image composed of zero and non-zero values. The pixels corresponding to the segment may have a predetermined value, such as 1, while other pixels (e.g., pixels corresponding to other parts of the human body) have a zero value. Determining the intersection may also include calculating the intersection between the obtained masks. The calculation of the intersection may include determining the pixel-by-pixel intersection. For example, if at least one pixel at a given position of a given mask (at the x-y position of the mask) has the same value as the pixel at the same position of the given mask, it may be determined that the two binary masks intersect. The merging of the segments may be performed only for segments for which the intersection between the obtained masks is non-empty. In other words, if the intersection is empty, the bounding box may be maintained. This results in a more refined determination of the intersection, and thus more accurate.
[0079] Now refer to Figures 4 to 9 Discuss the example.
[0080] Figure 4An example of a pipeline for implementing the method is shown.
[0081] The pipeline 4100 may include obtaining a set of medical images by reading, for example, DICOM CT scan images or NIFTI images taken from a patient (e.g., during an examination) 4110. DICOM refers to a file format dedicated to storing medical images. The DICOM format may include several modalities (e.g.: CT, RMI, US), and some of such modalities may include text reports or descriptions of geometric objects. The method may be illustrated by steps 4120 to 4140, where step 4120 corresponds to the application S20 of the trained neural network to the set of medical images, and steps 4130 and 4140 refer to steps S30 to S50 of the method. The measurement results may be output in DICOM format 4150.
[0082] Interestingly, method steps S10 to S50 may be performed by a computerized system that interfaces between a device that produces images (e.g., a CT scanner) and a device that takes the images as input to display them to a radiologist (e.g., a viewer). No adaptation of the two devices is required. In other words, the computer system that performs method steps S10 to S50 may be conceived as a software plug-in connected at the output of the device that produces images and at the input of the display device.
[0083] An example set of medical images 4200 includes five medical images 4210 to 4250. Each medical image provides a 2D representation of a part of the human body, for example, in a transverse view or an axial view. The medical images 4210 - 4250 sample body tissues 4260 (shown as a group for illustration only). The images in the set are sorted according to a standard z-axis (not shown). The medical image 4220 shows a case where two body tissues 4270 and 4280 that actually belong to the same body tissue 4260 are represented. Determining the intersection between bounding box pairs (e.g., the bounding boxes respectively surrounding the segments on images 4220 and 4230), determining S420 and merging S440 result in a bounding box surrounding the segments of the images (e.g., the segments of images 4220 and 4230) that allows determining that the two body tissues 4270 and 4280 belong to the same body tissue 4260.
[0084] Now an example of a method including iteratively determining the intersection on bounding box pairs is discussed.
[0085] The method converts the segmentation into individual lesions, and it should be understood that the method is applicable to any human tissue. Then, the method iterates over all pairs of lesions to see if they intersect. If so, they are merged together; otherwise, other lesions are processed. After the method has completely finished comparing all lesions, the method verifies that the newly merged lesions do not have new intersections. Thus, the method loops back to all lesions. When no more merges occur, this aggregation process ends. At this point, the method outputs the lesions.
[0086] Reference Figure 5 , which shows an example of the iterative determination S40.
[0087] The iterative determination S40 begins at 5000 and selects the obtained segments, for example, as the 2D segmentation segments 5010 recognized as a result of the application S20 of a trained neural network to a medical image set. The method can transform 5020 the 2D segmentation segments into 3D segmentation segments (also referred to as "3D lesions") to give the segments 3D positions 5030. Here, the term "lesion" refers to a medical image used to measure the potential increase or decrease of a lesion. However, any type of human tissue can be used. The transformation of 2D segments to 3D segments is intended to provide an artificial thickness to the 2D segments. In an example, the thickness can be obtained by calculating the bounding box enclosing the 2D segment, and the bounding box has a thickness that is at least equal to the distance between two consecutive images in the image set; it should be understood that this distance is the same for each pair of consecutive images in the medical image set. The thickness of the bounding box can be equal to the space between two consecutive images in the image set. For example, if the resolution of the imaging device that generates the medical image set is 1 mm (i.e., the distance between two adjacent images - also referred to as slices - in the Z-axis in the medical image set), then the thickness of the bounding box used to obtain the 3D lesion is 1 mm. Interestingly, calculating (obtaining) the 3D bounding box is cost-effective in terms of computing resources and memory. Experiments have shown that using the resolution of the imaging device as the thickness of the bounding box provides the best results.
[0088] The iteration can continue with the step of generating all pairs 5040 of 3D segments (or 3D lesions), which defines the segments that can be obtained from the application of the trained neural network or merged after the iteration.
[0089] The method can retrieve a pair of segments (or "lesion pair" l1, l2) 5050 at the first iteration (or the next lesion pair in subsequent iterations). The method can determine 5060 whether the segments intersect by performing a determination of the intersection S410, and subsequently determining the intersection S420 between the pair of segments. If the segments intersect, the method can proceed 5070 to merge S430 segments l1 and l2 by calculating to generate a bounding box. If the segments do not intersect at S420, the method maintains the pair of bounding boxes enclosing each segment. The method can determine 5080 whether there is a non-empty intersection between all pairs of bounding boxes from which the bounding boxes are generated to determine whether the segments intersect. If there is 5090 a non-empty intersection, the method can repeat step 5040 and the following steps. If there is no non-empty intersection, the method determines that the merging between all segments has not occurred 5090, and the method terminates 5100 iteration S40.
[0090] Now discuss an example of detecting the intersection between segments.
[0091] Refer to Figure 5 , which shows an example of determining the intersection between a pair of bounding boxes, subsequently determining the intersection between 2D bounding boxes, and then determining the intersection between the segments enclosed by the pair of bounding boxes.
[0092] Figure 6 The example of Figure 5 discussed starts with two 3D segments (also referred to as the "3D lesions" Figure 5 discussed) l1 and l2 6000. The method determines the intersection between a pair of bounding boxes 6010. If the intersection 6020 is empty, the method determines that there is no lesion intersection 6040 and the process terminates. If the intersection 6020 of a pair of bounding boxes is non-empty, the method retrieves the segments included in the intersecting bounding boxes 6030. The method determines whether the segments included in the intersecting bounding boxes are separated from the reference axis (in this case the z-axis) by less than a predetermined threshold max_distance 6060, for example by comparing information about the intervals along the reference axis for the corresponding images. The max_distance can be selected according to the resolution of the scan that produced the image. The max_distance can be selected such that it is at most equal to the space between two consecutive images in the medical image set. The predetermined threshold max_distance 6060 can be selected by performing an exploration of hyperparameters. Alternatively, the predetermined threshold max_distance 6060 can be selected as a function of the expected size of the lesion, for example 15 mm or greater.
[0093] In the case of separation below a predetermined threshold, the method calculates a 2D bounding box for each identified segment of at least two images from an image set. The method determines an intersection 6070 between at least two of the calculated 2D bounding boxes, for example by determining whether the 2D bounding boxes overlap. If the intersection between the calculated 2D bounding boxes is empty, the method determines that there is no lesion intersection 6080 and the process terminates. In the case where the intersection between the calculated 2D bounding boxes is non-empty, the method determines an intersection between the segments 6090 enclosed by the bounding box pair, for example by obtaining the mask of each segment and determining whether the corresponding masks intersect. If the intersection between the masks of the segments is non-empty, the method merges the segments by calculating a generated bounding box that encloses the segments. When two segments are to be merged, a new segment is created to contain the segments from the two original lesions. The 3D bounding box is recalculated accordingly. Then the previous lesions are discarded.
[0094] The method determines whether all segments have been examined 6110. If not, the method returns to step 6050. Otherwise, the method terminates. It should be understood that, as is known in the art, the mask of the segment is calculated / obtained. For example, MaskR-CNN generates an accurate segmentation mask for each segment in the output.
[0095] Thus, the method detects as early as possible when the lesions do not intersect. As the method moves forward to calculate the intersection between the masks, the method goes from a rough approximation to a fine-grained calculation: First, the method looks at the 3D bounding box, then the method compares the 2D bounding boxes of each segment pair, and finally compares the 2D masks.
[0096] Figure 7 A pre-trained neural network is shown.
[0097] The method may use a pre-trained 2D deep convolutional neural network (Mask R-CNN) 7020, which has been trained to process CT scan images 7010.
[0098] The network takes the slice image 7011 and its 3D context as input and outputs zero or more detected 2D lesions 7030, for example from the liver 7031. Then, each detected 2D lesion is filtered out or stacked with other nearby detected lesions to form a 3D lesion. The method outputs the RECIST measurement of each detected 3D lesion.
[0099] Figure 8 An inference of the measurement of a RAW CT scan is shown. The method may produce 3D segments accompanied by measurements 820, or stack them together as 3D lesions 830.
[0100] Now discuss the calculation of RECIST measurements. Refer to Figure 9 , Figure 9 which shows the RECIST measurements calculated on segment 9000.
[0101] The method locates two specific diameters of the segment of interest: the long diameter 9010 and the short diameter 9020. Each segment must be measured at most once, so the method ensures that the selected segment for determining the measurement is the segment with the largest diameter. To this end, the method measures the diameters of all the slices that make up the segment and retains only the maximum value as the long diameter.
[0102] To calculate the diameters, the method establishes the following geometric approach: for the long diameter 9010, the method calculates all the distances between all pairs of points on the contour and retains the maximum value. For the short diameter 9020, the method passes through all the points on the contour of section 9030 and considers the line 9040 that passes through this point and is orthogonal to the long diameter 9010 just calculated. The method measures the segment located within the contour and retains the maximum value, which is line 9020.
Claims
1. A computer-implemented method for measuring human tissue from a set of medical images representing the human tissue, the method comprising: - obtaining (S10) a trained neural network, wherein the trained neural network is configured to output a human tissue segment according to a medical image in the medical image set; - applying (S20) the trained neural network to the medical image set, thereby identifying one or more human tissue segments from at least two images in the medical image set; - calculating (S30) a bounding box surrounding each fragment; - determining (S410) intersections between pairs of bounding boxes; - if the intersection of the pair of bounding boxes is non-empty, determining (S420) the intersection between the segments enclosed by the pair of bounding boxes; - if the intersection between the segments is non-empty, merging (S430) the segments by calculating a resulting bounding box enclosing the segments; and - measuring (S50) the size of the fragment included in said generated bounding box.
2. The method according to claim 1, before determining the intersection between the segments, further comprising: - for each segment identified from the at least two images in the medical image set, calculating a 2D bounding box of the segment, each 2D bounding box enclosing the identified segment; as well as - determining an intersection between at least two of the calculated 2D bounding boxes, the determining the intersection between the segments being performed only for segments for which the intersection between the 2D bounding boxes is non-empty.
3. The method according to claim 1 or 2, further comprising: The intersections (S410) on the bounding box pairs are iteratively determined (S40) by repeating the determining (S420) and merging (S430) for each remaining pair of bounding boxes.
4. The method according to any one of claims 1 to 3, further comprising: If the intersection between the segments is empty, the pair of bounding boxes is kept enclosing each segment.
5. The method according to any one of claims 1 to 4, further comprising: The distance between the images including the segments surrounded by the pair of bounding boxes is obtained, and the determining (S420) of the intersection between the segments surrounded by the pair of bounding boxes is performed only when the distance between the images including the segments is below a predetermined threshold.
6. The method according to any one of claims 1 to 5, wherein: The trained neural network is further configured to output a label for identifying a human tissue segment, and For segments having the same said label, a determination of intersection on pairs of bounding boxes is performed.
7. The method according to any one of claims 1 to 6, wherein: Determining the intersection between the segments includes: - obtain a mask for each fragment; and - Calculate the intersection between the obtained masks, and perform merging of fragments only for the fragments whose intersection between the obtained masks is non-empty.
8. The method according to any one of claims 2 to 7, wherein: The trained neural network is further configured to output a 2D bounding box surrounding the human tissue segment, and Computing the 2D bounding box is performed by the trained neural network.
9. The method according to any one of claims 1 to 8, wherein: Measuring the size of the fragments further comprises: The fragments included in the generated bounding box with the longest diameter are selected.
10. The method according to any one of claims 1 to 9, wherein: The medical image collection includes: A set of CT-SCAN images, A set of MRI images, PET scan images, or Ultrasound image.
11. The method according to any one of claims 1 to 10, wherein: The human tissue represented by the medical image set corresponds to a lesion in an organ, or to an aneurysm or tissue of an organ.
12. A method for identifying the evolution of human tissue, comprising: - obtaining a current set of medical images representing human tissue of a patient; - using the method of any one of claims 1 to 11 to obtain a current measurement result for the size of the fragment; - obtaining historical measurements obtained from a set of historical medical images representing human tissue of said patient; - calculating the difference between the current measurement and the historical measurements, thereby identifying the evolution of the human tissue.
13. A computer program comprising instructions which, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 11 and / or the method of claim 12.
14. A computer-readable storage medium having recorded thereon the computer program of claim 13.
15. A system comprising a processor coupled to a memory having recorded thereon the computer program of claim 13.