Medical image labeling device and medical image labeling method

By adopting a two-way reasoning mechanism in the medical image labeling device, the labeling of intermediate slice images is automatically generated using manual labeling data, which solves the problem of time-consuming and inconsistent manual labeling, and improves the labeling efficiency and accuracy.

CN120015248APending Publication Date: 2025-05-16HTC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411631539.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-16
Filing Date
2024-11-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

During the medical image labeling process, manually marking the area of ​​interest is time-consuming and inconsistent, which affects the efficiency of medical image analysis and diagnosis.

Method used

Design a medical image labeling device, including an interface, memory and processor, and automatically generates result labeling on intermediate slice images using limited manual labeling data through a two-way reasoning mechanism.

Benefits of technology

It improves the efficiency of semi-automatic labeling, enhances the accuracy of automatic labeling, and reduces the time and human error of manual labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015248A_ABST
    Figure CN120015248A_ABST
Patent Text Reader

Abstract

A medical image marking device comprises an interface, a memory and a processor. The memory is used for storing a plurality of serialized medical images. The processor is coupled to the interface and the memory. The processor is used for receiving a first manual label on a first slice image in the plurality of serialized medical images and a second manual label on a second slice image in the plurality of serialized medical images through the interface. The processor is configured to execute a bidirectional reasoning mechanism to generate a plurality of result annotation tags on a plurality of intermediate slice images of the plurality of serialized medical images according to the first manual annotation and the second manual annotation, respectively. The medical image labeling device is used for efficiently and accurately generating labels on serialized medical images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a medical imaging technology, and more specifically, to a method and device for annotating serialized medical images. Background Art

[0002] Medical imaging technologies, including X-rays, Magnetic Resonance Imaging (MRI) and Computed Tomography (CT), are vital tools in modern medical diagnosis. These technologies provide detailed images of the body's internal structures required for accurate diagnosis and treatment. However, the increase in the volume and complexity of image data poses a major challenge to medical professionals, especially in accurately identifying and annotating regions of interest (RoI).

[0003] Annotating regions of interest in large volumes of medical images is very time-consuming. In practice, accurate annotation of regions of interest usually relies on manual work by radiologists or radiology technicians. Manual processing of serialized images such as CT scans and MRI scans is particularly burdensome because it requires professionals to individually annotate each slice obtained in the scan. This process is not only time-consuming, but also leads to inconsistencies between annotations due to differences between annotators, which may affect the efficiency of medical image analysis and clinical diagnosis. Summary of the invention

[0004] One aspect of the disclosure document discloses a medical image annotation device, which includes an interface, a memory, and a processor. The memory is used to store a plurality of serialized medical images. The processor is coupled to the interface and the memory. The processor is used to receive a first manual annotation on a first slice image of the plurality of serialized medical images and a second manual annotation on a second slice image of the plurality of serialized medical images through the interface. The processor is used to execute a bidirectional reasoning mechanism to generate a plurality of result annotation labels on a plurality of intermediate slice images of the plurality of serialized medical images according to the first manual annotation and the second manual annotation.

[0005] In some embodiments, the bidirectional reasoning mechanism executed by the processor includes: executing a tracking model based on the first manual annotation on the first slice image to generate a plurality of forward prediction annotations in a forward order on the plurality of intermediate slice images of the plurality of serialized medical images; executing the tracking model based on the second manual annotation on the second slice image to generate a plurality of reverse prediction annotations in a reverse order on the plurality of intermediate slice images of the plurality of serialized medical images; and merging the plurality of forward prediction annotations and the plurality of reverse prediction annotations to generate the plurality of result annotation labels on the plurality of intermediate slice images.

[0006] In some embodiments, the tracking model includes a tracking any object model, which is used to detect, associate and track the first manual annotation on the first slice image between each of the multiple intermediate slice images in the forward order, and the tracking any object model is used to detect, associate and track the second manual annotation on the second slice image between each of the multiple intermediate slice images in the reverse order.

[0007] In some embodiments, corresponding to a Kth slice image among the multiple serialized medical images, a result annotation label on the Kth slice image is generated in the following manner: calculating a first distance between the first slice image and the Kth slice image and a second distance between the Kth slice image and the second slice image; when the first distance is shorter than the second distance, selecting a forward prediction annotation on the Kth slice image as the result annotation label on the Kth slice image; and when the second distance is shorter than the first distance, selecting a backward prediction annotation on the Kth slice image as the result annotation label on the Kth slice image.

[0008] In some embodiments, corresponding to a Kth slice image among the multiple serialized medical images, a result annotation label on the Kth slice image is generated in the following manner: calculating a first distance between the first slice image and the Kth slice image and a second distance between the Kth slice image and the second slice image; and generating the result annotation label on the Kth slice image based on the first distance and the second distance and in accordance with a weighted sum of a forward prediction annotation and a backward prediction annotation on the Kth slice image.

[0009] In some embodiments, the bidirectional reasoning mechanism executed by the processor includes: executing a bidirectional XMem model to generate the plurality of result annotation labels on the plurality of intermediate slice images according to the first manual annotation and the second manual annotation.

[0010] In some embodiments, the bidirectional XMem model includes a sensory memory, a working memory, and a long-term memory. The sensory memory is used to store a short-term forward hidden expression and a short-term reverse hidden expression. The working memory is used to store a forward memory key feature, a reverse memory key feature, a forward memory value feature, and a reverse memory value feature. The long-term memory is used to store a forward long-term memory key feature, a reverse long-term memory key feature, a forward long-term memory value feature, and a reverse long-term memory value feature.

[0011] In some embodiments, the bidirectional XMem model further includes a query encoder, a decoder, and a mask encoder. The query encoder is used to generate a forward input query about a forward input medical image and a reverse input query about a reverse input medical image. The decoder is used to generate a forward annotation mask about the forward input medical image and a reverse annotation mask about the reverse input medical image based on the forward input query, the reverse input query, a readout feature, the short-term forward hidden representation from the sensory memory, and the short-term reverse hidden representation. The mask encoder is used to generate the short-term forward hidden representation and the short-term reverse hidden representation based on the forward annotation mask and the reverse annotation mask.

[0012] In some embodiments, the readout feature is generated based on an association matrix and a readout value, the association matrix is ​​calculated by the processor based on the forward input query, the reverse input query, the forward memory key feature, the reverse memory key feature, the forward long-term memory key feature, and the reverse long-term memory key feature, and the readout value is calculated by the processor based on the forward memory value feature, the reverse memory value feature, the forward long-term memory value feature, and the reverse long-term memory value feature.

[0013] In some embodiments, the plurality of serialized medical images include a plurality of magnetic resonance imaging scan images or a plurality of computerized tomography scan images.

[0014] Another aspect of the present disclosure discloses a medical image annotation method, which includes: obtaining a plurality of serialized medical images; receiving a first manual annotation on a first slice image and a second manual annotation on a second slice image of the plurality of serialized medical images; and executing a bidirectional reasoning mechanism to generate a plurality of result annotation labels on a plurality of intermediate slice images of the plurality of serialized medical images according to the first manual annotation and the second manual annotation.

[0015] The medical image annotation device and method can effectively utilize a limited amount of user annotation data and automatically generate result annotation labels on the intermediate slice image based on bidirectional tracking. The medical image annotation device and method can improve the efficiency of semi-automatic annotation. Since the bidirectional reasoning mechanism can provide more reference information when tracking objects, the accuracy of automatic annotation when identifying unlabeled image data can be improved.

[0016] It should be noted that the above description and the following detailed description are illustrative of the present case in the form of embodiments, and are used to assist in the explanation and understanding of the invention content claimed in the present case. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to make the above and other objects, features and embodiments of the present disclosure more clearly understood, the accompanying drawings are described as follows:

[0018] Figure 1 is a schematic diagram of a medical image annotation device according to some embodiments of the present disclosure;

[0019] Figure 2 is a schematic diagram of a first implementation of a bidirectional reasoning mechanism executed by a processor according to the present disclosure;

[0020] Figure 3 is a schematic diagram illustrating how the tracking model generates forward prediction annotations;

[0021] Figure 4 is a schematic diagram illustrating how the tracking model generates the backward prediction annotation;

[0022] Figure 5 is a schematic diagram of a bidirectional XMem model executed by a processor according to some embodiments of the present disclosure; and

[0023] Figure 6 is a schematic diagram of a bidirectional XMem model executed by the processor based on another pair of slice images.

[0024] Explanation of symbols:

[0025] 100: Medical image annotation device

[0026] 120: Interface

[0027] 140: Processor

[0028] 142:Tracking Model

[0029] 144: Combiner

[0030] 146: Bidirectional XMem Model

[0031] 146a: Sensory Memory

[0032] 146b: Working memory

[0033] 146c: Long-term storage

[0034] 146d: Query encoder

[0035] 146e:Decoder

[0036] 146f: Mask encoder

[0037] 160: Memory

[0038] F RO : Read out features

[0039] h1,h2,h3: short-term forward hidden expression

[0040] h6,h7,h8: short-term reverse hidden expression

[0041] INF1: Forward information data

[0042] INF2: Reverse information data

[0043] Kc: Memory key

[0044] K LF : Forward long-term memory key features

[0045] K LB :Reverse long-term memory key features

[0046] K WF :Forward memory key feature

[0047] K WB :Reverse memory key feature

[0048] P2, P3, P4, P5, P6, P7: result labeling

[0049] P B2 ,P B3 ,P B4 ,P B5 ,P B6 ,P B7 : Backward Prediction Annotation

[0050] P F2 ,P F3 ,P F4 ,P F5 ,P F6 ,P F7 : Forward prediction annotation

[0051] P M1 ,P M8 : Manual annotation

[0052] QSL2 ,Q SL3 : Forward input query

[0053] Q SL6 ,Q SL7 :Reverse input query

[0054] SIMG: Serialized Medical Images

[0055] SL1,SL2,SL3,SL4: Slice images

[0056] SL5,SL6,SL7,SL8: Slice Image

[0057] Vc: Read value

[0058] V LB :Reverse long-term memory value feature

[0059] V LF : Forward long-term memory value characteristics

[0060] V WB :Reverse memory value feature

[0061] V WF : Forward memory value feature

[0062] W AM :Incidence Matrix DETAILED DESCRIPTION

[0063] The following disclosure provides many different embodiments or examples for implementing different features of the present disclosure. The elements and configurations in the specific examples are used to simplify the present disclosure in the following discussion. Any examples discussed are used for illustrative purposes only and do not limit the scope and significance of the present disclosure or its examples in any way. Where appropriate, the same reference numerals are used between the drawings and in the corresponding text descriptions to represent the same or similar elements.

[0064] See also Figure 1 , Figure 1 is a schematic diagram of a medical image annotation apparatus 100 according to some embodiments of the present disclosure. In some embodiments, the medical image annotation apparatus 100 is used to efficiently generate annotations on a plurality of serialized medical images SIMG.

[0065] In some embodiments, the serialized medical image SIMG includes a plurality of continuous magnetic resonance imaging (MRI) scan images or a plurality of continuous computed tomography (CT) scan images. Figure 1In the embodiment shown, the serialized medical image SIMG may be a series of abdominal magnetic resonance imaging / computerized tomography images. The serialized medical image SIMG may include a plurality of slice images captured in sequence. Figure 1 As shown, the serialized medical image SIMG includes multiple slice images SL1, SL2, SL3, SL4, SL5, SL6, SL7 and SL8. In some embodiments, these slice images SL1 to SL8 can be taken from the patient at successive time points by a nuclear magnetic resonance imaging / computer tomography scanner. For example, slice image SL1 can be taken first; slice image SL2 can be taken after slice image SL1; slice image SL3 can be taken after slice image SL2, and so on, and finally slice image SL8 can be taken. These slice images SL1 to SL8 can respectively represent continuous abdominal images located at different heights of the body.

[0066] SIMG for serialized medical images is not limited to Figure 1 The abdominal MRI / CT image shown. In other embodiments, the serialized medical image SIMG may be a series of brain MRI / CT images, a series of chest MRI / CT images, or other similar medical images. The medical image annotation device 100 may be used to process various types of serialized medical images SIMG, and may also be extended to medical films.

[0067] For the sake of simplicity, Figure 1 The serialized medical image SIMG shown contains a total of 8 slice images as an exemplary illustration. However, the number of serialized medical images SIMG in the present disclosure is not limited thereto. The serialized medical image SIMG may contain N slice images, where N is a positive integer greater than 3.

[0068] Since each serialized medical image SIMG is captured in sequence, adjacent slice images in the serialized medical image SIMG are related to each other. When an object (e.g., a body organ or tissue) appears in a slice image SL1 with a specific outline, the same object may also appear in a slice image SL2 with a similar outline. Therefore, if a manual annotation P about this object is provided on the slice image SL1, M1 , the manual annotation P on the slice image SL1 can be M1 is used as a hint to segment similar objects in subsequent slice images SL2 and SL3 according to this hint. Therefore, when another manual annotation P about this object is provided on slice image SL8, M8 When the manual annotation P on the slice image SL8 is M8 Used as a hint to segment similar objects in adjacent slice images SL7, SL6, etc.

[0069] The medical image annotation device 100 in the present disclosure aims to effectively utilize a limited amount of manually annotated data to improve the accuracy of the recognition algorithm in automatically annotating unannotated medical images. Figure 1 As shown, the medical image annotation device 100 includes an interface 120, a processor 140 and a memory 160. Figure 1 As shown, the memory 160 is used to store the serialized medical images SIMG (ie, slice images SL1 to SL8 ).

[0070] The interface 120 may include input-output elements (eg, a touch panel, a keyboard, a mouse, a microphone, a display). Figure 1 As shown, in some embodiments, the user can operate the interface 120 to input a manual annotation P on the first slice image (ie, slice image SL1) of the serialized medical image SIMG. M1 The user can operate the interface 120 to input another manual annotation P on the last slice image (ie, slice image SL8) of the serialized medical image SIMG. M8 .

[0071] like Figure 1 As shown, the processor 140 is coupled to the interface 120 and the memory 160. The processor 140 is used to receive the manual annotation P on the slice image SL1 in the serialized medical image SIMG through the interface 120. M1 , and receiving manual annotations P on slice images SL8 in serialized medical images SIMG M8 The processor 140 is used to execute a bidirectional reasoning mechanism to M1 and P M8 The result annotation labels P2, P3, P4, P5, P6 and P7 on the intermediate slice images (eg, slice images SL2, SL3, SL4, SL5, SL6 and SL7) of the serialized medical image SIMG are generated respectively. The detailed operation of the bidirectional reasoning mechanism will be further explained in the following embodiments.

[0072] In some embodiments, the two manual annotations are not limited to being marked on the first and last slice images SL1 and SL8. In other embodiments, the two manual annotations can also be marked on other slice image combinations, for example, on slice images SL2 and SL7, or on slice images SL3 and SL6.

[0073] Please also read Figure 2 , Figure 2 is a schematic diagram of a first embodiment of a bidirectional reasoning mechanism executed by processor 140 according to the present disclosure. Figure 2In the illustrated embodiment, the bidirectional reasoning mechanism executed by the processor 140 includes a tracking model 142 and a merger 144. The tracking model 142 is used to identify the manually labeled P on the slice image SL1. M1 , generate a forward prediction annotation P on the intermediate slice images (e.g., slice images SL2, SL3...SL7) of the serialized medical image SIMG F2 , P F3 , P F4 , P F5 , P F6 and P F7 At the same time, the tracking model 142 is used to manually mark the slice image SL8 according to the P M8 , generates a reverse prediction annotation P on the intermediate slice images (e.g., slice images SL7, SL6...SL2) of the serialized medical image SIMG B7 , P B6 , P B5 , P B4 , P B3 and P B2 .

[0074] Please also read Figure 3 , Figure 3 is to explain how the tracking model 142 generates the forward prediction annotation P F2 , P F3 , P F4 , P F5 , P F6 and P F7 In some embodiments, the tracking model 142 may be implemented by a Track-Anything-Model (TAM). The Track-Anything-Model (TAM) is used to detect, associate, and track the manual annotation P on the slice image SL1 between the intermediate slice images (e.g., slice images SL2, SL3, ... SL7) in a sequential order. M1 .

[0075] In some embodiments, the Tracking Any Object Model (TAM) is a computer vision algorithm for sequentially tracking or following an object or tag in a sequence of medical images SIMG. The Tracking Any Object Model (TAM) can be implemented by software program code, and can adopt a software architecture of a convolutional neural network (CNN), a recurrent neural network (RNN), or a transformer. The Tracking Any Object Model (TAM) is pre-trained on a large dataset to understand and predict the movement trajectory and appearance contours of various objects.

[0076] like Figure 3 As shown, the serialized medical images SIMG and the manually annotated P M1 The tracking model 142 (eg, any object tracking model) uses the manual annotation P on the slice image SL1 to track the model 142 (eg, any object tracking model). M1 As a hint, in order to generate a forward prediction annotation P about a similar object in the subsequent slice image SL2 F2 In this example, the manual annotation P M1 The liver region appearing in the slice image SL1 is marked so that the tracking model 142 can generate a forward prediction annotation P at the position of the potential liver region in the slice image SL2. F2 In a similar way, we manually annotate P M1 And the forward prediction label P F2 The hints may also be input to the tracking model 142 to generate forward prediction annotations P for similar objects in the subsequent slice image SL3. F3 In the same manner, the tracking model 142 can sequentially generate forward prediction annotations P on the intermediate slice images (eg, slice images SL2, SL3, SL4, SL5, SL6, and SL7) of the serialized medical image SIMG. F2 , P F3 , P F4 , P F5 , P F6 and P F7 .

[0077] Please also read Figure 4 , Figure 4 is to explain how the tracking model 142 generates the reverse prediction annotation P B7 , P B6 , P B5 , P B4 , P B3 and P B2 In some embodiments, the tracking model 142 may be implemented by a tracking any object model (TAM). The tracking any object model (TAM) is used to detect, associate and track the manual annotation P on the slice image SL1 between the intermediate slice images (e.g., slice images SL7, SL6, ... SL2) in reverse order. M1 .

[0078] like Figure 4 As shown, the serialized medical images SIMG and the manually annotated P M8 The tracking model 142 (eg, any object tracking model) is input to the tracking model 142 (eg, any object tracking model). The tracking model 142 (eg, any object tracking model) uses the manual annotations P on the slice image SL8. M8As a hint, a reverse prediction annotation P about similar objects is generated in the adjacent slice image SL7. B7 In this example, the manual annotation P M8 The liver region appearing in the slice image SL8 is marked so that the tracking model 142 can generate a reverse prediction annotation P in the potential liver region of the slice image SL7. B7 In a similar way, we manually annotate P M8 And reverse prediction annotation P B7 The hints may also be input to the tracking model 142 to generate reverse prediction annotations P for similar objects in the adjacent slice image SL6. B6 In the same manner, the tracking model 142 can sequentially generate reverse prediction annotations P on the intermediate slice images (eg, slice images SL7, SL6, SL5, SL4, SL3, and SL2) of the serialized medical image SIMG. B7 , P B6 , P B5 , P B4 , P B3 and P B2 .

[0079] like Figure 2 As shown, based on the above bidirectional tracking, the tracking model 142 will generate two predicted annotations for each intermediate slice image (e.g., slice images SL2 to SL7). M1 The generated forward prediction annotation P F2 And based on manual annotation P M8 The generated backward prediction annotation P B2 Regarding the slice image SL3, there is a forward prediction annotation P F3 And the reverse prediction annotation P B3 Regarding the slice image SL7, there is a forward prediction annotation P F7 And reverse prediction annotation P B7 .

[0080] The merger 144 is used to receive and merge the forward prediction annotation P related to the slice image SL2. F2 And reverse prediction annotation P B2 , and generates a result annotation label P2 on the slice image SL2. Similarly, the merger unit 144 is also used to receive and merge the forward prediction label P related to the slice image SL3. F3 And reverse prediction annotation P B3 , and generates a result annotation label P3 on the slice image SL3. Similarly, the merger 144 is also used to receive and merge the forward prediction label P about the slice image SL7. F7 And reverse prediction annotation P B7, and generate the result annotation label P7 on the slice image SL7.

[0081] In some embodiments, the result annotation labels P2 to P7 may be determined by the merger 144 according to the distance between the target slice image and the manually labeled image.

[0082] For example, regarding the Kth slice image of the serialized medical image SIMG, the result annotation label on the Kth slice image is generated by calculating the first distance between the slice image SL1 and the Kth slice image and the second distance between the Kth slice image and the slice image SL8. If the first distance is shorter than the second distance, the merger 144 selects the forward prediction annotation on the Kth slice image as the result annotation label on the Kth slice image. If the second distance is shorter than the first distance, the merger 144 selects the reverse prediction annotation on the Kth slice image as the result annotation label on the Kth slice image. Wherein K is a positive integer.

[0083] For example, regarding slice image SL2, the first distance between slice image SL1 and slice image SL2 is "1", and the second distance between slice image SL2 and slice image SL8 is "6". At this time, the first distance is shorter, in other words, slice image SL2 is closer to slice image SL1. In this case, the forward prediction label P will be selected. F2 As a result, the slice image SL2 is labeled with label P2.

[0084] Regarding the slice image SL3, the first distance between the slice image SL1 and the slice image SL3 is "2", and the second distance between the slice image SL3 and the slice image SL8 is "5". In this case, the forward prediction label P is selected. F3 As a result, a label P3 is annotated on the slice image SL3.

[0085] Regarding slice image SL7, the first distance between slice image SL1 and slice image SL7 is "6", and the second distance between slice image SL7 and slice image SL8 is "1". At this time, the second distance is shorter, in other words, slice image SL7 is closer to slice image SL8. In this case, the reverse prediction label P will be selected. B7 As a result, the slice image SL7 is labeled with label P7.

[0086] In some embodiments, the result annotation labels P2 to P7 are not limited to the forward prediction labels P F2 To P F7 And reverse prediction annotation P B2 To P B7 In other embodiments, the result labels P2 to P7 may be selected by the merger 144 according to the forward prediction label PF2 To P F7 And reverse prediction annotation P B2 To P B7 It is determined by the weighted sum of the two.

[0087] For example, regarding the Kth slice image of the serialized medical image SIMG, the result annotation label on the K slice image is obtained by calculating the first distance between the slice image SL1 and the Kth slice image and the second distance between the Kth slice image and the slice image SL8, and then generating a weighted sum between the forward prediction annotation and the backward prediction annotation on the Kth slice image based on the first distance and the second distance.

[0088] For example, regarding the slice image SL2, the first distance between the slice image SL1 and the slice image SL2 is "1", and the second distance between the slice image SL2 and the slice image SL8 is "6". The result annotation label P2 on the slice image SL2 can be obtained by the forward prediction annotation P F2 And reverse prediction annotation P B2 It is generated by the weighted sum between , as shown below:

[0089]

[0090] In some embodiments, the forward prediction label P F2 Contains probability values ​​(e.g., 0% to 100%) corresponding to each pixel on the slice image SL2. These probability values ​​represent whether each pixel is related to the target object or label. B2 The merger 144 is used to calculate the forward prediction label P F2 And reverse prediction annotation P B2 In some examples, if the weighted sum of a pixel exceeds 50%, the pixel will be selected as part of the result annotation label P2 on the slice image SL2. In some examples, if the weighted sum of a pixel is less than 50%, the pixel will be excluded from the result annotation label P2 on the slice image SL2.

[0091] Regarding the slice image SL3, the first distance between the slice image SL1 and the slice image SL3 is "2", and the second distance between the slice image SL3 and the slice image SL8 is "5". The result annotation label P3 on the slice image SL3 can be obtained by the forward prediction annotation P F3 And reverse prediction annotation P B3 It is generated by the weighted sum between , as shown below:

[0092]

[0093] Regarding the slice image SL7, the first distance between the slice image SL1 and the slice image SL7 is "6", and the second distance between the slice image SL7 and the slice image SL8 is "1". The result annotation label P7 on the slice image SL7 can be obtained by the forward prediction annotation P F7 And reverse prediction annotation P B7 It is generated by the weighted sum between , as shown below:

[0094]

[0095] Figure 2 The bidirectional reasoning mechanism shown includes a tracking model 142 and a merger 144. The bidirectional reasoning mechanism can be seamlessly integrated into the Tracking Any Object Model (TAM) framework. The medical image annotation apparatus 100 can effectively utilize a limited amount of user annotation data (i.e., manually annotated P M1 and P M8 ), and automatically generates result annotation labels P2 to P7 on the intermediate slice image according to the bidirectional tracking. The medical image annotation device 100 can improve the efficiency of semi-automatic annotation. Since the bidirectional reasoning mechanism can provide more reference information when tracking objects, it can improve the accuracy of automatic annotation when identifying unlabeled image data.

[0096] Manual annotation on serialized medical images SIMG M1 and P M8 And the resulting annotation labels P2 to P7 can be used as training materials to train medical-related models, such as organ segmentation models, medical image classification models, and diagnosis assistance models.

[0097] Figure 2 The bidirectional reasoning mechanism shown includes the tracking model 142 and the merger 144, which is a post-facto fusion (decision-level fusion) method with a simple architecture and easy implementation. In some embodiments, Figure 2 The bidirectional reasoning mechanism shown may lead to discontinuous shapes in the generated annotations, and sometimes further manual adjustments are required to refine the resulting annotation labels P2 to P7.

[0098] In order to overcome the potential problem of shape discontinuity caused by post-fusion when merging annotation information, this disclosure also provides another method for performing a bidirectional reasoning mechanism. Figure 5 , Figure 5 1 is a schematic diagram of a bidirectional XMem model 146 executed by a processor 140 according to some embodiments of the present disclosure. In some embodiments, the processor 140 is used to execute the bidirectional XMem model 146 as a way to implement a bidirectional reasoning mechanism, which is based on the manual annotation P M1 and P M8Result annotation labels P2, P3, P4, P5, P6 and P7 are generated on the intermediate slice images (eg, slice images SL2, SL3, SL4, SL5, SL6 and SL7).

[0099] XMem is a model proposed to achieve long-term video object segmentation (VOS), which refers to the Atkinson-Shiffrin memory model and consists of short-term and long-term memory systems.

[0100] like Figure 5 As shown, the bidirectional XMem model 146 is an improved version of XMem. The bidirectional XMem model 146 using the bidirectional mechanism can consider the information from the previous and next slices. This bidirectional memory retrieval enhances the contextual understanding of the bidirectional XMem model 146. The bidirectional XMem model 146 adopts a feature-level fusion strategy (rather than decision-level fusion). This strategy allows integration at the feature level, thereby making full use of the model's ability to consider the time and space dimensions in the feature representation. Through feature-level spatial fusion, the annotation information from different perspectives can be more closely and effectively combined to achieve the multi-dimensional analysis capability of the bidirectional XMem model 146.

[0101] In some embodiments, the bidirectional XMem model 146 is used to generate a first manual annotation (ie, a manual annotation P M1 ) and the second manual annotation (ie, manual annotation P M8 ), generating the resulting annotation labels on the intermediate slice images.

[0102] In some embodiments, the bidirectional XMem model 146 can process two paired slice images in the serialized medical image SIMG in each round to generate result annotation labels on the two paired slice images. For example, after receiving the serialized medical image SIMG (including slice images SL1 to SL8), the manual annotation P on the slice image SL1 M1 and manual annotation P on slice image SL8 M8 Afterwards, the bidirectional XMem model 146 may first generate result annotation labels P2 and P7 on the pair of slice images SL2 and SL7.

[0103] like Figure 5 As shown, the bidirectional XMem model 146 includes a sensory memory 146a, a working memory 146b, and a long-term memory 146c.

[0104] The sensory memory 146a is a short-term memory element that is responsible for processing the real-time data of the most recent slice from the serialized medical image SIMG. The sensory memory 146a captures and processes the real-time information that is critical to the current slice segmentation. Figure 5 As shown, the sensory memory 146a is used to store the short-term forward hidden expression h1 and the short-term backward hidden expression h8. The short-term forward hidden expression h1 is based on the manual annotation P M1 The short-term forward hidden expression h1 can reflect the manual annotation P on the slice image SL1. M1 The short-term reverse hidden representation h8 is based on the manual annotation P on the slice image SL8. M8 (Mask / Label) generated.

[0105] The working memory 146b and the long-term memory 146c are used to retain key information that may no longer exist in the short-term memory. The working memory 146b and the long-term memory 146c are used to maintain the correlation between the previous and next images on the serialized medical image SIMG. The working memory 146b can be regarded as a dynamic and flexible storage space for processing information for real-time processing tasks, and the working memory 146b is used to process the current and most recent frames, so that the system can quickly adapt to changes in input information. Compared with the working memory 146b, the long-term memory 146c is used to store important information that is considered to maintain the consistency and correlation between the previous and next images of the serialized medical image SIMG. The long-term memory 146c is used to retain key object features and previous and next image correlation information that lasts for a longer time scale.

[0106] like Figure 5 As shown, the working memory 146b is used to store the forward memory key feature K WF , reverse memory key feature K WB , forward memory value feature V WF And reverse memory value characteristic V WB The long-term memory 146c is used to store the forward long-term memory key feature K LF , reverse long-term memory key feature K LB , forward long-term memory value feature V LF And the reverse long-term memory value feature V LB .

[0107] In some embodiments, the forward memory key feature K in the working memory 146b WF And the forward memory value characteristic V WF Updated according to the forward information data INF1. In the initial stage, the forward information data INF1 includes the slice image SL1 and the manually marked P M1 On the other hand, the reverse memory key feature K in the working memory 146bWB and reverse memory value feature V WB Updated according to the reverse information data INF2. In the initial stage, the reverse information data INF2 includes the slice image SL8 and the manually marked P M8 .

[0108] Using the bidirectional mechanism, the data stored in the working memory 146b and the long-term memory 146c will take into account information from the front and back slices (eg, forward information data INF1 and backward information data INF2). This bidirectional memory retrieval enhances the front and back image understanding of the bidirectional XMem model 146.

[0109] like Figure 5 As shown, the bidirectional XMem model 146 further includes a query encoder 146d, a decoder 146e, and a mask encoder 146f. Figure 5 As shown, the query encoder 146d is used to generate a forward input query Q about the forward input medical image (ie, the slice image SL2). SL2 And the reverse input query Q about the reverse input medical image (i.e. slice image SL7) SL7 The query encoder 146d is responsible for converting the input slice images SL2 and SL7 into a high-dimensional feature representation (i.e., the forward input query Q SL2 And reverse input query Q SL7 ). Next, enter the query Q SL2 And reverse input query Q SL7 Used to query working memory 146b and long-term memory 146c to retrieve relevant past information.

[0110] like Figure 5 As shown, in some embodiments, the forward memory key feature K WF , reverse memory key feature K WB , forward long-term memory key feature K LF and reverse long-term memory key feature K LB The two-way XMem model 146 executed by the processor 140 is used to query Q according to the forward input SL2 And reverse input query Q SL7 The similarity between the memory key Kc and the affinity matrix W is calculated. AM In other words, the processor 140 queries Q according to the forward input SL2 , reverse input query Q SL7 , forward memory key feature K WF , reverse memory key feature K WB , forward long-term memory key feature K LF and reverse long-term memory key feature K LBThen calculate the correlation matrix W AM .

[0111] The bidirectional XMem model 146 is used to merge memory data retrieved at different times (mid-term and long-term) to obtain a consistent and comprehensive memory representation, such as the memory key Kc. Unlike the original XMem, the bidirectional XMem model 146 ensures that the shape and time information can be more completely expressed. If the shape information from the early, mid-term and late time points of the serialized images can assist each other, it is very helpful for continuous tracking of objects in medical images.

[0112] like Figure 5 As shown, the bidirectional XMem model 146 performs a memory read operation to obtain the memory according to the association matrix W AM and the read value Vc to generate the read feature F RO In some embodiments, the processor 140 generates a forward memory value feature V WF , reverse memory value feature V WB , forward long-term memory value feature V LF And the reverse long-term memory value feature V LB The above memory read operation is used to calculate the correlation between various past information in the memory and the features extracted by the query encoder 146d (correlation matrix W AM ) and identify which features or regions are most relevant to the target object.

[0113] Decoder 146e is used to query Q according to the forward input SL2 , reverse input query Q SL7 , read out feature F RO As well as the short-term forward hidden representation h1 and the short-term reverse hidden representation h8 from the sensory memory 146a, generate a forward annotation mask (i.e., the result annotation label P2) regarding the forward input medical image (i.e., slice image SL2) and a reverse annotation mask (i.e., the result annotation label P7) regarding the reverse input medical image (i.e., slice image SL7).

[0114] Decoder 146e is responsible for generating the result segmentation mask from the merged features obtained from query encoder 146d and memory modules (i.e., sensory memory 146a, working memory 146b, and long-term memory 146c). The input of decoder 146e combines the current slice features from query encoder 146d and the high-dimensional feature representation of the relevant historical data retrieved from the memory module. This combination ensures that both the current observation and the past context jointly form the image segmentation result (result annotation label P2 and result annotation label P7).

[0115] The mask encoder 146f is used to generate a short-term forward hidden expression h2 according to the forward labeling mask (i.e., the result labeling label P2), and to generate a short-term reverse hidden expression h7 according to the reverse labeling mask (i.e., the result labeling label P7). The short-term forward hidden expression h2 and the short-term reverse hidden expression h7 will be updated to the perception memory 146a so as to segment another pair of slice images in the future.

[0116] In this case, the result annotation label P2 on the slice image SL2 and the result annotation label P7 on the slice image SL7 can be obtained by the bidirectional XMem model 146 by referring to the manual annotation P on the slice image SL1. M1 and manual annotation P on slice image SL8 M8 Produce. Figure 5 As shown, the result annotation label P2 about the slice image SL2 will be added to the forward information data INF1, and the result annotation label P7 about the slice image SL7 will be added to the backward information data INF2 for subsequent segmentation.

[0117] Please also read Figure 6 , Figure 6 is a schematic diagram of a bidirectional XMem model 146 executed by the processor 140 according to another pair of slice images SL3 and SL6 .

[0118] like Figure 6 As shown, the perception memory 146a is used to store the short-term forward hidden expression h2 and the short-term reverse hidden expression h7. The short-term forward hidden expression h2 can reflect the distribution of the forward annotation mask (i.e., the result annotation label P2) on the slice image SL2. The short-term reverse hidden expression h7 can reflect the distribution of the reverse annotation mask (i.e., the result annotation label P7) on the slice image SL7.

[0119] In some embodiments, the forward memory key feature K in the working memory 146b is updated according to the forward information data INF1. WF and forward memory value characteristic V WF At the current stage, the forward information data INF1 includes slice images SL1, manually marked P M1 , slice image SL2 and result label P2. On the other hand, the reverse memory key feature K in the working memory 146b is updated according to the reverse information data INF2. WB and reverse memory value feature V WB At present, the reverse information data INF2 includes slice images SL8, manual annotation P M8 , slice image SL7 and result annotation label P7.

[0120] like Figure 6As shown, the query encoder 146d is used to generate a forward input query Q about the forward input medical image (ie, the slice image SL3). SL3 And the reverse input query Q about the reverse input medical image (i.e. slice image SL6) SL6 The query encoder 146d is responsible for converting the incoming slice images SL3 and SL6 into a high-dimensional feature representation (i.e., the forward input query Q SL3 and reverse input query Q SL6 ).

[0121] like Figure 6 As shown, in some embodiments, the forward memory key feature K WF , reverse memory key feature K WB , forward long-term memory key feature K LF and reverse long-term memory key feature K LB The two-way XMem model 146 executed by the processor 140 is used to query Q according to the forward input SL3 and reverse input query Q SL6 The similarity calculation association matrix W between the memory key Kc AM .

[0122] like Figure 6 As shown, the bidirectional XMem model 146 performs a memory read operation to obtain the memory according to the association matrix W AM and the read value Vc to generate the read feature F RO .

[0123] Decoder 146e is used to query Q according to the forward input SL3 , reverse input query Q SL6 , read out feature F RO As well as the short-term forward hidden representation h2 and the short-term reverse hidden representation h7 from the sensory memory 146a, a forward annotation mask (i.e., the result annotation label P3) regarding the forward input medical image (i.e., slice image SL3) and a reverse annotation mask (i.e., the result annotation label P6) regarding the reverse input medical image (i.e., slice image SL6) are generated.

[0124] Decoder 146e is responsible for generating the result segmentation mask from the merged features obtained from query encoder 146d and memory modules (sensory memory 146a, working memory 146b and long-term memory 146c). The input of decoder 146e combines the current slice features from query encoder 146d and the high-dimensional feature representation of the relevant historical data retrieved from the memory module. This combination ensures that the current observation and the past context jointly form the segmentation result (result annotation label P3 and result annotation label P6).

[0125] The mask encoder 146f is used to generate a short-term forward hidden expression h3 according to the forward labeling mask (i.e., the result labeling label P3), and to generate a short-term reverse hidden expression h6 according to the reverse labeling mask (i.e., the result labeling label P6). The short-term forward hidden expression h3 and the short-term reverse hidden expression h6 will be updated to the perception memory 146a for subsequent segmentation of another pair of slice images.

[0126] In this case, the result annotation label P3 for the slice image SL3 and the result annotation label P6 for the slice image SL6 can be generated by the bidirectional XMem model 146 with reference to the current input data and the historical data. Figure 6 As shown, the result annotation label P3 about the slice image SL3 will be added to the forward information data INF1, and the result annotation label P6 about the slice image SL6 will be added to the backward information data INF2 for subsequent segmentation.

[0127] Likewise, result annotation labels P4 and P5 for other slice images SL4 and SL5 may be generated by the bidirectional XMem model 146 .

[0128] In some embodiments, Figure 5 and Figure 6 The query encoder 146d, decoder 146e and mask encoder 146f shown may be composed of Figure 1 The software instructions executed by the processor 140 shown are implemented. In some embodiments, Figure 5 and Figure 6 The sensory memory 146a, working memory 146b and long-term memory 146c shown can be composed of Figure 1 The illustrated memory 160 may be implemented as a memory block defined in the memory 160 or may be implemented by a separate memory.

[0129] As shown in the above embodiments, the medical image annotation device 100 can simultaneously process various continuous data, such as serialized images and medical films, and can effectively utilize a variety of different annotation data to perform a bidirectional reasoning mechanism.

[0130] In some embodiments, the medical image annotation apparatus 100 may be implemented by a computer, a computing server or a medical image server. The processor 140 may be implemented by a central processing unit, a graphics processing unit, a tensor processing unit or an application-specific integrated circuit.

[0131] The medical image annotation method performed by the medical image annotation device 100 is also an embodiment of the present disclosure. The medical image annotation method comprises the following steps: obtaining a serialized medical image SIMG; receiving a first manual annotation (e.g., a manual annotation P) on a first slice image of the serialized medical image SIMG; M1 ) and a second manual annotation on the second slice image of the serialized medical image SIMG (e.g., a manual annotation P M8 ); perform bidirectional reasoning (e.g. Figure 2 The tracking model 142 and merger 144 shown, or Figure 5 and Figure 6 The bidirectional XMem model 146 shown in FIG. 146 is used to define the bidirectional XMem model 146 according to the first manual annotation (eg, the manual annotation P M1 ) and a second manual annotation (e.g., manual annotation P M8 ) respectively generate result annotation labels (eg, result annotation labels P2 to P7) on the intermediate slice images (eg, slice images SL2 to SL7) of the serialized medical image SIMG. The details of these steps have been discussed in the above embodiments and will not be repeated.

[0132] Although specific embodiments of the present disclosure have been disclosed with respect to the above-mentioned embodiments, these embodiments are not intended to limit the present disclosure. Various substitutions and improvements can be performed in the present disclosure by a person of ordinary skill in the relevant art without departing from the principles and spirit of the present disclosure. Therefore, the protection scope of the present disclosure is determined by the appended claims.

Claims

1. A medical image annotation device, characterized in that: Include: an interface; and A memory for storing a plurality of serialized medical images; and A processor coupled to the interface and the memory, the processor being configured to: Receiving, through the interface, a first manual annotation on a first slice image among the plurality of serialized medical images and a second manual annotation on a second slice image among the plurality of serialized medical images; and A bidirectional reasoning mechanism is executed to generate a plurality of result annotation labels on a plurality of intermediate slice images of the plurality of serialized medical images according to the first manual annotation and the second manual annotation.

2. The medical image annotation device as claimed in claim 1, wherein the bidirectional reasoning mechanism executed by the processor comprises: Executing a tracking model based on the first manual annotation on the first slice image to generate a plurality of forward prediction annotations on the plurality of intermediate slice images of the plurality of serialized medical images in a forward order; executing the tracking model based on the second manual annotation on the second slice image to generate a plurality of reverse prediction annotations on the plurality of intermediate slice images of the plurality of serialized medical images in a reverse order; and The plurality of forward prediction annotations and the plurality of backward prediction annotations are combined to generate the plurality of result annotation labels on the plurality of intermediate slice images.

3. The medical image annotation device as described in claim 2, wherein the tracking model includes a tracking any object model, the tracking any object model is used to detect, associate and track the first manual annotation on the first slice image between each of the multiple intermediate slice images in the forward order, and the tracking any object model is used to detect, associate and track the second manual annotation on the second slice image between each of the multiple intermediate slice images in the reverse order.

4. The medical image annotation device as claimed in claim 2, wherein corresponding to a K-th slice image in the plurality of serialized medical images, a result annotation label on the K-th slice image is generated by: Calculating a first distance between the first slice image and the Kth slice image and a second distance between the Kth slice image and the second slice image; When the first distance is shorter than the second distance, selecting a forward prediction label on the K-th slice image as the result label on the K-th slice image; and When the second distance is shorter than the first distance, a reverse prediction annotation on the K-th slice image is selected as the result annotation label on the K-th slice image.

5. The medical image annotation device as claimed in claim 2, wherein corresponding to a K-th slice image in the plurality of serialized medical images, a result annotation label on the K-th slice image is generated by: calculating a first distance between the first slice image and the Kth slice image and a second distance between the Kth slice image and the second slice image; and The result label on the K-th slice image is generated according to the first distance and the second distance and in accordance with a weighted sum of a forward prediction label and a backward prediction label on the K-th slice image.

6. The medical image annotation device as claimed in claim 1, wherein the bidirectional reasoning mechanism executed by the processor comprises: A bidirectional XMem model is executed to generate the plurality of result annotation labels on the plurality of intermediate slice images according to the first manual annotation and the second manual annotation.

7. The medical image annotation device as claimed in claim 6, wherein the bidirectional XMem model comprises: a sensory memory for storing a short-term forward hidden representation and a short-term backward hidden representation; a working memory for storing a forward memory key feature, a reverse memory key feature, a forward memory value feature, and a reverse memory value feature; and A long-term memory is used to store a forward long-term memory key feature, a reverse long-term memory key feature, a forward long-term memory value feature and a reverse long-term memory value feature.

8. The medical image annotation device as claimed in claim 7, wherein the bidirectional XMem model further comprises: a query encoder for generating a forward input query about a forward input medical image and a reverse input query about a reverse input medical image; a decoder for generating a forward-annotated mask for the forward-input medical image and a backward-annotated mask for the backward-input medical image according to the forward-input query, the backward-input query, a readout feature, the short-term forward hidden representation from the sensory memory, and the short-term backward hidden representation; and A mask encoder is used to generate the short-term forward hidden representation and the short-term backward hidden representation according to the forward labeled mask and the backward labeled mask.

9. The medical image annotation device as claimed in claim 8, wherein the readout feature is generated according to a correlation matrix and a readout value, The association matrix is ​​calculated by the processor according to the forward input query, the reverse input query, the forward memory key feature, the reverse memory key feature, the forward long-term memory key feature and the reverse long-term memory key feature, The readout value is calculated by the processor according to the forward memory value feature, the reverse memory value feature, the forward long-term memory value feature and the reverse long-term memory value feature.

10. A medical image annotation method, characterized in that: Include: Acquire multiple serialized medical images; receiving a first manual annotation on a first slice image of the plurality of serialized medical images and a second manual annotation on a second slice image of the plurality of serialized medical images; and A bidirectional reasoning mechanism is executed to generate a plurality of result annotation labels on a plurality of intermediate slice images of the plurality of serialized medical images according to the first manual annotation and the second manual annotation.