A medical data image labeling method based on eye tracker point labeling information
By using point annotation information based on eye trackers, medical images are automatically annotated, solving the problem of complex medical image annotation, achieving efficient image segmentation and rapid labeling, and improving segmentation accuracy.
Patent Information
- Application Number
- CN202310192494.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-02
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2043-03-02
AI Technical Summary
Existing medical image annotation methods require a lot of manual work and lack large-scale annotated image datasets, resulting in complex dataset annotation and poor generalization ability of model recognition algorithms.
By using eye-tracking-based point annotation information, medical images are automatically annotated through 3D reconstruction and eye-tracking data acquisition, reducing manual operations. Image segmentation thresholds are established using eye-tracking viewpoint tracking and positioning data, enabling automatic threshold segmentation.
It reduces the medical image annotation process, saves time, improves segmentation accuracy, assists doctors in quickly annotating large amounts of medical data, and reduces manual operations.
Smart Images

Figure CN116152486B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical imaging technology, and in particular to a method for annotating medical data images based on eye tracker point annotation information. Background Technology
[0002] In the era of big data, artificial intelligence (AI) technology is playing an increasingly important role in the development of smart healthcare. Computer-aided diagnosis, particularly in medical image processing, has consistently been at the forefront of AI application and holds immense potential. AI can perform efficient analysis, quickly assisting doctors in handling large amounts of simple, repetitive tasks, providing preliminary diagnoses or reasonable treatment plans, saving doctors' time, improving the efficiency of medical diagnosis and treatment, and enabling more rational and optimized use of medical resources.
[0003] Utilizing deep learning technology to train medical images using various feature information is a major application of artificial intelligence in the medical field. Medical image detection mainly includes three stages: region of interest (ROI) localization, ROI segmentation, and ROI recognition. Existing medical image detection methods often use deep algorithms such as Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). However, these strongly supervised learning-based algorithms often require a large amount of data for fitting and analysis to achieve high accuracy. Therefore, the requirements for training data are relatively high; the quantity, standardization, and accuracy of data annotation all affect the algorithm's training results.
[0004] In general image-related computer artificial intelligence algorithm research, images can be collected from ordinary users using labeled training datasets. However, this method cannot be applied to the field of medical images because medical image annotation requires extensive clinical expertise and can only be performed by doctors or related personnel with rich medical experience. For doctors, annotating medical images often requires manual methods, which are time-consuming and labor-intensive. Therefore, large-scale annotated image datasets are currently lacking in medical image analysis. Furthermore, current medical segmentation datasets are complex to annotate, and the generalization ability of model recognition algorithms is poor. Summary of the Invention
[0005] (a) Technical problems to be solved
[0006] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a medical data image annotation method based on eye tracker point annotation information, which solves the technical problem of complex annotation of segmented datasets in the current medical field.
[0007] (II) Technical Solution
[0008] To achieve the above objectives, the main technical solutions adopted by the present invention include:
[0009] This invention provides a method for annotating medical data images based on eye tracker annotation information, the method comprising:
[0010] S1. Analyze each DICOM medical data image to be labeled in the medical image sequence to obtain the three-dimensional reconstruction information corresponding to each DICOM medical data image;
[0011] DICOM (Digital Imaging and Communications in Medicine) is an international standard for medical images and related information.
[0012] S2. Based on the three-dimensional reconstruction information corresponding to each DICOM medical data image, perform three-dimensional reconstruction processing on the DICOM medical data images in the medical image sequence to obtain the medical image point cloud structure corresponding to the DICOM medical data images in the medical image sequence.
[0013] S3. Based on the medical image point cloud structure, obtain the coronal image, sagittal image and transverse image of the medical image point cloud structure respectively;
[0014] S4. Using an eye tracker, acquire the acquisition information corresponding to the coronal, sagittal, and transverse images of the medical image point cloud structure according to a pre-set first strategy.
[0015] The collected information includes: hotspot area bounding boxes and collection time;
[0016] S5. Based on the acquisition information corresponding to the coronal, sagittal, and transverse images of the medical image point cloud structure, the medical image point cloud structure is converted into a binary image point cloud structure according to a pre-set second strategy.
[0017] Preferably,
[0018] The three-dimensional reconstruction information includes: the image position of the DICOM medical image data, the image orientation of the DICOM medical image data, the total number of rows and columns of the DICOM medical image data, the physical spacing between pixel centers in the DICOM medical image data, the spacing between layers of the DICOM medical image data, and the pixel values of the DICOM medical image data.
[0019] Preferably, S4 specifically includes:
[0020] S41. Determine the initial position of the viewpoint in the scanned image of the medical image point cloud structure by the eye tracker, and start acquiring the position of the viewpoint in the scanned image at preset time intervals until the first event occurs;
[0021] The scanned image is a coronal, sagittal, or transverse image of the medical image point cloud structure.
[0022] The first event is when the eye tracker captures the user blinking twice consecutively within a preset time period;
[0023] S42. Obtain the viewpoint position in the scanned image within the acquisition time;
[0024] The acquisition time is the time from determining the initial position of the eye tracker's viewpoint in the scanned image of the medical image point cloud structure to the occurrence of the first event;
[0025] S43. Based on the viewpoint position in the scanned image within the acquisition time, obtain the hotspot region bounding box of the scanned image of the medical image point cloud structure, and combine the acquisition time and the obtained hotspot region bounding box of the scanned image to form acquisition information corresponding to the scanned image.
[0026] The hotspot region bounding box of the scanned image is the hotspot region bounding box of the coronal image, the hotspot region bounding box of the sagittal image, or the hotspot region bounding box of the transverse image.
[0027] Preferably,
[0028] The initial position of the viewpoint in the scanned image of the medical image point cloud structure is the position of the viewpoint captured by the eye tracker in the scanned image of the medical image point cloud structure where the viewpoint stays for more than 1 second.
[0029] Preferably,
[0030] If the scanned image is a coronal image, then the hot spot region box of the scanned image is the hot spot region box of the coronal image. The hot spot region box of the coronal image is a region box formed by the vertical line of the viewpoint position with the largest horizontal coordinate and the vertical line of the viewpoint position with the smallest horizontal coordinate in the coronal image in the first time time, and the horizontal line of the viewpoint position with the largest vertical coordinate and the horizontal line of the viewpoint position with the smallest vertical coordinate.
[0031] The first time is the acquisition time when the scanned image is a coronal image;
[0032] If the scanned image is a coronal image, then the hot spot region box of the scanned image is the hot spot region box of the sagittal image. The hot spot region box of the sagittal image is a region box formed by the vertical line of the viewpoint position with the largest x-coordinate and the vertical line of the viewpoint position with the smallest x-coordinate in the sagittal image during the second time period as the vertical boundary, and the horizontal line of the viewpoint position with the largest y-coordinate and the horizontal line of the viewpoint position with the smallest y-coordinate as the horizontal boundary.
[0033] The second time is the acquisition time when the scanned image is a sagittal bitmap;
[0034] If the scanned image is a transverse image, then the hot spot region of the scanned image is the hot spot region box of the transverse image. The hot spot region box of the transverse image is a region box formed by the vertical line of the viewpoint position with the largest horizontal coordinate and the vertical line of the viewpoint position with the smallest horizontal coordinate in the coordinates of the viewpoint position in the transverse image in the third time period, with the horizontal line of the viewpoint position with the largest vertical coordinate and the horizontal line of the viewpoint position with the smallest vertical coordinate as the horizontal boundary.
[0035] The third time is the acquisition time when the scanned image is a transverse image.
[0036] Preferably,
[0037] The eye tracker is pre-calibrated with a computer display screen used to display DICOM medical data images;
[0038] The eye tracker is also pre-calibrated through the display windows of the coronal, sagittal, and transverse images of the medical image point cloud structure, respectively.
[0039] Preferably, S5 includes:
[0040] S51. Based on the hotspot region bounding boxes of the coronal image, the sagittal image, and the transverse image of the medical image point cloud structure, obtain the hotspot region bounding boxes.
[0041] S52. Based on the first time, the second time, the third time, and the bounding box of the hotspot region, the medical image point cloud structure is converted into a point cloud structure of a binary image.
[0042] Preferably, S51 specifically includes:
[0043] S511. The hotspot region boxes of the coronal image, sagittal image, and transverse image of the medical image point cloud structure are transformed to the three-dimensional image coordinate system through the three-dimensional reconstruction transformation relationship in the three-dimensional reconstruction process to obtain the first region box corresponding to the hotspot region box of the coronal image in the three-dimensional image coordinate system, the second region box corresponding to the hotspot region box of the sagittal image in the three-dimensional image coordinate system, and the third region box corresponding to the hotspot region box of the transverse image in the three-dimensional image coordinate system.
[0044] S512. The hotspot bounding boxes formed by the vertical lines of the largest and smallest coordinate points on the x-axis, the horizontal lines of the largest and smallest coordinate points on the y-axis, and the first and second vertical lines perpendicular to the z-axis of the largest and smallest coordinate points on the z-axis in the three-dimensional image coordinate system within the first, second, and third region bounding boxes are defined in the three-dimensional image coordinate system.
[0045] Preferably, S52 specifically includes:
[0046] S521. Normalize the first time, the second time, and the third time respectively to obtain the normalized result of the first time, the normalized result of the second time, and the normalized result of the third time.
[0047] S522. Obtain the maximum value of the pixel points in the first region box, the second region box and the third region box respectively, and obtain the first value Threshold using formula (1);
[0048] The formula (1) is:
[0049] (CoronalPixelValue*W1+SagittalPixelValue*W2+AxialPixelValue
[0050] *W3)×45%=Threshold;
[0051] Wherein, W1 is the result of the first time normalization process, W1 = first time / (first time + second time + third time);
[0052] W2 is the result of the second time normalization process, W2 = second time / (first time + second time + third time);
[0053] W3 is the result of the third time normalization process, W3 = third time / (first time + second time + third time);
[0054] CoronalPixelValue is the maximum value of the pixels within the first region bounding box;
[0055] SagittalPixelValue is the maximum value of the pixels within the second region bounding box;
[0056] AxialPixelValue is the maximum value of the pixels within the third region bounding box;
[0057] S523. Based on the first value Threshold, the medical image point cloud structure is converted into a binary image point cloud structure using a pre-set conversion rule.
[0058] The pre-defined conversion rules include:
[0059] Image data points outside the bounding box of the hotspot region will have values set to 0;
[0060] The 3D reconstructed image data points within the bounding box of the hotspot region are assigned a value of 1 in the binary image when the value of the data point is greater than or equal to the first value Threshold, and a value of 0 in the binary image when the value of the data point is less than the first value Threshold.
[0061] Preferably,
[0062] The preset time interval is 30 milliseconds;
[0063] The preset time period is 500 milliseconds.
[0064] (III) Beneficial Effects
[0065] The beneficial effects of this invention are as follows: The medical image annotation method based on eye-tracking annotation information, by employing an eye-tracking device—a behavioral input information device—and using an interactive, mutually supportive approach, reduces the number of steps involved in medical image annotation, saving time. Furthermore, based on the eye-tracking gaze positioning data, a reference threshold for image segmentation is established. Using this threshold as a basis, automatic threshold segmentation is performed, improving segmentation accuracy. Additionally, it assists doctors in minimizing manual operations during image annotation. By recognizing and recording visual attention points through the eye tracker, regions of interest are automatically marked, eliminating the need for doctors to painstakingly browse medical images layer by layer to delineate regions of interest for image annotation. This reduces annotation time, allowing doctors to quickly annotate large amounts of medical data in a short period. Attached Figure Description
[0066] Figure 1 This is a flowchart of a medical data image annotation method based on eye tracker annotation information according to the present invention;
[0067] Figure 2 This is a cross-sectional image of the medical image point cloud structure in the embodiment;
[0068] Figure 3 This is a sagittal image of the medical image point cloud structure in the embodiment;
[0069] Figure 4 This is a coronal image of the medical image point cloud structure in the embodiment;
[0070] Figure 5 This is a flowchart of a medical data image annotation method based on eye tracker annotation information in Example 2. Detailed Implementation
[0071] To better explain and facilitate understanding of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0072] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present invention can be understood more clearly and thoroughly, and that the scope of the present invention can be fully conveyed to those skilled in the art.
[0073] Example 1
[0074] See Figure 1 This embodiment provides a medical data image annotation method based on eye tracker annotation information, the method comprising:
[0075] S1. Analyze each DICOM medical data image to be labeled in the medical image sequence to obtain the three-dimensional reconstruction information corresponding to each DICOM medical data image.
[0076] DICOM (Digital Imaging and Communications in Medicine) is an international standard for medical images and related information.
[0077] The three-dimensional reconstruction information includes: the image position of the DICOM medical image data, the image orientation of the DICOM medical image data, the total number of rows and columns of the DICOM medical image data, the physical spacing between pixel centers in the DICOM medical image data, the spacing between layers of the DICOM medical image data, and the pixel values of the DICOM medical image data.
[0078] S2. Based on the three-dimensional reconstruction information corresponding to each DICOM medical data image, perform three-dimensional reconstruction processing on the DICOM medical data images in the medical image sequence to obtain the medical image point cloud structure corresponding to the DICOM medical data images in the medical image sequence.
[0079] S3. Based on the aforementioned medical image point cloud structure, see [link / reference] Figure 2 , Figure 3 , Figure 4 The coronal, sagittal, and transverse images of the medical image point cloud structure are obtained respectively.
[0080] S4. Using an eye tracker, acquire the acquisition information corresponding to the coronal, sagittal, and transverse images of the medical image point cloud structure according to a pre-set first strategy.
[0081] The collected information includes: hotspot area bounding boxes and collection time.
[0082] In this embodiment, the eye tracker is pre-calibrated with a computer display screen used to display DICOM medical data images; the eye tracker is also pre-calibrated with the display windows of the coronal, sagittal, and transverse images of the medical image point cloud structure, respectively.
[0083] In practical applications of this embodiment, S4 specifically includes:
[0084] S41. Determine the initial position of the viewpoint in the scanned image of the medical image point cloud structure by the eye tracker, and start acquiring the position of the viewpoint in the scanned image at preset time intervals until the first event occurs; in practical applications, the preset time interval is 30 milliseconds.
[0085] The scanned image is a coronal, sagittal, or transverse image of the medical image point cloud structure.
[0086] The first event is when the eye tracker captures the user blinking twice consecutively within a preset time period; in specific applications, the preset time period is 500 milliseconds.
[0087] Specifically, the initial position of the viewpoint in the scanned image of the medical image point cloud structure is the position of the viewpoint captured by the eye tracker in the scanned image of the medical image point cloud structure where the viewpoint stays for more than 1 second.
[0088] S42. Obtain the viewpoint position in the scanned image during the acquisition time.
[0089] The acquisition time is the time from determining the initial position of the eye tracker's viewpoint in the scanned image of the medical image point cloud structure to the occurrence of the first event.
[0090] S43. Based on the viewpoint position in the scanned image within the acquisition time, obtain the hotspot region bounding box of the scanned image of the medical image point cloud structure, and combine the acquisition time and the obtained hotspot region bounding box of the scanned image to form acquisition information corresponding to the scanned image.
[0091] The hotspot region bounding box of the scanned image is the hotspot region bounding box of the coronal image, the hotspot region bounding box of the sagittal image, or the hotspot region bounding box of the transverse image.
[0092] If the scanned image is a coronal image, then the hotspot region box of the scanned image is the hotspot region box of the coronal image. The hotspot region box of the coronal image is a region box formed by the vertical line of the viewpoint position with the largest x-coordinate and the vertical line of the viewpoint position with the smallest x-coordinate in the coronal image in the first time time as the vertical boundary, and the horizontal line of the viewpoint position with the largest y-coordinate and the horizontal line of the viewpoint position with the smallest y-coordinate as the horizontal boundary.
[0093] The first time is the acquisition time when the scanned image is a coronal image.
[0094] If the scanned image is a coronal image, then the hotspot region box of the scanned image is the hotspot region box of the sagittal image. The hotspot region box of the sagittal image is a region box formed by the vertical line of the viewpoint position with the largest x-coordinate and the vertical line of the viewpoint position with the smallest x-coordinate in the sagittal image during the second time period as the vertical boundary, and the horizontal line of the viewpoint position with the largest y-coordinate and the horizontal line of the viewpoint position with the smallest y-coordinate as the horizontal boundary.
[0095] The second time is the acquisition time when the scanned image is a sagittal bitmap.
[0096] If the scanned image is a transverse image, then the hot spot region of the scanned image is the hot spot region box of the transverse image. The hot spot region box of the transverse image is a region box formed by the vertical line of the viewpoint position with the largest horizontal coordinate and the vertical line of the viewpoint position with the smallest horizontal coordinate in the coordinates of the viewpoint position in the transverse image in the third time period, with the horizontal line of the viewpoint position with the largest vertical coordinate and the horizontal line of the viewpoint position with the smallest vertical coordinate as the horizontal boundary.
[0097] The third time is the acquisition time when the scanned image is a transverse image.
[0098] S5. Based on the acquisition information corresponding to the coronal, sagittal, and transverse images of the medical image point cloud structure, the medical image point cloud structure is converted into a binary image point cloud structure according to a pre-set second strategy.
[0099] Specifically, S5 includes:
[0100] S51. Based on the hotspot region bounding boxes of the coronal image, the sagittal image, and the transverse image of the medical image point cloud structure, obtain the hotspot region bounding boxes.
[0101] S51 specifically includes:
[0102] S511. The hotspot region boxes of the coronal image, sagittal image, and transverse image of the medical image point cloud structure are transformed to the three-dimensional image coordinate system through the three-dimensional reconstruction transformation relationship in the three-dimensional reconstruction process to obtain the first region box corresponding to the hotspot region box of the coronal image in the three-dimensional image coordinate system, the second region box corresponding to the hotspot region box of the sagittal image in the three-dimensional image coordinate system, and the third region box corresponding to the hotspot region box of the transverse image in the three-dimensional image coordinate system.
[0103] S512. The hotspot bounding boxes formed by the vertical lines of the largest and smallest coordinate points on the x-axis, the horizontal lines of the largest and smallest coordinate points on the y-axis, and the first and second vertical lines perpendicular to the z-axis of the largest and smallest coordinate points on the z-axis in the three-dimensional image coordinate system within the first, second, and third region bounding boxes are defined in the three-dimensional image coordinate system.
[0104] S52. Based on the first time, the second time, the third time, and the bounding box of the hotspot region, the medical image point cloud structure is converted into a point cloud structure of a binary image.
[0105] Specifically, S52 includes:
[0106] S521. Normalize the first time, the second time, and the third time respectively to obtain the normalized result of the first time, the normalized result of the second time, and the normalized result of the third time.
[0107] S522. Obtain the maximum value of the pixel points in the first region box, the second region box and the third region box respectively, and obtain the first value Threshold using formula (1).
[0108] The formula (1) is:
[0109] (CoronalPixelValue*W1+SagittalPixelValue*W2+AxialPixelValue
[0110] *W3)×45%=Threshold.
[0111] Wherein, W1 is the result of the first time normalization process, W1 = first time / (first time + second time + third time).
[0112] W2 is the result of the second time normalization process, W2 = second time / (first time + second time + third time).
[0113] W3 is the result of the third time normalization process, W3 = third time / (first time + second time + third time).
[0114] CoronalPixelValue is the maximum value of the pixels within the first region bounding box.
[0115] SagittalPixelValue is the maximum value of the pixels within the second region box.
[0116] AxialPixelValue is the maximum value of the pixels within the third region bounding box.
[0117] S523. Based on the first value Threshold, the medical image point cloud structure is converted into a binary image point cloud structure using a pre-set conversion rule.
[0118] The pre-defined conversion rules include:
[0119] Image data points outside the bounding box of the hotspot area will have a value of 0.
[0120] The 3D reconstructed image data points within the bounding box of the hotspot region are assigned a value of 1 in the binary image when the value of the data point is greater than or equal to the first value Threshold, and a value of 0 in the binary image when the value of the data point is less than the first value Threshold.
[0121] This embodiment provides a medical image annotation method based on eye-tracking annotation information. By introducing an eye-tracking device, a behavioral input information device, and using an interactive, mutually supportive approach, the medical image annotation process can be reduced, saving time. Furthermore, based on the eye-tracking gaze positioning data, a reference threshold for image segmentation is established. Using this threshold as a basis, automatic threshold segmentation is performed, improving segmentation accuracy.
[0122] Example 2
[0123] See Figure 5 This embodiment provides a medical data image annotation method based on eye tracker annotation information. Before implementing this method, the eye tracker is first connected to the computer via USB. The calibration program built into the eye tracker is started, and the user of the eye tracker sequentially looks at the four corners and the center point of the computer display screen to complete the eye tracker's capture of the eye's gaze position, thus completing the eye tracker calibration.
[0124] The eye tracker is also calibrated separately with the display windows for coronal, sagittal, and transverse images used to display the point cloud structure of medical images.
[0125] This embodiment provides a medical data image annotation method based on eye tracker annotation information, including:
[0126] A1. Analyze each DICOM medical data image to be labeled in the medical image sequence to obtain the three-dimensional reconstruction information corresponding to each DICOM medical data image.
[0127] The three-dimensional reconstruction information includes: the image position of the DICOM medical image data, the image orientation of the DICOM medical image data, the total number of rows and columns of the DICOM medical image data, the physical spacing between pixel centers in the DICOM medical image data, the spacing between layers of the DICOM medical image data, and the pixel values of the DICOM medical image data.
[0128] A2. Based on the three-dimensional reconstruction information corresponding to each DICOM medical data image, perform three-dimensional reconstruction processing on the DICOM medical data images in the medical image sequence to obtain the medical image point cloud structure corresponding to the DICOM medical data images in the medical image sequence.
[0129] A3. Based on the aforementioned medical image point cloud structure, obtain the coronal, sagittal, and transverse images of the medical image point cloud structure, respectively. (See [reference]) Figure 2 , Figure 3 , Figure 4 .
[0130] A4. Using an eye tracker, according to a pre-set first strategy, acquire the hotspot region bounding boxes of the coronal image, the sagittal image, and the transverse image of the medical image point cloud structure, as well as the first time corresponding to the hotspot region bounding box of the coronal image, the second time corresponding to the hotspot region bounding box of the sagittal image, and the third time corresponding to the hotspot region bounding box of the transverse image.
[0131] In the practical application of this embodiment, A4 specifically includes:
[0132] Using an eye tracker, the hotspot region bounding box of the coronal view of the medical image point cloud structure and the first time corresponding to the hotspot region bounding box of the coronal view are obtained according to a pre-set first strategy.
[0133] In a specific application of this embodiment, the step of using an eye tracker to acquire the hotspot region bounding box of the coronal image of the medical image point cloud structure and the first time corresponding to the hotspot region bounding box of the coronal image according to a pre-set first strategy specifically includes:
[0134] The initial position of the eye tracker's viewpoint in the coronal image of the medical image point cloud structure is determined, and the position of the eye tracker's viewpoint in the coronal image of the medical image point cloud structure is acquired at preset time intervals until the first event occurs.
[0135] The preset time interval is 30 milliseconds.
[0136] The first event is when the eye tracker captures the user blinking twice consecutively within a preset time period; the preset time period is 500 milliseconds.
[0137] The initial position of the viewpoint in the coronal image of the medical image point cloud structure is the position of the viewpoint captured by the eye tracker in the coronal image of the medical image point cloud structure where the viewpoint stays for more than 1 second.
[0138] Obtain the viewpoint position in the coronal image at the first time interval.
[0139] The first time is from the initial position of the eye tracker in the coronal image of the medical image point cloud structure to the occurrence of the first event (that is, the acquisition time when the scanned image is a coronal image).
[0140] Based on the viewpoint position in the coronal image during the first time period, the hotspot region bounding box of the coronal image of the medical image point cloud structure is obtained.
[0141] The hotspot region bounding box of the coronal image of the medical image point cloud structure is a region bounding box formed by the vertical lines of the viewpoint positions in the coronal image at the first time time, with the vertical lines of the viewpoint positions with the largest x-coordinate and the smallest x-coordinate as the vertical boundaries, and the horizontal lines of the viewpoint positions with the largest y-coordinate and the smallest y-coordinate as the horizontal boundaries.
[0142] In the practical application of this embodiment, the eye tracker user controls their gaze to remain fixed at a position in the display window of the coronal image for 1 second, locking the eye tracker's viewpoint within this window and recording the initial position point information as A(x, y), where the coordinates of this initial position point are in the screen coordinate system. Then, starting from point A, the eye tracker's viewpoint moves along with the user's gaze. The position of the eye tracker's viewpoint in the window is recorded every 30 milliseconds. Once all the images to be segmented are within the hotspot region bounding box, the eye tracker user blinks twice consecutively to end the hotspot region drawing, and the drawing time is recorded. At this point, the maximum and minimum coordinate values along the X and Y axes of the screen coordinate system are used as boundaries to draw the hotspot region bounding box of the coronal image of the medical image point cloud structure.
[0143] Using an eye tracker, the hotspot region bounding box of the sagittal image of the medical image point cloud structure and the second time corresponding to the hotspot region bounding box of the sagittal image are obtained according to a pre-set first strategy.
[0144] Specifically, the step of using an eye tracker to acquire the hotspot region bounding box of the sagittal image of the medical image point cloud structure and the second time corresponding to the hotspot region bounding box of the sagittal image according to a pre-set first strategy includes:
[0145] The initial position of the eye tracker's viewpoint in the sagittal image of the medical image point cloud structure is determined, and the position of the eye tracker's viewpoint in the sagittal image of the medical image point cloud structure is acquired at preset time intervals until the first event occurs.
[0146] The first event is when the eye tracker captures the user blinking twice consecutively within a preset time period.
[0147] The initial position of the viewpoint in the sagittal image of the medical image point cloud structure, as captured by the eye tracker, is the position of the viewpoint in the sagittal image of the medical image point cloud structure where the viewpoint lingers for more than 1 second.
[0148] Obtain the viewpoint position in the coronal image at the second time interval.
[0149] The second time is from the initial position of the eye tracker in the sagittal image of the medical image point cloud structure to the occurrence of the first event (the second time is also the acquisition time when the scanned image is a sagittal map).
[0150] Based on the viewpoint position in the sagittal image during the second time period, the hotspot region bounding box of the sagittal image of the medical image point cloud structure is obtained.
[0151] The hotspot region bounding box of the sagittal image of the medical image point cloud structure is a region bounding box formed by the vertical lines of the viewpoint positions in the sagittal image at the second time interval, with the vertical lines of the viewpoint positions with the largest x-coordinate and the smallest x-coordinate as the vertical boundaries, and the horizontal lines of the viewpoint positions with the largest y-coordinate and the smallest y-coordinate as the horizontal boundaries.
[0152] Using an eye tracker, the hotspot region bounding box of the transverse section image of the medical image point cloud structure and the third time corresponding to the hotspot region bounding box of the transverse section image are obtained according to a pre-set first strategy.
[0153] In a specific implementation, the step of using an eye tracker to acquire the hotspot region bounding box of the transverse section image of the medical image point cloud structure and the third time corresponding to the hotspot region bounding box of the transverse section image according to a pre-set first strategy specifically includes:
[0154] The initial position of the eye tracker's viewpoint in the transverse image of the medical image point cloud structure is determined, and the position of the eye tracker's viewpoint in the transverse image of the medical image point cloud structure is acquired at preset time intervals until the first event occurs.
[0155] The first event is when the eye tracker captures the user blinking twice consecutively within a preset time period.
[0156] The initial position of the viewpoint in the transverse image of the medical image point cloud structure, as captured by the eye tracker, is the position of the viewpoint in the transverse image of the medical image point cloud structure where the viewpoint stays for more than 1 second.
[0157] Obtain the viewpoint position in the transverse image during the third time interval.
[0158] The third time is from the initial position of the eye tracker in the transverse image of the medical image point cloud structure to the occurrence of the first event (the third time is also the acquisition time when the scanned image is a transverse image).
[0159] Based on the viewpoint position in the transverse image during the third time period, the hotspot region bounding box of the transverse image of the medical image point cloud structure is obtained.
[0160] The hotspot region bounding box of the transverse image of the medical image point cloud structure is a region bounding box formed by the vertical lines of the viewpoint positions in the transverse image at the third time interval, with the vertical lines of the viewpoint positions with the largest and smallest horizontal coordinates as the vertical boundaries, and the horizontal lines of the viewpoint positions with the largest and smallest vertical coordinates as the horizontal boundaries.
[0161] A5. Based on the hotspot region bounding boxes of the coronal image, sagittal image, and transverse image of the medical image point cloud structure, obtain the hotspot region bounding boxes.
[0162] In the practical application of this embodiment, A5 specifically includes:
[0163] A51. The hotspot region boxes of the coronal image, sagittal image, and transverse image of the medical image point cloud structure are transformed to the three-dimensional image coordinate system through the three-dimensional reconstruction transformation relationship in the three-dimensional reconstruction process to obtain the first region box corresponding to the hotspot region box of the coronal image in the three-dimensional image coordinate system, the second region box corresponding to the hotspot region box of the sagittal image in the three-dimensional image coordinate system, and the third region box corresponding to the hotspot region box of the transverse image in the three-dimensional image coordinate system.
[0164] A52. The hotspot bounding boxes formed by the vertical lines of the largest and smallest coordinate points on the x-axis, the horizontal lines of the largest and smallest coordinate points on the y-axis, and the first and second vertical lines perpendicular to the z-axis of the largest and smallest coordinate points on the z-axis in the three-dimensional image coordinate system within the first, second, and third region bounding boxes are defined in the three-dimensional image coordinate system.
[0165] A6. Based on the first time, the second time, the third time, and the bounding box of the hotspot region, the medical image point cloud structure is converted into a point cloud structure of a binary image.
[0166] Specifically, A6 includes:
[0167] A61. Normalize the first time, the second time, and the third time respectively to obtain the normalized result of the first time, the normalized result of the second time, and the normalized result of the third time.
[0168] A62. Obtain the maximum value of the pixel points in the first region box, the second region box, and the third region box respectively, and use formula (1) to obtain the first value Threshold.
[0169] The formula (1) is:
[0170] (CoronalPixelValue*W1+SagittalPixelValue*W2+AxialPixelValue
[0171] *W3)×45%=Threshold.
[0172] Wherein, W1 is the result of the first time normalization process, W1 = first time / (first time + second time + third time).
[0173] W2 is the result of the second time normalization process, W2 = second time / (first time + second time + third time).
[0174] W3 is the result of the third time normalization process, W3 = third time / (first time + second time + third time).
[0175] CoronalPixelValue is the maximum value of the pixels within the first region bounding box.
[0176] SagittalPixelValue is the maximum value of the pixels within the second region box.
[0177] AxialPixelValue is the maximum value of the pixels within the third region bounding box.
[0178] A63. Based on the first value Threshold, the medical image point cloud structure is converted into a binary image point cloud structure using a pre-set conversion rule.
[0179] The pre-defined conversion rules include:
[0180] Image data points outside the bounding box of the hotspot area will have a value of 0.
[0181] The 3D reconstructed image data points within the bounding box of the hotspot region are assigned a value of 1 in the binary image when the value of the data point is greater than or equal to the first value Threshold, and a value of 0 in the binary image when the value of the data point is less than the first value Threshold.
[0182] In this embodiment, the point cloud structure of the binary image is also saved as a point cloud data set in NRRD format. This file is the annotation file. This completes one medical image annotation.
[0183] This embodiment provides a medical data image annotation method based on eye-tracking annotation information. It can assist doctors in minimizing manual operations during image annotation. By recognizing and recording visual focus points through an eye tracker, regions of interest are automatically marked, eliminating the need for doctors to painstakingly browse through medical images to delineate regions of interest for image annotation. This reduces annotation time, allowing doctors to quickly annotate large amounts of medical data in a short period for training artificial intelligence algorithm models. This significantly shortens the development time of artificial intelligence applications and the final implementation time of algorithms.
[0184] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0185] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions.
[0186] It should be noted that any reference numerals placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In claims that enumerate several means, several of these means may be embodied by the same hardware. The use of the terms first, second, third, etc., is merely for convenience of expression and does not indicate any order. These terms can be understood as part of the component names.
[0187] Furthermore, it should be noted that in the description of this specification, the terms "one embodiment," "some embodiments," "embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0188] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the claims should be interpreted to include both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0189] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, then this invention should also include these modifications and variations.
Claims
1. A method for annotating medical data images based on eye-tracking annotation information, characterized in that, The method includes: S1. Analyze each DICOM medical data image to be labeled in the medical image sequence to obtain the three-dimensional reconstruction information corresponding to each DICOM medical data image; S2. Based on the three-dimensional reconstruction information corresponding to each DICOM medical data image, perform three-dimensional reconstruction processing on the DICOM medical data images in the medical image sequence to obtain the medical image point cloud structure corresponding to the DICOM medical data images in the medical image sequence. S3. Based on the medical image point cloud structure, obtain the coronal image, sagittal image and transverse image of the medical image point cloud structure respectively; S4. Using an eye tracker, acquire the acquisition information corresponding to the coronal, sagittal, and transverse images of the medical image point cloud structure according to a pre-set first strategy. The collected information includes: hotspot region bounding boxes and collection time; S4 specifically includes: S41. Determine the initial position of the viewpoint in the scanned image of the medical image point cloud structure by the eye tracker, and start acquiring the position of the viewpoint in the scanned image at preset time intervals until the first event occurs; The scanned image is a coronal, sagittal, or transverse image of the medical image point cloud structure; the first event is when the eye tracker captures the user blinking twice consecutively within a preset time period. S42. Obtain the viewpoint position in the scanned image within the acquisition time; the acquisition time is the time from determining the initial position of the viewpoint in the scanned image of the medical image point cloud structure by the eye tracker to the occurrence of the first event; S43. Based on the viewpoint position in the scanned image during the acquisition time, obtain the hotspot region bounding box of the scanned image of the medical image point cloud structure, and combine the acquisition time and the obtained hotspot region bounding box of the scanned image to form acquisition information corresponding to the scanned image; the hotspot region bounding box of the scanned image is the hotspot region bounding box of the coronal image, the hotspot region bounding box of the sagittal image, or the hotspot region bounding box of the transverse image. S5. Based on the acquisition information corresponding to the coronal, sagittal, and transverse images of the medical image point cloud structure, respectively, the medical image point cloud structure is converted into a binary image point cloud structure according to a pre-set second strategy; S5 includes: S51. Based on the hotspot region bounding boxes of the coronal image, the sagittal image, and the transverse image of the medical image point cloud structure, obtain the hotspot region bounding boxes. S52. Based on the first time, the second time, the third time and the bounding box of the hotspot region, the medical image point cloud structure is converted into the point cloud structure of a binary image. The first time is the acquisition time when the scanned image is a coronal image; the second time is the acquisition time when the scanned image is a sagittal image; and the third time is the acquisition time when the scanned image is a transverse image.
2. The method according to claim 1, characterized in that, The three-dimensional reconstruction information includes: the image position of the DICOM medical image data, the image orientation of the DICOM medical image data, the total number of rows and columns of the DICOM medical image data, the physical spacing between pixel centers in the DICOM medical image data, the spacing between layers of the DICOM medical image data, and the pixel values of the DICOM medical image data.
3. The method according to claim 2, characterized in that, The initial position of the viewpoint in the scanned image of the medical image point cloud structure is the position of the viewpoint captured by the eye tracker in the scanned image of the medical image point cloud structure where the viewpoint stays for more than 1 second.
4. The method according to claim 3, characterized in that, If the scanned image is a coronal image, then the hot spot region box of the scanned image is the hot spot region box of the coronal image. The hot spot region box of the coronal image is a region box formed by the vertical line of the viewpoint position with the largest horizontal coordinate and the vertical line of the viewpoint position with the smallest horizontal coordinate in the coronal image in the first time time, and the horizontal line of the viewpoint position with the largest vertical coordinate and the horizontal line of the viewpoint position with the smallest vertical coordinate. If the scanned image is a coronal image, then the hot spot region box of the scanned image is the hot spot region box of the sagittal image. The hot spot region box of the sagittal image is a region box formed by the vertical line of the viewpoint position with the largest x-coordinate and the vertical line of the viewpoint position with the smallest x-coordinate in the sagittal image during the second time period as the vertical boundary, and the horizontal line of the viewpoint position with the largest y-coordinate and the horizontal line of the viewpoint position with the smallest y-coordinate as the horizontal boundary. If the scanned image is a transverse image, then the hot spot region of the scanned image is the hot spot region bounding box of the transverse image. The hot spot region bounding box of the transverse image is a region bounding box formed by the vertical lines of the viewpoint positions in the transverse image during the third time period, with the vertical lines of the viewpoint positions with the largest x-coordinate and the smallest x-coordinate as the vertical boundaries, and the horizontal lines of the viewpoint positions with the largest y-coordinate and the smallest y-coordinate as the horizontal boundaries.
5. The method according to claim 4, characterized in that, The eye tracker is pre-calibrated with a computer display screen used to display DICOM medical data images; The eye tracker is also pre-calibrated through the display windows of the coronal, sagittal, and transverse images of the medical image point cloud structure, respectively.
6. The method according to claim 5, characterized in that, S51 specifically includes: S511. The hotspot region boxes of the coronal image, sagittal image, and transverse image of the medical image point cloud structure are transformed to the three-dimensional image coordinate system through the three-dimensional reconstruction transformation relationship in the three-dimensional reconstruction process to obtain the first region box corresponding to the hotspot region box of the coronal image in the three-dimensional image coordinate system, the second region box corresponding to the hotspot region box of the sagittal image in the three-dimensional image coordinate system, and the third region box corresponding to the hotspot region box of the transverse image in the three-dimensional image coordinate system. S512. The hotspot bounding boxes formed by the vertical lines of the largest and smallest coordinate points on the x-axis, the horizontal lines of the largest and smallest coordinate points on the y-axis, and the first and second vertical lines perpendicular to the z-axis of the largest and smallest coordinate points on the z-axis in the three-dimensional image coordinate system within the first, second, and third region bounding boxes are defined in the three-dimensional image coordinate system.
7. The method according to claim 6, characterized in that, S52 specifically includes: S521. Normalize the first time, the second time, and the third time respectively to obtain the normalized result of the first time, the normalized result of the second time, and the normalized result of the third time. S522. Obtain the maximum value of the pixel points in the first region box, the second region box and the third region box respectively, and use formula (1) to obtain the first value Threshold; The formula (1) is: (CoronalPixelValue*W1+SagittalPixelValue*W2+AxialPixelValue *W3)×45%=Threshold; Wherein, W1 is the result of the first time normalization process, W1 = first time / (first time + second time + third time). W2 is the result of the second time normalization process, W2 = second time / (first time + second time + third time). W3 is the result of the third time normalization process, W3 = third time / (first time + second time + third time). CoronalPixelValue is the maximum value of the pixels within the first region bounding box; SagittalPixelValue is the maximum value of the pixels within the second region bounding box; AxialPixelValue is the maximum value of the pixels within the third region bounding box; S523. Based on the first value Threshold, the medical image point cloud structure is converted into a binary image point cloud structure using a pre-set conversion rule. The pre-defined conversion rules include: Image data points outside the bounding box of the hotspot region will have values set to 0; The 3D reconstructed image data points within the bounding box of the hotspot region are assigned a value of 1 in the binary image when the value of the data point is greater than or equal to the first value Threshold, and a value of 0 in the binary image when the value of the data point is less than the first value Threshold.
8. The method according to claim 7, characterized in that, The preset time interval is 30 milliseconds; The preset time period is 500 milliseconds.
Citation Information
Patent Citations
Data acquisition method / system based on behaviors of doctors, and medical image processing system
CN109887583A
Medical image labeling system
CN110993067A