AI-based medical image intelligent labeling method and system

By generating an initial mask using an AI model and combining it with user interaction, and utilizing a condition-guided generative repair model, the problem of low efficiency in traditional annotation is solved, achieving efficient and accurate intelligent annotation of medical images.

CN122117272APending Publication Date: 2026-05-29ZHEJIANG FEITU IMAGING TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG FEITU IMAGING TECH CO LTD
Filing Date
2026-01-19
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Traditional medical image annotation relies on AI pre-annotation results, which require manual correction by physicians. This is inefficient and produces inconsistent results. Existing technologies struggle to achieve efficient and accurate intelligent annotation.

Method used

An initial mask is generated through pre-annotation using an AI model and then overlaid on the image receiving user interaction in a semi-transparent form. This captures the rasterization correction intent of the brush strokes. A conditionally guided generative repair model is then used to generate a carefully selected mask candidate set for physicians to choose the final mask.

Benefits of technology

It improves the efficiency and interactive experience of medical image annotation, realizes an intelligent collaborative closed loop from user correction intent to accurate mask generation, and improves the accuracy and consistency of annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122117272A_ABST
    Figure CN122117272A_ABST
Patent Text Reader

Abstract

The application discloses an AI-based medical image intelligent labeling method and system, relating to the technical field of intelligent labeling, which captures and analyzes the user's simple interactive strokes, converts them into structured guiding information containing correction intentions. At the same time, the high-level intention is fused with the deep visual features of the original medical image to jointly drive a conditional guided generative repair model. The model explores the exploratory regeneration in the defect area based on the user's intention and actively provides a candidate set containing multiple high-quality correction schemes conforming to anatomical logic. Finally, the user's interaction mode changes from tedious pixel-level fine carving to efficient selection and confirmation of high-quality candidate items provided by the system. In this way, an intelligent collaborative closed loop from user correction intention to accurate mask generation is realized, thereby improving the efficiency and interaction experience of medical image intelligent labeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent annotation technology, and more specifically, to an AI-based intelligent annotation method and system for medical images. Background Technology

[0002] In modern clinical medicine, medical imaging technologies such as computed tomography (CT) and magnetic resonance imaging (MRI) have become indispensable core tools for disease diagnosis, treatment planning, and prognostic assessment. The accurate interpretation of massive amounts of medical image data relies on the precise segmentation and annotation of organs, tissues, or lesion regions within the images. However, traditional manual annotation methods heavily depend on physicians' professional knowledge and clinical experience, resulting in high labor intensity, extremely long processing times, and a lack of consistency and reproducibility due to subjective factors. This has become a bottleneck restricting clinical work efficiency and large-scale medical imaging research. To overcome this predicament, introducing artificial intelligence technology to achieve automated and intelligent analysis and processing of medical images has become an industry consensus and development trend.

[0003] To improve annotation efficiency, existing technologies have introduced AI-assisted annotation techniques. These techniques typically use deep learning models to pre-annotate or automatically segment medical images, generating an initial mask or contour. These approaches reduce the workload of manual annotation from scratch to some extent. However, limited by model generalization capabilities, differences in image quality, and the complex and varied morphology of lesions, the initial masks generated by AI often fail to meet the accuracy standards for clinical applications, frequently exhibiting flaws that contradict anatomical intuition. Examples include excessively smooth or rough boundary contours, voids within the target area, or topological errors such as incorrect adhesion between small tissues. Faced with these imperfect pre-annotation results, clinicians are forced to revert to traditional pixel-based tools like erasers or brushes for meticulous refinement. This correction process is not only cumbersome but also inefficient; when dealing with complex or minute structural errors, the workload can sometimes be no less than completely re-annotating manually, significantly diminishing the efficiency advantage of AI pre-annotation.

[0004] Therefore, there is an urgent need for an optimized AI-based intelligent annotation method and system for medical images. Summary of the Invention

[0005] This application is made in order to solve the above-mentioned technical problems.

[0006] According to one aspect of this application, an AI-based intelligent annotation method for medical images is provided, comprising: Acquiring medical images; Medical images are pre-annotated using an AI model to obtain an initial mask; The initial mask is overlaid on the medical image in a semi-transparent manner and user interaction is received. The user interaction is rasterized using strokes to obtain a revised intent map; Condition-guided generative mask repair of medical images based on modified intent maps is used to obtain a selected set of mask candidates. The user selects the most satisfactory candidate mask from the selected mask candidate set as the final mask.

[0007] According to another aspect of this application, an AI-based intelligent annotation system for medical images is provided, comprising: The medical image acquisition module is used to acquire medical images; The medical image pre-annotation module is used to pre-annotate medical images based on AI models to obtain initial masks; The initial mask overlay module is used to overlay the initial mask onto the medical image in a semi-transparent form and receive user interaction. The stroke rasterization module is used to rasterize user interactions with strokes to obtain a revised intent map; The generative mask repair module is used to perform condition-guided generative mask repair on medical images based on the modified intent map to obtain a selected set of mask candidates. The mask confirmation and output module is used to select the most satisfactory candidate mask from the selected mask candidate set as the final mask.

[0008] Compared to existing technologies, this application provides an AI-based intelligent medical image annotation method and system. It captures and analyzes the user's concise interactive strokes, transforming them into structured guiding information containing corrective intent. Simultaneously, this high-level intent is fused with the deep visual features of the original medical image to jointly drive a conditionally guided generative repair model. This model exploratoryly regenerates in flawed areas based on the user's intent and proactively provides a candidate set of high-quality correction schemes that conform to anatomical logic. Ultimately, the user's interaction mode shifts from tedious pixel-level refinement to efficient selection and confirmation of high-quality candidates provided by the system. This achieves an intelligent collaborative closed loop from user corrective intent to accurate mask generation, thereby improving the efficiency and interactive experience of intelligent medical image annotation. Attached Figure Description

[0009] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0010] Figure 1 This is a flowchart of an AI-based intelligent annotation method for medical images according to an embodiment of this application.

[0011] Figure 2 This is a data flow diagram of an AI-based intelligent annotation method for medical images according to an embodiment of this application.

[0012] Figure 3 This is a flowchart of sub-step S2 of the AI-based intelligent medical image annotation method according to an embodiment of this application.

[0013] Figure 4 This is a flowchart of sub-step S4 of the AI-based intelligent medical image annotation method according to an embodiment of this application.

[0014] Figure 5 This is a flowchart of sub-step S43 of the AI-based intelligent medical image annotation method according to an embodiment of this application.

[0015] Figure 6 This is a flowchart of sub-step S5 of the AI-based intelligent medical image annotation method according to an embodiment of this application.

[0016] Figure 7 This is a block diagram of an AI-based intelligent medical image annotation system according to an embodiment of this application. Detailed Implementation

[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0018] To address the problems mentioned above in the background technology, this application proposes an AI-based intelligent annotation method for medical images. Figure 1 This is a flowchart of an AI-based intelligent annotation method for medical images according to an embodiment of this application. Figure 2 This is a data flow diagram of an AI-based intelligent annotation method for medical images according to an embodiment of this application. For example... Figure 1and Figure 2 As shown, the AI-based intelligent medical image annotation method includes the following steps: S1, acquiring medical images; S2, performing pre-annotation on the medical images based on an AI model to obtain an initial mask; S3, overlaying the initial mask onto the medical images in a semi-transparent form and receiving user interaction; S4, rasterizing the user interaction to obtain a correction intent map; S5, performing condition-guided generative mask repair on the medical images based on the correction intent map to obtain a selected mask candidate set; S6, selecting the most satisfactory candidate mask from the selected mask candidate set as the final mask.

[0019] In the aforementioned AI-based intelligent medical image annotation method, step S1 involves acquiring medical images. It should be understood that a lack of compliant and high-quality image input will render subsequent pre-annotation and correction processes unreliable, hindering accurate annotation. Therefore, this application collects medical image data from actual clinical applications through compliant channels to ensure the authenticity, completeness, and validity of the anatomical information. This provides reliable raw materials for subsequent AI model pre-annotation, guarantees the quality of the initial mask generation, and provides accurate anatomical context support for user interaction correction and generative mask repair, laying the foundation for the entire intelligent annotation process.

[0020] Specifically, in one possible embodiment, step S1 is implemented as follows: First, the medical image data of the target patient is retrieved through the hospital's clinical image storage and transmission system. This system has completed the anonymization of patient information and complies with medical data security standards. Second, for images of different modalities such as CT and MRI, they are completely exported according to their original storage formats to ensure that key data such as image pixel information and scanning parameters are not lost. Finally, the exported image data is verified for integrity, checking whether the image sequence is continuous and whether the pixel resolution meets the annotation requirements. Abnormal images with data corruption or missing parameters are removed to ensure that the input medical images can meet the needs of subsequent standardized preprocessing and AI model inference.

[0021] In the aforementioned AI-based intelligent medical image annotation method, step S2 involves pre-annotating the medical image using an AI model to obtain an initial mask. It should be understood that traditional manual annotation relies on physicians performing pixel-by-pixel operations, which is time-consuming, labor-intensive, and the annotation results are easily affected by subjective factors, failing to meet the needs of large-scale medical image annotation. Therefore, this application further utilizes a pre-trained AI model to automatically segment and annotate medical images, thereby generating a preliminary lesion region mask. This transforms the annotation work from creating from scratch to correcting based on initial results, significantly reducing the amount of manual work required by physicians, while providing a clear framework for subsequent user-interactive corrections, improving the efficiency and consistency of the overall annotation process.

[0022] In particular, in one specific embodiment, Figure 3 This is a flowchart of sub-step S2 of the AI-based intelligent medical image annotation method according to an embodiment of this application. Figure 3 As shown, step S2 includes: S21, performing image data standardization preprocessing on the medical image to obtain a preprocessed medical image; S22, inputting the preprocessed medical image into a pre-trained deep learning segmentation network to obtain a probability map; S23, generating an initial mask based on the probability map.

[0023] Specifically, step S21 involves performing image data standardization preprocessing on the medical images to obtain preprocessed medical images. It should be understood that medical images generated by different devices and with different scanning parameters differ in format, grayscale range, and pixel resolution. Directly inputting these images into an AI model would lead to poor model inference stability, reduced generalization ability, and an inability to guarantee pre-annotation accuracy. Therefore, this application further standardizes the original medical images to unify the data format and feature distribution. This eliminates data heterogeneity from different sources, ensuring that the input data received by the AI ​​model meets the feature requirements during its training, improving the model's adaptability to different images, and thus guaranteeing the accuracy and reliability of subsequent pre-annotation results.

[0024] Specifically, in one possible embodiment, step S21 is implemented as follows: First, format conversion is performed to extract the original DICOM format medical image into a pixel data matrix, while irrelevant patient metadata (already desensitized) is removed. Second, grayscale standardization is performed by setting a fixed window width and level according to the corresponding anatomical location (e.g., lung, brain) to map the original grayscale values ​​to a preset effective observation range. Finally, resolution normalization is performed by scaling the image to the preset input resolution of the AI ​​model, such as 512×512 pixels, and normalizing the pixel values ​​to ensure that all pixel values ​​are within a uniform numerical range, ultimately obtaining a preprocessed medical image that meets the model's input requirements.

[0025] Specifically, in step S22, the preprocessed medical image is input into a pre-trained deep learning segmentation network to obtain a probability map. It should be understood that since the preprocessed medical image only achieves data format standardization, its anatomical features need to be analyzed by a professional model to distinguish lesions from the background. Furthermore, directly outputting the binarization result would lose pixel-level prediction confidence information, which is detrimental to subsequent accurate correction. Therefore, this application further inputs the preprocessed image into a pre-trained deep learning segmentation network to generate a probability map of each pixel belonging to the lesion region. This preserves the refined information of the model's predictions, providing a flexible basis for subsequent binarization to generate the initial mask, while also facilitating subsequent analysis of the model's prediction confidence for different regions, thus improving the accuracy of the initial mask generation.

[0026] Specifically, in one possible embodiment, step S22 is implemented as follows: First, a deep learning network suitable for the medical image segmentation task, such as U-Net or its variants, is selected. The pre-training process of this network is performed on a large-scale, diverse medical image dataset, where each image is accompanied by a pixel-level segmentation mask manually annotated by experienced physicians. The training objective is to enable the network to learn the precise mapping relationship from the original medical image to its corresponding standard mask. Typically, loss functions such as Dice loss or cross-entropy loss are used to optimize the network parameters, making its predicted probability map as close to the standard as possible. After training, the network possesses the ability to perform high-quality automatic segmentation of new images. In implementation, the pre-processed medical image is input into this pre-trained network. After processing by the encoder and decoder structure, the network outputs a probability map with the same size as the original image via a sigmoid activation function. The value of each pixel (between 0 and 1) in the map represents its confidence level in belonging to the target region.

[0027] Specifically, step S23 generates an initial mask based on the probability map. It should be understood that since the values ​​in the probability map are continuous confidence levels, they cannot directly distinguish between lesions and background areas, nor can they serve as interactive annotation results for user correction; therefore, they need to be converted into a clear binary mask. Thus, this application further processes the probability map based on preset rules to generate a clear binary initial mask. In a specific example of this application, step S23 includes: generating an initial mask using the following formula:

[0028] in, For the first Image samples in Predicted probability map at each pixel. The preset threshold, For the first Image samples in A binary initial mask is used at each pixel. This allows the model's probability prediction results to be transformed into intuitive region divisions, clearly marking the pre-labeled lesion range, providing users with clear interactive objects, and facilitating users to quickly identify deviations in the pre-labeling and make targeted corrections.

[0029] In the aforementioned AI-based intelligent annotation method for medical images, step S3 involves overlaying an initial mask onto the medical image in a semi-transparent form and receiving user interaction. It should be understood that since the initial mask is a pre-annotation result of the AI ​​model, it may have issues such as boundary deviations, region omissions, or misjudgments. Simply displaying the mask alone is insufficient for physicians to determine its match with the actual lesions; verification in conjunction with the original image is necessary. Therefore, this application further overlays the initial mask onto the medical image with a preset transparency and provides a dedicated interactive tool for physicians to operate. This allows physicians to intuitively compare the consistency between the mask and the lesions in the image and express their correction intentions through interaction. This provides physicians with clear judgment criteria, ensuring that interactive operations accurately target the areas requiring correction, avoiding correction deviations caused by information fragmentation, and laying an accurate interactive foundation for subsequent intent interpretation.

[0030] Specifically, in one possible embodiment, step S3 is implemented as follows: First, the overlay transparency of the initial mask is set to 50%, which allows the mask outline to be clearly displayed without obscuring lesion details in the image. Second, the overlaid image-mask combination is presented in a professional annotation interface, which simultaneously provides a positive correction brush (green), a negative correction brush (red), and brush thickness adjustment functions. Finally, the physician's mouse operations are monitored; when the physician clicks or drags on the interface using the specified brush, the coordinate sequence of the brush trajectory and the corresponding correction type (positive or negative) are recorded in real time, forming complete user interaction data.

[0031] In the aforementioned AI-based intelligent annotation method for medical images, step S4 involves rasterizing the user interaction strokes to obtain a correction intent map. It should be understood that since user interactions exist as a coordinate sequence of brush trajectories, they are unstructured data. Generative inpainting models cannot directly parse the implied correction intent; they need to be transformed into structured image data of the same dimension as the medical image and the initial mask. Therefore, this application further rasterizes the brush trajectories of the user interaction, determines reliable anchor point regions in conjunction with the initial mask, and then synthesizes multi-channel images to transform abstract interactive actions into concrete, pixel-level intent markers. This allows the construction of complete intent information including positive correction, negative correction, and reliable anchor points, providing clear and resolvable guidance signals for subsequent conditionally guided generative mask inpainting, ensuring the model accurately understands the physician's correction needs.

[0032] In particular, in one specific embodiment, Figure 4 This is a flowchart of sub-step S4 of the AI-based intelligent medical image annotation method according to an embodiment of this application. Figure 4As shown, step S4 includes: S41, extracting a positive correction stroke set and a negative correction stroke set from user interaction; S42, performing intention stroke rasterization and channel generation on the positive correction stroke set and the negative correction stroke set to obtain a positive correction channel map and a negative correction channel map; S43, performing anchor point region calculation and three-channel intention map synthesis on the initial mask, the positive correction channel map and the negative correction channel map to obtain a corrected intention map.

[0033] Specifically, step S41 involves extracting a set of positive correction strokes and a set of negative correction strokes from the user interaction. It should be understood that since the user interaction contains two opposing correction intentions—namely, supplementing lesion areas missed by the AI ​​(positive) and eliminating non-lesion areas misjudged by the AI ​​(negative)—if the strokes of these two intentions are stored together, the model will be unable to distinguish the correction direction in subsequent processing, leading to deviations in intention transmission. Therefore, this application further performs type identification on the brush operations during the user interaction process, classifying and storing the stroke trajectories of different intentions separately to clearly distinguish the interaction data of the two correction directions. This provides a classification basis for subsequent rasterization processing, ensuring that each correction intention can be parsed and labeled separately, avoiding subsequent repair errors caused by intention confusion, and improving the accuracy of intention transmission.

[0034] Specifically, in one possible embodiment, step S41 is implemented as follows: First, during the interactive annotation tool initialization phase, unique operation type identifiers are assigned to the positive correction brush and the negative correction brush, respectively. The positive correction brush identifier represents the supplementary area, and the negative correction brush identifier represents the removal area. These identifiers are then bound to the brush's visual style (e.g., green for positive, red for negative). Second, during user operation, the system monitors the brush's start, movement, and termination events in real time. When the brush starts, the corresponding operation type identifier is recorded. During brush movement, the screen coordinates of the brush's center point are collected at fixed time intervals (e.g., 15 milliseconds / time), forming a continuous coordinate sequence. Finally, the coordinate sequences are categorized according to the operation type identifiers. All coordinate sequences identified as supplementary areas are integrated into a positive correction stroke set, and all coordinate sequences identified as removal areas are integrated into a negative correction stroke set. Both sets are associated with the unique identifier of the currently processed medical image, ensuring clear data ownership.

[0035] Specifically, in step S42, the positive and negative correction stroke sets are rasterized and generated with intention strokes to obtain positive and negative correction channel maps. It should be understood that since the positive and negative correction stroke sets are vector data existing in the form of coordinate sequences, their data format is incompatible with the pixel matrix format of medical images and initial masks, and generative models cannot directly use vector data as guiding conditions. Therefore, this application further creates a blank binary image of the same size as the medical image, and rasterizes the coordinate sequences of the two stroke sets into the corresponding images, thereby converting vector strokes into pixel-level region markers. This results in two independent binary channel maps, accurately marking the lesion areas to be supplemented and the non-lesion areas to be removed, providing standardized pixel-level data for subsequent anchor point calculation and three-channel intention map synthesis, ensuring that the model can directly read the correction intention.

[0036] Specifically, in one possible embodiment, step S42 is implemented as follows: First, two blank single-channel images with the exact same resolution as the medical image are created, named the positive correction channel image and the negative correction channel image, respectively, with initial pixel values ​​set to 0. Second, each coordinate sequence in the positive correction stroke set is traversed, mapping each coordinate point in the sequence to a pixel coordinate of the image, and setting the value of these pixel coordinates in the positive correction channel image to 1. Subsequently, the negative correction stroke set is traversed in the same way, setting the value of the corresponding pixel coordinate in the negative correction channel image to 1. Finally, edge smoothing processing is performed on the two channel images to eliminate discrete pixels caused by the coordinate sampling interval, ensuring the continuity of the correction area.

[0037] Specifically, in step S43, anchor point regions are calculated and three-channel intent maps are synthesized from the initial mask, the positive correction channel map, and the negative correction channel map to obtain the corrected intent map. It should be understood that since the initial mask contains areas where the physician has not interacted, these areas are reliable parts with a high degree of matching with the actual lesions in the AI ​​pre-annotation. If these areas are not explicitly marked, the generative model may mistakenly modify them during repair. Furthermore, the positive and negative channels alone cannot fully convey the intent to "preserve reliable areas," requiring supplementary anchor point information. Therefore, this application further extracts anchor point regions from the initial mask through logical operations, and then stacks the anchor point channels with the positive and negative channels to construct complete correction information containing the three intents of addition, deletion, and retention. This allows the generative model to strictly follow the physician's correction instructions while preserving reliable areas in the initial mask, avoiding unnecessary modifications and improving the accuracy and stability of the repair results.

[0038] In particular, in one specific embodiment, Figure 5 This is a flowchart of sub-step S43 of the AI-based intelligent medical image annotation method according to an embodiment of this application. Figure 5 As shown, step S43 includes: S431, generating an interactive union mask based on the positive correction channel map and the negative correction channel map; S432, performing stable anchor point channel calculation on the initial mask and the interactive union mask to obtain an anchor point channel map; S433, synthesizing a multi-channel correction intent map from the positive correction channel map, the negative correction channel map, and the anchor point channel map to obtain a correction intent map.

[0039] More specifically, step S431 generates an interaction union mask based on the positive and negative correction channel maps. It should be understood that since the positive and negative correction channel maps respectively carry two opposite correction intentions, storing them separately makes it impossible to intuitively determine the total area covered by all user interactions. This lack of a unified interaction boundary basis when subsequently calculating anchor point regions (excluding interaction areas) can easily lead to anchor point regions containing pixels that need correction. Therefore, this application further performs a logical OR operation on the two channel maps to integrate all interaction areas, forming a single interaction range marker. This clearly defines the correction areas that the user is concerned with, providing an accurate exclusion basis for subsequent anchor point calculations, ensuring that anchor point regions only contain trusted parts that have not been interacted with, and avoiding confusion between anchor points and correction areas.

[0040] Specifically, in one possible embodiment, step S431 is implemented as follows: First, a blank binary image with the same size as the positive and negative correction channel images is created and set as an interactive union mask, with initial pixel values ​​of 0. Second, each pixel in the two channel images is traversed; if a pixel is 1 in the positive channel image or 1 in the negative channel image, the corresponding pixel value in the interactive union mask is set to 1. Finally, edge noise reduction processing is performed on the interactive union mask, that is, isolated points with an area of ​​less than 3 pixels are removed, and redundant marks caused by brush mis-touches are eliminated to obtain a complete interactive union mask.

[0041] More specifically, step S432 involves calculating stable anchor channels on the initial mask and the interaction union mask to obtain an anchor channel map. It should be understood that since the initial mask contains areas not touched by user interaction, these areas are reliable parts in the AI ​​pre-annotation consistent with the physician's expectations. If not separately labeled, the generative model may make unnecessary adjustments during repair, disrupting the integrity of correctly labeled lesion boundaries or regions. Therefore, this application further performs a logical AND operation on the logical NOT result of the initial mask and the interaction union mask to extract reliable areas not interacted with in the initial mask. In a specific example of this application, step S432 includes: calculating stable anchor channels on the initial mask and the interaction union mask using the following formula:

[0042] in, For the first Image samples in Negative correction channel map at the pixel. For the first Image samples in Positive correction channel map at the pixel For the first Image samples in Initial mask at pixel, Logical NOT, For logical AND, For logical OR, For the first Image samples in The anchor point channel map value at each pixel. That is, first, a logical OR operation is performed on the positive and negative correction channel maps to integrate all areas interacted with by the user. Then, the result is logically NOTed to obtain the areas not interacted with by the user. Finally, a logical AND operation is performed on this result with the initial mask. The resulting anchor point channel map value essentially filters out the areas in the initial mask that the user has not modified. This clearly marks the anchor points that the model needs to retain, ensuring that adjustments are only made to interactive areas during the repair process, maintaining the stability of reliable areas, reducing invalid modifications, and lowering the complexity of model repair.

[0043] More specifically, step S433 involves synthesizing a multi-channel correction intent map from the positive correction channel map, negative correction channel map, and anchor point channel map to obtain a correction intent map. It should be understood that since the positive, negative, and anchor point channel maps each carry a correction intent, if they are input into the generative model individually, the model needs to read the data multiple times and cannot connect the logical relationships between different intents, easily leading to fragmented intent understanding and an inability to fully grasp the physician's correction needs, thus affecting the repair effect. Therefore, this application further stacks the three channel maps along the channel dimension to synthesize a single multi-channel image, thereby integrating the three types of intent information: addition, deletion, and retention. This allows the model to obtain complete correction guidance through a single data read, ensuring that the correlation between different intents is accurately captured, improving the model's overall understanding of correction needs, and laying the foundation for accurate generative repair.

[0044] Specifically, in one possible embodiment, step S433 is implemented as follows: First, verify that the dimensions (height and width) of the three channel images are completely consistent with the medical image, and that the pixel format is uniformly 8-bit binary. Second, create a three-dimensional tensor with dimensions of height × width × 3. Assign the positive correction channel image to the first channel of the tensor, the negative correction channel image to the second channel, and the anchor point channel image to the third channel. Finally, perform normalization processing on the tensor, that is, scale the pixel values ​​to the [0,1] range to ensure that the data format meets the input requirements of the generative model, thus obtaining a complete correction intention map.

[0045] In the aforementioned AI-based intelligent medical image annotation method, step S5 involves performing condition-guided generative mask repair on the medical image based on the modified intent map to obtain a selected mask candidate set. It should be understood that after the initial mask is modified by user interaction, the composite information of AI pre-annotation and user intent needs to be transformed into an accurate final mask. Traditional segmentation models struggle to simultaneously accommodate multi-source intents and generate diverse results. Generative models possess the capability for multimodal generation under conditional guidance. Therefore, this application further uses the modified intent map as conditional guidance and the medical image as the basic input to drive the generative model to perform mask repair, thereby generating multiple mask candidates that conform to the modified intent. This covers the possibilities of different repair details, providing sufficient samples for physicians to select the optimal mask. Simultaneously, through the continuity and precision of generative modeling, the matching accuracy between the mask and the actual lesion is improved, overcoming the limitations of a single model output.

[0046] In particular, in one specific embodiment, Figure 6 This is a flowchart of sub-step S5 of the AI-based intelligent medical image annotation method according to an embodiment of this application. Figure 6 As shown, step S5 includes: S51, performing multimodal conditional encoding and fusion on the modified intention map and medical image to obtain a multi-scale conditional tensor; S52, performing guided iterative denoising on the multi-scale conditional tensor to obtain a predicted soft mask; S53, performing proposal binarization and multi-scheme collection on the predicted soft mask to obtain a selected mask candidate set.

[0047] Specifically, step S51 involves multimodal conditional encoding and fusion of the modified intent map and the medical image to obtain a multi-scale conditional tensor. It should be understood that since the modified intent map is a semantic-level correction guide (channelized intent), and the medical image is pixel-level anatomical information (grayscale / structural features), the two have significant modal differences and different feature levels. Directly inputting them into a generative model would lead to feature fragmentation, preventing the model from collaboratively understanding multi-source information. Therefore, this application further performs feature encoding on the two modalities separately, and then performs feature fusion at multiple scale levels to generate a unified multi-scale conditional tensor. This allows the generative model to simultaneously capture the semantic instructions of the modified intent and the anatomical details of the medical image, achieving precise association between intent and image at different feature levels. This provides comprehensive and hierarchical guidance signals for subsequent iterative denoising, improving the semantic consistency and anatomical accuracy of mask restoration.

[0048] Specifically, in one possible embodiment, step S51 is implemented as follows: First, multi-scale feature extraction is performed on the medical image, obtaining a sequence of feature maps from low to high (corresponding to different resolutions of the image) through a convolutional neural network (containing 3 downsampling layers). Second, semantic encoding is performed on the modified intent map, generating an intent feature map sequence that matches the size of the feature maps at each scale of the medical image through 1×1 convolution and upsampling operations. Finally, the medical image feature map and the intent feature map at each scale are fused element-wise, and then the fused multi-scale feature maps are concatenated along the channel dimension to form a multi-scale conditional tensor.

[0049] Specifically, in step S52, the multi-scale conditional tensor undergoes guided iterative denoising to obtain the predicted soft mask. It should be understood that since the core mechanism of generative models is to gradually transform random noise into a conditional target output through iterative denoising, and the multi-scale conditional tensor contains the layered features of the correction intention and medical image, these features cannot be transformed into a mask in the form of continuous probabilities without an iterative denoising process. Therefore, this application further inputs the multi-scale conditional tensor into a denoising network and performs multi-step iterative denoising operations to gradually optimize the noise tensor into a soft mask with pixel-level probabilities. In this way, the gradual optimization characteristics of the denoising process can be utilized to refine the mask's boundaries and regional integrity in each iteration. The resulting soft mask (with pixel values ​​of continuous probabilities from 0 to 1) retains the precision of the generative model while providing a flexible threshold adjustment space for subsequent binarization, improving the mask's quality and adjustability.

[0050] Specifically, in one possible embodiment, step S52 is implemented as follows: First, a random noise tensor with the same size as the medical image is initialized as the initial input for the backdiffusion process. Second, a pre-trained denoising U-Net network is loaded, which has been fully learned during the training phase to perform mask repair based on the user's correction intent (i.e., the multi-scale conditional tensor). In implementation, the system starts from the final time step and iteratively performs denoising for a preset number of steps (e.g., 50 steps). At each time step, a standard backdiffusion single-step calculation formula is used: this formula uses the denoising U-Net network to predict the global noise distribution based on the current noisy mask and the multi-scale conditional tensor as conditions, then removes it from the noisy mask, and injects a small amount of new random noise to calculate a clearer mask for the next time step. This standard formula is used to perform a uniform denoising operation on all pixels until a predicted soft mask is finally generated completely from the random noise.

[0051] In particular, in another preferred embodiment, a preset noise scheduling constant is used in the standard back-diffusion single-step calculation formula. and It is global, meaning that regardless of whether a pixel is located in the user-specified anchor region (i.e., the region with a pixel value of 1 in the anchor channel graph), it will be affected by the same intensity of new noise injection and the same intensity of model prediction noise cancellation during the denoising process. However, because the anchor channel graph represents highly reliable regions indirectly confirmed by the user without any operation, it is desirable for the model to preserve these regions to the greatest extent possible during the generation process, reducing unnecessary changes and noise intervention. Global noise scheduling obviously cannot effectively distinguish these high-confidence regions from other uncertain regions, which may lead to subtle jitter or morphological changes in the anchor region.

[0052] Therefore, it is necessary to introduce spatial adaptive noise scheduling for reverse diffusion denoising. This involves dynamically adjusting the denoising intensity and noise injection amount during the reverse diffusion process based on whether a pixel belongs to the anchor point channel map. Specifically, for pixels in anchor point regions, random noise injection is reduced or eliminated, and the impact of model prediction noise on their correction is decreased to maintain consistency with the initial mask. For pixels in non-anchor point regions, standard denoising and noise injection processes are maintained, allowing the generative model to fully utilize its generative capabilities in these uncertain regions and explore new forms that conform to the image and user intent.

[0053] Therefore, the single-step calculation formula for reverse diffusion is improved as follows:

[0054] in, Is Time step, pixel The denoised mask value, Is Time step, pixel The noisy mask value, It is a preset noise scheduling constant, representing the proportion of signal retained at each time step. It is a deep learning model Predicted in Add time step to The pixel values ​​of the noisy pixels are generated by a specially trained denoising UNet network, which takes the noisy mask of the current time step and the user's correction intention as input. It is conditional information, namely, the multi-scale conditional tensor. It is a noise scheduling constant, representing the noise scheduling constant. The intensity of random noise injected during the time-step reverse diffusion process Random noise sampled from the standard normal distribution Pixel value at that location, It is the first Image samples in The anchor point channel map value at the pixel; a value of 1 indicates that the pixel is an anchor point, and a value of 0 indicates that it is not. These are hyperparameters, typically ranging from [0,1], which control the anchor point region. Predict the degree of noise suppression, when hour, Factor will scale The predicted noise contribution, when When the prediction noise contribution is completely eliminated, the anchor point region is completely unaffected by the model prediction. When the predicted noise contribution remains unchanged, the formula is in standard form. This is another hyperparameter, typically ranging from [0,1], which controls the degree of suppression of random noise injection in the anchor point region. hour, The factor will scale random noise When the injection, When random noise injection is completely eliminated, the anchor point region will not introduce any new randomness during the denoising process. When random noise injection remains unchanged, the formula is in standard form.

[0055] That is, for the anchor point area , Multiply Reduce the prediction noise and its contribution. Multiply To reduce random noise injection, by adjusting and (For example, setting them all to 1) can keep the anchor point region highly stable during the denoising process, with almost no morphological changes. However, for non-anchor point regions... 1 multiplied by The predicted noise is reduced so that its contribution remains constant, and multiplied by 1. This keeps the random noise injection unchanged, allowing the generative model to freely explore and generate reasonable segmentation results in these user-undefined regions.

[0056] In this way, the pixels in the anchor region will maintain extremely high stability during the denoising iterations. The generated mask will be highly faithful to the areas that the user did not correct in the initial mask. This means that users can confidently focus their attention on those blurry or erroneous areas that truly need repair without worrying about unnecessary drift in the already identified areas. Simultaneously, by locking the anchor region, the generative model can more effectively concentrate its generative capabilities and attention on repairing non-anchor regions. The model does not need to waste computational resources and generative potential on already determined areas, which helps the model generate more reasonable, anatomically consistent, and smoothly contoured masks in uncertain areas, while also accelerating the convergence speed of the denoising process.

[0057] Specifically, in step S53, the predicted soft mask is subjected to proposal binarization and multiple scheme collection to obtain a selected mask candidate set. It should be understood that since the soft mask is a continuous probability value (0-1), it cannot be directly used as the final mask. Furthermore, different physicians have preferences for the boundary thresholds between lesions and the background. Outputting only a single threshold binarization result cannot meet diverse clinical needs. Therefore, this application further performs multi-threshold binarization on the soft mask and collects binary masks under different thresholds to generate multiple candidate schemes. This provides mask options covering different boundary judgment styles, allowing physicians to select the optimal scheme based on the actual lesion's morphology, density, and other characteristics. Simultaneously, by comparing multiple schemes, annotation errors caused by single threshold deviations are reduced, improving the clinical applicability of the final mask.

[0058] Specifically, in one possible embodiment, step S53 is implemented as follows: First, five candidate thresholds are determined, such as 0.3, 0.4, 0.5, 0.6, and 0.7. Second, the predicted soft mask is binarized for each threshold; a pixel value ≥ the threshold is set to 1, otherwise it is set to 0, resulting in five binary masks. Finally, morphological post-processing, such as opening a 3×3 kernel, is performed on each binary mask to eliminate minor noise and holes. The five processed binary masks are then combined to form a selected candidate mask set.

[0059] In the aforementioned AI-based intelligent medical image annotation method, step S6 involves the user selecting the most satisfactory candidate mask from the selected mask candidate set as the final mask. It should be understood that since the selected mask candidate set contains multiple mask schemes based on different generation logics, each scheme differs in lesion boundary refinement and regional integrity. Furthermore, determining the optimal mask in a clinical setting requires combining the physician's subjective clinical experience regarding lesion morphology and anatomical relationships. Generative models cannot completely replace the human final decision on clinical applicability. Therefore, this application further provides an interactive interface for physicians to evaluate and select candidate masks one by one, allowing physicians to determine the mask that best matches objective anatomy and clinical judgment based on actual clinical needs (such as the clinical significance of the lesion and the adaptability to subsequent diagnostic procedures). This ensures that the final mask fully meets the actual needs of the clinical scenario, avoids deviations between the generative repair results and the physician's expectations, improves the clinical credibility of medical image annotation, and enables it to be directly used for subsequent diagnostic support, research analysis, and other stages.

[0060] Specifically, in one possible embodiment, step S6 is implemented as follows: First, each candidate mask in the selected mask candidate set is sequentially overlaid and displayed with 50% transparency in the same view area of ​​the original medical image. Simultaneously, a thumbnail of each candidate and its generation parameters, such as the binarization threshold, are displayed in the sidebar of the interface. Second, the physician triggers a selection operation by clicking on the thumbnail or the candidate mask area within the interface, and the system highlights the selected candidate in real time. Finally, after the physician confirms the selection, the system determines the candidate mask as the final mask, automatically stores it in the medical image annotation database, associates it with patient information, image modality, and other metadata, and generates an annotation confirmation log, completing the annotation process loop.

[0061] In summary, the AI-based intelligent medical image annotation method based on the embodiments of this application is explained. It captures and analyzes the user's concise interactive strokes, transforming them into structured guiding information containing corrective intent. Simultaneously, this high-level intent is fused with the deep visual features of the original medical image to jointly drive a conditionally guided generative repair model. This model exploratoryly regenerates in the defective area based on the user's intent and proactively provides a candidate set containing multiple high-quality correction schemes conforming to anatomical logic. Ultimately, the user's interaction mode shifts from tedious pixel-level refinement to efficient selection and confirmation of high-quality candidates provided by the system. This achieves an intelligent collaborative closed loop from user corrective intent to accurate mask generation, thereby improving the efficiency and interactive experience of intelligent medical image annotation.

[0062] Figure 7 This is a block diagram of an AI-based intelligent medical image annotation system according to an embodiment of this application. Figure 7As shown, the AI-based intelligent medical image annotation system 100 according to an embodiment of this application includes: a medical image acquisition module 110 for acquiring medical images; a medical image pre-annotation module 120 for performing AI-based pre-annotation on the medical images to obtain an initial mask; an initial mask overlay module 130 for overlaying the initial mask on the medical images in a semi-transparent form and receiving user interaction; a stroke rasterization module 140 for performing stroke rasterization on the user interaction to obtain a correction intent map; a generative mask repair module 150 for performing condition-guided generative mask repair on the medical images based on the correction intent map to obtain a selected mask candidate set; and a mask confirmation and output module 160 for selecting the most satisfactory candidate mask from the selected mask candidate set as the final mask.

[0063] As described above, the AI-based medical image intelligent annotation system 100 according to the embodiments of this application can be implemented in various wireless terminals, such as servers with AI-based medical image intelligent annotation algorithms. In one possible implementation, the AI-based medical image intelligent annotation system 100 according to the embodiments of this application can be integrated into the wireless terminal as a software module and / or hardware module. For example, the AI-based medical image intelligent annotation system 100 can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the AI-based medical image intelligent annotation system 100 can also be one of many hardware modules of the wireless terminal.

[0064] Alternatively, in another example, the AI-based medical image intelligent annotation system 100 and the wireless terminal can also be separate devices, and the AI-based medical image intelligent annotation system 100 can connect to the wireless terminal via wired and / or wireless networks and transmit interactive information in accordance with an agreed data format.

[0065] Here, those skilled in the art will understand that the specific operations of each step in the above-described AI-based intelligent medical image annotation system have been referenced above. Figures 1 to 6 The AI-based intelligent annotation method for medical images has been described in detail, and therefore, its repeated description will be omitted.

Claims

1. An AI-based intelligent annotation method for medical images, characterized in that, include: Acquiring medical images; Medical images are pre-annotated using an AI model to obtain an initial mask; The initial mask is overlaid on the medical image in a semi-transparent manner and user interaction is received. The user interaction is rasterized using strokes to obtain a revised intent map; Condition-guided generative mask repair of medical images based on modified intent maps is used to obtain a selected set of mask candidates. The user selects the most satisfactory candidate mask from the selected mask candidate set as the final mask.

2. The AI-based intelligent annotation method for medical images according to claim 1, characterized in that, Medical images are pre-annotated using an AI model to obtain an initial mask, including: Medical images are subjected to image data standardization preprocessing to obtain preprocessed medical images; The pre-processed medical images are input into a pre-trained deep learning segmentation network to obtain a probability map; An initial mask is generated based on the probabilistic graph.

3. The AI-based intelligent annotation method for medical images according to claim 2, characterized in that, Based on the probabilistic graph, an initial mask is generated, including: generating the initial mask using the following formula, wherein the formula is: ; in, For the first Image samples in Predicted probability map at each pixel. The preset threshold, For the first Image samples in The initial binary mask at the pixel.

4. The AI-based intelligent annotation method for medical images according to claim 1, characterized in that, The user interaction is rasterized using strokes to obtain a revised intent map, including: Extract the positive and negative correction stroke sets from user interactions; Intended stroke rasterization and channel generation are performed on the positive correction stroke set and the negative correction stroke set to obtain the positive correction channel map and the negative correction channel map; Anchor point regions are calculated and three-channel intent maps are synthesized from the initial mask, positive correction channel map, and negative correction channel map to obtain the corrected intent map.

5. The AI-based intelligent annotation method for medical images according to claim 4, characterized in that, Anchor point regions are calculated and three-channel intent maps are synthesized from the initial mask, positive correction channel map, and negative correction channel map to obtain the corrected intent map, including: An interactive union mask is generated based on the positive and negative correction channel graphs. Stable anchor point channel calculations are performed on the initial mask and the interactive union mask to obtain the anchor point channel map; Multi-channel correction intention maps are synthesized from the positive correction channel map, negative correction channel map, and anchor point channel map to obtain the correction intention map.

6. The AI-based intelligent annotation method for medical images according to claim 5, characterized in that, To obtain an anchor channel map, stable anchor channel calculation is performed on the initial mask and the interactive union mask, including: calculating stable anchor channels on the initial mask and the interactive union mask using the following formula: ; in, For the first Image samples in Negative correction channel map at the pixel. For the first Image samples in Positive correction channel map at the pixel For the first Image samples in Initial mask at pixel, Logical NOT, For logical AND, For logical OR, For the first Image samples in Anchor point channel map value at the pixel.

7. The AI-based intelligent annotation method for medical images according to claim 1, characterized in that, Condition-guided generative masking of medical images based on modified intent maps yields a refined set of mask candidates, including: Multimodal conditional encoding and fusion of modified intent maps and medical images are performed to obtain multiscale conditional tensors; Guided iterative denoising is performed on the multi-scale conditional tensor to obtain the predicted soft mask; The predicted soft mask is subjected to proposal binarization and multiple scheme collection to obtain a refined mask candidate set.

8. The AI-based intelligent annotation method for medical images according to claim 7, characterized in that, Guided iterative denoising of the multi-scale conditional tensor to obtain the predicted soft mask includes: Inverse diffusion denoising is performed using spatial adaptive noise scheduling, and the single-step calculation formula for inverse diffusion is as follows: ; in, Is Time step, pixel The denoised mask value, Is Time step, pixel The noisy mask value, It is a preset noise scheduling constant, representing the proportion of signal retained at each time step. and These are different hyperparameters. It is a deep learning model Predicted in Add time step to The pixel values ​​of noise on the surface, where It is conditional information, namely, the multi-scale conditional tensor. It is a noise scheduling constant, representing the noise scheduling constant. The intensity of random noise injected during the time-step reverse diffusion process Random noise sampled from the standard normal distribution Pixel value at that location, It is the first Image samples in The anchor point channel map value at the pixel; a value of 1 indicates that the pixel is an anchor point, and a value of 0 indicates that it is not.

9. An AI-based intelligent annotation system for medical images, characterized in that, include: The medical image acquisition module is used to acquire medical images; The medical image pre-annotation module is used to pre-annotate medical images based on AI models to obtain initial masks; The initial mask overlay module is used to overlay the initial mask onto the medical image in a semi-transparent form and receive user interaction. The stroke rasterization module is used to rasterize user interactions with strokes to obtain a revised intent map; The generative mask repair module is used to perform condition-guided generative mask repair on medical images based on the modified intent map to obtain a selected set of mask candidates. The mask confirmation and output module is used to select the most satisfactory candidate mask from the selected mask candidate set as the final mask.