Automatic labeling method and system for medical image or video based on deep learning
Through the automatic labeling method based on deep learning, medical images or videos are labeled, which solves the problem of low labeling efficiency and accuracy caused by uneven density distribution, and achieves a fast and accurate labeling effect.
Patent Information
- Application Number
- CN202510124834.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-06-03
AI Technical Summary
In the prior art, when labeling medical images or videos, uneven density distribution leads to low labeling efficiency and accuracy.
The automatic labeling method based on deep learning is adopted. After obtaining medical images or videos and preprocessing them, they are input into the preset automatic labeling model. Components such as backbone network, dual intercorrelation fusion network and anchor-free position prediction head are used to automatically label objects, regions and behaviors in images or videos.
It realizes fast and accurate labeling of medical images or videos, improves labeling efficiency and accuracy, and is suitable for many types of medical images and videos.
Smart Images

Figure CN120088700A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data annotation, and particularly relates to an automatic annotation method, system, device and computer-readable storage medium for medical images or videos based on deep learning. Background Art
[0002] With the continuous progress of surgical robot technology, artificial intelligence and machine learning have played an increasingly important role in the research and development of surgical robots. Among them, medical image and video annotation, as a key link in training the surgical robot AI system, provides strong support for the efficient operation of surgical robots in complex environments.
[0003] Image and video annotation is an important step in providing understandable data for AI models. By annotating objects, regions and behaviors in images, robots can learn to identify key elements in the environment and make accurate responses.
[0004] Advanced video annotation techniques can create powerful robotic process automation datasets. These datasets significantly enhance the ability of robotic systems to perceive and interact with the environment.
[0005] Surgical robots are an important achievement of technological innovation in the medical field, and their precision and intelligence are inseparable from the support of high-quality image annotation data.
[0006] Currently, a related technology discloses an automatic annotation method applicable to multiple types of pathological images. This method uses a density-based clustering algorithm to classify and identify the blocks of pathological images using a trained deep learning model to obtain classification results and the position coordinates of each classification.
[0007] If the density distribution is uneven, the annotation of pathological images by the clustering algorithm will be inefficient and inaccurate.
[0008] Therefore, how to quickly and accurately annotate medical images or videos is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0009] An embodiment of this application provides an automatic annotation method, system, device and computer-readable storage medium for medical images or videos based on deep learning, which can quickly and accurately annotate medical images or videos.
[0010] In a first aspect, an embodiment of this application provides an automatic annotation method for medical images or videos based on deep learning, including:
[0011] Obtain a medical image or video;
[0012] Input the medical image or video into a preset automatic annotation model and output an automatic annotation result;
[0013] Among them, the automatic annotation results include: annotating objects, regions, and behaviors in medical images or videos.
[0014] Optionally, obtaining medical images includes:
[0015] The obtained medical images include: X-ray images, CT images, MR images, and ultrasound images;
[0016] Performing preprocessing on the medical images;
[0017] Among them, the preprocessing includes: denoising, image normalization, and image enhancement.
[0018] Optionally, performing denoising, image normalization, and image enhancement on the medical images includes:
[0019] Using Gaussian filtering, arithmetic mean filtering, and median filtering to denoise the medical images;
[0020] Performing image enhancement by adjusting the contrast of the denoised images;
[0021] Normalizing the pixel values of the medically enhanced images to a preset range.
[0022] Optionally, inputting a medical video into a preset automatic annotation model to output automatic annotation results, including:
[0023] Adding an initial target box to the first frame image of the medical video to calibrate the initial position of the target, using this first frame image as a template image, and using the other frame images of the medical video as search images;
[0024] Performing preprocessing on the template image and the search images to obtain preprocessed template images and search images;
[0025] Inputting the preprocessed template image and search images into a preset automatic annotation model to obtain the predicted position of the target in a certain frame image of the medical video;
[0026] Based on the predicted position of the target in a certain frame image of the medical video, achieving target tracking and automatic annotation;
[0027] Among them, the automatic annotation model includes: a backbone network, a dual cross-correlation fusion network, and an anchor-free position prediction head.
[0028] Optionally, it includes:
[0029] The backbone network is used to extract features from the preprocessed template image to obtain a template image feature map, and to extract features from the preprocessed search images to obtain search image feature maps;
[0030] A dual cross - correlation fusion network is used to match and fuse the pixel similarity of the template image feature map and the search image feature map globally to obtain the final classification map and regression map;
[0031] An anchor - free location prediction head is used to obtain the predicted location of the target in a certain frame of the medical video according to the final classification map and regression map.
[0032] Optionally, the dual cross - correlation fusion network includes:
[0033] A depth cross - correlation module, a pixel - by - pixel cross - correlation module, a spatial and channel fusion module, and a reverse activation module;
[0034] The template image feature map and the search image feature map are input into the pixel - by - pixel cross - correlation module, and after pixel - by - pixel correlation operation, a first feature map is obtained;
[0035] The search image feature map and the cropped template image feature map are input into the depth cross - correlation module, and after depth cross - correlation operation, a second feature map is obtained;
[0036] The second feature map is input into the spatial and channel fusion module to obtain a third feature map, and the third feature map is input into the reverse activation module to obtain a fourth feature map;
[0037] The first feature map passes through the SE layer and the 1×1 convolutional layer in sequence and then combines with the fourth feature map as the output of the dual cross - correlation fusion network.
[0038] Optionally, it includes:
[0039] The spatial and channel fusion module is composed of a 3×3 Conv - BN branch and a 1×1 Conv - BN branch;
[0040] The reverse activation module includes, in sequence, a first 1×1 Conv - BN branch, a non - linear Gelu activation function layer, and a second 1×1 Conv - BN branch.
[0041] In a second aspect, an embodiment of the present application provides an automatic annotation system for medical images or videos based on deep learning, including:
[0042] A medical data acquisition module is used to acquire medical images or videos;
[0043] An automatic annotation module is used to input the medical images or videos into a preset automatic annotation model and output an automatic annotation result;
[0044] Among them, the automatic annotation result includes: annotating objects, regions, and behaviors in the medical images or videos.
[0045] In a third aspect, an embodiment of the present application provides an electronic device, which includes: a processor and a memory storing computer program instructions;
[0046] When the processor executes the computer program instructions, an automatic annotation method for medical images or videos based on deep learning is implemented.
[0047] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, an automatic annotation method for medical images or videos based on deep learning is implemented.
[0048] The automatic annotation method, system, device, and computer-readable storage medium for medical images or videos based on deep learning according to the embodiments of the present application can annotate medical images or videos quickly and accurately.
[0049] The automatic annotation method for medical images or videos based on deep learning includes:
[0050] Obtain medical images or videos;
[0051] Input the medical images or videos into a preset automatic annotation model to output an automatic annotation result;
[0052] Among them, the automatic annotation result includes: annotating objects, regions, and behaviors in the medical images or videos. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0054] Figure 1 is a flowchart of an automatic annotation method for medical images or videos based on deep learning provided by an embodiment of the present application;
[0055] Figure 2 is a schematic structural diagram of an automatic annotation model provided by an embodiment of the present application;
[0056] Figure 3 is a schematic structural diagram of a spatial and channel fusion module provided by an embodiment of the present application;
[0057] Figure 4 is a schematic structural diagram of a reverse activation module provided by an embodiment of the present application;
[0058] Figure 5 It is a schematic structural diagram of an automatic annotation system for medical images or videos based on deep learning provided by an embodiment of the present application;
[0059] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0060] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of the present application by showing examples of the present application.
[0061] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "including..." do not exclude the presence of additional identical elements in the process, method, article or device including the said elements.
[0062] To solve the problems of the prior art, embodiments of the present application provide an automatic annotation method, system, device and computer-readable storage medium for medical images or videos based on deep learning. First, the automatic annotation method for medical images or videos based on deep learning provided by the embodiments of the present application will be introduced below.
[0063] Figure 1 It shows a schematic flow diagram of an automatic annotation method for medical images or videos based on deep learning provided by an embodiment of the present application. As Figure 1 shown, the automatic annotation method for medical images or videos based on deep learning includes:
[0064] S101. Obtain medical images or videos;
[0065] S102. Input a medical image or video into a preset automatic annotation model to output an automatic annotation result. The automatic annotation result includes annotating objects, regions, and behaviors in the medical image or video.
[0066] In one embodiment, obtaining a medical image includes:
[0067] The obtained medical image includes: X-ray image, CT image, MR image, ultrasound image;
[0068] Preprocess the medical image;
[0069] Among them, the preprocessing includes: denoising, image normalization, and image enhancement.
[0070] In one embodiment, denoising, image normalization, and image enhancement of the medical image include:
[0071] Use Gaussian filtering, arithmetic mean filtering, and median filtering to denoise the medical image;
[0072] Perform image enhancement by adjusting the contrast of the denoised image;
[0073] Normalize the pixel values of the medical image after image enhancement to a preset range.
[0074] In one embodiment, inputting a medical video into a preset automatic annotation model to output an automatic annotation result includes:
[0075] Add an initial target box to the first frame image of the medical video to calibrate the initial position of the target, use this first frame image as the template image, and use the other frame images of the medical video as search images;
[0076] Preprocess the template image and the search images to obtain the preprocessed template image and search images;
[0077] Input the preprocessed template image and search images into a preset automatic annotation model to obtain the predicted position of the target in a certain frame image of the medical video;
[0078] Based on the predicted position of the target in a certain frame image of the medical video, implement target tracking and automatic annotation;
[0079] Among them, the automatic annotation model includes: a backbone network, a dual cross-correlation fusion network, and an anchor-free position prediction head.
[0080] Figure 2 It is a schematic structural diagram of the automatic annotation model provided by an embodiment of the present application.
[0081] In one embodiment, the automatic annotation model includes:
[0082] The backbone network is used to extract features from the preprocessed template image to obtain the template image feature map, and to extract features from the preprocessed search image to obtain the search image feature map;
[0083] The dual cross-correlation fusion network is used to match and fuse the pixel similarity between the template image feature map and the search image feature map globally to obtain the final classification map and regression map;
[0084] The anchor-free position prediction head is used to obtain the predicted position of the target in a certain frame image of the medical video according to the final classification map and regression map.
[0085] In one embodiment, the dual cross-correlation fusion network includes:
[0086] The depth cross-correlation module, the per-pixel cross-correlation module, the spatial and channel fusion module, and the reverse activation module;
[0087] The template image feature map and the search image feature map are input into the per-pixel cross-correlation module, and after the per-pixel correlation operation, the first feature map is obtained;
[0088] The search image feature map and the cropped template image feature map are input into the depth cross-correlation module, and after the depth cross-correlation operation, the second feature map is obtained;
[0089] The second feature map is input into the spatial and channel fusion module to obtain the third feature map, and the third feature map is input into the reverse activation module to obtain the fourth feature map;
[0090] The first feature map passes through the SE layer and the 1×1 convolutional layer in sequence and then combines with the fourth feature map as the output of the dual cross-correlation fusion network.
[0091] Figure 3 It is the structural schematic diagram of the spatial and channel fusion module provided by an embodiment of the present application;
[0092] Figure 4 It is the structural schematic diagram of the reverse activation module provided by an embodiment of the present application.
[0093] In one embodiment, it includes:
[0094] The spatial and channel fusion module is composed of a 3×3 Conv-BN branch and a 1×1 Conv-BN branch;
[0095] The reverse activation module sequentially includes the first 1×1 Conv-BN branch, the non-linear Gelu activation function layer, and the second 1×1 Conv-BN branch.
[0096] Figure 5It is a schematic structural diagram of an automatic annotation system for medical images or videos based on deep learning provided by an embodiment of the present application.
[0097] The automatic annotation system for medical images or videos based on deep learning includes:
[0098] A medical data acquisition module 501, configured to acquire medical images or videos;
[0099] An automatic annotation module 502, configured to input the medical images or videos into a preset automatic annotation model and output an automatic annotation result;
[0100] Wherein, the automatic annotation result includes: annotating objects, regions and behaviors in the medical images or videos.
[0101] Figure 6 It shows a schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0102] The electronic device may include a processor 601 and a memory 602 storing computer program instructions.
[0103] Specifically, the above-mentioned processor 601 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0104] The memory 602 may include a mass storage for data or instructions. By way of example and not limitation, the memory 602 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 602 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 602 may be inside or outside the electronic device. In a specific embodiment, the memory 602 may be a non-volatile solid state memory.
[0105] In one embodiment, the memory 602 may be a read only memory (ROM). In one embodiment, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.
[0106] The processor 601 reads and executes the computer program instructions stored in the memory 602 to implement any one of the automatic annotation methods for medical images or videos based on deep learning in the above embodiments.
[0107] In one example, the electronic device may further include a communication interface 603 and a bus 610. Among them, as Figure 6 shown, the processor 601, the memory 602, and the communication interface 603 are connected through the bus 610 and complete communication with each other.
[0108] The communication interface 603 is mainly used to implement communication between various modules, systems, units, and / or devices in the embodiments of the present application.
[0109] The bus 610 includes hardware, software, or both, and couples the components of the electronic device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. In a suitable case, the bus 610 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.
[0110] In addition, in combination with the automatic annotation method for medical images or videos based on deep learning in the above embodiments, the embodiments of the present application may provide a computer-readable storage medium to implement. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, any one of the automatic annotation methods for medical images or videos based on deep learning in the above embodiments is implemented.
[0111] It should be clear that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0112] The functional modules shown in the above-described structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments for performing the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. A "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.
[0113] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or systems. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, can be different from the order in the embodiments, or several steps can be executed simultaneously.
[0114] Aspects of the present application have been described above with reference to the flowcharts and / or block diagrams of methods, systems (systems), and computer program products according to embodiments of the present application. It should be understood that each block in the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing system to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing system enable the implementation of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It should also be understood that each block in the block diagrams and / or flowcharts, and the combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0115] The above are only specific embodiments of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.
Claims
1. A method for automatic annotation of medical images or videos based on deep learning, characterized in that: include: Acquiring medical images or videos; Input the medical image or video into the preset automatic annotation model and output the automatic annotation result; Among them, the automatic labeling results include: labeling objects, areas and behaviors in medical images or videos.
2. The method for automatic annotation of medical images or videos based on deep learning according to claim 1, characterized in that: Acquisition of medical images, including: The acquired medical images include: X-ray images, CT images, MR images, and ultrasound images; Preprocess medical images; Among them, preprocessing includes: denoising, image normalization, and image enhancement.
3. The method for automatic annotation of medical images or videos based on deep learning according to claim 2, characterized in that: Denoising, image normalization, and image enhancement for medical images, including: Use Gaussian filtering, arithmetic mean filtering and median filtering to denoise medical images; Image enhancement is performed by adjusting the contrast of the denoised image; The pixel values of the medical image after image enhancement are normalized to a preset range.
4. The method for automatic annotation of medical images or videos based on deep learning according to claim 1, characterized in that: Input the medical video into the preset automatic annotation model and output the automatic annotation results, including: Adding an initial target frame to the first frame image of the medical video to calibrate the initial position of the target, using the first frame image as a template image, and using other frame images of the medical video as search images; Preprocessing the template image and the search image to obtain a preprocessed template image and a search image; The preprocessed template image and the search image are input into a preset automatic annotation model to obtain the predicted position of the target in a certain frame of the medical video; Based on the predicted position of the target in a certain frame of the medical video, the target is tracked and automatically labeled; Among them, the automatic labeling model includes: backbone network, dual cross-correlation fusion network and anchor-free position prediction head.
5. The method for automatic annotation of medical images or videos based on deep learning according to claim 4, characterized in that: Automatic annotation model, including: A backbone network is used to extract features from the preprocessed template image to obtain a template image feature map, and to extract features from the preprocessed search image to obtain a search image feature map; The dual cross-correlation fusion network is used to match and fuse the global pixel similarity of the template image feature map and the search image feature map to obtain the final classification map and regression map; The anchor-free position prediction head is used to obtain the predicted position of the target in a certain frame image of the medical video according to the final classification map and regression map.
6. The method for automatic annotation of medical images or videos based on deep learning according to claim 5, characterized in that: Dual cross-correlation fusion network, including: Depth cross-correlation module, pixel-by-pixel cross-correlation module, spatial and channel fusion module, and inverse activation module; The template image feature map and the search image feature map are input into a pixel-by-pixel cross-correlation module, and a first feature map is obtained after a pixel-by-pixel correlation operation; Input the search image feature map and the cropped template image feature map into the depth cross-correlation module, and obtain the second feature map after the depth cross-correlation operation; The second feature map is input into the space and channel fusion module to obtain a third feature map, and the third feature map is input into the inverse activation module to obtain a fourth feature map; The first feature map is sequentially passed through the SE layer and the 1×1 convolution layer and combined with the fourth feature map as the output of the dual cross-correlation fusion network.
7. The method for automatic annotation of medical images or videos based on deep learning according to claim 6, characterized in that: include: The spatial and channel fusion module consists of a 3×3 Conv-BN branch and a 1×1 Conv-BN branch; The reverse activation module includes a first 1×1 Conv-BN branch, a nonlinear Gelu activation function layer and a second 1×1 Conv-BN branch in sequence.
8. An automatic annotation system for medical images or videos based on deep learning, characterized in that: The system comprises: A medical data acquisition module, used to acquire medical images or videos; An automatic annotation module, used to input medical images or videos into a preset automatic annotation model and output automatic annotation results; Among them, the automatic labeling results include: labeling objects, areas and behaviors in medical images or videos.
9. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the automatic annotation method for medical images or videos based on deep learning as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by the processor, the method for automatic annotation of medical images or videos based on deep learning as described in any one of claims 1 to 7 is implemented.