Image Segmentation Data Annotation Method and Device

By acquiring and matching image sets and generating labels using specified backgrounds, the problem of low labeling efficiency in image segmentation is solved, and efficient and accurate automatic labeling is achieved.

CN114972881BActive Publication Date: 2025-05-30SHANGHAI MICROPORT MEDBOT (GRP) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210680078.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-16
Publication Date
2025-05-30
Estimated Expiration
2042-06-16

AI Technical Summary

Technical Problem

In the prior art, the image labeling efficiency required for image segmentation is low, resulting in high labor costs and dependent on labor on labeling accuracy.

Method used

By obtaining the original image set and the specified image set, and generating labels with the specified background image, matching the original image with the specified image for automatic annotation, improving data labeling efficiency and accuracy.

Benefits of technology

The automated image annotation process is realized, which significantly improves the efficiency and accuracy of data annotation, reduces manual intervention and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972881B_ABST
    Figure CN114972881B_ABST
Patent Text Reader

Abstract

This specification relates to the technical field of deep learning image segmentation, and specifically discloses an image segmentation data annotation method and apparatus. Among them, the method includes: obtaining an original image set, a specified image set, and a specified background image; generating first labels corresponding to each specified image in the multiple specified images based on the specified image set and the specified background image; the first labels are used to annotate the target objects in the specified images; matching the multiple original images with the multiple specified images according to the first positions corresponding to each original image in the multiple original images and the second positions corresponding to each specified image in the multiple specified images; and annotating the multiple original images based on the matching results and the first labels corresponding to each specified image in the multiple specified images to obtain multiple annotated original images. The above method realizes automatic data annotation for image segmentation, effectively improving the data annotation efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of deep learning image segmentation technology, and in particular to a method and device for labeling image segmentation data. Background Art

[0002] In the field of deep learning image segmentation, surgical instrument annotation in endoscopic images is mainly done by manual labeling. Currently, manual labeling is mainly done through open source tools (e.g., Labelme), which requires manual step-by-step labeling of the image to be labeled, similar to cutting out the image. Figure 1 For example, for a 2k image, finely labeling a surgical instrument image takes about 10 minutes. This is because deep learning image segmentation requires high-quality, high-precision data annotation. Manual annotation is labor-intensive, resulting in very low annotation efficiency and high labor costs.

[0003] To address the above issues, no effective solutions have been proposed so far. Summary of the Invention

[0004] The embodiments of this specification provide a method and apparatus for annotating image segmentation data to solve the problem of low efficiency of image annotation required for image segmentation in the prior art.

[0005] An embodiment of the present specification provides an image segmentation data annotation method, including: obtaining an original image set, a specified image set and a specified background image; the original image set includes a plurality of original images captured when the target object is in a plurality of first positions in a target scene, and the specified image set includes a plurality of specified images captured when the target object is in a plurality of second positions under a specified background; the plurality of first positions and the plurality of second positions have an intersection; based on the specified image set and the specified background image, generating a first label corresponding to each specified image in the plurality of specified images; the first label is used to annotate the target object in the specified image; matching the plurality of original images with the plurality of specified images according to the first position corresponding to each original image in the plurality of original images and the second position corresponding to each specified image in the plurality of specified images; annotating the plurality of original images based on the matching results and the first label corresponding to each specified image in the plurality of specified images to obtain a plurality of annotated original images.

[0006] An embodiment of the present specification provides an image segmentation data annotation method, including: acquiring an original image set, a designated image set, and a designated background image; the original image set includes a plurality of original images captured when the surgical instrument is in a plurality of first positions in an endoscopic surgical environment, and the designated image set includes a plurality of designated images captured when the surgical instrument is in a plurality of second positions under a designated background; the plurality of first positions partially overlap or completely overlap with the plurality of second positions; based on the designated image set and the designated background image, generating a first label corresponding to each designated image in the plurality of designated images; the first label is used to annotate the surgical instrument in the designated image; matching the plurality of original images with the plurality of designated images to determine a second label corresponding to each original image in the plurality of original images; the second label is used to annotate the surgical instrument in the original image.

[0007] In one embodiment, the target scene includes an endoscopic surgery environment, and the target object includes a surgical instrument for performing endoscopic surgery; or, the target scene includes an extracorporeal automated surgery environment, and the target object includes a surgical instrument for performing extracorporeal automated surgery.

[0008] In one embodiment, the endoscopic surgery environment includes at least one of the following: an endoscopic surgery environment under lighting conditions, an endoscopic surgery environment under smoke conditions, an endoscopic surgery environment under shadow conditions, and an endoscopic surgery environment under bleeding scenes.

[0009] In one embodiment, based on the designated image set and the designated background image, a first label corresponding to each designated image in the multiple designated images is generated, including: calculating the difference between the pixel value of each designated image in the multiple designated images and the pixel value of the designated background image to obtain the difference image corresponding to each designated image; and generating the first label corresponding to each designated image in the multiple designated images based on the difference image corresponding to each designated image.

[0010] In one embodiment, based on the difference image, a first label corresponding to each specified image in the multiple specified images is generated, including: regularizing the differences corresponding to the specified images to obtain the difference images corresponding to the specified images after regularization; converting the specified images after regularization into binary grayscale images to obtain the first labels corresponding to the specified images.

[0011] In one embodiment, the multiple original images are matched with the multiple designated images based on the first position corresponding to each original image in the multiple original images and the second position corresponding to each designated image in the multiple designated images, including: comparing the coordinate data of each first position in the multiple first positions with the coordinate data of each second position in the multiple second positions to determine a plurality of position pairs, each position pair in the multiple position pairs including a first position and a second position with the same coordinate data; and determining the original image corresponding to the first position in each position pair and the designated image corresponding to the second position as a successful match.

[0012] In one embodiment, the multiple original images are labeled based on the matching results and the first label corresponding to each specified image in the multiple specified images to obtain a plurality of labeled original images, including: determining the first label of the specified image matching the each original image as the second label corresponding to the each original image; and labeling the original image using the second label to obtain a plurality of labeled original images.

[0013] In one embodiment, after obtaining multiple annotated original images, it also includes: constructing a training sample set based on the multiple annotated original images; using the training sample set to train a preset image segmentation deep learning network to obtain a trained image segmentation deep learning network; the trained image segmentation deep learning network is used to identify the target object from the target image collected from the target scene.

[0014] In one embodiment, after obtaining multiple annotated original images, it also includes: constructing a verification sample set based on the multiple annotated original images; inputting the verification sample set into the trained image segmentation deep learning network to calculate the evaluation index of the trained image segmentation deep learning network; when the evaluation index meets the preset conditions, determining the trained image segmentation deep learning network as the target image segmentation deep learning network; wherein, the target image segmentation deep learning network is used to identify the target object from the target image collected from the target scene.

[0015] In one embodiment, after calculating the evaluation index of the trained image segmentation deep learning network, it also includes: when the evaluation index does not meet the preset conditions, optimizing the network structure of the image segmentation deep learning network; using the training sample set to train the optimized image segmentation deep learning network to obtain a newly trained image segmentation deep learning network.

[0016] In one embodiment, after calculating the evaluation index of the trained image segmentation deep learning network, it also includes: when the evaluation index does not meet the preset conditions, re-acquiring the original image set and multiple annotated original images corresponding to the original image set to reconstruct the training sample set; using the reconstructed training sample set to train the preset image segmentation deep learning network to obtain a newly trained image segmentation deep learning network.

[0017] In one embodiment, after the trained image segmentation deep learning network is determined as the target image segmentation deep learning network, it also includes: obtaining a target image captured in a target scene; inputting the target image into the target image segmentation deep learning network to identify a target object area from the target image; when a target object key point exists in the target object area, controlling the target object to perform a preset operation based on the position of the target object key point.

[0018] In one embodiment, the target object key points include a plurality of candidate target object key points; after identifying the target object region from the target image, the method further includes: filtering out target object key points that are not in the target object region.

[0019] In one embodiment, the designated background is a green screen.

[0020] An embodiment of the present specification also provides an image segmentation data annotation device, including: an acquisition module, used to acquire an original image set, a specified image set and a specified background image; the original image set includes multiple original images collected when the target object is in multiple first positions in the target scene, and the specified image set includes multiple specified images collected when the target object is in multiple second positions under the specified background; the multiple first positions and the multiple second positions have an intersection; a generation module, used to generate a first label corresponding to each specified image in the multiple specified images based on the specified image set and the specified background image; the first label is used to annotate the target object in the specified image; a matching module, used to match the multiple original images with the multiple specified images according to the first position corresponding to each original image in the multiple original images and the second position corresponding to each specified image in the multiple specified images; a labeling module, used to annotate the multiple original images based on the matching results and the first label corresponding to each specified image in the multiple specified images, to obtain multiple annotated original images.

[0021] An embodiment of this specification also provides a medical device, including a processor and a memory for storing processor-executable instructions, wherein when the processor executes the instructions, the steps of the image segmentation data labeling method described in any of the above embodiments are implemented.

[0022] An embodiment of this specification further provides a computer-readable storage medium having computer instructions stored thereon, which, when executed, implement the steps of the image segmentation data annotation method described in any of the above embodiments.

[0023] In an embodiment of the present specification, a method for labeling image segmentation data is provided, which can obtain an original image set, a specified image set and a specified background image, wherein the original image set includes multiple original images collected when the target object is in multiple first positions in a target scene, and the specified image set includes multiple specified images collected when the target object is in multiple second positions under a specified background; the multiple first positions and the multiple second positions have an intersection, and a first label corresponding to each specified image in the multiple specified images can be generated based on the specified image set and the specified background image, and the first label is used to label the target object in the specified image. The multiple original images can be matched with the multiple specified images according to the first position corresponding to each original image in the multiple original images and the second position corresponding to each specified image in the multiple specified images, and the multiple original images are labeled according to the matching results and the first label corresponding to each specified image in the multiple specified images to obtain multiple labeled original images. In the above scheme, by obtaining an image containing a target object in a real-time target scene and a specified background, since the specified background is fixed, the specified background image can be obtained in advance, and thus a first label for annotating the target object can be generated based on the specified background image and the specified image. Afterwards, since the multiple first positions partially overlap or completely overlap with the multiple second positions, the original image in the original image set can be matched with the specified image in the specified image set, and the multiple original images are annotated based on the matching results and the first labels corresponding to each of the multiple specified images. A plurality of annotated original images are obtained, which can provide high-quality data annotation samples for the image segmentation network to construct an image segmentation deep learning network, realize automatic annotation, and improve data annotation efficiency and accuracy. The above scheme solves the technical problem of low image annotation efficiency required for image segmentation in the prior art, and achieves the technical effect of effectively improving data annotation efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings described herein are used to provide a further understanding of this specification, constitute a part of this specification, and do not constitute a limitation of this specification. In the accompanying drawings:

[0025] Figure 1 A schematic diagram of an image segmentation data annotation method in the prior art is shown;

[0026] Figure 2 A flowchart of an image segmentation data annotation method according to an embodiment of the present invention is shown;

[0027] Figure 3 A schematic diagram of a surgical robot surgical device according to an embodiment of the present invention is shown;

[0028] Figure 4 A schematic diagram of an image acquisition process in an embodiment of the present specification is shown;

[0029] Figure 5 The whole process of endoscopic surgical instrument image segmentation in one embodiment of this specification is shown;

[0030] Figure 6 A flowchart of an image segmentation data annotation method according to an embodiment of the present invention is shown;

[0031] Figure 7 A schematic diagram showing data annotation using the endoscopic surgical instrument image segmentation method in an embodiment of this specification is shown;

[0032] Figure 8 A flowchart of an algorithm for generating labels in an embodiment of this specification is shown;

[0033] Figure 9 A schematic diagram of an algorithm for generating labels in an embodiment of this specification is shown;

[0034] Figure 10 A flowchart of increasing data diversity matching in an embodiment of this specification is shown;

[0035] Figure 11 A schematic diagram of increasing data diversity matching in an embodiment of this specification is shown;

[0036] Figure 12 A schematic diagram of a deep learning network for image segmentation in one embodiment of this specification is shown;

[0037] Figure 13 FIG1 shows a flow chart of a network training process in an embodiment of the present specification;

[0038] Figure 14 A schematic diagram showing a network model effect test standard in one embodiment of this specification is shown;

[0039] Figure 15 A schematic diagram showing the network model effect test results in an embodiment of this specification;

[0040] Figure 16 A flowchart showing the use of an image segmentation method in surgery according to an embodiment of the present specification is shown;

[0041] Figure 17 A schematic diagram of an image segmentation data labeling device according to an embodiment of the present specification is shown;

[0042] Figure 18 A schematic diagram of a medical device in an embodiment of this specification is shown. DETAILED DESCRIPTION

[0043] The principles and spirit of this specification will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided solely to enable those skilled in the art to better understand and implement this specification, and are not intended to limit the scope of this specification in any way. Rather, these embodiments are provided to make this specification more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.

[0044] Those skilled in the art will appreciate that the embodiments of this specification may be implemented as a system, device, method, or computer program product. Therefore, the disclosure herein may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0045] In the field of deep learning image segmentation, the annotation of surgical instruments in endoscopic images is mainly done manually. For example, the traditional annotation method can be achieved by manually using the labelme annotation tool, which is to manually mark the instruments to be labeled one by one, and finally form a closed loop. The images within the closed loop will be labeled. Please refer to Figure 1 , which shows a schematic diagram of an image segmentation data annotation method in the prior art. Figure 1 The left image in shows the acquired original image, ie, the original endoscopic image. Figure 1 The middle figure shows that the original image can be imported into the Lableme annotation tool, and the instruments that need to be annotated can be manually marked one by one to form a closed loop. Figure 1 The right figure shows how to generate labels for images in a closed loop. It can be seen that the efficiency of image segmentation data annotation methods in the prior art is too low.

[0046] In order to solve the above problems, an embodiment of this specification provides an image segmentation data annotation method. Figure 2A flowchart of an image segmentation data annotation method in one embodiment of this specification is shown. Although this specification provides method operation steps or device structures as shown in the following embodiments or drawings, more or fewer operation steps or module units may be included in the method or device based on routine or no creative labor. In the steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure described in the embodiments of this specification and shown in the drawings. When the method or module structure is applied to an actual device or terminal product, it can be connected in accordance with the method or module structure shown in the embodiments or drawings for sequential execution or parallel execution (for example, a parallel processor or multi-threaded processing environment, or even a distributed processing environment).

[0047] Specifically, if Figure 2 As shown, an image segmentation data annotation method provided in one embodiment of this specification may include the following steps:

[0048] Step S201, obtaining an original image set, a specified image set and a specified background image; the original image set includes a plurality of original images captured when the target object is in a plurality of first positions in a target scene, and the specified image set includes a plurality of specified images captured when the target object is in a plurality of second positions under a specified background; the plurality of first positions and the plurality of second positions have an intersection.

[0049] The image segmentation data annotation method in the embodiment of this specification can be applied to a computer device. An original image set, a specified image set, and a specified background image can be obtained. The original image set can include multiple original images captured when the target object is in multiple first positions in the target scene. The target object is the object to be annotated in the image, which can include various objects, such as an object or a part of an object, or an animal or a part of an animal. The target scene can be the scene in which the target object is located in actual application. For example, the target scene can be an endoscopic surgical environment, and the target object can be a surgical instrument in the endoscopic surgical environment. For another example, the target scene can be an in vitro automated surgical environment, and the target object can be a surgical instrument for performing in vitro automated surgery. For another example, the target scene can be a factory automated production environment, and the target object can be a machine instrument used for automated production.

[0050] The designated image set may include multiple designated images captured when the target object is at multiple second positions against a designated background. The designated background may be any static background with known pixels, preferably a background with a color significantly different from that of the target object, such as a solid color background or a simple patterned background. The designated images only include the designated background and the target object. The designated background images are images captured against the designated background. The designated images and the designated background images are captured under the same environment.

[0051] The first position and the second position are relative to the image acquisition device that acquires the image. Multiple first positions and multiple second positions may intersect, that is, the multiple first positions and the multiple second positions may partially overlap or completely overlap. For example, the multiple first positions may correspond one-to-one with the multiple second positions. For another example, the multiple first positions may include all or part of the multiple second positions. For another example, the multiple second positions may include all or part of the multiple first positions. The original image set and the designated image set may be acquired by the same image acquisition device or an image acquisition device with the same parameters. The shooting parameters for acquiring the original image set may be the same as the shooting parameters for acquiring the designated image set.

[0052] In one embodiment, the original image set may be a plurality of original images captured when the target object moves along a first motion trajectory in a target scene, and the designated image set may be a plurality of designated images captured when the target object moves along a second motion trajectory against a designated background. The first motion trajectory and the second motion trajectory may partially overlap or completely overlap. The first motion trajectory and the second motion trajectory may partially overlap or completely overlap in the three-dimensional coordinate system where the image acquisition device resides.

[0053] Step S202 : generating a first label corresponding to each designated image in the plurality of designated images based on the designated image set and the designated background image; the first label is used to mark a target object in the designated image.

[0054] After obtaining a designated image set and a designated background image, a first label corresponding to each designated image in the designated image set can be generated based on the designated image set and the designated background image. The first label is used to annotate a target object in the designated image. For example, the first label is the region corresponding to the target object in the designated image. Since the designated image includes both the designated background and the target object, a difference image corresponding to the target object can be obtained by subtracting the designated image from the designated background image.

[0055] Step S203 : matching the multiple original images with the multiple designated images according to the first position corresponding to each original image in the multiple original images and the second position corresponding to each designated image in the multiple designated images.

[0056] Since the multiple first positions partially or completely overlap with the multiple second positions, the multiple original images can be matched with the multiple designated images based on the first position corresponding to each of the multiple original images and the second position corresponding to each of the multiple designated images. That is, the original images with the designated images having the same target object position are paired.

[0057] In one embodiment, when capturing an image, the position number corresponding to the image can be associated and stored with the image. It can be determined which of the multiple first positions and multiple second positions are identical, and the position numbers of those positions that are identical can be associated. Thus, when matching an original image with a designated image, the position number of the original image can be used to search for the corresponding position number of the designated image to obtain the designated image that matches the original image.

[0058] In another embodiment, the first motion trajectory overlaps with the second motion trajectory, and the speed and acceleration during the motion are the same. When capturing the original image set, timing is started from the start of shooting, an original image is captured every preset time period, and the original images are numbered in sequence. Similarly, when capturing the specified image set, timing is started from the start of shooting, a specified image is captured every preset time period, and the specified images are numbered in sequence. In this way, when matching the original image set with the specified image set, the original image and the specified image with the same number can be determined as a successful match. Exemplarily, two implementation methods of image matching are given in the above two embodiments. It is understandable that this specification can also use other methods for matching.

[0059] Step S204 : annotating the plurality of original images according to the matching result and the first label corresponding to each of the plurality of designated images to obtain a plurality of annotated original images.

[0060] After the original image with the target object at the same position is paired with the designated image, the first label corresponding to the designated image can be used to annotate the original image that matches the designated image, thereby obtaining a plurality of annotated original images.

[0061] In the above embodiment, by acquiring images containing target objects in a real-time target scene and under a specified background, since the specified background is fixed, the specified background image can be acquired in advance, and thus a first label for labeling the target object can be generated based on the specified background image and the specified image. Afterwards, since the multiple first positions partially overlap or completely overlap with the multiple second positions, the original images in the original image set can be matched with the specified images in the specified image set. Based on the matching results and the first labels corresponding to each of the multiple specified images, the multiple original images are labeled to obtain multiple labeled original images. This can provide high-quality data labeling samples for the image segmentation network to construct an image segmentation deep learning network, realize automatic labeling, and improve data labeling efficiency and accuracy.

[0062] In some embodiments of this specification, the target scene may include an endoscopic surgical environment, and the target object may include surgical instruments used to perform endoscopic surgery. By applying the image segmentation data annotation method described in the above embodiments to an endoscopic surgical environment, the location of surgical instruments can be annotated in endoscopic images, achieving automatic annotation of image segmentation data, improving the accuracy and efficiency of data annotation, and providing a large number of high-quality samples for the construction of an image segmentation network, thereby improving the accuracy of the image segmentation network.

[0063] In some embodiments of the present specification, the target scene may include an in vitro automated surgery environment, and the target object may include surgical instruments used to perform in vitro automated surgery. Among them, in vitro automated surgery may include automated dental surgery, automated orthopedic surgery, and the like, for example, automatic tooth filling, automatic teeth cleaning, knee bone replacement, and the like. Through the image segmentation data annotation method of this embodiment, a large number of image samples annotated with surgical instruments in an in vitro automated surgery environment can be obtained, providing high-quality data annotation samples for the image segmentation network to construct an image segmentation deep learning network. In subsequent in vitro automated surgeries, automatic annotation of surgical instruments can be achieved based on this network, thereby improving the accuracy and efficiency of data annotation and thereby improving the degree of automation of in vitro surgery.

[0064] Furthermore, due to the efficiency of manual labeling, the labeled data often only has a single scene or only a few scenes, which ultimately leads to insufficient data collection and fails to fully consider unexpected situations during the operation, such as lighting, bleeding, fogging, occlusion, reflection, shadow, blur, etc. However, deep learning often requires a large amount of data from different scenes as training data, which is not ideal and has poor generalization performance. Therefore, in some embodiments of this specification, the endoscopic surgical environment includes at least one of the following: an endoscopic surgical environment under lighting conditions, an endoscopic surgical environment under smoke conditions, an endoscopic surgical environment under shadow conditions, and an endoscopic surgical environment under bleeding scenes.

[0065] Specifically, the original image set may include multiple original images collected when the surgical instrument is in multiple first positions in each of a plurality of endoscopic surgical environments; the plurality of endoscopic surgical environments may include at least one of the following: an endoscopic surgical environment under lighting conditions, an endoscopic surgical environment under smoke conditions, an endoscopic surgical environment under shadow conditions, and an endoscopic surgical environment under bleeding scenes.

[0066] The original images of the target object in the endoscopic surgical environment can be obtained under various lighting conditions. The original images of the target object in the endoscopic surgical environment can be obtained under various smoke conditions with different smoke concentrations. The original images of the target object in the endoscopic surgical environment can be obtained under different shadow conditions by adjusting the angle of the light source. The original images of the target object in the endoscopic surgical environment under bleeding conditions can be obtained when there is bleeding in the endoscopic surgical environment. By acquiring images in a variety of endoscopic surgical environments, images in diverse scenes can be provided. Since images in multiple scenes correspond to the same label, the complexity of calculating the label can be reduced. Data annotation of multiple scenes can provide richer samples for subsequent deep learning algorithms, thereby improving the generalization performance of the algorithm.

[0067] In some embodiments of the present specification, generating a first label corresponding to each specified image in the multiple specified images based on the specified image set and the specified background image may include: calculating the difference between the pixel value of each specified image in the multiple specified images and the pixel value of the specified background image to obtain a difference image corresponding to each specified image; and generating a first label corresponding to each specified image in the multiple specified images based on the difference image corresponding to each specified image.

[0068] Specifically, the difference between the pixel values ​​corresponding to each of the multiple specified images and the pixel values ​​of the specified background image can be calculated to obtain a difference image corresponding to each of the specified images. In one embodiment, when calculating the difference, a grayscale difference image can be calculated. After obtaining the difference image, a first label corresponding to each of the multiple specified images can be generated based on the difference image. In this manner, the label corresponding to the specified image can be conveniently and accurately obtained.

[0069] In some embodiments of the present specification, generating a first label corresponding to each specified image in the multiple specified images based on the difference image may include: performing regularization processing on the differences corresponding to the each specified image to obtain a difference image corresponding to the each specified image after regularization processing; converting the each specified image after regularization processing into a binary grayscale image to obtain a first label corresponding to the each specified image.

[0070] In this embodiment, the difference of the RGB three channels can be calculated, and then the difference image corresponding to each channel in the RGB three channels can be obtained. The obtained difference image can be subjected to L1 regularization processing to generate an RGB three-channel black and white image. Specifically, a threshold ε can be set, where x, y, and z are the pixel values ​​of the image RGB three channels, and ε is the set threshold. When |x|+|y|+|z|<ε, the pixel value is assigned to (0,0,0), and when |x|+|y|+|z|>ε, the pixel value is assigned to (255,255,255), and a three-channel black and white image is obtained. The three-channel black and white image is then converted into a binary grayscale image, and the first label corresponding to each specified image can be obtained.

[0071] In some embodiments of the present specification, matching the multiple original images with the multiple designated images based on the first position corresponding to each original image in the multiple original images and the second position corresponding to each designated image in the multiple designated images may include: comparing the coordinate data of each first position in the multiple first positions with the coordinate data of each second position in the multiple second positions to determine a plurality of position pairs, each position pair in the multiple position pairs including a first position and a second position with the same coordinate data; and determining the original image corresponding to the first position in each position pair and the designated image corresponding to the second position as a successful match.

[0072] Specifically, each original image in the original image set is associated with coordinate data of a corresponding position. The coordinate data of each first position in the plurality of first positions can be compared with the coordinate data of each second position in the plurality of second positions to determine a plurality of position pairs. Each position pair in the plurality of position pairs includes a first position and a second position having the same coordinate data. The original image corresponding to the first position in each position pair and the designated image corresponding to the second position can be determined as a successful match. In this manner, image matching can be performed directly based on the coordinate data of the position, which is convenient and efficient.

[0073] In some embodiments of the present specification, labeling the multiple original images based on the matching results and the first label corresponding to each specified image in the multiple specified images to obtain a plurality of labeled original images may include: determining the first label of the specified image that matches the each original image as the second label corresponding to the each original image; and labeling the original image using the second label to obtain a plurality of labeled original images.

[0074] Specifically, after matching multiple designated images with multiple original images, the first label of each designated image matched to each original image can be determined as the second label corresponding to each original image. The second label is used to annotate the target object in the original image. The target object in the original image can be annotated using the second label, thereby obtaining multiple annotated original images. In this manner, the original image can be annotated using the first label and the matching results.

[0075] In some embodiments of the present specification, after obtaining multiple annotated original images, the method may also include: constructing a training sample set based on the multiple annotated original images; using the training sample set to train a preset image segmentation deep learning network to obtain a trained image segmentation deep learning network; the trained image segmentation deep learning network is used to identify the target object from the target image collected from the target scene.

[0076] The set of second labels corresponding to each original image in the multiple original images in the original image set can be determined as a second label set. A training sample set can be constructed based on the original image set and the second label set corresponding to the original image set. The constructed training sample set can be used to train a preset image segmentation deep learning network to obtain a trained image segmentation deep learning network. The trained image segmentation deep learning network can be used to identify a target object from a target image corresponding to a target scene. The target image corresponding to the target scene refers to a target image including a target object taken in the target scene. In the case where the target scene is an endoscopic surgical environment, the target object is a surgical instrument. The trained image segmentation deep learning network can be used to identify surgical instruments from a target image taken in an actual endoscopic surgical environment, providing a basis for surgical automation.

[0077] In some embodiments of the present specification, after obtaining multiple annotated original images, the method may also include: constructing a verification sample set based on the multiple annotated original images; inputting the verification sample set into the trained image segmentation deep learning network to calculate the evaluation index of the trained image segmentation deep learning network; when the evaluation index meets the preset conditions, determining the trained image segmentation deep learning network as the target image segmentation deep learning network; wherein, the target image segmentation deep learning network is used to identify the target object from the target image captured from the target scene.

[0078] Furthermore, to ensure the recognition accuracy of the trained image segmentation deep learning network, a training sample set and a validation sample set can be constructed based on the original image set and a second label set corresponding to the original image set. After the image segmentation deep learning network is trained using the training sample set, the validation sample set can be input into the trained image segmentation deep learning network to calculate evaluation metrics for the trained image segmentation deep learning network. The evaluation metrics can assess the effectiveness of the image segmentation deep learning network. For example, the Dice coefficient (Dice similarity coefficient) and the intersection over union (IOU) of the image segmentation deep learning network can be calculated. The larger the values ​​of these two coefficients, the greater the overlap between the true result and the predicted result. It can be determined whether the calculated evaluation metrics meet preset conditions. In one embodiment, the preset condition is determined to be met when the Dice coefficient is greater than a first preset value. In another embodiment, the preset condition is determined to be met when the IOU is greater than a second preset value. In yet another embodiment, the preset condition is determined to be met when both the Dice coefficient is greater than the first preset value and the IOU is greater than the second preset value. If the calculated evaluation metrics meet the preset conditions, the trained image segmentation deep learning network is determined as the target image segmentation deep learning network. The target image segmentation deep learning network is used to identify the target object from the target image corresponding to the target scene. This method can further improve the accuracy of image segmentation.

[0079] In some embodiments of the present specification, after calculating the evaluation index of the trained image segmentation deep learning network, it may also include: when the evaluation index does not meet the preset conditions, optimizing the network structure of the image segmentation deep learning network; using the training sample set to train the optimized image segmentation deep learning network to obtain a newly trained image segmentation deep learning network.

[0080] If the evaluation index does not meet the preset conditions, the network structure of the image segmentation deep learning network can be optimized, and then the optimized image segmentation deep learning network can be trained using the training sample set to obtain a newly trained image segmentation deep learning network. In this way, the accuracy of image segmentation can be improved.

[0081] In some embodiments of the present specification, after calculating the evaluation index of the trained image segmentation deep learning network, it also includes: when the evaluation index does not meet the preset conditions, re-acquiring the original image set and multiple annotated original images corresponding to the original image set to reconstruct the training sample set; using the reconstructed training sample set to train the preset image segmentation deep learning network to obtain a newly trained image segmentation deep learning network.

[0082] If the evaluation metric is determined not to meet the preset conditions, a new set of original images and a corresponding second label set can be reacquired. A training sample set can be reconstructed using the previously acquired original image set and the newly acquired original image set, and the previously acquired second label set and the newly acquired second label set. The reconstructed training sample set can then be used to train the preset image segmentation deep learning network, resulting in a trained image segmentation deep learning network. In the above embodiment, increasing the number of training samples can further improve the accuracy of image segmentation.

[0083] In some embodiments of this specification, if the evaluation metric does not meet the preset conditions, it is determined whether the reason is an insufficient number of training samples. If so, the original image set and the second label set corresponding to the original image set are reacquired to reconstruct the training sample set, and the preset image segmentation deep learning network is trained using the reconstructed training sample set to obtain a newly trained image segmentation deep learning network. Otherwise, the network structure of the image segmentation deep learning network is optimized, and the optimized image segmentation deep learning network is trained using the training sample set to obtain a newly trained image segmentation deep learning network. In this way, the prediction accuracy of the image segmentation model can be improved.

[0084] In some embodiments of the present specification, after the trained image segmentation deep learning network is determined as the target image segmentation deep learning network, it may also include: obtaining a target image captured in a target scene; inputting the target image into the target image segmentation deep learning network to identify a target object area from the target image; when a target object key point exists in the target object area, controlling the target object to perform a preset operation based on the position of the target object key point.

[0085] In the target scene, the target object key points in the target image can be detected by image detection. For an endoscopic surgical environment, the target object key points can be surgical instrument key points, such as the blade of a scalpel, or the position of the rotating shaft of a surgical instrument. Through detection, multiple key points may be obtained. In order to further determine whether the detection of the key points is accurate, the image segmentation data labeling method in the embodiment of this specification can be used to identify the target object area, that is, the label. Afterwards, it can be determined whether the detected target object key points are in the identified target object area. In the case of determining that the target object key points exist in the target object area, the target object can be controlled to perform a preset operation based on the position of the target object key points. For example, in an endoscopic surgical environment, after determining the surgical instrument key points, the surgical instrument can be controlled to perform a preset surgical operation to improve the automation of endoscopic surgery.

[0086] In some embodiments of the present specification, the target object key points include multiple candidate target object key points; after identifying the target object area from the target image, the method may further include: filtering out target object key points that are not in the target object area.

[0087] If multiple target object key points are detected, target object key points that are not in the target object area can be filtered out. If multiple target object key points exist in the target object area, other technical means can be used to further determine the final target object key point. This method can assist in positioning from area to point, providing a basis for automatic operation of the target object.

[0088] In some embodiments of this specification, the designated background can be a green screen. Green screens are easier to separate from the foreground during image segmentation. They are also brighter and less likely to produce black edges. Furthermore, in most cameras, the Bayer array has twice as many green pixels as red and blue, resulting in the strongest signal and the least noise, encompassing the majority of brightness information. Therefore, setting a green screen as the designated background can reduce noise, facilitating subsequent determination of the label corresponding to the designated image.

[0089] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. For details, please refer to the description of the aforementioned related processing embodiments, and no further description is given here.

[0090] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0091] The above method is described below with reference to a specific embodiment. However, it should be noted that this specific embodiment is only for better illustrating this specification and does not constitute an improper limitation to this specification.

[0092] This specific embodiment provides a method for segmenting images of surgical instruments under endoscopy. Figure 3 , shows a schematic diagram of a surgical robot surgical device involved in an embodiment of this specification. Figure 3 As shown, the surgical device is a doctor console 301, which is mainly used for the doctor to manipulate surgical instruments and record the surgical motion trajectory through the doctor console 301. The patient console 302 includes an endoscope 303, which is used to collect images of human tissues and organs, surgical instruments, and surgical environment information.

[0093] Please refer to Figure 4 , shows a schematic diagram of an image acquisition in an embodiment of this specification. Figure 4 , a schematic diagram of a hand-drawn image under an endoscope is shown. Under the endoscopic field of view, surgical instruments, soft tissues in the abdominal cavity, and intracavitary fluid can be seen. The generated image is also the original image data sample used in this embodiment.

[0094] Please refer to Figure 5 , shows the full flow chart of the method for image segmentation of endoscopic surgical instruments in one embodiment of this specification. Figure 5 As shown, the image segmentation method in this specific embodiment may include the following steps:

[0095] A1. Collect endoscopic images. Collect color images of surgical instruments under the endoscope and save them in formats such as jpg, png, etc.

[0096] A2. Data annotation to generate labels: Through data annotation, corresponding binary labels are generated for the images collected in step A1.

[0097] A3. Data preprocessing to generate training data. The data collected in step A1 is divided into a training set and a validation set. Data augmentation, horizontal flipping, rotation, random cropping, random erasing, color jittering, and image brightness processing are also performed.

[0098] A4. Build a deep learning network for image segmentation. Based on the characteristics of the image, conduct research and analysis and use an appropriate image segmentation network, such as Unet, PSPNet, UperNet, DeepLabv3, etc.

[0099] A5. Train the network. Use the training set to train the image segmentation network to obtain a trained image segmentation network.

[0100] A6. Model performance test. Use the validation set to test the model's performance. Calculate the evaluation metrics of the trained image segmentation network and evaluate the model based on the network's evaluation metrics, such as IoU and Dice. If the model performs well, proceed to step A9. If not, proceed to step A7.

[0101] A7. Determine whether the poor performance is due to data volume issues. If so, return to step A1; otherwise, proceed to step A8.

[0102] A8. Optimize the network structure in the image segmentation network and return to step A4.

[0103] A9. Engineering: Apply the trained image segmentation network to actual endoscopic surgery scenarios.

[0104] Please refer to Figure 6 , shows a flow chart of an image segmentation data annotation method in one embodiment of this specification. Figure 6 As shown, the image segmentation data annotation method in the embodiment of this specification may include the following steps:

[0105] A21. Capture original endoscope images and record surgical motion trajectories. During soft tissue surgery on animals, capture images of surgical instruments against the background. The surgeon's console records the surgical movements and saves the trajectory.

[0106] A22. Repeat surgical action trajectory. Repeat the surgical action saved in step A21.

[0107] A23. Collect the endoscope image under the green cloth background. Replace the animal soft tissue with the green cloth and collect the image of the surgical instrument under the green cloth according to step A22.

[0108] A24. The difference method and L1 regularization method generate a three-channel black-and-white image. A threshold ε can be set, where x, y, and z are the pixel values ​​of the three RGB channels of the image, respectively, and ε is the threshold. When |x|+|y|+|z|<ε, the pixel value is assigned to (0,0,0); when |x|+|y|+|z|>ε, the pixel value is assigned to (255,255,255), resulting in a three-channel black-and-white image. The three-channel black-and-white image is then converted into a binary grayscale image to obtain the label corresponding to each image.

[0109] A25. Collecting diverse endoscopic images. In step A21, artificially increase lighting by adding fill lights, create shadows, create fog, add bleeding, and other diverse scenes to collect endoscopic images in diverse scenes.

[0110] A26. Matching: Match the images collected in steps A21 and A25 with the image collected in step A23 to obtain labels corresponding to the images collected in steps A21 and A25.

[0111] Please refer to Figure 7 , shows a schematic diagram of data annotation using the endoscopic surgical instrument image segmentation method in the embodiment of this specification. Figure 7 As shown in FIG, the original image under the endoscope is collected while the corresponding image after replacing the background is collected. Finally, the pixel value is filtered to generate a label.

[0112] Please refer to Figure 8 , shows a flow chart of the algorithm for generating labels in one embodiment of this specification. Figure 8 As shown in the figure, the image under the green cloth background is subtracted from the green cloth pixel values ​​to calculate the RGB pixel difference between the two images, and the difference is L1 regularized. Where x, y, and z are the pixel values ​​of the three RGB channels of the image, and ε is the set threshold. When |x|+|y|+|z|<ε, the pixel value is assigned to (0,0,0), and when |x|+|y|+|z|>ε, the pixel value is assigned to (255,255,255), resulting in a three-channel black and white image. Then, the value is assigned to convert it into a binary grayscale image.

[0113] Please refer to Figure 9 , shows a schematic diagram of the algorithm for generating labels in an embodiment of this specification. Figure 9 As shown, labels can be obtained based on the endoscope image under the green cloth background and the green cloth background image.

[0114] Please refer to Figure 10 , shows a flow chart of increasing data diversity matching in an embodiment of this specification. Figure 10As shown, based on the original image, you can add lighting, fog, shadow, and bleeding scenes. This will generate at least 5 images corresponding to one original image, thereby increasing data diversity. At the same time, these at least 5 images are matched with the same label, which can reduce the amount of calculation. In other implementations, other scenes can be added according to actual needs. Please refer to Figure 11 , shows a schematic diagram of increasing data diversity matching in an embodiment of this specification. Figure 11 As shown, the method in this embodiment reduces the complexity of calculating labels because images of multiple scenes correspond to the same label.

[0115] This embodiment uses a deep neural network to segment the location of surgical instruments in the image. Figure 12 , shows a schematic diagram of an image segmentation deep learning network in one embodiment of this specification. Figure 12 As shown in FIG, the neural network structure diagram of the deep learning model in this embodiment is shown, which describes the input and output data structure of the neural network and the process of how the encoder layer extracts surgical instrument features and how the decoder layer outputs labels. The neural network structure constructed is as follows Figure 12 As shown in the figure, the image to be inspected is input into the model and then passes through the encoder feature extraction layer to extract deep features and obtain a feature map. The feature map is then upsampled in the decoder layer and fused with the corresponding encoder layer features to improve the model's ability to extract surgical instrument features. Finally, an output tensor is generated to predict the position of the surgical instrument in the image.

[0116] Please refer to Figure 13 , shows a flow chart of the training network process in one embodiment of this specification. This figure is a step diagram of neural network training in deep learning algorithm. Figure 13 As shown, the network training process may include the following steps:

[0117] A51. Initialize all parameters w (coefficients) and b (intercept) in the neural network.

[0118] A52. Input the first set of samples, that is, convert the endoscopic images collected by the present invention into an array through the OpenCV function, and then input it into the neural network.

[0119] A53. Compare the output of the neural network with the true label and calculate the error between the two using the cross-entropy loss function.

[0120] A54. Based on the error obtained in A53, the reverse gradient update is performed through the chain derivation method to update the parameter values ​​w and b of each layer of the neural network.

[0121] A55. Continuously update through step A54 until the minimum error is reached or the maximum value of the epoch (the highest number of iterations) is reached, and the training ends.

[0122] Please refer to Figure 14 , which shows a schematic diagram of the network model effect test standard in an embodiment of this specification. The model effect test of the present invention mainly includes the following two indicators:

[0123]

[0124]

[0125] The larger the Dice and Iou values ​​are, the more the actual results overlap with the predictions, and the better the model effect is.

[0126] Figure 15 Schematic diagram showing the network model effect test results in one embodiment of this specification. Figure 15 As shown in Figure 1, the original image, the label corresponding to the original image, and the result A61 output by the model are shown respectively. Figure 15 As can be seen, the model output differs from the label in the white box. Due to the small object input from the device head, the model is difficult to fully learn. This can be further optimized by increasing the amount of data and tuning the network later.

[0127] Please refer to Figure 16 , shows a flow chart of the use of the image segmentation method in surgery in one embodiment of this specification. Figure 16 This paper mainly describes the application of image segmentation technology in the process of automatic surgery. Figure 16 As shown, the use of image segmentation methods in surgery includes the following steps:

[0128] Step 1: The endoscope collects images during the operation.

[0129] Step 2: Segment the captured image using image segmentation technology. The captured image can be input into the image segmentation network in the embodiment of this specification.

[0130] Step 3: Segment the surgical instrument area.

[0131] Step 4: Determine whether the instrument key point exists within the segmented instrument area. If so, proceed to step 5; otherwise, proceed to step 7.

[0132] Step 5: The key points of the instrument are positioned correctly.

[0133] Step 6, continue to perform automatic surgery (resection of the lesion) using instruments.

[0134] Step 7: Find the key points of the instrument again and return to step 4.

[0135] Automated surgery requires constant knowledge of the locations of key points on surgical instruments (e.g., the blade and the instrument's pivot) to accurately resect or suture tissue. The location of instrument key points often cannot be completely accurate, necessitating the image segmentation method described in the embodiments of this specification to aid in determining whether key points lie within the segmented area. This method can then filter out unqualified key points, thereby improving the accuracy of automated surgery.

[0136] In the embodiments of this specification, in order to provide high-quality data annotation samples for the image segmentation algorithm, a fast image data annotation method is provided, which uses a green cloth to achieve automatic annotation and improve efficiency. A screen emitting green light is used as the background to replace the original real animal soft tissue, and by setting a threshold, labels are automatically generated to replace manual annotation. The kinematic information of the surgical robot is used to repeat the movement repeatedly, and experimental scenes are added at the same time, such as blood and flesh, occlusion, reflection, etc., to expand the data diversity. The method in the embodiments of this specification can be used for the annotation of surgical instruments in images. By replacing the manual annotation method, the efficiency of data annotation in image segmentation is improved, and multiple scenes are added to ensure the diversification of collected data, which can greatly improve the robustness of the image segmentation algorithm. Thereby improving the efficiency and accuracy of key point identification of instruments in automatic surgery.

[0137] Based on the same inventive concept, an image segmentation data annotation device is also provided in the embodiments of this specification, as described in the following embodiments. Since the principle of solving the problem by the image segmentation data annotation device is similar to that of the image segmentation data annotation method, the implementation of the image segmentation data annotation device can refer to the implementation of the image segmentation data annotation method, and the repeated parts will not be repeated. As used below, the term "unit" or "module" can be a combination of software and / or hardware that implements the predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived. Figure 17 This is a structural block diagram of the image segmentation data labeling device according to the embodiment of this specification. Figure 17 As shown, it includes: an acquisition module 171, a generation module 172, a matching module 173 and a labeling module 174. The structure is described below.

[0138] The acquisition module 171 is used to acquire an original image set, a specified image set and a specified background image; the original image set includes multiple original images captured when the target object is in multiple first positions in the target scene, and the specified image set includes multiple specified images captured when the target object is in multiple second positions under the specified background; the multiple first positions and the multiple second positions have an intersection.

[0139] The generating module 172 is configured to generate a first label corresponding to each designated image in the plurality of designated images based on the designated image set and the designated background image; the first label is used to mark a target object in the designated image.

[0140] The matching module 173 is configured to match the multiple original images with the multiple designated images according to a first position corresponding to each original image in the multiple original images and a second position corresponding to each designated image in the multiple designated images.

[0141] The labeling module 174 is configured to label the plurality of original images according to the matching results and the first label corresponding to each of the plurality of designated images to obtain a plurality of labeled original images.

[0142] From the above description, it can be seen that the embodiments of this specification achieve the following technical effects: by acquiring images containing target objects in a real-time target scene and a specified background, since the specified background is fixed, the specified background image can be acquired in advance, and thus a first label for annotating the target object can be generated based on the specified background image and the specified image. Afterwards, since the multiple first positions partially overlap or completely overlap with the multiple second positions, the original images in the original image set can be matched with the specified images in the specified image set, and the multiple original images are annotated based on the matching results and the first labels corresponding to each of the multiple specified images, to obtain multiple annotated original images, which can provide high-quality data annotation samples for the image segmentation network to construct an image segmentation deep learning network, realize automatic annotation, and improve data annotation efficiency and accuracy. The above scheme solves the technical problem of low image annotation efficiency required for image segmentation in the prior art, and achieves the technical effect of effectively improving data annotation efficiency and accuracy.

[0143] This specification also provides a medical device. Figure 18 The schematic diagram of the structure of a medical device based on the image segmentation data annotation method provided in the embodiments of this specification is shown. The medical device may specifically include an input device 181, a processor 182, and a memory 183. The memory 183 is used to store processor-executable instructions. When the processor 182 executes these instructions, the steps of the image segmentation data annotation method described in any of the above embodiments are implemented.

[0144] In this embodiment, the input device can specifically be one of the primary devices for exchanging information between a user and a computer system. The input device can include a keyboard, mouse, camera, scanner, light pen, handwriting input tablet, voice input device, etc.; the input device is used to input raw data and programs for processing these data into the computer. The input device can also receive data transmitted from other modules, units, and devices. The processor can be implemented in any appropriate manner. For example, the processor can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers, etc. The memory can specifically be a memory device used to store information in modern information technology. The memory can include multiple levels. In digital systems, anything that can store binary data can be considered a memory device. In integrated circuits, a circuit with storage functionality that does not have a physical form is also called a memory device, such as a RAM or FIFO. In systems, a physical storage device is also called a memory device, such as a memory stick or a TF card.

[0145] In this embodiment, the specific functions and effects achieved by the medical device can be explained in comparison with other embodiments and will not be repeated here.

[0146] In an embodiment of the present specification, a computer storage medium based on the image segmentation data labeling method is further provided, wherein the computer storage medium stores computer program instructions, and when the computer program instructions are executed, the steps of the image segmentation data labeling method described in any of the above embodiments are implemented.

[0147] In this embodiment, the storage medium includes, but is not limited to, random access memory (RAM), read-only memory (ROM), cache, hard disk drive (HDD), or memory card. The memory can be used to store computer program instructions. The network communication unit can be an interface configured in accordance with the standards specified by the communication protocol for network connection communication.

[0148] In this embodiment, the functions and effects specifically implemented by the program instructions stored in the computer storage medium can be explained in comparison with other embodiments and will not be repeated here.

[0149] Obviously, those skilled in the art should understand that the various modules or steps of the above-mentioned embodiments of this specification can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed across a network composed of multiple computing devices. Alternatively, they can be implemented using program code executable by the computing device, so that they can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than herein, or they can be made into separate integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module for implementation. Thus, the embodiments of this specification are not limited to any specific combination of hardware and software.

[0150] It should be understood that the above description is intended to be illustrative and not limiting. Numerous embodiments and applications beyond the examples provided will be readily apparent to those skilled in the art upon reading the above description. Therefore, the scope of this specification should not be determined with reference to the above description, but rather with reference to the preceding claims, along with the full scope of equivalents to which such claims are entitled.

[0151] The above description is merely a preferred embodiment of this specification and is not intended to limit this specification. Those skilled in the art will readily appreciate that various modifications and variations to the embodiments of this specification are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this specification shall be within the scope of protection of this specification.

Claims

1. An image segmentation data annotation method, characterized in that, it includes: Obtain an original image set, a specified image set, and a specified background image; The original image set includes multiple original images collected when the target object is at multiple first positions in the target scene, and the specified image set includes multiple specified images collected when the target object is at multiple second positions in the specified background; there is an intersection between the multiple first positions and the multiple second positions; Based on the specified image set and the specified background image, generate a first label corresponding to each specified image in the multiple specified images; The first label is used to annotate the target object in the specified image; According to the first position corresponding to each original image in the multiple original images and the second position corresponding to each specified image in the multiple specified images, match the multiple original images with the multiple specified images; Based on the matching result and the first label corresponding to each specified image in the multiple specified images, annotate the multiple original images to obtain multiple annotated original images.

2. The image segmentation data annotation method according to claim 1, characterized in that, The target scene includes an endoscopic surgery environment, and the target object includes a surgical instrument for performing endoscopic surgery; or, The target scene includes an in vitro automated surgery environment, and the target object includes a surgical instrument for performing in vitro automated surgery.

3. The image segmentation data annotation method according to claim 2, characterized in that, The endoscopic surgery environment includes at least one of the following: an endoscopic surgery environment under light conditions, an endoscopic surgery environment under smoke conditions, an endoscopic surgery environment under shadow conditions, and an endoscopic surgery environment with a bleeding scene.

4. The image segmentation data annotation method according to claim 1, characterized in that, Based on the specified image set and the specified background image, generating a first label corresponding to each specified image in the multiple specified images includes: Calculate the difference between the pixel values of each specified image in the multiple specified images and the pixel values of the specified background image to obtain a difference image corresponding to each specified image; According to the difference image corresponding to each specified image, generate a first label corresponding to each specified image in the multiple specified images.

5. The image segmentation data annotation method according to claim 4, characterized in that, According to the difference image, generating a first label corresponding to each specified image in the multiple specified images includes: Perform regularization processing on the difference corresponding to each specified image to obtain a difference image corresponding to each specified image after regularization processing; Convert each specified image after regularization processing into a binary grayscale image to obtain a first label corresponding to each specified image.

6. The image segmentation data annotation method according to claim 1, characterized in that, According to the first position corresponding to each original image in the multiple original images and the second position corresponding to each specified image in the multiple specified images, matching the multiple original images with the multiple specified images includes: Compare the coordinate data of each first position among the multiple first positions with the coordinate data of each second position among the multiple second positions to determine a plurality of position pairs, where each position pair in the plurality of position pairs includes a first position and a second position with the same coordinate data; Determine that the original image corresponding to the first position in each position pair and the specified image corresponding to the second position are successfully matched.

7. The method for annotating image segmentation data according to claim 1, characterized in that, Based on the matching result and the first label corresponding to each specified image among the multiple specified images, annotate the multiple original images to obtain multiple annotated original images, including: Determine the first label of the specified image that matches each original image as the second label corresponding to each original image; Use the second label to annotate the original image to obtain multiple annotated original images.

8. The method for annotating image segmentation data according to claim 1, characterized in that, After obtaining multiple annotated original images, it further includes: Construct a training sample set based on the multiple annotated original images; Use the training sample set to construct an image segmentation model; the image segmentation model is used to identify a target object from a target image collected from a target scene.

9. The method for annotating image segmentation data according to claim 8, characterized in that, After obtaining multiple annotated original images, it further includes: Construct a validation sample set based on the multiple annotated original images; Input the validation sample set into the image segmentation model to calculate the evaluation index of the image segmentation model; When the evaluation index meets the preset conditions, determine the image segmentation model as the target image segmentation model.

10. The method for annotating image segmentation data according to claim 9, characterized in that, After calculating the evaluation index of the image segmentation model, it further includes: When the evaluation index does not meet the preset conditions, adjust the model parameters of the image segmentation model to obtain an adjusted image segmentation model; Use the training sample set to train the adjusted image segmentation model to obtain an optimized image segmentation model.

11. The method for annotating image segmentation data according to claim 9, characterized in that, After calculating the evaluation index of the image segmentation model, it further includes: When the evaluation index does not meet the preset conditions, re-obtain the original image set and the multiple annotated original images corresponding to the re-obtained original image set to re-construct a training sample set; Use the re-constructed training sample set to construct an optimized image segmentation model.

12. The method for annotating image segmentation data according to claim 9, characterized in that, After determining the image segmentation model as the target image segmentation model, it further includes: Obtain a target image collected from the target scene; Input the target image into the target image segmentation model to identify a target object area from the target image; When the target object key points exist in the target object area, control the target object to perform a preset operation based on the positions of the target object key points.

13. The image segmentation data annotation method according to claim 12, wherein, the key points of the target object include a plurality of candidate key points of the target object; after identifying the target object area from the target image, it further includes: filtering out the candidate key points of the target object that are not in the target object area to obtain the verified key points of the target object; controlling the target object to perform a preset operation based on the positions of the verified key points of the target object.

14. An image segmentation data annotation device, wherein, it includes: an acquisition module, configured to acquire an original image set, a specified image set, and a specified background image; the original image set includes a plurality of original images collected when the target object is at a plurality of first positions in a target scene, and the specified image set includes a plurality of specified images collected when the target object is at a plurality of second positions in a specified background; there is an intersection between the plurality of first positions and the plurality of second positions; a generation module, configured to generate a first label corresponding to each specified image in the plurality of specified images based on the specified image set and the specified background image; the first label is used to annotate the target object in the specified image; a matching module, configured to match the plurality of original images with the plurality of specified images according to the first position corresponding to each original image in the plurality of original images and the second position corresponding to each specified image in the plurality of specified images; a labeling module, configured to label the plurality of original images according to the matching result and the first label corresponding to each specified image in the plurality of specified images to obtain a plurality of labeled original images.

15. A medical device, wherein, it includes a processor and a memory for storing processor-executable instructions, and when the processor executes the instructions, it implements the steps of the image segmentation data annotation method according to any one of claims 1 to 13.

16. A computer-readable storage medium, on which computer instructions are stored, wherein, when the instructions are executed by a processor, it implements the steps of the image segmentation data annotation method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Image segmentation network processing method and device, image segmentation method and device and computer equipment

    CN112232355A

  • Image processing method and apparatus, server, and storage medium

    WO2020182036A1